9.1 Hazard Identification Principles, Weight of Evidence (WoE) & Systematic Review
Key Takeaways
- Hazard is the inherent capacity to cause a defined adverse effect; risk is the probability and severity of that effect under a specified exposure, dose–response, and population.
- Centrilobular hepatocellular hypertrophy with CYP induction, no necrosis, quiet leakage enzymes, and reversal is commonly adaptive (non-adverse); necrosis, inflammation, fibrosis, or lasting dysfunction is an adverse liver hazard.
- False-positive hazard calls over-restrict; false-negative calls leave a real effect unidentified—weight of evidence balances those errors across SAR/QSAR, in vitro and short-term tests, animal bioassays, and human epidemiology.
- Bradford Hill viewpoints (strength, consistency, specificity, temporality, biological gradient, plausibility, coherence, experiment, analogy) are judgments, not a nine-item pass/fail list; only temporality is logically required.
- Systematic review starts with a protocol and a PECO question, then uses structured risk-of-bias appraisal in a Cochrane-style process; test species must represent the target organism, including ecological receptors.
Hazard is not risk
Handbook III.A is hazard identification: can the agent cause a defined adverse effect in a biological system? Risk is the probability and magnitude of that effect under a specified exposure, in a specified population, after dose–response has been described. An IARC Group 1 evaluation, an EPA descriptor of “carcinogenic to humans,” or a positive two-year rodent bioassay is a hazard statement. It is not a prediction that every workplace air concentration or environmental residue will produce disease. Domain III later converts hazard plus dose–response plus exposure into risk characterization. Mixing those steps is a high-yield examination error: treating “hazard identified” as “unacceptable risk at any dose” skips two of the four risk-assessment legs.
Independent OpenExamPrep material in this section covers III.A.1–2 (principles and evidence streams, including species selection for the target organism) and introduces the systematic-review process that handbook III.D.2 B expects at a conceptual level. It is not an ABT product and does not claim official approval, review, or partnership with ABT, IARC, EPA, or Cochrane.
Adverse versus adaptive (non-adverse)
Not every statistically significant change is a hazard. Adverse effects impair function, reduce reserve to meet additional stress, or are precursors to frank pathology that matter for the organism’s health or survival. Adaptive changes (often labeled non-adverse in safety evaluation) maintain homeostasis without functional impairment.
The classic liver contrast is still the cleanest teaching pair. Centrilobular hepatocellular hypertrophy with cytochrome P450 (CYP) induction, increased smooth endoplasmic reticulum, a modest liver-weight increase, no necrosis, no inflammation, no fibrosis, leakage enzymes (ALT, GLDH, SDH) within concurrent-control ranges, preserved clinical condition, and reversal after dosing stops is commonly interpreted as adaptive enzyme induction, not hepatotoxicity. Contrast centrilobular necrosis, hemorrhage, inflammation, marked enzyme leakage, canalicular failure, fibrosis, or failure to reverse: that package is injury, and it is a hepatic hazard finding.
Hypertrophy is a morphologic word, not a verdict. The same description can be adaptive housekeeping or the first step in a proliferative sequence, depending on dose, duration, companion lesions, and recovery. Hazard identification asks whether the organism was harmed in a way that is relevant to the target organism—often humans, sometimes a wildlife receptor—not whether a p-value appeared in an organ-weight table.
False positives and false negatives
False-positive hazard calls label an agent as capable of an effect it does not produce under relevant conditions. Drivers include cytotoxicity artifacts in genotoxicity assays, particle lung overload in high-dose rodent inhalation studies, species-specific pathways that do not operate in the target organism, and multiple-comparison noise in wide clinical-pathology batteries. False positives over-restrict use and can divert attention from real hazards.
False-negative calls miss a real hazard. Drivers include underpowered epidemiology, animal studies that never reached an adequate top dose, the wrong species or metabolic pathway, insensitive endpoints, too-short duration, and reporting bias that hides positive studies. False negatives leave people or ecosystems unprotected.
Weight of evidence (WoE) is how a competent assessor balances those errors. A single positive Ames result at a precipitating, cytotoxic concentration does not automatically create a human mutagenic hazard. A single negative 90-day study in one sex of one species does not erase a coherent occupational cancer cluster with documented exposure. WoE is not a majority vote of papers and not “the newest article wins.” It asks how quality, consistency, biological gradient, and mechanistic coherence look when all streams are read together.
Data streams for hazard identification (III.A.2)
Handbook III.A.2 groups the evidence you actually use.
Structure–activity relationship (SAR) and quantitative SAR (QSAR). Structure flags alerts: an aromatic amine, an α,β-unsaturated carbonyl, a hexavalent chromium oxide. QSAR models—statistical or expert-rule—estimate mutagenicity, sensitization, or aquatic toxicity from descriptors. They are screening and gap-filling tools. They are not a two-year bioassay. Read the domain of applicability: a model trained on industrial organics may be silent or misleading for organometallics or nanomaterials.
In vitro and short-term in vivo tests. Receptor assays, cytotoxicity, Ames and mammalian genotoxicity, hERG, reconstructed skin irritation, 14- or 28-day range-finders. They identify mode of action (MOA) hypotheses and hazards that do not require lifetime dosing (irritation, sensitization, mutagenicity). Limitations: missing absorption–distribution–metabolism–excretion (ADME), missing intact immune and endocrine axes, and concentration artifacts.
Animal bioassays. Repeat-dose, reproductive and developmental, and carcinogenicity studies remain the backbone for many systemic and chronic hazards. Strengths: controlled exposure, histopathology, dose–response. Weaknesses: species differences, high-dose design, and limited power for rare events.
Human epidemiology. Occupational cohorts, case–control studies, accidental and environmental exposures, and clinical-trial adverse-event data for pharmaceuticals. Strength: the target species when the target is humans. Weaknesses: confounding, exposure misclassification, the healthy-worker effect, and latency for cancer.
Concordance across structure, in vitro MOA, animal target organs, and human disease is stronger than any one stream standing alone.
| Stream | What it can show | What it cannot replace |
|---|---|---|
| SAR / QSAR | Structural alerts and modeled hazard scores inside the model’s domain | A well-conducted in vivo study or human investigation |
| In vitro / short-term | MOA hypotheses; irritation, sensitization, genotoxicity, acute lethality clues | Intact ADME, chronic pathology, and population confounding |
| Animal bioassays | Controlled dose–response, target organs, developmental windows | Automatic human or ecological relevance without species logic |
| Human epidemiology | Effects in the target species | Controlled dose and histology; underpowered studies miss rare outcomes |
Bradford Hill viewpoints
Sir Austin Bradford Hill’s 1965 viewpoints remain the epidemiologic language of causal inference in toxicology. They are viewpoints, not a nine-item pass/fail checklist. Temporality—exposure before disease—is the only logically required condition.
- Strength — large relative risks or clear experimental effect sizes are harder to explain by modest confounding.
- Consistency — the association repeats in different populations, laboratories, or species.
- Specificity — a relatively particular exposure–disease pairing. Vinyl chloride and hepatic angiosarcoma is the textbook pairing. Specificity helps when present; many genuine causes are not specific.
- Temporality — cause precedes effect; essential for cancer latency and developmental windows.
- Biological gradient — dose–response or duration–response.
- Plausibility — a credible biological mechanism (it need not be fully known).
- Coherence — the association does not contradict what is already known of the natural history of the disease.
- Experiment — removal or reduction of exposure changes the outcome (workplace controls, smoking cessation, laboratory intervention).
- Analogy — similar chemicals or similar diseases behave similarly (a new organophosphate beside known acetylcholinesterase inhibitors).
On the examination, missing specificity is not proof of no causality, and a plausible cartoon mechanism is not proof when human and animal data disagree. Hill’s point was structured judgment under uncertainty.
Systematic review: protocol, PECO, risk of bias (III.D.2 B)
Modern chemical hazard identification increasingly uses systematic review rather than an undocumented expert narrative. Handbook III.D.2 B expects the process, not a Cochrane software tutorial.
Protocol first. Before searching, write the question, eligibility criteria, search strategy, bias tools, and synthesis plan. Changing inclusion rules after seeing the results is a bias. Public posting of a protocol documents that sequence.
PECO is the environmental-health analogue of clinical PICO:
- Population — humans, a laboratory species, an ecological receptor, a life stage (pregnant, neonate, elderly).
- Exposure — the agent, route, window, and metrics.
- Comparator — unexposed, lower exposed, vehicle control, background incidence.
- Outcome — the hazard endpoint (liver cancer, malformation, acetylcholinesterase inhibition).
A PECO that says “all toxicity of chemical X” is not a review question; it is a fishing expedition.
Risk of bias (RoB). Each included study is appraised for selection, confounding, exposure assessment, outcome assessment, attrition, selective reporting, and—in toxicology—adequacy of concurrent controls, randomization or blinding where applicable, and documentation of study conduct. NTP OHAT, EPA IRIS systematic-review methods, and the Navigation Guide are conceptual cousins of Cochrane review: predefined question, comprehensive search, dual screening, structured RoB, qualitative or quantitative synthesis, and a certainty statement. Meta-analysis is optional. Pool only when studies are similar enough in population, exposure metric, and outcome. A forest plot of incomparable doses and endpoints is decoration, not synthesis.
Cochrane-style process at the conceptual level: (1) question and protocol, (2) systematic search of multiple sources, including grey literature when relevant, (3) transparent inclusion and exclusion with recorded reasons, (4) extraction into tables, (5) RoB, (6) synthesis, (7) interpretation that separates study quality from study size.
Selecting the relevant test species (III.A.2 B)
Hazard in a rat is useful only if it informs the target organism. Handbook III.A.2 B asks you to choose species—and often strain, sex, and life stage—that can represent the biology you care about.
For human health, default laboratory species (rat, mouse, dog, rabbit, nonhuman primate) are starting points, not magic. Ask whether the species metabolizes the agent along a pathway that matters in humans (CYP2E1-dependent bioactivation is a teaching example for carbon tetrachloride-type solvents). Ask whether anatomy at the portal of entry is comparable (rodents have a forestomach; humans do not, so forestomach-only irritant tumors need extra interpretation). Ask whether the life stage matches the hazard (developmental toxicity needs a gestation and a postnatal brain-growth pattern you can defend). For ecological targets, the test organism should represent the receptor: fish and aquatic invertebrates for surface-water risk, birds for granular pesticide risk, bees for pollinator risk, plants for herbicide drift. A two-year rat cancer study does not identify a fish early-life-stage hazard.
When human data exist, they can outweigh discordant rodent findings if exposure and outcome are well measured. Rodent data can still identify hazards epidemiology is too underpowered to see. Species selection is a WoE input, not a preference for whichever species was negative.
Scenario
A new industrial intermediate shows a QSAR mutagenicity alert, a clean five-strain Ames assay ±S9, a 28-day rat study with centrilobular hypertrophy, induced CYPs, no necrosis, quiet ALT, and full reversal, and no epidemiology. WoE for a hepatic injury hazard is weak; WoE for enzyme induction is strong. Calling the file “hepatotoxic” because liver weight rose would confuse adaptive change with adversity. Calling the file “non-hazardous in every organ” would ignore that you have not yet asked chronic, reproductive, or human questions.
A second file is an organophosphate insecticide. In vitro AChE inhibition, cholinergic signs in the rat, red-cell AChE depression, and a small occupational case series with the same toxidrome tell one story. A negative Ames assay does not erase that neurotoxic hazard. The PECO for a systematic review would name the population (applicators), the exposure (the active ingredient and route), the comparator (unexposed crews), and the outcome (cholinergic illness or AChE inhibition)—not “toxicity in general.”
A third file is a herbicide destined for rice paddies. The sponsor submits only rodent oral studies. For the aquatic ecological target organism, you still need fish, invertebrate, and aquatic-plant evidence. The rat is the wrong receptor for that PECO.
Traps
- Equating a p-value with adversity, or equating CYP-induction hypertrophy with necrosis.
- Treating hazard identification as finished risk characterization.
- Using Bradford Hill as a nine-box checklist that fails if specificity is absent.
- Writing the inclusion criteria after looking at which studies were positive.
- Declaring ecological safety from mammalian cancer bioassays alone.
Which statement correctly distinguishes hazard from risk in chemical assessment?
Which liver package is most consistent with an adaptive, non-adverse response rather than a hepatic injury hazard?
A herbicide will be used in flooded rice paddies. Which species-selection statement matches handbook III.A.2 B logic for the ecological target organism?