14.3 Uncertainty & Sensitivity Analysis in Risk Characterization
Key Takeaways
- Parameter uncertainty is in the numbers (C, IR, BW, CSF, RfD). Model uncertainty is in the equations (linear versus threshold, dose addition versus independent action). Scenario uncertainty is in the story (which receptor, which future land use, which complete pathway).
- Monte Carlo sampling propagates input distributions into a distribution of HQ or risk. For a linear HQ, the mean of HQ equals HQ of the mean inputs; the useful output is still the spread, including how often HQ exceeds 1.
- Tornado or one-at-a-time sensitivity ranks which inputs swing the result; it is not a mixture-interaction model and not a substitute for a wrong conceptual site model.
- Bradford Hill viewpoints return in characterization to judge confidence in the causal story at the environmental dose, not only to decide hazard identification. Temporality remains logically required.
- Handbook III.D.2 expects structured evidence integration: systematic review (protocol, PECO, risk of bias, Cochrane/OHAT-style process) and transparent characterization in the spirit of NRC’s 2009 Silver Book (Science and Decisions).
The point estimate is not the whole characterization
Handbook III.D.2 asks how you integrate disparate evidence and how you show uncertainty so a manager is not handed a naked HQ. Independent OpenExamPrep teaching covers parameter versus model versus scenario uncertainty, Monte Carlo propagation, tornado / one-at-a-time sensitivity, qualitative uncertainty tables, Bradford Hill viewpoints used again at the characterization step (not only in hazard identification, Chapter 9), systematic review in the Cochrane / NTP OHAT spirit, and transparency as framed conceptually by the National Research Council’s 2009 Science and Decisions (“Silver Book”). It is not an ABT, NRC, Cochrane, or EPA product and does not claim official approval, review, or partnership with those bodies.
The 1983 Red Book named the four risk-assessment steps. The 1994 Science and Judgment (sometimes called the Blue Book) pressed uncertainty analysis. The Silver Book argued that characterization was often the weakest step because defaults were hidden and cancer/noncancer methods were split. On the examination, you need the distinctions and the tools, not a book-report.
Variability is real heterogeneity (children ingest more soil than adults; some workers overtime). More samples of the same group do not erase it; you describe it. Uncertainty is incomplete knowledge (is the CSF linear below the PoD? did we miss a pathway?). More relevant data can reduce it. Mixing the two words is a III.D.2 miss.
Parameter, model, and scenario uncertainty
Parameter uncertainty. The inputs to the equations you already chose: concentration C (nondetect handling, 95% UCL versus mean), intake rate IR, body weight BW, exposure frequency, RfD or CSF numerical values, TEFs. A different IR changes HQ even if the CSM and the linear model stay put.
Model uncertainty. The form of the equation: threshold RfD versus linear CSF; dose additivity versus independent action; default interspecies factor 10 versus a chemical-specific adjustment factor; using ADD versus LADD. Choosing dose addition for dissimilar MOAs is model error, not a noisy IR.
Scenario uncertainty. The story: future residential use of an industrial parcel; whether a well will be drilled; child trespasser versus indoor worker; climate-driven change in fish consumption. A precise Monte Carlo on an incomplete pathway is a precise estimate of a dose that does not occur (Chapter 11).
| Example in the file | Type | What more data or a clearer decision can do |
|---|---|---|
| Soil concentration 95% UCL versus mean | Parameter | More samples tighten the UCL; they do not pick the receptor |
| Linear CSF applied to a cytotoxicity-driven tumor MOA | Model | MOA work can support a nonlinear reference value instead |
| Cleanup goal assumes houses will be built on the parcel | Scenario | Land-use covenant or a different reasonably expected use |
| Default TEF year (WHO 2005 vs a later table) | Parameter (table choice) with a model flavor | Cite the TEF year; it is still dose addition |
| Ignoring dermal while quantifying only water | Scenario / CSM | Complete the pathway list before sampling IR |
Those rows are a qualitative uncertainty table: name the source, classify it, state the likely direction of bias (high or low HQ/risk), and state what would change the management call. That table is still characterization when you lack a probabilistic model.
Monte Carlo: a distribution of HQ or risk
Point-estimate HQ uses one C, one IR, one BW. Monte Carlo draws those parameters from distributions, recomputes HQ or extra risk thousands of times, and reports a distribution—median, mean, 95th percentile—not a single heroic digit.
Tiny discrete teaching sample (three equally likely intake rates, everything else fixed). Water C = 0.05 mg/L, BW = 70 kg, RfD = 0.002 mg/kg-day, IR = 1.0, 2.0, or 3.0 L/day, daily exposure so ADD = C × IR / BW.
| IR (L/day) | ADD (mg/kg-day) | HQ = ADD / 0.002 |
|---|---|---|
| 1.0 | 0.05 × 1 / 70 = 0.000714 | 0.357 |
| 2.0 | 0.05 × 2 / 70 = 0.001429 | 0.714 |
| 3.0 | 0.05 × 3 / 70 = 0.002143 | 1.071 |
Mean HQ = (0.357 + 0.714 + 1.071) / 3 = 0.714, which equals HQ at the mean IR of 2.0 L/day because HQ is linear in IR. The useful Monte Carlo lesson is not that mean: one of the three equally likely draws exceeds 1. A manager who only hears “HQ = 0.71” never sees that tail. If IR and BW are correlated (larger people drink more), sampling them independently misstates the tail—that is a Monte Carlo design error, not a reason to skip probabilistic analysis.
Monte Carlo does not repair a wrong model (linear CSF when MOA is nonlinear) and does not complete an incomplete CSM. It is a method for parameter (and sometimes variability) propagation. EPA RAGS Volume III is the Superfund-language cousin.
Tornado plots and sensitivity
Sensitivity analysis asks which input moves the output. A tornado diagram ranks parameters by the swing in HQ or risk when each is varied across a stated range (for example 5th to 95th percentile) one at a time, holding others at baseline. The longest bar is the most influential input—often C or IR, sometimes the CSF if cancer risk is the endpoint.
Tornado is not a full Monte Carlo (it misses joint sampling). It is an efficient way to see that refining a poorly measured soil concentration buys more than arguing about the third significant figure on body weight. If the tornado is dominated by model choice (linear versus threshold), more samples of C will not settle the file—that is model uncertainty, and the honest product is two characterizations, not one fake digit of precision.
Bradford Hill again—at the dose you are characterizing
Chapter 9 used Bradford Hill viewpoints for hazard identification. Characterization re-asks them at the environmental or clinical dose the HQ or extra-risk number represents:
- Temporality is still logically required: exposure before the effect.
- Strength: occupational relative risks of 5 do not automatically transfer to an extra-risk of 10^−6; small environmental associations may be real and still fragile.
- Biological gradient: supports using a slope; its absence at low dose may support a threshold model.
- Plausibility and coherence: does the MOA operate at this dose? Cytotoxicity-driven tumors in a bioassay may not operate at Superfund concentrations.
- Consistency and experiment: several study types and a positive intervention or cessation study raise confidence in the number’s meaning, not only in the hazard label.
- Specificity and analogy remain supporting viewpoints, not a nine-item pass/fail scorecard.
A high HQ on a weakly causal story is not the same characterization as a modest extra risk with a mutagenic MOA, concordance across species, and a clear gradient. III.D.2 is that confidence sentence sitting next to the ratio.
Systematic review and Cochrane-style process (III.D.2)
Integrating epidemiology, animal bioassays, and mechanistic streams by “reading whatever PDFs arrived this week” is not evidence integration. Systematic review—as practiced in Cochrane reviews and adapted in NTP OHAT and EPA IRIS protocols—imposes an order:
- Protocol before seeing the results (question, methods, conflict rules).
- PECO (population, exposure, comparator, outcome)—the toxicology cousin of PICO.
- Comprehensive search with stated databases and dates, not a convenience sample.
- Inclusion/exclusion applied in duplicate when possible.
- Risk-of-bias appraisal (Cochrane RoB tools; OHAT risk-of-bias questions).
- Synthesis: meta-analysis when studies are combinable; structured narrative when they are not.
- Confidence in the body of evidence, then translation into the RfD, CSF, or qualitative hazard that characterization will use.
That process is how disparate evidence becomes a transparent input to HQ or extra risk. It does not replace the arithmetic in sections 14.1–14.2. It decides whether those equations are worth trusting at the dose of interest.
Silver Book transparency
NRC Science and Decisions (2009) conceptually asked risk assessors to treat characterization as decision support: make defaults explicit, show variability and uncertainty separately, avoid a false single number when model uncertainty dominates, and do not bury science/policy choices (for example, which percentile of IR, whether 10^−6 is a screen) inside an unexplained HQ. You do not need to recite committee chapter numbers. You do need to show the calculation, name the decision context, classify uncertainty, and state what would change the call.
Scenario
A site HQ is 0.71 using mean IR = 2 L/day. The discrete Monte Carlo above shows a 1.07 tail when IR = 3 L/day. The tornado shows C and IR dominate; BW barely moves HQ. A second group prefers a linear CSF while the first used an RfD because they judged a threshold MOA. That disagreement is model uncertainty: report both results and the MOA basis, do not average 0.71 with 2 × 10^−5 extra risk into one “compromise HQ.” Systematic review of the liver endpoint finds high risk of bias in the key 90-day study (no randomization stated, concurrent controls poorly described). Bradford Hill gradient is weak below the bioassay PoD. The Silver Book-style characterization therefore says: screening HQ 0.71 (mean), tail HQ > 1 at high IR, confidence in the RfD is limited by study quality, and land-use (residential versus industrial) is still an open scenario choice—not “HQ = 0.71, no issue.”
Traps
- Calling every wide confidence interval “variability,” or calling children’s higher soil ingestion “uncertainty.”
- Running Monte Carlo on the wrong pathway, or treating the 95th percentile as automatically equal to a reasonable-maximum point estimate.
- Using a tornado to claim mixture synergism.
- Stopping Bradford Hill at IARC Group 1 and never asking whether the MOA holds at the characterized dose.
- Substituting a narrative pile of citations for a protocol-driven systematic review when III.D.2 asks how evidence was integrated.
- Handing managers a single HQ with defaults unnamed.
Soil lead 95% UCL versus mean, a linear CSF used where cytotoxicity is the hypothesized tumor driver, and a cleanup goal that assumes future houses on an industrial parcel: which uncertainty classification is correct?
Using C = 0.05 mg/L, BW = 70 kg, RfD = 0.002 mg/kg-day, and equally likely IR = 1, 2, or 3 L/day, which Monte Carlo and sensitivity statement is accurate?
Handbook III.D.2 evidence integration and transparent characterization are best described how?