2.3 Statistical Design, Power Analysis & Statistical Testing of Toxicological Data
Key Takeaways
- For control-versus-dose questions on approximately normal, equal-variance data, ANOVA plus Dunnett (each dose versus concurrent control) matches the NOAEL question; Tukey compares every group with every other group and is usually the wrong default.
- Non-normal or heterogeneous data typically move to Kruskal–Wallis plus a nonparametric control comparison (for example Steel); ordered dose–response may use Williams, Jonckheere–Terpstra, or Cochran–Armitage trend tests.
- Survival is summarized with Kaplan–Meier methods; tumor analysis with differential mortality uses Peto-type fatal-versus-incidental (context-sensitive) tests, not an unadjusted terminal-only chi-square.
- OECD TG 408 typically uses 10 animals/sex/group; TG 452 chronic studies normally at least 20/sex; TG 451 carcinogenicity at least 50/sex/group; TG 407’s 5/sex detects only large effects.
- A single p < 0.05 ALT in one animal without dose–response, companion enzymes, or histopathology is not automatically a NOAEL-defining adverse effect; large clinical-pathology batteries inflate false positives.
Domain I.A.1 B (statistical design of studies) and I.C.1 (interpretation of toxicological data) travel together on the DABT exam. One item may show a table of p-values and ask whether the NOAEL should move; another may ask which post-hoc test belongs on a four-group 90-day study. Independent OpenExamPrep teaching here stays conceptual: match the comparison to the question, respect distributional assumptions, plan for survival and tumors, and do not confuse a lone p < 0.05 with a hazard. This section does not paste software printouts or invented BMDS runs.
Power and sample size
Power is the chance of detecting a true effect of a stated size at a stated alpha. OECD sample sizes are regulatory minima, not a guarantee of 80% power for every endpoint.
| Study type | Typical minimum n (rodents) | What that n can and cannot do |
|---|---|---|
| OECD TG 407 28-day | 5/sex/group | Large, consistent effects; weak for small clinical-chemistry shifts |
| OECD TG 408 90-day | 10/sex/group | Core subchronic package; still modest power for rare lesions |
| OECD TG 452 chronic (usually 12 months) | 20/sex/group (4/sex for many non-rodent chronic designs) | Better precision than a 90-day; not a lifetime tumor study |
| OECD TG 451 / 453 carcinogenicity phase | 50/sex/group | Lifetime tumor detection; still limited for very rare tumors |
| OECD 408 recovery satellite | ≥5/sex in control and high dose | Reversibility, not a second full 90-day |
| Prenatal developmental (OECD 414) | About 20 litters/group | Litter is the unit |
Conceptually, a two-sample t-test aiming at a 1-standard-deviation mean shift, two-sided α = 0.05, and about 80% power needs on the order of the mid-teens per group. That is why n = 5 (TG 407) is a screen and n = 50 (TG 451) is required for carcinogenicity. OECD TG 451 also notes that a moderate increase above 50 adds relatively little power for rare tumors—dose placement and survival usually buy more interpretability than adding a handful of extra mice. Unequal allocation (more animals at low dose) is sometimes discussed to improve low-dose estimates; it does not rescue a high dose that kills the group by week 8.
The bar chart in this section plots those OECD minima per sex. Recovery satellites and TK extras are additional animals; they are not a substitute for the main-study n that feeds ANOVA and histopathology.
Multiple-group testing: ANOVA, Dunnett, Tukey, and nonparametric cousins
For continuous, approximately normal data with similar variance (body weight, many clinical-chemistry means), the workhorse is one-way ANOVA (or a linear model with sex and dose) followed by a pre-specified post-hoc comparison written into the protocol.
- Dunnett’s test compares each dose with the concurrent control. That is the default NOAEL question on a 90-day or chronic study. Multiplicity is built for those control comparisons, not for every pairwise curiosity.
- Tukey’s HSD (Tukey–Kramer with unequal n) compares every group with every other group. It spends the error budget on mid-versus-high contrasts that rarely define the NOAEL, and it is more conservative for the control-versus-dose tests you actually need. Use Tukey when there is truly no designated control (several formulations compared with each other).
- If a monotonic dose–response is assumed, Williams’ test (or a trend contrast) can be more sensitive at the low end than Dunnett.
- If residuals are clearly non-normal or variances are unequal, move the decision tree: Kruskal–Wallis globally, then a nonparametric control comparison (Steel for control-only; Steel–Dwass for all pairwise). Do not log-transform “until Dunnett looks significant” without a planned rule.
Binary lesion incidence uses other tools: Fisher’s exact or chi-square for two groups, Cochran–Armitage for a dose trend in proportions. Ordered continuous responses may use Jonckheere–Terpstra. Recovery groups that contain only control and former high-dose animals collapse Dunnett to an ordinary two-group comparison (t-test or Wilcoxon) because only two means remain.
Survival and tumors
Kaplan–Meier curves with a log-rank (or Wilcoxon) test compare time to death. They do not, by themselves, analyze tumors.
When mortality differs by dose, raw terminal tumor counts are biased. Animals that die early are not at risk for late-onset incidental tumors found at necropsy, and fatal tumors shorten life. Peto-type (IARC) context-sensitive analysis classifies tumors as fatal versus incidental (and sometimes mortality-independent) and tests a death-rate endpoint for fatal tumors and a prevalence endpoint for incidental tumors, with time on study taken into account. That is the standard long-term bioassay approach—not an unadjusted chi-square on terminal kills only, and not ANOVA on tumor counts as if they were ALT values.
Haseman-type decision rules (different p-value thresholds for common versus rare tumors) appear in NTP interpretation history as a way to manage false positives across many tissues. They are a policy-history point. They do not replace the protocol’s planned statistics or a pathologist’s reading of biological gradient.
Multiplicity, clinical pathology, and statistical versus biological significance
A 90-day study may emit 20–40 clinical-pathology parameters × 2 sexes × 3 doses. At α = 0.05, several “significant” cells are expected by chance. Interpretation uses a pattern, not a single asterisk:
- Dose–response, both sexes, or a TK reason for one sex
- Concordant enzymes (ALT and AST, or ALT with GLDH/SDH)
- Magnitude versus concurrent and matched historical ranges
- Matching histopathology (hepatocellular necrosis versus glycogen vacuolation)
Statistical significance without that pattern is a hypothesis, not a NOAEL hammer. A clear biological pattern with a p-value that sits near 0.06 is still a toxicologist’s finding.
Protocol scenario. Control male ALT values sit at 35–50 U/L. High-dose males: 38, 41, 44, 47, 49, 52, 55, 58, 61, and one animal at 210 U/L. The group mean is p < 0.05 versus control by Dunnett if the 210 U/L value is kept. Necropsy of that animal shows a handling bruise and no hepatic lesion; the other nine livers are normal; AST, bilirubin, and bile acids are unchanged. Treating 210 U/L as a group liver effect and dropping the NOAEL is poor toxicological reasoning—the signal is one outlier, not a treatment pattern. Conversely, a 30% ALT rise in 8/10 high-dose animals, a parallel AST rise, and centrilobular necrosis in 6/10 is an adverse liver effect even if a reviewer quibbles about exact p-values.
A second scenario on the same 90-day: the statistician, without a protocol amendment, runs Tukey on ALT because “it is stricter.” High versus control is then not significant, but mid versus high is. The NOAEL discussion becomes a pairwise tangle that nobody pre-specified. Re-run the pre-specified Dunnett control comparisons and interpret ALT with the slides.
Bioinformatics and modeling, conceptually
Benchmark-dose (BMD/BMDL) modeling and PBPK models are later interpretation tools in risk assessment. At the study-design stage, know that well-spaced doses, a true low-dose no-effect region, concurrent controls, and adequate n are what those models need. Omics batteries multiply the false-positive problem unless false-discovery control and a hypothesized mechanism are in place. This chapter does not walk through software screens or fabricated output files; the exam tests whether you know why design choices determine whether a model is fittable at all.
Traps
Using Tukey “because it is more rigorous” and then missing a real control-versus-high-dose change. Declaring a NOAEL from p-values without looking at histopathology. Analyzing fetuses as independent n. Ignoring differential survival in a 2-year tumor table. Equating n = 10 (90-day) with n = 50 (carcinogenicity) as if power were identical. Reporting a BMD from a study whose high dose died at week 6.
A 90-day rat study has a concurrent vehicle control and three dose groups. The pre-specified question is whether any treated group differs from concurrent control on serum ALT, assuming approximately normal equal-variance data. Which post-hoc procedure matches that question?
OECD Test Guideline 451 carcinogenicity studies in rodents specify that each dose group and the concurrent control should contain at least how many animals of each sex?
A 2-year rat study has higher late mortality in the high-dose group. Unadjusted tumor rates based only on animals alive at terminal kill ignore decedents. Which analysis approach accounts for survival and for whether a tumor was fatal or incidental?