6.2 Reproductive & Developmental Toxicity Testing (ICH S5, Segments I, II, III)

Key Takeaways

  • Historical Segments map to ICH S5(R3) study types: Segment I is fertility and early embryonic development (FEED, stages A–B); Segment II is embryo-fetal development (EFD, stages C–D; OECD 414 in rats plus rabbits); Segment III is pre- and postnatal development (PPND, stages C–F).
  • Fertility index is pregnant females divided by mated females, not live fetuses divided by implants and not pregnant females divided by all females placed for cohabitation.
  • OECD 414 requires the litter, not the individual fetus, as the statistical unit because fetuses in one uterus are not independent observations.
  • Reduced fetal weight and delayed ossification at maternally toxic doses are often developmental delay, not malformations; they must be interpreted against dam body-weight loss, litter size, and gestational timing.
  • ICH S5(R3) allows a documented weight of evidence, including qualified alternative assays in defined scenarios, to defer or reduce some in vivo EFD work; ICH S11 uses a separate weight of evidence to decide whether a juvenile animal study is needed for pediatric use.
Last updated: September 2026

Why reproductive packages are a DABT interpretation topic

Domain I.C.2 includes reproductive and developmental endpoints because the same finding—small fetuses, missing sternebrae, fewer pregnancies—can mean a primary developmental hazard or a secondary effect of a sick dam. OpenExamPrep independent study material for this section covers ICH S5(R3) for human pharmaceuticals and the OECD chemical guidelines that share the same biology. Historical U.S. Segment I, II, and III labels still appear in protocols and on exam items; S5(R3) names the same work fertility and early embryonic development (FEED), embryo-fetal development (EFD), and pre- and postnatal development (PPND).

The design rule is simple: in total, the studies must leave no gaps between reproductive stages A through F (gametogenesis and mating through sexual maturity of the offspring), even if combination designs are used to cut animal numbers.

Segment I / FEED (stages A–B)

Segment I treats males and/or females before mating, through cohabitation, and through implantation. The aim is effects on gonadal function, estrous cyclicity, mating behavior, fertilization, tubal transport, and implantation—not organogenesis malformations. Rodents are usual. Repeat-dose toxicity of at least two weeks often supplies dose-selection data, but two weeks can miss slower testicular effects; sperm counts, motility, morphology, and testicular histopathology from general toxicology still inform whether treated males should be paired with untreated females (and the reverse) so you can assign the affected sex.

Typical endpoints: precoital interval, mating confirmation (plug or sperm-positive smear), pregnancy status, corpora lutea, implants, pre- and post-implantation loss, and reproductive organ histopathology. This study does not replace EFD visceral and skeletal examinations.

Fertility index: teach the denominator

Laboratories report several related percentages. They are not interchangeable.

  • Mating index = mated females / females paired (cohabited) × 100. “Mated” means evidence of mating (plug, sperm, or equivalent), not pregnancy.
  • Fertility index = pregnant females / mated females × 100. Pregnancy is confirmed by implants, viable embryos, or equivalent evidence. If 20 females are paired, 18 are sperm-positive, and 15 are pregnant, fertility index is 15/18, not 15/20 and not live fetuses/implants.
  • Gestation (pregnancy outcome) index = females with live offspring / pregnant females × 100.
  • Pre-implantation loss uses corpora lutea versus implants; post-implantation loss uses implants versus live fetuses. Those are litter-composition measures, not the fertility index.

Using paired females in the fertility denominator mixes mating failure with conception failure. Using fetuses in the fertility denominator mixes implantation biology with later embryo survival.

Segment II / EFD (stages C–D) and OECD 414

Segment II doses the pregnant dam during major organogenesis and examines fetuses near term for death, growth, and structural abnormality. OECD 414 is the chemical prenatal developmental toxicity guideline; pharmaceutical EFD studies follow the same biology under S5(R3). Two species are the default: a rodent (usually rat) plus a nonrodent, usually rabbit. Mice are not the default second species. Dosing windows are species-specific (rat organogenesis is not a copy-paste of rabbit gestation days).

Fetal examinations include external, visceral, and skeletal evaluations. Group size is set to yield enough litters for interpretation (commonly on the order of 20 pregnant animals per group in OECD 414-style work; follow the protocol’s evaluable-litter target).

The litter is the statistical unit

OECD 414 states that numerical results are evaluated using the litter as the unit for data analysis. Fetuses sharing a uterus, a dam’s pharmacokinetics, and a common nutrition supply are not independent. Treating 12 fetuses from one dam as 12 independent malformations inflates the sample size and can create false positives. Report both percent fetuses affected and percent litters affected, and run statistics on litter means or litter incidences. The same litter logic applies to fertility studies and to OECD 443.

Segment III / PPND (stages C–F)

Segment III doses from implantation through weaning and follows offspring through sexual maturity because some effects (behavior, reproduction of F1) appear only after birth. Endpoints include parturition, pup viability, growth, sensory and reflex development, often a functional observation battery, and F1 mating in many designs. The rodent is usual. For biopharmaceuticals with limited rodent pharmacology, an enhanced PPND (ePPND) in nonhuman primates may replace separate fertility and EFD studies that those species cannot support.

EOGRTS (OECD 443) for chemicals

The extended one-generation reproductive toxicity study (EOGRTS, OECD 443) is the REACH-era chemical design that can replace a two-generation study (OECD 416) when triggered. Parental animals are dosed at least two weeks before mating, through mating, gestation, and weaning. At weaning, F1 animals are assigned to cohort 1 (reproductive/developmental toxicity; 1B may produce F2), cohort 2 (developmental neurotoxicity), and cohort 3 (developmental immunotoxicity) when those cohorts are triggered. EOGRTS is not a synonym for ICH Segment II. A chemical program that ran only OECD 414 has not evaluated postnatal reproductive function the way 443 does.

Maternal toxicity confounding fetal weight and ossification

Handbook-style interpretation (including delayed ossification as a developmental-delay finding) requires you to read the dam and the fetus together. At doses that cause maternal body-weight loss, reduced food intake, or clinical toxicity, fetuses are often smaller. Delayed ossification—unossified or incompletely ossified sternebrae, phalanges, skull bones—is common in smaller fetuses and in larger litters. It is usually classified as a variation or developmental delay, not a malformation, unless the pattern is frank absence of a bone with other dysmorphology.

You cannot automatically dismiss every skeletal finding as “just maternal toxicity,” and you cannot call delayed ossification a teratogenic syndrome solely because it is statistically increased at a high dose that also crushed maternal weight. Ask: Did fetal weight fall in parallel? Is gestational age at cesarean section comparable? Do historical controls show the same ossification variants? Are there malformations at doses without maternal toxicity? Adverse fetal effects remain adverse for hazard identification, but mode and human relevance change if they appear only with marked maternal illness.

ICH S5(R3) weight of evidence and study reduction

S5(R3) still expects the reproductive cycle to be covered, but it is not a mandatory three-study stamp for every molecule. Combination FEED/EFD/PPND designs are acceptable if stages do not gap. Qualified alternative assays (in vitro/ex vivo systems listed in the S5(R3) annex, used inside their chemical and biological domain) can, in defined scenarios, support deferral of definitive EFD or contribute to replacing one species when combined with an enhanced preliminary EFD in the remaining species. ICH M3 already allows limited enrollment of women of childbearing potential on preliminary two-species EFD data; S5(R3) adds options that pair a qualified alternative with a preliminary EFD, or that enrich a GLP preliminary EFD with more litters and skeletal exams. None of that is a blank waiver. Limitations, exposure, and residual uncertainty stay in the risk narrative. Definitive in vivo studies still carry more weight than alternative assays when they exist.

Juvenile animal studies when pediatric use is planned

Pediatric plans are ICH S11, not Segment II. A juvenile animal study (JAS) is not automatic because a label might include children. S11 uses a weight of evidence: youngest intended age, developing-organ risk, pharmacology, existing adult and PPND data, treatment duration, and whether clinical monitoring can manage residual risk. When a JAS is warranted, it is customized (growth, long-bone length, sexual development, clinical pathology, toxicokinetics, plus concern-driven endpoints) and generally a single species, preferably rodent. Pediatric-first or pediatric-only development without prior adult data generally needs two JAS (rodent and nonrodent) if feasible. ICH S9 governs whether oncology products need a JAS; S11 still informs design if one is done.

Realistic scenario

A rat EFD study at the high dose shows 12% maternal body-weight decrement, fetal weight −18%, and an increase in incomplete sternebral ossification, with no visceral malformations and no effects at the mid dose where dams gained weight normally. The study director who calls the high dose a teratogenic NOAEL failure without discussing maternal toxicity and fetal weight is misreading the skeleton. The study director who discards the finding entirely because “dams were sick” is also wrong: it remains a high-dose developmental delay that may still inform the margin. The mid-dose no-effect level for both dam and fetus is the more useful point of departure if the human dose sits far below that exposure.

Traps

  • Computing fertility as live fetuses/implants or as pregnant/paired females.
  • Analyzing each fetus as an independent statistical unit.
  • Equating OECD 414 with EOGRTS because fetuses were weighed.
  • Treating delayed ossification at maternally toxic doses as equivalent to a major malformation syndrome.
  • Starting a juvenile study solely because a pediatric indication exists, without an S11 weight-of-evidence justification.
Test Your Knowledge

In a rat FEED study, 24 females are paired, 20 have copulatory plugs (mated), 16 have implants (pregnant), and mean live fetuses per litter is 12. What is the fertility index?

A
B
C
D
Test Your Knowledge

An OECD 414 rabbit EFD study finds a malformation in 8 fetuses, all from 2 of 20 litters at the high dose. Which statistical approach is appropriate?

A
B
C
D
Test Your Knowledge

High-dose rat EFD dams lose 15% body weight versus controls. Their fetuses are lighter, with increased incomplete ossification of sternebrae and phalanges and no visceral malformations. How should that pattern be interpreted?

A
B
C
D