2.2 Control Groups (Negative, Vehicle, Positive, Sham) & Confounding Factors

Key Takeaways

  • Concurrent controls from the same protocol, strain, room, and calendar interval are the primary comparison; historical controls from the same laboratory and a recent, matched database are supporting only.
  • Vehicle controls receive the vehicle at the highest volume used among treated groups; corn-oil gavage itself can raise pancreatic acinar proliferative lesions and lower mononuclear-cell leukemia in male F344 rats (NTP TR-426).
  • Positive controls show assay sensitivity—cyclophosphamide in the mammalian erythrocyte micronucleus test; sodium azide (−S9) and 2-aminoanthracene (+S9) in Ames—and they do not replace vehicle or negative controls.
  • Sham surgery or air-only nose-only restraint isolates procedure stress from test-article effects; pair-feeding isolates palatability and caloric restriction from chemical toxicity.
  • Stratified randomization, mixed cage locations, adequate n, and treating the litter (not the fetus) as the unit in developmental studies keep housing and kinship from masquerading as treatment.
Last updated: September 2026

Every treated group is interpreted against something. Domain I.A.1 C of the DABT practice analysis (control groups and confounding factors) is where otherwise well-dosed studies fail: the “control” was the wrong procedure, the vehicle was biologically active, cages were stacked by dose, or fetuses were analyzed as if they were independent adults. Independent OpenExamPrep coverage here maps negative, vehicle, untreated, positive, and sham controls, then the design habits that keep those controls honest.

Concurrent versus historical controls

Concurrent controls are animals from the same protocol, strain, supplier, diet, water, room, and calendar interval as the treated groups. They are the primary statistical comparison. OECD chronic and carcinogenicity guidelines require a concurrent control. If a vehicle is used, that control receives the vehicle at the highest volume used among the dose groups, and is handled identically except for the test chemical.

Historical controls are the laboratory’s own recent control database—ideally the same strain, supplier, diet, diagnostic criteria, and a defined recent time window. They are supporting. They help when the finding is a rare tumor and the concurrent control happened to be 0/50 while the laboratory background is 0–3%; when the concurrent control looks unusually high or low (a pituitary-tumor cluster in one study); or when a statistically significant shift still sits inside ordinary biological noise. They do not replace concurrent controls. Mixing NTP, CRO A, and CRO B tables across decades, diets (NIH-07 versus NTP-2000), and diagnostic conventions is how false “within historical range” arguments are built.

Negative, vehicle, untreated, and positive controls

Control typeWhat animals receiveWhy it existsTeaching example
Negative / vehicleVehicle at the high-dose volume, no test articleIsolates formulation effects from the chemical0.5% methylcellulose gavage; corn oil; aqueous DMSO mixtures
UntreatedNo gavage and no vehicleDetects vehicle toxicity or procedure stressExtra untreated arm when oil gavage or frequent restraint is used
PositiveA substance known to produce the endpointProves the assay can detect a response (sensitivity)Cyclophosphamide in OECD 474; sodium azide (−S9) and 2-AA (+S9) in OECD 471
ShamThe procedure without test article, atmosphere, or implantIsolates surgery, restraint, or device-placement traumaSham laparotomy; nose-only air-only restraint; empty pump

Vehicle effects are not hypothetical. NTP Technical Report 426 showed that chronic corn-oil gavage in male F344/N rats at 2.5, 5, and 10 mL/kg (5 days/week for 2 years) increased pancreatic acinar hyperplasia and adenoma and decreased mononuclear-cell leukemia relative to untreated controls. Safflower oil and tricaprylin produced related pancreatic findings. If the only control is untreated, oil-dosed groups can look like pancreatic carcinogens for reasons that have nothing to do with the test article. That is why a volume-matched vehicle control is required, and why an extra untreated group is sometimes added when the vehicle itself is under suspicion.

DMSO is the workhorse organic solvent in Ames tests (commonly on the order of 100 μL/plate) but is a poor neat repeated-dose gavage vehicle: it is irritant, can dominate clinical signs, and in-life laboratories typically keep it to a small percentage of an aqueous/organic mix rather than 100% DMSO every day. Residual-solvent permitted daily exposures in finished drug product (ICH Q3C) are a manufacturing specification. They are not a license to use DMSO as the daily toxicology vehicle.

When possible, aqueous vehicles (water, saline, methylcellulose, carboxymethylcellulose) are preferred over oils. OECD TG 474 lists water, saline, methylcellulose, CMC, olive oil, and corn oil among compatible micronucleus vehicles and warns that an atypical solvent needs historical or published evidence that it does not itself raise micronuclei.

Positive controls in genetic toxicology demonstrate proficiency and, where relevant, metabolic activation. They are not a substitute for the vehicle/negative control and they are not part of a standard OECD 408 90-day general-toxicity design.

  • Ames (OECD 471): Sodium azide is a classic direct-acting positive control for Salmonella typhimurium TA1535 and TA100 without S9. 2-Aminoanthracene (2-AA) is widely used with S9 to show that the exogenous activation system can convert a promutagen. Strain-specific direct-acting controls (2-nitrofluorene for TA98, 9-aminoacridine for TA1537) complete a competent panel. An Ames study whose +S9 2-AA plates look like solvent plates does not “prove the test article is non-mutagenic”; it proves the S9 mix or the method failed.
  • Micronucleus (OECD 474): Cyclophosphamide is a canonical in-life inducer of micronuclei (often given intraperitoneally in the positive-control arm). Mitomycin C is another. The vehicle control still anchors the spontaneous micronucleus rate against the laboratory historical range.

Sham procedures, inhalation restraint, and pair-feeding

Sham surgery belongs in protocols for implants, infusion pumps, and surgically placed devices. Otherwise adhesions, infection, analgesic effects, and postsurgical ileus are credited to the material. Sham restraint belongs in nose-only inhalation: tube restraint is stressful and can cut food intake; an air-only sham group shows what the procedure does before any test atmosphere is added. Whole-body chambers trade that restraint stress for a different confounder—animals groom deposited material and add an oral dose they were not supposed to have.

Pair-feeding is the control for palatability collapse. When high-dose dietary animals eat 30% less, they lose weight, drop thymus and liver weight, delay puberty, and look “toxic.” Pair-fed controls receive the same amount of untreated diet that the treated animals consumed. If pair-fed controls reproduce the finding, caloric restriction is in the dock, not necessarily the chemical. OECD TG 452 notes that an additional pair-fed control may be useful when reduced dietary intake is driven by palatability.

Cage location, litter effects, n, and randomization

Cage-location bias is real. Temperature, light intensity, vibration, and technician traffic differ from rack top to bottom and from the door to the back wall. If every high-dose cage sits on the top shelf under a fluorescent bank, reduced body weight or apparent phototoxicity can be housing. Randomize cage position independently of dose; some laboratories also rotate positions on a schedule written into the SOP.

Litter effects dominate developmental toxicology. In OECD TG 414 prenatal developmental studies and TG 443 extended one-generation studies, pups in a litter share genetics, maternal nutrition, and uterine position. The litter is the statistical unit, not the fetus. Treating 12 dams × 12 fetuses as n = 144 independent observations inflates false positives for fetal weight and malformations.

Sample size and randomization complete the control story. Concurrent control n should match treated groups (10/sex in OECD 408; 50/sex in OECD 451). Randomization is typically stratified by body weight so one group is not packed with runts. Without randomization, assignment order (first ten animals off the truck become controls) confounds supplier-shipping stress with treatment.

Why n matters: a beautifully randomized vehicle control with n = 3 per sex cannot distinguish a 20% ALT shift from noise. Why randomization matters: a large n allocated by cage row still estimates the row effect, not the test article.

Protocol scenario

You review a 2-year male F344 rat gavage bioassay. Treated groups received the test article in 5 mL/kg corn oil. The only control listed is “untreated.” High-dose males show an increase in pancreatic acinar adenomas (p < 0.05 versus untreated) and a decrease in mononuclear-cell leukemia. NTP oil-vehicle studies already associate those two shifts with corn oil itself. The missing vehicle control at 5 mL/kg corn oil is a design defect: you cannot separate test-article carcinogenicity from vehicle biology. Historical oil-gavage controls from the same laboratory, strain, and era can support that discussion; they cannot retroactively create the concurrent vehicle arm. A second defect on the same protocol: dose groups were housed by row on one rack. Even a correct vehicle control would not fully rescue a location confounder. The repair for the next study is mechanical: volume-matched vehicle concurrent control, body-weight-stratified randomization, rack positions mixed by dose, and historical ranges used only as supporting narrative.

A second, shorter illustration: an OECD 474 micronucleus study reports a negative test article, a quiet vehicle control, and no cyclophosphamide (or other) positive-control response. That study is uninterpretable, not reassuring. Sensitivity was not demonstrated in that experimental run.

Test Your Knowledge

A 2-year rat study shows 3 hepatocellular adenomas in high-dose males versus 0 in concurrent controls. The same laboratory’s last 10 control groups in that strain, diet, and diagnostic system ranged from 0 to 4 adenomas per 50 males. What is the scientifically primary comparison?

A
B
C
D
Test Your Knowledge

A laboratory is running an OECD 474 mammalian erythrocyte micronucleus test. Which positive-control choice is appropriate to show the in-life assay can detect micronucleus induction?

A
B
C
D
Test Your Knowledge

A dietary 90-day study shows high-dose rats eat about 30% less feed because of poor palatability and lose weight, with thymic atrophy. Which control strategy best separates starvation effects from chemical toxicity?

A
B
C
D