4.2 In Vitro, In Silico (QSAR, Read-Across, PBPK, AI/ML), Field, Clinical & Epidemiologic Characterization
Key Takeaways
- Primary hepatocytes retain the most physiologic Phase I/II activity but are short-lived and donor-variable; differentiated HepaRG is typically more cytochrome P450-competent than standard HepG2; unmodified HepG2 is convenient but metabolically limited.
- Ames (OECD 471) detects bacterial gene mutation with and without S9; mammalian cell assays are needed for clastogens and aneugens; hERG under ICH S7B is a functional IKr endpoint, not a cytotoxicity assay.
- Neutral red uptake and MTT estimate viability; functional assays (channel current, beating, transporter inhibition) can be positive at non-cytotoxic concentrations; adding S9 is a bioactivation design choice, not a universal improvement.
- QSAR is statistical or expert-rule and needs an applicability domain; read-across needs analogue justification; physiologically based pharmacokinetic models support route and species internal-dose extrapolation; artificial intelligence and machine learning require fit-for-purpose validation. No specific commercial system is required by ABT.
- Mesocosm and field studies characterize ecological community change; Phase 1 characterizes human safety and pharmacokinetics; cohort, case-control, and cross-sectional epidemiology characterize associations with different bias structures—not risk-assessment arithmetic.
Why non-animal and human evidence streams sit beside in vivo work
Handbook tasks I.B.3 (in vitro), I.B.4 (in silico, including quantitative structure–activity relationships, read-across, physiologically based pharmacokinetics, and emerging artificial intelligence and machine learning), and I.B.5 (field, clinical, and epidemiologic characterization) ask you to pick the right evidence stream for the decision, not to run a rat 90-day study by habit. Independent OpenExamPrep teaching for this section treats these methods as characterization tools with named endpoints and named limitations. Quantitative risk-assessment arithmetic (reference doses, margins of exposure, cancer slope factors) belongs in later chapters; here you must know what each stream can actually measure.
A useful first question is always: what must be true in the biology, and is a living intact mammal the only system that can show it? Local skin corrosion, bacterial gene mutation, delayed ventricular repolarization at a channel, metabolite identification, community-level aquatic effects, first-in-human adverse events, and occupational cancer incidence are different characterization problems.
In vitro hepatic systems: primary hepatocytes, HepG2, and HepaRG
Primary hepatocytes are freshly isolated from human or animal liver and plated as monolayer or sandwich cultures. They retain the most physiologic Phase I oxidation (cytochrome P450 enzymes) and Phase II conjugation for a limited window—often hours to a few days—before de-differentiation. They are the usual first choice for intrinsic clearance, metabolite identification, and hepatotoxicity that depends on bioactivation. Limitations are donor variability, scarce tissue, plating artifacts, and loss of non-parenchymal immune cells unless you build a co-culture.
HepG2 is a human hepatocellular carcinoma line. It is robust, inexpensive, and easy to transfect. Basal cytochrome P450 expression is low compared with primary cells, so metabolism-dependent toxicity can be a false negative unless you add an exogenous activation system or engineer specific enzymes. Treat HepG2 as a convenient human hepatic-derived screen, not as a drop-in primary hepatocyte.
HepaRG is a human progenitor hepatoma line that can be differentiated into hepatocyte-like and biliary-like cells. After differentiation, activities such as CYP3A4, CYP1A2, and CYP2B6 and some Phase II pathways are typically closer to primary cells than standard HepG2, with better lot-to-lot control than primary human donors. HepaRG is still not a complete liver: immune cells, zonal architecture, and in vivo hemodynamics are missing unless the model is expanded.
| System | Metabolic competence | Best characterization use | Main limitation |
|---|---|---|---|
| Primary hepatocytes | Highest constitutive Phase I/II in the short term | Clearance, metabolite ID, bioactivation-dependent hepatotoxicity | Donor variability; rapid de-differentiation |
| Differentiated HepaRG | Intermediate to high CYP after differentiation | Repeatable human-like metabolism screens | Still a cell line; not a full liver |
| HepG2 (unmodified) | Low basal CYP | Cytotoxicity, transfection, some transporter work | False negatives for bioactivation |
hERG and ICH S7B
The human Ether-à-go-go-Related Gene (hERG) encodes Kv11.1, which carries the rapid delayed-rectifier potassium current IKr. Block delays ventricular repolarization, raising concern for QT prolongation and torsade de pointes. ICH S7B is the nonclinical strategy for delayed ventricular repolarization: an in vitro IKr/hERG assay plus in vivo QT in a suitable species, later integrated with clinical electrocardiography under ICH E14. Manual patch clamp remains the reference electrophysiology method; automated patch and binding assays exist as screens.
hERG is a functional endpoint (current block at a concentration), not a cytotoxicity endpoint. A compound can inhibit IKr at nanomolar free concentrations while leaving cells alive in an MTT plate at micromolar concentrations. Characterization compares an IC50 (or similar potency) with free therapeutic exposure—this chapter does not complete torsade risk math.
Ames versus mammalian-cell mutagenicity
The bacterial reverse-mutation assay (Ames, OECD 471) uses Salmonella typhimurium tester strains and often Escherichia coli WP2 to detect point mutations (base-pair substitution and frameshift) as reversion to histidine or tryptophan prototrophy. Run the assay with and without exogenous metabolic activation. Ames is strong for DNA-reactive mutagenicity. It is not a chromosome-damage assay; bacteria lack mammalian chromosome structure and many repair pathways.
Mammalian-cell mutagenicity and cytogenetics cover what Ames cannot:
- OECD 476 in vitro hypoxanthine-guanine phosphoribosyltransferase (HPRT) or xanthine-guanine phosphoribosyltransferase gene mutation
- OECD 490 mouse lymphoma thymidine kinase (Tk) assay (mutations and some clastogenic events via small versus large colonies)
- OECD 473 in vitro chromosomal aberration
- OECD 487 in vitro micronucleus (clastogens and aneugens)
An ICH S2(R1) pharmaceutical battery typically pairs Ames with a mammalian system (in vitro and/or in vivo). A clean Ames plate does not characterize aneugenicity. A positive in vitro micronucleus with a negative Ames points you toward chromosome damage, not bacterial gene mutation.
Three-dimensional skin and eye models
Reconstructed human epidermis models are used for skin irritation (OECD 439) and skin corrosion (OECD 431). They characterize local injury (cell viability in a three-dimensional construct after topical exposure), not systemic 90-day toxicity. Eye characterization uses OECD 492 reconstructed human cornea-like epithelium, OECD 437 bovine corneal opacity and permeability, and related isolated-eye methods. These streams exist so local irritation or corrosion can be characterized without a Draize test as the first default where the method is accepted. Brand names of commercial tissues are examples of a model class; ABT does not require a named commercial insert.
Cytotoxicity (NRU, MTT) versus functional endpoints
Neutral red uptake (NRU) measures accumulation of the dye in intact lysosomes of living cells; loss of uptake indicates membrane or lysosomal injury. OECD 432 3T3 NRU is the standard phototoxicity in vitro method. MTT and related tetrazolium assays measure reduction by cellular reductases and are often interpreted as metabolic viability. MTT can mislead if the chemical itself reduces the dye or if metabolism changes without death.
Functional endpoints ask whether a living cell can still do a job: hERG current, cardiomyocyte beating (impedance), bile-salt export pump inhibition, neuronal firing, cytochrome P450 activity, cytokine release. Report an LC50 only when the decision is lethality in the dish. If the decision is channel block, transporter failure, or phototoxicity, a viability number is supporting context, not the primary characterization.
Metabolic competence and S9 as a design choice
S9 is the post-mitochondrial supernatant from induced rodent liver (classically Aroclor 1254 or phenobarbital plus β-naphthoflavone), fortified with NADPH, used as extracellular metabolic activation. Adding S9 is a design choice, not a quality badge:
- Needed to detect many pro-mutagens (benzo[a]pyrene, cyclophosphamide as teaching examples) whose DNA-reactive metabolites form only after oxidation.
- Can detoxify other structures, so a response only without S9 still matters.
- S9 is not a hepatocyte: incubations are short, enzymes sit outside the cell, conjugation is incomplete, and transporter-mediated uptake is absent.
Choose primary hepatocytes or differentiated HepaRG when the question is a human metabolite profile, time-dependent inactivation, or hepatocyte-specific transporters. Choose S9-plus Ames when the question is bacterial mutagenicity of oxidative metabolites under OECD 471. Do not add S9 to a hERG assay as if ICH S7B required liver enzymes in the patch-clamp bath.
In silico characterization: QSAR, read-across, PBPK, and AI/ML
Quantitative structure–activity relationship (QSAR) models relate chemical structure to a predicted activity.
Statistical (Q)SAR is trained on measured data (regression, random forests, other learners). Performance counts only inside the applicability domain (structural coverage, descriptor range, analogue density). The OECD (Q)SAR validation principles expect a defined endpoint, an unambiguous algorithm, a defined domain, measures of goodness-of-fit, robustness, and predictivity, and a mechanistic interpretation when possible.
Expert-rule (knowledge-based) systems encode structural alerts (aromatic amine, epoxide, alkylating function) with written reasoning. Commercial examples you will see named in laboratories include Derek Nexus (expert rules across several toxicity endpoints) and Sarah Nexus (a statistical mutagenicity system). Treat them as illustrations of the two model classes. ABT does not require any specific commercial product. A documented public workflow using the OECD QSAR Toolbox (profilers, category formation, data-gap filling) can be equally legitimate.
Read-across predicts a target chemical from source analogues that have data. Analogue justification must cover structural similarity (scaffold, functional groups, likely metabolites), physicochemical similarity (for example log Kow and pKa), toxicokinetic similarity (same bioactivation map?), quality of the source data, and uncertainty—including whether you are interpolating or taking a worst-case analogue. A Tanimoto coefficient alone is not a complete justification.
Physiologically based pharmacokinetic (PBPK) models use organ blood flows, partition coefficients, and clearance to estimate internal dose (plasma or tissue area under the curve, Cmax). Characterization uses include route-to-route extrapolation (oral animal data toward an inhaled workplace scenario) and species extrapolation (rat to human) when physiology and absorption, distribution, metabolism, and excretion parameters are adequate. PBPK is not a hazard-identification assay by itself. It needs a stated purpose, sensitivity analysis, and comparison to observed pharmacokinetic data when they exist.
Artificial intelligence and machine learning (AI/ML)—handbook I.B.4 B—are emerging tools for property prediction, image analysis, and literature mining. Limits are black-box predictions, biased training sets, overfitting, undefined domains, and weak mechanistic narrative. Fit-for-purpose validation (sensitivity and specificity for the actual decision, documented domain, independent test chemicals) is required before an AI score characterizes a hazard. An impressive area-under-the-receiver-operating-characteristic curve on the training set does not replace Ames, hERG, or a 90-day study when those experiments are the decision-relevant endpoints.
Field, clinical, and epidemiologic characterization
Laboratory ecotoxicology (for example OECD 203 fish acute, OECD 202 Daphnia immobilization, algal growth tests) characterizes single-species apical effects under controlled water quality. Mesocosms (outdoor or large indoor multi-trophic experimental ecosystems) and true field studies characterize community-level endpoints: abundance, taxonomic diversity, functional groups, and recovery after a pulse such as a pesticide application. These higher-tier aquatic or terrestrial studies ask whether an assemblage changes under a realistic exposure regime. Limitations are weather and immigration confounding, exposure-measurement error, low replication, and cost. They do not produce a rat no-observed-adverse-effect level.
Phase 1 clinical studies are usually first-in-human safety, tolerability, and pharmacokinetics in healthy volunteers or, in oncology, in patients. Single ascending dose and multiple ascending dose designs collect adverse events, laboratories, vital signs, electrocardiograms, and plasma concentrations. This is human in-life characterization, not efficacy. Small n, short duration, and selected healthier organs mean absence of events is not chronic safety.
Observational epidemiology characterizes associations in people already exposed:
- Cohort: define exposure, follow forward (prospective) or reconstruct a historical cohort for incidence. Strength: temporality. Weaknesses: cost for rare outcomes, loss to follow-up, healthy-worker effect.
- Case-control: start with cases of disease and controls, compare odds of past exposure. Strength: efficiency for rare diseases. Weaknesses: recall bias, control-selection bias, sloppy exposure timing.
- Cross-sectional: exposure and outcome at one time (prevalence). Strength: speed, hypothesis generation. Weaknesses: reverse causation, prevalence-incidence bias, no follow-up.
These designs characterize associations. They are not yet the risk-characterization math of later chapters.
Evidence streams compared
| Evidence stream | Typical endpoints | Main limitation |
|---|---|---|
| In vivo laboratory toxicology | Clinical signs, body weight, FOB, ophthalmology, pathology, TK | Interspecies relevance, cost, animal use |
| In vitro (hepatic, genotoxicity, local, safety pharmacology) | Ames, mammalian cell mutation or micronucleus, hERG current, NRU/MTT, 3D skin/eye viability, hepatocyte clearance | Metabolic competence and in vivo dose context |
| In silico | QSAR calls, Toolbox profilers, analogue read-across, PBPK internal dose, fit-for-purpose ML scores | Applicability domain, analogue justification, validation |
| Field / mesocosm ecotoxicology | Community abundance, diversity, functional groups, recovery | Environmental confounding and exposure reconstruction |
| Phase 1 clinical | Adverse events, laboratories, ECG, human pharmacokinetics | Small n, short duration, selected subjects |
| Observational epidemiology | Incidence, odds ratios, prevalence | Confounding, bias, exposure misclassification |
Scenario: choosing a characterization stream
A data-poor industrial aromatic amine needs a mutagenicity picture before anyone designs a 90-day dietary study: start with documented QSAR or Toolbox profiling plus analogue read-across, then Ames with and without S9 and an appropriate mammalian cell assay—not a mesocosm and not a Phase 1 trial. A herbicide already measured in ditch water, with a question about invertebrate community recovery, needs a mesocosm or field study; HepG2 MTT will not answer it. A small-molecule drug approaching first-in-human dosing needs ICH S7B hERG (with in vivo QT as required by that strategy) and then Phase 1 human safety and pharmacokinetics; a cross-sectional worker survey is the wrong species and the wrong time. An occupational bladder-cancer question in a mature workforce is a cohort or case-control epidemiology problem; PBPK can support internal-dose reconstruction but does not by itself establish the association.
High-yield traps
- Treating unmodified HepG2 as metabolically equivalent to primary hepatocytes.
- Using only MTT to dismiss a hERG concern.
- Claiming that ABT requires Derek Nexus, Sarah Nexus, or any other named commercial engine.
- Adding S9 to every in vitro assay as if more metabolism were always more valid.
- Using a Tanimoto score as a complete read-across justification.
- Treating a cross-sectional prevalence snapshot as proof of temporality for a chronic disease.
- Using PBPK as a substitute for analogue similarity when the method is read-across, or as a substitute for Phase 1 when the method is human safety.
Which statement correctly contrasts common hepatic in vitro systems used to characterize metabolism-informed toxicity?
A data-poor industrial chemical needs a mutagenicity characterization plan that may include computational methods. Which approach is scientifically justified?
Workers with quantitative inhalation exposure records are followed for 20 years and bladder-cancer incidence is compared with an unexposed worker group. Which characterization method is this, and what is its main epidemiologic strength?