11.4 Molecular Epidemiology & Pathogen Typing

Key Takeaways

  • Molecular epidemiology applies genotypic characterization to track infectious disease transmission dynamics, investigate hospital outbreaks, identify point-source contaminations, and monitor pathogen evolution.
  • Pulsed-field gel electrophoresis (PFGE) evaluates macrorestriction fragment patterns using the Tenover criteria: 0 band differences indicate indistinguishable strains, 2–3 indicate closely related strains, and 7 or more indicate unrelated strains.
  • Multi-Locus Sequence Typing (MLST) sequences internal fragments of 6 to 8 housekeeping genes to generate standardized allelic profiles defining discrete Sequence Types (STs) and Clonal Complexes.
  • Single-locus repeat-based typing methods, such as Staphylococcus aureus spa typing and Multi-Locus Variable Number Tandem Repeat Analysis (MLVA), provide rapid, high-resolution discriminatory power.
  • Whole-Genome Sequencing (WGS), through whole-genome SNP analysis (wgSNP) and core-genome MLST (cgMLST), represents the definitive high-resolution reference standard for outbreak epidemiology.
Last updated: August 2026

11.4 Molecular Epidemiology & Pathogen Typing

Quick Summary: Molecular epidemiology utilizes nucleic acid-based strain typing (genotypic fingerprinting) to establish clonality, track hospital-acquired infection (HAI) transmission chains, detect contaminated food/water distribution networks, and monitor antimicrobial-resistant clones. Core methodologies include macrorestriction analysis by Pulsed-Field Gel Electrophoresis (PFGE) interpreted via the classic Tenover criteria, Multi-Locus Sequence Typing (MLST) of conserved housekeeping genes, repeat-based typing (spa typing and MLVA), and modern Whole-Genome Sequencing (WGS) utilizing whole-genome single nucleotide polymorphism (wgSNP) and core-genome MLST (cgMLST) bioinformatics pipelines. Mastery of discriminatory index calculations, band pattern interpretation rules, and comparative typing resolutions is critical for the ASCP MB examination.


1. Principles of Molecular Typing Systems & Performance Metrics

A molecular typing system must accurately distinguish between outbreak-related isolates (clonally identical progeny of a single ancestral cell) and unrelated background community strains.

+---------------------------------------------------------------------------------------------------------+
|                                TYPING SYSTEM PERFORMANCE CRITERIA & METRICS                             |
+---------------------+-----------------------------------------------------------------------------------+
| Metric              | Operational Definition & Validation Standard                                      |
+---------------------+-----------------------------------------------------------------------------------+
| **Discriminatory**  | Ability of the method to distinguish between genetically unrelated isolates.      |
| **Power ($D$)**     | Quantified statistically by **Simpson's Index of Diversity ($D$)**;               |
|                     | clinical and epidemiological typing systems should achieve **$D \ge 0.95$**.      |
+---------------------+-----------------------------------------------------------------------------------+
| **Typeability**     | Percentage of isolates within a target species that generate an unambiguous,      |
|                     | interpretable typing profile (ideal standard: **$>95–98\%$**).                    |
+---------------------+-----------------------------------------------------------------------------------+
| **Reproducibility** | Ability of the technique to yield identical typing results when the same isolate  |
|                     | is tested repeatedly under identical or modified laboratory conditions ($>99\%$).|
+---------------------+-----------------------------------------------------------------------------------+
| **Stability**       | Biological property ensuring the molecular marker remains invariant during in vitro|
|                     | subculturing and within reasonable epidemiologic transmission timeframes.         |
+---------------------+-----------------------------------------------------------------------------------+
| **Epidemiological** | Concordance between the molecular typing clusters and established epidemiological|
| **Concordance**     | contact tracing, geographical distribution, and outbreak timelines.               |
+---------------------+-----------------------------------------------------------------------------------+

Simpson's Index of Diversity ($D$) Formula

Discriminatory power is mathematically evaluated using Simpson's Index of Diversity ($D$): D=11N(N1)i=1sni(ni1)D = 1 - \frac{1}{N(N - 1)} \sum_{i=1}^{s} n_i(n_i - 1)

Where:

  • $N = \text{Total number of bacterial isolates in the test population}$
  • $s = \text{Total number of distinct typing patterns (strains/types) resolved}$
  • $n_i = \text{Number of isolates belonging to the } i^{\text{th}} \text{ typing pattern}$

(An index of $D = 1.0$ indicates that every single isolate in the test population is resolved as a unique type; an index of $D = 0.0$ indicates that all isolates are indistinguishable).


2. Macrorestriction & Pulsed-Field Gel Electrophoresis (PFGE)

For three decades, Pulsed-Field Gel Electrophoresis (PFGE) served as the CDC PulseNet gold standard for foodborne pathogen surveillance and hospital outbreak tracking.

                           PULSED-FIELD GEL ELECTROPHORESIS (PFGE)
                           
   1. Cell Suspension in Agarose Plugs ---> Lysozyme & Proteinase K in situ Lysis
      (Prevents mechanical DNA shearing during preparation)
                                      |
                                      v
   2. Rare-Cutting Restriction Digestion ---> Digested with XbaI, SmaI, or NotI
      (Yields 10 to 30 large macro-fragments: 30 kb to 1,000+ kb)
                                      |
                                      v
   3. CHEF Electrophoresis Box ---> Alternating 120° Electric Fields with Ramping Switch Times
      (e.g., switch times ramped from 2 to 40 seconds over 20 hours at 14°C)
                                      |
                                      v
   4. Banding Pattern Analysis ---> Comparison of 15-25 visible ethidium-stained macro-bands

The Tenover Criteria for Outbreak Strain Relatedness

In 1995, Tenover and colleagues established standardized criteria for interpreting PFGE restriction patterns during an acute epidemiologic outbreak investigation (spanning $<1–3\text{ months}$):

+---------------------------------------------------------------------------------------------------------+
|                                  THE TENOVER CRITERIA FOR PFGE INTERPRETATION                           |
+---------------------+-------------------+-----------------------------------+---------------------------+
| Category            | Number of Band    | Underlying Genetic Events         | Epidemiological           |
|                     | Differences       | (Mutations / Indels)              | Conclusion                |
+---------------------+-------------------+-----------------------------------+---------------------------+
| **Indistinguishable**| **0 band**        | **Zero genetic differences.**     | **Outbreak strain.**      |
|                     | **differences**   | Identical restriction sites       | Part of the outbreak      |
|                     | (100% concordant) | throughout the whole chromosome.  | transmission cluster.     |
+---------------------+-------------------+-----------------------------------+---------------------------+
| **Closely Related** | **2 to 3 band**   | **Consistent with 1 single        | **Probably part of the    |
|                     | **differences**   | genetic event** (e.g., a single   | **outbreak.**             |
|                     |                   | point mutation altering 1 enzyme  | Minor subtype of the      |
|                     |                   | cleavage site, or 1 small indel). | outbreak clone.           |
+---------------------+-------------------+-----------------------------------+---------------------------+
| **Possibly Related**| **4 to 6 band**   | **Consistent with 2 independent   | **Possibly part of the    |
|                     | **differences**   | genetic events** (e.g., two       | **outbreak.**             |
|                     |                   | independent mutations or indels). | Additional contact tracing|
|                     |                   |                                   | required.                 |
+---------------------+-------------------+-----------------------------------+---------------------------+
| **Different /**     | **$\ge 7$ band** | **Consistent with $\ge 3$ genetic | **NOT part of the         |
| **Unrelated**       | **differences**   | events.** Unrelated background    | **outbreak.**             |
|                     |                   | community isolate.                | Independent strain origin.|
+---------------------+-------------------+-----------------------------------+---------------------------+

The Mathematical "Three-Band Difference" Rule

Why does a single point mutation cause up to 3 band differences on a PFGE gel?

  • If a single-nucleotide point mutation destroys an internal restriction endonuclease recognition site, two smaller adjacent DNA restriction fragments merge into one single large restriction fragment.
  • The resulting gel loses two existing bands and gains one new higher-molecular-weight band—producing a total difference of 3 bands from a single genetic point mutation!
   Original Wild-Type Genome:        ---[ Site 1 ]-------------( Site 2 )-------------[ Site 3 ]---
                                         |         Fragment A       |       Fragment B      |
                                         +--------------------------+-----------------------+
                                         (Yields 2 bands on PFGE: Band A + Band B)
   
   Point Mutation Destroys Site 2:   ---[ Site 1 ]------------------------------------[ Site 3 ]---
                                         |                    Fragment C                    |
                                         +--------------------------------------------------+
                                         (Yields 1 new large band: Band C)
                                         
   NET RESULT: 2 bands disappear, 1 new band appears = 3 BAND DIFFERENCES from 1 MUTATION!

3. Sequence-Based Typing: Multi-Locus Sequence Typing (MLST)

While PFGE patterns can vary slightly due to minor electrophoresis condition differences, sequence-based methods generate unambiguous digital data that can be shared globally.

+---------------------------------------------------------------------------------------------------------+
|                                MULTI-LOCUS SEQUENCE TYPING (MLST) ARCHITECTURE                          |
+---------------------------------------------------------------------------------------------------------+
|  1. Select 6 to 8 Conserved Housekeeping Genes (Essential metabolic genes under neutral purifying selection)
|     (e.g., for E. coli: adk, fumC, gyrB, icd, mdh, purA, recA)
|                                         |
|                                         v
|  2. PCR Amplify & Sanger/NGS Sequence Internal ~450–500 bp Fragments of Each Locus
|                                         |
|                                         v
|  3. Assign Unique Integer Allele Numbers to Distinct Nucleotide Sequences
|     (e.g., adk: 53, fumC: 40, gyrB: 47, icd: 13, mdh: 36, purA: 28, recA: 29)
|                                         |
|                                         v
|  4. Generate Allelic Profile -> Defines Specific Sequence Type (ST)
|     (e.g., Allelic Profile [53-40-47-13-36-28-29] = Sequence Type ST131 - Global Pandemic Clone!)
|                                         |
|                                         v
|  5. eBURST / PHYLOViZ Analysis -> Group Related STs into Clonal Complexes (CCs)
|     (STs sharing identity at 5 of 7 or 6 of 7 loci belong to the same Clonal Complex)
+---------------------------------------------------------------------------------------------------------+

Why Housekeeping Genes are Used for MLST:

  • Neutral Evolution: Housekeeping genes encode core metabolic enzymes (e.g., adenylate kinase adk, fumarate hydratase fumC, DNA gyrase gyrB). Because mutations that alter active-site amino acids are lethally selected against, nucleotide variations accumulate slowly through silent, neutral synonymous mutations.
  • Avoidance of Hypervariable Antigenic Genes: Genes encoding surface antigens or virulence factors (such as flaA flagellin or ompA outer membrane proteins) are under intense host immune selection pressure (diversifying selection) and horizontal gene transfer. Using them for phylogenetic typing produces distorted lineage trees.
  • Major Clinical Pandemic Clones Defined by MLST:
    • Escherichia coli ST131: Global pandemic uropathogenic and bloodstream clone carrying blaCTX-M-15 ESBL and fluoroquinolone resistance.
    • Klebsiella pneumoniae ST258: Worldwide healthcare-associated clone carrying blaKPC carbapenemase.
    • Staphylococcus aureus ST8 (USA300): Predominant North American community-associated MRSA clone.

4. Single-Locus Repeat Typing & MLVA

+---------------------------------------------------------------------------------------------------------+
|                                REPEAT-BASED TYPING: SPA TYPING & MLVA                                   |
+---------------------+-----------------------------------+-----------------------------------------------+
| Methodology         | Target Locus & Biochemical Basis  | Diagnostic Utility & Applications             |
+---------------------+-----------------------------------+-----------------------------------------------+
| ***spa* Typing**    | **Polymorphic X-region** of the   | **Gold standard single-locus typing for**     |
| (*S. aureus*)       | Staphylococcal **Protein A gene** | ***S. aureus***. Rapid single-reaction Sanger  |
|                     | (***spa***), containing variable  | sequencing. Standardized Ridom nomenclature   |
|                     | 24-bp tandem repeats.             | (e.g., **t008** corresponds to ST8/USA300).   |
+---------------------+-----------------------------------+-----------------------------------------------+
| **MLVA**            | **Multiple-Locus Variable Number  | High discriminatory power for **monomorphic** |
| (Multiple-Locus     | Tandem Repeat Analysis;** PCR sizes| **bacterial pathogens** with low point-       |
| VNTR Analysis)      | 5 to 10 variable tandem repeat    | mutation rates (e.g., ***Bacillus anthracis***,|
|                     | loci via capillary electrophoresis.| ***Salmonella enterica***, *Yersinia pestis*).|
+---------------------+-----------------------------------+-----------------------------------------------+

5. Whole-Genome Sequencing (WGS): wgSNP vs. cgMLST

The clinical and public health microbiology community has transitioned from PFGE and 7-gene MLST to high-resolution Next-Generation Sequencing (NGS) platforms.

                           WHOLE-GENOME TYPING METHODOLOGIES
                           
         Raw NGS Short Reads (Illumina / Ion Torrent / PacBio / Nanopore)
                                  |
                                  +---------------------------------------+
                                  |                                       |
                                  v                                       v
          Whole-Genome SNP Pipeline (wgSNP)               Core-Genome MLST (cgMLST)
          ---------------------------------               -------------------------
          • Map reads against Reference Genome            • Assemble reads de novo into contigs
          • Call high-confidence single SNPs              • Query 1,500 to 4,000 conserved core
            across entire non-repetitive core               genes against standardized database
          • Filter recombination & prophage regions       • Assign allele numbers to every gene
          • Build maximum-likelihood SNP trees            • Build standardized allele distance trees
          • Outbreak threshold: <= 5-10 SNPs              • Database: PulseNet NextGen / NCBI Pathogens

Whole-Genome SNP Analysis (wgSNP) vs. Core-Genome MLST (cgMLST)

FeatureWhole-Genome SNP (wgSNP) AnalysisCore-Genome MLST (cgMLST / wgMLST)
Analysis BasisSingle nucleotide mutations across mapped reference genomeAllelic variation across 1,500–4,000 defined core genes
Reference DependencyHighly dependent on choice of reference genomeReference-independent; queries standardized gene allele databases
Database PortabilityLow portability; difficult to compare SNP tables across different pipelinesUltra-high portability; standardized integer allele strings universally sharable
Resolution PowerUltimate single-base resolutionNear-equivalent to wgSNP; captures gene-level indels and mutations
Outbreak ThresholdTypically $\le 5–10$ SNP differences indicates a clonal foodborne outbreakTypically $\le 5–10$ allelic differences across thousands of core loci

6. Comprehensive Comparative Synthesis of Pathogen Typing Methods

+---------------------------------------------------------------------------------------------------------+
|                               COMPARISON OF PATHOGEN TYPING TECHNOLOGIES                                |
+---------------------+-------------------+-------------------+-------------------+-----------------------+
| Typing Method       | Discriminatory    | Reproducibility   | Portability &     | Current Clinical &    |
|                     | Power ($D$)       | & Ease of Use     | Standardization   | Public Health Utility |
+---------------------+-------------------+-------------------+-------------------+-----------------------+
| **PFGE**            | High              | Moderate; labor-  | Moderate; TIFF    | Historical benchmark; |
| (Macrorestriction)  | ($D \approx 0.90–0.95$)| intensive (24–48 h)| gel images stored | widely replaced by WGS|
|                     |                   | technical setup   | in PulseNet DB    | in modern laboratories|
+---------------------+-------------------+-------------------+-------------------+-----------------------+
| **Classic MLST**    | Moderate          | High; simple PCR  | **High;** public  | Global evolutionary   |
| (7 Housekeeping)    | ($D \approx 0.85–0.90$)| and sequencing    | PubMLST databases | lineages & clonal     |
|                     |                   |                   | of integer alleles| complex surveillance  |
+---------------------+-------------------+-------------------+-------------------+-----------------------+
| ***spa* Typing**    | Moderate to High  | High; single-tube | **High;** Ridom   | Rapid, low-cost MRSA  |
| (*S. aureus*)       | ($D \approx 0.90–0.94$)| Sanger sequencing | SpaServer digital | hospital surveillance |
|                     |                   |                   | repeat sequences  | and outbreak screening|
+---------------------+-------------------+-------------------+-------------------+-----------------------+
| **MLVA**            | Very High         | High; multi-well  | Moderate to High; | Monomorphic outbreak  |
| (VNTR Sizing)       | ($D \approx 0.95–0.98$)| capillary CE sizing| allele repeat counts| investigations      |
|                     |                   |                   | in local databases| (*B. anthracis*)      |
+---------------------+-------------------+-------------------+-------------------+-----------------------+
| **WGS (wgSNP /**    | **Maximum**       | High; automated   | **Maximum;** NCBI | **Current definitive  |
| **cgMLST)**         | ($D > 0.99$)      | NGS library prep  | Pathogen Detection| reference standard**  |
|                     |                   | & bioinformatic QC| global databases  | for epidemiology      |
+---------------------+-------------------+-------------------+-------------------+-----------------------+
Loading diagram...
Molecular Epidemiology Pathogen Typing Hierarchy and Resolution
Test Your Knowledge

During an acute hospital-acquired infection outbreak investigation of Pseudomonas aeruginosa in an intensive care unit, a clinical laboratory performs Pulsed-Field Gel Electrophoresis (PFGE) on four patient isolates. When compared to the primary outbreak index pattern, isolate P-04 exhibits exactly 2 band differences. Based on the Tenover criteria, how should this isolate be categorized?

A
B
C
D
Test Your Knowledge

Why does classical Multi-Locus Sequence Typing (MLST) target internal fragments of 6 to 8 housekeeping genes rather than genes encoding cell-surface virulence factors or outer membrane antigens?

A
B
C
D
Test Your Knowledge

Why is intact bacterial genomic DNA embedded in low-melting-point agarose plugs prior to in situ cell lysis during sample preparation for Pulsed-Field Gel Electrophoresis (PFGE)?

A
B
C
D