1.4 Evidence-Based Medicine, Study Design, & Clinical Statistics

Key Takeaways

  • Systematic reviews of randomized controlled trials (RCTs) represent the highest level of evidence for therapeutic interventions.
  • Sensitivity rules OUT disease when negative (SnNOut); Specificity rules IN disease when positive (SpPIn).
  • Positive and Negative Predictive Values (PPV/NPV) depend heavily on disease prevalence, whereas Likelihood Ratios (LR+/LR-) are independent of prevalence.
  • Absolute Risk Reduction (ARR) dictates the Number Needed to Treat (NNT = 1 / ARR); NNT calculations must always be rounded UP to the nearest integer.
  • Screening programs are vulnerable to lead-time bias (earlier detection mistaken for prolonged survival) and length-time bias (disproportionate detection of indolent cases).
Last updated: July 2026

1.4 Evidence-Based Medicine, Study Design, & Clinical Statistics

EBM Core Principle: Clinical decision-making on the MCCQE Part I requires integrating the best available research evidence with clinical expertise and patient values. Candidates must evaluate study designs, calculate diagnostic and therapeutic statistics, and identify methodological biases.


Hierarchy of Evidence & Study Designs

Medical evidence is structured hierarchically based on susceptibility to bias:

                  ▲  [ Level 1: Systematic Reviews & Meta-Analyses ]
                 ╱ ╲
                ╱   ╲ [ Level 2: Randomized Controlled Trials (RCTs) ]
               ╱     ╲
              ╱       ╲ [ Level 3: Cohort Studies (Prospective/Retrospective) ]
             ╱         ╲
            ╱           ╲ [ Level 4: Case-Control Studies (Retrospective) ]
           ╱             ╲
          ╱               ╲ [ Level 5: Cross-Sectional Studies / Case Series ]
         └─────────────────┘

Epidemiological Study Design Summary

Study DesignTemporal DirectionKey MetricPrimary AdvantageMajor Limitation
Randomized Controlled Trial (RCT)ProspectiveRelative Risk (RR) / ARREstablishes causality; minimizes confounding via randomization.Expensive; ethical constraints.
Cohort StudyProspective or RetrospectiveRelative Risk (RR) / IncidenceExcellent for rare exposures; assesses temporal sequence.Vulnerable to loss to follow-up and confounding.
Case-Control StudyRetrospectiveOdds Ratio (OR)Excellent for rare diseases; fast and inexpensive.Highly vulnerable to recall bias and selection bias.
Cross-Sectional StudySingle point in timePrevalenceFast; generates hypotheses.Cannot establish causality or temporal sequence.

Diagnostic Testing Metrics & The 2x2 Table

Diagnostic testing performance is evaluated using a standard 2x2 contingency table comparing test results against a gold-standard reference.

Disease Present ($D+$)Disease Absent ($D-$)Total
Test Positive ($T+$)True Positive ($TP$)False Positive ($FP$)$TP + FP$
Test Negative ($T-$)False Negative ($FN$)True Negative ($TN$)$FN + TN$
Total$TP + FN$$FP + TN$$N$

Diagnostic Test Formulas

  • Sensitivity (Sn): $\frac{TP}{TP + FN}$ — Probability that a test is positive given the disease is present. High sensitivity tests have few false negatives (SnNOut: Negative test rules OUT disease). Ideal for screening.
  • Specificity (Sp): $\frac{TN}{TN + FP}$ — Probability that a test is negative given the disease is absent. High specificity tests have few false positives (SpPIn: Positive test rules IN disease). Ideal for confirmation.
  • Positive Predictive Value (PPV): $\frac{TP}{TP + FP}$ — Probability of having disease if test is positive. PPV increases as disease prevalence increases.
  • Negative Predictive Value (NPV): $\frac{TN}{TN + FN}$ — Probability of being disease-free if test is negative. NPV decreases as disease prevalence increases.
  • Positive Likelihood Ratio ($LR+$): $\frac{\text{Sensitivity}}{1 - \text{Specificity}} = \frac{\text{TP Rate}}{\text{FP Rate}}$. An $LR+ > 10$ strongly confirms disease.
  • Negative Likelihood Ratio ($LR-$): $\frac{1 - \text{Sensitivity}}{\text{Specificity}} = \frac{\text{FN Rate}}{\text{TN Rate}}$. An $LR- < 0.1$ strongly rules out disease.
  • Key Distinction: Likelihood Ratios are independent of disease prevalence.

Interventional & Risk Statistics Formulas

When evaluating therapeutic clinical trials, risks must be expressed in absolute and relative terms:

  • Risk in Experimental Group ($R_e$): $\frac{a}{a + b}$
  • Risk in Control Group ($R_c$): $\frac{c}{c + d}$
  • Relative Risk (RR): $\frac{R_e}{R_c}$
  • Relative Risk Reduction (RRR): $\frac{R_c - R_e}{R_c} = 1 - RR$
  • Absolute Risk Reduction (ARR): $|R_c - R_e|$
  • Number Needed to Treat (NNT): $\frac{1}{\text{ARR}} = \frac{1}{|R_c - R_e|}$
    • Rule: Always round UP to the next whole integer (e.g., $NNT = 14.2 \rightarrow 15$).
  • Number Needed to Harm (NNH): $\frac{1}{\text{ARI}}$ (where ARI = Absolute Risk Increase).
    • Rule: Always round DOWN to the next whole integer.

Biases, Confounding, & Screening Fallacies

1. Selection & Information Biases

  • Selection Bias: Non-random sampling leading to a study sample unrepresentative of the target population (e.g., Berkson bias in hospitalized patients).
  • Recall Bias: Patients with adverse outcomes recall past exposures more intensely than controls (major threat in case-control studies).
  • Hawthorne Effect: Participants alter behavior because they know they are being observed.

2. Screening Program Fallacies

  • Lead-Time Bias: Screening detects a disease earlier in its course without altering the natural history or time of death, creating an illusion of prolonged survival.
  • Length-Time Bias: Screening preferentially detects slow-growing, indolent lesions with a favorable prognosis, while missing rapidly progressive cases, creating an illusion of improved treatment efficacy.

3. Confounding vs Effect Modification

  • Confounding: A third variable linked to both exposure and outcome that distorts the true association. Can be controlled by stratification, matching, or multivariable regression.
  • Effect Modification: A third variable that naturally alters the magnitude of effect across strata (e.g., a drug works in men but not women). It is a biological reality, not a bias, and should be reported rather than controlled.

Common Exam Traps & Clinical Scenarios

⚠️ EXAM TRAP: Misinterpreting RRR vs ARR

Pharmaceutical trials often highlight a "50% reduction in mortality" (Relative Risk Reduction). However, if baseline control risk is 2% and treatment risk is 1%, the Absolute Risk Reduction (ARR) is only 1% ($0.02 - 0.01 = 0.01$). The NNT is $1 / 0.01 = 100$. Always calculate ARR and NNT on the exam to determine true clinical impact!

🚨 CLINICAL SCENARIO CALLOUT

Vignette: In a randomized controlled trial of 1,000 patients with heart failure, 50 out of 500 patients receiving standard care died over 3 years (10%), compared to 25 out of 500 patients receiving a novel SGLT2 inhibitor (5%).

Step-by-Step Calculations:

  1. Control Risk ($R_c$): $50 / 500 = 0.10$ (10%).
  2. Experimental Risk ($R_e$): $25 / 500 = 0.05$ (5%).
  3. Absolute Risk Reduction (ARR): $0.10 - 0.05 = 0.05$ (5%).
  4. Number Needed to Treat (NNT): $1 / 0.05 = 20$.
  • Interpretation: You need to treat 20 heart failure patients with the SGLT2 inhibitor for 3 years to prevent 1 additional death.
Loading diagram...
Diagnostic Testing & 2x2 Contingency Matrix
Test Your Knowledge

A new therapeutic agent reduces the 5-year risk of stroke from 8% in the control group to 4% in the treatment group. What is the Number Needed to Treat (NNT) to prevent one stroke over 5 years?

A
B
C
D
Test Your Knowledge

Which epidemiological study design selects participants based on the presence or absence of a disease outcome and retrospectively evaluates past exposures, generating an Odds Ratio (OR)?

A
B
C
D
Test Your Knowledge

An implementation of a new low-dose CT screening program for lung cancer identifies early-stage indolent tumors that would never have caused symptoms during the patients' natural lifespans, creating an artificial improvement in survival statistics. Which bias is demonstrated?

A
B
C
D