18.3 Epidemiology Measures & Population Health
Key Takeaways
- Incidence = new cases / population at risk / time (flow); prevalence = existing cases / population (stock); prevalence ≈ incidence × duration when steady state holds.
- Mortality rates, case fatality, and attack rates answer different questions; attack rate is cumulative incidence in an outbreak-defined population.
- Herd immunity threshold for a perfectly mixed population approximates 1 − 1/R0; outbreak investigation follows verify → define/count → describe (time/place/person) → hypothesize → test → control → communicate.
- Primary prevention prevents disease onset; secondary detects early disease; tertiary reduces complications of established disease; risk factors are attributes increasing disease probability, while determinants of health include broader social/structural drivers.
- p < α (often 0.05) addresses Type I error under a null, not clinical importance; power = 1 − β controls Type II error; confidence intervals estimate precision of effect sizes and whether null values are excluded.
18.3 Epidemiology Measures & Population Health
Quick Answer: Incidence = new cases over time; prevalence = existing cases at a time (≈ I × duration). Attack rate = ill / exposed in an outbreak. Herd immunity threshold ≈ 1 − 1/R0. Primary/secondary/tertiary prevention map to before disease / early detection / limit disability. p-values ≠ effect size; power = 1 − β; CIs show precision and null exclusion.
This section ties classical epidemiology measures to prevention, outbreak response, disparities framing, and the inferential statistics tested alongside biostats on CBSE/PBLI-style items.
Incidence vs Prevalence
| Measure | Definition | Units / form | Rises when… |
|---|---|---|---|
| Incidence (cumulative) | New cases / population at risk over a period | Dimensionless proportion over time window | More new disease |
| Incidence rate (density) | New cases / person-time at risk | Cases per person-year | More new disease or better ascertainment |
| Prevalence (point) | Existing cases / population at a time | Proportion | Higher incidence or longer duration (better survival, chronicity) |
Relationship (steady state, low disease): Prevalence ≈ incidence × average duration.
Worked example 1 — Incidence and prevalence
City of 100,000 with no disease at t = 0. Over 1 year, 500 new cases of a chronic condition develop; none die or leave; all remain diseased.
- Annual cumulative incidence ≈ 500/100,000 = 0.5%
- End-of-year prevalence ≈ 500/100,000 = 0.5% (first year)
If average duration becomes 10 years with steady 500 new cases/year, prevalence approaches 500 × 10 / 100,000 = 5%. Improved treatment that prolongs life with disease increases prevalence even if incidence is flat—classic exam trap.
Worked example 2 — Person-time rate
200 patients followed: 50 for 1 year, 100 for 2 years, 50 for 4 years without the event until 12 events occur.
Person-years = 50·1 + 100·2 + 50·4 = 50 + 200 + 200 = 450
Incidence rate = 12/450 = 0.0267 per person-year (≈ 2.67 per 100 person-years)
Mortality Rates and Related Metrics
| Metric | Formula (teaching form) | Use |
|---|---|---|
| Crude mortality rate | Deaths / mid-year population / time | Overall death burden |
| Cause-specific mortality | Deaths from cause / population / time | Burden of one disease |
| Age-specific mortality | Deaths in age band / population in band | Compare age groups |
| Case fatality rate (CFR) | Deaths from disease / cases of disease | Severity among the ill |
| Proportionate mortality | Deaths from cause / all deaths | Share of deaths, not risk |
CFR ≠ cause-specific mortality. CFR uses cases in the denominator; mortality uses population.
Worked example 3 — CFR vs mortality
Population 1,000,000; 2,000 myocardial infarctions; 400 MI deaths in a year.
- CFR = 400/2000 = 20%
- Cause-specific MI mortality = 400/1,000,000 = 0.04% per year (40 per 100,000)
Attack Rate
Attack rate is cumulative incidence in a clearly defined exposed group during an outbreak window:
Attack rate = number who became ill / number exposed (at risk)
Often expressed as percent. Food-specific attack rates compare foods at a picnic.
Worked example 4 — Outbreak table
| Ate potato salad | Ill | Well | Total | Attack rate |
|---|---|---|---|---|
| Yes | 45 | 15 | 60 | 45/60 = 75% |
| No | 10 | 40 | 50 | 10/50 = 20% |
Risk ratio = 0.75/0.20 = 3.75. Potato salad is the leading hypothesis for the vehicle. Secondary attack rate among household contacts uses new cases among contacts / susceptible contacts.
Herd Immunity Threshold (Concept)
Basic reproduction number R0 = average secondary cases from one infectious person in a fully susceptible population.
Effective reproduction number Rt falls as immunity rises.
With homogeneous mixing and solid immunity, herd immunity threshold (HIT):
HIT ≈ 1 − 1/R0
Worked example 5
If R0 = 5, HIT ≈ 1 − 1/5 = 0.80 (80%) immune needed so each case produces <1 secondary case on average.
If R0 = 15 (highly transmissible agent), HIT ≈ 93%.
Imperfect vaccines require higher coverage: roughly HIT / vaccine effectiveness.
Real populations are not perfectly mixed; clustering and waning immunity complicate estimates—but the formula is standard exam material.
Outbreak Investigation Steps
A practical sequence (order may overlap):
- Prepare / verify diagnosis — confirm the disease is real (lab, clinical criteria); rule out pseudo-outbreaks.
- Confirm outbreak — cases exceed expected baseline.
- Case definition — person, place, time, clinical/lab criteria (possible/probable/confirmed).
- Find and count cases — active surveillance, line list.
- Descriptive epidemiology — time (epidemic curve), place (spot map), person (age, sex, exposures).
- Generate hypotheses — source, mode, vehicle, risk factors.
- Test hypotheses — analytic epidemiology (cohort/case-control attack rates), environmental/lab testing.
- Implement control and prevention — remove source, isolate, vaccinate, educate—often start control before all analytics finish when risk is high.
- Communicate findings — public, clinicians, agencies.
- Maintain surveillance — ensure outbreak ends; evaluate response.
Epidemic curve shapes: point source (single peak), continuous common source (prolonged plateau), propagated (serial peaks at incubation intervals).
Levels of Prevention
| Level | Goal | Examples |
|---|---|---|
| Primary | Prevent disease onset | Immunization, smoking never-start, seat belts, clean water, PrEP |
| Secondary | Detect early / preclinical disease to improve outcomes | Pap smear, mammography, BP screening, newborn metabolic screen |
| Tertiary | Reduce complications/disability of established disease | Cardiac rehab post-MI, glycemic control to prevent retinopathy, stroke rehabilitation |
Primordial prevention (sometimes tested) targets underlying social/environmental conditions that give rise to risk factors (e.g., urban design reducing default sedentary life).
Exam trap: treating severe disease complications is tertiary, not primary—even if heroic.
Risk Factors vs Determinants of Health
- Risk factor: measurable attribute/exposure associated with increased probability of disease (hypertension → stroke; smoking → lung cancer). May be causal or marker.
- Determinants of health: broader set of personal, social, economic, and environmental factors that influence health status—income, education, housing, racism/discrimination, access to care, built environment, health behaviors, genetics.
Health disparities (basic-science framing on CBSE): systematic, potentially avoidable differences in health outcomes between groups defined by social advantage. Mechanisms taught at the interface of biology and environment include chronic stress physiology (HPA/sympathetic activation), differential toxic exposures, nutrition access, infectious exposure density, and unequal treatment access—not "biological race as destiny." Disparities questions often ask which intervention level or structural factor best explains a gradient after individual risk factors are listed.
Statistical Significance vs Clinical Significance
| Concept | Meaning |
|---|---|
| Null hypothesis (H0) | Usually "no association / no difference" |
| p-value | Probability of data as extreme as observed if H0 true (with model assumptions); not P(H0 true), not effect size |
| α (Type I error rate) | Prespecified false-positive threshold (commonly 0.05) |
| Statistically significant | p < α (by convention)—may be tiny and meaningless clinically if n is huge |
| Clinical significance | Effect size large enough to matter to patients/clinicians (ARR, NNT, QoL) |
Worked example 6 — Significant but trivial
Drug lowers systolic BP by 0.5 mm Hg; n = 100,000; p = 0.001. Statistically significant, clinically negligible for individuals. Conversely, a 15 mm Hg drop in a small pilot with p = 0.08 may be clinically promising but not "significant" at α = 0.05—need more power, not dismissal of the effect size.
Type I Error, Type II Error, and Power
| H0 true | H0 false | |
|---|---|---|
| Reject H0 | Type I error (false positive), probability α | Correct detection (power) |
| Fail to reject H0 | Correct | Type II error (false negative), probability β |
Power = 1 − β = probability of detecting a true effect of a specified size.
Power rises with: larger effect size, larger sample size, higher α, lower variance, more precise measurements, one-sided tests (when appropriate).
Worked example 7 — Interpreting β
Study designed with 80% power (β = 0.20) at α = 0.05 to detect ARR of 5%. If the true ARR is 5%, there is still a 20% chance of a nonsignificant result. A negative underpowered study does not prove no effect.
Confidence Intervals
A 95% confidence interval (CI) for a mean difference or RR is a range from a procedure that, in repeated sampling, covers the true parameter 95% of the time (frequentist teaching). Practical exam uses:
- Width → precision (narrower = more precise, often larger n)
- Whether null is inside → for difference, null is 0; for ratio, null is 1. If 95% CI for RR is 0.6–0.9, excludes 1 → consistent with statistical significance at α = 0.05 (two-sided) under standard correspondence
- Clinical bounds → even if CI excludes null, if entire CI is within a trivial effect range, clinical importance remains doubtful
Worked example 8 — CI reading
RR = 1.20; 95% CI 0.95–1.52 → not statistically significant at 0.05; compatible with modest harm, null, or modest benefit—imprecise.
RR = 0.70; 95% CI 0.65–0.76 → precise protective association; still check absolute risks for decision-making.
Integration for CBSE Vignettes
- Chronic disease more prevalent despite flat incidence → longer duration/survival.
- Outbreak picnic → attack rates by food; highest AR and AR ratio drive hypothesis.
- "What percent must be immune?" → 1 − 1/R0.
- Vaccine before exposure → primary prevention; mammography → secondary; post-stroke rehab → tertiary.
- Tiny p with tiny ARR → emphasize clinical vs statistical significance.
- Nonsignificant trial with n = 20 → suspect low power, not proof of equivalence.
- RR 0.5 (0.4–0.6) → significant and relatively precise; still compute NNT when rates given.
Population measures describe burden and spread; prevention levels organize interventions; inferential statistics police how we claim knowledge. Together they complete the 4–6% biostats/epidemiology band on the CBSE.
A stable population has incidence of a disease of 2 per 1,000 per year and average disease duration of 10 years. What is the approximate prevalence under the steady-state relationship prevalence ≈ incidence × duration?
For an infection with R0 = 4 in a freely mixing population with lifelong solid immunity after infection or perfect vaccination, what is the approximate herd immunity threshold?
A large RCT reports a 0.3 mm Hg blood-pressure difference favoring drug Y over placebo (p = 0.01). Which interpretation is most accurate?