6.2 Biostatistics: ARR, RRR, NNT, NNH, Confidence Intervals & P-values

Key Takeaways

  • Absolute Risk Reduction (ARR=∣CER−EER∣ARR = |CER - EER|) quantifies the true clinical magnitude of treatment benefit in the population, whereas Relative Risk Reduction (RRR=ARRCER=1−RRRRR = \frac{ARR}{CER} = 1 - RR) inflates clinical perception by expressing proportional reduction regardless of baseline risk.

  • Number Needed to Treat (NNT=1ARRNNT = \frac{1}{ARR}) is calculated by dividing 1 by the ARR in decimal form and must always be rounded UP to the next whole integer to prevent overestimating clinical benefit over a specified treatment timeframe.

  • Number Needed to Harm (NNH=1ARINNH = \frac{1}{ARI}) quantifies adverse event incidence from an Absolute Risk Increase (ARI=EER−CERARI = EER - CER) and must always be rounded DOWN to the previous whole integer to avoid underestimating toxic harm.

  • In hypothesis testing, Type I error (α\alpha, typically set at 0.050.05) represents the probability of falsely rejecting a true null hypothesis (false positive), whereas Type II error (β\beta, typically 0.100.10 to 0.200.20) represents failing to reject a false null hypothesis (false negative), giving statistical power (1−β1 - \beta, ≥80%\ge 80\%).

  • A 95% Confidence Interval (CI) indicates statistical significance at p<0.05p < 0.05 if it excludes the value of no effect: for difference metrics (ARRARR, mean difference), the CI must not cross 0; for ratio metrics (RRRR, OROR, HRHR), the CI must not cross 1.0.

Last updated: October 2026

Biostatistics: ARR, RRR, NNT, NNH, Confidence Intervals & P-values

Biostatistics provides the quantitative tools required to interpret therapeutic trial results, separate true pharmacological effects from random variation, and translate research data into actionable patient care decisions. For Canadian pharmacists, evaluating drug promotional materials, manufacturer claims, and clinical guidelines requires fluent mastery of risk metrics, hypothesis testing, error boundaries, and confidence interval interpretation.


Quantitative Clinical Trial Metrics: The 2x2 Table

Most binary clinical trial outcomes (such as stroke vs. no stroke, mortality vs. survival, or adverse reaction vs. no reaction) can be organized into a standard 2×22 \times 2 contingency table:

Treatment GroupPrimary Event OccurredPrimary Event Did Not OccurTotal ParticipantsEvent Rate Formula
Experimental Armaabba+ba + bExperimental Event Rate (EER)=aa+b\text{Experimental Event Rate } (EER) = \frac{a}{a+b}
Control / Placebo Armccddc+dc + dControl Event Rate (CER)=cc+d\text{Control Event Rate } (CER) = \frac{c}{c+d}

From these foundational event rates, clinical biostatisticians derive relative and absolute effect metrics.


Relative Risk, Absolute Risk Reduction, and Relative Risk Reduction

1. Relative Risk (Risk Ratio, RRRR)

Relative Risk is the ratio of the probability of an outcome occurring in the experimental group compared to the control group:

RR=EERCERRR = \frac{EER}{CER}

  • Interpretation:
    • RR=1.0RR = 1.0: There is no difference in risk between the experimental and control arms (the null value).
    • RR<1.0RR < 1.0: The experimental intervention decreases risk (protective effect).
    • RR>1.0RR > 1.0: The experimental intervention increases risk (harmful effect).

2. Absolute Risk Reduction (ARRARR)

Absolute Risk Reduction (also termed Risk Difference) represents the absolute arithmetic difference in event rates between the two groups:

ARR=∣CER−EER∣=CER−EER(when CER>EER)ARR = |CER - EER| = CER - EER \quad (\text{when } CER > EER)

ARRARR directly reflects the true clinical impact in the population, answering the question: "Out of 100 patients treated with this therapy, how many will be spared an adverse clinical event?"

3. Relative Risk Reduction (RRRRRR)

Relative Risk Reduction represents the proportional reduction in risk achieved by the experimental intervention relative to baseline control risk:

RRR=1−RR=CER−EERCER=ARRCERRRR = 1 - RR = \frac{CER - EER}{CER} = \frac{ARR}{CER}

Important

The Baseline Risk Fallacy (Why RRR Distorts Clinical Reality): Pharmaceutical marketing frequently highlights RRRRRR rather than ARRARR because RRRRRR yields impressively large percentages that remain constant regardless of whether the baseline risk is massive or negligible.

Consider two clinical trials, both reporting an identical RRR=25%RRR = 25\%:

  • High-Risk Population: CER=20.0%CER = 20.0\%, EER=15.0%EER = 15.0\%.
    • ARR=20.0%−15.0%=5.0%ARR = 20.0\% - 15.0\% = 5.0\%
    • NNT=10.05=20NNT = \frac{1}{0.05} = 20
  • Low-Risk Population: CER=0.20%CER = 0.20\%, EER=0.15%EER = 0.15\%.
    • ARR=0.20%−0.15%=0.05%ARR = 0.20\% - 0.15\% = 0.05\%
    • NNT=10.0005=2,000NNT = \frac{1}{0.0005} = 2,000

While both therapies achieve the exact same "25%25\% relative risk reduction," treating high-risk patients requires treating only 20 individuals to prevent an event, whereas treating low-risk patients requires treating 2,000 individuals. Pharmacists must always calculate ARRARR and NNTNNT to discern true clinical utility.

Number Needed to Treat (NNT) and Number Needed to Harm (NNH)

Number Needed to Treat (NNTNNT)

NNTNNT is the number of patients who must receive a specific therapeutic intervention for a defined period of time for one additional patient to experience a beneficial clinical outcome (or avoid an adverse event) compared to control:

NNT=1ARR=1CER−EERNNT = \frac{1}{ARR} = \frac{1}{CER - EER}

(Note: ARRARR must be entered in decimal form, e.g., 5%=0.055\% = 0.05)

Mandatory Rounding Convention for NNTNNT:

Always round UP to the next whole integer.

  • Mathematical Rationale: If NNTNNT calculates to 23.123.1 or 23.823.8, treating 23 patients will prevent fewer than 1 full event (23×0.042=0.96623 \times 0.042 = 0.966 events prevented). To ensure at least one full event is prevented, 24 patients must be treated. Rounding down would overestimate therapeutic benefit.
  • Temporal Requirement: NNTNNT is clinically meaningless without an explicit timeframe. An NNTNNT of 25 over 1 year represents far greater clinical benefit than an NNTNNT of 25 over 10 years.

Absolute Risk Increase (ARIARI) and Number Needed to Harm (NNHNNH)

When an intervention increases the incidence of an adverse drug event or toxicity, the difference in event rates is the Absolute Risk Increase (ARIARI):

ARI=EER−CER(when EER>CER)ARI = EER - CER \quad (\text{when } EER > CER)

Number Needed to Harm (NNHNNH) indicates the number of patients who must be exposed to the intervention for a defined period before one additional patient experiences the adverse event:

NNH=1ARI=1EER−CERNNH = \frac{1}{ARI} = \frac{1}{EER - CER}

Mandatory Rounding Convention for NNHNNH:

Always round DOWN to the previous whole integer.

  • Clinical Safety Rationale: If NNHNNH calculates to 52.852.8, treating 52 patients already exposes the cohort to the toxic threshold. To maintain conservative vigilance and avoid underestimating harm, round down to 52. Rounding up to 53 would falsely suggest that patients can tolerate higher exposure before experiencing toxicity.

Step-by-Step Worked Clinical Biostatistics Problem

Clinical Trial Scenario: A 3-year randomized, double-blind trial investigates a novel SGLT2 inhibitor versus placebo in 5,000 patients with heart failure with reduced ejection fraction (2,5002,500 patients per arm).

  • Primary Efficacy Outcome (CV Death or HF Hospitalization):
    • Placebo arm: 345345 events out of 2,5002,500 patients.
    • Treatment arm: 240240 events out of 2,5002,500 patients.
  • Primary Safety Outcome (Symptomatic Mycotic Genital Infections):
    • Treatment arm: 6868 events out of 2,5002,500 patients.
    • Placebo arm: 2020 events out of 2,5002,500 patients.

Step-by-Step Efficacy Calculations:

  1. Calculate Control and Experimental Event Rates: CER=3452,500=0.138(13.8%implication)CER = \frac{345}{2,500} = 0.138 \quad (13.8\% implication) EER=2402,500=0.096(9.6%implication)EER = \frac{240}{2,500} = 0.096 \quad (9.6\% implication)
  2. Calculate Relative Risk (RRRR): RR=EERCER=0.0960.138=0.6957≈0.70RR = \frac{EER}{CER} = \frac{0.096}{0.138} = 0.6957 \approx 0.70
  3. Calculate Absolute Risk Reduction (ARRARR): ARR=CER−EER=0.138−0.096=0.042(4.2%)ARR = CER - EER = 0.138 - 0.096 = 0.042 \quad (4.2\%)
  4. Calculate Relative Risk Reduction (RRRRRR): RRR=ARRCER=0.0420.138=0.3043(30.43%≈30.4%)RRR = \frac{ARR}{CER} = \frac{0.042}{0.138} = 0.3043 \quad (30.43\% \approx 30.4\%)
  5. Calculate NNTNNT: NNT=1ARR=10.042=23.81NNT = \frac{1}{ARR} = \frac{1}{0.042} = 23.81
    • Applying the round-up rule: NNT=24NNT = 24 over 3 years.

Step-by-Step Safety Calculations:

  1. Calculate Safety Event Rates: EERharm=682,500=0.0272(2.72%)EER_{\text{harm}} = \frac{68}{2,500} = 0.0272 \quad (2.72\%) CERharm=202,500=0.0080(0.80%)CER_{\text{harm}} = \frac{20}{2,500} = 0.0080 \quad (0.80\%)
  2. Calculate Absolute Risk Increase (ARIARI): ARI=EERharm−CERharm=0.0272−0.0080=0.0192(1.92%)ARI = EER_{\text{harm}} - CER_{\text{harm}} = 0.0272 - 0.0080 = 0.0192 \quad (1.92\%)
  3. Calculate NNHNNH: NNH=1ARI=10.0192=52.08NNH = \frac{1}{ARI} = \frac{1}{0.0192} = 52.08
    • Applying the round-down rule: NNH=52NNH = 52 over 3 years.

Clinical Synthesis: Over 3 years, for every 24 patients treated with the novel drug instead of placebo, 1 cardiovascular death or heart failure hospitalization is prevented (NNT=24NNT = 24), while 1 additional mycotic infection occurs for every 52 patients treated (NNH=52NNH = 52).


Odds Ratio (OR) vs. Hazard Ratio (HR)

Odds Ratio (OROR)

In retrospective case-control studies where total population denominator counts are unavailable, incidence cannot be measured, precluding calculation of Relative Risk. Instead, researchers compute the Odds Ratio (OROR):

OR=Odds of exposure among casesOdds of exposure among controls=a/cb/d=a×db×cOR = \frac{\text{Odds of exposure among cases}}{\text{Odds of exposure among controls}} = \frac{a / c}{b / d} = \frac{a \times d}{b \times c}

Note

The Rare Disease Assumption: When the clinical outcome of interest is rare in the general population (incidence <5%< 5\% to 10%10\%), the Odds Ratio closely approximates the Relative Risk (OR≈RROR \approx RR). However, when the outcome is common, the Odds Ratio substantially diverges and overstates the Relative Risk (exaggerating both protective and harmful effects).

Hazard Ratio (HRHR)

A Hazard Ratio is derived from survival (time-to-event) analysis, typically using Cox proportional hazards regression modeling. Unlike Relative Risk—which simply counts whether an event occurred by the end of a trial—the Hazard Ratio compares the instantaneous event rate occurring at any given point in time during follow-up, accounting for:

  1. Varying durations of patient follow-up;
  2. Censoring (participants who complete the study without an event, withdraw, or are lost to follow-up);
  3. The timing of events (detecting whether a drug delays events early even if lifetime cumulative incidence converges).
  • Interpretation:
    • HR=1.0HR = 1.0: Equivalent instantaneous event hazard in both arms (null value).
    • HR<1.0HR < 1.0: Experimental therapy reduces the hazard rate over time (protective).
    • HR>1.0HR > 1.0: Experimental therapy accelerates the hazard rate over time (hazardous).

Statistical Hypothesis Testing: Type I Error, Type II Error, and Power

Clinical trials utilize inferential statistics to test formal hypotheses regarding whether observed differences reflect true pharmacological reality or mere random sampling fluctuation.

                            THE HYPOTHESIS TESTING MATRIX

                                         TRUE STATE OF REALITY
                                ┌───────────────────────┬───────────────────────┐
                                │   H₀ is TRUE          │   H₀ is FALSE         │
                                │ (No True Difference)  │ (True Difference Exists)│
┌───────────────────────────────┼───────────────────────┼───────────────────────┤
│ REJECT H₀                     │   TYPE I ERROR (α)    │   CORRECT DECISION    │
│ (Conclude Difference Exists)  │   "False Positive"    │   POWER (1 - β)       │
├───────────────────────────────┼───────────────────────┼───────────────────────┤
│ FAIL TO REJECT H₀             │   CORRECT DECISION    │   TYPE II ERROR (β)   │
│ (Conclude No Difference)      │       (1 - α)         │   "False Negative"    │
└───────────────────────────────┴───────────────────────┴───────────────────────┘

1. The Hypotheses

  • Null Hypothesis (H0H_0): There is no true difference between the experimental treatment and control (RR=1.0RR = 1.0, ARR=0ARR = 0).
  • Alternative Hypothesis (H1H_1): A true difference exists between the treatments (RR≠1.0RR \neq 1.0, ARR≠0ARR \neq 0).

2. Type I Error (α\alpha) — "False Positive"

Occurs when researchers reject the null hypothesis when H0H_0 is actually true, falsely concluding that an ineffective drug is effective. In biomedical research, the alpha threshold is conventionally set at α=0.05\alpha = 0.05 (5%5\%), meaning researchers accept a maximum 5%5\% probability that a positive finding occurred purely by chance.

3. Type II Error (β\beta) — "False Negative"

Occurs when researchers fail to reject the null hypothesis when H0H_0 is actually false, missing a true therapeutic effect. Beta is conventionally set between β=0.10\beta = 0.10 and 0.200.20 (10%10\% to 20%20\%).

4. Statistical Power (1−β1 - \beta)

Statistical power is the probability that a trial will detect a statistically significant difference if a true difference of a given magnitude actually exists. Standard clinical trial design demands power of at least 80%80\% (1−0.201 - 0.20) or 90%90\% (1−0.101 - 0.10).

  • Key Determinants of Statistical Power:
    • Sample Size (NN): Larger sample size increases power.
    • Effect Size: Detecting a large therapeutic difference requires fewer patients than detecting a subtle difference.
    • Baseline Event Rate: In trials with rare clinical events, total event count—not just patient count—dictates power. If fewer events occur than projected, power plummets.
    • Significance Level (α\alpha): A more stringent alpha (e.g., 0.010.01 vs. 0.050.05) reduces power.

5. The pp-value: Definition and Misconceptions

The pp-value is the probability of obtaining an effect equal to or more extreme than the observed trial result, assuming that the null hypothesis is true.

  • If p<0.05p < 0.05: The result is deemed statistically significant, and H0H_0 is rejected.
  • Crucial Clinical Truths:
    • A pp-value does not indicate the probability that the hypothesis is true.
    • A pp-value does not reflect the clinical magnitude or importance of the effect. With massive sample sizes (N>50,000N > 50,000), a clinically trivial difference (e.g., systolic blood pressure reduction of 0.4 mmHg0.4\text{ mmHg}) can achieve p<0.001p < 0.001.
    • A pp-value ≥0.05\ge 0.05 does not prove that treatments are equivalent; it merely indicates that the study lacked sufficient evidence to reject the null hypothesis ("absence of evidence is not evidence of absence").

Confidence Intervals: Interpretation of Difference vs. Ratio Metrics

A 95% Confidence Interval (95% CI95\%\text{ CI}) represents the range of plausible values within which the true population effect parameter lies with 95%95\% certainty. Confidence intervals convey both statistical significance and the precision of the estimate (narrow intervals reflect high precision; wide intervals indicate low precision due to small sample size or sparse events).

                  HOW TO INTERPRET A 95% CONFIDENCE INTERVAL

 DIFFERENCE METRICS (ARR, Mean Difference)        RATIO METRICS (RR, OR, HR)
 ─────────────────────────────────────────        ──────────────────────────
 Null Value = 0                                   Null Value = 1.0

 CI: [0.02 to 0.08]                               CI: [0.65 to 0.88]
 ────●──────●──────►                              ────●──────●──────►
 0                                                0          1.0
 Excludes 0 ──► STATISTICALLY SIGNIFICANT         Excludes 1.0 ──► STATISTICALLY SIGNIFICANT
               (p < 0.05)                                         (p < 0.05)

 CI: [-0.01 to 0.06]                              CI: [0.82 to 1.15]
 ────●───0──●──────►                              ────●───1.0──●────►
         ▲                                                ▲
 Crosses 0 ──► NOT STATISTICALLY SIGNIFICANT      Crosses 1.0 ──► NOT STATISTICALLY SIGNIFICANT
               (p ≥ 0.05)                                         (p ≥ 0.05)

Difference Metrics vs. Ratio Metrics Summary:

  1. Difference Metrics (ARRARR, Absolute Risk Increase, Mean Difference in Blood Pressure or HbA1c):
    • Null Value is 00.
    • Statistically significant (p<0.05p < 0.05) if the 95% CI95\%\text{ CI} excludes 00 (both boundaries are positive or both are negative).
    • Not statistically significant (p≥0.05p \ge 0.05) if the 95% CI95\%\text{ CI} crosses 00 (lower bound is negative, upper bound is positive).
  2. Ratio Metrics (RRRR, Odds Ratio, Hazard Ratio):
    • Null Value is 1.01.0.
    • Statistically significant (p<0.05p < 0.05) if the 95% CI95\%\text{ CI} excludes 1.01.0 (both boundaries <1.0< 1.0 or both >1.0> 1.0).
    • Not statistically significant (p≥0.05p \ge 0.05) if the 95% CI95\%\text{ CI} crosses 1.01.0 (e.g., 0.850.85 to 1.201.20).

Clinical Significance vs. Statistical Significance

A statistically significant finding (p<0.05p < 0.05, CI excluding null) does not guarantee that the intervention is clinically meaningful. Clinicians must compare the confidence interval bounds against the Minimal Clinically Important Difference (MCID)—the smallest change in outcome that patients or clinicians perceive as beneficial. If a confidence interval falls entirely below the MCID, the drug's effect is clinically trivial despite statistical significance.

Test Your Knowledge

A 3-year multicentre randomized controlled trial evaluates the efficacy of a novel SGLT2 inhibitor versus placebo in preventing cardiovascular death or heart failure hospitalization in patients with heart failure with reduced ejection fraction. The trial enrolls 5,000 patients randomized equally into two arms (2,500 patients per arm). At 3 years, 345 patients in the placebo group and 240 patients in the treatment group experience the primary composite endpoint. In the safety analysis, 68 patients in the treatment group and 20 patients in the placebo group develop symptomatic mycotic genital infections. What are the Absolute Risk Reduction (ARR), Relative Risk Reduction (RRR), Number Needed to Treat (NNT), and Number Needed to Harm (NNH) over 3 years?

A

ARR = 4.2%, RRR = 43.8%, NNT = 23, NNH = 53 over 3 years

B

ARR = 4.2%, RRR = 30.4%, NNT = 24, NNH = 52 over 3 years

C

ARR = 4.2%, RRR = 30.4%, NNT = 23, NNH = 53 over 3 years

D

ARR = 4.2%, RRR = 30.4%, NNT = 24, NNH = 53 over 3 years

Test Your Knowledge

A clinical trial evaluates a novel direct oral anticoagulant (DOAC) versus adjusted-dose warfarin for stroke prevention in 14,000 patients with non-valvular atrial fibrillation over a median follow-up of 2.5 years. The primary efficacy outcome (ischemic or hemorrhagic stroke or systemic embolism) occurred at a Hazard Ratio HR = 0.79 (95% CI: 0.66 to 0.94, p = 0.008). The primary safety outcome of major gastrointestinal bleeding occurred at HR = 1.24 (95% CI: 0.96 to 1.59, p = 0.098). How should a clinical pharmacist interpret these statistical findings?

A

The DOAC demonstrates a statistically significant 21% reduction in the hazard of stroke or systemic embolism because the 95% confidence interval excludes 1.0, while the increase in major gastrointestinal bleeding is not statistically significant because its 95% confidence interval spans across 1.0.

B

The stroke prevention benefit is statistically insignificant because the 95% confidence interval fails to cross 0, whereas gastrointestinal bleeding is statistically significant because its p-value exceeds 0.05.

C

Both efficacy and safety endpoints achieve statistical significance because the trial enrolled a massive cohort of 14,000 patients, which automatically renders all point estimates definitive.

D

The DOAC shows no statistically significant efficacy benefit because the lower bound of the hazard ratio confidence interval is less than 1.0, while major gastrointestinal bleeding is proven significantly elevated because the point estimate exceeds 1.0.

Test Your Knowledge

A phase III trial investigating a new lipid-lowering agent was designed with alpha = 0.05 and an anticipated statistical power of 80% (beta = 0.20) requiring 1,200 participants to detect a 15% reduction in major adverse cardiovascular events (MACE). Due to funding termination, the study closed prematurely after recruiting only 450 participants, reporting a non-statistically significant reduction in MACE (HR = 0.84, 95% CI: 0.65 to 1.08, p = 0.17). What statistical concept best describes the primary vulnerability of this study's conclusion?

A

The trial has excess statistical power, causing it to detect clinically trivial differences as statistically significant.

B

The trial has an excessive Type I error rate (alpha), meaning there is a severe risk of a false-positive conclusion that the drug is superior to placebo for preventing cardiovascular events.

C

The trial is underpowered: a high Type II error risk means a true effect may have been missed.

D

The trial proved that the experimental drug is therapeutically equivalent to placebo because a p-value of 0.17 confirms the null hypothesis is true.

Sections you finish are checked off in the contents.