14.3 Statistical Significance, Risk Measures & Evidence-Based Medicine
Key Takeaways
- Absolute Risk Reduction (ARR = Risk_control - Risk_treatment) determines the Number Needed to Treat (NNT = 1 / ARR), rounding up to the nearest whole integer.
- Relative Risk (RR = Risk_exposed / Risk_unexposed) measures strength of association in cohort studies; Relative Risk Reduction (RRR = 1 - RR = ARR / Risk_control).
- Type I error (alpha, typically set at 0.05) is rejecting a true null hypothesis (false positive); Type II error (beta) is failing to reject a false null hypothesis (false negative).
- Statistical Power (1 - beta, target >= 0.80) increases with larger sample size, larger effect size, lower measurement variance, and higher significance threshold (alpha).
- Confidence intervals (95% CI) that include 1.0 for ratios (RR, OR, HR) or 0.0 for mean differences indicate lack of statistical significance at the alpha = 0.05 level.
14.3 Statistical Significance, Risk Measures & Evidence-Based Medicine
Measures of Risk & Clinical Impact
In clinical trial appraisal, quantifying treatment effect requires understanding both relative and absolute risk metrics. Relative measures express proportional change, whereas absolute measures incorporate baseline disease frequency, providing the true clinical impact of an intervention.
Event Rates
- Experimental Event Rate (EER) or $I_e$: Incidence of the outcome in the intervention group $= \frac{a}{a + b}$.
- Control Event Rate (CER) or $I_c$: Incidence of the outcome in the control group $= \frac{c}{c + d}$.
Relative Risk (RR) and Relative Risk Reduction (RRR)
- Relative Risk (RR): Ratio of disease incidence in the treated/exposed group to incidence in the control group:
- $RR = 1.0$: No difference in risk between groups.
- $RR < 1.0$: Intervention reduces risk (protective effect).
- $RR > 1.0$: Intervention increases risk (harmful effect).
- Relative Risk Reduction (RRR): Proportional reduction in risk achieved by the intervention:
Absolute Risk Reduction (ARR) & Number Needed to Treat (NNT)
- Absolute Risk Reduction (ARR): The true difference in absolute disease risk between control and intervention groups:
- Number Needed to Treat (NNT): The number of patients who must receive the treatment for one additional patient to experience a beneficial outcome (or prevent one adverse event):
- Rounding Rule: Always round NNT UP to the next whole integer to avoid overestimating clinical benefit (e.g., $NNT = 14.2 \rightarrow 15$).
Number Needed to Harm (NNH)
When an exposure or therapy increases adverse events, the Absolute Risk Increase (ARI) is $EER - CER$. The Number Needed to Harm (NNH) represents the number of patients exposed to the risk factor for one additional patient to experience harm:
- Rounding Rule: Always round NNH DOWN to the next whole integer to avoid underestimating risk of harm.
| Metric | Formula | Clinical Interpretation |
|---|---|---|
| Relative Risk (RR) | $EER / CER$ | Proportional risk ratio; independent of baseline incidence |
| Relative Risk Reduction (RRR) | $(CER - EER) / CER$ | Percentage reduction in risk relative to baseline control rate |
| Absolute Risk Reduction (ARR) | $CER - EER$ | True absolute difference in disease incidence between arms |
| Number Needed to Treat (NNT) | $1 / ARR$ | Patients treated to prevent 1 event; round UP |
| Number Needed to Harm (NNH) | $1 / (EER - CER)$ | Patients exposed for 1 adverse event to occur; round DOWN |
Hypothesis Testing: Errors & Power
Hypothesis testing evaluates whether observed sample differences reflect true population effects or mere random chance.
- Null Hypothesis ($H_0$): Statement of no difference or no association between groups.
- Alternative Hypothesis ($H_1$): Statement that a true difference or association exists.
| Study Conclusion | Real-World Truth: $H_0$ True (No Effect) | Real-World Truth: $H_1$ True (Real Effect) |
|---|---|---|
| Reject $H_0$ (Find Effect) | Type I Error ($\alpha$) (False Positive) | Correct Decision ($1 - \beta$) (Power) |
| Fail to Reject $H_0$ (No Effect) | Correct Decision ($1 - \alpha$) | Type II Error ($\beta$) (False Negative) |
Type I Error ($\alpha$)
A Type I error occurs when researchers reject the null hypothesis when $H_0$ is actually true (finding a false positive difference). The significance level $\alpha$ (typically set at $0.05$ or $5%$) is the maximum acceptable probability of committing a Type I error. The calculated $p$-value represents the probability of obtaining results at least as extreme as observed, assuming $H_0$ is true. If $p \le \alpha$, $H_0$ is rejected.
Type II Error ($eta$) & Statistical Power
A Type II error occurs when researchers fail to reject $H_0$ when a true treatment difference exists (a false negative study).
- Statistical Power ($1 - \beta$): The probability that a study will correctly detect a true effect of a specified size. Target power in clinical trials is usually $\ge 0.80$ ($80%$).
- Determinants of Power:
- Sample Size ($n$): Larger sample size reduces standard error, directly increasing power.
- Effect Size: Larger true differences between groups are easier to detect, increasing power.
- Measurement Variance ($\sigma^2$): Lower data scatter/SD increases power.
- Significance Level ($\alpha$): Increasing $\alpha$ (e.g., $0.01 \rightarrow 0.05$) increases power (though increases Type I error risk).
Confidence Intervals & Statistical Significance
A 95% Confidence Interval (CI) provides a range of plausible values for a population parameter, constructed such that $95%$ of such intervals contain the true parameter value.
Determining Statistical Significance from CIs
- For Ratio Metrics (RR, Odds Ratio, Hazard Ratio):
- The null value of no effect is 1.0.
- If the 95% CI includes 1.0 (e.g., $RR = 0.82$, $95%\text{ CI: } 0.65 - 1.08$), the result is not statistically significant ($p > 0.05$).
- If the 95% CI excludes 1.0 (e.g., $RR = 0.74$, $95%\text{ CI: } 0.58 - 0.94$), the result is statistically significant ($p < 0.05$).
- For Mean Differences / Absolute Risk Differences (ARR):
- The null value of no effect is 0.0.
- If the 95% CI includes 0.0 (e.g., Mean Diff $= -1.4$, $95%\text{ CI: } -3.2 \text{ to } +0.4$), the result is not statistically significant ($p > 0.05$).
- If the 95% CI excludes 0.0 (e.g., Mean Diff $= -3.8$, $95%\text{ CI: } -5.6 \text{ to } -2.0$), the result is statistically significant ($p < 0.05$).
Selecting the Correct Statistical Test
Selecting appropriate statistical tests requires identifying the variable types (continuous vs. categorical) and group arrangements.
Statistical Test Selection Guide
- Comparing 2 Independent Group Means:
- Two-Sample $t$-test (Unpaired $t$-test): Compares means of a continuous dependent variable across 2 categorical groups (e.g., mean blood pressure in treatment vs. placebo).
- Comparing $\ge 3$ Independent Group Means:
- ANOVA (Analysis of Variance): Compares continuous means across 3 or more categorical groups (e.g., mean HbA1c across placebo, low-dose, and high-dose drug arms).
- Comparing Paired Continuous Measurements:
- Paired $t$-test: Compares continuous measurements before and after an intervention in the same individuals (e.g., pre- vs post-treatment serum cholesterol).
- Comparing Categorical Proportions:
- Chi-Square ($\chi^2$) Test: Evaluates associations between 2 categorical variables (e.g., proportion of smokers vs non-smokers who develop stroke).
- Fisher Exact Test: Used instead of Chi-Square when expected cell counts in a $2 \times 2$ table are small ($< 5$).
- Correlation & Regression:
- Pearson Correlation Coefficient ($r$): Measures linear association between 2 continuous variables (ranges from $-1.0$ to $+1.0$).
- Linear Regression: Predicts a continuous outcome from continuous/categorical predictors.
- Logistic Regression: Predicts a binary categorical outcome (e.g., dead/alive) from predictors.
| Independent Variable | Dependent Variable | Number of Groups | Recommended Statistical Test |
|---|---|---|---|
| Categorical | Continuous | 2 Independent groups | Two-Sample (Unpaired) $t$-test |
| Categorical | Continuous | $\ge 3$ Independent groups | ANOVA (Analysis of Variance) |
| Categorical | Continuous | Paired (Pre/Post in same subject) | Paired $t$-test |
| Categorical | Categorical | 2 or more groups (cell counts $\ge 5$) | Chi-Square ($\chi^2$) test |
| Categorical | Categorical | 2 groups (cell counts $< 5$) | Fisher Exact test |
| Continuous | Continuous | N/A | Pearson Correlation ($r$) / Linear Regression |
Evidence-Based Medicine Decision Algorithm
- Calculate Clinical Risk Reduction:
- Given $EER$ and $CER$ $\rightarrow$ Calculate $ARR = CER - EER$, then $NNT = \frac{1}{ARR}$ (round UP).
- Evaluate 95% Confidence Interval:
- Ratio metric $\rightarrow$ Check if interval crosses 1.0. If crossed $\rightarrow p > 0.05$.
- Difference metric $\rightarrow$ Check if interval crosses 0.0. If crossed $\rightarrow p > 0.05$.
- Choose Statistical Test:
- Continuous outcome across 2 groups $\rightarrow$ $t$-test.
- Continuous outcome across $\ge 3$ groups $\rightarrow$ ANOVA.
- Categorical outcome across groups $\rightarrow$ Chi-Square.
A randomized controlled trial evaluates a new SGLT2 inhibitor in patients with heart failure with reduced ejection fraction. Over a 3-year follow-up period, the primary composite endpoint of cardiovascular death or heart failure hospitalization occurs in 12% of patients in the placebo group and 8% of patients in the treatment group. What is the Number Needed to Treat (NNT) to prevent one primary composite endpoint over 3 years?
A clinical trial compares four distinct antihypertensive regimens regarding their effect on reducing stroke risk. The reported Relative Risks (RR) and 95% Confidence Intervals (CI) for stroke prevention compared to placebo are as follows: Regimen A: RR 0.65 (95% CI: 0.42–0.98); Regimen B: RR 0.70 (95% CI: 0.48–1.02); Regimen C: RR 0.85 (95% CI: 0.72–0.99); Regimen D: RR 0.90 (95% CI: 0.78–1.04). Which antihypertensive regimens fail to demonstrate a statistically significant reduction in stroke risk at the alpha = 0.05 level?
A clinical trial evaluates the efficacy of three different oral hypoglycemic agents (Drug A, Drug B, and Drug C) versus placebo on lowering hemoglobin A1c. A total of 400 patients are randomized equally into four treatment groups, and the mean reduction in hemoglobin A1c (a continuous numeric variable) is measured for each group after 24 weeks. Which statistical test is most appropriate to evaluate whether a significant difference exists among the four group means?