13.2 Hypothesis Testing, Power, Type I/II Errors, and Confidence Intervals

Key Takeaways

  • The null hypothesis (H0H_0) states that there is no true difference or association between study groups; the P-value is the probability of observing a test statistic as extreme as or more extreme than the observed data, assuming H0H_0 is true.

  • A Type I error (probability α\alpha, typically 0.050.05) occurs when a true null hypothesis is falsely rejected; a Type II error (probability β\beta, typically 0.100.10 to 0.200.20) occurs when a false null hypothesis is mistakenly retained.

  • Statistical power (1−β1 - \beta, target ≥80%\ge 80\%) is the probability of correctly rejecting a false null hypothesis; it is augmented by larger sample size (nn), greater anticipated effect size (Δ\Delta), higher alpha threshold (α\alpha), and lower population variance (σ2\sigma^2).

  • A 95%95\% confidence interval is a range of values compatible with the data: if a study were repeated many times, 95%95\% of such intervals would contain the true population parameter; for differences between means or proportions, an interval crossing 00 denotes non-significance (p≥0.05p \ge 0.05), whereas for ratio metrics (RR, OR, HR), an interval spanning 1.01.0 indicates lack of statistical significance.

  • Statistical significance does not equal clinical significance: very large trials can yield minute, clinically meaningless differences with p<0.001p < 0.001, while underpowered trials may fail to achieve statistical significance (p≥0.05p \ge 0.05) despite a large, life-saving clinical treatment effect.

Last updated: October 2026

13.2 Hypothesis Testing, Power, Type I/II Errors, and Confidence Intervals

Inferential statistics enables clinicians to draw valid scientific conclusions about broad patient populations using finite clinical study samples. Mastery of hypothesis testing, statistical error types, sample size determination, and confidence interval estimation is indispensable for independent critical appraisal.


1. The Inferential Framework: Null and Alternative Hypotheses

Modern biomedical research relies on the hypothetico-deductive paradigm established by Karl Popper and formalized statistically by Ronald Fisher, Jerzy Neyman, and Egon Pearson. In this framework, empirical science cannot prove a proposition conclusively true; it can only attempt to falsify a pre-specified null hypothesis.

The Null Hypothesis (H0H_0)

  • Definition: The baseline proposition that there is no true difference, treatment effect, or association between the examined study interventions in the target population.
    • Mathematical Formulation: For two continuous means, H0:μ1−μ2=0H_0: \mu_1 - \mu_2 = 0; for binary event ratios, H0:RR=1.0H_0: RR = 1.0 or OR=1.0OR = 1.0.
  • All statistical test calculations begin by assuming that H0H_0 is absolute truth.

The Alternative Hypothesis (H1H_1 or HAH_A)

  • Definition: The proposition that a genuine treatment effect, difference, or association exists between study groups.
  • Two-Tailed (Bilateral) Hypothesis: Postulates that the groups differ without pre-specifying directionality (H1:μ1≠μ2H_1: \mu_1 \neq \mu_2). This is the scientific gold standard in clinical anaesthesia trials because an intervention may prove unexpectedly superior, neutral, or harmful.
  • One-Tailed (Unilateral) Hypothesis: Postulates an effect in a single pre-specified direction (H1:μ1>μ2H_1: \mu_1 > \mu_2 or H1:μ1<μ2H_1: \mu_1 < \mu_2). One-tailed tests double the nominal significance level in the chosen tail, making statistical significance easier to achieve. However, they completely ignore the possibility of harm and are rarely acceptable in regulatory or clinical trials unless a unidirectional effect is physiologically mandated.

2. The P-value: Mathematical Definition and Clinical Fallacies

Rigorous Definition

The P-value is the conditional probability of obtaining a test statistic at least as extreme as, or more extreme than, the value observed in the study data, calculated under the assumption that the null hypothesis (H0H_0) is strictly true: P=P(Data as extreme or more extreme∣H0 is true)P = P(\text{Data as extreme or more extreme} \mid H_0\text{ is true})

If the calculated P-value falls below a pre-determined significance threshold (α\alpha, conventionally 0.050.05), the observed difference is deemed unlikely to have arisen solely through random sampling variation, and H0H_0 is rejected.

The Five P-value Fallacies Tested in Examinations

  1. The Fallacy of the Null Probability: The P-value is NOT the probability that the null hypothesis is true. P(Data∣H0)≠P(H0∣Data)P(\text{Data} \mid H_0) \neq P(H_0 \mid \text{Data}). Calculating P(H0∣Data)P(H_0 \mid \text{Data}) requires Bayesian inference incorporating prior probability.
  2. The Fallacy of Chance: A P-value of 0.040.04 does NOT mean there is a 4%4\% probability that the study results occurred by chance. The null hypothesis was already assumed to be true during test calculation.
  3. The Fallacy of Effect Magnitude: A very small P-value (e.g., p<0.0001p < 0.0001) does NOT indicate a large or clinically meaningful treatment effect. In huge cohorts, trivial physiological differences achieve extreme statistical significance.
  4. The Fallacy of Equivalence ("Absence of Evidence"): A non-significant P-value (p≥0.05p \ge 0.05) does NOT prove that the treatments are equivalent or that the null hypothesis is true. It merely indicates that the study failed to detect sufficient evidence to reject H0H_0, frequently due to inadequate sample size.
  5. The Fallacy of Replicability: 1−P1 - P does NOT represent the probability that a repeat trial will achieve statistical significance.

3. Statistical Error Typology: Type I (α\alpha) vs Type II (β\beta) Errors

When testing a statistical hypothesis, the investigator's decision to reject or retain H0H_0 can intersect with biological reality in four distinct ways.

                                BIOLOGICAL TRUTH
                           H0 True             H0 False
                     (No True Effect)    (True Effect Exists)
                   +-------------------+-------------------+
    Reject H0      |   TYPE I ERROR    | CORRECT DECISION  |
  (Claim Effect)   |    (Alpha, α)     | Statistical Power |
DECISION           |  False Positive   |     (1 - Beta)    |
                   +-------------------+-------------------+
  Fail to Reject   | CORRECT DECISION  |   TYPE II ERROR   |
        H0         |    Confidence     |    (Beta, β)      |
   (No Effect)     |    (1 - Alpha)    |  False Negative   |
                   +-------------------+-------------------+

Type I Error (α\alpha, False Positive)

  • Definition: Rejecting the null hypothesis when H0H_0 is actually true in reality; concluding that a treatment effect exists when there is no genuine biological difference.
  • Significance Level (α\alpha): The maximum acceptable probability of committing a Type I error, set a priori by the researcher (conventionally α=0.05\alpha = 0.05 or 5%5\%).
  • Clinical Consequence: Adopting an ineffective, costly, or potentially toxic drug into clinical practice under the false impression of therapeutic superiority.

Type II Error (β\beta, False Negative)

  • Definition: Failing to reject the null hypothesis when H0H_0 is actually false in reality; failing to detect a genuine clinical treatment effect that truly exists.
  • Type II Error Rate (β\beta): The probability of committing a Type II error, standardly targeted at β=0.10\beta = 0.10 (10%10\%) or β=0.20\beta = 0.20 (20%20\%).
  • Clinical Consequence: Discarding or abandoning a truly effective, life-saving therapeutic agent because a trial had inadequate statistical power to prove its efficacy.

The Inverse Trade-off

For a fixed sample size, α\alpha and β\beta are inversely linked. Lowering α\alpha (e.g., from 0.050.05 to 0.010.01) reduces the false-positive risk but inevitably increases β\beta (the false-negative risk), thereby suppressing statistical power. The only mechanism to simultaneously reduce both α\alpha and β\beta is to increase sample size (nn).


4. Statistical Power and Sample Size Calculation

Statistical Power (1−β1 - \beta)

Statistical power is the probability of correctly rejecting a false null hypothesis when a true treatment effect of a specified magnitude exists in the population: Power=1−β\text{Power} = 1 - \beta

In clinical research, the standard minimum acceptable power is 80%80\% (0.800.80), with higher-stakes trials powered to 90%90\% (0.900.90).

The Four Interdependent Pillars of Sample Size Determination

To calculate the sample size (nn) required for a trial, four variables must be specified. Modifying any one variable alters the remaining parameters:

  1. Significance Level (α\alpha): The false-positive threshold. Setting a stricter alpha (e.g., 0.010.01 instead of 0.050.05) requires a larger sample size.
  2. Statistical Power (1−β1 - \beta): Targeting higher power (e.g., 90%90\% instead of 80%80\%) requires a larger sample size.
  3. Anticipated Effect Size (Δ\Delta or δ=μ1−μ2\delta = \mu_1 - \mu_2): The magnitude of difference the trial is designed to detect. Detecting subtle, small differences requires substantially larger cohorts than detecting dramatic clinical effects (n∝1/Δ2n \propto 1 / \Delta^2).
  4. Population Variance (σ2\sigma^2 or SDSD): The background biological scatter. Greater variance obscures the treatment signal, demanding a larger sample size (n∝σ2n \propto \sigma^2).

Simplified Sample Size Formula (Continuous Means)

For a two-arm trial with equal allocation (1:11:1) comparing two independent continuous means with variance σ2\sigma^2: n=2(Zα/2+Zβ)2⋅σ2Δ2n = \frac{2 \left( Z_{\alpha/2} + Z_\beta \right)^2 \cdot \sigma^2}{\Delta^2}

Where:

  • For a two-tailed α=0.05\alpha = 0.05, Zα/2=1.96Z_{\alpha/2} = 1.96.
  • For 80%80\% power (β=0.20\beta = 0.20), Zβ=0.84Z_\beta = 0.84.
  • For 90%90\% power (β=0.10\beta = 0.10), Zβ=1.28Z_\beta = 1.28.

Multiplicity and Bonferroni Correction: When a trial evaluates multiple primary endpoints or performs repeated interim looks, the cumulative probability of at least one Type I error escalates rapidly (1−(1−α)k1 - (1 - \alpha)^k). To preserve an overall family-wise error rate of 0.050.05 across kk comparisons, the Bonferroni correction sets the significance threshold for each individual test to α∗=α/k\alpha^* = \alpha / k.


5. Confidence Intervals (CI): Derivation and Null Value Rules

While a P-value provides a binary test of hypothesis against an arbitrary cut-point, a Confidence Interval (CI) conveys both the point estimate of effect size and the degree of measurement precision.

Mathematical Formulation

For normally distributed continuous data, a 95%95\% confidence interval around a sample mean is calculated as: 95% CI=xˉ±(1.96×SEM)=xˉ±1.96×SDn95\%\text{ CI} = \bar{x} \pm (1.96 \times SEM) = \bar{x} \pm 1.96 \times \frac{SD}{\sqrt{n}}

For a 99%99\% confidence interval, the critical multiplier expands to 2.582.58.

Frequentist Interpretation

The correct frequentist interpretation is: If the study were repeated an infinite number of times under identical experimental conditions, 95%95\% of the calculated 95%95\% confidence intervals would contain the true, unknown population parameter. It does not mean that there is a 95%95\% probability that the true mean lies within the specific numbers of a single published interval (in classical frequentist statistics, the true parameter is a fixed constant, not a random variable).

The Null Hypothesis Equivalence Rules

A confidence interval serves as an immediate test of statistical significance at the corresponding α\alpha level (a 95% CI95\%\text{ CI} corresponds to α=0.05\alpha = 0.05):

                         DIFFERENCE METRICS (Null = 0)
             Statistically Significant          Non-Significant
             [-----]                            [------0------]
        <----+------+------+------+---->   <----+------+------+------+---->
            -2     -1      0     +1            -2     -1      0     +1
                     (Excludes 0)                      (Crosses 0)

                           RATIO METRICS (Null = 1.0)
             Statistically Significant          Non-Significant
             [-----]                                  [------1.0------]
        <----+------+------+------+---->   <----+------+------+------+---->
            0.2    0.6    1.0    1.4           0.2    0.6    1.0    1.4
                    (Excludes 1.0)                    (Crosses 1.0)
  1. Additive / Difference Measures (Mean Difference, Risk Difference, ARR):
    • The null value of no effect is 00.
    • If the 95% CI95\%\text{ CI} excludes 00 (e.g., mean difference in extubation time −12.4 min-12.4\text{ min} to −3.2 min-3.2\text{ min}), the finding is statistically significant (p<0.05p < 0.05).
    • If the 95% CI95\%\text{ CI} includes 00 (e.g., −2.1 min-2.1\text{ min} to +4.5 min+4.5\text{ min}), the finding is not statistically significant (p≥0.05p \ge 0.05).
  2. Multiplicative / Ratio Measures (Relative Risk, Odds Ratio, Hazard Ratio):
    • The null value of no effect is 1.01.0.
    • If the 95% CI95\%\text{ CI} excludes 1.01.0 (e.g., RR=0.65RR = 0.65, 95% CI 0.48−0.8895\%\text{ CI } 0.48 - 0.88), the finding is statistically significant (p<0.05p < 0.05).
    • If the 95% CI95\%\text{ CI} spans or touches 1.01.0 (e.g., OR=1.15OR = 1.15, 95% CI 0.92−1.4495\%\text{ CI } 0.92 - 1.44), the finding is not statistically significant (p≥0.05p \ge 0.05).

6. Statistical Significance versus Clinical Relevance

A critical objective of the EDAIC examination is distinguishing mathematical significance from bedside clinical value.

The Overpowered Trial Trap

When a study enrolls tens of thousands of participants, standard errors shrink toward zero (SEM∝1/nSEM \propto 1/\sqrt{n}). Consequently, clinically imperceptible physiological differences achieve extreme statistical significance.

  • Example: An ultra-large multicentre trial (n=40,000n = 40{,}000) evaluates a new vasopressor and finds a mean arterial pressure increase of 1.1 mmHg1.1\text{ mmHg} (p=0.0002p = 0.0002). While statistically irrefutable, an elevation of 1.1 mmHg1.1\text{ mmHg} is clinically irrelevant at the bedside and does not justify the cost or adverse risk of changing practice.

The Underpowered Trial Trap (Type II Error)

Small clinical studies frequently fail to achieve statistical significance (p≥0.05p \ge 0.05) despite a large, life-saving clinical treatment effect.

  • Example: A pilot study (n=20n = 20 per group) of a new antiemetic reduces postoperative vomiting from 40%40\% to 20%20\%. Because the sample size was tiny, the chi-squared test yields p=0.16p = 0.16. Concluding that "the antiemetic has no effect" is a severe error of interpretation; the trial was simply underpowered to reject H0H_0.

The Minimally Clinically Important Difference (MCID)

The MCID is the smallest change in an outcome measure that clinicians and patients perceive as clinically beneficial and that would mandate an active change in clinical management. Clinical trials must be powered specifically to detect the MCID rather than an arbitrarily chosen mathematical difference.

Test Your Knowledge

An investigator plans a randomized controlled trial comparing a new antiemetic against ondansetron for the prevention of postoperative nausea and vomiting (PONV). Which modification to the trial design will increase the statistical power (1 - β) of the study?

A

Decreasing the pre-specified significance level (α) from 0.05 to 0.01

B

Selecting a patient cohort with greater biological heterogeneity and higher baseline outcome variance (σ²)

C

Reducing the pre-specified targeted minimally clinically important difference (Δ) from 20% to 10% without altering sample size

D

Increasing the total sample size (n) recruited into the study groups

Test Your Knowledge

A multicentre randomized controlled trial evaluates the efficacy of prophylactic dexamethasone versus placebo in reducing 30-day surgical site infections after colorectal surgery. The study reports a Relative Risk (RR) of 0.72 with a 95% confidence interval of 0.54 to 0.96 (p = 0.024). How should this confidence interval be interpreted regarding statistical significance and clinical effect?

A

It is statistically significant at the 5% level because the 95% CI for a ratio measure excludes 1.0

B

The finding is not statistically significant because the 95% confidence interval does not contain 0

C

There is a 95% probability that the true relative risk in the global population is exactly 0.72

D

Dexamethasone reduces the absolute rate of surgical site infections by exactly 28% in every treated patient

Test Your Knowledge

In a randomized trial of a novel short-acting opioid for day-case anaesthesia, the primary outcome of discharge readiness time showed no statistically significant difference between the novel agent and remifentanil (p = 0.18). The trial enrolled 20 patients per group, though the initial power calculation indicated 85 patients per group were required to detect the targeted 15-minute difference with 80% power. Which statistical error has most likely occurred?

A

A Type I error (α), because the study falsely rejected a true null hypothesis

B

A Type II error (β): the underpowered trial failed to reject a possibly false null hypothesis

C

Selection bias, because the pre-specified significance level was set at α = 0.05 rather than α = 0.01

D

Homoscedasticity failure, because the standard error of the mean was utilized to calculate confidence intervals

Sections you finish are checked off in the contents.