15.1 Biostatistics & Hypothesis Testing: Power, Errors & Interpretations
Key Takeaways
Continuous variables require verification of normal Gaussian distribution (summarized by mean and standard deviation) versus skewed non-parametric distribution (summarized by median and interquartile range) before selecting statistical tests.
Hypothesis testing balances Type I error (alpha, false-positive threshold, standard 0.05) against Type II error (beta, false-negative rate, standard 0.10–0.20), where statistical power (1 - beta, standard 80%–90%) depends on sample size, alpha, effect size, and variance.
A p-value measures the probability of obtaining results at least as extreme as observed assuming the null hypothesis is true; it does not measure effect size, precision, or clinical relevance (Minimal Clinically Important Difference [MCID]).
The 95% confidence interval defines the range containing the true population parameter: difference metrics (mean difference, ARR) are non-significant if they cross 0, whereas ratio metrics (RR, OR, HR) are non-significant if they cross 1.0.
Non-inferiority trials evaluate whether an investigational drug is not clinically worse than standard care by more than a predefined margin (delta); non-inferiority must be demonstrated in both Per-Protocol (PP) and Intention-to-Treat (ITT) populations.
15.1 Biostatistics & Hypothesis Testing: Power, Errors & Interpretations
Note
Independent study resource provided by OpenExamPrep. Content is organized around official Board of Pharmacy Specialties (BPS) Emergency Medicine Pharmacy examination specifications.
Variable Taxonomy & Measurement Scales
In emergency medicine pharmacotherapy, choosing an appropriate statistical test begins with accurate variable classification. Analyzing a clinical variable using the wrong statistical model invalidates trial findings and distorts bedside clinical translation.
TAXONOMY OF CLINICAL VARIABLES
┌─────────────────────────────────────────────────────────────────────────────┐
│ 1. Categorical / Qualitative Variables │
│ ├─ Nominal: Unordered mutually exclusive categories │
│ │ ├─ Dichotomous (Binary): Survival to discharge (yes/no), ROSC (yes/no)│
│ │ └─ Polychotomous: Blood type (A/B/AB/O), Toxidrome classification │
│ └─ Ordinal: Ranked categories with non-uniform quantitative intervals │
│ └─ Glasgow Coma Scale (3–15), RASS (-5 to +4), ESI Triage Level (1–5)│
├─────────────────────────────────────────────────────────────────────────────┤
│ 2. Continuous / Quantitative Variables │
│ ├─ Interval: Ordered scale with equal intervals, arbitrary zero point │
│ │ └─ Temperature (°C or °F) │
│ └─ Ratio: Ordered scale with equal intervals and an absolute true zero │
│ └─ Blood pressure (mmHg), Serum lactate (mmol/L), Weight (kg), Time │
└─────────────────────────────────────────────────────────────────────────────┘
Nominal Variables
Nominal variables describe categorical attributes lacking inherent quantitative ranking:
- Dichotomous (Binary): Exactly two possible categories. Examples include return of spontaneous circulation (ROSC achieved vs not achieved), in-hospital mortality (survived vs died), 90-day functional independence (modified Rankin Scale [mRS] 0–2 vs 3–6), and acute kidney injury occurrence (yes vs no).
- Polychotomous (Multinomial): Three or more unordered categories. Examples include ABO blood groups, primary shock etiology (septic, cardiogenic, hypovolemic, obstructive, neurogenic), and initial cardiac arrest rhythm (VF/pVT, asystole, PEA).
Ordinal Variables
Ordinal variables consist of ranked categories where the relative order is clinically meaningful, but the quantitative distance between consecutive ranks is unequal, inconsistent, or arbitrary:
- Emergency Medicine Examples: Glasgow Coma Scale (GCS 3–15), Richmond Agitation-Sedation Scale (RASS -5 to +4), Emergency Severity Index (ESI Levels 1–5), Numeric Pain Rating Scale (0–10), and the National Institutes of Health Stroke Scale (NIHSS 0–42).
- Biostatistical Rule: Ordinal data cannot be analyzed using standard arithmetic means or standard deviations. For example, a patient with a GCS of 6 is not precisely "half as conscious" as a patient with a GCS of 12, nor is the clinical change from GCS 4 to 5 biologically equivalent to the change from GCS 14 to 15. Ordinal variables must be summarized using medians and interquartile ranges (IQR) and analyzed using non-parametric tests.
Continuous Variables: Interval vs Ratio
Continuous variables represent measurements taken along a continuum that can assume any numerical value within a given range:
- Interval Variables: Quantitative scales with equal distances between increments but an arbitrary zero point. Zero does not represent complete absence of the measured property. A classic example is temperature (0°C does not represent the absence of heat, and 40°C is not "twice as hot" as 20°C).
- Ratio Variables: Quantitative scales with equal distances between increments and an absolute, non-arbitrary true zero. At zero, the quantity ceases to exist. Ratios between values are mathematically valid (e.g., a serum lactate of 8.0 mmol/L is precisely double a lactate of 4.0 mmol/L). Emergency medicine ratio variables include mean arterial pressure (MAP in mmHg), serum lactate, blood glucose, arterial pH, duration of vasopressor infusion (hours), and door-to-needle thrombolytic time (minutes).
Descriptive Statistics: Parametric vs Non-Parametric Distributions
Before selecting an inferential statistical test, continuous data must be tested for normality (using Shapiro-Wilk or Kolmogorov-Smirnov tests, or visual histogram/Q-Q plot inspection).
Gaussian (Normal) Distribution & Parametric Metrics
A normal distribution is symmetric and bell-shaped, where the mean, median, and mode are identical:
- Central Tendency & Dispersion: Described using the Mean (arithmetic average) and Standard Deviation (SD).
- Empirical 68-95-99.7 Rule:
- of all observations fall within .
- (commonly approximated as using ) fall within .
- of all observations fall within .
- Standard Error of the Mean (SEM): Calculated as . SEM quantifies the precision of the sample mean as an estimate of the true population mean, not the variability of the sample data. Reporting SEM instead of SD artificially minimizes apparent data variability and is considered misleading in clinical drug trials.
Skewed Distributions & Non-Parametric Metrics
Emergency department data frequently exhibit asymmetric, non-normal distributions with prolonged tails:
- Positively Skewed (Right-Skewed): The distribution tail extends toward higher positive values. In positive skewness, extreme high values pull the mean upward: Mean > Median > Mode.
- Clinical Examples: ED length of stay (hours), initial serum troponin I concentrations, total cumulative fentanyl dose in trauma resuscitation, ICU mechanical ventilation days.
- Negatively Skewed (Left-Skewed): The distribution tail extends toward lower values, pulling the mean downward: Mean < Median < Mode.
- Clinical Examples: Patient age in ischemic stroke registries, baseline oxygen saturation in acute COPD exacerbations.
- Central Tendency & Dispersion: Skewed data must be described using the Median (50th percentile) and Interquartile Range (IQR), defined as the difference between the 75th percentile () and the 25th percentile (). The median is resistant to extreme outliers, whereas the arithmetic mean is heavily distorted by extreme values.
The Hypothesis Testing Framework: Errors & Statistical Power
Clinical trials evaluate whether observed differences between pharmacotherapies reflect true therapeutic effects or random sampling variation.
The Null and Alternative Hypotheses
- Null Hypothesis (): States that there is no true difference, association, or effect between the comparative treatments in the underlying target population (e.g., or ).
- Alternative Hypothesis ( or ): States that a true difference or effect exists (e.g., for a two-sided hypothesis).
Decision Matrix: Type I vs Type II Errors
TRUE STATE OF NATURE
Null Hypothesis True Null Hypothesis False
┌─────────────────────────┬─────────────────────────┐
Reject Null │ TYPE I ERROR │ CORRECT DECISION │
Hypothesis │ (Alpha, α) │ Statistical Power │
(Statistically │ False Positive │ (1 - β) │
Significant) │ Standard: α = 0.05 │ Standard: 80% to 90% │
├─────────────────────────┼─────────────────────────┤
Fail to Reject │ CORRECT DECISION │ TYPE II ERROR │
Null Hypothesis │ Confidence Level │ (Beta, β) │
(Not Statistically │ (1 - α) │ False Negative │
Significant) │ Standard: 1 - α = 0.95 │ Standard: β = 0.10–0.2│
└─────────────────────────┴─────────────────────────┘
- Type I Error (, False Positive): Rejecting the null hypothesis when is actually true. Concluding that a new drug is superior or active when, in reality, it is ineffective. The allowable threshold is established prior to trial initiation as the significance level (standard or ).
- Type II Error (, False Negative): Failing to reject the null hypothesis when is actually false. Concluding that a drug has no beneficial effect when a clinically meaningful difference truly exists. The acceptable threshold is typically set at to ( to ).
- Statistical Power (): The mathematical probability of correctly rejecting a false null hypothesis—that is, detecting a true clinical difference of a specified magnitude if one genuinely exists. Standard clinical trial design targets a power of to ( to ).
Four Determinants of Statistical Power
- Sample Size (): Power is directly proportional to sample size. Increasing patient enrollment increases the statistical precision of effect estimates, shrinks the standard error, and expands power.
- Significance Level (): Setting a more lenient alpha (e.g., relaxing from to ) increases power because the critical threshold for rejecting is lower, though this increases Type I error risk.
- Effect Size ( or Delta): The magnitude of the anticipated clinical difference between treatment arms. Detecting a massive treatment effect (e.g., absolute mortality reduction) requires a substantially smaller sample size than detecting a subtle benefit (e.g., absolute mortality reduction).
- Population Variance (): Power is inversely related to data variance. High variability or heterogeneous baseline clinical characteristics increase noise, reducing statistical power. Minimizing measurement variability increases power.
Important
An underpowered clinical trial (e.g., power , high ) that yields a non-significant result () cannot prove that two emergency interventions are equivalent. It merely demonstrates that the study lacked adequate sample size to detect a true difference, committing a Type II error.
The -Value & Clinical vs Statistical Significance
Rigorous Definition of the -Value
The -value is the probability of obtaining test results at least as extreme as the observed data, assuming that the null hypothesis () is strictly true.
Common Misconceptions to Avoid on the BCEMP Exam
- Misconception 1: The -value is NOT the probability that the null hypothesis is true.
- Misconception 2: The -value is NOT the probability that the observed results occurred by chance alone.
- Misconception 3: The -value does NOT reflect the magnitude, direction, or clinical relevance of a drug's effect.
Statistical Significance vs Clinical Significance (MCID)
- Statistical Significance: Simply denotes that the observed test statistic crossed a predefined mathematical threshold (, typically ). In massive clinical trials (), tiny, clinically meaningless differences easily reach .
- Clinical Significance & MCID: The Minimal Clinically Important Difference (MCID) represents the smallest change in a clinical outcome that patients or clinicians identify as meaningful. For example, in an ED trial of 15,000 patients, an investigational antiemetic reduced nausea visual analog scores by 0.3 mm on a 100 mm scale (). While statistically significant, an improvement of 0.3 mm falls far below the accepted MCID for acute nausea (typically mm), making the drug clinically inert.
Statistical Test Selection Matrix
The following matrix outlines the decision rules for selecting inferential statistical tests on the BCEMP examination:
| Variable Type & Distribution | Groups / Comparison Structure | Parametric Test | Non-Parametric Alternative | Emergency Pharmacotherapy Example |
|---|---|---|---|---|
| Continuous, Normal | 2 Independent groups | Student's independent two-sample t-test | — | Mean MAP achieved at 1 hour: Norepinephrine vs Vasopressin |
| Continuous, Skewed / Ordinal | 2 Independent groups | — | Mann-Whitney U test (Wilcoxon rank-sum) | Median ED length of stay (hours): Early pharmacy intervention vs standard care |
| Continuous, Normal | 2 Paired / Matched groups | Paired Student's t-test | — | Mean heart rate: Pre- vs post-diltiazem bolus in rapid atrial fibrillation |
| Continuous, Skewed / Ordinal | 2 Paired / Matched groups | — | Wilcoxon signed-rank test | Median pain score (0–10): Pre- vs 15-min post-IV fentanyl administration |
| Continuous, Normal | Independent groups | One-Way ANOVA (Analysis of Variance) | — | Mean time to sedation across 3 RSI induction arms: Etomidate vs Ketamine vs Propofol |
| Continuous, Skewed / Ordinal | Independent groups | — | Kruskal-Wallis test | Median ICU delirium days across 3 delirium regimens: Haloperidol vs Quetiapine vs Placebo |
| Continuous, Normal | Repeated measures over time | Repeated Measures ANOVA | Friedman test | Serial mean serum lactate at 0h, 2h, 4h, and 6h during septic shock resuscitation |
| Categorical / Nominal | 2 or more groups (Large sample, all expected cell counts ) | — | Pearson's Chi-Square () test of independence | 30-day all-cause mortality (yes/no): Tenecteplase vs Alteplase in acute PE |
| Categorical / Nominal | 2 groups (Small sample, any expected cell count ) | — | Fisher's Exact test | Severe angioedema occurrence (yes/no) after thrombolysis in a pilot safety study () |
Tip
When One-Way ANOVA or Kruskal-Wallis detects a statistically significant difference () among groups, it only indicates that at least one group differs from the others. It does not identify which specific pairs differ. A post-hoc multiple comparison test (e.g., Tukey's HSD or Bonferroni correction for ANOVA; Dunn's test for Kruskal-Wallis) must be performed to maintain the overall family-wise Type I error rate at .
Confidence Intervals & Clinical Interpretation
A Confidence Interval (95% CI) provides a range of plausible values for the true unknown population parameter. Mathematically, if an experiment is repeated 100 times under identical sampling conditions, of the generated confidence intervals will encompass the true population effect.
Interpreting the Null Value in Confidence Intervals
95% CONFIDENCE INTERVAL NULL VALUES
┌─────────────────────────────────────────────────────────────────────────────┐
│ 1. Difference Metrics (Null Value = 0.0) │
│ Mean Difference, Absolute Risk Reduction (ARR), Risk Difference │
│ ├─ CI: [-4.2 to +1.8] --> Crosses 0.0 --> NOT Statistically Significant│
│ └─ CI: [+1.5 to +6.8] --> Excludes 0.0 --> STATISTICALLY SIGNIFICANT │
├─────────────────────────────────────────────────────────────────────────────┤
│ 2. Ratio Metrics (Null Value = 1.0) │
│ Relative Risk (RR), Odds Ratio (OR), Hazard Ratio (HR) │
│ ├─ CI: [0.82 to 1.15] --> Crosses 1.0 --> NOT Statistically Significant│
│ └─ CI: [0.54 to 0.88] --> Excludes 1.0 --> STATISTICALLY SIGNIFICANT │
└─────────────────────────────────────────────────────────────────────────────┘
- Precision: The width of the confidence interval reflects statistical precision. Narrow intervals reflect large sample sizes and high precision; wide intervals indicate small sample sizes, high variance, and clinical imprecision.
Non-Inferiority Clinical Trial Design
In emergency medicine pharmacotherapy, novel agents are frequently developed not to demonstrate superior efficacy over standard-of-care agents, but to offer distinct secondary clinical advantages—such as oral administration, single-dose administration, elimination of laboratory monitoring, or reduced organ toxicity.
The Non-Inferiority Margin ( or Delta)
- Definition: The non-inferiority margin () is the maximum clinically acceptable amount by which the new investigational therapy can be worse than the active standard-of-care control while still being considered therapeutically equivalent.
- Setting the Margin: is established prior to trial initiation based on historical data showing the active comparator's benefit over placebo. If is set too wide, an ineffective or harmful drug could falsely be declared non-inferior.
- Hypotheses in Non-Inferiority:
- : Difference (Standard - Investigational) (Investigational agent is inferior by margin or more).
- : Difference (Standard - Investigational) (Investigational agent is non-inferior).
- Statistical Testing: Uses a one-sided hypothesis test with a standard significance threshold of .
Intention-to-Treat vs Per-Protocol Analysis in Non-Inferiority
- Superiority Trials: Intention-to-Treat (ITT) is the gold standard because dropouts, crossovers, and non-compliance dilute treatment differences toward the null, providing a conservative estimate and guarding against Type I error.
- Non-Inferiority Trials: Dilution toward the null makes two comparative treatments appear more similar than they actually are. In non-inferiority trials, an ITT analysis can falsely drive the treatment difference toward zero, inappropriately causing the trial to reject and declare non-inferiority (causing a Type I error).
- Regulatory Standard: In non-inferiority trials, non-inferiority must be demonstrated conclusively in both the Per-Protocol (PP) population and the Intention-to-Treat (ITT) population to ensure validity.
A clinical specialist in emergency medicine is evaluating a randomized trial comparing duration of vasopressor support (measured in hours, non-normally distributed with pronounced right-skew) between septic shock patients receiving norepinephrine plus vasopressin versus norepinephrine monotherapy. What is the most appropriate statistical test to compare the primary endpoint between these two independent groups, and how should central tendency be reported?
Paired Student's t-test; central tendency reported as mean difference with 95% confidence interval.
Mann-Whitney U test (Wilcoxon rank-sum test); central tendency reported as median with interquartile range (IQR).
Independent two-sample Student's t-test; central tendency reported as mean with standard deviation (SD).
Kruskal-Wallis test; central tendency reported as median with standard error of the mean (SEM).
An emergency department investigative team is designing a randomized superiority trial comparing a novel push-dose vasopressor against phenylephrine for transient hypotension during procedural sedation. The investigators set alpha = 0.05 and calculate a required sample size of 320 patients to achieve 80% power (beta = 0.20) assuming a true absolute difference of 12% in successful hemodynamics. If the investigators decide to increase their target statistical power to 90% (beta = 0.10) while maintaining alpha = 0.05 and the expected effect size, which change occurs to the trial parameters?
The non-inferiority margin must be widened to ensure statistical significance across the primary hemodynamic endpoint.
The required sample size decreases because a higher power threshold lowers the risk of Type I false-positive error.
The required sample size must increase to reduce the probability of committing a Type II error (beta) from 20% down to 10%.
The alpha level automatically contracts from 0.05 to 0.01 to offset the higher statistical power calculation.
A multicenter randomized controlled trial evaluates a new rapid-acting antiarrhythmic agent versus IV amiodarone for the acute termination of stable wide-complex monomorphic ventricular tachycardia in the ED. The investigators report the primary outcome of successful sinus conversion within 15 minutes as an odds ratio of 1.48 (95% CI, 0.94 to 2.32; p = 0.09). The secondary endpoint of hospital discharge survival is reported as an absolute risk reduction of 4.2% (95% CI, 1.1% to 7.3%; p = 0.008). How should an emergency medicine clinical specialist interpret these statistical findings?
The primary endpoint failed to achieve statistical significance because the 95% confidence interval for the odds ratio includes 1.0, whereas the secondary survival outcome is statistically significant because its 95% confidence interval does not cross 0.
The primary endpoint achieved statistical significance because the 95% confidence interval point estimate exceeds 1.0, whereas the secondary endpoint is non-significant because the CI includes positive numbers.
The trial is invalid because odds ratios cannot be calculated in randomized prospective trials and must be converted to hazard ratios prior to hypothesis testing.
Both primary and secondary outcomes demonstrate statistically significant superiority because both point estimates favor the investigational antiarrhythmic agent.
Sections you finish are checked off in the contents.