6.3 Research Study Methodology, Literature Evaluation & Biostatistics
Key Takeaways
- The hierarchy of evidence ranks systematic reviews/meta-analyses and RCTs above observational designs (cohort, case-control, cross-sectional).
- Hypothesis testing evaluates the null hypothesis ($H_0$), balancing Type I error (alpha, false positive) and Type II error (beta, false negative, statistical power).
- $p$-values $<0.05$ indicate statistical significance; confidence intervals including 1.0 (for ratios) or 0 (for differences) indicate non-significant results.
- Sensitivity (SNOUT) rules out disease, Specificity (SPIN) rules in disease; PPV increases directly with higher disease prevalence.
- Critical appraisal requires evaluating internal validity (controlling selection, measurement, and confounding bias) and external validity (generalisability).
Evidence-based infection prevention relies on critical scientific evaluation and statistical literacy. The Infection Preventionist (IP) must analyze published research, assess methodological rigor, evaluate statistical significance versus clinical relevance, and apply biostatistical formulas to diagnostic testing and surveillance data.
The Hierarchy of Scientific Evidence
Scientific study designs possess varying degrees of protection against bias and confounding. The pyramid of evidence, ordered from highest to lowest strength of causality, guides clinical decision-making:
| Evidence Level & Study Design | Temporal Direction | Key Measures of Effect | Primary Advantages | Major Vulnerabilities / Biases |
|---|---|---|---|---|
| Systematic Reviews & Meta-Analyses | Retrospective summary of literature | Pooled Relative Risk (RR), Pooled Odds Ratio (OR) | Combines sample sizes from multiple studies to maximize statistical power and precision | Subject to publication bias and heterogeneity among included primary studies. |
| Randomized Controlled Trials (RCTs) | Prospective experimental | Relative Risk (RR), Absolute Risk Reduction (ARR), NNT | Randomization minimizes selection bias and confounding; gold standard for causality | High cost, ethical constraints in withholding proven infection prevention measures. |
| Cohort Studies | Prospective or Retrospective observational | Relative Risk (RR), Cumulative Incidence Rate Ratio | Tracks exposed vs. unexposed groups over time; excellent for measuring incidence | Vulnerable to loss to follow-up, healthy worker effect, and unmeasured confounding. |
| Case-Control Studies | Retrospective observational | Odds Ratio (OR) | Efficient for rare diseases or acute outbreak investigations; compares cases vs. controls | High susceptibility to recall bias, selection bias, and non-comparable control groups. |
| Cross-Sectional Studies | Snapshot / Single point in time | Point Prevalence, Prevalence Odds Ratio | Inexpensive and rapid; assesses overall disease burden or baseline compliance | Cannot establish temporal sequence or causality (egg-or-chicken dilemma). |
| Case Series / Case Reports | Retrospective descriptive | Descriptive statistics, proportions | First signal of novel pathogens or unique outbreak presentations | No control group; cannot test hypotheses or establish statistical association. |
Core Biostatistical Concepts & Hypothesis Testing
Interpreting research literature requires a rigorous understanding of biostatistical hypothesis testing:
Hypotheses and Errors
- Null Hypothesis (H0): The baseline assumption that no true statistical difference, association, or effect exists between comparison groups.
- Alternative Hypothesis (Ha): The claim that a true difference or effect exists between groups.
- Type I Error (α / False Positive): Rejecting the null hypothesis when it is actually true (concluding a difference exists when there is none). The standard alpha threshold is set at α = 0.05.
- Type II Error (β / False Negative): Failing to reject the null hypothesis when it is actually false (failing to detect a true clinical difference).
- Statistical Power (1 - β): The probability of correctly rejecting a false null hypothesis (detecting a true difference if one exists). Standard power is targeted at 80% (0.80) or 90%, influenced by sample size, effect size, and significance level.
p-Values and Confidence Intervals
- p-Value: The probability of obtaining a test statistic at least as extreme as the observed result, assuming the null hypothesis is true. A p-value <0.05 indicates statistical significance, but does not measure the magnitude or clinical importance of the effect.
- 95% Confidence Interval (CI): A range of values constructed around a point estimate that has a 95% probability of containing the true population parameter.
- For ratio measures (Relative Risk or Odds Ratio), if the 95% CI includes 1.0 (e.g., RR = 0.85, 95% CI: 0.65-1.12), the finding is not statistically significant at the α = 0.05 level.
- For difference measures (Mean Difference), if the 95% CI includes 0.0, the finding is not statistically significant.
- A narrower confidence interval indicates greater statistical precision, usually driven by a larger study sample size.
Diagnostic Testing Metrics & Calculations
IPs evaluate diagnostic and screening assays using four foundational biostatistical performance metrics:
Sensitivity and Specificity (Intrinsic Test Properties)
- Sensitivity: The proportion of individuals with the disease who test positive: High sensitivity minimizes false negatives. A highly sensitive test is used for screening (SNOUT: Sensitive test Negative rules OUT).
- Specificity: The proportion of individuals without the disease who test negative: High specificity minimizes false positives. A highly specific test is used for confirmation (SPIN: Specific test Positive rules IN).
Predictive Values (Prevalence-Dependent Metrics)
- Positive Predictive Value (PPV): The probability that a patient with a positive test result actually has the disease: PPV increases directly as disease prevalence increases in the population being tested.
- Negative Predictive Value (NPV): The probability that a patient with a negative test result actually is free of the disease: NPV decreases as disease prevalence increases.
Critical Appraisal of Infection Control Literature
When reviewing published studies to update institutional protocols, the IP conducts a critical appraisal evaluating:
- Internal Validity: The degree to which the study design, execution, and analysis minimize bias and confounding within the sample.
- Selection Bias: Non-random sampling or systematic differences between comparison groups.
- Measurement / Information Bias: Misclassification of exposures or outcomes (e.g., unblinded assessors evaluating surgical site infections).
- Confounding: A distortion caused by a third variable associated with both the exposure and outcome (controlled via restriction, matching, stratification, or multivariable logistic regression).
- External Validity (Generalizability): The degree to which study findings apply to the IP’s specific healthcare setting, patient demographics, and operational workflows.
An Infection Preventionist is evaluating a novel rapid antigen test for Respiratory Syncytial Virus (RSV). The test exhibits a sensitivity of 95% and a specificity of 98%. If this test is deployed during a peak winter outbreak when RSV prevalence is high, how will the Positive Predictive Value (PPV) compare to testing during mid-summer when RSV prevalence is low?
To investigate a sudden outbreak of Clostridioides difficile infections in an orthopedic ward, an IP identifies 20 infected patients ("cases") and selects 40 uninfected orthopedic patients admitted during the same timeframe ("controls"). The IP then reviews historical medical records to compare prior exposure to specific fluoroquinolone antibiotics. Which study design is being utilized?
A multi-center study evaluates a new chlorhexidine gluconate (CHG) bathing protocol for reducing central line-associated bloodstream infections (CLABSI). The reported Relative Risk (RR) is 0.72 with a 95% Confidence Interval (CI) of 0.48 to 1.08 (p = 0.11). How should the IP interpret this finding?
An IP conducts a study comparing two disinfectant wipes. The statistical analysis yields a p-value of 0.03, leading the IP to reject the null hypothesis and conclude that Wipe A is superior to Wipe B. However, in reality, Wipe A and Wipe B have identical antimicrobial efficacy. What statistical error has occurred?