9.3 Critical Appraisal of Infection Prevention Literature
Key Takeaways
- Epidemiological study designs dictate causal strength: Randomized Controlled Trials (RCTs) minimize internal confounding, while observational Cohort and Case-Control studies evaluate real-world exposures and rare HAI outbreaks.
- A p-value < 0.05 indicates statistical significance, but 95% Confidence Intervals (CI) communicate clinical precision; a 95% CI for a ratio metric (RR, OR, HR) that includes 1.0 indicates a non-statistically significant result.
- Relative Risk (RR) measures relative incidence in prospective cohort studies while Odds Ratio (OR) estimates relative odds in retrospective case-control designs, and Number Needed to Treat (NNT = 1 / Absolute Risk Reduction) quantifies how many patients must receive an intervention to prevent one additional adverse outcome.
- Methodological threats such as selection bias, measurement bias, the Hawthorne effect, and confounding must be controlled through multivariable regression modeling and evaluated using evidence hierarchies like GRADE.
- Quality improvement applies existing evidence locally and generally does not require IRB review, whereas research seeks generalizable knowledge and does; intent to publish is the point at which most facilities request an IRB determination.
9.3 Critical Appraisal of Infection Prevention Literature
Quick Answer: Critical appraisal of infection prevention literature requires matching clinical research questions to appropriate study designs (RCTs, prospective cohort, retrospective case-control, interrupted time series) and accurately interpreting biostatistical metrics—including $p$-values, 95% Confidence Intervals (CI), Relative Risk (RR), Odds Ratio (OR), Absolute Risk Reduction (ARR), and Number Needed to Treat (NNT)—while controlling for bias, the Hawthorne effect, and confounding using evidence frameworks like GRADE.
Infection Preventionists must base hospital policies, product selections, and clinical practice bundles on robust scientific evidence rather than habit or vendor claims. Critically appraising peer-reviewed literature is an essential core competency for the IPC professional.
Epidemiological Study Designs in IPC Research
Study designs fall into two broad categories: Experimental (where the investigator assigns the intervention) and Observational (where the investigator observes natural exposures and outcomes).
Primary Study Designs Comparison
| Study Design | Type | Group Allocation | Primary Metric | Major Strength | Principal Limitation |
|---|---|---|---|---|---|
| Randomized Controlled Trial (RCT) | Experimental | Random assignment to intervention vs. control | Relative Risk (RR) / Hazard Ratio (HR) | Minimizes confounding; establishes causality | High cost; ethical limits in outbreak settings |
| Prospective Cohort Study | Observational | Follows exposed vs. unexposed over time | Relative Risk (RR) | Establishes temporal sequence; good for rare exposures | Susceptible to loss to follow-up & confounding |
| Case-Control Study | Observational | Selects cases (HAI) and controls (no HAI); looks back | Odds Ratio (OR) | Efficient for rare outcomes or rapid outbreak investigations | Prone to recall bias & selection bias |
| Cross-Sectional Study | Observational | Measures exposure & outcome simultaneously | Prevalence Ratio / Odds Ratio | Fast, low cost; excellent hypothesis generation | Cannot establish temporality or causality |
| Quasi-Experimental / ITS | Observational / Non-random | Compares pre- and post-intervention hospital trends | Change in level & trend slope | Evaluates system-wide policy changes without randomization | Vulnerable to co-interventions & secular trends |
Study Design Applications in IPC
- Randomized Controlled Trial (RCT): Evaluates the efficacy of daily 2% chlorhexidine gluconate (CHG) bathing versus plain soap bathing in preventing central line bloodstream infections among intensive care patients.
- Case-Control Study: Investigates a sudden cluster of 6 patients with Burkholderia cepacia bloodstream infections. IPs compare the 6 cases to 18 matched control patients without infection on the same unit to retroactively assess exposure to contaminated mouthwash batches.
- Interrupted Time Series (ITS): Evaluates the impact of a hospital-wide mandatory ultraviolet (UV-C) room disinfection policy implemented on January 1st by tracking monthly facility-onset Clostridioides difficile rates for 12 months before and 12 months after implementation.
Key Statistical Measures & Biostatistical Interpretation
Evaluating research publications requires understanding statistical metrics that convey both statistical significance and clinical magnitude.
Hypotheses & p-Values
- Null Hypothesis ($H_0$): The assumption that there is no true difference between the intervention group and control group.
- $p$-value: The probability of obtaining results at least as extreme as the observed data, assuming the null hypothesis is true. By convention, a $p$-value < 0.05 indicates statistical significance (less than a 5% probability that the result occurred by random chance).
95% Confidence Intervals (95% CI)
While a $p$-value provides a binary test of significance, the 95% Confidence Interval (95% CI) provides a range of values within which the true population parameter is expected to fall with 95% certainty. It communicates both precision (narrower interval = higher precision) and clinical significance.
Interpreting 95% Confidence Intervals for Ratio Metrics (RR, OR, HR):
├── If 95% CI crosses 1.0 (e.g., RR = 0.75 [95% CI: 0.52 - 1.15]) ──> NOT Statistically Significant
└── If 95% CI does NOT cross 1.0 (e.g., RR = 0.62 [95% CI: 0.44 - 0.88]) ──> Statistically Significant Reduction
Critical Exam Rule: For ratio metrics (Relative Risk, Odds Ratio, Hazard Ratio), if the 95% Confidence Interval includes the null value of 1.0, the findings are not statistically significant, regardless of how impressive the point estimate appears.
Relative Risk (RR) vs. Odds Ratio (OR)
- Relative Risk (RR): Compares the risk (incidence) of an event occurring in an exposed group to the unexposed group. Used in cohort studies and RCTs.
- Odds Ratio (OR): Compares the odds of exposure among cases (diseased) to the odds of exposure among controls (non-diseased). Used in case-control studies.
Absolute Risk Reduction (ARR) & Number Needed to Treat (NNT)
- Control Event Rate (CER): Proportion of control group experiencing the outcome.
- Experimental Event Rate (EER): Proportion of intervention group experiencing the outcome.
- Absolute Risk Reduction (ARR): The absolute percentage difference in event rates between groups.
- Number Needed to Treat (NNT): The number of patients who must receive an intervention to prevent one additional adverse event (e.g., one HAI).
Example Calculation: In a trial of skin antiseptic preps, CAUTI occurs in 8% of standard-care patients ($ ext{CER} = 0.08$) and 3% of intervention-prep patients ($ ext{EER} = 0.03$).
- $ ext{ARR} = 0.08 - 0.03 = 0.05$ (or 5% absolute reduction).
- $ ext{NNT} = 1 / 0.05 = 20$.
- Interpretation: Treating 20 patients with the new antiseptic prep prevents 1 CAUTI case.
Bias, Confounding, & Methodological Vulnerabilities
Methodological flaws threaten a study's internal validity (whether the study accurately measures what it intended) and external validity (generalizability to other healthcare settings).
Major Types of Bias in IPC Studies
- Selection Bias: Systematic error in how study subjects are recruited or selected. (e.g., enrolling healthier non-ICU patients into the new central line bundle group).
- Information / Measurement Bias: Errors in data collection or outcome classification (e.g., unblinded auditors diagnosing fewer SSIs in the intervention arm).
- Hawthorne Effect: A phenomenon where healthcare personnel temporarily alter or improve their behavior simply because they know they are being observed.
- Impact on IPC: Hand hygiene compliance routinely spikes from 45% baseline to 95% when a visible auditor holding a clipboard enters the unit.
- Publication Bias: The tendency for academic journals to publish studies showing positive, statistically significant results while suppressing negative or inconclusive studies.
Confounding & Methods of Control
A Confounder is a third variable associated with both the exposure and the outcome that distorts the true relationship.
Methods for Controlling Confounding
- In Study Design:
- Randomization: Random assignment balances known and unknown confounders equally between arms.
- Restriction: Restricting study entry to a specific sub-group (e.g., enrolling only adult surgical patients).
- Matching: Matching cases and controls on confounding variables (e.g., age, unit type, comorbidity score).
- In Statistical Analysis:
- Stratification: Analyzing data within sub-strata (e.g., analyzing outcomes separately in ventilated vs. non-ventilated patients).
- Multivariable Regression Analysis: Logistic regression or Cox proportional hazards modeling to mathematically adjust for multiple confounders simultaneously.
- Propensity Score Matching: Summarizing multiple baseline confounders into a single propensity score to pair intervention and control subjects.
Evidence Hierarchies & Systematic Appraisal (GRADE)
Not all published literature carries equal weight. IPs evaluate scientific evidence using established Evidence Pyramids and appraisal frameworks.
Hierarchy of Evidence Pyramid:
├── Level 1: Systematic Reviews & Meta-Analyses (PRISMA Guidelines)
├── Level 2: Large Randomized Controlled Trials (RCTs)
├── Level 3: Prospective Cohort Studies
├── Level 4: Retrospective Case-Control Studies
├── Level 5: Quasi-Experimental & Interrupted Time Series
├── Level 6: Cross-Sectional Studies & Case Series
└── Level 7: Expert Opinion, Narrative Reviews, & Benchmark Summaries
The GRADE Approach
The GRADE (Grading of Recommendations Assessment, Development, and Evaluation) system classifies quality of evidence into four levels:
- High Quality: Further research is very unlikely to change confidence in the estimate of effect.
- Moderate Quality: Further research is likely to have an important impact on confidence.
- Low Quality: Further research is very likely to have an important impact on confidence.
- Very Low Quality: Any estimate of effect is highly uncertain.
Recommendations are subsequently categorized as Strong (desirable consequences clearly outweigh undesirable consequences) or Weak (conditional upon patient values, feasibility, or local resources).
From Appraisal to Action: Identifying and Reporting Research Opportunities
The blueprint does not stop at reading the literature. Two further associate-level tasks are to identify opportunities for research and to report applicable research findings to relevant personnel.
Spotting a knowledge or practice gap
A research opportunity is simply a question your own data raise that the literature does not answer. Recurring sources:
| Source of the gap | Example |
|---|---|
| Surveillance data that contradict the benchmark | Your CAUTI rate is at goal but your device utilization ratio is far above peers — why is the rate low? |
| An intervention that worked here and nowhere else, or vice versa | A bundle with strong published effect produces no change on your unit |
| A guideline built on weak evidence | A recommendation graded as expert opinion that drives significant cost or workload |
| A population the literature ignores | Most device-associated evidence comes from adult ICUs; behavioral health, home infusion, and rural critical access settings are thinly studied |
| A repeated near-miss or workaround | Staff consistently deviate from a policy in the same way, suggesting the policy does not fit the work |
| An unexplained pseudo-outbreak or novel reservoir | A transmission route not previously described |
Distinguish quality improvement from research. QI applies known evidence locally to improve care and does not usually require Institutional Review Board (IRB) approval. Research aims to produce generalizable knowledge and does require IRB review. The intent to publish alone does not convert QI into research, but it is the point at which most facilities ask the IRB for a determination — and asking early is far easier than asking after data collection.
Reporting findings to the people who can use them
A finding that stays with the infection preventionist changes nothing. Route it deliberately:
| Audience | Vehicle |
|---|---|
| Frontline staff | Unit huddles, one-page summaries, just-in-time teaching at the point of the practice |
| Physicians and pharmacy | Medical staff and stewardship committee presentations; a new diagnostic method or resistance pattern is directly actionable for them |
| Infection Prevention Committee | Formal presentation with an appraisal of evidence quality and a specific recommendation |
| Executive leadership | Implications for cost, regulatory standing, and resources |
| Public health department | Emerging pathogens, novel resistance mechanisms, unusual clusters |
| The wider field | APIC or SHEA abstracts, chapter presentations, peer-reviewed manuscripts |
When reporting, state the strength of the evidence alongside the finding. "One retrospective single-center study suggests" and "a Cochrane review of 14 randomized trials shows" justify very different actions, and an infection preventionist who blurs that distinction loses credibility the first time a weakly supported recommendation fails.
A clinical trial evaluating a novel silver-impregnated vascular dressing reports a Relative Risk (RR) of 0.65 for CLABSI prevention with a 95% Confidence Interval of 0.42 to 1.12. How should an Infection Preventionist interpret this result?
In a study of surgical wound irrigations, the Surgical Site Infection (SSI) rate was 10% in the control group and 5% in the experimental irrigation group. What is the Number Needed to Treat (NNT) to prevent one additional SSI?
An Infection Preventionist needs to investigate an abrupt, localized cluster of 4 cases of postoperative endophthalmitis following cataract surgeries. Which observational study design is most appropriate and efficient for rapidly identifying potential exposure sources?
During direct covert auditing of bedside hand hygiene compliance, an IP notes that compliance measures 92% when auditors are visible on the unit, but drops to 48% when monitored by automated electronic sensors. What methodological phenomenon explains this discrepancy?
You've completed this section
Continue exploring other exams