16.2 Critical Literature Evaluation and Biostatistical Interpretation
Key Takeaways
In non-inferiority trials, the per-protocol (PP) population is methodologically conservative because treatment crossovers, dropouts, and non-adherence dilute true differences toward the null, which anti-conservatively favors non-inferiority in intention-to-treat (ITT) cohorts.
The non-inferiority margin (delta) must be pre-specified based on clinical judgment and historical active-control trial data to ensure that the new therapy preserves an acceptable fraction of the active comparator's proven efficacy over placebo.
Interrupted time series (ITS) with segmented regression is the gold standard quasi-experimental design in antimicrobial stewardship, differentiating immediate level changes (beta-2) from ongoing slope alterations (beta-3) while controlling for baseline secular trends and autocorrelation.
Likelihood ratios are prevalence-independent diagnostic metrics that translate pre-test odds to post-test odds via Fagan nomograms, whereas Positive and Negative Predictive Values fluctuate directly with disease prevalence.
Advanced observational techniques—including multivariable Cox proportional hazards modeling, propensity score matching, and time-dependent covariates—are vital to control for confounding by indication and immortal time bias in infectious diseases studies.
Critical Literature Evaluation and Biostatistical Interpretation
Evaluating infectious diseases pharmacotherapy literature requires sophisticated biostatistical and epidemiological literacy. Unlike chronic disease states where superiority trials with placebo controls predominate, infectious diseases registration trials and stewardship studies routinely employ non-inferiority designs, quasi-experimental interrupted time series, and complex observational cohorts. Antimicrobial stewardship pharmacists must master these analytical designs, effect measures, diagnostic test characteristics, and methods for mitigating systematic biases.
Randomized Controlled Trials: Superiority vs. Non-Inferiority
Randomized controlled trials (RCTs) represent the highest level of primary clinical evidence. In infectious diseases, trial designs diverge fundamentally based on their underlying clinical hypotheses.
SUPERIORITY VS. NON-INFERIORITY HYPOTHESES
Superiority Trial: Non-Inferiority Trial:
H0: P_new - P_control ≤ 0 H0: P_new - P_control ≤ -Δ
H1: P_new - P_control > 0 H1: P_new - P_control > -Δ
-Δ 0 +Δ
──────────────┼─────────────────┼─────────────────┼──────────────►
│ │ │ (New - Control)
│ │ │
│ [───●───] │ │ Inconclusive
│ │ │
│ [───●───] │ Non-Inferior (excludes -Δ)
│ │ │
│ │ [───●───] │ Superior (excludes 0)
The Rationale for Non-Inferiority (NI) Trials in ID
In life-threatening systemic bacterial infections (e.g., nosocomial pneumonia, bacteremia, complicated intra-abdominal infections), utilizing a placebo control is unethical because effective standard-of-care antimicrobial therapy is lifesaving. Consequently, novel antimicrobials are compared against established active comparators. Establishing that a new drug is superior in clinical cure is rarely feasible or expected when the active comparator already cures 80% to 90% of patients. Instead, registration trials aim to demonstrate non-inferiority—confirming that the new agent is not clinically worse than the active comparator by more than a pre-specified clinical margin (, delta), while potentially offering secondary advantages (e.g., activity against resistant strains, reduced nephrotoxicity, improved oral bioavailability, or less frequent dosing).
The Non-Inferiority Margin () and Confidence Intervals
- Defining the Margin (): The margin is established a priori based on clinical consensus and historical active-comparator trials vs. placebo. The margin must preserve a substantial fraction (typically ) of the active comparator's historical treatment effect over placebo. In modern antimicrobial registration trials, is typically set between 10% and 12.5% (or 0.10 to 0.125).
- Hypothesis Testing: Non-inferiority is demonstrated if the lower bound of the two-sided 95% confidence interval (or one-sided 97.5% CI) for the difference in cure rates () lies entirely to the right of (i.e., is greater than ). If the lower bound is , non-inferiority is established. If the lower bound also exceeds 0, superiority can be declared without alpha penalty.
Analysis Populations: ITT vs. mITT vs. Per-Protocol (PP)
| Analysis Population | Definition | Role in Superiority Trials | Role in Non-Inferiority Trials |
|---|---|---|---|
| Intention-to-Treat (ITT) | Includes all randomized subjects in the group to which they were assigned, regardless of adherence, withdrawal, or protocol violations. | Conservative: Non-compliance, dropouts, and drop-ins dilute treatment differences toward the null, guarding against false-positive claims of superiority (Type I error). | Anti-Conservative: Diluting differences toward the null makes two therapies look identical, artificially favoring the demonstration of non-inferiority. |
| Modified ITT (mITT) | Subjects randomized who received at least one dose of study medication and have confirmed baseline microbiological criteria (e.g., microbiologically confirmed infection). | Common primary efficacy cohort in infectious diseases trials (e.g., micro-ITT). | Balances clinical applicability with microbiological confirmation of target infection. |
| Per-Protocol (PP) | Restricts analysis strictly to subjects who met all eligibility criteria, completed full therapy, adhered to protocol, and had evaluable test-of-cure outcomes. | Anti-conservative: Excludes non-adherent subjects, potentially inflating treatment differences. | Conservative: Maximizes true pharmacological differences between treatments. If the new drug is truly inferior, PP analysis unmasks this failure. |
Important
The Conservative Benchmark in Non-Inferiority: Regulatory agencies (FDA and EMA) mandate that in non-inferiority trials, both the ITT/mITT and the Per-Protocol populations must independently demonstrate non-inferiority. If an antimicrobial achieves non-inferiority in the ITT population but fails in the PP population, non-inferiority cannot be confirmed, because the ITT finding may be an artifact of treatment dilution toward the null.
Observational and Quasi-Experimental Study Designs
When randomized trials are unfeasible, stewardship researchers rely on observational and quasi-experimental architectures.
Cohort and Case-Control Studies
- Prospective and Retrospective Cohort Studies:
- Subjects are categorized by exposure status (e.g., receipt of vancomycin AUC-guided dosing vs. trough-guided dosing) and followed forward (or historically) for the development of outcomes (e.g., acute kidney injury, 30-day mortality).
- Primary Effect Measure: Relative Risk (RR) or Risk Ratio ().
- Case-Control Studies:
- Subjects are selected based on the presence of the disease/outcome ("cases", e.g., patients with Clostridioides difficile infection) and compared to subjects without the outcome ("controls"). Past exposures (e.g., prior fluoroquinolone therapy) are evaluated retrospectively.
- Primary Effect Measure: Odds Ratio (OR) ().
- Rare Disease Assumption: When the outcome prevalence is low (< 5% to 10%), the OR approximates the RR (). In high-incidence settings, the OR significantly overstates the relative risk.
Interrupted Time Series (ITS) Analysis: The Stewardship Gold Standard
Antimicrobial stewardship interventions (e.g., implementing prospective audit and feedback, pre-authorization, or a novel clinical pathway) are implemented across entire hospital units, precluding individual patient randomization. Interrupted Time Series (ITS) with segmented regression analysis is the gold-standard quasi-experimental design for evaluating such interventions.
INTERRUPTED TIME SERIES SEGMENTED REGRESSION
Outcome Metric
(e.g., DOT/1000 PD)
▲
│ Pre-Intervention Secular Trend (β1)
│ ───
│ ───
│ ───
│ ───
│ Immediate Level Drop (β2)
│ ▼
│ ●
│ ───
│ ─── Post-Intervention Slope Change (β3)
│ ───
└───────────────────┬─────────────────────────────────►
Intervention Point Time (Months)
Segmented Regression Mathematical Formulation
The segmented linear regression model is expressed as:
- : The aggregated outcome metric at month (e.g., Days of Therapy [DOT] per 1,000 Patient-Days, or incidence of MRSA bacteremia per 10,000 bed-days).
- : Baseline level of the outcome at the start of the pre-intervention period.
- : Baseline secular trend (pre-intervention slope) per unit time, accounting for whether antibiotic use was already rising or falling prior to the intervention.
- : Continuous time variable from the beginning of the observation period.
- : Immediate level change occurring immediately following the stewardship intervention (step change).
- : Indicator dummy variable (0 before intervention, 1 after intervention).
- : Slope change representing the difference between the post-intervention slope and the pre-intervention slope (longitudinal trajectory change).
- : Continuous time variable elapsed since the implementation of the intervention (0 before intervention, after intervention).
- : Random error term.
Note
Why Simple Pre/Post Comparisons Fail: Simple "before-and-after" aggregate comparisons (e.g., Student's t-test comparing average antibiotic use in 2024 vs. 2025) commit profound attribution errors by ignoring pre-existing secular trends, seasonality (e.g., winter surges in viral respiratory infections), and autocorrelation (where data points close in time correlate with one another). ITS models control for these factors, isolating the true causal effect of the stewardship policy.
Biostatistical Parameters, Hypothesis Testing, and Effect Measures
Hypothesis Testing and Errors
- Null Hypothesis (): States there is no true difference between study arms ().
- Type I Error (): Rejecting the null hypothesis when it is actually true (false positive). Standardly set at .
- Type II Error (): Failing to reject the null hypothesis when it is actually false (false negative). Standardly set at to .
- Statistical Power (): The probability of detecting a statistically significant difference when a true difference exists (typically 80% to 90%). Underpowered trials risk prematurely dismissing effective therapies.
- P-Value vs. Confidence Interval: A p-value conveys only whether an arbitrary significance threshold was crossed; a 95% Confidence Interval (CI) conveys both statistical significance and the precision of the effect estimate, allowing clinicians to judge clinical significance.
Effect Measures: Risk Reductions, NNT, and NNH
2x2 CLINICAL CONTINGENCY MATRIX
Outcome Present (+) Outcome Absent (-)
Experimental (E) a b Total E = a + b
Control (C) c d Total C = c + d
Experimental Event Rate (EER) = a / (a + b)
Control Event Rate (CER) = c / (c + d)
- Absolute Risk Reduction (ARR):
- Relative Risk Reduction (RRR):
- Number Needed to Treat (NNT): The number of patients who must receive the experimental treatment rather than the control to prevent one additional adverse outcome: Rounding Rule: Always round UP to the nearest whole integer (e.g., ).
- Number Needed to Harm (NNH): The number of patients exposed to a therapy for one additional patient to experience a specific adverse event: Rounding Rule: Always round DOWN to the nearest whole integer (e.g., ) to avoid underestimating clinical risk.
Diagnostic Test Statistics and ROC Curves
Infectious diseases practice relies heavily on rapid molecular diagnostics, antigen detection assays, and serology.
Sensitivity, Specificity, and Predictive Values
DIAGNOSTIC ACCURACY 2x2 TABLE
Disease Present (D+) Disease Absent (D-)
Test Positive (T+) True Pos (TP) False Pos (FP) PPV = TP/(TP+FP)
Test Negative (T-) False Neg (FN) True Neg (TN) NPV = TN/(TN+FN)
Sensitivity = Specificity =
TP / (TP + FN) TN / (TN + FP)
- Sensitivity (Sn): Probability that the test is positive in a patient who has the disease (). A test with high sensitivity has few false negatives; useful for screening to rule out disease (SnNOut).
- Specificity (Sp): Probability that the test is negative in a patient who is free of disease (). A test with high specificity has few false positives; useful for confirmation to rule in disease (SpPIn).
- Positive Predictive Value (PPV): Proportion of patients with a positive test who truly have the infection ().
- Negative Predictive Value (NPV): Proportion of patients with a negative test who are truly free of infection ().
Warning
Prevalence Fluctuation of Predictive Values: Sensitivity and Specificity are intrinsic mathematical properties of the diagnostic test that remain constant across populations. In contrast, PPV and NPV depend heavily on disease prevalence. As disease prevalence decreases (e.g., testing asymptomatic patients for C. difficile toxins):
- PPV falls precipitously (most positive results are false positives).
- NPV rises toward 100%.
Likelihood Ratios (LR+ and LR-)
Likelihood ratios quantify how much a test result shifts the post-test probability of disease, completely independent of population prevalence.
- Positive Likelihood Ratio (): Ratio of the probability of a positive test in diseased individuals to that in healthy individuals: Interpretation: An provides strong, decisive diagnostic evidence to rule in disease.
- Negative Likelihood Ratio (): Ratio of the probability of a negative test in diseased individuals to that in healthy individuals: Interpretation: An provides strong, decisive diagnostic evidence to rule out disease.
- Bayes' Theorem Application:
Receiver Operating Characteristic (ROC) Curves and AUC
- The ROC curve plots the True Positive Rate (Sensitivity) on the y-axis against the False Positive Rate () on the x-axis across continuous diagnostic cut-offs (e.g., serum procalcitonin levels, C-reactive protein concentrations).
- Area Under the Curve (AUC / c-statistic): Measures overall diagnostic discrimination across all possible cut-points.
- : No discriminatory value (equivalent to a coin flip).
- : Acceptable discrimination.
- : Excellent discrimination.
- : Outstanding discrimination.
Controlling for Confounding and Bias in ID Literature
Observational studies evaluating antimicrobials are vulnerable to profound systematic biases that can distort findings.
EPIDEMIOLOGICAL BIASES IN ID RESEARCH
┌────────────────────────────────────────────────────────────────────────┐
│ CONFOUNDING BY INDICATION (CHANNELLING BIAS) │
│ • Sicker patients (septic shock, ICU, multiple comorbidities) receive │
│ the newer/broader reserve antibiotic; healthier patients receive SoC.│
│ • Unadjusted analysis falsely shows higher mortality for new drug. │
└───────────────────────────────────┬────────────────────────────────────┘
│ Statistical Mitigation
┌───────────────────────────────────▼────────────────────────────────────┐
│ PROPENSITY SCORE MATCHING (PSM) & IPTW │
│ • Calculates conditional probability of receiving treatment given │
│ baseline covariates; balances treatment and control groups. │
└───────────────────────────────────┬────────────────────────────────────┘
│ Survival Time Correction
┌───────────────────────────────────▼────────────────────────────────────┐
│ IMMORTAL TIME BIAS │
│ • Time window between culture draw and antibiotic switch (e.g., 48h). │
│ • Patients MUST SURVIVE 48h to receive the drug; survivor advantage │
│ falsely attributed to drug efficacy unless modeled as time-dependent.│
└────────────────────────────────────────────────────────────────────────┘
1. Confounding by Indication (Channelling Bias)
Occurs when the clinical indication for prescribing a drug is itself associated with the study outcome. In ID, clinicians preferentially prescribe newer, broader-spectrum reserve antimicrobials (e.g., ceftazidime-avibactam, cefiderocol) to patients with advanced sepsis, multi-organ failure, or prior treatment failures. In crude analyses, the new drug appears associated with higher mortality solely because it was channeled to sicker patients.
- Mitigation: Multivariable logistic regression, multivariable Cox proportional hazards modeling, Propensity Score Matching (PSM), and Inverse Probability of Treatment Weighting (IPTW).
2. Immortal Time Bias
Occurs when there is a time interval during follow-up in which the study outcome (death) cannot occur for one treatment group. For example, comparing patients switched to oral step-down therapy versus those maintained on IV therapy: a patient must survive long enough (e.g., 3–5 days) to become clinically stable and receive the oral agent. If the researcher counts follow-up from hospital admission, the step-down cohort receives "immortal time," introducing an artificial survival bias that drastically inflates the apparent benefit of oral step-down.
- Mitigation: Time-dependent Cox regression modeling, landmark analysis, or nested case-control designs.
3. Selection Bias and Loss to Follow-Up
Systematic differences between participants who complete a study and those who drop out. In HIV clinical trials, differential attrition due to intolerable adverse effects or virologic failure can distort long-term durability comparisons if non-completers are ignored (mitigated by classifying non-completers as treatment failures in "Missing = Failure" ITT analyses).
A phase 3 randomized clinical trial evaluates a novel intravenous cephalosporin against meropenem for complicated intra-abdominal infections using a non-inferiority margin of 10% (delta = -0.10). In the modified intention-to-treat (mITT) population, clinical cure was achieved in 84.0% (420/500) of the novel cephalosporin group and 86.0% (430/500) of the meropenem group (treatment difference: -2.0%; 95% CI: -6.2% to +2.2%). In the per-protocol (PP) population, clinical cure was 81.0% (324/400) versus 93.0% (372/400), respectively (treatment difference: -12.0%; 95% CI: -16.8% to -7.2%). How should an antimicrobial stewardship specialist interpret these findings?
Non-inferiority is conclusively established because the mITT analysis represents the gold standard and its lower confidence limit (-6.2%) does not cross the -10% margin
The novel cephalosporin is superior to meropenem because the upper confidence limit in the mITT population (+2.2%) exceeds zero
The trial demonstrates equivalence between the two agents because both achieved over 80% cure rates in the primary efficacy cohorts
Non-inferiority cannot be confirmed because the per-protocol analysis failed to exclude the -10% margin, indicating protocol non-adherence biased the mITT result
A clinical laboratory implements a novel molecular multiplex PCR assay for detecting Clostridioides difficile in unformed stool. The assay exhibits a sensitivity of 95% and a specificity of 90%. An antimicrobial stewardship team evaluates this assay across two distinct hospital units: Unit A (where C. difficile prevalence among tested symptomatic patients is 30%) and Unit B (where excessive ordering in mild, non-specific diarrhea results in a prevalence of only 2%). Which of the following statements correctly describes the diagnostic performance of this test across these units?
The test's sensitivity and specificity will both drop substantially in Unit B because predictive values and test sensitivity always move in the same direction
The positive predictive value will be far higher in Unit A (about 75–80%) than Unit B (about 16%), because PPV varies with disease prevalence
The positive likelihood ratio (LR+) will be significantly higher in Unit A than in Unit B because likelihood ratios depend heavily on the underlying prevalence of infection
The negative predictive value (NPV) will be substantially lower in Unit B than in Unit A because fewer patients in Unit B truly have the infection
An infectious diseases pharmacist designs an interrupted time series (ITS) study with segmented regression to evaluate the institutional impact of an antimicrobial stewardship policy restricting empiric anti-pseudomonal carbapenem use. The model is specified as: Y_t = β0 + β1(T_t) + β2(D_t) + β3(P_t). The regression yields: β1 = +0.8 (p = 0.04), β2 = -18.5 (p < 0.001), and β3 = -1.2 (p = 0.01), where Y_t represents meropenem Days of Therapy (DOT) per 1,000 patient-days. What is the correct interpretation of the parameter β2?
The baseline secular trend in meropenem consumption prior to the policy intervention
The month-over-month trajectory change in meropenem consumption sustained over the entire post-intervention period
The immediate, acute reduction (level change) in meropenem consumption observed in the first month following policy implementation
The total percentage reduction in hospital-acquired carbapenem-resistant Pseudomonas aeruginosa infections at the study's conclusion
Sections you finish are checked off in the contents.