12.2 Clinical Trial Evaluation and Information Source Reliability
Key Takeaways
- Information source reliability is judged by author credentials, peer review, funding source, conflicts of interest, publication date, and journal quality; predatory journals, preprints, and conference abstracts require extra scrutiny before citation.
- RCT validity hinges on adequate randomization, allocation concealment, blinding, and appropriate control; non-inferiority trials additionally require a pre-specified margin, assay sensitivity, and the constancy assumption.
- Internal validity threats — selection, performance, detection, attrition, and reporting bias — are mitigated by randomization, blinding, ITT analysis, and protocol registration; CONSORT standardizes trial reporting.
- Effect size interpretation requires distinguishing relative risk reduction from absolute risk reduction and NNT; confidence intervals convey both magnitude and precision, while p-values alone do not establish clinical significance.
- Meta-analysis quality depends on heterogeneity (I²), choice of fixed versus random effects models, and assessment of publication bias via funnel plot asymmetry; surrogate endpoints and composite endpoints require careful clinical interpretation.
12.2 Clinical Trial Evaluation and Information Source Reliability
Quick Answer: Evaluating evidence demands two skills: judging whether a source is trustworthy and judging whether a trial is valid and clinically meaningful. Source reliability turns on peer review, funding, conflicts of interest, and freedom from predatory publishing. Trial appraisal examines randomization, blinding, allocation concealment, and control choice for internal validity; external validity asks whether the result applies to your patient. Effect size interpretation distinguishes RRR, ARR, NNT, HR, OR, and confidence intervals from p-values alone; meta-analyses add heterogeneity (I²) and publication-bias concerns.
Information Source Reliability
Before appraising a trial, judge the source. The criteria are concrete: author credentials (relevant expertise, prior published work), peer review status, funding source (industry-funded trials tend to favor the sponsor's drug), conflicts of interest (financial, intellectual, institutional), publication date (medicine moves fast — a 1998 trial may be obsolete), and journal quality (impact factor is a rough proxy, not a guarantee).
Predatory Journals, Preprints, and Conference Abstracts
Predatory journals charge publication fees without legitimate peer review; the Beall's List legacy and the Directory of Open Access Journals (DOAJ) quality-control vetting help distinguish them. Symptoms include rapid acceptance, generic emails, and absent indexing in MEDLINE. Preprints (bioRxiv, medRxiv) are not peer-reviewed and should be tagged as such when cited; treat their conclusions as provisional. Conference abstracts often never reach full publication and may differ substantially from the eventual paper — verify against the full text before acting.
Grey Literature and Regulatory Documents
Grey literature — FDA clinical and statistical reviews, EMA European Public Assessment Reports (EPARs), ClinicalTrials.gov registry entries, and health technology assessment reports — often contains data never published in journals. Regulatory reviews are especially valuable because they include unpublished trials and patient-level data. Always cross-check a journal article against the corresponding ClinicalTrials.gov record for outcome reporting consistency.
Media, Promotion, and Anecdote
Media coverage routinely exaggerates effect sizes and omits absolute risks. Drug promotion — even when compliant with FDA rules — is selective by design; independent evidence (Cochrane reviews, independent guidelines, peer-reviewed RCTs) should anchor decisions. Anecdote and single-case experience generate hypotheses but cannot establish causation; systematic observation in controlled trials is the gold standard.
Clinical Trial Design Appraisal
Randomized Controlled Trials (RCTs)
The RCT is the gold-standard design for causal inference. Appraise four features:
- Randomization method — simple, block, stratified, or cluster; computer-generated sequence preferred over quasi-random methods (alternation, chart number) which allow prediction.
- Allocation concealment — ensures enrolling clinicians cannot foresee the next assignment; central phone/web randomization or sealed opaque envelopes. Without concealment, selection bias can occur despite "randomization."
- Blinding — double-blind (participant and investigator) is strongest; open-label, single-blind, or unblinded evaluator designs may be necessary but introduce performance and detection bias.
- Control — placebo (best for efficacy), active comparator (best for comparative effectiveness), usual care, or no-treatment. Analysis sets include intention-to-treat (ITT, all randomized), per-protocol (PP, completers only), and modified ITT; ITT is preferred for superiority trials, PP for non-inferiority.
Non-Inferiority Trials
A non-inferiority (NI) trial asks whether a new treatment is "not unacceptably worse" than an established active control. Three assumptions are essential: (1) a pre-specified non-inferiority margin (M) based on historical placebo-controlled effect of the active control; (2) assay sensitivity (the trial could have detected a difference if one existed); and (3) the constancy assumption (the active control's historical effect persists in the new setting). NI margins are easily manipulated — a too-wide margin can declare inferior drugs non-inferior.
Adaptive, Pragmatic, and Equivalence Designs
Adaptive designs allow pre-specified modifications (sample size re-estimation, dropping a dose, stopping for efficacy/futility) based on interim analyses, regulated by FDA and ICH E9(R1) guidance. Pragmatic trials (e.g., TASTE, SPRINT-REP) test interventions in real-world settings with broad eligibility and routine care delivery, contrasting with explanatory trials that maximize internal control. Equivalence trials seek to show two treatments are within a defined symmetric margin of each other.
Validity Assessment
| Validity Threat | Type | Mitigation |
|---|---|---|
| Selection bias | Internal | Adequate randomization with allocation concealment |
| Performance bias | Internal | Blinding of participants and personnel; standardized protocols |
| Detection bias | Internal | Blinded outcome assessors; objective endpoints |
| Attrition bias | Internal | ITT analysis; high follow-up rates; multiple imputation |
| Reporting bias | Internal/External | Protocol registration; CONSORT reporting; full protocol publication |
| Limited generalizability | External | Pragmatic eligibility; representative settings; subgroup transparency |
The CONSORT (Consolidated Standards of Reporting Trials) statement standardizes RCT reporting; journal endorsement is near-universal among high-impact medical journals.
Effect Size Interpretation
A statistically significant p-value does not mean a clinically meaningful effect. The FPGEE tests the math.
- Relative risk reduction (RRR) = (control event rate − treatment event rate) / control event rate.
- Absolute risk reduction (ARR) = control event rate − treatment event rate — the same arithmetic but on an absolute scale; typically smaller and more clinically honest.
- Number needed to treat (NNT) = 1 / ARR (expressed as a decimal). A smaller NNT is more powerful.
- Hazard ratio (HR) from time-to-event analyses; HR 0.75 means a 25% reduction in the instantaneous hazard.
- Odds ratio (OR) from case-control and logistic regression; OR approximates RR only when events are rare (< 10%).
- Confidence intervals (CI) communicate both magnitude and precision; a 95% CI that excludes 1.0 for RR/HR/OR or 0 for ARR is conventionally "statistically significant," but the lower bound tells you the smallest plausible effect — what matters clinically. The minimally clinically important difference (MCID) is the smallest effect patients perceive as beneficial.
Worked Example
A trial reports cardiovascular events in 10% of placebo patients and 7% of those on a new drug over 3 years. RRR = (10 − 7) / 10 = 30%. ARR = 10% − 7% = 3% (or 0.03). NNT = 1 / 0.03 = 33 — meaning 33 patients must be treated for 3 years to prevent one cardiovascular event. The RRR sounds impressive; the NNT communicates the real-world effort.
Subgroup, Composite, and Surrogate Endpoints
Subgroup analyses are credible only when pre-specified in the protocol, supported by a biological rationale, and accompanied by a formal interaction test. Post-hoc subgroup findings are hypothesis-generating only and prone to false positives from multiplicity (require Bonferroni or similar correction). Composite endpoints (e.g., MACE = cardiovascular death, MI, stroke) increase statistical power but can be dominated by the softest component, masking a true effect on the hardest endpoint. Surrogate endpoints (LDL reduction, HbA1c, blood pressure, tumor shrinkage) allow smaller, shorter trials but require validation — torcetrapib raised HDL and lowered LDL yet increased mortality, a classic surrogate failure.
Meta-Analysis
A meta-analysis pools trial results for a precise summary effect. Quality depends on heterogeneity measured by I² (0% none, 25% low, 50% moderate, 75% substantial) — high I² should prompt a random-effects model and exploration of causes (subgroup, sensitivity analyses). A fixed-effects model assumes a single true effect across trials; random-effects allows true effects to vary. Publication bias — tendency of positive trials to be published — is detected with a funnel plot; asymmetry suggests missing negative studies. Cochrane Risk of Bias 2 tool formalizes within-study risk assessment for pooled analyses.
Applying Trial Results to Patients
A valid trial still may not apply to your patient. Ask: Is my patient similar to enrolled patients in age, sex, comorbidities, and disease severity? Was the comparator and setting similar to my environment? Are the NNT and number needed to harm (NNH) acceptable to this patient given their values? Shared decision making pairs the trial data with patient preferences — a 1% absolute benefit may matter to one patient and not another. Always reconcile NNT with NNH: if a drug prevents one stroke per 33 treated but causes one major bleed per 50, the patient deserves both numbers.
FDA Approval as Evidence Evaluation
The FDA approval process itself is a structured evidence evaluation: Phase I (safety, dosing in healthy volunteers), Phase II (efficacy signal, dose-ranging in patients), Phase III (pivotal efficacy and safety RCTs), and Phase IV (post-marketing surveillance). Accelerated approval permits earlier marketing based on a surrogate endpoint reasonably likely to predict clinical benefit, with confirmatory trials required. Breakthrough designation expedites review for serious conditions with preliminary substantial evidence of improvement over existing therapy. Pharmacists should know the evidence base at approval is often thinner than post-marketing experience reveals.
A non-inferiority trial of a new antibiotic versus vancomycin for MRSA bacteremia reports a pre-specified margin of 10% and a treatment difference of −3% (95% CI −7% to +1%). Which assumption, if violated, most undermines the conclusion that the new drug is non-inferior?
A trial reports that a new statin reduces cardiovascular events from 12% to 9% over 5 years (p = 0.04). Which calculation gives the number needed to treat (NNT)?
Which of the following best distinguishes a predatory journal from a legitimate peer-reviewed journal?