4.1 Study Designs & Evidence Hierarchy
Key Takeaways
- The Oxford Centre for Evidence-Based Medicine positions systematic reviews and meta-analyses of randomized controlled trials at Level 1a, followed by individual RCTs (1b), prospective cohort studies (2b), case-control studies (3b), and case series (4).
- Allocation concealment shields the assignment sequence prior to participant enrollment and can always be implemented, whereas blinding prevents performance and ascertainment bias after assignment but may be unfeasible in surgical or procedural trials.
- Intention-to-treat (ITT) analysis evaluates subjects in their randomized groups regardless of protocol adherence or withdrawal, preserving baseline prognostic balance and avoiding overly optimistic efficacy estimates in superiority trials.
- Per-protocol (PP) analysis evaluates only fully compliant subjects, reflecting maximum biological efficacy; however, it is essential in non-inferiority trials where ITT's tendency to dilute differences toward the null falsely inflates claims of non-inferiority.
- Confounding factors must be associated with the exposure, independently predict the outcome, and not lie on the causal pathway; they are controlled at the design stage (randomization, matching, restriction) or analysis stage (stratification, multivariable regression).
[!NOTE] Curriculum Focus: The MRCP(UK) Part 1 examination rigorously tests clinical research methodology. Candidates must be able to identify study designs from clinical vignettes, evaluate trial safeguards (such as allocation concealment versus blinding), distinguish intention-to-treat from per-protocol analysis, and diagnose subtle biases in published clinical trials.
Evidence-based medicine (EBM) integrates clinical expertise with the best external clinical evidence. Understanding the strengths, limitations, and vulnerability to bias of each study design allows physicians to critically appraise trial literature and translate findings into internal medicine practice.
The Hierarchy of Clinical Evidence
The Oxford Centre for Evidence-Based Medicine (CEBM) ranks clinical research designs based on their susceptibility to systematic bias and their ability to establish causality:
| Level | Study Design | Primary Purpose & Features | Key Limitations |
|---|---|---|---|
| 1a | Systematic Review & Meta-Analysis of RCTs | Quantitative pooling of homogenous RCTs; minimizes random error | Vulnerable to publication bias and clinical heterogeneity |
| 1b | Individual Randomized Controlled Trial (RCT) | Random allocation to intervention vs control; establishes causality | High cost, limited generalisability, artificial protocol adherence |
| 2a | Systematic Review of Cohort Studies | Synthesis of prospective observational data | Subject to confounding across pooled cohorts |
| 2b | Individual Prospective Cohort Study | Tracks exposed and unexposed individuals forward in time to measure incidence | Expensive, long follow-up, prone to attrition bias (loss to follow-up) |
| 3a | Systematic Review of Case-Control Studies | Synthesis of retrospective observational investigations | Vulnerable to shared retrospective biases |
| 3b | Individual Case-Control Study | Compares past exposures between diseased 'cases' and disease-free 'controls' | High risk of recall bias and selection bias; cannot calculate true incidence |
| 4 | Case Series & Poor-Quality Cohort/Case-Control | Descriptive summary of clinical characteristics in a single patient cohort | No comparative control group; cannot determine causality |
| 5 | Expert Opinion & Mechanistic Reasoning | Consensus opinions, physiological reasoning, bench-laboratory animal models | High risk of subjective bias; frequently contradicted by clinical trials |
Meta-Analyses and Heterogeneity Assessment
A meta-analysis pools quantitative data from multiple independent studies to produce a single pooled effect estimate. In the MRCP examination, meta-analysis interpretation focuses on two graphical and statistical tools:
-
The Forest Plot:
- Each individual study is represented by a horizontal line (indicating the 95% confidence interval) and a marker box (point estimate), where the area of the box reflects the study's statistical weight (inversely proportional to sample variance).
- The pooled summary estimate is depicted at the base as a diamond (the lateral points represent the 95% confidence interval; the central vertical axis is the point estimate).
- If the diamond crosses the line of no effect (0 for differences, 1.0 for risk/odds/hazard ratios), the pooled result is not statistically significant.
-
Statistical Heterogeneity:
- Cochran's Q Test: A chi-squared test evaluating whether differences between study results exceed random sampling error. A p-value < 0.10 indicates significant heterogeneity.
- I² Statistic: Quantifies the percentage of total variation across studies attributable to genuine heterogeneity rather than chance:
- I² < 25%: Low heterogeneity (fixed-effects model appropriate).
- I² 25% – 50%: Moderate heterogeneity.
- I² > 50%: Substantial heterogeneity (a random-effects model must be applied, or subgroup analysis conducted).
-
Publication Bias & Funnel Plots:
- Funnel plots graph study precision (sample size or inverse standard error on the vertical y-axis) against treatment effect size (horizontal x-axis).
- In the absence of bias, studies scatter symmetrically in an inverted funnel shape around the true effect.
- An asymmetric funnel plot (specifically missing small studies with negative or null findings in the lower quadrant) indicates publication bias (positive trials are preferentially published). Egger's linear regression test objectively confirms funnel asymmetry.
Observational Study Architectures
When random assignment is unethical (e.g., evaluating tobacco carcinogenicity) or impractical, observational studies are employed:
+-----------------------------------------------------------------------------------------+
| Observational Study Continuum |
+-----------------------------------------------------------------------------------------+
| Direction of Inquiry: |
| |
| [Past: Past Exposure] <============ [Present: Cases vs Controls] (Case-Control) |
| |
| [Present: Exposure & Outcome] (Cross-Sectional) |
| |
| [Present: Exposed vs Unexposed] ===> [Future: Incidence / Outcome] (Prospective Cohort) |
+-----------------------------------------------------------------------------------------+
1. Cohort Studies (Prospective vs Retrospective)
- Architecture: A defined group of subjects free of the outcome of interest is classified by exposure status (e.g., smokers vs non-smokers) and followed over time to compare the incidence of disease.
- Primary Metric: Relative Risk (Risk Ratio, RR) and Incidence Rate.
- Strengths: Preserves clear temporal sequence (exposure precedes outcome); ideal for evaluating rare exposures; can evaluate multiple clinical outcomes from a single exposure.
- Weaknesses: Inefficient for rare diseases (requires tens of thousands of person-years); vulnerable to attrition bias (loss to follow-up over extended observation periods).
- Retrospective Cohort: Both exposure and disease have already occurred when the study begins, but the cohort is defined and reconstructed using historical workplace or hospital records before outcomes are evaluated.
2. Case-Control Studies
- Architecture: Subjects are selected on the basis of disease status—patients with the disease (cases) and individuals without the disease (controls)—and historical records or interviews are examined retrospectively to compare exposure frequencies.
- Primary Metric: Odds Ratio (OR). Because the investigator fixes the number of cases and controls, true population incidence and relative risk cannot be directly calculated.
- Strengths: Highly efficient for rare diseases (e.g., pheochromocytoma, angiosarcoma) and diseases with prolonged latency periods (e.g., mesothelioma); rapid and inexpensive.
- Weaknesses: High vulnerability to recall bias (patients diagnosed with serious illness recall past exposures more meticulously than healthy controls) and selection bias in choosing appropriate control groups.
3. Cross-Sectional Studies (Prevalence Surveys)
- Architecture: Assesses both exposure and disease status simultaneously in a population sample at a single, discrete point in time.
- Primary Metric: Prevalence and Prevalence Odds Ratio.
- Limitations: Cannot determine temporal sequence (whether exposure preceded outcome); vulnerable to reverse causality (e.g., finding that physical inactivity is associated with depression does not reveal whether inactivity causes depression or depression causes inactivity).
Randomized Controlled Trial (RCT) Design Concepts
Randomized controlled trials represent the benchmark for establishing therapeutic efficacy because randomization distributes known and unknown confounding variables equally across trial arms.
Randomization Modalities
- Simple Randomization: Unrestricted random assignment (analogous to flipping a coin). In small trials (n < 100), it can produce unequal group sizes or prognostic imbalances by chance.
- Block Randomization: Randomizes participants in predetermined block sizes (e.g., blocks of 4, 6, or 8) to ensure equal participant numbers across comparative arms throughout the trial.
- Stratified Block Randomization: Patients are first grouped into homogenous strata based on critical prognostic variables (e.g., age < 65 vs >= 65; presence of chronic kidney disease), and then block randomization is executed within each stratum. This guarantees balance for high-impact confounding covariates.
- Cluster Randomization: Pre-existing groups (e.g., entire medical wards, GP surgeries, geographic villages) are randomized as clusters rather than individuals, preventing treatment contamination across participants.
Allocation Concealment vs Blinding (Masking)
The MRCP Part 1 frequently tests this distinction:
+------------------------------------------------------------------------------------------+
| Allocation Concealment vs Blinding Timeline |
+------------------------------------------------------------------------------------------+
| [Patient Enrolled] --> [ALLOCATION CONCEALMENT] --> [Treatment Assigned] |
| • Prevents Selection Bias | |
| • Always Feasible v |
| [BLINDING / MASKING] --> [Outcome] |
| • Prevents Performance/Observer Bias |
| • Not Always Feasible (e.g., Surgery) |
+------------------------------------------------------------------------------------------+
- Allocation Concealment: Protects the randomization schedule before the patient enters the trial and is assigned a treatment. It prevents recruiters from predicting which group the next eligible patient will enter (e.g., using centralized telephone/web-based randomization services or sequentially numbered, opaque, sealed envelopes [SNOSE]). It eliminates selection bias and can always be performed.
- Blinding (Masking): Conceals the assigned intervention after allocation has occurred:
- Single-blind: Patients are unaware of treatment assignment.
- Double-blind: Both patients and treating healthcare providers/outcome assessors are unaware of assignment. Prevents performance bias (differential clinical care outside the protocol) and detection/ascertainment bias (differential measurement of outcomes).
- Triple-blind: Patients, clinicians, outcome assessors, and the data safety monitoring statisticians remain masked until final database lock.
Intention-to-Treat (ITT) vs Per-Protocol (PP) Analysis
When trial participants discontinue medications, deviate from protocols, or cross over to alternative arms, analytical strategy dictates validity:
- Intention-to-Treat (ITT):
- Rule: "Once randomized, always analyzed." Every patient is evaluated within their assigned group, regardless of non-compliance, adverse events, or complete treatment discontinuation.
- Rationale: Preserves the prognostic baseline balance established by randomization; reflects pragmatic real-world clinical effectiveness; provides a conservative estimate of therapeutic superiority.
- Gold Standard: Mandated by regulatory agencies (NICE, EMA, FDA) for superiority trials.
- Per-Protocol (PP):
- Rule: Evaluates only those patients who completed the pre-specified treatment regimen with full protocol adherence.
- Rationale: Assesses maximum biological efficacy under ideal conditions.
- Danger in Superiority Trials: Destroys randomization balance, introducing severe attrition and confounding biases.
- Mandatory Role in Non-Inferiority Trials: In non-inferiority trials, ITT dilutes treatment differences toward the null (making two different therapies look identically mediocre due to non-adherence). Thus, both ITT and PP analyses must demonstrate non-inferiority to confirm genuine therapeutic equivalence.
- As-Treated Analysis: Re-classifies subjects according to the treatment they actually received. This completely obliterates randomization and reduces an RCT to an observational cohort.
Crossover Trials and Washout Periods
In a crossover trial, each subject receives intervention A for a defined period, followed by intervention B, allowing each patient to act as their own control. This removes between-patient confounding and achieves high statistical power with smaller sample sizes.
- Strict Prerequisites: The condition must be chronic, stable, and incurable (e.g., stable angina, mild asthma, essential hypertension). It cannot be used for acute infections, surgical cures, or rapidly progressive malignancies.
- Carryover Effect: The residual physiological action of the first drug persisting into the second phase. To prevent this, an adequate washout period (typically >= 5 half-lives of the drug) is mandatory.
Superiority vs Non-Inferiority Trials
- Superiority Trial: Aims to demonstrate that a novel intervention is statistically significantly better than an active comparator or placebo (H0: Treatment A = Treatment B).
- Non-Inferiority Trial: Aims to establish that a new intervention is not unacceptably worse than current standard of care by more than a pre-specified margin (-Delta). Typically conducted when the new drug offers secondary advantages (e.g., oral administration instead of daily intravenous infusions, lower cost, reduced hepatotoxicity).
Systematic Bias and Confounding
Bias represents systematic (non-random) error that distorts study results away from the truth.
Key Biases in Clinical Research
- Selection Bias:
- Berkson's Bias (Admission Rate Bias): Hospitalized patients have higher rates of multiple comorbidities and exposure profiles compared to the general public, distorting apparent associations.
- Healthy Worker Effect: Employed populations exhibit lower baseline morbidity and mortality than the general population (which includes the chronically disabled).
- Measurement & Information Bias:
- Recall Bias: Differential accuracy of historical recall between diseased cases and healthy controls.
- Hawthorne Effect: Modification of participant behavior simply because they know they are being observed.
- Observer / Detection Bias: Clinician assessors evaluate subjective outcomes more favorably in patients receiving the experimental therapy (mitigated by double-blinding).
- Screening Biases (High-Yield for MRCP):
- Lead-Time Bias: Early diagnosis via screening creates the mathematical illusion of prolonged survival without altering the natural history or calendar date of death. Diagnosed by evaluating population mortality rates rather than 5-year survival from diagnosis.
- Length-Time Bias: Screening disproportionately identifies indolent, slow-progressing lesions with prolonged asymptomatic phases, while aggressive, rapidly fatal cases present symptomatically between screening intervals. This creates an overestimation of screening benefit.
Confounding and Methods of Control
A confounder is an extraneous variable that is:
- Independently associated with the exposure.
- An independent risk factor for the outcome.
- Not an intermediate step on the causal pathway between exposure and outcome.
[ Confounder: Cigarette Smoking ]
/ \
/ \
v v
[ Exposure: Coffee Drinking ] ------> [ Outcome: Pancreatic Cancer ]
(No Causal Link)
Smoking is associated with coffee drinking and independently causes cancer; coffee drinking does not cause cancer.
Methods to Control Confounding
- At Study Design Phase:
- Randomization: The only method that balances both known and unknown confounders (applicable only to experimental trials).
- Restriction: Restricting enrollment criteria (e.g., studying only non-smokers). Limits sample size and generalisability.
- Matching: Matching cases and controls for key confounders (e.g., age, sex, socioeconomic status).
- At Data Analysis Phase:
- Stratification: Analyzing data within discrete subgroups (e.g., Mantel-Haenszel pooling across age brackets). If the crude relative risk differs from the stratum-specific risks, confounding was present.
- Multivariable Regression: Mathematical models adjusting for multiple confounders simultaneously (multiple linear regression for continuous outcomes; multivariable logistic regression for binary outcomes; Cox proportional hazards regression for time-to-event outcomes).
A multicentre surgical trial compares laparoscopic hemicolectomy with open resection for locally advanced colon cancer. Because of visible surgical scars, operating surgeons and patients cannot be blinded to the allocated intervention. Which trial design safeguard is most critical to prevent selection bias prior to participant enrollment?
A national healthcare authority reviews outcomes following the introduction of a low-dose computed tomography (LDCT) lung cancer screening initiative. The screened cohort exhibits a median survival of 5.2 years from diagnosis compared with 2.8 years in an unscreened historical cohort. However, regional age-adjusted 10-year lung cancer mortality rates remain entirely unchanged between screened and unscreened populations. Which form of bias best explains this finding?
A phase III randomized trial evaluates whether a novel oral Janus kinase (JAK) inhibitor is non-inferior to subcutaneous adalimumab in active rheumatoid arthritis. The protocol specifies a non-inferiority margin of -10% in the ACR50 response rate. During trial monitoring, 18% of patients in the oral JAK inhibitor arm fail to take their medication consistently. Why must a per-protocol (PP) analysis be executed alongside an intention-to-treat (ITT) analysis in this non-inferiority study?