15.2 Clinical Study Designs & Measures of Effect

Key Takeaways

  • The hierarchy of clinical evidence ascends from expert opinion and observational designs up to randomized controlled trials (RCTs) and systematic reviews/meta-analyses, balancing observational efficiency against interventional control of confounding.

  • RCT designs include parallel-group, factorial, and crossover designs (requiring sufficient drug washout of at least 5 half-lives to eliminate carryover effects); crossover designs are contraindicated in acute, curable, or rapidly evolving emergency conditions.

  • Cohort studies follow exposed and unexposed populations over time to calculate Relative Risk (RR), whereas Case-Control studies sample based on outcome presence to calculate Odds Ratios (OR), representing the optimal observational design for rare clinical diseases and severe adverse drug reactions.

  • Absolute Risk Reduction (ARR = CER - EER) determines the Number Needed to Treat (NNT = 1 / ARR), which must always be rounded UP to the next integer; Number Needed to Harm (NNH = 1 / ARI) must always be rounded DOWN to the preceding integer.

  • Diagnostic test accuracy evaluates Sensitivity (SnNOut) and Specificity (SpPIn) as intrinsic properties, while Positive and Negative Predictive Values depend directly on pre-test disease prevalence; Likelihood Ratios (>10 or <0.1) drive substantial shifts in post-test disease probability.

Last updated: October 2026

15.2 Clinical Study Designs & Measures of Effect

Note

Independent study resource provided by OpenExamPrep. Content is organized around official Board of Pharmacy Specialties (BPS) Emergency Medicine Pharmacy examination specifications.

The Hierarchy of Clinical Evidence

Evidence-based emergency medicine ranks clinical research designs based on their inherent ability to control for confounding, eliminate systematic bias, and establish causal relationships between pharmacotherapy and clinical outcomes.

                               HIERARCHY OF CLINICAL EVIDENCE
                                             ▲
                                            / \
                                           /   \
                                          /     \
                                         / Meta- \
                                        /Analyses \
                                       / & System- \
                                      /atic Reviews \
                                     /───────────────\
                                    /   Randomized    \
                                   /Controlled Trials  \
                                  /─────────────────────\
                                 /     Cohort Studies    \
                                /  (Prospective/Retro)    \
                               /───────────────────────────\
                              /     Case-Control Studies    \
                             /───────────────────────────────\
                            /     Cross-Sectional Surveys     \
                           /───────────────────────────────────\
                          /      Case Series & Case Reports     \
                         /───────────────────────────────────────\
                        / In Vitro Models, Animal Studies, Opinion\
                       └───────────────────────────────────────────┘
  1. Systematic Reviews & Meta-Analyses: Collate and synthesize all available clinical trials meeting predefined eligibility criteria. Meta-analyses quantitatively pool data from homogeneous RCTs using fixed-effect or random-effects models. Results are displayed via forest plots, and study heterogeneity is quantified using the I2I^2 statistic (<25%<25\% low, 25%–50%25\%–50\% moderate, >50%>50\% high heterogeneity).
  2. Randomized Controlled Trials (RCTs): Interventional design allocating subjects to active or control arms via strict random sequence generation. The gold standard for proving pharmacotherapeutic causality.
  3. Cohort Studies: Observational longitudinal design following exposed and unexposed populations over time to assess incidence of outcomes. Establishes temporal sequence; primary measure is Relative Risk (RR).
  4. Case-Control Studies: Observational retrospective design identifying patients with an outcome (Cases) and matching them with unaffected subjects (Controls) to evaluate prior drug exposure. Primary measure is Odds Ratio (OR).
  5. Cross-Sectional Studies: Observational assessment of exposure and disease prevalence at a single discrete point in time. Cannot determine temporality or causality.
  6. Case Series and Case Reports: Descriptive, uncontrolled observational reports of novel drug interactions, toxicities, or off-label rescue therapies. Hypothesis-generating only.
  7. In Vitro Research, Animal Models & Expert Opinion: Biological plausibility foundation; lowest clinical certainty.

Randomized Controlled Trial (RCT) Architectures

Parallel-Group Design

Patients are randomized to two or more distinct arms (e.g., Investigational Drug vs Standard Active Comparator or Placebo) and remain in their assigned treatment groups concurrently for the entire duration of the study. This is the standard architecture for emergency medicine resuscitation and critical care trials.

Crossover Design

Each patient receives both comparative treatments sequentially, acting as their own matched control:

  • Sequence: Group 1 receives Drug A →\rightarrow Washout Period →\rightarrow Drug B. Group 2 receives Drug B →\rightarrow Washout Period →\rightarrow Drug A.
  • The Washout Period: A mandatory drug-free interval designed to eliminate the carryover effect (residual pharmacologic or clinical effects of the first drug persisting into the second treatment phase). The washout period must span at least 5 elimination half-lives (5×t1/25 \times t_{1/2}) of the initial drug and its active metabolites.
  • Emergency Medicine Contraindication: Crossover designs are fundamentally invalid in acute, unstable, or curable conditions (e.g., septic shock, cardiac arrest, acute ischemic stroke, status epilepticus, polytrauma). Crossover trials require a stable, chronic baseline condition (e.g., stable chronic asthma, mild persistent pain).

Factorial Design

A factorial design evaluates two or more distinct interventions simultaneously within a single trial population. In a classic 2×22 \times 2 factorial trial, patients are randomized into four distinct groups:

  1. Treatment A + Treatment B
  2. Treatment A + Placebo B
  3. Placebo A + Treatment B
  4. Placebo A + Placebo B

This design allows investigators to evaluate the independent primary effects of both drugs while assessing potential synergistic or antagonistic pharmacological interactions.

Blinding Methodologies

  • Open-Label: Both clinicians and patients know the assigned therapy. Prone to severe performance and ascertainment bias.
  • Single-Blind: Only the patient is unaware of treatment assignment.
  • Double-Blind: Both patient and treating clinical investigators are blinded. Prevents differential co-interventions.
  • Triple-Blind: Patient, bedside clinical team, and independent outcome adjudication/statistical analysis committees remain blinded until final database lock.

Analysis Populations: Intention-to-Treat vs Per-Protocol

                             CLINICAL TRIAL ANALYSIS POPULATIONS
  ┌─────────────────────────────────────────────────────────────────────────────┐
  │ Intention-to-Treat (ITT) Analysis                                           │
  │ • "Once randomized, always analyzed" in originally assigned group          │
  │ • Analyzes all subjects regardless of adherence, crossover, or withdrawal   │
  │ • Preserves prognostic baseline balance achieved by randomization           │
  │ • Prevents attrition bias; reflects real-world clinical effectiveness       │
  │ • Gold standard for SUPERIORITY trials (conservative estimate)              │
  ├─────────────────────────────────────────────────────────────────────────────┤
  │ Per-Protocol (PP) / On-Treatment Analysis                                   │
  │ • Analyzes ONLY subjects who completed full protocol without violations     │
  │ • Excludes non-compliant patients, dropouts, and treatment crossovers       │
  │ • Introduces significant selection bias and breaks randomization balance    │
  │ • Reflects ideal biological efficacy under perfect compliance               │
  │ • Required alongside ITT in NON-INFERIORITY trials to prevent false claims │
  └─────────────────────────────────────────────────────────────────────────────┘

Observational Study Designs: Cohort vs Case-Control vs Cross-Sectional

When randomized trials are unethical, unfeasible, or logistically impractical, observational study designs provide vital clinical insight:

Study FeatureProspective CohortRetrospective CohortCase-ControlCross-Sectional
Starting PointDefined by Exposure statusDefined by Exposure statusDefined by Outcome / Disease statusPopulation sample at single timepoint
Direction of InquiryForward in time (Follow-up)Backward in existing recordsBackward in time (Exposure history)Simultaneous assessment
Primary MeasureRelative Risk (RR)Relative Risk (RR)Odds Ratio (OR)Prevalence Ratio
Optimal Clinical UseCommon exposures; establishing incidence & natural historyRapid evaluation of past exposures using electronic health recordsRare diseases (<5%<5\%) and rare adverse drug reactionsEpidemiologic disease surveillance & health resource planning
Key StrengthsEstablishes temporality; minimizes recall bias; measures multiple outcomesFaster & cheaper than prospective; excellent for ED registriesHighly efficient for rare outcomes; small sample size requiredRapid, inexpensive; generates clinical hypotheses
VulnerabilitiesExpensive; attrition bias; confounding by indicationSelection bias; missing historical chart dataSevere recall bias; cannot directly calculate incidence or RRCannot establish temporality or causality
Emergency ExampleFollowing sepsis patients given balanced crystalloids vs salineChart review of ketamine vs etomidate for post-intubation hypotensionMatching ED patients with Stevens-Johnson Syndrome to controlsAssessing point prevalence of methicillin-resistant S. aureus

Important

On the BCEMP exam, remember that Case-Control studies cannot measure Relative Risk (RR) because the investigator artificially fixes the proportion of cases and controls (e.g., 1 case to 4 controls). Only the Odds Ratio (OR) can be calculated.


Quantitative Measures of Association & Effect Size

Clinical trials and observational studies summarize binary outcome data using a standard 2×22 \times 2 contingency table:

                                  $2 \times 2$ CONTINGENCY MATRIX
                                Disease / Outcome Present   Disease / Outcome Absent     Total
  Exposed / Investigational                a                           b                 a + b
  Unexposed / Control                      c                           d                 c + d
  Total                                  a + c                       b + d             a + b + c + d
  • Experimental Event Rate (EER): Risk in experimental group =aa+b= \frac{a}{a + b}
  • Control Event Rate (CER): Risk in control group =cc+d= \frac{c}{c + d}

Relative Risk (Risk Ratio, RR)

The ratio of the probability of an outcome occurring in the exposed/treated group compared to the unexposed/control group:

RR=EERCER=a/(a+b)c/(c+d)\text{RR} = \frac{\text{EER}}{\text{CER}} = \frac{a / (a + b)}{c / (c + d)}

  • RR=1.0\text{RR} = 1.0: No difference in risk between groups.
  • RR<1.0\text{RR} < 1.0: Exposure/treatment reduces risk of outcome (protective effect).
  • RR>1.0\text{RR} > 1.0: Exposure/treatment increases risk of outcome (harmful effect).

Odds Ratio (OR)

The ratio of the odds of exposure among cases compared to the odds of exposure among controls (or odds of event vs non-event):

OR=Odds in ExposedOdds in Unexposed=a/bc/d=a×db×c\text{OR} = \frac{\text{Odds in Exposed}}{\text{Odds in Unexposed}} = \frac{a / b}{c / d} = \frac{a \times d}{b \times c}

  • The Rare Disease Assumption: When an outcome is rare in the underlying population (typically incidence <5%<5\% to 10%10\%), aa is negligible relative to bb, and cc is negligible relative to dd. Under these conditions, aa+b≈ab\frac{a}{a+b} \approx \frac{a}{b} and cc+d≈cd\frac{c}{c+d} \approx \frac{c}{d}, meaning OR≈RR\text{OR} \approx \text{RR}.

Relative Risk Reduction (RRR)

The proportional reduction in risk achieved by the experimental intervention relative to the baseline control risk:

RRR=CER−EERCER=1−RR\text{RRR} = \frac{\text{CER} - \text{EER}}{\text{CER}} = 1 - \text{RR}

Warning

Pharmaceutical literature frequently highlights RRR because it yields large, impressive percentages that can mislead clinicians. For instance, if mortality drops from 2%2\% to 1%1\%, the RRR is 50%50\%, but the true absolute benefit is only 1%1\%. Always calculate the Absolute Risk Reduction (ARR).

Absolute Risk Reduction (ARR) & Absolute Risk Increase (ARI)

The true arithmetic difference in event rates between groups:

ARR=CER−EER(for beneficial interventions)\text{ARR} = \text{CER} - \text{EER} \quad (\text{for beneficial interventions}) ARI=EER−CER(for harmful adverse effects)\text{ARI} = \text{EER} - \text{CER} \quad (\text{for harmful adverse effects})

Number Needed to Treat (NNT) & Number Needed to Harm (NNH)

NNT=1ARR=1CER−EER\text{NNT} = \frac{1}{\text{ARR}} = \frac{1}{\text{CER} - \text{EER}} NNH=1ARI=1EER−CER\text{NNH} = \frac{1}{\text{ARI}} = \frac{1}{\text{EER} - \text{CER}}

Critical Biostatistical Rounding Rules

  • NNT Rounding Rule: NNT must always be rounded UP to the next whole integer, regardless of the decimal value. Rounding down would overestimate treatment efficacy. For example, if ARR=0.043\text{ARR} = 0.043 (4.3%4.3\%), NNT=1/0.043=23.255→\text{NNT} = 1 / 0.043 = 23.255 \rightarrow NNT = 24.
  • NNH Rounding Rule: NNH must always be rounded DOWN to the preceding whole integer, regardless of the decimal value. Rounding up would underestimate the risk of drug harm. For example, if ARI=0.038\text{ARI} = 0.038 (3.8%3.8\%), NNH=1/0.038=26.315→\text{NNH} = 1 / 0.038 = 26.315 \rightarrow NNH = 26.

Diagnostic Test Evaluation: 2×22 \times 2 Matrix & Clinical Utility

Emergency medicine relies heavily on rapid diagnostic tests (point-of-care ultrasound, high-sensitivity troponin, D-dimer, venous blood gas). Test performance is evaluated against a definitive "gold standard" reference:

                               DIAGNOSTIC $2 \times 2$ CONTINGENCY MATRIX
                                      Disease Present            Disease Absent
  Positive Index Test             True Positive (TP)         False Positive (FP)
  Negative Index Test             False Negative (FN)        True Negative (TN)
  Total                            TP + FN (Total Diseased)   FP + TN (Total Non-Diseased)

Sensitivity (SnNOut)

The proportion of patients with the disease who test positive:

Sensitivity=TPTP+FN\text{Sensitivity} = \frac{\text{TP}}{\text{TP} + \text{FN}}

  • Clinical Application (SnNOut): A test with very High Sensitivity, when Negative, rules Out disease. Tests with high sensitivity have very few false negatives. A negative high-sensitivity cardiac troponin or high-sensitivity D-dimer safely excludes acute myocardial infarction or pulmonary embolism in low-to-intermediate risk patients.

Specificity (SpPIn)

The proportion of patients without the disease who test negative:

Specificity=TNTN+FP\text{Specificity} = \frac{\text{TN}}{\text{TN} + \text{FP}}

  • Clinical Application (SpPIn): A test with very High Specificity, when Positive, rules In disease. Highly specific tests have very few false positives. Non-contrast CT brain demonstrating hyperdense intracranial blood has near-100%100\% specificity, definitively ruling in hemorrhagic stroke.

Predictive Values: The Influence of Disease Prevalence

  • Positive Predictive Value (PPV): Probability that a patient with a positive test truly has the disease: PPV=TPTP+FP\text{PPV} = \frac{\text{TP}}{\text{TP} + \text{FP}}
  • Negative Predictive Value (NPV): Probability that a patient with a negative test is truly free of disease: NPV=TNTN+FN\text{NPV} = \frac{\text{TN}}{\text{TN} + \text{FN}}

Important

Sensitivity and Specificity are intrinsic operating properties of a diagnostic test that remain relatively constant across patient cohorts. In contrast, PPV and NPV depend directly on disease prevalence (pre-test probability) in the tested population:

  • As disease prevalence increases: PPV increases, while NPV decreases.
  • As disease prevalence decreases: NPV increases, while PPV decreases.

Likelihood Ratios (Prevalence-Independent Test Power)

Likelihood ratios quantify the magnitude of shift from pre-test probability to post-test probability, independent of population disease prevalence:

Positive Likelihood Ratio (LR+)=Sensitivity1−Specificity=True Positive RateFalse Positive Rate\text{Positive Likelihood Ratio (LR+)} = \frac{\text{Sensitivity}}{1 - \text{Specificity}} = \frac{\text{True Positive Rate}}{\text{False Positive Rate}} Negative Likelihood Ratio (LR-)=1−SensitivitySpecificity=False Negative RateTrue Negative Rate\text{Negative Likelihood Ratio (LR-)} = \frac{1 - \text{Sensitivity}}{\text{Specificity}} = \frac{\text{False Negative Rate}}{\text{True Negative Rate}}

Likelihood RatioMagnitude of Shift in Post-Test ProbabilityClinical Utility in Emergency Decision-Making
LR+ >10>10Substantial increase (>45%>45\%)Often definitive; virtually confirms diagnosis
LR+ 55 to 1010Moderate increase (30%–45%30\%–45\%)Clinically meaningful shift; usually warrants therapy
LR+ 22 to 55Small increase (15%–30%15\%–30\%)Modest contribution; rarely definitive alone
LR =1.0= 1.0No change (0%0\%)Completely uninformative test; post-test prob = pre-test prob
LR- 0.20.2 to 0.50.5Small decrease (−15%-15\% to −30%-30\%)Modest contribution; rarely sufficient to rule out
LR- 0.10.1 to 0.20.2Moderate decrease (−30%-30\% to −45%-45\%)Clinically meaningful reduction in probability
LR- <0.1<0.1Substantial decrease (>45%>45\%)Often definitive; safely rules out acute disease
Test Your Knowledge

In a large multicenter emergency trial evaluating early administration of tranexamic acid (TXA, 1 g IV bolus over 10 min followed by 1 g IV infusion over 8 h) versus placebo in severe traumatic hemorrhagic shock, all-cause 28-day mortality was 14.5% in the TXA group and 16.0% in the placebo group. Concurrently, the incidence of non-fatal thrombotic events (deep vein thrombosis or pulmonary embolism) was 2.1% in the TXA group and 1.5% in the placebo group. What are the calculated Number Needed to Treat (NNT) to prevent one death and the Number Needed to Harm (NNH) for one additional thrombotic event, applying standard biostatistical rounding conventions?

A

NNT = 66 (rounded up); NNH = 166 (rounded down).

B

NNT = 66 (rounded down); NNH = 167 (rounded up).

C

NNT = 67 (rounded up); NNH = 167 (rounded up).

D

NNT = 67 (rounded up); NNH = 166 (rounded down).

Test Your Knowledge

An emergency department clinical researcher wishes to investigate whether recent exposure to second-generation antipsychotics is associated with the rare development of neuroleptic malignant syndrome (NMS), a condition occurring in less than 0.02% of psychiatric presentations to the ED. What is the most methodologically rigorous and logistically efficient study design for this investigation, and what measure of association should be reported?

A

Cross-sectional survey assessing concurrent antipsychotic prescriptions and elevated creatine kinase levels; measure of association is Absolute Risk Reduction (ARR).

B

Crossover randomized controlled trial administering antipsychotic therapy followed by a 2-week washout period; measure of association is Number Needed to Harm (NNH).

C

Case-control study identifying patients presenting with NMS (cases) and matching them to ED psychiatric patients without NMS (controls); measure of association is the Odds Ratio (OR).

D

Prospective cohort study following 50,000 patients prescribed antipsychotics forward in time over 5 years; measure of association is Relative Risk (RR).

Test Your Knowledge

A point-of-care rapid diagnostic assay for acute pulmonary embolism is evaluated in the emergency department against computed tomography pulmonary angiography (CTPA). The test demonstrates a sensitivity of 96% and a specificity of 85%. In a low-acuity fast-track ED cohort where the pre-test disease prevalence of PE is 3%, compared to a high-acuity resuscitation bay where the PE prevalence is 30%, how will the diagnostic performance metrics change?

A

Sensitivity, specificity, PPV, and NPV are intrinsic test characteristics that remain entirely constant regardless of clinical prevalence.

B

Positive Predictive Value (PPV) will be markedly higher in the high-acuity resuscitation bay, while Negative Predictive Value (NPV) will remain higher in the low-acuity fast-track cohort.

C

Sensitivity and specificity will both increase substantially in the high-acuity resuscitation bay due to spectrum bias.

D

Negative Predictive Value (NPV) will increase in the high-acuity bay because high specificity ensures complete exclusion of non-diseased patients.

Sections you finish are checked off in the contents.