13.3 Research Methodology, Biostatistics & Evidence-Based Practice
Key Takeaways
- Study designs span a hierarchy from case reports (weakest) to randomized controlled trials and systematic reviews/meta-analyses (strongest); blinding, randomization, and allocation concealment reduce bias in RCTs.
- Sensitivity (true positive rate) and specificity (true negative rate) characterize a test independent of prevalence; positive/negative predictive values depend on disease prevalence.
- A p-value is the probability of observing data as extreme as observed if the null hypothesis is true; a 95% confidence interval that excludes the null (e.g., 0 for differences, 1 for relative risks/odds ratios) indicates statistical significance.
- Evidence-based practice integrates the best research evidence, clinical expertise, and patient values; GRADE and levels of evidence rate recommendation strength.
Research Methodology, Biostatistics & Evidence-Based Practice
Research and statistics is a Domain A applied-science topic that recurs in item stems requiring interpretation of study design, diagnostic test performance, and effect estimates.
Study Design Hierarchy
Case report/series < Case-control < Cohort < Randomized Controlled Trial < Systematic Review / Meta-analysis
| Design | Strength | Key Feature / Limitation |
|---|---|---|
| Case-control | Observational, retrospective | Odds ratio; prone to recall and selection bias; efficient for rare diseases |
| Cohort | Observational, prospective/retrospective | Relative risk; temporal relationships; confounding |
| RCT | Experimental | Randomization, blinding, allocation concealment reduce bias; gold standard for therapy |
| Systematic review/meta-analysis | Synthesis | Pooled estimates; heterogeneity must be assessed (I-squared) |
Bias to recognize: selection, recall, measurement/observer, publication (in meta-analysis). Confounding is mitigated by randomization (RCT) or multivariable adjustment (observational).
Diagnostic Test Performance
| Metric | Definition |
|---|---|
| Sensitivity (Sn) | Proportion with disease who test positive (true positive rate); high sensitivity → rules OUT (SnNOUT) |
| Specificity (Sp) | Proportion without disease who test negative (true negative rate); high specificity → rules IN (SpPIN) |
| PPV | Proportion testing positive who truly have disease; rises with prevalence |
| NPV | Proportion testing negative who truly lack disease; falls with prevalence |
| Likelihood ratio | LR+ = sensitivity/(1-specificity); LR- = (1-sensitivity)/specificity; combine with pretest probability |
Disease + Disease -
Test + a (TP) b (FP)
Test - c (FN) d (TN)
Sensitivity = a/(a+c); Specificity = d/(b+d); PPV = a/(a+b); NPV = d/(c+d)
Sensitivity and specificity are intrinsic to the test and do not change with prevalence; PPV and NPV do change with prevalence.
Hypothesis Testing & Confidence Intervals
- Null hypothesis (H0): no effect/difference; alternative (H1): an effect exists.
- p-value: probability of observing results as extreme as observed if H0 is true; a small p (commonly <0.05) leads to rejecting H0, but does not measure effect size or clinical importance.
- Type I error (alpha): false positive (finding an effect that does not exist); Type II error (beta): false negative (missing a real effect); power = 1 - beta.
- 95% confidence interval: range that, over repeated samples, would contain the true parameter 95% of the time; if it excludes the null (0 for absolute differences, 1 for ratios), the result is statistically significant at the 0.05 level.
- Effect size: clinical importance differs from statistical significance—a large trial can detect trivial differences; report and interpret effect sizes and CIs.
Common Statistical Tests
| Comparison | Test |
|---|---|
| Two group means (parametric) | t-test |
| >2 group means | ANOVA |
| Categorical association | Chi-square / Fisher exact (small samples) |
| Paired/repeated measures | Paired t-test, repeated-measures ANOVA |
| Correlation | Pearson (linear), Spearman (nonparametric) |
| Time-to-event | Kaplan-Meier, log-rank, Cox regression |
Evidence-Based Practice (EBP)
EBP integrates (1) the best available research evidence, (2) clinical expertise, and (3) patient values and preferences. GRADE rates the certainty of evidence and strength of recommendations, considering risk of bias, inconsistency, indirectness, imprecision, publication bias, and magnitude of effect. Appraisal tools (CASP) structure critical reading. Candidates should distinguish statistical significance from clinical importance and interpret CIs and LRs.
Observational vs Experimental Designs & Confounding
Observational studies (cohort, case-control) cannot fully exclude confounding even with adjustment; RCTs, through randomization, balance known and unknown confounders. When RCTs are unethical or impractical (e.g., effect of a harmful exposure), well-designed observational studies with techniques like propensity-score matching provide evidence, interpreted with caution. Recognizing residual confounding is essential to applying evidence to individual patients.
Likelihood Ratios & Clinical Application
Likelihood ratios combine test performance with pretest probability to generate post-test probability, independent of prevalence-driven PPV/NPV limitations:
Pretest Probability ──► Pretest Odds ──► x LR ──► Post-test Odds ──► Post-test Probability
An LR+ >10 or LR- <0.1 meaningfully shifts probability. This framework helps the physiatrist decide when an additional test (e.g., imaging before injection) changes management, versus when it is low-value.
Effect Size & Clinical Significance
Statistical significance (p<0.05) does not imply clinical importance. A large trial can detect a trivial between-group difference; conversely, a small trial may miss a clinically important effect (underpowered). Candidates should interpret effect sizes (mean difference, relative risk, odds ratio, NNT) and confidence intervals together, distinguishing "statistically detectable" from "clinically meaningful." Number-needed-to-treat (NNT = 1/absolute risk reduction) communicates practical impact.
Critical Appraisal & Applying Evidence
| Appraisal Question | What to Check |
|---|---|
| Validity | Randomization, blinding, allocation concealment, ITT analysis |
| Importance | Effect size, confidence interval, NNT, clinical relevance |
| Applicability | Patient similarity, setting, feasibility, patient values |
Evidence-based practice integrates appraisal with clinical expertise (judging applicability, technical skill) and patient values (preferences, goals, risk tolerance). GRADE rates certainty of evidence (high/moderate/low/very low) and strength of recommendations (strong/weak), incorporating consistency, directness, precision, and magnitude. The physiatrist applies this to rehab interventions—many of which are complex, multimodal, and benefit from individualized application of group-level evidence.
A screening test has very high sensitivity but low specificity. Which statement is most accurate?
A randomized trial reports a treatment effect with a 95% confidence interval of (0.8, 1.4) for a relative risk. What is the correct interpretation?
Which study design provides the strongest evidence for the efficacy of a rehabilitation intervention?