7.4 Evidence-Based Research Appraisal
Key Takeaways
- Match design to question: a therapy question wants an RCT (ideally intention-to-treat); prognosis wants a cohort; rare harm or etiology may use case-control; expert opinion is the weakest therapy evidence.
- A 50% relative-risk reduction from 4% to 2% is a 2-percentage-point absolute reduction; NNT is about 50. Relative risk without absolute risk is a marketing voice.
- A p-value below 0.05 is not the same as clinical importance. Read effect size, confidence intervals, and whether the outcome matters to this patient.
- Intention-to-treat keeps patients in the group to which they were randomized and answers effectiveness; per-protocol can exaggerate benefit.
- Apply the paper to this patient — age, comorbidity, and SDOH — and inspect predatory journals and industry framing before you change the plan.
Official TCO skill in Domain III Planning: appraise design, results, and clinical applicability. You will not run a biostatistics seminar at Prometric. You will look at a stem that quotes a paper — or describes a design without naming it — and decide whether the evidence is strong enough to change this patient's plan. That is Planning, not journal club as a hobby.
Match the design to the question
Levels of evidence are a ladder, not a personality test.
| Question you want answered | Strongest practical design | Weaker cousins |
|---|---|---|
| Does this therapy work? | Randomized controlled trial; best when pooled in a systematic review or meta-analysis of RCTs | Cohort “we gave it to whoever showed up”; case series; expert opinion |
| What is this patient's prognosis? | Cohort that starts with the condition and follows forward | Case-control (backwards); anecdote |
| Did this exposure cause a rare harm? | Case-control (efficient for rare outcomes); large cohort if the harm is common enough | Case reports |
| Does this test discriminate disease? | Cross-sectional (or cohort) against an independent reference standard | Using the new test as its own gold standard |
| What do experts believe? | Useful for values and rare situations | Not the top of the ladder for a therapy question |
RCT: investigators assign the intervention by chance. Randomization, if it works, balances known and unknown confounders. Blinding reduces bias in how outcomes are ascertained. Allocation concealment is not the same as blinding — concealment happens before assignment.
Cohort: people are grouped by exposure (smoking, an A1c, a drug already chosen in clinic) and followed for an outcome. Prospective is cleaner than retrospective. Confounding is the tax you pay for not randomizing.
Case-control: start with people who have the outcome (cases) and people who do not (controls), then look backward at exposures. Odds ratios approximate risk when the outcome is rare. You cannot read incidence off a case-control table.
Expert opinion / consensus / pathophysiology: the bottom of the therapy ladder. Necessary when trials do not exist (pregnancy emergencies, rare inborn errors). Wrong answer when a good RCT already answered the question.
ANCC-style item: which study design answers a therapy question? The answer is an RCT (or a meta-analysis of RCTs), not a case-control of hobby preference and not a single case report.
Absolute risk, relative risk, NNT
Industry adjectives live in relative risk. Clinicians should also compute absolute risk.
- Control event rate (CER) = how often the outcome happens on placebo or usual care
- Experimental event rate (EER) = how often it happens on the new drug
- Absolute risk reduction (ARR) = CER − EER
- Relative risk reduction (RRR) = (CER − EER) / CER
- Number needed to treat (NNT) = 1 / ARR (use the decimal form of ARR)
Worked example. Stroke occurs in 4% on placebo and 2% on drug. RRR = (0.04 − 0.02) / 0.04 = 50%. That 50% will be the press-release number. ARR = 0.02, so NNT ≈ 50 — treat about 50 people for the trial's time horizon to prevent one stroke. If the same 50% RRR is 0.2% versus 0.1%, NNT is 1,000. Always ask “50% of what?”
Number needed to harm (NNH) uses absolute risk increase the same way. A drug that cuts strokes with NNT 50 but causes serious bleed with NNH 80 is a different conversation than NNT 50 and NNH 800.
Relative risk and odds ratios that include 1.0 in the confidence interval are compatible with no effect. Do not treat a point estimate of 0.80 as proven benefit if the interval is 0.40–1.60.
P-value versus clinical importance
A p-value is the probability of data as extreme as observed (or more extreme) if the null hypothesis were true. It is not the probability that the null is true. It is not an effect size.
- p < 0.05 (the conventional threshold) means the result is unlikely under the null, given the model's assumptions. It does not mean the drug matters to your patient.
- A tiny p-value on a 0.1 mm Hg blood-pressure change in 40,000 people is statistically significant and clinically trivial.
- A well-done trial that barely misses p = 0.05 on a large, patient-important reduction may still be more useful than a “significant” surrogate.
Read effect size, confidence intervals, and whether the outcome is patient-important (death, hospitalization, stroke, function, quality of life) versus a convenient lab. Surrogate endpoints (a biomarker, a millimeter of plaque) need a chain of evidence before they rewrite primary care.
Intention-to-treat
Intention-to-treat (ITT) analyzes patients in the group to which they were randomized, whether or not they took the drug or finished the protocol. ITT preserves randomization and answers a practical effectiveness question: what happens if I prescribe this in the real world, dropouts included?
Per-protocol / as-treated keeps only people who complied. That can exaggerate benefit and re-introduce confounding (the people who stay on the drug are different). For superiority trials, ITT is the more conservative and usually the preferred primary analysis. For noninferiority trials, both ITT and per-protocol are inspected because ITT can bias toward “no difference.” If the stem asks which analysis answers “does offering this therapy help the average patient I intend to treat?”, choose ITT.
How to read a forest plot conceptually
A forest plot is the picture inside a meta-analysis.
- Each horizontal line is one study. The square (or tick) is that study's effect estimate; square size usually tracks weight (larger studies pull harder).
- The horizontal whiskers are the confidence interval. If a study's line crosses the vertical null line (RR or OR = 1, or mean difference = 0), that individual study is not statistically significant.
- The diamond at the bottom is the pooled estimate. The diamond's width is the pooled confidence interval. If the diamond crosses the null, the combined result is compatible with no effect.
- Notice whether the studies point the same direction (consistent) or scatter (heterogeneity). You do not need an I² formula on exam day, but you should not average apples and engines.
You are not asked to draw the plot. You are asked not to confuse “one big square on the benefit side” with “every patient in my panel will benefit.”
Apply the paper to THIS patient
Appraisal fails if you stop at “the trial was positive.” Ask PICO against the person in the room:
- Population: age, sex, pregnancy, frailty, eGFR, setting (clinic versus ICU versus nursing home). An RCT in adults 40–75 with diabetes and eGFR >30 does not automatically apply to an 89-year-old nursing-home resident with eGFR 18.
- Intervention and comparator: was the comparator today's usual care or a straw-man dose?
- Outcome: did they measure something the patient cares about, over a time horizon that matches remaining life expectancy?
- SDOH: can this person obtain the drug, get the lab, take time off work, or store a refrigerated pen? A brilliant ARR of 2% is theoretical if the copay means the bottle stays at the pharmacy.
Applicability is a Planning decision: start, do not start, start a different option, or refer to the trial's enrollment world (specialty clinic) before you copy the protocol.
Predatory journals and industry bias
Predatory journals solicit manuscripts, charge fees, and skip real peer review. Warning signs: aggressive email, fake impact factors, an editorial board that does not know it is listed, acceptance in days, and a website that cannot name a legitimate publisher. A citation is not a credential. If the only paper supporting a trendy supplement lives in a journal you cannot verify, it does not outrank ADA or USPSTF.
Industry bias is not a reason to throw away every funded RCT — most large cardiovascular trials are industry-funded and still change practice when methods are strong. It is a reason to read past the abstract: surrogate endpoints, noninferiority against a weak comparator, short follow-up, run-in periods that drop intolerant patients, ghostwriting, and selective publication of “positive” trials. Conflicts of interest belong in the appraisal, not in a conspiracy footnote.
Vignette. A detailer quotes “50% fewer hospitalizations, p = 0.03” for a new heart-failure pill. The paper is an RCT, ITT analyzed, but it enrolled adults 45–70 with HFrEF and eGFR >40, excluded SGLT2 inhibitors, and the 50% is 4% versus 2% at 12 months (NNT 50). Your patient is 84, eGFR 22, already on an SGLT2 inhibitor, and cannot afford another brand-name copay. Design is right for a therapy question. Result is statistically significant and absolutely modest. Applicability to this patient is poor. You do not start the drug because a p-value was printed on a mug. You also do not call every industry trial a lie. You match design → result → this person, which is the TCO skill.
When a stem asks “which design best answers whether drug X reduces hospitalizations?”, pick the RCT. When it asks what the 50% means, compute ARR and NNT. When it asks whether to change Mrs. Chen's plan, apply population, comorbidity, and SDOH.
Which study design best answers whether a new oral drug reduces heart-failure hospitalization compared with usual care?
A trial reports a 50% relative reduction in stroke. Event rates are 4% on placebo and 2% on drug. Which interpretation is correct?
A statin RCT enrolled adults 40–75 with diabetes and excluded frailty, eGFR below 30, and nursing-home residence. Your patient is an 89-year-old nursing-home resident with eGFR 18. What is the best appraisal?
Which statement about evidence appraisal is correct?