11.1 Study Designs and the Hierarchy of Evidence
Key Takeaways
- Case-control studies can only report odds ratios because they cannot measure incidence, whereas cohort studies and randomised trials can report relative risk.
- Allocation concealment prevents subversion of randomisation before assignment, while blinding prevents bias after assignment; the two are not interchangeable.
- Case-control designs are the efficient choice for rare outcomes such as oral cancer, and cohort designs for rare exposures.
- An I-squared statistic above about 50% in a meta-analysis indicates substantial heterogeneity and may make pooling inappropriate.
- Funnel plot asymmetry, with small negative studies missing, is the standard visual indicator of publication bias.
Why This Is Examined
"Statistics and research methods" is a named Paper A blueprint topic, and Preparing for Practice outcome 1.1.1 requires registrants to "explain, evaluate and apply the principles of an evidence-based approach to learning, clinical and professional practice and decision making", while outcome 1.1.2 requires them to "critically appraise approaches to dental research". Outcome 1.1.12 adds epidemiology. These are the outcomes behind every question about study design, bias and statistics.
Most candidates lose marks here not because the statistics are hard but because they never revised the topic at all.
The Hierarchy of Evidence
Systematic review and meta-analysis of RCTs
▲
Randomised controlled trial (RCT)
▲
Cohort study
▲
Case-control study
▲
Cross-sectional survey
▲
Case series and case report
▲
Expert opinion, bench research
The hierarchy ranks designs by their resistance to bias and confounding, not by their importance. A well-conducted cohort study answers questions an RCT cannot ethically address — you cannot randomise patients to smoke.
The Main Designs
| Design | Direction | Measure of effect | Strengths | Weaknesses |
|---|---|---|---|---|
| Randomised controlled trial | Prospective, allocation controlled | Relative risk, risk difference | Randomisation balances known and unknown confounders; only design supporting causal inference | Expensive, sometimes unethical, may lack external validity |
| Cohort study | Prospective, exposure to outcome | Relative risk, incidence | Can measure incidence and multiple outcomes; good for rare exposures | Long, expensive, loss to follow-up, confounding |
| Case-control study | Retrospective, outcome to exposure | Odds ratio only | Efficient for rare outcomes such as oral cancer; quick and cheap | Recall bias, selection bias, cannot measure incidence |
| Cross-sectional survey | Snapshot | Prevalence | Ideal for prevalence, for example national dental epidemiology surveys | Cannot establish temporal sequence |
| Case series or report | Descriptive | None | Signals novel phenomena, for example the first MRONJ reports | No comparison group |
Randomisation, Allocation Concealment and Blinding
These three are separate and are frequently confused.
- Randomisation is the process of assigning participants to groups by chance, so that known and unknown confounders are distributed evenly.
- Allocation concealment prevents the recruiting clinician from knowing which group the next participant will enter. Without it, randomisation can be subverted — a clinician may steer sicker patients away from the experimental arm.
- Blinding (masking) prevents participants, clinicians or assessors from knowing the allocation after it has occurred, and protects against performance and detection bias.
Allocation concealment is always possible; blinding sometimes is not. A trial of a surgical technique cannot blind the operator, but it can and should blind the outcome assessor.
Systematic Reviews and Meta-Analysis
A systematic review uses a pre-specified protocol, an explicit search strategy, defined inclusion criteria and formal risk-of-bias assessment. A narrative review does none of these and sits near the bottom of the hierarchy despite often being written by an eminent author.
A meta-analysis statistically pools results from several studies. Its output is usually a forest plot: each study appears as a square sized by its weight with a horizontal confidence interval, and the pooled estimate appears as a diamond. If the diamond crosses the line of no effect, the pooled result is not statistically significant.
Heterogeneity describes variation between studies beyond chance and is quantified by the I² statistic. Conventionally I² above 50% indicates substantial heterogeneity, and pooling may then be inappropriate.
Publication bias — the tendency for positive results to be published and null results not to be — is assessed visually with a funnel plot. Asymmetry, with small negative studies missing from one corner, suggests publication bias.
UK link. The Cochrane Oral Health Group is the principal source of systematic reviews in dentistry, and NICE guidance and the SDCEP and BSP guidelines that dominate Paper B are themselves built from this evidence hierarchy. Understanding how a guideline is constructed is part of understanding why it says what it says.
Matching the Design to the Question
The best study design depends on the question being asked, and examiners test that pairing directly. For a question about the effectiveness of an intervention, a randomised controlled trial sits at the top because randomisation balances known and unknown confounders. For a question about the cause of a rare disease or of a rare outcome, a case-control study is efficient because it starts with the outcome and looks backwards. For a question about incidence and natural history, a cohort study is appropriate. For a question about prevalence — how much disease exists now — a cross-sectional survey is correct, and the national Adult Dental Health and Child Dental Health surveys are the UK examples. For a question about diagnostic accuracy, a cross-sectional study comparing the index test against a reference standard in a consecutive, clinically relevant sample is required.
Where Randomised Trials Fail
An SBA that asks for the "best" design without context is usually a trap, because randomised controlled trials have real limitations. They are expensive and slow, so they are rarely long enough to detect the outcomes dentistry most cares about, such as tooth survival over decades. They recruit selected populations under ideal conditions, so their efficacy results may not reflect effectiveness in ordinary practice. They are unethical when the exposure is harmful — no one will randomise patients to smoke — so causal evidence about tobacco, alcohol and betel quid rests on observational studies. And some interventions cannot be blinded: you cannot blind an operator to whether a tooth has been extracted, which is why outcome assessor blinding and objective endpoints matter so much in surgical and restorative trials.
Reading the Hierarchy Critically
The hierarchy ranks designs by their vulnerability to bias, not by their certainty. A poorly conducted randomised trial with inadequate allocation concealment, high attrition and selective outcome reporting is weaker evidence than a large, well-conducted cohort study. Likewise, a systematic review is only as good as the studies it contains — the GRADE approach exists precisely to rate certainty in the evidence independently of the design label. The examinable message is that design sets a ceiling on credibility; conduct determines whether the study reaches it.
Investigators wish to examine whether long-term use of a particular drug is associated with medication-related osteonecrosis of the jaw, a rare outcome. Which study design is the most efficient starting point, and which measure of effect can it generate?