16.4 Evidence-Based Dentistry & Research Methods
Key Takeaways
- Evidence-based dentistry integrates the best research evidence with clinical expertise and patient values; the hierarchy of evidence ranks systematic reviews/meta-analyses and randomised controlled trials above observational and case studies.
- The PICO framework structures a clinical question: Population, Intervention, Comparison, Outcome — the basis for a focused literature search.
- Study designs include RCTs (gold standard for therapy), cohort studies (risk/prognosis), case-control studies (rare diseases/exposures), cross-sectional (prevalence), and case series; each is suited to different questions and has characteristic biases.
- Key validity concepts: randomisation and allocation concealment, blinding, intention-to-treat analysis, and the distinction between statistical significance (p-value) and clinical importance (effect size).
- Critical appraisal tools (CASP, CONSORT, PRISMA, STROBE) provide checklists to assess the quality and applicability of evidence.
What Is Evidence-Based Dentistry?
Evidence-based dentistry (EBD) is the integration of:
- The best research evidence (from well-designed studies).
- Clinical expertise (the clinician's knowledge and judgement).
- Patient values and circumstances (their preferences and situation).
EBD does not replace clinical judgement — it informs it. The dentist asks a focused question, finds the best evidence, appraises it, and applies it to the individual patient.
The Hierarchy of Evidence
Evidence is ranked by its susceptibility to bias. A common pyramid:
| Level | Study type | Strength |
|---|---|---|
| 1 | Systematic review / meta-analysis of RCTs | Strongest |
| 2 | Randomised controlled trial (RCT) | Strong |
| 3 | Cohort study (prospective) | Moderate |
| 4 | Case-control study | Moderate-lower |
| 5 | Cross-sectional study | Lower |
| 6 | Case series / case report | Weakest |
| 7 | Expert opinion / bench research | Lowest |
A systematic review with meta-analysis sits at the top because it synthesises multiple studies, reducing the bias of any single trial. The pyramid is a guide — a well-conducted cohort may answer some questions better than a flawed RCT, and the design must match the question.
The PICO Framework
A focused clinical question uses PICO:
- Population/patient — who is the patient (age, condition, setting)?
- Intervention — what treatment/exposure is being considered?
- Comparison — what is it being compared with (placebo, another treatment, no treatment)?
- Outcome — what matters to the patient (cure, pain relief, adverse effects, survival)?
A fifth element, Timing and Study type (PICOTS), is sometimes added. PICO structures both the question and the literature search.
Example: In adults with chronic periodontitis (P), does full-mouth disinfection (I) compared with quadrant-wise scaling (C) reduce probing depth (O) at 6 months (T)?
Study Designs and Their Uses
| Design | Best for | Key feature / bias |
|---|---|---|
| RCT | Therapy/intervention efficacy | Random allocation, blinding; gold standard but costly, ethical limits |
| Cohort (prospective) | Prognosis, incidence, risk factors | Follows exposed vs unexposed groups forward in time; recall bias avoided, but confounding |
| Case-control | Rare diseases or exposures | Compares cases vs matched controls retrospectively; recall and selection bias; suited to rare disease |
| Cross-sectional | Prevalence | Snapshot; cannot infer causation |
| Case series/report | Rare presentations | Descriptive only; no control group |
| Systematic review | Synthesising evidence | Pre-defined protocol, comprehensive search; risk of publication bias if not well done |
| Meta-analysis | Combining data | Quantitatively pools results; needs homogeneity (assessed by I² statistic) |
Choosing the Right Design
- Therapy question → RCT.
- Harm/risk question → cohort or case-control.
- Prognosis question → cohort.
- Diagnosis question → cross-sectional with a reference standard.
- Prevalence → cross-sectional.
Threats to Validity
Bias
- Selection bias — systematic differences between groups (reduced by randomisation and allocation concealment).
- Information/measurement bias — errors in measuring exposure or outcome (reduced by blinding and standardised tools).
- Recall bias — differential recall between cases and controls.
- Performance bias — differences in care apart from the intervention (reduced by blinding).
- Attrition bias — differential loss to follow-up (assess by intention-to-treat analysis).
- Publication bias — positive studies more likely published (assessed by funnel plot in reviews).
Confounding
A confounder is a third variable linked to both exposure and outcome that distorts the apparent association (e.g. smoking confounds the link between coffee and oral cancer). Controlled by randomisation (in RCTs), matching, stratification, and multivariable adjustment.
Statistical Significance and Clinical Importance
The p-value
The p-value is the probability of observing the data (or more extreme) if the null hypothesis were true. A conventional threshold p < 0.05 means a less than 5% probability of such a result by chance — but it does not measure effect size, clinical importance, or the probability that the null is true.
Confidence Intervals
A 95% confidence interval (CI) is the range within which the true effect is estimated to lie with 95% certainty. A CI that excludes the null value (0 for differences, 1 for ratios) indicates statistical significance; a narrow CI indicates precision; a wide CI indicates uncertainty.
Clinical vs Statistical Significance
A large study can find a statistically significant result of trivial clinical size. Always consider the effect size (e.g. mean difference, relative risk, number needed to treat) and the CI — not just the p-value.
Effect Size Measures
| Measure | Use |
|---|---|
| Relative risk (risk ratio) | Ratio of risk in intervention vs control |
| Odds ratio | Used in case-control studies |
| Number needed to treat (NNT) | Patients to treat to prevent one outcome — intuitive and clinical |
| Absolute risk reduction (ARR) | Difference in risk between groups |
NNT = 1 / ARR. A smaller NNT is more effective.
Intention-to-Treat Analysis
In an RCT, intention-to-treat (ITT) analysis includes every participant in the group to which they were randomised, regardless of whether they completed or received the assigned treatment. ITT preserves the benefit of randomisation and reflects 'real-world' effectiveness. Per-protocol analysis (only those who completed the protocol) can introduce bias and overestimate efficacy.
Appraisal Tools
| Tool | Use |
|---|---|
| CASP (Critical Appraisal Skills Programme) | Generic checklists for RCTs, cohort, case-control, qualitative |
| CONSORT | Reporting RCTs |
| PRISMA | Reporting systematic reviews and meta-analyses |
| STROBE | Reporting observational studies (cohort, case-control, cross-sectional) |
| GRADE | Grading the certainty of evidence (high/moderate/low/very low) |
Levels of Certainty (GRADE)
GRADE classifies the certainty (quality) of a body of evidence as high, moderate, low, or very low — based on study design and factors that downgrade (risk of bias, inconsistency, indirectness, imprecision, publication bias) or upgrade (large effect, dose-response) the evidence. This moves beyond a single 'level' to a nuanced judgement of how much trust to place in a finding.
Applying Evidence
Find the best evidence → appraise it for validity and applicability → integrate with clinical judgement and the patient's values → monitor the outcome. The process is iterative: the question may need refining, and the evidence should be revisited as new studies appear.
Which study design sits at the top of the hierarchy of evidence for evaluating the efficacy of a new therapeutic intervention?
A dentist asks: 'In adult patients with chronic periodontitis, does full-mouth disinfection compared with quadrant-wise scaling reduce probing depth at six months?' Which element of the PICO framework is 'quadrant-wise scaling'?
A trial reports a p-value of 0.04 for a reduction in bleeding on probing, with a 95% confidence interval of 0.1 to 0.3 mm. Which interpretation is most appropriate?
In an RCT comparing a new sealant with standard care, 4% of children in the control group and 1% in the sealant group developed caries over two years. What is the number needed to treat (NNT)?
Intention-to-treat (ITT) analysis in an RCT preserves the benefit of randomisation. Which statement best explains why ITT is preferred to per-protocol analysis?