22.3 Evidence-Based Dentistry & Biostatistics

Key Takeaways

  • Evidence-based dentistry integrates best research evidence, clinical expertise, and patient values/preferences—not papers alone and not anecdote alone.
  • Sensitivity detects disease when present; specificity excludes disease when absent; PPV and NPV depend on prevalence as well as test accuracy.
  • Study designs rank from case reports through case–control and cohort studies to randomized controlled trials and systematic reviews/meta-analyses for interventional questions.
  • Bias (selection, information, confounding, recall, publication) systematically distorts results; random error is addressed with sample size and statistics.
  • P-values and confidence intervals require careful interpretation: statistical significance is not the same as clinical importance or proof of no bias.
Last updated: July 2026

22.3 Evidence-Based Dentistry & Biostatistics

Quick Answer: EBD = best evidence + clinical expertise + patient values. Frame questions with PICO, then match study design to the question. For tests, know sensitivity, specificity, PPV, NPV. For results, distinguish bias from chance, and never confuse a small p-value with automatic clinical importance.

This section is pure AFK reasoning skill: many stems describe a study or a diagnostic test and ask what it measures, what it cannot prove, or which design is strongest. Prevention claims from 22.1–22.2 are ideal material for appraisal practice.

What Evidence-Based Dentistry Is (and Is Not)

ComponentRole
Best available research evidenceSystematic, preferably patient-important outcomes
Clinical expertiseDiagnosis skill, technical ability, judgment of applicability
Patient values and circumstancesPreferences, culture, costs, access, risk tolerance

Not EBD: “I always do it this way,” or “one viral post said,” or “the paper’s p < 0.05 so every patient must receive the product.”

Also not EBD: ignoring patient goals to follow a guideline blindly when comorbidities or values differ—guidelines inform, they do not replace consent.

The 5 A’s cycle (practical)

  1. Ask a focused clinical question.
  2. Acquire evidence (guidelines, systematic reviews, primary studies).
  3. Appraise validity, importance, applicability.
  4. Apply with the patient.
  5. Assess outcomes and improve.

PICO (and PICOT)

LetterMeaningExample
PPatient / Population / ProblemHigh-caries-risk children aged 6–8
IIntervention (or exposure)School sealant program
CComparisonNo sealant / usual care
OOutcomeNew DMFS over 24 months
TTime (optional)24 months

Diagnostic variant: P = patients suspected of disease; I = index test; C = reference standard; O = accuracy metrics.

Hierarchy of Evidence (Interventional Questions)

Higher designs better control bias when well conducted—quality still matters.

Level (typical teaching ladder)DesignStrength for therapy questions
HighestSystematic review / meta-analysis of RCTsSynthesizes multiple trials
HighRandomized controlled trial (RCT)Randomization balances known/unknown confounders
ModerateCohort study (prospective/retrospective)Good for prognosis, harms, when RCT unethical/impractical
LowerCase–control studyEfficient for rare outcomes; recall/selection issues
LowerCross-sectional studyPrevalence and associations at one time—weak for causality
LowestCase series / case reportHypothesis generating
FoundationExpert opinion / pathophysiologyUseful but weakest alone

Guidelines may be high impact clinically but quality varies—check whether recommendations are evidence-linked.

Matching design to question

Question typePreferred evidence
Therapy / prevention effectivenessRCT → systematic review of RCTs
Harm / etiologyCohort; case–control for rare harms
DiagnosisCross-sectional accuracy studies vs reference standard
PrognosisInception cohort
PrevalenceCross-sectional survey

AFK trap: a large case series of implant success without controls “proves” superiority of a brand—it does not.

Diagnostic Test Metrics

2×2 structure

Disease presentDisease absent
Test positiveTrue positive (TP)False positive (FP)
Test negativeFalse negative (FN)True negative (TN)

Definitions and formulas

MetricFormulaPlain languageWhen you care most
SensitivityTP / (TP + FN)Of those with disease, % testing positiveRule out when high (SnNOut): Negative test → less likely disease
SpecificityTN / (TN + FP)Of those without disease, % testing negativeRule in when high (SpPIn): Positive test → more likely disease
PPVTP / (TP + FP)Of positive tests, % truly diseasedDepends on prevalence—higher prevalence → higher PPV
NPVTN / (TN + FN)Of negative tests, % truly non-diseasedHigher when disease is rare (often)
Accuracy(TP+TN)/allOverall correctCan mislead with imbalanced prevalence
Likelihood ratiosCombine sensitivity/specificityHow much a result shifts oddsAdvanced but elegant

Prevalence dependence—critical AFK concept

Sensitivity and specificity are often treated as properties of the test (in a defined population and threshold). PPV and NPV change with prevalence:

ScenarioEffect
Screening low-risk population with imperfect specificityMany false positives → PPV drops
Testing high-risk clinic populationPPV rises for the same test
Very rare diseaseEven specific tests can yield low PPV

Example narrative: A pulp vitality test with good sensitivity still needs clinical context; a positive “caries detection device” in a sealed fissure may be FP noise.

ROC thinking (recognition)

Changing the positivity threshold trades sensitivity vs specificity. ROC curves plot that tradeoff; area under curve summarizes discrimination.

Core Biostatistics for AFK

Types of data

TypeExamplesTypical stats
ContinuousPocket depth mm, DMFS countMeans, t-tests, regression
BinaryCaries yes/noProportions, chi-square, logistic regression
OrdinalPain scale 0–10 categoriesNonparametric or ordinal models
SurvivalTime to implant failureKaplan–Meier, Cox models

Central tendency and spread

MeasureUse
MeanSymmetric continuous data
MedianSkewed data (income, some cost data)
SDSpread around mean
IQRSpread around median
95% CIInterval likely to contain the true parameter with 95% confidence under model assumptions

Hypothesis testing language

ConceptMeaning
Null hypothesis (H₀)Typically “no difference / no association”
p-valueProbability of data as extreme as observed if H₀ were true—not the probability H₀ is true
α (alpha)Pre-set type I error rate (often 0.05)
Type I errorFalse positive: reject true H₀
Type II errorFalse negative: fail to reject false H₀
Power1 − β; probability of detecting a true effect
Clinical vs statistical significanceTiny irrelevant difference can be “significant” in huge samples; important effects can be nonsignificant if underpowered

Confidence interval pearl: a mean difference of 0.1 mm CAL with 95% CI 0.05–0.15 may be statistically significant but clinically trivial; a large effect with wide CI crossing the null is uncertain.

Relative vs absolute measures

MeasureIdeaTrap
Relative risk (RR)Risk_intervention / Risk_control“50% reduction” can hide small absolute change
Odds ratio (OR)Odds_exposed / Odds_unexposed≈ RR when outcomes rare; otherwise not identical
Absolute risk reduction (ARR)Risk_control − Risk_interventionMore transparent for patients
NNT1 / ARRPatients needed to treat for one extra good outcome
NNHFor harmsNumber for one extra harm

Prevention example: If sealants reduce 2-year occlusal caries from 10% to 5%, RR = 0.5 (50% relative reduction), ARR = 5%, NNT = 20.

Bias and Validity

Internal vs external validity

ValidityQuestion
InternalAre the study’s conclusions correct for the participants studied? (bias/confounding minimized)
External (generalizability)Do results apply to my patient / my setting?

Major bias types

BiasWhat goes wrongDental example
Selection biasWho enters or remains in the study is non-comparableClinic volunteers for a “new whitening gel” are healthier and more compliant
Information / measurement biasExposures or outcomes measured differently across groupsExaminer knows which group got sealant and scores caries more leniently
Recall biasCases remember exposures differently than controlsOral cancer cases over-report alcohol vs controls
Interviewer biasQuestioning differs by case status
Lead-time biasEarly detection looks like longer survival without changing courseScreening “improves survival” only by diagnosing earlier
Length-time biasScreening finds slowly progressive disease preferentiallyOverestimates benefit of screening
Publication biasPositive studies publish more readilyMeta-analyses overstate effects
ConfoundingThird factor associated with exposure and outcome distorts effectCoffee “causes” oral cancer if smoking not controlled
Attrition biasDifferential loss to follow-upFailures drop out of implant follow-up
Spectrum biasDiagnostic study uses extreme sick vs healthy onlyOverstates real-world accuracy

Controlling confounding

MethodNotes
RandomizationBalances confounders in RCTs
RestrictionStudy only non-smokers
MatchingPair on age/sex
Stratification / multivariable analysisStatistical adjustment—residual confounding possible

Appraising a Prevention Claim (Worked Pattern)

Claim: “Product X toothpaste cuts cavities 40%.”

  1. PICO: Who? vs what? measured how?
  2. Design: RCT of children? industry-funded?
  3. Outcomes: DMFS increments? patient-important?
  4. Absolute numbers: 40% of what baseline? NNT?
  5. Bias: blinding of examiners? attrition?
  6. Applicability: fluoride background, diet, Canadian setting?
  7. Harms/cost: fluorosis, price, adherence?
  8. Patient values: taste, cost, preferences.

Ethics of Evidence Use

  • Do not overstate certainty to sell procedures.
  • Disclose material alternatives (including monitoring/prevention).
  • Research participation needs consent and REB oversight (jurisprudence crossover).
  • Industry relationships require transparent appraisal.

Rapid review list

  • EBD triad: evidence + expertise + patient values
  • PICO structures searchable questions
  • Therapy: prefer well-done RCTs and systematic reviews
  • Sensitivity/specificity vs PPV/NPV (prevalence matters for PPV/NPV)
  • SnNOut / SpPIn mnemonics
  • p-value ≠ probability the null is true; ≠ clinical importance
  • ARR and NNT communicate absolute benefit
  • Selection, information, confounding, recall, publication, lead-time, length-time biases
  • Internal validity first, then external applicability

Mastery of Chapters 22.1–22.3 means you can deliver fluoride and sealants correctly, speak the language of population oral health and DMFT, and critically appraise the evidence behind any prevention or treatment claim—the full prevention and EBD slice of the AFK blueprint.

Test Your Knowledge

Evidence-based dentistry is best defined as the integration of:

A
B
C
D
Test Your Knowledge

A new caries-detection device is evaluated: of 100 teeth truly with dentin caries, 90 test positive; of 100 truly sound teeth, 80 test negative. Sensitivity and specificity are:

A
B
C
D
Test Your Knowledge

Which study design generally provides the strongest evidence that a preventive intervention causes a reduction in new caries, assuming high methodological quality?

A
B
C
D
Test Your Knowledge

Positive predictive value (PPV) of a diagnostic test is most accurately described as:

A
B
C
D