11.3 Diagnostic Test Performance

Key Takeaways

  • Sensitivity is the proportion of diseased people correctly identified, and a negative result from a highly sensitive test rules the condition out.
  • Specificity is the proportion of disease-free people correctly excluded, and a positive result from a highly specific test rules the condition in.
  • Predictive values depend on disease prevalence, so the same test has a much lower positive predictive value in a low-prevalence population.
  • Likelihood ratios are independent of prevalence, with a positive ratio above 10 or a negative ratio below 0.1 being clinically decisive.
  • Cohen's kappa measures inter-examiner agreement corrected for chance, with values above 0.8 indicating very good agreement.
Last updated: September 2026

The 2 x 2 Table

Every diagnostic accuracy question reduces to the same table, comparing a test against a reference standard.

Disease presentDisease absent
Test positiveTrue positive (TP)False positive (FP)
Test negativeFalse negative (FN)True negative (TN)

The Four Core Measures

Sensitivity=TPTP+FN\text{Sensitivity} = \frac{TP}{TP + FN}

Sensitivity is the proportion of people with the disease whom the test correctly identifies. A highly sensitive test has few false negatives, so a negative result rules the condition out (SnNOut).

Specificity=TNTN+FP\text{Specificity} = \frac{TN}{TN + FP}

Specificity is the proportion of people without the disease whom the test correctly excludes. A highly specific test has few false positives, so a positive result rules the condition in (SpPIn).

Positive Predictive Value (PPV)=TPTP+FPNegative Predictive Value (NPV)=TNTN+FN\text{Positive Predictive Value (PPV)} = \frac{TP}{TP + FP} \qquad \text{Negative Predictive Value (NPV)} = \frac{TN}{TN + FN}

Predictive values answer the question the patient actually asks: given this result, what is the chance I have the disease? Unlike sensitivity and specificity, predictive values depend on prevalence. The same test applied in a high-risk oral cancer clinic and in a low-risk general practice population has identical sensitivity and specificity but a very different positive predictive value.

A Worked Dental Example

A new caries detection device is evaluated against histological sectioning on 1,000 extracted teeth, of which 200 truly have dentinal caries.

Caries present (200)Caries absent (800)
Device positive18080
Device negative20720
  • Sensitivity = 180 / (180 + 20) = 90%
  • Specificity = 720 / (720 + 80) = 90%
  • PPV = 180 / (180 + 80) = 69%
  • NPV = 720 / (720 + 20) = 97%

Now apply the same device in a low-caries adult population where true prevalence is 2%. Sensitivity and specificity are unchanged, but out of 1,000 patients there are 20 with caries and 980 without. The device finds 18 true positives and 98 false positives, so PPV collapses to about 16%. Five out of six positive results would be false. This is the single most important concept in diagnostic testing and it is why screening a low-prevalence population with an imperfect test causes harm.

Likelihood Ratios and ROC Curves

Positive Likelihood Ratio=Sensitivity1Specificity\text{Positive Likelihood Ratio} = \frac{\text{Sensitivity}}{1 - \text{Specificity}}

A positive likelihood ratio above 10 substantially increases the probability of disease; a negative likelihood ratio below 0.1 substantially decreases it. Likelihood ratios are useful because, unlike predictive values, they are independent of prevalence.

A receiver operating characteristic (ROC) curve plots sensitivity against 1 − specificity across all possible cut-off values. The area under the curve (AUC) summarises overall accuracy: 0.5 is no better than chance and 1.0 is perfect. Moving the cut-off always trades sensitivity against specificity; there is no threshold that improves both.

Reliability Versus Validity

  • Validity (accuracy) is whether a test measures what it claims to measure.
  • Reliability (precision) is whether it gives the same answer on repetition.

A test can be highly reliable and completely invalid — a consistently miscalibrated probe reads the same wrong number every time.

Cohen's kappa measures agreement between two observers, or between the same observer on two occasions, corrected for agreement expected by chance. Values above 0.8 indicate very good agreement, 0.6 to 0.8 good, 0.4 to 0.6 moderate, and below 0.4 poor. Kappa is the statistic quoted whenever examiners are calibrated for indices such as DMFT, BPE or IOTN.

Exam link. Radiographic caries diagnosis is highly specific but only moderately sensitive: approximately 30% to 40% of mineral must be lost before a lesion becomes visible. A positive bitewing therefore rules caries in, but a negative bitewing does not rule it out — which is exactly why visual examination and bitewings are used together rather than either alone.

Prevalence, Predictive Values and the Screening Problem

The most examinable property of predictive values is that they depend on prevalence, while sensitivity and specificity do not. Take a test with 90 per cent sensitivity and 90 per cent specificity. In a high-risk clinic where 30 per cent of patients have the disease, the positive predictive value is around 79 per cent. In a general population where 1 per cent have the disease, the same test gives a positive predictive value of only about 8 per cent — more than nine in ten positives are false. This single arithmetic fact explains why population screening for rare conditions generates so much unnecessary anxiety and investigation, and why opportunistic case-finding in a high-risk group is more efficient than untargeted screening.

The dental application is oral cancer screening. Visual examination has reasonable sensitivity but, applied to an unselected population with a low prevalence of malignancy, produces many false positives. This is why UK practice is opportunistic examination of every patient combined with targeted vigilance in tobacco and alcohol users, rather than a formal national screening programme, and why adjunctive devices such as toluidine blue and tissue fluorescence have not been adopted — their specificity is too low to improve on careful clinical examination.

Test Your Knowledge

A caries detection device has a sensitivity of 90% and a specificity of 90%. It is validated in a high-risk population where caries prevalence is 20%, giving a positive predictive value of 69%. The device is then marketed for use in a low-risk adult population where true prevalence is 2%. What happens to its performance characteristics?

A
B
C
D