5.2 Physical Examination Framework & Diagnostic Accuracy
Key Takeaways
- HOPS is History, Observation, Palpation, Special tests, with the comparable uninjured side used as the baseline after life/limb and neurovascular threats are cleared
- High sensitivity supports ruling a condition out when the test is negative (SnNOut); high specificity supports ruling a condition in when the test is positive (SpPIn)
- Likelihood ratios move pretest to post-test probability; positive and negative predictive values change with prevalence, so a good test in a high-risk locker room is not equally predictive in a low-risk clinic
- Clusters outperform single provocative tests: Ottawa Ankle Rules (Bachmann 2003 pooled sensitivity 97.6%, median specificity 31.5%) screen for fracture; Lachman (Benjaminse 2006 pooled sensitivity 85%, specificity 94%) outperforms anterior drawer for ACL
- A test with 70% sensitivity cannot rule pathology out when negative; generate a differential (diagnostic timeout) and then choose tests that discriminate, using AROM then PROM then resisted testing and 0–5 MMT grades
5.2 Physical Examination Framework & Diagnostic Accuracy
PA8 task 0202 is to perform a physical examination using appropriate diagnostic techniques. The BOC does not want a random walk through every eponym you memorized. It wants a repeatable order, a comparable baseline, and enough diagnostic-accuracy literacy to know what a positive or negative test actually moved.
Region-specific tests (Hawkins-Kennedy, Thessaly, talar tilt, and the rest) are Chapters 6–8. This section is the chassis those tests bolt onto.
HOPS and the Order That Protects Patients
HOPS is History, Observation, Palpation, Special tests. Some texts say HIPS (inspection) or fold the same content into SOAP. The letters matter less than the discipline:
- Rule out life and limb threats. Unstable airway, uncontrolled bleeding, suspected cervical-spine injury, pulseless or cold extremity, open fracture, and the red flags from section 5.1 are not "special tests later."
- Neurovascular screen of the involved region. Distal pulses, capillary refill, sensation, and key myotomes. Document them before you aggressively palpate or stress a joint, and again after you do.
- Observation of the involved and uninvolved sides: deformity, effusion, ecchymosis, atrophy, attitude of the limb, skin, and how the athlete undresses or walks.
- Palpation of bony landmarks then soft tissue: temperature, point tenderness, defect, crepitus, pulses, lymph nodes when infection is in the differential.
- Range of motion and strength, then selected special tests, then functional tests if the tissue can tolerate them.
Comparable uninjured side first when the athlete can tolerate it. Starkey and Magee both use the uninvolved limb to set the patient's normal range, end-feel, laxity, and strength. Exceptions: do not delay a primary survey to "get a baseline hamstring grade," and do not grind through a complete contralateral exam while an obvious dislocation waits. In clinic, the uninjured side is how you stop calling physiologic recurvatum an ACL tear.
| Step | Purpose | Common error |
|---|---|---|
| History (0201) | Build the differential and pretest probability | Skipping MOI |
| Observation | Deformity, swelling, gait, willingness | Starting with thumbs on the joint line |
| Neurovascular | Limb threat and nerve-root screen | Checking pulses only after a long special-test list |
| Palpation | Localize tissue, temperature, defect | Palpating so hard you create a new positive |
| AROM → PROM → RROM | Contractile vs inert; irritability | Forcing PROM before the athlete shows you what they will do |
| Special tests | Discriminate among remaining hypotheses | Running a 12-test battery "so the chart looks complete" |
| Functional tests | Sport and life tasks | Hop-testing an acutely unstable knee |
Diagnostic Accuracy: Sensitivity, Specificity, and the Two Mnemonics
Every special test is a filter. Filters leak.
Sensitivity (Sn) is the true-positive rate: of all people who have the pathology, what fraction test positive? Sn = TP / (TP + FN). A highly sensitive test has few false negatives. SnNOut: a Sensitive test, when Negative, helps rule Out disease. Screening tools and decision rules that must not miss a fracture or a cord compression need high sensitivity.
Specificity (Sp) is the true-negative rate: of all people who do not have the pathology, what fraction test negative? Sp = TN / (TN + FP). A highly specific test has few false positives. SpPIn: a Specific test, when Positive, helps rule In disease. Confirmatory tests (classically a pivot-shift that is clearly positive in a relaxed athlete) live here.
A test can be sensitive but not specific (Ottawa Ankle Rules: almost nobody with a fracture is missed, but many without a fracture still "fail" the rule and get a radiograph). A test can be specific but not sensitive (a clearly positive pivot shift is compelling for ACL deficiency; a negative pivot shift misses many tears because guarding and swelling kill the maneuver).
Likelihood Ratios and Why Prevalence Matters
Sensitivity and specificity are properties of the test in a studied population. Clinicians need a property of this athlete in front of you.
- Positive likelihood ratio (LR+) = Sn / (1 − Sp). How much a positive test raises the odds of disease.
- Negative likelihood ratio (LR−) = (1 − Sn) / Sp. How much a negative test lowers the odds of disease.
A conventional interpretation (Jaeschke, Guyatt, and Sackett) used throughout orthopedic physical diagnosis:
| LR+ | LR− | Approximate shift |
|---|---|---|
| >10 | <0.1 | Large, often conclusive |
| 5–10 | 0.1–0.2 | Moderate |
| 2–5 | 0.2–0.5 | Small but sometimes useful |
| 1–2 | 0.5–1 | Slight; rarely changes a decision |
Positive predictive value (PPV) is the chance the athlete actually has the disease if the test is positive. Negative predictive value (NPV) is the chance they are disease-free if the test is negative. Both move with prevalence (pretest probability). In a high-school training room the day after a non-contact pop and hemarthrosis, ACL prevalence is high, so a positive Lachman is very believable (high PPV). If you performed Lachman tests on every sore knee at a middle-school health fair, even a good test would generate more false positives and a lower PPV. NPV moves the other way: a negative test is more reassuring when the condition is rare, and less reassuring when the story already screams ACL.
This is why you never quote a single percentage as if it were magic. Pretest probability comes from history. The test only nudges it.
Verified Numbers You Can Defend (and What They Teach)
Do not memorize a zoo of unstable point estimates. Memorize a few replicated examples and the concept they illustrate.
Ottawa Ankle Rules as a screen (high Sn, modest Sp). Bachmann, Kolb, Koller, Steurer, and ter Riet (BMJ 2003) pooled 27 studies (15,581 patients). Overall sensitivity was 97.6% (95% CI 96.4–98.9) with median specificity 31.5% (IQR 23.8–44.4). Pooled LR− was 0.08 for the ankle assessment. Applied to a typical ~15% fracture prevalence, a negative rule left a less than about 1.4% chance of fracture in the analyzed subgroups. A 2017 British Journal of Sports Medicine meta-analysis (66 studies) found pooled ankle-rule sensitivity 99.4% (97.9–99.8) and specificity 35.3% (28.8–42.3). Use: a negative rule helps rule out a radiograph-worthy fracture. A positive rule does not prove a fracture; it only says imaging is indicated. Full decision-rule teaching continues in Chapter 8; here the lesson is SnNOut versus a low-specificity confirm.
Lachman versus anterior drawer for ACL. Benjaminse, Gokeler, and van der Schans (J Orthop Sports Phys Ther / meta-analysis 2006, widely cited in Magee and sports-medicine reviews) reported pooled Lachman sensitivity 85% (83–87) and specificity 94% (92–95). Anterior drawer is consistently less sensitive, especially in the acute, guarded knee, because hamstrings fire at 90° of flexion and a tense hemarthrosis blocks the position. A 2022 Medicine meta-analysis (18 studies, 2,031 participants) pooled Lachman sensitivity 76%, specificity 89%, LR+ 5.65, LR− 0.28, versus anterior drawer sensitivity 64%, specificity 87%, LR+ 3.57, LR− 0.44. Pivot shift is typically specific (a positive test is meaningful) and insensitive (a negative test is not a clearance). McGee's Evidence-Based Physical Diagnosis synthesis similarly treats a positive Lachman as a large LR+ and a negative anterior drawer as a weak rule-out.
The 70% sensitivity trap. If Sn = 70%, then 30 of every 100 athletes who truly have the pathology will test negative. You cannot SnNOut with that test. McMurray's meniscus test is the classic athletic-training example: later meta-analyses often land near the 60–70% sensitivity / similar specificity neighborhood, with wide study-to-study scatter. A negative McMurray never "clears the meniscus." Combine history (locking, twisting MOI, delayed effusion), joint-line tenderness, and other tests, and still image or refer when the story is strong.
Clusters Beat Hero Tests
One provocative maneuver almost never "proves" a diagnosis. Clusters raise specificity (and LR+) as you require more positives, and they raise sensitivity (and lower LR−) as you accept any one positive as a screen.
Wainner cluster for cervical radiculopathy (Spine 2003): Spurling, cervical distraction, ipsilateral rotation less than 60°, and upper-limb neurodynamic test A (median). Four of four positives: specificity about 99%, LR+ about 30. Three of four: specificity about 94%, LR+ about 6. Upper-limb tension test A alone is the sensitive screen (historically ~97% in that derivation); a negative ULNT-A argues against radiculopathy more than a negative Spurling does.
Cook myelopathy cluster (J Man Manip Ther 2010): gait deviation, Hoffmann sign, inverted supinator sign, Babinski sign, age greater than 45. One of five positives: sensitivity 0.94, LR− 0.18 (useful screen). Three of five: specificity 0.99, LR+ 30.9 (useful confirm in that sample). Section 5.3 uses this clinically; here it is the poster child for clustering.
ACL in practice: pop + immediate effusion + giving way already creates a high pretest probability. Lachman then pivot shift (when swelling allows) is a cluster, not a religion around a single drawer test.
Reliability: Can Two ATs Agree?
Intra-rater reliability is the same examiner repeating the test. Inter-rater reliability is two examiners. Continuous measures (goniometry, girth) use intraclass correlation coefficients; yes/no tests use kappa. A test with beautiful published sensitivity is still junk in your hands if you cannot reproduce the starting position. Guarding, effusion, athlete muscle size versus examiner hand size, and inadequate practice all wreck Lachman and anterior-drawer agreement. If the finding will change clearance, have a second skilled clinician confirm it or use an instrumented device when available.
Diagnostic Timeout
Before your hands become a special-test slot machine, pause:
- Name three to five competing hypotheses (example: ACL tear vs patellar dislocation vs MCL vs meniscus vs bone bruise / occult fracture).
- Choose tests that split those hypotheses, not tests that are positive in all of them.
- Stop when further testing will not change management (already referring for MRI and ortho, athlete too irritable, or red flag already tripped).
This is the opposite of "I always do the full knee battery."
Palpation, ROM, MMT, Girth, Function
Palpation localizes after observation. Compare temperature (infection, complex regional pain, inflammatory arthropathy), effusion versus extra-articular swelling, and a palpable defect (Achilles, pectoralis major, rectus femoris). Palpate the bony Ottawa points when an ankle injury might need imaging, but do not pretend palpation of the ATFL is a highly specific ligament test; ATFL palpation is sensitive and not specific in lateral-ankle literature.
Range of motion follows Cyriax logic still taught in athletic training programs:
- Active ROM (AROM) first: willingness, painful arc, substitution, and what the athlete will actually use in sport.
- Passive ROM (PROM) next: end-feel (capsular, bony, springy block, spasm, empty, soft-tissue approximation) and whether PROM exceeds AROM (contractile inhibition vs inert restraint).
- Resisted ROM / break tests last in the mobility sequence, then formal manual muscle testing (MMT).
If AROM is full and painless, aggressive PROM may add little. If AROM is limited and PROM is full and painless, think weakness, inhibition, or poor effort rather than a locked joint.
MMT grades (Kendall / Daniels and Worthingham scale) — BOC items expect the numbers:
| Grade | Name | Definition |
|---|---|---|
| 0 | Zero | No palpable contraction |
| 1 | Trace | Palpable contraction, no joint motion |
| 2 | Poor | Full ROM with gravity eliminated |
| 3 | Fair | Full ROM against gravity |
| 4 | Good | Full ROM against gravity plus moderate resistance |
| 5 | Normal | Full ROM against gravity plus maximal resistance |
Plus/minus modifiers (3+, 4−) appear in clinics; if the item stem uses whole numbers, stay on 0–5. Pain-limited weakness is not the same as neurologic weakness; note which you observed. Hold isometric tests about 5 seconds so fatigable myotomal weakness can appear.
Girth is a tape measure with a standardized landmark (for example 5 cm and 10 cm proximal to the medial joint line, or mid-patella for effusion). Compare sides. Rapid increase is swelling; chronic decrease is atrophy (vastus medialis wasting after knee injury is a functional finding, not a cosmetic one).
Functional tests (squat, step-down, hop battery, sport-specific cutting) come when tissue healing and irritability allow. They are not a substitute for a neurovascular exam, and they are not day-of-injury entertainment.
Worked Example: The 70% Sensitivity Trap
An athlete has joint-line pain after a twist, occasional locking, and a small delayed effusion. McMurray is negative. The intern says, "Meniscus is ruled out; sensitivity is about 70%, so we are probably fine." That sentence is the error. Seventy percent sensitivity means false negatives are common. Specificity near 70% also means false positives are common. You do not clear, and you do not rush to arthroscopy, on that one maneuver. You keep the meniscus on the differential, add joint-line palpation and a loaded rotation test the athlete can tolerate, treat the irritability, and refer for imaging if mechanical symptoms persist — which is diagnosis (0203) built on a literate 0202, not on a single eponym.
Trap: treating a special test with 70% sensitivity as if a negative result rules the pathology out. It does not. High sensitivity rules out; modest sensitivity only slightly lowers probability, and history still owns the pretest.
A provocative special test has a reported sensitivity of 70% and specificity of 92%. The test is negative. What is the best interpretation?
After a focused history in an athlete with a subacute, non-emergent knee injury, which examination sequence is most appropriate?
Bachmann and colleagues' systematic review found the Ottawa Ankle Rules have pooled sensitivity of about 97.6% and median specificity of about 31.5%. How should a negative result be used?