6.2 Word Recognition Testing (WRS), PB-Max & Rollover
Key Takeaways
- Word Recognition Testing (WRS) evaluates suprathreshold speech intelligibility and phonemic clarity, expressed as the percentage of phonetically balanced (PB) monosyllabic words correctly identified at conversational or sensation levels.
- Standardized word lists include CID W-22 and Northwestern University NU-6 (CNC words), administered using a carrier phrase ('Say the word...') and standardized recorded speech to eliminate clinician accent, pacing, and vocal intensity variability.
- The Performance-Intensity (PI) function identifies PB-Max, the maximum achievable word recognition score; in normal and conductive ears, PB-Max reaches 100% and plateaus, whereas cochlear pathology reduces PB-Max and shifts it to higher sensation levels.
- The Rollover Phenomenon is a definitive marker of retrocochlear pathology (vestibular schwannoma); a Rollover Index RI = (PB_max - PB_min) / PB_max ≥ 0.45 (or ≥ 0.40 on NU-6) indicates neural conduction failure and mandates urgent medical referral.
- Clinical masking for WRS is mandatory whenever the presentation level in the test ear exceeds the best bone conduction threshold in the non-test ear by the interaural attenuation value (Level_TE - BC_NTE ≥ IA), which occurs frequently due to high suprathreshold test levels.
6.2 Word Recognition Testing (WRS), PB-Max & Rollover
[!IMPORTANT] Pure-tone thresholds and speech reception thresholds (SRT) measure auditory sensitivity—the minimum sound level a listener can detect. In contrast, Word Recognition Testing (WRS) measures auditory clarity and processing fidelity—the ability of the cochlea, auditory nerve, and central pathways to resolve and decode acoustic phonemes at comfortable, audible suprathreshold intensities. Two patients may share an identical 50 dB HL pure-tone loss, yet one achieves a 96% word recognition score (thriving with basic amplification), while the other achieves 44% (suffering from severe phonemic regression and distortion). Understanding WRS, PB-Max, and the retrocochlear rollover phenomenon is a core competency on the NBC-HIS examination.
Word Recognition Score (WRS): Definition and Materials
Clinical Definition
The Word Recognition Score (WRS)—historically termed Speech Discrimination Score (SDS)—is defined as the percentage of standardized monosyllabic speech tokens correctly identified, repeated, or transcribed by a listener when presented at a suprathreshold listening level.
Unlike threshold tests where results are recorded in decibels (dB HL), the WRS is recorded as a percentage of correct responses (0% to 100%).
Phonetically Balanced (PB) Monosyllabic Word Lists
To ensure scientific standardization and cross-test validity, test materials consist of phonetically balanced (PB) monosyllabic words. Phonetic balance requires that the relative frequency of individual speech sounds (vowels, consonants, diphthongs, and consonant clusters) within each word list directly mirrors their statistical occurrence in spoken American English.
The two standard validated word corpora used in modern clinical practice are:
- CID Auditory Test W-22 (Central Institute for the Deaf):
- Developed by Hirsh et al. (1952).
- Consists of four lists of 50 monosyllabic words (Lists 1, 2, 3, and 4), each scrambled into four randomized presentations (A, B, C, D).
- Words were selected from the Thorndike-Lorge lexical frequency tables to ensure high familiarity across adult English speakers (e.g., yard, wire, camp, bell, shoe, smart, give, there).
- Northwestern University Auditory Test No. 6 (NU-6):
- Developed by Tillman and Carhart (1966).
- Consists of four lists of 50 monosyllabic words (Lists 1, 2, 3, and 4), each available in four scrambled randomizations.
- Formatted strictly around a Consonant-Nucleus-Consonant (CNC) phonemic architecture (initial consonant, vowel nucleus, final consonant), providing rigorous acoustic-phonetic control (e.g., bar, kiss, pool, tape, thumb, void, knock, white).
PB LIST PHONEMIC ARCHITECTURE (NU-6 CNC Word)
Initial Consonant Vowel Nucleus Final Consonant
(C) (N) (C)
┌──────────────┐ ┌──────────────┐ ┌──────────────┐
│ /t/ │ + │ /eɪ/ │ + │ /p/ │ = "TAPE"
└──────────────┘ └──────────────┘ └──────────────┘
Word List Length: 50 Words vs. 25 Words (Binomial Variability)
While time-pressured clinicians frequently administer abbreviated half-lists (25 words), psychometric science demonstrates that word recognition follows a binomial distribution (Thornton & Raffin, 1978):
- In a 50-word list, each word accounts for 2% of the score. The 95% critical difference confidence interval around a score of 80% spans from approximately 68% to 90% (a 22% spread).
- In a 25-word list, each word accounts for 4% of the score. The 95% critical difference confidence interval around an 80% score widens dramatically to 60% to 94% (a 34% spread).
[!CAUTION] For routine adult screening where the first 10 to 12 consecutive words are identified with 100% accuracy, stopping at 25 words is clinically defensible. However, whenever retrocochlear pathology is suspected, when medical-legal determinations are involved, or when the initial 25-word score falls below 88%, the specialist must administer the full 50-word list to ensure statistical reliability.
Standardization and Administration Methodology
The Carrier Phrase Requirement
Every test word must be preceded by a standardized carrier phrase, most commonly:
- "Say the word... [target word]"
- "You will say... [target word]"
The carrier phrase serves three indispensable clinical functions:
- Alerting Function: Conditions the patient's executive attention, signaling that the target test token is imminent.
- Acoustic Context: Sets the natural vocal pitch, cadence, and acoustic volume for the target monosyllable, mimicking fluent speech.
- VU Meter Calibration in Monitored Live Voice: When testing via live voice, the clinician articulates the carrier phrase ("Say the word...") so that the carrier phrase peaks precisely at 0 VU (Volume Unit) on the audiometer's display. The subsequent target word is spoken with natural vocal inflection, typically peaking slightly below 0 VU.
AUDIOMETER VU METER
-20 -10 -7 -5 -3 -1 0 +1 +2 +3
┌───┬───┬───┬───┬───┬───┬───┼───┬───┬───┐
│ │ │ │ │ │ │ │ ▲ │ │ │
└───┴───┴───┴───┴───┴───┴───┴─│─┴───┴───┘
│
Carrier phrase ("Say the word...")
must peak precisely at 0 VU
Recorded Speech vs. Monitored Live Voice (MLV)
Clinical standards strongly dictate the delivery transducer medium:
| Feature | Digitally Recorded Speech (CD / WAV) | Monitored Live Voice (MLV) |
|---|---|---|
| Acoustic Standardization | Gold Standard: Identical frequency spectrum, vocal intensity, and timing on every presentation | Highly variable; subject to vocal strain, pitch drift, and microphone distance shifts |
| Inter-Examiner Reliability | Absolute; test results can be compared across clinics, providers, and years | Extremely poor; results cannot be objectively compared across providers |
| Dialect and Accent Bias | Calibrated standard General American dialect | Introduces regional clinician accents, pronunciation idiosyncrasies, and co-articulation differences |
| VU Meter Calibration | Precise 1000 Hz calibration tone matched to 0 VU | Requires continuous visual monitoring and vocal control on every carrier phrase |
| Clinical Flexibility | Rigid pacing; cannot pause mid-word without manual intervention | High; allows rapid pacing, pausing for coughing, or slowing for frail elderly patients |
| Diagnostic Rollover Validity | Mandatory for identifying retrocochlear lesions and calculating Rollover Indices | Invalid; vocal variations mask or artificially create rollover |
NBC-HIS guidelines emphasize that recorded speech is the professional standard of care. Monitored Live Voice should be reserved exclusively for patients with severe cognitive impairment, marked attention deficits, or physical limitations requiring customized pacing.
Presentation Level Protocols
Selecting the presentation level dictates the diagnostic validity of the Word Recognition Score. Clinicians utilize three primary presentation protocols depending on the diagnostic objective:
┌─────────────────────────────────────────────────────────────────────────────┐
│ WRS PRESENTATION LEVEL PROTOCOLS │
├─────────────────────────────────────────────────────────────────────────────┤
│ 1. SENSATION LEVEL (SL): SRT + 30 to 40 dB SL (Optimizes Speech Audibility) │
│ 2. MOST COMFORTABLE LOUDNESS (MCL): Patient Self-Selected Comfort Level │
│ 3. CONVERSATIONAL LEVEL: Fixed 50 to 60 dB HL (Tests Real-World Adequacy) │
└─────────────────────────────────────────────────────────────────────────────┘
1. Sensation Level Protocol (SRT + 30 to 40 dB SL)
- Rationale: Presenting words at a sensation level (dB SL) between 30 dB and 40 dB above the patient's SRT (most commonly $\text{SRT} + 35\text{ dB SL}$ or $\text{SRT} + 40\text{ dB SL}$) lifts the acoustic speech spectrum completely above the patient's pure-tone threshold curve across all critical speech frequencies (500 to 4000 Hz).
- Verification: The clinician must ensure that the calculated level ($\text{SRT} + 40\text{ dB}$) does not exceed the patient's Uncomfortable Loudness Level (UCL). If the level approaches the UCL, the presentation level must be lowered to 5 dB below the UCL.
2. Most Comfortable Loudness (MCL) Protocol
- Rationale: The patient is presented with running speech and asked to identify the level where speech sounds "most comfortable—neither too soft nor too loud." Words are then delivered at this subjective level.
- Application: Highly useful in patients with severe sensorineural recruitment where the dynamic range is compressed, making an arbitrary $+40\text{ dB SL}$ presentation intolerable.
3. Fixed Conversational Level (50 to 60 dB HL)
- Rationale: Presenting words at a fixed level of 50 to 60 dB HL (equivalent to approximately 65 dB SPL) evaluates how well the patient understands speech at average conversational loudness without hearing aid amplification.
- Application: Essential for patient counseling and demonstrating the functional communication deficit to family members. A patient who scores 92% at $+40\text{ dB SL}$ (85 dB HL) but drops to 52% at 55 dB HL vividly understands why amplification is required.
The Performance-Intensity (PI) Function and PB-Max
Defining the PI-PB Curve
The Performance-Intensity Function for Phonetically Balanced words (PI-PB curve) is a graphic plot illustrating word recognition percentage (y-axis, 0% to 100%) as a function of speech presentation level in dB HL (x-axis).
100% ┼───────────────────────╭──────────────────────── (Normal: PB-Max 100% Plateau)
│ ╭╯ │
│ ╭╯ │
Word │ ╭╯ ╰─────────────────────── (Cochlear: Slight Dip/Plateau)
Percent │ ╭╯
Correct │ ╭╯
│ ╭╯ PB-Max
│ ╭╯ │
│ ╭╯ ▼
│ ╭╯ ╭───────────╮
│ ╭╯ │ ╰───────╮ (Retrocochlear Rollover)
│ ╭╯ │ ╰───────┐
0% ┼─┴─────────────────────────┴───────────────────────────┴─
0 10 20 30 40 50 60 70 80 90 100 110
Presentation Level (dB HL)
Defining PB-Max
PB-Max is defined as the highest word recognition score achievable by a patient at any presentation level. To identify PB-Max, the clinician may need to test at multiple intensity levels, systematically raising the volume until the score reaches its peak and begins to plateau or decline.
PI-PB Behavior Across Auditory Pathologies
| Diagnostic Category | PB-Max Score Range | Presentation Level Required for PB-Max | Curve Characteristics at High Intensities (85-100 dB HL) |
|---|---|---|---|
| Normal Hearing | 98% to 100% | 25 to 30 dB SL<br/>(~30 to 40 dB HL) | Asymptotic Plateau: Score maintains 100% without degrading even at maximum output |
| Conductive Hearing Loss | 98% to 100% | SRT + 35 to 40 dB SL<br/>(Level shifted by ABG) | Shifted Plateau: Curve is shifted horizontally to the right by the air-bone gap; once audibility is restored, score reaches 100% |
| Cochlear / Sensory Loss | 50% to 90%<br/>(Depressed) | SRT + 30 to 40 dB SL<br/>(~70 to 85 dB HL) | Reduced Plateau / Minimal Rollover: Score peaks below 100% due to outer/inner hair cell loss; may show very slight rollover ($RI < 0.25$) |
| Retrocochlear Pathology<br/>(Acoustic Neuroma / CN VIII) | Highly Variable<br/>(Often 40% to 80%) | Moderate levels<br/>(~60 to 70 dB HL) | Severe Rollover ($RI \ge 0.45$): Score drops precipitously as intensity rises into the high suprathreshold range |
Clinical Classification of Word Recognition Scores
To interpret findings and establish realistic amplification expectations, clinicians utilize standardized performance categories:
| WRS Percentage | Clinical Interpretation | Typical Auditory Pathology | Amplification Prognosis & Counseling |
|---|---|---|---|
| 90% – 100% | Excellent | Normal hearing, mild loss, or pure conductive loss | Optimal: Restoring audibility via hearing instruments provides near-flawless speech clarity. Exceptional fitting prognosis. |
| 78% – 88% | Good | Mild to moderate sensorineural hearing loss | Good: Patient experiences substantial communicative benefit; minor difficulty decodable consonants in reverberation or noise. |
| 66% – 76% | Fair | Moderate sensorineural hearing loss | Satisfactory: Hearing aids restore conversational audibility in quiet; directional microphones and noise-reduction DSP are mandatory for noisy settings. |
| 54% – 64% | Poor | Moderately-severe sensorineural hearing loss | Guarded: Amplification restores loudness, but speech remains distorted. Requires assistive listening systems, remote microphones, and visual lip-reading cues. |
| < 50% | Very Poor | Severe-to-profound loss / severe phonemic regression | Very Guarded: Patient complains "I can hear that you are speaking, but I cannot understand the words." High risk of device abandonment; evaluate for cochlear implantation. |
The Retrocochlear Indicator: Rollover Phenomenon & Rollover Index
The Neurophysiology of Rollover
In a healthy ear or an ear with purely sensory (cochlear) hair cell damage, increasing speech presentation level to high intensities (85 to 95 dB HL) either maintains word clarity or produces a negligible reduction in recognition due to mechanical cochlear saturation.
However, in retrocochlear pathology—specifically space-occupying lesions such as a vestibular schwannoma (acoustic neuroma) or cerebellopontine angle (CPA) meningioma compressing the eighth cranial nerve (vestibulocochlear nerve)—the neural architecture fails under high acoustic loads:
- Neural Refractory Fatigue: Cranial nerve VIII consists of approximately 30,000 afferent nerve fibers. A compressed nerve suffers from focal demyelination and axonal ischemia.
- Conduction Block and Desynchronization: At low-to-moderate intensities, surviving nerve fibers maintain sufficient temporal phase-locking and synchronous discharge to transmit phonemic acoustic cues. However, when high-intensity sound floods the cochlea, the massive neural volley overwhelms the compromised nerve fibers.
- Catastrophic Intelligibility Collapse: The damaged fibers enter neural conduction block, discharge timing becomes disorganized, and acoustic speech cues are completely scrambled. The patient perceives high-volume speech as a loud, incomprehensible buzz. This paradoxical drop in word recognition at elevated sound intensities is termed the Rollover Phenomenon.
The Rollover Index Formula
To objectively quantify rollover, Jerger and Jerger (1971) formulated the Rollover Index ($RI$):
Where:
- $PB_{\text{max}}$ is the maximum word recognition score achieved at any presentation level on the PI function.
- $PB_{\text{min}}$ is the lowest (minimum) word recognition score obtained at any presentation level higher (louder) than the presentation level where $PB_{\text{max}}$ was achieved.
ROLLOVER INDEX FORMULA
PB_max - PB_min
RI = ───────────────────
PB_max
PB_max = Maximum score on PI-PB function (e.g., at 70 dB HL)
PB_min = Minimum score obtained at a HIGHER intensity (e.g., at 95 dB HL)
Diagnostic Decision Criteria
- Rollover Index $\ge 0.45$ (CID W-22 lists): Indicates significant rollover diagnostic of retrocochlear pathology (8th nerve lesion).
- Rollover Index $\ge 0.40$ (NU-6 lists): On Northwestern University NU-6 lists, the diagnostic criterion is slightly lower ($0.40$ to $0.45$) due to tighter phonemic balancing.
- Action Required: A positive rollover index is an absolute medical referral red flag. The specialist must immediately refer the patient to an otolaryngologist (ENT) or neurotologist for a contrast-enhanced Magnetic Resonance Imaging (MRI) of the internal auditory canals and brainstem.
Quantitative Clinical Case Example
A 56-year-old patient presents with unilateral high-frequency sensorineural hearing loss in the left ear. The clinician performs word recognition testing at multiple levels:
- Tested at $65\text{ dB HL} \implies \text{Score} = 76%$
- Tested at $75\text{ dB HL} \implies \text{Score} = 80% \implies PB_{\text{max}}$
- Tested at $90\text{ dB HL} \implies \text{Score} = 36% \implies PB_{\text{min}}$
-
Calculate the Rollover Index ($RI$):
-
Diagnostic Interpretation: The calculated $RI$ is $0.55$, which significantly exceeds the critical diagnostic cutoff of $0.45$. This patient demonstrates severe retrocochlear rollover, indicating Eighth Nerve pathology (suspected vestibular schwannoma), mandating urgent physician referral.
Clinical Masking for Word Recognition Testing
Why Masking is More Critical for WRS than Threshold Testing
During pure-tone threshold testing and SRT testing, stimuli are delivered at relatively low decibel levels near threshold. In contrast, Word Recognition Testing is delivered at high suprathreshold intensities (typically 70 to 90 dB HL). Consequently, the likelihood that sound will exceed interaural attenuation and stimulate the non-test cochlea via bone conduction is vastly higher during WRS than during any other audiometric test.
Masking Rule for Suprathreshold Speech
Contralateral masking in the non-test ear (NTE) is mandatory whenever the presentation level in the test ear exceeds the bone-conduction threshold of the non-test ear by the interaural attenuation of the transducer:
Where:
- $\text{Presentation Level}_{\text{TE}}$ is the decibel level (in dB HL) at which words are delivered to the test ear.
- $\text{Best BC}_{\text{NTE}}$ is the most sensitive (lowest) bone-conduction threshold of the non-test ear across 500, 1000, and 2000 Hz.
- $\text{IA}$ is the minimum interaural attenuation:
- $40\text{ dB}$ for supra-aural earphones (TDH-39/50).
- $60\text{ dB}$ for insert earphones (ER-3A/5A).
Transducer Advantage: Insert Earphones
Notice the profound clinical benefit of insert earphones during WRS:
- With supra-aural earphones ($\text{IA} = 40\text{ dB}$), delivering speech at $80\text{ dB HL}$ requires masking whenever non-test ear bone conduction is $40\text{ dB HL}$ or better ($80 - 40 = 40\text{ dB}$). Virtually every patient tested with supra-aural earphones requires masking during WRS!
- With insert earphones ($\text{IA} = 60\text{ dB}$), delivering speech at $80\text{ dB HL}$ only requires masking if non-test ear bone conduction is $20\text{ dB HL}$ or better ($80 - 60 = 20\text{ dB}$). Insert earphones eliminate the need for masking in over 60% of routine clinical WRS presentations.
Masking Noise and Level Selection
- Masking Stimulus: Speech-Spectrum Noise delivered to the non-test ear.
- Setting the Masking Level: Or more simply, delivered at an effective sensation level of 20 to 30 dB above the non-test ear's SRT or PTA, provided that level does not cross back and overmask the test ear ($M_{\text{NTE}} - \text{IA} < \text{BC}_{\text{TE}}$).
A 54-year-old patient undergoing audiometric evaluation achieves a maximum word recognition score of 80% at 70 dB HL (PB-Max) in the right ear using recorded NU-6 lists. When presented with words at 90 dB HL in the same ear, the score drops to 36% (PB-Min). What is the patient's Rollover Index (RI), and what clinical action is required?
Why is recorded speech presentation using standardized digital media considered mandatory over Monitored Live Voice (MLV) when evaluating a patient for retrocochlear rollover?
A clinician prepares to perform Word Recognition Testing on a patient's left ear at a presentation level of 80 dB HL using supra-aural earphones (IA = 40 dB). The right ear (non-test ear) exhibits normal bone-conduction thresholds with a best bone-conduction threshold of 10 dB HL. Which statement correctly identifies whether clinical masking is required in the right ear?