6.1 Speech Reception (SRT), Detection (SDT) & Threshold Agreement
Key Takeaways
- The Speech Recognition / Reception Threshold (SRT) identifies the lowest decibel level (dB HL) at which a patient correctly identifies 50% of standardized spondaic words from a closed or open set.
- The Speech Detection Threshold (SDT) or Speech Awareness Threshold (SAT) identifies the lowest level where speech is detected 50% of the time, typically presenting 8 to 10 dB lower (better) than the SRT.
- The SRT-PTA Agreement Rule requires the SRT to match the 3-frequency Pure Tone Average (PTA at 500, 1000, 2000 Hz) or Fletcher 2-frequency PTA within ±6 dB; discrepancies of 7-11 dB are questionable, while ≥12 dB represent significant incongruence.
- An SRT significantly better than the PTA (>10-12 dB) is the primary audiometric hallmark of pseudohypacusis (functional or non-organic hearing loss), driven by the steep psychometric intelligibility curve of spondees.
- Clinical masking for speech thresholds is mandatory whenever the SRT in the test ear exceeds the best bone conduction threshold in the non-test ear by the interaural attenuation value (SRT_TE - BC_NTE ≥ IA), using speech-spectrum noise.
6.1 Speech Reception (SRT), Detection (SDT) & Threshold Agreement
[!NOTE] Pure-tone audiometry quantifies auditory sensitivity across discrete octave frequencies, but it cannot evaluate how the human auditory system processes complex, acoustically dynamic linguistic signals. Speech audiometry bridges the laboratory measurement of pure tones with real-world communicative functioning. On the NBC-HIS National Competency Examination (NCE), speech threshold testing and pure-tone threshold agreement serve as vital diagnostic reliability checks, cross-validating the audiometric test battery and exposing functional or non-organic hearing loss.
Speech Reception Threshold (SRT)
Clinical Definition and Purpose
The Speech Reception Threshold (SRT)—also termed the Speech Recognition Threshold under American Speech-Language-Hearing Association (ASHA) standards—is defined as the lowest hearing level (in decibels Hearing Level, dB HL) at which an individual can correctly recognize, repeat, or identify 50% of a standardized set of spondaic words.
The clinical objectives of establishing the SRT include:
- Validating Pure-Tone Thresholds: Serving as an internal cross-check of pure-tone air-conduction thresholds across the core speech frequencies (500, 1000, and 2000 Hz).
- Establishing a Reference Baseline for Suprathreshold Testing: Providing the baseline decibel reference level from which sensation levels (dB SL) are calculated for suprathreshold Word Recognition Testing (WRS) and Most Comfortable Loudness (MCL) measurements.
- Determining Hearing Aid Candidacy and Gain Requirements: Quantifying the minimum sound pressure level required for speech detection and recognition to guide prescriptive amplification targets (e.g., NAL-NL2, DSL v5).
Spondaic Words (Spondees)
The standardized speech materials utilized for SRT determination are spondaic words (commonly called spondees). A spondee is a two-syllable word pronounced with equal vocal effort and stress on both syllables.
SPONDAIC STRESS PATTERN (Spondee)
Syllable 1 Syllable 2
┌──────────────┐ ┌──────────────┐
│ BASE (50%) │ + │ BALL (50%) │ Equal Acoustic Energy
└──────────────┘ └──────────────┘
TROCHAIC STRESS PATTERN (Trochee)
Syllable 1 Syllable 2
┌──────────────┐ ┌──────────────┐
│ WA- (85%) │ + │ -ter (15%) │ Unequal (Strong-Weak)
└──────────────┘ └──────────────┘
In classical poetic meter, a spondee represents a metrical foot consisting of two long or stressed syllables. Common examples of standardized English spondees include:
- Baseball, hotdog, ice cream, airplane, cowboy, toothbrush, northwest, sidewalk, sunset, railroad, grandson, playground, mushroom, oatmeal, cupcake, doorway, pancake, drawbridge, blackboard, workshop.
Spondees must be carefully distinguished from trochees (words with strong-weak stress patterns, such as water, table, mother) and iambs (words with weak-strong stress patterns, such as today, guitar, above). Because both syllables of a spondee carry identical acoustic energy, the word provides a homogeneous acoustic stimulus that yields an exceptionally steep psychometric intelligibility growth function.
The Psychometric Function of Spondees
The psychometric function of spondaic words is the steepest of any speech material used in clinical audiometry. Intelligibility increases at a rate of approximately 8% to 10% per decibel increase in presentation level. Consequently, the transition from 0% recognition to 100% recognition occurs within a very narrow intensity window of only 10 to 12 dB.
100% ┼──────────────────────────────╭─────── (100% Intelligibility)
│ ╭╯
│ ╭╯ Spondee Slope: ~8-10% per dB
Word │ ╭╯ Narrow dynamic transition
Percent │ ╭╯ range (10 - 12 dB)
Correct │ ╭╯
50% ┼────────────────────────┼─────── SRT (50% Threshold)
│ ╭╯
│ ╭╯
│ ╭╯
0% ┼────────────────────┴───────────────────
Intensity (dB HL)
This rapid rise makes spondees the ideal stimulus for threshold determination, as small changes in stimulus level produce decisive changes in patient response accuracy, minimizing threshold ambiguity.
Familiarization Protocol
ASHA standards mandate that the patient must be familiarized with the spondee word list prior to threshold search. Familiarization is conducted at a suprathreshold, clearly audible level (typically 40 to 50 dB SL, or 50 to 60 dB HL for normal-hearing listeners):
- Auditory and Vocabulary Verification: The clinician presents the spondees via air-conduction transducers or monitored live voice. The patient listens and repeats each word, or reads along from a printed sheet.
- Elimination of Unfamiliar Tokens: Any word that the patient cannot articulate clearly, hesitates over, or fails to recognize due to cognitive limitations, dialectal variation, or non-native English background is permanently eliminated from that patient's test set.
- Response Task Conditioning: Familiarization ensures the patient understands the response paradigm and confirms that the test measures auditory sensitivity rather than lexical vocabulary knowledge.
Familiarization improves speech threshold sensitivity by approximately 4 to 5 dB. Skipping familiarization artificially elevates the SRT, creating false discrepancies with pure-tone thresholds.
Threshold Search Protocols
Once familiarized, the SRT is established using an adaptive psychophysical tracking procedure:
- Modified Hughson-Westlake Protocol (Ascending-Descending):
- Present the first spondee at a clearly audible level (e.g., 20 dB above estimated threshold).
- If correct, descend in 10 dB steps until the patient misses a word.
- Upon the first incorrect response, ascend in 5 dB steps until the patient responds correctly.
- Employ the "down 10 dB, up 5 dB" bracketing rule. Threshold is defined as the lowest attenuation level at which the patient correctly identifies at least 50% of the spondees (typically defined as 3 out of 5 presentations, or 2 out of 3 presentations at that specific decibel level).
- ASHA (1988) Bracketed Descending Protocol: Presents a fixed block of words (e.g., 4 or 5 spondees) per decibel level, starting suprathreshold and decreasing systematically in 5 dB or 2 dB steps, utilizing statistical stopping rules to calculate the 50% threshold.
Speech Detection Threshold (SDT) / Speech Awareness Threshold (SAT)
Definition and Diagnostic Scope
The Speech Detection Threshold (SDT)—synonymously termed the Speech Awareness Threshold (SAT)—is defined as the lowest sound pressure level (in dB HL) at which a listener can merely detect the presence of an acoustic speech signal 50% of the time.
Unlike the SRT, the SDT does not require the patient to recognize, identify, decode, or repeat speech tokens. The patient simply indicates awareness that speech sound energy is present (e.g., by raising a hand, pressing a response button, or turning their head toward a sound source).
Stimuli and Test Materials
- Continuous Discourse ("Cold Running Speech"): The clinician reads continuous, non-emotional factual prose (such as a news article or encyclopedia excerpt) in a rapid, monotonous voice, systematically lowering the attenuator.
- Repetitive Syllabic Trains: Presenting rapid nonsense syllables such as "ba-ba-ba", "da-da-da", or "puh-puh-puh".
- Pediatric Calling Signals: Utilizing the child's own name or familiar infant vocalizations ("uh-oh", "bye-bye").
Clinical Indications for SDT/SAT
The SDT is not administered routinely when an SRT can be reliably obtained. It is specifically indicated for patient populations unable to perform the cognitive-linguistic task of spondee repetition:
- Pediatric Patients: Infants, toddlers, and young children evaluated via Visual Reinforcement Audiometry (VRA) or Conditioned Play Audiometry (CPA) who lack the expressive language skills to repeat words.
- Cognitive and Neurological Impairments: Individuals with advanced dementia, severe intellectual disability, traumatic brain injury (TBI), or post-stroke expressive/receptive aphasia.
- Severe Language Barriers: Non-English-speaking patients when validated speech materials in their native language are unavailable.
- Profound Sensorineural Hearing Loss: Patients with profound cochlear damage where phonemic regression and severe distortion preclude speech recognition at any decibel level, yet residual low-frequency acoustic awareness remains intact.
The Quantitative Relationship: SRT vs. SDT
In standard clinical audiology, a precise physiological and acoustic relationship exists between the SDT and SRT:
The SDT is typically 8 to 10 dB lower (better / more sensitive) than the SRT.
Acoustic Dimension SDT / SAT SRT
┌──────────────────────┬────────────────────────────┬────────────────────────────┐
│ Cognitive Demand │ Awareness of Sound Only │ Cognitive Recognition & │
│ │ (No linguistic decoding) │ Phonemic Identification │
├──────────────────────┼────────────────────────────┼────────────────────────────┤
│ Primary Frequency │ Low Frequencies (250-500Hz)│ Core Speech Range │
│ Energy Utilized │ First Formants (F1) │ (500, 1000, 2000 Hz) │
├──────────────────────┼────────────────────────────┼────────────────────────────┤
│ Typical Level │ ~8 to 10 dB lower │ Baseline reference (0 dB) │
│ (Relative dB HL) │ (More sensitive) │ │
├──────────────────────┼────────────────────────────┼────────────────────────────┤
│ Agreement Standard │ Within ±5 dB of best pure- │ Within ±6 dB of 3-freq PTA │
│ │ tone threshold (250-4000Hz)│ (or Fletcher 2-freq PTA) │
└──────────────────────┴────────────────────────────┴────────────────────────────┘
The Acoustic-Phonetic Mechanism
Why is speech detected at an intensity 8 to 10 dB lower than it can be recognized? The human voice distributes acoustic power unevenly across the frequency spectrum. The vast majority of acoustic energy in conversational speech is concentrated in the fundamental frequency ($F_0$) and first formants ($F_1$) of vowels, located primarily between 250 Hz and 700 Hz.
- To establish an SDT, the listener only needs to detect this high-energy, low-frequency acoustic envelope. The peripheral auditory system senses the presence of sound without extracting phonemic boundaries.
- To establish an SRT, the listener must not only detect the low-frequency vowel energy but also perceive the higher-frequency second and third formants ($F_2, F_3$) and consonant acoustic cues (bursts, frication noise, formant transitions) necessary to differentiate "baseball" from "airplane" or "ice cream". This phonemic identification requires greater signal-to-noise ratio and higher decibel intensity.
Clinical Rule: The SDT should agree within ±5 dB of the best (lowest) pure-tone air-conduction threshold between 250 Hz and 4000 Hz (most commonly matching the threshold at 250 Hz or 500 Hz).
The SRT-PTA Agreement Rule
Calculating the Pure Tone Average (PTA)
The standard Three-Frequency Pure Tone Average (PTA) is calculated as the arithmetic mean of pure-tone air-conduction thresholds at 500 Hz, 1000 Hz, and 2000 Hz:
The Clinical Agreement Standard
Under clinical practice standards and NBC-HIS testing guidelines, the SRT must agree within $\pm 6\text{ dB}$ of the Pure Tone Average:
- Good Agreement: Discrepancy between $0\text{ and }6\text{ dB}$. Confirms high audiometric reliability and valid behavioral responses.
- Questionable / Fair Agreement: Discrepancy between $7\text{ and }11\text{ dB}$. Demands immediate re-instruction of the patient and verification of headphone placement, calibration, or pure-tone thresholds.
- Poor / Incongruent Agreement: Discrepancy $\ge 12\text{ dB}$. Indicates significant clinical pathology, technical error, or pseudohypacusis.
The Fletcher Two-Frequency Pure Tone Average
In audiograms exhibiting a flat, gently sloping, or mildly rising configuration, the standard 3-frequency PTA reliably matches the SRT. However, when an audiogram displays a steeply sloping high-frequency hearing loss—or a precipitous drop across the speech frequencies—the standard 3-frequency PTA substantially overestimates the hearing handicap and creates a false discrepancy with the SRT.
Frequency (Hz) 500 1000 2000 4000
┌───────┬───────┬───────┬───────┐
20 dB HL ─── │ X │ │ │ │
30 dB HL ─── │ │ X │ │ │ ◄── Slope: ≥ 20 dB drop
40 dB HL ─── │ │ │ │ │ between adjacent octaves
50 dB HL ─── │ │ │ │ │
60 dB HL ─── │ │ │ │ │
70 dB HL ─── │ │ │ │ │
80 dB HL ─── │ │ │ X │ X │
└───────┴───────┴───────┴───────┘
The Fletcher Rule
When there is a difference of $\ge 20\text{ dB}$ between any two adjacent frequencies among 500 Hz, 1000 Hz, and 2000 Hz, the clinician must calculate the Fletcher Two-Frequency PTA (also termed the Two-Best Frequency PTA):
Quantitative Clinical Example
A patient presents with the following air-conduction thresholds:
- $500\text{ Hz} = 20\text{ dB HL}$
- $1000\text{ Hz} = 30\text{ dB HL}$
- $2000\text{ Hz} = 75\text{ dB HL}$
- Measured $\text{SRT} = 25\text{ dB HL}$
-
Standard 3-Frequency PTA Calculation: Comparing SRT to 3-frequency PTA: Without further analysis, this 16.7 dB gap falsely suggests malingering or invalid test data.
-
Applying the Fletcher Protocol: Notice the precipitous drop between 1000 Hz (30 dB) and 2000 Hz (75 dB): the inter-octave difference is $75 - 30 = 45\text{ dB}$ (well exceeding the $\ge 20\text{ dB}$ threshold criterion). The clinician selects the two best thresholds: 500 Hz (20 dB) and 1000 Hz (30 dB). Comparing SRT to Fletcher PTA: The SRT exhibits perfect threshold agreement with the Fletcher PTA! The patient's auditory system utilized the preserved low-frequency speech cues to correctly identify spondees at 25 dB HL.
Diagnostic Significance of SRT-PTA Discrepancies
┌──────────────────────────────┐
│ EVALUATE |SRT - PTA| GAP │
└──────────────┬───────────────┘
│
┌──────────────────────────┴──────────────────────────┐
▼ ▼
┌───────────────────────┐ ┌───────────────────────┐
│ SRT Significantly │ │ SRT Significantly │
│ BETTER than PTA │ │ WORSE than PTA │
│ (SRT << PTA by ≥12dB)│ │ (SRT >> PTA by ≥10dB)│
└───────────┬───────────┘ └───────────┬───────────┘
│ │
▼ ▼
┌───────────────────────┐ ┌───────────────────────┐
│ PSEUDOHYPACUSIS │ │ CENTRAL / AUDITORY │
│ (Non-Organic Hearing │ │ NEUROPATHY / LANGUAGE │
│ Loss / Malingering) │ │ (CAPD, ANSD, Aphasia) │
└───────────────────────┘ └───────────────────────┘
Discrepancy Type 1: SRT Significantly Better Than PTA (> 10 to 12 dB)
When the SRT is substantially lower (better) than the PTA by $12\text{ dB}$ or more (e.g., pure-tone PTA is 55 dB HL, but the patient repeats spondees at 20 dB HL), the finding is pathognomonic for Pseudohypacusis—also designated as functional hearing loss, non-organic hearing loss, or malingering.
Why Spondees Defeat the Malingerer
An individual attempting to feign or exaggerate a hearing loss (for financial compensation, disability claims, or psychological secondary gain) typically uses loudness reference anchoring during pure-tone testing. Because pure tones rise slowly in subjective loudness, the malingerer can choose an internal loudness anchor (e.g., "I will only push the button when the tone sounds loud") and respond consistently at 50 or 60 dB HL.
However, when the clinician introduces spondaic words, this deception collapses due to the exceptionally steep psychometric function of spondees:
- At 5 dB above detection, spondees are already intelligible.
- Spondees presented at 20 or 25 dB HL sound clear, effortless, and overwhelmingly loud to a person with normal organic hearing.
- The malingerer perceives the spondees as impossible to ignore and assumes that if words sound that loud and distinct, anyone with a hearing loss would hear them. Consequently, they repeat the words, inadvertently revealing their true organic hearing threshold.
Diagnostic Behavioral Hallmarks
- Half-Word Spondee Responses: The patient repeats only one syllable of the spondee (e.g., Clinician presents "baseball", patient responds "base..."; clinician presents "hotdog", patient responds "...dog"). In authentic organic hearing loss, a patient either hears both syllables or misses the entire word; repeating half a spondee indicates that the stimulus was heard with sufficient audibility to decode phonemes, but the patient deliberately withheld the full response.
- Unexplained Acoustic Reflexes: Acoustic reflex thresholds occurring at or below the reported pure-tone thresholds (e.g., patient reports pure-tone thresholds of 70 dB HL, but bilateral acoustic reflexes are present at 75 dB HL).
- Ascending vs. Descending Threshold Gaps: Discrepancies $\ge 10-15\text{ dB}$ when pure tones are tracked ascending from inaudibility versus descending from suprathreshold audibility.
Discrepancy Type 2: SRT Significantly Worse Than PTA (> 6 to 10 dB)
When the SRT is poorer (higher in decibels) than the pure-tone average by $10\text{ dB}$ or more (e.g., pure-tone PTA is 20 dB HL, but the SRT is 45 dB HL), the patient exhibits an inability to recognize speech despite possessing normal or near-normal peripheral pure-tone sensitivity.
Primary Diagnostic Etiologies
- Central Auditory Processing Disorder (CAPD): Deficits in the central auditory nervous system (brainstem, thalamus, or auditory cortex) that impair phonemic decoding, temporal patterning, and binaural integration.
- Auditory Neuropathy Spectrum Disorder (ANSD): Characterized by intact outer hair cell function (present otoacoustic emissions and/or cochlear microphonic) paired with dyssynchronous cranial nerve VIII firing. Patients hear pure-tone sounds but cannot process speech timing cues.
- Cognitive Impairment / Early Dementia: Decline in executive working memory and linguistic processing speed.
- Receptive / Expressive Dysphasia: Neurological vascular insults (e.g., stroke affecting Wernicke's or Broca's area) impairing speech comprehension or motor speech repetition.
- Severe Language / Dialectal Barrier: Non-native English speakers unfamiliar with English phonology.
Clinical Masking for Speech Threshold Audiometry
Cross-Hearing Mechanics for Speech
Just as in pure-tone audiometry, acoustic energy presented to the test ear (TE) can cross the skull via bone conduction and stimulate the cochlea of the non-test ear (NTE). If the speech stimulus reaches the non-test cochlea at or above its bone-conduction threshold, cross-hearing occurs, producing a false shadow threshold.
CROSS-HEARING MECHANICS
Test Ear (TE) Non-Test Ear (NTE)
┌─────────────────┐ ┌─────────────────┐
│ Speech Stimulus │ ──── Transcranial Bone ───► │ Opposite │
│ Presented at │ Transmission │ Cochlea │
│ Attenuator Level│ (Loss = IA dB) │ Stimulated │
└─────────────────┘ └─────────────────┘
Masking Criteria for Speech Thresholds
Clinical masking must be delivered to the non-test ear whenever the speech presentation level in the test ear exceeds the bone-conduction sensitivity of the non-test ear by the interaural attenuation (IA) of the transducer:
Where:
- $\text{SRT}_{\text{TE}}$ is the measured unmasked speech reception threshold in the test ear.
- $\text{Best BC}_{\text{NTE}}$ is the most sensitive bone-conduction threshold in the non-test ear across 500, 1000, and 2000 Hz.
- $\text{IA}$ is the minimum interaural attenuation value for the transducer used:
- Supra-aural Earphones (TDH-39, TDH-50): $\text{IA} = 40\text{ dB}$
- Insert Earphones (ER-3A, ER-5A): $\text{IA} = 60\text{ dB}$
[!IMPORTANT] If bone-conduction thresholds for the non-test ear have not yet been established, the clinician must substitute the air-conduction Pure Tone Average of the non-test ear ($\text{PTA}_{\text{NTE}}$) under the assumption that bone conduction could be equal to air conduction. However, if an air-bone gap exists in the non-test ear, bone-conduction thresholds must be used to prevent severe undermasking.
Transducer Comparison: Interaural Attenuation for Speech
| Transducer Type | Coupling Method | Minimum IA for Speech | Clinical Advantage |
|---|---|---|---|
| Supra-Aural Earphones<br/>(TDH-39 / TDH-50) | Cushions pressed against the external ear / cranial bones | 40 dB | Large dynamic surface area; higher risk of canal collapse; requires masking frequently |
| Insert Earphones<br/>(ER-3A / ER-5A) | Expanding foam tip deeply seated in the cartilaginous ear canal | 60 dB | Drastically reduces acoustic contact with skull; increases IA by 20 dB; prevents canal collapse; rarely requires masking |
The Masking Noise: Speech-Spectrum Noise (SSN)
Standard white noise or narrow-band noise is acoustically inappropriate for masking speech stimuli:
- White Noise: Possesses equal acoustic energy per cycle across all frequencies ($0\text{ dB/octave}$ slope). Because speech energy is concentrated in the low frequencies, white noise over-amplifies high-frequency energy, delivering excessive acoustic power that causes patient discomfort and central masking without efficiently masking speech.
- Narrow-Band Noise (NBN): Tailored specifically for pure tones. Its narrow bandwidth allows speech formants outside the band to leak through unmasked.
- Speech-Spectrum Noise (Speech-Shaped Noise): The psychophysically validated masking stimulus for speech audiometry. It provides flat spectrum energy from 100 Hz to 1000 Hz, followed by a 12 dB per octave roll-off above 1000 Hz, precisely mirroring the Long-Term Average Speech Spectrum (LTASS). This shape delivers maximum masking efficiency with minimum total sound energy.
Masking Protocol for SRT (Hood Plateau Method for Speech)
When masking is indicated:
- Introduce speech-spectrum noise to the non-test ear at an initial effective masking level:
- Re-establish the SRT in the test ear.
- Increase masking noise in 5 dB steps across three consecutive increments (a 15 dB plateau) while verifying that the SRT in the test ear remains stable, confirming the true masked threshold without undermasking or overmasking.
A patient's pure-tone air-conduction thresholds are 25 dB HL at 500 Hz, 35 dB HL at 1000 Hz, and 80 dB HL at 2000 Hz. The measured unmasked Speech Reception Threshold (SRT) is 30 dB HL. How should the Hearing Instrument Specialist evaluate the relationship between the SRT and the pure-tone findings?
A 45-year-old worker undergoing a disability assessment demonstrates pure-tone air-conduction thresholds yielding a three-frequency PTA of 65 dB HL bilaterally. During speech audiometry using insert earphones (IA = 60 dB), the patient readily and consistently repeats spondaic words at 25 dB HL, frequently repeating only one syllable (e.g., responding 'dog' when presented with 'hotdog'). What is the primary clinical interpretation of these findings?
Under what clinical condition is a clinician justified in establishing a Speech Detection Threshold (SDT) instead of a Speech Reception Threshold (SRT), and what is the expected quantitative relationship between the two measurements in the same ear?