5.1 Validated Questionnaires & Screening Instruments
Key Takeaways
The Epworth Sleepiness Scale (ESS) rates the chance of dozing in 8 situations (0–3 each, total 0–24); a total above 10 indicates excessive daytime sleepiness, and 16–24 is severe.
The STOP-Bang questionnaire is an 8-item screening tool exhibiting high sensitivity (>90%) for moderate-to-severe OSA, categorizing patients into low (0-2), intermediate (3-4), or high risk (5-8, or 2+ STOP items plus specific demographic risk factors).
The Berlin Questionnaire stratifies OSA risk into high or low across three symptom categories (snoring/apneas, daytime sleepiness/fatigue, and hypertension/BMI >30 kg/m²), requiring positive findings in at least two categories for high risk.
The Insomnia Severity Index (ISI, clinical cutoff ≥15) and Pittsburgh Sleep Quality Index (PSQI, global cutoff >5) provide validated metrics for quantifying insomnia severity, overall sleep disturbance, and longitudinal treatment response.
The FOSQ-10 scores how sleepiness affects daily function from 5 to 20 (higher is better); 17.9 or higher is commonly treated as normal.
5.1 Validated Questionnaires & Screening Instruments
Clinical sleep assessment begins with systematic triage using standardized, validated questionnaires. These instruments quantify subjective symptom burdens, stratify patient risk for obstructive sleep apnea (OSA) and insomnia, and establish objective baselines against which therapeutic efficacy can be measured over time. While questionnaires cannot substitute for objective diagnostic sleep testing, their structured implementation optimizes clinical resource allocation, identifies occult sleep pathology in ambulatory and perioperative settings, and reveals discrepancies between subjective patient perceptions and objective physiological parameters.
Epworth Sleepiness Scale (ESS)
Developed by Dr. Murray Johns in 1991, the Epworth Sleepiness Scale (ESS) is the worldwide standard for quantifying subjective daytime sleep propensity. Rather than assessing generalized physical fatigue, lethargy, or tiredness, the ESS specifically measures a patient's likelihood of falling asleep in active and passive everyday situations.
Structure and Administration
The instrument presents eight distinct daily scenarios. Patients rate their likelihood of dozing or nodding off on a 4-point Likert scale (0 to 3) based on their recent life experience:
- 0 = Would never doze
- 1 = Slight chance of dozing
- 2 = Moderate chance of dozing
- 3 = High chance of dozing
The eight standardized situational items are:
- Sitting and reading
- Watching television
- Sitting inactive in a public place (such as a theater, lecture hall, or business meeting)
- As a passenger in a car for an hour without a break
- Lying down to rest in the afternoon when circumstances permit
- Sitting and talking to someone
- Sitting quietly after a lunch without alcohol
- In a car, while stopped for a few minutes in traffic
Scoring and Clinical Stratification
Individual item scores are summed to generate a cumulative score ranging from 0 to 24:
- 0 to 5: Lower normal daytime sleepiness
- 6 to 10: Higher normal daytime sleepiness
- 11 to 12: Mild excessive daytime sleepiness (EDS)
- 13 to 15: Moderate excessive daytime sleepiness
- 16 to 24: Severe excessive daytime sleepiness
These bands are the ones published by the scale's developer, Dr. Johns.
A total score above 10 is the conventional threshold for excessive daytime sleepiness that warrants further evaluation.
Clinical Interpretation and Limitations
The ESS measures a persistent trait—the general daytime sleep propensity—rather than a fluctuating acute state. However, clinical sleep health specialists must recognize several critical limitations:
- Subjective Recall Bias and Habituation: Patients with chronic, longstanding sleep disorders often habituate to daytime somnolence, misinterpreting severe sleepiness as normal aging or occupational stress.
- Deliberate Underreporting: Individuals in safety-critical occupations (such as commercial motor vehicle drivers, airline pilots, and heavy equipment operators) frequently underreport symptoms on the ESS to avoid regulatory disqualification or license suspension.
- Fatigue vs. Sleepiness Confusion: Patients frequently conflate physical exhaustion, low energy, and depressive apathy with true physiological sleepiness.
- Discordance with Objective Metrics: The ESS demonstrates only weak-to-moderate correlation with objective measures of daytime sleep propensity, such as the Multiple Sleep Latency Test (MSLT). A patient with severe sleep-disordered breathing may record an ESS score under 10, while a patient with an elevated ESS may demonstrate normal objective latency on testing.
STOP-Bang Questionnaire
Originally developed and validated by Dr. Frances Chung and colleagues for preoperative anesthesia risk stratification, the STOP-Bang Questionnaire has become the primary screening tool for obstructive sleep apnea across primary care, cardiology, bariatrics, and sleep clinics.
The Eight Dichotomous Components
The questionnaire consists of eight concise, objective Yes/No questions (1 point per affirmative answer, range 0 to 8):
| Acronym | Clinical Domain | Screening Criterion |
|---|---|---|
| S | Snoring | Do you snore loudly (louder than talking or loud enough to be heard through closed doors)? |
| T | Tired | Do you often feel tired, fatigued, or sleepy during daytime? |
| O | Observed | Has anyone observed you stop breathing, choking, or gasping during your sleep? |
| P | Pressure | Do you have or are you being treated for high blood pressure? |
| B | BMI | Body Mass Index greater than 35 kg/m²? |
| A | Age | Age older than 50 years? |
| N | Neck | Neck size large: shirt collar 17 inches (43 cm) or larger for men, or 16 inches (41 cm) or larger for women? |
| G | Gender | Male gender assigned at birth? |
Scoring and Risk Stratification
Patients are stratified into three distinct risk tiers based on cumulative score:
- Low Risk for OSA: Score of 0 to 2
- Intermediate Risk for OSA: Score of 3 to 4
- High Risk for OSA: Score of 5 to 8
Additionally, validated alternative criteria classify a patient as High Risk even with a lower cumulative score if they meet:
- 2 or more STOP items PLUS Male gender; OR
- 2 or more STOP items PLUS BMI >35 kg/m²; OR
- 2 or more STOP items PLUS neck circumference of 17 in or more (male) or 16 in or more (female).
Diagnostic Performance Profile
The primary clinical strength of STOP-Bang lies in its exceptional sensitivity:
- Sensitivity exceeds 90% to 93% for moderate-to-severe OSA (Apnea-Hypopnea Index [AHI] ≥15 events/hour).
- Sensitivity approaches 96% to 100% for severe OSA (AHI ≥30 events/hour).
- Specificity is modest (30% to 45%), reflecting a trade-off designed to minimize false negatives during clinical triage.
- Because of its high negative predictive value (NPV), a score of 0 to 2 reliably rules out moderate-to-severe OSA in general surgical and clinic populations.
Berlin Questionnaire
The Berlin Questionnaire is an epidemiological and clinical screening instrument developed during the 1996 Conference on Sleep in Primary Care in Berlin, Germany. It assesses OSA risk across three distinct categorical symptom domains.
Structure and Scoring Categories
The instrument comprises 10 self-administered questions organized into three clinical categories:
-
Category 1: Snoring and Witnessed Apneas (5 questions)
- Evaluates snoring presence, loudness, frequency, disruptiveness to others, and witnessed breathing cessations.
- Positive Score: High-risk responses in 2 or more items within this category.
-
Category 2: Daytime Sleepiness and Fatigue (4 questions)
- Evaluates feeling tired or unrested upon awakening, daytime fatigue/tiredness frequency, and nodding off while driving.
- Positive Score: High-risk responses in 2 or more items within this category.
-
Category 3: Cardiovascular Comorbidity and Obesity (1 question)
- Evaluates diagnosed systemic hypertension or a measured Body Mass Index >30 kg/m².
- Positive Score: Presence of hypertension OR a BMI exceeding 30 kg/m².
Risk Classification
- High Risk for OSA: Positive scores in 2 or more categories (e.g., Categories 1 and 3, or Categories 1, 2, and 3).
- Low Risk for OSA: Positive score in only 1 category or 0 categories.
In primary care settings, the Berlin Questionnaire demonstrates a sensitivity of approximately 86% and a specificity of 77% for predicting an AHI >5, making it a reliable referral filter for general medical practitioners.
Insomnia Severity Index (ISI)
Developed by Dr. Charles Morin, the Insomnia Severity Index (ISI) is a brief, 7-item patient-reported outcome measure evaluating the nature, severity, and daytime impact of insomnia symptoms over the preceding two weeks.
Evaluated Domains and Scoring
Each item is scored on a 5-point Likert scale (0 = none/not at all, to 4 = very severe/very much), yielding a global score of 0 to 28:
- Severity of sleep onset difficulty (initial insomnia)
- Severity of sleep maintenance difficulty (middle insomnia)
- Severity of early morning awakening problems (terminal insomnia)
- Current satisfaction with sleep pattern
- Noticeability of sleep impairment to others regarding quality of life
- Level of distress or worry caused by current sleep difficulties
- Degree of daytime interference with occupational and social functioning
Clinical Thresholds and Interpretation
| Global Score Range | Clinical Category | Clinical Meaning |
|---|---|---|
| 0 to 7 | No clinically significant insomnia | Normal sleep patterns; absence of clinically relevant pathology |
| 8 to 14 | Subthreshold insomnia | Mild insomnia symptoms; monitor or provide preventative sleep hygiene education |
| 15 to 21 | Clinical insomnia (moderate severity) | Clinically significant insomnia warranting behavioral therapy or formal clinical evaluation |
| 22 to 28 | Clinical insomnia (severe) | Severe insomnia associated with profound daytime disability and psychological distress |
Longitudinal Tracking in Behavioral Sleep Medicine
The ISI is the most widely used instrument for tracking patient response to Cognitive Behavioral Therapy for Insomnia (CBT-I) and pharmacotherapy. Clinical significance benchmarks include:
- Clinically Meaningful Improvement: A reduction of ≥6 points from baseline.
- Clinical Remission: Achieving a post-treatment score of ≤7 points.
Pittsburgh Sleep Quality Index (PSQI)
Developed in 1989 by Dr. Daniel Buysse and colleagues at the University of Pittsburgh, the Pittsburgh Sleep Quality Index (PSQI) evaluates broad, multi-dimensional sleep quality and sleep disturbances over a 1-month recall window.
Instrument Architecture
The questionnaire includes 19 self-administered questions (and 5 optional bed partner/roommate questions used for qualitative clinical insight). The 19 items combine to form 7 clinical component scores, each rated from 0 (no difficulty) to 3 (severe difficulty):
- Subjective sleep quality (overall self-rating)
- Sleep latency (minutes to fall asleep and frequency of taking >30 minutes)
- Sleep duration (hours of actual nocturnal sleep achieved)
- Habitual sleep efficiency (ratio of total sleep time to time in bed)
- Sleep disturbances (nocturnal awakenings, coughing, nocturia, pain, temperature)
- Use of sleep medications (prescription and over-the-counter sleep aids)
- Daytime dysfunction (trouble staying awake and maintaining enthusiasm)
Global Scoring and Diagnostic Cutoff
The seven component scores are summed to generate a Global PSQI Score ranging from 0 to 21:
- Global Score ≤5: Good sleep quality
- Global Score >5: Poor sleep quality
A global score cutoff >5 demonstrates a sensitivity of 89.6% and a specificity of 86.5% in distinguishing good sleepers from poor sleepers with clinical sleep disorders.
Functional Outcomes of Sleep Questionnaire (FOSQ)
The blueprint names the Functional Outcomes of Sleep Questionnaire (FOSQ) among the tools a CCSH must administer and interpret. Developed by Terri Weaver and colleagues in 1997, the FOSQ measures how sleepiness affects everyday functioning, not just how likely a person is to doze.
Structure
- FOSQ-30: the original 30-item version with five subscales: Activity Level, Vigilance, Intimacy and Sexual Relationships, General Productivity, and Social Outcome.
- FOSQ-10: a 10-item short form (Chasens and colleagues, 2009) that keeps the same five subscales and tracks the full version closely; it is the version most clinics use.
- Each item asks whether the person has difficulty with an activity because they are sleepy or tired, rated from 1 (extreme difficulty) to 4 (no difficulty), with an option for "I don't do this activity for other reasons."
Scoring and Interpretation
- Subscale scores combine into a total score from 5 to 20. Higher scores mean better functioning.
- A total of 17.9 or higher is commonly used as the threshold for normal functioning, based on Weaver's normative work.
- A change of about 2 points is generally treated as clinically meaningful.
Clinical Use
The FOSQ captures outcomes patients care about, such as staying alert while driving, getting work done and intimacy. It improves with effective treatment, so programs report it alongside PAP adherence data to show benefit. Its direction is the opposite of the ESS: a rising FOSQ is good, while a rising ESS is bad. Read score changes carefully before documenting them.
Tip
The sleep diary is also a named blueprint tool. The sleep history section shows how to calculate time in bed, total sleep time and sleep efficiency from a two-week diary.
Comprehensive Questionnaire Comparison Table
| Screening Tool | Primary Target Domain | Number of Items & Scoring Range | Validated Diagnostic Cutoff | Sensitivity / Specificity Profile | Clinical Utility & Limitations |
|---|---|---|---|---|---|
| Epworth Sleepiness Scale (ESS) | Subjective daytime sleep propensity across daily life situations | 8 items; scored 0–3 per item (Total: 0–24) | Score >10 indicates Excessive Daytime Sleepiness (EDS) | Modest sensitivity; low-to-moderate correlation with objective MSLT | Excellent triage tool; vulnerable to subjective underreporting and habituation bias |
| STOP-Bang | Preoperative and ambulatory screening for OSA risk | 8 dichotomous items; scored 0 or 1 (Total: 0–8) | 0–2: Low Risk; 3–4: Intermediate Risk; 5–8: High Risk (or 2+ STOP + risk factor) | Sensitivity >90% for moderate-to-severe OSA; specificity 30–45% | Outstanding triage sensitivity and negative predictive value; high false-positive rate due to low specificity |
| Berlin Questionnaire | Primary care identification of obstructive sleep apnea risk | 10 items grouped into 3 distinct categories | Positive score in ≥2 of 3 categories defines High Risk | Sensitivity ~86%, specificity ~77% in ambulatory populations | Well-validated in primary care; categorizes risk rather than providing numerical severity score |
| Insomnia Severity Index (ISI) | Subjective insomnia severity, daytime impact, and distress | 7 items; scored 0–4 per item (Total: 0–28) | ≥15: Moderate clinical insomnia; ≥22: Severe clinical insomnia; ≤7: Remission | A score of about 10–11 or higher best detected insomnia in validation studies | Gold standard for evaluating CBT-I outcomes; clinically meaningful response is ≥6-point decrease |
| Pittsburgh Sleep Quality Index (PSQI) | Multi-dimensional sleep quality and disturbance over past month | 19 self-rated items; 7 component scores (Total: 0–21) | Global Score >5 indicates poor sleep quality | Sensitivity 89.6%, specificity 86.5% for poor sleep quality | Comprehensive global measure; 1-month recall window limits utility for acute night-to-night tracking |
| Functional Outcomes of Sleep Questionnaire (FOSQ-10) | Effect of sleepiness on daily functioning | 10 items in 5 subscales (Total: 5–20) | Total ≥17.9 treated as normal functioning | Responsive to treatment; about 2 points is a meaningful change | Higher is better (opposite of ESS); measures function, not dozing |
A 54-year-old male commercial truck driver is referred to an accredited sleep center for evaluation. His STOP-Bang screening reveals a score of 6 (positive for loud snoring, witnessed pauses, hypertension, age >50, neck circumference 17.5 inches, and male gender), but his Epworth Sleepiness Scale (ESS) score is 4 out of 24. What is the most clinically sound interpretation of these discordant findings?
High OSA risk; the low ESS likely reflects underreporting or unrecognized sleepiness, so testing is still needed
The low ESS rules out clinically significant sleep apnea, so diagnostic sleep testing is not needed
The findings point to psychophysiological insomnia rather than collapse of the upper airway
STOP-Bang is invalid in commercial drivers because it relies on self-reported body measurements and history alone
A clinical sleep health specialist evaluates an adult patient completing the Insomnia Severity Index (ISI) at the initiation of Cognitive Behavioral Therapy for Insomnia (CBT-I). The patient's baseline score is 19. Following six sessions of CBT-I, the score decreases to 11. How should the specialist interpret this clinical change?
Full remission, because any score below 15 means that no residual insomnia remains
A meaningful response (an 8-point drop) from moderate to subthreshold insomnia, but not remission
Treatment failure, because any score above 7 means the patient did not respond at all
Progression to severe insomnia that now requires immediate hypnotic medication in addition to CBT-I
When scoring the Berlin Questionnaire to assess an adult patient's risk profile for obstructive sleep apnea, which combination of findings meets the threshold for high-risk categorization?
Positive scores in at least two of the three categories, such as snoring and hypertension/BMI
A positive score in the snoring and witnessed-apnea category alone, regardless of other findings
A positive score in the hypertension/BMI category alone, based on a BMI of 32 kg/m²
At least 8 affirmative answers in total across any of the questionnaire's categories
A patient's FOSQ-10 total score rises from 13.5 to 18.2 after three months of CPAP. How should the clinical sleep health specialist interpret this change?
Functioning improved meaningfully and now reaches the commonly used normal threshold of 17.9
Daytime sleepiness worsened, because higher FOSQ scores mean more impairment, as on the ESS
The change is too small to matter, because the FOSQ needs a 10-point change to be meaningful
The result is invalid, because the FOSQ can only be used before treatment, not to track it
Sections you finish are checked off in the contents.