5.1 Validated Questionnaires & Screening Instruments

Key Takeaways

  • The Epworth Sleepiness Scale (ESS) rates the chance of dozing in 8 situations (0–3 each, total 0–24); a total above 10 indicates excessive daytime sleepiness, and 16–24 is severe.

  • The STOP-Bang questionnaire is an 8-item screening tool exhibiting high sensitivity (>90%) for moderate-to-severe OSA, categorizing patients into low (0-2), intermediate (3-4), or high risk (5-8, or 2+ STOP items plus specific demographic risk factors).

  • The Berlin Questionnaire stratifies OSA risk into high or low across three symptom categories (snoring/apneas, daytime sleepiness/fatigue, and hypertension/BMI >30 kg/m²), requiring positive findings in at least two categories for high risk.

  • The Insomnia Severity Index (ISI, clinical cutoff ≥15) and Pittsburgh Sleep Quality Index (PSQI, global cutoff >5) provide validated metrics for quantifying insomnia severity, overall sleep disturbance, and longitudinal treatment response.

  • The FOSQ-10 scores how sleepiness affects daily function from 5 to 20 (higher is better); 17.9 or higher is commonly treated as normal.

Last updated: October 2026

5.1 Validated Questionnaires & Screening Instruments

Clinical sleep assessment begins with systematic triage using standardized, validated questionnaires. These instruments quantify subjective symptom burdens, stratify patient risk for obstructive sleep apnea (OSA) and insomnia, and establish objective baselines against which therapeutic efficacy can be measured over time. While questionnaires cannot substitute for objective diagnostic sleep testing, their structured implementation optimizes clinical resource allocation, identifies occult sleep pathology in ambulatory and perioperative settings, and reveals discrepancies between subjective patient perceptions and objective physiological parameters.


Epworth Sleepiness Scale (ESS)

Developed by Dr. Murray Johns in 1991, the Epworth Sleepiness Scale (ESS) is the worldwide standard for quantifying subjective daytime sleep propensity. Rather than assessing generalized physical fatigue, lethargy, or tiredness, the ESS specifically measures a patient's likelihood of falling asleep in active and passive everyday situations.

Structure and Administration

The instrument presents eight distinct daily scenarios. Patients rate their likelihood of dozing or nodding off on a 4-point Likert scale (0 to 3) based on their recent life experience:

  • 0 = Would never doze
  • 1 = Slight chance of dozing
  • 2 = Moderate chance of dozing
  • 3 = High chance of dozing

The eight standardized situational items are:

  1. Sitting and reading
  2. Watching television
  3. Sitting inactive in a public place (such as a theater, lecture hall, or business meeting)
  4. As a passenger in a car for an hour without a break
  5. Lying down to rest in the afternoon when circumstances permit
  6. Sitting and talking to someone
  7. Sitting quietly after a lunch without alcohol
  8. In a car, while stopped for a few minutes in traffic

Scoring and Clinical Stratification

Individual item scores are summed to generate a cumulative score ranging from 0 to 24:

  • 0 to 5: Lower normal daytime sleepiness
  • 6 to 10: Higher normal daytime sleepiness
  • 11 to 12: Mild excessive daytime sleepiness (EDS)
  • 13 to 15: Moderate excessive daytime sleepiness
  • 16 to 24: Severe excessive daytime sleepiness

These bands are the ones published by the scale's developer, Dr. Johns.

A total score above 10 is the conventional threshold for excessive daytime sleepiness that warrants further evaluation.

Clinical Interpretation and Limitations

The ESS measures a persistent trait—the general daytime sleep propensity—rather than a fluctuating acute state. However, clinical sleep health specialists must recognize several critical limitations:

  • Subjective Recall Bias and Habituation: Patients with chronic, longstanding sleep disorders often habituate to daytime somnolence, misinterpreting severe sleepiness as normal aging or occupational stress.
  • Deliberate Underreporting: Individuals in safety-critical occupations (such as commercial motor vehicle drivers, airline pilots, and heavy equipment operators) frequently underreport symptoms on the ESS to avoid regulatory disqualification or license suspension.
  • Fatigue vs. Sleepiness Confusion: Patients frequently conflate physical exhaustion, low energy, and depressive apathy with true physiological sleepiness.
  • Discordance with Objective Metrics: The ESS demonstrates only weak-to-moderate correlation with objective measures of daytime sleep propensity, such as the Multiple Sleep Latency Test (MSLT). A patient with severe sleep-disordered breathing may record an ESS score under 10, while a patient with an elevated ESS may demonstrate normal objective latency on testing.

STOP-Bang Questionnaire

Originally developed and validated by Dr. Frances Chung and colleagues for preoperative anesthesia risk stratification, the STOP-Bang Questionnaire has become the primary screening tool for obstructive sleep apnea across primary care, cardiology, bariatrics, and sleep clinics.

The Eight Dichotomous Components

The questionnaire consists of eight concise, objective Yes/No questions (1 point per affirmative answer, range 0 to 8):

AcronymClinical DomainScreening Criterion
SSnoringDo you snore loudly (louder than talking or loud enough to be heard through closed doors)?
TTiredDo you often feel tired, fatigued, or sleepy during daytime?
OObservedHas anyone observed you stop breathing, choking, or gasping during your sleep?
PPressureDo you have or are you being treated for high blood pressure?
BBMIBody Mass Index greater than 35 kg/m²?
AAgeAge older than 50 years?
NNeckNeck size large: shirt collar 17 inches (43 cm) or larger for men, or 16 inches (41 cm) or larger for women?
GGenderMale gender assigned at birth?

Scoring and Risk Stratification

Patients are stratified into three distinct risk tiers based on cumulative score:

  • Low Risk for OSA: Score of 0 to 2
  • Intermediate Risk for OSA: Score of 3 to 4
  • High Risk for OSA: Score of 5 to 8

Additionally, validated alternative criteria classify a patient as High Risk even with a lower cumulative score if they meet:

  • 2 or more STOP items PLUS Male gender; OR
  • 2 or more STOP items PLUS BMI >35 kg/m²; OR
  • 2 or more STOP items PLUS neck circumference of 17 in or more (male) or 16 in or more (female).

Diagnostic Performance Profile

The primary clinical strength of STOP-Bang lies in its exceptional sensitivity:

  • Sensitivity exceeds 90% to 93% for moderate-to-severe OSA (Apnea-Hypopnea Index [AHI] ≥15 events/hour).
  • Sensitivity approaches 96% to 100% for severe OSA (AHI ≥30 events/hour).
  • Specificity is modest (30% to 45%), reflecting a trade-off designed to minimize false negatives during clinical triage.
  • Because of its high negative predictive value (NPV), a score of 0 to 2 reliably rules out moderate-to-severe OSA in general surgical and clinic populations.

Berlin Questionnaire

The Berlin Questionnaire is an epidemiological and clinical screening instrument developed during the 1996 Conference on Sleep in Primary Care in Berlin, Germany. It assesses OSA risk across three distinct categorical symptom domains.

Structure and Scoring Categories

The instrument comprises 10 self-administered questions organized into three clinical categories:

  1. Category 1: Snoring and Witnessed Apneas (5 questions)

    • Evaluates snoring presence, loudness, frequency, disruptiveness to others, and witnessed breathing cessations.
    • Positive Score: High-risk responses in 2 or more items within this category.
  2. Category 2: Daytime Sleepiness and Fatigue (4 questions)

    • Evaluates feeling tired or unrested upon awakening, daytime fatigue/tiredness frequency, and nodding off while driving.
    • Positive Score: High-risk responses in 2 or more items within this category.
  3. Category 3: Cardiovascular Comorbidity and Obesity (1 question)

    • Evaluates diagnosed systemic hypertension or a measured Body Mass Index >30 kg/m².
    • Positive Score: Presence of hypertension OR a BMI exceeding 30 kg/m².

Risk Classification

  • High Risk for OSA: Positive scores in 2 or more categories (e.g., Categories 1 and 3, or Categories 1, 2, and 3).
  • Low Risk for OSA: Positive score in only 1 category or 0 categories.

In primary care settings, the Berlin Questionnaire demonstrates a sensitivity of approximately 86% and a specificity of 77% for predicting an AHI >5, making it a reliable referral filter for general medical practitioners.


Insomnia Severity Index (ISI)

Developed by Dr. Charles Morin, the Insomnia Severity Index (ISI) is a brief, 7-item patient-reported outcome measure evaluating the nature, severity, and daytime impact of insomnia symptoms over the preceding two weeks.

Evaluated Domains and Scoring

Each item is scored on a 5-point Likert scale (0 = none/not at all, to 4 = very severe/very much), yielding a global score of 0 to 28:

  1. Severity of sleep onset difficulty (initial insomnia)
  2. Severity of sleep maintenance difficulty (middle insomnia)
  3. Severity of early morning awakening problems (terminal insomnia)
  4. Current satisfaction with sleep pattern
  5. Noticeability of sleep impairment to others regarding quality of life
  6. Level of distress or worry caused by current sleep difficulties
  7. Degree of daytime interference with occupational and social functioning

Clinical Thresholds and Interpretation

Global Score RangeClinical CategoryClinical Meaning
0 to 7No clinically significant insomniaNormal sleep patterns; absence of clinically relevant pathology
8 to 14Subthreshold insomniaMild insomnia symptoms; monitor or provide preventative sleep hygiene education
15 to 21Clinical insomnia (moderate severity)Clinically significant insomnia warranting behavioral therapy or formal clinical evaluation
22 to 28Clinical insomnia (severe)Severe insomnia associated with profound daytime disability and psychological distress

Longitudinal Tracking in Behavioral Sleep Medicine

The ISI is the most widely used instrument for tracking patient response to Cognitive Behavioral Therapy for Insomnia (CBT-I) and pharmacotherapy. Clinical significance benchmarks include:

  • Clinically Meaningful Improvement: A reduction of ≥6 points from baseline.
  • Clinical Remission: Achieving a post-treatment score of ≤7 points.

Pittsburgh Sleep Quality Index (PSQI)

Developed in 1989 by Dr. Daniel Buysse and colleagues at the University of Pittsburgh, the Pittsburgh Sleep Quality Index (PSQI) evaluates broad, multi-dimensional sleep quality and sleep disturbances over a 1-month recall window.

Instrument Architecture

The questionnaire includes 19 self-administered questions (and 5 optional bed partner/roommate questions used for qualitative clinical insight). The 19 items combine to form 7 clinical component scores, each rated from 0 (no difficulty) to 3 (severe difficulty):

  1. Subjective sleep quality (overall self-rating)
  2. Sleep latency (minutes to fall asleep and frequency of taking >30 minutes)
  3. Sleep duration (hours of actual nocturnal sleep achieved)
  4. Habitual sleep efficiency (ratio of total sleep time to time in bed)
  5. Sleep disturbances (nocturnal awakenings, coughing, nocturia, pain, temperature)
  6. Use of sleep medications (prescription and over-the-counter sleep aids)
  7. Daytime dysfunction (trouble staying awake and maintaining enthusiasm)

Global Scoring and Diagnostic Cutoff

The seven component scores are summed to generate a Global PSQI Score ranging from 0 to 21:

  • Global Score ≤5: Good sleep quality
  • Global Score >5: Poor sleep quality

A global score cutoff >5 demonstrates a sensitivity of 89.6% and a specificity of 86.5% in distinguishing good sleepers from poor sleepers with clinical sleep disorders.


Functional Outcomes of Sleep Questionnaire (FOSQ)

The blueprint names the Functional Outcomes of Sleep Questionnaire (FOSQ) among the tools a CCSH must administer and interpret. Developed by Terri Weaver and colleagues in 1997, the FOSQ measures how sleepiness affects everyday functioning, not just how likely a person is to doze.

Structure

  • FOSQ-30: the original 30-item version with five subscales: Activity Level, Vigilance, Intimacy and Sexual Relationships, General Productivity, and Social Outcome.
  • FOSQ-10: a 10-item short form (Chasens and colleagues, 2009) that keeps the same five subscales and tracks the full version closely; it is the version most clinics use.
  • Each item asks whether the person has difficulty with an activity because they are sleepy or tired, rated from 1 (extreme difficulty) to 4 (no difficulty), with an option for "I don't do this activity for other reasons."

Scoring and Interpretation

  • Subscale scores combine into a total score from 5 to 20. Higher scores mean better functioning.
  • A total of 17.9 or higher is commonly used as the threshold for normal functioning, based on Weaver's normative work.
  • A change of about 2 points is generally treated as clinically meaningful.

Clinical Use

The FOSQ captures outcomes patients care about, such as staying alert while driving, getting work done and intimacy. It improves with effective treatment, so programs report it alongside PAP adherence data to show benefit. Its direction is the opposite of the ESS: a rising FOSQ is good, while a rising ESS is bad. Read score changes carefully before documenting them.

Tip

The sleep diary is also a named blueprint tool. The sleep history section shows how to calculate time in bed, total sleep time and sleep efficiency from a two-week diary.


Comprehensive Questionnaire Comparison Table

Screening ToolPrimary Target DomainNumber of Items & Scoring RangeValidated Diagnostic CutoffSensitivity / Specificity ProfileClinical Utility & Limitations
Epworth Sleepiness Scale (ESS)Subjective daytime sleep propensity across daily life situations8 items; scored 0–3 per item (Total: 0–24)Score >10 indicates Excessive Daytime Sleepiness (EDS)Modest sensitivity; low-to-moderate correlation with objective MSLTExcellent triage tool; vulnerable to subjective underreporting and habituation bias
STOP-BangPreoperative and ambulatory screening for OSA risk8 dichotomous items; scored 0 or 1 (Total: 0–8)0–2: Low Risk; 3–4: Intermediate Risk; 5–8: High Risk (or 2+ STOP + risk factor)Sensitivity >90% for moderate-to-severe OSA; specificity 30–45%Outstanding triage sensitivity and negative predictive value; high false-positive rate due to low specificity
Berlin QuestionnairePrimary care identification of obstructive sleep apnea risk10 items grouped into 3 distinct categoriesPositive score in ≥2 of 3 categories defines High RiskSensitivity ~86%, specificity ~77% in ambulatory populationsWell-validated in primary care; categorizes risk rather than providing numerical severity score
Insomnia Severity Index (ISI)Subjective insomnia severity, daytime impact, and distress7 items; scored 0–4 per item (Total: 0–28)≥15: Moderate clinical insomnia; ≥22: Severe clinical insomnia; ≤7: RemissionA score of about 10–11 or higher best detected insomnia in validation studiesGold standard for evaluating CBT-I outcomes; clinically meaningful response is ≥6-point decrease
Pittsburgh Sleep Quality Index (PSQI)Multi-dimensional sleep quality and disturbance over past month19 self-rated items; 7 component scores (Total: 0–21)Global Score >5 indicates poor sleep qualitySensitivity 89.6%, specificity 86.5% for poor sleep qualityComprehensive global measure; 1-month recall window limits utility for acute night-to-night tracking
Functional Outcomes of Sleep Questionnaire (FOSQ-10)Effect of sleepiness on daily functioning10 items in 5 subscales (Total: 5–20)Total ≥17.9 treated as normal functioningResponsive to treatment; about 2 points is a meaningful changeHigher is better (opposite of ESS); measures function, not dozing
Test Your Knowledge

A 54-year-old male commercial truck driver is referred to an accredited sleep center for evaluation. His STOP-Bang screening reveals a score of 6 (positive for loud snoring, witnessed pauses, hypertension, age >50, neck circumference 17.5 inches, and male gender), but his Epworth Sleepiness Scale (ESS) score is 4 out of 24. What is the most clinically sound interpretation of these discordant findings?

A

High OSA risk; the low ESS likely reflects underreporting or unrecognized sleepiness, so testing is still needed

B

The low ESS rules out clinically significant sleep apnea, so diagnostic sleep testing is not needed

C

The findings point to psychophysiological insomnia rather than collapse of the upper airway

D

STOP-Bang is invalid in commercial drivers because it relies on self-reported body measurements and history alone

Test Your Knowledge

A clinical sleep health specialist evaluates an adult patient completing the Insomnia Severity Index (ISI) at the initiation of Cognitive Behavioral Therapy for Insomnia (CBT-I). The patient's baseline score is 19. Following six sessions of CBT-I, the score decreases to 11. How should the specialist interpret this clinical change?

A

Full remission, because any score below 15 means that no residual insomnia remains

B

A meaningful response (an 8-point drop) from moderate to subthreshold insomnia, but not remission

C

Treatment failure, because any score above 7 means the patient did not respond at all

D

Progression to severe insomnia that now requires immediate hypnotic medication in addition to CBT-I

Test Your Knowledge

When scoring the Berlin Questionnaire to assess an adult patient's risk profile for obstructive sleep apnea, which combination of findings meets the threshold for high-risk categorization?

A

Positive scores in at least two of the three categories, such as snoring and hypertension/BMI

B

A positive score in the snoring and witnessed-apnea category alone, regardless of other findings

C

A positive score in the hypertension/BMI category alone, based on a BMI of 32 kg/m²

D

At least 8 affirmative answers in total across any of the questionnaire's categories

Test Your Knowledge

A patient's FOSQ-10 total score rises from 13.5 to 18.2 after three months of CPAP. How should the clinical sleep health specialist interpret this change?

A

Functioning improved meaningfully and now reaches the commonly used normal threshold of 17.9

B

Daytime sleepiness worsened, because higher FOSQ scores mean more impairment, as on the ESS

C

The change is too small to matter, because the FOSQ needs a 10-point change to be meaningful

D

The result is invalid, because the FOSQ can only be used before treatment, not to track it

Sections you finish are checked off in the contents.