15.1 Spirometry & Lung Volume Reliability
Key Takeaways
- Domain III.B reliability asks whether accepted spirometry and lung-volume numbers are trustworthy for clinical use—not merely whether a maneuver met minimal software flags.
- Technical unreliability (poor effort, leak, incomplete pant, non-equilibration) must be separated from true physiologic abnormality before pattern labeling.
- Within-session repeatability supports confidence in a single visit; day-to-day change requires serial context and does not automatically equal disease progression.
- Effort-dependent PEF can look obstructive when effort is soft, while effort-independent mid-expiratory plateaus and consistent FEF patterns support true obstruction when trials are acceptable.
- For static volumes (III.B.4), leaks, incomplete pant loops, and dilution non-equilibration bias FRC/TLC/RV and can invent restriction or hyperinflation on paper.
Why Domain III.B Starts After Validity Flags
Domain II.C taught you to score acceptability and repeatability of maneuvers. Domain III.B — Evaluate reliability of results asks the next clinical question: Can these numbers be trusted for decisions? On the NBRC PFT Examination (high cut / RPFT), III.B.3.a/b (spirometry reliability) and III.B.4.a/b (lung-volume reliability) sit in Data Management. A waveform can pass automated end-of-test checks and still be unreliable if effort was soft, coaching incomplete, or the volume method never truly equilibrated.
Reliability judgment is analysis-level work: compare trials, compare indices that should move together, and decide whether a “bad-looking” result is disease or technique.
Technical Reliability vs Clinical Abnormality
| Finding on report | Prefer technical unreliability when… | Prefer true clinical abnormality when… |
|---|---|---|
| Low FVC / early termination look | Abrupt end-of-test, variable expiratory time, cough | Multiple acceptable full efforts, consistent plateau |
| Low FEV1 / “obstruction” | Soft peak, inconsistent PEF, rising FEV1 with coaching | Acceptable peaks, repeatable FEV1, concave MEFV |
| High RV / air trapping | Leak on dilution, incomplete pant, shutter artifact | Box and dilution agree; consistent with obstruction |
| Low TLC / “restriction” | Under-pant, mouth leak, early shutter | Repeatable TGV/FRC, consistent VC linkage |
| Variable PEF | Effort ramp incomplete; language barrier; pain | PEF stable across best efforts; disease pattern elsewhere |
Rule of thumb for exam stems: If one index is extreme while related indices and trial graphs disagree, search for technique first. If multiple independent signals converge (spirometry shape + volumes + history), clinical abnormality is more likely—after technical gates are clean.
Never reverse the order: do not invent disease to “explain” a single poor peak flow when the volume-time curve shows hesitation and the next trial improves 15% with coaching.
Within-Session Repeatability vs Day-to-Day Change
Within-session (same visit)
Within-session reliability is the primary NBRC concern for a single report:
- Spirometry: Best efforts cluster; FEV1 and FVC differences among acceptable trials stay within ATS/ERS-aligned repeatability limits used by the lab.
- Volumes: Replicate FRC/TGV trials agree; TLC derived from a stable FRC + reliable VC is internally consistent.
- Cross-check: FVC from spirometry and VC used in volume calculations should not tell incompatible stories without comment.
If the best three FEV1 values fan out widely after many attempts, reliability is low even if the single largest FEV1 is reported. Document variability; do not over-interpret a lonely “best” number.
Day-to-day / visit-to-visit
Day-to-day change includes biologic variation, medication timing, infection, and lab method differences. A drop in FEV1 between visits is not automatically progression and is not automatically technical error. Reliability thinking for serial data asks:
- Were both visits technically acceptable and repeatable?
- Same reference set, same BTPS handling, similar bronchodilator state?
- Is the change larger than expected biologic/test variability for that patient?
Exam items often contrast within-session scatter (fix/reliability of this session) with between-visit change (clinical tracking only after both sessions are solid).
When Spirometry “Looks Obstructive” Because Effort Failed
Obstruction is a pattern, not a PEF alone. Soft effort preferentially damages effort-dependent portions of the flow-volume loop.
Effort-dependent vs effort-independent cues
| Signal | Effort dependence | Soft-effort artifact | True obstruction (with good effort) |
|---|---|---|---|
| PEF | High | Low, variable trial-to-trial | May be reduced but more stable across best efforts |
| Early peak / blast | High | Rounded, delayed peak | Sharp peak then rapid fall |
| Mid-expiratory flows / plateau shape | Lower (more effort-independent once flow limited) | May look “scooped” artifactually if volume incomplete | Concavity with repeatable FEV1/FVC ratio |
| FEV1 repeatability | — | Improves markedly with coaching | Stays reduced despite maximal, repeatable peaks |
| Inspiratory limb | High | Weak inspiration, small IC | Usually preserved effort if coached |
Classic stem pattern: Low PEF, low FEV1, improving PEF and FEV1 after vigorous coaching and visual feedback, with rising peak sharpness → technical, not new COPD diagnosis.
Contrasting stem: Acceptable peaks, repeatable FEV1, concave expiratory limb, FEV1/FVC reduced on multiple best efforts → true airflow limitation (still report severity only after reliability is established).
Incomplete exhalation can also fake restriction (low FVC) or distort ratios. Reliability review always includes volume-time (plateau, expiratory time) and flow-volume together—not PEF in isolation.
Lung Volume Reliability (III.B.4)
Static volumes are only as reliable as FRC method quality plus VC linkage.
Body plethysmography pitfalls
| Problem | How it shows up | Reliability action |
|---|---|---|
| Mouth leak / nose leak | Open Pmouth–Pbox loops; TGV scatter; patient “feels air” | Reseat interface; reject trials; do not average leaks |
| Incomplete pant | Tiny pressures, noisy slope, shutter not truly closed | Re-coach gentle pant; verify shutter; retest |
| Wrong shutter timing | TGV not at FRC; TLC/RV nonsense vs spirometry | Restart at end-expiration per protocol |
| Thermal / door nonequilibrium | Baseline Pbox drift, shifting loops | Re-equilibrate cabin; delay testing |
| Glottic closure / straining | Nonphysiologic pressure spikes | Soften pant; reject non-loop trials |
Unreliable TGV propagates to TLC and RV. A “restricted” TLC from under-measured FRC is a reliability failure, not interstitial disease, until repeats confirm.
Gas dilution pitfalls
| Problem | Effect on FRC/TLC | Reliability cue |
|---|---|---|
| System or patient leak | Usually underestimates FRC (tracer loss) | Failure to plateau; rising/falling tracer unexpectedly |
| Incomplete equilibration | Wrong FRC (often low in poorly mixed / obstructed lungs if stopped early) | Tracer still changing at “end”; short wash-in |
| Poor switch-in timing | Starting volume error | Inconsistent replicates |
| O2 consumption / CO2 handling faults (method-specific) | Drift, unstable end criteria | QC and analyzer flags |
In severe obstruction, dilution FRC may be lower than box FRC because trapped gas is poorly accessible—this can be physiologic method difference, not always technologist error. Reliability language still requires: equilibration criteria met, no leak, and clear method labeling when box and dilution disagree.
Exam Scenarios (Reliability Framing)
Scenario A — Soft peak “obstruction.” New patient FEV1/FVC 0.62, PEF 45% predicted, highly variable peaks; after coaching, PEF rises to 90% predicted and ratio normalizes. Reliable conclusion: prior ratios were effort-contaminated; do not finalize obstructive pattern on the soft set.
Scenario B — Repeatable obstruction. Three acceptable efforts, FEV1 within repeatability, concave loops, PEF stable. Reliable conclusion: airflow limitation pattern is supportable for clinical interpretation (III.C comes next).
Scenario C — Dilution “restriction.” He dilution FRC low, TLC low, but patient has known bullae and box TGV much higher with good pant loops. Reliability: dilution may be incomplete for trapped gas; do not report isolated dilution TLC as definitive restriction without method context.
Scenario D — Within-session chaos. FVC differs by large absolute volume across “accepted” trials; patient fatigues. Action: more rest/coaching or limited report with explicit unreliability—not a confident severity grade.
Decision Workflow for the RPFT
- Confirm Domain II acceptability for the maneuvers you will use.
- Inspect within-session clustering of key indices (FEV1, FVC, FRC/TGV).
- Cross-check effort-dependent vs effort-independent signals for obstruction claims.
- For volumes, verify no leak, complete pant or equilibration, and sensible TLC = FRC + IC (or method-equivalent linkage).
- If unreliable: retest, limit the report, or flag—do not silently promote artifact to disease.
- Only then hand clean data to III.C clinical implications.
Domain III.B.3 and III.B.4 reward technologists who protect patients from both false disease and missed true obstruction by treating reliability as a separate, scored judgment—not a rubber stamp on the largest number printed.
A spirometry session shows FEV1/FVC reduced with very low, highly variable PEF. After vigorous coaching, PEF and FEV1 rise substantially and the ratio normalizes on repeatable efforts. What is the best reliability conclusion?
Which finding most strongly supports true airflow limitation rather than poor peak effort alone?
During N2 washout, tracer concentration is still changing when the system stops, and replicate FRC values disagree widely. Per III.B.4 reliability thinking, what should the technologist do?
How should within-session spirometry repeatability be used differently from day-to-day FEV1 change on the exam?