13.5 Serial Pulmonary Function Testing, Clinical History, Demographics, and Laboratory Quality Management
Key Takeaways
- Serial pulmonary function testing, clinical history and demographics, and laboratory quality management are Domain III items 17, 18, and 19, each examined for calculation, reliability, and clinical implication.
- A change between visits is meaningful only when it exceeds within-subject biological and measurement variability, conventionally about 12% or 200 mL for FEV1 within a day and about 15% between years.
- A DLCO change is considered significant when it exceeds roughly 10% within a day or about 20% between years, reflecting the larger inherent variability of the measurement.
- Demographic inputs drive the predicted value, so an entry error in height, age, or sex distorts percent predicted and z-score without changing a single measured value; the stadiometer, scale, and calipers are Domain I equipment that require their own quality control schedule.
- Laboratory quality management under the content outline includes standard operating procedures, inventory control, and customer satisfaction alongside technical quality control.
13.5 Serial Pulmonary Function Testing, Clinical History, Demographics, and Laboratory Quality Management
Three Domain III topics look administrative and are not: serial pulmonary function testing — trending a single patient (item 17), clinical history and demographics (item 18), and laboratory quality management (item 19). Each is examined under calculate, evaluate reliability, and evaluate clinical implications. Because Domain III carries no recall items, these appear as judgment problems: is this change real?
Serial Testing: Distinguishing Real Change from Noise
A patient whose FEV$_1$ was 2.40 L in March and 2.22 L in September has not necessarily declined. Every measurement carries within-subject biological variation and measurement variation, and a difference must exceed both before it means anything clinically.
Conventional Significant-Change Thresholds
| Parameter | Within one day | Week to week | Year to year |
|---|---|---|---|
| FEV$_1$ | $> 12%$ and $> 200$ mL | $> 15%$ | $> 15%$ |
| FVC | $> 12%$ and $> 200$ mL | $> 15%$ | $> 15%$ |
| TLC | — | — | $> 10%$ |
| DLCO | $> 10%$ | — | $> 20%$ |
Two properties fall out of this table and are worth stating explicitly:
- DLCO is noisier than spirometry. It requires a longer breath hold, depends on hemoglobin and carboxyhemoglobin, and involves gas analysis, so its year-to-year threshold is wider than spirometry's.
- Longer intervals require larger changes, because true biological drift accumulates alongside measurement error.
Separately, healthy adults lose FEV$_1$ at roughly 25–30 mL per year after the peak at age 20–25. An accelerated decline — commonly defined as more than about 40–60 mL/year sustained over several years — is itself the finding, even when no single interval crosses a threshold. Plotting absolute values against age against the predicted decline curve reveals this; comparing two isolated visits does not.
Worked example. FEV$_1$ 2.40 L in March, 2.22 L in September. Absolute change 0.18 L; percent change 0.18/2.40 = 7.5%. Neither the 200 mL nor the 12% criterion is met, so this is within measurement variability and should not be reported as a decline. Report the trend, note the direction, and recommend continued surveillance.
What Must Be Constant for a Trend to Exist
- Same laboratory and, ideally, the same device. A change of spirometer, sensor type, or software version resets the baseline; document it prominently on the report.
- Same reference equation set. Switching from NHANES III to GLI, or from race-specific to race-neutral equations, changes percent predicted and z-score without any change in the patient. When a laboratory migrates equations, re-express historical values in the new set before trending.
- Same time of day and medication state. Diurnal variation and bronchodilator timing both move FEV$_1$.
- Same posture, same effort quality. Compare the quality grade alongside the value; a Grade A session following a Grade D session may show an apparent "improvement" that is purely technical.
- Same hemoglobin adjustment for DLCO. An uncorrected DLCO in a patient whose hemoglobin dropped from 14 to 9 g/dL will appear to have fallen when diffusion is unchanged.
Report the technologist comment. A note reading "patient coughing throughout, best of 8 efforts, Grade D" is the single most valuable piece of context for whoever compares this study to the next one.
Clinical History and Demographics
The Detailed Content Outline names the elements: age, race, sex, smoking history, medication, clinical indication. These are not clerical fields — they are inputs to the predicted value, and an error in them produces a wrong interpretation from perfectly good measurements.
| Field | How an Error Propagates |
|---|---|
| Standing height | The dominant predictor. Entering 175 cm for a 165 cm patient inflates predicted values and manufactures an apparent restrictive defect. Measure without shoes at each visit; use arm span or ulnar length for spinal deformity and document the estimate. |
| Age | Predicted values decline with age; an error shifts percent predicted in the same direction. |
| Sex | Separate regression coefficients; a mis-entry produces a large, systematic error. |
| Race / ethnicity | ATS/ERS 2022 recommends race-neutral equations such as GLI Global. Whichever the laboratory uses must be stated on the report, because switching sets changes every percent predicted. |
| Smoking history (pack-years) | Does not enter the equation but drives interpretation, and directly affects DLCO through carboxyhemoglobin. |
| Medication and last dose | Determines whether a study is a true baseline; a bronchodilator taken two hours before "pre-bronchodilator" spirometry invalidates the reversibility assessment. |
| Clinical indication | Determines which tests are selected and establishes medical necessity for billing. |
Example: 1.5 packs per day for 24 years = 36 pack-years.
The Body Habitus Equipment Itself Needs Quality Control
Domain I lists body habitus equipment (stadiometer, body weight scale, caliper) as equipment category 1 under set-up/maintain/calibrate, troubleshoot, and perform quality control. It is the only "instrument" in the laboratory whose output is a predicted value rather than a measured one, and it is almost never on a QC schedule.
- Stadiometer: verify against a rigid reference rod or a marked wall standard on a documented schedule. Check that the headpiece slides freely and sits perpendicular, that the base is level and against a wall, and that a floor mat has not been added under it. Position the patient with heels together, buttocks and scapulae touching the upright, and the head in the Frankfort horizontal plane (the line from the external auditory meatus to the lower orbital margin parallel to the floor), then measure at the end of a quiet inspiration.
- Weight scale: verify with certified test weights spanning the clinical range, and re-verify after any move — a scale relocated onto carpet or an uneven floor reads incorrectly.
- Caliper / segmental measures: used for arm span and ulnar length when standing height cannot be obtained. Document which estimate was used and the conversion applied, because switching between measured height and an estimate mid-series creates a false trend.
A stadiometer that reads 2 cm high produces a systematically inflated predicted value for every patient measured on it, which shows up as an apparent restrictive pattern across the whole laboratory rather than in one chart — the reason this belongs on a QC schedule rather than in a chart review.
Two habits prevent most demographic errors: verify identity and height at the point of care rather than trusting the order, and reconcile any implausible percent predicted against the raw numbers before the report leaves the laboratory. A percent predicted that jumps 20 points between visits with unchanged raw volumes is a demographic entry error until proven otherwise.
Laboratory Quality Management
Item 19 gives three examples — customer satisfaction, inventory control, standard operating procedures — which signals that the exam treats quality management as operational, not just statistical.
Standard Operating Procedures
A written procedure manual is required by accrediting bodies and should cover, for every test offered: indications and contraindications, patient preparation, step-by-step performance, acceptability and repeatability criteria, calculations and reference equations in use, reporting format, calibration and quality control schedules with acceptance limits, and corrective action pathways. Procedures require documented annual review, version control with effective dates, and a record that each technologist has read the current version. When a procedure changes, the change itself becomes a serial-testing confounder and should be dated in the record.
Inventory Control
Consumables are a quality issue, not merely a budget issue:
- Calibration and test gases — verify certified concentration, lot number, and expiration on receipt; a DLCO gas cylinder whose stated CO concentration is wrong biases every study run against it. Track cylinder pressure and rotate stock first-expired, first-out.
- Absorbents — Drierite and soda lime are replaced on hours of use, not appearance, because soda lime regenerates its indicator color at rest.
- Filters, mouthpieces, nose clips, electrodes — maintain par levels; an expired or dried-out ECG electrode is a data-quality failure.
- Reagents and quality control material — record lot numbers so that a shift on a Levey-Jennings chart can be correlated with a lot change.
Personnel and Competency
Track credentials, BLS status, annual competency assessment for each procedure performed, and biological quality control performance by technologist. Between-technologist variability is a real source of drift: if one technologist's biological control values run consistently lower, the issue is coaching technique, not the instrument.
Customer Satisfaction and Service Metrics
The outline explicitly includes satisfaction. Practical measures a laboratory tracks: turnaround time from test completion to signed report, appointment availability and wait time, cancellation and no-show rates, repeat-test rate driven by unacceptable quality, patient experience survey results, and referring-clinician feedback on report clarity. A rising repeat-test rate is simultaneously a satisfaction problem, a cost problem, and a technical quality signal.
Continuous Quality Improvement
Close the loop: define a metric, set a target, measure, intervene, re-measure. A laboratory with a Grade D-or-worse rate of 15% has an actionable coaching problem; one that never audits its grade distribution does not know whether it has one.
A patient's FEV1 was 2.40 L six months ago and is 2.22 L today, with both sessions graded A on the same device using the same reference equations. How should this be reported?
A patient's raw FVC and FEV1 are essentially unchanged from last year, yet percent predicted has risen by 18 points. What is the most likely explanation?
Under laboratory quality management, why must soda lime and desiccant absorbents be replaced according to documented hours of use rather than by inspecting their indicator color?
You've completed this section
Continue exploring other exams