3.8 Data Validation & Surveillance Program Evaluation
Key Takeaways
- Surveillance data validation compares reported findings against an independent re-review of source records to measure sensitivity, specificity, and positive predictive value of case finding.
- Under-reporting from missed cases is the most common validation failure and artificially lowers the standardized infection ratio, creating false reassurance.
- Denominator errors are as damaging as numerator errors, because an inflated device-day count lowers the rate without any change in patient care.
- CDC's framework for evaluating surveillance systems assesses simplicity, flexibility, data quality, acceptability, sensitivity, predictive value positive, representativeness, timeliness, and stability.
- The surveillance plan should be re-evaluated at least annually against the facility risk assessment, retiring measures that no longer inform action.
3.8 Data Validation & Surveillance Program Evaluation
Quick Answer: Validation asks are these numbers right? — usually by re-reviewing a sample of records against the reported result. Evaluation asks is this surveillance worth doing? — using CDC's attributes: simplicity, flexibility, data quality, acceptability, sensitivity, predictive value positive, representativeness, timeliness, and stability.
Surveillance data drive public reporting, CMS payment adjustments, resource allocation, and clinical practice change. The blueprint therefore lists both validate surveillance data and participate in the evaluation of surveillance plans and goal progress as distinct associate-level tasks.
Why Validation Matters More Than It Seems
A facility that misses half its CLABSIs reports an excellent standardized infection ratio. Leadership concludes the program is working, redirects the infection preventionist to other priorities, and patients continue to be harmed by a problem nobody is measuring. Under-reporting is dangerous precisely because it looks like success.
The reverse error also occurs. Over-reporting — counting contaminated blood cultures as CLABSIs, or counting community-onset infections as healthcare-associated — triggers investigations, punitive attention on units doing nothing wrong, and eventually staff disengagement from a system they believe is unfair.
Validation Methods
| Method | How it works | Strength |
|---|---|---|
| Internal re-review | A second reviewer independently applies the surveillance definition to a sample of records | Ongoing, inexpensive, builds reviewer skill |
| Targeted sampling | Deliberately pull high-yield records — all positive blood cultures, all patients on specified antimicrobials, all ICU patients with devices | Efficient at finding missed cases |
| Random sampling | Randomly select records regardless of risk | Unbiased estimate of overall accuracy |
| External validation | State health department or CMS-contracted reviewers audit against source records | Independent; carries regulatory consequence |
| Inter-rater reliability testing | Two or more reviewers apply the definition to the same set of cases and agreement is measured | Detects definitional drift between reviewers |
The performance metrics
Applying the standard 2×2 framework, with expert re-review as the reference standard:
- Sensitivity = of the true cases, the proportion the surveillance system found. Low sensitivity = missed cases.
- Specificity = of the true non-cases, the proportion correctly excluded.
- Positive predictive value = of the cases reported, the proportion that are truly cases. Low PPV = over-reporting.
Exam Tip: For HAI surveillance, sensitivity is the priority. A missed infection is an uninvestigated harm event; a falsely included case is corrected on review. Case-finding screens are therefore designed to be broad, with confirmation applied afterward.
Common Errors, by Where They Live
Numerator errors
| Error | Effect |
|---|---|
| Missed cases from incomplete case finding | Falsely low rate and SIR |
| Applying clinical diagnosis instead of the surveillance definition | Both directions; the physician's note is not the criterion |
| Misapplying the infection window or repeat infection timeframe | Duplicate or missing events |
| Wrong attribution — assigning an infection to the wrong unit or the wrong admission | Correct facility total, wrong unit blamed |
| Counting present-on-admission infections as healthcare-associated | Falsely high rate |
| Counting contaminated blood cultures as bloodstream infections | Falsely high rate |
Denominator errors
Denominator errors are underappreciated because they never appear in a case review.
| Error | Effect |
|---|---|
| Counting a patient with two lumens as two central lines | Inflated device days, falsely low rate |
| Counting the device on the day of removal inconsistently | Small systematic bias |
| Missing patient days for observation or short-stay patients | Falsely high rate |
| Electronic device-day extraction not matching the manual counting rule | Discontinuity when the method changes |
Exam Tip: If a unit's infection count is unchanged but its rate dropped sharply, suspect the denominator before celebrating.
Definitional drift
When NHSN revises a definition — a new VAE criterion, a changed infection window — historical comparison breaks. Any trend line crossing a definitional change must be annotated, and pre- and post-change data must not be pooled silently.
Evaluating the Surveillance Program Itself
CDC's framework for evaluating public health surveillance systems supplies the attributes to assess:
| Attribute | The question |
|---|---|
| Simplicity | Can it be operated without disproportionate effort? |
| Flexibility | Can it absorb a new organism or a new definition without redesign? |
| Data quality | Are fields complete and valid? |
| Acceptability | Do clinicians and units participate willingly? |
| Sensitivity | What proportion of true cases is detected? |
| Predictive value positive | What proportion of detected cases are real? |
| Representativeness | Does it describe the whole population at risk? |
| Timeliness | Do results arrive early enough to act on? |
| Stability | Is it reliable and available without frequent failure? |
Alongside these, evaluate usefulness: has any decision or practice actually changed because of this measure? Surveillance that no one acts on is data collection, not surveillance.
The Annual Cycle
The surveillance plan should be reviewed at least annually against the facility risk assessment, asking:
- Did we meet the goals set for the year, and can we show the trend?
- Do the measured processes and outcomes still match the facility's top risks, given new service lines, new populations, or new organisms?
- Which measures have been stable at goal long enough to move to periodic rather than continuous surveillance, freeing time for a rising risk?
- Are new regulatory or accreditation requirements captured?
- Is time being spent where the risk is, or where the habit is?
Retiring a measure is a legitimate outcome. Surveillance capacity is finite, and continuing to count something simply because it has always been counted is one of the most common ways an IPC program runs out of time for the risk that actually matters.
A validation review finds that the surveillance program identified 12 of 20 true CLABSIs confirmed by independent expert re-review. What does this indicate and what is the consequence?
A unit's CLABSI count is identical to last quarter, but its reported rate per 1,000 central line days fell by half. What should be investigated first?
Which attribute from CDC's surveillance system evaluation framework asks whether results arrive early enough for infection prevention action to be taken?
During the annual review of the surveillance plan, a measure has been stable at goal for three consecutive years and has prompted no practice change. What is an appropriate action?