The Six Safety Data Quality Attributes

Key Takeaways

  • Define each quality metric’s numerator, denominator, baseline, and period.

  • Field completeness and event capture are different.

  • Faster data are not automatically more accurate or complete.

  • Quality limitations should change the interpretation or method, not merely appear in a footnote.

Last updated: October 2026

The Six Safety Data Quality Attributes

Quality is fitness for the decision

A large database can be unsuitable for a decision if locations are wrong, injuries are inconsistently coded, or reports arrive too late. Quality management asks whether the available information is usable for the specific analysis and how shortcomings influence conclusions. A complete statewide total might support a trend but be insufficient for an intersection diagnosis if location precision is poor.

NHTSA's six attributes are timeliness, accuracy, completeness, uniformity, integration, and accessibility. Assess them across the relevant systems, not only police crash reports. The same roadway inventory can be accurate in its dimensions but incomplete on local roads. NHTSA traffic-records quality framework

Timeliness: information arrives when needed

Timeliness concerns the interval between an event or change and its availability for use. Example measures include median report-processing days or the percentage available within a stated target. Define the starting and ending timestamps: crash occurrence to agency receipt differs from receipt to validated publication.

Long delays can prevent timely diagnosis or make a current dashboard appear artificially low. Faster entry alone is not enough if reports remain in a validation queue. Inspect the entire workflow and compare periods with the same completeness cutoff. Recent counts that are still accumulating should be labeled preliminary.

Accuracy: values represent the intended condition

Accuracy concerns correctness. Compare coordinates, maneuver codes, dates, volumes, or injury classifications with suitable reference information. Automated validity checks catch impossible values but cannot prove that a plausible value is correct. A location inside the state may still be assigned to the wrong road.

Measures can include error proportions from an audited sample or agreement with a verified reference. Define the reference and sampling method. Police suspected injury status and a clinical diagnosis are different measures; disagreement can reveal limits of roadside assessment, but the clinical record also has its own population and linkage constraints. Avoid assuming every difference is an officer error.

Completeness: needed events and fields are present

Completeness has event and field dimensions. A report may have every required field while the system misses many reportable crashes. Alternatively, a system may capture events but leave critical variables unknown. Measure event capture against an appropriate reference and field completeness against the required or applicable records.

Suppose hospitals indicate injuries on a corridor with no matching crash records. Investigate reportability, dates, location, linkage, and system capture before concluding the repository missed every event. The discrepancy is evidence for a completeness review, not an exact count of unreported crashes without reconciliation.

Uniformity: definitions are comparable

Uniformity concerns consistent definitions and formats across agencies or periods. Standardized variables support comparable analysis. NHTSA's MMUCC is a model guideline for crash data, not a guarantee that every state uses identical forms or implements every element.

Review changes in injury definitions, property-damage thresholds, mode categories, and geographic assignment. An apparent drop in reported PDO crashes can follow a higher reporting threshold instead of a safety improvement. Mapping old to new categories can help, but may lose detail; document what can and cannot be compared.

Integration: sources can be connected appropriately

Integration concerns linking relevant systems to support analysis. Crash, roadway, exposure, driver, vehicle, citation, and health records can answer more together than individually. Measures might include the percentage of records with usable location links or a validated match rate under a defined method.

More links are not always better. Incorrect matches can introduce false relationships, and missing links may not be random across populations. Review match quality, duplicates, coverage, and authorization. A common identifier or linear reference helps only if it is stable and correctly applied.

Accessibility: authorized users can use the data

Accessibility concerns timely, usable access for legitimate analytical needs. It includes documentation, tools, formats, training, and a clear request process. A downloadable file without definitions can be technically available but functionally difficult to use. A detailed database accessible to only one analyst can create a delivery bottleneck.

Protect privacy and restricted information while supplying appropriate access. Public dashboards can show aggregated outcomes; authorized analysts may need more detail under controlled conditions. Accessibility is not equivalent to publishing every identifiable record without limits.

Interpret measures together

Consider an electronic reporting upgrade. Median processing time falls from 30 to 10 days in an illustrative comparison. That is a timeliness improvement, but verify whether mandatory-field completeness, coding accuracy, geolocation, and coverage also improved. Faster incorrect data can harm screening decisions.

Define numerator, denominator, baseline, target, and period for each measure. “Data quality improved 20%” is ambiguous without specifying whether errors fell, completeness rose, or processing time changed. Distinguish percentage-point changes from relative percentage changes. For example, completeness increasing from 80% to 90% is 10 percentage points, or 12.5% relative to the initial value.

Quality findings should change the analysis

  • Incorrect locations constrain site-level screening.
  • Changed definitions constrain trend comparisons.
  • Missing exposure constrains rate interpretation.
  • Delayed reports constrain recent-period completeness.

Use quality findings in the safety decision

If crash locations are uncertain, use a screening method and spatial scale appropriate to that precision or prioritize location improvement. If injury coding changed, avoid claiming a continuous serious-injury trend without reconciliation. If exposure is missing, do not present a frequency comparison as a rate comparison.

Document limitations alongside results and identify a practical improvement. A useful quality statement explains how the defect affects the decision, not only that a field is missing. This connects quality management directly to better programs, projects, and investment priorities.

Test Your Knowledge

A reporting upgrade cuts processing delay but leaves locations incorrect. Which conclusion is sound?

A

All quality dimensions improved

B

Crash risk necessarily fell

C

Timeliness improved; accuracy still needs attention

D

No measure can be used

Test Your Knowledge

Completeness rises from 80% to 90%. What is the change?

A

90% relative improvement

B

10 percentage points, or 12.5% relative to the initial value

C

Exactly 10% relative improvement

D

No improvement

Sections you finish are checked off in the contents.