Collecting Quantitative and Qualitative Safety Data

Key Takeaways

  • Quantitative and qualitative information complement each other but support different claims.

  • Define exposure, observation opportunities, and ambiguous cases before field collection.

  • Record linkage needs validation and protection against duplicate encounters.

  • Document missing values, revisions, and conditions so the analysis can be repeated.

Last updated: October 2026

Collecting Quantitative and Qualitative Safety Data

Select measures before collecting information

Quantitative data describe counts, magnitudes, proportions, or measured values: crashes, injuries, speeds, volumes, response times, or observed yielding. Qualitative information describes meaning and experience through narratives, interviews, open comments, and field notes. Both can help explain a problem, but they support different claims.

A coded category can be analyzed quantitatively, while the original narrative remains qualitative. Converting comments into categories requires documented definitions and consistent coding. Counting comments does not automatically measure population prevalence. The sampling frame, collection method, and reporting opportunity still matter.

Write the question, unit of observation, population, period, and intended analysis before gathering data. For a yielding study, define who has the right-of-way, what constitutes an opportunity, what counts as yielding, and how ambiguous events are handled. Without those definitions, two observers can produce incompatible rates while watching the same crossing.

Plan exposure and operating observations

Traffic counts should capture relevant modes and movements. A total AADT may support a segment model but cannot describe turning conflicts or pedestrian crossing opportunities by itself. Short-duration counts need consideration of time, day, season, weather, and unusual events. Record the observation conditions and any adjustment method.

Speed observations should describe the location, direction, collection method, and traffic state. A free-flow speed sample answers a different question from peak-hour congested operation. Report the distribution or relevant proportions rather than selecting one percentile without explaining why. Avoid conflating posted, design, and observed speeds.

For conflict or near-miss observations, define the interaction and metric consistently. Observer judgment, camera position, occlusion, and algorithm settings can affect the result. A surrogate measure may help identify a mechanism or intermediate change, but it is not automatically a validated estimate of future fatal crashes.

Collect narratives and community experience

Interviews can reveal why users select a route, avoid a crossing, misunderstand a device, or encounter maintenance barriers. Use questions that invite description rather than steering respondents toward the team's favored treatment. Record whose perspectives were sought and whose may be missing.

Accessible participation matters. Offer usable formats, appropriate languages, locations and times, and ways to contribute without attending one formal meeting. People who avoid a road because of risk may not appear in on-site user samples. A study of current users alone can miss suppressed demand and barriers to participation.

Field notes should distinguish observation from interpretation. “Vehicles parked within the sightline during the evening observation” is an observation. “Parking caused every crash” is a causal conclusion requiring further evidence. Photographs, sketches, time records, and measurement methods can make an observation reproducible.

Combine records carefully

Joining crash and roadway files requires consistent identifiers, location referencing, and dates. A geocoded crash can be placed on the wrong segment if the map changed or coordinates were imprecise. Check direction, intersection assignment, route version, and whether a record appears more than once.

Joining police and health records can reveal injuries that roadside classifications miss. Deterministic linkage uses suitable matching identifiers or rules; probabilistic linkage uses patterns of agreement and disagreement where an exact common identifier is absent. Both require validation of false matches and missed matches. One person can have several clinical encounters, so repeated admissions should not automatically become separate crash events.

NHTSA supported CODES from 1992 through 2013, after which responsibility continued at the state level. CODES is a useful historical model of linked crash and health data, not a current uniform national repository that automatically links every state. NHTSA CODES description

Manage definitions, privacy, and revisions

A data dictionary should define variables, categories, units, valid values, and missing-value codes. Unknown, not applicable, and not collected are different states. Treating all three as zero creates false absence. Preserve original records and document derived fields so another analyst can reproduce the result.

Use only the information needed and authorized for the analysis. Person-level linkage may require agreements and protected access, while public reporting may use aggregated or de-identified results. Accessibility for authorized users does not mean unlimited public release of identifiable records. Manage retention, roles, and disclosure according to the governing requirements.

Record revisions and source dates. A preliminary fatality estimate can later change; an inventory can be updated; a coding definition can be revised. A time series should identify these changes rather than presenting them as a continuous, unchanging measurement system.

Specify the observation protocol

  • Define the event or opportunity being counted.
  • Select relevant users, periods, and operating conditions.
  • Record ambiguous cases, missing information, and revisions.
  • Validate joins and prevent duplicate events or encounters.

Work a collection plan

Imagine recurring near misses where a bicycle route crosses a driveway beside a bus stop. Collect turning movements, bicycle and pedestrian counts, bus activity, vehicle speeds, sight obstructions, and systematically defined conflicts. Obtain crash narratives and maintenance information. Ask users about confusing transitions and route choices.

Select periods covering relevant commuting and transit conditions. Use consistent camera or observer methods and note occlusion. Review whether the same interaction is counted by several observers or reported repeatedly. Compare patterns among sources: a narrative can suggest a mechanism, observations can test how often it occurs, and inventory can identify a physical condition.

Finally, report what the collection supports. The data may justify further investigation or a candidate treatment, but a short sample may not estimate annual crash risk reliably. State limits, needed follow-up, and how the same measures will be collected after implementation. This creates a usable baseline and prevents later evaluation from comparing incompatible observation methods.

Test Your Knowledge

Which is a defensible qualitative observation?

A

Vehicles blocked the observed sightline during the evening visit

B

Every near miss predicts a fatality

C

Parking caused every recorded crash

D

The entire population avoids the road because two people said so

Test Your Knowledge

Why validate a probabilistic linkage?

A

Identifiers are never useful

B

Every approximate match is exact

C

It eliminates the need for privacy controls

D

False and missed matches can distort injury findings

Sections you finish are checked off in the contents.