Planning Evaluations and Collecting Credible Evidence

Key Takeaways

  • Plan the question, outcomes, comparison, data, and responsibilities before treatment.

  • Use compatible exposure, crash definitions, periods, and location identifiers.

  • Document actual implementation and concurrent changes; package effects are not component effects.

  • Report uncertainty, adverse effects, deviations, and lessons for future decisions.

Last updated: October 2026

Planning Evaluations and Collecting Credible Evidence

Design the evaluation before implementation

A good evaluation begins while a project or program is being planned. Decide what change is expected, why it should occur, how it will be measured, and what comparison will help attribute it to the intervention. Waiting until the project is complete can leave missing baseline data, no suitable comparison group, or ambiguous implementation dates. The FHWA CMF development guide provides a foundation for matching study design to the safety question.

Write an explicit question: “Does this crossing treatment reduce pedestrian injury crashes at comparable uncontrolled crossings?” is more useful than “Did safety improve?” Specify the treatment, users, crash types, severity, sites, period, and intended mechanism. A program aimed at seat-belt use may initially measure observed use, while its long-term safety objective concerns occupant injuries. Those outcomes are related but not interchangeable.

Distinguish a project-level treatment effect from a broader program's delivery performance. Evaluating a portfolio can reveal whether agencies delivered projects on time and reached priority networks, but identifying one treatment's causal effect may require a separate study. Define each purpose rather than forcing one dashboard to answer every question.

Collect a stable baseline and comparison

Obtain crash records, exposure data, roadway and traffic-control characteristics, and contextual information for the required periods. Preserve definitions and identifiers so that records can be consistently located and classified. Verify crash severity, geocoding, study boundaries, reporting thresholds, and data completeness. Document changes in police reporting or injury coding that could create an apparent trend.

Exposure should match the outcome: entering vehicles for an intersection, vehicle-miles for a segment, and relevant pedestrian or bicycle activity where available. Collect seasonal and time-of-day patterns when they affect the mechanism. A one-day fair-weather pedestrian count may poorly represent annual activity near a school. If exposure is unavailable, describe the limitation rather than manufacturing an exact risk rate.

Select untreated comparison sites before learning which ones produce the preferred result. Check their treatment histories and changes in traffic, land use, or reporting. For EB studies, verify SPF applicability, calibration, variables, and prediction periods. Keep treatment and comparison data in compatible units; annual predictions cannot be combined directly with multi-year observed totals without conversion.

Record what was actually implemented

Plans describe intended treatment; evaluations need the delivered treatment. Record installation dates, locations, dimensions, operational settings, maintenance interruptions, and deviations from design. A campaign's scheduled media plan is not the same as measured message exposure. A camera program's authorized extent is not the same as its operating hours or actual enforcement locations.

For a corridor redesign, record whether speed limits, signals, crossings, lighting, and enforcement changed together. If the package's effect is evaluated, label it as a package effect. Without a study design that separates components, do not assign the entire reduction to one feature or claim independent CMFs for every component. Treatment fidelity also matters: an installed but nonfunctioning beacon does not provide the intended intervention.

Separate construction periods and transition effects from the intended stable after period when the study method warrants it. Construction can alter exposure, routes, and reporting. A temporary decrease caused by closing the road should not be attributed to the permanent safety design. State exclusion decisions in advance or justify them transparently.

Choose periods and measures with enough information

Fatal and serious-injury events are relatively rare at individual locations. Very short studies may have too little information to distinguish an effect from random variation. Larger samples, multiple appropriate sites, or longer periods can improve precision, but longer periods can also introduce more background changes. There is no universal rule that every evaluation requires exactly three before years and three after years.

Use intermediate measures where they are relevant, while recognizing their limits. Speeds, yielding, conflicts, seat-belt use, and emergency response intervals can test parts of a causal pathway. A reduction in measured conflicts supports a behavioral observation; it does not by itself establish a particular fatal-crash CMF. Ensure the measurement protocol is consistent, observers or devices are trained and checked, and the metric's definition is recorded.

Define success using both a meaningful magnitude and uncertainty, rather than declaring success after any decrease. For instance, an estimated small reduction with a very wide interval may justify further monitoring rather than a definitive effectiveness claim. Record adverse effects, other crash categories, nearby displacement, and impacts on different users.

Preserve the evidence before construction

  • Define the primary question, outcome, and comparison.
  • Collect compatible baseline crashes, exposure, and context.
  • Record delivered treatment, dates, and concurrent changes.
  • Plan analysis, review, privacy, and reporting responsibilities.

Manage analysis, transparency, and learning

Assign responsibility for data collection, analysis, independent review, and publication. Budget for the evaluation, including access agreements and privacy safeguards. Set a schedule that allows complete after-period data to become available. Do not confuse a project completion date with the date when injury outcomes can be reliably evaluated.

Predefine key analysis choices to reduce selective reporting: primary outcomes, sites, periods, comparison method, and handling of missing data. Explain deviations. An evaluation that tests many outcomes and reports only the most favorable one can mislead even when individual calculations are correct. Preserve reproducible records and calculations without releasing personally identifiable information.

Report estimates and limitations in language the decision maker can use. If a study finds improved yielding but inconclusive injury effects, present both results. If the package underperforms, distinguish a weak underlying intervention from incomplete delivery or changed conditions. Feed the findings into treatment selection, design standards where appropriate, training, and future program priorities. Evaluation is most useful when it changes the next decision rather than merely closing a grant file.

Test Your Knowledge

Why should comparison sites and primary outcomes be specified before examining results?

A

It reduces selective choices that could exaggerate apparent benefits

B

It makes all studies randomized

C

It eliminates the need for exposure data

D

It guarantees statistical significance

Test Your Knowledge

A redesign adds lighting, changes signals, and reduces speeds simultaneously. What can a study of the combined project generally estimate?

A

The combined package effect, unless the design can separately identify component effects

B

No useful outcome at all

C

An independent CMF for each component without additional analysis

D

Only the lighting effect

Sections you finish are checked off in the contents.