4.3 Evaluation Designs, Measurement Tools & Data Validity
Key Takeaways
Evaluation research designs exist on a hierarchy of causal rigor: Experimental (RCTs), Quasi-Experimental (comparison groups and interrupted time series), and Non-Experimental (single-group pre/post and post-test only).
Threats to internal validity—including history, maturation, testing effects, instrumentation, regression to the mean, and attrition—can produce deceptive changes that mimic or mask true program impact.
A defensible measurement instrument must exhibit high psychometric reliability (consistency across test-retest, internal consistency, and inter-rater agreement) and validity (face, construct, and predictive accuracy).
Response-shift bias occurs when participants' internal calibration changes as they gain knowledge; retrospective pretests (post-then-pre) effectively mitigate this distortion.
Utilizing standardized surveillance instruments (YRBSS, MTF, NSDUH) and federal GPRA measures allows local prevention coalitions to benchmark results against state and national datasets.
Evaluation Designs, Measurement Tools & Data Validity
Core Principle: Demonstrating a correlation between program participation and positive community changes does not prove causation. Rigorous evaluation research designs and validated psychometric instruments are essential to eliminate competing rival explanations and confirm that the prevention intervention was the actual driver of community improvement.
In prevention science, demonstrating effectiveness requires answering a fundamental scientific question: "Would these positive changes have occurred even if our coalition had done nothing?" Adolescents naturally mature, media headlines influence public opinions, and law enforcement practices fluctuate. If a community observes a 5% decline in underage drinking following an educational campaign, that decline could be driven by the campaign—or it could be the result of a statewide alcohol excise tax increase, natural demographic shifts, or economic inflation.
To rule out rival explanations (threats to internal validity), prevention specialists must understand the strengths and limitations of various evaluation research designs, select psychometrically sound measurement instruments, and adhere to standardized federal reporting standards.
The Spectrum of Evaluation Research Designs
Evaluation designs vary in their ability to establish causal inference. In prevention practice, designs are grouped into three primary classifications:
1. Experimental Designs (Randomized Controlled Trials - RCTs)
Experimental designs represent the gold standard for establishing definitive causality. In an RCT, eligible individuals, classrooms, or community units are randomly assigned either to the intervention group (which receives the prevention program) or the control group (which receives standard programming or a placebo):
Intervention group (R): O1 ---> X ---> O2
Control group (R): O1 ---------> O2
(R = random assignment; O = observation/measurement; X = intervention)
- Causal Rigor: Random assignment mathematically equalizes known and unknown confounding variables across both groups before the intervention begins. Any statistically significant post-test differences (O2) between the groups can be confidently attributed to the intervention (X).
- Community Prevention Realities: While common in academic university trials, RCTs face severe ethical and practical barriers in community coalitions. Withholding an evidence-based intervention from high-risk adolescents in a control school raises ethical concerns. Furthermore, community-wide environmental strategies (such as billboard bans or alcohol tax increases) cannot be randomly assigned to individuals within the same city due to unavoidable contamination and spillover effects.
2. Quasi-Experimental Designs
Quasi-experimental designs provide substantial causal rigor without requiring random assignment. Instead, researchers use non-random comparison groups or structured longitudinal measurement:
- Pretest-Posttest with Non-Equivalent Comparison Group:
- Compares a community or school receiving the program against a carefully matched comparison community or school with similar demographic, socioeconomic, and baseline substance profiles.
Intervention group: O1 ---> X ---> O2 Comparison group: O1 ---------> O2 (not randomly assigned)- Key Challenge: Selection bias. Because groups were not randomly assigned, subtle unmeasured differences (e.g., higher parent involvement in the volunteer school) may influence outcomes.
- Interrupted Time Series Design:
- Collects multiple sequential observations over an extended timeframe prior to the intervention, followed by multiple sequential observations post-intervention:
O1 -> O2 -> O3 -> O4 [X] O5 -> O6 -> O7 -> O8- Ideal Use in Prevention: Evaluating environmental and policy interventions. For example, tracking monthly alcohol-related motor vehicle crashes for 24 months before and 24 months after the passage of a municipal social host ordinance. The multi-point baseline establishes the pre-existing trend, allowing evaluators to determine whether the policy caused an immediate step-change or a slope divergence, effectively ruling out natural maturation and seasonal spikes.
3. Non-Experimental Designs
Non-experimental designs lack a control or comparison group, focusing exclusively on the group receiving the intervention. While common due to budget and resource constraints, they possess the weakest causal validity:
- One-Group Pretest-Posttest Design (O1 → X → O2):
- Measures participants before the program and immediately following completion.
- Limitation: Highly vulnerable to internal validity threats (e.g., Did youth improve because of the curriculum, or because they simply matured over the school year?).
- Post-Test Only Design (X → O1):
- Collects data only after program completion.
- Limitation: Scientifically incapable of measuring change from baseline. It only measures post-program status.
- Retrospective Pretest Design (Post-then-Pre):
- Administered to participants at the conclusion of the program. For each item, participants rate their current knowledge/skill level and simultaneously rate what they retrospectively realize their knowledge/skill level was before participating.
- Key Advantage: Eliminates response-shift bias. In traditional pretests, naive participants often suffer from the Dunning-Kruger effect—overestimating their baseline competence because they "do not know what they do not know." Once educated, their internal metric recalibrates. A retrospective pretest captures this cognitive shift accurately.
Threats to Internal Validity: The Evaluator's Checklist
When evaluating a prevention initiative, a certified prevention specialist must systematically examine seven classic threats to internal validity that can distort findings:
| Threat to Validity | Operational Definition in Prevention | Real-World Community Example | Mitigation Strategy |
|---|---|---|---|
| History | External events occurring between pre- and post-testing that influence participant behaviors independent of the intervention. | A prominent national celebrity dies of an accidental fentanyl overdose during a coalition's high school opioid campaign, causing a surge in adolescent perception of risk. | Utilize a matched comparison group; both communities experience the external historical event, isolating program impact. |
| Maturation | Natural biological, psychological, or cognitive growth occurring within participants simply as time passes. | 6th-grade students complete an anger management program over an 18-month window; their executive functioning and impulse control naturally improve with biological brain development. | Use a control or comparison group of the same chronological age cohort to benchmark natural developmental trends. |
| Testing Effects (Sensitization) | The process of taking a pretest sensitizes participants, altering their performance or awareness on the post-test. | Completing a pretest on alcohol pharmacology alerts youth to specific facts, causing them to research or pay closer attention even without curriculum impact. | Use Solomon Four-Group designs, retrospective pretests, or post-test only comparison designs. |
| Instrumentation | Unintended changes in measurement instruments, survey wording, data collection administration, or observer scoring standards between pre- and post-tests. | Moving from anonymous paper pencil surveys in Year 1 to identified online school tablet surveys in Year 2 causes youth to under-report illicit drug use out of privacy fears. | Standardize survey administration protocols, maintain identical question phrasing, and ensure consistent survey modalities. |
| Statistical Regression to the Mean | The mathematical tendency for extreme initial scores to naturally gravitate closer to the statistical population average upon re-measurement. | A coalition selects students with the highest suspension rates for an intensive intervention; their disciplinary incidents decline partly due to statistical regression. | Avoid selecting participants solely on the basis of extreme, single-point outlier scores; utilize randomized comparison controls. |
| Attrition (Experimental Mortality) | The non-random loss of participants between pretest and post-test data collection points. | The highest-risk adolescents in a prevention program drop out or move away before the post-test, leaving only compliant, low-risk youth in the sample and creating a false appearance of program success. | Track attrition demographics carefully, conduct intent-to-treat analyses, and implement retention incentives. |
| Selection Bias | Pre-existing systematic differences between intervention and comparison groups prior to program delivery. | Comparing youth who voluntarily enrolled in an after-school leadership program against the general student body; the leadership group was already more motivated and less prone to substance use. | Use random assignment whenever possible; when using comparison groups, conduct rigorous statistical propensity score matching. |
Psychometric Properties: Reliability and Validity
Evaluation instruments must demonstrate scientific credibility through established psychometric properties. A measurement tool must be both reliable (consistent) and valid (truthful).
Reliability (Measurement Consistency)
Reliability reflects the degree to which an instrument produces consistent, stable results across repeated administrations under identical conditions:
- Test-Retest Reliability: Stability of scores when the same individual completes the identical instrument at two different points in time (assuming the underlying trait has not changed).
- Internal Consistency: The degree to which multiple survey items intended to measure the same underlying psychological construct correlate with one another. Typically quantified using Cronbach's alpha (α), where values of 0.70 or higher indicate acceptable internal reliability (e.g., an 8-item scale measuring "Perceived Risk of Cannabis Use").
- Inter-Rater Reliability: The degree of agreement between two or more independent observers evaluating the same behavior or program session. Commonly quantified using Cohen's Kappa (κ), where values above 0.75 indicate strong inter-rater concordance (essential for fidelity monitoring checklists).
Validity (Measurement Accuracy)
Validity reflects whether an instrument actually measures what it purports to measure:
- Face Validity: A superficial, subjective judgment regarding whether survey questions appear relevant and transparent to respondents and stakeholders.
- Construct Validity: The degree to which an instrument truly captures the abstract psychological or social construct it claims to assess (e.g., does a survey scale truly measure "refusal self-efficacy," or is it merely measuring social desirability?). Includes convergent validity (correlating with similar constructs) and discriminant validity (not correlating with unrelated constructs).
- Criterion-Related Validity: The extent to which an instrument's scores correlate with or predict an external benchmark:
- Concurrent Validity: The instrument correlates strongly with an existing, validated gold-standard measure administered at the same time.
- Predictive Validity: The instrument accurately forecasts future behavior (e.g., low scores on a middle school "Perception of Peer Disapproval" scale successfully predict past 30-day binge drinking in high school).
Standardized Surveillance Instruments & Federal GPRA Measures
Rather than creating custom, unvalidated questionnaires from scratch, prevention specialists rely on standardized national surveillance instruments. Utilizing established tools ensures high psychometric reliability, protects against amateur survey wording flaws, and allows local coalitions to benchmark their community data directly against county, state, and national prevalence rates.
Leading Surveillance Instruments
- Youth Risk Behavior Surveillance System (YRBSS): CDC national school-based survey monitoring health-risk behaviors, dietary habits, physical activity, and substance use among 9th- through 12th-grade students.
- Monitoring the Future (MTF): National Institute on Drug Abuse (NIDA) longitudinal study tracking substance use trends, perceived risk, and disapproval among 8th, 10th, and 12th graders across the United States.
- National Survey on Drug Use and Health (NSDUH): SAMHSA household interview survey providing annual national and state estimates on tobacco, alcohol, illicit drug use, and mental health indicators across civilian populations aged 12 and older.
Government Performance and Results Act (GPRA) Measures
All recipients of federal prevention grant funding (such as SAMHSA Strategic Prevention Framework - Partnerships for Success grants) must collect and report standardized GPRA measures. These standardized indicators ensure accountability and allow federal oversight agencies to aggregate local grant performance nationwide. Core GPRA indicators evaluate:
- Past 30-Day Substance Consumption: Frequency and quantity of alcohol, cannabis, tobacco/nicotine, and illicit or non-medical prescription drug use.
- Perception of Risk / Harm: Perceived physical, legal, and social danger associated with daily or weekly substance use.
- Perception of Peer Disapproval: Degree to which youth believe their close friends would disapprove of them using substances.
- Perception of Parental Disapproval: Youth perception of whether their parents/caregivers would disapprove of their substance use.
A community coalition selects the 20 middle school students with the highest number of disciplinary referrals and substance-related school suspensions during the first semester for an intensive mentoring intervention. At the end of the second semester, disciplinary referrals among these 20 students declined by 40%. The coalition claims the mentoring program was an unqualified success. What primary threat to internal validity undermines this conclusion?
Instrumentation
Statistical regression to the mean
Testing effects
Contamination
A prevention coalition is evaluating the impact of a newly enacted countywide social host ordinance that holds property owners liable for underage drinking on their premises. To evaluate the policy, the coalition analyzes monthly juvenile alcohol-related citations and emergency department intoxication admissions for 24 months before the ordinance and 24 months after enactment. Which research design is the coalition utilizing?
Randomized controlled trial (RCT)
Post-test only comparison design
Interrupted time series design
One-group pretest-posttest design
Two independent trained evaluators observe a prevention specialist deliver a curriculum session. Both evaluators utilize a standardized 15-item fidelity rubric to rate facilitator adherence, resulting in a 92% concordance rate (Cohen's Kappa = 0.84). What psychometric property of the measurement system does this demonstrate?
High inter-rater reliability
Strong predictive validity
High test-retest reliability
Robust construct validity
Sections you finish are checked off in the contents.