5.5 Practice Evaluation, Outcome Monitoring, and Program Research
Key Takeaways
- Single-system designs (AB, ABA, ABAB, multiple baseline) use repeated measurement across phases to evaluate individual client progress.
- Goal Attainment Scaling (GAS) individualizes outcomes on a -2 to +2 scale and supports aggregated evaluation across clients.
- Feedback-Informed Treatment (FIT) uses session-by-session measures (ORS/SRS, OQ-45) to detect deterioration early and improve outcomes.
- Agency program evaluation methods include needs assessment, formative and summative evaluation, process evaluation, cost-effectiveness and cost-benefit analysis, and outcomes assessment.
- Reliability (consistency) and validity (accuracy) are required properties of any standardized instrument used in practice evaluation.
Practice Evaluation, Outcome Monitoring, and Program Research
Accountability in social work practice requires that social workers systematically evaluate the effectiveness of their interventions and the programs delivering them. The ASWB Masters blueprint (IIIC Assessment Practices and the broader IIIC/IIID boundary) tests four applied knowledge statements: techniques for evaluating intervention progress, instruments for evaluating practice, methods for evaluating agency programs, and basic/applied research design. Mastery of this content safeguards clients from ineffective care and ensures compliance with NASW Standard 1.04(c) - social workers should monitor practice outcomes and revise practice accordingly.
Techniques for Evaluating Progress and Effectiveness of Interventions
Single-system (single-subject) design is the foundational method for evaluating individual client progress. The social worker repeatedly measures a target behavior across phases:
| Design | Structure | Use Case |
|---|---|---|
| AB | Baseline (A), then intervention (B) | Simplest; cannot fully rule out history |
| ABA / ABAB | Withdraw intervention to confirm effect | Behaviorally focused problems |
| Multiple baseline | Same intervention across behaviors/settings/subjects | Demonstrates functional relationship |
Goal Attainment Scaling (GAS): Client and social worker collaboratively define expected outcomes on a -2 to +2 scale (much less than expected, less than expected, expected, more than expected, much more than expected). GAS supports individualized measurement while permitting aggregation.
Routine outcome monitoring: Standardized measures administered at every session (or regular intervals) to detect deterioration early. Common tools include the PHQ-9 (depression), GAD-7 (anxiety), ORS/SRS (Outcome Rating Scale / Session Rating Scale - Miller, Duncan, Sparks), and OQ-45 (OQ-45.2). Feedback-Informed Treatment (FIT) cuts deterioration rates and improves outcomes by alerting clinicians when progress stalls.
Progress notes and service plans should reference measurable objectives and document client-reported and observed changes. Social workers should distinguish status measures (point-in-time symptoms) from change measures (difference from baseline).
Methods and Instruments Used to Evaluate Social Work Practice
Practice evaluation draws from both quantitative and qualitative traditions:
- Standardized instruments: Reliable and valid tools such as the Beck Depression Inventory-II, Child Behavior Checklist, Connor-Davidson Resilience Scale, and PTSD Checklist (PCL-5). Social workers must select instruments with demonstrated reliability (consistency) and validity (accuracy) for the relevant population.
- Qualitative methods: Semi-structured interviews, focus groups, and case studies capture client experience and context. Useful for complex interventions where standardized measures miss meaning.
- Mixed methods: Combine quantitative outcome scores with qualitative client narratives.
- Direct observation: Behavioral coding in home, school, or clinic settings.
- Client satisfaction surveys: Useful for program feedback but prone to social desirability bias and not a substitute for outcome measurement.
Methods for Evaluating Agency Programs
Agency-level evaluation answers broader questions: Does the program achieve its mission? Is it cost-effective? Should funders continue support?
| Method | Question Addressed | Example |
|---|---|---|
| Needs assessment | What does the community require? | Survey of unhoused youth service gaps |
| Formative evaluation | How can the program improve during implementation? | Mid-cycle staff feedback on intake process |
| Summative evaluation | Did the program achieve its goals? | End-of-year outcomes report |
| Process evaluation | Is the program being delivered as designed? | Dosage, fidelity checks |
| Cost-effectiveness analysis | Which program achieves outcomes at lower cost? | Compare two foster-care models per placement stability outcome |
| Cost-benefit analysis | Do monetary benefits exceed costs? | ROI of supported employment program |
| Outcomes assessment | What changed for participants? | Pre/post measures, comparison groups |
Logic models connect inputs (resources), activities (services), outputs (units delivered), and outcomes (short-term, intermediate, long-term). Social workers should be able to read and construct a logic model for grant proposals and program evaluation reports.
Basic and Applied Research Design
Social workers must critically appraise research and apply findings to practice (NASW Standard 4.01). Research literacy requires understanding:
- Basic vs. applied research: Basic research tests theory; applied research solves practice problems.
- Quantitative design: Experimental (randomized controlled trial), quasi-experimental (nonequivalent control group, time series), and non-experimental (correlational, descriptive) designs.
- Qualitative design: Phenomenology, grounded theory, ethnography, case study.
- Data collection: Surveys, interviews, focus groups, record review, observation, administrative data.
- Sampling: Probability (simple random, stratified, cluster) vs. non-probability (convenience, purposive, snowball). Probability sampling supports generalizability.
- Reliability: Consistency of measurement - types include test-retest, inter-rater, internal consistency (Cronbach's alpha).
- Validity: Accuracy of measurement - content validity, criterion validity (concurrent/predictive), construct validity (convergent/discriminant).
- Statistical literacy: Mean, median, mode, standard deviation, correlation (Pearson r), p-value, confidence interval, effect size. A statistically significant result (p < .05) is not necessarily clinically significant; effect size (e.g., Cohen's d) communicates practical importance.
- Evidence-based practice hierarchy: Systematic reviews and meta-analyses (highest), RCTs, quasi-experimental, cohort/case-control, case series, expert opinion. Social workers select the best available evidence, not automatically the highest tier - client preferences, clinical expertise, and context moderate.
Ethical and Cultural Considerations in Evaluation
- Informed consent for evaluation: Clients must know they are participating in outcome monitoring, with privacy protections.
- Cultural validity: Standardized instruments may lack norms for diverse populations. Social workers verify cultural appropriateness and use culturally validated translations.
- Survivorship bias and selection bias: Clients who drop out may differ systematically from completers; intent-to-treat analysis mitigates this.
- Conflicts of interest: Program evaluators should be independent from program staff when possible.
Clinical Case Study: Building an Evaluation Plan
A school-based social worker implements a 12-week trauma-focused CBT group for adolescents exposed to community violence. To evaluate effectiveness, she uses a single-system AB design with multiple baselines across three participants. Each adolescent completes the UCLA PTSD Reaction Index at baseline, weekly during treatment, and 1 month post-treatment. She supplements quantitative data with a focus group at program end. For agency-level evaluation, she builds a logic model linking inputs (clinician time, training, materials), activities (weekly group sessions, parent meetings), outputs (number of sessions delivered, attendance), and outcomes (reduced PTSD scores at post-treatment, improved school attendance). She compares per-student cost to alternative referral options in a cost-effectiveness summary for the school board, applying intent-to-treat analysis to include two adolescents who dropped out.
A social worker implements a new intervention with a client and wants to evaluate its effectiveness using the simplest single-system design. Which design is most appropriate?
An agency wants to determine whether its supported employment program for clients with severe mental illness produces monetary benefits that exceed program costs. Which evaluation method is most appropriate?
A social worker selects a standardized depression measure for use with adolescent clients. Which properties must the instrument demonstrate for the relevant population to support valid interpretation?