12.5 Program Evaluation and Needs Assessment

Key Takeaways

  • Formative evaluation improves a program while it runs; summative evaluation judges whether it worked.
  • Process evaluation examines implementation fidelity; outcome evaluation examines participant change.
  • Cost-effectiveness compares cost per unit of outcome; cost-benefit converts both costs and outcomes to dollars.
  • Needs assessment methods include key informant interviews, community forums, surveys, rates-under-treatment, and social indicators, each with distinct biases.
  • Evaluation involving human subjects requires informed consent, voluntary participation, and confidentiality protections under NASW Standard 5.02.
Last updated: September 2026

The Vocabulary the Exam Tests

The applied knowledge statement names five things explicitly: needs assessment, formative assessment, summative assessment, cost-effectiveness analysis, cost-benefit analysis, and outcomes assessment. Each is a distinct concept, and items typically test the distinction.

TermQuestion it answersTiming
Needs assessmentWhat does this population need, and is a program warranted?Before the program
Formative evaluationHow can we improve this program while it operates?During; feedback for improvement
Process evaluationIs the program being delivered as designed, to whom, at what dose?During
Summative evaluationDid the program work, and should it continue?At the end or at a decision point
Outcome evaluationDid participants change?End of an episode or follow-up
Impact evaluationDid the change result from the program rather than something else?Requires a comparison condition
Cost-effectiveness analysisWhat is the cost per unit of outcome?Compares alternatives producing the same outcome
Cost-benefit analysisDo the monetized benefits exceed the monetized costs?Expresses both in dollars

Two distinctions students confuse most:

  • Formative versus summative. Formative feeds back into the program to improve it while it runs; summative renders a verdict for a funding or continuation decision. A mid-year staff debrief that changes the curriculum is formative; a final report determining whether the contract is renewed is summative.
  • Cost-effectiveness versus cost-benefit. Cost-effectiveness holds the outcome constant and compares cost — for example, cost per client housed across two programs. Cost-benefit converts the outcome itself into dollars — for example, jail days avoided valued at their cost — and can therefore compare programs with different outcomes. Cost-benefit is more powerful and more contestable, because monetizing human outcomes requires value judgments.

Needs Assessment Methods

MethodDescriptionStrengthsBiases and limits
Key informantInterviews with people who know the community or systemFast, inexpensive, rich contextReflects informants' vantage point; may miss the unserved
Community forumOpen public meetingBroad input, builds engagementAttendees are self-selected; the loudest voices dominate
Survey of residentsStructured sample surveyGeneralizable if the sample is soundCostly; low response rates; excludes people without stable contact
Rates under treatmentCounts of people currently servedEasy to obtain from existing dataSystematically misses people who never reached services
Social indicatorsCensus, vital statistics, administrative dataObjective, comparable across areasProxies rather than direct need; may lag
Focus groupsFacilitated small-group discussionDepth, surfaces unanticipated issuesNot generalizable; group effects shape responses

The most-tested limitation is the rates-under-treatment problem: measuring need by counting current service users guarantees that unmet need is invisible. A county that concludes demand is low because its clinic is not full has measured its own access barriers, not need.

Good practice triangulates — combining several methods so that each one's blind spot is covered by another — and distinguishes normative need (defined by experts), felt need (what people say they want), expressed need (demand actually presented, such as a waiting list), and comparative need (relative to a similar population elsewhere).

Evaluation Designs

DesignStructureCausal strength
Post-test onlyMeasure after the programWeakest; no baseline
Pretest-post-test, single groupMeasure before and afterShows change but cannot attribute it
Nonequivalent comparison groupCompare with a similar untreated groupModerate; selection differences remain
Randomized controlled trialRandom assignment to conditionStrongest; frequently infeasible or ethically constrained in service settings
Single-system design (AB, ABAB, multiple baseline)Repeated measures on one client or unitStrong for individual practice evaluation
Qualitative and mixed methodsInterviews, observation, document reviewExplains mechanism and context

Threats to internal validity that exam items reference: history (an outside event affects everyone), maturation (people change simply with time), testing (the measure itself produces change), instrumentation (the measure or rater changes), regression to the mean (extreme scores drift toward average on retest), selection (groups differed at the start), and attrition (dropouts differ systematically from completers).

Regression to the mean deserves particular attention: programs that enroll people at their worst moment will show improvement even if the program does nothing, which is why single-group pretest-post-test findings are weak evidence.

Single-system designs are the practical bridge between practice and evaluation for generalists. AB establishes a baseline then introduces the intervention. ABAB withdraws and reintroduces it, providing stronger causal evidence — and is ethically unacceptable when withdrawal would expose the client to harm, such as with self-injury or suicidality. Multiple baseline designs stagger the start of the intervention across behaviors, settings, or clients, providing causal evidence without withdrawal, which makes them the ethical alternative.

Measurement

Selecting measures: prefer standardized instruments with established reliability (consistency) and validity (measuring what is intended), and check whether the instrument was normed on a population resembling your clients. Combine self-report with behavioral or administrative indicators where possible, keep the burden on participants low, and choose measures sensitive to change over the program's actual duration.

Distinguish once more between outputs — units of service delivered — and outcomes — changes in participants' knowledge, behavior, condition, or status. Funders increasingly require outcomes, and reporting outputs in their place is the most common weakness in agency reporting.

Ethics of Evaluation

NASW Standard 5.02 governs evaluation and research. Its requirements, tested regularly:

  • Monitor and evaluate policies, program implementation, and practice interventions, and contribute to the knowledge base.
  • Voluntary, informed consent without implied or actual deprivation of services for declining, and without undue inducement.
  • Appropriate consent procedures for participants unable to consent, including assent and permission from a legally authorized proxy.
  • Protect participants from unwarranted distress, and provide appropriate support afterward.
  • Confidentiality of data, with participants informed about limits, disposal, and any planned disclosures.
  • Report findings accurately, disclose conflicts of interest, and never fabricate or falsify results.

Two practical consequences: a client must never be told or led to believe that services depend on participating in an evaluation, and evaluation data must be reported honestly even when the findings are unflattering to the program. Suppressing a negative finding to protect funding is a falsification, and an option describing it is always wrong.

Using Findings

Evaluation that ends in a filed report changes nothing. Utilization-focused evaluation builds in the intended users and intended uses from the beginning: involve staff and participants in designing the questions, report in a usable format and timeframe, present negative findings as improvement information rather than as indictment, and close the loop with the people who supplied the data. A program that learns it is not producing outcomes and adjusts has done exactly what evaluation is for.

Test Your Knowledge

A county concludes that demand for adolescent behavioral health services is low because its clinic has open appointment slots each week. Which needs assessment limitation does this reasoning illustrate?

A
B
C
D
Test Your Knowledge

A researcher wants to establish that a behavioral intervention is responsible for reducing a client's self-injurious behavior, but withdrawing the intervention would expose the client to serious harm. Which single-system design is the appropriate choice?

A
B
C
D
Test Your Knowledge

A program evaluation shows that participants had worse outcomes than a comparison group. The executive director asks the social worker to omit the comparison data from the funder report. What does NASW Standard 5.02 require?

A
B
C
D