7.3 Systems Change & Program Evaluation
Key Takeaways
The National Implementation Research Network (NIRN) Active Implementation Framework delineates four recursive stages of systems change: Exploration, Installation, Initial Implementation, and Full Implementation.
The 'Implementation Dip'—a temporary drop in perceived competence and performance during Initial Implementation—must be anticipated and mitigated through ongoing coaching, behavioral modeling, and supportive administration rather than premature abandonment of the innovation.
Implementation fidelity must be systematically evaluated across five core dimensions: adherence, dosage/exposure, quality of delivery, participant responsiveness, and program differentiation.
Program evaluation requires a dual methodology: Formative evaluation provides continuous process-oriented feedback during development and execution to guide real-time iterations, while Summative evaluation evaluates end-state efficacy, outcome achievement, and systemic cost-benefit impact.
Logic models serve as essential structural blueprints for system-wide initiatives, mapping linear and recursive linkages among Inputs (resources), Activities (practices delivered), Outputs (tangible counts), Short-term Outcomes (skill/knowledge acquisition), Intermediate Outcomes (behavior change), and Long-term Impact (systemic change).
Systems Change & Program Evaluation
School psychologists serve not merely as individual clinicians assessing individual children, but as organizational leaders and systems-change agents (NASP Practice Model Domain 5: School-Wide Practices to Promote Learning, and Domain 9: Research and Evidence-Based Practice). Educational history is littered with well-intentioned reforms, curricula, and initiatives that were enthusiastically introduced, poorly supported, and quickly abandoned. Bridging the 'science-to-practice gap' requires a sophisticated understanding of Implementation Science—the empirical study of methods that promote the systematic uptake of research findings and evidence-based practices into routine educational operations (Fixsen et al., 2005).
The NIRN Active Implementation Framework
The National Implementation Research Network (NIRN) framework provides the preeminent architecture for educational systems reform. NIRN identifies that interventions do not succeed simply because they possess empirical evidence; they succeed when an evidence-based practice is supported by effective implementation drivers within a receptive organizational climate.
If any one of these factors is zero, student outcomes will be zero. NIRN outlines four sequential, overlapping stages of implementation:
[Stage 1: EXPLORATION] ──► [Stage 2: INSTALLATION]
• Needs assessment • Acquire resources & tools
• Hexagon Tool review • Deliver pre-service PD
• Assess organizational fit• Establish data systems
• Staff buy-in (PBIS ~80%) • Build coaching infrastructure
│ │
▼ ▼
[Stage 3: INITIAL IMPL.] ──► [Stage 4: FULL IMPL.]
• Pilot testing cohorts • >=50% of staff at fidelity
• Overcome 'Implementation • Deeply embedded in SOP
Dip' via coaching • Faded coaching to maintenance
• Rapid PDSA cycles • Expected student outcomes
Detailed Breakdown of the 4 Stages of Implementation
| Stage | Core Objectives | Key Milestones & Diagnostic Gates | School Psychologist Leadership Role |
|---|---|---|---|
| 1. Exploration | Determine community/student needs; evaluate potential evidence-based interventions; assess readiness, fit, and feasibility. | • Conduct epidemiological needs screening.; • Apply the NIRN Hexagon Tool (current version: Evidence, Usability, Supports, Need, Fit, Capacity; the original 2013 tool used Needs, Fit, Resource Availability, Evidence, Readiness for Replication, and Capacity).; • Secure staff buy-in (many SWPBIS adoption processes look for about 80% staff agreement before launch). | Lead data audits; facilitate stakeholder focus groups; synthesize research evidence on candidate interventions. |
| 2. Installation | Build organizational and structural capacity before delivering the intervention to students. | • Secure permanent budget allocations.; • Purchase curricular materials and testing software.; • Establish scheduling adjustments.; • Train staff through high-quality pre-service professional development.; • Create fidelity check instruments. | Design fidelity observation protocols; establish data-collection infrastructure; lead staff professional development. |
| 3. Initial Implementation | Roll out the intervention to students (often via pilot grades/cohorts); navigate logistical friction; provide coaching. | • First cohort receives intervention.; • Provide in-vivo modeling and coaching.; • Navigate the 'Implementation Dip'.; • Run PDSA (Plan-Do-Study-Act) rapid-cycle problem-solving loops. | Conduct live classroom coaching; collect weekly fidelity probes; reassure anxious educators; eliminate logistical hurdles. |
| 4. Full Implementation | The practice becomes standard operating procedure; NIRN marks full implementation when at least 50% of intended practitioners use it with fidelity and good outcomes. | • Intervention embedded into general operating budget.; • Standardized onboarding for new staff.; • Anticipated student outcomes documented and sustained over multiple academic years. | Conduct summative evaluations; analyze return on investment (ROI); monitor for 'program drift'; disseminate outcomes. |
Navigating the 'Implementation Dip'
A ubiquitous phenomenon during Stage 3 (Initial Implementation) is the Implementation Dip (Fullan, 2001). When teachers transition from old, familiar habits to complex new behavioral or instructional practices, their perceived self-efficacy, comfort, and operational speed temporarily plummet. Classroom disruptions may temporarily rise as students test new boundaries, and teachers feel clumsy and overburdened.
Systems-Level Exam Focus: A critical test item on systems change describes educators demanding the immediate abandonment of a new framework 4 to 8 weeks after launch because 'it feels awkward and discipline hasn't improved.' The school psychologist must recognize this as the classic Implementation Dip. The correct systems response is never to abandon the model, nor to reprimand teachers as insubordinate; rather, the team must increase active classroom coaching, modeling, and administrative support, while validating teacher stress.
Implementation Fidelity & Dosage: The 5 Core Dimensions
When a program fails to achieve desired student outcomes, the team must answer the foundational question: Did the intervention fail, or did the implementation fail?
- Intervention Failure: The intervention was delivered with 100% fidelity and full dosage, but failed to alter student behavior (indicating the practice was ill-matched or ineffective for this population).
- Implementation Failure: The intervention was never delivered as designed; sessions were cancelled, materials omitted, or protocols altered, meaning the empirical model was never genuinely tested.
To differentiate between these outcomes, school psychologists systematically measure the Five Dimensions of Implementation Fidelity (Dane & Schneider, 1998):
┌─────────────────────────────────────────┐
│ 5 CORE DIMENSIONS OF FIDELITY │
└────────────────────┬────────────────────┘
┌──────────────┬──────────────────┼──────────────┬──────────────┐
▼ ▼ ▼ ▼ ▼
[1. Adherence] [2. Dosage] [3. Quality] [4. Receptive] [5. Contrast]
Steps delivered Frequency, Enthusiasm, Participant Differentiation
per protocol duration, count pacing, skill engagement from control
- Adherence: The extent to which the practitioner delivers the specific operational components, scripts, and procedures specified in the intervention manual.
- Dosage / Exposure: The quantitative units of intervention received by the student (e.g., number of sessions attended, minutes per session, total weeks completed). A 12-week SEL group delivered for only 4 sessions represents inadequate dosage.
- Quality of Delivery: The pedagogical skill, clinical warmth, pacing, and enthusiasm with which the practitioner delivers the intervention (e.g., adult-student rapport, behavioral modeling quality).
- Participant Responsiveness: The degree to which participating students actively engage with, attend to, and internalize the intervention activities.
- Program Differentiation: Evidence that the active ingredients of the intervention were distinctly present in the treatment condition and completely absent from the comparison/control condition (preventing contamination).
Addressing Barriers to Change
Change efforts fail for predictable reasons. The ETS outline asks candidates to know how to plan and evaluate change activities, monitor fidelity, and address barriers to change at both the individual and systems levels.
| Common Barrier | Strategies |
|---|---|
| Low buy-in or unclear rationale | Share local data showing the need; involve staff in selecting the practice; start with volunteers and show early results |
| Too many competing initiatives | Complete an initiative inventory; align or retire overlapping programs before adding new ones |
| Insufficient training and coaching | Pair training with ongoing coaching and performance feedback, which change practice far more than one-time workshops |
| Lack of time and resources | Protect meeting time, secure funding in the budget, and choose practices that fit existing capacity (the Hexagon Tool) |
| Leadership or staff turnover | Build a team-based implementation structure, document procedures, and train new staff as part of onboarding |
| Weak data feedback | Review fidelity and outcome data regularly and act on them with PDSA cycles |
At the individual level, people move through predictable stages of concern when asked to adopt something new. The Concerns-Based Adoption Model (Hall & Hord) describes them as moving from awareness and information needs, to personal concerns ("How will this affect me?"), to management concerns ("How do I fit this in?"), and later to concerns about consequences for students, collaboration, and refinement. Support should match the stage: information early, practical help with logistics next, and outcome data later. Rogers' diffusion of innovations research adds that adoption spreads from innovators and early adopters to the majority, so visible early successes and respected peer champions speed change.
Program Evaluation Methodologies: Formative vs. Summative
Program evaluation is the systematic collection and analysis of information about the activities, characteristics, and outcomes of programs to make judgments about program effectiveness, improve operations, and inform future programming decisions (NASP Practice Model Domain 9).
Formative Evaluation vs. Summative Evaluation
| Evaluation Dimension | Formative Evaluation (Process Evaluation) | Summative Evaluation (Outcome Evaluation) |
|---|---|---|
| Primary Timing | Conducted during program design, piloting, and active implementation. | Conducted at the conclusion of an intervention cycle or during mature full implementation. |
| Core Purpose | Quality improvement, iterative troubleshooting, and mid-course workflow corrections. | Determining overall efficacy, outcome attainment, scalability, and long-term sustainability. |
| Core Question | Is the program being delivered as intended, and how can we optimize operations? | Did the program achieve its intended student outcomes, and was it worth the investment? |
| Primary Metrics | Fidelity checklists, weekly attendance logs, teacher satisfaction probes, dosage counts. | Standardized CBM gains, ODR reduction percentages, suspension drops, graduation rates, effect sizes (d). |
| Analogy (Robert Stake) | "When the cook tastes the soup, that's formative..." | "...when the guests taste the soup, that's summative." |
Program Evaluation Architecture: The Logic Model
A Logic Model is a visual, conceptual blueprint that articulates the underlying 'theory of change' connecting institutional investments to ultimate community outcomes. It maps how specific activities systematically generate measurable short-term, intermediate, and long-term results.
[INPUTS] ──► [ACTIVITIES] ──► [OUTPUTS] ──► [SHORT-TERM] ──► [INTERMEDIATE] ──► [LONG-TERM]
Resources Practices Tangible Knowledge, Behavioral Systemic
Invested Delivered Counts Attitudes Changes Impact
Complete Systems Logic Model for a Schoolwide Tier 2 Behavioral Initiative
| Logic Model Component | Technical Definition | Concrete Schoolwide Tier 2 CICO Example |
|---|---|---|
| Inputs (Resources) | The raw human, financial, physical, and organizational assets invested in the initiative. | • Title I grant funding ($15,000).; • 0.5 FTE School Psychologist & 1 Paraprofessional.; • DPR card materials and SWIS electronic license.; • Administrative release time for weekly team meetings. |
| Activities (Practices) | The actual evidence-based actions, clinical processes, and events conducted by staff. | • Facilitating 8-hour staff professional development on CICO.; • Conducting daily morning check-ins and afternoon check-outs.; • Delivering period-by-period DPR teacher feedback.; • Bi-weekly PBIS data team problem-solving meetings. |
| Outputs (Direct Products) | The immediate, tangible, countable units of service or products generated by the activities. | • 35 staff members trained.; • 28 at-risk students enrolled in CICO.; • 1,450 DPR rating cards completed across 12 weeks.; • 6 bi-weekly data meetings held. |
| Short-Term Outcomes | Immediate alterations in participant knowledge, skills, perceptions, or attitudes (Weeks 1–6). | • Participating teachers master behavior-specific praise delivery.; • Enrolled students articulate their daily behavioral goals.; • Caregivers report feeling informed regarding daily student progress. |
| Intermediate Outcomes | Changes in actual participant behaviors, disciplinary habits, and academic performance (Months 2–9). | • 82% of enrolled students achieve ≥ 80% of daily DPR points.; • Class-wide instructional off-task behavior drops by 40%.; • ODRs for enrolled students decline by an average of 55%. |
| Long-Term Impact | Enduring, systemic, institutional, or societal transformation sustained over years. | • Grade-level retention drops by 20%.; • High school graduation rate increases by 8 percentage points.; • Racial discipline disproportionality is eliminated across the district. |
Important
Exam Trap: Conflating Outputs with Outcomes
Evaluators frequently commit the grave error of reporting Outputs (e.g., 'We trained 50 teachers and held 20 parent workshops') and claiming that the program was successful. Training counts and distributed workbooks demonstrate only effort and dosage; they provide zero evidence that student knowledge, behavior, or institutional climate actually improved. Always demand Outcome data (behavior change, skill mastery) to prove efficacy.
Stakeholder Engagement & Culturally Responsive Evaluation
A program evaluation cannot be conducted in a clinical vacuum. The Joint Committee on Standards for Educational Evaluation (JCSEE) identifies that rigorous evaluation requires deep stakeholder engagement:
- Involving Diverse Constituents: Engaging students, families, general educators, and community members in defining what 'success' looks like, ensuring that outcome metrics reflect the lived values of the school community.
- Culturally Responsive Evaluation: Ensuring evaluation instruments are linguistically accessible, culturally unbiased, and normed on populations matching the demographic composition of the student body.
Case Vignette: District-wide Rollout of Restorative Behavioral MTSS
Setting: Central Unified School District (12 schools, 8,500 students) implemented a multi-tiered restorative discipline framework to replace zero-tolerance suspensions.
Implementation Science in Action:
- Stage 1 (Exploration): The district leadership team spent 9 months reviewing suspension data, surveying staff readiness, and utilizing the NIRN Hexagon Tool. The team secured an 84% voluntary faculty vote in favor of adopting restorative practices, above the roughly 80% agreement many schools seek before launch.
- Stage 2 (Installation): Before working with students, the district hired 3 restorative coaches, modified daily master schedules to include 20-minute morning restorative circles, and built an online fidelity tracking rubric.
- Stage 3 (Initial Implementation): At week 6, elementary and middle school teachers reported extreme stress, stating circles took too much instructional time and student disruptions persisted (The Implementation Dip). The school psychology team provided live classroom co-facilitation, weekly problem-solving circles for teachers, and administrative reassurance. By week 16, teacher fidelity surpassed 80%.
- Stage 4 (Full Implementation & Evaluation): At year 2, a comprehensive evaluation revealed:
- Formative Fidelity Probes: Restorative circle adherence averaged 88% across classrooms.
- Summative Outcomes: Out-of-school suspensions decreased by 48% across the district, instructional seat time gained totaled 3,200 hours, and disaggregated climate surveys demonstrated significant increases in student perceptions of fairness.
Six weeks following the district-wide rollout of a new Tier 1 restorative practices and positive classroom management framework, several elementary teachers express intense frustration during a staff meeting, reporting that the new protocols 'take too much time, feel awkward, and student disruptions have temporarily increased.' Several grade-level chairs propose immediately abandoning the new model and returning to traditional punitive referral practices. Based on the National Implementation Research Network (NIRN) framework, how should the school psychologist interpret and respond to this situation?
Agree to terminate the restorative practices program immediately because an initial increase in student disruptions proves the intervention is empirically flawed for this population.
Conclude that the teachers are willfully insubordinate and recommend formal administrative disciplinary evaluations for any educator refusing to comply.
Recognize this phenomenon as the classic 'Implementation Dip' during Initial Implementation, and provide targeted peer coaching, behavioral modeling, and administrative reassurance rather than abandoning the initiative.
Transition immediately into the Full Implementation stage by removing coaching supports and requiring teachers to self-evaluate their own compliance.
A school district conducts an annual program evaluation of its newly implemented multi-tiered Social-Emotional Learning (SEL) curriculum. The evaluation report lists the following findings: (1) 85 staff members completed 12 hours of professional development, (2) 1,200 student workbooks were distributed, and (3) 420 SEL lesson blocks were delivered across 18 classrooms. The district superintendent concludes from these metrics that 'the SEL program has demonstrated outstanding student outcomes.' Within a formal logic model evaluation framework, how should the school psychologist evaluate the superintendent's conclusion?
The superintendent's conclusion is valid because high counts of training hours and distributed workbooks directly confirm positive student emotional growth.
The superintendent has confused Inputs with Activities, as professional development and workbooks represent financial expenditures rather than school actions.
The reported metrics are Intermediate Outcomes because they reflect measurable behavioral changes occurring within the educational environment.
The superintendent has conflated program Outputs (quantifiable units of service delivered) with student Outcomes (demonstrated gains in student social-emotional skills, attitudes, or behavioral competence).
At the conclusion of a school year, a high school PBIS team reviews outcome data for an intensive Tier 2 mentoring and behavioral goal-setting program for ninth-grade students at risk of academic failure. The summative data indicate that participants showed no statistically significant reduction in disciplinary referrals or improvement in grade-point averages compared to a matched control group. Before concluding that the mentoring intervention is ineffective, what critical dimension of implementation science must the school psychologist investigate?
The cost-per-pupil financial expenditure of the curriculum materials.
Implementation fidelity and dosage, to determine whether the intervention was delivered as planned (intervention failure vs. implementation failure).
The intellectual quotient (IQ) distribution of all participating ninth-grade students.
The political alignment and ideological perspectives of the school board members who funded the initiative.
Sections you finish are checked off in the contents.