11.1 Kirkpatrick's Four Levels of Evaluation Framework

Key Takeaways

  • The Kirkpatrick Four-Level Model, originated by Donald Kirkpatrick in 1959 and modernized into the New World Kirkpatrick Model by Jim and Wendy Kirkpatrick, provides the foundational taxonomy for evaluating learning across Reaction (Level 1), Learning (Level 2), Behavior (Level 3), and Results (Level 4).
  • Modern Level 1 evaluation moves decisively beyond superficial satisfaction 'smile sheets' to rigorously capture learner engagement and perceived utility/relevance, which serve as early predictors of learner motivation to apply skills on the job.
  • The New World model expands Level 2 beyond declarative knowledge and technical skills to encompass learner attitude, confidence (self-efficacy), and commitment (intent to apply), recognizing these psychological constructs as essential prerequisites for behavioral transfer.
  • Level 3 (Behavior) represents the critical execution bridge where most programs fail; sustaining on-the-job transfer requires establishing 'required drivers'—reinforcing, encouraging, monitoring, and rewarding critical behaviors while dismantling transfer obstacles.
  • The 'Chain of Evidence' demonstrates the unbroken causal pathway from Level 1 to Level 4; Robert Brinkerhoff's Success Case Method (SCM) complements this by investigating extreme outliers (top 5-10% success cases vs. bottom non-users) to diagnose organizational enablers and systemic barriers.
Last updated: September 2026

11.1 Kirkpatrick's Four Levels of Evaluation Framework

CPTD Exam Focus: The Kirkpatrick Four-Level Model is the most frequently tested evaluation framework on the CPTD examination. Candidates must master both Donald Kirkpatrick's classical 1959 taxonomy and the modernized New World Kirkpatrick Model developed by Jim and Wendy Kirkpatrick. Critical exam themes include the shift from satisfaction 'smile sheets' to perceived utility in Level 1, the expanded dimensions of Level 2 (confidence and commitment), the orchestration of Level 3 'required drivers' to overcome the knowing-doing gap, the distinction between leading indicators and lagging business outcomes in Level 4, and the strategic application of Robert Brinkerhoff's Success Case Method (SCM) to isolate organizational transfer enablers and systemic blockers.


1. Evolution: Classical vs. New World Kirkpatrick Model

In 1959, Donald Kirkpatrick published a landmark series of four articles in the Journal of the American Society of Training Directors (ASTD, now ATD), introducing a systematic taxonomy for evaluating training programs. For decades, organizations implemented this model as a post-training linear checklist, starting at Level 1 and rarely progressing beyond Level 2.

In the 2010s, Jim Kirkpatrick and Wendy Kayser Kirkpatrick modernized the framework into the New World Kirkpatrick Model. This contemporary iteration addresses the greatest historical vulnerability of corporate talent development: designing instruction in isolation and failing to ensure on-the-job execution. The New World model introduces three foundational shifts:

  1. Planning in Reverse: High-impact talent development practitioners do not design Level 1 first. Instead, they begin at Level 4 (Business Results), identifying the organization's strategic objectives and key performance indicators (KPIs). They then work backward to determine the Level 3 critical behaviors required to move those metrics, the Level 2 learning and psychological enablers needed to perform those behaviors, and finally the Level 1 learning environment that fosters engagement and perceived relevance.
  2. The Primacy of Level 3 Execution: The New World model asserts that the execution of training occurs after the event, on the job. Training alone produces zero organizational impact; business value is realized only when newly acquired capabilities are consistently applied in the workplace.
  3. The Chain of Evidence: Evaluation must construct an unbroken narrative and empirical thread connecting learner reaction, learning acquisition, on-the-job application, and ultimate business results.
 The Kirkpatrick Evaluation Continuum
 ┌────────────────────────────────────────────────────────────────────────┐
 │  REVERSE PLANNING: Start with Strategic Business Results               │
 │  Level 4: Results      ─── What organizational KPIs must improve?      │
 │         ▲                                                              │
 │  Level 3: Behavior     ─── What critical behaviors drive those KPIs?   │
 │         ▲                                                              │
 │  Level 2: Learning     ─── What knowledge, confidence, & commitment?   │
 │         ▲                                                              │
 │  Level 1: Reaction     ─── What engagement & utility must we deliver?  │
 │                                                                        │
 │  FORWARD EXECUTION: Deliver, Support, and Measure Chain of Evidence    │
 └────────────────────────────────────────────────────────────────────────┘

2. Level 1: Reaction (Moving Beyond "Smile Sheets")

Historically, Level 1: Reaction was derided as the "smile sheet"—a superficial end-of-course survey asking whether participants liked the instructor, enjoyed the catering, or found the classroom comfortable. Research in organizational psychology (such as meta-analyses by Alliger et al.) has repeatedly demonstrated that superficial affective satisfaction has virtually zero correlation with learning retention or on-the-job behavioral transfer.

The New World Kirkpatrick Model elevates Level 1 by dissecting reaction into three distinct, measurable components:

The Three Sub-Dimensions of Level 1

  1. Customer Satisfaction: The degree to which participants found the learning environment, instructor professionalism, technology interface, and pacing satisfactory. While necessary for basic comfort, it is the least predictive of transfer.
  2. Engagement: The degree to which participants actively contributed, maintained cognitive focus, interacted with peers, and engaged with instructional exercises. High engagement reflects active cognitive processing.
  3. Perceived Utility / Relevance: The degree to which participants perceive that the training directly addresses authentic job challenges and can be applied immediately to their daily workflows. Perceived utility is the single most important predictor in Level 1; adult learners who view training as highly relevant demonstrate significantly higher intrinsic motivation to overcome transfer obstacles.

Modern Level 1 Instrumentation

Rather than asking vague Likert-scale questions ("I enjoyed the instructor's presentation style"), high-impact Level 1 instruments evaluate actionable utility:

  • "The tools and frameworks practiced today will directly increase my efficiency in executing [Specific Job Task]." (Strongly Disagree to Strongly Agree)
  • "I have identified at least two operational workflows where I will implement these techniques within the next 14 days." (Yes / No / Specific Action Plan)
  • Net Promoter Score (NPS) for L&D: Asking learners: "On a scale from 0 to 10, how likely are you to recommend this program to a colleague who needs to master this capability?" NPS=% Promoters (Scores 9–10)% Detractors (Scores 0–6)\text{NPS} = \% \text{ Promoters (Scores 9–10)} - \% \text{ Detractors (Scores 0–6)}

3. Level 2: Learning (The Expanded Five Dimensions)

In classical evaluation, Level 2: Learning was confined to assessing declarative knowledge (memorized facts) and procedural skills (demonstrations). The New World Kirkpatrick Model dramatically expands Level 2 to encompass five interdependent dimensions:

 New World Level 2 Learning Components
 ┌────────────────────────────────────────────────────────────────────────┐
 │ 1. Knowledge   ─── What do I know? (Declarative & Conceptual)          │
 │ 2. Skill       ─── Can I do it in a safe environment? (Procedural)    │
 │ 3. Attitude    ─── Do I believe this is valuable? (Affective Mindset) │
 │ 4. Confidence  ─── Do I believe I can succeed on the job? (Efficacy)  │
 │ 5. Commitment  ─── Do I intend to apply this on the job? (Pledge)     │
 └────────────────────────────────────────────────────────────────────────┘

Deconstructing the Five Level 2 Dimensions

  • Knowledge: Understanding of principles, frameworks, technical terminology, and operational guidelines. Measured via pre- and post-instruction knowledge checks, scenario quizzes, and cognitive problem-solving exercises.
  • Skill: The demonstrated behavioral capability to execute a procedure, operate software, or handle an interpersonal conversation within the training environment. Measured via behavioral simulations, roleplays with rubrics, work sample evaluations, and technical labs.
  • Attitude: The affective disposition toward the new methodology. If an employee possesses the skill to execute a new customer intake process but harbors a cynical attitude ("Management is just micromanaging us"), behavioral transfer will not occur.
  • Confidence (Self-Efficacy): Grounded in Albert Bandura's self-efficacy theory, confidence measures the learner's personal conviction that they have the capability to execute the skill successfully amidst real-world workplace pressures. Measured on self-efficacy rating scales (e.g., "Rate your confidence from 0% to 100% in handling an escalated customer complaint without managerial escalation").
  • Commitment (Intent to Apply): The learner's explicit dedication and personal accountability to apply the newly acquired capabilities on the job. Commitment is often operationalized through an Action Plan or Application Contract where the learner specifies what behaviors they will execute, with whom, by what target date, and what anticipated barriers they will navigate.

4. Level 3: Behavior (The Execution Bridge & Required Drivers)

Level 3: Behavior evaluates the degree to which participants apply what they learned during training when they return to the operational workplace. This is the execution bridge where talent development programs succeed or fail. The persistent disconnect between high classroom scores and zero workplace application is known as the Knowing-Doing Gap.

Critical Behaviors

To measure Level 3, practitioners must first define Critical Behaviors—the vital few, observable, high-leverage workplace actions that performers must execute consistently on the job to drive business results. If critical behaviors are defined vaguely ("Be a better communicator"), Level 3 cannot be evaluated or coached. They must be operationalized behaviorally ("Conduct weekly 1-on-1 coaching sessions using the GROW model, documenting agreed developmental actions in the HRIS within 24 hours").

The Four Required Drivers

In the New World Kirkpatrick Model, the most profound insight is that instructional designers and trainers cannot achieve Level 3 alone. Sustained behavioral change requires organizational infrastructure termed Required Drivers—processes and execution systems that monitor, encourage, reinforce, and reward the performance of critical behaviors:

Required DriverOperational DefinitionCorporate Talent Development Example
ReinforcingStructured mechanisms that refresh, deepen, and scaffold skills post-training.Automated microlearning prompts, job aids, spaced-retrieval scenario apps, and job-embedded toolkits.
EncouragingRelational, managerial, and peer support that fosters psychological safety and motivation.Weekly managerial 1-on-1 check-ins, peer coaching cohorts, and executive town hall endorsements.
MonitoringSystematic observation, auditing, and measurement of behavioral application over time.Supervisory behavioral checklists, customer interaction quality audits, and CRM workflow telemetry.
RewardingIntrinsic and extrinsic recognition, praise, and incentives tied directly to behavioral execution.Spotlight recognition in team meetings, performance appraisal alignment, and developmental milestone rewards.

Transfer Obstacles: Broad & Newstrom and Baldwin & Ford

When Level 3 behavioral transfer fails, practitioners must analyze the organizational ecosystem. Mary Broad and John Newstrom's Transfer of Training framework highlights the temporal roles of three key stakeholders: the learner, the trainer, and the manager across three time phases (before, during, and after training).

  • The Broad & Newstrom Discovery: Empirical research proves that the manager's actions before and after training have the single greatest impact on behavioral transfer. Yet, historically, organizations focus almost all resources on the trainer during the event. When managers fail to discuss learning objectives before a program or fail to provide coaching and feedback after the program, transfer rates plummet.
  • Baldwin & Ford's Transfer Model: Outlines three input factors influencing transfer: Trainee Characteristics (ability, self-efficacy, motivation), Training Design (identical elements, principles, behavioral modeling), and Work Environment (manager support, peer climate, opportunity to perform). The opportunity to perform—giving employees immediate, real-world opportunities to practice newly learned skills without fear of penalty—is an absolute prerequisite for retention.

5. Level 4: Results (Business Outcomes & Leading Indicators)

Level 4: Results measures the targeted organizational outcomes that occur as a result of the application of critical behaviors. Historically, Level 4 was synonymous with lagging macroeconomic or operational indicators: profit margins, customer retention, revenue growth, defect reduction, or compliance violations.

Lagging Outcomes vs. Leading Indicators

A fundamental contribution of the New World model is the operationalization of Leading Indicators:

  • Lagging Outcomes: High-level organizational KPIs that take months or years to materialize and are influenced by numerous macro-environmental factors (e.g., annual market share, enterprise operating income, annual employee turnover).
  • Leading Indicators: Short-term, observable operational checkpoints and early milestones that demonstrate critical behaviors are taking hold and progressing toward the ultimate business result.
 The Chain of Evidence Architecture
 ┌────────────────────────────────────────────────────────────────────────┐
 │ Level 1: High engagement & clear perceived utility                     │
 │     │                                                                  │
 │     ▼ [Enables]                                                        │
 │ Level 2: Mastery of skills, high self-efficacy, & commitment           │
 │     │                                                                  │
 │     ▼ [Applied via Required Drivers]                                   │
 │ Level 3: Consistent on-the-job execution of Critical Behaviors         │
 │     │                                                                  │
 │     ▼ [Produces]                                                       │
 │ Level 4 (Leading Indicators): Early operational milestones (30-60 days)│
 │     │                                                                  │
 │     ▼ [Culminates in]                                                  │
 │ Level 4 (Lagging Results): Strategic enterprise business outcomes      │
 └────────────────────────────────────────────────────────────────────────┘

Constructing the Chain of Evidence

Talent development professionals must provide executive sponsors with a compelling Chain of Evidence. If an executive asks, "How did our $500,000 sales training investment increase revenue by $4 million?" the practitioner does not simply show a before-and-after revenue chart. They trace the unbroken causal narrative:

  1. Level 1: 92% of sales engineers rated the solution-selling methodology as directly relevant to their complex enterprise accounts.
  2. Level 2: Post-instruction simulations verified that 89% demonstrated mastery of consultative discovery questioning and scored over 85% in confidence.
  3. Level 3: Managerial observation audits (monitoring driver) showed that within 60 days, 78% of sales engineers consistently executed discovery questionnaires prior to product demonstrations (critical behavior).
  4. Level 4 (Leading Indicator): Qualified deal pipeline velocity increased by 22% in the first quarter post-training.
  5. Level 4 (Lagging Result): Enterprise software annual contract value closed up 14% over baseline, contributing $4 million in new revenue.

6. Robert Brinkerhoff's Success Case Method (SCM)

While the Kirkpatrick framework provides a hierarchical evaluation structure, Robert Brinkerhoff's Success Case Method (SCM) offers a highly agile, diagnostic methodology for investigating why training works or fails in complex corporate systems. Published in 2003, SCM avoids the trap of aggregating post-training survey data into bland averages that conceal the operational truth.

The Core Philosophy of SCM

Brinkerhoff recognized that when training fails, the training content itself is rarely the primary culprit. Instead, the failure typically stems from a breakdown in the organizational ecosystem: unsupportive managers, conflicting incentives, obsolete software, or lack of tools. Rather than surveying an entire population to compute average satisfaction, SCM deliberately studies the extreme outliers:

  • The Success Cases (Top 5–10%): Those individuals who embraced the training, applied it tenaciously in their workflow, and achieved spectacular, measurable business results.
  • The Non-Users (Bottom 5–10%): Those individuals who attended the exact same training, possessed the cognitive ability to pass, but failed to apply any of the skills on the job.

The Two-Part SCM Implementation Process

  1. Step 1: The Broad Screening Survey: A brief, targeted survey administered to the entire participant cohort (e.g., 500 sales reps or 200 customer care agents). The survey asks simple, objective screening questions: "Have you used the new diagnostic framework in client engagements?" and "What measurable business outcomes have you realized?" Based on verified responses, the practitioner identifies the top 5–10% extreme success cases and the bottom 5–10% non-users.
  2. Step 2: In-Depth Qualitative Interviews: The practitioner conducts structured, rigorous investigative interviews with both cohorts and their respective supervisors:
    • Interviewing Success Cases: The evaluator explores: What specific behaviors did you execute? What obstacles did you encounter, and how did you navigate them? What tools or manager behaviors supported you? What measurable business value was produced?
    • Interviewing Non-Users: The evaluator explores: What specific workplace factors prevented you from using these skills? Did your supervisor discuss the program with you? Were software tools available? Were you rewarded or penalized for trying new behaviors?

Strategic Utility of SCM for Talent Development

Brinkerhoff's SCM delivers actionable intelligence to executive leaders. It proves that when the organizational ecosystem supports learning, the training intervention produces demonstrable business value. Concurrently, it pinpoints the exact systemic bottlenecks (e.g., "70% of non-users cited that branch managers instructed them to ignore the new protocol in favor of end-of-month quotas"), allowing leadership to intervene and fix operational friction rather than wasting money re-training employees.

Loading diagram...
New World Kirkpatrick Architecture: Reverse Planning vs. Forward Execution
Empirical Workplace Barriers to Level 3 Behavioral Transfer
Test Your Knowledge

A global manufacturing enterprise implements an intensive technical safety training program for 300 plant supervisors. Post-training Level 1 evaluations show an average rating of 4.8 out of 5.0 for instructor quality and engagement. Level 2 post-tests demonstrate that 91% of supervisors achieved mastery of the compliance safety procedures. However, an operational safety audit conducted six months later reveals that only 14% of supervisors are actively executing the mandatory pre-shift safety inspections on the factory floor, and workplace incident rates remain unchanged. According to the New World Kirkpatrick Model, which systemic failure most directly explains this lack of impact?

A
B
C
D
Test Your Knowledge

A healthcare system launches an enterprise customer service and empathy communication initiative for 1,200 outpatient clinic staff members. Hospital executives demand proof of program success within 45 days of program completion, specifically asking whether patient satisfaction scores have improved. The talent development director knows that organizational patient satisfaction ratings are published quarterly with significant reporting lag. What should the director present to leadership to demonstrate empirical progress along the Chain of Evidence?

A
B
C
D
Test Your Knowledge

A financial services organization implements a consultative wealth-management program for 250 financial advisors. Six months post-training, portfolio expansion across the entire cohort shows a statistically ambiguous increase of only 1.2%. Rather than accepting that the program was ineffective or simply calculating average survey scores, the Head of Talent Development wants to uncover why the program succeeded dramatically in certain advisory branches while failing in others. Which evaluation methodology is specifically designed to address this challenge by studying extreme performers?

A
B
C
D