19.2 Learning Outcomes and Evaluation Data
Key Takeaways
Pre/post comparisons need comparable outcomes, populations and scoring conditions.
Normalized gain is descriptive and does not prove causality or individual mastery.
A pre-test score of 100 makes the normalized-gain denominator zero.
Level 2: Learning — Objective Assessment and Pre/Post Testing
Level 2 concerns the intended learning. Use the knowledge, reasoning or physical-performance evidence appropriate to the objectives. Do not require every course to combine every instrument, and do not infer physical skill from a written test alone.
Pre-Test and Post-Test Methodology
Administering identical or psychometrically parallel assessments before and after training provides empirical baseline comparison. However, raw percentage point increases () distort true instructional effectiveness because students with high baseline scores have limited room for numerical improvement—a psychometric artifact known as the ceiling effect. Conversely, very difficult tests can generate floor effects, masking foundational gains.
Normalized gain as a descriptive calculation
A normalized gain expresses observed score improvement as a fraction of the available improvement from the pre-test score. It can be useful descriptively but does not neutralize all baseline differences or prove training caused the gain. It is undefined for a pre-test score of 100 and can be negative.
Do not apply universal low, medium and high gain categories to every occupational course. Interpret the result with test quality, sample, baseline, task requirements and other evidence.
Worked Calculation Example
A cohort of 24 chemical process technicians completes an OSHA Process Safety Management (PSM) training module.
- Baseline cohort pre-test average:
- Post-instruction cohort post-test average:
For these hypothetical cohort averages, the observed normalized gain is 0.60. This does not mean every learner gained equally or that workplace behavior improved. Review individual critical competencies and the design of the comparison.
Controlling Threats to Internal Validity
When conducting pre/post testing, trainers must recognize and control threats to internal validity:
- Pre-Test Sensitization (Testing Effect): Taking a pre-test alerts learners to key exam concepts, prompting them to pay disproportionate attention to those topics during class. Control: Utilize item-pool randomization or counterbalanced split-half testing (Group A receives Form 1 pre and Form 2 post; Group B receives Form 2 pre and Form 1 post).
- Instrumentation Drift: Changing test difficulty, scoring rubrics, or proctor leniency between administrations. Control: Use standardized, validated rubrics and psychometrically parallel items.
- Statistical Regression Toward the Mean: Extreme low scorers naturally tend to score closer to the average on re-testing simply due to random variance. Control: Include an uninstructed control group or calculate normalized gains across matched pairs.
Note
A physical performance objective needs appropriate performance evidence. A multiple-choice knowledge gain or positive reaction alone cannot establish it.
Collecting data that supports a comparison
Specify the learner population, instruments, dates and conditions before collecting results. A pre/post comparison needs comparable evidence of the intended outcome. Identical questions can introduce familiarity; alternate forms can differ in difficulty. Neither choice automatically solves the problem. Document the design and its limitations, and use additional performance evidence where the outcome requires it.
Pair learner records appropriately when examining individual change. An average pre-test from one group and average post-test from a partly different group can reflect composition rather than learning. Report missing data and interruptions. Preserve privacy with suitable identifiers and avoid displaying identifiable scores in a class-wide report.
| Data check | Reason |
|---|---|
| Same outcome and scoring criteria | Comparison has a common meaning |
| Matched population or explained differences | Composition changes are visible |
| Missing responses reported | Results do not silently exclude difficulties |
| Critical criteria examined individually | Averages do not conceal essential gaps |
Consider a hypothetical course with a mean score increase from 40 to 76. The normalized gain is . That describes the aggregate change under the stated conditions. It does not mean every person gained 36 points or that training alone caused the difference. Examine individual results, instrument quality and other explanations.
Quality and completion-time measures also require comparable tasks and conditions. A faster task performed with different equipment may not indicate learning; a high quality score from an easier specimen set may not be comparable. Define the measure, denominator and expected interpretation before using it in a report.
Combine quantitative and qualitative evidence. Counts identify patterns; learner comments and observations can suggest causes. Distinguish an observed finding from a hypothesis, then test the hypothesis through appropriate course or environment changes. Do not declare success solely from a favorable average when the required performance evidence remains incomplete.
Reporting individual and aggregate evidence
Use aggregate results to identify program patterns, and individual evidence for learner decisions. A class average is not a substitute for checking each person's essential criteria. State how many learners completed both measures and how many still need follow-up.
An evaluation report should also describe the test's scope. Improvement on terminology questions establishes a different outcome from improvement on an observed inspection. Label the evidence accurately and avoid converting one favorable measure into a claim about all learning domains.
Key takeaways
- Pre/post comparisons need comparable outcomes, populations and scoring conditions.
- Normalized gain is descriptive and does not prove causality or individual mastery.
- A pre-test score of 100 makes the normalized-gain denominator zero.
A cohort mean rises from 40 to 76 on a 100-point scale. What is normalized gain?
0.60
0.36
0.76
0.90
Sections you finish are checked off in the contents.