18.2 Item Analysis and Assessment Data
Key Takeaways
Item difficulty is the proportion correct for the stated population.
Upper-minus-lower discrimination can flag problems but does not diagnose their cause.
The worked counts reconcile to 100 learners, 60 correct and discrimination about minus 0.22.
Psychometric Evaluation: Item Analysis and Statistical Metrics
Following assessment administration, instructional trainers evaluate exam performance through classical test theory metrics to verify item health.
Item Difficulty (p-value): p = R / N
Where: R = Number of examinees answering correctly, N = Total number of examinees
Item Discrimination Index (D): D = P_U - P_L
Where: P_U = Proportion correct in upper 27% group, P_L = Proportion correct in lower 27% group
Item difficulty
The proportion correct is . An item answered correctly by 60 of 100 candidates has . A larger value means an easier item for that group. There is no universal acceptable range; an essential safety item may reasonably be easy after effective instruction. Consider the objective and sample before deciding whether to revise it.
Item discrimination
For a selected upper and lower group, . The group selection method must be stated. A negative result prompts investigation of the key, wording, instruction and scoring; it does not prove a particular cause. In a small sample, results may be unstable. Resolve scoring concerns fairly before using the result for consequential decisions.
Distractor analysis
Check which options learners select and whether the pattern matches a plausible misconception. An unused distractor may be weak, but a fixed percentage is not a universal retirement rule. Review content, language and opportunity to learn before revising the item.
Calculating and interpreting a worked item analysis
Suppose 100 learners take a knowledge test. For one item, 60 answer correctly. The difficulty index is . Despite its name, a larger value means the item was easier for this population. Do not interpret this as an intrinsic property independent of teaching or learner preparation.
For an illustrative discrimination calculation, select the highest-scoring 27 and lowest-scoring 27 learners based on overall test performance. Six upper-group learners and twelve lower-group learners answer this item correctly. The remaining 46 learners include 42 correct responses, so the total remains 60. The upper proportion is , the lower is , and , approximately . These internally consistent counts make a useful warning case.
The negative discrimination means this item was answered correctly less often by the selected upper group. It does not identify the cause by itself. Check the answer key, ambiguity, multiple defensible options, technical premise, scoring transfer and what was actually taught. A miskey is possible, but so are other problems. Investigate before using the questionable result in a consequential decision.
| Finding | Investigation |
|---|---|
| High proportion correct | Essential mastery, easy cue or weak distractors? |
| Low proportion correct | Difficult objective, instruction gap or ambiguous item? |
| Negative discrimination | Key, wording, group selection and scoring correct? |
| Distractor rarely chosen | Plausible misconception or obviously irrelevant option? |
Linking analysis with content review
Statistics are supporting evidence. A high discrimination value does not prove technical accuracy; an item can separate groups while testing irrelevant vocabulary. An essential safety item that everyone answers correctly may still belong in a qualification assessment. Review the objective and content alongside the data.
For distractors, record counts by option and examine learner explanations where feasible. If many learners choose a plausible but incorrect action, determine whether it reflects a misconception that needs clearer instruction. If an option is absurd, replace it with a defensible novice error. Avoid making a revised distractor so similar to the key that two answers become valid.
Small samples require caution. A few responses can change an index substantially, and tied scores can complicate upper/lower grouping. Record the sample and grouping method. Do not claim psychometric equivalence merely because two forms have the same number of questions or similar averages. Form comparability requires appropriate content and measurement evidence.
Fair correction and reporting
If review confirms a key or scoring error, follow the applicable correction policy and communicate the effect on decisions. Preserve the original result, corrected result and rationale where the record policy requires it. Avoid concealing a defect by deleting a difficult item without investigating which learners or decisions were affected.
For a safety qualification, ensure unresolved measurement issues do not produce unsupported authorization. Other valid performance evidence may clarify competence, but a revised written score cannot stand in for a missing practical check. Plan remediation for genuine learning gaps separately from correction of the instrument.
Use a report that identifies the item, objective, sample, observed indices, review findings and action owner. After revision, gather new evidence rather than assuming the problem disappeared. This joins statistical analysis with technical review and instructional improvement.
Checking arithmetic before interpretation
Upper, lower and remaining groups must reconcile with the total sample. In this example, 27 plus 27 plus 46 equals 100, and six plus twelve plus 42 equals 60 correct responses. If the counts do not reconcile, correct the data before discussing item quality.
Keep rounding until the final reported value. The exact discrimination is minus six twenty-sevenths, approximately minus 0.22. A rounded index should not be mixed with incompatible percentage claims. Report the underlying counts where they help readers audit the result.
Key takeaways
- Item difficulty is the proportion correct for the stated population.
- Upper-minus-lower discrimination can flag problems but does not diagnose their cause.
- The worked counts reconcile to 100 learners, 60 correct and discrimination about minus 0.22.
Six of 27 upper-group learners and twelve of 27 lower-group learners answer correctly. What is the discrimination index?
Approximately −0.22
Approximately 0.22
Approximately 0.60
Approximately −0.75
Sections you finish are checked off in the contents.