11.1 Item Analysis, Difficulty, Discrimination & Reliability
Key Takeaways
- Item analysis turns raw exam scores into evidence about each item’s difficulty (p-value), discrimination (often point-biserial), and distractor performance so faculty revise tests rather than only grade students.
- Classroom tests often aim for item p-values roughly 0.30–0.70 and discrimination indices commonly above about 0.20–0.30, with KR-20 or Cronbach’s alpha often targeted at ≥0.70—treat these as useful guidelines, not universal laws.
- Negatively discriminating items and non-functioning distractors threaten score meaning; a master workflow is: blueprint review → score → compute stats → interpret with content experts → revise, discard, or reteach.
- Reliability estimates consistency of scores; validity concerns whether scores support intended use—high reliability does not prove a test measures the right outcomes.
- CNE traps include keeping bad items because “we always used them,” interpreting stats without context (small N, mastery tests), and ignoring distractor analysis.
Why Item Analysis Matters on the CNE
Domain 3 of the NLN CNE Detailed Test Blueprint—Use Assessment and Evaluation Strategies—includes tasks on analyzing assessment data and using results to improve teaching and learning. After you write and administer a test, the work is not finished. Item analysis is the systematic review of how each question performed for a given group of examinees. Without it, faculty grade students on instruments they never inspect, keep weak items year after year, and miss teaching gaps hidden in the score pattern.
On CNE items, stems often present a p-value, a discrimination index, or a distractor pattern and ask what the educator should do next: revise the stem, rewrite options, discard the item, keep it, or reteach content. The correct move is almost always data-informed judgment—statistics plus content expertise—not blind trust in numbers or blind loyalty to tradition.
Quick Answer: Difficulty (p) describes how many students got the item right; discrimination describes how well the item separates higher from lower scorers; reliability estimates overall score consistency. Use guideline ranges for classroom tests, then act: revise, discard, or reteach—never file the report and do nothing.
Difficulty Index (p-Value)
Definition and calculation
For dichotomously scored items (correct/incorrect), the difficulty index or p-value is the proportion of examinees who answered correctly:
p = number correct ÷ number of examinees
| p-value | Interpretation (general classroom language) |
|---|---|
| Near 1.0 | Very easy—almost everyone correct |
| Near 0.0 | Very hard—almost everyone incorrect |
| ~0.50 | Moderate difficulty for a traditional multi-choice test |
Desirable ranges—guidelines, not law
For many classroom nursing exams designed to differentiate achievement across a range of performance, educators and measurement texts often cite a desirable difficulty band of roughly 0.30–0.70 (sometimes framed more tightly around 0.40–0.60 for maximum discrimination potential on four- or five-option items). Items outside that band can still be justified:
- Mastery / competency tests (e.g., must-know safety content) may intentionally show high p-values after effective teaching.
- Very hard items may be appropriate for advanced synthesis if discrimination remains acceptable and the blueprint requires that depth.
- Very easy items may be acceptable for foundational checks early in a course or as confidence builders—if they do not dominate high-stakes decisions.
CNE nuance: An item with p = 0.95 is not automatically “bad.” Ask: Was this a critical safety competency everyone should know? Did teaching succeed? Or is the item so transparent that it wastes test time and fails to discriminate?
Actions by difficulty pattern
| Pattern | Possible meanings | Educator actions |
|---|---|---|
| Very high p (e.g., >0.90) | Easy content; excellent teaching; giveaway wording | Keep if mastery expected; revise if too obvious; reduce weight if it inflates grades without adding information |
| Very low p (e.g., <0.30) | Hard content; poor teaching; ambiguous stem; wrong key; content not taught | Content review first; check key; revise wording; reteach if gap is real |
| Mid-range p (~0.30–0.70) | Often useful for differentiating performance | Still inspect discrimination and distractors |
Discrimination Index and Point-Biserial
What discrimination means
Discrimination answers: Do students who score high on the whole test tend to get this item right more often than students who score low? A good item should be answered correctly more often by stronger examinees.
Common indices:
- Item discrimination index (D) — often based on upper vs lower scoring groups (e.g., top and bottom 27%): D ≈ (proportion correct upper) − (proportion correct lower).
- Point-biserial correlation (r_pbis) — correlation between item score (0/1) and total test score (often total score excluding the item itself to reduce inflation).
Desirable ranges—again, guidelines
Many classroom assessment references treat discrimination above about 0.20–0.30 as desirable for typical course exams, with higher values preferred when stakes and sample size support finer interpretation. Rough interpretive language used in faculty development:
| Discrimination (approx.) | Common faculty language |
|---|---|
| ≥0.40 | Excellent / very good |
| 0.30–0.39 | Good |
| 0.20–0.29 | Acceptable / fair—review |
| 0.00–0.19 | Weak—revise or rewrite |
| Negative | Problematic—reverse relationship |
Negatively discriminating items are answered correctly more often by lower scorers than by higher scorers. Causes include wrong answer key, ambiguous wording that “tricks” high performers, multiple defensible answers, or content that rewards testwise guessing over knowledge. Do not keep negatively discriminating items without urgent review; they actively harm score validity.
Relationship of difficulty and discrimination
Items that are extremely easy or extremely hard have limited room to discriminate (almost everyone right or wrong). Moderate difficulty often supports higher discrimination if the item is well written and aligned to taught outcomes. Do not chase a p-value of 0.50 while ignoring content fidelity.
Distractor Analysis
For multiple-choice items, distractor analysis examines how often each incorrect option is chosen—overall and by high vs low scorers.
| Distractor finding | Meaning | Action |
|---|---|---|
| Never or almost never chosen | Non-functioning distractor | Rewrite to be plausible for those who lack the concept |
| Chosen mainly by low scorers | Functioning distractor | Keep if content-valid misconception |
| Chosen mainly by high scorers | Ambiguity, dual key, or keyed error | Fix stem/options/key immediately |
| Even spread across wrong options with low discrimination | Confusing item or under-taught content | Revise + consider reteach |
| One wrong option absorbs almost all errors | Common misconception | Teaching opportunity; may keep if intentional |
Plausible distractors based on real clinical misconceptions both improve discrimination and generate formative teaching data. Implausible distractors waste option slots and inflate p-values without measuring understanding.
Test-Level Reliability: KR-20 and Cronbach’s Alpha
Concepts
- Reliability ≈ consistency of scores (would similar results occur under similar conditions?).
- KR-20 (Kuder–Richardson 20) estimates internal consistency for dichotomously scored tests.
- Cronbach’s alpha is a more general internal-consistency coefficient used for multi-item scales and often reported for exams as well.
For many classroom tests, faculty handbooks and measurement texts often cite α or KR-20 ≥ 0.70 as a reasonable target for low-to-moderate stakes course exams, with higher values preferred as stakes rise. Certification and licensure exams aim much higher; classroom tests rarely match that bar and should not be held to the same absolute standard.
Factors that raise or lower reliability estimates
| Tends to increase reliability estimate | Tends to decrease it |
|---|---|
| More quality items | Very short tests |
| Adequate score variance | Homogeneous performance (everyone similar) |
| Clear, discriminating items | Ambiguous, negatively discriminating items |
| Consistent scoring | Inconsistent partial-credit scoring |
| Appropriate difficulty mix | Floor/ceiling effects |
Limits of reliability numbers
- High reliability ≠ validity. A consistently measured wrong construct is still wrong.
- Small class sizes produce unstable item stats—interpret cautiously; look for patterns across offerings.
- Mastery tests with little score variance can show low reliability estimates even when teaching succeeded.
- Do not “fix” reliability by adding redundant easy items that do not map to outcomes.
Master Test Analysis Workflow
Use a repeatable faculty process after each major exam:
- Before scoring decisions finalize: Confirm keys, version codes, and any invalidated items (typos discovered during administration).
- Generate item statistics: p-values, discrimination (point-biserial or D), distractor frequencies, and test reliability if available from the LMS/testing platform.
- Flag items: Negative discrimination; extreme difficulty inconsistent with intent; non-functioning distractors; items with student challenges during the exam.
- Content expert review: Does the item still match the blueprint and taught content? Is the key correct? Is wording biased or ambiguous?
- Decide per item: Keep as-is; minor edit for next form; major rewrite; retire; or accept for mastery reasons with documentation.
- Score policy decision: For clearly flawed items on this administration, faculty/policy may credit all students, drop the item, or accept dual keys—document rationale for fairness and consistency.
- Close the loop to teaching: Cluster weak items by content area → plan reteach, future emphasis, or clinical application practice.
- Archive: Store analysis with the exam form for curriculum review and accreditation evidence.
| Flag | Prefer action |
|---|---|
| Wrong key confirmed | Rescore; fix bank |
| Negative discrimination + ambiguity | Drop or dual-credit this term; rewrite |
| Low p + low discrimination + not taught | Do not punish students; fix curriculum/teaching |
| Low p + high discrimination + was taught | Hard but fair—keep or slightly revise; consider more practice next term |
| High p + low discrimination + critical safety | Acceptable mastery item—keep |
| High p + low discrimination + trivial content | Replace with better-sampled outcome |
Connecting Analysis to Fairness and Legal Defensibility
Item analysis supports due process and fairness: students deserve instruments that measure published outcomes, not faculty “gotchas.” Documented review processes also support program evaluation and continuous quality improvement. On the CNE, options that say “ignore stats and keep tradition” or “curve everything without examining items” are weak educator practice.
Common CNE Traps
| Trap | Why it fails | Better move |
|---|---|---|
| Keep bad items forever | Accumulates invalid variance | Retire/rewrite on evidence |
| Treat 0.30–0.70 or α≥0.70 as absolute law | Context (mastery, N, stakes) matters | Interpret with purpose and content |
| Only look at total mean score | Hides item-level failure | Full item + distractor review |
| Drop every hard item after student complaint | May gut rigorous outcomes | Use discrimination + content review |
| Assume high α means “good test” | Reliability ≠ validity/alignment | Check blueprint and outcomes |
| Never act on analysis | Data theater | Revise items and teaching |
Bottom Line for Domain 3 Task F (Analysis)
Read difficulty, discrimination, reliability, and distractors as decision tools. Target ranges commonly cited for classroom tests—p roughly 0.30–0.70, discrimination often >0.20–0.30, KR-20/α often ≥0.70—are guidelines. Negative discriminators and non-functioning distractors demand action. Master the workflow: analyze → interpret with content expertise → revise instruments and instruction.
A unit exam item has a p-value of 0.22 and a point-biserial of −0.18. Student comments and faculty review show two options could be defended as correct. What is the best immediate faculty response?
For a typical multi-topic classroom nursing exam intended to differentiate student achievement, which item difficulty (p-value) range is most often cited as generally desirable?
An item’s correct answer is chosen mostly by low scorers, while high scorers prefer a particular distractor. Which interpretation is most accurate?
A skills-theory quiz yields Cronbach’s alpha of 0.72 in a class of 48 students. Which statement best reflects CNE-level interpretation?