11.1 Item Analysis, Difficulty, Discrimination & Reliability

Key Takeaways

  • Item analysis turns raw exam scores into evidence about each item’s difficulty (p-value), discrimination (often point-biserial), and distractor performance so faculty revise tests rather than only grade students.
  • Classroom tests often aim for item p-values roughly 0.30–0.70 and discrimination indices commonly above about 0.20–0.30, with KR-20 or Cronbach’s alpha often targeted at ≥0.70—treat these as useful guidelines, not universal laws.
  • Negatively discriminating items and non-functioning distractors threaten score meaning; a master workflow is: blueprint review → score → compute stats → interpret with content experts → revise, discard, or reteach.
  • Reliability estimates consistency of scores; validity concerns whether scores support intended use—high reliability does not prove a test measures the right outcomes.
  • CNE traps include keeping bad items because “we always used them,” interpreting stats without context (small N, mastery tests), and ignoring distractor analysis.
Last updated: August 2026

Why Item Analysis Matters on the CNE

Domain 3 of the NLN CNE Detailed Test Blueprint—Use Assessment and Evaluation Strategies—includes tasks on analyzing assessment data and using results to improve teaching and learning. After you write and administer a test, the work is not finished. Item analysis is the systematic review of how each question performed for a given group of examinees. Without it, faculty grade students on instruments they never inspect, keep weak items year after year, and miss teaching gaps hidden in the score pattern.

On CNE items, stems often present a p-value, a discrimination index, or a distractor pattern and ask what the educator should do next: revise the stem, rewrite options, discard the item, keep it, or reteach content. The correct move is almost always data-informed judgment—statistics plus content expertise—not blind trust in numbers or blind loyalty to tradition.

Quick Answer: Difficulty (p) describes how many students got the item right; discrimination describes how well the item separates higher from lower scorers; reliability estimates overall score consistency. Use guideline ranges for classroom tests, then act: revise, discard, or reteach—never file the report and do nothing.

Difficulty Index (p-Value)

Definition and calculation

For dichotomously scored items (correct/incorrect), the difficulty index or p-value is the proportion of examinees who answered correctly:

p = number correct ÷ number of examinees

p-valueInterpretation (general classroom language)
Near 1.0Very easy—almost everyone correct
Near 0.0Very hard—almost everyone incorrect
~0.50Moderate difficulty for a traditional multi-choice test

Desirable ranges—guidelines, not law

For many classroom nursing exams designed to differentiate achievement across a range of performance, educators and measurement texts often cite a desirable difficulty band of roughly 0.30–0.70 (sometimes framed more tightly around 0.40–0.60 for maximum discrimination potential on four- or five-option items). Items outside that band can still be justified:

  • Mastery / competency tests (e.g., must-know safety content) may intentionally show high p-values after effective teaching.
  • Very hard items may be appropriate for advanced synthesis if discrimination remains acceptable and the blueprint requires that depth.
  • Very easy items may be acceptable for foundational checks early in a course or as confidence builders—if they do not dominate high-stakes decisions.

CNE nuance: An item with p = 0.95 is not automatically “bad.” Ask: Was this a critical safety competency everyone should know? Did teaching succeed? Or is the item so transparent that it wastes test time and fails to discriminate?

Actions by difficulty pattern

PatternPossible meaningsEducator actions
Very high p (e.g., >0.90)Easy content; excellent teaching; giveaway wordingKeep if mastery expected; revise if too obvious; reduce weight if it inflates grades without adding information
Very low p (e.g., <0.30)Hard content; poor teaching; ambiguous stem; wrong key; content not taughtContent review first; check key; revise wording; reteach if gap is real
Mid-range p (~0.30–0.70)Often useful for differentiating performanceStill inspect discrimination and distractors

Discrimination Index and Point-Biserial

What discrimination means

Discrimination answers: Do students who score high on the whole test tend to get this item right more often than students who score low? A good item should be answered correctly more often by stronger examinees.

Common indices:

  • Item discrimination index (D) — often based on upper vs lower scoring groups (e.g., top and bottom 27%): D ≈ (proportion correct upper) − (proportion correct lower).
  • Point-biserial correlation (r_pbis) — correlation between item score (0/1) and total test score (often total score excluding the item itself to reduce inflation).

Desirable ranges—again, guidelines

Many classroom assessment references treat discrimination above about 0.20–0.30 as desirable for typical course exams, with higher values preferred when stakes and sample size support finer interpretation. Rough interpretive language used in faculty development:

Discrimination (approx.)Common faculty language
≥0.40Excellent / very good
0.30–0.39Good
0.20–0.29Acceptable / fair—review
0.00–0.19Weak—revise or rewrite
NegativeProblematic—reverse relationship

Negatively discriminating items are answered correctly more often by lower scorers than by higher scorers. Causes include wrong answer key, ambiguous wording that “tricks” high performers, multiple defensible answers, or content that rewards testwise guessing over knowledge. Do not keep negatively discriminating items without urgent review; they actively harm score validity.

Relationship of difficulty and discrimination

Items that are extremely easy or extremely hard have limited room to discriminate (almost everyone right or wrong). Moderate difficulty often supports higher discrimination if the item is well written and aligned to taught outcomes. Do not chase a p-value of 0.50 while ignoring content fidelity.

Distractor Analysis

For multiple-choice items, distractor analysis examines how often each incorrect option is chosen—overall and by high vs low scorers.

Distractor findingMeaningAction
Never or almost never chosenNon-functioning distractorRewrite to be plausible for those who lack the concept
Chosen mainly by low scorersFunctioning distractorKeep if content-valid misconception
Chosen mainly by high scorersAmbiguity, dual key, or keyed errorFix stem/options/key immediately
Even spread across wrong options with low discriminationConfusing item or under-taught contentRevise + consider reteach
One wrong option absorbs almost all errorsCommon misconceptionTeaching opportunity; may keep if intentional

Plausible distractors based on real clinical misconceptions both improve discrimination and generate formative teaching data. Implausible distractors waste option slots and inflate p-values without measuring understanding.

Test-Level Reliability: KR-20 and Cronbach’s Alpha

Concepts

  • Reliability ≈ consistency of scores (would similar results occur under similar conditions?).
  • KR-20 (Kuder–Richardson 20) estimates internal consistency for dichotomously scored tests.
  • Cronbach’s alpha is a more general internal-consistency coefficient used for multi-item scales and often reported for exams as well.

For many classroom tests, faculty handbooks and measurement texts often cite α or KR-20 ≥ 0.70 as a reasonable target for low-to-moderate stakes course exams, with higher values preferred as stakes rise. Certification and licensure exams aim much higher; classroom tests rarely match that bar and should not be held to the same absolute standard.

Factors that raise or lower reliability estimates

Tends to increase reliability estimateTends to decrease it
More quality itemsVery short tests
Adequate score varianceHomogeneous performance (everyone similar)
Clear, discriminating itemsAmbiguous, negatively discriminating items
Consistent scoringInconsistent partial-credit scoring
Appropriate difficulty mixFloor/ceiling effects

Limits of reliability numbers

  • High reliability ≠ validity. A consistently measured wrong construct is still wrong.
  • Small class sizes produce unstable item stats—interpret cautiously; look for patterns across offerings.
  • Mastery tests with little score variance can show low reliability estimates even when teaching succeeded.
  • Do not “fix” reliability by adding redundant easy items that do not map to outcomes.

Master Test Analysis Workflow

Use a repeatable faculty process after each major exam:

  1. Before scoring decisions finalize: Confirm keys, version codes, and any invalidated items (typos discovered during administration).
  2. Generate item statistics: p-values, discrimination (point-biserial or D), distractor frequencies, and test reliability if available from the LMS/testing platform.
  3. Flag items: Negative discrimination; extreme difficulty inconsistent with intent; non-functioning distractors; items with student challenges during the exam.
  4. Content expert review: Does the item still match the blueprint and taught content? Is the key correct? Is wording biased or ambiguous?
  5. Decide per item: Keep as-is; minor edit for next form; major rewrite; retire; or accept for mastery reasons with documentation.
  6. Score policy decision: For clearly flawed items on this administration, faculty/policy may credit all students, drop the item, or accept dual keys—document rationale for fairness and consistency.
  7. Close the loop to teaching: Cluster weak items by content area → plan reteach, future emphasis, or clinical application practice.
  8. Archive: Store analysis with the exam form for curriculum review and accreditation evidence.
FlagPrefer action
Wrong key confirmedRescore; fix bank
Negative discrimination + ambiguityDrop or dual-credit this term; rewrite
Low p + low discrimination + not taughtDo not punish students; fix curriculum/teaching
Low p + high discrimination + was taughtHard but fair—keep or slightly revise; consider more practice next term
High p + low discrimination + critical safetyAcceptable mastery item—keep
High p + low discrimination + trivial contentReplace with better-sampled outcome

Connecting Analysis to Fairness and Legal Defensibility

Item analysis supports due process and fairness: students deserve instruments that measure published outcomes, not faculty “gotchas.” Documented review processes also support program evaluation and continuous quality improvement. On the CNE, options that say “ignore stats and keep tradition” or “curve everything without examining items” are weak educator practice.

Common CNE Traps

TrapWhy it failsBetter move
Keep bad items foreverAccumulates invalid varianceRetire/rewrite on evidence
Treat 0.30–0.70 or α≥0.70 as absolute lawContext (mastery, N, stakes) mattersInterpret with purpose and content
Only look at total mean scoreHides item-level failureFull item + distractor review
Drop every hard item after student complaintMay gut rigorous outcomesUse discrimination + content review
Assume high α means “good test”Reliability ≠ validity/alignmentCheck blueprint and outcomes
Never act on analysisData theaterRevise items and teaching

Bottom Line for Domain 3 Task F (Analysis)

Read difficulty, discrimination, reliability, and distractors as decision tools. Target ranges commonly cited for classroom tests—p roughly 0.30–0.70, discrimination often >0.20–0.30, KR-20/α often ≥0.70—are guidelines. Negative discriminators and non-functioning distractors demand action. Master the workflow: analyze → interpret with content expertise → revise instruments and instruction.

Test Your Knowledge

A unit exam item has a p-value of 0.22 and a point-biserial of −0.18. Student comments and faculty review show two options could be defended as correct. What is the best immediate faculty response?

A
B
C
D
Test Your Knowledge

For a typical multi-topic classroom nursing exam intended to differentiate student achievement, which item difficulty (p-value) range is most often cited as generally desirable?

A
B
C
D
Test Your Knowledge

An item’s correct answer is chosen mostly by low scorers, while high scorers prefer a particular distractor. Which interpretation is most accurate?

A
B
C
D
Test Your Knowledge

A skills-theory quiz yields Cronbach’s alpha of 0.72 in a class of 48 students. Which statement best reflects CNE-level interpretation?

A
B
C
D