10.2 Item Writing Best Practices (MCQ & Alternative Formats)
Key Takeaways
- Strong selected-response items use a clear stem, a single best answer, and plausible distractors grounded in common learner errors—not trick wording or absolute language.
- Avoid classic flaws: negatives/double negatives, window dressing, all/none-of-the-above, implausible distractors, clang associations, and testing trivia.
- Clinical vignettes and data-rich stems support application and analysis; NGN-style case elements can strengthen teaching tests when aligned to clinical judgment outcomes.
- Alternative formats (select-all with care, short answer, performance rubrics, clinical rating scales) must match the outcome domain and include explicit scoring rules.
- Note: the CNE exam itself uses three-option MCQs, but faculty often write four-option student exams—item-writing principles apply in both contexts.
Item Writing Is a Core Educator Competency
Once the blueprint exists, someone must write the tasks. Domain 3 expects academic nurse educators to create methods that yield interpretable evidence—poor items produce noise, unfairness, and teaching-to-tricks. On CNE scenarios, you may be asked which revision improves an item, why a distractor fails, or which format fits an outcome.
Context note for CNE candidates: The NLN CNE examination uses three-option multiple-choice items (150 total; 130 scored) with emphasis on application and analysis. In practice, many nursing faculty still write four-option student exams. Principles below apply to both: clarity, single best answer framing, plausible options, and congruence with blueprinted levels. Do not confuse “how the CNE is written” with “how you must format every classroom quiz,” but do internalize quality standards either way.
Quick Answer: Write a focused stem that poses one problem, offer one best answer and distractors that reflect realistic misconceptions, avoid linguistic tricks and absolute terms, and prefer clinical contexts for application/analysis. For alternative formats, publish scoring rules and match format to domain.
Anatomy of a Quality Multiple-Choice Item
| Element | Standard |
|---|---|
| Stem | Clear, complete problem; most words in the stem, not options; positive wording preferred |
| Lead-in | Direct question or clear incomplete statement; one central idea |
| Correct option (key) | Unequivocally best given stem; accurate; parallel form |
| Distractors | Plausible to the underprepared; based on common errors; homogeneous with key |
| Cognitive demand | Matches blueprint cell (recall vs apply/analyze) |
| Sensitivity | Free of bias, stereotypes, unnecessary cultural or linguistic load |
Stem writing rules
- One problem per item — Do not combine unrelated decisions.
- Front-load clinical data — Age, setting, relevant history, vital signs, labs that matter; omit window dressing (irrelevant details that only inflate reading load).
- Avoid teaching in the stem — The stem assesses, it does not re-lecture a paragraph of new content not in the course.
- Prefer positive phrasing — “Which action is the priority?” beats “Which is not inappropriate?”
- If negation is essential, emphasize it carefully (and sparingly)—double negatives are almost never defensible.
- Independent of other items — Do not require Item 12’s answer to solve Item 13 (item dependence).
Option writing rules
- Single best answer — Even when more than one option has partial merit, the key is clearly best by nursing standards taught.
- Parallel grammar and length — Avoid making the key obviously longest/most detailed (clang or length cues).
- Homogeneous options — All interventions, or all assessments, or all values—not mixed types that cue the answer.
- Plausible distractors — Drawn from common student errors, partial knowledge, or priority mistakes—not absurd jokes.
- Avoid absolute terms as automatic giveaways or tricks: always, never, only, all, none—unless the science truly supports an absolute and distractors are fair.
- Avoid all-of-the-above / none-of-the-above in most local nursing exams—they reward testwiseness and complicate partial knowledge interpretation.
- Avoid overlapping options (ranges that include each other) and duplicate meaning.
- Logical order when numeric (ascending labs/doses) to avoid position patterns as the only structure—and still randomize keys across the test.
Clinical vignettes for higher cognitive levels
To hit apply/analyze blueprint cells:
- Present a client scenario with cues (data).
- Ask for priority action, best interpretation, most appropriate teaching, or next assessment.
- Align with clinical judgment language used in your curriculum (e.g., noticing, interpreting, responding, reflecting / CJMM elements) without turning every item into jargon soup.
- Include only necessary data; extra red herrings can be intentional for analysis items but must be fair and taught as a skill—not pure trickery.
Weak recall stem: “What is the normal adult serum potassium range?”
Stronger application stem: “A client receiving furosemide has muscle weakness and a serum potassium of 2.9 mEq/L. Which action is the priority?”
NGN-style case elements in teaching tests (use thoughtfully)
Next Generation NCLEX (NGN) popularized case-based item sets and partial-credit clinical judgment formats for licensure. Academic educators may adapt case-based clusters, unfolding data, and multi-step judgment tasks for course assessment when outcomes target clinical judgment. For CNE purposes:
- Use NGN-inspired design when it matches your outcomes and scoring capacity.
- Do not assume every classroom quiz must mimic licensure technology.
- Ensure faculty can score reliably; complex formats without training increase error.
- Still blueprint: a flashy case that only tests trivia is still trivia.
Classic Item Flaws (CNE Recognition List)
| Flaw | Why it hurts validity | Example pattern |
|---|---|---|
| Trick wording | Measures reading games, not nursing knowledge | Subtle double meaning; “except” buried |
| Absolute terms | Cue-sensitive; rarely true clinically | “Always,” “never” as only difference |
| Implausible distractors | Inflates difficulty discrimination poorly; lucky guessing among real options | Obviously wrong brand-new nonsense |
| Window dressing | Construct-irrelevant reading load; bias risk | Long irrelevant social history |
| All/none of the above | Testwiseness; ambiguous partial knowledge | A, B, C, all of the above |
| Negative stems | Higher misread rate | “Which is not…” repeatedly |
| Clang / grammar cues | Answer form matches stem uniquely | Stem article “an” → only option starting with vowel sound |
| Heterogeneous options | Examinees eliminate by type | Mix of drugs, labs, and phone numbers |
| Opinion without standard | No defensible key | “What would you feel is nicest?” |
| Trivia | Underrepresents important outcomes | Obsolete protocol minutiae |
| True/false overuse | 50% guessing; hard to write unambiguously | Entire exam T/F |
Alternative Formats Faculty Create
Select-all-that-apply (multiple response)—use carefully
Strengths: Can assess multiple correct actions or findings.
Risks: Higher cognitive and scoring complexity; student anxiety; if poorly written, becomes a disguised multi-true/false with harsh dichotomous scoring.
Design rules:
- Stem must clearly allow multiple selections.
- Each option independently true or false relative to the stem—no interdependent tricks.
- Prefer scoring models your LMS/policy can defend (e.g., partial credit rules published in advance).
- Limit frequency; do not convert an entire exam to select-all without blueprint rationale and student practice.
- Still avoid absolute language and trivia.
Short answer / fill-in
Best for: Specific calculations, named lab critical values when exactness matters, brief justifications.
Rules: Acceptable alternate answers list; clear units; avoid ultra-obscure phrasing; decide in advance whether spelling variants count for drug names.
Essay / constructed response
Best for: Analysis, EBP critique, ethics reasoning, teaching plans.
Requires: Analytic or holistic rubric, anchor papers, and preferably double scoring for high stakes.
Performance checklists and rubrics (psychomotor / complex performance)
| Tool | Use |
|---|---|
| Checklist | Discrete steps; critical safety items (e.g., patient ID, asepsis breaks) |
| Analytic rubric | Separate criteria (accuracy, safety, communication, timeliness) with level anchors |
| Holistic rubric | Single global rating with rich anchors—faster but less diagnostic |
| OSCE station guide | Standardized patient instructions + rater form |
Critical steps: Some skills use automatic failure for critical safety errors (e.g., contaminating sterile field, failing to identify patient). Publish these rules before the assessment.
Rating scales for clinical evaluation
Clinical tools often use leveled scales (e.g., dependent → supervised → independent; or course-specific benchmarks). Creation standards:
- Behavioral anchors for each level (what “3” looks like in observable terms)
- Same tool interpretation across faculty (calibration)
- Midterm formative use + final summative judgment
- Space for anecdotal documentation supporting ratings
- Explicit professionalism/safety criteria—not vague “attitude”
Matching, ordering, hot-spot, and tech-enhanced items
Use when they fit the outcome (e.g., sequence of triage actions; identify correct landmark on an image). Apply the same anti-flaw rules: clear directions, plausible foils, accessibility, and blueprint alignment. Technology novelty is not validity.
Item Review Process Before High-Stakes Use
- Author self-check against blueprint cell and flaw checklist.
- Peer review by another content expert.
- Sensitivity/bias review.
- Key verification against current guidelines/texts used in the course.
- After administration: item analysis (difficulty, discrimination, distractor function)—detailed in later Domain 3 sections—to revise or retire items.
Common CNE Traps
| Trap | Why it fails | Better move |
|---|---|---|
| Trick items to “separate A from B students” | Measures cynicism/testwiseness | Discriminate via important content difficulty |
| Testing trivia | Weak content validity | Need-to-know screen |
| All-of-the-above comfort blanket | Flawed measurement | Independent best-answer options |
| Negatives and double negatives | Random error | Positive stems |
| Select-all without scoring rules | Unfair, unreliable | Publish model; train students |
| Rubric-free essays/clinical ratings | Subjectivity, grievances | Anchored criteria in advance |
| Four absurd distractors | False sense of difficulty | Plausible misconceptions |
| Ignoring CNE’s three-option format awareness | Confusion on exam items about “best practice” | Know CNE format; apply principles to local four-option exams |
Bottom Line for Item & Format Creation
Write items that sample the blueprint fairly: clear stems, single best answers, realistic distractors, clinical contexts for judgment, and alternative formats with explicit scoring. Eliminate tricks and trivia. On CNE questions, choose revisions that improve clarity, plausibility, and congruence—not options that make the item “harder” by word games.
Which revision best improves a flawed classroom MCQ stem that currently reads: “Regarding heart failure, which of the following is not an unlikely finding?”
A faculty writer includes three options that are obviously impossible and one correct answer so that “everyone who studied even a little will pass the item.” What is the main psychometric/educational problem?
Which statement correctly relates CNE exam format to faculty item writing for students?
Faculty want to assess whether students can perform sterile tracheostomy suctioning safely. Which creation choice is most congruent?