10.1 Creating Assessment Methods to Evaluate Outcomes
Key Takeaways
- Build assessments backward from published learning outcomes and a content blueprint (table of specifications)—not from leftover lecture slides or last year’s test file.
- A table of specifications maps content areas and cognitive levels to item/task counts and weights so sampling is intentional and defensible.
- Weighting should reflect outcome importance and decision stakes; high-stakes progression measures need broader sampling and stronger design than low-stakes practice checks.
- Congruent creation means the method’s domain (cognitive, psychomotor, affective) and cognitive demand match the outcome—not merely that the topic “was taught.”
- CNE traps include testing trivia, oversampling easy recall, under-sampling clinical judgment, and building exams without a blueprint then defending grades by habit.
Creating Assessments Is Design Work, Not File Recycling
Domain 3 of the NLN CNE Detailed Test Blueprint—Use Assessment and Evaluation Strategies—includes tasks that move from selecting methods (covered earlier in Domain 3) to creating methods that actually measure intended outcomes. Roughly 13.8% of scored items (about 18 of 130) address assessment and evaluation overall; items on creation and congruence often present a mismatched exam or a faculty member who “writes questions from the PowerPoint” and ask what should change.
Creating an assessment method means deliberately designing the instrument, tasks, scoring rules, and sampling plan so results support a stated educational decision (formative feedback, unit grade, skills clearance, clinical pass). It is the educator’s applied form of content validity work: ensuring the test or performance task represents the construct defined by the outcomes.
Quick Answer: Start with outcomes and decision stakes → build a blueprint (table of specifications) that samples content and cognitive/performance levels → write or select tasks that match domain and level → set weights and scoring rules → pilot/review before high-stakes use. Never reverse the process by inventing items first and reverse-engineering “objectives” afterward.
Backward Design Applied to Assessment
Wiggins and McTighe’s backward design (also used in instructional design more broadly) is the CNE-friendly sequence:
- Identify desired results — course/unit outcomes; leveling to program outcomes when relevant.
- Determine acceptable evidence — what performance would convince a competent faculty member the outcome is met?
- Plan learning experiences — teaching that prepares learners for that evidence (not teaching only what is easy to test).
When faculty skip step 2 and jump from “I taught X” to “I tested Y trivia about X,” congruence collapses. Creation work is step 2 made concrete.
| Creation step | Faculty questions | Product |
|---|---|---|
| Clarify outcomes | What must learners know, do, value? At what level? | Outcome list with domain/level tags |
| Clarify decision | Formative check, grade, progression, certification of skill? | Stakes and use statement |
| Blueprint | Which content domains? How many tasks? What weight? | Table of specifications |
| Method/task design | MCQ bank, OSCE stations, rubric-scored paper, clinical tool? | Draft instrument + scoring |
| Review | Peer item review, content expert check, pilot stats if available | Revised instrument |
| Implementation rules | Time, accommodations, security, feedback plan | Administration plan |
From Outcomes to a Table of Specifications
A table of specifications (TOS)—also called a test blueprint—is a two-way grid that links content areas (rows) with cognitive or performance levels (columns) and assigns numbers of items or percentage weights to cells. It is the primary tool for intentional sampling.
Why blueprinting matters
- Prevents construct underrepresentation (critical clinical judgment barely tested; definitions dominate).
- Prevents construct-irrelevant overemphasis (one pet lecture topic eats 40% of points).
- Supports fairness arguments when students challenge grades (“the exam matched the published outcomes and weights”).
- Guides item writers and clinical tool developers toward balanced coverage.
- Documents content-related validity evidence for program review and accreditation narratives.
Example blueprint structure (cognitive unit exam)
| Content area (unit outcomes) | Remember/Understand | Apply | Analyze/Evaluate | Total items | % of exam |
|---|---|---|---|---|---|
| Fluid & electrolyte concepts | 2 | 4 | 2 | 8 | 20% |
| IV therapy & complications | 1 | 5 | 2 | 8 | 20% |
| Shock states prioritization | 0 | 4 | 6 | 10 | 25% |
| Pharmacologic support | 2 | 4 | 2 | 8 | 20% |
| Patient teaching & safety | 1 | 3 | 2 | 6 | 15% |
| Total | 6 (15%) | 20 (50%) | 14 (35%) | 40 | 100% |
Notice the deliberate tilt toward application and analysis—appropriate for nursing courses that claim clinical reasoning outcomes, and consistent with the cognitive emphasis of the CNE exam itself (application/analysis for educator competencies). A 40-item exam that is 70% pure recall would be poorly congruent with “prioritize care for the client in hypovolemic shock.”
Blueprints for non-MCQ methods
Blueprints are not only for multiple-choice tests:
| Method | Blueprint / sampling plan looks like… |
|---|---|
| Skills checkoffs | Curriculum skill map: which skills, which course, formative vs summative gate, critical safety steps |
| OSCE | Station-by-station map to outcomes (communication, assessment, skill, teaching) |
| Clinical evaluation tool | Competency dimensions with leveling (novice→competent expectations by course) |
| Paper / project | Analytic rubric criteria mapped to outcomes (EBP appraisal, writing, ethics) |
| Simulation assessment | Scenario objectives and scored behaviors mapped to clinical judgment elements |
| Portfolio | Required artifacts mapped to program outcomes with minimum evidence rules |
Weighting: Importance, Stakes, and Honesty
Weighting is how much each content area, item type, or course component contributes to a score or decision.
Principles
- Importance, not convenience — Weight what is essential for safe practice and stated outcomes, not what is easiest to auto-grade.
- Stakes proportionality — A progression-deciding clinical evaluation needs broader sampling and clearer criteria than a 5-point formative poll.
- Syllabus honesty — Published weights must match how grades are actually calculated; hidden reweighting fuels grievances.
- Avoid single-point catastrophe when avoidable — One 100% final with no earlier signal is weak pedagogy for most nursing courses (though some competency gates appropriately require demonstration).
- Component balance — Cognitive exams, skills, clinical, and professional behavior each need enough weight to matter if they are claimed outcomes—token 2% “professionalism” with no criteria is window dressing.
| Course component | Weak design | Stronger design |
|---|---|---|
| Unit exams | 90% of grade; no skills evidence | Substantial but balanced with performance |
| Skills | Optional practice only | Formative practice + required checkoff |
| Clinical | Pass/fail with vague “attitude” | Leveled tool + midterm formative conference |
| Participation | Subjective pure attendance | Criteria-based engagement or drop low-stakes points |
| Final | Surprises on untaught minutiae | Blueprinted cumulative sampling of priorities |
Content Sampling: Enough Evidence, Right Evidence
No assessment can test everything. Sampling is inevitable—and must be planned.
Sampling rules of thumb for CNE judgment
- Sample critical/safety content more densely than peripheral enrichment.
- Sample at the cognitive level of the outcome (application vignettes for apply/analyze outcomes).
- For psychomotor competence, sample enough occasions or skills that a one-lucky-day pass is less likely for high-stakes gates.
- For clinical judgment, sample across contexts (age groups, acuity, settings) over the curriculum—not one favorite disease.
- For affective/professional outcomes, sample patterns of behavior with documentation, not a single personality impression.
Under-sampling vs over-testing trivia
| Problem | Example | Fix |
|---|---|---|
| Under-sampling judgment | 50 recall items, zero prioritization | Blueprint forces apply/analyze cells |
| Trivia overload | Brand names of discontinued products; obscure statistics | Peer review: “need to know for outcome?” |
| Teaching-to-last-year’s-test | Items unchanged while curriculum updated | Annual blueprint and item retirement |
| Clinical opportunity inequity | Only some students see a skill | Sim/lab gates + equitable clinical planning |
Creating Multi-Method Evaluation Plans
Congruent creation often means a system, not one instrument:
- Formative cognitive checks blueprinted lightly for high-yield concepts.
- Summative unit/final exams with full TOS and item review.
- Skills validations with critical-step checklists and rater training.
- Clinical tool aligned to course-level expectations and program outcomes.
- Simulation for rare/high-risk judgment when clinical opportunity is thin.
- Written work only when outcomes require writing, synthesis, or EBP—not as busywork filler.
Each method should answer: Which outcome cells does this fill? If the answer is “none—we’ve always done it,” redesign or retire it.
Process Quality When Creating Methods
High-quality creation includes social and technical process:
- Peer review of items and rubrics before use
- Content expert check for accuracy and currency (guidelines change)
- Sensitivity review for bias, stereotype threat, and unnecessary cultural load
- Accommodation planning (extended time, alternate formats) designed in, not improvised under stress
- Security and integrity plans proportionate to stakes (especially online)
- Feedback design built in: when will rationales, rubric notes, or debrief occur?
Common CNE Traps
| Trap | Why it fails | Better move |
|---|---|---|
| Items written from slides only | Samples teaching artifacts, not outcome construct | Blueprint from outcomes first |
| No table of specifications | Unbalanced, indefensible sampling | TOS with content × level weights |
| Testing trivia / “gotchas” | Construct-irrelevant; destroys trust | Need-to-know peer review |
| Weighting by ease of grading | Domain underrepresentation | Weight by outcome importance |
| One method for all domains | MCQ cannot certify sterile technique alone | Multi-method map |
| Copying commercial items blindly | May not match your outcomes/level | Align purchased banks to local blueprint |
| Creating after teaching ends | Rushed, unreviewed instruments | Design evidence plan before term |
Bottom Line for Domain 3 Tasks D & E (Creation Side)
Create assessments as engineered evidence systems. Publish outcomes, build a blueprint, sample content and levels intentionally, weight what matters, and match method domain to outcome domain. On CNE items, prefer the faculty action that starts with outcomes and a table of specifications over recycling last year’s test or packing an exam with obscure facts.
A faculty member begins a new medical-surgical unit by opening last year’s 50-item exam and replacing three outdated drug names. No outcomes review or blueprint is used. Which critique best reflects Domain 3 expectations for creating assessments?
What is the primary purpose of a table of specifications (test blueprint) when creating a unit exam?
Course outcomes emphasize analyzing cues and prioritizing nursing actions for deteriorating clients, yet the new exam blueprint allocates 70% of items to definition recall and 5% to prioritization vignettes. What is the main congruence problem?
Which weighting decision best aligns with outcome importance and Domain 3 judgment for a skills-heavy fundamentals course that also claims sterile technique competence?