10.1 Creating Assessment Methods to Evaluate Outcomes

Key Takeaways

  • Build assessments backward from published learning outcomes and a content blueprint (table of specifications)—not from leftover lecture slides or last year’s test file.
  • A table of specifications maps content areas and cognitive levels to item/task counts and weights so sampling is intentional and defensible.
  • Weighting should reflect outcome importance and decision stakes; high-stakes progression measures need broader sampling and stronger design than low-stakes practice checks.
  • Congruent creation means the method’s domain (cognitive, psychomotor, affective) and cognitive demand match the outcome—not merely that the topic “was taught.”
  • CNE traps include testing trivia, oversampling easy recall, under-sampling clinical judgment, and building exams without a blueprint then defending grades by habit.
Last updated: August 2026

Creating Assessments Is Design Work, Not File Recycling

Domain 3 of the NLN CNE Detailed Test Blueprint—Use Assessment and Evaluation Strategies—includes tasks that move from selecting methods (covered earlier in Domain 3) to creating methods that actually measure intended outcomes. Roughly 13.8% of scored items (about 18 of 130) address assessment and evaluation overall; items on creation and congruence often present a mismatched exam or a faculty member who “writes questions from the PowerPoint” and ask what should change.

Creating an assessment method means deliberately designing the instrument, tasks, scoring rules, and sampling plan so results support a stated educational decision (formative feedback, unit grade, skills clearance, clinical pass). It is the educator’s applied form of content validity work: ensuring the test or performance task represents the construct defined by the outcomes.

Quick Answer: Start with outcomes and decision stakes → build a blueprint (table of specifications) that samples content and cognitive/performance levels → write or select tasks that match domain and level → set weights and scoring rules → pilot/review before high-stakes use. Never reverse the process by inventing items first and reverse-engineering “objectives” afterward.

Backward Design Applied to Assessment

Wiggins and McTighe’s backward design (also used in instructional design more broadly) is the CNE-friendly sequence:

  1. Identify desired results — course/unit outcomes; leveling to program outcomes when relevant.
  2. Determine acceptable evidence — what performance would convince a competent faculty member the outcome is met?
  3. Plan learning experiences — teaching that prepares learners for that evidence (not teaching only what is easy to test).

When faculty skip step 2 and jump from “I taught X” to “I tested Y trivia about X,” congruence collapses. Creation work is step 2 made concrete.

Creation stepFaculty questionsProduct
Clarify outcomesWhat must learners know, do, value? At what level?Outcome list with domain/level tags
Clarify decisionFormative check, grade, progression, certification of skill?Stakes and use statement
BlueprintWhich content domains? How many tasks? What weight?Table of specifications
Method/task designMCQ bank, OSCE stations, rubric-scored paper, clinical tool?Draft instrument + scoring
ReviewPeer item review, content expert check, pilot stats if availableRevised instrument
Implementation rulesTime, accommodations, security, feedback planAdministration plan

From Outcomes to a Table of Specifications

A table of specifications (TOS)—also called a test blueprint—is a two-way grid that links content areas (rows) with cognitive or performance levels (columns) and assigns numbers of items or percentage weights to cells. It is the primary tool for intentional sampling.

Why blueprinting matters

  • Prevents construct underrepresentation (critical clinical judgment barely tested; definitions dominate).
  • Prevents construct-irrelevant overemphasis (one pet lecture topic eats 40% of points).
  • Supports fairness arguments when students challenge grades (“the exam matched the published outcomes and weights”).
  • Guides item writers and clinical tool developers toward balanced coverage.
  • Documents content-related validity evidence for program review and accreditation narratives.

Example blueprint structure (cognitive unit exam)

Content area (unit outcomes)Remember/UnderstandApplyAnalyze/EvaluateTotal items% of exam
Fluid & electrolyte concepts242820%
IV therapy & complications152820%
Shock states prioritization0461025%
Pharmacologic support242820%
Patient teaching & safety132615%
Total6 (15%)20 (50%)14 (35%)40100%

Notice the deliberate tilt toward application and analysis—appropriate for nursing courses that claim clinical reasoning outcomes, and consistent with the cognitive emphasis of the CNE exam itself (application/analysis for educator competencies). A 40-item exam that is 70% pure recall would be poorly congruent with “prioritize care for the client in hypovolemic shock.”

Blueprints for non-MCQ methods

Blueprints are not only for multiple-choice tests:

MethodBlueprint / sampling plan looks like…
Skills checkoffsCurriculum skill map: which skills, which course, formative vs summative gate, critical safety steps
OSCEStation-by-station map to outcomes (communication, assessment, skill, teaching)
Clinical evaluation toolCompetency dimensions with leveling (novice→competent expectations by course)
Paper / projectAnalytic rubric criteria mapped to outcomes (EBP appraisal, writing, ethics)
Simulation assessmentScenario objectives and scored behaviors mapped to clinical judgment elements
PortfolioRequired artifacts mapped to program outcomes with minimum evidence rules

Weighting: Importance, Stakes, and Honesty

Weighting is how much each content area, item type, or course component contributes to a score or decision.

Principles

  1. Importance, not convenience — Weight what is essential for safe practice and stated outcomes, not what is easiest to auto-grade.
  2. Stakes proportionality — A progression-deciding clinical evaluation needs broader sampling and clearer criteria than a 5-point formative poll.
  3. Syllabus honesty — Published weights must match how grades are actually calculated; hidden reweighting fuels grievances.
  4. Avoid single-point catastrophe when avoidable — One 100% final with no earlier signal is weak pedagogy for most nursing courses (though some competency gates appropriately require demonstration).
  5. Component balance — Cognitive exams, skills, clinical, and professional behavior each need enough weight to matter if they are claimed outcomes—token 2% “professionalism” with no criteria is window dressing.
Course componentWeak designStronger design
Unit exams90% of grade; no skills evidenceSubstantial but balanced with performance
SkillsOptional practice onlyFormative practice + required checkoff
ClinicalPass/fail with vague “attitude”Leveled tool + midterm formative conference
ParticipationSubjective pure attendanceCriteria-based engagement or drop low-stakes points
FinalSurprises on untaught minutiaeBlueprinted cumulative sampling of priorities

Content Sampling: Enough Evidence, Right Evidence

No assessment can test everything. Sampling is inevitable—and must be planned.

Sampling rules of thumb for CNE judgment

  • Sample critical/safety content more densely than peripheral enrichment.
  • Sample at the cognitive level of the outcome (application vignettes for apply/analyze outcomes).
  • For psychomotor competence, sample enough occasions or skills that a one-lucky-day pass is less likely for high-stakes gates.
  • For clinical judgment, sample across contexts (age groups, acuity, settings) over the curriculum—not one favorite disease.
  • For affective/professional outcomes, sample patterns of behavior with documentation, not a single personality impression.

Under-sampling vs over-testing trivia

ProblemExampleFix
Under-sampling judgment50 recall items, zero prioritizationBlueprint forces apply/analyze cells
Trivia overloadBrand names of discontinued products; obscure statisticsPeer review: “need to know for outcome?”
Teaching-to-last-year’s-testItems unchanged while curriculum updatedAnnual blueprint and item retirement
Clinical opportunity inequityOnly some students see a skillSim/lab gates + equitable clinical planning

Creating Multi-Method Evaluation Plans

Congruent creation often means a system, not one instrument:

  1. Formative cognitive checks blueprinted lightly for high-yield concepts.
  2. Summative unit/final exams with full TOS and item review.
  3. Skills validations with critical-step checklists and rater training.
  4. Clinical tool aligned to course-level expectations and program outcomes.
  5. Simulation for rare/high-risk judgment when clinical opportunity is thin.
  6. Written work only when outcomes require writing, synthesis, or EBP—not as busywork filler.

Each method should answer: Which outcome cells does this fill? If the answer is “none—we’ve always done it,” redesign or retire it.

Process Quality When Creating Methods

High-quality creation includes social and technical process:

  • Peer review of items and rubrics before use
  • Content expert check for accuracy and currency (guidelines change)
  • Sensitivity review for bias, stereotype threat, and unnecessary cultural load
  • Accommodation planning (extended time, alternate formats) designed in, not improvised under stress
  • Security and integrity plans proportionate to stakes (especially online)
  • Feedback design built in: when will rationales, rubric notes, or debrief occur?

Common CNE Traps

TrapWhy it failsBetter move
Items written from slides onlySamples teaching artifacts, not outcome constructBlueprint from outcomes first
No table of specificationsUnbalanced, indefensible samplingTOS with content × level weights
Testing trivia / “gotchas”Construct-irrelevant; destroys trustNeed-to-know peer review
Weighting by ease of gradingDomain underrepresentationWeight by outcome importance
One method for all domainsMCQ cannot certify sterile technique aloneMulti-method map
Copying commercial items blindlyMay not match your outcomes/levelAlign purchased banks to local blueprint
Creating after teaching endsRushed, unreviewed instrumentsDesign evidence plan before term

Bottom Line for Domain 3 Tasks D & E (Creation Side)

Create assessments as engineered evidence systems. Publish outcomes, build a blueprint, sample content and levels intentionally, weight what matters, and match method domain to outcome domain. On CNE items, prefer the faculty action that starts with outcomes and a table of specifications over recycling last year’s test or packing an exam with obscure facts.

Test Your Knowledge

A faculty member begins a new medical-surgical unit by opening last year’s 50-item exam and replacing three outdated drug names. No outcomes review or blueprint is used. Which critique best reflects Domain 3 expectations for creating assessments?

A
B
C
D
Test Your Knowledge

What is the primary purpose of a table of specifications (test blueprint) when creating a unit exam?

A
B
C
D
Test Your Knowledge

Course outcomes emphasize analyzing cues and prioritizing nursing actions for deteriorating clients, yet the new exam blueprint allocates 70% of items to definition recall and 5% to prioritization vignettes. What is the main congruence problem?

A
B
C
D
Test Your Knowledge

Which weighting decision best aligns with outcome importance and Domain 3 judgment for a skills-heavy fundamentals course that also claims sterile technique competence?

A
B
C
D