2.3 Screening, Standardized & Non-Standardized Assessment Selection

Key Takeaways

  • Clinical screening is a rapid, preliminary process used to identify if a full occupational therapy evaluation is indicated; it cannot be used to diagnose pathology, establish functional goals, or plan interventions.
  • Standardized assessments are classified as norm-referenced (comparing individual performance against a standardized normative sample) or criterion-referenced (measuring mastery against specific functional criteria or cutoffs).
  • Essential psychometric metrics include test-retest reliability, inter-rater reliability (ICC / Cohen's kappa), internal consistency (Cronbach's alpha), validity (construct, content, criterion), sensitivity, and specificity.
  • Top-down evaluation approaches prioritize the client's occupational performance, meaningful roles, and environmental context first, whereas bottom-up approaches focus primarily on discrete client factors and body function deficits.
  • Standardized administration requires rigorous procedural fidelity, strict adherence to basal and ceiling rules, and caution regarding non-standard accommodations that invalidate normative score comparisons.
Last updated: August 2026

Screening & Standardized Assessment Selection

Core Principle: Assessment selection is an exercise in diagnostic precision and clinical reasoning. A skilled occupational therapist selects tools with proven psychometric reliability and validity, applying a top-down occupational framework that evaluates real-world participation before drilling down into underlying impairment metrics.


Clinical Screening vs. Comprehensive Evaluation

A critical distinction on the NBCOT exam is the operational boundary between a Screening and a Comprehensive Evaluation:

+-----------------------------------------------------------------------------------+
|                         SCREENING vs. COMPREHENSIVE EVALUATION                    |
|                                                                                   |
|  CLINICAL SCREENING                        COMPREHENSIVE OT EVALUATION            |
|  ------------------                        ---------------------------            |
|  * Purpose: Determine IF a full OT         * Purpose: Diagnose performance        |
|    evaluation is warranted.                  breakdowns, establish baseline,      |
|  * Scope: Brief, hands-off review of         develop individualized plan of care. |
|    records, screening tool, or brief       * Scope: Complete Occupational         |
|    observation (e.g., MoCA, TUG).            Profile, performance analysis,       |
|  * Regulatory Rule: NEVER used to            standardized testing, goal setting.  |
|    establish goals, plan treatment,        * Mandatory: Required prior to         |
|    or bill for therapy intervention.         initiating any treatment.            |
+-----------------------------------------------------------------------------------+

Roles of OTR and COTA in Screening & Evaluation:

  • Occupational Therapist Registered (OTR): Legally and professionally responsible for the entire evaluation process. The OTR determines the screening scope, selects specific assessment tools, interprets all assessment data, formulates the clinical diagnosis, and writes the intervention plan.
  • Certified Occupational Therapy Assistant (COTA): May contribute to the screening and evaluation process by gathering specific assessment data or administering standardized tools only after demonstrating service competency to the supervising OTR. The COTA cannot independently evaluate, interpret assessment results, or synthesize goals.

Psychometric Properties of Standardized Assessments

Understanding assessment psychometrics is essential for selecting valid tools and defending clinical decision-making:

                    ┌──────────────────────────────────────────────┐
                    │         PSYCHOMETRIC FOUNDATIONS             │
                    └──────────────────────┬───────────────────────┘
                                           │
                ┌──────────────────────────┴──────────────────────────┐
                │                                                     │
                ▼                                                     ▼
┌────────────────────────────────────────┐  ┌────────────────────────────────────────┐
│         RELIABILITY (Consistency)      │  │           VALIDITY (Accuracy)          │
├────────────────────────────────────────┤  ├────────────────────────────────────────┤
│ * Test-Retest: Score stability over    │  │ * Construct: Measures the theoretical  │
│   time (r ≥ 0.80 desired)              │    trait/construct it claims to measure   │
│ * Inter-Rater: Consistency across two  │  │ * Content: Items represent entire      │
│   different raters (Kappa/ICC > 0.80)  │    domain of performance                  │
│ * Intra-Rater: Consistency of same     │  │ * Criterion-Concurrent: Correlates     │
│   rater over repeated trials           │    with established gold standard test    │
│ * Internal Consistency: Item harmony   │  │ * Criterion-Predictive: Predicts future│
│   (Cronbach's alpha 0.70 - 0.95)       │    real-world outcome (e.g., fall risk)   │
└────────────────────────────────────────┘  └────────────────────────────────────────┘

Norm-Referenced vs. Criterion-Referenced Tests

DimensionNorm-Referenced AssessmentCriterion-Referenced Assessment
DefinitionCompares individual performance against a representative normative sample of peersMeasures individual performance against an established standard, cutoff score, or skill mastery criteria
Score InterpretationStandard scores, Z-scores, T-scores, percentiles, age equivalentsPass/fail, cutoff scores, percentage mastery, functional independence levels
Primary PurposeDiagnostic ranking, determining eligibility for specialized services (e.g., school-based OT)Measuring functional capability, tracking clinical progress over time, discharge readiness
NBCOT ExamplesBruininks-Oseretsky Test of Motor Proficiency (BOT-2), Peabody Developmental Motor Scales (PDMS-2)Performance Assessment of Self-Care Skills (PASS), Section GG / FIM, School Function Assessment (SFA) criterion score

Diagnostic Accuracy Metrics: Sensitivity and Specificity

  • Sensitivity (True Positive Rate): The ability of a screening tool to correctly identify individuals who have the condition (Sensitivity=TPTP+FN\text{Sensitivity} = \frac{TP}{TP + FN}). High sensitivity is vital when missing a pathology carries severe consequences (e.g., dysphagia or fall risk screen). SnNout: High Sensitivity, Negative result rules OUT the condition.
  • Specificity (True Negative Rate): The ability of a tool to correctly identify individuals who do not have the condition (Specificity=TNTN+FP\text{Specificity} = \frac{TN}{TN + FP}). High specificity ensures individuals without impairment are not subjected to unnecessary interventions. SpPin: High Specificity, Positive result rules IN the condition.

Top-Down vs. Bottom-Up vs. Contextual Assessment Frameworks

+-----------------------------------------------------------------------------------+
|                       EVALUATION PHILOSOPHY PARADIGMS                             |
|                                                                                   |
|  TOP-DOWN APPROACH (OTPF-4 Preferred)       BOTTOM-UP APPROACH (Biomedical)       |
|  ------------------------------------       -------------------------------       |
|  1. Focus on client's desired roles,        1. Focus on specific underlying       |
|     occupational performance, & context.       body structures and functions.     |
|  2. Assess actual task execution            2. Measures isolated impairments      |
|     (e.g., AMPS, PASS, Barthel, COPM).         (e.g., goniometry, MMT, visual).   |
|  3. Examine underlying client factors       3. Inactive/impaired components       |
|     ONLY as they disrupt the task.             assumed to explain disability.     |
|  * Clinical Rationale: Ensures therapy is   * Clinical Limitation: Normal ROM or  |
|    immediately relevant and functional.       strength does not equal function.   |
+-----------------------------------------------------------------------------------+

Assessment Selection Across Clinical Settings

Assessment selection must align with patient diagnosis, clinical setting, cognitive capacity, and acute medical stability:

Practice SettingClinical PrioritiesPrimary Assessment Battery
Acute Care HospitalDischarge safety, basic ADL status, acute delirium/cognition, tolerance to upright postureSection GG, AM-PAC "6-Clicks", CAM (Confusion Assessment Method), Barthel Index
Inpatient Rehabilitation (IRF)Intensive ADL/IADL recovery, motor control, executive functioningSection GG Self-Care & Mobility, Fugl-Meyer Assessment (stroke), MoCA, PASS, Executive Function Performance Test (EFPT)
Skilled Nursing Facility (SNF)Long-term functional maintenance, safe transfers, cognitive baseline for feeding/dressingSection GG, Allen Cognitive Level Screen (ACLS-5), Functional Reach Test, Katz Index of ADLs
Home HealthHome environmental safety, fall hazards, community IADLs, medication managementOASIS, SAFER-HOME, Westmead Home Safety, Timed Up and Go (TUG), Lawton IADL Scale
Outpatient Upper ExtremityBiomechanical range, strength, pinch, fine motor dexterity, edema, sensory mappingGoniometry, Hand Dynamometer, Pinch Gauge, Semmes-Weinstein Monofilaments, DASH / QuickDASH, 9-Hole Peg Test
School-Based PracticeEducational participation, peer social play, handwriting, classroom navigationSchool Function Assessment (SFA), BOT-2, Beery VMI, Sensory Processing Measure (SPM-2)

Standardized Administration Ethics: Basal and Ceiling Rules

Standardized tests enforce strict administration protocols to maintain validity against normative populations:

                        BASAL AND CEILING RULES EXAMPLE

     TEST START (Item 10 based on chronological age)
          │
          ├── Item 10: Correct [1]
          ├── Item 11: Correct [1]  <─── BASAL ESTABLISHED (e.g., 3 consecutive correct)
          ├── Item 12: Correct [1]       (All prior items 1-9 automatically scored correct)
          ├── Item 13: Correct [1]
          ├── Item 14: Incorrect [0]
          ├── Item 15: Incorrect [0]
          └── Item 16: Incorrect [0] <── CEILING REACHED (e.g., 3 consecutive incorrect)
                                         (Testing terminates immediately)

Key Testing Principles:

  1. Basal Rule: The designated entry-level score criteria that allows the examiner to assume all items prior to the basal would be performed successfully. If the client fails any initial basal item, the examiner tests backward in reverse order until the basal threshold is met.
  2. Ceiling Rule: The stopping point representing the upper limit of the client's ability. Once the ceiling threshold is reached (e.g., 3 consecutive zero scores), testing is discontinued to prevent client fatigue, frustration, and invalid measurement.
  3. Fidelity vs. Accommodations: Standard accommodations (e.g., wearing corrective eyeglasses, utilizing hearing aids) are permitted if specified in the manual. However, unstandardized modifications (e.g., rephrasing prompts, translating non-validated languages, offering extra demonstration) invalidate the standardized normative comparison; results must be documented qualitatively rather than as standard scores.

Non-Standardized (Informal) Assessment: Purpose, Advantages & Limitations

The blueprint asks for the administration, purpose, indications, advantages, and limitations of standardized and non-standardized tools. Everything above covers the standardized half. The other half is not a lesser fallback — it is a deliberate clinical choice, and NBCOT writes items in which the standardized option is the wrong answer.

A non-standardized (informal) assessment has no fixed administration protocol and no normative comparison sample. It includes structured or unstructured clinical observation of task performance in the natural context, unstructured and semi-structured interviews with the client and caregivers, chart and record review, therapist-constructed checklists, and trial use of adaptive equipment or assistive technology.

When Non-Standardized Assessment Is the Correct Choice

  • The client cannot follow a standardized protocol — severe cognitive impairment, receptive aphasia, low arousal, acute medical instability, behavioral dysregulation, or a young child who will not sit for structured testing.
  • The norms do not fit the client — the client falls outside the normative sample's age, culture, language, or diagnostic group, so the standard score would be meaningless or discriminatory.
  • The clinical question is contextual — "can this person make breakfast in their own kitchen, with their own equipment" is not answerable from a normed subtest score.
  • Time and setting constrain testing — an acute-care therapist with a 20-minute window and a discharge decision to inform observes an actual transfer rather than administering a full battery.
  • A standardized test was modified — the moment administration is altered beyond the accommodations the manual permits, the result becomes non-standardized data. Report the qualitative findings; never report standard scores or percentiles from a modified administration.

Advantages vs. Limitations

AdvantagesLimitations
Ecological validity — measures real performance in the real context, where transfer of training actually mattersNo normative comparison — cannot produce a standard score when an eligibility team, payer, or local policy requires a norm-referenced result
Flexible and gradable in real time — the therapist can cue, adapt, and probe to find why performance broke downVulnerable to examiner bias and to seeing what the therapist expects to see
Identifies the specific OTPF-4 performance skill that failed, not just that a score was lowWeak inter-rater reliability — two therapists may reach different conclusions from the same session
No basal/ceiling constraint and no fatigue from irrelevant itemsHard to quantify change precisely, which weakens progress documentation and appeals
Low cost, minimal equipment, no manual requiredNot reproducible across therapists, settings, or time points

The Exam Decision Rule

Read what the stem is asking the score to do:

  • Stem says establish eligibility, compare to same-age peers, document measurable progress for the payer, or defend the finding → choose the standardized tool.
  • Stem says the client cannot follow directions, norms do not exist for this population, or the question concerns real-world contextual performance → choose the observation-based, non-standardized option.

Best practice is not either/or. Skilled evaluation pairs a standardized measure for the defensible, comparable baseline with non-standardized observation and interview for the contextual picture that actually drives the intervention plan.

Loading diagram...
Top-Down Assessment Selection Clinical Algorithm
Test Your Knowledge

An occupational therapist in an inpatient rehabilitation facility evaluates a client who sustained a right middle cerebral artery stroke with left hemiparesis. Following a top-down evaluation framework, what action should the therapist take FIRST?

A
B
C
D
Test Your Knowledge

A rehabilitation hospital is selecting a new standardized screening tool to evaluate fall risk in community-dwelling older adults. The clinical team requires a tool with high sensitivity. What is the primary clinical advantage of selecting an assessment with high sensitivity?

A
B
C
D
Test Your Knowledge

During the administration of a standardized pediatric developmental motor test, the examiner begins testing at the child's chronological age entry point (Item 15). The child fails Item 15 and Item 16. According to standard testing administration rules, what action must the occupational therapist take next?

A
B
C
D
Test Your Knowledge

An occupational therapist must evaluate a 79-year-old client admitted with delirium superimposed on dementia to inform a discharge recommendation within the next two days. The client cannot sustain attention through multi-step directions, becomes agitated when placed at a table for structured testing, and speaks a language for which no validated translation of the department's standardized battery exists. Which evaluation approach is MOST appropriate?

A
B
C
D