9.1 Assessment Types: Screening, Diagnostic, Progress Monitoring, & Outcome

Key Takeaways

  • Universal screening is administered to 100% of students three times per year (Fall/BOY, Winter/MOY, Spring/EOY) using brief (1–3 minute) standardized probes to identify students at risk of reading failure before deficits compound.
  • Diagnostic assessments are in-depth, targeted measures administered selectively to at-risk students to pinpoint precise phonological, orthographic, morphological, fluency, or comprehension deficits to design prescriptive interventions.
  • Progress monitoring employs standardized, alternate equivalent Curriculum-Based Measurement (CBM) forms administered weekly (Tier 3) or biweekly (Tier 2) to calculate the rate of improvement (slope) against an established aim line.
  • Assessment reliability measures the consistency and stability of scores (test-retest, inter-rater, internal consistency, alternate-form), whereas validity measures whether the test accurately assesses the specific intended theoretical construct.
  • Norm-referenced tests compare student performance against a national normative sample using percentiles and standard scores, whereas criterion-referenced tests evaluate mastery against predetermined objective standards and benchmark cut scores.
Last updated: August 2026

The Four Pillars of a Comprehensive Reading Assessment System

An effective, evidence-based elementary reading assessment system does not rely on a single instrument. Instead, it coordinates four distinct assessment types—universal screening, diagnostic assessments, progress monitoring, and outcome/summative assessments—to drive instructional decisions across general education and tiered intervention frameworks. Understanding the specific purpose, administrative timing, format, and clinical utility of each assessment type is essential for the Praxis 5205 examination.

┌─────────────────────────────────────────────────────────────────────────────┐
│                     FOUR PURPOSES OF READING ASSESSMENT                    │
├──────────────────────┬──────────────────────┬───────────────────────────────┤
│ Assessment Type      │ Administration Timing│ Primary Instructional Purpose │
├──────────────────────┼──────────────────────┼───────────────────────────────┤
│ Universal Screening  │ 3x / Year (BOY/MOY/EOY)│ Early identification of risk  │
│ Diagnostic           │ As needed / Selective│ Pinpoint specific skill gaps  │
│ Progress Monitoring  │ Weekly or Biweekly   │ Gauge intervention rate/slope │
│ Outcome / Summative  │ End-of-Unit / EOY    │ Evaluate standards mastery    │
└──────────────────────┴──────────────────────┴───────────────────────────────┘

1. Universal Screening Assessments

Universal screening is a brief, standardized assessment administered to 100% of students across a grade level three times per academic year: Beginning of Year (BOY / Fall), Middle of Year (MOY / Winter), and End of Year (EOY / Spring).

  • Primary Purpose: To identify students who are at risk of reading difficulties or reading failure early in the developmental trajectory, long before academic failure becomes entrenched (preventing the "wait-to-fail" model).
  • Characteristics: Brief (typically 1 to 3 minutes per probe), highly standardized, and validated against long-term literacy outcomes. Common universal screening batteries include DIBELS (Dynamic Indicators of Basic Early Literacy Skills), Acadience Reading, aimswebPlus, and FASTBridge.
  • Performance Classifications: Screening results classify students into risk tiers based on empirical cut scores:
    • Well Below Benchmark (High Risk / Tier 3 Candidate): Significant probability of reading failure without intensive, immediate intervention.
    • Below Benchmark (Some Risk / Tier 2 Candidate): Elevated probability of reading failure; requires supplemental, targeted small-group intervention.
    • At Benchmark (Low Risk / Tier 1): High probability of achieving end-of-year grade-level reading goals with core instruction.
    • Above Benchmark (Negligible Risk / Tier 1): Exceeds grade-level expectations; candidate for enrichment.

2. Diagnostic Reading Assessments

Diagnostic assessments are detailed, comprehensive, and in-depth evaluations administered selectively to students identified as at-risk by universal screening or teacher observation.

  • Primary Purpose: To pinpoint the exact phonological, orthographic, morphological, decoding, fluency, vocabulary, or comprehension deficits underlying a student's reading difficulty to design targeted, prescriptive instructional interventions.
  • Characteristics: Untimed or multi-component, highly granular, and criterion-referenced. Diagnostic tools isolate sub-skills that screening probes measure only globally.
  • Representative Diagnostic Tools:
    • Phonological/Phonemic Processing: Comprehensive Test of Phonological Processing (CTOPP-2), Phonological Awareness Screening Test (PAST).
    • Phonics and Decoding: CORE Phonics Survey, Diagnostic Assessment of Reading (DAR), Quick Phonics Screener (QPS).
    • Qualitative Spelling: Words Their Way Primary/Elementary Spelling Inventories (Bear et al.).
    • Comprehensive Inventories: Informal Reading Inventories (IRIs) such as the Qualitative Reading Inventory (QRI) or Flynt-Cooter Reading Inventory.

3. Progress Monitoring Assessments

Progress monitoring refers to the frequent, ongoing measurement of student performance to evaluate whether an instructional intervention is accelerating the student's learning trajectory toward grade-level benchmarks.

  • Primary Purpose: To evaluate student responsiveness to intervention, determine the rate of improvement (ROI) or growth slope, and guide timely adjustments to instructional strategies, group size, or intervention intensity.
  • Characteristics: Brief (1-minute probes), administered frequently (weekly for Tier 3 intensive interventions; biweekly for Tier 2 targeted interventions; monthly for borderline Tier 1 students).
  • Curriculum-Based Measurement (CBM): Progress monitoring must use standardized alternate equivalent forms of equal difficulty (e.g., standardized Oral Reading Fluency passages or Nonsense Word Fluency probes). Using the exact same passage introduces practice effects, while using unstandardized texts introduces invalid variability.

4. Outcome and Summative Assessments

Outcome assessments (or summative assessments) evaluate student achievement and standards mastery at the conclusion of an instructional cycle, such as the end of an instructional unit, semester, or academic school year.

  • Primary Purpose: To determine the extent to which students have mastered grade-level English Language Arts (ELA) standards and evaluate overall school, district, or curricular program effectiveness.
  • Characteristics: Comprehensive, broad in scope, standardized, and high-stakes. Examples include annual state-mandated reading assessments, end-of-year standardized achievement tests (e.g., Stanford Achievement Test, TerraNova), and unit summative evaluations.

Psychometric Properties of Reading Assessments

Every standardized assessment must possess robust psychometric properties to yield dependable, legally defensible instructional decisions. The two primary pillars of psychometrics are reliability and validity.

       ┌────────────────────────────────────────────────────────┐
       │           PSYCHOMETRIC INTEGRITY OF ASSESSMENTS         │
       ├──────────────────────────┬─────────────────────────────┤
       │  RELIABILITY             │  VALIDITY                   │
       │  (Consistency of Score)  │  (Accuracy of Measurement)  │
       ├──────────────────────────┼─────────────────────────────┤
       │ • Test-Retest            │ • Construct Validity        │
       │ • Inter-Rater            │ • Content Validity          │
       │ • Alternate-Form         │ • Criterion (Concurrent)    │
       │ • Internal Consistency   │ • Criterion (Predictive)    │
       └──────────────────────────┴─────────────────────────────┘

Reliability: The Consistency of Measurement

Reliability refers to the degree to which an assessment tool yields consistent, stable, and reproducible results across repeated administrations. A reliable assessment minimizes measurement error.

  1. Test-Retest Reliability: Consistency of scores when the identical assessment is administered to the same group of students on two separate occasions over a short time interval without intervening instruction.
  2. Inter-Rater / Inter-Scorer Reliability: The degree of agreement or consistency between two or more independent examiners scoring the same student performance (e.g., two teachers scoring the same oral reading fluency audio recording).
  3. Alternate-Form (Parallel-Form) Reliability: The consistency of results obtained when students complete two different, equivalent forms of the same assessment designed to measure the same construct with identical difficulty levels. This is critical for CBM progress monitoring.
  4. Internal Consistency: The degree to which individual items within a single test correlate with one another, confirming that all items measure the same unified domain (frequently indexed by Cronbach's alpha).

Validity: The Accuracy of Measurement

Validity refers to the degree to which an assessment instrument measures what it purports to measure and the appropriateness of the inferences made based on test scores. An assessment can be highly reliable (producing consistent results) without being valid (measuring the wrong construct).

  1. Construct Validity: The extent to which an assessment accurately measures the underlying theoretical psychological or linguistic construct it claims to assess (e.g., ensuring a phonemic awareness probe isolates auditory sound manipulation rather than general intelligence, expressive vocabulary, or visual memory).
  2. Content Validity: The extent to which test items representatively sample the entire domain of knowledge, skills, or state curricular standards being assessed (e.g., a 2nd-grade phonics test containing all six syllable types rather than only open syllables).
  3. Criterion-Related Validity: The degree to which scores on the assessment correlate with or predict performance on an established external criterion measure:
    • Concurrent Validity: How strongly assessment scores correlate with an established, validated benchmark test administered at the same approximate time point.
    • Predictive Validity: How accurately an early assessment score forecasts future reading achievement (e.g., how accurately Kindergarten Phoneme Segmentation Fluency scores predict Grade 1 Oral Reading Fluency performance).

Norm-Referenced vs. Criterion-Referenced Assessments

Standardized assessments are classified into two broad measurement frameworks based on how scores are interpreted and reported:

Assessment DimensionNorm-Referenced Tests (NRTs)Criterion-Referenced Tests (CRTs)
Core PurposeCompare an individual student's performance against a representative peer group (normative sample)Measure an individual student's performance against predetermined, objective performance standards or learning criteria
Score InterpretationRelative rank ordering ("How did the student perform compared to national peers of the same age/grade?")Absolute mastery ("Has the student mastered the silent-e decoding rule with 85% accuracy?")
Common Score TypesPercentile Ranks (1–99), Stanines (1–9, mean 5), Standard Scores (mean 100, SD 15), Normal Curve Equivalents (NCEs)Percentage correct, Cut scores, Benchmark performance levels (Meets, Approaching, Does Not Meet)
Curricular LinkBroad domain sampling; may not align directly with local classroom curriculumDirect alignment with specific state standards, scope and sequence, or targeted reading objectives
Primary Educational UseIdentifying outliers, determining eligibility for special education or gifted programs, program evaluationGuiding daily instructional pacing, forming skill-based small groups, determining mastery of specific phonics patterns
Representative ExamplesWoodcock-Johnson Tests of Achievement, CTOPP-2, Iowa Assessments, TerraNovaPhonics mastery checks, DIBELS benchmark cutoffs, End-of-Unit reading tests, State Criterion Assessments

Deconstructing Score Types on the Praxis 5205

  • Percentile Rank (PR): Indicates the percentage of students in the normative comparison group who scored at or below the student's score. A student at the 65th percentile scored equal to or higher than 65% of national peers. Crucial exam distinction: Percentile rank does NOT equal percentage correct.
  • Stanine (Standard Nine): A standardized score scale dividing the normal distribution into nine equal intervals with a mean of 5 and standard deviation of 2. Stanines 1–3 represent below-average performance; 4–6 represent average; 7–9 represent above-average.
  • Standard Score (SS): A transformed score with a predetermined mean (typically 100) and standard deviation (typically 15). A score of 85 is exactly one standard deviation below the mean (16th percentile), representing a common cutoff for clinical concern.
  • Grade-Equivalent Score (GE): Represents the grade and month of the normative group that earned the same raw score (e.g., 4.2 indicates 4th grade, 2nd month). Praxis Warning: GE scores are frequently misinterpreted. If a 2nd grader scores 4.2 on a 2nd-grade reading test, it does NOT mean the child can read 4th-grade instructional materials; it means the 2nd grader performed as well on 2nd-grade items as an average 4th grader would perform on that exact same 2nd-grade test.

Common Educator Misconceptions & Diagnostic Traps

Misconception 1: Using Screening Assessments for Daily Progress Monitoring

Universal screeners are broad and brief. Using a screening instrument designed for three-times-per-year administration as a weekly progress monitoring tool creates measurement instability and invalidates the norming benchmarks. Educators must use calibrated, equivalent alternate CBM progress monitoring forms.

Misconception 2: Confusing Reliability with Validity

Novice educators assume that because an assessment consistently produces identical scores (high reliability), it must be measuring the desired reading skill (high validity). For instance, an oral reading test where a student recites a memorized storybook has high test-retest reliability, but zero construct validity as an assessment of independent decoding.

Diagnostic Classroom Scenario

Scenario: During the September universal screening (BOY), a 2nd-grade student, Lucas, scores at the 14th percentile on Oral Reading Fluency (ORF) with a reading rate of 18 Words Correct Per Minute (WCPM), placing him in the "Well Below Benchmark" risk category. His teacher wants to immediately assign him to an intensive comprehension intervention.

Clinical Diagnostic Critique: The teacher made an invalid instructional leap. The universal screener successfully identified that Lucas is at high risk of reading failure, but the screener does NOT provide diagnostic information regarding why his reading rate is suppressed. Administering a targeted diagnostic battery (e.g., CORE Phonics Survey and a Qualitative Spelling Inventory) is required. If the diagnostic assessment reveals Lucas lacks mastery of consonant blends and closed syllable patterns, his low ORF score stems from word-level decoding deficits, not a linguistic comprehension deficit. Prescribing a comprehension intervention without addressing foundational decoding would fail to resolve his primary reading barrier.

Loading diagram...
Comprehensive Reading Assessment and Decision-Making Cycle
Test Your Knowledge

A 1st-grade teacher administers a 1-minute Phoneme Segmentation Fluency (PSF) probe to all students in September, January, and May. What primary assessment purpose does this practice represent?

A
B
C
D
Test Your Knowledge

Two independent reading specialists listen to an audio recording of a 3rd-grade student reading a standardized passage aloud. Both specialists record exactly 78 Words Correct Per Minute (WCPM) and classify the identical three miscue errors. This scenario directly demonstrates high:

A
B
C
D
Test Your Knowledge

When interpreting standardized assessment data to parents, which of the following statements correctly articulates the difference between norm-referenced and criterion-referenced test scores?

A
B
C
D