9.1 Screening, Diagnostic, Progress Monitoring & Summative Assessments
Key Takeaways
- Reading assessments are systematically categorized into four distinct functional purposes: Universal Screening (early identification of risk), Diagnostic Assessment (pinpointing specific component skill deficits), Progress Monitoring (evaluating rate of growth and intervention responsiveness), and Summative Assessment (evaluating cumulative mastery against standards).
- Universal screeners are brief, standardized, administered 3x per year (Fall, Winter, Spring) to 100% of students within an MTSS/RTI framework to identify students requiring Tier 2 or Tier 3 intervention.
- Diagnostic assessments are in-depth, fine-grained, and administered post-screening to analyze specific instructional strengths and deficits to build prescriptive instructional plans.
- Progress monitoring utilizes frequent, brief, alternate-form Curriculum-Based Measurement (CBM) probes (weekly for Tier 3, bi-weekly for Tier 2) to track a student's trendline against an established aimline.
- Criterion-referenced assessments measure an individual's performance against fixed mastery standards, whereas norm-referenced assessments rank an individual's score relative to a normative peer comparison group.
9.1 Screening, Diagnostic, Progress Monitoring & Summative Assessments
Core Principle: Assessment in evidence-based reading instruction is not a single, monolithic event; it is an interconnected, data-driven cycle designed to prevent reading failure, inform targeted instruction, evaluate intervention efficacy, and measure academic achievement. On the Foundations of Reading Test (FoRT Objective 0008 & Subarea IV Foundations of Reading Case Studies), candidates must demonstrate deep mastery of the four functional purposes of reading assessment, differentiate between criterion-referenced and norm-referenced instruments, interpret growth trendlines within Multi-Tiered Systems of Support (MTSS/RTI), and select appropriate assessment tools matched to instructional needs.
The Four Purposes of Reading Assessment
In a scientifically based reading framework, assessments are classified by their primary instructional purpose. Using an assessment for a purpose other than that for which it was validated leads to inaccurate instructional decisions and misallocated resources.
┌────────────────────────────────────────────────────────────────────────┐
│ THE FOUR PURPOSES OF READING ASSESSMENT │
├───────────────────┬────────────────────────────────────────────────────┤
│ 1. Universal │ Brief, standardized battery administered 3x/year │
│ Screening │ (Fall, Winter, Spring) to identify at-risk students│
├───────────────────┼────────────────────────────────────────────────────┤
│ 2. Diagnostic │ In-depth, fine-grained assessment administered to │
│ Assessment │ pinpoint exact skill deficits and guide instruction│
├───────────────────┼────────────────────────────────────────────────────┤
│ 3. Progress │ Brief, frequent probes (weekly/bi-weekly) tracking │
│ Monitoring │ rate of improvement (trendline vs. goal line) │
├───────────────────┼────────────────────────────────────────────────────┤
│ 4. Summative │ Cumulative evaluation administered at unit/year-end │
│ Assessment │ to measure mastery against grade-level standards │
└───────────────────┴────────────────────────────────────────────────────┘
1. Universal Screening (Risk Identification)
- Primary Purpose: To proactively identify students who are at risk of reading difficulty or failure before severe academic deficits emerge, and to evaluate the overall health of the core Tier 1 reading curriculum.
- Administration Cadence: Administered to 100% of students three times per school year: Beginning of Year (BOY / Fall benchmark), Middle of Year (MOY / Winter benchmark), and End of Year (EOY / Spring benchmark).
- Core Characteristics:
- Brief and Efficient: Typically takes 1 to 3 minutes per student per subtest.
- Standardized and Normed: Rigidly timed, standard administration protocols, standardized scoring rubrics, and research-validated national benchmark cut scores.
- High Sensitivity: Prioritizes identifying all potentially at-risk students (minimizing false negatives), ensuring that struggling readers receive immediate diagnostic attention.
- Exemplar Instruments: DIBELS 8th Edition (Dynamic Indicators of Basic Early Literacy Skills), Acadience Reading, AIMSweb Plus, easyCBM, FastBridge.
- Instructional Decision-Making: If fewer than 80% of all students meet Tier 1 benchmark cut scores, the core Tier 1 general education curriculum and delivery require structural enhancement. Individual students scoring below the 20th–25th percentile are flagged for immediate diagnostic follow-up and Tier 2/Tier 3 intervention.
2. Diagnostic Assessment (Pinpointing Instructional Deficits)
- Primary Purpose: To identify specific instructional strengths, skill deficits, and underlying cognitive-linguistic weaknesses in students identified as at-risk by universal screening.
- Administration Cadence: Administered post-screening upon identification of risk, for newly enrolled struggling readers, or when a student demonstrates non-responsiveness to intervention.
- Core Characteristics:
- In-Depth and Granular: Evaluates discrete, component-level reading sub-skills across the foundational hierarchy (e.g., segmenting 4-phoneme blends, decoding closed vs. silent-e syllables, distinguishing inflectional from derivational morphemes).
- Untimed or Detailed Qualitative Scoring: Focuses on how and why the student makes errors rather than pure processing speed.
- Instructionally Actionable: Provides the specific blueprint for differentiated small-group instruction or Tier 2/3 intervention planning.
- Exemplar Instruments: Phonological Awareness Screening Test (PAST by David Kilpatrick), CORE Phonics Survey, Really Great Reading Diagnostic Decoding Surveys, Informal Decoding Inventories, Diagnostic Reading Scales, Qualitative Reading Inventory (QRI).
- Instructional Decision-Making: A teacher does not simply conclude 'this student struggles with decoding'; the diagnostic assessment pinpoints that the student has mastered initial consonants and short vowels but fails on r-controlled vowels and consonant digraphs, directly dictating the instructional entry point.
3. Progress Monitoring (Evaluating Growth & Intervention Responsiveness)
- Primary Purpose: To evaluate a student's rate of growth (Rate of Improvement / ROI), determine whether an intervention is working in real time, and inform timely pedagogical adjustments before months of instructional time are lost.
- Administration Cadence:
- Tier 2 (Targeted Intervention): Administered every 1 to 2 weeks.
- Tier 3 (Intensive Intervention): Administered weekly.
- Tier 1 (Borderline / Monitor): Administered monthly.
- Core Characteristics:
- Curriculum-Based Measurement (CBM): Uses brief (1-minute), standardized, alternate equivalent forms of equal difficulty that sample the entire annual curriculum domain.
- Plotted on a Visual Graph: Performance is plotted along an Aimline (Goal Line) connecting baseline performance to the end-of-year benchmark target.
- Trendline Analysis: A calculated trendline represents the student's actual trajectory of progress.
- Decision Rules for Progress Monitoring (The 3–4 Data Point Rule):
- Below the Aimline: If 3 to 4 consecutive data points fall below the aimline (or if the slope of the trendline is shallower than the aimline), the intervention is deemed insufficient. The teacher must modify instruction: increase dosage (time/frequency), reduce small-group size, enhance explicitness, or change the instructional program.
- Above the Aimline: If 3 to 4 consecutive data points fall significantly above the aimline, the student has accelerated their growth. The teacher can adjust the goal upward or plan the systematic fading of intervention back into Tier 1.
4. Summative Assessment (Evaluating Mastery Against Standards)
- Primary Purpose: To evaluate cumulative student learning, measure mastery of grade-level reading standards, determine program efficacy, and provide accountability data to stakeholders.
- Administration Cadence: Administered at the conclusion of a defined instructional block (e.g., end-of-unit basal tests, end-of-semester exams, annual state accountability assessments).
- Core Characteristics:
- Cumulative and Broad: Covers a wide spectrum of grade-level standards (phonics, vocabulary, literary analysis, informational text comprehension).
- Evaluative Rather Than Formative: Evaluates what was learned over a past period; too distal in time to inform immediate, daily instructional adjustments for that cohort.
- Exemplar Instruments: State End-of-Grade Reading Assessments, End-of-Unit Basal Reading Tests, Stanford Achievement Test (SAT-10), Iowa Assessments.
Master Reference Table: The Four Assessment Purposes
| Assessment Type | Primary Question Answered | Target Population | Typical Frequency | Nature of Measures | Action Triggered by Results |
|---|---|---|---|---|---|
| Universal Screening | "Which students are at risk of reading failure?" | 100% of all students | 3x per year (Fall, Winter, Spring) | Brief (1-3 min), standardized, highly sensitive CBM probes | Allocate students to Tier 1, Tier 2, or Tier 3 MTSS; trigger diagnostic assessment. |
| Diagnostic Assessment | "What specific skill deficits are causing the struggle?" | Flagged at-risk students & non-responders | As needed (post-screening or intake) | Untimed, in-depth, granular sub-skill checklists & inventories | Design prescriptive small-group instruction and select specific intervention curricula. |
| Progress Monitoring | "Is the student making sufficient growth to reach the benchmark?" | Students in Tier 2 & Tier 3 interventions | Weekly (Tier 3) to Bi-weekly (Tier 2) | Brief (1 min), standardized alternate equivalent CBM forms | Adjust intervention intensity, change instructional strategy, or fade support if goal met. |
| Summative Assessment | "Has the student mastered grade-level standards?" | 100% of all students | End of unit, semester, or school year | Comprehensive, cumulative, high-stakes standardized tests | Grade assignment, state accountability, curriculum evaluation, annual placement. |
Criterion-Referenced vs. Norm-Referenced Assessments
A critical competency on the Foundations of Reading Test is distinguishing between the psychometric frameworks underlying test construction and score interpretation.
┌────────────────────────────────────────────────────────────────────────┐
│ ASSESSMENT MEASUREMENT FRAMEWORKS │
├───────────────────┬────────────────────────────────────────────────────┤
│ CRITERION-REFERENCED │ NORM-REFERENCED │
├───────────────────┼────────────────────────────────────────────────────┤
│ • Compares performance against a │ • Compares performance against a │
│ predetermined standard/mastery │ representative normative sample │
│ • Score expressed as percentage, │ • Score expressed as Percentile │
│ mastery level, or benchmark │ Rank, Stanine, or NCE │
│ • Answers: "What can the student │ • Answers: "Where does the student │
│ actually do?" │ rank compared to peers?" │
│ • Diagnostic & classroom-focused │ • Program eligibility & placement │
└───────────────────┴────────────────────────────────────────────────────┘
1. Criterion-Referenced Assessments
- Core Mechanism: An individual student's performance is measured against a fixed, predetermined objective, standard, or criterion of performance. The performance of other students has zero impact on an individual's score.
- Scoring Metrics: Percentage correct (e.g., 90%), benchmark levels (Mastery / Proficient / Needs Improvement), or rubric score (e.g., 4 out of 5).
- Primary Pedagogical Value: Direct instructional utility. It tells the teacher exactly what skills the student has mastered and which skills require remediation (e.g., "Maria decoded 18 out of 20 r-controlled vowel words, demonstrating 90% mastery on this criterion").
- Typical Examples: Phonics inventory subtests, spelling mastery tests, end-of-unit curriculum checks, state standards-based assessments with cut-score proficiency levels.
2. Norm-Referenced Assessments
- Core Mechanism: An individual student's performance is compared and ranked against a large, nationally representative sample of students of the same age or grade level (the normative sample or norm group). The test is designed to yield a bell curve (normal distribution) of scores.
- Key Scoring Metrics:
- Percentile Rank (PR): A score from 1 to 99 indicating the percentage of students in the norm group who scored at or below that individual's score. A student at the 65th percentile scored equal to or higher than 65% of peer test-takers. (Note: Percentile rank is NOT percentage correct.)
- Stanine (Standard Nine): A standardized scale dividing the normal distribution into 9 bands (mean = 5, standard deviation = 2). Stanines 1–3 represent below average, 4–6 represent average, and 7–9 represent above average performance.
- Normal Curve Equivalent (NCE): An equal-interval scale ranging from 1 to 99 with a mean of 50 and standard deviation of 21.06. Unlike percentile ranks, NCEs can be mathematically averaged and compared across years.
- Grade Equivalent (GE) Scores & Their Critical Misinterpretation: A GE score (e.g., 4.2) indicates that the student obtained the same raw score as an average 4th-grade student in the second month of school would have obtained on that specific test.
Crucial FoRT Warning: A GE score of 4.2 earned by a 2nd grader does NOT mean the 2nd grader is ready for 4th-grade instructional materials. It merely means the 2nd grader performed exceptionally well on 2nd-grade test items, performing as an average 4th grader would perform if given that 2nd-grade test.
- Typical Examples: Woodcock-Johnson Tests of Achievement, Wechsler Individual Achievement Test (WIAT), Kaufman Test of Educational Achievement (KTEA), Iowa Assessments, TerraNova.
Formative vs. Summative Assessment: The Instructional Distinction
| Dimension | Formative Assessment ("Assessment FOR Learning") | Summative Assessment ("Assessment OF Learning") |
|---|---|---|
| Temporal Position | Embedded during the active instructional process | Conducted after instructional cycle completion |
| Primary Function | Inform, guide, and adapt real-time instruction | Evaluate learning, assign grades, verify accountability |
| Teacher Action | Reteach, scaffold, adjust pacing, regroup students | Record grades, report to parents/state, evaluate program |
| Student Role | Receive specific, actionable feedback; self-correct | Demonstrate cumulative mastery on evaluative tasks |
| Typical Examples | Running records, exit tickets, oral checks, miscue analysis | State accountability exams, end-of-unit tests, final exams |
┌─────────────────────────────────────────┐
│ THE MTSS / RTI ASSESSMENT CONTINUUM │
└────────────────────┬────────────────────┘
│
┌────────────────────────────────┼────────────────────────────────┐
▼ ▼ ▼
┌─────────────────┐ ┌─────────────────┐ ┌─────────────────┐
│ TIER 1: CORE │ │ TIER 2: TARGETED│ │ TIER 3: INTENSIVE│
│ • 100% students │ │ • 15-20% at-risk│ │ • 5% persistent │
│ • Screen 3x/yr │ │ • Diagnostic dx │ │ • Comprehensive │
│ • Formative obs │ │ • PM bi-weekly │ │ • PM weekly │
└─────────────────┘ └─────────────────┘ └─────────────────┘
Data-Based Decision Making in MTSS / RTI
When analyzing progress monitoring data within an MTSS framework, teachers must make systematic, rule-governed decisions based on empirical trendlines:
- Establish the Baseline & Aimline: Collect 3 baseline data points, take the median, and draw a straight line (aimline) from the baseline to the target benchmark score at the end of the intervention period (e.g., 20 weeks).
- Collect Frequent Data: Administer brief CBM probes under identical standardized conditions at regular intervals (weekly/bi-weekly).
- Evaluate the Trendline:
- Trendline Steeper than Aimline: The student is accelerating faster than expected. Action: Maintain intervention or transition student toward lower tier.
- Trendline Parallel to Aimline: The student is making progress at the expected rate but not closing the gap. Action: Maintain or slightly increase intensity to catch up to grade level.
- Trendline Flatter than Aimline / 3-4 Consecutive Points Below: The intervention is failing to accelerate learning. Action: Mandatory instructional change. Teachers must systematically alter one or more dimensions: group size (e.g., reduce from 5:1 to 3:1), instructional time (e.g., increase from 20 to 30 min/day), instructional explicitness, or target prerequisite skill deficits.
At the beginning of the school year, an elementary school administers a standardized 1-minute oral reading fluency measure to all second-grade students. The primary purpose of this assessment is to:
A reading specialist conducts bi-weekly progress monitoring with a third-grade student participating in a Tier 2 phonics intervention. After six weeks, four consecutive data points fall significantly below the established aimline (goal line). According to MTSS/RTI data-based decision-making protocols, which of the following actions should the specialist take?
A second-grade student takes a standardized norm-referenced reading achievement test and receives a Grade Equivalent (GE) score of 4.5. Which of the following statements represents the correct interpretation of this score?
A universal screening assessment identifies a first-grade student as performing well below the benchmark cut score in basic early literacy. To design an effective, targeted intervention plan for this student, the reading specialist should next administer: