7.1 Assessment Typology: Universal Screening, Diagnostic, Progress Monitoring & Summative

Key Takeaways

  • Universal screening is administered to all students three times per year (fall, winter, spring) to identify at-risk learners early, relying on cut scores, high sensitivity to minimize false negatives, and high specificity to minimize false positives.
  • Diagnostic assessments are targeted, deep-dive instruments administered to at-risk students to pinpoint fine-grained component deficits in phonemic awareness, phonics, fluency, vocabulary, or comprehension to drive explicit instruction.
  • Progress monitoring utilizes frequent, brief, standardized Curriculum-Based Measurement (CBM) probes administered weekly or bi-weekly to plot trendlines against aimlines and apply standardized decision rules (such as the four-point rule).
  • Summative assessments are cumulative, retrospective evaluations administered at the end of instructional units or school years (e.g., Georgia Milestones) to evaluate grade-level standards mastery and curriculum efficacy rather than guide immediate lesson adjustments.
Last updated: September 2026

7.1 Assessment Typology: Universal Screening, Diagnostic, Progress Monitoring & Summative

GACE Blueprint Focus: Objective 0005 (Subarea II: Assessment, Evaluation, Curriculum, and Instruction - 33% weight) assesses an educator's ability to select, administer, interpret, and act upon diverse reading assessments. Candidates must master the four core assessment typologies—Universal Screening, Diagnostic Assessment, Progress Monitoring, and Summative/Outcome Evaluation—and understand how to integrate their data within a Multi-Tiered System of Supports (MTSS) framework.


The Assessment Hierarchy in Reading Education

In a scientifically grounded literacy framework, assessment is not an isolated, periodic event used merely to assign grades. Rather, assessment is an iterative, continuous decision-making architecture that fuels data-driven instruction. The Science of Reading (SoR) and the Multi-Tiered System of Supports (MTSS) framework categorize reading assessments into four distinct functional typologies based on their primary instructional purpose, target population, administration frequency, and psychometric construction:

  1. Universal Screening: Broad-gauge, brief assessments administered to the entire student cohort to identify learners at risk for reading failure before severe academic difficulties manifest.
  2. Diagnostic Assessment: In-depth, component-specific instruments administered to identified at-risk students to pinpoint exact sub-skill deficits and establish explicit instructional targets.
  3. Progress Monitoring: Brief, frequent, repeatable standardized probes administered to students receiving targeted or intensive interventions to evaluate real-time responsiveness and inform instructional adjustments.
  4. Summative / Outcome Assessment: Comprehensive, cumulative evaluations administered at the conclusion of an instructional cycle (unit, semester, or academic year) to determine overall mastery of grade-level reading standards.
+-------------------------------------------------------------------------+
|                 THE FOUR READING ASSESSMENT TYPOLOGIES                  |
+-------------------------------------------------------------------------+
|                                                                         |
|  1. UNIVERSAL SCREENING     --> All Students | 3x Per Year (Benchmark)  |
|     "Who is at risk?"       --> DIBELS, Acadience, AIMSweb              |
|                                                                         |
|  2. DIAGNOSTIC ASSESSMENT   --> Flagged Students | Post-Screening       |
|     "What exact sub-skills  --> PAST, CTOPP-2, Core Phonics Survey      |
|      need intervention?"                                                |
|                                                                         |
|  3. PROGRESS MONITORING     --> Tier 2/3 Students | Weekly / Bi-Weekly  |
|     "Is the intervention    --> CBM Probes (ORF, NWF), Aimlines,        |
|      accelerating growth?"      Trendlines, 4-Point Decision Rule       |
|                                                                         |
|  4. SUMMATIVE ASSESSMENT    --> All Students | End of Unit / Year       |
|     "Did students master    --> Georgia Milestones (GMAS EOG/EOC),      |
|      the state standards?"      Cumulative End-of-Course Exams          |
+-------------------------------------------------------------------------+

1. Universal Screening: Early Identification and Risk Prevention

Universal screening serves as the primary prevention mechanism in literacy education. Grounded in the public health model of disease prevention, universal screening operates on the premise that identifying reading vulnerabilities in kindergarten and first grade is exponentially more efficacious and cost-effective than attempting remediation in third or fourth grade after the "Matthew Effect" (Stanovich, 1986)—the phenomenon wherein proficient readers read more and widen the gap while struggling readers read less and fall further behind—has taken hold.

Administration Protocol

  • Target Population: Universal; administered to 100% of enrolled students across general education, special education, and English learner populations.
  • Frequency: Administered three times per academic year during designated benchmark windows: Fall (beginning of year - BOY), Winter (middle of year - MOY), and Spring (end of year - EOY).
  • Structure: Extremely brief (typically 1 to 3 minutes per probe), standardized, individually administered General Outcome Measures (GOMs) or computer-adaptive assessments.
  • Prominent Instruments: Dynamic Indicators of Basic Early Literacy Skills (DIBELS 8th Edition), Acadience Reading, AIMSweb Plus, and FastBridge.

Psychometric Mechanics: Cut Scores and Thresholds

Universal screening tools rely on empirically validated cut scores (benchmark thresholds) derived from large national normative samples. These cut scores stratify students into distinct risk classifications:

  • Well Above Benchmark / Negligible Risk: Expected to achieve grade-level reading standards with core Tier 1 instruction alone (typically scoring at or above the 60th percentile).
  • At Benchmark / Low Risk: Adequate foundational skills; high probability of meeting grade-level standards with robust Tier 1 instruction (typically 40th to 59th percentile).
  • Below Benchmark / Moderate Risk (Some Risk): Requires supplemental, targeted Tier 2 small-group intervention in addition to Tier 1 core instruction (typically 20th to 39th percentile).
  • Well Below Benchmark / High Risk (At Risk): Demonstrates severe foundational deficits; requires immediate, intensive Tier 3 intervention and comprehensive diagnostic follow-up (typically below the 20th percentile).

Sensitivity vs. Specificity in Screening

A high-quality universal screener must exhibit balanced psychometric classification accuracy, governed by two fundamental statistical properties:

  • Sensitivity (True Positive Rate): The ability of the screener to correctly identify students who are genuinely at risk for reading failure or dyslexia. A test with high sensitivity produces very few false negatives (students who are at risk but pass the screener undetected). In early literacy screening, high sensitivity is paramount: failing to detect a dyslexic learner in kindergarten deprives that child of critical neurodevelopmental intervention windows.
  • Specificity (True Negative Rate): The ability of the screener to correctly identify typically developing students as not at risk. A test with high specificity produces very few false positives (students who are typically developing but erroneously flagged as at risk). High specificity ensures that school intervention resources (reading specialists, Tier 2 intervention slots) are not overburdened by non-at-risk students.

Sensitivity=True Positives (Identified At-Risk)True Positives+False Negatives (Missed At-Risk)\text{Sensitivity} = \frac{\text{True Positives (Identified At-Risk)}}{\text{True Positives} + \text{False Negatives (Missed At-Risk)}}

Specificity=True Negatives (Identified Typical)True Negatives+False Positives (Over-Identified)\text{Specificity} = \frac{\text{True Negatives (Identified Typical)}}{\text{True Negatives} + \text{False Positives (Over-Identified)}}


2. Diagnostic Assessments: Pinpointing Underlying Component Deficits

While universal screeners act as a "thermometer" signaling that a student has an academic "fever," they do not specify the underlying pathogen. A universal screener might show that a second-grade student scored in the 8th percentile on an Oral Reading Fluency probe. However, that low score could stem from weak phonemic awareness, deficient letter-sound correspondence, an inability to decode multisyllabic words, slow sight word recognition, or limited vocabulary knowledge. Diagnostic assessments provide the deep, granular inventory required to isolate the exact cognitive root cause.

Administration Protocol

  • Target Population: Sub-cohort of students flagged by universal screening as "Below Benchmark" or "Well Below Benchmark," or students who demonstrate inadequate responsiveness to Tier 2 intervention.
  • Frequency: Administered post-screening prior to intervention design, and selectively updated when a student reaches a plateau or demonstrates non-responsiveness.
  • Structure: In-depth, criterion-referenced or norm-referenced diagnostic batteries consisting of untimed, discrete subtests that evaluate specific micro-skills along the developmental reading trajectory.

Prominent Reading Diagnostic Instruments

Assessment InstrumentPrimary Reading Domain AssessedTarget Micro-Skills / Subtests
PAST (Phonological Awareness Screening Test - David Kilpatrick)Advanced Phonemic AwarenessSyllable levels, onset-rime, basic phoneme segmentation, and advanced phonemic manipulation (deletion, substitution) with automaticity timing.
CTOPP-2 (Comprehensive Test of Phonological Processing)Phonological Processing BatteryPhonological Awareness (elision, blending), Phonological Memory (memory for digits, non-word repetition), and Rapid Automatized Naming (RAN digits/letters).
Core Phonics Survey (Consortium on Reaching Excellence)Phonics & Structural Word AnalysisAlphabet knowledge, letter sounds, short vowels, consonant blends, digraphs, vowel teams, r-controlled vowels, and multisyllabic decoding.
Quick Phonics Screener (QPS - Jan Hasbrouck)Phonics & Word IdentificationGraded phonetic word lists progressing systematically from closed CVC syllables through complex multisyllabic morphemic structures.
Diagnostic Reading Assessment (DRA / DRP)Holistic Diagnostic Reading ProfileOrthographic decoding, oral reading fluency, retellings, and guided comprehension checks across graduated textual levels.

Informing Exact Instructional Targets

Diagnostic assessment data translates directly into instructional objectives. For example, if a diagnostic phonics screener reveals that a student decodes 10/10 closed CVC syllables (e.g., mat, bed, sip) but scores 1/10 on words containing initial consonant blends (e.g., flag, step, drum), the literacy educator does not subject the student to broad phonics drills. Instruction is surgically targeted at blending four-phoneme CCVC and CVCC words using explicit, multi-sensory Elkonin sound boxes and magnetic letter tiles.


3. Progress Monitoring: Curriculum-Based Measurement (CBM) and Decision Rules

Progress monitoring is the engine of the MTSS framework. Once an educator implements an evidence-based reading intervention, progress monitoring provides ongoing, empirical verification of whether the intervention is closing the achievement gap. Rather than waiting months for the next seasonal universal screening window, educators utilize frequent, standardized probes to monitor learning trajectories in real time.

Curriculum-Based Measurement (CBM)

Progress monitoring relies on Curriculum-Based Measurement (CBM) (Deno, 1985). Unlike teacher-made quizzes that evaluate mastery of an isolated, short-term classroom skill, CBMs are General Outcome Measures (GOMs) that evaluate overall proficiency in the comprehensive annual reading curriculum. Key properties of CBM probes include:

  • Standardized Administration and Scoring: Strict administrative protocols, verbatim scripted directions, and rigorous timing rules (typically 1-minute probes).
  • Alternate-Form Equivalence: Multiple parallel versions of equal difficulty and readability to allow repeated administration without practice or memory effects.
  • High Reliability and Predictive Validity: Consistently correlates strongly with performance on state summative reading tests.
  • Common CBM Probes: Oral Reading Fluency (ORF / PRF), Nonsense Word Fluency (NWF - Correct Letter Sounds and Whole Words Read), and Word Identification Fluency (WIF).

Monitoring Frequency within MTSS Tiers

  • Tier 1 (Core General Education): Monitored via the three seasonal universal screening benchmark windows (fall, winter, spring).
  • Tier 2 (Targeted Small-Group Intervention): Monitored bi-weekly (every two weeks) using grade-level or instructional-level CBM probes.
  • Tier 3 (Intensive, Individualized Intervention): Monitored weekly using standardized CBM probes to ensure immediate responsiveness to specialized instruction.

Graphing Dynamics: Baselines, Aimlines, and Trendlines

To interpret CBM data objectively, educators plot student scores on a standardized progress monitoring chart:

  1. Establishing the Baseline: Before intervention begins, the educator administers three alternate CBM probes across several days. The median score (not the mean) of these three probes establishes the baseline starting point, neutralizing single-day performance anomalies.
  2. Setting the Goal and Aimline: The educator establishes an end-of-intervention goal (e.g., reaching 75 WCPM by Week 16). A straight line—the aimline (goal line)—is drawn connecting the baseline median score to the targeted goal score. The slope of the aimline represents the expected Rate of Improvement (ROI) per week.
  3. Plotting Data and Calculating the Trendline: Each week, a new probe is administered and plotted. After collecting 6 to 8 data points, a trendline is fitted to the data (using the split-middle technique or linear regression) to represent the student's actual growth trajectory.
  WCPM
   90 |                                                    [GOAL: 75]
   80 |                                                    / *
   70 |                                              *   /   *
   60 |                                          *     /
   50 |                                      *       /   AIMLINE (Expected)
   40 |                              *   *         /
   30 |                      *   *               /       * = Actual Data Points
   20 |              *   *                     /
   10 | [BASELINE] /                         /
    0 +----+----+----+----+----+----+----+----+----+----+----+----+--> Weeks
      B1   B2   B3   W1   W2   W3   W4   W5   W6   W7   W8   W9

Standardized Decision Rules: The Four-Point Rule

To eliminate subjective teacher bias, MTSS teams govern instructional changes through empirically validated decision rules:

  • The Four-Point Rule (Under-Performance): If four consecutive data points fall below the aimline, the intervention is deemed ineffective in its current form. The educator must immediately execute an instructional modification (e.g., increase intervention frequency/time, reduce group size, change the instructional program, or conduct further diagnostic assessment).
  • Accelerated Growth (Goal Exceeded): If four consecutive data points fall consistently above the aimline, the student is outperforming expectations. The team should either increase the end-of-year goal to accelerate expectations or initiate a planned step-down to a less restrictive intervention tier.
  • Parallel Growth (Hovering on Aimline): If data points fluctuate closely around the aimline, the intervention is working as intended; maintain current instruction without alteration.

4. Summative / Outcome Assessments: Cumulative Standards Mastery

Summative assessments (also termed outcome assessments) are administered at the culmination of a designated instructional sequence—such as the end of a thematic unit, the end of a semester, or the end of an academic school year. Their primary educational purpose is evaluative and accountability-driven: to ascertain the degree to which students have attained mastery of grade-level Georgia Standards of Excellence (GSE) or national college- and career-readiness benchmarks.

Key Characteristics

  • Target Population: Universal; all enrolled students at specified grade levels participate.
  • Frequency: Infrequent; typically administered once annually (spring state testing) or at unit terminations.
  • Standardized Accountability: Formally administered under standardized testing conditions with strict security protocols. Results are reported to state departments of education, school boards, and families.
  • Prominent Instruments: The Georgia Milestones Assessment System (GMAS) End-of-Grade (EOG) assessments in English Language Arts (grades 3–8), GMAS End-of-Course (EOC) in American Literature and Composition, and publisher-developed cumulative end-of-unit tests.

Instructional Limitations for Daily Teaching

While summative assessments are vital for institutional accountability, program evaluation, resource allocation, and curriculum validation, they possess severe limitations for daily classroom guidance:

  • Distal Feedback: Results are returned weeks or months after administration, long after the student has moved on to subsequent content or another grade level.
  • Broad Grain Size: Summative tests report broad scaled scores and performance levels (e.g., Beginning Learner, Developing Learner, Proficient Learner, Distinguished Learner), but fail to provide the granular, sub-skill diagnostic data required to plan tomorrow morning's phonics or vocabulary intervention.

Comparative Analysis of the Four Assessment Typologies

The following matrix provides a comprehensive comparison of the four reading assessment typologies evaluated on the GACE reading educator examination:

Assessment TypologyPrimary PurposeTarget PopulationAdministration FrequencyCore Assessment InstrumentsPsychometric & Design FocusInstructional Decision / Action
Universal ScreeningEarly identification of students at risk for reading failure or dyslexia; primary prevention.All enrolled students (100% cohort).3 times per academic year (Fall, Winter, Spring benchmark windows).DIBELS 8th, Acadience Reading, AIMSweb Plus, FastBridge.High sensitivity (low false negatives); high specificity; standardized benchmark cut scores.Place students into MTSS tiers (Tier 1 core vs. Tier 2/3 intervention); flag for diagnostic follow-up.
Diagnostic AssessmentPinpoint specific underlying cognitive and component deficits in reading sub-skills.Students flagged as at-risk by screeners, or non-responsive to intervention.Post-screening, prior to intervention; periodically as needed.PAST (Kilpatrick), CTOPP-2, Core Phonics Survey, Quick Phonics Screener (QPS).Criterion-referenced, deep-dive subtests; untimed mastery of discrete developmental micro-skills.Establish precise instructional objectives; select targeted intervention programs and multi-sensory materials.
Progress MonitoringEvaluate real-time responsiveness to intervention; track rate of improvement (ROI).Students receiving Tier 2 targeted or Tier 3 intensive interventions.Weekly (Tier 3 intensive) or Bi-weekly (Tier 2 targeted).Standardized CBM Probes (Oral Reading Fluency, Nonsense Word Fluency, Word ID).Alternate-form equivalence; General Outcome Measures (GOM); aimlines, trendlines, decision rules.Modify intervention if 4 points fall below aimline; raise goal or step down tier if 4 points exceed aimline.
Summative / OutcomeMeasure cumulative mastery of grade-level reading standards; school accountability.All enrolled students at mandated grade levels.End of unit, semester, or academic school year (annual spring window).Georgia Milestones (GMAS EOG/EOC), end-of-unit comprehensive exams.Standardized scaled scores; criterion-referenced to Georgia Standards of Excellence (GSE).Evaluate overall curriculum efficacy; assign final course grades; programmatic and school-wide resource allocation.

Realistic Classroom Scenario: Diagnostic Application

The Classroom Context

Ms. Vance teaches a second-grade general education class of 22 students at Oglethorpe Elementary School in Savannah, Georgia. In late August, she administers the Fall Universal Screening battery (DIBELS 8th Edition). The screening results reveal that four students score well below the 20th percentile cut score on the Composite Score, specifically exhibiting depressed scores on both Oral Reading Fluency (ORF) and Nonsense Word Fluency (NWF).

Among these students are Jaden and Aaliyah:

  • Jaden: ORF = 14 WCPM (Cut score = 35 WCPM); NWF = 18 Correct Letter Sounds (CLS).
  • Aaliyah: ORF = 12 WCPM (Cut score = 35 WCPM); NWF = 15 CLS.

Diagnostic Investigation

Recognizing that the universal screener acts as a thermometer rather than a diagnostic blueprint, Ms. Vance avoids placing Jaden and Aaliyah into an identical generic intervention. Instead, she schedules immediate diagnostic testing during intervention blocks:

  1. Diagnostic Testing for Jaden:
    • Assessment Administered: Core Phonics Survey.
    • Findings: Jaden displays mastery of short vowels and closed CVC syllables (15/15), and initial/final consonant blends (15/15). However, he scores 0/10 on vowel-consonant-e syllables (silent e) and 1/10 on vowel digraph teams (ea, oa, ai). When prompted to assess advanced phonemic awareness via the PAST, Jaden scores 100% on phonemic substitution and deletion.
    • Diagnostic Conclusion: Jaden possesses intact, robust phonological and phonemic awareness. His reading breakdown is strictly an orthographic decoding deficit centered on advanced vowel patterns.
  2. Diagnostic Testing for Aaliyah:
    • Assessments Administered: Core Phonics Survey and the PAST (Kilpatrick).
    • Findings: On the Core Phonics Survey, Aaliyah struggles across all syllable types, including CVC words with short vowels. On the PAST, Aaliyah demonstrates profound breakdowns at the basic and advanced phonemic levels: she cannot isolate final phonemes, and when asked to say "slip" without the /l/, she responds "lip" or "sip" only after 8 seconds of struggle.
    • Diagnostic Conclusion: Aaliyah's low fluency is symptomatic of a fundamental core phonemic awareness deficit. She cannot map letters to sounds because she cannot manipulate the underlying speech sounds in spoken words.

Targeted Action Plan and Progress Monitoring

Ms. Vance differentiates their Tier 2 interventions:

  • Jaden's Intervention: Placed in a targeted group focusing on systematic, explicit orthographic mapping of vowel teams and silent e syllable patterns using decodable text practice.
  • Aaliyah's Intervention: Placed in an intensive group focusing on oral articulatory phonemic segmentation and deletion routines (Kilpatrick's one-minute drills) paired simultaneously with multi-sensory Elkonin boxes and letter-sound mapping.
  • Progress Monitoring Protocol: Ms. Vance establishes bi-weekly progress monitoring using alternate forms of DIBELS NWF and ORF. She plots baseline medians, draws aimlines targeting grade-level winter benchmarks, and applies the four-point decision rule. After six weeks, Jaden's scores track directly along his aimline (growth confirmed). Aaliyah records four consecutive scores below her aimline, prompting Ms. Vance and the MTSS team to adjust her intervention by reducing group size to 2 students and increasing phonemic manipulation frequency to daily 20-minute sessions.

Common GACE Exam Traps & Misconceptions

[!WARNING] Trap 1: Using Summative Data to Plan Daily Foundational Phonics Instruction. A perennial GACE question asks what assessment a teacher should consult to design a phonics intervention for a struggling reader. Exam distractors frequently offer state summative test reports (such as a Georgia Milestones ELA score or an aggregate Lexile band). These are incorrect. Summative scores provide a broad, retrospective outcome measure; teachers must administer or consult a diagnostic assessment (such as a phonics survey or phonemic inventory) to identify exact sub-skill gaps.

[!WARNING] Trap 2: Conflating Universal Screening Frequency with Progress Monitoring Frequency. Do not confuse the schedules of screening and progress monitoring. Universal screening is administered three times per year (fall, winter, spring) to all students. Progress monitoring is administered weekly or bi-weekly only to students receiving Tier 2 or Tier 3 interventions. An answer choice suggesting that an entire general education class should undergo weekly progress monitoring probes is incorrect and operationally unfeasible.

[!WARNING] Trap 3: Misunderstanding Sensitivity vs. Specificity (False Negatives vs. False Positives). GACE test items frequently test your understanding of screening accuracy. Remember: Sensitivity catches the students who need help (avoiding false negatives). If a screening tool has low sensitivity, struggling readers will slip through the cracks without intervention. Specificity correctly clears typically developing students (avoiding false positives). If specificity is low, too many typical readers are misidentified, overwhelming intervention staff.

[!WARNING] Trap 4: Modifying Interventions Based on a Single CBM Data Point. When a student in Tier 2 or Tier 3 scores significantly lower on a single weekly CBM probe, novice teachers often panic and immediately change the instructional program. GACE scenarios penalize this knee-jerk reaction. Standardized MTSS decision rules require three to four consecutive data points below the aimline, or a calculated trendline clearly diverging from the goal, before altering an intervention. A single drop can be caused by illness, fatigue, or momentary distraction.

Test Your Knowledge

An elementary school's literacy leadership team is selecting a new kindergarten universal screening instrument for early reading and dyslexia risk. The team is evaluating two competing instruments. Tool A boasts a 96% sensitivity rate and an 82% specificity rate, while Tool B boasts an 81% sensitivity rate and a 97% specificity rate. Why should the leadership team select Tool A for universal screening?

A
B
C
D
Test Your Knowledge

A third-grade student receiving Tier 2 reading fluency intervention participates in weekly Oral Reading Fluency (ORF) Curriculum-Based Measurement (CBM) progress monitoring. The student's baseline median was 42 WCPM, and the intervention aimline targets 66 WCPM by Week 12. During Weeks 5, 6, 7, and 8, the student's recorded scores are 46, 45, 47, and 44 WCPM, all falling substantially below the designated aimline. Based on standardized MTSS decision rules, what immediate action should the interventionist take?

A
B
C
D
Test Your Knowledge

During the fall benchmark administration, a first-grade teacher identifies that 6 out of 24 students scored 'Well Below Benchmark' on the DIBELS Nonsense Word Fluency (NWF) probe. To plan explicit, differentiated small-group instruction for these six students, what should the teacher do next?

A
B
C
D
Test Your Knowledge

Which of the following statements most accurately contrasts a summative reading assessment (such as the Georgia Milestones EOG) with a Curriculum-Based Measurement (CBM) progress monitoring probe?

A
B
C
D