2.1 Formal vs. Informal Assessments and Testing Purposes
Key Takeaways
- Formal assessments are standardized, norm- or criterion-referenced instruments with strict administration protocols used for eligibility, placement, and annual reviews, whereas informal assessments are flexible, classroom-grounded measures used to inform daily instructional adjustments.
- The special education assessment continuum encompasses four distinct testing purposes: universal screening (identifying risk early), diagnostic assessment (identifying specific processing deficits for eligibility and IEP planning), formative assessment (guiding in-process instructional adaptations), and summative assessment (evaluating cumulative mastery against standards).
- Norm-referenced tests compare an individual student's performance against a representative national peer sample to determine relative standing, whereas criterion-referenced tests measure performance against an absolute mastery benchmark or curricular objective.
- Dynamic assessment employs a test-teach-retest model grounded in Vygotsky's Zone of Proximal Development to assess cognitive modifiability and learning potential, making it invaluable for culturally and linguistically diverse learners.
- Classroom literacy diagnostics—including running records, miscue analysis, and informal reading inventories—provide critical qualitative and quantitative data regarding a student's reliance on semantic, syntactic, and graphophonic cueing systems.
2.1 Formal vs. Informal Assessments and Testing Purposes
Quick Focus: Special education teachers must select, administer, and interpret both formal standardized tests and informal classroom measures to drive the evaluation and instructional cycle under IDEA. This section examines the four foundational assessment purposes—universal screening, diagnostic evaluation, formative progress monitoring, and summative evaluation—alongside norm-referenced, criterion-referenced, and authentic assessment methodologies.
Assessment in special education is not an isolated administrative event; it is an ongoing, decision-making continuum governed by federal mandates under the Individuals with Disabilities Education Act (IDEA 34 CFR § 300.304) and state regulations such as Georgia Department of Education (GaDOE) Rule 160-4-7-.04. Every assessment administered to a student with an exceptionality must serve a defensible educational purpose, produce actionable data, and adhere to strict technical standards.
Formal vs. Informal Assessments
Educational assessments fall into two overarching categories: formal and informal. Understanding the distinct characteristics, procedural requirements, and clinical trade-offs of each is essential for both daily classroom practice and the GACE Special Education exam.
+-----------------------------------------------------------------------------------------+
| ASSESSMENT CONTINUUM |
+-----------------------------------------------------------------------------------------+
| FORMAL ASSESSMENTS INFORMAL ASSESSMENTS |
| - Standardized administration & scoring - Flexible, non-standardized administration|
| - Empirical reliability & validity data - High ecological & curricular validity |
| - Norm- or standardized criterion-referenced - Teacher-made, rubric- or task-based |
| - Purpose: Eligibility, placement, state reporting - Purpose: Day-to-day instructional adjustments|
| - Examples: WISC-V, WJ IV, Ga. Milestones - Examples: Running records, exit tickets |
+-----------------------------------------------------------------------------------------+
Formal Assessments
Formal assessments are systematically designed instruments characterized by standardized administration procedures, predetermined scoring criteria, and published technical data documenting reliability and validity. When administering a formal assessment, the examiner must adhere strictly to the standardized script, time constraints, presentation sequence, and allowable prompting rules outlined in the examiner's manual. Any deviation from these standard procedures invalidates the published norms.
- Primary Functions: Determining initial eligibility for special education services under IDEA; conducting mandatory triennial reevaluations; establishing baseline standard scores for comprehensive multidisciplinary evaluations; and evaluating district- or state-wide academic accountability (such as the Georgia Milestones Assessment System).
- Advantages: Provides objective, empirically validated comparative data; minimizes examiner subjectivity; possesses established psychometric properties; and provides legally defensible evidence for individualized education program (IEP) team decisions.
- Limitations: Typically administered in an artificial, decontextualized testing environment; sensitive to student test anxiety, fatigue, and language differences; yields minimal actionable data regarding daily instructional pacing; and may exhibit cultural or linguistic bias if the norming sample is non-representative.
Informal Assessments
Informal assessments are flexible, classroom-based procedures designed to evaluate a student's current academic performance, behavioral functioning, and learning processes in natural educational settings. Unlike formal measures, informal assessments do not rely on rigid administration scripts or standardized comparison groups.
- Primary Functions: Guiding immediate pedagogical modifications; identifying specific skill deficits; monitoring short-term mastery of IEP objectives; and providing contextualized, authentic work samples for Present Levels of Academic Achievement and Functional Performance (PLAAFP) statements.
- Advantages: High instructional utility; immediate feedback for teacher and student; easily embedded into daily classroom routines; flexible accommodation of student needs; and low-stress for the learner.
- Limitations: Susceptible to teacher bias and subjective scoring; lacks comparative national norms; cannot alone establish special education eligibility under IDEA; and often exhibits lower statistical reliability across different raters.
The Four Core Testing Purposes in P-12 Special Education
Every educational assessment serves one of four distinct purposes within the P-12 Multi-Tiered System of Supports (MTSS) and special education framework. Conflating these purposes leads to instructional failure and procedural non-compliance.
1. Universal Screening
Universal screening involves brief, low-burden assessments administered to all students within a grade or school—typically three times per academic year (fall, winter, and spring)—to identify individuals who are at risk of academic failure or behavioral difficulties. Universal screeners function like medical triage: they possess high sensitivity (correctly flagging students who need intervention) to ensure struggling learners receive immediate Tier 2 or Tier 3 support before deficits become intractable.
Examples: DIBELS 8th Edition (Dynamic Indicators of Basic Early Literacy Skills), easyCBM, AIMSweb, and broad-band behavioral screeners such as the SRSS (Student Risk Screening Scale).
2. Diagnostic Assessment
While screening answers who is struggling, diagnostic assessment answers why the student is struggling and what specific underlying deficits are causing the difficulty. Diagnostic evaluations conduct a comprehensive, deep-dive examination of an individual student's cognitive processing, phonological awareness, receptive/expressive language, memory, fine motor control, or specific academic micro-skills.
Special Education Application: Diagnostic assessments are indispensable during the initial comprehensive evaluation process to establish IDEA Part B eligibility (e.g., diagnosing a Specific Learning Disability in basic reading skills) and to formulate precise baseline data for drafting measurable annual IEP goals.
3. Formative Assessment (Progress Monitoring)
Formative assessment is an active, in-process evaluation administered frequently—weekly, bi-weekly, or during daily lessons—to monitor student learning and inform ongoing instruction. Formative assessments provide real-time feedback that enables educators to adapt instructional delivery, adjust scaffolding, reteach prerequisite skills, or accelerate pacing while learning is occurring.
Special Education Application: Curriculum-Based Measurement (CBM), such as weekly oral reading fluency or math calculation probes plotted against an aim line, is the gold standard for formative evaluation in special education.
4. Summative Assessment
Summative assessment evaluates cumulative student learning, skill acquisition, and academic achievement against predetermined grade-level standards or annual IEP goals at the conclusion of a defined instructional period (e.g., end of a unit, semester, or academic school year).
Special Education Application: Summative assessments determine whether a student has mastered annual IEP goals, achieved course credit, or met state accountability standards. Examples include end-of-unit examinations, final course projects, the general Georgia Milestones End-of-Grade (EOG) and End-of-Course (EOC) assessments, and the Georgia Alternate Assessment 2.0 (GAA 2.0) for students with significant cognitive disabilities.
Norm-Referenced vs. Criterion-Referenced Assessments
A critical distinction on the GACE Special Education General Curriculum examination is the structural difference between norm-referenced and criterion-referenced testing.
| Assessment Dimension | Norm-Referenced Assessment | Criterion-Referenced Assessment |
|---|---|---|
| Primary Question Answered | "How does this student's performance compare to same-age or same-grade peers across the nation?" | "What specific knowledge, skills, or standards has this student mastered?" |
| Comparison Standard | Relative standard: A representative national norming group (standardization cohort). | Absolute standard: A predetermined cut-score, curricular criterion, or learning standard. |
| Reported Score Formats | Standard scores ($M=100, SD=15$), percentile ranks ($1-99$), stanines, scaled scores. | Percentage correct, pass/fail, rubric mastery levels (e.g., proficient, advanced), raw mastery counts. |
| Primary Special Ed Application | IDEA eligibility determination, identification of cognitive/achievement discrepancies, triennial reviews. | Baseline data collection for PLAAFP, writing measurable annual IEP goals, tracking discrete skill acquisition. |
| Instructional Utility | Low: Informs whether a deficit exists relative to peers, but does not identify which specific items to teach. | High: Pinpoints exactly which prerequisite skills are unmastered to guide targeted lesson planning. |
| Prominent Examples | WISC-V, Woodcock-Johnson IV (WJ IV), KTEA-3, BASC-3 behavioral scales. | Brigance Comprehensive Inventory of Basic Skills, state standards tests, teacher-made mastery rubrics. |
Alternative and Authentic Assessment Methodologies
Traditional paper-and-pencil standardized tests often fail to capture the true capabilities of students with moderate-to-severe disabilities, sensory impairments, complex communication needs, or diverse linguistic backgrounds. Special educators deploy alternative and authentic assessment frameworks to gather ecologically valid performance data.
1. Performance-Based Assessment
Performance-based assessments require students to actively demonstrate their knowledge and skills by constructing a tangible response, executing an authentic task, or solving a real-world problem. Rather than selecting an answer from multiple choices, students synthesize information (e.g., conducting a hands-on science lab experiment, calculating change in a mock grocery store during Community-Based Instruction, or delivering an oral presentation using augmentative communication).
2. Portfolio Assessment
A portfolio assessment is a purposeful, systematic collection of student work gathered over an extended timeframe that directly exhibits the student's efforts, academic progress, and mastery of targeted IEP objectives. Effective portfolios include diverse work samples (drafts, completed projects, audio/video recordings), analytic scoring rubrics, and explicit student self-reflection entries.
3. Dynamic Assessment (Test-Teach-Retest)
Grounded in Lev Vygotsky's sociocultural theory and the Zone of Proximal Development (ZPD), dynamic assessment departs from static testing by actively embedding instruction into the evaluation process. It follows a structured Test-Teach-Retest sequence:
- Pretest: The educator assesses the student's independent performance on a novel task without assistance.
- Teach (Mediation): The examiner provides mediated learning experiences, offering targeted prompts, cognitive modeling, strategy instruction, and scaffolding to observe how the student responds to instruction.
- Posttest: The examiner reassesses the student on a parallel task to quantify learning modifiability and responsiveness.
Critical GACE Implication: Dynamic assessment is the most effective assessment methodology for differentiating between a language/cultural difference and an actual learning disability in Culturally and Linguistically Diverse (CLD) students and English Learners (ELs). Static IQ tests measure accumulated past knowledge; dynamic assessments measure current learning potential and cognitive modifiability.
4. Classroom Literacy Diagnostic Tools
Accurate diagnosis of reading failure requires specialized informal assessment procedures that reveal the cognitive and linguistic mechanisms underlying reading breakdowns:
- Running Records: Developed by Marie Clay, running records allow teachers to listen to a student read a leveled text aloud (100–150 words) while systematically coding oral reading behaviors in real time. The teacher records correct words, substitutions, omissions, insertions, teacher assists, and self-corrections to calculate:
- Independent Level: 95% to 100% accuracy (suitable for unguided independent reading).
- Instructional Level: 90% to 94% accuracy (optimal zone for teacher-led guided reading instruction).
- Frustrational Level: Below 90% accuracy (the text is too difficult; comprehension collapses; do not instruct at this level).
- Miscue Analysis: Originating from Kenneth Goodman's psycholinguistic reading model, miscue analysis analyzes the qualitative nature of a reader's oral errors (miscues) to determine which of the three cueing systems the student relies upon or neglects:
- Semantic Cueing System (Meaning - "Does it make sense?"): Did the error preserve the sentence's contextual meaning (e.g., reading "The horse ran through the field" instead of "pasture")?
- Syntactic Cueing System (Structure/Grammar - "Does it sound right?"): Did the error match the correct grammatical part of speech (e.g., substituting a noun for a noun)?
- Graphophonic Cueing System (Visual/Sound-Symbol - "Does it look right?"): Did the student decode based on phonetic letter-sound correspondences (e.g., reading "party" for "pretty")?
- Informal Reading Inventories (IRI): Comprehensive diagnostic reading batteries (such as the Qualitative Reading Inventory [QRI] or Flynt-Cooter) that combine graded word lists, graded narrative/expository passages, running records, and literal/inferential comprehension questions to identify word recognition accuracy, reading rate, and comprehension profiles across grade levels.
A multidisciplinary evaluation team is reviewing assessment data for a third-grade student experiencing persistent reading difficulties. To determine whether the student meets eligibility criteria for a Specific Learning Disability under IDEA and to identify specific cognitive processing deficits, which assessment type should the school psychologist primarily administer?
A special education teacher is writing the Present Levels of Academic Achievement and Functional Performance (PLAAFP) and draft annual goals for an eighth-grade student's IEP in mathematics. The teacher needs to identify the exact math calculation and algebraic skills the student has mastered and which specific prerequisite skills remain unmastered. Which assessment type is most appropriate for this purpose?
A fifth-grade English Learner who recently immigrated to Georgia is struggling significantly with complex multi-step science tasks. The teacher suspects a possible cognitive processing disability, but the student's bilingual evaluation team is concerned that traditional standardized cognitive testing will penalize the student's emerging English proficiency. Which assessment approach would best allow the team to evaluate the student's true learning potential and cognitive modifiability?