3.1 Assessment Terminology, Formal/Informal Measures & Psychometric Principles

Key Takeaways

  • Standard scores (mean 100, SD 15) and scaled scores (mean 10, SD 3) allow valid psychometric comparisons across assessment batteries, whereas age and grade equivalents are non-linear ordinal extrapolations that must never be used to establish instructional levels or determine placement.
  • The Standard Error of Measurement (SEM) quantifies expected score fluctuation due to unreliability; Committee on Special Education (CSE) members must interpret test performance using confidence intervals (typically 90% or 95%) rather than isolated point estimates.
  • Reliability denotes measurement consistency across time, forms, items, or raters, whereas validity verifies that an instrument measures its intended theoretical construct; high reliability is a prerequisite for validity, but reliability alone does not guarantee validity.
  • Dynamic assessment evaluates a student's learning potential and cognitive modifiability using a test-teach-retest paradigm grounded in Vygotsky's Zone of Proximal Development (ZPD), making it particularly valuable for culturally and linguistically diverse learners.
  • Ecological assessment examines student performance across authentic physical, social, and instructional environments through environmental inventories to identify contextual barriers and support systemic intervention planning.
Last updated: September 2026

3.1 Assessment Terminology, Formal/Informal Measures & Psychometric Principles

Quick Summary: Special education assessment is anchored in rigorous psychometric standards designed to ensure that clinical, diagnostic, and instructional placement decisions reflect authentic student abilities rather than measurement artifact or cultural bias. Special educators in New York must master the statistical foundations of reliability, validity, normal distribution metrics, the Standard Error of Measurement (SEM), and multi-method evaluation paradigms to serve effectively on the Committee on Special Education (CSE).

Assessment in special education serves multiple high-stakes functions: screening, determining eligibility, guiding individualized instructional design, monitoring academic progress, and evaluating program efficacy. Under the Individuals with Disabilities Education Act (IDEA) and Part 200 of the Regulations of the Commissioner of Education of the State of New York (8 NYCRR Part 200), assessment procedures must be technically sound, non-discriminatory, and administered by qualified personnel. To interpret multidisciplinary evaluation reports accurately, special educators must possess a sophisticated understanding of psychometrics—the science of psychological and educational measurement.


Psychometric Fundamentals: Reliability, Validity, and Standardization

Every standardized measurement tool rests upon two foundational psychometric pillars: reliability and validity. An assessment cannot provide legally defensible or educationally meaningful data without established technical adequacy across both dimensions.

Reliability: The Consistency of Measurement

Reliability refers to the degree to which an assessment tool produces stable, consistent, and replicable results across repeated administrations, varying test forms, differing raters, and internal item samples. Reliability coefficients ($r$) range from 0.00 to 1.00. For high-stakes special education decisions—such as determining eligibility or classification under IDEA—instruments must demonstrate a reliability coefficient of at least $r \ge 0.90$. Screening instruments typically require at least $r \ge 0.80$.

Reliability TypeCore Measurement QuestionPrimary Source of Error ControlledSpecial Education Application
Test-Retest ReliabilityDoes the instrument produce stable scores across time?Temporal fluctuations, fatigue, practice effects, maturational variance.Administering the Woodcock-Johnson IV Achievement battery at a two-week interval to confirm score stability.
Alternate-Form (Parallel-Form)Do two equivalent forms of the test measure the same construct identically?Item specificity, content sampling differences between Form A and Form B.Curriculum-Based Measurement (CBM) progress monitoring probes administered bi-weekly to evaluate reading growth.
Internal ConsistencyDo all items within a single subtest measure the identical underlying construct?Content heterogeneity, poorly calibrated test items.Calculating Cronbach's alpha ($\alpha$) or split-half reliability on the WISC-V Working Memory Index.
Inter-Rater (Inter-Observer)Do two independent evaluators score the student's performance identically?Scorer subjectivity, ambiguous scoring rubrics, observer drift.Two evaluators independently scoring a Functional Behavioral Assessment (FBA) direct observation log or a standardized writing rubric.

Validity: The Truthfulness of Measurement

Validity is the degree to which an instrument measures precisely what it purports to measure and the appropriateness of the inferences, decisions, and actions based on the test scores. It is critical to remember that reliability is a necessary but insufficient condition for validity: an assessment can be extraordinarily reliable (producing identical erroneous scores repeatedly) while remaining completely invalid for a specific diagnostic purpose.

  • Content Validity: The extent to which the items on an assessment representatively sample the broader domain, curriculum, or behavioral construct being evaluated. A fourth-grade mathematics achievement test lacks content validity if it measures third-grade arithmetic or assesses reading decoding through convoluted word problems rather than pure mathematical reasoning.
  • Construct Validity: The degree to which an assessment measures a theoretical, psychological construct or trait (such as intelligence, executive functioning, phonological awareness, or anxiety). Construct validity is established through:
    • Convergent Validity: Demonstrating that scores correlate strongly with other established tests measuring the same or related constructs (e.g., strong correlation between the WISC-V and the Stanford-Binet-5 Full Scale IQ).
    • Discriminant (Divergent) Validity: Demonstrating that scores do not correlate significantly with tests measuring unrelated psychological traits (e.g., an attention-deficit rating scale should not correlate highly with an expressive vocabulary test).
  • Criterion-Related Validity: The extent to which test scores correlate with an external standard, criterion, or outcome:
    • Concurrent Validity: The test score corresponds closely with an established criterion measure administered at approximately the same time (e.g., a newly developed reading screener administered concurrently with the DIBELS 8th Edition benchmark).
    • Predictive Validity: The test score accurately forecasts future performance on a relevant criterion (e.g., a kindergarten phonemic awareness screening score predicting third-grade performance on the New York State ELA assessment).

Standardization: Preserving Technical Adequacy

Standardization requires that an instrument be administered, scored, and interpreted under strictly uniform, invariant conditions. Standardized protocols specify exact verbatim directions, precise timing constraints, uniform demonstration items, permitted prompting, and strict rules for basals (the point below which the examiner assumes the student would answer all easier items correctly) and ceilings (the point at which testing stops because the student has reached a designated threshold of consecutive incorrect responses).

Furthermore, standardized tests rely on a normative sample (the "norm group"). To be psychometrically valid for a student in New York, the norm sample must be representative of the national population across key demographic variables: racial and ethnic composition, geographic region, parent educational attainment, socioeconomic status, and proportional representation of students with disabilities, matched to recent U.S. Census data.


The Normal Distribution, Standard Scores, and Derived Scores

Standardized norm-referenced tests rely on the theoretical properties of the normal distribution curve (the bell curve). In a normal distribution, scores cluster symmetrically around a central arithmetic mean, with predictable proportions of the population falling within specific standard deviations (SD) from that mean:

  • 68.26% of all scores fall within $\pm 1$ standard deviation of the mean ($+1\sigma$ to $-1\sigma$).
  • 95.44% of all scores fall within $\pm 2$ standard deviations of the mean ($+2\sigma$ to $-2\sigma$).
  • 99.74% of all scores fall within $\pm 3$ standard deviations of the mean ($+3\sigma$ to $-3\sigma$).
                               Normal Distribution Curve
                                     Mean (100)
                                         |
                                       .---.
                                      /     \
                                     /   |   \
                                    /    |    \
                                  .'     |     '.
                               .-'       |       '-.
                           _.-'          |          '-._
                      _..-'              |              '-.._
     ______________.-'                   |                   '-.______________
          -3 SD        -2 SD           -1 SD           +1 SD        +2 SD        +3 SD
          (55)         (70)            (85)            (115)        (130)        (145)
            |            |               |               |            |            |
            |--- 2.14% --|---- 13.59% ---|---- 34.13% ---|--- 34.13% -|--- 13.59% -|
            |                                                                      |
            |<----------------------- Average Range (85-115) --------------------->|

Because raw scores (the actual number of items answered correctly) lack interpretive meaning across differing tests, psychometricians transform raw scores into derived scores. The primary derived score formats used in special education multidisciplinary evaluations include:

Derived Score TypeMean ($\mu$)Standard Deviation ($\sigma$)Normal / Average RangePrimary Clinical / Educational Application
Standard Score (SS)1001585 – 115Global cognitive ability (Full Scale IQ) and composite academic achievement (WJ-IV, WISC-V, WIAT-4, KTEA-3). Standard scores maintain equal intervals across the entire scale.
Scaled Score1038 – 12Standardized subtest performance (e.g., WISC-V Matrix Reasoning, WIAT-4 Pseudoword Decoding). Scaled scores 4–7 indicate below-average functioning; scores 1–3 reflect significant impairment.
T-Score501040 – 60Behavioral, emotional, and social rating scales (e.g., BASC-3, Conners-4). In behavioral measures, elevated T-scores (60–69 = At-Risk; $\ge 70$ = Clinically Significant) indicate pathology.
Z-Score01-1.0 to +1.0Statistical baseline score indicating the exact number of standard deviation units a raw score deviates from the mean ($Z = \frac{X - \mu}{\sigma}$).
Stanine ("Standard Nine")524 – 6Coarse scale dividing the normal curve into 9 bands. Stanines 1–3 represent below average, 4–6 average, and 7–9 above average. Useful for broad group comparisons but lacks precision.
Percentile Rank (PR)50 (Median)N/A16th – 84thIndicates the percentage of individuals in the normative reference group who scored at or below the student's score (range 1–99). Non-linear intervals.

Percentile Ranks vs. Equal-Interval Standard Scores

A critical psychometric trap in CSE meetings is conflating percentile ranks with equal-interval scores. Percentile ranks are ordinal ranks, not equal-interval measurements. Because scores cluster densely around the center of the normal distribution, a small difference in standard scores near the mean yields a massive change in percentile rank. Conversely, at the extreme tails of the curve, a large difference in standard score produces almost no change in percentile rank:

  • Moving from a Standard Score of 95 to 100 (a 5-point increase near the mean) moves a student from the 37th percentile to the 50th percentile (a 13-percentile jump).
  • Moving from a Standard Score of 65 to 70 (the identical 5-point increase at the lower tail) moves a student from the 1st percentile to the 2nd percentile (only a 1-percentile jump).

Special educators must prevent CSE teams from misinterpreting small percentile changes at the extremes as indicators of zero instructional progress.

The Critical Fallacy of Age and Grade Equivalents

Age Equivalents (AE) and Grade Equivalents (GE) represent the most frequently misinterpreted derived scores in education, and professional measurement bodies (including the American Psychological Association and National Council on Measurement in Education) strongly discourage their use in high-stakes placement decisions.

Psychometric Warning on Grade Equivalents (GE): A Grade Equivalent of 4.2 does not mean that a student is performing proficiently at a fourth-grade level, nor does it mean that the student has mastered fourth-grade curriculum. It simply means that the student answered the same raw number of items correctly on that specific test as an average student in the second month of fourth grade in the norming sample.

Why age and grade equivalents must never guide instructional placement:

  1. Extrapolated and Interpolated Data: Test developers do not test every grade level every month. Most GE scores are mathematically interpolated or extrapolated, assuming linear growth that does not reflect real childhood cognitive development.
  2. False Equivalence Across Grades: If an eighth-grade student with a reading disability obtains a GE of 4.0 on an eighth-grade reading test, that eighth grader did not take a fourth-grade test. The eighth grader answered easier eighth-grade items correctly (perhaps using compensatory adult vocabulary), whereas a fourth grader earning that raw score answered fourth-grade questions using emerging decoding skills. The eighth grader cannot be handed a fourth-grade basal reader and expected to learn appropriately.
  3. Unequal Intervals: Growth across grades is not constant. A one-year difference in grade equivalent between kindergarten and first grade (learning to decode) represents an enormous developmental leap, whereas a one-year difference between tenth and eleventh grade represents a minor refinement in reading comprehension rate.

Standard Error of Measurement (SEM) and Confidence Intervals

In classical test theory, any observed score ($X$) consists of two components: the student's hypothetical True Score ($T$) and the inevitable Error Score ($E$):

X=T+EX = T + E

Because no educational or psychological instrument possesses perfect reliability ($r = 1.00$), error is present in every administration due to internal factors (illness, anxiety, motivation, attention lapses) and external factors (examiner pacing, room temperature, ambient noise). The Standard Error of Measurement (SEM) quantifies the standard deviation of these error scores around the true score:

SEM=σ1r\text{SEM} = \sigma \sqrt{1 - r}

Where $\sigma$ is the standard deviation of the test and $r$ is the test's reliability coefficient. As test reliability increases, the SEM decreases, yielding a more precise measurement.

Confidence Intervals in Practice

Because a student's observed score is merely an estimate, psychometrically sound reports present scores surrounded by a Confidence Interval (confidence band). The confidence interval defines a statistical range within which the student's true score is expected to fall with a specified probability:

  • 68% Confidence Interval: Observed Score $\pm 1.00 \times \text{SEM}$
  • 90% Confidence Interval: Observed Score $\pm 1.65 \times \text{SEM}$
  • 95% Confidence Interval: Observed Score $\pm 1.96 \times \text{SEM}$

CSE Clinical Implication: The Danger of Cut-Off Scores

Consider an evaluation where a third-grader achieves a Full Scale Standard Score of 69 on a cognitive battery with an SEM of 3. If a district rigidly enforces an IQ cut-off of 70 for Intellectual Disability classification, an uninformed team might inappropriately classify the student. However, the 95% confidence interval spans from 63 to 75 ($69 \pm [1.96 \times 3]$). Because the confidence band clearly encompasses scores well within the borderline-to-low-average range, the observed score cannot be treated as an absolute or infallible finding.


Formal vs. Informal Assessment Paradigms

A comprehensive special education evaluation synthesizes data from both formal, standardized tools and informal, classroom-based assessments.

Standardized Norm-Referenced Assessments

Norm-referenced tests compare an individual student's performance against a large, nationally representative normative sample of age- or grade-matched peers. Their primary function is diagnostic classification, eligibility determination, and broad cognitive/academic profiling.

  • Comprehensive Cognitive Batteries:
    • Wechsler Intelligence Scale for Children, Fifth Edition (WISC-V): Ages 6:0 to 16:11. Generates a Full Scale IQ (FSIQ) and five primary index scores: Verbal Comprehension (VCI), Visual Spatial (VSI), Fluid Reasoning (FRI), Working Memory (WMI), and Processing Speed (PSI).
    • Woodcock-Johnson IV Tests of Cognitive Abilities (WJ-IV COG): Grounded in the Cattell-Horn-Carroll (CHC) theory of cognitive abilities, assessing broad abilities such as Comprehension-Knowledge ($Gc$), Fluid Reasoning ($Gf$), Short-Term Working Memory ($Gwm$), Visual Processing ($Gv$), Auditory Processing ($Ga$), and Processing Speed ($Gs$).
    • Stanford-Binet Intelligence Scales, Fifth Edition (SB-5): Evaluates verbal and nonverbal intelligence across five cognitive factors: Fluid Reasoning, Knowledge, Quantitative Reasoning, Visual-Spatial Processing, and Working Memory.
  • Comprehensive Standardized Achievement Batteries:
    • Woodcock-Johnson IV Tests of Achievement (WJ-IV ACH): Measures core academic domains: Reading (Letter-Word Identification, Passage Comprehension, Word Attack), Mathematics (Calculation, Applied Problems, Math Facts Fluency), and Written Language (Spelling, Writing Samples, Sentence Writing Fluency).
    • Wechsler Individual Achievement Test, Fourth Edition (WIAT-4): Evaluates reading, writing, mathematics, and oral language with detailed error analyses and dyslexia index composites.
    • Kaufman Test of Educational Achievement, Third Edition (KTEA-3): Co-normed achievement battery featuring comprehensive qualitative error analysis across academic domains.

Criterion-Referenced Assessments

Unlike norm-referenced tests, criterion-referenced assessments do not compare a student to peers. Instead, they measure performance against an absolute, pre-determined standard, mastery benchmark, or curriculum competency (e.g., "Student correctly solves two-digit addition problems with regrouping with 80% accuracy").

  • Brigance Comprehensive Inventory of Basic Skills (CIBS-II / Brigance Early Childhood): Identifies specific academic and developmental skill sequences mastered, emerging, or unmastered to establish baseline data for Individualized Education Program (IEP) goals.
  • State Learning Standards Checklists: Evaluating mastery of specific New York State Next Generation Learning Standards to guide direct instruction.

Diagnostic, Formative, and Summative Assessments

Assessment CategoryTiming & FrequencyPrimary PurposeHigh/Low StakesSpecial Education Example
Diagnostic AssessmentPrior to instruction or during evaluation.Pinpoints specific foundational deficits, cognitive gaps, or subskill breakdowns.Moderate StakesAdministering the Comprehensive Test of Phonological Processing (CTOPP-2) to evaluate phonological memory and rapid naming.
Formative AssessmentOngoing throughout the instructional unit ("assessment for learning").Informs daily instructional adjustments, identifies misconceptions, provides immediate student feedback.Low StakesExit tickets, running records, dynamic checks for understanding during direct instruction.
Summative AssessmentAt the conclusion of a unit, semester, or school year ("assessment of learning").Evaluates cumulative knowledge acquisition and program effectiveness against standards.High StakesNew York State Grades 3–8 ELA and Mathematics tests; Regents Examinations.

Dynamic Assessment

Rooted in Lev Vygotsky's concept of the Zone of Proximal Development (ZPD), dynamic assessment departs from traditional static testing (which measures only what a student can do independently without assistance). It employs a structured Test-Teach-Retest paradigm:

  1. Pretest: Evaluate independent baseline performance on a cognitive or academic task.
  2. Mediated Learning Experience (Teach): The examiner provides explicit instruction, scaffolds cognitive strategies, and models self-monitoring techniques.
  3. Posttest: Re-evaluate performance to assess cognitive modifiability—the degree to which the student incorporates feedback and improves under scaffolded conditions.

Dynamic assessment is critical when evaluating culturally and linguistically diverse learners, as it distinguishes between a lack of prior educational opportunity (which responds rapidly to mediated learning) and an intrinsic learning disability (which exhibits persistent processing difficulties despite mediation).

Ecological Assessment

An ecological assessment evaluates the student across their natural physical, social, and instructional environments (classroom, cafeteria, playground, hallway, bus, and community). Evaluators conduct an environmental inventory:

  1. Identify the specific physical, instructional, and behavioral demands of each environment.
  2. Observe the student's performance within each environment.
  3. Perform a discrepancy analysis comparing environmental demands against the student's demonstrated skills.
  4. Identify environmental triggers, structural barriers, or sensory inputs that disrupt learning, informing positive behavioral supports and physical accommodations.

Curriculum-Based Assessment (CBA) vs. Curriculum-Based Measurement (CBM)

While frequently conflated, CBA is an overarching umbrella term, whereas CBM is a specific, standardized empirical methodology:

  • Curriculum-Based Assessment (CBA): Any informal or criterion-referenced assessment method that directly utilizes materials from the student's local classroom curriculum (e.g., end-of-chapter math quizzes, teacher-made spelling inventories).
  • Curriculum-Based Measurement (CBM): A standardized, empirically validated subset of CBA characterized by brief (1–3 minute), timed probes of general outcome indicators (e.g., Oral Reading Fluency, Maze comprehension, Math computation), alternate forms of equivalent difficulty, standardized administration, and repeated administration to track rates of improvement over time.

Practical NY CSE Scenario & Psychometric Data Interpretation

Case Profile: Marcus (Grade 4, Age 9:8)

Marcus was referred to the Committee on Special Education (CSE) following 20 weeks of Tier 2 and Tier 3 reading interventions. The multidisciplinary evaluation generated the following standardized assessment profile:

Assessment Battery / SubtestRaw ScoreStandard Score (SS)Percentile Rank95% Confidence IntervalGrade Equivalent (GE)
WJ-IV Cognitive General Intellectual Ability (GIA)--10255th97 – 1074.8
WJ-IV Auditory Processing ($Ga$)--744th68 – 801.8
WJ-IV Visual Processing ($Gv$)--11075th104 – 1166.2
WJ-IV Short-Term Working Memory ($Gwm$)--8110th75 – 872.5
WJ-IV ACH Letter-Word Identification--765th71 – 812.1
WJ-IV ACH Word Attack (Pseudoword Decoding)--713rd66 – 761.7
WJ-IV ACH Passage Comprehension--8414th78 – 902.8
WJ-IV ACH Calculation--9947th93 – 1054.2
Brigance CIBS-II Criterion Assessment--Mastery of CVC and CCVC words; 20% accuracy on consonant digraphs and vowel teams.

CSE Multidisciplinary Synthesis

  1. Cognitive Profile: Marcus demonstrates average overall intellectual potential (GIA 102), but exhibits an acute processing deficit in Auditory Processing ($Ga = 74$) and secondary weaknesses in Working Memory ($Gwm = 81$), contrasted with superior Visual Processing ($Gv = 110$).
  2. Academic Profile: In basic reading skills, Marcus exhibits severe deficits in phonemic decoding (Word Attack SS = 71, 3rd percentile) and sight word identification (Letter-Word ID SS = 76). His reading comprehension (SS = 84) is partially sustained by his average general intelligence and contextual reasoning.
  3. Refuting the Grade Equivalent Trap: A general education committee member asserts: "Marcus has a Grade Equivalent of 1.7 in Word Attack, so we should place him in first-grade phonics materials." The special education teacher refutes this: Marcus is a fourth grader whose raw score matched the median raw score of beginning first graders on this test, but his developmental profile, visual processing strengths, and vocabulary are at a fourth-grade level. Placing him in juvenile first-grade materials would harm his self-concept and instructional engagement. Instead, the Brigance criterion data reveals the exact point of instructional breakdown: he has mastered CVC and CCVC patterns but requires systematic, multisensory structured literacy instruction targeting consonant digraphs and vowel teams using age-appropriate materials.
  4. SEM and Confidence Band Integration: The team notes that Marcus's Word Attack confidence interval ($66–76$) remains entirely within the below-average-to-impaired range, confirming a true, pervasive Specific Learning Disability in basic reading (Dyslexia) rather than an isolated testing artifact.
Test Your Knowledge

During a Committee on Special Education (CSE) meeting, a general education teacher reviews a fourth-grade student's standardized reading achievement report and notes a Grade Equivalent (GE) score of 2.2 on the reading comprehension subtest. The teacher recommends placing the student in second-grade reading curriculum materials. How should the special education teacher interpret this psychometric data to the committee?

A
B
C
D
Test Your Knowledge

A school psychologist and special education teacher evaluate a third-grade multilingual learner who recently relocated to New York from another country. Standardized norm-referenced cognitive tests yield below-average processing scores, but the evaluation team suspects that unfamiliarity with standardized testing formats and English linguistic complexity may be confounding the results. Which assessment methodology is most appropriate to evaluate the student's authentic learning potential?

A
B
C
D
Test Your Knowledge

A certified school psychologist presents an initial psychological evaluation report indicating that a fifth-grade student achieved a Full Scale Standard Score of 72 on a standardized cognitive assessment with a Standard Error of Measurement (SEM) of 3.5 points. The report presents the score with a 95% confidence interval of 65 to 79. When the CSE reviews this report, what is the most accurate psychometric interpretation of these data?

A
B
C
D