2.1 Assessment Types, Terminology, and Psychometric Properties
Key Takeaways
- Norm-referenced assessments compare a student's performance against a standardized peer sample, whereas criterion-referenced assessments evaluate mastery of specific curriculum objectives or standards.
- Standard scores (Mean = 100, SD = 15) and T-scores (Mean = 50, SD = 10) transform raw scores into equal-interval scales that allow valid comparisons across different tests.
- Reliability indicates the consistency or stability of assessment results (e.g., test-retest, split-half, inter-rater), while validity ensures the instrument measures what it claims to measure (content, criterion, construct).
- The Standard Error of Measurement (SEM) quantifies expected score variation due to measurement error; teams must use confidence intervals rather than single point scores when interpreting evaluation data.
- Age and grade equivalents lack equal-interval properties and encourage misleading developmental misinterpretations, making standard scores and percentile ranks the psychometrically preferred reporting metrics.
Evaluation in special education relies on accurate, psychometrically sound data to make life-altering decisions regarding student eligibility, instructional placement, and intervention design. For the ILTS Learning Behavior Specialist I (290) exam, educators must possess a thorough understanding of assessment classifications, score transformations, statistical distributions, and psychometric properties.
1. Classifications of Assessment Tools
Special educators utilize a variety of assessment instruments, each designed to answer specific educational and diagnostic questions. Selecting the appropriate tool requires understanding the fundamental distinctions among assessment paradigms.
Norm-Referenced vs. Criterion-Referenced Assessments
The distinction between norm-referenced and criterion-referenced testing lies in the frame of reference used to interpret student performance:
| Feature | Norm-Referenced Assessments (NRA) | Criterion-Referenced Assessments (CRA) |
|---|---|---|
| Primary Purpose | Compare a student's performance to a representative peer sample (norm group). | Evaluate a student's mastery of specific skills, objectives, or curricular standards. |
| Score Interpretation | Relative standing (e.g., percentile ranks, standard scores). | Absolute performance (e.g., 85% accuracy, mastery/non-mastery). |
| Instructional Utility | High for eligibility determination and macro-level diagnostic classification. | High for day-to-day instructional planning, task analysis, and IEP goal tracking. |
| Examples | Wechsler Individual Achievement Test (WIAT-IV), Woodcock-Johnson IV (WJ-IV). | Brigance Diagnostic Inventory, Teacher-made rubrics, State Content Benchmarks. |
Exam Tip: If the goal is to determine if a student performs significantly below age- or grade-level peers for special education eligibility, use a norm-referenced assessment. If the goal is to pinpoint exact skills a student has mastered or needs to learn next, use a criterion-referenced assessment.
Additional Assessment Dichotomies
- Formal vs. Informal Assessments:
- Formal Assessments follow standardized administration procedures, strict timing, explicit scoring scripts, and validated psychometric properties (e.g., standardized intelligence tests).
- Informal Assessments are non-standardized measures embedded in instruction, such as observational checklists, running records, work sample analyses, and teacher-made probes.
- Formative vs. Summative Assessments:
- Formative Assessment occurs during the instructional process to monitor learning and guide real-time instructional adjustments (e.g., exit tickets, weekly spelling probes).
- Summative Assessment occurs after a unit or period of instruction to evaluate cumulative learning outcomes (e.g., end-of-unit exams, annual state accountability tests).
- Aptitude vs. Achievement vs. Diagnostic Testing:
- Aptitude Tests measure potential or cognitive capacity to learn (e.g., WISC-V, Stanford-Binet 5).
- Achievement Tests measure acquired knowledge and academic skills resulting from instruction (e.g., KTEA-3).
- Diagnostic Tests break down specific academic domains (e.g., phonological processing, error analysis in math computation) to isolate underlying cognitive or skill deficits.
2. Psychometric Properties: Reliability and Validity
For an assessment to yield meaningful clinical data, it must demonstrate robust psychometric properties. Educators must evaluate both reliability (consistency) and validity (accuracy).
Reliability (Consistency of Measurement)
Reliability refers to the degree to which an assessment tool produces stable, consistent, and dependable results across administrations, forms, or raters. Reliability coefficients range from 0.00 (no reliability) to 1.00 (perfect reliability). Standardized diagnostic tests should possess a reliability coefficient of r = 0.90 or higher.
- Test-Retest Reliability: Measures stability over time by administering the same test to the same group at two different time points.
- Alternate-Form (Parallel-Form) Reliability: Measures equivalence across different versions of the test containing similar items targeting the same domain.
- Internal Consistency Reliability: Measures the degree to which items within the same subtest measure the same construct. Calculated using Split-Half Reliability (correlating scores on odd vs. even items) or Cronbach's Alpha.
- Inter-Rater Reliability: Measures agreement between two or more independent observers or scorers evaluating the same performance (critical for subjective measures like behavioral observations or essay scoring).
Validity (Accuracy of Measurement)
Validity refers to the extent to which a test measures what it purports to measure. A test cannot be valid unless it is reliable; however, a test can be highly reliable while being completely invalid for a specific purpose.
[ HIGH RELIABILITY ] + [ ACCURATE INTENT ] = [ VALID MEASURE ]
Consistent Results Measures Target Construct Trustworthy Clinical Data
- Content Validity: The extent to which test items representatively sample the complete domain or construct being assessed (e.g., a 4th-grade math test covering all state math standards, not just addition).
- Criterion-Related Validity: The extent to which scores on a test correlate with an established external criterion or outcome:
- Concurrent Validity: Test scores correlate strongly with a criterion measure administered at the same time (e.g., a new brief reading screener matching WIAT-IV reading comprehension scores).
- Predictive Validity: Test scores accurately predict future performance on a criterion measure (e.g., high school readiness assessments predicting postsecondary academic success).
- Construct Validity: The overarching degree to which an instrument accurately measures an abstract psychological construct or trait (e.g., intelligence, executive functioning, adaptive behavior). Established through convergent validity (correlating with similar constructs) and discriminant validity (showing no correlation with unrelated constructs).
3. Standardized Scores and the Normal Bell Curve
To evaluate where a student's score falls relative to the general population, raw scores (total points correct) must be converted into standardized scores. Standardized scores map onto the Normal Distribution (Bell Curve).
NORMAL DISTRIBUTION CURVE
Mean
|
|--------|--------|
-3SD -1SD +1SD +3SD
Percentiles: 0.1st 16th 84th 99.9th
Standard Scores: 55 85 115 145
T-Scores: 20 40 60 80
Z-Scores: -3.0 -1.0 +1.0 +3.0
Key Standard Score Metrics
- Mean and Standard Deviation (SD):
- In a standard normal distribution, 68.26% of scores fall within 1 SD of the mean.
- 95.44% of scores fall within 2 SDs of the mean.
- 99.72% of scores fall within 3 SDs of the mean.
- Standard Scores (SS): Commonly used on cognitive and achievement batteries (e.g., WISC-V, Woodcock-Johnson). Set with a Mean of 100 and an SD of 15.
- Average Range: 90 to 109 (within 1 SD).
- Below Average / Significant Deficit: Scores below 70 (more than 2 SDs below the mean).
- T-Scores: Frequently used in behavioral, emotional, and adaptive rating scales (e.g., BASC-3, Conners 3). Set with a Mean of 50 and an SD of 10.
- T-scores of 60-69 indicate "At-Risk" levels; T-scores of 70 or above indicate "Clinically Significant" concerns.
- Scaled Scores: Used for individual subtests within comprehensive batteries (e.g., WISC-V subtests). Set with a Mean of 10 and an SD of 3.
- Z-Scores: Express performance directly in standard deviation units from the mean (Mean = 0, SD = 1). A score 1.5 SDs below the mean equals a Z-score of -1.5.
- Percentile Ranks (PR): Indicate the percentage of students in the norm group who scored at or below a given score. A percentile rank of 25 means the student scored equal to or higher than 25% of peers. Note: Percentile ranks cluster densely around the median and spread out at the extremes, meaning they are non-linear.
The Danger of Age and Grade Equivalents
Age Equivalents (AE) and Grade Equivalents (GE) represent the age or grade level at which a given raw score was the median score in the norming sample.
Warning: Special educators must avoid using AE and GE for eligibility decisions. AE and GE scores do not represent equal intervals of performance, encourage false assumptions of developmental equivalence, and extrapolate performance outside tested content. Standard scores and percentiles must always be reported instead.
4. Standard Error of Measurement (SEM) and Confidence Intervals
No test score is perfectly accurate; every obtained score represents a combination of the student's true score and measurement error (e.g., fatigue, distraction, lighting, item sampling).
Calculating and Applying SEM
The Standard Error of Measurement (SEM) estimates the amount of error associated with a test instrument. SEM is inversely related to reliability: as the reliability coefficient increases, SEM decreases.
When interpreting evaluation results, multidisciplinary teams must apply a Confidence Interval (CI) around the obtained score rather than relying on a single point estimate.
- A 90% Confidence Interval is calculated as: Obtained Score ± (1.65 × SEM)
- A 95% Confidence Interval is calculated as: Obtained Score ± (1.96 × SEM)
Clinical Example
If a student earns a Standard Score of 68 on a cognitive test with an SEM of 3, the team can state with 95% confidence that the student's true cognitive score lies between 62 and 74 (68 ± 6). Because the confidence interval spans above 70, the team must examine secondary data and adaptive behavior measures before making an intellectual disability determination.
What is the fundamental psychometric relationship between test reliability and test validity?
On a standardized cognitive assessment with a Mean of 100 and a Standard Deviation of 15, a student obtains a Standard Score of 70. How should the multidisciplinary team interpret this performance relative to the normal curve?
Why do special education assessment guidelines strongly caution against using grade equivalents (GE) and age equivalents (AE) when reporting evaluation results to eligibility teams?
An evaluator administers an achievement test with a Standard Error of Measurement (SEM) of 4. If a student obtains a standard score of 82, what is the 95% confidence interval for this student's true score?