6.3 Reliability, Validity, Norm- vs Criterion-Referenced Tools
Key Takeaways
- Reliability means a test produces consistent results; common types include test-retest, alternate-form, internal consistency, and inter-rater reliability, each addressing a different source of inconsistency
- Validity means a test measures what it claims to measure and supports the specific inferences and decisions being made from its scores; validity is decision-specific, not a fixed property of a test
- The standard error of measurement (SEM) conceptually describes the margin of error around any observed score, which is why scores are best interpreted as a range or band rather than one exact number
- Norm-referenced tests compare a student's performance to a representative peer group (percentile ranks, standard scores); criterion-referenced tools compare performance to a fixed standard or skill mastery level, independent of how peers perform
- Choosing between norm-referenced and criterion-referenced tools depends on the decision being made: eligibility and broad standing questions favor norm-referenced data, while day-to-day instructional planning favors criterion-referenced and curriculum-based data
6.3 Reliability, Validity, Norm- vs Criterion-Referenced Tools
Quick Answer: Reliability asks whether a test produces consistent results across time, forms, items, or raters. Validity asks whether a test actually measures what it claims to measure and supports the specific decision being made from it. The standard error of measurement (SEM) reminds us that every observed score carries some margin of error, so scores should be interpreted as a band, not an exact point. Norm-referenced tools compare a student to a peer group; criterion-referenced tools compare a student to a fixed skill standard. These measurement concepts underlie every assessment decision an ESE teacher makes, from eligibility to daily instructional planning.
Reliability: Consistency of Measurement
Reliability refers to the consistency, dependability, or repeatability of test scores. A reliable test produces similar results under similar conditions; an unreliable test produces scores that bounce around due to measurement error rather than true changes in the student's skill. There are several distinct types of reliability, each addressing a different potential source of inconsistency:
| Type of Reliability | What It Measures | Example |
|---|---|---|
| Test-retest | Consistency of scores across two administrations of the same test over time | A student takes the same math test twice, two weeks apart, and scores are similar |
| Alternate-form (parallel-form) | Consistency between two equivalent versions of a test | A student's score on Form A closely matches their score on Form B of the same CBM probe set |
| Internal consistency | Consistency of performance across items within a single test | Items measuring the same construct correlate with each other (commonly reported as a coefficient such as Cronbach's alpha) |
| Inter-rater (inter-observer) | Agreement between two or more people scoring or observing the same performance | Two teachers independently scoring the same writing sample, or two observers coding the same behavior episode, arrive at similar results |
Inter-rater reliability deserves special attention in ESE because so much assessment relies on human judgment — behavior observation, rubric-scored writing samples, portfolio evaluation. A tool with poor inter-rater reliability produces scores that depend more on who is scoring than on what the student actually did, which undermines its usefulness for high-stakes decisions like eligibility or progress-monitoring conclusions.
Validity: Measuring What You Intend to Measure
Validity is the degree to which a test measures the construct it is intended to measure and supports the specific inferences and decisions made from its scores. Unlike reliability, which is a property that can be estimated somewhat independently of purpose, validity is fundamentally tied to how the score will be used — a test can be valid for one purpose and invalid for another. Common types of validity evidence include:
- Content validity — the test items adequately sample the full domain or curriculum they claim to represent (a reading comprehension test that only asks vocabulary-definition questions has weak content validity for "comprehension")
- Construct validity — the test genuinely measures the underlying psychological or academic construct it claims to measure (an "anxiety" scale that actually just measures general negative mood has weak construct validity)
- Criterion-related validity — scores correlate with an external, independently measured outcome; this includes concurrent validity (correlating with a currently available measure) and predictive validity (correlating with a future outcome, such as a screening tool's score predicting later reading failure)
A critical relationship to remember: a test can be reliable without being valid, but it cannot be valid without being reasonably reliable. A bathroom scale that consistently reads five pounds too heavy is highly reliable (consistent) but not valid (inaccurate) for measuring true body weight. A test whose scores bounce around randomly, by contrast, cannot possibly be measuring anything consistently, so it cannot be a valid measure of a stable trait or skill. Exam items sometimes describe a highly consistent test that measures the wrong construct for the decision at hand — the issue there is validity, not reliability, even though the test is technically dependable.
Standard Error of Measurement (SEM): Scores Are Bands, Not Points
Every observed test score is only an estimate of a student's "true" score, because no test is perfectly reliable. The standard error of measurement (SEM) is the statistic that describes how much an observed score is likely to vary from that true score due to ordinary measurement error. Conceptually, a larger SEM means more measurement error and less precision around any single score; a smaller SEM means the observed score is a more precise estimate of the true score.
The practical implication for ESE practice is significant: a single test score should be interpreted as a range (a confidence band around the score), not as one exact, immovable number. Two students with observed scores that are only a few points apart may not have a meaningfully different true skill level once measurement error is accounted for. This is precisely why the multiple-measures requirement discussed in Section 6.2 exists — relying on one score's exact position, especially near an eligibility cut point, risks treating measurement noise as a real difference in ability. Teams should never invent or assume a specific numeric SEM value for a given test without the test's technical manual; the exam tests the conceptual understanding that scores carry error, not memorized SEM figures for particular instruments.
Norm-Referenced vs. Criterion-Referenced Tools
These two categories differ in what they compare a student's performance to:
| Feature | Norm-Referenced | Criterion-Referenced |
|---|---|---|
| Compares student to | A representative peer group (the norming sample) | A fixed standard, skill, or mastery level |
| Typical scores reported | Percentile ranks, standard scores, age/grade equivalents | Percent correct, mastery/non-mastery, pass/fail against a benchmark |
| Answers the question | "How does this student compare to same-age or same-grade peers?" | "Has this student mastered this specific skill or standard?" |
| Best used for | Eligibility determination, identifying a significant discrepancy from typical development, broad standing questions | Instructional planning, tracking mastery of specific IEP goals or curriculum objectives |
| Example | A cognitive ability test yielding a standard score and percentile rank | A curriculum-based mastery check on two-digit subtraction with regrouping |
Norm-referenced tools are essential when a decision genuinely requires comparing a student to typical peer development — this is exactly the kind of comparison IDEA eligibility categories require (is this student's performance significantly below what is typical for same-age peers?). Criterion-referenced tools, including CBM probes and curriculum-based mastery checks, are the better fit for day-to-day instructional planning because they tell a teacher exactly which specific skills a student has or has not yet mastered, information a percentile rank alone cannot provide. A percentile rank of 10 tells you a student scored below 90% of peers, but it does not tell you which specific skills to teach next — only criterion-referenced, curriculum-based data does that.
Choosing the Right Tool for the Decision
Putting reliability, validity, and the norm- versus criterion-referenced distinction together, an ESE teacher facing an assessment decision should ask three questions in sequence: (1) What decision am I trying to make — eligibility/broad standing, or specific instructional next steps? (2) Does the available evidence show this tool is reliable and valid for that specific decision? (3) Should I use a norm-referenced tool (for peer comparison) or a criterion-referenced/curriculum-based tool (for mastery and instructional planning) — or, as is most often correct, some combination of both alongside observation and existing data? On the FTCE exam, resist the temptation to treat any one score as a complete, error-free answer; the correct response usually reflects the more cautious, multiple-measures, band-not-point interpretation of assessment data.
Two teachers independently score the same student writing sample using the same rubric and arrive at very similar scores. Which type of reliability does this best illustrate?
A bathroom scale always reads exactly five pounds heavier than a person's true weight, every single time it is used. What does this scenario illustrate about the relationship between reliability and validity?
A teacher wants to determine exactly which specific two-digit subtraction skills a student has and has not yet mastered in order to plan tomorrow's small-group lesson. Which type of assessment tool is the better fit for this instructional planning decision?