6.3 Reliability, Validity, Norm- vs Criterion-Referenced Tools

Key Takeaways

  • Reliability means a test produces consistent results; common types include test-retest, alternate-form, internal consistency, and inter-rater reliability, each addressing a different source of inconsistency
  • Validity means a test measures what it claims to measure and supports the specific inferences and decisions being made from its scores; validity is decision-specific, not a fixed property of a test
  • The standard error of measurement (SEM) conceptually describes the margin of error around any observed score, which is why scores are best interpreted as a range or band rather than one exact number
  • Norm-referenced tests compare a student's performance to a representative peer group (percentile ranks, standard scores); criterion-referenced tools compare performance to a fixed standard or skill mastery level, independent of how peers perform
  • Choosing between norm-referenced and criterion-referenced tools depends on the decision being made: eligibility and broad standing questions favor norm-referenced data, while day-to-day instructional planning favors criterion-referenced and curriculum-based data
Last updated: August 2026

6.3 Reliability, Validity, Norm- vs Criterion-Referenced Tools

Quick Answer: Reliability asks whether a test produces consistent results across time, forms, items, or raters. Validity asks whether a test actually measures what it claims to measure and supports the specific decision being made from it. The standard error of measurement (SEM) reminds us that every observed score carries some margin of error, so scores should be interpreted as a band, not an exact point. Norm-referenced tools compare a student to a peer group; criterion-referenced tools compare a student to a fixed skill standard. These measurement concepts underlie every assessment decision an ESE teacher makes, from eligibility to daily instructional planning.

Reliability: Consistency of Measurement

Reliability refers to the consistency, dependability, or repeatability of test scores. A reliable test produces similar results under similar conditions; an unreliable test produces scores that bounce around due to measurement error rather than true changes in the student's skill. There are several distinct types of reliability, each addressing a different potential source of inconsistency:

Type of ReliabilityWhat It MeasuresExample
Test-retestConsistency of scores across two administrations of the same test over timeA student takes the same math test twice, two weeks apart, and scores are similar
Alternate-form (parallel-form)Consistency between two equivalent versions of a testA student's score on Form A closely matches their score on Form B of the same CBM probe set
Internal consistencyConsistency of performance across items within a single testItems measuring the same construct correlate with each other (commonly reported as a coefficient such as Cronbach's alpha)
Inter-rater (inter-observer)Agreement between two or more people scoring or observing the same performanceTwo teachers independently scoring the same writing sample, or two observers coding the same behavior episode, arrive at similar results

Inter-rater reliability deserves special attention in ESE because so much assessment relies on human judgment — behavior observation, rubric-scored writing samples, portfolio evaluation. A tool with poor inter-rater reliability produces scores that depend more on who is scoring than on what the student actually did, which undermines its usefulness for high-stakes decisions like eligibility or progress-monitoring conclusions.

Validity: Measuring What You Intend to Measure

Validity is the degree to which a test measures the construct it is intended to measure and supports the specific inferences and decisions made from its scores. Unlike reliability, which is a property that can be estimated somewhat independently of purpose, validity is fundamentally tied to how the score will be used — a test can be valid for one purpose and invalid for another. Common types of validity evidence include:

  • Content validity — the test items adequately sample the full domain or curriculum they claim to represent (a reading comprehension test that only asks vocabulary-definition questions has weak content validity for "comprehension")
  • Construct validity — the test genuinely measures the underlying psychological or academic construct it claims to measure (an "anxiety" scale that actually just measures general negative mood has weak construct validity)
  • Criterion-related validity — scores correlate with an external, independently measured outcome; this includes concurrent validity (correlating with a currently available measure) and predictive validity (correlating with a future outcome, such as a screening tool's score predicting later reading failure)

A critical relationship to remember: a test can be reliable without being valid, but it cannot be valid without being reasonably reliable. A bathroom scale that consistently reads five pounds too heavy is highly reliable (consistent) but not valid (inaccurate) for measuring true body weight. A test whose scores bounce around randomly, by contrast, cannot possibly be measuring anything consistently, so it cannot be a valid measure of a stable trait or skill. Exam items sometimes describe a highly consistent test that measures the wrong construct for the decision at hand — the issue there is validity, not reliability, even though the test is technically dependable.

Standard Error of Measurement (SEM): Scores Are Bands, Not Points

Every observed test score is only an estimate of a student's "true" score, because no test is perfectly reliable. The standard error of measurement (SEM) is the statistic that describes how much an observed score is likely to vary from that true score due to ordinary measurement error. Conceptually, a larger SEM means more measurement error and less precision around any single score; a smaller SEM means the observed score is a more precise estimate of the true score.

The practical implication for ESE practice is significant: a single test score should be interpreted as a range (a confidence band around the score), not as one exact, immovable number. Two students with observed scores that are only a few points apart may not have a meaningfully different true skill level once measurement error is accounted for. This is precisely why the multiple-measures requirement discussed in Section 6.2 exists — relying on one score's exact position, especially near an eligibility cut point, risks treating measurement noise as a real difference in ability. Teams should never invent or assume a specific numeric SEM value for a given test without the test's technical manual; the exam tests the conceptual understanding that scores carry error, not memorized SEM figures for particular instruments.

Norm-Referenced vs. Criterion-Referenced Tools

These two categories differ in what they compare a student's performance to:

FeatureNorm-ReferencedCriterion-Referenced
Compares student toA representative peer group (the norming sample)A fixed standard, skill, or mastery level
Typical scores reportedPercentile ranks, standard scores, age/grade equivalentsPercent correct, mastery/non-mastery, pass/fail against a benchmark
Answers the question"How does this student compare to same-age or same-grade peers?""Has this student mastered this specific skill or standard?"
Best used forEligibility determination, identifying a significant discrepancy from typical development, broad standing questionsInstructional planning, tracking mastery of specific IEP goals or curriculum objectives
ExampleA cognitive ability test yielding a standard score and percentile rankA curriculum-based mastery check on two-digit subtraction with regrouping

Norm-referenced tools are essential when a decision genuinely requires comparing a student to typical peer development — this is exactly the kind of comparison IDEA eligibility categories require (is this student's performance significantly below what is typical for same-age peers?). Criterion-referenced tools, including CBM probes and curriculum-based mastery checks, are the better fit for day-to-day instructional planning because they tell a teacher exactly which specific skills a student has or has not yet mastered, information a percentile rank alone cannot provide. A percentile rank of 10 tells you a student scored below 90% of peers, but it does not tell you which specific skills to teach next — only criterion-referenced, curriculum-based data does that.

Choosing the Right Tool for the Decision

Putting reliability, validity, and the norm- versus criterion-referenced distinction together, an ESE teacher facing an assessment decision should ask three questions in sequence: (1) What decision am I trying to make — eligibility/broad standing, or specific instructional next steps? (2) Does the available evidence show this tool is reliable and valid for that specific decision? (3) Should I use a norm-referenced tool (for peer comparison) or a criterion-referenced/curriculum-based tool (for mastery and instructional planning) — or, as is most often correct, some combination of both alongside observation and existing data? On the FTCE exam, resist the temptation to treat any one score as a complete, error-free answer; the correct response usually reflects the more cautious, multiple-measures, band-not-point interpretation of assessment data.

Test Your Knowledge

Two teachers independently score the same student writing sample using the same rubric and arrive at very similar scores. Which type of reliability does this best illustrate?

A
B
C
D
Test Your Knowledge

A bathroom scale always reads exactly five pounds heavier than a person's true weight, every single time it is used. What does this scenario illustrate about the relationship between reliability and validity?

A
B
C
D
Test Your Knowledge

A teacher wants to determine exactly which specific two-digit subtraction skills a student has and has not yet mastered in order to plan tomorrow's small-group lesson. Which type of assessment tool is the better fit for this instructional planning decision?

A
B
C
D