5.6 Norm-Referenced and Criterion-Referenced Assessment

Key Takeaways

  • Norm-referenced tests compare a student to a norming sample; criterion-referenced tests compare performance to a fixed standard.
  • Percentiles, stanines, and grade equivalents are norm-referenced score types; mastery levels and cut scores are criterion-referenced.
  • Norm-referenced results are invalid for English learners when the norming sample excluded them.
  • Criterion-referenced results tell a teacher which specific skills were and were not mastered, which is what instructional planning requires.
  • Grade-equivalent scores are the most misinterpreted score type and should never be read as a recommended placement level.
Last updated: August 2026

5.6 Norm-Referenced and Criterion-Referenced Assessment

The ETS content outline names this distinction explicitly: the candidate "knows the difference between norm-referenced and criterion-referenced assessments and how they are used with ELs," and recognizes both the consequences of norm-referenced comparison for English learners and the instructional utility of criterion-referenced information. The distinction is not about the test items — it is about what the score is compared to.


1. The Core Contrast

Norm-referenced (NRT)Criterion-referenced (CRT)
Answers the question"How does this student compare to other students?""What can this student actually do?"
Reference pointA norming sample of test-takersA fixed standard, objective, or cut score
Score typesPercentile rank, stanine, normal curve equivalent, grade equivalent, standard scorePercent correct, mastery/non-mastery, performance level, proficiency-level descriptor
Item designItems chosen to spread scores out; very easy and very hard items are discardedItems chosen to represent the standard; everyone can, in principle, score 100%
DistributionDesigned to produce a bell curve; half of all students are below average by constructionNo fixed distribution; all students could reach mastery
Primary usesScreening, program eligibility, comparison to a populationInstructional planning, progress monitoring, grading, standards accountability
ExamplesCognitive ability batteries; nationally normed achievement testsUnit tests, state standards assessments, ELP performance-level determinations, rubric-scored performance tasks

The bell-curve consequence

A norm-referenced test is engineered so that scores spread. That means "below average" is a statistical position, not a diagnosis. Fifty percent of the norming population scores below the median by design. A parent told their child is "below the 50th percentile" has been given a ranking, not a description of what the child knows.


2. Decoding Norm-Referenced Score Types

Score typeWhat it meansCommon misreading
Percentile rankPercent of the norming group this student outscored (e.g., 62nd percentile = outscored 62%)Mistaken for percent correct
StanineNine-point band; 4-6 is the average rangeTreated as a fine-grained measure
Normal curve equivalent (NCE)Percentile rescaled to equal intervals so scores can be averagedConfused with percentile
Standard scoreDistance from the mean in standard-deviation units (often mean 100, SD 15)Read as a percentage
Grade equivalent (GE)The grade level at which this raw score would be the medianThe most dangerous score type on any report

The grade-equivalent trap

A 4th grader who earns a grade equivalent of 7.2 has not demonstrated 7th-grade skills. It means her raw score equals the median raw score of 7th graders on the 4th-grade test — a test containing no 7th-grade content. Grade equivalents are extrapolated, unequal-interval estimates and should never be used to set placement, select texts, or write goals. The same logic runs the other way: a 6th-grade EL with a GE of 2.4 has not been shown to "read like a 2nd grader."


3. Why Norm-Referenced Tests Distort English Learner Performance

This is the point the exam presses hardest, and it follows directly from the norming-sample bias discussed in Section 5.2.

  1. The comparison group is inappropriate. If the norming sample consisted overwhelmingly of native English speakers, then comparing a developing multilingual learner to it produces a ranking against a population the student is not a member of. The resulting percentile measures English proficiency as much as the target construct.
  2. Language becomes construct-irrelevant variance. On a nationally normed science or mathematics test delivered in English, an EL's score confounds content knowledge with English reading ability (Section 5.3).
  3. Consequences compound. Low norm-referenced scores drive placement into remedial tracks, exclusion from gifted screening (Section 4.7), and over-referral to special education (Section 5.5) — outcomes documented as disproportionality, not as accurate measurement.
  4. Cultural content bias. Items presuppose U.S.-specific background knowledge that is unrelated to the skill being measured.

Legitimate uses remain. Norm-referenced data is not useless: it can flag a student for further examination, and it is required for some eligibility determinations. The rule is that a norm-referenced score alone may never drive a high-stakes decision about an English learner. It must be triangulated with criterion-referenced data, native-language performance, work samples, and observation over time.


4. Why Criterion-Referenced Data Drives Instruction

Criterion-referenced results answer the question a teacher actually needs answered on Monday morning.

Report saysInstructional value
"38th percentile in reading"None. It ranks the student; it names no skill
"Mastered main idea and sequencing; not yet mastered inference from context clues"Directly actionable — the next lesson writes itself

This is why progress monitoring, rubric-scored performance tasks, and standards-based grading are criterion-referenced by design (Section 5.5), and why English language proficiency assessments report performance-level descriptorsEntering, Emerging, Developing, Expanding, Bridging, Reaching — rather than percentiles. A descriptor states what the student can do with language; a percentile states only where they rank.

A note on ELP assessments

Annual ELP assessments such as ACCESS for ELLs and ELPA21 are criterion-referenced against English language development standards. That is precisely why their results can be used for reclassification decisions (Section 5.5): a proficiency-level determination is a statement about attainment of a defined standard, not a ranking against other test-takers.


5. Choosing the Right Tool

PurposeUse
Decide what to teach tomorrowCriterion-referenced (formative, rubric, mastery check)
Determine whether a student met a standardCriterion-referenced
Make a reclassification/exit decisionCriterion-referenced ELP determination plus multiple measures
Screen a large population for follow-upNorm-referenced, with caution
Establish eligibility requiring population comparisonNorm-referenced, with native-language and nonverbal data alongside
Report growth to a familyCriterion-referenced descriptors and work samples, not percentiles
Test Your Knowledge

A 6th-grade multilingual learner's report shows a grade equivalent of 2.4 in reading. What does this score actually indicate?

A
B
C
D
Test Your Knowledge

A district uses a nationally normed English reading achievement test, normed on a predominantly native-English-speaking sample, as the sole basis for placing multilingual learners into remedial tracks. What is the primary technical objection?

A
B
C
D
Test Your Knowledge

Which assessment result is criterion-referenced and directly actionable for instructional planning?

A
B
C
D
Test Your Knowledge

Why do annual English language proficiency assessments such as ACCESS for ELLs report performance-level descriptors rather than percentile ranks?

A
B
C
D