5.6 Norm-Referenced and Criterion-Referenced Assessment
Key Takeaways
- Norm-referenced tests compare a student to a norming sample; criterion-referenced tests compare performance to a fixed standard.
- Percentiles, stanines, and grade equivalents are norm-referenced score types; mastery levels and cut scores are criterion-referenced.
- Norm-referenced results are invalid for English learners when the norming sample excluded them.
- Criterion-referenced results tell a teacher which specific skills were and were not mastered, which is what instructional planning requires.
- Grade-equivalent scores are the most misinterpreted score type and should never be read as a recommended placement level.
5.6 Norm-Referenced and Criterion-Referenced Assessment
The ETS content outline names this distinction explicitly: the candidate "knows the difference between norm-referenced and criterion-referenced assessments and how they are used with ELs," and recognizes both the consequences of norm-referenced comparison for English learners and the instructional utility of criterion-referenced information. The distinction is not about the test items — it is about what the score is compared to.
1. The Core Contrast
| Norm-referenced (NRT) | Criterion-referenced (CRT) | |
|---|---|---|
| Answers the question | "How does this student compare to other students?" | "What can this student actually do?" |
| Reference point | A norming sample of test-takers | A fixed standard, objective, or cut score |
| Score types | Percentile rank, stanine, normal curve equivalent, grade equivalent, standard score | Percent correct, mastery/non-mastery, performance level, proficiency-level descriptor |
| Item design | Items chosen to spread scores out; very easy and very hard items are discarded | Items chosen to represent the standard; everyone can, in principle, score 100% |
| Distribution | Designed to produce a bell curve; half of all students are below average by construction | No fixed distribution; all students could reach mastery |
| Primary uses | Screening, program eligibility, comparison to a population | Instructional planning, progress monitoring, grading, standards accountability |
| Examples | Cognitive ability batteries; nationally normed achievement tests | Unit tests, state standards assessments, ELP performance-level determinations, rubric-scored performance tasks |
The bell-curve consequence
A norm-referenced test is engineered so that scores spread. That means "below average" is a statistical position, not a diagnosis. Fifty percent of the norming population scores below the median by design. A parent told their child is "below the 50th percentile" has been given a ranking, not a description of what the child knows.
2. Decoding Norm-Referenced Score Types
| Score type | What it means | Common misreading |
|---|---|---|
| Percentile rank | Percent of the norming group this student outscored (e.g., 62nd percentile = outscored 62%) | Mistaken for percent correct |
| Stanine | Nine-point band; 4-6 is the average range | Treated as a fine-grained measure |
| Normal curve equivalent (NCE) | Percentile rescaled to equal intervals so scores can be averaged | Confused with percentile |
| Standard score | Distance from the mean in standard-deviation units (often mean 100, SD 15) | Read as a percentage |
| Grade equivalent (GE) | The grade level at which this raw score would be the median | The most dangerous score type on any report |
The grade-equivalent trap
A 4th grader who earns a grade equivalent of 7.2 has not demonstrated 7th-grade skills. It means her raw score equals the median raw score of 7th graders on the 4th-grade test — a test containing no 7th-grade content. Grade equivalents are extrapolated, unequal-interval estimates and should never be used to set placement, select texts, or write goals. The same logic runs the other way: a 6th-grade EL with a GE of 2.4 has not been shown to "read like a 2nd grader."
3. Why Norm-Referenced Tests Distort English Learner Performance
This is the point the exam presses hardest, and it follows directly from the norming-sample bias discussed in Section 5.2.
- The comparison group is inappropriate. If the norming sample consisted overwhelmingly of native English speakers, then comparing a developing multilingual learner to it produces a ranking against a population the student is not a member of. The resulting percentile measures English proficiency as much as the target construct.
- Language becomes construct-irrelevant variance. On a nationally normed science or mathematics test delivered in English, an EL's score confounds content knowledge with English reading ability (Section 5.3).
- Consequences compound. Low norm-referenced scores drive placement into remedial tracks, exclusion from gifted screening (Section 4.7), and over-referral to special education (Section 5.5) — outcomes documented as disproportionality, not as accurate measurement.
- Cultural content bias. Items presuppose U.S.-specific background knowledge that is unrelated to the skill being measured.
Legitimate uses remain. Norm-referenced data is not useless: it can flag a student for further examination, and it is required for some eligibility determinations. The rule is that a norm-referenced score alone may never drive a high-stakes decision about an English learner. It must be triangulated with criterion-referenced data, native-language performance, work samples, and observation over time.
4. Why Criterion-Referenced Data Drives Instruction
Criterion-referenced results answer the question a teacher actually needs answered on Monday morning.
| Report says | Instructional value |
|---|---|
| "38th percentile in reading" | None. It ranks the student; it names no skill |
| "Mastered main idea and sequencing; not yet mastered inference from context clues" | Directly actionable — the next lesson writes itself |
This is why progress monitoring, rubric-scored performance tasks, and standards-based grading are criterion-referenced by design (Section 5.5), and why English language proficiency assessments report performance-level descriptors — Entering, Emerging, Developing, Expanding, Bridging, Reaching — rather than percentiles. A descriptor states what the student can do with language; a percentile states only where they rank.
A note on ELP assessments
Annual ELP assessments such as ACCESS for ELLs and ELPA21 are criterion-referenced against English language development standards. That is precisely why their results can be used for reclassification decisions (Section 5.5): a proficiency-level determination is a statement about attainment of a defined standard, not a ranking against other test-takers.
5. Choosing the Right Tool
| Purpose | Use |
|---|---|
| Decide what to teach tomorrow | Criterion-referenced (formative, rubric, mastery check) |
| Determine whether a student met a standard | Criterion-referenced |
| Make a reclassification/exit decision | Criterion-referenced ELP determination plus multiple measures |
| Screen a large population for follow-up | Norm-referenced, with caution |
| Establish eligibility requiring population comparison | Norm-referenced, with native-language and nonverbal data alongside |
| Report growth to a family | Criterion-referenced descriptors and work samples, not percentiles |
A 6th-grade multilingual learner's report shows a grade equivalent of 2.4 in reading. What does this score actually indicate?
A district uses a nationally normed English reading achievement test, normed on a predominantly native-English-speaking sample, as the sole basis for placing multilingual learners into remedial tracks. What is the primary technical objection?
Which assessment result is criterion-referenced and directly actionable for instructional planning?
Why do annual English language proficiency assessments such as ACCESS for ELLs report performance-level descriptors rather than percentile ranks?