5.3 Score Interpretation: Criterion- vs. Norm-Referenced Metrics
Key Takeaways
- Criterion-Referenced Tests (CRTs) measure student performance against predefined standards or cut scores (absolute mastery), whereas Norm-Referenced Tests (NRTs) rank students against a representative norm group.
- Raw scores represent the simple count of correct items, percentage scores reflect the proportion of correct items, and scaled scores mathematically transform raw scores to account for varying test form difficulty.
- Percentile ranks (PR 1–99) indicate the percentage of students in the norm group who scored at or below a student's score; a PR of 82 does NOT mean 82% correct.
- Stanines (Standard Nines) compress scores into a 9-point scale with a mean of 5 and a standard deviation of 2, where stanines 1–3 are below average, 4–6 are average, and 7–9 are above average.
- Grade-Equivalent (GE) scores indicate that a student performed as an average student in that grade would have performed ON THE SAME GRADE-LEVEL MATERIAL; a 4th grader scoring GE 7.2 should NOT be moved to 7th grade.
5.3 Score Interpretation: Criterion- vs. Norm-Referenced Metrics
Interpreting assessment scores accurately is an essential competency for Florida educators. Under Competency 4 of the FTCE Professional Education Test, teachers must analyze diagnostic, classroom, and standardized assessment reports to diagnose student learning needs, differentiate instruction, and communicate performance accurately to parents, administrators, and educational specialists.
Misinterpreting standardized derived scores—particularly percentile ranks and grade equivalents—can lead to flawed pedagogical placements and inappropriate parent expectations.
1. Criterion-Referenced vs. Norm-Referenced Assessment Systems
The fundamental distinction in educational measurement lies in the frame of reference used to interpret a student's score: whether performance is judged against an absolute standard of mastery (Criterion-Referenced) or against the relative performance of a peer group (Norm-Referenced).
+-----------------------------------------------------------------------------------------+
| CRITERION- VS. NORM-REFERENCED FRAMEWORK |
| |
| [CRITERION-REFERENCED TESTS (CRTs)] [NORM-REFERENCED TESTS (NRTs)] |
| • Focus: What can the student do? • Focus: How does the student compare? |
| • Standard: Fixed benchmarks / cut scores • Standard: Normal distribution curve |
| • Ranking: No student-to-student comparison • Ranking: Rank-orders students (1-99) |
| • Scoring: Pass/Fail, Mastery Tiers (1 to 5) • Scoring: Percentiles, Stanines, NCEs |
| • Examples: Florida FAST, FTCE Exam, EOCs • Examples: SAT, ACT, Stanford 10, WISC |
+-----------------------------------------------------------------------------------------+
Comprehensive Comparison: CRTs vs. NRTs
| Assessment Dimension | Criterion-Referenced Tests (CRTs) | Norm-Referenced Tests (NRTs) |
|---|---|---|
| Core Purpose | Determine whether a student has achieved mastery of specific learning standards or behavioral criteria. | Discriminate between high and low performers and rank students across a broad continuum of capability. |
| Reference Standard | Fixed, predetermined objective standards (e.g., "Can solve two-step linear equations with 80% accuracy"). | A representative sample of peers who took the same test under identical conditions (the Norm Group). |
| Score Interpretation | Absolute: A student's score is independent of how classmates or peers performed. In theory, 100% of students could achieve mastery. | Relative: A student's score depends entirely on peer performance. Half the test-takers must fall below the 50th percentile by definition. |
| Item Selection | Items sample targeted state benchmarks directly; items where all students answer correctly are retained if aligned. | Items with broad difficulty (p = 0.50) that maximize variance and spread out student scores across a bell curve are prioritized. |
| Primary Educational Use | Classroom grading, competency certification, diagnostic mastery checks, state accountability (FAST, FTCE). | College admissions (SAT/ACT), gifted program screening, national norm benchmarking, special education eligibility. |
| Florida / Real-World Example | Florida FAST (Levels 1–5), FTCE Professional Education Test (Passing scaled score ≥ 200), Driver's license road test. | SAT Reasoning Test, ACT, Iowa Assessments, Stanford Achievement Test 10 (SAT-10), Woodcock-Johnson. |
[!NOTE] FTCE Exam Passing Metric: The FTCE Professional Education Test is a Criterion-Referenced Test (CRT). Candidates must achieve a scaled score of 200 or higher to pass. Your score is based solely on your demonstrated mastery of the 8 Competencies, not on how other test-takers perform on that test day.
2. Types of Assessment Scores: Raw, Percentage, and Scaled
When scoring assessments, three distinct score formats are commonly generated:
+-----------------------------------------------------------------------------------------+
| THE THREE PRIMARY SCORE FORMATS |
| |
| [1. RAW SCORE] ---> Simple count of correctly answered items (e.g., 42 / 50). |
| Unadjusted; incomparable across different test versions. |
| |
| [2. PERCENTAGE SCORE] ---> Proportion of total items answered correctly (e.g., 84%). |
| Reflects content proportion, but ignores item difficulty. |
| |
| [3. SCALED SCORE] ---> Mathematical conversion placing raw scores onto a standardized|
| common scale (e.g., FTCE 200, SAT 200-800, FAST Levels 1-5).|
| Adjusts statistically for variations in test form difficulty.|
+-----------------------------------------------------------------------------------------+
Why Standardized Tests Use Scaled Scores (Equating)
Different versions (forms) of a standardized test administered across different testing windows vary slightly in difficulty despite rigorous authoring blueprints. Test equating mathematically converts raw scores into scaled scores so that a scaled score of 200 on Form A represents the exact same level of underlying student competence as a 200 on a slightly harder Form B.
3. Standardized Derived Metrics: Percentiles, Stanines, and Grade Equivalents
Standardized score reports provide several derived metrics that translate student performance into normative terms.
+-----------------------------------------------------------------------------------------+
| THE NORMAL DISTRIBUTION SCORE CONVERSION |
| |
| Percent of Cases: 2.14% | 13.59% | 34.13% | 34.13% | 13.59% | 2.14% |
| |
| Standard Deviations: -3σ -2σ -1σ 0 +1σ +2σ +3σ |
| |---------|---------|---------|---------|---------|----------| |
| z-Scores: -3.0 -2.0 -1.0 0 +1.0 +2.0 +3.0 |
| T-Scores: 20 30 40 50 60 70 80 |
| Stanines (1-9): 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 |
| Percentile Ranks: 0.1% 2% 16% 50% 84% 98% 99.9%|
+-----------------------------------------------------------------------------------------+
Detailed Analysis of Standardized Metrics
| Derived Metric | Scale & Properties | Pedagogical Interpretation & Meaning | Common Misconception to Avoid |
|---|---|---|---|
| Percentile Rank (PR) | Scale: 1 to 99<br>Median: 50 | Indicates the percentage of students in the normative comparison group who scored at or below the student's score. (e.g., PR 75 = student scored higher than or equal to 75% of norm peers). | CRITICAL ERROR: A percentile rank of 75 does NOT mean the student answered 75% of questions correctly. It is a rank, not a percentage score. |
| Stanine (Standard Nine) | Scale: 1 to 9<br>Mean: 5<br>SD: 2 | Compresses the normal distribution into 9 discrete bands:<br>• Stanines 1–3: Below Average (bottom 23%)<br>• Stanines 4–6: Average (middle 54%)<br>• Stanines 7–9: Above Average (top 23%) | Stanines are broad score bands. Small raw score shifts within the same stanine (e.g., moving within Stanine 5) do not indicate meaningful change. |
| Grade-Equivalent (GE) | Format: Grade.Month (e.g., 4.6 = 4th grade, 6th month) | Indicates that the student's raw score is equal to the average raw score earned by a norm student in that grade/month when taking that same test. | MOST TESTED TRAP: A 4th grader scoring GE 7.2 does NOT mean they can do 7th-grade math; it means they scored as an average 7th grader would on a 4th-grade test. |
| z-Score | Mean: 0<br>SD: 1 | Indicates the exact number of standard deviation units a raw score lies above (+) or below (-) the population mean: z = (X - μ) / σ. | z-Scores include negative numbers and decimals, making them unsuited for parent score reports. |
| T-Score | Mean: 50<br>SD: 10 | Linear transformation of z-scores (T = 50 + 10z) that eliminates negative numbers and decimals. Widely used in behavioral assessments (e.g., BASC-3). | A T-score of 60 represents exactly +1 standard deviation above the mean (84th percentile). |
[!CAUTION] The Grade-Equivalent (GE) Parent Conference Trap on FTCE: Scenario questions frequently present an elementary student (e.g., 3rd grade) who earns a high Grade-Equivalent score (e.g., GE 6.4) on a norm-referenced math test. The parent demands that the child be double-promoted to 6th-grade math. The correct teacher response is to explain that the score demonstrates superior mastery of 3rd-grade math benchmarks, not readiness for 6th-grade algebra, and provide within-class differentiated enrichment and deeper application tasks.
During a parent-teacher conference, the mother of a 4th-grade student reviews her child's standardized norm-referenced reading score report. The report indicates a Grade-Equivalent (GE) score of 7.4. The parent excitedly requests that her child be immediately accelerated into a 7th-grade English Language Arts classroom. What is the most pedagogically accurate explanation the teacher should provide?
A middle school guidance counselor is reviewing standardized assessment results for an incoming 6th-grade student. The student's score profile shows a Stanine of 8 in Mathematics Concepts and a Stanine of 2 in Reading Comprehension. How should the counselor interpret these performance bands?
Which of the following testing scenarios best illustrates the administration and interpretation of a Criterion-Referenced Assessment?