14.3 Mean, Median, Mode, Standard Deviation & Derived Scores
Key Takeaways
The mean is sensitive to extreme scores, so the median better represents a typical student when scores are skewed.
Standard deviation describes spread; in a normal distribution about 68% of scores fall within one standard deviation of the mean.
A percentile rank is the percentage of the norm group scoring at or below a student, never the percentage of items answered correctly.
Stanines have a mean of 5 and a standard deviation of 2; z-scores have a mean of 0 and SD of 1; T-scores have a mean of 50 and SD of 10.
A grade-equivalent score of 10.2 for a 7th grader means the student did as well as a typical 10th grader would on the 7th-grade test, not that the student has mastered 10th-grade content.
Why Score Terminology Matters
ETS asks candidates to define and explain terms related to testing and scoring: validity, reliability, raw score, scaled score, percentile, standard deviation, mean, mode, median, grade-equivalent scores, and age-equivalent scores. Validity and reliability are covered in the previous section. This section starts with the basic statistics and then turns to derived scores.
Central Tendency and Variability: A Worked Example
Seven students score 62, 70, 75, 75, 80, 84, and 100 on a quiz.
| Statistic | Definition | Result |
|---|---|---|
| Mean | Sum of scores divided by the number of scores: | 546 ÷ 7 = 78 |
| Median | Middle score when scores are in order | 75 (the fourth of seven scores) |
| Mode | Most frequent score | 75 |
| Range | Highest minus lowest | 100 − 62 = 38 |
| Standard deviation | Typical distance of scores from the mean | About 11.1 (dividing by n; about 12.0 if dividing by n − 1) |
Notice that the single score of 100 pulls the mean (78) above the median (75). Without that score, the other six average about 74.3 while the median stays at 75. The mean is sensitive to extreme scores, so the median better represents a typical student when scores are skewed. The mode is most useful for the most common response or category.
Standard deviation tells how spread out the scores are. A small standard deviation means scores cluster near the mean; a large one means they vary widely. In a roughly normal distribution, about 68 percent of scores fall within one standard deviation of the mean and about 95 percent within two.
Skew matters on classroom tests. When most students score high and a few score low (common on a mastery quiz), the distribution is negatively skewed and the mean falls below the median. When most score low and a few score high, it is positively skewed and the mean rises above the median.
Raw Scores vs. Derived and Scaled Scores
Raw Scores and Their Inherent Limitations
A raw score is the unadjusted, direct count of points earned or questions answered correctly on an assessment (e.g., 42 out of 50 correct). While raw scores provide basic classroom information, they possess fatal psychometric limitations when used for comparative analysis:
- Lack of Meaningful Comparison: A raw score of 35 on a physics exam cannot be compared to a raw score of 35 on an English exam, because the tests differ in total items, item difficulty, and cognitive complexity.
- Form Inequivalence: If Form A of a state biology test contains slightly more demanding genetics items than Form B, an identical raw score of 40 represents different levels of underlying student ability.
Scaled Scores: Statistical Equating Across Test Forms
To eliminate form inequivalence, psychometricians transform raw scores into scaled scores through a mathematical process known as equating. Equating adjusts for slight differences in difficulty across various test administrations, placing scores onto a continuous, invariant numerical scale (e.g., the SAT 200–800 scale or state assessment scales ranging from 300 to 900). A scaled score of 650 reflects the exact same level of proficiency regardless of whether the student took the exam in May or November.
Deconstructing Derived Scores and Measurement Scales
Derived scores convert raw performance into meaningful comparative metrics. Secondary educators must master four primary derived score types:
[ Raw Score ]
|
+----------------+----------------+
| |
[ Scaled Scores ] [ Derived Scores ]
(Equated across forms; (Comparative interpretations)
e.g., SAT 200-800) |
+------------------+------------------+
| | |
[ Percentile ] [ Stanines ] [ Standard Scores ]
Ranks (1-99) (1-9) (z-scores, T-scores)
1. Percentile Ranks (PR)
A percentile rank indicates the percentage of students in the normative comparison group who scored at or below a given student's score. Percentile ranks range from the 1st percentile to the 99th percentile (a percentile of 100 does not exist because a student cannot score higher than 100% of the sample, which includes themselves).
Important
The Cardinal Rule of Percentile Ranks: A percentile rank is never the percentage of questions answered correctly on the test. If a 10th grader scores in the 88th percentile on a standardized geometry test, it means the student performed equal to or better than 88% of the national normative comparison group. It does not mean the student answered 88% of the geometry questions correctly (their raw score might have been only 65% on an extraordinarily difficult test).
- Non-Linear Interval Warning: Percentile ranks are ordinal, not interval, metrics. Because test scores cluster heavily in the middle of a normal distribution, a small change in raw score near the 50th percentile causes a massive jump in percentile rank (e.g., from the 48th to the 58th percentile). Conversely, at the extreme tails of the distribution, a large increase in raw score produces only a tiny change in percentile rank (e.g., from the 97th to the 98th percentile). Educators must never average percentile ranks arithmetically.
2. Stanines ("Standard Nines")
The stanine (short for "standard nine") scale divides the normal distribution into nine standardized numerical units ranging from 1 to 9. The scale is constructed with a theoretical mean of 5 and a standard deviation of 2.
- Stanine Bands:
- Stanines 1, 2, and 3: Below Average performance (representing the lowest 23% of test-takers).
- Stanines 4, 5, and 6: Average performance (representing the middle 54% of test-takers, with Stanine 5 spanning the exact middle 20%).
- Stanines 7, 8, and 9: Above Average performance (representing the top 23% of test-takers, with Stanine 9 representing the top 4%).
- Practical Utility: Stanines condense granular raw scores into broad, manageable bands. This prevents educators and parents from over-interpreting minor, statistically insignificant fluctuations in raw test scores.
3. Standard Scores: z-Scores and T-Scores
Standard scores express an individual's distance from the mean in standardized standard deviation units, allowing educators to compare student performance across completely different subject tests.
- z-Scores: The fundamental baseline standard score. A z-score has a mean (μ) of 0 and a standard deviation (σ) of 1.
Formula:
z = (X - μ) / σWhereXis the raw score,μis the group mean, andσis the standard deviation. A student with a z-score of+1.5performed 1.5 standard deviations above the average test-taker.- Limitation: z-scores involve negative numbers (for below-average performances) and decimals, which are confusing and distressing when communicated to parents and adolescents.
- T-Scores: A transformed standard score designed specifically to eliminate negative numbers and decimals. A T-score has a mean of 50 and a standard deviation of 10.
Formula:
T = 50 + 10(z)- A z-score of
0converts to a T-score of50(exactly average). - A z-score of
+1.0converts to a T-score of60(one standard deviation above the mean, roughly the 84th percentile). - A z-score of
-2.0converts to a T-score of30(two standard deviations below the mean, roughly the 2nd percentile). - T-scores are widely used in secondary psychological evaluations, behavioral assessments (such as the BASC), and clinical diagnostics.
- A z-score of
The Normal Distribution Curve and the Empirical Rule
The normal distribution (bell curve) is a mathematically symmetrical, unimodal distribution where the mean, median, and mode all coincide at the exact center (0 standard deviations).
The 68-95-99.7 Empirical Rule
When data conforms to a normal distribution, the distribution of scores across standard deviation units (σ) follows invariant proportions:
- 68.2% of all scores fall within one standard deviation of the mean (between -1σ and +1σ; between T = 40 and T = 60).
- 95.4% of all scores fall within two standard deviations of the mean (between -2σ and +2σ; between T = 30 and T = 70).
- 99.7% of all scores fall within three standard deviations of the mean (between -3σ and +3σ; between T = 20 and T = 80).
THE NORMAL DISTRIBUTION
MEAN
z = 0
T = 50
PR = 50
Stanine 5
|
+---+---+
/| | |\
/ | | | \
/ | | | \
/ | | | \
/ | | | \
/34.1%|34.1% \
/ | | \
--+-------+-------+------+--
13.6% 13.6%
2.1% 2.1%
-------------------------------------------------------------------------
Standard Deviations: -3σ -2σ -1σ 0 +1σ +2σ +3σ
z-Score: -3.0 -2.0 -1.0 0 +1.0 +2.0 +3.0
T-Score: 20 30 40 50 60 70 80
Percentile Rank: 0.1 2 16 50 84 98 99.9
Stanines: 1 2 3 4 5 6 7 8 9
Stanine Range: Lowest 23% Middle 54% Top 23%
Comparative Metric Alignment
| Standard Deviations (σ) | z-Score | T-Score | Percentile Rank | Stanine | Descriptive Classification |
|---|---|---|---|---|---|
| +3.0σ | +3.0 | 80 | 99.9th | 9 | Exceptionally High / Gifted Range |
| +2.0σ | +2.0 | 70 | 97.7th (~98th) | 9 | Well Above Average (Top 2%) |
| +1.5σ | +1.5 | 65 | 93.3rd (~93rd) | 8 | Above Average |
| +1.0σ | +1.0 | 60 | 84.1st (~84th) | 7 | High Average (Top 16%) |
| 0.0σ (Mean) | 0.0 | 50 | 50.0th | 5 | Exactly Average |
| -1.0σ | -1.0 | 40 | 15.9th (~16th) | 3 | Low Average (Bottom 16%) |
| -2.0σ | -2.0 | 30 | 2.3rd (~2nd) | 1 | Well Below Average (Bottom 2%) |
| -3.0σ | -3.0 | 20 | 0.1st | 1 | Exceptionally Low / Clinical Deficit |
Grade-Equivalent (GE) and Age-Equivalent (AE) Scores: The Pervasive Pitfall
A Grade-Equivalent (GE) score is a developmental metric expressed as a decimal number representing grade level and month of the school year (e.g., 8.4 corresponds to an 8th-grade student in the fourth month of school, typically December).
The Fatal Misinterpretation: Mastery vs. Relative Performance
Grade-equivalent scores are the single most misinterpreted psychometric metric in secondary education. Parents, administrators, and inexperienced educators routinely make the catastrophic error of assuming that a GE score reflects curricular placement:
Caution
The Classic GE Trap: A 7th-grade student takes a 7th-grade standardized mathematics test in October and receives a Grade-Equivalent score of 10.2.
Incorrect Interpretation: "The student has mastered 10th-grade geometry and algebra, is doing 10th-grade work, and should immediately be accelerated into a high school sophomore math class."
Correct Psychometric Interpretation: "The 7th-grade student answered the 7th-grade math questions as accurately as an average 10th-grade student in the second month of school would perform if that 10th grader were given the exact same 7th-grade math test."
Why GE Scores Must Never Guide Curricular Acceleration
- The Student Never Saw High School Content: The 7th-grade test contained questions on ratios, proportional relationships, and basic integer operations. It contained zero items on quadratic equations, geometric proofs, or polynomial factoring. The student simply demonstrated extraordinary mastery of 7th-grade standards.
- Statistical Extrapolation and Artificial Curves: Test publishers rarely administer 7th-grade tests to actual 10th-grade students. Instead, GE scores above the grade level are generated through mathematical extrapolation—projecting lines outward along an assumed developmental trajectory. These projected scores are statistical artifacts, not empirical measurements of advanced curricular mastery.
- Age-Equivalent (AE) Scores: Age-equivalent scores (e.g., 14.6, meaning 14 years, 6 months) suffer from the exact same psychometric limitation. A 10-year-old scoring an AE of 15 on a reading test read 5th-grade text with the speed and comprehension of an average 15-year-old; they were not tested on 10th-grade literary theory.
In January, a 7th-grade student completes a standardized, nationally normed reading comprehension exam and receives a Grade-Equivalent (GE) score of 10.4. During an academic conference, the student's guardian insists that the child be immediately transferred to a 10th-grade Advanced Placement English course. Which response represents the most psychometrically sound and professional explanation the educator should provide?
Agree with the guardian's request, because a GE score of 10.4 provides empirical evidence that the student has mastered high school literature and composition curricula
Advise the guardian that grade-equivalent scores are psychometrically invalid metrics that should be completely disregarded in favor of raw scores
Explain that the score indicates the student performed on 7th-grade reading passages as effectively as an average 10th grader would on that same 7th-grade test, not that the student has mastered 10th-grade literary standards
Inform the guardian that the student scored in the top 10.4% of all 7th-grade students nationwide on the reading assessment
A high school sophomore takes a standardized norm-referenced achievement test with a normal distribution. On the biology subtest, the student earns a T-score of 70. Given that the T-score distribution has a mean of 50 and a standard deviation of 10, which statement accurately characterizes the student's performance?
The student answered exactly 70% of the questions on the biology examination correctly
The student scored one standard deviation above the normative national mean, placing them in Stanine 5
The student scored two standard deviations below the mean, placing them in the bottom 2% of test-takers
The student scored two standard deviations above the mean, placing them at approximately the 98th percentile
Seven students score 62, 70, 75, 75, 80, 84, and 100 on a quiz. Which statement is correct?
The mean is 78 and the median is 75, and the score of 100 pulls the mean above the median
The mean and median are both 75 because 75 is the mode
The median is 80 because it is the middle of the range from 62 to 100
The mean is 75 because extreme scores do not affect the mean
Sections you finish are checked off in the contents.