4.3 Score Interpretation, Norms & Testing Fairness

Key Takeaways

  • The normal bell curve distribution is symmetrical and unimodal with Mean = Median = Mode, governed by the empirical 68-95-99.7 rule; skewness pulls the mean toward the tail (Positive Skew: Mode < Median < Mean; Negative Skew: Mean < Median < Mode).
  • Standard scores convert raw scores into standardized metrics expressing distance from the mean in standard deviation units: z-scores (M=0, SD=1), T-scores (M=50, SD=10), Wechsler Deviation IQ (M=100, SD=15), and Stanines (M=5, SD=2).
  • Percentile ranks represent the percentage of examinees in the norm group scoring at or below a given raw score; because percentiles represent an ordinal scale with non-linear spacing, raw score shifts near the median create exaggerated percentile rank changes compared to identical shifts in the tails.
  • Norm-referenced tests (NRT) evaluate relative standing against a representative peer group, while criterion-referenced tests (CRT) evaluate mastery against an absolute performance cutoff or domain standard.
  • Landmark legal cases (Larry P. v. Riles, Diana, Griggs) and the ACA Code of Ethics Section E establish strict mandates regarding testing fairness, cultural bias, testing accommodations, test user qualification levels (A, B, C), and the restricted release of raw test data.
Last updated: September 2026

4.3 Score Interpretation, Norms & Testing Fairness

Quick Summary: Converting raw assessment data into meaningful, ethically defensible clinical interpretations is a vital competency for professional counselors. Counselors must master the mathematical properties of the normal bell curve distribution, understand how distribution skewness affects central tendency, and fluently translate between various standard score metrics (z-scores, T-scores, Wechsler Deviation IQ, stanines) and percentile ranks. Furthermore, clinicians must adhere to legal and ethical standards governing assessment fairness, mitigating cultural and linguistic test bias, implementing ADA accommodations, and strictly enforcing the American Counseling Association (ACA) Code of Ethics Section E standards on user competence and test security.


The Normal Bell Curve Distribution (Gaussian Distribution)

The normal distribution—frequently termed the bell curve or Gaussian distribution—is a theoretical, symmetrical mathematical probability distribution that underpins classical psychometric theory. Many biological, physiological, and psychological attributes (e.g., height, blood pressure, general cognitive ability) distribute normally across large, unselected populations.

Mathematical Properties of the Normal Curve

  1. Symmetry: The curve is perfectly bilateral; the left half is an exact mirror image of the right half.
  2. Unimodal Alignment of Central Tendency: In a perfectly normal distribution, the Mean, Median, and Mode are exactly identical and coincide at the exact center of the distribution ($z = 0$).
  3. Asymptotic Tails: The tails of the curve extend infinitely in both directions, approaching the horizontal baseline but never actually touching it ($-\infty$ to $+\infty$).
  4. Area Under the Curve: The total area under the normal curve represents 100% (or a probability of 1.00) of all observed scores.

The Empirical Rule (68-95-99.7 Rule)

Psychometricians rely on the empirical rule to determine the percentage of scores falling within specific standard deviation ($SD$) units from the mean:

  • 68.26% of all scores fall within $\pm 1$ Standard Deviation of the mean ($34.13%$ between the mean and $+1 SD$, and $34.13%$ between the mean and $-1 SD$).
  • 95.44% of all scores fall within $\pm 2$ Standard Deviations of the mean ($13.59%$ between $+1 SD$ and $+2 SD$, and $13.59%$ between $-1 SD$ and $-2 SD$).
  • 99.73% of all scores fall within $\pm 3$ Standard Deviations of the mean ($2.14%$ between $+2 SD$ and $+3 SD$, and $2.14%$ between $-2 SD$ and $-3 SD$). Only 0.13% of scores fall in each extreme outer tail beyond $\pm 3 SD$.
                  Percent of Scores Under the Normal Curve
                       
                              34.13%   34.13%
                           ┌─────────┬─────────┐
                    13.59% │         │         │ 13.59%
                 ┌─────────┤         │         ├─────────┐
           2.14% │         │         │         │         │ 2.14%
        ──┬──────┴─────────┴─────────┼─────────┴─────────┴──────┬──
         -3 SD    -2 SD     -1 SD   Mean     +1 SD     +2 SD   +3 SD

Skewness: Asymmetry in Score Distributions

When an assessment distribution departs from symmetry, it is skewed. On the NCE exam, skewness is always defined by the direction in which the elongated tail points:

1. Positively Skewed Distribution (Right-Skewed)

  • Visual Shape: The elongated tail points toward the positive (right) end of the horizontal axis. The vast majority of scores cluster at the lower (left) end of the scale.
  • Clinical Example: An exceptionally difficult examination where almost all examinees score poorly, or population distributions of annual income where most individuals earn modest wages while a few billionaires pull the tail to the right.
  • Central Tendency Relationship: Outliers in the positive tail pull the sensitive arithmetic mean to the right. Therefore:

Mode<Median<Mean\mathbf{Mode < Median < Mean}

The Mode represents the peak, the Median represents the 50th percentile, and the Mean is pulled furthest into the positive tail.

2. Negatively Skewed Distribution (Left-Skewed)

  • Visual Shape: The elongated tail points toward the negative (left) end of the horizontal axis. The vast majority of scores cluster at the higher (right) end of the scale.
  • Clinical Example: An overly easy classroom mastery test where almost all students earn 95% or 100%, with a small handful of unprepared students pulling the tail to the left.
  • Central Tendency Relationship: Outliers in the negative tail pull the arithmetic mean to the left. Therefore:

Mean<Median<Mode\mathbf{Mean < Median < Mode}

The Mean is pulled furthest into the negative tail, followed by the Median, with the Mode located at the high peak.

[!TIP] NCE Skewness Memory Rule: The tail tells the tale. If the tail points to the right (positive numbers), it is positively skewed. If the tail points to the left (negative numbers), it is negatively skewed. The Mean always follows the tail!

Loading diagram...
Distribution Skewness and Central Tendency Relationships

Kurtosis: Peakedness and Tail Weight

Kurtosis describes the steepness, peakedness, and tail thickness of a distribution relative to the normal curve:

  • Mesokurtic: The standard, baseline normal bell curve distribution ($Kurtosis = 0$).
  • Leptokurtic: A distribution that is tall, sharply peaked, with thin shoulders and heavy tails ($Kurtosis > 0$). Scores cluster tightly around the mean with minimal variance.
  • Platykurtic: A distribution that is flat, broad, and shallow ($Kurtosis < 0$). Scores are widely dispersed across the continuum.

Raw Scores vs. Standardized Scores

The Inadequacy of Raw Scores

A raw score is the immediate arithmetic count of items answered correctly, points earned, or rating responses summed on an assessment (e.g., scoring 38 out of 50 on an anxiety checklist). Raw scores are inherently uninterpretable in isolation because they provide zero contextual information regarding test difficulty, group variability, or how an examinee compares to relevant peers.

Standard Scores

To make scores clinically meaningful, raw scores are converted into standard scores. Standard scores transform raw data into a fixed, standardized distribution with a predetermined Mean and Standard Deviation, expressing an examinee's exact distance from the normative mean in standard deviation units.

1. z-Scores

The z-score is the fundamental building block of all standard scores. It has a Mean of 0 and a Standard Deviation of 1:

z=XMSDz = \frac{X - M}{SD}

where $X$ is the examinee's raw score, $M$ is the normative group mean, and $SD$ is the standard deviation. A $z$-score directly indicates how many standard deviations an individual falls above ($+z$) or below ($-z$) the mean. A score equal to the mean has $z = 0$. While mathematically elegant, $z$-scores carry clinical disadvantages: they contain negative numbers (which can distress clients) and decimal points.

2. T-Scores

Developed by educational psychologist William A. McCall, the T-score transforms $z$-scores into a user-friendly scale with a Mean of 50 and a Standard Deviation of 10:

T=50+10(z)T = 50 + 10(z)

T-scores eliminate negative numbers and decimals entirely. Widely used on the MMPI-2-RF/MMPI-3, MCMI, BASC, and symptom inventories. A $z$-score of $+1.5$ corresponds to $T = 50 + 10(1.5) = 65$ (the MMPI clinical elevation threshold).

3. Wechsler Deviation IQ Scores

Wechsler transformed cognitive ability scores into a standard distribution with a Mean of 100 and a Standard Deviation of 15:

Wechsler IQ=100+15(z)\text{Wechsler IQ} = 100 + 15(z)

  • Average Range: IQ 85 to 115 (within $\pm 1 SD$, encompassing $68.26%$ of the population).
  • Intellectual Disability Boundary: IQ $\le 70$ (2 SD below the mean, representing the bottom $2.27%$ of the population).
  • Giftedness Threshold: IQ $\ge 130$ (2 SD above the mean, representing the top $2.27%$ of the population).
  • (Note: The Stanford-Binet 5 also utilizes $M = 100, SD = 15$, whereas older Stanford-Binet editions utilized $SD = 16$.)

4. Stanines ("Standard Nines")

Developed by the U.S. Army Air Forces during World War II, stanines divide the normal curve into nine discrete integer units ranging from 1 to 9, with a Mean of 5 and a Standard Deviation of 2:

  • Stanine 5 represents the exact average middle 20% of the distribution (spanning from $z = -0.25$ to $+0.25$).
  • Stanines 4, 5, and 6 represent the broad average band (encompassing 54% of scores).
  • Stanine 1 represents the bottom 4% of performers ($z < -1.75$).
  • Stanine 9 represents the top 4% of performers ($z > +1.75$).
  • Widely used in K-12 standardized educational achievement reporting (e.g., Iowa Assessments).

5. Sten Scores ("Standard Tens")

Utilized primarily on Raymond Cattell's Sixteen Personality Factor Questionnaire (16PF), sten scores divide the normal curve into 10 integer bands (1 to 10) with a Mean of 5.5 and a Standard Deviation of 2.


Percentile Ranks (PR)

A percentile rank indicates the exact percentage of individuals in a normative reference group who scored at or below a specific raw score:

PR=(B+0.5EN)×100PR = \left( \frac{B + 0.5E}{N} \right) \times 100

where $B$ is the number of scores below the client, $E$ is the number of examinees tied at the score, and $N$ is total sample size. Percentile ranks range from 1 to 99 (the 0th and 100th percentiles do not exist in continuous distributions, as an individual cannot score below 0% or above 100% of everyone including themselves).

Normal Curve Percentile Equivalents

  • $z = -3.0 \rightarrow \text{Wechsler IQ } 55 \rightarrow T = 20 \rightarrow \mathbf{0.1\text{st Percentile}}$
  • $z = -2.0 \rightarrow \text{Wechsler IQ } 70 \rightarrow T = 30 \rightarrow \mathbf{2\text{nd Percentile}} \ (2.27%)$
  • $z = -1.0 \rightarrow \text{Wechsler IQ } 85 \rightarrow T = 40 \rightarrow \mathbf{16\text{th Percentile}} \ (15.87%)$
  • $z = 0.0 \rightarrow \text{Wechsler IQ } 100 \rightarrow T = 50 \rightarrow \mathbf{50\text{th Percentile}}$ (Median)
  • $z = +1.0 \rightarrow \text{Wechsler IQ } 115 \rightarrow T = 60 \rightarrow \mathbf{84\text{th Percentile}} \ (84.13%)$
  • $z = +2.0 \rightarrow \text{Wechsler IQ } 130 \rightarrow T = 70 \rightarrow \mathbf{98\text{th Percentile}} \ (97.72%)$
  • $z = +3.0 \rightarrow \text{Wechsler IQ } 145 \rightarrow T = 80 \rightarrow \mathbf{99.9\text{th Percentile}}$

The Crucial Psychometric Trap: Percentiles are an Ordinal Scale

[!CAUTION] The Percentile Rank Distortion Trap: Percentile ranks are an ordinal scale, NOT an interval scale. In a normal distribution, scores bunch up densely around the mean. Consequently, the distance between percentile ranks is non-linear:

  • Near the median (center of the curve), a tiny 2-point increase in raw score can leap a client from the 45th to the 55th percentile (a dramatic 10-percentile jump).
  • Out in the extreme tails (e.g., between IQ 135 and 137), an identical 2-point raw score increase might only move a client from the 99.0th to the 99.3rd percentile (a negligible 0.3-percentile shift).

Never average, add, or subtract percentile ranks! Doing so produces severe mathematical distortion. Standard scores ($z, T, IQ$) must be used for statistical operations.

Loading diagram...
Standard Score and Percentile Normal Curve Conversions

Standard Score Metric Comparison Matrix

Score MetricMean ($M$)Standard Deviation ($SD$)Formula / DerivationScale TypeClinical Elevation / Landmark ScorePrimary Clinical / Educational Uses
z-Score01$z = \frac{X - M}{SD}$Interval$+1.0$ (84th %ile); $-2.0$ (2nd %ile)Psychometric research; foundation for all standard transformations
T-Score5010$T = 50 + 10(z)$Interval$\ge 65$ (Clinical elevation on MMPI, 93rd %ile)MMPI-2-RF/MMPI-3, MCMI, BASC, behavioral checklists
Wechsler IQ10015$IQ = 100 + 15(z)$Interval$\le 70$ (Intellectual disability); $\ge 130$ (Gifted)WAIS-IV, WISC-V, WPPSI-IV, Stanford-Binet 5
Stanine52Integer 1–9Interval-like$5$ (Middle 20%); $1$ (Bottom 4%); $9$ (Top 4%)K-12 standardized achievement tests; military classification
Sten5.52Integer 1–10Interval-like$5–6$ (Average); $1–3$ (Low); $8–10$ (High)Cattell's 16PF Questionnaire
Percentile Rank50 (Median)N/A$% \le \text{Raw Score}$Ordinal$16\text{th} (-1 SD); 50\text{th} (M); 84\text{th} (+1 SD); 98\text{th} (+2 SD)$Communicating results clearly to clients, parents, and schools

Norm-Referenced vs. Criterion-Referenced Interpretation

Counselors must choose the appropriate interpretive framework based on the assessment's intended purpose:

1. Norm-Referenced Tests (NRT)

Norm-referenced interpretation compares an individual examinee's performance to the relative performance of a representative comparison cohort—the normative reference group (e.g., national age-matched peers, gender-specific cohorts).

  • Core Question Answered: "How does this examinee perform relative to others?"
  • Key Metrics: Percentile ranks, z-scores, T-scores, deviation IQ.
  • Clinical Examples: WAIS-IV Full Scale IQ, SAT, ACT, MMPI-3 clinical profiles.

2. Criterion-Referenced Tests (CRT)

Criterion-referenced interpretation (also termed domain-referenced or mastery testing) compares an examinee's performance against an absolute, predetermined standard of mastery, behavioral cutoff, or knowledge domain, completely independent of how other test-takers perform.

  • Core Question Answered: "What specific knowledge or skills has this examinee mastered?" or "Did the examinee achieve the required competency cutoff?"
  • Key Metrics: Percentage correct, pass/fail status, cut-scores established via empirical standard-setting methodologies (such as the Angoff Method, where a panel of experts estimates the probability that a minimally competent candidate will answer each item correctly).
  • Clinical Examples: The National Counselor Examination (NCE; candidates must achieve the absolute passing cut-score regardless of whether 90% or 40% of peers pass), state driver's license exams, mastery learning checklists.

Cultural, Racial, Linguistic, and Disability Test Bias

Psychometric fairness requires that assessment instruments function with equivalent accuracy across diverse demographic groups. Counselors must distinguish between statistical test bias and systemic unfairness.

Psychometric Bias Defined

In psychometrics, bias is not merely a subjective feeling of unfairness; it is a demonstrable statistical property where an assessment yields systematic errors in measurement or prediction for members of an identifiable demographic subgroup:

  • Differential Item Functioning (DIF): An item displays DIF if individuals from different demographic subgroups (e.g., racial, gender, linguistic) who possess the exact same underlying ability level have significantly different probabilities of answering the item correctly.
  • Construct Underrepresentation: The test fails to sample important, culturally salient dimensions of the construct, narrowing the evaluation unfairly.
  • Construct-Irrelevant Difficulty: The assessment introduces extraneous, irrelevant task requirements that impede certain subgroups from demonstrating their true ability (e.g., using complex English syntax and idiomatic expressions on a nonverbal spatial reasoning test).

Landmark Legal Cases Governing Testing Fairness

Counselors must know the legal precedents that shaped modern assessment ethics:

1. Larry P. v. Riles (9th Cir. 1979, upheld 1984)

  • The Case: A landmark class-action lawsuit filed on behalf of African American schoolchildren in California who were disproportionately placed in segregated classes for the Educable Mentally Retarded (EMR) based on standardized individual intelligence tests.
  • The Ruling: Federal District Judge Robert Peckham ruled that standard IQ tests (including the WISC and Stanford-Binet) were racially and culturally biased against Black children, having been normed exclusively on white, middle-class populations. The court permanently enjoined California public schools from utilizing standardized intelligence tests to place African American students into special education EMR programs.

2. Diana v. California State Board of Education (N.D. Cal. 1970)

  • The Case: Spanish-speaking Mexican American children were placed in mentally retarded classrooms based on English-language Stanford-Binet IQ scores.
  • The Ruling: Mandated that children must be tested in their primary native language or evaluated using culture-reduced nonverbal assessment instruments to prevent linguistic misdiagnosis.

3. Griggs v. Duke Power Co. (U.S. Supreme Court 1971)

  • The Case: An employer required a high school diploma and passing scores on two general aptitude tests (Wonderlic Personnel Test) for employment promotion, resulting in a disparate impact that excluded Black applicants.
  • The Ruling: Chief Justice Warren Burger established under Title VII of the Civil Rights Act that employment testing producing disparate impact is unlawful unless the employer proves the test is demonstrably job-related and required by business necessity.

4. Debra P. v. Turlington (5th Cir. 1981)

  • The Case: High school graduation competency exit exams in Florida that disproportionately disqualified Black students who had experienced decades of segregated, unequal schooling.
  • The Ruling: Established the doctrine of curricular validity; a state cannot deny a high school diploma based on a standardized competency test unless the school system proves it provided students with an adequate, equal instructional opportunity to learn the tested material.

Testing Accommodations for Diverse Examinees

Governed by the Americans with Disabilities Act (ADA) and the Individuals with Disabilities Education Act (IDEA), testing accommodations remove construct-irrelevant barriers without altering the fundamental construct being measured:

  • Core Rule: Accommodations level the playing field; they do not lower the standard.
  • Standard Accommodations: Extended testing time (time-and-a-half, double time), separate distraction-free private testing rooms, human readers or screen reading software, scribes or speech-to-text software, large print or Braille materials, and frequent structured breaks.

Ethical Standards in Assessment: ACA Code of Ethics Section E

Section E of the American Counseling Association (ACA) Code of Ethics (2014) governs the evaluation, assessment, and interpretation of tests:

Key Standards Breakdown

Standard E.1: General Client Welfare

  • Assessments must be utilized exclusively to promote client welfare, self-understanding, and evidence-based treatment planning.
  • Counselors respect the client's right to receive transparent, developmentally accessible explanations of assessment results and implications.

Standard E.2: Competence to Use and Interpret Assessments

  • Counselors administer only those assessments for which they have received specialized graduate didactic coursework, supervised practicum training, and documented clinical competence.
  • Test User Qualification Levels (APA / Publisher Standards):
    • Level A: No advanced graduate degree required. Tests can be administered by following the test manual (e.g., basic vocational interest inventories, simple symptom rating checklists).
    • Level B: Requires a master's degree in counseling, psychology, or related field, including formal graduate coursework in psychometrics, test construction, and statistical interpretation (e.g., BDI-II, Myers-Briggs, 16PF, Strong Interest Inventory).
    • Level C: Requires a doctoral degree in psychology or related field, or state licensure authorizing independent psychological diagnostic testing, along with comprehensive training in individual cognitive and personality administration (e.g., WAIS-IV/WISC-V, Stanford-Binet 5, MMPI-3, MCMI-IV, Rorschach).

Standard E.3: Informed Consent in Assessment

  • Prior to administering any test, counselors must obtain informed consent explaining the assessment's nature, purpose, normative sample, costs, and intended recipients of the resulting report, communicated in understandable language.

Standard E.4: Release of Data to Qualified Professionals

  • Strict Release Restrictions: Counselors release raw test data (e.g., scored protocols, client answer sheets, computerized item responses) ONLY with written client release or a formal judicial court order, and ONLY to professionals qualified to interpret the data.
  • Never Release Raw Test Data to Clients: Releasing raw test protocols or item sheets directly to clients, parents, or unqualified third parties is an ethical violation because raw scores are easily misinterpreted and release compromises proprietary test security.

Standard E.5: Diagnosis of Mental Disorders & Cultural Sensitivity

  • Counselors must carefully evaluate the socioeconomic, cultural, and linguistic background of the client when formulating diagnoses.
  • Counselors may refrain from making or reporting a formal psychiatric diagnosis if they judge that doing so would cause tangible harm to the client or if the diagnostic system lacks cultural validity for that individual.

Standard E.8: Multicultural Issues in Assessment

  • Counselors select instruments that have been normed on populations representative of the client's cultural, racial, and linguistic identity, explicitly acknowledging normative limitations in written evaluation reports.

Standard E.10: Assessment Security

  • Counselors maintain the integrity and security of copyrighted test manuals, protocols, scoring algorithms, and stimulus items, refraining from reproducing or disclosing test items in public forums or coaching clients on test questions.

On the NCE Exam: Key Tips and Traps

  • Skewness Order Trap: Memorize the direction of the mean! On a positively skewed distribution, the tail is to the right, so Mode < Median < Mean. On a negatively skewed distribution, the tail is to the left, so Mean < Median < Mode.
  • Release of Test Data: If an exam vignette asks whether a counselor should hand over raw MMPI or WISC protocols to an adult client or parent who demands to see their answers, the answer is NO. Raw test data is released only to qualified professionals with client consent or under judicial court order (ACA Standard E.4).
  • Percentiles Cannot Be Averaged: If a question describes a school counselor attempting to average a student's percentile ranks across reading, math, and science, remember that this is a psychometric error because percentiles represent an ordinal scale with unequal intervals.
  • Larry P. v. Riles: If asked which landmark legal case banned the use of standardized IQ tests for placing African American children in EMR classes in California due to cultural bias, immediately select Larry P. v. Riles.
Test Your Knowledge

A vocational counselor administers a newly developed occupational aptitude test to a cohort of 500 job applicants. The resulting distribution of raw scores reveals that the arithmetic Mean is 68, the Median is 74, and the Mode is 82. Based on this statistical configuration, what is the shape of the score distribution?

A
B
C
D
Test Your Knowledge

An adult client who recently completed a court-ordered psychoeducational evaluation for a disability claim requests that the counselor provide them with a full copy of their raw test protocols, marked item response sheets, and WAIS-IV subtest scoring sheets. Under Standard E.4 of the ACA Code of Ethics, what is the counselor's ethical obligation?

A
B
C
D
Test Your Knowledge

Which landmark federal court decision permanently enjoined California public schools from utilizing standardized individual intelligence tests to place African American children into classes for the Educable Mentally Retarded (EMR), ruling that the tests were culturally and racially biased?

A
B
C
D