7.3 Psychometric Properties, Test Interpretation & Communicating Data to Stakeholders
Key Takeaways
- Reliability denotes the consistency and stability of test results across administrations, forms, and raters; validity denotes the degree to which an assessment measures what it purports to measure. A test can be reliable without being valid, but can never be valid without being reliable.
- Norm-referenced tests compare individual performance to a normative peer distribution (reporting percentile ranks and stanines), whereas criterion-referenced tests measure individual performance against an absolute mastery benchmark of specific standards.
- Grade-equivalent (GE) scores are psychometrically flawed metrics that reflect the grade level at which an average student earned the same raw score; they do not indicate that a student can successfully read or master higher grade-level curriculum.
- Ethical, transparent communication of reading data requires translating complex psychometric terminology into accessible, asset-based language, presenting visual MTSS/IEP aimline data, and providing concrete, actionable home literacy routines.
7.3 Psychometric Properties, Test Interpretation & Communicating Data to Stakeholders
GACE Blueprint Focus: Objective 0005 requires candidates to demonstrate sound psychometric literacy, interpret diverse standardized and classroom score formats, evaluate test quality (reliability and validity), and ethically communicate assessment findings to parents, general educators, special educators, and multidisciplinary MTSS/IEP teams using asset-based, accessible language.
Psychometric Foundations: Reliability and Validity
Educational assessment data is only as sound as the psychometric architecture of the instruments used to collect it. When literacy educators select assessment tools or interpret evaluation reports, they must evaluate two foundational psychometric properties: reliability (consistency of measurement) and validity (accuracy of inference).
+-------------------------------------------------------------------------+
| RELIABILITY VERSUS VALIDITY |
+-------------------------------------------------------------------------+
| |
| RELIABILITY = CONSISTENCY VALIDITY = TRUTHFULNESS |
| - Can we replicate the score? - Are we measuring what we think? |
| - Is measurement error minimized? - Are inferences justifiable? |
| |
| KEY AXIOM: A test can be highly RELIABLE without being VALID, |
| but an assessment can NEVER be VALID without being RELIABLE.|
+-------------------------------------------------------------------------+
The Four Forms of Reliability
Reliability refers to the degree to which an assessment instrument yields dependable, stable, and repeatable results across time, testing conditions, evaluators, and test forms. A reliable test minimizes measurement error ($e$ in classical test theory: $\text{Observed Score} = \text{True Score} + \text{Error}$).
- Test-Retest Reliability (Temporal Stability): Evaluates whether an identical test administered to the same cohort of students at two different points in time produces consistent scores. A high test-retest correlation coefficient (e.g., $r \ge 0.85$) indicates that scores are not subject to random temporal fluctuations.
- Alternate-Form (Parallel / Equivalent-Form) Reliability: Evaluates the consistency of scores across two distinct, equated versions of the same assessment designed to measure the identical construct. This is the bedrock psychometric requirement for Curriculum-Based Measurement (CBM) progress monitoring probes (e.g., DIBELS ORF passages). If Form A is significantly easier than Form B, fluctuations in weekly scores reflect form discrepancy rather than genuine student growth.
- Inter-Rater (Inter-Scorer / Inter-Observer) Reliability: Evaluates the degree of agreement or concordance between two or more independent examiners scoring the same student performance. High inter-rater reliability is critical for subjective reading measures, such as qualitative running records, open-ended oral retellings, and holistic writing rubrics.
- Internal Consistency Reliability: Evaluates the extent to which all individual items within a single test or subtest measure the same underlying cognitive construct. It is typically calculated via Cronbach's alpha ($\alpha$) or Split-Half correlation. For high-stakes screening or special education eligibility batteries (e.g., CTOPP-2, Woodcock-Johnson), internal consistency should exceed $\alpha = 0.90$.
The Three Core Dimensions of Validity
Validity is the most fundamental psychometric consideration. While reliability focuses on the test instrument itself, validity focuses on the soundness, truthfulness, and appropriateness of the inferences and actions based on test scores (Messick, 1989). A test does not possess general "validity"; rather, it possesses validity for a specific, intended purpose.
- Content Validity: The extent to which an assessment comprehensively and proportionally samples the specific curricular domain, learning standards, or cognitive processes it claims to assess. For example, a third-grade ELA assessment aligned with the Georgia Standards of Excellence (GSE) exhibits content validity only if its items represent the full breadth of grade-level informational, literary, foundational, and language standards, rather than over-sampling basic recall.
- Construct Validity: The extent to which test scores reflect the theoretical psychological or cognitive construct the test purports to measure. For example, if a test designed to measure "reading comprehension" is heavily loaded with archaic cultural references or complex geometric vocabulary, it may inadvertently measure background cultural schema or math vocabulary rather than reading comprehension, undermining construct validity.
- Criterion-Related Validity: The degree to which test scores correlate with an external, established criterion measure of the same ability. Criterion validity is bifurcated into:
- Concurrent Validity: How strongly scores on the new assessment correlate with an established "gold standard" assessment administered at the same time.
- Predictive Validity: The statistical power of an early assessment to accurately forecast future performance on a distal outcome measure. In early literacy, predictive validity is vital: kindergarten DIBELS phonemic segmentation fluency must demonstrate strong predictive validity for third-grade Georgia Milestones reading achievement.
The Relationship Between Reliability and Validity
A cornerstone principle of psychometrics tested on the GACE is the directional relationship between reliability and validity:
- Reliability is a necessary, but not sufficient, condition for validity.
- An assessment can be perfectly reliable without being valid. For example, measuring a student's head circumference with a laser scanner every morning for a month will yield almost 100% reliability (identical measurements every day). However, using head circumference to infer reading comprehension capacity has zero validity.
- Conversely, an assessment can NEVER be valid without being reliable. If a reading test produces erratic, wildly fluctuating scores from day to day or rater to rater (unreliable), the scores are clouded by random measurement error and cannot legitimately measure the target construct (invalid).
Score Interpretation Formats: Norm-Referenced vs. Criterion-Referenced Tests
Educational assessment scores must be contextualized within an established measurement framework to carry instructional meaning. Standardized reading assessments fall into two primary score interpretation paradigms:
+-------------------------------------------------------------------------+
| NORM-REFERENCED VERSUS CRITERION-REFERENCED |
+-------------------------------------------------------------------------+
| |
| NORM-REFERENCED TESTS (NRT) CRITERION-REFERENCED TESTS (CRT) |
| - Compares student to PEERS. - Compares student to STANDARDS. |
| - Normal Bell Curve distribution. - Absolute Mastery Benchmark. |
| - Percentile Ranks, Stanines. - Percent Correct, Standards Level. |
| - "Where does the student stand - "What specific knowledge and |
| relative to the national group?" skills has the student mastered?" |
+-------------------------------------------------------------------------+
Comparative Matrix: NRT vs. CRT
| Assessment Dimension | Norm-Referenced Tests (NRT) | Criterion-Referenced Tests (CRT) |
|---|---|---|
| Core Philosophy | Relative standing; compares an individual's performance to a representative national or state normative sample. | Absolute mastery; compares an individual's performance to a predetermined standard, learning objective, or cutoff criterion. |
| Score Reporting Formats | Percentile Ranks (1–99), Stanines (1–9), Standard Scores ($M=100, SD=15$), Normal Curve Equivalents (NCE). | Performance levels (e.g., Beginning, Developing, Proficient, Distinguished), raw percentage correct, pass/fail thresholds. |
| Distribution of Scores | Designed to yield a normal bell-shaped curve, intentionally maximizing score variance and differentiation between test-takers. | No required normal distribution; theoretically, 100% of students could achieve mastery if provided adequate instruction. |
| Exemplary Instruments | Woodcock-Johnson IV (WJ-IV), CTOPP-2, Iowa Assessments, Stanford Achievement Test (SAT-10). | Georgia Milestones Assessment System (GMAS EOG/EOC), teacher unit tests, criterion-referenced phonics surveys. |
| Primary Educational Use | Universal screening, gifted identification, special education eligibility determination, identifying relative deficits. | Standards-based grading, determining curriculum mastery, certifying graduation/promotion, school accountability. |
Statistical Metrics: Percentiles, Stanines, Scaled Scores & The Grade-Equivalent Fallacy
To interpret evaluation reports and explain test scores to parents and multidisciplinary teams, educators must understand the underlying statistical properties of common score reporting metrics:
+-------------------------------------------------------------------------+
| THE NORMAL DISTRIBUTION BELL CURVE |
+-------------------------------------------------------------------------+
| |
| PERCENTILE: 1st 16th 50th 84th 99th |
| STANINE: 1 2 3 4 5 6 7 8 9 |
| STD SCORE: 70 85 100 115 130 |
| |
| Below Average Above |
| Average (54% of pop) Average |
| (23% of pop) (23% of pop) |
+-------------------------------------------------------------------------+
1. Percentile Ranks (PR)
- Definition: A percentile rank indicates the percentage of students in the national normative comparison group who scored at or below a given student's raw score. Percentile ranks range on an ordinal scale from the 1st to the 99th percentile, with the 50th percentile representing the national median/average.
- Crucial Distinction for GACE: A percentile rank is NOT the percentage of questions answered correctly. If a student scores at the 75th percentile on an oral reading fluency test, it means the student read more words correctly than 75% of same-grade peers in the national norm group; it does not mean the student achieved 75% accuracy.
- Non-Linear Ordinal Metric: Percentile ranks are clustered around the median of the normal bell curve. A 5-point percentile difference near the middle of the distribution (e.g., between the 48th and 53rd percentile) represents a tiny difference in raw score points, whereas a 5-point difference at the extreme tails (e.g., between the 94th and 99th percentile) represents a massive difference in raw score points. Consequently, percentile ranks cannot be added, subtracted, or averaged.
2. Stanines (Standard Nine Scale)
- Definition: A standardized nine-point scale that divides the normal distribution into nine bands, with a mean of 5 and a standard deviation of 2.
- Bands:
- Stanines 1, 2, 3: Below Average performance (representing the lowest 23% of the distribution).
- Stanines 4, 5, 6: Average performance (representing the middle 54% of the distribution; Stanine 5 is exactly average).
- Stanines 7, 8, 9: Above Average performance (representing the top 23% of the distribution).
- Educational Utility: Stanines condense complex psychometric distributions into easily understood single-digit performance bands, preventing over-interpretation of minor score fluctuations.
3. Scaled Scores and Standard Scores
- Standard Scores (SS): Linear transformations of raw scores that preserve equal intervals (typically with a mean of 100 and a standard deviation of 15). Standard scores between 85 and 115 fall within the average range. Standard scores allow mathematically valid comparisons across distinct cognitive subtests and longitudinal tracking.
- Scaled Scores: Continuous developmental scales that allow tracking of academic growth. On the Georgia Milestones Assessment System, every grade and subject uses the same two achievement cut points — 475 separates Beginning from Developing Learner and 525 separates Developing from Proficient Learner — but the lowest and highest obtainable scale scores differ by grade and by subject (grade 3 ELA runs roughly 180–830, grade 8 ELA roughly 225–730). Never read a GMAS scale score as a percentage: a 500 does not mean 50 percent correct.
- Standard Error of Measurement (SEM) & Confidence Intervals: No test is a 100% error-free measurement. Psychometricians calculate the Standard Error of Measurement (SEM) to quantify the margin of error. Educators report scores within a Confidence Interval (e.g., "We are 95% confident that Ethan's true reading standard score falls between 88 and 96"), acknowledging measurement error.
4. The Grade-Equivalent (GE) & Age-Equivalent (AE) Fallacy
One of the most persistent and dangerous misconceptions in educational testing involves Grade-Equivalent (GE) scores (e.g., 4.2 representing 4th grade, 2nd month) and Age-Equivalent (AE) scores (e.g., 9-4 representing 9 years, 4 months). The GACE reading examination consistently evaluates an educator's understanding of why GE scores are psychometrically flawed and clinically misleading.
[!CAUTION] The Grade-Equivalent Misconception: If a second-grade student takes a second-grade reading assessment in October and receives a Grade-Equivalent score of 5.2, what does this mean?
- INCORRECT Interpretation: The student is reading at a fifth-grade level and should be accelerated to fifth-grade reading materials and textbooks.
- CORRECT Psychometric Interpretation: The second-grader obtained the same raw score on a second-grade test that an average fifth-grader would obtain if given that same second-grade test. It simply means the second-grader has exceptionally strong mastery of second-grade reading material. It does NOT indicate that the student has the vocabulary, background knowledge, figurative language comprehension, or emotional maturity to navigate fifth-grade expository and literary curriculum.
Psychometric Flaws of Grade Equivalents:
- Lack of Equal-Interval Scaling: Growth between GE 2.0 and 3.0 represents a massive developmental leap in foundational phonics, whereas growth between GE 10.0 and 11.0 represents minor shifts in reading rate or vocabulary.
- Extrapolation and Interpolation: Most GE scores at the upper and lower extremes are never actually tested; they are mathematically extrapolated by test publishers based on statistical projections.
- Professional Standard: Both the International Literacy Association (ILA) and the American Educational Research Association (AERA) explicitly discourage the use of Grade-Equivalent scores for instructional placement or special education qualification.
Ethical, Asset-Based Communication of Assessment Data to Stakeholders
Assessment data is useless if it remains sequestered in cumulative folders or communicated in dense, impenetrable psychometric jargon that alienates families and educators. Literacy specialists must serve as skilled translators, communicating assessment findings with transparency, ethical integrity, and empathy.
Core Principles for Communicating with Parents and Caregivers
- Demystify Jargon into Everyday Language: Avoid opaque psychometric terms ("heteroscedasticity," "construct under-representation," "standard deviation"). Instead of stating, "Your daughter scored at the 16th percentile on the DIBELS ORF probe, placing her at -1.5 standard deviations below the mean," state: "When reading grade-level stories aloud, Maria reads 35 words per minute accurately. An average second-grader in the fall reads about 55 words per minute. Her accuracy is strong, but she is spending so much mental energy sounding out words that it slows her down. Our goal is to build her automatic word recognition so reading feels effortless."
- Adopt an Asset-Based (Strengths-Based) Orientation: Reject deficit-laden labeling that reduces a child to a test score ("low reader," "remedial," "Tier 3 kid"). Always begin data conferences by celebrating the student's demonstrated competencies, cognitive strengths, and effort before detailing target intervention areas.
- Provide Visual, Accessible Data Displays: Share visual CBM progress monitoring charts displaying the student's baseline, aimline, and weekly trendline. Visual graphs provide families with an intuitive, transparent picture of whether the intervention is accelerating growth toward the benchmark.
- Deliver Concrete, Actionable Home Literacy Guidance: Parents frequently ask, "What can I do to help at home?" Never offer generic, unhelpful advice like "Just read more books." Provide specific, structured, low-stress routines:
- Paired / Echo Reading: The parent and child read a sentence or paragraph together aloud to model expressive prosody and phrasing.
- Oral Phonemic Word Games: Engaging in playful 2-minute sound substitution games while driving or cooking ("What word do we get if we take the /s/ out of 'spin'? Pin!").
- Dialogic Reading: Asking open-ended inferential questions during bedtime shared reading ("Why do you think the character made that choice? What would you have done?").
Collaborative MTSS and IEP Team Communication
When presenting data at Multi-Tiered System of Supports (MTSS) problem-solving meetings or Individualized Education Program (IEP) committee reviews, literacy educators must:
- Triangulate Multiple Data Sources: Never base high-stakes instructional or placement decisions on a single assessment. Synthesize universal screening data, fine-grained diagnostic inventories, running records, and classroom performance samples.
- Report Rate of Improvement (ROI): Present empirical growth metrics calculating weekly words correct per minute gained (e.g., "Ethan is growing at 1.4 WCPM per week, exceeding the typical peer growth rate of 0.8 WCPM per week, demonstrating positive responsiveness to Tier 2 small-group intervention").
- Foster Interdisciplinary Alignment: Coordinate data interpretation across general education classroom teachers, special education case managers, English learner (EL) specialists, and Speech-Language Pathologists (SLPs) to ensure cohesive, unified instructional targets.
Realistic IEP & Parent Conference Scenario: Collaborative Data Translation
The Conference Setting
Dr. Bennett, the reading specialist at Sweetwater Elementary School in Gwinnett County, Georgia, convenes an MTSS Tier 3 data review conference for Ethan, an eight-year-old third-grader. In attendance are Ethan's general education teacher (Ms. Morales), the special education lead teacher (Mr. Vance), and Ethan's parents (Mr. and Mrs. Hayes).
The Data Portfolio
Dr. Bennett compiles Ethan's comprehensive reading portfolio:
- Fall Universal Screening (DIBELS 8th): Composite Score at the 12th percentile (Well Below Benchmark); ORF = 38 WCPM (Benchmark = 70 WCPM).
- Diagnostic Phonics Battery (Core Phonics Survey): 100% mastery on short vowels, consonant blends, and silent e words; 40% mastery on complex multisyllabic words with inflectional and derivational affixes (unbreakable, nonperishable).
- Diagnostic Phonological Processing (CTOPP-2): Phonological Awareness Standard Score = 94 (35th percentile, Average); Rapid Automatized Naming (RAN) Standard Score = 72 (3rd percentile, Well Below Average).
- Classroom Informal Reading Inventory (IRI): Oral reading at instructional 2nd-grade level; Listening Comprehension Capacity at 5th-grade level (90% comprehension on passages read aloud to him).
- Commercial Achievement Test (administered by previous private tutor): Reports a Grade-Equivalent score of 1.8 in reading fluency.
Addressing Parental Confusion and Anxiety
Mrs. Hayes opens the meeting visibly distraught: "The tutor told us Ethan has the reading brain of a first-grader because his grade-equivalent was 1.8. Does this mean he should be retained back to first grade? How can he ever pass third-grade Georgia Milestones?"
Dr. Bennett executes an asset-based, psychometrically sound data translation:
- Dismantling the Grade-Equivalent Fallacy: Dr. Bennett gently reassures the parents: "Mrs. Hayes, that 1.8 grade-equivalent score is one of the most widely misunderstood numbers in education. It does not mean Ethan has the brain of a first-grader, nor does it mean he belongs in a first-grade classroom. It simply means that on that specific fluency test, Ethan read basic words at the same speed that an average first-grader would in the eighth month of school. Ethan's listening comprehension assessment shows that when stories are read to him, his comprehension reaches the fifth-grade level. He has the intellectual curiosity, vocabulary, and advanced reasoning of an older child; what he has is a specific mechanical bottleneck in naming speed and multisyllabic decoding."
- Explaining the Double-Deficit Dynamic: Dr. Bennett explains Ethan's CTOPP-2 scores: his phonemic awareness is intact (Standard Score 94), but his Rapid Automatized Naming (RAN) is depressed (Standard Score 72). She uses the analogy of a high-speed computer processor with a narrow internal cable: Ethan knows the sounds, but retrieving the visual symbol-to-sound connection takes extra milliseconds per syllable.
- Presenting Graphed CBM Progress Monitoring: Dr. Bennett displays Ethan's weekly ORF chart from Tier 2 intervention. Over the past 8 weeks, Ethan's trendline slope shows a growth rate of 0.5 WCPM/week, which falls below his aimline of 1.2 WCPM/week (verifying non-responsiveness to general Tier 2 instruction under the four-point rule).
- Collaborative Action Plan: The team agrees to intensify Ethan's intervention to Tier 3 (Intensive Multi-Sensory Structured Literacy) in a 1:2 student-to-teacher setting 45 minutes daily, focusing on explicit syllable division strategies (morpheme-based decoding for prefixes and suffixes) and repeated paired reading to build retrieval automaticity. Dr. Bennett provides Mr. and Mrs. Hayes with a structured home paired-reading routine and audiobooks for content-area social studies and science texts, ensuring Ethan continues to expand his advanced conceptual schema while his mechanical decoding is remediated.
Common GACE Exam Traps & Misconceptions
[!WARNING] Trap 1: Confusing Percentile Rank with Percentage Correct. Do not confuse these two metrics on the GACE. If a question states that a student scored at the 88th percentile on an oral reading fluency test, distractors will suggest the student "answered 88% of questions correctly" or "read 88% of the words correctly." The correct interpretation is that the student performed better than 88% of students in the national normative comparison group.
[!WARNING] Trap 2: Using Grade-Equivalent Scores for Grade Placement or Acceleration. Whenever a GACE scenario depicts an educator making an instructional decision based on a Grade-Equivalent score—such as placing a 2nd grader with a 4.5 GE into a 4th-grade reading group, or retaining a 3rd grader with a 1.9 GE—it is always incorrect. Professional standards dictate that GE scores must never be used for curriculum placement, tracking, or grade retention decisions.
[!WARNING] Trap 3: Assuming High Reliability Guarantees Validity. A favorite GACE conceptual trap tests whether candidates understand that reliability does not ensure validity. A test can yield identical, rock-solid results day after day (perfect reliability) while failing completely to measure the target educational construct (zero validity). Reliability is necessary for validity, but reliability alone never guarantees validity.
[!WARNING] Trap 4: Communicating Deficit-Laden Labels or Raw Statistics to Families. In constructed-response scenarios and multiple-choice questions evaluating family communication, avoid options that use cold educational jargon ("Your child is at -2.0 standard deviations and belongs in Tier 3 remediation"). Correct responses always feature asset-based language, clear translation of data into observable reading behaviors, visual progress displays, and concrete, manageable home literacy partnerships.
A school psychologist and reading specialist are reviewing the psychometric properties of a newly published diagnostic reading comprehension assessment. The test manual reports a test-retest reliability coefficient of r = 0.92 and an internal consistency alpha of 0.94. However, independent construct validity studies demonstrate that student scores on the assessment correlate at r = 0.88 with general IQ non-verbal puzzle solving, but only at r = 0.31 with established reading comprehension measures. Which of the following conclusions is psychometrically accurate?
During a parent-teacher conference, a third-grade teacher shares the results of a norm-referenced standardized reading test administered in the fall. The student earned a Grade-Equivalent (GE) score of 6.2 in reading comprehension. The parent enthusiastically asks if the student should be moved immediately into sixth-grade reading curriculum and textbooks. How should the teacher professionally and accurately respond?
A second-grade student completes a standardized, norm-referenced reading assessment at the end of the school year. The student's score report indicates a Percentile Rank (PR) of 68. Which of the following statements provides the most accurate and psychometrically sound interpretation of this score?
A reading specialist is preparing to communicate assessment data to the parents of a first-grade student who has exhibited severe decoding difficulties and has been flagged for Tier 3 intensive intervention. Which of the following communication approaches represents the most ethical, asset-based, and effective practice?