2.1 Measurement Theory, Reliability, & Validity

Key Takeaways

  • Classical Test Theory establishes that every observed score consists of a hypothetical true score plus random measurement error (X = T + E), meaning observed scores are estimates rather than absolute values.

  • The Standard Error of Measurement (SEM = SD * √(1 - r_xx)) quantifies score precision, demonstrating an inverse mathematical relationship with test reliability.

  • Confidence intervals (typically 68%, 90%, or 95%) must accompany reported scores to communicate measurement uncertainty and prevent over-interpretation of isolated test results during eligibility meetings.

  • Internal consistency metrics differ by item type: Cronbach's alpha assesses multi-point, continuous, or rating scale items, whereas Kuder-Richardson 20 (KR-20) is required for dichotomously scored items.

  • Construct validity functions as the unifying evidentiary umbrella for score interpretation, supported by convergent and discriminant validity and threatened by construct underrepresentation and construct-irrelevant variance.

Last updated: September 2026

2.1 Measurement Theory, Reliability, & Validity

Core Principle: Psychoeducational assessment scores never represent immutable, error-free reflections of a student's inner cognitive or academic capability. Every assessment score is an observed estimate governed by Classical Test Theory, subject to random and systematic measurement errors. Competent school psychologists must evaluate the reliability and validity of assessment tools to interpret test performance defensibly and communicate confidence bands to multidisciplinary evaluation teams.


1. Classical Test Theory (CTT) & Measurement Error

Classical Test Theory (CTT), often termed true-score theory, serves as the mathematical foundation for the vast majority of standardized cognitive, achievement, and behavioral rating batteries used in school psychology.

The Mathematical Formulation

CTT asserts that any score an examiner obtains during an evaluation—the Observed Score (X)—is composed of two additive, independent components:

X = T + E

Where:

  • X (Observed Score): The actual number of raw points or derived standardized score earned by the examinee during a specific testing administration.
  • T (True Score): The theoretical value that would be obtained if the examinee were assessed an infinite number of times with identical or perfectly parallel instruments, assuming no memory, practice, fatigue, or developmental changes occurred. It represents the student's true, unadulterated ability level on that construct.
  • E (Error Score): The cumulative discrepancy between the observed score and the true score resulting from unsystematic, random measurement fluctuations.

Core Assumptions of Classical Test Theory

For CTT formulas to function mathematically, test developers and psychometricians rely on several explicit postulates:

  1. Expected Value of Error is Zero: Across an infinite number of repeated administrations, random errors cancel each other out: Expected Value of Error E(E) = 0. True score is mathematically defined as the expected value (mean) of observed scores: T = E(X).
  2. Zero Correlation Between True Score and Error: The magnitude of measurement error is completely independent of the examinee's true ability level (r_TE = 0). A student with an exceptionally high true cognitive ability is just as likely to experience positive or negative random error as a student with a lower ability level.
  3. Zero Correlation Between Errors Across Administrations: Measurement errors occurring during one test administration have zero correlation with errors occurring during a second administration or on an alternate test (r_E1,E2 = 0).

Systematic Error vs. Random Error

Understanding the distinction between random error and systematic error is essential for legal and psychometric defensibility:

FeatureRandom ErrorSystematic Error
DefinitionUnpredictable, chance fluctuations affecting an individual test administration.Consistent, predictable bias that shifts scores systematically in a single direction.
ExamplesTransient examinee fatigue, minor illness, ambient hallway noise, examiner stopwatch slip, momentary lapse in attention.Testing an English Learner using a test with high English linguistic demand, uncalibrated audiometer, cultural test bias, strict time limits on a conceptual math test.
Psychometric ImpactDirectly deflates test reliability; creates wider confidence intervals around observed scores.Directly invalidates test validity; does not necessarily reduce reliability (a biased test can measure the wrong construct consistently).
RemediationUse tests with high internal consistency; calculate Standard Error of Measurement (SEM) and confidence intervals.Select non-discriminatory, culturally and linguistically appropriate instruments; adhere to standardized administration procedures.

2. Types of Reliability Coefficients

In psychometrics, reliability refers to the consistency, dependability, and reproducibility of test scores across time, test forms, item samples, and raters. Mathematically, the reliability coefficient (r_xx) represents the proportion of observed score variance (σ²_X) that is attributable to true score variance (σ²_T):

r_xx = σ²_T / σ²_X = 1 - (σ²_E / σ²_X)

Reliability coefficients range from 0.00 (all observed variance is random error) to 1.00 (perfect measurement with zero error). In psychoeducational evaluation, different types of reliability isolate distinct sources of measurement error.

Stability Across Time: Test-Retest Reliability

  • Mechanics: The identical test is administered to the same group of students on two separate occasions, and the resulting score pairs are correlated to calculate the coefficient of stability.
  • Error Source Captured: Time-sampling error (fluctuations in examinee state, alertness, mood, or environment over time).
  • Interval Considerations:
    • Too Short (e.g., 1–3 days): Results in artificially inflated reliability due to memory, carryover, and practice effects.
    • Too Long (e.g., 6–12 months): Conflates measurement instability with authentic developmental maturation, instruction, or intervention effects.
    • Optimal Window: A 2- to 4-week retest window is standard in commercial cognitive and academic battery standardization manuals.

Equivalence Across Forms: Alternate / Parallel Form Reliability

  • Mechanics: Two distinct forms of a test (Form A and Form B) designed to parallel specifications (identical content domains, item counts, difficulty parameters, means, and variances) are administered to the same sample.
  • Error Source Captured: Content-sampling error (specificity of particular items) when administered simultaneously; both content-sampling and time-sampling error when administered with a time delay.
  • Clinical Application: Vital for progress monitoring tools (e.g., Curriculum-Based Measurement probes in reading and math) where students must be retested frequently without memorizing test items.

Internal Consistency Reliability

Internal consistency evaluates how well the individual items within a single test administration correlate with each other, capturing content-sampling error and item heterogeneity.

1. Split-Half Reliability & The Spearman-Brown Formula

  • An examiner splits a test into two equivalent halves (commonly by assigning odd-numbered items to Half 1 and even-numbered items to Half 2 to counterbalance fatigue and progressive item difficulty).
  • Correlating Half 1 with Half 2 (r_hh) yields the reliability of a test half as long. Because test length directly affects reliability, psychometricians apply the Spearman-Brown Prophecy Formula to estimate full-test reliability:
r_sb = (2 * r_hh) / (1 + r_hh)
  • General Spearman-Brown Formula: For altering test length by a factor of k (where k = new length / old length):
r_kk = (k * r_xx) / [1 + (k - 1) * r_xx]

Exam Key Takeaway: Adding homogeneous items of equal quality systematically increases test reliability. Shortening a test to save administration time systematically decreases its reliability.

2. Cronbach's Alpha (Coefficient α)

  • Developed by Lee Cronbach, Coefficient Alpha represents the mathematical mean of all possible split-half combinations for an instrument.
  • Item Suitability: Required for tests containing multi-point, continuous, or Likert-scale items (e.g., behavioral rating scales like the BASC-3, Conners 4, BRIEF-2, or cognitive subtests scored on 0, 1, or 2-point rubrics).
α = [k / (k - 1)] * [1 - (Σ σ²_i / σ²_X)]

Where k is the number of items, σ²_i is the variance of each individual item, and σ²_X is the total test variance.

3. Kuder-Richardson Formula 20 (KR-20)

  • A specialized algebraic simplification of Cronbach's alpha restricted strictly to dichotomously scored items (scored exclusively as correct [1] or incorrect [0], such as multiple-choice tests or standard cognitive item passes/fails):
KR-20 = [k / (k - 1)] * [1 - (Σ p_i * q_i / σ²_X)]

Where p_i is the proportion of examinees passing item i, and q_i = 1 - p_i is the proportion failing.

Inter-Rater / Inter-Scorer Reliability

  • Error Source Captured: Scorer subjectivity, clerical grading variations, and inconsistent application of scoring rubrics.
  • Metrics:
    • Cohen's Kappa (κ): Used for categorical, nominal, or diagnostic classification data between two raters. Critically, Kappa statistically corrects for agreement that would occur by chance alone:
κ = (P_o - P_e) / (1 - P_e)

Where P_o is the observed proportion of agreement, and P_e is the expected chance agreement. Widely used when validating direct classroom behavioral observation protocols (e.g., momentary time sampling of off-task behavior).

  • Intraclass Correlation Coefficient (ICC): Used when multiple raters assign continuous scores or ratings (e.g., evaluating inter-rater agreement on the Vineland-3 or BASC-3 teacher forms across three classroom teachers).

Professional Standards for Reliability Thresholds

In school psychology practice, acceptable reliability thresholds vary depending on how scores are utilized:

Purpose / ContextMinimum Acceptable Reliability (r_xx)Psychometric Rationale
High-Stakes Individual Decisions (Special Education Eligibility, Placement, Intellectual Disability diagnosis)r_xx ≥ 0.90 (Preferably ≥ 0.95)Minimizes random error; ensures confidence intervals are sufficiently narrow to prevent false-positive or false-negative eligibility decisions. Global composites (FSIQ, Total Achievement) typically meet this standard.
Universal Screening / Group Decision-Makingr_xx ≥ 0.80Acceptable for broad initial triage where false positives will be screened again in Tier 2 progress monitoring.
Individual Subtest Level Analysisr_xx ≥ 0.80 (Often 0.80 - 0.89)Subtests have fewer items and narrower sampling; isolated subtests must never be used as sole criteria for diagnostic decisions due to higher SEM.
Unacceptable for Clinical Practicer_xx < 0.70High error variance makes scores uninterpretable and legally indefensible.

3. Standard Error of Measurement (SEM) & Confidence Intervals

The Standard Error of Measurement (SEM) is the practical operationalization of Classical Test Theory. It quantifies the standard deviation of observed scores an individual would obtain if tested an infinite number of times under identical conditions.

Mathematical Formula

SEM = SD * √(1 - r_xx)

Where:

  • SD = Standard deviation of the standardized assessment tool (e.g., 15 for standard scores, 3 for scaled scores, 10 for T-scores).
  • r_xx = Published reliability coefficient of the scale.

Core Psychometric Law: There is an inverse mathematical relationship between test reliability and the Standard Error of Measurement. As test reliability increases toward 1.00, the SEM approaches 0. Conversely, as reliability drops, the SEM expands, increasing measurement uncertainty.

Calculating and Interpreting Confidence Intervals

Because an observed score is merely an estimate of a student's true capability, school psychologists report confidence intervals (typically 68%, 90%, or 95%) around the score.

Confidence intervals are calculated using the standard normal distribution critical values (z):

Confidence Interval = Observed Score ± (z_crit * SEM)
Confidence Levelz_crit ValueCommon Clinical ApproximationInterpretation
68%1.00± 1.0 * SEMApproximately 68 out of 100 theoretical retests will yield an observed score within this interval.
90%1.65± 1.65 * SEMCommon alternative level offered by cognitive-test scoring software alongside 95%.
95%1.96± 2.0 * SEMRecommended standard for high-stakes special education eligibility and legal defensibility.

Clinical Case Vignette: Evaluating Confidence Bands in Intellectual Disability Eligibility

Case Scenario: Marcus, a 9-year-old student, is evaluated for special education eligibility under the Intellectual Disability (ID) category. On a standardized cognitive battery (Mean = 100, SD = 15, r_xx = 0.96), Marcus earns an observed Full Scale IQ (FSIQ) of 69.

Psychometric Calculation:

  1. SEM = 15 * √(1 - 0.96) = 15 * √(0.04) = 15 * 0.20 = 3.0
  2. 95% Confidence Interval = 69 ± (1.96 * 3.0) = 69 ± 5.88 = [63.12, 74.88] → [63 to 75]

Clinical Interpretation at the IEP Table: Under state criteria, an intellectual disability requires cognitive performance falling at or below approximately 2 standard deviations below the mean (SS ≤ 70), accompanied by concurrent deficits in adaptive functioning across environments. The school psychologist must explain to the team that Marcus's observed score of 69 cannot be treated as an absolute pinpoint value. His 95% confidence interval spans from 63 to 75. Because this confidence interval crosses the clinical cutoff of 70, the team must not base eligibility solely on this single score. The evaluation team must carefully weigh adaptive behavior assessment data (e.g., Vineland-3 or ABAS-3), developmental history, and classroom observational data to determine whether pervasive intellectual and adaptive limitations exist.

True Score Centering vs. Observed Score Centering

Technically, Classical Test Theory demonstrates regression to the mean: an individual who scores far from the population mean (Mean = 100) will have an estimated true score (T') that is slightly closer to the population mean than their observed score (X):

T' = Mean + r_xx * (X - Mean)

If a student scores an observed X = 70 on a test with r_xx = 0.90 (Mean = 100):

T' = 100 + 0.90 * (70 - 100) = 100 - 27 = 73

While advanced computer scoring platforms construct asymmetric confidence intervals centered around estimated true scores (T'), traditional psychoeducational reports frequently center confidence intervals around the observed score for clarity when presenting to multidisciplinary teams.


4. Factors That Influence Reliability Coefficients

School psychologists must critically evaluate technical manuals to understand how sample characteristics and test design artificially inflate or deflate reported reliability:

  1. Test Length: As established by the Spearman-Brown formula, adding homogeneous items increases reliability. Brief or "abbreviated" test forms inevitably sacrifice reliability.
  2. Sample Heterogeneity vs. Range Restriction: Reliability coefficients are correlation-based and depend directly on sample variance. When a test is normed on a wide, heterogeneous general population sample (e.g., all 8-year-olds), reliability appears high. If the test is evaluated on a homogeneous, restricted clinical group (e.g., only students diagnosed with severe reading disabilities), score variance shrinks, drastically deflating the observed reliability coefficient.
  3. Test-Taker Characteristics: Student fatigue, anxiety, lack of motivation, or bilingualism introduce extraneous random variance that lowers reliability.
  4. Administration Consistency: Deviations from standardized instructions, non-standard prompts, or irregular timing introduce unquantified error into observed scores.
  5. The Speeded Test Trap: On a speeded test (where performance reflects rate of completion rather than item difficulty), split-half reliability (odd-even) is spuriously and artificially inflated. Students answer nearly all attempted items correctly before time expires, yielding virtually identical half-scores. For speeded measures, only test-retest or alternate-form reliability provides valid estimates.

5. Modern Validity Theory: From Trinitarian to Unitary Construct Validity

Historically, psychometrics viewed validity through a "trinitarian" model: content validity, criterion validity, and construct validity were treated as separate, independent types of validity.

Under modern standards established by the Standards for Educational and Psychological Testing (AERA, APA, & NCME), validity is conceptualized as a unitary construct. Validity is not a property of the test itself; rather, it refers to the degree to which accumulated empirical evidence and theoretical rationales support the adequacy and appropriateness of interpretations and actions based on test scores.

                          CONSTRUCT VALIDITY
                 (The Overarching Evidentiary Umbrella)
                                  │
         ┌────────────────────────┼────────────────────────┐
         ▼                        ▼                        ▼
  Content-Related        Criterion-Related         Internal Structure
     Evidence                 Evidence                  Evidence
  • Domain sampling        • Concurrent evidence    • Factor analysis (EFA/CFA)
  • Expert panel review    • Predictive evidence    • Multitrait-Multimethod
  • Alignment to standards • Standard error of est. • Invariance testing

1. Content-Oriented Validity Evidence

  • Definition: The extent to which the items, tasks, and questions on a test adequately sample and represent the complete universe of the domain the test is intended to measure.
  • Establishment: Built into the instrument during development via detailed test blueprints, curricular alignment analyses, and expert review panels (e.g., Lawshe's Content Validity Ratio).
  • Key Threats:
    • Construct Underrepresentation: The test is too narrow; it fails to capture essential, critical aspects of the domain (e.g., assessing a child's reading comprehension using only single-word vocabulary definitions, completely omitting paragraph-level inference).
    • Construct-Irrelevant Variance: The test measures extraneous, irrelevant variables that contaminate test performance (e.g., a math problem-solving test featuring complex, paragraph-long linguistic narratives that penalize students with reading disabilities or English Learners, rather than purely measuring mathematical reasoning).

2. Criterion-Related Validity Evidence

Criterion validity evaluates the extent to which test scores correlate with a meaningful external behavioral or diagnostic outcome (the criterion).

  • Concurrent Validity: The new assessment and the established external criterion are administered at the same point in time.
    • Example: Correlating scores on a newly developed 15-minute dyslexia screener with scores on the comprehensive Woodcock-Johnson IV Tests of Achievement administered in the same week.
  • Predictive Validity: The assessment scores predict performance on a criterion measured at a future point in time.
    • Example: Evaluating whether a kindergarten universal screening assessment (e.g., DIBELS/Acadience Phonemic Segmentation Fluency) accurately predicts 3rd-grade performance on state-mandated high-stakes reading tests.
  • Standard Error of Estimate (SE_est): While SEM quantifies error around an observed score on the same test, SE_est quantifies the margin of error when using test score X to predict performance on criterion Y:
SE_est = SD_Y * √(1 - r_XY²)

3. Construct-Oriented Validity Evidence: The Core Umbrella

Construct validity examines whether an assessment truly measures the theoretical psychological trait, construct, or attribute it claims to measure. Key empirical methods include:

Convergent Validity

  • Evaluates whether the test correlates strongly with existing, validated instruments that measure the same or theoretically related constructs.
  • Example: The Working Memory Index of a new cognitive battery should demonstrate high positive correlations (r ≥ 0.70) with the Working Memory Index of the WISC-V.

Discriminant (Divergent) Validity

  • Evaluates whether the test exhibits low or near-zero correlations with instruments measuring theoretically unrelated, distinct constructs.
  • Example: A newly developed measure of childhood depression should correlate near zero (r ≤ 0.15) with a measure of phonological awareness.

Campbell & Fiske's Multitrait-Multimethod Matrix (MTMM)

  • A rigorous experimental validation framework that crosses multiple traits (e.g., Inattention, Hyperactivity, Anxiety) with multiple measurement methods (e.g., Teacher Rating Scale, Parent Rating Scale, Direct Classroom Observation).
  • Core Expectation: Correlations between the same trait measured across different methods (monotrait-heteromethod) must be substantially higher than correlations between different traits measured via the same method (heterotrait-monomethod). If method variance exceeds trait variance, the test suffers from severe method bias.

Factor Analytic Evidence (Internal Structure)

  • Exploratory Factor Analysis (EFA) & Confirmatory Factor Analysis (CFA): Statistical procedures that evaluate the internal covariance structure of subtests.
  • Example: CFA evidence in the WISC-V technical manual confirms that a 5-factor model (Verbal Comprehension, Visual Spatial, Fluid Reasoning, Working Memory, Processing Speed) provides a statistically superior fit to student response data compared to a 1-factor general intelligence (g) or 4-factor model.

6. Common Exam Pitfalls & Psychometric Fallacies

Psychometric FallacyThe Reality for Practice
"High reliability guarantees high validity."False. Reliability is a necessary but insufficient condition for validity. A broken scale that consistently reads 10 pounds light is perfectly reliable (r = 1.00), but completely invalid for measuring true weight. A test can be highly reliable while measuring the wrong construct.
"Confusing SEM with Standard Deviation (SD)."False. SD reflects the variability of scores across different people in a population (SD = 15). SEM reflects the variability of scores for a single person across repeated testing due to error (SEM ≈ 3). SEM will always be smaller than SD whenever r_xx > 0.
"Reporting isolated point scores as absolute truths."False. Professional ethics (NASP 2020) and testing standards mandate reporting confidence intervals to acknowledge random measurement error and avoid unwarranted diagnostic conclusions.
"Using split-half reliability on timed fluency probes."False. Split-half procedures spuriously inflate reliability on speeded tests. Alternate-form or test-retest procedures must be used instead.
Loading diagram...
Classical Test Theory and Measurement Precision Architecture
Test Your Knowledge

A school psychologist evaluates a 9-year-old student referred for academic difficulties. On a standardized cognitive battery with an overall standard deviation of 15 and a published internal consistency reliability coefficient of r_xx = 0.96, the student earns an observed Full Scale IQ standard score of 85. What is the approximate 95% confidence interval for this student's score, and what does this interval clinically convey to the multidisciplinary team?

A

[79 to 91]; the team can be 95% confident that the student's true score falls between 79 and 91, which accounts for random measurement error (SEM = 3.0).

B

[82 to 88]; it conveys that the student's true cognitive potential is guaranteed to be within one standard error of measurement of the observed score.

C

[70 to 100]; it conveys that the student's score fluctuates by a full standard deviation due to standard error of the estimate.

D

[85 to 95]; it conveys that practice effects will raise the student's retest score by at least two standard errors upon re-evaluation.

Test Your Knowledge

A multidisciplinary evaluation team is reviewing a newly purchased social-emotional screening rating scale. The technical manual reports that the correlation between the screener's internalizing scale and an established anxiety measure is r = 0.81, whereas its correlation with an established measure of expressive vocabulary is r = 0.12. How should the school psychologist interpret these psychometric findings?

A

The screener demonstrates criterion-related concurrent validity through high item homogeneity and high split-half reliability.

B

The screener suffers from significant construct underrepresentation because expressive language failed to correlate with internalizing symptoms.

C

The screener demonstrates strong construct validity via convergent evidence (r = 0.81 with anxiety) and discriminant evidence (r = 0.12 with vocabulary).

D

The screener exhibits high inter-rater reliability across distinct clinical raters but poor predictive validity for academic achievement.

Test Your Knowledge

A school district adopts an online timed math fact fluency assessment to screen elementary students for math intervention. The test developer reports an internal consistency reliability coefficient of r = 0.98 derived from an odd-even split-half analysis administered within a strict 2-minute testing window. Why should a school psychologist view this reported reliability coefficient with extreme psychometric caution?

A

Split-half coefficients cannot be calculated for tests that use dichotomous correct/incorrect scoring protocols.

B

Odd-even split-half reliability systematically underestimates true score variance on timed curriculum-based measures.

C

Speeded tests violate Classical Test Theory assumptions because error variance must always be larger than true score variance.

D

Split-half reliability is artificially and spuriously inflated on speeded tests because items completed are virtually all answered correctly before time expires.

Sections you finish are checked off in the contents.