16.1 Psychometric Theory: Classical Test Theory, Reliability Types, and Validity Constructs
Key Takeaways
Classical Test Theory (CTT) decomposes observed test scores into true score and random error components (X = T + E), partitioning total variance into true and error variance (σ²_X = σ²_T + σ²_E) to define reliability as r_xx' = σ²_T / σ²_X.
The Standard Error of Measurement (SEM = s_x √(1 - r_xx')) quantifies observed score fluctuations around an examinee's true score, establishing confidence intervals where higher reliability directly shrinks measurement error.
Reliability evaluates distinct dimensions of consistency: temporal stability (test-retest), form equivalence (alternate-form), internal consistency (split-half with Spearman-Brown prophecy, Cronbach's α, KR-20), and rater agreement (Cohen's κ, Intraclass Correlation).
Validity establishes whether an instrument measures its target theoretical construct, encompassing content representation, criterion prediction (concurrent vs. predictive, corrected for attenuation by r_T_X T_Y = r_xy / √(r_xx' · r_yy')), and construct networks examined via Campbell & Fiske's Multitrait-Multimethod (MTMM) matrix.
Item analysis evaluates item difficulty (p-value, optimal ~0.50) and discrimination (D = P_U - P_L), while Item Response Theory (IRT) models probabilistic performance along latent trait θ using 1PL (Rasch difficulty), 2PL (discrimination + difficulty), and 3PL (pseudo-guessing) Item Characteristic Curves.
Psychometric Theory: Classical Test Theory, Reliability Types, and Validity Constructs
Psychometrics is the scientific discipline concerned with the theory and techniques of psychological measurement. Because psychological attributes—such as intelligence, extraversion, neuroticism, or perceptual speed—cannot be directly observed or physicalized like mass or length, psychometricians construct formal mathematical models to evaluate whether assessment instruments measure target constructs reliably, validly, and with minimal measurement error.
1. Classical Test Theory (CTT) and the True Score Model
Classical Test Theory (CTT), often termed the true score model, was formalized by Charles Spearman and later refined by Harold Gulliksen and Frederic Lord. CTT posits that every observed test score () is an additive composite of two independent, unobservable quantities: an individual's latent true score () and unsystematic random measurement error ().
Theoretical Assumptions of CTT
Classical Test Theory rests upon several foundational axiomatic assumptions regarding random error:
- Zero Expected Error: Across an infinite number of repeated administrations of parallel tests, random errors cancel each other out. The expected value (mean) of error is zero:
- Uncorrelated True Scores and Errors: An examinee's true score bears no relationship to the magnitude or sign of the error affecting their score. High-ability and low-ability individuals are equally subject to positive or negative random error:
- Uncorrelated Errors Across Administrations: The measurement error affecting a person on one test form () is completely uncorrelated with the measurement error affecting them on a second test form ():
- Definition of the True Score: A person's true score is formally defined as the expected value (theoretical mean) of observed scores obtained if the individual were tested an infinite number of times under identical, independent conditions:
Variance Partitioning and the Definition of Reliability
Because true scores and random errors are assumed to be uncorrelated (), the total observed score variance () across a population of examinees partitions cleanly into the sum of true score variance () and error variance ():
From this variance decomposition, the reliability coefficient () is mathematically defined as the proportion of observed score variance that is attributable to genuine true score variance:
- When a test has zero measurement error (), reliability reaches unity (), meaning all observed variance reflects true differences among individuals.
- When all observed differences reflect pure random noise (), reliability drops to zero ().
2. The Standard Error of Measurement (SEM)
While the reliability coefficient () provides a single population-level metric of test consistency, clinicians and educational diagnosticians need to know how much uncertainty surrounds an individual examinee's observed score. This uncertainty is quantified by the Standard Error of Measurement ().
Mathematical Formulation
The represents the standard deviation of observed scores an individual would obtain if tested repeatedly across an infinite number of parallel testing sessions. It is derived directly from the sample standard deviation () and the reliability coefficient ():
Relationship Between Reliability (r_xx') and SEM (Assuming sx = 15):
Reliability (r_xx') √(1 - r_xx') SEM = 15 · √(1 - r_xx')
────────────────────────────────────────────────────────────
1.00 0.000 0.00 (Zero measurement error)
0.96 0.200 3.00 (Typical of IQ composite scores)
0.91 0.300 4.50
0.84 0.400 6.00
0.75 0.500 7.50
0.00 1.000 15.00 (Pure error; SEM equals test SD)
Constructing Confidence Intervals Around Observed Scores
Because CTT assumes measurement errors are normally distributed around the true score with a mean of zero and a standard deviation equal to the , diagnosticians construct confidence intervals (CIs) to express the margin of measurement precision around an observed score ():
- Confidence Interval ():
- Confidence Interval ():
- Confidence Interval ():
Note
Consider an examinee who obtains a Full Scale IQ score of on a test with standard deviation and reliability . The is . A confidence interval around the examinee's score is , or approximately . This informs the diagnostician that the examinee's true intellectual capacity likely falls between 104 and 116.
3. Typology of Reliability Estimates
Reliability is not an inherent, immutable property of a test; rather, it is a property of test scores obtained from a specific sample under specific administration conditions. Different reliability methodologies isolate different sources of unwanted error variance.
| Reliability Type | Method of Administration | Primary Source of Error Variance Isolated | Statistical Index Used | Primary Clinical / Research Limitation |
|---|---|---|---|---|
| Test-Retest | Same test administered to the same examinees at two distinct time points. | Time-sampling error (temporal fluctuations, mood, fatigue, illness). | Pearson (termed the coefficient of stability). | Vulnerable to carryover effects, practice effects, memory retrieval, and genuine developmental maturation across intervals. |
| Alternate-Form (Parallel-Form) | Two equivalent forms of a test containing matched content, means, and variances administered to the same group. | Content-sampling error (if given immediately); Content- and time-sampling error (if delayed). | Pearson (termed the coefficient of equivalence). | Highly expensive and laborious to construct two genuinely parallel forms; examinees may find one form systematically more difficult. |
| Split-Half | Single test administration divided into two equivalent halves (e.g., odd-numbered vs. even-numbered items). | Content-sampling error across test items. | Pearson between halves, corrected via the Spearman-Brown Prophecy formula. | Test length is effectively halved during correlation, requiring statistical correction; splitting arbitrary items can yield conflicting coefficients. |
| Internal Consistency (Cronbach's ) | Single test administration; evaluates covariance among all individual items simultaneously. | Content-sampling error and item heterogeneity. | Cronbach's Coefficient Alpha (). | Assumes tau-equivalence (equal item factor loadings); inflated by scale length; underestimates reliability if items are multidimensional. |
| Internal Consistency (KR-20) | Single test administration for tests with dichotomously scored items (right/wrong, ). | Content-sampling error across binary items. | Kuder-Richardson Formula 20 (KR-20). | Restricted exclusively to dichotomous items; cannot be used with polytomous Likert-type rating scales. |
| Inter-Rater (Inter-Scorer) | Two or more independent raters score the exact same behavioral performance or clinical protocol. | Rater-sampling error (observer bias, idiosyncratic rating severity, halo effects). | Cohen's Kappa () for categorical ratings; Intraclass Correlation (ICC) for continuous scores. | Requires intensive rater training and standardized behavioral coding rubrics. |
The Spearman-Brown Prophecy Formula
Because reliability is directly proportional to test length—longer tests sample a broader domain and allow random errors to cancel out—splitting a test into halves to calculate internal consistency artificially depresses the reliability coefficient. In 1910, Charles Spearman and William Brown independently derived the Spearman-Brown Prophecy formula to estimate the reliability of a test if its length is altered by a factor of :
- Correcting Split-Half Reliability: When estimating the reliability of the full test from two half-tests, the test is doubled in length, so . The formula simplifies to:
- If the correlation between the odd and even halves of a test is , the estimated reliability of the entire test is:
Cronbach's Alpha () and Kuder-Richardson Formula 20 (KR-20)
Lee Cronbach generalized split-half reliability in 1951 by demonstrating that Cronbach's Alpha () equals the mathematical mean of all possible split-half reliability coefficients for a test. Its computational formula is:
Where is the total number of items, is the variance of item , and is the total observed test score variance.
When items are scored dichotomously (coded for correct and for incorrect), item variance simplifies to (where is the proportion answering correctly and ). Substituting this into Cronbach's formula yields Kuder-Richardson Formula 20 (KR-20):
Inter-Rater Reliability: Cohen's Kappa ()
When two independent clinical raters categorize psychiatric patients into nominal diagnostic categories (e.g., Schizophrenia, Bipolar I, Major Depressive Disorder), calculating simple percentage agreement is misleading because raters will agree on a substantial percentage of cases purely by chance. Jacob Cohen developed Cohen's Kappa () to correct for chance agreement:
Where is the observed proportion of agreement, and is the expected proportion of agreement predicted by marginal probabilities under chance alone. A kappa of reflects perfect agreement; indicates agreement no better than chance; and reflects agreement worse than chance.
4. Validity Constructs: Content, Criterion-Related, and Construct Validity
While reliability establishes that a test measures something consistently, validity evaluates whether a test measures the specific construct it was designed to assess, and whether clinical or educational inferences drawn from test scores are scientifically justifiable. In classic psychometrics, validity is organized around the "trinitarian" model: content validity, criterion-related validity, and construct validity.
Content Validity
Content validity reflects the degree to which items on a test adequately and representatively sample the entire domain of knowledge, skills, or behaviors the test is intended to measure.
- Evaluation Method: Content validity cannot be summarized by a single correlation coefficient; it is established qualitatively during test construction by assembling panels of Subject Matter Experts (SMEs) who review test specifications, item blueprints, and alignment matrices.
- Lawshe's Content Validity Ratio (CVR): SMEs rate whether each item is "essential," "useful but not essential," or "not necessary." Items rated essential by a statistically significant majority are retained.
- Face Validity vs. Content Validity: Face validity is not a technical psychometric validity construct. It refers merely to whether a test looks valid to the untrained examinee on superficial inspection. A test can possess high scientific content validity while lacking face validity (e.g., subtle projective stimuli or empirically keyed personality items).
Criterion-Related Validity
Criterion-related validity evaluates how well scores on a predictor test () correlate with an independent, external outcome measure or standard (), termed the criterion.
- Concurrent Validity: The predictor test and criterion are measured at approximately the same point in time. Example: Administering a newly designed 15-minute depression screening scale and correlating its scores with a concurrent 3-hour structured diagnostic clinical interview conducted that same afternoon.
- Predictive Validity: The predictor test is administered first, and the criterion outcome is measured at a future point in time. Example: Administering the SAT or GRE to predict cumulative undergraduate or graduate GPA four years later.
Standard Error of Estimate ()
In criterion-related validity, the accuracy of predicting a criterion score () from an observed test score () is quantified by the Standard Error of Estimate ():
Where is the standard deviation of the criterion and is the criterion validity coefficient. Note the structural distinction: the uses test reliability () to estimate true score error, whereas the uses criterion validity () to estimate prediction error in the criterion.
The Correction for Attenuation
In the real world, both the predictor test () and the criterion measure () contain unsystematic measurement error. Because random error attenuates (dampens) linear relationships, the observed correlation between two imperfect tests () is systematically lower than the true theoretical correlation between the underlying constructs ().
Charles Spearman derived the correction for attenuation formula to estimate what the correlation between two variables would be if both were measured with perfect reliability ( and ):
- Upper Bound of Validity: A vital corollary tested frequently on the GRE is that a test's criterion validity cannot exceed the square root of its reliability when the criterion is perfectly reliable ():
If a test has a reliability of , its maximum possible theoretical correlation with any external criterion is .
5. Construct Validity and the Multitrait-Multimethod (MTMM) Matrix
Construct validity is the overarching psychometric framework assessing whether an assessment instrument truly measures a theoretical psychological construct (e.g., anxiety, intelligence, working memory). It subsumes all other forms of validity evidence and is demonstrated through an ongoing process of hypothesis testing within a theoretical nomological network.
Convergent vs. Discriminant Validity
- Convergent Validity: Demonstrates that the test correlates strongly with alternative measures designed to assess the same or theoretically similar constructs. (e.g., a new trait anxiety scale correlates with the State-Trait Anxiety Inventory).
- Discriminant (Divergent) Validity: Demonstrates that the test has low, near-zero, or negligible correlations with measures designed to assess theoretically distinct constructs. (e.g., a new trait anxiety scale correlates with a measure of mechanical reasoning ability).
Campbell & Fiske's Multitrait-Multimethod (MTMM) Matrix
In 1959, Donald T. Campbell and Donald W. Fiske revolutionized construct validation by introducing the Multitrait-Multimethod (MTMM) Matrix. They argued that every psychological measurement reflects two distinct sources of variance: the trait of interest and the method used to measure it (e.g., self-report inventory, peer rating, projective test, behavioral observation). To validate a construct, one must simultaneously measure at least two distinct traits using at least two distinct measurement methods.
Hypothetical MTMM Matrix (2 Traits: Trait A & Trait B; 2 Methods: M1 & M2)
Method 1 (Self-Report) Method 2 (Peer Rating)
Trait A1 Trait B1 Trait A2 Trait B2
───────────────────────────────────────────────────────────────────────────
Method 1 Trait A1 ( .88 )
Trait B1 .32 ( .85 )
Method 2 Trait A2 [ .68 ] .25 ( .84 )
Trait B2 .22 [ .64 ] .30 ( .82 )
───────────────────────────────────────────────────────────────────────────
Key to MTMM Coefficients:
( .88 ) = Monotrait-Monomethod (Reliability Diagonal)
[ .68 ] = Monotrait-Heteromethod (Convergent Validity Diagonal)
.32 = Heterotrait-Monomethod (Method Variance / Common-Method Bias)
.22 = Heterotrait-Heteromethod (Discriminant Validity / Construct Uniqueness)
Decoding the Four MTMM Correlations
- Monotrait-Monomethod: Same trait, same method. Located on the main diagonal. These are reliability coefficients (e.g., and ). They should represent the highest values in the entire matrix.
- Monotrait-Heteromethod: Same trait, different methods. These are convergent validity coefficients (e.g., and ). They reflect whether Trait A measured by self-report correlates strongly with Trait A measured by peer rating. They must be statistically significant and substantially greater than zero.
- Heterotrait-Monomethod: Different traits, same method (e.g., and ). These correlations reflect common-method bias (method variance). If these correlations are high, it indicates that shared measurement format (e.g., self-report response sets) artificially inflates relationships between unrelated constructs.
- Heterotrait-Heteromethod: Different traits, different methods (e.g., and ). These represent pure construct uniqueness and provide evidence for discriminant validity. They should be the lowest coefficients in the matrix.
Important
MTMM Construct Validity Requirements: For construct validity to be confirmed:
- Convergent validity coefficients (Monotrait-Heteromethod) must be substantially higher than discriminant validity coefficients (Heterotrait-Heteromethod).
- Convergent validity coefficients must be higher than method bias coefficients (Heterotrait-Monomethod).
- The pattern of trait interrelationships should remain consistent across methods.
6. Classical Item Analysis vs. Item Response Theory (IRT)
Assessment quality ultimately depends on the psychometric performance of individual test items. Psychometricians analyze items using either Classical Test Theory or Item Response Theory.
Classical Item Analysis
Classical item analysis evaluates two primary parameters: item difficulty and item discrimination.
1. Item Difficulty (-Value)
In classical testing, item difficulty is indexed by the proportion of examinees who answer the item correctly:
- Counterintuitive Direction: A higher -value indicates an easier item (e.g., means got it right), whereas a lower -value denotes a harder item (e.g., ).
- Optimal Difficulty: To maximize item variance and test score dispersion, the optimal average item difficulty is approximately .
- Correction for Guessing: On multiple-choice tests, optimal difficulty is adjusted halfway between the guessing probability () and :
For a 4-option multiple-choice item (), the optimal difficulty is .
2. Item Discrimination Index ()
Item discrimination measures how effectively an item differentiates between high-scoring and low-scoring examinees. The most common classical metric is the extreme groups index ():
Where is the proportion of examinees in the upper scoring bracket (top ) answering correctly, and is the proportion in the lower scoring bracket (bottom ) answering correctly.
- ranges from to .
- : Excellent discrimination; item cleanly distinguishes high performers.
- : Zero discrimination; item fails to differentiate high and low performers.
- : Negative discrimination; more low performers answered correctly than high performers. Such items are psychometrically defective, indicating confusing distractors, typographical errors, or incorrect answer keys, and must be eliminated.
Item Response Theory (IRT) and Item Characteristic Curves
A major limitation of Classical Test Theory is that item parameters ( and ) are sample-dependent (an item appears hard in a low-ability sample and easy in a high-ability sample), and examinee ability scores are test-dependent (scores depend on the specific difficulty of the test administered).
Item Response Theory (IRT) overcomes these limitations by modeling the nonlinear mathematical relationship between an individual's unobservable latent trait level (denoted by Greek letter theta, , where and ) and the probability of correctly answering a specific item (). This relationship is depicted graphically as an S-shaped Item Characteristic Curve (ICC) or trace line.
Item Characteristic Curve (3PL Model)
Probability P(θ)
1.0 │ ● Upper Asymptote
│ ..-''
0.8 │ _.-'
│ .-'
0.6 │ .-' <-- Steepest Slope = Discrimination (a)
│ .-'
0.4 │ .-' <-- Inflection Point θ = Difficulty (b)
│ _..-'
c --> 0.2 │...........-------'
(Guessing) │
0.0 └─────────────────────────────────────────────────────────────
-3 -2 -1 0 +1 +2 +3
Latent Trait Level (θ)
The Three Logistic IRT Models
- One-Parameter Logistic (1PL / Rasch Model):
- Assumes all items have equal discrimination () and zero guessing ().
- Models only item difficulty (), defined as the latent trait level at which an examinee has exactly a probability of answering correctly ().
- Two-Parameter Logistic (2PL Model):
- Models both item difficulty () and item discrimination ().
- The -parameter represents the slope of the ICC at the inflection point (). Steeper slopes reflect greater discrimination among examinees near the threshold .
- Three-Parameter Logistic (3PL Model):
- Incorporates difficulty (), discrimination (), and a pseudo-guessing parameter ().
- The -parameter represents the lower horizontal asymptote of the ICC—the probability that an examinee with infinitely low ability () will endorse or guess the item correctly purely by chance.
Sample Invariance Property
The defining triumph of IRT over CTT is parameter invariance: item parameters () are invariant across different examinee populations, and examinee trait estimates () are invariant across different item subsets. This mathematical foundation enables Computerized Adaptive Testing (CAT), where items are dynamically tailored in real time to the examinee's emerging estimate.
A school psychologist administers a standardized spatial reasoning battery (Mean = 100, SD = 15) with an established reliability coefficient of r_xx' = .84. A student achieves an observed score of 115. What is the standard error of measurement (SEM) of this assessment, and what is the approximate 95% confidence interval for the student's score?
SEM = 12.60; 95% CI is [90.3, 139.7]
SEM = 4.80; 95% CI is [105.6, 124.4]
SEM = 2.40; 95% CI is [110.3, 119.7]
SEM = 6.00; 95% CI is [103.2, 126.8]
An educational psychometrician develops a 40-item diagnostic math inventory and calculates a split-half reliability coefficient of r = .60 between the 20 odd-numbered items and the 20 even-numbered items. If the psychometrician uses the Spearman-Brown prophecy formula to estimate the reliability of the complete 40-item test, what is the resulting coefficient?
.68
.75
.80
.85
In Campbell and Fiske's Multitrait-Multimethod (MTMM) matrix, which of the following correlation coefficients provides direct empirical evidence for convergent validity, and what structural condition must it satisfy?
Monotrait-Heteromethod; it must be significant and clearly greater than zero across different measurement methods
Monotrait-Monomethod; it must be the highest value in the matrix to confirm internal consistency
Heterotrait-Heteromethod; it must be substantially higher than the reliability diagonal to prove construct uniqueness
Heterotrait-Monomethod; it must be zero to demonstrate that different traits do not share common method bias
A psychometrician inspects the Item Characteristic Curve (ICC) generated by a three-parameter logistic (3PL) Item Response Theory model for a challenging multiple-choice question. What does the lower horizontal asymptote of this curve represent?
The pseudo-guessing parameter (c): the chance that an examinee with extremely low ability answers correctly
The total error variance of the item derived from classical test theory
The item difficulty parameter (b), representing the trait level at which 50% of examinees respond correctly
The item discrimination parameter (a), representing the slope of the curve at its point of inflection
Sections you finish are checked off in the contents.