16.1 Psychometric Theory: Classical Test Theory, Reliability Types, and Validity Constructs

Key Takeaways

  • Classical Test Theory (CTT) decomposes observed test scores into true score and random error components (X = T + E), partitioning total variance into true and error variance (σ²_X = σ²_T + σ²_E) to define reliability as r_xx' = σ²_T / σ²_X.

  • The Standard Error of Measurement (SEM = s_x √(1 - r_xx')) quantifies observed score fluctuations around an examinee's true score, establishing confidence intervals where higher reliability directly shrinks measurement error.

  • Reliability evaluates distinct dimensions of consistency: temporal stability (test-retest), form equivalence (alternate-form), internal consistency (split-half with Spearman-Brown prophecy, Cronbach's α, KR-20), and rater agreement (Cohen's κ, Intraclass Correlation).

  • Validity establishes whether an instrument measures its target theoretical construct, encompassing content representation, criterion prediction (concurrent vs. predictive, corrected for attenuation by r_T_X T_Y = r_xy / √(r_xx' · r_yy')), and construct networks examined via Campbell & Fiske's Multitrait-Multimethod (MTMM) matrix.

  • Item analysis evaluates item difficulty (p-value, optimal ~0.50) and discrimination (D = P_U - P_L), while Item Response Theory (IRT) models probabilistic performance along latent trait θ using 1PL (Rasch difficulty), 2PL (discrimination + difficulty), and 3PL (pseudo-guessing) Item Characteristic Curves.

Last updated: October 2026

Psychometric Theory: Classical Test Theory, Reliability Types, and Validity Constructs

Psychometrics is the scientific discipline concerned with the theory and techniques of psychological measurement. Because psychological attributes—such as intelligence, extraversion, neuroticism, or perceptual speed—cannot be directly observed or physicalized like mass or length, psychometricians construct formal mathematical models to evaluate whether assessment instruments measure target constructs reliably, validly, and with minimal measurement error.

1. Classical Test Theory (CTT) and the True Score Model

Classical Test Theory (CTT), often termed the true score model, was formalized by Charles Spearman and later refined by Harold Gulliksen and Frederic Lord. CTT posits that every observed test score (XX) is an additive composite of two independent, unobservable quantities: an individual's latent true score (TT) and unsystematic random measurement error (EE).

X=T+EX = T + E

Theoretical Assumptions of CTT

Classical Test Theory rests upon several foundational axiomatic assumptions regarding random error:

  1. Zero Expected Error: Across an infinite number of repeated administrations of parallel tests, random errors cancel each other out. The expected value (mean) of error is zero:

E(E)=0E(E) = 0

  1. Uncorrelated True Scores and Errors: An examinee's true score bears no relationship to the magnitude or sign of the error affecting their score. High-ability and low-ability individuals are equally subject to positive or negative random error:

rTE=0r_{TE} = 0

  1. Uncorrelated Errors Across Administrations: The measurement error affecting a person on one test form (E1E_1) is completely uncorrelated with the measurement error affecting them on a second test form (E2E_2):

rE1E2=0r_{E_1 E_2} = 0

  1. Definition of the True Score: A person's true score is formally defined as the expected value (theoretical mean) of observed scores obtained if the individual were tested an infinite number of times under identical, independent conditions:

T=E(X)T = E(X)

Variance Partitioning and the Definition of Reliability

Because true scores and random errors are assumed to be uncorrelated (rTE=0r_{TE} = 0), the total observed score variance (σX2\sigma_X^2) across a population of examinees partitions cleanly into the sum of true score variance (σT2\sigma_T^2) and error variance (σE2\sigma_E^2):

σX2=σT2+σE2\sigma_X^2 = \sigma_T^2 + \sigma_E^2

From this variance decomposition, the reliability coefficient (rxx′r_{xx'}) is mathematically defined as the proportion of observed score variance that is attributable to genuine true score variance:

rxx′=σT2σX2=σX2−σE2σX2=1−σE2σX2r_{xx'} = \frac{\sigma_T^2}{\sigma_X^2} = \frac{\sigma_X^2 - \sigma_E^2}{\sigma_X^2} = 1 - \frac{\sigma_E^2}{\sigma_X^2}

  • When a test has zero measurement error (σE2=0\sigma_E^2 = 0), reliability reaches unity (rxx′=1.0r_{xx'} = 1.0), meaning all observed variance reflects true differences among individuals.
  • When all observed differences reflect pure random noise (σT2=0\sigma_T^2 = 0), reliability drops to zero (rxx′=0.0r_{xx'} = 0.0).

2. The Standard Error of Measurement (SEM)

While the reliability coefficient (rxx′r_{xx'}) provides a single population-level metric of test consistency, clinicians and educational diagnosticians need to know how much uncertainty surrounds an individual examinee's observed score. This uncertainty is quantified by the Standard Error of Measurement (SEMSEM).

Mathematical Formulation

The SEMSEM represents the standard deviation of observed scores an individual would obtain if tested repeatedly across an infinite number of parallel testing sessions. It is derived directly from the sample standard deviation (sxs_x) and the reliability coefficient (rxx′r_{xx'}):

SEM=sx1−rxx′SEM = s_x \sqrt{1 - r_{xx'}}

Relationship Between Reliability (r_xx') and SEM (Assuming sx = 15):

Reliability (r_xx')   √(1 - r_xx')   SEM = 15 · √(1 - r_xx')
────────────────────────────────────────────────────────────
1.00                  0.000          0.00  (Zero measurement error)
0.96                  0.200          3.00  (Typical of IQ composite scores)
0.91                  0.300          4.50
0.84                  0.400          6.00
0.75                  0.500          7.50
0.00                  1.000         15.00  (Pure error; SEM equals test SD)

Constructing Confidence Intervals Around Observed Scores

Because CTT assumes measurement errors are normally distributed around the true score with a mean of zero and a standard deviation equal to the SEMSEM, diagnosticians construct confidence intervals (CIs) to express the margin of measurement precision around an observed score (XX):

Confidence Interval=X±(zcritical×SEM)\text{Confidence Interval} = X \pm (z_{\text{critical}} \times SEM)

  • 68%68\% Confidence Interval (z=±1.00z = \pm 1.00): X±1.00(SEM)X \pm 1.00(SEM)
  • 95%95\% Confidence Interval (z=±1.96z = \pm 1.96): X±1.96(SEM)X \pm 1.96(SEM)
  • 99%99\% Confidence Interval (z=±2.58z = \pm 2.58): X±2.58(SEM)X \pm 2.58(SEM)

Note

Consider an examinee who obtains a Full Scale IQ score of X=110X = 110 on a test with standard deviation sx=15s_x = 15 and reliability rxx′=.96r_{xx'} = .96. The SEMSEM is 151−.96=15.04=15(0.20)=3.015\sqrt{1 - .96} = 15\sqrt{.04} = 15(0.20) = 3.0. A 95%95\% confidence interval around the examinee's score is 110±1.96(3.0)=110±5.88110 \pm 1.96(3.0) = 110 \pm 5.88, or approximately [104,116][104, 116]. This informs the diagnostician that the examinee's true intellectual capacity likely falls between 104 and 116.

3. Typology of Reliability Estimates

Reliability is not an inherent, immutable property of a test; rather, it is a property of test scores obtained from a specific sample under specific administration conditions. Different reliability methodologies isolate different sources of unwanted error variance.

Reliability TypeMethod of AdministrationPrimary Source of Error Variance IsolatedStatistical Index UsedPrimary Clinical / Research Limitation
Test-RetestSame test administered to the same examinees at two distinct time points.Time-sampling error (temporal fluctuations, mood, fatigue, illness).Pearson rr (termed the coefficient of stability).Vulnerable to carryover effects, practice effects, memory retrieval, and genuine developmental maturation across intervals.
Alternate-Form (Parallel-Form)Two equivalent forms of a test containing matched content, means, and variances administered to the same group.Content-sampling error (if given immediately); Content- and time-sampling error (if delayed).Pearson rr (termed the coefficient of equivalence).Highly expensive and laborious to construct two genuinely parallel forms; examinees may find one form systematically more difficult.
Split-HalfSingle test administration divided into two equivalent halves (e.g., odd-numbered vs. even-numbered items).Content-sampling error across test items.Pearson rr between halves, corrected via the Spearman-Brown Prophecy formula.Test length is effectively halved during correlation, requiring statistical correction; splitting arbitrary items can yield conflicting coefficients.
Internal Consistency (Cronbach's α\alpha)Single test administration; evaluates covariance among all individual items simultaneously.Content-sampling error and item heterogeneity.Cronbach's Coefficient Alpha (α\alpha).Assumes tau-equivalence (equal item factor loadings); inflated by scale length; underestimates reliability if items are multidimensional.
Internal Consistency (KR-20)Single test administration for tests with dichotomously scored items (right/wrong, 0/10/1).Content-sampling error across binary items.Kuder-Richardson Formula 20 (KR-20).Restricted exclusively to dichotomous items; cannot be used with polytomous Likert-type rating scales.
Inter-Rater (Inter-Scorer)Two or more independent raters score the exact same behavioral performance or clinical protocol.Rater-sampling error (observer bias, idiosyncratic rating severity, halo effects).Cohen's Kappa (κ\kappa) for categorical ratings; Intraclass Correlation (ICC) for continuous scores.Requires intensive rater training and standardized behavioral coding rubrics.

The Spearman-Brown Prophecy Formula

Because reliability is directly proportional to test length—longer tests sample a broader domain and allow random errors to cancel out—splitting a test into halves to calculate internal consistency artificially depresses the reliability coefficient. In 1910, Charles Spearman and William Brown independently derived the Spearman-Brown Prophecy formula to estimate the reliability of a test if its length is altered by a factor of kk:

rSB=k⋅rxy1+(k−1)rxyr_{SB} = \frac{k \cdot r_{xy}}{1 + (k - 1)r_{xy}}

Where k=Number of items in the new testNumber of items in the original test\text{Where } k = \frac{\text{Number of items in the new test}}{\text{Number of items in the original test}}

  • Correcting Split-Half Reliability: When estimating the reliability of the full test from two half-tests, the test is doubled in length, so k=2k = 2. The formula simplifies to:

rfull=2r1/21+r1/2r_{\text{full}} = \frac{2 r_{1/2}}{1 + r_{1/2}}

  • If the correlation between the odd and even halves of a test is r1/2=.70r_{1/2} = .70, the estimated reliability of the entire test is:

rfull=2(.70)1+.70=1.401.70≈.824r_{\text{full}} = \frac{2(.70)}{1 + .70} = \frac{1.40}{1.70} \approx .824

Cronbach's Alpha (α\alpha) and Kuder-Richardson Formula 20 (KR-20)

Lee Cronbach generalized split-half reliability in 1951 by demonstrating that Cronbach's Alpha (α\alpha) equals the mathematical mean of all possible split-half reliability coefficients for a test. Its computational formula is:

α=kk−1(1−∑i=1ksi2sX2)\alpha = \frac{k}{k - 1} \left(1 - \frac{\sum_{i=1}^k s_i^2}{s_X^2}\right)

Where kk is the total number of items, si2s_i^2 is the variance of item ii, and sX2s_X^2 is the total observed test score variance.

When items are scored dichotomously (coded 11 for correct and 00 for incorrect), item variance simplifies to piqip_i q_i (where pip_i is the proportion answering correctly and qi=1−piq_i = 1 - p_i). Substituting this into Cronbach's formula yields Kuder-Richardson Formula 20 (KR-20):

KR-20=kk−1(1−∑i=1kpiqisX2)KR\text{-}20 = \frac{k}{k - 1} \left(1 - \frac{\sum_{i=1}^k p_i q_i}{s_X^2}\right)

Inter-Rater Reliability: Cohen's Kappa (κ\kappa)

When two independent clinical raters categorize psychiatric patients into nominal diagnostic categories (e.g., Schizophrenia, Bipolar I, Major Depressive Disorder), calculating simple percentage agreement is misleading because raters will agree on a substantial percentage of cases purely by chance. Jacob Cohen developed Cohen's Kappa (κ\kappa) to correct for chance agreement:

κ=Po−Pe1−Pe\kappa = \frac{P_o - P_e}{1 - P_e}

Where PoP_o is the observed proportion of agreement, and PeP_e is the expected proportion of agreement predicted by marginal probabilities under chance alone. A kappa of κ=1.0\kappa = 1.0 reflects perfect agreement; κ=0\kappa = 0 indicates agreement no better than chance; and κ<0\kappa < 0 reflects agreement worse than chance.

4. Validity Constructs: Content, Criterion-Related, and Construct Validity

While reliability establishes that a test measures something consistently, validity evaluates whether a test measures the specific construct it was designed to assess, and whether clinical or educational inferences drawn from test scores are scientifically justifiable. In classic psychometrics, validity is organized around the "trinitarian" model: content validity, criterion-related validity, and construct validity.

Content Validity

Content validity reflects the degree to which items on a test adequately and representatively sample the entire domain of knowledge, skills, or behaviors the test is intended to measure.

  • Evaluation Method: Content validity cannot be summarized by a single correlation coefficient; it is established qualitatively during test construction by assembling panels of Subject Matter Experts (SMEs) who review test specifications, item blueprints, and alignment matrices.
  • Lawshe's Content Validity Ratio (CVR): SMEs rate whether each item is "essential," "useful but not essential," or "not necessary." Items rated essential by a statistically significant majority are retained.
  • Face Validity vs. Content Validity: Face validity is not a technical psychometric validity construct. It refers merely to whether a test looks valid to the untrained examinee on superficial inspection. A test can possess high scientific content validity while lacking face validity (e.g., subtle projective stimuli or empirically keyed personality items).

Criterion-Related Validity

Criterion-related validity evaluates how well scores on a predictor test (XX) correlate with an independent, external outcome measure or standard (YY), termed the criterion.

  1. Concurrent Validity: The predictor test and criterion are measured at approximately the same point in time. Example: Administering a newly designed 15-minute depression screening scale and correlating its scores with a concurrent 3-hour structured diagnostic clinical interview conducted that same afternoon.
  2. Predictive Validity: The predictor test is administered first, and the criterion outcome is measured at a future point in time. Example: Administering the SAT or GRE to predict cumulative undergraduate or graduate GPA four years later.

Standard Error of Estimate (SEestSE_{\text{est}})

In criterion-related validity, the accuracy of predicting a criterion score (YY) from an observed test score (XX) is quantified by the Standard Error of Estimate (SEestSE_{\text{est}}):

SEest=sy1−rxy2SE_{\text{est}} = s_y \sqrt{1 - r_{xy}^2}

Where sys_y is the standard deviation of the criterion and rxyr_{xy} is the criterion validity coefficient. Note the structural distinction: the SEMSEM uses test reliability (sx1−rxx′s_x \sqrt{1 - r_{xx'}}) to estimate true score error, whereas the SEestSE_{\text{est}} uses criterion validity (sy1−rxy2s_y \sqrt{1 - r_{xy}^2}) to estimate prediction error in the criterion.

The Correction for Attenuation

In the real world, both the predictor test (XX) and the criterion measure (YY) contain unsystematic measurement error. Because random error attenuates (dampens) linear relationships, the observed correlation between two imperfect tests (rxyr_{xy}) is systematically lower than the true theoretical correlation between the underlying constructs (rTXTYr_{T_X T_Y}).

Charles Spearman derived the correction for attenuation formula to estimate what the correlation between two variables would be if both were measured with perfect reliability (rxx′=1.0r_{xx'} = 1.0 and ryy′=1.0r_{yy'} = 1.0):

rTXTY=rxyrxx′⋅ryy′r_{T_X T_Y} = \frac{r_{xy}}{\sqrt{r_{xx'} \cdot r_{yy'}}}

  • Upper Bound of Validity: A vital corollary tested frequently on the GRE is that a test's criterion validity cannot exceed the square root of its reliability when the criterion is perfectly reliable (ryy′=1.0r_{yy'} = 1.0):

rxy(max⁡)=rxx′r_{xy(\max)} = \sqrt{r_{xx'}}

If a test has a reliability of rxx′=.64r_{xx'} = .64, its maximum possible theoretical correlation with any external criterion is .64=.80\sqrt{.64} = .80.

5. Construct Validity and the Multitrait-Multimethod (MTMM) Matrix

Construct validity is the overarching psychometric framework assessing whether an assessment instrument truly measures a theoretical psychological construct (e.g., anxiety, intelligence, working memory). It subsumes all other forms of validity evidence and is demonstrated through an ongoing process of hypothesis testing within a theoretical nomological network.

Convergent vs. Discriminant Validity

  • Convergent Validity: Demonstrates that the test correlates strongly with alternative measures designed to assess the same or theoretically similar constructs. (e.g., a new trait anxiety scale correlates r=.82r = .82 with the State-Trait Anxiety Inventory).
  • Discriminant (Divergent) Validity: Demonstrates that the test has low, near-zero, or negligible correlations with measures designed to assess theoretically distinct constructs. (e.g., a new trait anxiety scale correlates r=.12r = .12 with a measure of mechanical reasoning ability).

Campbell & Fiske's Multitrait-Multimethod (MTMM) Matrix

In 1959, Donald T. Campbell and Donald W. Fiske revolutionized construct validation by introducing the Multitrait-Multimethod (MTMM) Matrix. They argued that every psychological measurement reflects two distinct sources of variance: the trait of interest and the method used to measure it (e.g., self-report inventory, peer rating, projective test, behavioral observation). To validate a construct, one must simultaneously measure at least two distinct traits using at least two distinct measurement methods.

          Hypothetical MTMM Matrix (2 Traits: Trait A & Trait B; 2 Methods: M1 & M2)

                       Method 1 (Self-Report)        Method 2 (Peer Rating)
                       Trait A1      Trait B1        Trait A2      Trait B2
 ───────────────────────────────────────────────────────────────────────────
 Method 1  Trait A1    ( .88 )       
           Trait B1      .32         ( .85 )

 Method 2  Trait A2    [ .68 ]         .25           ( .84 )
           Trait B2      .22         [ .64 ]           .30         ( .82 )
 ───────────────────────────────────────────────────────────────────────────
 Key to MTMM Coefficients:
 ( .88 ) = Monotrait-Monomethod (Reliability Diagonal)
 [ .68 ] = Monotrait-Heteromethod (Convergent Validity Diagonal)
   .32   = Heterotrait-Monomethod (Method Variance / Common-Method Bias)
   .22   = Heterotrait-Heteromethod (Discriminant Validity / Construct Uniqueness)

Decoding the Four MTMM Correlations

  1. Monotrait-Monomethod: Same trait, same method. Located on the main diagonal. These are reliability coefficients (e.g., (.88)(.88) and (.85)(.85)). They should represent the highest values in the entire matrix.
  2. Monotrait-Heteromethod: Same trait, different methods. These are convergent validity coefficients (e.g., [.68][.68] and [.64][.64]). They reflect whether Trait A measured by self-report correlates strongly with Trait A measured by peer rating. They must be statistically significant and substantially greater than zero.
  3. Heterotrait-Monomethod: Different traits, same method (e.g., .32.32 and .30.30). These correlations reflect common-method bias (method variance). If these correlations are high, it indicates that shared measurement format (e.g., self-report response sets) artificially inflates relationships between unrelated constructs.
  4. Heterotrait-Heteromethod: Different traits, different methods (e.g., .22.22 and .25.25). These represent pure construct uniqueness and provide evidence for discriminant validity. They should be the lowest coefficients in the matrix.

Important

MTMM Construct Validity Requirements: For construct validity to be confirmed:

  1. Convergent validity coefficients (Monotrait-Heteromethod) must be substantially higher than discriminant validity coefficients (Heterotrait-Heteromethod).
  2. Convergent validity coefficients must be higher than method bias coefficients (Heterotrait-Monomethod).
  3. The pattern of trait interrelationships should remain consistent across methods.

6. Classical Item Analysis vs. Item Response Theory (IRT)

Assessment quality ultimately depends on the psychometric performance of individual test items. Psychometricians analyze items using either Classical Test Theory or Item Response Theory.

Classical Item Analysis

Classical item analysis evaluates two primary parameters: item difficulty and item discrimination.

1. Item Difficulty (pp-Value)

In classical testing, item difficulty is indexed by the proportion of examinees who answer the item correctly:

p=NcorrectNtotalp = \frac{N_{\text{correct}}}{N_{\text{total}}}

  • Counterintuitive Direction: A higher pp-value indicates an easier item (e.g., p=.90p = .90 means 90%90\% got it right), whereas a lower pp-value denotes a harder item (e.g., p=.15p = .15).
  • Optimal Difficulty: To maximize item variance and test score dispersion, the optimal average item difficulty is approximately p≈.50p \approx .50.
  • Correction for Guessing: On multiple-choice tests, optimal difficulty is adjusted halfway between the guessing probability (gg) and 1.001.00:

Optimal p=1.00+g2\text{Optimal } p = \frac{1.00 + g}{2}

For a 4-option multiple-choice item (g=.25g = .25), the optimal difficulty is 1.00+.252=.625\frac{1.00 + .25}{2} = .625.

2. Item Discrimination Index (DD)

Item discrimination measures how effectively an item differentiates between high-scoring and low-scoring examinees. The most common classical metric is the extreme groups index (DD):

D=PU−PLD = P_U - P_L

Where PUP_U is the proportion of examinees in the upper scoring bracket (top 27%27\%) answering correctly, and PLP_L is the proportion in the lower scoring bracket (bottom 27%27\%) answering correctly.

  • DD ranges from −1.00-1.00 to +1.00+1.00.
  • D≥.40D \ge .40: Excellent discrimination; item cleanly distinguishes high performers.
  • D≈0.00D \approx 0.00: Zero discrimination; item fails to differentiate high and low performers.
  • D<0.00D < 0.00: Negative discrimination; more low performers answered correctly than high performers. Such items are psychometrically defective, indicating confusing distractors, typographical errors, or incorrect answer keys, and must be eliminated.

Item Response Theory (IRT) and Item Characteristic Curves

A major limitation of Classical Test Theory is that item parameters (pp and DD) are sample-dependent (an item appears hard in a low-ability sample and easy in a high-ability sample), and examinee ability scores are test-dependent (scores depend on the specific difficulty of the test administered).

Item Response Theory (IRT) overcomes these limitations by modeling the nonlinear mathematical relationship between an individual's unobservable latent trait level (denoted by Greek letter theta, θ\theta, where μ=0\mu = 0 and σ=1\sigma = 1) and the probability of correctly answering a specific item (P(θ)P(\theta)). This relationship is depicted graphically as an S-shaped Item Characteristic Curve (ICC) or trace line.

                   Item Characteristic Curve (3PL Model)

    Probability P(θ)
         1.0 │                                             ● Upper Asymptote
             │                                        ..-''
         0.8 │                                    _.-'
             │                                 .-'
         0.6 │                              .-'  <-- Steepest Slope = Discrimination (a)
             │                           .-'
         0.4 │                        .-'  <-- Inflection Point θ = Difficulty (b)
             │                   _..-'
   c --> 0.2 │...........-------'
  (Guessing) │
         0.0 └─────────────────────────────────────────────────────────────
            -3          -2          -1           0          +1          +2          +3
                                    Latent Trait Level (θ)

The Three Logistic IRT Models

  1. One-Parameter Logistic (1PL / Rasch Model):
    • Assumes all items have equal discrimination (aa) and zero guessing (c=0c = 0).
    • Models only item difficulty (bb), defined as the latent trait level θ\theta at which an examinee has exactly a 50%50\% probability of answering correctly (P(θ)=.50P(\theta) = .50).
  2. Two-Parameter Logistic (2PL Model):
    • Models both item difficulty (bb) and item discrimination (aa).
    • The aa-parameter represents the slope of the ICC at the inflection point (bb). Steeper slopes reflect greater discrimination among examinees near the threshold θ\theta.
  3. Three-Parameter Logistic (3PL Model):
    • Incorporates difficulty (bb), discrimination (aa), and a pseudo-guessing parameter (cc).
    • The cc-parameter represents the lower horizontal asymptote of the ICC—the probability that an examinee with infinitely low ability (θ→−∞\theta \rightarrow -\infty) will endorse or guess the item correctly purely by chance.

Sample Invariance Property

The defining triumph of IRT over CTT is parameter invariance: item parameters (a,b,ca, b, c) are invariant across different examinee populations, and examinee trait estimates (θ\theta) are invariant across different item subsets. This mathematical foundation enables Computerized Adaptive Testing (CAT), where items are dynamically tailored in real time to the examinee's emerging θ\theta estimate.

Test Your Knowledge

A school psychologist administers a standardized spatial reasoning battery (Mean = 100, SD = 15) with an established reliability coefficient of r_xx' = .84. A student achieves an observed score of 115. What is the standard error of measurement (SEM) of this assessment, and what is the approximate 95% confidence interval for the student's score?

A

SEM = 12.60; 95% CI is [90.3, 139.7]

B

SEM = 4.80; 95% CI is [105.6, 124.4]

C

SEM = 2.40; 95% CI is [110.3, 119.7]

D

SEM = 6.00; 95% CI is [103.2, 126.8]

Test Your Knowledge

An educational psychometrician develops a 40-item diagnostic math inventory and calculates a split-half reliability coefficient of r = .60 between the 20 odd-numbered items and the 20 even-numbered items. If the psychometrician uses the Spearman-Brown prophecy formula to estimate the reliability of the complete 40-item test, what is the resulting coefficient?

A

.68

B

.75

C

.80

D

.85

Test Your Knowledge

In Campbell and Fiske's Multitrait-Multimethod (MTMM) matrix, which of the following correlation coefficients provides direct empirical evidence for convergent validity, and what structural condition must it satisfy?

A

Monotrait-Heteromethod; it must be significant and clearly greater than zero across different measurement methods

B

Monotrait-Monomethod; it must be the highest value in the matrix to confirm internal consistency

C

Heterotrait-Heteromethod; it must be substantially higher than the reliability diagonal to prove construct uniqueness

D

Heterotrait-Monomethod; it must be zero to demonstrate that different traits do not share common method bias

Test Your Knowledge

A psychometrician inspects the Item Characteristic Curve (ICC) generated by a three-parameter logistic (3PL) Item Response Theory model for a challenging multiple-choice question. What does the lower horizontal asymptote of this curve represent?

A

The pseudo-guessing parameter (c): the chance that an examinee with extremely low ability answers correctly

B

The total error variance of the item derived from classical test theory

C

The item difficulty parameter (b), representing the trait level at which 50% of examinees respond correctly

D

The item discrimination parameter (a), representing the slope of the curve at its point of inflection

Sections you finish are checked off in the contents.