5.2 Validity, Reliability, and Measurement Error

Key Takeaways

  • Validity refers to the accuracy, truthfulness, and appropriateness of the inferences and decisions made from assessment scores (measuring what is intended to be measured).
  • Reliability refers to the consistency, stability, and reproducibility of assessment results across time, test forms, and evaluators.
  • A foundational psychometric axiom states: An assessment can be highly reliable without being valid, but it CANNOT be valid without being reliable.
  • The Standard Error of Measurement (SEM) quantifies measurement uncertainty around an observed score, establishing confidence intervals (score bands) within which a student's true score likely resides.
  • Construct-irrelevant variance (e.g., overly complex reading vocabulary on a math exam) and test bias compromise validity by introducing systematic measurement error that disadvantages specific student populations.
Last updated: August 2026

5.2 Validity, Reliability, and Measurement Error

For classroom assessments and standardized examinations to provide meaningful, defensible data, they must adhere to rigorous psychometric standards. Under Competency 4 of the FTCE Professional Education Test, teachers must understand how measurement qualities are evaluated, how errors distort test results, and how to construct assessments that yield accurate, equitable inferences regarding student achievement.

Two essential technical concepts govern all educational measurement: Validity (accuracy of inference) and Reliability (consistency of score).


1. Test Validity: Accuracy and Meaningfulness of Inferences

In modern measurement theory (anchored by Samuel Messick and the Standards for Educational and Psychological Testing), validity is not an intrinsic property of a test instrument itself, but rather the degree to which empirical evidence and theoretical rationales support the adequacy and appropriateness of interpretations and actions based on test scores.

In simple terms: Does the assessment measure what it purports to measure, and are the resulting pedagogical conclusions sound?

+-----------------------------------------------------------------------------------------+
|                              THE FOUR FACETS OF TEST VALIDITY                           |
|                                                                                         |
|   [CONTENT VALIDITY]        ---> Representative sampling of the targeted domain/standard |
|                                  Constructed via a Table of Specifications (Test Blueprint)|
|                                                                                         |
|   [CONSTRUCT VALIDITY]      ---> Accurate measurement of an underlying psychological     |
|                                  trait, theory, or unobservable attribute (e.g., logic) |
|                                                                                         |
|   [CRITERION-RELATED]       ---> Correlation with an external standard or future outcome  |
|     • Concurrent Validity   ---> Correlates with an established measure given simultaneously|
|     • Predictive Validity   ---> Accurately forecasts future performance (e.g., SAT/ACT)  |
|                                                                                         |
|   [FACE VALIDITY]           ---> Superficial appearance of test relevance to lay observers|
|                                  (Weakest form; non-statistical)                        |
+-----------------------------------------------------------------------------------------+

Deep Breakdown of Validity Types

Validity TypeOperational DefinitionClassroom Construction / Verification MethodConcrete Example
Content ValidityThe extent to which assessment items comprehensively and proportionally sample the entire content domain and cognitive skills defined by learning benchmarks.Developing a Table of Specifications (Test Blueprint) that cross-references state standards with Bloom's/DOK cognitive levels before item writing.A 50-item semester exam in Florida World History that allocates exactly 20% of questions to Ancient Greece if that unit comprised 20% of instructional time.
Construct ValidityThe degree to which an assessment accurately captures an unobservable, theoretical psychological construct (e.g., reading comprehension, creativity, spatial reasoning) without contamination.Statistical factor analysis; establishing convergent validity (high correlation with similar constructs) and discriminant validity (low correlation with unrelated traits).A mathematical word problem test that measures algebraic reasoning rather than student reading decoding ability.
Criterion-Related: ConcurrentHow well test scores correlate with an established, validated external measure administered at approximately the same time.Calculating correlation coefficients (r) between a newly developed classroom screening probe and a validated standardized diagnostic battery.A new 10-minute digital reading screener yielding scores that correlate r = 0.88 with the established DIBELS assessment.
Criterion-Related: PredictiveThe extent to which assessment scores accurately forecast a student's future academic performance or behavior.Longitudinal tracking correlating initial assessment scores with subsequent performance metrics (e.g., GPA, college completion).High school GPA and SAT scores predicting first-year university grade performance.
Face ValidityThe superficial, subjective impression of whether the test items look relevant, fair, and professional to examinees and parents.Informal review of test formatting, font clarity, terminology, and contextual framing.A chemistry exam featuring realistic laboratory scenarios rather than questions framed around video game trivia.

[!NOTE] The Table of Specifications (Test Blueprint): The single most effective tool for establishing content validity in classroom assessment is the Table of Specifications (TOS). A TOS is a two-way grid that aligns content topics (rows) with cognitive levels (columns), ensuring that the test does not over-sample low-level recall items or ignore critical unit benchmarks.


2. Test Reliability: Consistency and Precision

Reliability refers to the consistency, stability, and repeatability of assessment results across repeated administrations, varying test forms, or different evaluators. A reliable test produces stable results that are free from erratic random measurement error.

+-----------------------------------------------------------------------------------------+
|                            THE FOUR TYPES OF TEST RELIABILITY                           |
|                                                                                         |
|   [TEST-RETEST RELIABILITY]     ---> Stability of scores across time intervals           |
|                                      (Administer same test at Time 1 and Time 2)        |
|                                                                                         |
|   [INTERNAL CONSISTENCY]        ---> Homogeneity and coherence of items within one test |
|                                      (Cronbach's Alpha, Split-Half, Kuder-Richardson)    |
|                                                                                         |
|   [INTER-RATER RELIABILITY]     ---> Consistency across independent scorers / evaluators|
|                                      (Scoring rubrics, calibration training, moderation) |
|                                                                                         |
|   [PARALLEL / ALTERNATE FORMS]  ---> Equivalence across two distinct versions of a test |
|                                      (Form A vs. Form B measuring identical blueprint)  |
+-----------------------------------------------------------------------------------------+

Detailed Breakdown of Reliability Types

Reliability TypeMethod of CalculationThreat / Source of ErrorTeacher Best Practice to Maximize
Test-RetestPearson correlation (r) between scores from the same group taking the same test on two distinct occasions.Memory carryover (if time is too short); maturation/learning (if time is too long).Allow 1–2 weeks between administrations when testing stable traits; ensure testing environment remains identical.
Internal ConsistencyCronbach's Alpha (α), Split-Half correlation (Spearman-Brown formula), or KR-20 (for dichotomous items).Flawed, ambiguous, or multidimensional items that measure disconnected skills.Increase the number of well-written, aligned items; ensure all items within a subscale measure the targeted standard.
Inter-Rater (Inter-Scorer)Percentage agreement or Cohen's Kappa (κ) between two or more independent evaluators scoring the same student work.Subjective grading bias, halo effect, fatigue, or vague scoring criteria.Utilize standardized analytic rubrics, conduct blind scoring, and hold anchor-paper calibration sessions prior to grading.
Parallel / Alternate FormsCorrelation between scores obtained by the same students on two distinct forms (Form A and Form B) built from the same blueprint.Differences in item difficulty, vocabulary load, or cognitive complexity between forms.Generate items simultaneously from an identical Table of Specifications matching item difficulty (p-values).

3. The Interdependence of Validity and Reliability

A critical, frequently tested psychometric relationship on the FTCE examination is the relationship between validity and reliability:

+-----------------------------------------------------------------------------------------+
|                     THE DARTBOARD ANALOGY: RELIABILITY VS. VALIDITY                     |
|                                                                                         |
|        🎯 [Reliable, Not Valid]       🎯 [Not Reliable, Not Valid]     🎯 [Reliable & Valid]     |
|        (Tightly clustered off-target) (Scattered all over)        (Tightly clustered bullseye)|
|                                                                                         |
|        Hits the exact same wrong spot  Hits random locations;      Hits the intended target  |
|        every single time.              no pattern or accuracy.     consistently every time.  |
+-----------------------------------------------------------------------------------------+

The Fundamental Psychometric Axiom:

Reliability is a NECESSARY, but NOT SUFFICIENT, condition for Validity.

  • A test CAN be highly reliable without being valid: Example: A broken digital scale that consistently records exactly 10 pounds under a person's actual weight is perfectly reliable (identical readings every time), but entirely invalid (does not report true weight). Classroom Example: A 100-item multiple-choice spelling test administered to evaluate mathematical problem-solving will yield highly consistent, reliable scores, but it has zero validity for assessing mathematical competence.
  • A test CANNOT be valid without being reliable: If an assessment produces wildly erratic, unstable scores due to massive measurement error, those scores cannot accurately represent the student's true mastery of the standard.

4. Classical Test Theory & Standard Error of Measurement (SEM)

According to Classical Test Theory (CTT), every score a student earns on an assessment (the Observed Score, X) is composed of two theoretical components: the student's actual underlying capability (the True Score, T) and random error introduced by the testing situation (the Error Score, E):

X=T+Ewhere Observed Score=True Score±Measurement ErrorX = T + E \quad \text{where } \text{Observed Score} = \text{True Score} \pm \text{Measurement Error}

Sources of Measurement Error

  • Student Factors (Internal): Fatigue, test anxiety, illness, hunger, medication changes, fleeting motivational lapses.
  • Environmental Factors (External): Noise disruptions, room temperature extremes, flickering lighting, malfunctioning technology.
  • Test / Administration Factors: Ambiguously worded stems, confusing formatting, strict unfair time limits, subjective scoring.

The Standard Error of Measurement (SEM)

The Standard Error of Measurement (SEM) quantifies the standard deviation of error scores. It reflects the amount of spread or uncertainty around an observed score. The SEM is mathematically related to the test's standard deviation (sx) and reliability coefficient (rxx):

SEM=sx×1rxxSEM = s_x \times \sqrt{1 - r_{xx}}

[!IMPORTANT] Psychometric Rule: As test reliability (rxx) increases, the Standard Error of Measurement (SEM) decreases (approaching zero error). Conversely, low reliability yields high SEM and wide uncertainty.

+-----------------------------------------------------------------------------------------+
|                        CONFIDENCE INTERVALS AND SCORE BANDS                             |
|                                                                                         |
|                              Observed Score = 85 (SEM = 3)                              |
|                                                                                         |
|                       79          82          85          88          91                |
|                       |-----------|-----------|-----------|-----------|                 |
|                                   [   68% Confidence Band   ]                           |
|                                   (85 ± 1 SEM = 82 to 88)                               |
|                       [             95% Confidence Band             ]                   |
|                       (85 ± 2 SEM = 79 to 91)                                           |
+-----------------------------------------------------------------------------------------+

Constructing and Interpreting Confidence Intervals (Score Bands):

Educators should never view a student's test score as an absolute point, but rather as a confidence band:

  • 68% Confidence Interval: Observed Score ± (1 × SEM) (There is a 68% probability that the student's True Score lies within this interval).
  • 95% Confidence Interval: Observed Score ± (2 × SEM) (There is a 95% probability that the student's True Score lies within this interval).

Classroom Application: If a student earns a score of 80 on a benchmark with an SEM of 4, the 68% confidence band is 76–84. A small change in score from 78 to 82 on a re-test falls within normal measurement error and does not necessarily indicate meaningful academic decline or growth.


5. Test Bias, Construct-Irrelevant Variance, and Equity

Two major threats systematically distort assessment validity and harm educational equity:

+-----------------------------------------------------------------------------------------+
|                         THREATS TO VALIDITY & MEASUREMENT EQUITY                        |
|                                                                                         |
|   [CONSTRUCT-IRRELEVANT VARIANCE]                                                       |
|   • Test measures extraneous factors unrelated to the target standard.                  |
|   • Example: Using complex, archaic English idioms in a 5th-grade science test,         |
|     penalizing English Language Learners (ELLs) on science knowledge.                   |
|                                                                                         |
|   [CONSTRUCT UNDERREPRESENTATION]                                                       |
|   • Assessment is too narrow; fails to sample critical dimensions of the standard.      |
|   • Example: Assessing a standard on 'conducting scientific investigations' solely       |
|     through a 10-item multiple-choice matching test.                                    |
|                                                                                         |
|   [CULTURAL / LINGUISTIC BIAS]                                                          |
|   • Items assume background experiential knowledge specific to a dominant culture.      |
|   • Example: Framing math word problems around polo rules or skiing vacations.          |
+-----------------------------------------------------------------------------------------+

Strategies to Mitigate Bias and Ensure Fair Measurement:

  1. Universal Design for Assessment (UDA): Ensure font legibility, uncluttered white space, clear graphics, and plain-language directions.
  2. Linguistic Simplification: Eliminate non-essential vocabulary, idioms, and double negatives from math and science word problem stems without lowering the cognitive mathematical demand.
  3. Accommodation Fidelity: Provide statutory testing accommodations (extended time, audio amplification, bilingual glossaries) specified in Individualized Education Programs (IEPs) and ELL plans without modifying the underlying test construct.
Loading diagram...
Psychometric Measurement Framework: CTT, Reliability, Validity, and SEM
Test Your Knowledge

A physics teacher constructs a 50-item final exam. To verify that the exam yields consistent results, the teacher administers it to the same group of students on two consecutive Mondays without intervening instruction. The resulting test-retest correlation coefficient is r = 0.94. However, a curriculum review reveals that 40 of the 50 items assessed rote historical trivia about famous scientists rather than the targeted state physics standards. How should this assessment be evaluated in terms of reliability and validity?

A
B
C
D
Test Your Knowledge

A student receives a score of 78 on a standardized reading assessment that has a standard deviation of 8 and a published reliability coefficient that results in a Standard Error of Measurement (SEM) of 3. When communicating these results during a parent conference, which interpretation most accurately reflects psychometric principles?

A
B
C
D
Test Your Knowledge

A middle school social studies department is designing a common end-of-year summative examination. Which procedure should the department follow first to ensure high content validity for the test?

A
B
C
D