5.2 Technical Assessment Principles: Validity, Reliability, and Fairness
Key Takeaways
- Construct validity ensures an assessment accurately measures the targeted theoretical construct (e.g., reading comprehension or math reasoning) without distortion from unintended variables.
- Reliability measures test consistency across time (test-retest), raters (inter-rater), and internal test items (internal consistency/Cronbach's alpha).
- Measurement bias occurs when test items systematically penalize ELLs due to cultural ethnocentrism, complex non-target language, or inappropriate standardization norming samples.
- Construct-irrelevant variance arises when extraneous factors—such as dense English reading load on a science test—interfere with measuring true target knowledge.
- High-stakes testing produces washback (or backwash); positive washback aligns curriculum with communicative goals, whereas negative washback narrows instruction to repetitive test-drilling.
5.2 Technical Assessment Principles: Validity, Reliability, and Fairness
For an assessment to yield meaningful and equitable inferences regarding an English Language Learner’s knowledge or abilities, it must satisfy rigorous technical standards. Psychometric evaluation rests upon three foundational pillars: validity, reliability, and fairness. When testing multilingual learners, educators must vigilantly guard against construct-irrelevant variance and measurement bias, while understanding how assessment policies exert washback on classroom instruction.
1. The Validity Triad
Validity refers to the degree to which empirical evidence and theoretical rationales support the adequacy and appropriateness of inferences and actions based on test scores. Validity is not a static property of a test itself, but of the interpretations derived from test results in a specific context.
┌── Construct Validity (Measures theoretical target construct)
│
Validity Evidence Matrix ─┼── Content Validity (Aligns with domain standards/blueprint)
│
└── Criterion Validity ┬── Concurrent Validity (Correlates with existing measure)
└── Predictive Validity (Predicts future performance)
Construct Validity
Construct validity is the overarching psychometric requirement that an assessment accurately measures the specific theoretical construct it claims to measure (e.g., mathematical problem-solving, phonological awareness, or academic reading comprehension). If a math word problem requires an advanced level of English syntactic parsing that obscures the mathematical calculation, the assessment suffers from compromised construct validity.
Content Validity
Content validity assesses how comprehensively and accurately test items represent the targeted academic or linguistic domain. In ESOL education, content validity requires full alignment between assessment items and established standards, such as state ELP standards (e.g., WIDA ELD Standards) or grade-level academic content standards. A content-valid test covers the full depth and breadth of the domain specifications without omitting critical sub-skills.
Criterion-Related Validity
Criterion validity measures how strongly test performance correlates with an external benchmark or outcome:
- Concurrent Validity: Demonstrating that scores on a newly developed ELP instrument correlate highly with an established, validated test administered at the same time.
- Predictive Validity: The degree to which an assessment score accurately forecasts future performance (e.g., predicting whether an ELL's score on a end-of-year reading exam predicts academic success in un-scaffolded general education courses).
2. Reliability Principles
Reliability refers to the consistency, stability, and replicability of assessment results across repeated administrations, different test forms, or different raters. A test cannot be valid if it is not reliable; however, a reliable test can still be invalid if it consistently measures the wrong construct.
| Reliability Type | Method of Determination | Critical Consideration for ELLs |
|---|---|---|
| Test-Retest Reliability | Administering the same instrument to the same students at two points in time. | Language acquisition occurs rapidly; real proficiency gains must not be confused with measurement instability. |
| Inter-Rater Reliability | Degree of agreement between independent raters scoring the same subjective performance. | Essential when scoring speaking and writing rubrics; raters require rigorous calibration to avoid bias against accented speech or non-native syntactic structures. |
| Internal Consistency | Evaluating item correlation within a single test (e.g., split-half, Kuder-Richardson, Cronbach's $\alpha$). | Items must consistently sample the target construct without introducing secondary linguistic distractors. |
Standard Error of Measurement (SEM)
Every test score contains an inherent degree of measurement error ($X = T + E$, where Observed Score = True Score + Error). The Standard Error of Measurement (SEM) quantifies the expected variation in a student's score if retested. Educators must interpret ELL test scores as score ranges (confidence intervals) rather than absolute, flawless numbers.
3. Measurement Bias and Norming Sample Flaws
Measurement bias occurs when test items contain extraneous factors that systematically misrepresent the true abilities of specific subgroups of test-takers, such as multilingual learners.
Dimensions of Assessment Bias
- Cultural Bias: Items that presuppose familiarity with specific cultural background knowledge, mainstream American traditions, regional idioms, or specialized pop culture (e.g., test items referencing baseball scoring rules, suburban farming equipment, or specific holiday customs). ELLs may fail the item due to lack of cultural exposure rather than lack of cognitive or analytical skill.
- Linguistic Bias: Using overly complex, non-essential syntactic structures (e.g., double negatives, dense passive voice, unnecessary idioms) in items intended to test non-linguistic content (e.g., science hypotheses or math operations).
- Norming Sample Bias: Standardized norm-referenced tests compare individual student scores against a representative norming sample. If the norming population consisted exclusively of native English speakers, applying those norm tables to ELLs is psychometrically invalid. ELL performance is evaluated against a distribution that fails to account for second language acquisition trajectories.
4. Construct-Irrelevant Variance
Construct-irrelevant variance is the primary threat to assessment fairness for ELLs. It refers to systemic error introduced into test scores by factors that are irrelevant to the construct being evaluated.
When an ELL takes a science examination written in dense academic text with low-frequency vocabulary, the test measures two things simultaneously: science content knowledge (the target construct) and English reading proficiency (construct-irrelevant variance). As a result, the score reflects language difficulty rather than scientific mastery, rendering the test invalid for its intended purpose.
5. Washback (Backwash) Effects of High-Stakes Testing
Washback (also called backwash) describes the direct and indirect impact of high-stakes testing on curriculum design, instructional practices, and educational priorities.
Positive vs. Negative Washback
- Positive Washback: Occurs when high-stakes assessments align with authentic communicative language goals. For example, if a state ELP exam requires authentic speaking tasks and communicative writing, teachers prioritize real-world conversation, collaborative projects, and functional writing in daily instruction.
- Negative Washback: Occurs when high-stakes tests rely on isolated multiple-choice grammar items or decontextualized reading passages. Teachers feel pressured to 'teach to the test,' narrowing the curriculum to mechanical test-taking strategies, rote grammar drills, and memorization, thereby depriving ELLs of rich, communicative language acquisition opportunities.
A 6th-grade mathematics assessment contains a word problem requiring students to calculate speed, but the problem is framed within an elaborate 150-word narrative filled with obscure nautical terminology and complex passive voice syntax. An ELL fails the item despite having mastered the underlying math formula. Which psychometric flaw is demonstrated?
A standardized reading test is developed and normed using a sample composed entirely of native English speakers from affluent suburban school districts. When this test is administered to newly arrived immigrant ELLs, what is the primary technical concern?
Three independent raters score an ELL student's oral language presentation using a holistic rubric. Rater A assigns a score of 90%, Rater B assigns 55%, and Rater C assigns 70%. What technical psychometric property is severely lacking in this assessment scenario?
Because a state's high-stakes annual ELL assessment consists entirely of decontextualized multiple-choice grammar questions, local ESOL teachers spend eight weeks drilling isolated grammar rules and bubble-sheet strategies rather than engaging students in meaningful communicative dialogue. This phenomenon is an example of: