12.2 Validity, Reliability, Practicality & Testing Washback
Key Takeaways
Construct validity is the central psychometric quality in language testing, threatened by construct under-representation (omitting vital skills) and construct-irrelevant variance (contamination by extraneous factors).
Assessment reliability requires minimizing measurement error through standardized test conditions, objective rubric descriptors, and rigorous inter-rater and intra-rater calibration.
The Bachman & Palmer Test Usefulness Framework models assessment design as an ongoing optimization among construct validity, reliability, authenticity, interactiveness, impact, and practicality.
Testing washback (backwash) refers to the systemic influence of tests on classroom instruction; positive washback enriches communicative curricula, while negative washback narrows teaching to test cramming.
Creating high-washback classroom tests involves using direct communicative tasks, criterion transparency, detailed diagnostic feedback, and authentic real-world language tasks.
12.2 Validity, Reliability, Practicality & Testing Washback
Note
In language assessment, test scores have no intrinsic meaning in isolation; their legitimacy rests upon psychometric qualities ensuring fairness, consistency, and pedagogical value. Educators must balance validity and reliability with practical constraints while maximizing positive testing washback.
The Assessment Triad: Validity, Reliability, Practicality
Evaluating any language assessment requires balancing three foundational psychometric pillars:
- Validity: The degree to which empirical evidence and theoretical rationales support the adequacy of inferences drawn from test scores (Messick, 1989): Does the test measure what it claims to measure?
- Reliability: The consistency and dependability of scores across administrations, test forms, and raters: Does the test yield stable, replicable results?
- Practicality: The operational feasibility of designing, administering, and scoring an assessment within constraints of time, budget, and staffing.
Lyle Bachman and Adrian Palmer (1996) synthesized these qualities into their Test Usefulness Framework:
These attributes exist in constant dynamic tension: standardized multiple-choice tests maximize reliability and administrative practicality but often sacrifice construct validity and task authenticity, whereas direct communicative tasks achieve high validity but demand extensive scoring time.
Exploring Dimensions of Assessment Validity
Validity pertains to score inferences rather than the test itself. Key facets include:
- Construct Validity: Fidelity to the underlying theoretical language ability being measured (e.g., communicative competence). It faces two major threats:
- Construct Under-representation: The test omits vital aspects of the construct (e.g., assessing oral conversational ability solely through written multiple-choice grammar items).
- Construct-Irrelevant Variance: Scores are distorted by extraneous factors unrelated to language proficiency (e.g., assessing reading comprehension through passages requiring complex mathematical calculations or obscure cultural trivia).
- Content Validity: Representative sampling of the instructional domain based on a structured test blueprint aligned with course learning objectives.
- Criterion-Related Validity: Statistical relationship to external criteria, encompassing concurrent validity (correlation with an established current test) and predictive validity (forecasting future academic or workplace performance).
- Face Validity: Perceived fairness and relevance to non-expert stakeholders (students, parents, employers), strongly influencing test-taker motivation.
Validity Types and Threats in Language Assessment
| Validity Type | Focus / Definition | Testing Threats | Mitigation Strategy |
|---|---|---|---|
| Construct Validity | Theoretical fidelity to target construct | Construct under-representation; irrelevant cognitive load | Define clear construct boundaries; use direct performance tasks |
| Content Validity | Representative sampling of curriculum | Over-sampling favorite topics; omitting productive skills | Construct an explicit test blueprint mapping items to syllabus goals |
| Concurrent Validity | Correlation with established current measure | Inconsistencies arising from differing test formats | Pilot classroom instruments against validated proficiency scales |
| Predictive Validity | Accurate forecasting of future performance | Discrete-point tests with no communicative transfer | Embed scenario-based tasks mirroring real-world communication |
| Face Validity | Perceived fairness to stakeholders | Abstract puzzle items, trick questions, obscure formatting | Utilize authentic text genres, clear prompts, transparent directions |
Understanding Dimensions of Assessment Reliability
Reliability addresses measurement error and score consistency across testing conditions:
- Test-Retest Reliability: Score stability across repeated administrations under comparable conditions.
- Parallel-Forms Reliability: Equivalence between two distinct test versions constructed to identical specifications.
- Internal Consistency: Item homogeneity within a test, measured by split-half methods or Cronbach's alpha ().
- Inter-Rater Reliability: Scoring consistency between independent raters evaluating subjective performances (interviews, essays).
- Intra-Rater Reliability: The internal consistency of a single evaluator applying identical standards across papers over time.
Reliability Dimensions and Threats Matrix
| Reliability Dimension | Core Question | Primary Threats / Error Sources | Remediation Strategy |
|---|---|---|---|
| Test-Retest | Are scores stable across time? | Learner fatigue, illness, practice effects | Standardize environment; ensure adequate recovery intervals |
| Parallel-Forms | Are alternate forms equivalent? | Discrepant passage difficulty, unequal vocabulary frequency | Calibrate forms via item analysis and readability metrics |
| Internal Consistency | Do all items measure one construct? | Ambiguous distractors, poor item discrimination | Conduct item facility and point-biserial discrimination analyses |
| Inter-Rater | Do different scorers agree? | Subjective bias, halo effect, diverging criteria interpretations | Conduct calibration workshops; use benchmark exemplars and double scoring |
| Intra-Rater | Does one scorer maintain standards? | Scorer fatigue, shifting expectations, grading order effects | Employ detailed analytic rubrics; score question-by-question |
Testing Practicality: Operational Feasibility
Practicality evaluates whether an assessment can be realistically deployed within educational constraints:
- Financial Cost: Expenditures for proprietary testing materials, software licenses, or external scorers.
- Time Constraints: Hours required to write, administer, score, and return assessments with meaningful feedback.
- Administrative Feasibility: Physical facility space, recording equipment, staffing, and invigilation demands.
- Interpretative Ease: The speed and clarity with which educators translate scores into instructional next steps.
Testing Washback / Backwash: Systemic Impact on Teaching
Washback (or backwash) denotes the systemic impact testing exerts on curriculum and classroom teaching. High-stakes tests inevitably drive instruction because educators and learners prioritize what is evaluated:
- Negative Washback: Occurs when test formats distort instruction into drill-and-kill test preparation, rote grammar memorization, and isolated vocabulary lists, marginalizing authentic oral communication.
- Positive Washback: Occurs when assessment drives sound communicative practices. When assessments demand authentic spoken interaction, listening synthesis, and persuasive writing, classroom activities naturally center on discussion, audio analysis, and writing workshops.
Washback Optimization Strategies Guide
| Assessment Design Lever | High-Washback Implementation | Negative Washback Trap |
|---|---|---|
| Task Authenticity | Use communicative tasks (interviews, emails, presentations) mirroring real life | Over-relying on discrete multiple-choice items that invite guessing |
| Curricular Alignment | Derive test tasks directly from communicative syllabus learning outcomes | Allowing commercial test-prep books to dictate daily classroom lessons |
| Criterion Transparency | Share descriptive scoring rubrics and exemplar models prior to testing | Keeping grading criteria opaque, fostering student anxiety and speculation |
| Feedback Richness | Return assessments with diagnostic analytic commentary and error logs | Returning only raw percentage scores without actionable qualitative advice |
| Learner Agency | Incorporate self-assessment and peer review into the evaluation cycle | Positioning assessment as punitive teacher-centered surveillance |
Important
The most powerful mechanism for generating positive washback is authentic task alignment. When assessment formats mirror communicative curricular objectives, preparing for the test becomes identical to developing genuine language proficiency.
A teacher designs an achievement test intended to measure intermediate ESL reading comprehension. However, the reading passage features dense mathematical word problems requiring advanced algebraic calculations to answer the comprehension questions. Several advanced English learners fail the test because of mathematical confusion rather than linguistic comprehension breakdowns. In psychometrics, this test flaw exemplifies:
Construct-irrelevant variance that undermines score validity
Negative washback resulting from subjective rater scoring bias
Low test-retest reliability caused by environmental testing distractions
Construct under-representation caused by omitting vocabulary items
A university ESL department notices that two different instructors grading the same student writing portfolio award radically different scores (one gives an A- while the other gives a C+). To establish consistent scoring across the faculty, the department organizes a calibration workshop where instructors review benchmark student essays, discuss scoring criteria, and practice until their grading achieves statistical consensus. Which psychometric quality is the department striving to improve?
Parallel-forms reliability through algorithmic item difficulty balancing
Predictive validity through longitudinal correlation with grade point average
Practicality through streamlining scoring administrative procedures
Inter-rater reliability through rater standardization and rubric calibration
A national education ministry reforms its secondary English exit exam, replacing a 100-item discrete-point multiple-choice grammar and vocabulary test with an integrated communicative assessment requiring a paired conversational interview, an authentic listening synthesis, and a persuasive argumentative essay. Over the subsequent academic year, classroom teachers replace fill-in-the-blank drill sheets with interactive group discussions, audio analysis projects, and structured peer-writing workshops. This transformation in instructional practice demonstrates:
Positive washback resulting from authentic communicative task alignment
Negative washback resulting in construct under-representation
Increased test practicality resulting from simplified testing logistics
High intra-rater reliability resulting from standardized scoring machines
Sections you finish are checked off in the contents.