Formative & Summative Assessment Principles
Key Takeaways
- Formative assessments provide continuous, low-stakes feedback to adjust instruction, while summative assessments evaluate overall learning at unit completion.
- Criterion-referenced assessments measure mastery against specific standards (TEKS), whereas norm-referenced tests compare students against peer groups.
- Cultural test bias occurs when test items assume background knowledge specific to the dominant culture.
- Linguistic test bias occurs when unnecessarily complex language turns a content assessment into a de facto reading test.
- Authentic assessments (performance tasks, portfolios) offer valid, non-discriminatory alternatives for measuring EB academic progress.
Formative & Summative Assessment Principles
Assessing Emergent Bilinguals (EBs) requires a rigorous understanding of educational measurement principles. Educators must distinguish between a student's underlying academic content knowledge and their developing English language proficiency. Standardized assessments engineered for native English speakers often fail to yield valid scores for EBs because language barriers introduce measurement errors. To conduct fair, non-discriminatory evaluations, teachers must apply valid assessment principles, recognize cultural and linguistic test bias, and utilize authentic, linguistically accommodated evaluation tools.
Formative vs. Summative Assessment
- Formative Assessment (Assessment FOR Learning): Continuous, informal, low-stakes checks for understanding integrated seamlessly into daily instruction. Formative assessments provide immediate feedback to both teacher and student, allowing real-time instructional adjustments. Examples include exit tickets with sentence stems, Think-Pair-Share observations, digital response polls, mini-whiteboards, and thumbs-up checks.
- Summative Assessment (Assessment OF Learning): Formal, high-stakes evaluations administered at the conclusion of an instructional unit, semester, or academic year to measure cumulative mastery. Examples include chapter exams, benchmark assessments, STAAR state criterion-referenced exams, and annual TELPAS evaluations.
Criterion-Referenced vs. Norm-Referenced Tests
- Criterion-Referenced Tests: Measure student performance against a predetermined, absolute set of academic standards or learning objectives (such as the Texas Essential Knowledge and Skills - TEKS). Student scores indicate mastery of specific skills regardless of peer performance. STAAR exams are criterion-referenced.
- Norm-Referenced Tests: Compare an individual student's performance against a national sample norming group, reporting results in percentiles or grade-equivalent scores (e.g., Iowa Assessments, SAT). Caution for EBs: Norm-referenced tests are inherently biased against EBs if the original norming sample population did not include a representative proportion of Emergent Bilinguals at similar English proficiency levels.
Assessment Validity, Reliability, and Construct Validity
- Validity: The extent to which an assessment accurately measures the specific construct it claims to measure. For EBs, construct-irrelevant variance poses the single greatest threat to validity. When a mathematics test features dense, complex English paragraph prompts, it no longer measures mathematical competence; instead, it inadvertently measures English reading comprehension.
- Reliability: The degree to which an assessment produces consistent, stable results across multiple administrations under similar conditions.
Identifying and Mitigating Test Bias
Test bias occurs when test items contain systemic features that result in lower scores for a specific group of students for reasons unrelated to the academic construct being evaluated.
1. Cultural Bias
Cultural bias occurs when test items assume background experiences, socio-cultural norms, or prior knowledge unique to the dominant culture.
- Biased Test Item: A word problem requiring students to calculate bowling scores or evaluate financial rules for a suburban country club. An EB may fail the item not because they lack math skills, but because they are unfamiliar with bowling scoring or country club norms.
- Mitigation: Ensure test items utilize culturally neutral contexts (e.g., calculating distance traveled or measuring school garden plots).
2. Linguistic Bias
Linguistic bias occurs when assessment items incorporate unnecessarily complex Tier 2 academic vocabulary, double negatives, obscure idioms, or complex passive voice structures that are not essential to the content being tested.
- Biased Item: "Which catalyst precipitated an uncharacteristic surge in agrarian production?"
- Accommodated Item: "Which factor caused a rapid increase in crop production?"
3. Format Bias
Format bias occurs when test layouts, answer sheets, or digital testing interfaces are confusing or unfamiliar to EBs, causing errors due to procedural confusion rather than content knowledge deficits.
Authentic Assessment Modalities for EBs
To obtain valid data regarding an EB's true academic capabilities, educators employ authentic assessment alternatives:
Performance-Based Assessments
Students demonstrate mastery by performing an authentic, hands-on task (e.g., conducting a science experiment, constructing a geometric model, or delivering a scaffolded oral presentation). Performance tasks evaluate applied knowledge while reducing reliance on dense written text.
Portfolio Assessment
A structured, longitudinal collection of authentic student work gathered over time across content areas. Portfolios showcase growth in listening, speaking, reading, and writing, providing a comprehensive profile of student progress that single snapshot standardized tests cannot capture.
Observational Rubrics (e.g., SOLOM)
Teachers utilize standardized observational instruments such as the Student Oral Language Observation Matrix (SOLOM) to rate EB oral language development across five domains: comprehension, fluency, vocabulary, pronunciation, and grammar during authentic classroom interactions.
| Assessment Dimension | Standardized Paper-Pencil Test | Authentic / Accommodated Assessment |
|---|---|---|
| Primary Construct Measured | Content knowledge + English reading ability | Pure content knowledge via visual/hands-on tasks |
| Linguistic Load | High; unedited dense academic text | Scaffolding applied (sentence stems, visual aids) |
| Cultural Context | Often dominant-culture centric | Culturally responsive and context-neutral |
| Actionable Feedback | Delayed summative score report | Immediate formative feedback for instructional adjustment |
| Assessment Type | Core Purpose | Example for EBs |
|---|---|---|
| Formative | Adjust daily instruction | Exit ticket with sentence frame & mini-whiteboards |
| Summative | Evaluate unit mastery | Criterion-referenced TEKS unit exam with designated supports |
| Performance-Based | Authentic demonstration of skills | Hands-on science lab demonstration with rubric |
| Portfolio | Longitudinal growth tracking | Dual-language writing collection evaluated over time |
A middle school math test problem asks students to calculate average scores from a baseball statistics table. Recently arrived EBs struggle with the problem. This assessment item is problematic primarily due to:
Which of the following represents a formative assessment designed to guide daily classroom instruction for EBs?