15.3 Formative vs. Summative Assessment, Rubrics, STAAR Context & Data-Driven Differentiation
Key Takeaways
Diagnostic assessments identify prerequisite readiness, formative assessments inform real-time instructional adjustments, and summative assessments measure cumulative standards mastery.
Criterion-referenced assessments measure student performance against predefined objective standards (such as TEKS), unlike norm-referenced tests that rank students against a normative cohort.
Analytic rubrics evaluate distinct dimensions of mathematical performance—conceptual understanding, procedural execution, reasoning, and representation—providing actionable diagnostic feedback.
Texas STAAR grades 3–8 mathematics tests (scheduled to be replaced by through-year tests in 2027–28) use varied item types and four performance levels: Did Not Meet, Approaches, Meets, and Masters Grade Level.
Data-driven differentiation leverages tiered tasks, flexible grouping, and systematic Multi-Tiered System of Supports (MTSS/RTI) to remediate misconceptions while maintaining high cognitive demand for all students.
15.3 Formative vs. Summative Assessment, Rubrics, STAAR Context & Data-Driven Differentiation
Assessment in middle-school mathematics is an ongoing, dynamic process integrated into every phase of learning rather than an isolated terminal event. The National Council of Teachers of Mathematics (NCTM) Assessment Principle asserts that assessment should support the learning of mathematics and furnish useful information to both teachers and students. To realize this goal, educators must deploy a comprehensive assessment system that balances diagnostic, formative, and summative instruments, utilize objective criterion-referenced rubrics, interpret state standardized testing data (STAAR), and translate assessment evidence into differentiated, equitable classroom interventions.
The Tripartite Assessment Architecture: Diagnostic, Formative & Summative
[Diagnostic Assessment] → [Formative Assessment Loop] → [Summative Evaluation]
(Pre-Unit Readiness) (Daily Checks & Hinge Items) (Unit Mastery & Projects)
↓ ↓ ↓
Identify prerequisite gaps Real-time instruction tweaks Document terminal achievement
1. Diagnostic / Pre-Assessments
Administered prior to introducing a new unit or concept, diagnostic assessments determine students' baseline knowledge, prerequisite operational skills, and pre-existing misconceptions.
- Purpose: Identifies specific learning gaps without assigning punitive evaluative grades (e.g., assessing whether incoming 6th graders possess fraction equivalence concepts before introducing fraction division).
- Classroom Tools: Brief 3–5 question readiness screeners, concept maps, or diagnostic error-interview prompts. The data informs unit pacing and the design of prerequisite reviews.
2. Formative Assessment ("Assessment for Learning")
Formative assessment is an ongoing, low-stakes diagnostic process conducted during instruction to monitor student learning and provide immediate, actionable feedback to both students and teachers.
- Purpose: Informs instructional adjustments in real time. If 40% of students exhibit an error on a mid-lesson check, the teacher halts individual work to deliver targeted small-group re-teaching or whole-class clarification.
- Key Formative Strategies:
- Exit Tickets: Concise, 1–2 question prompts completed at the close of class targeting the core objective of the day (e.g., "Write an equation for a line with slope -3 passing through (0, 4)").
- Hinge Questions: A multiple-choice question posed at a critical conceptual juncture in the lesson. Every student responds simultaneously (via individual dry-erase whiteboards or digital polling). The question is designed so that each incorrect option reflects a specific conceptual misconception.
- Error Analysis Tasks: Students are presented with a solved mathematical problem containing an intentional, realistic student error. Students must identify the error, explain the flawed thinking, and provide the correct mathematical solution.
- Four Corners & Think-Pair-Share: Kinesthetic or collaborative checks where students select and defend mathematical claims.
3. Summative Assessment ("Assessment of Learning")
Summative assessment evaluates cumulative student learning, skill acquisition, and academic achievement at the conclusion of a defined instructional period.
- Purpose: Documents student performance, determines grade placement, and verifies mastery of grade-level TEKS standards.
- Instruments: End-of-unit chapter exams, cumulative semester exams, comprehensive performance tasks, and state standardized assessments.
Criterion-Referenced vs. Norm-Referenced Assessment
Understanding the distinction between criterion-referenced and norm-referenced measurement is essential for interpreting test results and educator certification standards.
Criterion-Referenced Tests (CRTs)
- Measurement Basis: Evaluates an individual student's performance against a fixed, predefined standard or criterion of mastery (e.g., mastering 80% of grade 7 proportional reasoning TEKS standards).
- Independence: A student's score is independent of other examinees' scores. Theoretically, 100% of students can achieve mastery if all meet the standard.
- Examples: Texas STAAR assessments, TExES Educator Certification exams (passing score of 240), driver's license exams, and teacher-created unit tests.
- Diagnostic Utility: High; allows teachers to report exactly which mathematical standards a student has mastered or not mastered.
Norm-Referenced Tests (NRTs)
- Measurement Basis: Compares and ranks an individual student's performance against the performance of a statistically representative normative cohort of peers who took the identical test.
- Distribution: Scores are intentionally distributed along a normal bell curve. By design, a predetermined percentage of examinees will score above and below the median.
- Metrics Reported: Percentile ranks (e.g., 78th percentile), stanines (1–9), and standard scores (e.g., NCEs).
- Examples: Iowa Assessments (ITBS), Northwest Evaluation Association Measures of Academic Progress (NWEA MAP), SAT, and ACT.
- Diagnostic Utility: Low for specific standard mastery; indicates relative standing rather than what specific mathematical skills a student can execute.
Evaluating the Quality of an Assessment
Competency 019 expects teachers to judge assessment methods and materials for validity, reliability, absence of subjectivity (bias), clarity of language and appropriateness of mathematical level. It also expects assessments to be consistent with what is taught and how it is taught.
| Quality | Question to ask | Example of a problem |
|---|---|---|
| Validity | Does the task measure the intended mathematics? | A proportional-reasoning item buried in a long, unfamiliar story ends up measuring reading, not ratios |
| Reliability | Would the results be consistent across scorers, forms or occasions? | Two teachers score the same open-response item very differently because there is no rubric |
| Objectivity / freedom from bias | Could background knowledge unrelated to the mathematics advantage some students? | A probability item assumes every student knows the rules of a particular card game |
| Clarity of language | Is the wording as simple as the mathematics allows? | Double negatives, idioms or vague terms confuse emergent bilingual students |
| Appropriateness of level | Does the item match the grade-level expectations and the numbers students have learned? | A grade 6 unit-rate check that requires operations with irrational numbers |
| Consistency with instruction | Does the format and content match how the topic was taught? | Students explored slope with tables and graphs, but the test asks only for the formula |
Improving reliability: Use analytic rubrics with anchor papers, score student work blind to names, and have two scorers calibrate on a sample. Improving validity: Write items directly from the learning goal, and check that every distractor reflects a real misconception, not a trick.
Assessment quality also matters for fairness. If an item is unclear or culturally loaded, a low score may say more about the item than about the student's mathematics.
Designing Scoring Rubrics for Middle School Mathematics
Rubrics provide objective, transparent evaluative criteria for scoring complex mathematical investigations, open-ended problem solving, and performance tasks.
Holistic Rubrics
A holistic rubric assigns a single, global score (e.g., 4, 3, 2, 1) based on an overall impression of the student's complete work.
- Advantages: Quick to score; provides a macro-level assessment of general competence.
- Disadvantages: Lacks diagnostic granularity. A student receiving a score of "2" does not know whether their deficiency was computational inaccuracy, flawed conceptual strategy, or poor communication.
Analytic Rubrics
An analytic rubric deconstructs mathematical performance into distinct, independent evaluative dimensions, scoring each trait on an explicit developmental continuum.
| Evaluative Dimension | Exemplary / Advanced (4) | Proficient (3) | Developing (2) | Novice / Beginning (1) |
|---|---|---|---|---|
| 1. Conceptual Understanding | Demonstrates complete understanding of underlying mathematical concepts; selects optimal mathematical models. | Demonstrates understanding of major concepts; minor conceptual gaps do not impede solution. | Exhibits partial understanding; misapplies key concepts or chooses inefficient models. | Demonstrates severe conceptual misunderstandings; unable to model the problem. |
| 2. Procedural Accuracy | All calculations, algebraic manipulations, and algorithms executed with complete precision. | Minor computational or transcription slip; procedural execution is fundamentally sound. | Multiple arithmetic or procedural errors; incomplete algorithm execution. | Pervasive computational errors; incorrect algorithms applied throughout. |
| 3. Communication & Representation | Translates fluidly across symbols, graphs, tables, and verbal text; uses precise mathematical terminology. | Clear representations with minor labeling errors; explanations are understandable and coherent. | Incomplete representations; missing labels, ambiguous graphs, or informal colloquial language. | Lacks mathematical representations; work is disorganized, illegible, or unexplained. |
| 4. Justification & Reasoning | Provides rigorous, deductive justification; verifies reasonableness of results; evaluates alternatives. | Valid reasoning provided; minor omissions in logical chain of argument. | Incomplete justification; asserts conclusions without mathematical backing. | No mathematical justification provided; offers unsubstantiated guesses. |
The Texas Standardized Testing Context: STAAR Grades 4–8 Mathematics
The State of Texas Assessments of Academic Readiness (STAAR) program evaluates public school students' mastery of the TEKS curriculum standards. Middle-school mathematics educators must understand STAAR's structure, item formats, and performance classifications.
Redesigned Item Types
Since the STAAR redesign (first administered in spring 2023), TEA caps multiple-choice questions at 75% of each test and uses new question types in mathematics. The old paper "griddable" items were replaced by typed entry:
- Multiple choice: Traditional selected-response items.
- Equation editor: Students type numbers, fractions, expressions, equations, or inequalities.
- Multiselect: Students select every correct answer from a list.
- Inline choice: Students choose terms or values from drop-down menus inside a sentence.
- Drag and drop, hot spot, and match table grid: Students place, click, or classify items.
- Graphing, number line, and fraction model: Students plot points and lines, mark solution sets, or shade models.
Transition ahead: Texas House Bill 8 (signed in September 2025) replaces the grades 3–8 STAAR tests with three shorter beginning-, middle-, and end-of-year assessments starting in the 2027–28 school year. Teachers should expect STAAR through 2026–27 and then the new through-year system.
STAAR Performance Level Descriptors
Student performance on STAAR is classified into four distinct performance tiers:
- Did Not Meet Grade Level: Student performance falls below passing standards; students exhibit a lack of sufficient understanding of grade-level TEKS and require significant, ongoing academic intervention.
- Approaches Grade Level: The baseline passing standard. Performance indicates that students are likely to succeed in the next grade or course with targeted academic intervention; students recognize and apply mathematical concepts in familiar contexts.
- Meets Grade Level: Student performance indicates a high likelihood of academic success in the subsequent grade or course without ongoing intervention. Students demonstrate critical-thinking ability and can apply mathematical concepts in familiar and unfamiliar contexts.
- Masters Grade Level: The highest achievement tier. Performance indicates deep, comprehensive mastery of grade-level TEKS; students demonstrate advanced critical thinking, can apply mathematical concepts in novel, complex contexts, and are well prepared for advanced postsecondary coursework.
Data-Driven Differentiation: Tiered Tasks & MTSS / RTI
Differentiating mathematics instruction means adjusting content, process, or products to meet individual student readiness, interests, and learning profiles while maintaining high cognitive demand for all students.
1. Tiered Assignments
Tiered instruction addresses the same essential TEKS standard and big idea for all students, but varies the level of complexity, abstractness, and instructional scaffolding.
- Tier 1 (Foundation / Support): For students needing prerequisite support. Focuses on the core concept using concrete manipulatives, visual graphic organizers, or smaller integer values, while maintaining the same mathematical reasoning requirements.
- Tier 2 (Core / Grade-Level): Focuses on standard grade-level numbers, multiple representations (tables, graphs, equations), and standard application problems.
- Tier 3 (Extension / Acceleration): For advanced learners. Involves open-ended investigations, non-routine problem solving, formulating algebraic generalizations, or analyzing non-linear extensions, avoiding mere "more of the same" busywork.
2. Multi-Tiered System of Supports (MTSS) / Response to Intervention (RTI)
MTSS provides a structured, multi-layered continuum of evidence-based instructional support:
- Tier 1: High-Quality Universal Core Instruction (100% of students): Effective, research-based TEKS instruction delivered in the general education classroom, incorporating universal screening, differentiated grouping, and ongoing formative assessment.
- Tier 2: Targeted Small-Group Intervention (10–15% of students): Supplemental small-group instruction (3–5 students, 20–30 minutes, 2–3 times per week) addressing specific identified prerequisite gaps (e.g., integer subtraction or ratio reasoning) using targeted visual models and progress monitoring.
- Tier 3: Intensive Individualized Intervention (1–5% of students): Highly specialized, intensive intervention (individual or 1–2 students, daily) targeting severe computational or conceptual deficits, utilizing diagnostic assessments and frequent progress monitoring.
3. Accommodating Diverse Learners
- Accommodations vs. Modifications: An accommodation alters how a student accesses information or demonstrates mastery without lowering or changing the learning standard (e.g., extra time, oral reading of word problems, access to formula charts or basic calculation aids). A modification fundamentally alters what the student is expected to learn (e.g., lowering the cognitive complexity or eliminating TEKS standards); modifications are reserved strictly for students with significant cognitive disabilities documented in an Individualized Education Program (IEP).
- Supporting Emergent Bilinguals (English Learners): Mathematics is not language-free. Educators support English learners by utilizing multimodal visual representations, bilingual glossaries, sentence frames and stems ("The graph shows a positive slope because as x increases by __, y increases by __"), interactive word walls with pictorial definitions, and allowing students to discuss problems in their home languages before translating to academic English.
Comprehensive Comparison of Mathematics Assessment Types
| Assessment Type | Primary Educational Purpose | Administration Timing | Stakes & Grading Impact | Typical Middle School Classroom Example |
|---|---|---|---|---|
| Diagnostic | Identify prior knowledge, prerequisite skill gaps, and misconceptions. | Prior to instruction or unit launch. | Very low stakes; non-graded diagnostic feedback. | A 4-question screener on fraction division before starting unit rates in grade 6. |
| Formative | Monitor ongoing learning; provide immediate feedback; adjust instruction. | Continuously during daily instruction. | Low stakes; informal grading, completion, or rubric checks. | Exit ticket on solving two-step equations; individual whiteboard hinge question. |
| Summative | Evaluate cumulative mastery of standards and assign terminal scores. | End of unit, semester, or school year. | High stakes; formal grade recording and accountability. | Chapter 5 exam on linear functions; STAAR Grade 8 Mathematics Assessment. |
| Criterion-Referenced | Measure student performance against fixed, predefined curriculum standards (TEKS). | Periodic unit tests and annual state tests. | Varies from medium (classroom test) to high (state testing). | STAAR Mathematics test where passing requires achieving "Approaches Grade Level." |
| Norm-Referenced | Rank students relative to a national normative peer comparison group. | Annually or triennially. | Medium to high; used for program placement or honors screening. | NWEA MAP Growth test reporting student percentile rank (e.g., 84th percentile). |
Step-by-Step Worked Assessment Data Analysis & Intervention Scenario
Classroom Context: Grade 8 Mathematics Unit on Linear Functions
Following an interim formative assessment on graphing linear relationships (TEKS 8.4B–C and 8.5B, proportional and non-proportional relationships in the form ), the teacher analyzes student work from a 25-student class on the problem: "Graph the linear function and state the coordinates of its - and -intercepts."
Step 1: Error Categorization & Data Clustering
The teacher reviews the student work and identifies three distinct performance clusters:
- Cluster 1: Complete Mastery (11 students - 44%): Correctly plot , apply a negative rate of change (down 2 units, right 3 units) to find and , identify the -intercept as and -intercept as , and explain the negative slope.
- Cluster 2: Directional Slope Error (9 students - 36%): Correctly plot the -intercept , but interpret the slope as "up 2, right 3" (positive slope), plotting points and . They calculate the -intercept algebraically by setting , but their graph contradicts their algebraic work.
- Cluster 3: Intercept & Variable Inversion (5 students - 20%): Invert the slope and intercept, plotting the point on the -axis as their starting point, or treating as the slope and as the -intercept.
Step 2: Formulating Targeted Differentiated Interventions
Rather than re-teaching the entire lesson to the whole class (which would disengage Cluster 1 and fail to address the specific confusion of Cluster 3), the teacher plans differentiated small-group stations:
- Station 1 (Extension for Cluster 1): Students work on an extension-level tiered task: Investigating how changing the parameters and to variable values ( and ) produces perpendicular lines, using dynamic Desmos sliders to formulate conjectures regarding opposite reciprocal slopes.
- Station 2 (Targeted Re-Teaching for Cluster 2): The teacher leads a targeted small-group mini-lesson focusing on the directional sign of slope. Students use physical slope triangles on dry-erase grid boards, physically walking or tracing lines from left to right. The teacher uses revoicing: "If the slope is negative, what must happen to the y-value as we move forward in time?" Students verify their negative slopes by building a table of values (or using a graphing tool's table view).
- Station 3 (Intensive Prerequisite Remediation for Cluster 3): The teacher provides more intensive structural scaffolding using color-coded equation mats. The constant term is highlighted in green ("Starting Value / Vertical Intercept on the -axis "), while the coefficient is highlighted in yellow ("Rate of Change / Directional Movement"). Students construct equations and graphs using concrete pegboards before attempting pencil-and-paper graphing.
Step 3: Progress Monitoring (Re-Assessment)
At the end of the intervention session, all students complete a new exit ticket with . Cluster 2 and Cluster 3 students demonstrate 86% mastery on the re-check, verifying that targeted data-driven intervention successfully resolved the underlying conceptual misconceptions.
A middle school mathematics department is reviewing assessment data to evaluate student readiness for eighth-grade algebra. Which of the following correctly characterizes the fundamental difference between a criterion-referenced test (such as the STAAR Mathematics test) and a norm-referenced test (such as the TerraNova or Iowa Assessments)?
A criterion-referenced test measures a student's performance against an absolute standard of specific curriculum objectives (TEKS), whereas a norm-referenced test measures performance relative to a national cohort of peers.
A criterion-referenced test forces student scores into a normal bell-shaped distribution, whereas a norm-referenced test allows all students to achieve the highest possible score.
A criterion-referenced test is utilized exclusively for assigning letter grades, whereas a norm-referenced test is utilized exclusively for daily formative instructional adjustments.
A criterion-referenced test reports student mastery solely through national percentile rankings, whereas a norm-referenced test identifies specific standard deficiencies.
A teacher is evaluating students' open-ended performance on a complex multi-step task involving ratio, proportion, and financial modeling. To provide actionable diagnostic feedback that pinpoints whether a student's struggle stems from computational errors, conceptual misinterpretation, or poor written justification, which scoring instrument should the teacher utilize?
A single-point checklist that records only whether the final numerical answer is correct or incorrect.
A norm-referenced grading curve that calculates standard deviations from the classroom mean score.
A holistic scoring rubric that assigns a single global score from 1 to 4 based on overall presentation.
An analytic scoring rubric that evaluates independent criteria such as conceptual understanding, procedural accuracy, mathematical communication, and justification.
Following a formative check on solving two-step linear equations, a seventh-grade teacher discovers that 8 out of 24 students consistently add the constant term to both sides when the equation contains a positive constant (e.g., in 2x + 6 = 18, they write 2x = 24). Within a Multi-Tiered System of Supports (MTSS) framework, what is the most appropriate instructional response?
Assign all 24 students twenty additional practice problems of the exact same type for homework to reinforce procedural repetition.
Provide Tier 2 targeted small-group intervention for the 8 students using algebra tiles and balance mats to model inverse operations and zero pairs, while the remaining students engage in independent application or extension tasks.
Move directly to the next curriculum unit on inequalities, assuming that students will naturally master two-step equations as they advance.
Refer all 8 students immediately for special education evaluation and Tier 3 intensive individualized placement.
Sections you finish are checked off in the contents.
You've completed this section
Continue exploring other exams