Designing Aligned Diagnostic, Formative & Summative Assessments

Key Takeaways

  • Validity concerns evidence supporting the intended interpretation and use.

  • Reliability concerns consistency and does not alone prove validity.

  • Criterion scores and normed comparisons answer different questions.

Last updated: October 2026

Assessment is not merely the final, administrative step of a unit where grades are recorded; it is the compass that guides every instructional decision a teacher makes. In Texas classrooms, assessment must be seamlessly integrated into the learning cycle. Competency 003 of the TExES PPR requires educators to design and administer a balanced continuum of assessments that accurately measure student progress toward the TEKS, provide actionable feedback, and guide responsive reteaching.

The Assessment Continuum: Timing, Purpose, and Action

Assessments fall into three major modalities based on where they occur within the instructional cycle:

                     THE CLASSROOM ASSESSMENT CONTINUUM

  BEFORE INSTRUCTION            DURING INSTRUCTION             AFTER INSTRUCTION
┌────────────────────┐       ┌────────────────────┐       ┌────────────────────┐
│     DIAGNOSTIC     │       │     FORMATIVE      │       │     SUMMATIVE      │
│  (Pre-Assessment)  │──────→│ (Checks for Under- │──────→│ (Evaluation of     │
│                    │       │     standing)      │       │     Mastery)       │
└────────────────────┘       └────────────────────┘       └────────────────────┘
  • Identifies prior           • Real-time data to          • Final determination
    knowledge & gaps             adjust teaching              of standards mastery
  • Informs grouping &         • Low stakes; guides         • High stakes; grades
    differentiation              reteaching                   recorded

1. Diagnostic Assessment (Assessment Before Learning)

Diagnostic assessments are administered prior to instruction to establish a baseline of student knowledge, identify prerequisite skill deficits, and reveal misconceptions. Diagnostic data allows teachers to plan tiered instruction, form flexible small groups, and differentiate content before whole-group instruction begins.

  • Common Instruments: Pre-tests, KWL charts (What I Know, What I Want to Know, What I Learned), anticipation guides, diagnostic running records, and concept maps.

  • Pedagogical Rule: Diagnostic assessments should rarely be graded for report cards; their sole purpose is diagnostic information.

  • Use evidence for an instructional response: Examine individual errors and choose timely feedback, explanation, practice, or a different representation as needed. A fixed class percentage does not prescribe the same action for every learner. Required supports remain in place while ready students extend learning and others receive targeted help.

3. Summative Assessment (Assessment of Learning)

Summative assessment occurs at the conclusion of an instructional cycle (unit, semester, or school year). It evaluates cumulative student learning against predetermined TEKS standards and assigns a formal evaluative score or grade.

  • Common Instruments: Unit exams, cumulative projects, authentic performance tasks, STAAR assessments, and End-of-Course (EOC) examinations.
Assessment TypeWhen AdministeredPrimary PurposeImpact on InstructionSample Classroom Tools
DiagnosticBefore a unit or lesson beginsIdentify baseline competencies, entry skills, and prerequisite misconceptionsGuides initial lesson planning, differentiation strategies, and flexible groupingSkill pre-test, anticipation guide, KWL chart, entry inventory
FormativeContinuously during the lesson cycleMonitor student comprehension in real time and provide immediate corrective feedbackTriggers instant instructional adjustments, mid-course pivots, and targeted small-group reteachingExit tickets, individual whiteboards, digital polls, four corners
SummativeAt the conclusion of a unit or termMeasure cumulative student mastery against grade-level TEKS benchmarksInforms final grades, curriculum effectiveness, and campus intervention placementEnd-of-unit exam, research paper, science lab practical, capstone

Psychometric Principles: Validity, Reliability, and Fairness

For an assessment to provide trustworthy data, it must satisfy two core psychometric criteria while eliminating bias:

Content Validity (Alignment)

Validity answers the question: Does the assessment measure what it claims to measure? In Texas, validity is synonymous with TEKS alignment:

  • Content Alignment: The assessment items must test the exact knowledge specified in the student expectation.
  • Cognitive Rigor Alignment: The assessment items must match the cognitive demand of the TEKS verb. If the TEKS standard requires students to critique an author's bias, a multiple-choice test that merely asks students to identify definitions of literary terms lacks content validity.

Reliability (Consistency)

Reliability answers the question: Does the assessment produce stable, consistent results across time, testing environments, and different raters? An assessment is reliable if a student possessing a certain level of mastery would achieve a comparable score if they took the test on Tuesday or Thursday, or if their essay was scored by Teacher A or Teacher B. Reliable scoring of open-ended tasks requires structured, descriptive rubrics.

Eliminating Assessment Bias and Ensuring Equity

An assessment is biased if it contains cultural, linguistic, or socioeconomic references that advantage or disadvantage specific student groups independent of their content mastery:

  • Linguistic and Cultural Bias: Using colloquial idioms, obscure sports metaphors, or culturally specific scenarios (e.g., asking elementary students to calculate ski trip expenses when many have never seen snow) creates barriers for English Language Learners and economically disadvantaged students.
  • Accommodations: Teachers must faithfully implement all testing accommodations mandated in a student's Individualized Education Program (IEP) under IDEA or Section 504 plan—such as oral administration, graphic organizers, extra time, or frequent breaks.

Note

Validity Trumps Reliability An assessment can be highly reliable while being completely invalid. For example, a scale that consistently weighs an object five pounds too light is perfectly reliable (consistent), but completely invalid (inaccurate). In the classroom, an easy multiple-choice recall quiz may be easy to grade with high reliability, but if the TEKS required creative synthesis, the test is invalid.


Criterion-Referenced vs. Norm-Referenced Assessment

Criterion-referenced interpretation compares performance with defined learning criteria or standards. Norm-referenced interpretation compares performance with an identified reference group. A percentage correct is not a percentile rank: 80% correct means eight of ten items, while a percentile describes relative standing in the specified comparison distribution. Understand the actual score report and purpose before interpreting a result.

STAAR content results use criterion-related performance standards, but Texas schools also use assessments and norms for other purposes. Do not claim that every assessment in public education is exclusively criterion-referenced. Likewise, a normed test does not force a fixed percentage of the current class to fail; interpretation depends on the reference group. A classroom grade should communicate the intended achievement under applicable policy rather than impose a failure quota unrelated to demonstrated learning.

A Comparison in Practice

In a fictional vocabulary check, a student answers sixteen of twenty items correctly, producing 80% correct. Whether that meets the classroom criterion depends on the announced purpose and standards. If a separate score report gives the student the 70th percentile in its reference group, that is not 70% correct and does not automatically convert to a letter grade. The teacher reviews which items were missed and whether their demands match the intended skill.

A criterion score can guide reteaching, but a summary alone may hide different errors. A normed comparison can signal relative performance, but it does not diagnose a disability or explain the cause. Combine appropriate reports with student work and observation. Document scale, reference group, conditions, and missing data when those affect interpretation. Use the statistic for the question it answers rather than treat all numerical results as interchangeable.


Performance Assessments and Analytic Rubric Construction

While traditional multiple-choice questions can assess foundational knowledge efficiently, complex higher-order thinking skills (Bloom's Levels 4–6) are best evaluated through performance assessments—authentic, hands-on demonstrations such as science laboratory experiments, persuasive speeches, oral defenses, or multi-step engineering designs.

To evaluate performance assessments with high validity and inter-rater reliability, teachers construct rubrics:

Analytic vs. Holistic Rubrics

  • Holistic Rubrics: Provide a single overall score based on an impressionistic evaluation of the entire product. While fast to grade, holistic rubrics provide poor diagnostic feedback to students.
  • Analytic Rubrics: Decompose the task into discrete, independent dimensions (e.g., Thesis, Evidence & Textual Support, Logical Organization, Conventions) and evaluate each dimension across descriptive performance levels (e.g., Exemplary, Proficient, Developing, Beginning).
                    ANALYTIC RUBRIC ARCHITECTURE

CRITERIA         BEGINNING (1)    DEVELOPING (2)    PROFICIENT (3)    ADVANCED (4)
──────────────────────────────────────────────────────────────────────────────────
Evidence &       Provides no      Cites text but    Cites at least    Skillfully weaves
Textual Support  textual support  evidence is       two relevant      compelling textual
                 for claims       irrelevant or     quotes that       quotes that deeply
                                  unexplained       support claims    illuminate analysis
──────────────────────────────────────────────────────────────────────────────────
Analysis &       Merely retells   Summarizes plot   Explains how      Insightfully critiques
Reasoning        plot events      with minimal      evidence proves   underlying themes
                                  interpretation    the thesis        and author craft
  1. Supports more consistent judgment: Clear descriptive criteria can reduce inconsistency when teachers calibrate scoring with work samples. A rubric does not eliminate judgment, fatigue, halo effects, or bias by itself; review the criteria and compare scoring decisions.

Realistic Texas Classroom Scenario: Integrated Assessment Architecture in 8th Grade ELAR

These are fictional teaching examples. Counts, timings, and outcomes illustrate decisions; they are not research findings or promised effects.

Ms. Garza, an 8th-grade ELAR teacher at San Jacinto Middle School, designed a four-week unit on persuasive writing aligned to the illustrative learning target Rather than waiting until the final week to administer a test, Ms. Garza embedded assessments throughout the unit cycle:

  1. Diagnostic Stage: On Day 1, Ms. Garza administered a 10-minute diagnostic prompt: "Take a position on school uniforms and provide one reason." Analyzing the responses, she discovered that while students could easily state an opinion, 70% could not formulate a counterargument or cite external evidence.
  2. Formative Stage: Throughout Weeks 2 and 3, Ms. Garza instituted daily exit tickets. When an exit ticket revealed that students struggled to distinguish between a credible statistic and anecdotal opinion, she immediately paused the next day's schedule to deliver a 15-minute targeted mini-lesson on source reliability, using sample tweets and academic journals.
  3. Peer Evaluation with Rubric: Prior to drafting, Ms. Garza distributed an analytic rubric detailing four distinct dimensions: Thesis Clarity, Counterargument Refutation, Textual Evidence, and Sentence Variety. Students used the rubric to conduct structured peer conferences.
  4. Summative Stage: Students submitted their polished persuasive essays, which Ms. Garza scored using the analytic rubric.

Because her assessment system was fully aligned, diagnostic data prevented blind spots, formative checks drove real-time instructional pivots, and the analytic rubric ensured objective, transparent grading that fostered profound student growth.

Test Your Knowledge

During a 6th-grade math lesson on dividing fractions, a teacher gives each student an individual dry-erase whiteboard and prompts them to solve 2/3 ÷ 1/4. Upon scanning the held-up whiteboards, the teacher observes that approximately 65% of students inverted the dividend rather than the divisor. Which instructional response exemplifies effective formative assessment pedagogy?

A

Record a failing grade in the gradebook for all students who made the calculation error and assign twenty homework problems.

B

Proceed immediately to the planned independent practice worksheet, assuming students will correct the mistake on their own.

C

Halt independent practice, explicitly reteach the reciprocal concept using a visual fraction model, and conduct another whiteboard check before proceeding.

D

Tell the students that division is difficult and skip the fraction division standard entirely for the rest of the unit.

Test Your Knowledge

A high school English IV teacher wants to assess students' ability to compose a college-level research paper. To ensure students receive targeted feedback on specific aspects of their writing (such as thesis formulation, source integration, synthesis of arguments, and MLA citation mechanics), which assessment tool should the teacher implement?

A

An analytic rubric that defines discrete performance criteria and explicit descriptive levels of achievement for each writing component.

B

A holistic rubric that assigns a single overall letter grade based on the general impression of the essay.

C

A norm-referenced bell curve that ranks students against their classmates' essays regardless of individual quality.

D

A true/false quiz testing whether students remember the rules of MLA documentation format.

Test Your Knowledge

A Texas middle school administrator reviews a novice teacher's grading policy and notices that the teacher grades all unit assessments on a strict bell curve, ensuring that only the top 10% of students earn an 'A' and the bottom 15% automatically receive failing grades, regardless of their actual test scores. Why is this practice inconsistent with Texas pedagogical standards?

A

State regulations require that 100% of enrolled students receive an 'A' in every course regardless of academic effort.

B

Bell-curve grading is illegal under federal copyright law governing standardized testing materials.

C

Middle school students should never be evaluated using numerical grades or percentage scores under district policies.

D

Texas public education relies on criterion-referenced assessment, where grades must reflect student mastery of specific TEKS standards rather than relative ranking against peers.

Sections you finish are checked off in the contents.