4.1 Assessment Types in a Standards-Aligned System
Key Takeaways
- PECT 0002.1 scores whether you match authentic, screening, diagnostic, formative, summative, and benchmark assessment to the decision each one can support in Pennsylvania's Standards Aligned System.
- Screening is a brief, universal flag for who needs a closer look; diagnostic assessment is a deeper, targeted map of specific skills — a screener never diagnoses a disability by itself.
- Formative assessment happens during instruction and changes what you teach next; summative assessment judges learning after a defined period of teaching.
- Benchmark assessment is a periodic standards checkpoint (often fall, winter, spring), not a daily exit ticket and not a full diagnostic inventory.
- SAS treats Assessments as one of six achievement elements beside Standards, Curriculum Framework, Instruction, Materials & Resources, and school climate — data that never change grouping, time, or reteaching are not yet decisions.
4.1 Assessment Types in a Standards-Aligned System
PECT lens: Objective 0002.1 (Module 1, Subarea I) asks you to name the types of assessment used in a standards-aligned system — authentic, screening, diagnostic, formative, summative, and benchmark — and to use each type for the right data-based decision. Pennsylvania's organizer is SAS, the Standards Aligned System at pdesas.org. The exam scores the decision the teacher makes next, not a slogan about "using data."
A PreK–4 classroom that is standards-aligned does not merely post Pennsylvania Core Standards, Pennsylvania Academic Standards, or the Pennsylvania Learning Standards for Early Childhood on a bulletin board. Adults collect evidence of what children know and can do, then change grouping, materials, time, and reteaching. SAS names six elements that affect achievement: Standards, Assessments, Curriculum Framework, Instruction, Materials & Resources, and the Office of School Climate and Well-Being. Assessment is the element that tells you whether instruction is actually moving children toward the standard. If you skip it, you are guessing.
Pennsylvania SAS materials often highlight four state-named categories — formative, benchmark, diagnostic, and summative. PECT 0002.1 also names authentic and screening. Learn all six. They answer different questions, and the distractors on this objective are usually a true description of the wrong type.
The six types PECT expects you to separate
| Type | Core question | Typical timing | PreK–4 snapshot | What it is not |
|---|---|---|---|---|
| Authentic | Can the child use the skill in a real task? | Whenever genuine performance is the evidence | A kindergartner writes the lunch count; Grade 3 students measure a garden bed | A multiple-choice packet that never asks children to do the work |
| Screening | Who might need a closer look? | Brief, universal (whole class or grade) | Every kindergartner completes a short literacy screener in September | A diagnosis, an IEP, or a report-card grade |
| Diagnostic | Exactly which skills are missing, and why is this hard? | After a flag, or when planning targeted instruction | A first grader who flagged on phonics gets a sound-by-sound inventory | A two-minute whole-class check |
| Formative | What should I teach tomorrow — or in the next ten minutes? | During instruction, frequent and low-stakes | Exit tickets, anecdotal notes, a running record used to regroup | The unit test that becomes the grade |
| Summative | Did they meet the standard after teaching? | After a unit, term, or year | Unit rubric, report-card evidence, PSSA in tested grades | Daily checks that still change this week's teaching |
| Benchmark | Are we on track toward grade-level standards at this checkpoint? | Periodic (often three times a year) | Midyear Grade 2 literacy checkpoint aligned to PA Core | A sticky-note observation during centers |
Authentic assessment is a method: evidence comes from a task that resembles real use of the standard. A PreK child who sorts snack cups to set the table is showing one-to-one correspondence more honestly than a coloring sheet of numerals. Authentic work can be formative (you change the next invitation) or summative (you score a rubric at the end of the unit). The word names the nature of the evidence, not the calendar. On items, a pumpkin investigation that asks children to weigh, count seeds, and explain a finding is authentic if it maps to a named standard. A pumpkin collage that only looks seasonal is not assessment of a standard.
Screening is a gate, not a verdict. Good screeners are brief, economical in time, given to everyone, and sensitive enough to catch children who may need more. A child who scores below a screener's concern range is not thereby labeled with a disability. The professional next step is more information — often a diagnostic probe, classroom work samples, and family conversation. The professional error is treating a three-minute literacy check as a diagnosis or skipping screening and waiting until May to notice a child cannot blend CVC words. Universal screening is for all children in the group, including children who already have IEPs and children who look typically developing. You screen the class; you diagnose the concern.
Diagnostic assessment is narrow and deep. Pennsylvania's Classroom Diagnostic Tools (CDT) on SAS are an example of online diagnostics in eligible grades and subjects; classroom teachers also diagnose with letter-sound inventories, miscue analysis, math interviews, and targeted probes. Diagnostic results should name specific skills ("short-vowel decoding; blending three phonemes") so instruction can change. They are not a replacement for daily formative checks, and they are not a universal screener. If a stem says the teacher needs to know which phonemes or which computation strategies are missing, the answer is diagnostic, not screening and not benchmark.
Formative assessment is assessment for learning. You collect evidence while teaching so you can adjust instruction. The same running record is formative if you use it to pull a decoding group tomorrow and wasted if you only file it. PECT cares about use. Formative evidence is often ungraded or lightly marked. Thumbs-up checks, whiteboard responses, conference notes, and center anecdotal records all count when they change teaching. Public comparison charts that shame children are not "formative"; they are a climate and ethics problem (Section 4.4).
Summative assessment is assessment of learning. After a defined period of instruction, you judge achievement against the standard. In Pennsylvania, the Pennsylvania System of School Assessment (PSSA) is a statewide summative measure in tested grades — English language arts and mathematics beginning in Grade 3, with science in the elementary grade PDE currently specifies. PreK, kindergarten, Grade 1, and Grade 2 do not take the PSSA. A PreK–2 summative is more often a unit rubric, a portfolio conference, or an end-of-term standards checklist. Summative data evaluate a period of teaching. They are usually too late to change that unit, but they should change the next unit, summer planning, and school-level supports. Do not treat a kindergarten play-based documentation panel as "not real assessment" just because it is not a bubble test.
Benchmark assessment is the type candidates confuse with formative. A benchmark is a periodic checkpoint against grade-level standards — often fall, winter, and spring — so a team can see who is on track, who needs intervention, and whether the core program is working for the grade. It is not a daily exit ticket. It is not a deep diagnostic of every subskill. Think of it as a standards pulse, not a lesson pulse. SAS Assessment Center tools that let teachers build periodic standards-aligned checks sit in this family when they are used as grade-level checkpoints rather than as tomorrow's reteach notes.
Data-based decision-making, not data theater
SAS assessment exists so teachers can make decisions. A useful loop:
- Clarify the standard (PA Core, Early Learning Standards, STEELS, or other PA Academic Standards).
- Choose the type that answers the decision you actually have.
- Collect evidence with a tool that matches the type.
- Interpret in light of more than one source when the decision is high-stakes (Section 4.3).
- Act — regroup, reteach, enrich, or refer for further assessment.
- Recheck to see whether the action worked.
If a "data meeting" produces a color-coded spreadsheet and no change in instruction, it was not data-based decision-making. PECT stems often hide the right type in the next teacher move. "Reteach the mini-lesson tomorrow" points to formative use. "Identify who needs a follow-up inventory" points to screening. "Name the missing phonics patterns" points to diagnostic. "Assign a report-card proficiency level after the unit" points to summative. "See whether the grade is on track in January" points to benchmark.
Scenario: kindergarten September (screening, then diagnostic)
Ms. Patel screens every kindergartner with a brief letter-sound and phonological-awareness measure. Four children fall below the concern range. She does not call families to announce a disability. She gives a diagnostic letter-name/sound inventory, watches those children at the writing center, and talks with families about home languages. Two children need a small-group letter-sound cycle; one is a dual language learner whose English oral language is still emerging; one was exhausted on screening day and looks typical on the diagnostic. Screening found them; diagnosis explained them.
Scenario: Grade 1 reading group (formative)
During guided reading, Mr. Alvarez takes a running record. He notices slow blending on short vowels. That night he rebuilds tomorrow's word-work from CVC short a and i. The running record was formative because it changed instruction. If he only filed it for an April portfolio without changing teaching, he wasted the tool — even though a running record can also feed a portfolio or a later summative judgment.
Scenario: Grade 3 after a fractions unit (summative and authentic)
The class has spent three weeks on equal shares. Ms. Chen gives a summative performance task: partition a real pan of cornbread and explain why two-fourths equals one-half, scored with a rubric aligned to PA Core mathematics. That task is also authentic. The PSSA later in the year is a statewide summative of a broader Grade 3 set; it does not replace her unit evidence, and her unit task is not "practice PSSA" just because Grade 3 is a tested grade.
Scenario: Grade 2 winter (benchmark)
The grade-level team administers a midyear benchmark aligned to Grade 2 reading standards. The results are not used to rewrite tomorrow's mini-lesson for one child (that is formative). They are used to ask: Is core instruction moving the grade toward year-end standards? Which students need a diagnostic look? Which classrooms need coaching?
Exam traps for 0002.1
- Screening = diagnostic. Screening is brief and universal; diagnostic is targeted and deep.
- Formative = any test that is not the PSSA. Formative is defined by when and how you use the evidence — during instruction, to change teaching.
- Benchmark = daily formative. Benchmark is a periodic standards check.
- Authentic = ungraded play with no standard. Authentic tasks still map to a standard.
- PSSA as the only assessment that counts. Young children are assessed constantly with observation, work samples, and classroom measures; statewide summative tests begin in tested grades.
When a stem describes an assessment, ask: What decision is this teacher trying to make? Who needs a closer look (screening), which skill is broken (diagnostic), what do I teach next (formative), did they meet the standard after teaching (summative), are we on track this term (benchmark), or can they do the real work (authentic)? Match the type to the decision.
A kindergarten teacher administers a five-minute letter-sound measure to every child in September. Several children score below the concern range. What is the most appropriate interpretation?
A Grade 2 team gives the same standards-aligned literacy measure in September, January, and May to see who is on track for year-end Pennsylvania Core reading standards. This assessment is best classified as:
During a fractions mini-lesson, a Grade 3 teacher asks students to show one-half on whiteboards, notices three children drawing two-thirds, and immediately reteaches with fraction strips before independent practice. This is primarily: