4.2 Informal and Formal Assessment Tools
Key Takeaways
- PECT 0002.2 tests whether you can pick, build, and use informal and formal tools — systematic observation, portfolios, peer and group assessment, curriculum-based assessment/measurement, criterion-referenced tests, and norm-referenced tests — for a stated purpose.
- A criterion-referenced test (CRT) measures mastery of a taught standard or criterion; a norm-referenced test (NRT) ranks a student against a comparison group and is the wrong tool for 'did they master this week's standard?'
- Curriculum-based measurement (CBM) repeatedly samples skills from the local curriculum so progress can be graphed; it is not a one-shot national ranking.
- Systematic observation, anecdotal notes, running records, and rubrics are only as good as their focus, dated evidence, and criteria — watching without a purpose is not systematic.
- Portfolios, peer assessment, and group assessment can show authentic growth or collapse into popularity contests; the limitation is usually the quality of the criteria, not the folder itself.
4.2 Informal and Formal Assessment Tools
PECT lens: Objective 0002.2 asks you to know the characteristics, uses, advantages, and limitations of informal and formal tools — systematic observation, portfolio, peer assessment, group assessment, curriculum-based assessment/measurement, criterion-referenced tests, and norm-referenced tests — and to select, construct, and use them for different purposes and needs. Running records, anecdotal notes, and rubrics are the classroom crafts that make those tools real in PreK–4.
Informal does not mean unplanned or invalid. Informal tools are often classroom-embedded: observation, anecdotal notes, running records, work samples, conferences. Formal tools are typically standardized in administration and scoring: published norm-referenced tests (NRTs), many criterion-referenced tests (CRTs), and statewide summative tests. Both families can be high-quality. Both can be misused. PECT items usually hide the right tool in the purpose of the stem, not in the word "test."
Tools at a glance: advantages and limits
| Tool | What it typically shows | Advantage | Limitation | Best use |
|---|---|---|---|---|
| Systematic observation | What the child does in context | Captures play, process, and language you would miss on a worksheet | Time; observer bias if notes are vague or undated | PreK centers, social skills, process in science |
| Anecdotal notes | A dated, factual snapshot of one incident | Quick; builds a pattern across days | Useless if they are judgments ("lazy") instead of evidence | Collecting examples for conferences and portfolios |
| Running record | Oral reading accuracy, error types, and often fluency/self-correction | Fine-grained literacy evidence you can use tomorrow | Narrow if used as the only reading measure; administration skill matters | Formative grouping in Grade 1–2 reading |
| Rubric | Performance against described criteria | Makes a complex task scorable and teachable | Vague descriptors ("good job") collapse reliability | Authentic writing, math explanations, science investigations |
| Portfolio | Growth over time in selected work | Shows process and product; powerful with families | Time; only as valid as selection criteria | Documenting ELS domains or a writing standard across months |
| Peer assessment | How classmates judge work against criteria | Builds audience awareness and self-monitoring | Popularity, social power, and unclear criteria | Grade 2–4 writing or project checklists after teaching the criteria |
| Group assessment | What a team produced together | Mirrors collaborative work | Masks individual skill unless you also collect individual evidence | A center investigation — never the sole basis for one child's proficiency |
| CBA / CBM | Performance on the actual curriculum; repeated brief probes | Sensitive to growth; tied to what you taught | A single probe is a snapshot, not a diagnosis | Weekly oral-reading or math-fact probes to graph progress |
| CRT | Mastery of a defined standard or criterion | Answers "did they meet this standard?" | Tells you little about standing in a national sample | Unit tests, standards checklists, mastery rubrics |
| NRT | Standing relative to a norm group | Useful for some program or large-scale comparison questions | Wrong tool for mastery of a taught standard | Not your weekly "did they learn fractions?" check |
Informal tools you must be able to construct
Systematic observation is planned watching with a focus. Before centers, a PreK teacher decides, "Today I am collecting evidence of one-to-one correspondence and turn-taking in the block area." She uses a class list, codes, or a simple checklist, and she dates the notes. Wandering the room hoping something interesting happens is not systematic. Advantage: you see children using skills in play, which is often the most valid window in PreK–K. Limitation: you cannot observe everyone on every skill every day, and an observer who only writes "off task" has collected a judgment, not evidence.
Anecdotal notes are short, dated, factual records ("9/14, sand table: Maya said 'full' and 'empty' while pouring, then counted six cups"). They become a body of evidence when you collect several over time. Construct them without labels about character. Use them to feed portfolios, family conferences, and later diagnostic hunches. They are a poor standalone tool for a high-stakes decision.
Running records (including related oral-reading records) capture what a child does with text: substitutions, omissions, self-corrections, and, when timed, rate. Construct one by sitting beside the child with a familiar or benchmark text, coding errors, and then analyzing the error types (meaning, visual, structure) so instruction changes. A running record used to regroup is formative (Section 4.1). A running record dropped in a folder and never analyzed is busywork. Limitation: it does not, by itself, measure listening comprehension, vocabulary depth, or writing.
Rubrics turn a complex performance into shared criteria. A strong PreK–4 rubric names observable levels ("explains why the two shares are equal using the words half or fourths") rather than mushy labels ("excellent / good / poor"). You construct a rubric from the standard, share it with children in kid language, and score from evidence. Rubrics support authentic tasks, portfolios, and CRT-style mastery judgments. They fail when every child lands in the middle because the descriptors do not discriminate, or when the teacher scores from memory instead of from the work.
Portfolios, peers, and groups
A portfolio is a purposeful collection of work over time — dated writing samples, photos of block structures with transcribed child comments, scored rubrics, a self-selected "best piece." Advantage: families can see growth; teachers can see a child's process, not one Friday quiz. Limitation: bulky, time-consuming, and only as valid as the selection rules. A scrapbook of cute photos with no link to a standard is not an assessment portfolio. Construct portfolios with a short list of standards, dated entries, and a caption that says what the piece shows.
Peer assessment asks children to give feedback using taught criteria ("Does the story have a beginning, middle, and end?"). It can strengthen audience awareness in Grade 2–4 writing. It is a terrible tool for grading individuals, for ranking friends, or for PreK children who are still learning that critique is about work, not about the person. Teach the checklist first; keep stakes low; never use peer scores as the sole report-card evidence.
Group assessment scores a team's product — a Grade 4 science poster, a kindergarten mural plan. Collaboration is a real standard-adjacent skill, but a group grade hides who cannot yet measure in centimeters. Select group assessment when the learning target is collaboration or a shared investigation, and add individual checks (an exit ticket, a conference, a photo of one child's contribution). On PECT, the distractor is often "the group got an A, so every child has mastered the standard."
CBA, CBM, CRT, and NRT — the formal core
Curriculum-based assessment (CBA) means you assess what you actually taught from the local curriculum and Pennsylvania standards, not a generic national skill list. A teacher-made fractions quiz aligned to this week's PA Core target is CBA. Curriculum-based measurement (CBM) is a specific, research-based subset: brief, repeatable probes sampled from the curriculum (for example, oral reading of grade-level passages, or math computation sheets) given under similar conditions so you can graph progress over weeks. Advantage: sensitive to small growth; useful for progress monitoring (Section 4.3). Limitation: a probe is not a full diagnostic of every subskill, and a flattened CBM line tells you to change instruction, not to keep giving the same probe and hoping.
A criterion-referenced test (CRT) compares the student to a predefined criterion — usually a standard or a cut score the teacher or state set for mastery, not to other children's ranks. "Can this child independently write a complete sentence with capital and period?" is a criterion question. Unit tests, standards checklists, many SAS-built items, and well-written rubrics live here. Advantage: they answer the question Pennsylvania teachers actually have after teaching a standard. Limitation: a poorly written CRT with trivial items will "pass" children who cannot do the authentic work.
A norm-referenced test (NRT) compares the student to a norm group (often a national sample) and reports standing as a percentile, stanine, or similar rank. Advantage: useful when the question is "How does this performance compare with a large reference group?" — some program-eligibility or large-scale questions. Limitation — and the exam trap: an NRT is not designed to tell you whether the child mastered this week's taught Pennsylvania standard. A child can sit at the 60th percentile nationally and still have not mastered equivalent fractions you just taught; another child can master the taught standard and still sit below a national mean. Do not use an NRT as a mastery measure of a taught standard. That is a CRT (or rubric/CBA) job.
Scenario: PreK blocks (observation + anecdotal notes)
Mr. Cole wants evidence of spatial language and persistence. He sets a two-column checklist, watches the block area for fifteen minutes, and writes: "11/3, Jordan: said 'under' and 'beside'; rebuilt the ramp three times after it collapsed." That is systematic observation plus an anecdotal note. Scoring Jordan with a published NRT of "school readiness" would not capture the ramp persistence the Early Learning Standards actually ask him to see.
Scenario: Grade 1 reading (running record + CBM)
Ms. Lee takes a running record on Monday (error analysis → short-vowel reteach) and a one-minute oral-reading CBM probe each Friday to graph words read correctly over time. The running record explains the errors; the CBM tracks whether instruction is moving the skill. Neither is an NRT.
Scenario: Grade 4 fractions (CRT vs NRT trap)
The principal asks whether students mastered equivalent fractions. Mr. Diaz should use a CRT or rubric aligned to the PA Core standard — a set of items and an authentic partitioning task with a mastery criterion. Pulling last year's national achievement-battery percentile is the PECT-wrong move: that NRT ranks against a norm group; it does not certify mastery of the taught standard.
How to select, construct, and use
- Name the purpose (screen, diagnose, formatively regroup, judge mastery, track weekly growth, compare to a norm group).
- Match the tool using the table above.
- Construct or select so the items or observation focus match the standard and the child's developmental level (a 40-item bubble test is a poor PreK tool; a one-item worksheet is a poor Grade 4 mastery check).
- Use the results for the purpose you named. Do not convert a peer checklist into a special-education referral, or an NRT percentile into a unit grade.
Exam traps for 0002.2
- NRT to measure mastery of a taught standard. Use a CRT, rubric, or CBA.
- Portfolio = junk drawer. Without criteria and dates, it is not assessment.
- Group score = individual mastery.
- Observation without a focus. Not systematic.
- Peer assessment as a high-stakes grade for young children.
- CBM as a once-a-year test. CBM's power is repeated probes.
If you can say, in one sentence, "I am using this tool to answer this question," you are inside 0002.2. If you cannot, you picked the tool because it was in the cabinet.
A Grade 4 teacher wants to know whether students have mastered the taught Pennsylvania Core standard on equivalent fractions. Which tool is designed for that purpose?
Which statement correctly contrasts a curriculum-based measurement (CBM) probe with a one-time norm-referenced achievement test?
Ms. Rivera keeps dated work samples, photos of block structures with transcribed child comments, and two scored rubrics for each PreK child. The main advantage of this portfolio is: