4.2 Informal and Formal Assessment Tools

Key Takeaways

  • PECT 0002.2 tests whether you can pick, build, and use informal and formal tools — systematic observation, portfolios, peer and group assessment, curriculum-based assessment/measurement, criterion-referenced tests, and norm-referenced tests — for a stated purpose.
  • A criterion-referenced test (CRT) measures mastery of a taught standard or criterion; a norm-referenced test (NRT) ranks a student against a comparison group and is the wrong tool for 'did they master this week's standard?'
  • Curriculum-based measurement (CBM) repeatedly samples skills from the local curriculum so progress can be graphed; it is not a one-shot national ranking.
  • Systematic observation, anecdotal notes, running records, and rubrics are only as good as their focus, dated evidence, and criteria — watching without a purpose is not systematic.
  • Portfolios, peer assessment, and group assessment can show authentic growth or collapse into popularity contests; the limitation is usually the quality of the criteria, not the folder itself.
Last updated: August 2026

4.2 Informal and Formal Assessment Tools

PECT lens: Objective 0002.2 asks you to know the characteristics, uses, advantages, and limitations of informal and formal tools — systematic observation, portfolio, peer assessment, group assessment, curriculum-based assessment/measurement, criterion-referenced tests, and norm-referenced tests — and to select, construct, and use them for different purposes and needs. Running records, anecdotal notes, and rubrics are the classroom crafts that make those tools real in PreK–4.

Informal does not mean unplanned or invalid. Informal tools are often classroom-embedded: observation, anecdotal notes, running records, work samples, conferences. Formal tools are typically standardized in administration and scoring: published norm-referenced tests (NRTs), many criterion-referenced tests (CRTs), and statewide summative tests. Both families can be high-quality. Both can be misused. PECT items usually hide the right tool in the purpose of the stem, not in the word "test."

Tools at a glance: advantages and limits

ToolWhat it typically showsAdvantageLimitationBest use
Systematic observationWhat the child does in contextCaptures play, process, and language you would miss on a worksheetTime; observer bias if notes are vague or undatedPreK centers, social skills, process in science
Anecdotal notesA dated, factual snapshot of one incidentQuick; builds a pattern across daysUseless if they are judgments ("lazy") instead of evidenceCollecting examples for conferences and portfolios
Running recordOral reading accuracy, error types, and often fluency/self-correctionFine-grained literacy evidence you can use tomorrowNarrow if used as the only reading measure; administration skill mattersFormative grouping in Grade 1–2 reading
RubricPerformance against described criteriaMakes a complex task scorable and teachableVague descriptors ("good job") collapse reliabilityAuthentic writing, math explanations, science investigations
PortfolioGrowth over time in selected workShows process and product; powerful with familiesTime; only as valid as selection criteriaDocumenting ELS domains or a writing standard across months
Peer assessmentHow classmates judge work against criteriaBuilds audience awareness and self-monitoringPopularity, social power, and unclear criteriaGrade 2–4 writing or project checklists after teaching the criteria
Group assessmentWhat a team produced togetherMirrors collaborative workMasks individual skill unless you also collect individual evidenceA center investigation — never the sole basis for one child's proficiency
CBA / CBMPerformance on the actual curriculum; repeated brief probesSensitive to growth; tied to what you taughtA single probe is a snapshot, not a diagnosisWeekly oral-reading or math-fact probes to graph progress
CRTMastery of a defined standard or criterionAnswers "did they meet this standard?"Tells you little about standing in a national sampleUnit tests, standards checklists, mastery rubrics
NRTStanding relative to a norm groupUseful for some program or large-scale comparison questionsWrong tool for mastery of a taught standardNot your weekly "did they learn fractions?" check

Informal tools you must be able to construct

Systematic observation is planned watching with a focus. Before centers, a PreK teacher decides, "Today I am collecting evidence of one-to-one correspondence and turn-taking in the block area." She uses a class list, codes, or a simple checklist, and she dates the notes. Wandering the room hoping something interesting happens is not systematic. Advantage: you see children using skills in play, which is often the most valid window in PreK–K. Limitation: you cannot observe everyone on every skill every day, and an observer who only writes "off task" has collected a judgment, not evidence.

Anecdotal notes are short, dated, factual records ("9/14, sand table: Maya said 'full' and 'empty' while pouring, then counted six cups"). They become a body of evidence when you collect several over time. Construct them without labels about character. Use them to feed portfolios, family conferences, and later diagnostic hunches. They are a poor standalone tool for a high-stakes decision.

Running records (including related oral-reading records) capture what a child does with text: substitutions, omissions, self-corrections, and, when timed, rate. Construct one by sitting beside the child with a familiar or benchmark text, coding errors, and then analyzing the error types (meaning, visual, structure) so instruction changes. A running record used to regroup is formative (Section 4.1). A running record dropped in a folder and never analyzed is busywork. Limitation: it does not, by itself, measure listening comprehension, vocabulary depth, or writing.

Rubrics turn a complex performance into shared criteria. A strong PreK–4 rubric names observable levels ("explains why the two shares are equal using the words half or fourths") rather than mushy labels ("excellent / good / poor"). You construct a rubric from the standard, share it with children in kid language, and score from evidence. Rubrics support authentic tasks, portfolios, and CRT-style mastery judgments. They fail when every child lands in the middle because the descriptors do not discriminate, or when the teacher scores from memory instead of from the work.

Portfolios, peers, and groups

A portfolio is a purposeful collection of work over time — dated writing samples, photos of block structures with transcribed child comments, scored rubrics, a self-selected "best piece." Advantage: families can see growth; teachers can see a child's process, not one Friday quiz. Limitation: bulky, time-consuming, and only as valid as the selection rules. A scrapbook of cute photos with no link to a standard is not an assessment portfolio. Construct portfolios with a short list of standards, dated entries, and a caption that says what the piece shows.

Peer assessment asks children to give feedback using taught criteria ("Does the story have a beginning, middle, and end?"). It can strengthen audience awareness in Grade 2–4 writing. It is a terrible tool for grading individuals, for ranking friends, or for PreK children who are still learning that critique is about work, not about the person. Teach the checklist first; keep stakes low; never use peer scores as the sole report-card evidence.

Group assessment scores a team's product — a Grade 4 science poster, a kindergarten mural plan. Collaboration is a real standard-adjacent skill, but a group grade hides who cannot yet measure in centimeters. Select group assessment when the learning target is collaboration or a shared investigation, and add individual checks (an exit ticket, a conference, a photo of one child's contribution). On PECT, the distractor is often "the group got an A, so every child has mastered the standard."

CBA, CBM, CRT, and NRT — the formal core

Curriculum-based assessment (CBA) means you assess what you actually taught from the local curriculum and Pennsylvania standards, not a generic national skill list. A teacher-made fractions quiz aligned to this week's PA Core target is CBA. Curriculum-based measurement (CBM) is a specific, research-based subset: brief, repeatable probes sampled from the curriculum (for example, oral reading of grade-level passages, or math computation sheets) given under similar conditions so you can graph progress over weeks. Advantage: sensitive to small growth; useful for progress monitoring (Section 4.3). Limitation: a probe is not a full diagnostic of every subskill, and a flattened CBM line tells you to change instruction, not to keep giving the same probe and hoping.

A criterion-referenced test (CRT) compares the student to a predefined criterion — usually a standard or a cut score the teacher or state set for mastery, not to other children's ranks. "Can this child independently write a complete sentence with capital and period?" is a criterion question. Unit tests, standards checklists, many SAS-built items, and well-written rubrics live here. Advantage: they answer the question Pennsylvania teachers actually have after teaching a standard. Limitation: a poorly written CRT with trivial items will "pass" children who cannot do the authentic work.

A norm-referenced test (NRT) compares the student to a norm group (often a national sample) and reports standing as a percentile, stanine, or similar rank. Advantage: useful when the question is "How does this performance compare with a large reference group?" — some program-eligibility or large-scale questions. Limitation — and the exam trap: an NRT is not designed to tell you whether the child mastered this week's taught Pennsylvania standard. A child can sit at the 60th percentile nationally and still have not mastered equivalent fractions you just taught; another child can master the taught standard and still sit below a national mean. Do not use an NRT as a mastery measure of a taught standard. That is a CRT (or rubric/CBA) job.

Scenario: PreK blocks (observation + anecdotal notes)

Mr. Cole wants evidence of spatial language and persistence. He sets a two-column checklist, watches the block area for fifteen minutes, and writes: "11/3, Jordan: said 'under' and 'beside'; rebuilt the ramp three times after it collapsed." That is systematic observation plus an anecdotal note. Scoring Jordan with a published NRT of "school readiness" would not capture the ramp persistence the Early Learning Standards actually ask him to see.

Scenario: Grade 1 reading (running record + CBM)

Ms. Lee takes a running record on Monday (error analysis → short-vowel reteach) and a one-minute oral-reading CBM probe each Friday to graph words read correctly over time. The running record explains the errors; the CBM tracks whether instruction is moving the skill. Neither is an NRT.

Scenario: Grade 4 fractions (CRT vs NRT trap)

The principal asks whether students mastered equivalent fractions. Mr. Diaz should use a CRT or rubric aligned to the PA Core standard — a set of items and an authentic partitioning task with a mastery criterion. Pulling last year's national achievement-battery percentile is the PECT-wrong move: that NRT ranks against a norm group; it does not certify mastery of the taught standard.

How to select, construct, and use

  1. Name the purpose (screen, diagnose, formatively regroup, judge mastery, track weekly growth, compare to a norm group).
  2. Match the tool using the table above.
  3. Construct or select so the items or observation focus match the standard and the child's developmental level (a 40-item bubble test is a poor PreK tool; a one-item worksheet is a poor Grade 4 mastery check).
  4. Use the results for the purpose you named. Do not convert a peer checklist into a special-education referral, or an NRT percentile into a unit grade.

Exam traps for 0002.2

  • NRT to measure mastery of a taught standard. Use a CRT, rubric, or CBA.
  • Portfolio = junk drawer. Without criteria and dates, it is not assessment.
  • Group score = individual mastery.
  • Observation without a focus. Not systematic.
  • Peer assessment as a high-stakes grade for young children.
  • CBM as a once-a-year test. CBM's power is repeated probes.

If you can say, in one sentence, "I am using this tool to answer this question," you are inside 0002.2. If you cannot, you picked the tool because it was in the cabinet.

Loading diagram...
Pick CRT, NRT, or CBM from the question you actually have
Test Your Knowledge

A Grade 4 teacher wants to know whether students have mastered the taught Pennsylvania Core standard on equivalent fractions. Which tool is designed for that purpose?

A
B
C
D
Test Your Knowledge

Which statement correctly contrasts a curriculum-based measurement (CBM) probe with a one-time norm-referenced achievement test?

A
B
C
D
Test Your Knowledge

Ms. Rivera keeps dated work samples, photos of block structures with transcribed child comments, and two scored rubrics for each PreK child. The main advantage of this portfolio is:

A
B
C
D