7.2 Interpreting Formal & Informal Assessment Results

Key Takeaways

  • Formal, norm-referenced scores come in several forms — standard scores, percentile ranks, stanines, and grade/age equivalents — and each communicates a different kind of comparison; grade/age equivalents are the most frequently misinterpreted and should never be treated as a target grade placement
  • Because every score carries measurement error (the standard error of measurement introduced in Section 6.3), scores should be interpreted as a band or range rather than a single exact point, especially when comparing a pre-test to a post-test
  • Informal assessment data — running records, curriculum-based probes, rubric-scored work samples, structured interviews, and observation — supplies the instructional detail formal scores cannot, showing specifically which skills, strategies, or conditions are working
  • Present Levels of Academic and Functional Performance (PLAAFP) are built by triangulating formal and informal data together, and measurable annual IEP goals should trace directly back to the specific skill gaps that triangulated data reveals
  • Common interpretation pitfalls include treating a grade-equivalent score as an instructional placement, making high-stakes decisions from a single score, and comparing scores from different tests or different norm groups as if they were directly equivalent
Last updated: August 2026

7.2 Interpreting Formal & Informal Assessment Results

Quick Answer: Interpreting assessment data well means matching the right score type to the right question. Formal, norm-referenced scores — standard scores, percentile ranks, stanines, and grade/age equivalents — each say something different, and grade/age equivalents are the score type most often misread as a placement recommendation, which it is not. Every score carries measurement error, so it should be read as a band, not a single exact point. Informal data — running records, curriculum-based probes, rubrics, interviews, and observation — fills in the instructional detail formal scores cannot provide. Combining, or triangulating, formal and informal data is what produces an accurate present level of performance and a measurable IEP goal that actually targets the student's real skill gap.

From Scores to Decisions: Why Interpretation Is a Skill

Collecting an assessment score is easy; interpreting it correctly to inform instruction and IEP goals is the actual professional skill this competency tests. Two teachers can look at the identical score report and reach very different — and not equally correct — conclusions, because one understands what the number represents and one does not. The FTCE ESE exam is far more likely to present you with a data point or a short data table and ask what it means, or what a team should do next, than to ask you to define a term in isolation.

Reading Formal, Norm-Referenced Score Types

Building on the norm-referenced versus criterion-referenced distinction from Section 6.3, norm-referenced tests report scores in several different forms, and each answers a slightly different question:

Score TypeWhat It Tells YouCommon Misread
Standard scoreDistance from the mean of the norm group, in standardized units (e.g., mean of 100, standard deviation of 15)Treating small point differences near a cutoff as a meaningful, guaranteed real difference
Percentile rankPercentage of the norm group the student scored at or aboveConfusing a percentile rank with "percent correct" — a percentile rank of 60 does not mean the student got 60% of items right
StanineA 1–9 band grouping percentile ranges, with 5 being averageTreating a one-stanine difference between two testings as automatically meaningful growth
Grade or age equivalentThe grade/age level at which the average student earns this raw scoreAssuming a fourth grader with a "7.2 grade equivalent" reading score can handle seventh-grade text, or should be moved to a higher grade

Grade and age equivalents deserve special caution because they are the type most often over-interpreted by families and even by educators outside special education. A grade equivalent describes where the raw score falls relative to the average performance of typically developing students at that grade — it does not describe the specific skills the student has or has not mastered, and it says nothing about whether grade-level content at that equivalent would actually be appropriate or accessible for the student. An exam item describing a parent (or a teacher) who insists a student reading at a "6.5 grade equivalent" should be placed in sixth-grade content is describing a classic misinterpretation of this score type.

Scores as Bands, Not Points

Section 6.3 introduced the standard error of measurement (SEM) as the conceptual margin of error around any observed score. That concept becomes directly actionable here: when interpreting a single score, or comparing a pre-test score to a post-test score, the appropriate question is not "did the number go up?" but "did the number move by more than would be expected from ordinary measurement error alone?" Two scores that differ by only a few points, especially on a test with a wide margin of error, may not represent a real change in the student's true skill level at all. This is exactly why eligibility and progress decisions should never hinge on one score sitting a single point above or below a cut point — the safer, more defensible interpretation treats every score as a band and looks for a pattern of evidence, not one number in isolation.

Interpreting Informal Assessment Data

Informal assessment — running records, curriculum-based measurement probes, rubric-scored writing or project samples, structured interviews with the student or family, and direct observation — does not produce the same standardized comparison to a norm group, but it supplies something formal scores cannot: specific, actionable instructional detail.

  • A running record of oral reading does not just produce an accuracy percentage; it shows which kinds of errors a student makes (meaning-based substitutions versus visual/phonetic confusions), which points directly to the next instructional strategy.
  • A rubric-scored writing sample shows exactly which trait — organization, elaboration, conventions — is holding a score down, information a single "writing composite" standard score cannot provide.
  • A structured interview with a family or a student can surface conditions under which a behavior or skill is stronger or weaker (time of day, setting, task type) that no standardized test session would ever reveal.
  • Direct observation across multiple natural settings shows whether a skill demonstrated in a one-on-one testing room actually shows up during independent classroom work — a question formal testing alone cannot answer.

The interpretive skill with informal data is pattern recognition: looking across several samples or observations for a consistent instructional signal, rather than treating any single informal data point as conclusive on its own.

Triangulating Data into Present Levels of Performance

A Present Level of Academic and Functional Performance (PLAAFP), sometimes referred to as present levels of performance (PLOP), is not a copy-paste of the most recent formal test score. It is a synthesis built by triangulating formal scores, informal classroom data, and observation/input from teachers and family into a coherent, specific description of what the student can currently do, under what conditions, and where the gap to grade-level or functional expectations actually lies. A present level built from a single data source — one standardized test score, or one teacher's informal impression — is professionally and legally thin; a present level built by triangulating multiple sources is both more accurate and more defensible if ever challenged.

From Present Levels to Measurable Annual Goals

A measurable annual IEP goal should be traceable directly back to a specific gap identified in the triangulated present level — not to a generic disability category or a boilerplate goal bank entry. If the present level, built from CBM data and classroom work samples, shows a student decoding multisyllabic words inaccurately while comprehension of orally-presented material is age-appropriate, the resulting goal should target multisyllabic decoding specifically, with a measurable criterion and timeline, rather than a vague "will improve reading" statement disconnected from what the data actually showed.

Common Interpretation Pitfalls

  • Grade-equivalent overreach — treating a grade/age equivalent as an instructional placement recommendation rather than a norm-comparison statistic
  • Single-score decisions — using one formal score, especially near a cut point, as if it carries no measurement error
  • Cross-test comparison — comparing scores from two different tests, or the same test with two different norm groups, as though they are on an identical scale
  • Ignoring context — interpreting a formal score without considering language proficiency, sensory access, testing conditions, or cultural/background factors that might affect performance independent of the construct being measured
  • Overlooking informal data entirely — relying only on standardized numbers when the classroom-level informal evidence tells a fuller, sometimes contradictory, story

Bringing It Together

On exam day, when a scenario hands you a score or a small data table, resist jumping to "good" or "bad." Ask what type of score this is, what comparison it actually supports, whether it should be read as a band rather than a point, and what informal data would need to be triangulated with it before a sound instructional or IEP decision could be made. That sequence — identify the score type, respect its margin of error, and triangulate with other evidence — is the interpretive habit the exam is built to reward.

Test Your Knowledge

A fourth-grade student earns a grade-equivalent score of 7.2 on a reading test, and a parent asks that the student be placed in seventh-grade reading material. What is the correct response to this request?

A
B
C
D
Test Your Knowledge

A student's pre-test and post-test standard scores differ by only two points on a test with a wide standard error of measurement. What is the most appropriate interpretation of this difference?

A
B
C
D
Test Your Knowledge

A teacher writes an IEP present level of performance using only the student's most recent standardized reading composite score, with no reference to classroom work samples, running records, or teacher observation. What is the main weakness of this present level?

A
B
C
D