6.6 Evaluating Research Evidence and Applying Data-Informed Practice

Key Takeaways

  • Competency 5, Skill 2 requires educators to analyze and apply data-informed research to improve instruction and student achievement.
  • Correlation does not establish causation; only a design that controls for alternative explanations, such as a randomized or well-matched comparison, supports a causal claim about an instructional practice.
  • Effect size expresses the magnitude of an intervention's impact in standard deviation units, which allows comparison across studies that used different measures.
  • The Every Student Succeeds Act defines four evidence tiers, from strong evidence based on well-implemented experimental studies down to research-based rationale, and Tier 4 is a rationale rather than a finding.
  • Practitioner claims should be checked for population match, implementation conditions, comparison group, and publication source before a practice is adopted for a specific group of students.
Last updated: August 2026

6.6 Evaluating Research Evidence and Applying Data-Informed Practice

Competency 5, Skill 2 asks educators to "analyze and apply data-informed research to improve instruction and student achievement." Section 6.1 covered action research — inquiry the teacher conducts in their own classroom. This section covers the complementary skill: appraising research that someone else conducted and deciding whether it should change your practice.


1. Correlation, Causation, and the Third Variable

ClaimDesign that supports itDesign that does not
"Students who read more score higher."Correlational — describes an associationAny causal claim about what reading does
"This reading program raises scores."Randomized or well-matched comparison with a control groupA single-group before-and-after comparison
"This practice works for our students."Replication in a similar population and settingA study in a very different population

The third-variable problem: students who read more may also have more books at home, quieter study space, and more educated caregivers. Any of those could produce the score difference. A correlation is real information — it tells you where to look — but it cannot tell you what to change.

Common confounds in education research:

  • Selection effects: students or teachers who volunteer for a program differ from those who do not.
  • Novelty effects: short-term gains from anything new, which fade.
  • Implementation differences: the study's version of the practice was more faithful than yours will be.
  • Regression to the mean: the lowest-scoring group improves on retest regardless of intervention.
  • Teacher effects: the enthusiastic developer taught the treatment classes.

2. Effect Size: How Big Is the Effect?

Statistical significance answers whether an effect is likely to be non-zero. It does not say whether the effect is large enough to matter. Effect size answers that.

Effect size (Cohen's d)Conventional labelPractical reading
0.20SmallDetectable, may still be worthwhile if cheap and easy
0.50MediumMeaningful classroom difference
0.80 and aboveLargeSubstantial difference

Two cautions the exam rewards:

  • A very large sample can produce a statistically significant result with a trivially small effect size.
  • Effect sizes must be compared within similar contexts. An effect measured on a narrow researcher-designed test is typically larger than the same practice measured on a broad standardized assessment.

3. The Every Student Succeeds Act Evidence Tiers

Federal law defines four tiers that districts use when selecting interventions, particularly for school improvement funding.

+-----------------------------------------------------------------------------+
|                   ESSA LEVELS OF EVIDENCE                                   |
|                                                                             |
|   TIER 1 - STRONG EVIDENCE                                                  |
|            At least one well-designed, well-implemented experimental        |
|            (randomized controlled) study showing a favorable effect.        |
|                                                                             |
|   TIER 2 - MODERATE EVIDENCE                                                |
|            At least one well-designed, well-implemented quasi-experimental  |
|            study showing a favorable effect.                                |
|                                                                             |
|   TIER 3 - PROMISING EVIDENCE                                               |
|            At least one well-designed correlational study with statistical  |
|            controls for selection bias.                                     |
|                                                                             |
|   TIER 4 - DEMONSTRATES A RATIONALE                                         |
|            A well-specified logic model informed by research or evaluation,  |
|            with an effort under way to study the effects.                   |
+-----------------------------------------------------------------------------+

[!IMPORTANT] Tier 4 is not a finding. "Demonstrates a rationale" means there is a plausible theory of action and a plan to study it — no evidence of effect yet. An FTCE distractor that treats a Tier 4 designation as proof that a program works is incorrect.

Useful clearinghouses for checking a claim include the federal What Works Clearinghouse, the National Center on Intensive Intervention, and Evidence for ESSA.


4. Appraising a Research Claim Before You Act

QuestionWhy it mattersRed flag
Who was studied?Findings may not transfer across grade band, language proficiency, or disability status"Proven for all learners"
Compared with what?A treatment beats no instruction easily; beating strong existing instruction is the real testNo comparison group described
How was it implemented?Dosage, duration, group size, and trainer expertise often exceed what your school can provideImplementation details omitted
How big was the effect?Significance without magnitude is not actionableOnly p-values reported
Who paid for and conducted it?Vendor-conducted studies of vendor products warrant more scrutinyThe only evidence is on the vendor's own site
Has it replicated?Single studies overturn regularly"Groundbreaking new study proves"
What is the cost, in time and money?Opportunity cost is real; time spent here is time not spent elsewhereCost never mentioned

5. Practices With Consistently Strong Evidence

These recur across syntheses and are safe defaults when a stem asks for a research-supported practice:

PracticeWhat it isWhy it works
Retrieval practiceLow-stakes recall from memory rather than reviewStrengthens the retrieval pathway itself
Spaced practiceDistributing study across timeReduces forgetting more than the same total time massed
InterleavingMixing problem types rather than blocking themForces discrimination among strategies
Formative assessment with feedbackFrequent checks that change instructionCloses the gap while it is still closable
Explicit instruction for novel contentModeling and guided practice before independenceManages cognitive load for novices
Metacognitive strategy instructionTeaching students to plan, monitor, and evaluateTransfers across subjects
Reciprocal teaching and comprehension strategy instructionStructured strategy routinesLarge effects for struggling readers

6. Turning Research Into a Local Decision

  1. Name the problem with data. Which students, which standard, how far behind?
  2. Search for evidence at the right tier, checking a clearinghouse before a vendor site.
  3. Check population and implementation match against your actual constraints.
  4. Pilot with a comparison. Two similar classes, one practice, same assessment.
  5. Measure implementation fidelity, not just outcomes; a practice never delivered has not been tested.
  6. Decide with the team, and document what you would accept as evidence of failure before you begin.
Test Your Knowledge

A vendor advertises that its reading program 'demonstrates a rationale' under the Every Student Succeeds Act evidence framework. What does this designation actually indicate?

A
B
C
D
Test Your Knowledge

A study reports that students who participated in an after-school tutoring program scored significantly higher than students who did not participate. Which limitation most threatens a causal interpretation of this result?

A
B
C
D
Test Your Knowledge

A study of 40,000 students finds a statistically significant advantage for a new instructional app, with an effect size of 0.04. How should a teacher interpret this finding?

A
B
C
D