6.6 Evaluating Research Evidence and Applying Data-Informed Practice
Key Takeaways
- Competency 5, Skill 2 requires educators to analyze and apply data-informed research to improve instruction and student achievement.
- Correlation does not establish causation; only a design that controls for alternative explanations, such as a randomized or well-matched comparison, supports a causal claim about an instructional practice.
- Effect size expresses the magnitude of an intervention's impact in standard deviation units, which allows comparison across studies that used different measures.
- The Every Student Succeeds Act defines four evidence tiers, from strong evidence based on well-implemented experimental studies down to research-based rationale, and Tier 4 is a rationale rather than a finding.
- Practitioner claims should be checked for population match, implementation conditions, comparison group, and publication source before a practice is adopted for a specific group of students.
6.6 Evaluating Research Evidence and Applying Data-Informed Practice
Competency 5, Skill 2 asks educators to "analyze and apply data-informed research to improve instruction and student achievement." Section 6.1 covered action research — inquiry the teacher conducts in their own classroom. This section covers the complementary skill: appraising research that someone else conducted and deciding whether it should change your practice.
1. Correlation, Causation, and the Third Variable
| Claim | Design that supports it | Design that does not |
|---|---|---|
| "Students who read more score higher." | Correlational — describes an association | Any causal claim about what reading does |
| "This reading program raises scores." | Randomized or well-matched comparison with a control group | A single-group before-and-after comparison |
| "This practice works for our students." | Replication in a similar population and setting | A study in a very different population |
The third-variable problem: students who read more may also have more books at home, quieter study space, and more educated caregivers. Any of those could produce the score difference. A correlation is real information — it tells you where to look — but it cannot tell you what to change.
Common confounds in education research:
- Selection effects: students or teachers who volunteer for a program differ from those who do not.
- Novelty effects: short-term gains from anything new, which fade.
- Implementation differences: the study's version of the practice was more faithful than yours will be.
- Regression to the mean: the lowest-scoring group improves on retest regardless of intervention.
- Teacher effects: the enthusiastic developer taught the treatment classes.
2. Effect Size: How Big Is the Effect?
Statistical significance answers whether an effect is likely to be non-zero. It does not say whether the effect is large enough to matter. Effect size answers that.
| Effect size (Cohen's d) | Conventional label | Practical reading |
|---|---|---|
| 0.20 | Small | Detectable, may still be worthwhile if cheap and easy |
| 0.50 | Medium | Meaningful classroom difference |
| 0.80 and above | Large | Substantial difference |
Two cautions the exam rewards:
- A very large sample can produce a statistically significant result with a trivially small effect size.
- Effect sizes must be compared within similar contexts. An effect measured on a narrow researcher-designed test is typically larger than the same practice measured on a broad standardized assessment.
3. The Every Student Succeeds Act Evidence Tiers
Federal law defines four tiers that districts use when selecting interventions, particularly for school improvement funding.
+-----------------------------------------------------------------------------+
| ESSA LEVELS OF EVIDENCE |
| |
| TIER 1 - STRONG EVIDENCE |
| At least one well-designed, well-implemented experimental |
| (randomized controlled) study showing a favorable effect. |
| |
| TIER 2 - MODERATE EVIDENCE |
| At least one well-designed, well-implemented quasi-experimental |
| study showing a favorable effect. |
| |
| TIER 3 - PROMISING EVIDENCE |
| At least one well-designed correlational study with statistical |
| controls for selection bias. |
| |
| TIER 4 - DEMONSTRATES A RATIONALE |
| A well-specified logic model informed by research or evaluation, |
| with an effort under way to study the effects. |
+-----------------------------------------------------------------------------+
[!IMPORTANT] Tier 4 is not a finding. "Demonstrates a rationale" means there is a plausible theory of action and a plan to study it — no evidence of effect yet. An FTCE distractor that treats a Tier 4 designation as proof that a program works is incorrect.
Useful clearinghouses for checking a claim include the federal What Works Clearinghouse, the National Center on Intensive Intervention, and Evidence for ESSA.
4. Appraising a Research Claim Before You Act
| Question | Why it matters | Red flag |
|---|---|---|
| Who was studied? | Findings may not transfer across grade band, language proficiency, or disability status | "Proven for all learners" |
| Compared with what? | A treatment beats no instruction easily; beating strong existing instruction is the real test | No comparison group described |
| How was it implemented? | Dosage, duration, group size, and trainer expertise often exceed what your school can provide | Implementation details omitted |
| How big was the effect? | Significance without magnitude is not actionable | Only p-values reported |
| Who paid for and conducted it? | Vendor-conducted studies of vendor products warrant more scrutiny | The only evidence is on the vendor's own site |
| Has it replicated? | Single studies overturn regularly | "Groundbreaking new study proves" |
| What is the cost, in time and money? | Opportunity cost is real; time spent here is time not spent elsewhere | Cost never mentioned |
5. Practices With Consistently Strong Evidence
These recur across syntheses and are safe defaults when a stem asks for a research-supported practice:
| Practice | What it is | Why it works |
|---|---|---|
| Retrieval practice | Low-stakes recall from memory rather than review | Strengthens the retrieval pathway itself |
| Spaced practice | Distributing study across time | Reduces forgetting more than the same total time massed |
| Interleaving | Mixing problem types rather than blocking them | Forces discrimination among strategies |
| Formative assessment with feedback | Frequent checks that change instruction | Closes the gap while it is still closable |
| Explicit instruction for novel content | Modeling and guided practice before independence | Manages cognitive load for novices |
| Metacognitive strategy instruction | Teaching students to plan, monitor, and evaluate | Transfers across subjects |
| Reciprocal teaching and comprehension strategy instruction | Structured strategy routines | Large effects for struggling readers |
6. Turning Research Into a Local Decision
- Name the problem with data. Which students, which standard, how far behind?
- Search for evidence at the right tier, checking a clearinghouse before a vendor site.
- Check population and implementation match against your actual constraints.
- Pilot with a comparison. Two similar classes, one practice, same assessment.
- Measure implementation fidelity, not just outcomes; a practice never delivered has not been tested.
- Decide with the team, and document what you would accept as evidence of failure before you begin.
A vendor advertises that its reading program 'demonstrates a rationale' under the Every Student Succeeds Act evidence framework. What does this designation actually indicate?
A study reports that students who participated in an after-school tutoring program scored significantly higher than students who did not participate. Which limitation most threatens a causal interpretation of this result?
A study of 40,000 students finds a statistically significant advantage for a new instructional app, with an effect size of 0.04. How should a teacher interpret this finding?