6.6 Evaluating Research Evidence and Applying Data-Informed Practice
Key Takeaways
Competency 5, Skill 2 requires educators to analyze and apply data-informed research to improve instruction and student achievement.
Correlation does not establish causation; only a design that controls for alternative explanations, such as a randomized or well-matched comparison, supports a causal claim about an instructional practice.
Effect size expresses the magnitude of an intervention's impact in standard deviation units, which allows comparison across studies that used different measures.
The Every Student Succeeds Act defines four evidence tiers, from strong evidence based on well-implemented experimental studies down to research-based rationale, and Tier 4 is a rationale rather than a finding.
Practitioner claims should be checked for population match, implementation conditions, comparison group, and publication source before a practice is adopted for a specific group of students.
6.6 Evaluating Research Evidence and Applying Data-Informed Practice
Competency 5, Skill 2 asks educators to "analyze and apply data-informed research to improve instruction and student achievement." Section 6.1 covered action research — inquiry the teacher conducts in their own classroom. This section covers the complementary skill: appraising research that someone else conducted and deciding whether it should change your practice.
1. Correlation, Causation, and the Third Variable
| Claim | Design that supports it | Design that does not |
|---|---|---|
| "Students who read more score higher." | Correlational — describes an association | Any causal claim about what reading does |
| "This reading program raises scores." | Randomized or well-matched comparison with a control group | A single-group before-and-after comparison |
| "This practice works for our students." | Replication in a similar population and setting | A study in a very different population |
The third-variable problem: students who read more may also have more books at home, quieter study space, and more educated caregivers. Any of those could produce the score difference. A correlation is real information — it tells you where to look — but it cannot tell you what to change.
Common confounds in education research:
- Selection effects: students or teachers who volunteer for a program differ from those who do not.
- Novelty effects: short-term gains from anything new, which fade.
- Implementation differences: the study's version of the practice was more faithful than yours will be.
- Regression to the mean: the lowest-scoring group improves on retest regardless of intervention.
- Teacher effects: the enthusiastic developer taught the treatment classes.
2. Effect Size: How Big Is the Effect?
Statistical significance answers whether an effect is likely to be non-zero. It does not say whether the effect is large enough to matter. Effect size answers that.
| Effect size (Cohen's d) | Conventional label | Practical reading |
|---|---|---|
| 0.20 | Small | Detectable, may still be worthwhile if cheap and easy |
| 0.50 | Medium | Meaningful classroom difference |
| 0.80 and above | Large | Substantial difference |
Two cautions the exam rewards:
- A very large sample can produce a statistically significant result with a trivially small effect size.
- Effect sizes must be compared within similar contexts. An effect measured on a narrow researcher-designed test is typically larger than the same practice measured on a broad standardized assessment.
3. The Every Student Succeeds Act Evidence Tiers
Federal law defines four tiers that districts use when selecting interventions, particularly for school improvement funding.
+-----------------------------------------------------------------------------+
| ESSA LEVELS OF EVIDENCE |
| |
| TIER 1 - STRONG EVIDENCE |
| At least one well-designed, well-implemented experimental |
| (randomized controlled) study showing a favorable effect. |
| |
| TIER 2 - MODERATE EVIDENCE |
| At least one well-designed, well-implemented quasi-experimental |
| study showing a favorable effect. |
| |
| TIER 3 - PROMISING EVIDENCE |
| At least one well-designed correlational study with statistical |
| controls for selection bias. |
| |
| TIER 4 - DEMONSTRATES A RATIONALE |
| A well-specified logic model informed by research or evaluation, |
| with an effort under way to study the effects. |
+-----------------------------------------------------------------------------+
Important
Tier 4 is not a finding. "Demonstrates a rationale" means there is a plausible theory of action and a plan to study it — no evidence of effect yet. An FTCE distractor that treats a Tier 4 designation as proof that a program works is incorrect.
Useful clearinghouses for checking a claim include the federal What Works Clearinghouse, the National Center on Intensive Intervention, and Evidence for ESSA.
4. Appraising a Research Claim Before You Act
| Question | Why it matters | Red flag |
|---|---|---|
| Who was studied? | Findings may not transfer across grade band, language proficiency, or disability status | "Proven for all learners" |
| Compared with what? | A treatment beats no instruction easily; beating strong existing instruction is the real test | No comparison group described |
| How was it implemented? | Dosage, duration, group size, and trainer expertise often exceed what your school can provide | Implementation details omitted |
| How big was the effect? | Significance without magnitude is not actionable | Only p-values reported |
| Who paid for and conducted it? | Vendor-conducted studies of vendor products warrant more scrutiny | The only evidence is on the vendor's own site |
| Has it replicated? | Single studies overturn regularly | "Groundbreaking new study proves" |
| What is the cost, in time and money? | Opportunity cost is real; time spent here is time not spent elsewhere | Cost never mentioned |
5. Practices With Consistently Strong Evidence
These recur across syntheses and are safe defaults when a stem asks for a research-supported practice:
| Practice | What it is | Why it works |
|---|---|---|
| Retrieval practice | Low-stakes recall from memory rather than review | Strengthens the retrieval pathway itself |
| Spaced practice | Distributing study across time | Reduces forgetting more than the same total time massed |
| Interleaving | Mixing problem types rather than blocking them | Forces discrimination among strategies |
| Formative assessment with feedback | Frequent checks that change instruction | Closes the gap while it is still closable |
| Explicit instruction for novel content | Modeling and guided practice before independence | Manages cognitive load for novices |
| Metacognitive strategy instruction | Teaching students to plan, monitor, and evaluate | Transfers across subjects |
| Reciprocal teaching and comprehension strategy instruction | Structured strategy routines | Large effects for struggling readers |
6. Turning Research Into a Local Decision
- Name the problem with data. Which students, which standard, how far behind?
- Search for evidence at the right tier, checking a clearinghouse before a vendor site.
- Check population and implementation match against your actual constraints.
- Pilot with a comparison. Two similar classes, one practice, same assessment.
- Measure implementation fidelity, not just outcomes; a practice never delivered has not been tested.
- Decide with the team, and document what you would accept as evidence of failure before you begin.
A vendor advertises that its reading program 'demonstrates a rationale' under the Every Student Succeeds Act evidence framework. What does this designation actually indicate?
The program has a well-specified logic model informed by research and an effort under way to study its effects, but no demonstrated favorable effect yet.
The program has been shown effective in at least one well-designed randomized controlled trial.
The program has been shown effective in at least one quasi-experimental study.
The program has been shown effective in a correlational study that controlled for selection bias.
A study reports that students who participated in an after-school tutoring program scored significantly higher than students who did not participate. Which limitation most threatens a causal interpretation of this result?
The reported effect size was 0.60, which is only a medium effect.
Participation was voluntary, so students and families who opted in may differ systematically from those who did not, which is a selection effect rather than an effect of tutoring.
The study was conducted over a full school year rather than a single semester.
The outcome was measured with a standardized assessment rather than a program-specific test.
A study of 40,000 students finds a statistically significant advantage for a new instructional app, with an effect size of 0.04. How should a teacher interpret this finding?
The result is trustworthy and the app should be adopted, because the sample size is unusually large.
The result should be dismissed entirely, because statistical significance is never meaningful in education research.
The effect is statistically detectable but practically trivial, so the app's cost in time and money should be weighed against alternatives with larger effects.
The result cannot be interpreted without knowing the p-value, which determines the size of the effect.
Sections you finish are checked off in the contents.