9.5 Evaluating Research Quality & Basic Statistics
Key Takeaways
Measurement scales are nominal, ordinal, interval, and ratio; the scale determines which statistics are appropriate.
The coefficient of determination (r squared) is the proportion of variance two variables share; r = .50 means 25% shared variance.
A Type I error is a false positive (rejecting a true null hypothesis); a Type II error is a false negative (missing a real effect).
Statistical power increases with larger samples, larger effects, and a less strict alpha level.
Evidence-based practice combines the best research evidence with professional expertise and the characteristics, values, and context of the student and family.
Why Research Literacy Is Tested
The ETS outline's Research and Evidence-Based Practice section asks candidates to know how to evaluate research quality and interpret outcomes, how to determine the relevance of research and apply it to practice, types of research designs and basic statistics, and how to analyze, interpret, and use research-based and evidence-based practices at the individual, group, and systems levels. Section 9.3 covers designs, validity threats, effect sizes, and the What Works Clearinghouse. This section covers the statistics underneath them and a practical way to judge studies.
Scales of Measurement
| Scale | Properties | Examples | Appropriate Statistics |
|---|---|---|---|
| Nominal | Categories with no order | Disability category, school, yes/no | Mode, frequencies, chi-square |
| Ordinal | Ordered categories with unequal intervals | Percentile ranks, class rank, Likert ratings (strictly) | Median, percentiles, Spearman correlation |
| Interval | Equal intervals, no true zero | Standard scores, T-scores, temperature in Fahrenheit | Mean, standard deviation, Pearson correlation |
| Ratio | Equal intervals with a true zero | Words read correctly per minute, number of office referrals, minutes on task | All of the above, and ratios ("twice as many") |
Percentile ranks are ordinal (Section 2.2), which is why they should not be averaged or subtracted.
Descriptive Statistics
- Central tendency: the mean (average), median (middle score), and mode (most frequent score). In a skewed distribution, the mean is pulled toward the tail; the median is the better summary for skewed data such as household income or referral counts.
- Variability: the range (highest minus lowest), variance, and standard deviation (average distance from the mean). Two classes with the same mean can have very different spreads.
- Distribution shape: skewness and kurtosis (Section 2.2).
Correlation and Regression
The Pearson correlation coefficient (r) ranges from −1.00 to +1.00. The sign shows direction; the absolute value shows strength.
- Coefficient of determination (r squared): the proportion of variance two variables share. An r of .50 means .25, or 25%, shared variance; an r of .30 means only 9%.
- Correlation does not show causation. A third variable may cause both, or the direction may be reversed. For example, homework time and grades may both reflect family resources.
- Restriction of range: correlations shrink when a sample covers only part of the range, such as a gifted-only group.
- Spearman's rho is used for ordinal data.
- Regression uses one or more predictors to estimate an outcome (for example, predicting third-grade reading from kindergarten screening), and its accuracy is described by the standard error of estimate (Section 2.1).
Hypothesis Testing, Errors, and Power
A study tests a null hypothesis (no effect or no difference) against an alternative hypothesis. The researcher sets alpha (often .05), the risk of a false positive they will accept.
| Null Hypothesis Is Actually True | Null Hypothesis Is Actually False | |
|---|---|---|
| Reject the null | Type I error (false positive), probability = alpha | Correct decision (power) |
| Fail to reject the null | Correct decision | Type II error (false negative), probability = beta |
Statistical power (1 − beta) is the probability of detecting a real effect. Power increases with:
- A larger sample size,
- A larger true effect,
- A less strict alpha (for example, .05 rather than .01), and
- More reliable measures (less measurement error).
Small studies often lack power, so a nonsignificant result from a small sample does not prove an intervention has no effect.
Common Statistical Tests
| Test | Use | Example |
|---|---|---|
| t-test | Compare two means | Treatment versus control reading scores |
| Analysis of variance (ANOVA) | Compare three or more means | Three intervention groups |
| Analysis of covariance (ANCOVA) | Compare means while statistically controlling a covariate | Posttest scores controlling for pretest (common in quasi-experiments) |
| Chi-square | Compare frequencies in categories | Suspension rates by group |
| Correlation and regression | Relationships and prediction | Screening scores predicting later outcomes |
Confidence intervals around a result show its precision. A wide interval signals uncertainty even when the result is statistically significant. Always pair p-values with effect sizes (Section 9.3).
Meta-Analysis
A meta-analysis statistically combines effect sizes from many studies. It gives a more stable estimate than any single study but can be distorted by publication bias (studies with null results are less likely to be published) and by combining very different studies.
Evaluating Research Quality
Use questions like these when reading a study:
| Question | What to Check |
|---|---|
| Design | Randomized experiment, quasi-experiment, single-case design, correlational study, or case report? Stronger designs support causal claims (Section 9.3). |
| Sample | Size, how participants were selected, attrition, and whether they resemble your students (age, language, disability, setting). |
| Measures | Reliability and validity of outcome measures; whether outcomes matter educationally, not only on researcher-made tests. |
| Implementation | Was fidelity measured? Was the intervention described well enough to replicate? |
| Results | Statistical significance and effect size, confidence intervals, and whether effects lasted at follow-up. |
| Replication | Have independent researchers found similar results? |
| Conflicts of interest | Did the program developer or publisher fund or run the study? |
| Peer review | Published in a peer-reviewed journal, or only in promotional materials? |
A rough hierarchy of evidence for intervention effectiveness runs from systematic reviews and meta-analyses of well-designed experiments, to single well-designed randomized trials and strong single-case design series, to quasi-experiments, to correlational studies, to case reports and expert opinion. Qualitative research serves different purposes, such as understanding experiences and implementation; it is judged by credibility and transferability rather than by statistical power.
Applying Research to Practice
Evidence-based practice integrates three elements: the best available research, professional expertise, and the student's and family's characteristics, values, and context. To decide whether a study applies locally:
- Relevance: Were the students, setting, and problem similar to yours?
- Feasibility: Do you have the time, training, materials, and staff the intervention requires?
- Fit: Does it match the school's needs, culture, and existing systems (Hexagon Tool, Section 7.3)?
- Core components: What must be kept for fidelity, and what can be adapted?
- Local evaluation: Monitor progress and outcomes to confirm the practice works here (practice-based evidence).
Useful clearinghouses and registries include the What Works Clearinghouse, the National Center on Intensive Intervention's tools charts for screening, progress monitoring, and interventions, Blueprints for Healthy Youth Development, and CASEL's program guide. Registries use different criteria, so read how each one rates evidence.
At each level of service:
- Individual: single-case design logic with baseline and progress monitoring (Sections 1.2 and 9.3).
- Group: pre- and post-measures with comparison groups when possible.
- System: program evaluation with logic models and fidelity data (Section 7.3).
A study reports a correlation of r = .50 between a kindergarten screening score and third-grade reading achievement. What proportion of the variance in third-grade reading is shared with the screening score?
50%
25%
5%
100%, because the correlation is statistically significant
A small pilot study of a new social skills program finds no statistically significant difference between the program group (n = 8) and a comparison group (n = 8), although the program group improved more. Which conclusion is most appropriate?
The program has been proven ineffective and should be abandoned.
The study may have lacked statistical power, so a real effect could have been missed (a possible Type II error); examine the effect size and consider a larger study.
The study made a Type I error by finding a false positive.
The program must be effective because the program group improved more.
A district reports the "average" number of office discipline referrals per student. Most students have 0 or 1 referral, but a few have more than 20. Which statistic best represents the typical student?
The mean, because it uses every score
The median, because it is less affected by the extreme scores in this skewed distribution
The range, because it shows the highest score
The standard deviation, because it describes the average student
Sections you finish are checked off in the contents.