9.5 Evaluating Research Quality & Basic Statistics

Key Takeaways

  • Measurement scales are nominal, ordinal, interval, and ratio; the scale determines which statistics are appropriate.

  • The coefficient of determination (r squared) is the proportion of variance two variables share; r = .50 means 25% shared variance.

  • A Type I error is a false positive (rejecting a true null hypothesis); a Type II error is a false negative (missing a real effect).

  • Statistical power increases with larger samples, larger effects, and a less strict alpha level.

  • Evidence-based practice combines the best research evidence with professional expertise and the characteristics, values, and context of the student and family.

Last updated: September 2026

Why Research Literacy Is Tested

The ETS outline's Research and Evidence-Based Practice section asks candidates to know how to evaluate research quality and interpret outcomes, how to determine the relevance of research and apply it to practice, types of research designs and basic statistics, and how to analyze, interpret, and use research-based and evidence-based practices at the individual, group, and systems levels. Section 9.3 covers designs, validity threats, effect sizes, and the What Works Clearinghouse. This section covers the statistics underneath them and a practical way to judge studies.

Scales of Measurement

ScalePropertiesExamplesAppropriate Statistics
NominalCategories with no orderDisability category, school, yes/noMode, frequencies, chi-square
OrdinalOrdered categories with unequal intervalsPercentile ranks, class rank, Likert ratings (strictly)Median, percentiles, Spearman correlation
IntervalEqual intervals, no true zeroStandard scores, T-scores, temperature in FahrenheitMean, standard deviation, Pearson correlation
RatioEqual intervals with a true zeroWords read correctly per minute, number of office referrals, minutes on taskAll of the above, and ratios ("twice as many")

Percentile ranks are ordinal (Section 2.2), which is why they should not be averaged or subtracted.

Descriptive Statistics

  • Central tendency: the mean (average), median (middle score), and mode (most frequent score). In a skewed distribution, the mean is pulled toward the tail; the median is the better summary for skewed data such as household income or referral counts.
  • Variability: the range (highest minus lowest), variance, and standard deviation (average distance from the mean). Two classes with the same mean can have very different spreads.
  • Distribution shape: skewness and kurtosis (Section 2.2).

Correlation and Regression

The Pearson correlation coefficient (r) ranges from −1.00 to +1.00. The sign shows direction; the absolute value shows strength.

  • Coefficient of determination (r squared): the proportion of variance two variables share. An r of .50 means .25, or 25%, shared variance; an r of .30 means only 9%.
  • Correlation does not show causation. A third variable may cause both, or the direction may be reversed. For example, homework time and grades may both reflect family resources.
  • Restriction of range: correlations shrink when a sample covers only part of the range, such as a gifted-only group.
  • Spearman's rho is used for ordinal data.
  • Regression uses one or more predictors to estimate an outcome (for example, predicting third-grade reading from kindergarten screening), and its accuracy is described by the standard error of estimate (Section 2.1).

Hypothesis Testing, Errors, and Power

A study tests a null hypothesis (no effect or no difference) against an alternative hypothesis. The researcher sets alpha (often .05), the risk of a false positive they will accept.

Null Hypothesis Is Actually TrueNull Hypothesis Is Actually False
Reject the nullType I error (false positive), probability = alphaCorrect decision (power)
Fail to reject the nullCorrect decisionType II error (false negative), probability = beta

Statistical power (1 − beta) is the probability of detecting a real effect. Power increases with:

  • A larger sample size,
  • A larger true effect,
  • A less strict alpha (for example, .05 rather than .01), and
  • More reliable measures (less measurement error).

Small studies often lack power, so a nonsignificant result from a small sample does not prove an intervention has no effect.

Common Statistical Tests

TestUseExample
t-testCompare two meansTreatment versus control reading scores
Analysis of variance (ANOVA)Compare three or more meansThree intervention groups
Analysis of covariance (ANCOVA)Compare means while statistically controlling a covariatePosttest scores controlling for pretest (common in quasi-experiments)
Chi-squareCompare frequencies in categoriesSuspension rates by group
Correlation and regressionRelationships and predictionScreening scores predicting later outcomes

Confidence intervals around a result show its precision. A wide interval signals uncertainty even when the result is statistically significant. Always pair p-values with effect sizes (Section 9.3).

Meta-Analysis

A meta-analysis statistically combines effect sizes from many studies. It gives a more stable estimate than any single study but can be distorted by publication bias (studies with null results are less likely to be published) and by combining very different studies.

Evaluating Research Quality

Use questions like these when reading a study:

QuestionWhat to Check
DesignRandomized experiment, quasi-experiment, single-case design, correlational study, or case report? Stronger designs support causal claims (Section 9.3).
SampleSize, how participants were selected, attrition, and whether they resemble your students (age, language, disability, setting).
MeasuresReliability and validity of outcome measures; whether outcomes matter educationally, not only on researcher-made tests.
ImplementationWas fidelity measured? Was the intervention described well enough to replicate?
ResultsStatistical significance and effect size, confidence intervals, and whether effects lasted at follow-up.
ReplicationHave independent researchers found similar results?
Conflicts of interestDid the program developer or publisher fund or run the study?
Peer reviewPublished in a peer-reviewed journal, or only in promotional materials?

A rough hierarchy of evidence for intervention effectiveness runs from systematic reviews and meta-analyses of well-designed experiments, to single well-designed randomized trials and strong single-case design series, to quasi-experiments, to correlational studies, to case reports and expert opinion. Qualitative research serves different purposes, such as understanding experiences and implementation; it is judged by credibility and transferability rather than by statistical power.

Applying Research to Practice

Evidence-based practice integrates three elements: the best available research, professional expertise, and the student's and family's characteristics, values, and context. To decide whether a study applies locally:

  1. Relevance: Were the students, setting, and problem similar to yours?
  2. Feasibility: Do you have the time, training, materials, and staff the intervention requires?
  3. Fit: Does it match the school's needs, culture, and existing systems (Hexagon Tool, Section 7.3)?
  4. Core components: What must be kept for fidelity, and what can be adapted?
  5. Local evaluation: Monitor progress and outcomes to confirm the practice works here (practice-based evidence).

Useful clearinghouses and registries include the What Works Clearinghouse, the National Center on Intensive Intervention's tools charts for screening, progress monitoring, and interventions, Blueprints for Healthy Youth Development, and CASEL's program guide. Registries use different criteria, so read how each one rates evidence.

At each level of service:

  • Individual: single-case design logic with baseline and progress monitoring (Sections 1.2 and 9.3).
  • Group: pre- and post-measures with comparison groups when possible.
  • System: program evaluation with logic models and fidelity data (Section 7.3).
Loading diagram...
Type I and Type II Errors
Test Your Knowledge

A study reports a correlation of r = .50 between a kindergarten screening score and third-grade reading achievement. What proportion of the variance in third-grade reading is shared with the screening score?

A

50%

B

25%

C

5%

D

100%, because the correlation is statistically significant

Test Your Knowledge

A small pilot study of a new social skills program finds no statistically significant difference between the program group (n = 8) and a comparison group (n = 8), although the program group improved more. Which conclusion is most appropriate?

A

The program has been proven ineffective and should be abandoned.

B

The study may have lacked statistical power, so a real effect could have been missed (a possible Type II error); examine the effect size and consider a larger study.

C

The study made a Type I error by finding a false positive.

D

The program must be effective because the program group improved more.

Test Your Knowledge

A district reports the "average" number of office discipline referrals per student. Most students have 0 or 1 referral, but a few have more than 20. Which statistic best represents the typical student?

A

The mean, because it uses every score

B

The median, because it is less affected by the extreme scores in this skewed distribution

C

The range, because it shows the highest score

D

The standard deviation, because it describes the average student

Sections you finish are checked off in the contents.