9.3 Research Designs, Statistics, & Evidence-Based Evaluation
Key Takeaways
Single-Case Experimental Designs (SCED) allow participants to serve as their own controls; experimental control is demonstrated through repeated measurement, systematic manipulation of the independent variable, and replication of effects.
While the ABAB reversal design rigorously establishes functional relations, it cannot be ethically applied to severe or dangerous behaviors, nor to irreversible learned skills (e.g., reading decoding).
Multiple Baseline designs (across participants, settings, or behaviors) demonstrate experimental control by staggering intervention introduction without requiring a withdrawal phase, making them the gold-standard SCED in schools.
Statistical significance (p < .05) only indicates that an observed effect is unlikely due to chance under the null hypothesis; practical significance requires calculating effect sizes such as Cohen's d, Hedges' g, or Percentage of Non-Overlapping Data (PND).
Under What Works Clearinghouse (WWC) evidence standards, well-executed Randomized Controlled Trials (RCTs) with low attrition can receive the highest rating ('Meets Standards Without Reservations'), whereas quasi-experimental designs can at best achieve 'Meets Standards With Reservations'.
Research Designs, Statistics, & Evidence-Based Evaluation
School psychologists function as local scientists and clinical evaluators who design, implement, and interpret educational research (NASP Practice Model Domain 9: Research and Evidence-Based Practice). Practitioners must critically evaluate peer-reviewed research, assess the validity of published curricular programs, implement single-case experimental designs to monitor individualized interventions, and compute effect sizes to determine real-world clinical significance. Navigating this domain requires mastery of experimental controls, validity threats, psychometric metrics, and federal research clearinghouse standards.
Single-Case Experimental Designs (SCED)
Single-Case Experimental Designs (SCED)—also called single-subject designs or N-of-1 studies—are empirical methodologies in which the individual participant serves as their own control. Rather than comparing group averages, SCED relies on repeated, continuous measurement of a target behavior across time to demonstrate a functional relation between the independent variable (the intervention) and the dependent variable (the student's behavior or academic skill).
Single-Case Experimental Designs
│
┌────────────────────────┬───────────┴────────────┬────────────────────────┐
▼ ▼ ▼ ▼
ABAB REVERSAL MULTIPLE BASELINE ALTERNATING TREATMENTS CHANGING CRITERION
• A1: Baseline • Across Participants • Rapid alternation of • Stepwise criteria
• B1: Intervention • Across Settings 2+ interventions • Gradual skill
• A2: Withdrawal • Across Behaviors • No baseline required shaping
• B2: Reinstatement • NO withdrawal needed • Fast comparative data • Evaluates rate
• Ethical/learning • Preferred school • Multi-treatment or reduction
limits apply design interference risk
1. ABAB Reversal / Withdrawal Design
- Structure:
- Phase A₁ (Baseline): Repeated baseline data collected until stability is achieved (minimum 3–5 points).
- Phase B₁ (Intervention): Introduction of the independent variable; data tracked until a distinct trend or level shift emerges.
- Phase A₂ (Withdrawal/Reversal): The intervention is systematically withdrawn, returning environmental conditions to baseline.
- Phase B₂ (Reinstatement): Reintroduction of the intervention.
- Demonstrating Experimental Control: A functional relation is verified when the target behavior systematically changes only during intervention phases (B₁ and B₂) and reverts toward baseline levels during the withdrawal phase (A₂).
- Critical Limitations (The Two Fatal Flaws of ABAB in Schools):
- Ethical Prohibition: It is clinically and ethically unacceptable to withdraw an intervention for severe, destructive, or dangerous behaviors (e.g., self-injurious head-banging, physical aggression, fire-setting).
- Irreversibility / Non-Reversible Learning: Skills acquired through instruction cannot be "unlearned" upon withdrawal. If a student is taught to decode consonant-vowel-consonant words or solve two-digit multiplication algorithms, removing the instructional prompt will not cause the student to unlearn the cognitive skill.
2. Multiple Baseline Designs
The Multiple Baseline Design is the primary, gold-standard SCED utilized in educational and school-based settings because it demonstrates experimental control without requiring a withdrawal phase.
- Mechanism: The researcher establishes two or more independent baselines concurrently. The intervention is introduced to the first baseline while all other baselines remain in their pre-intervention state. Once a treatment effect is established in the first baseline, the intervention is introduced in a staggered fashion to the second baseline, and subsequently to the third.
- Demonstrating Experimental Control: Control is established when each baseline changes if and only if the intervention is introduced, while untreated baselines continue their stable, pre-intervention trajectory.
- Three Variations:
- Across Participants: The same intervention is applied to the same target behavior across 3 or more distinct students in the same setting (staggered across time).
- Across Settings: The same intervention is applied to 1 student with the same target behavior across 3 distinct environments (e.g., math class, physical education, cafeteria).
- Across Behaviors: The same intervention is applied to 1 student across 3 distinct, functionally independent behaviors in the same environment (e.g., on-task attention, hand-raising, assignment completion).
3. Alternating Treatments Design (ATD) / Multi-Element Design
- Mechanism: Rapid, random, or semi-random alternation of two or more distinct intervention conditions within the same time period (e.g., Intervention A in the morning, Intervention B in the afternoon, counterbalanced across days).
- Advantages: Does not require an extended baseline; does not require withdrawal; allows rapid, direct comparison of two competing interventions to determine relative clinical efficacy.
- Vulnerability: Susceptible to multiple-treatment interference (the effects of Intervention A bleeding into or carrying over to Intervention B).
4. Changing Criterion Design
- Mechanism: Following baseline, the intervention is introduced with a specific performance criterion for reinforcement. Once performance stabilizes at that criterion, the criterion is adjusted in stepwise increments (either increasing or decreasing).
- Utility: Highly effective for shaping gradual behavioral changes or adjusting quantitative performance rates (e.g., increasing minutes of sustained reading from 5 to 10 to 15 to 20 minutes, or reducing daily vocal out-calls from 12 to 8 to 5 to 1).
- Demonstrating Control: Experimental control is demonstrated when the student's empirical performance closely tracks each stepwise criterion shift.
Group Experimental & Quasi-Experimental Research Designs
When evaluating systems-level programs or broad curricula, researchers utilize group designs.
Group Research Designs
│
┌────────────────────────────┴────────────────────────────┐
▼ ▼
TRUE EXPERIMENTAL DESIGNS QUASI-EXPERIMENTAL DESIGNS
• Randomized Controlled Trial (RCT) • Non-Equivalent Control Groups
• Random Assignment (R) to conditions • Intact Classrooms / Schools
• High Internal Validity • High Selection Bias Risk
• Controls for known & unknown confounders • Requires Baseline Equivalence checks
1. Randomized Controlled Trials (RCTs)
- The Gold Standard for Causal Inference: Participants are randomly assigned (R) to either an experimental treatment group or a control/comparison group (e.g., business-as-usual or active control).
- Mechanism: Random assignment equates groups on both known and unknown confounding variables at baseline, isolating the independent variable as the sole cause of observed posttest differences.
2. Quasi-Experimental Designs (QEDs)
- Characteristics: Evaluates treatment and comparison groups without random assignment.
- Educational Reality: Schools rarely permit individual students to be randomly assigned to experimental conditions. Instead, researchers use intact units (e.g., Classroom A receives the new phonics curriculum; Classroom B continues core instruction).
- Vulnerability: Vulnerable to selection bias and pre-existing group differences. Researchers must demonstrate baseline equivalence (documenting that groups did not differ significantly on pretests) and statistically control for covariates using ANCOVA or propensity score matching.
Threats to Validity: Donald Campbell & Julian Stanley Framework
Understanding methodological rigor requires analyzing threats that undermine study conclusions.
Threats to Internal Validity
Internal validity is the degree to which changes in the dependent variable can be unequivocally attributed to the independent variable, rather than extraneous confounding factors.
| Threat to Internal Validity | Psychometric Definition | Concrete School-Based Example |
|---|---|---|
| History | Specific external events occurring between pretest and posttest outside the researcher's control. | A school adopts a schoolwide digital reading app at home during a targeted phonics intervention study. |
| Maturation | Biological, physiological, or psychological growth within participants resulting from the passage of time. | First-grade students improve in motor control and attention simply because they grew 8 months older during the study. |
| Testing (Practice Effects) | Improved performance on posttests resulting from repeated prior exposure to the test instrument itself. | A student achieves higher scores on a standardized math test solely due to test-taking familiarity from the pretest. |
| Instrumentation | Shifts in measurement tools, calibration decay, or changes in observer scoring rubrics over time. | Behavioral observers become more lenient or fatigued in their rating standards during week 12 compared to week 1. |
| Statistical Regression (Regression to the Mean) | Tendency for extreme scores (very high or very low) at baseline to naturally move closer to the mean on retesting due to measurement error. | Selecting students who scored in the bottom 2nd percentile on a single screening day; their retest scores improve even without effective intervention. |
| Selection Bias | Systematic differences in pre-existing participant characteristics between treatment and control groups. | Comparing an after-school voluntary tutoring group (highly motivated families) to a non-attending comparison group. |
| Attrition / Mortality | Differential, non-random dropout of participants across study conditions during the intervention period. | The lowest-performing 30% of students in the experimental group drop out due to frustration, falsely inflating the posttest mean. |
Threats to External Validity
External validity represents the extent to which research findings can be generalized to other populations, educational settings, treatment providers, and outcome measures.
- Hawthorne Effect (Reactivity): Participants alter their performance or behavior simply because they are aware they are being observed.
- Multiple-Treatment Interference: Cumulative effects of prior interventions influencing responsiveness to a new intervention.
- Novelty Effect: Initial enthusiasm for an intervention producing short-term gains that fade once the novelty dissipates.
Statistical Metrics & Effect Size Calculations
School psychologists must look beyond simple p-values to evaluate the substantive educational importance of research outcomes.
1. Statistical Significance (p-value) vs. Practical Significance
- Statistical Significance (p < .05): If the null hypothesis (no true difference) were true, a difference at least this large would occur less than 5% of the time. Limitation: With large sample sizes (N > 1,000), even tiny, educationally meaningless differences achieve p < .001.
- Practical / Clinical Significance: The real-world magnitude, meaningfulness, and functional utility of the intervention effect, quantified via effect sizes.
2. Group Design Effect Sizes: Cohen's d & Hedges' g
Cohen's d expresses the difference between two group means in standard deviation units:
Where s(pooled) is the pooled standard deviation across groups:
Cohen's Empirical Benchmarks:
- d = 0.20: Small effect (noticeable only through statistical analysis; minimal real-world classroom impact).
- d = 0.50: Medium effect (visible to the naked eye of an experienced educator; meaningful educational gain).
- d = 0.80: Large effect (grossly observable, transformative educational intervention).
Psychometric Nuance (Hedges' g): When study sample sizes are small (N < 20), Cohen's d systematically overestimates the true population effect size. Researchers apply Hedges' g, which incorporates a small-sample correction factor (J = 1 − 3 ÷ (4df − 1)) to eliminate this upward bias.
3. Single-Case Effect Size: Percentage of Non-Overlapping Data (PND)
The Percentage of Non-Overlapping Data (PND) (Scruggs & Mastropieri, 1998) is the most widely cited non-parametric effect size metric for single-case experimental designs.
Worked PND Example (Words Correct Per Minute; behavior to increase)
Baseline (5 points): 42 44 41 46 48 Highest baseline point = 48
Intervention (7 points): 48 51 54 58 62 64 66
Points strictly above 48: 51, 54, 58, 62, 64, 66 = 6 points
The first intervention point (48) ties the baseline maximum, so it does NOT count.
PND = 6 / 7 x 100 = 85.7% (the "effective" range, 70%-90%)
Procedural Calculation Rules:
- For a behavior to increase (e.g., reading fluency, on-task behavior): Identify the single highest data point in the baseline phase. Count how many data points in the intervention phase fall strictly above this baseline extreme. Divide by the total number of intervention data points and multiply by 100.
- For a behavior to decrease (e.g., aggression, disruptive out-calls): Identify the single lowest data point in the baseline phase. Count how many intervention data points fall strictly below this baseline extreme. Divide by total intervention points and multiply by 100.
Scruggs & Mastropieri PND Interpretation Standards:
- PND > 90%: Highly Effective Intervention.
- PND = 70% to 90%: Moderately Effective Intervention.
- PND = 50% to 70%: Questionable / Mild Effectiveness.
- PND < 50%: Ineffective Intervention.
What Works Clearinghouse (WWC) Evidence Standards
Established by the U.S. Department of Education's Institute of Education Sciences (IES), the What Works Clearinghouse (WWC) provides the definitive federal framework for vetting educational interventions under the Every Student Succeeds Act (ESSA).
| WWC Rating Category | Essential Methodological Criteria | Highest Attainable ESSA Tier |
|---|---|---|
| Meets WWC Standards Without Reservations | Well-implemented Randomized Controlled Trials (RCTs); demonstrated low overall and differential sample attrition; no fatal confounding factors. | Tier 1: Strong Evidence (when the study also shows a statistically significant positive effect) |
| Meets WWC Standards With Reservations | Rigorous Quasi-Experimental Designs (QEDs) that establish baseline equivalence; or RCTs with high attrition that successfully demonstrate post-attrition baseline equivalence. | Tier 2: Moderate Evidence |
| Does Not Meet WWC Standards | Studies featuring fatal confounders (e.g., one teacher delivering treatment and one delivering control), unequivalent baseline groups without statistical controls, or severe instrumentation failure. | Not Tier 1 or 2; well-designed correlational studies with statistical controls can reach Tier 3: Promising Evidence, and a well-specified logic model with ongoing evaluation meets Tier 4: Demonstrates a Rationale |
A school psychologist designs an intervention to reduce a second-grade student's severe self-injurious head-banging behavior. The practitioner needs to demonstrate experimental control to verify that the replacement sensory intervention is functionally responsible for the behavioral reduction. Which research design is most ethically and clinically appropriate?
An ABAB reversal design, as withdrawing the sensory intervention will conclusively prove behavioral causation.
A pretest-posttest control group design comparing the student to randomly selected classroom peers.
A randomized controlled trial across multiple school districts.
A multiple baseline design across settings (e.g., classroom, occupational therapy clinic, playground).
A school psychologist tracks a student's on-task behavior during an intensive Tier 3 academic intervention. Across 5 baseline sessions, the student's on-task percentages are 20%, 25%, 22%, 30%, and 28%. Following the introduction of a self-monitoring intervention, the psychologist collects 10 intervention sessions: 32%, 35%, 34%, 38%, 42%, 40%, 45%, 48%, 50%, and 52%. What is the Percentage of Non-Overlapping Data (PND), and what does it indicate about the intervention's efficacy?
PND = 70.0%; indicates the self-monitoring intervention is moderately effective.
PND = 100.0%; indicates the self-monitoring intervention is highly effective.
PND = 30.0%; indicates the self-monitoring intervention is ineffective.
PND = 50.0%; indicates questionable or mild intervention efficacy.
A school district pilot-tests a proprietary remedial reading program by selecting the 20 elementary students with the absolute lowest reading scores in the district (all scoring below the 2nd percentile on a single fall screening). After 6 weeks of computer-based practice, the students are retested and show an average gain of 12 percentile points (p < .01). The district leadership concludes the software is highly effective. As a scientific problem solver, which threat to internal validity should the school psychologist highlight as the primary confound?
Instrumentation decay, because the software algorithm altered scoring standards between pretest and posttest.
Hawthorne effect, because the students were aware they were participating in an experimental study.
Attrition, because high-performing students dropped out of the remedial program.
Statistical regression to the mean, because selecting participants based on extreme initial scores guarantees an upward shift upon retesting due to measurement error.
Sections you finish are checked off in the contents.