6.3 Threats to Internal Validity & Social/External Validity
Key Takeaways
Internal validity reflects the extent to which observed changes in the dependent variable are directly and exclusively attributable to the independent variable rather than extraneous variables.
Extraneous variables represent uncontrolled ambient factors that add measurement noise, whereas confounding variables covary systematically with the independent variable, establishing plausible rival causal hypotheses.
The six classic threats to internal validity in behavior analysis comprise history, maturation, testing/reactivity, instrumentation (observer drift), statistical regression, and diffusion of treatment.
Montrose Wolf (1978) established that social validity must be evaluated across three essential dimensions: the social significance of target behaviors, the social appropriateness of procedures, and the social importance of behavioral outcomes.
External validity in behavior analysis is established not through statistical population sampling or inferential aggregation, but through systematic direct and conceptual replication across individuals, settings, and response classes.
The Architecture of Experimental Validity
In applied behavior analysis, experimental investigations and clinical evaluations are judged by the integrity and credibility of their experimental validity. Experimental validity answers two fundamental scientific questions: (1) "Did the intervention truly cause the observed behavior change?" (Internal Validity), and (2) "Will this intervention produce comparable outcomes with other individuals, in other settings, or with other behaviors?" (External Validity).
Additionally, because applied behavior analysis is dedicated to producing socially meaningful improvements in human lives, investigators must satisfy a third, equally vital standard: Social Validity. A study may possess airtight internal validity, but if the procedures are perceived as coercive, dehumanizing, or socially unacceptable, the intervention fails the fundamental standards of the discipline.
Defining Internal Validity and Scientific Control
Internal validity is the degree to which an experiment demonstrates that changes in the dependent variable (the target behavior) are a direct, causal function of the systematic manipulation of the independent variable, and were not the result of uncontrolled, extraneous, or confounding variables.
An experiment possesses high internal validity when the investigator can confidently declare that a true functional relation has been established. Conversely, when an experiment has low internal validity, plausible rival explanations exist: an observed improvement in a student's reading fluency could have resulted from a new medication, natural physical maturation, classroom curriculum changes, or biased observer scoring, rather than the behavior analyst's token economy.
Extraneous Variables vs. Confounding Variables
To master experimental methodology, BCaBAs must maintain a rigorous conceptual distinction between extraneous variables and confounding variables:
- Extraneous Variables: Any aspect of the experimental environment or participant characteristics that is not part of the independent variable, but could potentially affect the dependent variable if left unmonitored. Examples include room temperature, ambient fluorescent lighting, fluctuating noise levels in an adjacent hallway, mild fatigue, or minor variations in meal times. Extraneous variables typically introduce unsystematic "noise" or background variability into the data, making trends harder to detect, but they do not necessarily systematically bias the outcome.
- Confounding Variables: A specific, fatal subtype of extraneous variable that systematically covaries with the independent variable. When an uncontrolled variable shifts at the exact same moment that the independent variable is introduced, modified, or withdrawn, it provides a plausible rival hypothesis for the observed behavior change. Confounding variables completely destroy internal validity because the analyst cannot separate the effects of the intervention from the effects of the confound.
Clinical Vignette: The Medication Confound
Consider a clinical example frequently encountered in school-based practice. A BCaBA establishes a stable 10-day baseline of disruptive calling-out behavior (averaging 35 episodes per hour) for a third-grade student diagnosed with ADHD. On Monday of Week 3, the BCaBA introduces a differential reinforcement of other behavior (DRO) procedure paired with visual cue cards. Disruptive calling-out immediately plummets to 2 episodes per hour and remains near zero for the next two weeks.
Upon reviewing the data, the clinician is thrilled by the apparent success of the DRO program. However, during a team meeting, the student's mother casually reveals that over the preceding weekend, the child's pediatric psychiatrist initiated a trial of extended-release methylphenidate (a central nervous system stimulant for ADHD). In this scenario, the initiation of stimulant medication is a confounding variable. Because the medication adjustment coincided precisely with the introduction of the DRO protocol, the clinician cannot determine whether the behavior change was caused by the DRO, by the pharmacological enhancement of dopamine and norepinephrine, or by an interaction between the two. The internal validity of the behavioral intervention is completely compromised.
Six Classic Threats to Internal Validity in Behavior Analysis
Donald Campbell and Julian Stanley (1963) outlined the foundational threats to internal validity across scientific research. Within single-case experimental designs, behavior analysts must actively anticipate, recognize, and experimentally control for six classic threats:
1. History
A history threat refers to any environmental event occurring outside of the experimental arrangement during the course of the study that coincides with the introduction, alteration, or withdrawal of the independent variable. History involves events in the participant's broader ecological environment—such as a parent divorce, the death of a family pet, moving to a new home, a change in classroom teacher, the introduction of a new medical diet, or an adjustment in psychotropic medication.
- How SCED Controls History: In an ABAB reversal design, history is ruled out because an external historical event is extraordinarily unlikely to occur, reverse, and reoccur in exact synchronization with the four experimental phase shifts. In a multiple baseline design, history is controlled because when an external event occurs in the community, it should theoretically impact all tiers simultaneously; if untreated tiers remain stable while only the treated tier changes, history is ruled out.
2. Maturation
A maturation threat refers to biological, physiological, or neurological changes within the participant that occur naturally as a function of the passage of time, independent of experimental interventions. Maturation encompasses physical growth spurts, motor coordination development, hormonal shifts during puberty, the resolution of a physical injury, or temporary states such as acute physical fatigue or circadian sleepiness during long assessment sessions.
- How SCED Controls Maturation: Biological maturation is a continuous, gradual process. In contrast, behavioral interventions in SCED typically produce rapid, step-level shifts in responding that correspond precisely with phase transitions. Furthermore, maturational gains do not spontaneously deteriorate when an intervention is withdrawn in a reversal design; if behavior reverses upon withdrawal, maturation cannot account for the effect.
3. Testing and Observer Reactivity
Testing threats involve changes in a participant's responding that occur purely as a consequence of repeated exposure to the assessment procedures, measurement instrumentation, or test materials. Repeated testing can lead to practice effects (improved fluency due to repeated drills) or testing fatigue and extinction (declining motivation and performance due to unreinforced, repetitive questioning).
A primary manifestation of this threat in ABA is observer reactivity: the temporary alteration of a participant's behavior caused by their awareness that an observer is present, recording data, or pointing a video camera. When a behavior analyst enters a classroom holding a clipboard, students often sit up straight and reduce disruptive behavior simply because the observer is present. Over several sessions, the observer's presence becomes a neutral antecedent, and behavior returns to typical levels.
- How SCED Controls Testing/Reactivity: Behavior analysts control for reactivity by allowing observers to habituate to the setting (sitting quietly in the room without recording data until the participant ignores them), using unobtrusive data collection (video recording, permanent products), and employing multiple probe designs to minimize unreinforced exposure to test stimuli.
4. Instrumentation and Observer Drift
An instrumentation threat refers to unintended changes in the calibration, accuracy, or standards of the measurement system over time, rather than a genuine change in the participant's behavior. Instrumentation threats arise from equipment breakdown (e.g., drifting electronic timers, faulty counter clickers) or human observer error.
The most pervasive instrumentation threat in behavior analysis is observer drift. Observer drift is the systematic, unintended shift over time in how an observer interprets and applies an operational definition. Over weeks or months of data collection, an observer may unconsciously become more lenient (overlooking borderline infractions), become more stringent (recording behaviors that do not fit the original definition), or develop idiosyncratic interpretations that drift away from the original training criteria. Observer drift creates the false appearance of a behavior change when, in reality, only the observer's scoring standards have shifted.
- How SCED Controls Instrumentation: Behavior analysts control for observer drift by: (1) conducting frequent, scheduled Interobserver Agreement (IOA) assessments (targeting agreement across at least 20% to 33% of sessions), (2) holding regular observer retraining sessions where observers re-score standardized benchmark video samples, and (3) utilizing blind observers who are unaware of the study's current phase or experimental hypotheses.
5. Statistical Regression to the Mean
Statistical regression to the mean is the mathematical tendency for extreme, outlier scores on an initial measurement to move closer to the true population or individual mean upon subsequent measurements, purely as a function of random measurement error.
In applied settings, clients are almost always referred for behavioral services on days when their problem behavior reaches an uncharacteristically severe peak (e.g., a student is referred on the single day they emit 60 aggressive hits, even though their true average is 15 hits per day). If a clinician collects an immediate 1-day baseline and introduces an intervention, the behavior will naturally drop toward the mean over subsequent sessions, creating the false illusion of treatment efficacy.
- How SCED Controls Regression: The steady-state strategy directly neutralizes regression to the mean. By requiring an extended baseline that demonstrates steady-state responding across multiple consecutive sessions, single-case methodology washes out single-session outliers, ensuring that the intervention is evaluated against a true, stable baseline level.
6. Diffusion of Treatment
Diffusion of treatment (treatment contamination) occurs when the independent variable is inadvertently, prematurely, or accidentally delivered during baseline phases, to untreated control tiers, or to untreated participants.
For example, in a multiple baseline design across classrooms, Teacher A in Classroom 1 receives training on a proactive token economy. Teacher A becomes enthusiastic about the system and describes the procedures to Teacher B in Classroom 2 during lunch. Teacher B immediately begins praising students and implementing informal tokens in Classroom 2 while Classroom 2 is still supposed to be under baseline conditions. When Classroom 2's baseline data shift prematurely, experimental control across tiers is destroyed.
- How SCED Controls Diffusion: Mitigating diffusion requires strict treatment integrity (procedural fidelity) monitoring, using precise procedural checklists, training staff not to share materials across settings, keeping intervention materials secure, and utilizing independent reliability observers.
Threats to Internal Validity Matrix
The following matrix summarizes the six primary threats to internal validity, their operational definitions, their confounding mechanisms, and the single-case experimental strategies used to prevent or control them:
| Threat to Internal Validity | Definition & Operational Mechanism | How It Confounds Results | Method of Prevention / Control in SCED |
|---|---|---|---|
| History | External environmental events occurring outside the experiment that coincide with the introduction or removal of the IV. | The historical event provides a plausible rival explanation for the observed change in the dependent variable. | Staggered implementation in Multiple Baseline Designs; repeated phase reversals in ABAB designs. |
| Maturation | Natural biological, neurological, or physiological growth within the participant over the course of the study. | Behavior change may be due to natural development or fatigue rather than the independent variable. | Continuous time-series measurement; steady-state responding; rapid step-level shifts coinciding with phase changes. |
| Testing & Reactivity | Changes in responding caused by repeated exposure to test materials or awareness of being observed (observer reactivity). | The participant's behavior alters due to observation awareness or test practice, masking true operant rates. | Observer habituation periods; unobtrusive data collection (permanent products); multiple probe designs. |
| Instrumentation & Observer Drift | Changes in the measurement system over time, particularly observers shifting their interpretation of operational definitions. | Apparent changes in behavioral level reflect shifting observer standards rather than true client behavior change. | Frequent Interobserver Agreement (IOA) checks (); periodic observer retraining on video benchmarks; blind observers. |
| Statistical Regression to the Mean | Extreme, outlier baseline scores naturally drift closer to the central tendency on repeated measurements. | A natural statistical return to the mean is falsely attributed to the therapeutic efficacy of the intervention. | Applying the steady-state strategy; requiring extended multi-session baselines rather than single-point baselines. |
| Diffusion of Treatment | Accidental or premature delivery of the independent variable during baseline phases or to untreated tiers/settings. | Untreated baselines shift prematurely, destroying the opportunity to demonstrate verification and replication. | Explicit staff training; strict procedural fidelity checklists; securing intervention materials; monitoring integrity. |
Social Validity: Montrose Wolf (1978)
In 1978, Montrose Wolf published a landmark article titled "Social Validity: The Case for Subjective Measurement or How Applied Behavior Analysis Is Finding Its Heart." Wolf argued that applied behavior analysis cannot evaluate its success solely through internal validity, graphs, and statistical percentages. Because ABA is an applied human science, the ultimate value of our technology depends upon its social validity—the extent to which consumers, caregivers, and society value the work.
Wolf established that social validity must be rigorously evaluated across three distinct levels:
[ THREE LEVELS OF SOCIAL VALIDITY ]
(Montrose Wolf, 1978)
|
+-----------------------------------+-----------------------------------+
| | |
[ SOCIAL SIGNIFICANCE ] [ SOCIAL APPROPRIATENESS ] [ SOCIAL IMPORTANCE ]
OF GOALS OF PROCEDURES OF OUTCOMES
Are the target behaviors Are the intervention methods Did the behavior change
what consumers really need? acceptable, humane, & ethical? make a real-world difference?
1. The Social Significance of Target Behaviors (Goals)
- Core Question: "Are the specific behavioral goals really what society, the client, and their significant others want and need to change?"
- Operational Standard: The selection of target behaviors must prioritize habilitation—maximizing the individual's long-term access to reinforcers and minimizing access to punishers. Targeting trivial compliance behaviors (e.g., forcing an autistic child to sit with hands folded in their lap for 45 minutes) lacks social significance, whereas teaching functional communication to replace dangerous aggression possesses profound social significance.
- Assessment Methods: Administering consumer preference surveys to clients and families, conducting ecological assessments of typical community expectations, and comparing client baselines to normative peer performance.
2. The Social Appropriateness of Procedures (Interventions)
- Core Question: "Do the participants, caregivers, and implementers find the intervention procedures acceptable, humane, dignified, and culturally responsive?"
- Operational Standard: Even if an intervention is 100% effective in suppressing problem behavior, it is invalid if the procedures are perceived as excessively punitive, restrictive, painful, or humiliating. Practitioners must evaluate the cost-benefit ratio, staff burden, and ethical acceptability of the intervention. Reinforcement-based procedures (DRA, FCT, token economies) almost universally possess higher social acceptability than punishment-based procedures (overcorrection, contingent exercise, response cost).
- Assessment Methods: Administering standardized treatment acceptability rating scales (e.g., the Treatment Acceptability Rating Form-Revised [TARF-R]), interviewing direct-care staff and consumers, and employing concurrent-chains preference assessments where the client is given a direct choice between experiencing Intervention A or Intervention B.
3. The Social Importance of Behavior Changes (Outcomes / Effects)
- Core Question: "Did the quantitative change in behavior make a genuine, recognizable, real-world difference in the client's everyday functioning and quality of life?"
- Operational Standard: An intervention may achieve statistical significance () or an 80% reduction on a graph, but if the client still cannot attend general education classes, hold a community job, or live safely without continuous physical restraint, the clinical outcome lacks social validity. Conversely, a modest 30% reduction that allows a teenager to remain in their family home rather than being institutionalized possesses immense social importance.
- Assessment Methods:
- Social Comparison Method: Comparing the participant's post-intervention performance against a normative sample of typically developing peers in the same natural environment. If the participant's rate falls within the normative peer range, outcome validity is confirmed.
- Subjective Evaluation Method: Soliciting structured ratings and qualitative feedback from key stakeholders (parents, teachers, employers, peers) who interact daily with the client, asking whether they perceive a meaningful, positive difference in the individual's life.
External Validity in Behavior Analysis
External validity refers to the degree to which an experimental finding, intervention effect, or functional relation can be generalized across other participants, settings, interventionists, and response classes.
Contrasting SCED External Validity with Group Inferential Statistics
A frequent source of confusion on certification examinations is how external validity is established in behavior analysis compared to between-subject group research:
- Group Research Model: Traditional psychology and education attempt to achieve external validity through random population sampling. Researchers randomly select 100 individuals from a broader population, apply inferential statistics, and mathematically generalize the sample mean back to the population. However, because group research averages performance, it cannot predict how any single individual will respond.
- Behavior Analytic Model: Behavior analysis rejects the premise that external validity can be established through a single group experiment. As Murray Sidman (1960) articulated, external validity in behavior analysis is established exclusively through replication. A single single-case study with three participants does not claim universal generality; rather, generality is proven empirically across successive experiments.
Two Levels of Replication Establishing External Validity
- Direct Replication: The investigator repeats the exact experimental protocol with the same participant (intra-subject direct replication) or across identical participants in the same setting (inter-subject direct replication). Direct replication demonstrates the reliability of the behavior-change technology.
- Systematic Replication: The investigator intentionally varies one or more non-critical parameters of the original experiment—such as testing the intervention with a different age group, a different diagnostic population, in a different clinical setting (e.g., home vs. clinic vs. school), using different therapists, or targeting different response topographies—while holding the core independent variable constant. When an intervention is repeatedly validated across dozens of systematic replications by independent research teams, robust external validity is undeniably established.
Common BCaBA Exam Traps: Validity and Evaluation
- Trap 1: Confusing Confounding Variables with Extraneous Variables: An extraneous variable is any uncontrolled environmental factor that adds noise. A confounding variable is an extraneous variable that covaries systematically with the independent variable, creating a specific rival causal explanation.
- Trap 2: Confusing Observer Drift with Observer Reactivity: Observer drift is a change in the observer's scoring standards over time (an instrumentation threat). Observer reactivity is a change in the participant's behavior caused by being observed (a testing threat).
- Trap 3: Assuming High Internal Validity Guarantees Social Validity: A research design may demonstrate flawless experimental control with zero confounds (), but if the intervention relies on harsh physical restraint that caregivers refuse to implement at home, it possesses zero social validity.
- Trap 4: Believing Single-Case Designs Have No External Validity: Single-case designs do not lack external validity; they establish external validity through systematic direct and conceptual replication across time and settings, rather than through mathematical inferential sampling.
A BCaBA begins implementing a visual activity schedule and token reinforcement system to increase academic on-task behavior for a student with ADHD. Baseline data showed on-task behavior stable at 25%. On the exact Monday that the token system is introduced, the student's parents inform the school that the child's physician doubled their daily dose of stimulant medication over the weekend. Within two days, on-task behavior jumps to 85%. Why does this scenario undermine the internal validity of the behavioral intervention?
The medication change is a confounding variable that coincided with the independent variable, so the cause of the change cannot be isolated.
The visual schedule produced a testing effect that sensitized the student to the classroom observer's presence.
The abrupt improvement represents statistical regression to the mean resulting from an uncharacteristically low baseline period and observer bias.
The medication adjustment represents an instrumentation threat, because the measurement system's calibration was altered by the teacher.
A behavior analyst conducts a six-month evaluation of an intervention designed to reduce vocal stereotypy in a classroom. Over the course of the study, interobserver agreement (IOA) between the primary data collector and the reliability observer gradually drops from 92% to 68%. A review of video recordings reveals that the primary observer gradually expanded their interpretation of the operational definition to count quiet throat-clearing and audible humming, which were explicitly excluded in the original measurement protocol. This breakdown in internal validity represents which threat?
Diffusion of treatment, because the intervention was accidentally delivered during baseline phases.
Observer drift, because the observer's use of the operational definition gradually shifted from the original standard.
Maturation, because the student's vocal tract developed physically across the six-month evaluation period.
Observer reactivity, because the student changed their vocal behavior after realizing that a second observer was recording data.
A behavior analyst implements an extinction and physical redirection procedure that successfully reduces a teenage client's public motor stereotypy (flapping hands while jumping) by 95% in an analog clinic setting. However, when the client's parents and high school teachers complete a follow-up assessment, they express deep distress: the client appears visibly agitated, the physical redirection caused bruising on the client's wrists, the client now avoids therapists, and the adolescent still cannot successfully order food or navigate the school cafeteria. According to Montrose Wolf's (1978) framework, which critical dimension of validity does this intervention primarily lack?
Construct validity, because motor stereotypy was measured using discontinuous partial-interval recording.
Internal validity, because the 95% reduction cannot be replicated across multiple baseline tiers.
Nonparametric validity, because the intervention failed to evaluate the presence versus absence of reinforcement across separate phases.
Social validity, because the procedures were unacceptable and the outcomes did not meaningfully improve the client's quality of life.
Sections you finish are checked off in the contents.