6.2 Core Single-Case Designs: Reversal, Multiple Baseline, Alternating Treatments, and Changing Criterion

Key Takeaways

  • The ABAB reversal design provides the gold standard demonstration of experimental control through prediction, verification, and replication, but is strictly contraindicated for irreversible repertoires and severe, life-threatening target behaviors.

  • The Multiple Baseline Design demonstrates experimental control across staggered independent tiers (behaviors, settings, participants) without requiring a withdrawal phase, verified when untreated tiers remain stable while the treated tier changes.

  • The Multiple Probe Design strategically substitutes intermittent probes for continuous daily baseline recording, eliminating testing fatigue, extinction effects, and observer reactivity during prolonged baselines.

  • The Alternating Treatments Design (ATD) allows rapid comparison of two or more distinct interventions within the same phase using distinct discriminative stimuli, eliminating baseline/withdrawal requirements while being susceptible to multi-treatment interference.

  • The Changing Criterion Design demonstrates experimental control when performance closely tracks successive, stepwise criterion shifts, strengthened by bidirectional criterion reversals and varied step lengths/magnitudes.

Last updated: October 2026

Overview of Core Single-Case Experimental Designs

Applied behavior analysts rely on four primary single-case experimental design families to establish functional relations between independent variables and socially significant behavior change: the reversal (withdrawal) design, the multiple baseline design, the alternating treatments design (multi-element design), and the changing criterion design. Each design possesses unique baseline logic mechanisms, operational requirements, clinical strengths, and ethical boundaries.

Selecting an appropriate design requires balancing scientific rigor (maximizing internal validity) against ethical mandates (protecting client safety and providing effective treatment without unnecessary delays or harmful withdrawals).


1. Reversal / Withdrawal Designs (ABAB and Variations)

The reversal design—conventionally designated as an ABAB design—is the classic and most powerful experimental architecture in behavior analysis for demonstrating a functional relation between an environmental manipulation and a target behavior.

The Baseline Logic of the ABAB Reversal Design

The ABAB design demonstrates experimental control through the four cumulative elements of baseline logic:

  1. Initial Baseline Phase (A1): The investigator repeatedly measures the dependent variable in the absence of the independent variable until steady-state responding is achieved. This establishes an objective description of current performance and allows prediction of future responding if no intervention occurs.
  2. Initial Intervention Phase (B1): The independent variable is introduced. A pronounced shift in level, trend, or variability away from the predicted baseline path confirms an intervention effect.
  3. Withdrawal / Return to Baseline Phase (A2): The independent variable is completely withdrawn or returned to baseline conditions. If the target behavior reverses back to levels approximating the initial baseline (A1), the investigator achieves verification. Verification accomplishes two critical scientific objectives: (a) it confirms that the initial baseline prediction was accurate, and (b) it rules out extraneous variables (such as history, maturation, or seasonal shifts) as the cause of the behavior change observed during B1.
  4. Reintroduction of Intervention Phase (B2): The independent variable is reintroduced. If the behavior shifts once again in the therapeutic direction, the investigator achieves replication. Replicating the intervention effect across two distinct phases provides undeniable evidence of experimental control, confirming a functional relation.
Phase A1 (Baseline)       Phase B1 (DRA)        Phase A2 (Withdrawal)     Phase B2 (Re-DRA)
   [ Prediction ]   -->   [ IV Effect ]   -->     [ Verification ]   -->   [ Replication ]

Variations of the Reversal Paradigm

  • BAB Reversal Design: The investigation begins immediately with the active intervention (B1), withdraws the intervention to baseline conditions (A), and reinstates the intervention (B2).
    • Clinical Indication: A BAB design is used when treatment must begin immediately and an initial baseline cannot be collected. It still requires a withdrawal (A) phase, so it is not appropriate when withdrawing an effective treatment would endanger the client or others.
    • Methodological Limitation: The BAB design is experimentally weaker than an ABAB design because it lacks an initial baseline against which B1 can be evaluated; it demonstrates verification and replication, but does not establish an initial baseline prediction.
  • Multiple-Treatment Reversal Designs (e.g., ABCAC, ABACAB): The investigator evaluates two or more distinct independent variables (B and C) against baseline (A) and against each other. For example, comparing baseline (A), noncontingent reinforcement (B), and differential reinforcement of alternative behavior (C).
    • Critical Vulnerability: Multiple-treatment designs are highly vulnerable to sequence effects (order effects)—the phenomenon where a participant's responding in Condition C is influenced by their prior exposure to Condition B rather than being an unconfounded reflection of Condition C alone.
  • Noncontingent Reinforcement (NCR) Reversal: Rather than withdrawing the reinforcing stimulus entirely during Phase A2, the experimenter delivers the identical reinforcer on a fixed-time or variable-time schedule independent of behavior. If the target behavior increases during the NCR phase, the experimenter definitively proves that the response-reinforcer contingency (differential reinforcement) was the active mechanism, rather than the mere noncontingent delivery of the stimulus.
  • DRO / DRI / DRA Reversals: The reversal phase delivers reinforcement contingent on the absence of the target behavior (DRO) or contingent on an incompatible behavior (DRI) rather than returning to a true baseline, cleanly isolating behavioral mechanisms.

Limitations and Ethical Contraindications of Reversal Designs

Despite its scientific power, the reversal design is constrained by two paramount limitations:

  1. Behavioral Irreversibility: An experimental reversal design requires that the target behavior be physically and functionally reversible. Once an individual learns an academic repertoire, motor skill, or cognitive strategy (e.g., decoding words, riding a bicycle, long division algorithms, or a social conversation script), the behavior cannot be unlearned simply by withdrawing instruction. Furthermore, once a behavior comes into contact with natural, unprogrammed community contingencies of reinforcement, withdrawing the experimental intervention will not return responding to baseline levels. Reversal designs are utterly useless for irreversible acquisition repertoires.
  2. Ethical Hazards of Treatment Withdrawal: Withdrawing an effective behavioral intervention that suppresses severe, dangerous, or harmful behavior—such as self-injurious behavior (SIB), violent aggression toward peers, property destruction, or pica—conflicts with the Ethics Code's requirements to provide effective treatment (Standard 2.01) and to minimize the risk of harm (Standard 2.15). Intentionally subjecting a client or clinical staff to renewed physical injury purely to demonstrate experimental verification is ethically indefensible. Additionally, caregivers and teachers who have experienced therapeutic relief frequently refuse to comply with a withdrawal phase.

2. Multiple Baseline Designs

The multiple baseline design (MBD) is the most widely utilized single-case experimental architecture in applied behavior analysis. It provides a robust demonstration of experimental control without requiring a withdrawal or reversal phase, making it the design of choice for irreversible skills and severe problem behaviors.

Structural Logic and Staggered Implementation

In a multiple baseline design, two or more independent baselines are established concurrently across different tiers. Once steady-state baseline responding is achieved across all tiers, the independent variable is introduced to Tier 1 while baseline conditions are systematically maintained in all remaining tiers. After steady-state responding is demonstrated in Tier 1, the independent variable is introduced to Tier 2, while baseline conditions continue in the remaining untreated tiers. This sequential, staggered rollout continues across all successive tiers.

Three Classic Variations of Multiple Baseline Designs

  1. Multiple Baseline Across Behaviors: The independent variable is applied sequentially across two or more functionally independent response classes emitted by the same participant within the same setting (e.g., sequentially teaching mands, tacts, and intraverbals; or sequentially reducing hair-pulling, nail-biting, and skin-picking).
  2. Multiple Baseline Across Settings: The independent variable is applied sequentially across two or more independent environments or situations for the same participant exhibiting the same target behavior (e.g., implementing an on-task token economy sequentially in the classroom, the library, and the cafeteria).
  3. Multiple Baseline Across Participants: The independent variable is applied sequentially across two or more different individuals who exhibit the same target behavior within the same environmental setting (e.g., introducing a peer-tutoring reading intervention sequentially to Student 1, Student 2, and Student 3 in a third-grade classroom).

Baseline Logic in Multiple Baseline Designs

Experimental control in a multiple baseline design is established through a distinct triad of baseline logic:

  • Prediction: The initial baseline in each tier establishes a stable level, trend, and variability that predicts future performance in that specific tier if no intervention is applied.
  • Verification: When the independent variable is introduced to Tier 1 and Tier 1 changes, untreated tiers (Tier 2 and Tier 3) remain at their predicted baseline levels. The fact that behavior in the untreated tiers does not change verifies that extraneous history, maturation, or testing effects were not responsible for the change observed in Tier 1. It proves that the change in Tier 1 was uniquely tied to the introduction of the independent variable.
  • Replication: When the independent variable is subsequently introduced to Tier 2 and produces a behavioral shift comparable to Tier 1, the intervention effect is replicated. Introducing the IV to Tier 3 replicates the effect once more. Convincing replication across three or more tiers confirms experimental control.

Methodological Variations: Multiple Probe and Delayed Baseline

  • Multiple Probe Design (Horner & Baer, 1978): In a standard multiple baseline design, baseline data must be recorded continuously across every tier for the duration of the study. In many applied settings, continuous daily measurement during extended baselines is problematic: it creates severe testing fatigue, exposes the learner to prolonged extinction, induces frustration, and consumes massive clinical resources. Furthermore, for chained instructional tasks (e.g., a 10-step vocational hand-washing or assembly task), testing a learner daily on steps they cannot possibly complete is clinically counterproductive. The multiple probe design solves this problem by substituting intermittent, scheduled probe sessions for continuous daily measurement. Probes are gathered: (1) initially across all tiers, (2) immediately prior to introducing the IV to any tier, and (3) after criterion is reached in a trained tier. This preserves learner engagement while cleanly demonstrating experimental control.
  • Delayed Multiple Baseline Design: An experimental arrangement where initial baselines are collected on Tier 1, and subsequent baselines on additional tiers are started at later calendar dates. This variation is utilized when new participants, behaviors, or settings become available after a study has already begun, or when resource constraints prevented concurrent initial data collection. While useful in clinical practice, delayed baselines provide weaker experimental verification than concurrent multiple baselines because early tiers lack concurrent control data.

Critical Methodological Assumptions & Vulnerabilities of MBD

  • Functional Independence of Tiers: The tiers must be functionally independent. If the independent variable is introduced to Tier 1 and behavior in Tier 2 changes prematurely before Tier 2 receives treatment, experimental control is severely compromised. This premature change (generalization or behavioral covariation) destroys the opportunity for verification. For example, in a multiple baseline across behaviors, if teaching a child to mand also causes an immediate, spontaneous surge in untaught tacts, the clinician cannot prove that the tacting surge was caused by the intervention.
  • Common Sensitivity to the IV: Although the tiers must be independent, they must remain similar enough to respond to the same independent variable. If Tier 1 responds to a token system but Tier 2 does not, replication fails.

3. Alternating Treatments Design (ATD / Multi-Element Design)

The alternating treatments design (ATD)—also called a multielement design—is characterized by the rapid, semi-random, or alternating presentation of two or more distinct treatments or conditions across sessions, days, or times of day.

Operational Architecture of the ATD

In an ATD, the participant is exposed to two or more independent variable conditions in rapid succession (e.g., Condition B in the morning, Condition C in the afternoon; or Condition B on Monday, Condition C on Tuesday). The sequence of conditions is counterbalanced or semi-randomized to prevent order effects.

A fundamental, mandatory requirement of the ATD is that each condition must be paired with an unambiguous, distinct discriminative stimulus (SDS^D). For example, Condition B might be conducted by Therapist 1 wearing a blue smock in Room A with blue materials, while Condition C is conducted by Therapist 2 wearing a green smock in Room B with green materials. Distinct SDS^Ds enable the participant to rapidly discriminate which environmental contingency is operative in each session.

Session:   1 (B)  -->  2 (C)  -->  3 (B)  -->  4 (C)  -->  5 (B)  -->  6 (C)
Stimulus:  [Blue]     [Green]     [Blue]     [Green]     [Blue]     [Green]
Outcome:   High R      Low R      High R      Low R      High R      Low R

Demonstrating Experimental Control in an ATD

Experimental control in an alternating treatments design is demonstrated through vertical visual separation (fractionation) between the plotted data paths. When the data path for Treatment B consistently and decisively separates from the data path for Treatment C without overlapping, experimental control is confirmed. Each successive data point in Condition B replicates the effect of earlier Condition B points, while simultaneously verifying that Condition C produces a distinctly different level of responding.

Variations of the Alternating Treatments Paradigm

  1. ATD Without a Baseline: Compares two or more active treatments immediately from session 1 without an initial baseline phase. Highly useful when an urgent clinical intervention must be selected immediately.
  2. ATD With an Initial Baseline: Baseline data are collected until steady-state responding is achieved, followed by the rapid alternation phase. This allows the clinician to determine not only which treatment is superior, but also whether both treatments are superior to baseline.
  3. ATD With Baseline and a Final Best-Treatment Phase: Baseline is collected, followed by the alternating treatments comparison phase. Once visual fractionation clearly establishes that one treatment is superior, the clinician eliminates the inferior treatment and implements the winning intervention exclusively across all sessions to solidify clinical gains.

Advantages and Clinical Strengths of the ATD

  • No Baseline or Withdrawal Required: Experimental control can be demonstrated without ever withholding or withdrawing treatment, making it highly ethical for severe problem behaviors.
  • Speed and Efficiency: Yields rapid comparative answers in a fraction of the time required by sequential reversal or multiple baseline designs.
  • Insensitive to History, Maturation, and Instrumentation: Because all treatment conditions are evaluated concurrently across the same time window, extraneous historical events, biological maturation, and observer drift affect all conditions equally, reducing their threat to internal validity.
  • Applicable to Irreversible Behaviors: Can be used to compare acquisition strategies (e.g., comparing sight-word learning via flashcards vs. tablet apps) across matched, equivalent sets of instructional items.

Methodological Limitations & Vulnerabilities of the ATD

  • Multi-Treatment Interference (Carryover Effects): The paramount vulnerability of the ATD. Multi-treatment interference occurs when the effects of one treatment condition influence, contaminate, or bleed into the participant's responding in an adjacent treatment condition. If Condition B produces intense behavioral fatigue, contrast, or emotional escalation, responding in a subsequent Condition C session conducted an hour later may be artificially skewed.
  • Artificiality of Rapid Contingency Shifting: Real-world environments rarely shift reinforcement contingencies rapidly from hour to hour. Highly complex discriminations may confuse some learners.
  • Limited to Rapidly Reversible Behaviors: When evaluating single behavioral excesses (e.g., aggression), the behavior must be capable of shifting rapidly up and down from session to session as discriminative stimuli alternate.

4. Changing Criterion Design

The changing criterion design is an experimental architecture in which, following an initial baseline phase, the independent variable is implemented across successive, stepwise phases that require incremental adjustments in performance criteria to access reinforcement or avoid punishment.

Demonstrating Experimental Control

Experimental control in a changing criterion design is demonstrated when the participant's performance closely, systematically, and consistently tracks each stepwise criterion shift. If the criterion is raised, behavior increases to meet it; if the criterion is held constant, behavior stabilizes; if the criterion is lowered, behavior decreases accordingly. The closer the correspondence between the data points and the prescribed criterion lines across successive phases, the stronger the demonstration of experimental control.

Baseline       Phase 1 (Step 1)    Phase 2 (Step 2)    Phase 3 (Mini-Reversal)    Phase 4 (Step 3)
[Mean: 10] -->  [Criterion: 15] -->  [Criterion: 20] -->     [Criterion: 12]     -->  [Criterion: 25]

Four Methodological Rules for Valid Changing Criterion Designs

To construct an airtight changing criterion design that satisfies board-level standards, four operational parameters must be systematically controlled:

  1. Length of Phases: Each phase must be of sufficient duration to achieve steady-state responding before advancing the criterion. Changing criteria prematurely before responding stabilizes destroys experimental control. Furthermore, varying the length of phases (e.g., Phase 1 lasts 5 days, Phase 2 lasts 8 days, Phase 3 lasts 6 days) strengthens experimental control by proving that behavior change is linked to the criterion shift rather than a predictable chronological schedule.
  2. Magnitude of Criterion Changes: The magnitude of each stepwise shift must be carefully calibrated. If the criterion step is too small, changes will be lost in natural baseline variability and visual inspection will fail to detect experimental control. If the criterion step is too large, it may exceed the learner's capabilities, causing ratio strain, extinction bursts, or behavioral collapse. Varying the magnitude of shifts (e.g., increasing by 5 units, then by 8 units, then by 4 units) enhances the credibility of experimental control.
  3. Number of Criterion Steps: A changing criterion design should include several criterion changes (at least three is a common recommendation). Two criterion steps represent a weak demonstration that could easily be dismissed as coincidence.
  4. Bidirectional Changes (Mini-Reversals): Incorporating at least one bidirectional change—a temporary phase where the performance criterion is shifted in the opposite, counter-therapeutic direction—provides the ultimate demonstration of experimental control. If an intervention aims to increase steps walked per day (e.g., baseline: 2,000; Phase 1: 3,500; Phase 2: 5,000), temporarily lowering the criterion back to 3,500 in Phase 3 causes the participant's step count to drop cleanly to 3,500. When the criterion is subsequently raised to 6,500 in Phase 4, step count surges again. This bidirectional tracking convincingly proves that the criterion contingencies functionally govern behavior, completely ruling out general physical conditioning, maturational growth, or seasonal weather as confounds.

Clinical Scope and Appropriate Repertoires

The changing criterion design is exclusively suited for shaping or modulating the rate, frequency, duration, latency, or magnitude of behaviors that are ALREADY in the learner's repertoire. Common clinical exemplars include:

  • Gradually increasing daily distance walked or physical exercise repetitions.
  • Gradually increasing minutes of sustained independent seatwork.
  • Gradually increasing correct words read fluently per minute.
  • Gradually reducing daily cigarettes smoked or caffeine consumed.
  • Gradually reducing minutes of vocal stereotypy during free play.

Critical Exam Trap: A changing criterion design can NEVER be used for novel skill acquisition or complex chained tasks (such as teaching a child how to tie shoes, ride a bicycle, or complete a multi-step math problem). A non-existent behavioral repertoire cannot be adjusted in incremental criterion steps; novel repertoires require prompting hierarchies, shaping, or chaining evaluated via multiple baseline or multiple probe designs.


Single-Case Experimental Designs Comparative Matrix

The following comprehensive matrix details the core mechanisms, experimental control demonstrations, clinical strengths, and contraindications across the primary single-case architectures:

DesignCore ArchitectureMethod of Demonstrating Experimental ControlMajor Clinical StrengthsMajor Limitations & Contraindications
ABAB ReversalAlternating baseline (A) and intervention (B) across sequential phases.Prediction in A1, IV effect in B1, Verification in A2, Replication in B2.Gold standard of experimental control; cleanest proof of functional relation.Strictly contraindicated for irreversible repertoires (academics) and dangerous behaviors (severe SIB/aggression).
BAB ReversalBegins immediately with intervention (B1), withdraws to baseline (A), reinstates (B2).Verification in A, Replication in B2 (lacks initial baseline prediction).Starts treatment immediately when an initial baseline cannot be collected.Weaker than ABAB (no initial baseline); still requires a withdrawal phase, so it is unsuitable when withdrawal is dangerous.
Multiple BaselineStaggered implementation of IV across 2 or more independent tiers (behaviors, settings, participants).Untreated tiers remain at baseline (Verification) while treated tier changes; replicated across tiers.No withdrawal required; ideal for irreversible skills and severe problem behaviors.Vulnerable to tier interdependence (generalization across tiers ruins verification); prolonged baselines.
Multiple ProbeIntermittent baseline probes scheduled across tiers rather than continuous daily recording.Intermittent baseline probes verify stability prior to staggered IV introduction across tiers.Eliminates testing fatigue, extinction, and observer reactivity during extended baselines; ideal for chains.Does not provide continuous daily time-series data during baseline phases.
Alternating Treatments (ATD)Rapid, semi-random alternation of 2 or more treatments paired with distinct SDS^Ds.Vertical visual separation (fractionation) between plotted data paths without overlap.Highly efficient; no baseline or withdrawal required; less vulnerable to history and maturation threats.Highly susceptible to multi-treatment interference (carryover effects); requires rapidly reversible behaviors.
Changing CriterionStepwise, incremental changes in performance criteria following baseline.Behavior closely, consistently tracks each stepwise criterion shift across successive phases.Reinforces gradual shaping; no withdrawal of treatment required; highly motivating.Restricted strictly to adjusting rates of existing repertoires; completely contraindicated for novel skill acquisition.

Clinical Decision-Making: Selecting the Appropriate Single-Case Design

When selecting a single-case experimental design for clinical evaluation or research, assistant behavior analysts must navigate a structured decision hierarchy:

  1. Is the target behavior dangerous or life-threatening (e.g., severe SIB, concussive head-banging, intense aggression)?
    • Yes: Avoid any design that requires withdrawing an effective treatment (ABAB or BAB). Select a Multiple Baseline Design or, when comparing treatments, an Alternating Treatments Design.
  2. Is the target behavior irreversible (e.g., novel academic skill, motor chain, reading fluency)?
    • Yes: Reversal designs are impossible. Select a Multiple Baseline Across Behaviors/Participants or a Multiple Probe Design.
  3. Is the clinical goal to compare two or more distinct interventions head-to-head rapidly?
    • Yes: Select an Alternating Treatments Design (ATD) with distinct discriminative stimuli.
  4. Is the clinical goal to gradually shape or adjust the rate, duration, or magnitude of a behavior already in the repertoire?
    • Yes: Select a Changing Criterion Design with bidirectional shifts and varied step lengths.
  5. Is continuous daily baseline testing likely to induce testing fatigue, extinction bursts, or severe frustration on an unlearned chain?
    • Yes: Select a Multiple Probe Design.
Loading diagram...
Clinical Decision Framework for Single-Case Experimental Design Selection
Test Your Knowledge

An assistant behavior analyst is designing an instructional intervention to teach a 6-step vocational food preparation sequence to three young adults with intellectual disabilities. The intervention will be introduced sequentially across the three participants. Because the learners currently possess zero steps of the complex chained task, the BCaBA is concerned that conducting continuous, daily baseline testing across all steps for several weeks will cause extreme frustration, testing fatigue, and potential problem behavior during untreated baselines. Which experimental design should the BCaBA select to address this methodological and clinical concern?

A

A changing criterion design, because criteria for reinforcement can be adjusted in small steps without requiring baseline data.

B

An ABAB reversal design, because withdrawing instruction after Step 1 verifies experimental control before proceeding to Step 2.

C

A multiple probe design, because intermittent probe sessions can replace continuous daily baseline testing across the three participants.

D

An alternating treatments design, because it allows rapid alternation between food preparation and vocational cleaning tasks without distinct stimuli.

Test Your Knowledge

A behavior analyst implements an intervention to systematically increase the daily step count of a sedentary adult with developmental disabilities. Following a stable baseline averaging 2,500 steps per day, the analyst establishes an initial criterion of 3,500 steps, which the client reliably achieves. Successive phases set criteria at 4,500 steps, 5,500 steps, 4,000 steps, and 6,500 steps. The client's daily step count closely mirrors each prescribed level. What is the methodological purpose of setting the criterion to 4,000 steps during the fourth intervention phase?

A

To implement a component analysis by removing the pedometer feedback while maintaining the reinforcement contingency.

B

To prevent multi-treatment interference by alternating between aerobic and anaerobic exercise criteria.

C

To make a bidirectional change, showing control because behavior follows the criterion even in the counter-therapeutic direction.

D

To evaluate whether the target behavior has reached an irreversible plateau that prevents any further physical conditioning or progress.

Test Your Knowledge

A 14-year-old resident in an intensive crisis stabilization facility engages in severe, high-intensity head-banging against concrete walls, resulting in open tissue wounds and concussive injury. The clinical team immediately implements a protective helmet and a continuous differential reinforcement of alternative behavior (DRA) procedure, which successfully reduces head-banging to near-zero rates. The facility's research coordinator suggests removing the helmet and DRA contingency for two weeks to establish a formal baseline (A) and complete a standard ABAB reversal design for publication. How should the assistant behavior analyst respond?

A

Recommend transitioning immediately to a changing criterion design where head-banging is shaped down in stepwise increments across two weeks.

B

Oppose the withdrawal, because removing an effective treatment for life-threatening self-injury is unacceptable; use a design without withdrawal.

C

Propose an alternating treatments design comparing the helmet against a contingent shock apparatus without an initial baseline.

D

Support the withdrawal phase because demonstrating verification and replication through a complete ABAB reversal is ethically mandated prior to making any clinical recommendations.

Sections you finish are checked off in the contents.