10.1 Differential Reinforcement Procedures (DRA, DRI, DRO, DRL, DRH)
Key Takeaways
Differential reinforcement is an operant mechanism that simultaneously combines reinforcement for a desirable or acceptable response class with extinction for an undesirable response class.
Differential Reinforcement of Alternative Behavior (DRA) replaces problem behavior with a functionally equivalent, socially valid alternative response; Functional Communication Training (FCT) is a premier application of DRA.
Differential Reinforcement of Incompatible Behavior (DRI) is a subtype of DRA that reinforces a response topographically incapable of occurring simultaneously with the target problem behavior.
Differential Reinforcement of Other Behavior (DRO) reinforces the complete absence (zero rate) of the problem behavior during or at the end of an interval; it does not teach a replacement behavior and risks reinforcing undesirable behaviors occurring in the behavioral vacuum.
Differential Reinforcement of Low Rates (DRL) reduces excessive frequencies of behaviors acceptable in moderation and is strictly contraindicated for dangerous behaviors, whereas DRH increases response rates to foster behavioral fluency.
Foundational Mechanics and Conceptual Basis of Differential Reinforcement
In applied behavior analysis, differential reinforcement is a foundational, evidence-based behavior-change strategy defined by the simultaneous application of two distinct operant principles: reinforcement for a designated, desirable response class and operant extinction (withholding the maintaining reinforcer) for a designated problem behavior or non-target response class (Catania, 1998; Cooper, Heron, & Heward, 2020; Skinner, 1938).
The dual operant mechanism functions as a dynamic system:
- Reinforcement: Increases or maintains the future frequency, rate, duration, or fluency of the adaptive, functional, or acceptable response class.
- Extinction: Systematically withholds the functional reinforcer maintaining the challenging behavior, thereby decreasing its future probability and weakening the motivating relation.
When differential reinforcement is implemented across varied environmental conditions correlated with specific antecedent stimuli, it establishes stimulus discrimination (differential responding in the presence versus absence of a discriminative stimulus, ). When differential reinforcement is applied to specific variations within or across response forms, it produces response differentiation (the emergence of a distinct, refined response class while unreinforced response variations extinguish). Together, these mechanisms allow behavior analysts to eliminate severe problem behaviors while concurrently building socially valid, habilitative repertoires without resorting to primary aversive contingencies.
Differential Reinforcement of Alternative Behavior (DRA)
Core Mechanics and Function Matching
Differential Reinforcement of Alternative Behavior (DRA) involves reinforcing a desirable, socially valid alternative behavior that serves the identical function as the problem behavior, while placing the problem behavior on extinction. DRA is widely recognized as the gold-standard intervention for socially maintained problem behaviors because it directly adheres to the Fair Pair Rule (White & Haring, 1980): whenever a challenging behavior is reduced or eliminated, an adaptive replacement behavior must be explicitly taught and strengthened to replace it.
A fundamental clinical mandate of DRA is functional equivalence. The alternative response must access the exact functional reinforcer maintaining the challenging behavior, as identified through a Functional Behavior Assessment (FBA) or experimental functional analysis (FA). For example:
- If a student engages in physical aggression maintained by escape from difficult academic tasks (social negative reinforcement), the alternative behavior must produce escape or task reduction (e.g., requesting a brief break, asking for assistance, or negotiating task modifications). Reinforcing quiet sitting with edible candies (social positive reinforcement) fails to address the underlying establishing operation (EO) of task aversiveness and will routinely fail.
- If a child engages in screaming maintained by caregiver attention (social positive reinforcement), the alternative behavior must produce adult attention (e.g., tapping the caregiver's shoulder, using a polite vocal greeting, or holding up an attention card), while screams are placed on attention extinction (planned ignoring).
Functional Communication Training (FCT)
Pioneered by Carr and Durand (1985), Functional Communication Training (FCT) is an application of DRA where the alternative response is an explicit communicative mand (e.g., vocal word, manual sign, picture exchange icon, or speech-generating device button). FCT restructures the client's communicative environment so that the functional reinforcer is delivered immediately and densely on a continuous schedule (Fixed Ratio 1, or FR1) contingent upon the communicative mand, while the problem behavior produces zero reinforcement.
The Matching Law in DRA Design
Herrnstein's (1961) Matching Law dictates that when two or more concurrent response alternatives are available, the relative rate of responding matches the relative rate, immediacy, quality, and effort of reinforcement associated with each alternative:
To ensure that the alternative replacement behavior outcompetes the problem behavior, the behavior analyst must deliberately bias reinforcement parameters in favor of the alternative response:
- Schedule / Rate of Reinforcement: The alternative behavior must contact reinforcement on a denser schedule (initially FR1) than the problem behavior.
- Immediacy of Reinforcement: The alternative behavior must produce instant reinforcement (zero latency), whereas reinforcement for problem behavior must be eliminated or substantially delayed.
- Quality of Reinforcement: The alternative behavior must access higher-potency, preferred variations of the maintaining reinforcer.
- Response Effort: The physical and cognitive effort required to emit the alternative behavior must be lower than or equal to the effort required to emit the problem behavior. If vocalizing "May I please have a three-minute break from this worksheet?" requires high cognitive effort while slamming a textbook onto the floor is effortless, the client will allocate responding to slamming under the matching law. The alternative response must be engineered to be effortless (e.g., touching a single break card).
Differential Reinforcement With and Without Extinction
Task G.18 asks you to implement differential reinforcement with and without extinction. With extinction, problem behavior no longer produces its reinforcer, which is the most reliable arrangement. Sometimes extinction is impossible (peers keep laughing at call-outs) or unsafe (severe self-injury cannot simply be ignored). Differential reinforcement can still work if the reinforcement parameters strongly favor the alternative behavior: the alternative produces a longer, higher-quality, more immediate, or more frequent reinforcer than the problem behavior, which still contacts some reinforcement. Athens and Vollmer (2010) reduced problem behavior this way without extinction by manipulating the duration, quality, and delay of reinforcement. The trade-off is that the effect depends on keeping that advantage; if reinforcement for the alternative weakens, problem behavior can return quickly.
Differential Reinforcement of Incompatible Behavior (DRI)
Anatomical and Topographical Incompatibility
Differential Reinforcement of Incompatible Behavior (DRI) is a specialized subtype of DRA in which the alternative response class is topographically and anatomically incompatible with the target problem behavior. The learner physically cannot emit both behaviors at the exact same instant in time (Cooper et al., 2020).
The defining feature of DRI is physical impossibility. Incompatibility must be rooted in biomechanics rather than subjective interpretation:
- Problem Behavior: Hand-to-head self-injurious hitting Incompatible Behavior: Keeping both hands clasped firmly in lap, sitting on hands, or holding a weighted ball with both hands.
- Problem Behavior: Out-of-seat running in a classroom Incompatible Behavior: Sitting with buttocks flat on seat cushion and feet touching the floor.
- Problem Behavior: Mouthing non-food objects / pica Incompatible Behavior: Chewing sugar-free gum or biting on an oral motor chewy tube.
- Problem Behavior: Scratching, pinching, or gouging staff members Incompatible Behavior: Keeping hands inside zippered jacket pockets or interlocking fingers.
Clinical Strengths and Limitations of DRI
- Clinical Strengths: DRI provides an immediate safety buffer. Because the incompatible response physically prevents the emission of the problem behavior, it reduces the need for intrusive physical blocking or mechanical restraints. It provides clear, unambiguous observation criteria for direct-care staff and Registered Behavior Technicians (RBTs).
- Limitations: In daily living, identifying an incompatible behavior that is both anatomically exclusive and functionally equivalent to the problem behavior is often challenging. For example, sitting with hands in pockets does not provide functional escape from a difficult reading task; thus, DRI must often be combined with functional communication to address the maintaining establishing operation.
Differential Reinforcement of Other Behavior (DRO - Zero-Rate Responding)
Core Contingency and Definitions
Differential Reinforcement of Other Behavior (DRO) delivers a specified reinforcer contingent upon the complete absence (zero occurrences) of the target problem behavior throughout a specified temporal interval or at a specific momentary point in time. DRO is technically a misnomer; it does not reinforce a specific "other" behavior, but rather reinforces the non-occurrence (omission) of the target response class. For this reason, it is frequently termed differential reinforcement of zero responding or omission training (Reynolds, 1961).
Primary DRO Subtypes
Behavior analysts employ four primary variations of DRO, structured along two methodological axes: interval versus momentary and fixed versus variable.
- Fixed-Interval DRO (FI-DRO): A fixed temporal duration (e.g., 5 minutes) is established. If the learner does not engage in the target problem behavior at any point during the entire 5-minute interval, reinforcement is delivered immediately at the conclusion of the interval. If the problem behavior occurs at any point during the interval, reinforcement is withheld.
- Variable-Interval DRO (VI-DRO): Reinforcement is delivered following the absence of the target behavior across intervals of varying durations that revolve around a predetermined mathematical mean (e.g., VI 4 minutes, consisting of randomized intervals such as 2, 4, 6, 3, and 5 minutes). Varying the interval makes the end of each interval less predictable for the learner.
- Fixed-Momentary DRO (FM-DRO): The observer schedules a fixed temporal interval (e.g., every 5 minutes). Reinforcement is delivered contingent on the learner not engaging in the target behavior at the exact moment the interval expires (e.g., at minute 5:00). If the behavior occurs at minute 2:30 or 4:15, but is completely absent at minute 5:00, the learner earns reinforcement.
- Variable-Momentary DRO (VM-DRO): The observer checks for the presence or absence of the target behavior at variable, unpredictable moments throughout the day. If the learner is not emitting the behavior at that specific momentary observation point, reinforcement is delivered.
Interval Resetting vs. Non-Resetting Mechanics
When implementing interval DRO (FI-DRO or VI-DRO), clinicians must specify how the interval operates when a problem behavior occurs:
- Resetting Interval: The instant the target problem behavior occurs, the timer is immediately stopped and reset to zero, beginning the interval anew. This establishes a strict contingency where every occurrence of problem behavior directly postpones reinforcement by the full interval duration.
- Non-Resetting Interval: If the target behavior occurs, the timer continues running until the interval ends. Reinforcement is withheld at the interval's conclusion, and the subsequent interval begins automatically at the scheduled time. While resetting intervals generally produce faster behavioral deceleration, non-resetting intervals are significantly easier for classroom teachers and vocational supervisors managing group schedules.
Setting and Thinning Initial DRO Intervals
A critical, frequently tested rule of thumb governs initial DRO interval calibration: The initial DRO interval must be equal to or slightly less than the learner's baseline mean interresponse time (IRT):
Clinicians typically set the initial interval between and of the baseline mean IRT. For example, if a client emits 20 instances of severe self-injurious skin-picking during a 60-minute baseline observation, the mean IRT is 3 minutes (180 seconds). The initial DRO interval should be set between 90 seconds and 140 seconds.
The Trap of Setting Intervals Too Long: If the clinician sets the initial DRO interval at 10 minutes when the baseline mean IRT is 3 minutes, the client will almost certainly emit the behavior before the interval expires. The client will experience prolonged extinction without ever contacting reinforcement, precipitating severe extinction bursts, emotional aggression, and treatment failure.
Systematic Schedule Thinning: Once the learner consistently earns reinforcement across consecutive intervals (e.g., achieving of available intervals over 3 consecutive sessions), the interval is systematically thinned. Thinning can occur by increasing the interval duration by fixed increments (e.g., adding 30 seconds or lengthening by to ) or by transitioning from an interval DRO to a momentary DRO.
The "Behavioral Vacuum" and Critical Limitations of DRO
While DRO rapidly suppresses problem behavior, it possesses two major, inherent clinical vulnerabilities:
- Zero Alternative Skill Acquisition: DRO reinforces the absence of a behavior, meaning it fails to teach, strengthen, or promote an adaptive replacement repertoire. It violates the Fair Pair Rule if implemented in isolation.
- The Behavioral Vacuum and Accidental Reinforcement: When a high-rate problem behavior is suppressed, a "behavioral vacuum" is created. Because the DRO contingency reinforces whatever behavior is occurring when the timer expires, it can accidentally reinforce novel, secondary, or dangerous problem behaviors. If a student refrains from head-banging for 4 minutes and 59 seconds, but kicks a teacher or screams profanities at minute 5:00, delivering reinforcement at the 5-minute mark directly reinforces kicking or screaming. Therefore, DRO should rarely be used as a standalone intervention; it must be coupled with DRA, DRI, or skill-acquisition programming.
Differential Reinforcement of Low Rates of Responding (DRL)
Clinical Rationale and Scope
Differential Reinforcement of Low Rates of Responding (DRL) delivers reinforcement contingent upon the target behavior occurring at a rate lower than a predetermined, specified criterion. DRL is uniquely indicated when the target behavior is acceptable, appropriate, or socially valid in moderation, but problematic, disruptive, or counterproductive when emitted with excessive frequency.
Common clinical applications include:
- Asking a teacher for help or clarification.
- Raising one's hand during classroom instruction.
- Requesting permission to use the restroom or drink water.
- Eating rapidly (shoveling food during meals).
- Washing hands or checking door locks.
Three Subtypes of DRL
- Full-Session DRL: Reinforcement is delivered at the end of an entire instructional session or observation period if the total number of responses emitted during that entire session is equal to or less than a specified criterion (e.g., student earns a preferred activity at the end of a 60-minute period if they asked to leave their seat 3 or fewer times).
- Interval DRL: The total session is divided into equal, smaller temporal intervals (e.g., four 15-minute segments in an hour). Reinforcement is delivered at the conclusion of each interval if responses during that specific interval do not exceed the criterion (e.g., no more than 1 question per 15 minutes).
- Spaced-Responding DRL: Reinforcement is delivered immediately following an occurrence of the target behavior, provided that a specified minimum interresponse time (IRT) has elapsed since the preceding response (). If the response occurs prior to the criterion IRT, it is placed on extinction, and the IRT timer resets to zero. As with DRO, set the initial criterion at or slightly below the baseline mean IRT and lengthen it gradually.
Absolute Contraindication of DRL
DRL must NEVER be used for severe problem behaviors, including self-injurious behavior (SIB), physical aggression, property destruction, or elopement. In DRL, responses that meet the schedule criterion are explicitly reinforced; reinforcing even a "low rate" of head banging, biting, or violent assault is a dangerous and unethical clinical practice. Zero-tolerance behaviors mandate DRA, DRI, or DRO.
Differential Reinforcement of High Rates of Responding (DRH)
Differential Reinforcement of High Rates of Responding (DRH) delivers reinforcement contingent upon the target behavior occurring at or above a predetermined rate or frequency criterion, or with interresponse times shorter than a specified duration (). DRH is designed to build behavioral fluency, increase the speed of academic or vocational task execution, and promote rapid, active student responding (ASR).
Variations parallel DRL:
- Full-Session DRH: Reinforcement delivered if total responses across the session meet or exceed criterion (e.g., completing at least 40 math facts during a 2-minute timing).
- Interval DRH: Reinforcement delivered if responses in a discrete block meet or exceed criterion (e.g., reading at least 25 sight words in a 30-second interval).
- Spaced-Responding DRH: Reinforcement delivered immediately following a response if the IRT since the preceding response is shorter than a specified maximum duration (e.g., assembling two mechanical components within 4 seconds of each other).
Differential Reinforcement Procedures Matrix
The following matrix contrasts the operational contingencies, target behavioral effects, skill acquisition properties, clinical indications, and primary risks across the differential reinforcement family:
| Procedure | Reinforcement Delivery Contingency | Target Behavioral Effect | Replacement Behavior Taught? | Ideal Clinical Application | Major Risk / Clinical Trap |
|---|---|---|---|---|---|
| Differential Reinforcement of Alternative Behavior (DRA) | Delivered contingent on emission of a specified alternative, functionally equivalent response; withheld for problem behavior. | Decreases problem behavior; increases socially valid alternative behavior. | Yes (explicitly teaches a functionally equivalent replacement response). | Socially maintained problem behavior (escape, attention, tangible access); Functional Communication Training (FCT). | Inadvertently selecting an alternative behavior with higher response effort than problem behavior, causing treatment failure under Matching Law. |
| Differential Reinforcement of Incompatible Behavior (DRI) | Delivered contingent on a response that is anatomically/physically impossible to emit simultaneously with problem behavior. | Decreases problem behavior; increases incompatible motor response. | Yes (teaches a motorically incompatible replacement response). | High-intensity motor behaviors threatening immediate safety (eye-gouging, scratching, biting, out-of-seat running). | Incompatible behavior may not be functionally equivalent to problem behavior; fails to address establishing operations. |
| Differential Reinforcement of Other Behavior (DRO) | Delivered contingent on the complete absence (zero rate) of problem behavior across an interval or at a momentary point. | Decreases problem behavior to zero or near-zero levels. | No (reinforces the absence of behavior; creates a behavioral vacuum). | Rapid suppression of severe, high-rate problem behavior (severe SIB, aggression) where immediate reduction is paramount. | Accidentally reinforcing secondary undesirable behaviors occurring when the interval expires; setting initial interval longer than baseline IRT. |
| Differential Reinforcement of Low Rates (DRL) | Delivered when response rate is below criterion (full-session/interval) or when IRT exceeds minimum threshold (spaced-responding). | Moderates response rate; maintains behavior at an acceptable, tolerable frequency. | No (maintains existing behavior at reduced frequency). | Behaviors acceptable or desirable in moderation (asking questions, hand-raising, requesting help, eating pace). | Misapplying to dangerous behaviors (aggression, SIB, pica); slow behavioral reduction pace. |
| Differential Reinforcement of High Rates (DRH) | Delivered when response rate meets/exceeds criterion or when IRT is shorter than maximum threshold. | Accelerates response rate; fosters rapid, fluent execution. | No (increases rate of an already mastered behavior). | Behavioral fluency, active student responding (ASR), rapid task completion, athletic and vocational speed. | Setting criteria too high too quickly, causing ratio strain and behavioral extinction. |
Clinical Troubleshooting and Common Exam Pitfalls
- Pitfall 1: Confusing DRO with Extinction: Operant extinction involves withholding the maintaining reinforcer whenever the target behavior occurs. DRO involves delivering a reinforcer contingent on the non-occurrence of the target behavior. Extinction is a single operant process; DRO is a schedule of reinforcement.
- Pitfall 2: Setting DRO Intervals Based on Clinical Hope Rather Than Baseline IRT: Setting a 10-minute initial DRO interval for a client who engages in aggression every 2 minutes guarantees treatment failure. The initial interval must be calculated using baseline mean IRT (Initial Interval Baseline Mean IRT).
- Pitfall 3: Failing to Combine Extinction with DRA: If a teacher provides a 3-minute break when a student says "break, please", but also allows the student to escape the task when they throw a textbook, extinction is absent. Under the Matching Law, both behaviors remain in concurrent competition, and the problem behavior will persist if it requires less effort.
- Pitfall 4: Utilizing DRL for Self-Injurious Behavior or Aggression: Board exams routinely test this trap. DRL explicitly reinforces low rates of behavior. Never select DRL for behaviors that pose physical danger.
A BCaBA is designing a behavior intervention plan for an 8-year-old student who engages in high-rate vocal call-outs during independent academic work (baseline data indicate a mean IRT of 45 seconds). A functional assessment confirms the call-outs are maintained by teacher attention. The classroom goal is to eliminate disruptive call-outs and teach the student to raise their hand quietly and wait to be called upon. Which differential reinforcement procedure and initial parameter should the BCaBA implement?
Implement Differential Reinforcement of Incompatible Behavior (DRI) reinforcing sitting in the classroom chair, while continuing to provide verbal reprimands following call-outs.
Use DRA with functional communication: reinforce quiet hand-raising with teacher attention on an FR1 schedule while placing vocal call-outs on extinction.
Implement Spaced-Responding DRL with an initial criterion IRT of 15 seconds, reinforcing any call-out that occurs after at least 15 seconds have elapsed.
Implement a Fixed-Interval DRO with an initial interval duration of 5 minutes, delivering teacher attention if the student emits zero call-outs for the entire 5 minutes.
An assistant behavior analyst is consulting at an adult day program for a client who engages in chronic, severe hand-to-head hitting. Baseline observational data reveal a mean IRT of 90 seconds. The interdisciplinary clinical team suggests using a Fixed-Interval Differential Reinforcement of Other Behavior (FI-DRO) schedule. Which of the following represents the most critical technical limitation and clinical consideration when implementing FI-DRO in this scenario?
FI-DRO requires continuous physical restraint whenever the client attempts to emit the target response during the interval duration.
FI-DRO does not teach a replacement behavior and may accidentally reinforce other problem behavior occurring when the interval ends.
FI-DRO can only be implemented if the maintaining reinforcer is automatically mediated and cannot be utilized for socially maintained problem behaviors.
The initial interval duration must be set at 10 minutes to ensure that the client experiences sufficient establishing operations before accessing the reinforcer.
A high-school student with a developmental disability asks the job coach for the current time every 2 to 3 minutes throughout a 4-hour community vocational shift. Asking for the time is an acceptable and functional skill, but the high frequency severely disrupts work productivity. What is the most appropriate differential reinforcement procedure to implement?
Differential Reinforcement of High Rates of Responding (DRH) requiring at least 10 inquiries per hour to contact reinforcement.
Differential Reinforcement of Alternative Behavior (DRA) teaching the student to write down the question on an index card instead of asking vocally.
Spaced-responding DRL, starting the IRT criterion near the baseline average (about 2 to 3 minutes) and lengthening it gradually.
Differential Reinforcement of Other Behavior (DRO) with a 30-minute interval, resetting the timer whenever the student asks for the time.
Sections you finish are checked off in the contents.