5.4 Positive & Negative Reinforcement Principles
Key Takeaways
- Reinforcement is functionally defined as an immediate stimulus change following a response that results in an increase in the future frequency, rate, duration, or intensity of that behavior in similar conditions.
- Positive reinforcement (SR+) involves the presentation or addition of a stimulus, while negative reinforcement (SR-) involves the removal, termination, reduction, or postponement of an aversive stimulus; both operations increase future behavior.
- Negative reinforcement encompasses two distinct operational contingencies: escape (terminating an ongoing aversive stimulus) and avoidance (preventing or delaying an oncoming aversive stimulus before it occurs).
- Generalized Conditioned Reinforcers (GCRs), such as tokens and money, are paired with multiple backup reinforcers, making them highly resistant to satiation and independent of specific motivating operations.
- The clinical efficacy of reinforcement is governed by four primary parameters: immediacy of delivery, reinforcer magnitude and quality, contingency (conditional probability), and current motivating operations.
5.4 Positive & Negative Reinforcement Principles
Quick Answer: In Applied Behavior Analysis, reinforcement is strictly defined by its functional effect on future behavior: an environmental stimulus change immediately following a response that increases the future frequency, rate, duration, or intensity of that behavior class under similar conditions. Positive Reinforcement (SR+) involves the addition or presentation of a stimulus, whereas Negative Reinforcement (SR-) involves the removal, reduction, or postponement of a stimulus; both processes increase behavior. Negative reinforcement is divided into escape contingencies (terminating an existing aversive stimulus) and avoidance contingencies (preventing an oncoming aversive stimulus). Generalized Conditioned Reinforcers (GCRs), such as tokens, derive their power from being exchangeable for multiple backup reinforcers, rendering them immune to single-MO satiation.
Supervisors must recognize that reinforcement is not a reward, a social bribe, or a moral judgment. It is an objective natural process. If a consequence does not demonstrably increase the future rate of a target behavior, it cannot be classified as a reinforcer, regardless of clinician intent.
The Formal Functional Definition of Reinforcement
A behavioral consequence can only be defined as a reinforcer through empirical observation and data analysis. The formal operational definition requires three criteria:
- Temporal Contiguity: The stimulus change occurs immediately following the emission of the target response (ideally within 0 to 1 seconds).
- Functional Effect: The future probability, frequency, latency, or intensity of that response class increases or is maintained over time.
- Stimulus Context: The increase in responding occurs under similar antecedent stimulus conditions.
THE BEHAVIORAL FUNCTION OF REINFORCEMENT
Target Response Emitted ──> Immediate Stimulus Change ──> Future Frequency
["Break please"] [Worksheet removed] of Mand Increases
▲
│
Function: Verified Reinforcement (SR-)
The Functional vs. Topographical Fallacy (Reward vs. Reinforcer)
One of the most persistent clinical errors in autism intervention is treating "reward" as a synonym for "reinforcer":
- Reward: An arbitrary item, verbal praise, or activity given to an individual after an accomplishment, based on social convention or subjective assumption. A reward carries no empirical proof that it will alter future behavior.
- Reinforcer: An environmental consequence that has been empirically demonstrated through data collection to increase the future rate of the behavior it follows.
- Clinical Example: A technician gives an autistic child a piece of chocolate every time the child sits in his chair. Over two weeks, the data show that out-of-seat behavior actually increases, because the child loves running around and finds chocolate unpalatable. In this case, chocolate is a reward in the technician's mind, but functionally it is not a reinforcer—it may even be an aversive stimulus.
Positive Reinforcement (SR+): Stimulus Addition
Positive Reinforcement (SR+) occurs when a response is immediately followed by the presentation, addition, or increase in intensity of a stimulus, resulting in an increase in the future frequency of that response class under similar conditions.
Classes of Positive Reinforcers
- Edible Reinforcers: Consumables (e.g., fruit snacks, crackers, juice). Primarily effective during early developmental stages or when biological hunger (EO) is present.
- Tangible Reinforcers: Physical objects (e.g., toy cars, vibrating sensory balls, bubbles, stickers).
- Activity / Privilege Reinforcers: Access to preferred actions (e.g., swinging, jumping on a trampoline, watching a 2-minute video clip, being line leader).
- Social Reinforcers: Interpersonal responses from others (e.g., vocal praise, high-fives, smiles, tickles, clapping).
- Sensory Reinforcers: Direct sensory stimulation (e.g., visual spinning lights, auditory music, tactile textured fabrics).
The Premack Principle ("Grandma's Rule")
Formulated by David Premack (1959), the Premack Principle states that making the opportunity to engage in a high-probability behavior (a highly preferred activity) contingent upon completing a low-probability behavior (a non-preferred task) will function as positive reinforcement for the low-probability behavior.
- Formula: "First [Low-Probability Behavior], Then [High-Probability Behavior]."
- Clinical Application: An autistic student loves spinning the wheels on toy cars (high-probability behavior) but resists sorting flashcards (low-probability behavior). The QASP-S structures the program so that sorting 5 flashcards is immediately reinforced with 60 seconds of spinning car wheels.
Negative Reinforcement (SR-): Stimulus Removal and Relief
Negative Reinforcement (SR-) occurs when a response is immediately followed by the removal, termination, reduction, or postponement of a stimulus (typically an aversive or non-preferred event), resulting in an increase in the future frequency of that response class under similar conditions.
┌────────────────────────────────────────────────────────────────────────┐
│ THE GOLDEN RULE OF NEGATIVE REINFORCEMENT │
├────────────────────────────────────────────────────────────────────────┤
│ • Negative Reinforcement is NOT punishment! │
│ • Positive Reinforcement (+) ──> ADD stimulus ──> INCREASES rate │
│ • Negative Reinforcement (-) ──> REMOVE stimulus ──> INCREASES rate │
│ • Positive Punishment (+P) ──> ADD stimulus ──> DECREASES rate │
│ • Negative Punishment (-P) ──> REMOVE stimulus ──> DECREASES rate │
└────────────────────────────────────────────────────────────────────────┘
Escape vs. Avoidance Contingencies
Negative reinforcement is clinically operationalized into two distinct paradigms based on whether the aversive stimulus is present when the behavior occurs:
┌────────────────────────────────────────────────────────────────────────┐
│ ESCAPE VS. AVOIDANCE CONTINGENCIES │
├────────────────────────────────────────────────────────────────────────┤
│ 1. Escape Contingency │
│ • The aversive stimulus is CURRENTLY PRESENT in the environment. │
│ • Emitting the behavior TERMINATES or REDUCES the ongoing stimulus. │
│ • Example: Fire alarm blares ──> Student covers ears or exits room. │
├────────────────────────────────────────────────────────────────────────┤
│ 2. Avoidance Contingency │
│ • The aversive stimulus is NOT YET PRESENT in the environment. │
│ • Emitting the behavior PREVENTS or POSTPONES its occurrence. │
│ • Example: Student sees teacher holding math binder ──> Student │
│ requests a bathroom pass, PREVENTING the worksheet delivery. │
└────────────────────────────────────────────────────────────────────────┘
Discriminated Avoidance vs. Free-Operant Avoidance
- Discriminated Avoidance: An explicit warning stimulus (Sd or CMO-R) occurs, signaling that an aversive event is scheduled to follow. Responding in the presence of the warning signal prevents the onset of the aversive event. For example, a classroom timer beeps (warning signal), prompting the student to put his hands in his lap, avoiding a loss of recess tokens.
- Free-Operant Avoidance: No warning stimulus is presented. Responding at any point during an interval delays or postpones the onset of an aversive event according to a set schedule. For example, a driver taps the brake pedal periodically while descending an icy mountain road to prevent skidding, without any warning chime.
Primary, Secondary, and Generalized Conditioned Reinforcers
Reinforcers are classified according to how they acquired their reinforcing properties:
┌────────────────────────────────────────────────────────────────────────┐
│ REINFORCER CLASSIFICATION BY ORIGIN │
├────────────────────────────────────────────────────────────────────────┤
│ 1. Primary (Unconditioned) Reinforcers (SR+ unconditioned) │
│ • Phylogenic, innate biological survival value. │
│ • No learning history required. │
│ • Examples: Food, water, warmth, oxygen, sleep, relief from pain. │
├────────────────────────────────────────────────────────────────────────┤
│ 2. Secondary (Conditioned) Reinforcers (SR+ conditioned) │
│ • Ontogenic; neutral stimuli paired with existing reinforcers. │
│ • Examples: High-fives, stickers, musical chimes, checkmarks. │
├────────────────────────────────────────────────────────────────────────┤
│ 3. Generalized Conditioned Reinforcers (GCR) │
│ • Conditioned reinforcers paired with MANY backup reinforcers. │
│ • Highly resistant to satiation; independent of specific MOs. │
│ • Examples: Tokens in a token economy, money, generalized praise. │
└────────────────────────────────────────────────────────────────────────┘
The Clinical Power of Generalized Conditioned Reinforcers (GCRs)
A single conditioned reinforcer (such as a sticker or a specific toy car) quickly loses its effectiveness if the client becomes satiated on that specific item (Abolishing Operation). In contrast, a Generalized Conditioned Reinforcer (GCR) is paired with an extensive variety of backup reinforcers (edibles, tangibles, physical activities, computer time, breaks).
- Because a GCR can be exchanged for almost anything, it does not depend on any single, specific motivating operation for its effectiveness.
- Token Economies: In an autism classroom, tokens function as GCRs. A student who is not hungry can exchange tokens for sensory swinging; a student who is tired can exchange tokens for a quiet break; a student who is bored can exchange tokens for an iPad game. This multi-backup pairing makes the token system exceptionally durable and effective across changing motivational states.
Socially Mediated vs. Automatic Reinforcement
Reinforcement is also categorized by whether another human being delivers the consequence:
┌────────────────────────────────────────────────────────────────────────┐
│ SOCIALLY MEDIATED VS. AUTOMATIC REINFORCEMENT │
├────────────────────────────────────────────────────────────────────────┤
│ Socially Mediated Reinforcement │ The consequence is delivered by │
│ │ another person (teacher, parent, peer)│
│ │ following the target response. │
├────────────────────────────────────────────────────────────────────────┤
│ Automatic Reinforcement │ The consequence is produced DIRECTLY │
│ │ by the physical response itself, with │
│ │ zero social interaction or mediation. │
└────────────────────────────────────────────────────────────────────────┘
Diagnosing Automatic Reinforcement
Many repetitive behaviors common in autism spectrum disorder—such as vocal scripting, finger flapping, rocking, skin-picking, or humming—are maintained by automatic reinforcement (either positive sensory stimulation or negative sensory pain relief).
- Functional Analysis Signature: During a standard Functional Analysis (FA), if a behavior persists at high, sustained rates across the Alone (or No-Interaction) condition, in the complete absence of social attention, demands, or tangible items, the behavior is confirmed to be maintained by automatic reinforcement.
- Clinical Supervisory Rule: Automatic reinforcement cannot be extinguished by simply ignoring the client. Extinction of automatic reinforcement requires sensory extinction (e.g., placing carpet over a hard floor to eliminate the auditory feedback of plate spinning, or placing a padded glove on a client to prevent the tactile sensation of scratching).
Critical Parameters Governing Reinforcement Efficacy
The clinical effectiveness of any reinforcement protocol is governed by four empirical parameters:
- Immediacy (Temporal Contiguity): A reinforcer must be delivered within seconds (ideally under 1 second) following the response. Delays introduce the risk of adventitious reinforcement—accidentally reinforcing an intervening, unwanted behavior emitted during the delay interval.
- Magnitude and Quality: The quantity, duration, and intensity of the reinforcer must match the effort required by the response (The Matching Law). High-effort tasks (e.g., tolerating a blood draw) require massive, high-preference reinforcers; low-effort tasks (e.g., pointing to an icon) require modest reinforcers.
- Contingency (Conditional Probability): The reinforcer must be delivered if and only if the target response is emitted. If the client can access the reinforcer non-contingently for free, the contingency is corrupted, destroying the reinforcer's effectiveness.
- Motivating Operations (EO/AO): The current level of deprivation or satiation directly dictates whether the stimulus possesses reinforcing value at that moment.
Comparative Matrix: Reinforcement & Operant Contingencies
| Contingency Type | Stimulus Operation | Functional Effect on Future Behavior | Motivating Operation (MO) Required | Autism Clinical Practice Example |
|---|---|---|---|---|
| Positive Reinforcement (SR+) | Stimulus Added / Presented (e.g., praise, edible, token, toy) | Increases future rate of responding. | Establishing Operation for the added item (Deprivation). | Emitting vocal mand "ball" immediately produces a bouncy ball, increasing future requesting. |
| Negative Reinforcement: Escape (SR-) | Stimulus Removed / Terminated while currently ongoing. | Increases future rate of responding. | Ongoing aversive stimulus present (EO for escape). | Client signs "Finished" during loud music; staff turns off music, increasing signing. |
| Negative Reinforcement: Avoidance (SR-) | Stimulus Prevented / Postponed before it occurs. | Increases future rate of responding. | Threat CMO-R (warning signal of scheduled demand). | Client puts on shoes when door chime rings, preventing parent from repeating loud prompts. |
| Generalized Conditioned Reinforcer (GCR) | Token / Currency Added exchangeable for diverse backups. | Increases future rate of responding. | Independent of single MOs due to multiple backup items. | Delivering plastic tokens for completing discrete trials, exchanged for preferred activities. |
| Automatic Reinforcement | Direct physical/sensory consequence produced without social mediation. | Increases or maintains future responding. | Sensory deprivation (EO) or physical discomfort. | Hand flapping producing visual and proprioceptive sensory feedback during free time. |
Supervisory Decision Rules & Exam Traps
Exam Trap 1: The "Negative Reinforcement = Punishment" Trap
By far the most common error on behavioral certification exams is equating negative reinforcement with punishment. Remember the absolute mathematical rule:
- Reinforcement ALWAYS increases behavior.
- Punishment ALWAYS decreases behavior.
- Positive ALWAYS means adding a stimulus.
- Negative ALWAYS means removing/terminating a stimulus. Negative reinforcement strengthens functional escape and avoidance skills; it does not punish or suppress behavior.
Exam Trap 2: Assuming Reinforcers are Universal
There is no such thing as a "universal reinforcer." Edibles like M&Ms, verbal praise like "Good job!", and iPad screen time are not reinforcers by default. For some clients with sensory defensiveness, verbal praise is an aversive conditioned stimulus that evokes escape. A QASP-S must ensure that preference assessments (e.g., MSWO, Paired Stimulus) and reinforcer assessments are conducted empirically for each individual client.
Exam Trap 3: Overlooking Negative Reinforcement in Aggression
When evaluating severe problem behavior (aggression, property destruction, elopement), novice practitioners often assume the behavior is maintained by positive attention ("he just wants someone to look at him"). Empirical functional analysis literature demonstrates that in autism populations, the majority of severe problem behaviors are maintained by negative reinforcement (escape from instructional, social, or sensory demands). Overlooking escape contingencies leads to dangerous treatment failures.
An autistic 8-year-old student is working on a challenging handwriting task. The student suddenly throws his pencil across the room and screams. The classroom aide immediately walks over, picks up the pencil, and tells the student: 'You are clearly overwhelmed. Take a 10-minute break in the library corner.' Over the next three weeks, data show that the student's rate of throwing pencils and screaming during handwriting increases from 2 times per week to 12 times per week. What behavioral principle is demonstrated here?
A QASP-S is developing a token economy system for an early intervention classroom. The behavior technician asks why the team should use tokens that can be exchanged for multiple different items (edibles, playground time, sensory toys, bubbles) rather than simply handing the child a sticker each time he completes a task. How should the supervisor clinically justify the token economy?
During a Functional Analysis (FA) of severe skin-picking in an autistic adult, the supervisor observes that skin-picking occurs at consistently elevated rates during the Alone condition, where the client is in a barren room with no therapists, no demands, no social interaction, and no materials present. Rates during the Alone condition are virtually identical to rates observed in the Demand and Attention conditions. What is the most clinically sound conclusion?