3.4 Reinforcement and Punishment Schedules
Key Takeaways
- Positive reinforcement involves adding a stimulus following a response to increase future behavior frequency; negative reinforcement involves removing an aversive stimulus to increase future behavior frequency.
- Positive punishment adds an aversive stimulus to decrease future behavior; negative punishment removes a preferred stimulus to decrease future behavior.
- Continuous reinforcement (CRF/FR1) is optimal for initial skill acquisition, whereas intermittent reinforcement schedules maintain resistance to extinction.
- Variable ratio (VR) schedules generate extremely high, steady response rates without post-reinforcement pauses and are highly resistant to extinction.
- Abrupt increases in schedule requirements produce ratio strain, whereas discontinuing reinforcement produces extinction bursts.
3.4 Reinforcement and Punishment Schedules
In Applied Behavior Analysis (ABA), consequences are systematically categorized based on their functional impact on future behavior frequency. An Applied Behavior Analysis Technician (ABAT) must master the core principles of reinforcement, punishment, and their corresponding intermittent schedules to implement clinical treatment plans effectively.
Core Contingencies: The Four Consequence Quadrants
Consequences are classified along two dimensions:
- Effect on Behavior: Does future behavior frequency increase (Reinforcement) or decrease (Punishment)?
- Environmental Change: Is a stimulus added/presented (Positive) or removed/withdrawn (Negative)?
┌───────────────────────────┬───────────────────────────┐
│ INCREASES BEHAVIOR (S-R) │ DECREASES BEHAVIOR (S-P) │
┌─────────────────────┼───────────────────────────┼───────────────────────────┤
│ STIMULUS ADDED │ Positive Reinforcement │ Positive Punishment │
│ (Positive +) │ (S-R+) │ (S-P+) │
├─────────────────────┼───────────────────────────┼───────────────────────────┤
│ STIMULUS REMOVED │ Negative Reinforcement │ Negative Punishment │
│ (Negative -) │ (S-R-) │ (S-P-) │
└─────────────────────┴───────────────────────────┴───────────────────────────┘
1. Positive Reinforcement ($S^{R+}$)
Positive reinforcement occurs when a response is immediately followed by the presentation of a stimulus, resulting in an increase in the future frequency of that response under similar conditions.
- Clinical Example: A child hands an picture card of an apple to the ABAT ($R$), the ABAT immediately hands the child an apple slice ($S^{R+}$), and the child's frequency of handing picture cards increases in future sessions.
2. Negative Reinforcement ($S^{R-}$)
Negative reinforcement occurs when a response is immediately followed by the removal, termination, reduction, or postponement of a stimulus, resulting in an increase in the future frequency of that response.
- Escape Contingency: Behavior terminates an ongoing aversive stimulus (e.g., putting on sunglasses terminates bright glare).
- Avoidance Contingency: Behavior prevents an impending aversive stimulus from occurring (e.g., turning on windshield wipers before driving into rain).
- Clinical Example: A student completes 5 math problems ($R$), the ABAT removes the remaining work folder ($S^{R-}$), and the student's task completion increases in future sessions.
3. Positive Punishment ($S^{P+}$)
Positive punishment occurs when a response is followed by the presentation of an aversive stimulus, resulting in a decrease in the future frequency of that response.
- Examples: Overcorrection, contingent exercise, or verbal reprimands.
4. Negative Punishment ($S^{P-}$)
Negative punishment occurs when a response is followed by the removal of a preferred stimulus, resulting in a decrease in the future frequency of that response.
- Response Cost: Loss of a specific amount of earned positive reinforcers (e.g., losing 2 tokens for throwing items).
- Time-Out from Positive Reinforcement: Temporary loss of access to all positive reinforcement contingent on target behavior.
Unconditioned vs. Conditioned Reinforcers
- Unconditioned Reinforcers (Primary / Primary Reinforcers): Stimuli that function as reinforcers without any prior learning history due to biological necessity (e.g., food, water, warmth, oxygen, sexual stimulation).
- Conditioned Reinforcers (Secondary / Learned Reinforcers): Previously neutral stimuli that acquired reinforcing properties through paired association with other reinforcers (e.g., tokens, money, praise, checkmarks, social approval).
- Generalized Conditioned Reinforcer (GCSR): A conditioned reinforcer that has been paired with many unconditioned and conditioned reinforcers and does not depend on a specific MO to be effective (e.g., money or token economy stars).
Intermittent Schedules of Reinforcement
A schedule of reinforcement is a rule specifying which instances of a response will be reinforced.
Continuous Reinforcement (CRF / FR1)
Reinforcement is delivered after every single instance of the target behavior. CRF is used primarily during the acquisition phase when teaching a brand-new skill.
Intermittent Schedules (INT)
Reinforcement is delivered for some, but not all, instances of the target behavior. Intermittent schedules are used to maintain established behaviors and make them resistant to extinction.
Intermittent schedules are arranged across two main axes:
- Ratio vs. Interval: Ratio schedules depend on the number of responses; Interval schedules depend on the passage of time plus one correct response.
- Fixed vs. Variable: Fixed schedules have a constant requirement; Variable schedules have an average requirement.
┌───────────────────────────┬───────────────────────────┐
│ RATIO │ INTERVAL │
│ (Number of Responses) │ (Time Elapsed) │
┌────────────────────────┼───────────────────────────┼───────────────────────────┤
│ FIXED │ Fixed Ratio (FR) │ Fixed Interval (FI) │
│ (Constant Requirement) │ (High rate; Pause after) │ (Scalloped rate pattern) │
├────────────────────────┼───────────────────────────┼───────────────────────────┤
│ VARIABLE │ Variable Ratio (VR) │ Variable Interval (VI) │
│ (Average Requirement) │ (Steep, steady; No pause) │ (Moderate, constant rate) │
└────────────────────────┴───────────────────────────┴───────────────────────────┘
Detailed Breakdown of the 4 Basic Intermittent Schedules
1. Fixed Ratio (FR)
Reinforcement is delivered after a fixed, predetermined number of correct responses.
- Pattern: High rate of responding followed by a post-reinforcement pause (PRP) (a brief break before responding resumes).
- Example: FR5 schedule—delivering a token after every 5th completed discrete trial.
2. Variable Ratio (VR)
Reinforcement is delivered after an average, unpredictable number of correct responses.
- Pattern: Extremely high, steady rate of responding with no post-reinforcement pauses.
- Example: VR5 schedule—delivering reinforcement after 3, 7, 2, and 8 responses (average = 5). Slot machines operate on VR schedules.
3. Fixed Interval (FI)
Reinforcement is delivered for the first correct response emitted after a fixed duration of time has elapsed.
- Pattern: Scalloped pattern—slow response rate immediately following reinforcement, accelerating sharply as the end of the interval approaches.
- Example: FI 5-minute schedule—checking an oven timer when baking cookies; checking accelerates as 5 minutes nears.
4. Variable Interval (VI)
Reinforcement is delivered for the first correct response emitted after an average, unpredictable duration of time has elapsed.
- Pattern: Moderate, steady, highly consistent rate of responding with no pauses.
- Example: VI 3-minute schedule—unannounced pop quizzes or supervisory quality checks evoke consistent daily studying.
Intermittent Schedules Comparison Table
| Schedule | Requirement | Rate of Responding | Pattern / Pauses | Resistance to Extinction | Clinical Example |
|---|---|---|---|---|---|
| Fixed Ratio (FR) | Fixed # of responses | High | Post-reinforcement pause | Moderate | Token delivered after every 4th correctly assembled puzzle piece (FR4). |
| Variable Ratio (VR) | Average # of responses | Very High | Steep, steady, no pauses | Highest | Reinforcing vocal requests on average every 3rd attempt (VR3). |
| Fixed Interval (FI) | 1st response after fixed time | Low-to-Moderate | Scalloped (slow then fast) | Moderate | Student looks at clock every 10 minutes near recess bell (FI 10-min). |
| Variable Interval (VI) | 1st response after average time | Moderate | Steady, uniform rate | High | Technician delivers praise for on-task behavior on average every 5 min (VI5). |
Schedule Thinning, Ratio Strain, and Extinction
Schedule Thinning
Schedule thinning is the gradual process of systematically increasing the response requirements or interval durations (e.g., moving from FR1 to FR3, then VR5, then VR10) to fade artificial reinforcers into naturally occurring contingencies.
Ratio Strain
Ratio strain occurs when schedule requirements are increased too abruptly or raised beyond the client's current behavioral capacity. Symptoms of ratio strain include avoidance, aggression, fatigue, and sudden cessation of responding. If ratio strain occurs, the ABAT must immediately back up to a denser reinforcement schedule.
Extinction Burst & Spontaneous Recovery
When a previously reinforced behavior is placed on operant extinction:
- Extinction Burst: A temporary, sharp increase in the frequency, intensity, duration, and variability of the target behavior immediately following the withdrawal of reinforcement.
- Spontaneous Recovery: The re-emergence of a previously extinguished behavior after a period of time has passed without reinforcement.
A behavior technician is working on manding skills. The technician provides a bite of snack on average every 4th correct independent mand (e.g., after 2 mands, then 6 mands, then 3 mands, then 5 mands). What schedule of reinforcement is being implemented?
A student receives a break after working quietly for 15 minutes. Upon returning, the student engages in high work output right before the 15-minute mark, but shows very low output immediately after the break ends, forming a scalloped pattern of responding. Which schedule produces this pattern?
During session, an ABAT abruptly changes a client's token reinforcement schedule from FR2 straight to FR25. The client immediately stops responding, pushes materials off the table, and displays extreme agitation. What behavioral phenomenon has occurred?