7.2 Operant Conditioning & Schedules of Reinforcement
Key Takeaways
- Reinforcement always increases behavior frequency; punishment always decreases behavior frequency.
- Positive indicates adding a stimulus; negative indicates subtracting or removing a stimulus.
- Negative reinforcement includes escape learning (stopping ongoing pain) and avoidance learning (preventing upcoming pain).
- Variable-ratio reinforcement schedules produce the highest response rates and are most resistant to extinction.
- Instinctive drift occurs when conditioned operant responses revert toward innate, species-typical behaviors.
7.2 Operant Conditioning & Schedules of Reinforcement
Operant conditioning is a form of associative learning in which the likelihood of a voluntary behavior being repeated is modified by its consequences. First conceptualized by Edward Thorndike through his Law of Effect (which states that behaviors followed by satisfying outcomes are more likely to recur, while those followed by unpleasant outcomes are less likely to recur), operant conditioning was systematically developed by B.F. Skinner using the operant conditioning chamber ("Skinner Box").
Reinforcement vs. Punishment: The Four Quadrants
In operant conditioning, consequences are defined strictly by their effect on behavior frequency, not by subjective pleasantness or intent.
- Reinforcement: ANY consequence that increases or maintains the rate of a behavior.
- Punishment: ANY consequence that decreases or suppresses the rate of a behavior.
- Positive (+): Adding or introducing a stimulus to the environment.
- Negative (-): Removing or subtracting a stimulus from the environment.
Behavior Increases (Reinforcement) Behavior Decreases (Punishment)
+------------------------------------+------------------------------------+
Positive (+) | POSITIVE REINFORCEMENT | POSITIVE PUNISHMENT |
(Add Stimulus) | Add pleasant stimulus | Add unpleasant stimulus |
| Example: Give a reward for study | Example: Give a fine for speeding |
+------------------------------------+------------------------------------+
Negative (-) | NEGATIVE REINFORCEMENT | NEGATIVE PUNISHMENT |
(Remove Stimulus| Remove unpleasant stimulus | Remove pleasant stimulus |
| Example: Take aspirin for headache | Example: Revoke driver's license |
+------------------------------------+------------------------------------+
The Two Subtypes of Negative Reinforcement: Escape vs. Avoidance Learning
Negative reinforcement is frequently tested on the MCAT and is subdivided based on timing:
- Escape Learning: The organism performs a behavior to terminate an ongoing, currently existing aversive stimulus.
- Example: An individual turns on an air conditioner to stop feeling uncomfortably hot in a stifling room, or a rat presses a lever to turn off an active electric floor shock.
- Avoidance Learning: The organism performs a behavior to prevent a future, anticipated aversive stimulus from occurring in response to a signal or discriminative stimulus.
- Example: An individual checks weather reports and carries an umbrella before leaving the house to avoid getting wet in a forecasted rainstorm, or a rat presses a lever when a warning light turns on to prevent an electric shock from ever turning on.
| Operant Contingency | Action Taken | Stimulus Manipulation | Impact on Behavior | Clinical / Everyday Example |
|---|---|---|---|---|
| Positive Reinforcement | Add Stimulus | Desirable stimulus introduced | Behavior Increases | Verbal praise given for completing medical chart |
| Negative Reinforcement (Escape) | Remove Stimulus | Ongoing aversive stimulus ended | Behavior Increases | Taking ibuprofen to eliminate active tension headache |
| Negative Reinforcement (Avoidance) | Remove Stimulus | Impending aversive stimulus prevented | Behavior Increases | Studying weeks in advance to prevent exam failure anxiety |
| Positive Punishment | Add Stimulus | Undesirable stimulus introduced | Behavior Decreases | Receiving a formal reprimand for tardiness |
| Negative Punishment | Remove Stimulus | Desirable stimulus revoked | Behavior Decreases | Suspending privileges after a policy violation |
Schedules of Reinforcement
Schedules of reinforcement dictate the timing and frequency with which behavioral responses produce reinforcers. Schedules are classified along two dimensions:
- Ratio vs. Interval: Ratio schedules reinforce based on the number of behavioral responses executed. Interval schedules reinforce the first response after a specified time interval has elapsed.
- Fixed vs. Variable: Fixed schedules maintain a constant, predictable criterion. Variable schedules maintain an unpredictable criterion based on an average value.
The Four Basic Partial Reinforcement Schedules
1. Fixed-Ratio (FR) Schedule
Reinforcement is delivered after a set, predictable number of correct responses.
- Response Pattern: High rate of response with a characteristic post-reinforcement pause (a brief period of inactivity immediately after receiving reinforcement before responding resumes).
- Real-World Example: A garment factory worker paid $15 for every 10 shirts stitched (FR-10).
2. Variable-Ratio (VR) Schedule
Reinforcement is delivered after an unpredictable number of responses, fluctuating around a target average.
- Response Pattern: Highest overall response rate of all schedules, with no post-reinforcement pause. Behavior is steady, rapid, and most resistant to extinction.
- Real-World Example: Slot machines at a casino (VR-average), or a door-to-door salesperson making sales.
3. Fixed-Interval (FI) Schedule
Reinforcement is delivered for the first response executed after a fixed, predictable duration of time has elapsed.
- Response Pattern: Produces a distinct "scalloped" response curve. Response rates drop immediately after reinforcement and accelerate dramatically as the end of the time interval approaches.
- Real-World Example: Studying intensively for weekly scheduled quizzes (FI-7 days), or checking an oven as the timer approaches zero.
4. Variable-Interval (VI) Schedule
Reinforcement is delivered for the first response executed after an unpredictable duration of time has elapsed, averaging around a specific time frame.
- Response Pattern: Produces a moderate, steady rate of response without significant pauses or rapid acceleration.
- Real-World Example: Checking email for an important response, or unannounced pop quizzes (VI-average).
| Reinforcement Schedule | Reinforcement Basis | Response Rate | Pattern of Responding | Extinction Resistance |
|---|---|---|---|---|
| Fixed-Ratio (FR) | Set number of responses | High | Post-reinforcement pause then rapid burst | Moderate |
| Variable-Ratio (VR) | Average number of responses | Highest | Steepest, continuous, steady | Highest (Most resistant) |
| Fixed-Interval (FI) | Set time elapsed | Low to Moderate | Scalloped (sharp increase near interval end) | Low |
| Variable-Interval (VI) | Average time elapsed | Moderate | Smooth, steady, predictable rate | High |
Advanced Operant Concepts & Behavioral Phenomena
Shaping and Chaining
- Shaping: The process of reinforcing successive approximations of a desired target behavior. Used to establish complex or novel behaviors that an organism would never perform spontaneously (e.g., training a dog to turn on a light switch by rewarding looking at the switch, walking toward it, touching it, and finally pushing it).
- Chaining: Linking a sequence of discrete, individually shaped behaviors into a complex behavior chain, where each response acts as a discriminative stimulus for the next response and a conditioned reinforcer for the preceding one.
Extinction Burst
When a previously reinforced behavior is suddenly no longer reinforced, the organism temporarily exhibits a sharp, dramatic increase in the frequency, intensity, and variability of the behavior before it ultimately declines. For example, if a vending machine fails to drop a snack after money is inserted, a person will repeatedly and aggressively press the button (extinction burst) before giving up.
Discriminative Stimulus ($S^D$)
A discriminative stimulus ($S^D$) is a specific environmental signal or cue that indicates that reinforcement is currently available for a given behavior. Conversely, $S^\Delta$ (S-delta) indicates that reinforcement is not available. For example, a "Green Light" on a rat box signals that lever pressing will produce food pellets.
Instinctive Drift
Discovered by Marian and Keller Breland (students of B.F. Skinner), instinctive drift is the tendency of an animal to revert from a conditioned operant behavior back to an innate, biologically hardwired species-typical instinctual behavior—especially when the operant behavior closely resembles or competes with natural foraging or defensive instincts.
- Classic Example: Breland & Breland tried to train pigs to deposit wooden coins into a piggy bank for food reinforcers. Over time, instead of carrying the coins directly to the bank, the pigs began dropping the coins on the ground, rooting them with their snouts, and pushing them around. Rooting is an innate foraging instinct in pigs that interfered with and ultimately disrupted the conditioned operant behavior.
A patient taking a prescribed medication experiences unpleasant gastrointestinal side effects whenever they take it on an empty stomach. The patient learns to eat a meal immediately before taking the medication, which completely prevents the gastrointestinal distress from occurring. What operant conditioning mechanism is demonstrated by the behavior of eating before taking medication?
A researcher records the lever-pressing responses of four groups of rats trained under different reinforcement schedules. Group 4 exhibits a response pattern characterized by low activity immediately following reinforcement, followed by a sharp, dramatic acceleration in lever pressing as the end of a 2-minute timer approaches, producing a "scalloped" cumulative response curve. Which reinforcement schedule was assigned to Group 4?
Animal trainers attempt to condition raccoons to pick up two wooden tokens and deposit them into a container to receive a food reward. Although the raccoons learn to pick up the tokens, over repeated sessions they begin rubbing the tokens together, dipping them into the container, and holding onto them rather than dropping them in. This failure of operant conditioning is best explained by which concept?
When a toddler throws a tantrum in a grocery store because they want candy, the parent initially ignores the child. Unexpectedly, the toddler begins screaming significantly louder, flailing on the floor, and banging their head against the cart for two minutes before finally stopping. What operant conditioning phenomenon occurred during the sudden escalation of the tantrum?