4.2 Operant Conditioning: Reinforcement Schedules, Punishment, and Behavior Modification
Key Takeaways
Thorndike's Law of Effect established that behaviors followed by satisfying states of affairs are stamped in, whereas Skinner distinguished respondent from emitted operant behavior governed by the three-term contingency ().
The four operant contingencies are defined by whether an appetitive or aversive stimulus is presented or removed, with negative reinforcement functioning to increase behavior through escape or active avoidance.
Intermittent reinforcement schedules generate distinct cumulative response profiles: Fixed Ratio produces a post-reinforcement pause, Variable Ratio yields high steady rates, Fixed Interval produces a scalloped curve, and Variable Interval maintains moderate steady responding.
The Partial Reinforcement Extinction Effect (PREE) is explained by the discrimination, frustration, and sequential hypotheses.
Biological boundaries in operant learning are exemplified by Breland & Breland's instinctive drift, autoshaping, and sign-tracking.
Operant Conditioning: Reinforcement Schedules, Punishment, and Behavior Modification
Operant (instrumental) conditioning examines how organisms modify the frequency, intensity, and topography of emitted behaviors based on the environmental consequences those behaviors produce. While classical conditioning focuses on involuntary reflexive antecedents (), operant conditioning governs voluntary, goal-directed behavior operated upon the environment (). Mastery of operant terminology, schedules of reinforcement, behavior modification strategies, and biological boundaries is essential for the GRE Subject Test in Psychology.
1. Historical Foundations: Thorndike and Skinner
Edward Thorndike's Law of Effect
In 1898, Edward L. Thorndike placed hungry cats inside wooden puzzle boxes equipped with interior latches, treadles, or loop strings that opened an escape door to access food outside. Thorndike recorded the escape latency across consecutive trials:
- Instead of exhibiting sudden, reasoned cognitive insight, the cats engaged in chaotic motor trial-and-error thrashing—clawing, biting, and squeezing through bars. Eventually, a cat accidentally struck the release mechanism.
- Across successive trials, unrewarded random behaviors were gradually pruned away, while successful release responses showed a progressive, continuous reduction in latency.
From these data, Thorndike formulated the Law of Effect:
Responses followed immediately by a satisfying state of affairs become more firmly connected to the situation, so that when it recurs, they are more likely to recur; responses followed by an annoying state of affairs have their connections weakened, so that they are less likely to recur.
Thorndike conceptualized learning as the direct, mechanical stamping-in of an associative bond between the environmental stimulus and the motor response ( habit), operating without requiring conscious mental comprehension.
B.F. Skinner and the Experimental Analysis of Behavior
B.F. Skinner refined Thorndike's instrumental paradigm into operant conditioning, rejecting inner hypothetical constructs (e.g., 'satisfaction') in favor of an objective, functional analysis of observable behavior. Skinner introduced critical theoretical distinctions:
- Respondent Behavior: Involuntary, reflexive actions automatically elicited by preceding environmental stimuli (Pavlovian classical conditioning).
- Operant Behavior: Voluntary, spontaneous actions emitted by the organism that operate upon the environment to generate consequences. Operants are defined not by their physical motor topography, but by their functional effect on the environment.
The Three-Term Contingency (The ABCs of Operant Conditioning)
Skinner formalized operant behavior as a dynamic three-term unit:
- Discriminative Stimulus (): An environmental antecedent cue that sets the occasion for responding by signaling that a specific consequence is available contingent upon the emission of response .
- Operant Response (): The emitted behavioral unit.
- Reinforcing or Punishing Stimulus ( / / ): The subsequent environmental change that alters the future probability of that response class.
- (S-Delta): An antecedent stimulus signaling that reinforcement is unavailable; responses emitted in the presence of an undergo extinction.
2. The Four Operant Contingencies and Avoidance Mechanisms
Operant contingencies are classified along two orthogonal dimensions: (1) whether the consequence involves presenting or removing a stimulus, and (2) whether the future probability of the response increases (reinforcement) or decreases (punishment).
STIMULUS PRESENTED STIMULUS REMOVED
┌──────────────────────────────┬──────────────────────────────┐
BEHAVIOR INCREASES │ POSITIVE REINFORCEMENT │ NEGATIVE REINFORCEMENT │
│ (Deliver appetitive stimulus)│ (Remove/terminate aversive) │
│ e.g., Lever press -> Food │ e.g., Lever press -> Stops shock│
├──────────────────────────────┼──────────────────────────────┤
BEHAVIOR DECREASES │ POSITIVE PUNISHMENT │ NEGATIVE PUNISHMENT │
│ (Deliver aversive stimulus) │ (Remove appetitive stimulus)│
│ e.g., Barking -> Citronella │ e.g., Tantrum -> Loses tokens│
└──────────────────────────────┴──────────────────────────────┘
Escape vs. Active Avoidance Conditioning
Negative reinforcement operates through two distinct behavioral paradigms:
- Escape Conditioning: The organism emits an operant response that terminates an ongoing, physically present aversive stimulus (e.g., pressing a lever to turn off a loud 90-dB klaxon).
- Active Avoidance Conditioning: The organism emits an operant response in the presence of an antecedent warning cue () prior to the onset of the aversive stimulus, successfully preventing its delivery altogether.
Mowrer's Two-Process Theory of Avoidance (1947)
Orval Hobart Mowrer formulated the Two-Process Theory to solve a fundamental theoretical paradox: In active avoidance, how can the absence or non-occurrence of an event (shock) serve as a functional reinforcer to sustain responding over hundreds of trials?
- Process 1 (Classical Conditioning): The neutral warning signal (e.g., tone) is repeatedly paired with the noxious unconditioned stimulus (shock). The warning signal becomes an aversive conditioned stimulus () that elicits visceral fear and sympathetic arousal ().
- Process 2 (Operant Conditioning): The organism emits an instrumental motor response (e.g., jumping a hurdle) that immediately terminates the warning . The termination of the fear-inducing CS provides immediate negative reinforcement via fear reduction.
Note
The Avoidance Paradox & Conservation of Fear: If avoidance responses prevent shock entirely, why doesn't the undergo classical extinction? Solomon and Wynne demonstrated the conservation of fear: well-trained organisms respond so rapidly following onset that they terminate the cue before significant fear can be experienced, preserving the underlying excitatory trace indefinitely.
3. Schedules of Reinforcement and Cumulative Record Patterns
In natural environments, behavior is rarely reinforced on a Continuous Reinforcement (CRF / FR-1) schedule. Instead, consequences occur intermittently according to schedules of reinforcement, documented by Skinner and C.B. Ferster (1957) using mechanical cumulative recorders.
Cumulative
Responses
▲ Variable Ratio (VR) Fixed Ratio (FR)
│ / / (Post-reinforcement pause)
│ / /‾/
│ / /‾/
│ / /‾/
│ / /‾/
│ / /
│ / Variable Interval (VI) Fixed Interval (FI)
│ / / _.-'
│ / / _.-' (Scallop pattern)
│ / / _.-'
│ / / _.-'
└─┴───────────┴────────────────────┴────────────────────────► Time
Characteristics of the Four Simple Intermittent Schedules
| Schedule Type | Delivery Contingency | Response Rate & Steady-State Pattern | Resistance to Extinction |
|---|---|---|---|
| Fixed Ratio (FR) | Delivered after a fixed, invariant number of responses () | High run rate with Post-Reinforcement Pause (PRP); cumulative record shows a 'stop-and-go' pattern; PRP length is directly proportional to ratio requirement | Moderate; susceptible to ratio strain if requirement abruptly escalated |
| Variable Ratio (VR) | Delivered after an unpredictable, variable number of responses around an average () | Very high, steady, constant rate without pauses; fastest steady-state operant rate (e.g., slot machines, athletic casting) | Extremely high; creates immense persistence |
| Fixed Interval (FI) | Delivered for the first response emitted after a fixed, invariant duration of time () | The 'FI Scallop': prolonged pause following reinforcement, followed by an accelerating, upward-curving surge in responding as interval expiration nears (e.g., cramming before exams) | Moderate |
| Variable Interval (VI) | Delivered for the first response emitted after an unpredictable, variable duration averaging () | Moderate, remarkably stable, uniform rate without post-reinforcement pauses (e.g., checking email/texts, pop quizzes) | Very high; highly resistant to extinction |
The Partial Reinforcement Extinction Effect (PREE)
Behaviors maintained on intermittent (partial) reinforcement schedules are substantially more resistant to extinction than behaviors established under continuous reinforcement (CRF). Three primary theories account for PREE:
- The Discrimination Hypothesis (Mowrer): Extinction is immediately discriminable following CRF because the first non-reinforced response represents a stark departure from baseline. On intermittent schedules, the onset of extinction is indistinguishable from the long runs of unreinforced responses normally encountered during training.
- The Frustration Theory (Abram Amsel): Non-reinforcement elicits an innate, aversive unconditioned emotional state termed primary frustration. In partial schedules, the internal proprioceptive cues of frustration are repeatedly followed by eventual reinforcement; frustration thus becomes a conditioned discriminative stimulus () that energizes continued responding.
- The Sequential Theory (E.J. Capaldi): Focuses on memory traces. During intermittent schedules, sequences of non-reinforced trials leave lingering short-term memory traces (). When a reinforced trial occurs, these non-reward memory traces become conditioned cues for responding, driving behavior during extended extinction.
4. Behavior Modification, Differential Reinforcement, and Shaping
Applied Behavior Analysis (ABA) utilizes operant principles to establish new repertoires and remediate maladaptive behaviors without relying on aversive punishment.
Shaping by Successive Approximations
When a desired target behavior has a baseline operant rate of zero (e.g., a pigeon pecking a high illuminated disc, a child with autism verbalizing a complete sentence), the terminal response cannot be reinforced directly. Skinner developed shaping: the differential reinforcement of successive, progressive approximations toward the terminal target behavior, while placing earlier approximations on extinction.
Behavioral Chaining
Complex behavioral sequences (e.g., assembling an apparatus, washing hands) consist of discrete component responses linked into a functional chain:
- In a behavioral chain, each completed response produces a distinct environmental transition that functions simultaneously as a conditioned reinforcer for the preceding response and a discriminative stimulus () for the subsequent response.
- Forward Chaining: Training starts with the initial step in the chain. Once mastered, the second step is added, moving chronologically forward.
- Backward Chaining: Training begins with the final step in the sequence. Because the learner immediately contacts the terminal primary reinforcer upon completing the final step, backward chaining generates powerful motivational momentum and is particularly effective for learners with severe developmental delays.
Differential Reinforcement Modalities
• DRA (Alternative Behavior): Reinforce on-task hand raising; Extinguish blurting out.
• DRI (Incompatible Behavior): Reinforce hands in lap; Extinguish self-injurious face scratching.
• DRO (Other Behavior): Reinforce every 5 minutes spent with ZERO aggressive outbursts.
• DRL (Low Rates): Reinforce asking questions only if separated by at least 15 minutes.
- DRA (Differential Reinforcement of Alternative Behavior): Reinforcement is delivered for a specific desirable alternative response, while the problem behavior is placed on extinction.
- DRI (Differential Reinforcement of Incompatible Behavior): A rigorous subtype of DRA where the alternative behavior is topographically and physically incompatible with the maladaptive response (e.g., an organism cannot physically scratch its face while sitting on its hands).
- DRO (Differential Reinforcement of Other Behavior): Reinforcement is delivered strictly contingent upon the complete absence (zero occurrence) of the target problem behavior during a specified time interval, regardless of what other behaviors occur.
- DRL (Differential Reinforcement of Low Rates of Responding): Reinforcement is delivered only when the rate of responding falls below a predetermined threshold, used when the behavior is acceptable in moderation but disruptive when emitted too frequently.
5. Principles of Relative Reinforcement: Premack and Response Deprivation
Classical behaviorists struggled with circular definitions of reinforcement: a reinforcer was defined as anything that strengthened behavior, while strengthened behavior was cited as proof of reinforcement.
The Premack Principle (Differential-Probability Hypothesis)
In 1959, David Premack resolved this circularity by shifting from static stimulus objects (e.g., food pellets) to response probabilities. Premack established that any behavior can be ranked along a probability hierarchy based on the duration of time an organism freely allocates to it when all constraints are removed:
The Premack Principle ('Grandma's Rule'): A high-probability behavior can serve as an effective reinforcer for any lower-probability behavior; conversely, requiring a low-probability behavior will not reinforce a high-probability behavior.
Premack demonstrated this experimentally in children: children who preferred playing pinball over eating candy could be induced to eat candy by making pinball access contingent upon candy consumption. Conversely, children who preferred candy could be induced to play pinball by making candy contingent upon pinball play.
Timberlake & Allison's Response Deprivation Hypothesis (1974)
William Timberlake and James Allison expanded Premack's insight into the Response Deprivation Hypothesis. They demonstrated that relative baseline probability is not the critical determinant. Instead, any activity can function as a reinforcer if an operant contingency restricts the organism's access to that activity below its preferred baseline unconstrained equilibrium (termed the organism's behavioral bliss point):
- Even an intrinsically low-probability baseline behavior (e.g., running on a wheel for an indolent rat) will serve as a potent reinforcer for a high-probability behavior (drinking sweet water) if wheel-running is restricted below its normal baseline free-choice level.
6. Biological Constraints on Operant Conditioning
Keller and Marian Breland: 'The Misbehavior of Organisms' (1961)
Students of B.F. Skinner, Keller and Marian Breland founded Animal Behavior Enterprises to train commercial animal exhibits using operant shaping. Over decades and thousands of animals, they documented widespread, systematic failures of operant conditioning, published in their landmark paper The Misbehavior of Organisms:
- Raccoons and the Piggy Bank: The Brelands shaped a raccoon to pick up wooden coins and deposit them into a metal container for food reinforcement. When required to deposit two coins simultaneously, the operant conditioning collapsed: the raccoon refused to release the coins, incessantly clutching them, dipping them into the container, pulling them out, and rubbing them together for minutes at a time.
- Pigs Rooting Coins: Pigs shaped to carry coins similarly abandoned the efficient operant response: they dropped the coins, pushed them along the dirt with their snouts, tossed them in the air, and rooted them.
Operant Reinforcement Contingency: Coin pickup ──► Deposit ──► Food Reinforcer
│
▼
[INSTINCTIVE DRIFT]
│
▼
Phylogenetic Food-Procuring Reflexes: Raccoon: Coin rubbing / Washing reflex
Pig: Ground rooting reflex
The Brelands coined the term instinctive drift: as an animal gains extensive experience with an operant task involving food reinforcement, learned operant behaviors progressively drift away from the trained contingency and revert toward innate, species-typical, phylogenetic food-procuring instinctive motor patterns.
Autoshaping and Sign-Tracking vs. Goal-Tracking
In 1968, Paul Brown and Herbert Jenkins placed naive pigeons in an operant chamber with an illuminated response key. Every 60 seconds, the key was illuminated with a white light for 8 seconds, followed immediately by access to grain, with no response required (a pure classical Pavlovian pairing: key light food):
- Rather than waiting passively by the hopper, the pigeons spontaneously approached and began pecking the illuminated key at high rates. This phenomenon was termed autoshaping (automated shaping).
Hearst and Jenkins (1974) named this cue-directed behavior sign-tracking; Boakes (1977) described the contrasting goal-tracking pattern, and later work by Flagel and Robinson treated the two as stable individual differences:
- Sign-Trackers: Organisms that direct attention, approach, and physical contact toward the predictive conditional cue (the light), treating the cue as if it were the consummatory object itself. Sign-tracking is mediated by robust dopamine surges in the nucleus accumbens core and predicts vulnerability to cue-induced substance addiction, impulsive gambling, and relapse.
- Goal-Trackers: Organisms that use the predictive cue merely as an informational signal, immediately approaching and waiting at the food hopper where the reward will be delivered.
A clinical behavior analyst implements a treatment plan for a patient with intellectual disabilities who engages in frequent skin-picking. The analyst instructs staff to deliver social praise and access to an iPad contingent upon the patient sitting with both hands holding a rubber stress-ball. What specific differential reinforcement procedure is being utilized?
Differential Reinforcement of Low rates of responding (DRL)
Differential Reinforcement of Other behavior (DRO)
Differential Reinforcement of Alternative behavior (DRA)
Differential Reinforcement of Incompatible behavior (DRI)
An animal in an operant chamber produces a cumulative record characterized by prolonged pauses following the delivery of each reinforcer, followed by a sudden, rapid burst of responding until the next reinforcer is attained. What reinforcement schedule generated this record?
Variable Ratio (VR)
Fixed Ratio (FR)
Variable Interval (VI)
Fixed Interval (FI)
According to the Response Deprivation Hypothesis developed by Timberlake and Allison, under what condition can a baseline low-probability activity function as a reinforcer for a baseline high-probability activity?
When the operant contingency restricts the organism's access to the low-probability behavior below its free-baseline level
When the high-probability behavior undergoes classical extinction before the contingency is introduced in a second phase
When the low-probability behavior is paired with a secondary generalized reinforcer such as tokens
When the two behaviors are linked in a backward chain terminating in an unconditioned biological reinforcer
A pig is trained through operant shaping to retrieve large wooden tokens and deposit them into an automated bank to receive food. Over hundreds of successful trials, the pig's latency to drop the token gradually increases as it begins repeatedly dropping the coin, rooting it into the floor, and tossing it with its snout, causing it to miss scheduled feedings. What behavioral phenomenon accounts for this breakdown in operant performance?
The peak shift phenomenon during stimulus discrimination
Ratio strain resulting from an overly lean intermittent schedule
The Partial Reinforcement Extinction Effect (PREE)
Instinctive drift toward species-typical phylogenetic feeding reflexes
Sections you finish are checked off in the contents.