1.3 Reinforcement Schedules and Stimulus Control
Key Takeaways
- Reinforcement is defined functionally by its effect: a stimulus change following a behavior that increases or maintains the future frequency of that behavior.
- Positive reinforcement involves adding a stimulus; negative reinforcement involves removing or attenuating an aversive stimulus (via escape or avoidance).
- Intermittent schedules of reinforcement (FR, VR, FI, VI) produce characteristic patterns of responding, with Variable Ratio (VR) producing high, steady rates most resistant to extinction.
- Herrnstein's Matching Law demonstrates that on concurrent schedules, the relative rate of responding matches the relative rate of reinforcement provided by each alternative.
- Stimulus control is established through differential reinforcement; prompt fading procedures (e.g., time delay, most-to-least) systematically transfer stimulus control from artificial prompts to natural discriminative stimuli.
Reinforcement Schedules and Stimulus Control
Reinforcement is the single most fundamental principle of operant behavior. For a Qualified Behavior Analyst, understanding how reinforcement operates across basic schedules, choice environments, and antecedent stimulus arrangements is necessary for designing effective skill acquisition and behavior reduction plans.
Principles of Positive and Negative Reinforcement
In behavior analysis, reinforcement is defined strictly by its functional outcome on behavior, never by formal intentions or subjective pleasantness. Reinforcement has occurred only if a stimulus change immediately follows a response and results in an increase or maintenance of the future frequency, rate, or intensity of that response.
1. Positive Reinforcement ($S^{R+}$)
Occurs when a response is followed immediately by the presentation or addition of a stimulus, resulting in an increase in the future frequency of that response under similar conditions.
2. Negative Reinforcement ($S^{R-}$)
Occurs when a response is followed immediately by the termination, removal, reduction, or postponement of a stimulus, resulting in an increase in the future frequency of that response. Negative reinforcement is divided into two clinical mechanisms:
- Escape Behavior: The response terminates an ongoing, currently present aversive stimulus (e.g., opening an umbrella while standing in active rain).
- Avoidance Behavior: The response prevents or postpones the delivery of an impending aversive stimulus before it occurs (e.g., checking weather radar and staying indoors before rain begins).
Primary vs. Secondary Reinforcers & Generalized Conditioned Reinforcers
Reinforcers are classified by how they acquire their reinforcing properties:
- Unconditioned (Primary) Reinforcers: Stimuli that function as reinforcers without any prior learning history (phylogenic origin). Examples include food, water, warmth, oxygen, and sexual stimulation.
- Conditioned (Secondary) Reinforcers: Neutral stimuli that acquire reinforcing properties through repeated pairing with primary or established secondary reinforcers (ontogenic origin). Examples include sticker charts, points, and verbal praise.
- Generalized Conditioned Reinforcers (GCSRs): A type of conditioned reinforcer that has been paired with many unconditioned and conditioned reinforcers across diverse contexts. GCSRs do not depend on a specific MO state to be effective. The classic clinical example is money or tokens in a token economy, which can be exchanged for a wide array of back-up reinforcers.
Intermittent Schedules of Reinforcement
A schedule of reinforcement is a rule specifying which occurrences of a target behavior will be reinforced. While continuous reinforcement (CRF / FR1) is optimal for initial skill acquisition, intermittent schedules are used to maintain established behaviors and build resistance to extinction.
The four basic intermittent schedules are defined along two axes: Ratio (based on number of responses) vs. Interval (based on time elapsed before a response is reinforced), and Fixed (predictable requirement) vs. Variable (unpredictable average requirement):
| Schedule Type | Rule / Requirement | Performance Pattern | Extinction Resistance | Clinical Application |
|---|---|---|---|---|
| Fixed Ratio (FR) | Reinforcement delivered after a specific, constant number of responses. | High response rate with a post-reinforcement pause (PRP) immediately following reinforcement delivery. | Low to Moderate | Task completion quotas (e.g., FR5 = complete 5 math problems for a break). |
| Variable Ratio (VR) | Reinforcement delivered after an average, unpredictable number of responses. | High, steady, uninterrupted rate of responding; virtually no post-reinforcement pause. | Highest (Most resistant to extinction) | Slot machines, sports play, fluency building in token systems. |
| Fixed Interval (FI) | Reinforcement delivered for the first correct response after a fixed amount of time has elapsed. | FI Scallop: Low responding early in the interval, accelerating to a rapid rate as the interval end approaches. | Low | Checking an oven timer; studying heavily right before scheduled exams. |
| Variable Interval (VI) | Reinforcement delivered for the first correct response after an average, unpredictable time interval. | Moderate, steady, continuous rate of responding without pauses. | High | Surprise pop quizzes; checking email; supervisor spot-checks. |
Herrnstein's Matching Law and Concurrent Schedules
In natural clinical environments, individuals rarely encounter isolated single schedules. Instead, they choose between multiple simultaneously available options—a concurrent schedule of reinforcement.
Formulated by Richard Herrnstein (1961), The Matching Law states that on concurrent schedules of reinforcement, the relative rate of responding matches the relative rate of reinforcement provided by each schedule:
Where $R_1$ and $R_2$ represent the rates of response on options 1 and 2, and $r_1$ and $r_2$ represent the rates of reinforcement delivered by options 1 and 2.
Clinical Implication for QBAs
If a student can obtain reinforcement on a VI 30-second schedule for challenging behavior (screaming) and a VI 120-second schedule for functional communication (manding), the student will allocate approximately 80% of their responding to screaming. To reduce challenging behavior, the QBA must alter reinforcer parameters (rate, quality, delay, effort) so that alternative appropriate behavior receives a significantly higher density of reinforcement.
Stimulus Control, Generalization, and Discrimination
Stimulus control occurs when the rate, latency, duration, or amplitude of a response is altered in the presence of an antecedent stimulus. Stimulus control is established through differential reinforcement:
- Discriminative Stimulus ($S^D$): Reinforcement is delivered for the response in its presence.
- Stimulus Delta ($S^\Delta$): Reinforcement is withheld (extinction) or punished when the response occurs in its presence.
Discrimination vs. Generalization
- Stimulus Discrimination: The organism responds only in the presence of the $S^D$ and not in the presence of $S^\Delta$ (tight stimulus control).
- Stimulus Generalization: The response occurs in the presence of novel antecedent stimuli that share similar physical properties with the original $S^D$. A stimulus generalization gradient shows that as physical similarity to the $S^D$ decreases, response rate drops.
Prompting Strategies and Transfer of Stimulus Control
When a learner does not yet respond to the natural $S^D$, clinicians use prompts—supplemental antecedent stimuli used to evoke the correct response.
Prompt Types
- Response Prompts: Operate directly on the response (Vocal, Modeling, Physical guidance).
- Stimulus Prompts: Operate directly on the antecedent stimuli (Movement, Position, Redundancy/exaggeration).
Transfer of Stimulus Control (Prompt Fading)
To achieve true independence, stimulus control must be transferred from the prompt to the natural $S^D$. Common fading procedures include:
- Most-to-Least Prompting: Begins with maximum physical guidance and systematically fades to lesser prompts as mastery is achieved (prevents errors during acquisition).
- Least-to-Most Prompting: Gives the learner an opportunity to respond to the natural $S^D$ independently, delivering progressively intrusive prompts only if errors or no responses occur.
- Graduated Guidance: Provides physical prompts as needed moment-by-moment and immediately shadows/fades within a trial.
- Time Delay (Constant or Progressive): Inserts a fixed or gradually increasing time interval between the natural $S^D$ and the delivery of the prompt, allowing independent responding to emerge.
A student in a special education classroom exhibits high rates of hand-raising under a Variable Ratio 5 (VR5) schedule of teacher praise. The teacher decides to change the schedule to a Fixed Ratio 15 (FR15). Immediately after receiving praise, the student stops raising their hand for extended periods before resuming. This post-reinforcement pause is a characteristic pattern of which schedule?
A client is working on a concurrent schedule where alternative A provides reinforcement every 30 seconds (VI 30s) and alternative B provides reinforcement every 90 seconds (VI 90s). According to Herrnstein's Matching Law, what percentage of total responses will the client allocate to alternative A?
A QBA is teaching a child to touch a picture of a dog when presented with the spoken word 'Dog'. Initially, the QBA points directly to the correct picture immediately upon giving the vocal prompt. Over trials, the QBA waits 0 seconds, then 2 seconds, then 4 seconds between the vocal prompt and pointing. What procedure is being used to transfer stimulus control?