2.2 Positive and Negative Reinforcement & Schedules of Reinforcement

Key Takeaways

  • Reinforcement is defined solely by its functional effect: a contingent stimulus change that increases the future probability, frequency, or rate of the target response class.

  • Positive reinforcement contingently presents a stimulus, whereas negative reinforcement contingently removes, reduces, or postpones an aversive stimulus through escape or avoidance contingencies.

  • Generalized conditioned reinforcers (e.g., money, tokens) are paired with diverse backup reinforcers, so their effectiveness does not depend on any single motivating operation and they are less susceptible to satiation.

  • Intermittent schedules produce signature response patterns: Fixed Ratio yields post-reinforcement pauses and high run rates; Variable Ratio produces high, steady rates; Fixed Interval yields scalloped acceleration; Variable Interval yields moderate, stable rates.

  • Socially mediated reinforcement is delivered by another person (attention, items, escape from demands), whereas automatic reinforcement is produced directly by the response itself (e.g., sensory stimulation or relief from discomfort).

Last updated: October 2026

The Functional Definition of Reinforcement

In applied behavior analysis, reinforcement is not a thing, a reward, or an event; it is an empirical, functional relationship between an environmental change and behavior. Reinforcement occurs when a stimulus change immediately follows a response and results in an increase in the future frequency of that response class under similar antecedent conditions (other dimensions can also change in the reinforced direction; for example, latency can shorten or duration can lengthen).

Two critical principles govern reinforcement:

  1. The Functional Principle: A stimulus change is defined as a reinforcer solely by its demonstrated effect on future behavior, never by the intentions of the person delivering it, cultural norms, or subjective pleasantness. Giving a child a chocolate candy or high-five is not reinforcement unless reliable behavioral data confirm that the target behavior increases in future frequency. If handing a child a sticker results in no change or a decrease in future responding, the sticker is functionally ineffective or punitive, regardless of how "rewarding" the practitioner intended it to be.
  2. The Automaticity of Reinforcement: Operant conditioning does not require conscious awareness, verbal understanding, or intentional cognitive mediation by the organism. Reinforcement works automatically; whenever an environmental consequence is contingent on a response, that consequence selects the response class regardless of whether the individual can state the contingency.

Positive Reinforcement vs. Negative Reinforcement

The distinction between positive and negative reinforcement rests entirely on the nature of the stimulus change that immediately succeeds the response. Both positive and negative reinforcement produce the exact same behavioral outcome: an increase in the future frequency of the target behavior.

  • Positive Reinforcement (SR+S^{R+}): Occurs when a response is immediately followed by the presentation, addition, or increase in intensity of a stimulus, which increases the future frequency of that behavior in similar conditions (e.g., a child asks for juice, and the parent delivers juice; future asking behavior increases).
  • Negative Reinforcement (SR−S^{R-}): Occurs when a response is immediately followed by the removal, termination, reduction, attenuation, or postponement of an aversive stimulus, which increases the future frequency of that behavior in similar conditions (e.g., an individual turns down a blaring stereo to eliminate ear pain; future volume-lowering behavior increases).

Exam Trap Alert: Candidates frequently conflate negative reinforcement with punishment. The mathematical sign indicates the environmental operation, not the behavioral effect. Positive means stimulus added (+); Negative means stimulus removed (-). Reinforcement always means future behavior increases; punishment always means future behavior decreases. Negative reinforcement strengthens behavior.


Negative Reinforcement Subclasses: Escape vs. Avoidance

Negative reinforcement contingencies operate through two distinct functional subclasses: escape behavior and avoidance behavior.

1. Escape Behavior

In an escape contingency, the aversive stimulus is already physically present and ongoing in the environment before the response is emitted. The organism's response terminates, attenuates, or interrupts the active aversive stimulus.

  • Clinical Example: A client with sensory hypersensitivity enters a bustling cafeteria where overhead fluorescent lights are flickering violently and ambient chatter is deafening. The client immediately runs out the emergency side door to the quiet courtyard. The response (running out the door) terminates the ongoing sensory pain. Because the aversive condition was already active when the response occurred, this is an escape repertoire.

2. Avoidance Behavior

In an avoidance contingency, the response occurs before the aversive stimulus is experienced, thereby preventing, postponing, or delaying its onset entirely. Avoidance behavior is subdivided into two critical experimental and clinical paradigms:

  • Discriminated Avoidance: An explicit antecedent warning stimulus (a conditioned aversive stimulus or SDS^D) signals that an aversive event will occur unless an operant response is emitted within a specified time window. Emitting the response during the warning signal prevents the aversive stimulus from occurring.
    • Clinical Example: A student observes the teacher pick up a stack of red folders and announce, "Clear your desks for the weekly timed algebra quiz" (warning stimulus). The student immediately raises their hand and asks to visit the school nurse with a headache. The teacher excuses the student before the quiz is distributed. The student's behavior successfully prevented the presentation of the academic demand.
  • Free-Operant Avoidance (Sidman Avoidance): In free-operant avoidance (first demonstrated by Murray Sidman, 1953), responding can occur at any time in the complete absence of an explicit antecedent warning signal. Shocks or aversive events are programmed to occur periodically according to a set time interval (the shock-shock or S-S interval). Each time the organism emits the target response (e.g., pressing a lever), it resets an avoidance timer (the response-shock or R-S interval), postponing the next scheduled aversive event.
    • Applied Example: An automobile driver periodically checks their tire pressure and engine oil level every two weeks without any dashboard warning light illuminating. Each proactive maintenance response resets the risk interval, preventing sudden mechanical breakdown on the highway.

Automatic vs. Socially Mediated Contingencies

Task B.6 asks you to distinguish socially mediated contingencies from automatic contingencies. The distinction is about how the consequence is delivered, not about whether it is pleasant.

  • Socially mediated reinforcement: the consequence is delivered through the behavior of another person. A teacher's reprimand, a parent handing over a tablet, or a therapist removing a worksheet are socially mediated. If the other person stopped responding, the contingency would disappear.
  • Automatic reinforcement: the consequence is produced directly by the response itself, without anyone else's involvement. Humming produces sound, rocking produces vestibular stimulation, and scratching relieves an itch. The behavior would continue to produce the consequence even when the person is completely alone.

Each type can be positive or negative:

ContingencyConsequenceExample
Social positiveAnother person adds a stimulusScreaming produces adult attention or a toy
Social negativeAnother person removes a stimulusHitting leads the teacher to withdraw a demand
Automatic positiveThe response itself adds stimulationHand-flapping produces visual stimulation
Automatic negativeThe response itself reduces stimulationRubbing an aching ear attenuates pain

Punishment can also be automatic (touching a hot pan) or socially mediated (a reprimand). The distinction matters clinically because socially mediated behavior can be treated by changing what other people do (extinction, differential reinforcement of an alternative that produces the same social consequence), whereas automatically maintained behavior requires changing the stimulation itself (matched stimulation, sensory extinction, enrichment). On a functional analysis, responding that persists in the alone condition suggests automatic reinforcement.


Reinforcer Taxonomy: Unconditioned, Conditioned, and Generalized Reinforcers

Reinforcing stimuli are categorized according to their evolutionary and ontogenic development:

  1. Unconditioned Reinforcers (Primary / Unlearned): Stimuli whose reinforcing efficacy is an innate product of phylogenic evolution and requires no prior learning history. These stimuli possess biological survival value, such as food, water, thermal warmth, sleep, sexual stimulation, oxygen, and relief from physical pain or noxious pressure. Unconditioned reinforcers depend heavily on biological deprivation states (motivating operations).
  2. Conditioned Reinforcers (Secondary / Learned): Initially neutral stimuli that acquire reinforcing properties through repeated pairing with established unconditioned or conditioned reinforcers across an organism's lifetime. Examples include verbal praise, stickers, ringing chimes, social nods, and good grades.
  3. Generalized Conditioned Reinforcers (GCRs): A specialized and potent class of conditioned reinforcers that have been paired with multiple diverse backup reinforcers (both primary and secondary). Because GCRs are exchangeable for a wide variety of reinforcers, their reinforcing efficacy does NOT depend on any single motivating operation or state of deprivation. GCRs are exceptionally resistant to satiation.
    • Examples: Money ($100 bills), classroom tokens, arcade tickets, airline frequent-flyer miles, and social attention/approval (which has historically been paired with food, comfort, physical safety, and play).

Basic Schedules of Reinforcement

A schedule of reinforcement is an environmental rule specifying the environmental conditions (number of responses or temporal intervals) required to produce reinforcement. Schedules exist along a continuum from Continuous Reinforcement (CRF / FR1) to Extinction (EXT).

  • Continuous Reinforcement (CRF / FR1): Reinforcement is delivered for every single correct occurrence of the target behavior. CRF is utilized primarily during the acquisition phase of skill building. However, behavior maintained on CRF extinguishes rapidly once reinforcement ceases.
  • Intermittent Reinforcement (INT): Only some occurrences of the target behavior produce reinforcement. Intermittent schedules are used to maintain established behaviors, facilitate schedule thinning, and build resistance to extinction.

There are four basic intermittent schedules formed by crossing two dimensions: Ratio vs. Interval and Fixed vs. Variable.

1. Fixed Ratio (FR)

Reinforcement is delivered contingent on a set, unchanging number of responses.

  • Response Pattern: Characterized by a post-reinforcement pause (PRP) immediately following the delivery of reinforcement, followed by a rapid, steady "run rate" of responding until the next ratio requirement is completed (producing a "break-and-run" cumulative record shape). The length of the PRP is directly proportional to the magnitude of the ratio requirement.
  • Ratio Strain: If the ratio requirement is thinned too rapidly or set too high relative to the learner's reinforcement history, the behavior may collapse, resulting in extended pausing, avoidance, emotional outbursts, and schedule breakdown.

2. Variable Ratio (VR)

Reinforcement is delivered contingent on an average, variable number of responses across trials (e.g., on a VR5 schedule, reinforcement might be delivered after 3, 7, 2, and 8 responses, averaging 5 responses).

  • Response Pattern: Generates a consistent, steady, high rate of responding with no post-reinforcement pause. Because the organism cannot predict which specific response will yield reinforcement, responding remains uninterrupted.
  • Resistance to Extinction: Variable schedules (VR and VI) generally produce more resistance to extinction than fixed schedules or continuous reinforcement. Slot machines and recreational gambling operate on VR schedules.

3. Fixed Interval (FI)

Reinforcement is delivered for the first response emitted after a fixed, predetermined duration of time has elapsed since the previous reinforcer. The passage of time alone does not deliver the reinforcer; an operant response must be emitted following interval expiration.

  • Response Pattern: Produces a post-reinforcement pause during the initial portion of the interval, followed by an accelerating response rate as the end of the interval approaches. On a cumulative record, this generates a distinctive scalloped pattern (the "FI scallop").
  • Applied Example: Checking an oven for baking cookies when the timer is set for 20 minutes; responding begins slowly and accelerates dramatically during the 19th minute.

4. Variable Interval (VI)

Reinforcement is delivered for the first response emitted after a variable, unpredictable interval of time has elapsed, averaging a specific duration (e.g., on a VI 3-minute schedule, intervals between reinforcement availability vary around a 3-minute mean).

  • Response Pattern: Generates a moderate, highly steady and uniform rate of responding with no post-reinforcement pauses. On a cumulative record, VI responding displays a straight, moderate diagonal slope.
  • Applied Example: Checking email throughout the day or waiting for text messages; responses occur at a steady, consistent rate because delivery is unpredictable.

The Matching Law (Herrnstein, 1961)

In natural environments, organisms rarely operate under isolated schedules; they face concurrent schedules of reinforcement, where two or more response options with independent reinforcement schedules are available simultaneously. Richard Herrnstein (1961) discovered that organisms allocate their behavior across concurrent schedules in direct proportion to the relative rates of reinforcement delivered by each schedule.

The Mathematical Formulation

B1B1+B2=R1R1+R2\frac{B_1}{B_1 + B_2} = \frac{R_1}{R_1 + R_2}

Where:

  • B1B_1 and B2B_2 represent the rates of responding allocated to Alternative 1 and Alternative 2.
  • R1R_1 and R2R_2 represent the rates of reinforcement obtained from Alternative 1 and Alternative 2.

If Alternative 1 delivers 75% of the total reinforcement (R1/(R1+R2)=0.75R_1 / (R_1 + R_2) = 0.75), the organism will allocate approximately 75% of its total responses to Alternative 1 (B1/(B1+B2)=0.75B_1 / (B_1 + B_2) = 0.75).

Clinical Significance for BCaBAs

The Matching Law has profound implications for treating severe problem behavior. When implementing Differential Reinforcement of Alternative Behavior (DRA) or Functional Communication Training (FCT), teaching a replacement behavior (B1B_1) will fail if problem behavior (B2B_2) continues to access reinforcement on a richer, denser, or more immediate schedule (R2R_2). To eliminate problem behavior, behavior analysts must use the Matching Law strategically: make reinforcement for the appropriate alternative significantly denser, higher quality, more immediate, and lower in response effort than the reinforcement available for the problem behavior.


Schedules of Reinforcement Matrix

ScheduleContingency RuleCumulative Graph PatternPost-Reinforcement Pause?Resistance to ExtinctionClinical / Applied Example
Continuous (CRF / FR1)Reinforcer delivered after every single response.Moderate, steady rate during acquisition.No pause.Very low; rapid extinction upon discontinuation.Delivering an edible and praise every time a toddler independently matches an identical picture card.
Fixed Ratio (FR)Reinforcer delivered after a fixed number of NN responses.Rapid run rate with abrupt post-reinforcement pauses ("break-and-run").Yes; pause duration proportional to ratio size.Low to moderate.A factory piece-rate worker earning $10 for every 20 garments sewn (FR20).
Variable Ratio (VR)Reinforcer delivered after a variable number of responses averaging NN.Steep, unbroken, highly steady rate without pauses.No pause.High; variable schedules resist extinction more than fixed schedules or CRF.Asking prospective clients for sales contracts; slot machines; intermittent praise for completed math problems.
Fixed Interval (FI)Reinforcer delivered for first response after a fixed time TT elapses.Accelerating response rate as interval ends ("FI scallop").Yes; pause immediately following reinforcement.Moderate.Checking laundry machine when a 45-minute cycle is nearly finished; studying behavior before scheduled Friday quizzes.
Variable Interval (VI)Reinforcer delivered for first response after a variable time averaging TT.Moderate, stable, continuous rate without pausing.No pause.High; persistent steady responding.A supervisor dropping by unannounced at unpredictable intervals to provide verbal praise for on-task work.

Unwanted Effects of Reinforcement

Reinforcement-based procedures are preferred, but task H.5 expects you to anticipate their unwanted effects too:

  • Reinforcing the wrong behavior: delivering the reinforcer near problem behavior (or on a time-based schedule just after it) can strengthen the problem behavior.
  • Satiation: heavy use of one reinforcer lowers its value, so performance drops.
  • Dependence on contrived reinforcers: behavior that occurs only when tokens or edibles are present will not maintain in natural settings unless reinforcement is thinned and transferred to natural consequences.
  • Behavioral contrast: increasing reinforcement for a behavior in one setting can reduce it in another setting where reinforcement did not change.
  • Problems during thinning: thinning too fast can produce ratio strain, extinction-like bursts, or resurgence of previously reinforced problem behavior.

Plan for these effects up front: rotate reinforcers, pair contrived reinforcers with natural ones, thin gradually, and monitor data across settings.


Common BCaBA Exam Traps: Schedules and Reinforcement Principles

  • Trap 1: Interval vs. Ratio Timing: In interval schedules, reinforcement is NOT delivered automatically when the clock expires! The time interval simply sets the occasion for reinforcement; the organism must emit the operant response after the interval elapses to produce the reinforcer.
  • Trap 2: Praise as Reinforcement: Never assume social praise is reinforcing. For individuals with social anxiety, trauma history, or autism spectrum conditions, social praise may be neutral, an establishing operation for escape, or an aversive stimulus.
  • Trap 3: Ratio Strain vs. Extinction: When an FR schedule is thinned too quickly (e.g., jumping from FR2 to FR25), responding may drop to zero. This is ratio strain due to schedule thinning, not operant extinction (which requires withholding reinforcement completely).
Loading diagram...
Taxonomy and Classification of Basic Reinforcement Schedules
Test Your Knowledge

A warehouse logistics manager installs an audible alarm system. Whenever warehouse temperatures exceed 85 degrees Fahrenheit, a shrill bell sounds continuously. When workers pull an overhead cooling lever, the bell immediately turns off and the temperature rapidly cools. In response, workers pull the cooling lever as soon as the bell begins ringing. What behavioral process maintains the workers' lever-pulling behavior?

A

Free-operant avoidance, because workers pull the lever periodically in the absence of any environmental stimulus change.

B

Discriminated avoidance, because workers pull the lever before any aversive condition is experienced.

C

Escape behavior, because the aversive auditory stimulus was already present and ongoing, and pulling the lever terminated it.

D

Positive reinforcement, because pulling the lever delivers refreshing cool air to the warehouse.

Test Your Knowledge

A BCaBA observes a technician implementing a token economy with an adult learner in a vocational training center. The cumulative record shows that the learner works at an exceptionally high, unbroken rate, assembling electrical connectors without any pauses after token delivery. Reinforcement is delivered after varying numbers of completed assemblies—sometimes after 8 assemblies, sometimes after 16, and sometimes after 12—averaging 12 completed units. Which schedule of reinforcement is maintaining this client's performance?

A

Variable ratio 12 (VR12), which generates high, steady rates of responding without post-reinforcement pauses.

B

Variable interval 12-minute (VI12), which generates a moderate, steady rate of responding determined solely by the passage of time.

C

Fixed ratio 12 (FR12), which generates rapid rates of responding punctuated by regular post-reinforcement pauses.

D

Fixed interval 12-minute (FI12), which produces an accelerating scalloped pattern as the interval expires.

Test Your Knowledge

An adolescent client in a special education classroom displays two concurrent behaviors maintained by adult attention: asking questions appropriately (B1B_1) and disruptive desk-slamming (B2B_2). The behavior analyst collects baseline data and observes that desk-slamming receives adult attention on a dense schedule delivering 40 reinforcers per hour (R2=40R_2 = 40), while appropriate questions receive attention on a schedule delivering 20 reinforcers per hour (R1=20R_1 = 20). According to Herrnstein's Matching Law, what percentage of the student's total responding will be allocated to appropriate question-asking (B1B_1)?

A

100% of total responses, because concurrent interval schedules always lead to exclusive preference for the denser contingency.

B

Exactly 50.0% of total responses, because both behaviors are concurrently active within the same instructional environment.

C

Approximately 66.7% of total responses, because learners naturally prefer socially acceptable behaviors over disruptive topographies.

D

Approximately 33.3% of total responses, matching the relative rate of reinforcement obtained for appropriate questions.

Sections you finish are checked off in the contents.