4.2 Ratio, Interval & Variable Schedules in Dog Training

Key Takeaways

  • The four simple intermittent schedules are Fixed Ratio (FR), Variable Ratio (VR), Fixed Interval (FI), and Variable Interval (VI), each generating distinct cumulative record curves and response dynamics.
  • Fixed Ratio (FR) schedules produce a 'break-and-run' pattern with post-reinforcement pauses (PRPs) proportional to ratio size, whereas Variable Ratio (VR) schedules produce high, steady, pause-free response rates.
  • Fixed Interval (FI) schedules produce a characteristic 'scallop' where response rates accelerate as the time threshold approaches, whereas Variable Interval (VI) schedules maintain moderate, remarkably stable response rates.
  • Differential reinforcement schedules (DRH, DRL, DRD) control behavioral pacing, latency, and response velocity by reinforcing specific inter-response times (IRTs).
  • Herrnstein's Matching Law demonstrates that organisms allocate behavior across concurrent schedules in direct proportion to the relative rate, immediacy, and magnitude of reinforcement obtained from each alternative.
Last updated: September 2026

4.2 Ratio, Interval & Variable Schedules in Dog Training

Quick Answer: Intermittent schedules of reinforcement are categorized across two fundamental dimensions: Ratio vs. Interval (response count vs. elapsed time) and Fixed vs. Variable (predictable vs. randomized around a mean). The four simple schedules—Fixed Ratio (FR), Variable Ratio (VR), Fixed Interval (FI), and Variable Interval (VI)—generate distinct behavioral topographies, ranging from the 'break-and-run' pattern of FR to the high, pause-free responding of VR. When competing environmental reinforcers distract a dog, Herrnstein's Matching Law dictates that the dog will distribute its behavioral choices in direct proportion to the relative reinforcement rates delivered by each concurrent option.


The Two-by-Two Schedule Classification Matrix

Every simple intermittent schedule of reinforcement is defined by two foundational operational choices:

  1. The Controlling Metric (Ratio vs. Interval):
    • Ratio Schedule: Reinforcement delivery is strictly contingent upon the number of correct operant responses emitted by the organism. Elapsed time has zero bearing on reward delivery; if the animal responds faster, rewards are earned faster.
    • Interval Schedule: Reinforcement delivery is contingent upon the first correct operant response emitted after a specified duration of time has elapsed. Any responses emitted prior to the expiration of the time window produce no reinforcement.
  2. The Temporal Predictability (Fixed vs. Variable):
    • Fixed Schedule: The requirement (response count or time duration) remains identical and constant from one reinforcement delivery to the next.
    • Variable Schedule: The requirement varies from trial to trial, fluctuating unpredictably around a predetermined arithmetic mean ($n$).

Intersecting these two axes produces the classic two-by-two matrix of simple intermittent reinforcement schedules originally cataloged in Skinner's cumulative recorder experiments:

                    ┌───────────────────────┬───────────────────────┐
                    │         RATIO         │       INTERVAL        │
                    │    (Response Count)   │     (Elapsed Time)    │
┌───────────────────┼───────────────────────┼───────────────────────┤
│       FIXED       │   Fixed Ratio (FR)    │  Fixed Interval (FI)  │
│   (Predictable)   │   "Break-and-run"     │      "Scallop"        │
├───────────────────┼───────────────────────┼───────────────────────┤
│     VARIABLE      │  Variable Ratio (VR)  │ Variable Interval (VI)│
│   (Unpredictable) │ High, steady, no pause│ Moderate, steady rate │
└───────────────────┴───────────────────────┴───────────────────────┘

Detailed Analysis of the Four Simple Schedules

1. Fixed Ratio (FR) Schedules

In a Fixed Ratio (FR) schedule, a reinforcer is delivered contingent upon the completion of a predetermined, unvarying number of correct responses. For example, an FR5 schedule requires exactly 5 correct responses for every reward delivery.

  • Behavioral Topography & Cumulative Record: FR schedules produce a characteristic "break-and-run" pattern. Once the animal begins responding, it executes the responses at a rapid, uninterrupted rate until reinforcement is delivered (the "run"). Immediately following reinforcement, responding halts temporarily (the "break").
  • The Post-Reinforcement Pause (PRP): The pause immediately following reward delivery is the hallmark of fixed schedules. The duration of the PRP is directly proportional to the magnitude of the ratio requirement: an FR3 produces a barely perceptible pause, while an FR25 produces an extended pause. Research demonstrates that the PRP is not caused by physical fatigue; rather, it represents a period of anticipatory non-reinforcement (the animal has learned that the immediate next response never produces a reward).
  • Canine Training Applications: Requiring a dog to complete 3 agility jumps before receiving a toy tug, or heels for 10 paces before receiving a treat. Trainers must be cautious: high FR schedules breed anticipation and post-reinforcement procrastination.

2. Variable Ratio (VR) Schedules

In a Variable Ratio (VR) schedule, reinforcement is delivered after an unpredictable number of responses, fluctuating around a specific average. For instance, on a VR5 schedule, reward delivery might occur after 2 responses, then 7 responses, then 4 responses, then 8 responses, then 4 responses (arithmetic mean = 5).

  • Behavioral Topography & Cumulative Record: VR schedules generate the highest, steepest, and most consistent rates of responding of all simple schedules. Crucially, VR schedules display virtually zero post-reinforcement pauses.
  • The "Gambler's Schedule" / Slot Machine Effect: Because the animal cannot predict which specific response will trigger reinforcement—any single response could be the winning trial—the animal maintains continuous, energetic behavioral output. It represents the psychological engine underlying gambling, predatory search sequences, and canine ball-fetching addictions.
  • Canine Training Applications: The gold standard for maintaining established behaviors such as the formal obedience recall, scent discrimination indications, and loose-leash walking. The unpredictable delivery keeps the dog mentally alert and enthusiastically engaged.

3. Fixed Interval (FI) Schedules

In a Fixed Interval (FI) schedule, reinforcement is delivered for the first correct response emitted after a fixed, constant time interval has elapsed. For example, on an FI30s schedule, the feeder activates on the first response occurring after 30 seconds. Responses emitted at second 5, 12, or 28 produce no consequence and do not reset the timer.

  • Behavioral Topography & Cumulative Record: FI schedules generate the distinctive "FI scallop". Immediately following reinforcement delivery, the animal pauses (a long post-reinforcement pause). As the internal clock signals that the interval is drawing to a close, response rate begins to accelerate slowly, reaching a frenzied, rapid peak right as the interval expires.
  • Temporal Discriminations: Organisms develop remarkable internal timing estimations. If a dog on an FI60s schedule receives a treat, it will rest or engage in displacement sniffing for 45 seconds, then abruptly return to the handler and emit high-rate sits or paw-targets as the 60-second mark approaches.
  • Canine Training Applications: Stay exercises where an inexperienced handler returns to the dog at exact, predictable intervals (e.g., returning to feed every 15 seconds). The dog anticipates the return, exhibiting mounting restlessness, tail-wagging, and muscular tension (scalloping) just before the 15-second threshold.

4. Variable Interval (VI) Schedules

In a Variable Interval (VI) schedule, reinforcement is delivered for the first response emitted after an unpredictable duration of time has elapsed, varying around a specified average. On a VI20s schedule, intervals between reward availability might be 5s, 35s, 15s, 40s, and 5s (average = 20s).

  • Behavioral Topography & Cumulative Record: VI schedules produce a moderate, remarkably stable, uniform, and pause-free rate of responding. The cumulative recorder line is straight, smooth, and resistant to fluctuations.
  • Behavioral Calmness: Unlike VR schedules that encourage high-speed, energetic physical output, VI schedules foster calm, steady, persistent focus. Because rewards are governed by time rather than response count, sprinting through repetitions does not produce rewards any faster.
  • Canine Training Applications: Teaching a dog to settle quietly on a mat in public, maintaining sustained eye contact during heelwork, or holding an endurance down-stay while the handler cleans the room.

Differential Reinforcement Schedules of Rate

Beyond simple ratio and interval schedules, behavior analysts employ differential reinforcement schedules of rate to directly manipulate the speed, tempo, and inter-response times (IRTs) of behavior:

Differential Reinforcement of High Rates (DRH)

Reinforcement is delivered only if a specified number of responses are emitted within a strictly defined, brief time window, or if the inter-response time (IRT) between consecutive behaviors is shorter than a specified duration.

  • Applied Canine Goal: Elevating response velocity and enthusiasm.
  • Example: In competition agility or flyball, a trainer rewards a retrieve or recall only if the dog returns within 3.5 seconds of the cue. If the dog returns at 4.2 seconds, the behavior is acknowledged neutrally without reinforcement. DRH systematically drives speed.

Differential Reinforcement of Low Rates (DRL)

Reinforcement is delivered only if a minimum time duration has elapsed since the previous response. If the animal emits the response prematurely, the response is unreinforced and the timer resets.

  • Applied Canine Goal: Decreasing the frequency of frantic, hyperactive, or compulsive behaviors without eliminating them entirely.
  • Example: A dog that frantically paws at the handler's leg for attention. Under a DRL10s schedule, pawing is reinforced only if at least 10 seconds have elapsed since the last paw touch. Pacing out the behavior teaches impulse control and reduces behavioral intensity.

Differential Reinforcement of Diminishing Rates (DRD)

Reinforcement is delivered at the end of an interval if the total number of responses emitted is fewer than a specified criterion, with the threshold ratcheted progressively downward over successive sessions.

  • Applied Canine Goal: Systematic, graduated reduction of nuisance behaviors (e.g., reducing excessive alert barking from 20 barks per minute, down to 10, then 5, then 1).

Herrnstein's Matching Law & Concurrent Schedules

In natural environments, dogs are never presented with a single, isolated reinforcement schedule. Instead, they operate under concurrent schedules of reinforcement—simultaneous, independent schedules operating for two or more alternative behaviors.

In 1961, Harvard psychologist Richard Herrnstein formulated The Matching Law, demonstrating an elegant mathematical principle governing choice behavior in animals:

B1B1+B2=R1R1+R2\frac{B_1}{B_1 + B_2} = \frac{R_1}{R_1 + R_2}

Where:

  • $B_1$ and $B_2$ represent the relative rates of behavioral responding emitted across Alternative 1 and Alternative 2.
  • $R_1$ and $R_2$ represent the relative rates (or values) of reinforcement obtained from Alternative 1 and Alternative 2.

The Core Rule: An organism will allocate its behavioral choices across concurrent alternatives in direct proportion to the relative reinforcement obtained from each alternative.

[Alternative 1: Look at Handler] -------> Produces Dry Kibble (R1: Low Rate / Value)
                    VS.
[Alternative 2: Sniff Grass / Chase] ----> Produces High Dopamine Release (R2: Rich Environmental Value)

Result via Matching Law: Dog allocates 95% of behavior to Alternative 2!

Extended Parameters of the Matching Law

Later research expanded Herrnstein's equation beyond raw frequency to include three critical qualitative variables: Rate ($R$), Immediacy ($I$), and Magnitude/Quality ($M$) of reinforcement:

B1B1+B2=R1×I1×M1(R1×I1×M1)+(R2×I2×M2)\frac{B_1}{B_1 + B_2} = \frac{R_1 \times I_1 \times M_1}{(R_1 \times I_1 \times M_1) + (R_2 \times I_2 \times M_2)}

Resolving the Real-World Park Recall Dilemma

Consider the classic training breakdown: A handler calls their dog at an off-leash park. The dog looks at the handler, then turns and bolts toward a group of playing dogs.

  • Alternative 1 ($B_1$ - Recall to Handler):
    • Consequence ($R_1$): A dry biscuit delivered after a 5-second delay, followed by leash attachment and being locked in the car (leash clip functions as an aversive/punisher!).
  • Alternative 2 ($B_2$ - Play with Dogs):
    • Consequence ($R_2$): Immediate social chase, sensory stimulation, and predatory play motor release (massive magnitude, zero delay).

Under the Matching Law, the dog's choice is a mathematical certainty: the behavioral allocation to Alternative 2 will approach 99%. To tilt the Matching Law in the handler's favor, the certified trainer must systematically manipulate the variables:

  1. Increase Reinforcer Magnitude ($M_1$): Utilize real roast meat, warm tripe, or an energetic game of tug instead of dry biscuits.
  2. Eliminate Immediate Aversive Consequences: Clip the leash, feed jackpot treats, unclip the leash, and release the dog back to play (breaking the contingency between recall and park departure).
  3. Manage Competing Environmental Rates ($R_2$): Utilize a long line to prevent the dog from self-reinforcing on unauthorized environmental chases while the recall response history is being established.

Synthesis of Simple Intermittent Schedules

ScheduleDefining RuleCumulative Record ProfileResponse RatePost-Reinforcement Pause (PRP)Applied Canine Example
Fixed Ratio (FR)Reward delivered after $n$ unvarying responses"Break-and-run" stepped curveHigh rate during runYes; proportional to ratio sizeClick/treat after exactly 3 agility jumps
Variable Ratio (VR)Reward delivered after unpredictable responses averaging $n$Steep, straight lineVery high; maximal persistenceNo pause; steady continuous outputRewarding loose-leash walking every 2, 7, 3, 8 paces
Fixed Interval (FI)Reward for first response after $t$ unvarying elapsed timeScalloped curve; accelerating rateLow early; rapid near deadlineYes; prolonged pause post-rewardHandler returns to reward a stay at exactly 30s marks
Variable Interval (VI)Reward for first response after unpredictable elapsed time averaging $t$Flat, smooth, uniform lineModerate, steady rateNo pause; stable and calmIntermittent treat delivery for lying calmly on mat
Loading diagram...
Cumulative Record Response Profiles Across the Four Intermittent Schedules
Test Your Knowledge

A trainer records a dog's performance on an automated cumulative recorder during a duration stay task where reinforcement is delivered on a Fixed Interval 60-second schedule (FI60s). What visual pattern will the cumulative record curve exhibit?

A
B
C
D
Test Your Knowledge

Which simple intermittent schedule of reinforcement produces the highest overall rate of responding, exhibits virtually no post-reinforcement pause, and is commonly referred to as the 'gambler's schedule'?

A
B
C
D
Test Your Knowledge

An off-leash dog routinely ignores its handler's recall cues in an open dog park, choosing instead to chase running dogs. The handler attempts to reward the dog with standard dry kibble, but the dog continues to allocate 90% of its behavioral choices to chasing conspecifics. How does Herrnstein's Matching Law mathematically explain this training breakdown?

A
B
C
D