2.4 Unconditioned, Conditioned, and Generalized Reinforcers and Punishers
Key Takeaways
Unconditioned (primary) reinforcers and punishers work without a learning history because of the species' evolutionary history; examples include food, water, warmth, and painful stimulation.
Conditioned (secondary) reinforcers and punishers acquire their function by being paired with established reinforcers or punishers, such as praise paired with snacks or a warning tone paired with loss of access.
Generalized conditioned reinforcers (money, tokens, attention) and generalized conditioned punishers (disapproval, fines) are paired with many backup consequences, so they work across many motivating operations.
To establish a conditioned reinforcer, present the neutral stimulus just before an effective reinforcer so it reliably predicts it, confirm the effect with a reinforcer test, and keep pairing it occasionally so it keeps its value.
Why Stimulus Classifications Matter
Two TCO tasks meet in this section. Task B.7 asks you to identify and distinguish unconditioned, conditioned, and generalized reinforcers and punishers. Task G.2 asks you to develop and implement procedures that establish and use conditioned reinforcers. On the exam, you will be asked to classify a consequence from a scenario and to choose the procedure that would turn a neutral stimulus (a praise phrase, a token, a clicker) into a working reinforcer.
Remember that every label below describes function. A stimulus is a reinforcer only if it increases the behavior it follows and a punisher only if it decreases that behavior. "Conditioned" or "unconditioned" describes how it acquired that function.
The Six Categories
| Category | How it gets its function | Examples | Key exam point |
|---|---|---|---|
| Unconditioned reinforcer | Phylogenic (species) history; no learning needed | Food when hungry, water when thirsty, warmth when cold, sexual stimulation | Effectiveness depends heavily on motivating operations (deprivation and satiation) |
| Conditioned reinforcer | Pairing with an established reinforcer during the person's lifetime | Praise, a sticker, the sound of a clicker, a high five | Loses value if it stops being paired with backup reinforcers |
| Generalized conditioned reinforcer | Pairing with many different reinforcers | Money, tokens, points, social attention and approval | Less dependent on any one motivating operation, so less prone to satiation |
| Unconditioned punisher | Phylogenic history; no learning needed | Painful stimulation, extreme heat or cold, very loud noise | Functions the first time it is contacted |
| Conditioned punisher | Pairing with an established punisher or with the loss of reinforcers | A warning buzzer paired with response cost; "no" paired with loss of access | Weakens if it stops predicting the punisher or loss |
| Generalized conditioned punisher | Pairing with many different punishers or losses | Disapproval, a frown, fines, a failing grade | Works across many situations and motivating operations |
Unconditioned Punishers
Unconditioned punishers decrease behavior without any learning history: touching a hot stove, biting one's tongue, or stepping on a sharp object reduce the behaviors that produced them. Because they also elicit respondent reactions (pain, startle, fear), they carry the side effects described in section 2.3. In behavior plans, unconditioned punishers are essentially never the procedure of choice; you will meet them on the exam mostly as classification items and as automatic consequences in the natural environment.
Conditioned Punishers
A neutral stimulus becomes a conditioned punisher when it is reliably paired with an established punisher or with the loss of reinforcement. In a token system, a short tone that always precedes the removal of a token (response cost) can come to suppress behavior by itself. A parent's "No" that has repeatedly preceded removal of a toy can come to reduce grabbing even when no toy is removed. The pairing must continue occasionally; if the tone or "No" is never followed by a loss, it stops working.
A generalized conditioned punisher has been paired with many different punishers or losses, which is why disapproval from a teacher or employer can reduce a wide range of behavior in many motivational states.
Establishing Conditioned Reinforcers (Task G.2)
Many learners, especially young children and individuals with autism, do not initially respond to praise, tokens, or social approval as reinforcers. Teaching depends on these conditioned reinforcers because they can be delivered instantly, in any setting, without interrupting the activity. The procedure is a version of stimulus-stimulus pairing:
- Select a neutral stimulus that is easy to deliver and will be used later (a specific praise phrase, a token, a clicker sound).
- Present it immediately before an established reinforcer. The neutral stimulus should come first and should predict the reinforcer: deliver "Great job!" and then, within about a second, the preferred snack. Presenting the reinforcer first and the praise afterward (backward pairing), or presenting the praise often without the snack, weakens conditioning.
- Use a strong establishing operation for the backup reinforcer so the pairing involves a potent consequence.
- Consider response-contingent pairing. Delivering the neutral stimulus contingent on a simple response, followed by the backup reinforcer, can be more effective than pairing without any response requirement. Dozier and colleagues (2012) found that response-contingent pairing established praise as a reinforcer for several participants when stimulus-stimulus pairing alone did not.
- Test the result. Deliver the conditioned stimulus alone, contingent on a new, simple response, and compare with a baseline. If the response increases, the stimulus now functions as a reinforcer. Without a test, you only know the stimulus was paired, not that it works.
- Maintain it. Continue pairing the conditioned reinforcer with backup reinforcers intermittently. A conditioned reinforcer that is never again followed by a backup reinforcer undergoes extinction of its reinforcing function.
Applications
- Conditioning praise: pair a consistent praise phrase with preferred edibles or activities, then gradually thin the edible while keeping the praise.
- Token conditioning: begin with immediate exchange (one token traded for a backup reinforcer right away), then increase the number of tokens required and delay the exchange (section 10.4).
- Conditioning new leisure items: pairing a toy or book with existing reinforcers (and with adult attention) can broaden a learner's reinforcer pool, which helps when interests are very narrow.
Common Exam Traps
- Confusing "conditioned" with "less powerful." Conditioned reinforcers can be extremely effective; money is conditioned.
- Assuming praise is reinforcing. Praise is a conditioned reinforcer only for people with the right history; test it.
- Forgetting maintenance. Tokens that are never exchanged lose value.
- Mixing up generalized and simple conditioned stimuli. A buzzer paired only with loss of recess is conditioned; a frown paired with many different losses is generalized.
- Labeling by intent. A fine that does not decrease the behavior is not a punisher for that person, however unpleasant it was meant to be.
Across many trials, a therapist says "Nice!" about half a second before handing a child a preferred snack. Later, saying "Nice!" alone right after the child stacks blocks increases block stacking. What best explains the change?
"Nice!" became a conditioned reinforcer through pairing with the snack.
"Nice!" became a discriminative stimulus signaling that snacks are unavailable.
Block stacking increased through respondent conditioning of salivation to the praise word.
"Nice!" became an unconditioned reinforcer because it was repeated so often.
Which consequence is the best example of a generalized conditioned punisher?
Touching a hot pan, which reliably reduces future pan touching
A specific buzzer that has been paired only with the loss of recess minutes
A supervisor's frown that has preceded many different reprimands and losses
A sudden loud noise that reduces a behavior the very first time it follows it
Tokens in a classroom were originally exchanged for snacks and activities, but for three months no exchange periods have been held. Token-earning work has steadily declined. What is the most likely explanation?
The tokens now act as discriminative stimuli that signal more work is coming.
The snacks became unconditioned punishers after being withheld for three months.
The tokens lost conditioned value because pairing with backup reinforcers stopped.
The tokens became generalized reinforcers, which are immune to any loss of value.
Sections you finish are checked off in the contents.