9.4 Preference and Reinforcer Assessments & Social Validity
Key Takeaways
A stimulus preference assessment identifies candidate stimuli and generates a relative preference hierarchy; an empirical reinforcer assessment experimentally verifies whether contingent presentation of a preferred stimulus actually increases future response rates.
Direct stimulus preference assessment methods vary along presentation formats: Single-Stimulus (successive choice), Paired-Stimulus (forced-choice), Multiple Stimulus Without Replacement (MSWO), Multiple Stimulus With Replacement (MSW), and Free Operant.
Paired-stimulus assessments (Fisher et al., 1992) generate definitive, highly differentiated preference hierarchies by pairing every item with every other item, but require significant time ( trials) and can evoke problem behavior during item removal.
MSWO assessments (DeLeon & Iwata, 1996) provide a highly efficient, accurate rank-ordering of preferred items in 10 to 15 minutes by removing selected items on successive trials, making it the clinical standard for learners with scanning skills.
Social validity (Wolf, 1978) evaluates three distinct dimensions of applied interventions: the social significance of target behavioral goals, the social appropriateness and acceptability of procedures, and the social importance of behavioral outcomes.
Conceptual Distinction: Preference Assessment vs. Reinforcer Assessment
In applied behavior analysis, practitioners must never conflate a stimulus preference assessment with an empirical reinforcer assessment. While intimately related, they represent distinct scientific operations that answer fundamentally different behavioral questions:
- Stimulus Preference Assessment: A structured procedure designed to identify a collection of stimuli that an individual preferentially approaches, selects, or engages with, and to establish a relative preference hierarchy (high-preference [HP], medium-preference [MP], and low-preference [LP] stimuli). A preference assessment identifies candidate reinforcers. However, high preference represents a prediction, not an empirical proof of reinforcing efficacy.
- Reinforcer Assessment: An experimental evaluation that directly tests whether the contingent delivery of a preferred stimulus functions as a reinforcer by measuring its actual effect on the future rate, frequency, or maintenance of an operant response relative to baseline or control conditions. In short, preference predicts reinforcement, but only a reinforcer assessment demonstrates it.
A stimulus may be highly preferred during an assessment (e.g., a child loves holding a light-up spinner when zero demands exist), yet fail completely to reinforce academic work when task effort increases.
Indirect Preference Assessments
Indirect preference assessments gather information about potential reinforcers through informant interviews, surveys, and checklists administered to parents, teachers, and caregivers:
- Informant Surveys and Checklists: Questionnaires asking caregivers to rank items across sensory, edible, tangible, and social activity domains.
- Reinforcer Assessment for Individuals with Severe Disabilities (RAISD): Developed by Fisher, Piazza, Bowman, and Amari (1996), the RAISD is a structured interview protocol administered to caregivers that generates an individualized stimulus pool across sensory categories, scoring items by caregiver-ranked popularity.
- Utility and Limitations: Indirect methods are valuable for generating an initial pool of 10 to 15 candidate stimuli. However, caregiver rankings frequently exhibit low concordance with direct empirical choices; therefore, indirect nominations must always be confirmed through direct empirical testing.
Direct Empirical Stimulus Preference Assessment Methodologies
Behavior analysts utilize five primary direct empirical methodologies to evaluate preference hierarchies:
1. Single-Stimulus Preference Assessment (Successive Choice)
Developed by Pace, Ivancic, Edwards, Iwata, and Page (1985), the single-stimulus assessment involves presenting items to the individual one at a time in randomized order:
- Procedure: An item is placed in front of the learner. The observer records whether the learner approaches, touches, or consumes the item, and measures the duration of engagement across a fixed interval (e.g., 2 minutes).
- Scoring: Calculated as the percentage of trials the item was approached () or total duration of interaction.
- Strengths: Ideal for individuals who lack visual scanning repertoires, individuals who struggle to make choices between two or more items, or those who exhibit severe side biases.
- Limitations: Prone to false positives. Individuals often approach or touch nearly every item presented in isolation, resulting in uniformly high scores across all items and failing to differentiate a clear hierarchy between highly potent and mediocre stimuli.
2. Paired-Stimulus Preference Assessment (Forced-Choice)
Developed by Fisher et al. (1992), the paired-stimulus assessment presents two stimuli simultaneously on each trial:
- Procedure: The practitioner places two items equidistant from the learner. The learner is instructed to "Pick one." Upon reaching toward an item, access is granted for brief engagement (e.g., 15–30 seconds), while the unchosen item is immediately blocked or removed. If the learner reaches for both, both are blocked and the trial is re-presented.
- Trial Counterbalancing: Every stimulus in the assessment pool () is paired with every other stimulus exactly once (or twice with counterbalanced left/right positions to control for side bias). The total number of trials required for a single pairwise pairing is:
- Scoring: Items are rank-ordered based on the percentage of trials chosen:
- Strengths: Highly reliable; empirically proven to generate distinct, differentiated preference hierarchies that reliably predict reinforcer efficacy.
- Limitations: Extremely time-consuming (evaluating 8 items requires 28 trials; counterbalanced requires 56 trials). Item removal between trials can evoke severe problem behavior (tangible extinction bursts).
3. Multiple Stimulus Without Replacement (MSWO)
Developed by DeLeon and Iwata (1996), the MSWO assessment evaluates an array of stimuli simultaneously:
- Procedure: An array of 5 to 8 stimuli is placed in a horizontal line equidistant from the learner. The learner is instructed to select one item. Upon selection, the learner engages with the item for 15 to 30 seconds. Crucially, the chosen item is REMOVED from the array on subsequent trials.
- Trial Progression: The remaining items are rearranged in a randomized order and re-presented. The learner selects from the remaining items. This process continues until all items have been chosen or the learner makes no selection within a set latency (e.g., 30 seconds). Typically, 3 to 5 full sessions are conducted.
- Scoring: A percentage score is calculated across sessions based on average selection order, generating a clear rank-ordered hierarchy (First Chosen = HP; Middle = MP; Last/Unchosen = LP).
- Strengths: Highly efficient (completed in 10 to 15 minutes); produces a reliable, differentiated hierarchy comparable to paired-stimulus assessments.
- Limitations: Requires the learner to possess scanning skills across a large visual array; cannot be used with learners who exhibit severe aggression upon item removal.
4. Multiple Stimulus With Replacement (MSW)
In the MSW assessment, an array of items is presented simultaneously, but the chosen item is RETURNED to the array on the next trial:
- Procedure: The selected item is allowed brief engagement, returned to the array, the array is scrambled, and re-presented.
- Limitations: Vulnerable to the exclusive selection problem. If an individual has one dominant favorite item, they may select that same item on of trials. Consequently, the clinician gains zero information regarding the relative rankings of any remaining items.
5. Free Operant Preference Assessment
Developed by Roane, Vollmer, Ringdahl, and Marcus (1998), the free operant assessment provides continuous, unrestricted access to an array of items:
- Procedure: Multiple stimuli are placed in an accessible area. The learner is allowed to move freely and interact with any item for a set observation duration (e.g., 5 to 10 minutes). Absolutely zero items are removed, and zero demands are presented.
- Scoring: Cumulative duration of engagement with each stimulus is recorded via duration or momentary time sampling.
- Variants: Includes Contrived Free Operant (items placed intentionally in therapy room) and Naturalistic Free Operant (observing in natural living environment).
- Strengths: Rapid and non-aversive; zero risk of evoking problem behavior because items are never removed; excellent for brief pre-session preference probes.
- Limitations: May identify only the single highest-preference item if the client engages exclusively with one toy for the entire duration.
Empirical Reinforcer Assessment Methodologies
Once candidate stimuli are ranked via preference assessment, their actual reinforcing potency must be experimentally verified using one of three primary schedules:
1. Concurrent Schedules of Reinforcement
- Mechanism: Pits two or more distinct reinforcement contingencies against each other simultaneously for identical or different behaviors.
- Clinical Application: Evaluates relative reinforcing efficacy. For example, emitting Response A yields Stimulus 1 on an FR1 schedule, while emitting Response B yields Stimulus 2 on an FR1 schedule. The proportion of responding allocated to each schedule demonstrates the superior reinforcer according to the Matching Law (Herrnstein, 1961).
2. Multiple Schedules of Reinforcement
- Mechanism: Two or more basic schedules of reinforcement (e.g., FR1 and Extinction) alternate across time, each correlated with a distinct discriminative stimulus ().
- Clinical Application: Evaluates whether a stimulus increases response rates above an extinction baseline. An elevated response rate during the stimulus-correlated component relative to the extinction component confirms reinforcer efficacy.
3. Progressive-Ratio (PR) Schedules of Reinforcement
- Mechanism: The response requirement systematically increases following each reinforcer delivery according to an arithmetic or geometric progression (e.g., FR1, FR2, FR4, FR8, FR16, FR32...).
- The Breakpoint: Responding continues until the response requirement becomes so demanding that responding ceases entirely for a specified duration (e.g., 5 minutes). The final completed schedule requirement is the breakpoint.
- Clinical Significance: PR schedules quantify reinforcer potency and resilience under high response effort. A stimulus with a high breakpoint will maintain academic or vocational performance under demanding real-world workloads, whereas a stimulus with a low breakpoint will fail as soon as task difficulty increases.
Social Validity (Montrose Wolf, 1978)
Even an empirically flawless behavior-change program is a failure if it lacks social validity. In his foundational paper, Montrose Wolf (1978) established that applied behavior analysis must evaluate social validity across three distinct levels:
- The Social Significance of Target Behavioral Goals: Are the specific behavioral targets really what society and the client value? Assessment ensures that goals foster true habilitation, autonomy, and independence rather than mere administrative convenience or artificial compliance.
- The Social Appropriateness and Acceptability of Procedures: Do the consumers, implementers, and community consider the intervention procedures humane, acceptable, ethical, and cost-effective? Clinicians evaluate whether caregivers and direct staff are comfortable executing the procedures without experiencing undue stress or ethical conflict.
- The Social Importance of Behavioral Effects and Outcomes: Did the behavior change produce a real, noticeable difference in the client's everyday quality of life? Interventions must produce practical, meaningful improvements (e.g., making friends, returning to public school, accessing the community) rather than trivial, statistically isolated rate reductions.
Empirical Methods for Assessing Social Validity
- Social Comparison / Normative Comparisons: Comparing the client's post-intervention performance to the normative behavioral rates of typically developing peers in the same natural context (e.g., comparing a student's on-task percentage to the general education classroom peer average).
- Subjective Evaluations: Administering structured questionnaires and Likert rating scales (such as the Intervention Rating Profile-15 [IRP-15] or Treatment Acceptability Rating Form-Revised [TARF-R]) to parents, teachers, and clients.
- Blind Ratings by Naive Judges: Presenting before-and-after video samples in randomized order to judges who are unaware of which video represents baseline vs. treatment, asking them to rate qualitative competence.
Stimulus Preference Assessment Methods Matrix
The following table compares the administration format, scoring mechanics, time requirements, clinical strengths, and limitations of primary stimulus preference assessments:
| Method | Presentation Format | Scoring / Calculation | Time Requirement | Strengths | Limitations |
|---|---|---|---|---|---|
| Single-Stimulus (Pace et al., 1985) | Successive choice; one stimulus presented at a time in randomized order. | Percentage of trials approached () or total engagement duration. | Moderate (approx. 20–30 min). | Accommodates individuals lacking scanning or choice-making skills; eliminates side biases. | High false-positive rate; individuals approach nearly everything; fails to generate differentiated hierarchy. |
| Paired-Stimulus (Fisher et al., 1992) | Forced-choice; pairs of stimuli presented simultaneously; every item paired with every other item. | Percentage of trials chosen (). | High (30–60 min; trials). | Highly reliable; generates clear, differentiated hierarchy; strongly predictive of reinforcer potency. | Time-intensive; side-bias confounds; item removal can evoke severe problem behavior. |
| Multiple Stimulus Without Replacement (MSWO) (DeLeon & Iwata, 1996) | Array of 5–8 items presented simultaneously; chosen item is permanently removed on subsequent trials. | Rank-order percentage across repeated sessions based on selection sequence. | Low (10–15 min). | Highly efficient; generates accurate rank-order hierarchy; standard clinical choice for scanning learners. | Requires scanning across arrays; not suitable if individual resists item removal. |
| Multiple Stimulus With Replacement (MSW) | Array of items presented simultaneously; chosen item is returned to array on next trial. | Percentage of trials each item was selected. | Moderate (15–25 min). | Keeps all items available across all trials; less aversive than MSWO. | Exclusive selection problem: client selects top item on 100% of trials, yielding zero data on remaining items. |
| Free Operant (Roane et al., 1998) | Continuous unrestricted access to all array items simultaneously for 5–10 min. | Cumulative duration of engagement per item or momentary time sampling. | Very Low (5–10 min). | Zero item removal; zero behavioral escalation; ideal for brief pre-session preference probes. | May identify only single top item if client perseverates on one toy; fails to rank secondary items. |
Clinical Troubleshooting and Common Assessment Traps
- Trap 1: Assuming High Preference Equates to Strong Reinforcement: A practitioner identifies chocolate chips as the #1 high-preference item on an MSWO. However, when required to complete a 20-step vocational chain, the client refuses to work. Preference indicates relative liking in a demand-free environment; reinforcer efficacy must be empirically verified under actual task effort.
- Trap 2: Ignoring Motivating Operations Prior to Assessment: Conducting a food preference assessment immediately following lunch will yield suppressed engagement due to satiation (AO). Conducting assessments under appropriate deprivation states (EO) is essential for valid hierarchies.
- Trap 3: Neglecting Social Validity: A behavior analyst implements a technically effective overcorrection procedure that eliminates stereotypic hand-flapping. However, the classroom teacher finds the physical guiding procedure distressing and refuses to implement it. Because social acceptability was ignored, the intervention will suffer from treatment integrity failure.
A BCaBA is conducting an assessment to identify potential reinforcers for a 9-year-old student who has adequate visual scanning repertoires and readily relinquishes toys without emotional escalation. The school consultation team has only 15 minutes remaining before the student's scheduled bus departure to identify a preference hierarchy among 6 candidate educational games. Which stimulus preference assessment methodology is most appropriate for this clinical scenario?
Single-stimulus preference assessment (Pace et al., 1985)
Progressive-ratio schedule reinforcer assessment
MSWO preference assessment (DeLeon & Iwata, 1996)
Paired-stimulus preference assessment (Fisher et al., 1992)
An assistant behavior analyst conducts a reinforcer assessment using a progressive-ratio (PR) schedule to evaluate two potential edible reinforcers (fruit gummies vs. pretzel sticks) for an adult client in a vocational assembly training program. The schedule requires 1 response for the first reinforcer, 2 responses for the second, 4 for the third, 8 for the fourth, 16 for the fifth, and so forth. With pretzel sticks, the client ceases responding when the requirement reaches 16 responses (breakpoint = 8). With fruit gummies, the client continues responding until the requirement reaches 128 responses (breakpoint = 64). How should the BCaBA interpret this empirical finding?
Pretzel sticks and fruit gummies possess equivalent reinforcing value because both functioned as reinforcers on the initial FR1 schedule.
The progressive-ratio schedule was administered improperly because response requirements on operant schedules can never increase geometrically.
Fruit gummies are the more potent reinforcer under high response effort, so they better suit demanding vocational tasks.
Pretzel sticks are superior reinforcers because they reached the breakpoint faster, minimizing response fatigue.
A clinical team designs an intervention package that successfully reduces a client's stereotypic body rocking from 40 instances per hour to 5 instances per hour. However, the client's residential care staff report on a survey that the physical prompting and blocking procedures are excessively exhausting to implement, disrupt the group home's evening routine, and make them feel uncomfortable. According to Montrose Wolf's (1978) conceptual framework, which dimension of social validity has been violated?
The social significance of target behavioral goals
The empirical validity of internal single-case experimental control
The normative peer baseline comparison metric
The acceptability of the intervention procedures
Sections you finish are checked off in the contents.