4.3 Measurement Quality: Validity, Accuracy, Reliability, and IOA Calculations
Key Takeaways
Trustworthy measurement requires establishing validity (measuring the direct socially significant dimension), accuracy (matching the true physical value), and reliability (yielding consistent results, though high reliability does not guarantee accuracy).
Interobserver Agreement (IOA) assesses the degree of consistency between two independent, simultaneous observers, serving to detect observer drift, evaluate operational definitions, and establish data believability.
By research convention, IOA is collected in at least 20% of sessions (preferably 25%-33%) spread across all conditions, with 80% agreement as the usual minimum and 90% preferred.
Total count and total duration IOA provide crude, global estimates that obscure interval-level discrepancies and systematically overestimate interobserver agreement.
Scored-interval IOA is required for low-rate behaviors ( intervals) to prevent agreement inflation from joint non-occurrences, whereas unscored-interval IOA is required for high-rate behaviors ( intervals) to eliminate inflation from joint occurrences.
The Triad of Trustworthy Measurement: Validity, Accuracy, and Reliability
In applied behavior analysis, clinical conclusions and treatment adaptations depend entirely on the integrity of graphed data. If data are corrupted by measurement error, the clinician cannot determine whether an intervention produced genuine behavior change. To be scientifically trustworthy and clinically actionable, measurement must possess three fundamental psychometric properties: validity, accuracy, and reliability (Johnston & Pennypacker, 1993).
[ THE TRIAD OF TRUSTWORTHY DATA ]
|
+----------------------------+----------------------------+
| | |
v v v
[ VALIDITY ] [ ACCURACY ] [ RELIABILITY ]
Directly measures the Observed value matches Repeated measures of the
relevant behavior and the TRUE physical value same event yield the same
dimension in context of the event numerical value
| | |
Threat: Indirect surveys, Threat: Observer drift, High reliability does NOT
wrong dimension, artifacts calibration error, bias guarantee accuracy!
1. Validity
Measurement is valid when it directly measures a socially significant target behavior and a dimension of that behavior that is relevant to the clinical question, under conditions that are directly applicable to the intervention. Measurement lacks validity if:
- It measures an indirect proxy rather than the actual behavior (e.g., using a parent retrospective questionnaire or self-report scale to measure child aggression instead of direct observation).
- It measures the wrong dimensional quantity (e.g., measuring the count of fire evacuation steps rather than the latency to evacuate, or recording the frequency of peer-interaction episodes instead of total duration of cooperative play).
- It introduces measurement artifacts (e.g., using a 60-second partial-interval system to measure 1-second tics, producing artifactual 100% scores).
2. Accuracy
Measurement is accurate when the observed numerical value matches the true physical state or value of the event as it actually transpired in physical reality. Determining accuracy requires an independent, calibrated measurement protocol—a true value standard—that evaluates the phenomenon through error-free procedures (e.g., automated electronic sensors or frame-by-frame slow-motion video scoring reviewed by master observers).
3. Reliability
Measurement is reliable when repeated measurement of the identical physical event yields the same numerical values across independent observation occasions or independent observers. Reliability reflects the consistency, stability, and repeatability of the measurement procedure.
The Critical Interrelationship: Reliability Does Not Guarantee Accuracy
A foundational principle frequently evaluated on the BCaBA examination is that high reliability does not guarantee accuracy:
- Consider an analog kitchen scale that has a misaligned spring: every time an identical 1.00-kilogram weight is placed on the scale, it reads exactly 1.45 kilograms. The scale produces 100% reliability (perfect consistency across repeated measurements), but 0% accuracy (it deviates systematically from the true physical value).
- In applied settings, two observers can share an identical misunderstanding of an operational definition, consistently agreeing with 100% concordance to score non-examples as positive occurrences. Their data are highly reliable, yet utterly inaccurate.
- However, the converse holds true: low reliability guarantees low accuracy. If independent observers cannot even agree on what occurred, the data cannot be accurate.
Interobserver Agreement (IOA): Core Principles and Standards
Interobserver Agreement (IOA) is the degree to which two independent, simultaneous observers record the identical numerical values or interval determinations after observing the same behavioral event during the same observation window.
Clinical and Methodological Functions of IOA
- Determining Observer Competence: IOA data verify that direct care technicians, therapists, and teachers have mastered operational definitions and can execute measurement protocols with fidelity.
- Detecting Observer Drift: Observer drift is the gradual, unconscious shift in how an observer interprets and applies an operational definition over weeks or months of practice. Drift occurs when observers inadvertently expand or restrict definitions based on personal familiarity with the client. Periodic IOA checks identify drift early and trigger retraining.
- Verifying Operational Definition Clarity: If two well-trained observers repeatedly fail to achieve acceptable IOA, the primary flaw typically resides in the definition itself: it is subjective, ambiguous, or lacks explicit boundary conditions.
- Establishing Experimental Believability: High IOA assures peer reviewers, funding sources, and interdisciplinary teams that observed changes in level and trend reflect real behavior change rather than observer bias, expectancy effects, or recording idiosyncrasies.
Research Conventions for IOA
- Session Frequency: Research conventions call for IOA in at least 20% of sessions, preferably 25%–33% (Cooper, Heron, & Heward, 2020).
- Phase Distribution: IOA must not be clustered exclusively in baseline; it must be distributed systematically across all experimental conditions and phases (baseline, initial intervention, maintenance, generalization), across all participants, and across different times of day.
- Acceptable Benchmark Criterion: By convention, 80% agreement is the usual minimum and 90% or higher is preferred. The BACB does not set a numeric IOA standard; these are research conventions.
Mathematical Formulas and Step-by-Step IOA Calculations
Assistant behavior analysts must be proficient in calculating agreement across count-based, duration-based, trial-based, and interval-based measurement systems.
1. Count-Based IOA Methods
A. Total Count IOA
Total Count IOA is the simplest, most global count metric. It divides the smaller total count recorded by one observer by the larger total count recorded by the second observer, multiplied by 100.
- Worked Example: During a 30-minute session, Observer 1 records 16 instances of aggression; Observer 2 records 20 instances. .
- Limitation: Total Count IOA is crude and routinely overestimates true agreement. It tells you only that both observers counted a similar total, but provides zero confirmation that they recorded the behavior at the same times or during the same episodes.
B. Mean Count-per-Interval IOA
The observation session is divided into intervals. An agreement ratio is calculated separately for each individual interval (smaller count divided by larger count), the interval ratios are summed, and the sum is divided by the total number of intervals ().
(Note: If both observers score 0 in an interval, agreement for that interval is 1.0 or 100%).
C. Exact Count-per-Interval IOA
Exact Count-per-Interval IOA is the most conservative and rigorous count metric. It calculates the percentage of total intervals in which both observers recorded the exact identical integer count.
- Comparison Example: Consider the following 4-interval data:
- Interval 1: Obs 1 = 2, Obs 2 = 2 (Agreement = 2/2 = 1.0; Exact = YES)
- Interval 2: Obs 1 = 3, Obs 2 = 1 (Agreement = 1/3 = 0.333; Exact = NO)
- Interval 3: Obs 1 = 0, Obs 2 = 0 (Agreement = 1.0; Exact = YES)
- Interval 4: Obs 1 = 4, Obs 2 = 5 (Agreement = 4/5 = 0.80; Exact = NO)
- Mean Count-per-Interval IOA: .
- Exact Count-per-Interval IOA: .
2. Discrete Trial IOA: Trial-by-Trial IOA
For restricted operants or discrete trials where each trial yields a binary outcome (e.g., correct vs. incorrect, or prompt vs. independent), Trial-by-Trial IOA compares agreement trial-by-trial.
3. Duration-Based IOA Methods
A. Total Duration IOA
Like Total Count IOA, Total Duration IOA is crude and overestimates agreement by ignoring whether the observers timed the same episodes.
B. Mean Duration-per-Occurrence IOA
Calculates the agreement ratio for each separate episode of behavior (shorter duration divided by longer duration), sums the episode ratios, and divides by the total number of episodes (). Highly sensitive for duration data.
4. Interval-Based IOA Methods
A. Interval-by-Interval (Point-by-Point) IOA
Compares agreement across every single interval in a time-sampling system. An agreement occurs when both observers score an occurrence () OR both observers score a non-occurrence ().
The Critical Problem of Chance Agreement Inflation
Interval-by-Interval IOA is vulnerable to severe mathematical distortion based on the baseline prevalence of the behavior:
- In Very Low-Rate Behaviors ( of intervals): Observers will naturally agree on the overwhelming majority of non-occurrence intervals () strictly by chance. If a behavior occurs in only 2 of 100 intervals, two observers who record zeros almost everywhere can easily score 95%–98% agreement even if they completely disagreed on the 2 actual occurrences!
- In Very High-Rate Behaviors ( of intervals): Observers will naturally agree on occurrence intervals () strictly by chance, masking massive disagreement on non-occurrences.
To eliminate this distortion, behavior analysts utilize Scored-Interval IOA and Unscored-Interval IOA.
[ INTERVAL IOA SELECTION RULE ]
|
+--------------------------+--------------------------+
| |
v v
[ LOW-RATE BEHAVIOR (< 30% intervals) ] [ HIGH-RATE BEHAVIOR (> 70% intervals) ]
Utilize: SCORED-INTERVAL IOA Utilize: UNSCORED-INTERVAL IOA
Discards joint non-occurrences (-/-) Discards joint occurrences (+/+)
Prevents inflation from chance non-events Prevents inflation from chance omnipresence
B. Scored-Interval IOA (Occurrence IOA)
Scored-Interval IOA evaluates agreement only in intervals where at least one observer scored an occurrence (). Intervals where both observers scored a non-occurrence () are completely discarded from the calculation.
- Recommended Use: Report it whenever the target behavior occurs at low rates ( of intervals). Discarding joint non-occurrences exposes true observer disagreement.
C. Unscored-Interval IOA (Non-Occurrence IOA)
Unscored-Interval IOA evaluates agreement only in intervals where at least one observer scored a non-occurrence (). Intervals where both observers scored an occurrence () are completely discarded from the calculation.
- Recommended Use: Report it whenever the target behavior occurs at high rates ( of intervals). Discarding joint occurrences ensures agreement is tested on the rare non-occurrences.
IOA Calculation Formula and Application Guide
The following table outlines all major IOA methods, their mathematical formulas, clinical indications, and key methodological limitations:
| IOA Method | Mathematical Formula | When Indicated | Advantages | Critical Limitations / Biases |
|---|---|---|---|---|
| Total Count IOA | Low-precision count; quick spot-checks | Extremely fast; requires no interval timing devices | Crude; systematically overestimates agreement; masks timing disagreements | |
| Mean Count-per-Interval | Continuous count data divided into intervals | More sensitive than Total Count; evaluates consistency interval-by-interval | Does not confirm agreement on exact instances within the interval | |
| Exact Count-per-Interval | High-precision count across intervals | Most rigorous, conservative count measure | Highly stringent; yields low agreement scores in active sessions | |
| Trial-by-Trial IOA | Discrete trial training (DTT); restricted operants | Precise discrete opportunity agreement; easy to calculate | Dependent on explicit trial pacing and clear trial demarcations | |
| Total Duration IOA | Total session duration; gross time checks | Simple to calculate with standard stopwatches | Overestimates agreement; ignores whether same episodes were timed | |
| Mean Duration-per-Occurrence | Episode-by-episode continuous duration | Evaluates timing precision across each discrete behavioral episode | Demands dual stopwatches and exact onset/offset synchronization | |
| Interval-by-Interval IOA | Interval systems with moderate rate (30% to 70%) | Comprehensive; evaluates all observation intervals | Severely inflated by chance agreement in very high-rate or low-rate behaviors | |
| Scored-Interval IOA | Low-rate behaviors occurring in of intervals | Eliminates artificial inflation from joint non-occurrences () | Inapplicable to continuous or high-rate behaviors | |
| Unscored-Interval IOA | High-rate behaviors occurring in of intervals | Eliminates artificial inflation from joint occurrences () | Inapplicable to low-rate behaviors |
Two independent observers record the frequency of vocal stereotypy across four consecutive 5-minute intervals. Observer 1 records: Interval 1 = 2, Interval 2 = 3, Interval 3 = 0, Interval 4 = 4. Observer 2 records: Interval 1 = 2, Interval 2 = 1, Interval 3 = 0, Interval 4 = 5. What is the calculated Mean Count-per-Interval IOA for this session?
88.9%
78.3%
94.2%
50.0%
An assistant behavior analyst measures severe property destruction that occurs very infrequently, occurring in only 4 out of 100 observation intervals (a low-rate behavior). Two observers collect interval data. Interval-by-Interval (Point-by-Point) IOA is calculated at 94%, primarily because both observers agreed on 92 intervals where the behavior did not occur. However, on the 6 intervals where at least one observer scored an occurrence, they only agreed on 3 intervals. Why is Interval-by-Interval IOA misleading here, and what calculation should the analyst report?
Unscored-Interval IOA should be reported because it isolates non-occurrences and proves high observer agreement.
Total Count IOA should be reported because interval systems are invalid for low-rate responses.
Scored-Interval IOA, because it discards joint non-occurrences that inflate agreement for low-rate behavior.
Interval-by-Interval IOA is accurate and should be retained because 94% exceeds the standard 80% threshold.
A supervisor reviews video recordings of an intervention session where calibrated digital time-stamped video confirms that a client engaged in exactly 20 instances of physical aggression. During the live session, two independent behavioral technicians both recorded exactly 11 instances of aggression, yielding an Interobserver Agreement (IOA) score of 100%. How should the behavior analyst interpret these measurement results?
The measurement data demonstrate low reliability, because the technicians failed to capture nine instances of aggression.
The measurement data demonstrate high reliability (consistency between observers) but low accuracy (divergence from the true value).
The measurement data demonstrate an invalid operational definition, because accuracy cannot be assessed without continuous duration recording.
The measurement data demonstrate high validity and high accuracy, because both observers achieved perfect agreement.
Sections you finish are checked off in the contents.