3.4 Interobserver Agreement (IOA) Methods and Calculations
Key Takeaways
- Interobserver Agreement (IOA) measures the degree to which two independent observers report identical values, enhancing data believability and detecting observer drift across 20% to 30% of sessions.
- IOA measures agreement rather than accuracy (comparison to true values) or reliability (stability over time), meaning observers can achieve high IOA while misapplying definitions identically.
- Count IOA algorithms range from coarse Total Count IOA to interval-based metrics, with Exact Count-per-Interval IOA serving as the most stringent count agreement standard.
- Duration IOA algorithms evaluate overall temporal agreement via Total Duration IOA or average per-instance agreement using Mean Duration-per-Occurrence IOA.
- Interval IOA requires Scored-Interval (Occurrence) IOA for low-rate behaviors (<30%) and Unscored-Interval (Nonoccurrence) IOA for high-rate behaviors (>70%) to avoid chance inflation.
Interobserver Agreement (IOA) Methods and Calculations
In Applied Behavior Analysis, data believability is established primarily through Interobserver Agreement (IOA). IOA measures the degree to which two independent observers, using the same measurement system to monitor the same client during the same observation window, report identical values. High IOA indicates that the operational definition is clear and objective, observer training is adequate, and data collection is consistent.
Critical Conceptual Distinctions
On the QBA examination, candidates must clearly distinguish IOA from accuracy and reliability:
- IOA vs. Accuracy: Accuracy evaluates whether observed data match the true value of the event (measured using an unrefutable gold-standard criterion, such as a calibrated video recording). IOA evaluates agreement between observers. Two observers can achieve 100% IOA while being 100% inaccurate if both misapply an operational definition identically.
- IOA vs. Reliability: Reliability refers to the stability of measurement over time (e.g., whether a single observer scores a permanent product identically on two separate occasions).
- Observer Drift: The gradual, unintended shift in how an observer applies an operational definition over time. Regular IOA assessments detect and correct observer drift.
Professional Standards for IOA
- Frequency: IOA data must be collected across a minimum of 20% to 30% of sessions in each experimental phase.
- Threshold: Acceptable IOA is generally 80% or higher (or a Cohen's Kappa coefficient of 0.60+).
- Independence: Observers must collect data independently (without cues, communication, or looking at each other's data sheets).
IOA Calculation Formulas and Worked Examples
1. Count / Event IOA Algorithms
A. Total Count IOA
Calculates agreement based on the total count recorded across an entire session.
- Worked Example: Observer 1 tallies 15 instances of hand-raising; Observer 2 tallies 20 instances.
- Limitation: Overestimates true agreement because it does not confirm observers agreed on specific instances in time.
B. Mean Count-per-Interval IOA
Calculates the average agreement across individual intervals within a session.
- Worked Example: Across 3 intervals: Interval 1 (Obs A = 2, Obs B = 2; Agreement = 1.0); Interval 2 (Obs A = 1, Obs B = 3; Agreement = 0.33); Interval 3 (Obs A = 4, Obs B = 5; Agreement = 0.80).
C. Exact Count-per-Interval IOA
The most conservative count metric. Only intervals with 100% exact agreement are scored as agreements.
- Worked Example (using the 3 intervals above): Only Interval 1 had 100% agreement (1 out of 3 intervals).
2. Duration IOA Algorithms
A. Total Duration IOA
- Worked Example: Obs A records 45 minutes of crying; Obs B records 50 minutes.
B. Mean Duration-per-Occurrence IOA
Calculates duration agreement for each discrete occurrence of behavior and averages them across occurrences.
3. Interval Sampling IOA Algorithms
Interval IOA algorithms are evaluated using a 2x2 contingency matrix of interval scores:
| Observer B (+) | Observer B (-) | |
|---|---|---|
| Observer A (+) | Both (+) [Agreed Occurrence] | Obs A (+) / Obs B (-) [Disagreement] |
| Observer A (-) | Obs A (-) / Obs B (+) [Disagreement] | Both (-) [Agreed Nonoccurrence] |
A. Interval-by-Interval (Total) IOA
Includes all intervals (both agreed occurrences and agreed nonoccurrences).
- Limitation: Artificially inflates agreement for very high-rate or very low-rate behaviors due to chance agreement on nonoccurrences or occurrences.
B. Scored-Interval (Occurrence) IOA
Calculates agreement based ONLY on intervals where AT LEAST ONE observer scored an occurrence (+). Ignores intervals where both observers scored nonoccurrence (-).
- Mandatory Clinical Rule: Scored-Interval IOA must be used for LOW-RATE behaviors (occurring in <30% of intervals) to prevent agreements on nonoccurrences from masking low agreement on actual occurrences.
C. Unscored-Interval (Nonoccurrence) IOA
Calculates agreement based ONLY on intervals where AT LEAST ONE observer scored a nonoccurrence (-). Ignores intervals where both observers scored occurrence (+).
- Mandatory Clinical Rule: Unscored-Interval IOA must be used for HIGH-RATE behaviors (occurring in >70% of intervals) to prevent agreements on occurrences from masking low agreement on nonoccurrences.
Two behavior technicians conduct a 10-interval partial-interval recording IOA check. Their data sheets reveal: Both scored (+) in 2 intervals; Obs A scored (+) alone in 1 interval; Obs B scored (+) alone in 1 interval; Both scored (-) in 6 intervals. What is the calculated Scored-Interval (Occurrence) IOA?
Why is Scored-Interval (Occurrence) IOA specifically required when evaluating the believability of data collected on low-rate target behaviors?
A behavior analyst evaluates observer agreement across four 1-minute intervals. Interval 1: Obs A = 3, Obs B = 3 (1.0). Interval 2: Obs A = 2, Obs B = 4 (0.50). Interval 3: Obs A = 0, Obs B = 1 (0.0). Interval 4: Obs A = 5, Obs B = 5 (1.0). What is the calculated Mean Count-per-Interval IOA?