4.3 Measurement Quality: Validity, Accuracy, Reliability, and IOA Calculations

Key Takeaways

  • Trustworthy measurement requires establishing validity (measuring the direct socially significant dimension), accuracy (matching the true physical value), and reliability (yielding consistent results, though high reliability does not guarantee accuracy).

  • Interobserver Agreement (IOA) assesses the degree of consistency between two independent, simultaneous observers, serving to detect observer drift, evaluate operational definitions, and establish data believability.

  • By research convention, IOA is collected in at least 20% of sessions (preferably 25%-33%) spread across all conditions, with 80% agreement as the usual minimum and 90% preferred.

  • Total count and total duration IOA provide crude, global estimates that obscure interval-level discrepancies and systematically overestimate interobserver agreement.

  • Scored-interval IOA is required for low-rate behaviors (<30%<30\% intervals) to prevent agreement inflation from joint non-occurrences, whereas unscored-interval IOA is required for high-rate behaviors (>70%>70\% intervals) to eliminate inflation from joint occurrences.

Last updated: October 2026

The Triad of Trustworthy Measurement: Validity, Accuracy, and Reliability

In applied behavior analysis, clinical conclusions and treatment adaptations depend entirely on the integrity of graphed data. If data are corrupted by measurement error, the clinician cannot determine whether an intervention produced genuine behavior change. To be scientifically trustworthy and clinically actionable, measurement must possess three fundamental psychometric properties: validity, accuracy, and reliability (Johnston & Pennypacker, 1993).

                    [ THE TRIAD OF TRUSTWORTHY DATA ]
                                    |
       +----------------------------+----------------------------+
       |                            |                            |
       v                            v                            v
  [ VALIDITY ]                 [ ACCURACY ]                [ RELIABILITY ]
 Directly measures the         Observed value matches      Repeated measures of the
 relevant behavior and        the TRUE physical value     same event yield the same
 dimension in context         of the event                numerical value
       |                            |                            |
 Threat: Indirect surveys,    Threat: Observer drift,     High reliability does NOT
 wrong dimension, artifacts   calibration error, bias     guarantee accuracy!

1. Validity

Measurement is valid when it directly measures a socially significant target behavior and a dimension of that behavior that is relevant to the clinical question, under conditions that are directly applicable to the intervention. Measurement lacks validity if:

  • It measures an indirect proxy rather than the actual behavior (e.g., using a parent retrospective questionnaire or self-report scale to measure child aggression instead of direct observation).
  • It measures the wrong dimensional quantity (e.g., measuring the count of fire evacuation steps rather than the latency to evacuate, or recording the frequency of peer-interaction episodes instead of total duration of cooperative play).
  • It introduces measurement artifacts (e.g., using a 60-second partial-interval system to measure 1-second tics, producing artifactual 100% scores).

2. Accuracy

Measurement is accurate when the observed numerical value matches the true physical state or value of the event as it actually transpired in physical reality. Determining accuracy requires an independent, calibrated measurement protocol—a true value standard—that evaluates the phenomenon through error-free procedures (e.g., automated electronic sensors or frame-by-frame slow-motion video scoring reviewed by master observers).

3. Reliability

Measurement is reliable when repeated measurement of the identical physical event yields the same numerical values across independent observation occasions or independent observers. Reliability reflects the consistency, stability, and repeatability of the measurement procedure.

The Critical Interrelationship: Reliability Does Not Guarantee Accuracy

A foundational principle frequently evaluated on the BCaBA examination is that high reliability does not guarantee accuracy:

  • Consider an analog kitchen scale that has a misaligned spring: every time an identical 1.00-kilogram weight is placed on the scale, it reads exactly 1.45 kilograms. The scale produces 100% reliability (perfect consistency across repeated measurements), but 0% accuracy (it deviates systematically from the true physical value).
  • In applied settings, two observers can share an identical misunderstanding of an operational definition, consistently agreeing with 100% concordance to score non-examples as positive occurrences. Their data are highly reliable, yet utterly inaccurate.
  • However, the converse holds true: low reliability guarantees low accuracy. If independent observers cannot even agree on what occurred, the data cannot be accurate.

Interobserver Agreement (IOA): Core Principles and Standards

Interobserver Agreement (IOA) is the degree to which two independent, simultaneous observers record the identical numerical values or interval determinations after observing the same behavioral event during the same observation window.

Clinical and Methodological Functions of IOA

  1. Determining Observer Competence: IOA data verify that direct care technicians, therapists, and teachers have mastered operational definitions and can execute measurement protocols with fidelity.
  2. Detecting Observer Drift: Observer drift is the gradual, unconscious shift in how an observer interprets and applies an operational definition over weeks or months of practice. Drift occurs when observers inadvertently expand or restrict definitions based on personal familiarity with the client. Periodic IOA checks identify drift early and trigger retraining.
  3. Verifying Operational Definition Clarity: If two well-trained observers repeatedly fail to achieve acceptable IOA, the primary flaw typically resides in the definition itself: it is subjective, ambiguous, or lacks explicit boundary conditions.
  4. Establishing Experimental Believability: High IOA assures peer reviewers, funding sources, and interdisciplinary teams that observed changes in level and trend reflect real behavior change rather than observer bias, expectancy effects, or recording idiosyncrasies.

Research Conventions for IOA

  • Session Frequency: Research conventions call for IOA in at least 20% of sessions, preferably 25%–33% (Cooper, Heron, & Heward, 2020).
  • Phase Distribution: IOA must not be clustered exclusively in baseline; it must be distributed systematically across all experimental conditions and phases (baseline, initial intervention, maintenance, generalization), across all participants, and across different times of day.
  • Acceptable Benchmark Criterion: By convention, 80% agreement is the usual minimum and 90% or higher is preferred. The BACB does not set a numeric IOA standard; these are research conventions.

Mathematical Formulas and Step-by-Step IOA Calculations

Assistant behavior analysts must be proficient in calculating agreement across count-based, duration-based, trial-based, and interval-based measurement systems.

1. Count-Based IOA Methods

A. Total Count IOA

Total Count IOA is the simplest, most global count metric. It divides the smaller total count recorded by one observer by the larger total count recorded by the second observer, multiplied by 100.

Total Count IOA=Smaller CountLarger Count×100%\text{Total Count IOA} = \frac{\text{Smaller Count}}{\text{Larger Count}} \times 100\%

  • Worked Example: During a 30-minute session, Observer 1 records 16 instances of aggression; Observer 2 records 20 instances. Total Count IOA=1620×100%=80.0%\text{Total Count IOA} = \frac{16}{20} \times 100\% = 80.0\%.
  • Limitation: Total Count IOA is crude and routinely overestimates true agreement. It tells you only that both observers counted a similar total, but provides zero confirmation that they recorded the behavior at the same times or during the same episodes.

B. Mean Count-per-Interval IOA

The observation session is divided into intervals. An agreement ratio is calculated separately for each individual interval (smaller count divided by larger count), the interval ratios are summed, and the sum is divided by the total number of intervals (nn).

Mean Count-per-Interval IOA=∑i=1n(Smaller CountiLarger Counti)n×100%\text{Mean Count-per-Interval IOA} = \frac{\sum_{i=1}^{n} \left( \frac{\text{Smaller Count}_i}{\text{Larger Count}_i} \right)}{n} \times 100\%

(Note: If both observers score 0 in an interval, agreement for that interval is 1.0 or 100%).

C. Exact Count-per-Interval IOA

Exact Count-per-Interval IOA is the most conservative and rigorous count metric. It calculates the percentage of total intervals in which both observers recorded the exact identical integer count.

Exact Count-per-Interval IOA=Number of Intervals with 100% Identical CountTotal Number of Intervals×100%\text{Exact Count-per-Interval IOA} = \frac{\text{Number of Intervals with 100\% Identical Count}}{\text{Total Number of Intervals}} \times 100\%

  • Comparison Example: Consider the following 4-interval data:
    • Interval 1: Obs 1 = 2, Obs 2 = 2 (Agreement = 2/2 = 1.0; Exact = YES)
    • Interval 2: Obs 1 = 3, Obs 2 = 1 (Agreement = 1/3 = 0.333; Exact = NO)
    • Interval 3: Obs 1 = 0, Obs 2 = 0 (Agreement = 1.0; Exact = YES)
    • Interval 4: Obs 1 = 4, Obs 2 = 5 (Agreement = 4/5 = 0.80; Exact = NO)
    • Mean Count-per-Interval IOA: 1.0+0.333+1.0+0.804=3.1334×100%=78.3%\frac{1.0 + 0.333 + 1.0 + 0.80}{4} = \frac{3.133}{4} \times 100\% = 78.3\%.
    • Exact Count-per-Interval IOA: 24×100%=50.0%\frac{2}{4} \times 100\% = 50.0\%.

2. Discrete Trial IOA: Trial-by-Trial IOA

For restricted operants or discrete trials where each trial yields a binary outcome (e.g., correct vs. incorrect, or prompt vs. independent), Trial-by-Trial IOA compares agreement trial-by-trial.

Trial-by-Trial IOA=Number of Trials with AgreementTotal Number of Trials Presented×100%\text{Trial-by-Trial IOA} = \frac{\text{Number of Trials with Agreement}}{\text{Total Number of Trials Presented}} \times 100\%

3. Duration-Based IOA Methods

A. Total Duration IOA

Total Duration IOA=Shorter DurationLonger Duration×100%\text{Total Duration IOA} = \frac{\text{Shorter Duration}}{\text{Longer Duration}} \times 100\% Like Total Count IOA, Total Duration IOA is crude and overestimates agreement by ignoring whether the observers timed the same episodes.

B. Mean Duration-per-Occurrence IOA

Calculates the agreement ratio for each separate episode of behavior (shorter duration divided by longer duration), sums the episode ratios, and divides by the total number of episodes (nn). Highly sensitive for duration data.

Mean Duration-per-Occurrence IOA=∑i=1n(Shorter DurationiLonger Durationi)n×100%\text{Mean Duration-per-Occurrence IOA} = \frac{\sum_{i=1}^{n} \left( \frac{\text{Shorter Duration}_i}{\text{Longer Duration}_i} \right)}{n} \times 100\%

4. Interval-Based IOA Methods

A. Interval-by-Interval (Point-by-Point) IOA

Compares agreement across every single interval in a time-sampling system. An agreement occurs when both observers score an occurrence (+/++/+) OR both observers score a non-occurrence (−/−-/-).

Interval-by-Interval IOA=Number of Agreed Intervals (both + or both −)Total Number of Intervals×100%\text{Interval-by-Interval IOA} = \frac{\text{Number of Agreed Intervals (both }+\text{ or both }-)}{\text{Total Number of Intervals}} \times 100\%

The Critical Problem of Chance Agreement Inflation

Interval-by-Interval IOA is vulnerable to severe mathematical distortion based on the baseline prevalence of the behavior:

  • In Very Low-Rate Behaviors (<30%<30\% of intervals): Observers will naturally agree on the overwhelming majority of non-occurrence intervals (−/−-/-) strictly by chance. If a behavior occurs in only 2 of 100 intervals, two observers who record zeros almost everywhere can easily score 95%–98% agreement even if they completely disagreed on the 2 actual occurrences!
  • In Very High-Rate Behaviors (>70%>70\% of intervals): Observers will naturally agree on occurrence intervals (+/++/+) strictly by chance, masking massive disagreement on non-occurrences.

To eliminate this distortion, behavior analysts utilize Scored-Interval IOA and Unscored-Interval IOA.

                    [ INTERVAL IOA SELECTION RULE ]
                                   |
        +--------------------------+--------------------------+
        |                                                     |
        v                                                     v
 [ LOW-RATE BEHAVIOR (< 30% intervals) ]        [ HIGH-RATE BEHAVIOR (> 70% intervals) ]
 Utilize: SCORED-INTERVAL IOA                   Utilize: UNSCORED-INTERVAL IOA
 Discards joint non-occurrences (-/-)           Discards joint occurrences (+/+)
 Prevents inflation from chance non-events      Prevents inflation from chance omnipresence

B. Scored-Interval IOA (Occurrence IOA)

Scored-Interval IOA evaluates agreement only in intervals where at least one observer scored an occurrence (++). Intervals where both observers scored a non-occurrence (−/−-/-) are completely discarded from the calculation.

Scored-Interval IOA=Agreed Occurrence Intervals (both +)Total Intervals with at Least One + Scored×100%\text{Scored-Interval IOA} = \frac{\text{Agreed Occurrence Intervals (both }+)}{\text{Total Intervals with at Least One }+ \text{ Scored}} \times 100\%

  • Recommended Use: Report it whenever the target behavior occurs at low rates (<30%<30\% of intervals). Discarding joint non-occurrences exposes true observer disagreement.

C. Unscored-Interval IOA (Non-Occurrence IOA)

Unscored-Interval IOA evaluates agreement only in intervals where at least one observer scored a non-occurrence (−-). Intervals where both observers scored an occurrence (+/++/+) are completely discarded from the calculation.

Unscored-Interval IOA=Agreed Non-Occurrence Intervals (both −)Total Intervals with at Least One − Scored×100%\text{Unscored-Interval IOA} = \frac{\text{Agreed Non-Occurrence Intervals (both }-)}{\text{Total Intervals with at Least One }- \text{ Scored}} \times 100\%

  • Recommended Use: Report it whenever the target behavior occurs at high rates (>70%>70\% of intervals). Discarding joint occurrences ensures agreement is tested on the rare non-occurrences.

IOA Calculation Formula and Application Guide

The following table outlines all major IOA methods, their mathematical formulas, clinical indications, and key methodological limitations:

IOA MethodMathematical FormulaWhen IndicatedAdvantagesCritical Limitations / Biases
Total Count IOASmaller CountLarger Count×100%\frac{\text{Smaller Count}}{\text{Larger Count}} \times 100\%Low-precision count; quick spot-checksExtremely fast; requires no interval timing devicesCrude; systematically overestimates agreement; masks timing disagreements
Mean Count-per-Interval∑(Smalleri/Largeri)n×100%\frac{\sum (\text{Smaller}_i / \text{Larger}_i)}{n} \times 100\%Continuous count data divided into intervalsMore sensitive than Total Count; evaluates consistency interval-by-intervalDoes not confirm agreement on exact instances within the interval
Exact Count-per-IntervalIntervals with 100% Identical CountTotal Intervals×100%\frac{\text{Intervals with 100\% Identical Count}}{\text{Total Intervals}} \times 100\%High-precision count across intervalsMost rigorous, conservative count measureHighly stringent; yields low agreement scores in active sessions
Trial-by-Trial IOAAgreed TrialsTotal Trials×100%\frac{\text{Agreed Trials}}{\text{Total Trials}} \times 100\%Discrete trial training (DTT); restricted operantsPrecise discrete opportunity agreement; easy to calculateDependent on explicit trial pacing and clear trial demarcations
Total Duration IOAShorter DurationLonger Duration×100%\frac{\text{Shorter Duration}}{\text{Longer Duration}} \times 100\%Total session duration; gross time checksSimple to calculate with standard stopwatchesOverestimates agreement; ignores whether same episodes were timed
Mean Duration-per-Occurrence∑(Shorteri/Longeri)n×100%\frac{\sum (\text{Shorter}_i / \text{Longer}_i)}{n} \times 100\%Episode-by-episode continuous durationEvaluates timing precision across each discrete behavioral episodeDemands dual stopwatches and exact onset/offset synchronization
Interval-by-Interval IOAAgreed Intervals (both + or −)Total Intervals×100%\frac{\text{Agreed Intervals (both } + \text{ or } -)}{\text{Total Intervals}} \times 100\%Interval systems with moderate rate (30% to 70%)Comprehensive; evaluates all observation intervalsSeverely inflated by chance agreement in very high-rate or low-rate behaviors
Scored-Interval IOAAgreed +/+Intervals with at least one +×100%\frac{\text{Agreed } +/+}{\text{Intervals with at least one } +} \times 100\%Low-rate behaviors occurring in <30%< 30\% of intervalsEliminates artificial inflation from joint non-occurrences (−/−-/-)Inapplicable to continuous or high-rate behaviors
Unscored-Interval IOAAgreed −/−Intervals with at least one −×100%\frac{\text{Agreed } -/-}{\text{Intervals with at least one } -} \times 100\%High-rate behaviors occurring in >70%> 70\% of intervalsEliminates artificial inflation from joint occurrences (+/++/+)Inapplicable to low-rate behaviors
Loading diagram...
Interval-Based IOA Selection Algorithm
Test Your Knowledge

Two independent observers record the frequency of vocal stereotypy across four consecutive 5-minute intervals. Observer 1 records: Interval 1 = 2, Interval 2 = 3, Interval 3 = 0, Interval 4 = 4. Observer 2 records: Interval 1 = 2, Interval 2 = 1, Interval 3 = 0, Interval 4 = 5. What is the calculated Mean Count-per-Interval IOA for this session?

A

88.9%

B

78.3%

C

94.2%

D

50.0%

Test Your Knowledge

An assistant behavior analyst measures severe property destruction that occurs very infrequently, occurring in only 4 out of 100 observation intervals (a low-rate behavior). Two observers collect interval data. Interval-by-Interval (Point-by-Point) IOA is calculated at 94%, primarily because both observers agreed on 92 intervals where the behavior did not occur. However, on the 6 intervals where at least one observer scored an occurrence, they only agreed on 3 intervals. Why is Interval-by-Interval IOA misleading here, and what calculation should the analyst report?

A

Unscored-Interval IOA should be reported because it isolates non-occurrences and proves high observer agreement.

B

Total Count IOA should be reported because interval systems are invalid for low-rate responses.

C

Scored-Interval IOA, because it discards joint non-occurrences that inflate agreement for low-rate behavior.

D

Interval-by-Interval IOA is accurate and should be retained because 94% exceeds the standard 80% threshold.

Test Your Knowledge

A supervisor reviews video recordings of an intervention session where calibrated digital time-stamped video confirms that a client engaged in exactly 20 instances of physical aggression. During the live session, two independent behavioral technicians both recorded exactly 11 instances of aggression, yielding an Interobserver Agreement (IOA) score of 100%. How should the behavior analyst interpret these measurement results?

A

The measurement data demonstrate low reliability, because the technicians failed to capture nine instances of aggression.

B

The measurement data demonstrate high reliability (consistency between observers) but low accuracy (divergence from the true value).

C

The measurement data demonstrate an invalid operational definition, because accuracy cannot be assessed without continuous duration recording.

D

The measurement data demonstrate high validity and high accuracy, because both observers achieved perfect agreement.

Sections you finish are checked off in the contents.