5.2 Visual Analysis of Graphed Data: Level, Trend, Variability, and Immediacy of Effect
Key Takeaways
Visual analysis is a conservative, criterion-referenced evaluation method in single-case research that detects large, robust, and socially valid functional relations while minimizing Type I errors.
Level represents the central tendency of data within a phase, evaluated via the mean level line (suitable for normally distributed, low-variability data) or median level line (preferred when extreme outliers distort the arithmetic mean).
Trend characterizes the overall directional trajectory of responding (accelerating, decelerating, or zero trend), estimated visually or quantitatively through the split-middle method of trend estimation (quarter-intersect line).
Variability quantifies the bounce or dispersion of data around the central tendency; establishing steady-state responding with minimal variability is generally required before an experimental phase change.
Between-phase comparisons examine the immediacy of effect (abruptness of change following independent variable manipulation), percentage of non-overlapping data (PND), and consistency of data patterns across repeated experimental phases.
The Philosophy and Rationale of Visual Inspection
In single-case experimental methodology, the primary method for evaluating data is visual analysis (visual inspection). Rather than aggregating scores across large groups and calculating inferential statistics (-tests, ANOVAs, or -values), behavior analysts visually inspect an individual participant's repeated measurements across time to determine whether a functional relation exists between the independent variable and the dependent variable.
Baer, Wolf, and Risley (1968), Michael (1974), and Parsonson and Baer (1978) articulated the core rationale behind visual inspection:
- A Conservative Criterion for Discovery: Visual analysis represents a conservative standard. It is deliberately designed to detect only robust, potent, and obvious behavioral changes. By demanding that an experimental effect be clearly visible to the naked eye, visual analysis minimizes Type I errors (false positives: concluding that an intervention produced an effect when the change was actually an artifact of chance or extraneous noise).
- Social Validity over Statistical Significance: In group-design statistical research, an intervention that produces a minor, clinically negligible difference (e.g., reducing self-injurious head hits from 50 to 48 per hour) can be declared "statistically significant" at simply by increasing the sample size (). However, in clinical practice, a client whose head hits drop from 50 to 48 experiences zero meaningful clinical benefit. Behavior analysis demands clinical significance and social validity—the behavior change must fundamentally enhance the client's quality of life.
- Dynamic, Ongoing Experimental Decision-Making: Inferential statistics are post-hoc calculations performed after an entire clinical trial has ended. In contrast, visual analysis is performed session by session, empowering the assistant behavior analyst to detect trends, identify unprogrammed environmental intrusions, and decide in real time when to advance or modify an intervention.
Visual analysis is systematically divided into two interconnected domains: within-phase visual analysis and between-phase visual analysis.
Within-Phase Dimensions: Level, Trend, and Variability
Before comparing data across different experimental conditions, the behavior analyst must thoroughly analyze the data path within each individual phase. Within-phase analysis examines three primary dimensions: level, trend, and variability.
1. Level
Level refers to the position of the data set on the vertical ordinate; it characterizes the overall magnitude or central tendency of behavior within a specific experimental phase.
To quantify and visually depict level within a phase, behavior analysts utilize two primary metrics:
- Mean Level Line: The arithmetic average of all data points within a condition, calculated by summing all values and dividing by the total number of points (). The mean level line is drawn as a horizontal dashed line extending across the width of the phase. It is appropriate when data are relatively stable and follow an approximately normal distribution without extreme outliers.
- Median Level Line: The middle value of the data set when points are ranked in order of magnitude (or the average of the two middle points when is even). The median level line is drawn as a horizontal dashed line across the phase. The median level line is strongly preferred whenever a data set contains extreme outliers or skewed distributions, because the median is resistant to the distorting influence of anomalous data points.
Level Changes Within Phase
In addition to overall central tendency, clinicians evaluate the within-phase level change: the difference in behavioral magnitude between the first data point and the last data point of that condition. A substantial shift between the initial and terminal data points indicates that the behavior changed substantially across the duration of the phase.
2. Trend
Trend refers to the overall directional slope or trajectory of the data path over time. Trend answers the question: Is behavior systematically increasing, decreasing, or remaining flat?
Trend is characterized along two properties:
- Direction:
- Accelerating / Increasing: The data path exhibits an upward slope over time.
- Decelerating / Decreasing: The data path exhibits a downward slope over time.
- Zero Trend: The data path exhibits a flat, horizontal trajectory parallel to the abscissa.
- Magnitude / Steepness: The steepness of the slope, characterized qualitatively (e.g., steep, moderate, or gradual) or quantitatively (e.g., celeration rates on an SCC).
The Split-Middle Method of Trend Estimation (Quarter-Intersect Line)
While experienced behavior analysts often estimate trend visually using a "freehand" line of best fit, subjective estimation can introduce clinician bias. The split-middle method of trend estimation (White, 1971; Cooper, Heron, & Heward, 2020) provides an objective, mathematically replicable linear trend line:
- Step 1: Divide the Data Set in Half Vertically: Count the total number of data points () in the phase. Draw a vertical dashed line dividing the phase into two equal halves along the horizontal axis. If is odd, the vertical line falls directly on the middle data point (which can be assigned to both halves or omitted from calculation).
- Step 2: Find Mid-Dates and Mid-Rates for Each Half:
- For the left half, identify the median session number (mid-date) along the horizontal axis, and find the median response value (mid-rate) along the vertical axis. Mark the intersection of the mid-date and mid-rate as the first quarter-intersect point.
- Repeat this exact process for the right half to mark the second quarter-intersect point.
- Step 3: Connect the Two Quarter-Intersect Points: Draw a straight line passing through both quarter-intersect points and extending across the entire phase.
- Step 4: Adjust the Line to Balance Points (Split-Middle): Count how many data points fall strictly above the line and how many fall strictly below the line. If the line does not divide the data points equally, shift the line up or down parallel to itself until an equal number of data points fall on or above the line as fall on or below the line. This final line is the quarter-intersect split-middle trend line.
The split-middle trend line can be projected into subsequent experimental phases to establish an empirical baseline prediction against which intervention data are compared.
3. Variability
Variability refers to the degree of bounce, dispersion, or fluctuation of data points around the central tendency (level) or trend line.
- High Variability: The data points fluctuate wildly across sessions without a discernible pattern. High variability signals a lack of environmental control or the intrusion of uncontrolled extraneous variables (e.g., inconsistent sleep schedules, fluctuating caregiver implementation fidelity, acute medication changes, or unprogrammed reinforcement).
- Low Variability: The data points cluster tightly around the mean or trend line, exhibiting minimal bounce. Low variability demonstrates tight environmental control.
The Steady-State Strategy
In behavior analysis, single-case experimental control rests upon the steady-state strategy: the empirical practice of continuing an experimental phase until responding displays minimal variability and a stable level and trend (steady-state responding). Introducing an intervention or changing phases in the presence of excessive variability is methodologically hazardous, as the clinician cannot determine whether subsequent behavioral shifts are caused by the intervention or by ongoing uncontrolled noise.
Between-Phase Dimensions: Immediacy, Overlap, and Consistency
Once within-phase characteristics are documented, the behavior analyst performs between-phase visual analysis by comparing the final data points of one condition with the initial data points of the subsequent condition, and examining broader patterns across adjacent and repeated phases.
1. Immediacy of Effect
Immediacy of effect refers to the rapidity with which behavior changes following the introduction or withdrawal of the independent variable. It evaluates how quickly the level and trend shift between adjacent conditions.
- Assessment Method: Clinicians inspect the vertical discontinuity between the final data point of the preceding condition and the first one to three data points of the new condition.
- Immediate Effect: An abrupt, immediate shift in level or trend upon introducing the independent variable provides powerful evidence that the independent variable functionally caused the change.
- Delayed Effect: If responding changes only gradually over many sessions, confidence in experimental control is substantially weakened. A delayed effect introduces plausible competing explanations, such as maturation, seasonal shifts, or unprogrammed external environmental events. While some interventions (such as extinction or educational skill shaping) inherently produce gradual trends, experimental control must be rigorously demonstrated through repeated replications.
2. Overlap and Percentage of Non-Overlapping Data (PND)
Overlap refers to the proportion of data points in the intervention phase that fall within the range of values observed during the baseline phase.
- Clinical Significance: The less overlap that exists between baseline and intervention data, the higher the confidence that the independent variable produced a meaningful effect. Extensive overlap indicates weak or questionable intervention efficacy.
- Percentage of Non-overlapping Data (PND): A standardized, objective metric developed by Scruggs, Mastropieri, and Casto (1987) to quantify between-phase overlap:
Calculating PND for Behavior Reduction vs. Behavior Acceleration
- For Behavior Reduction Targets (e.g., Aggression, SIB): Identify the lowest data point in the baseline phase. Count the number of intervention data points that fall strictly below this lowest baseline value. Divide by the total number of intervention data points and multiply by 100.
- For Behavior Acceleration Targets (e.g., Academic Fluency, Mands): Identify the highest data point in the baseline phase. Count the number of intervention data points that fall strictly above this highest baseline value. Divide by the total number of intervention data points and multiply by 100.
Empirical Interpretation Guidelines (Scruggs & Mastropieri)
- PND > 90%: Highly effective intervention; robust functional relation.
- PND = 70% to 90%: Moderately effective intervention.
- PND = 50% to 69%: Questionable or marginal intervention effectiveness.
- PND < 50%: Ineffective intervention; failure to demonstrate behavioral control.
3. Consistency of Data Patterns Across Similar Phases
Consistency examines whether repeated implementations of the exact same experimental condition produce identical or highly similar data patterns across time or across participants:
- In an ABAB Reversal Design, does the data trajectory observed in the second baseline phase () replicate the pattern observed in the initial baseline phase ()? Does the second intervention phase () reproduce the level and trend achieved in the first intervention phase ()?
- In a Multiple Baseline Design, does introducing the independent variable produce the same abrupt shift in level and trend across Tier 1, Tier 2, and Tier 3?
High consistency across identical experimental conditions verifies the reliability of the behavioral technology.
Interpreting Challenging and Ambiguous Baseline Data Patterns
In applied settings, clinical data rarely resemble textbook perfection. Assistant behavior analysts must navigate three challenging baseline data trajectories when deciding whether to introduce an intervention:
1. Counter-Therapeutic Baseline Trend
A counter-therapeutic trend occurs when baseline responding is actively deteriorating over time (e.g., self-injury or property destruction is accelerating upward, or an essential communication skill is decelerating downward).
- Clinical Action: Introduce the intervention immediately.
- Methodological Rationale: Introducing an independent variable into a counter-therapeutic baseline does NOT compromise internal validity. If an intervention halts an escalating problem behavior and drives it downward, the change runs opposite to the baseline trajectory, which is strong preliminary evidence of an effect; replication is still needed to confirm a functional relation.
2. Therapeutic Baseline Trend
A therapeutic trend occurs when baseline responding is already improving on its own prior to intervention (e.g., aggressive outbursts are steadily dropping from 30 to 5 per day, or independent task completion is steadily climbing).
- Clinical Action: DO NOT introduce the intervention; hold baseline until responding stabilizes.
- Methodological Rationale: Introducing an independent variable during an improving baseline is a fatal experimental error. If the clinician introduces a token economy while aggression is already dropping, it is impossible to determine whether the continued drop was caused by the token economy or by whatever extraneous variable initiated the improvement during baseline (a confounding history or maturation effect).
3. Excessive Baseline Variability
When baseline data bounce wildly without stability, the clinician must not rush into intervention. Instead, the clinician should:
- Maintain baseline longer to determine whether a stable pattern or cyclical rhythm emerges.
- Systematically isolate and control extraneous environmental variables (e.g., standardize instructional demands, adjust seating, control noise levels, verify sleep schedules).
- Evaluate measurement fidelity: Confirm that observers are adhering strictly to the operational definition and collect interobserver agreement (IOA) data.
Dimensions of Visual Analysis Matrix
The following table synthesizes the within-phase and between-phase dimensions of visual analysis, their operational inspection methods, clinical significance, and common exam traps:
| Dimension | Formal Definition | Visual Inspection Method | What It Signals Clinically | Exam Pitfall / Trap |
|---|---|---|---|---|
| Level (Mean vs. Median) | The central tendency or magnitude of data points within an experimental condition. | Draw dashed horizontal line at arithmetic average (mean) or middle ranked value (median); inspect vertical position on ordinate. | Indicates overall behavioral severity or fluency; establishing change in level confirms intervention impact. | Using mean level lines when extreme outliers exist, which severely skews the visual representation of typical behavior. |
| Trend (Direction & Magnitude) | The overall directional slope or trajectory of the data path over time. | Apply the split-middle technique (quarter-intersect line) or freehand trend line; classify as accelerating, decelerating, or zero trend. | Predicts future performance in the absence of intervention; establishes whether behavior is naturally improving or deteriorating. | Introducing an intervention into a therapeutic baseline trend, which fatally confounds the demonstration of experimental control. |
| Variability (Bounce) | The degree of dispersion or scatter of data points around the mean or trend line. | Inspect vertical spread/envelope around the data path; evaluate stability of steady-state responding. | High variability signals uncontrolled extraneous variables or poor procedural control; low variability indicates experimental control. | Changing experimental phases prematurely in the presence of excessive baseline variability without isolating extraneous variables. |
| Immediacy of Effect | The rapidity with which behavior changes following independent variable introduction or withdrawal. | Inspect vertical discontinuity between the final data point of preceding phase and the first 1-3 points of the new phase. | Immediate change confirms tight functional control by the IV; delayed change suggests possible maturation or historical confounds. | Assuming delayed changes always indicate intervention failure; some interventions (extinction, shaping) inherently produce gradual trends. |
| Overlap (PND) | The proportion of intervention data points that fall within the range of values observed during baseline. | Calculate Percentage of Non-overlapping Data: divide points exceeding extreme baseline value by total intervention points . | Quantifies intervention potency (>90% = high efficacy; <50% = ineffective); establishes separation between phases. | Comparing intervention points against mean baseline level rather than the extreme baseline data point when calculating PND. |
| Consistency Across Phases | The extent to which identical experimental conditions produce similar data trajectories upon replication. | Compare level, trend, and variability across repeated phases (e.g., A1 vs. A2; B1 vs. B2) or across multiple baseline tiers. | High consistency verifies experimental reliability and internal validity; rules out idiosyncratic historical flukes. | Failing to recognize that lack of consistency across reversal phases indicates weak experimental control or irreversible skill acquisition. |
An assistant behavior analyst is conducting visual analysis on baseline data for a client's severe property destruction. The phase contains 10 sessions with the following response counts: 12, 14, 11, 13, 15, 12, 14, 13, 11, and 78. Session 10 coincided with an acute, documented ear infection. When preparing the clinical report to characterize the central tendency and trend of this baseline phase, which analytic methods should the clinician utilize?
Exclude the entire baseline phase from the report because any phase containing an outlier lacks validity and cannot be analyzed.
Use a median level line, because the median resists distortion from the extreme medical outlier, and the split-middle method to estimate trend.
Calculate the mean level line to ensure the acute medical episode is mathematically weighted into the client's long-term behavioral average.
Draw a best-fit regression line using ordinary least squares, because statistical algorithms supersede visual inspection in behavior analysis.
A behavior analyst is preparing to introduce a functional communication training (FCT) package to decrease a child's aggressive hitting. During the baseline phase, hitting occurs at the following rates across the last five sessions: 22, 18, 14, 9, and 4 instances per hour. The clinical team asks whether they should implement the intervention immediately tomorrow morning. How should the assistant behavior analyst guide the team based on single-case design principles?
Introduce the intervention immediately, but alter the operational definition of hitting to restore the baseline rate to 20 per hour.
Implement the intervention immediately, because the declining rate proves that the client is motivated to learn functional communication.
Withdraw the baseline condition and implement an immediate reversal to an alternating treatments design without collecting further data.
Postpone the intervention until hitting stabilizes, because intervening during a therapeutic baseline trend would confound the results.
A clinician evaluates an ABAB design targeting on-task academic engagement. In Baseline 1, on-task behavior ranges from 15% to 30% of intervals, with the final three sessions at 20%, 25%, and 22%. In the first session of Intervention 1 (token reinforcement), on-task engagement immediately jumps to 75% of intervals and remains between 70% and 85% across all 10 intervention sessions. When calculating the Percentage of Non-overlapping Data (PND) and evaluating between-phase characteristics, what should the clinician conclude?
The PND is 0%, indicating that the intervention failed to produce an experimental effect due to ceiling effects.
The PND is 100% and there is an immediate effect, indicating robust experimental control and high intervention effectiveness.
The PND is 75%, but the delayed latency of change suggests that extraneous historical variables caused the increase in engagement.
The PND cannot be calculated because single-case visual analysis strictly prohibits comparing baseline ranges to intervention data points.
Sections you finish are checked off in the contents.