11.3 Skewed Distributions, Outlier Effects & Bivariate Scatter Plots (Trend Lines & Correlation)
Key Takeaways
In symmetric unimodal distributions, the mean, median, and mode are approximately equal (Mean ≈ Median ≈ Mode).
Skewed distributions separate measures of center along the tail: in positively skewed data, Mode < Median < Mean; in negatively skewed data, Mean < Median < Mode.
Bivariate scatter plots assess the direction (positive, negative), form (linear, non-linear), and strength of associations between explanatory (x) and response (y) variables.
The slope of a trend line gives the predicted change in the response per unit of the explanatory variable, residuals (observed minus predicted) and residual plots test whether a linear model fits, and extrapolation beyond the observed domain is hazardous.
Correlation does not imply causation; an observed association between two variables may be driven by unmeasured confounding (lurking) variables or coincidental co-variation.
11.3 Skewed Distributions, Outlier Effects & Bivariate Scatter Plots (Trend Lines & Correlation)
A comprehensive foundation in statistics requires extending univariate analysis (examining single variables in isolation) into bivariate analysis (exploring relationships, associations, and predictions between two coupled variables). Understanding how distribution asymmetry influences univariate measures of center provides the bridge to bivariate modeling. When investigating bivariate relationships, educators guide students to construct scatter plots, estimate trend lines, interpret slope and intercept in contextual terms, and critically evaluate the boundaries of statistical claims—particularly distinguishing correlation from causation.
Distribution Shapes and Skewness
The geometric profile of a frequency distribution reveals how data is dispersed relative to its center. Deviations from symmetry dictate the spatial relationships among the mean, median, and mode.
1. Symmetric (Normal / Bell-Shaped) Distributions
In a perfectly symmetric unimodal distribution, the left and right halves are mirror images balanced around the center. The frequency of observations rises smoothly to a single peak and tapers off equally in both directions.
- Relationship of Center Measures: The arithmetic mean, median, and mode coincide at approximately the same numerical location:
- The physical fulcrum (mean) matches exactly with the 50th percentile rank (median) and the peak frequency (mode).
2. Positively Skewed (Skewed Right) Distributions
A distribution is positively skewed (skewed right) when the right tail stretches out much farther and thinner than the left tail. The bulk of observations cluster at lower values on the left side of the display.
- Mechanism: The extreme high values residing in the prolonged right tail exert powerful upward leverage on the arithmetic mean, pulling it rightward. The median, being resistant, shifts only slightly, while the mode remains anchored at the high peak on the left.
- Inequality Order:
- Authentic Contexts: Household income, corporate executive salaries, housing prices, vehicle mileage, and emergency room wait times.
3. Negatively Skewed (Skewed Left) Distributions
A distribution is negatively skewed (skewed left) when the left tail stretches out much farther and thinner than the right tail. The bulk of observations cluster at higher values on the right side of the display.
- Mechanism: The extreme low values residing in the prolonged left tail pull the non-resistant arithmetic mean downward toward the left. The median remains near the center of the ranked data, and the mode remains at the peak on the right.
- Inequality Order:
- Authentic Contexts: Student scores on an accessible mastery test (where most score high and a few score very low), age at retirement, or human gestation periods.
4. Uniform and Bimodal Distributions
- Uniform Distribution: Frequencies are approximately constant across all intervals; the distribution has no distinct mode, and mean and median coincide at the geometric midpoint.
- Bimodal Distribution: Features two distinct local peaks. In educational and biological contexts, a bimodal distribution often signals that two distinct underlying sub-populations have been pooled together (for example, combining adult male and female shoe sizes into a single distribution).
Reference Summary: Distribution Shapes and Summary Statistics
| Distribution Profile | Visual Tail Characteristic | Center Measure Inequality | Preferred Measure of Center | Preferred Measure of Spread |
|---|---|---|---|---|
| Symmetric (Bell-shaped) | Equal, balanced tails on both sides | Mean () | Standard Deviation () or MAD | |
| Positively Skewed (Right) | Prolonged, elongated tail to the right | Median () | Interquartile Range () | |
| Negatively Skewed (Left) | Prolonged, elongated tail to the left | Median () | Interquartile Range () | |
| Bimodal | Two distinct frequency peaks | Center depends on peak symmetry | Report both modes and subgroup medians | Subgroup standard deviations |
Bivariate Data Exploration: Scatter Plots
Bivariate data involves paired measurements collected on the same individual subjects or experimental units to investigate whether a relationship exists between two quantitative variables.
Coordinate Conventions
- Explanatory (Independent) Variable (): Placed on the horizontal -axis. It is the variable hypothesized to explain, predict, or influence changes in the response.
- Response (Dependent) Variable (): Placed on the vertical -axis. It is the outcome variable measured to assess the effect of the explanatory variable.
Visual Analysis Framework for Scatter Plots
When interpreting a scatter plot, educators evaluate four fundamental characteristics:
- Direction:
- Positive Association: As increases, tends to increase (points slope upward from lower-left to upper-right).
- Negative Association: As increases, tends to decrease (points slope downward from upper-left to lower-right).
- No Association: Changes in exhibit no systematic upward or downward trend in .
- Form:
- Linear: The point cloud approximates a straight-line trajectory.
- Non-Linear: The pattern follows a distinct curve (e.g., quadratic parabola, exponential growth/decay, or logarithmic curve).
- Strength:
- Refers to how tightly the points cluster around the underlying functional path. Relationships are classified qualitatively as strong, moderate, or weak based on the dispersion (vertical spread) of points from the trend line.
- Outliers and Influential Points:
- Points that depart substantially from the overall pattern of the rest of the data. An observation with an extreme -value (high leverage) can act as an influential point, dramatically altering the slope of a trend line if removed.
The Pearson Correlation Coefficient ()
The Pearson product-moment correlation coefficient () quantitatively measures the direction and strength of the linear relationship between two quantitative variables:
Mathematical Properties of
- Fixed Range: The correlation coefficient is strictly bounded between and :
- : Perfect positive linear correlation (all points fall exactly on a line with positive slope).
- : Perfect negative linear correlation (all points fall exactly on a line with negative slope).
- : Indicates the complete absence of a linear association.
- Dimensionless: The value of has no units of measurement because it is calculated from standardized -scores.
- Scale Invariance: Changing the units of measurement (such as converting height from inches to centimeters or temperature from Fahrenheit to Celsius) does not alter the numerical value of .
- Symmetry: Switching the roles of and does not change ; .
- Linearity Restriction: The correlation coefficient measures only linear relationships. A scatter plot displaying a perfect quadratic relationship () may have , yet the two variables are deterministically related. Always inspect the graphical display alongside numerical statistics.
Trend Lines and Lines of Best Fit (Linear Models)
In middle-grades mathematics, students transition from informal eyeball lines of best fit to formal linear regression models.
Informal Fitting Principles
When drawing an informal line of best fit through a scatter plot:
- The line must pass through the center of the cloud of points, balancing the number of points located above and below the line.
- The line should pass through the centroid (center of gravity) of the data, defined by the coordinates of the means: .
The Linear Model Equation
A linear trend line is expressed algebraically as:
where (read "y-hat") denotes the predicted value of the response variable for a given value of .
- Interpreting the Slope ( or ): Contextual Template: "For each additional 1-unit increase in the explanatory variable (), the model predicts that the response variable () will increase (or decrease) by an average of units."
- Interpreting the -Intercept ( or ): Contextual Template: "When the explanatory variable () is equal to 0, the model predicts that the response variable () will be units." Critical Thinking: Evaluate whether is practically meaningful or physically possible within the problem context (e.g., predicting a student's weight at height 0 cm is a mathematical artifact without physical reality).
Interpolation vs. Extrapolation
- Interpolation: Estimating a value of for an -value that falls within the domain of observed data (between and ). Interpolated predictions are generally reliable because the established linear relationship has been empirically observed across that interval.
- Extrapolation: Estimating a value of for an -value that lies well outside the domain of observed data. Extrapolation is extremely hazardous because there is no empirical evidence that the linear trend continues beyond the measured range. In physical, biological, and economic systems, linear trends routinely plateau, curve, or reverse outside observed bounds.
Least-Squares Regression and Residual Analysis
Competency 014 expects you to use scatter plots, regression lines, correlation coefficients and residual analysis to explore bivariate data and judge predictions.
The least-squares regression line
Of all possible lines, the least-squares regression line makes the sum of the squared vertical distances from the points to the line as small as possible. Two facts are worth memorizing:
- Slope: , where is the correlation and are the standard deviations.
- The line always passes through , the point of means.
Residuals
A residual is the prediction error for one data point:
- A positive residual means the point lies above the line, so the model under-predicted.
- A negative residual means the point lies below the line, so the model over-predicted.
- For the least-squares line, the residuals sum to zero.
Worked example. A model predicts quiz score from study time: . A student who studied hours scored 75. The prediction is , so the residual is . The student scored 4 points below the model's prediction.
Residual plots
A residual plot graphs residuals (vertical axis) against or against .
- Random scatter around 0 with no pattern: a linear model is appropriate.
- A curved (U-shaped or arched) pattern: the relationship is nonlinear, so a line is the wrong model even when is fairly large.
- A fan shape (spread growing with ): predictions become less reliable for large .
Correlation strength and
The square of the correlation, (the coefficient of determination), gives the fraction of the variation in that the linear model accounts for. If , then , so about 81% of the variation in is explained by the linear relationship with .
| Tool | Question it answers | Warning sign |
|---|---|---|
| Scatter plot | Is there an association? What direction, form and strength? | A curved form makes a line the wrong model |
| Correlation | How strong is the linear association? | near 0 does not rule out a strong curved pattern |
| Regression line | What is the predicted for a given ? | Extrapolating beyond the data range |
| Residual plot | Is a linear model appropriate? | A curved or fan-shaped pattern |
Correlation vs. Causation: The Core Statistical Principle
One of the most vital critical-thinking lessons in statistical education is: Correlation does not imply causation.
An observed strong correlation between variable and variable () does not prove that causes . Three primary alternative mechanisms can explain an observed correlation:
- Lurking / Confounding Variables: An unmeasured third variable () simultaneously influences both and , creating a spurious association.
- Classic Example: Ice cream sales () and swimming pool drownings () exhibit a strong positive correlation (). Does eating ice cream cause drowning? No. The lurking variable is outdoor summer temperature (); hot weather causes both higher ice cream consumption and increased swimming activity.
- Reverse Causality: Variable actually causes changes in variable , rather than the assumed direction.
- Coincidence / Spurious Correlation: In large data sets, two unrelated variables may exhibit a high correlation purely by random chance.
Establishing Causality: To prove a causal relationship, researchers must conduct a randomized controlled experiment where subjects are randomly assigned to treatment and control groups, actively isolating the variable of interest while controlling for confounding factors.
Worked Step-by-Step Mathematical Examples
Worked Example 1: Demonstrating the Center Inequality in Skewed Data
Problem: A burgeoning educational technology company employs 7 individuals. The annual salaries are:
Calculate the mean, median, and mode salary. Identify the shape of the distribution, verify the mathematical inequality among the measures of center, and determine which measure best communicates the "typical" employee salary.
Solution:
- Step 1: Calculate the Mode: The value $50,000 appears twice; all other values appear once.
- Step 2: Calculate the Median: The salaries are already ordered. The median is the 4th observation:
- Step 3: Calculate the Arithmetic Mean ():
- Step 4: Analyze Distribution Shape and Inequality: The extreme salary of the company founder ($240,000) creates an elongated tail to the right, producing a positively skewed (skewed right) distribution. Notice the relationship:
- Step 5: Contextual Interpretation: The mean salary of approximately $81,429 is higher than 6 out of the 7 employees' salaries (86% of the company)! Reporting the mean would falsely suggest that employees are well compensated near $80,000. The median of $50,000 represents a resistant, authentic summary of the typical employee's earnings.
Worked Example 2: Bivariate Modeling, Slope Interpretation, and Extrapolation Hazard
Problem: An eighth-grade science teacher guides students in tracking the relationship between the number of hours spent studying per week () and the resulting score on a 100-point comprehensive exam (). The study tracked student hours ranging from 2 to 10 hours per week. The calculated linear trend line is:
- Interpret the slope of the model in context.
- Interpret the -intercept of the model in context.
- Predict the exam score for a student who studies 6 hours per week, and classify this prediction as interpolation or extrapolation.
- Predict the exam score for a student who studies 20 hours per week, and explain why this calculation represents a hazardous extrapolation.
Solution:
-
Step 1: Interpret the Slope (): The slope of indicates that for each additional 1 hour of weekly study time, the model predicts an average increase of 4.5 points on the comprehensive exam score.
-
Step 2: Interpret the -Intercept (): The -intercept indicates that a student who engages in 0 hours of weekly study is predicted to achieve an exam score of 52 points.
-
Step 3: Prediction at Hours:
Because falls comfortably within the observed domain of 2 to 10 hours, this is an interpolation and represents a reliable statistical prediction.
- Step 4: Prediction at Hours and Extrapolation Hazard:
This calculation yields an impossible score of 142 points on a 100-point test! Because lies far outside the observed domain of 2 to 10 hours, this is an extrapolation. In reality, exam scores are physically bounded at 100%, and students encounter fatigue, diminishing returns, and maximum test ceilings. The linear relationship cannot hold indefinitely beyond the observed domain.
Diagnostic Misconceptions & Pedagogical Strategies
- Reversing the Direction of Skewness: Students routinely confuse "skewed right" with having the "bump on the right." Pedagogical Strategy: Teach students that skewness refers to the tail, not the peak. Emphasize: "The tail tells the tale." If the tail stretches to the right toward positive infinity, the distribution is skewed right (positively skewed). Use the visual image of pulling salt or dough across a table to show that the tail pulls the mean in its direction.
- The Causation Fallacy: Students frequently infer that a high correlation proves direct physical cause-and-effect. Pedagogical Strategy: Introduce absurd yet real spurious correlations (e.g., the correlation between per capita cheese consumption and the number of people who died by becoming tangled in their bedsheets). Have students identify plausible lurking variables (such as population growth or economic development) to cement the necessity of randomized experiments for proving causality.
- Blind Extrapolation: Students often plug arbitrary numbers into linear regression equations without checking domain bounds. Pedagogical Strategy: Have students calculate a linear model for height versus age from ages 2 to 10, and then ask them to predict a person's height at age 45. The resulting prediction of 12 feet tall immediately alerts students to the danger of extrapolating beyond observed data.
A state testing committee reviews scores on an advanced mathematics diagnostic exam administered to 8th-grade students. The distribution exhibits a prolonged left tail containing a small group of very low scores, while the vast majority of students scored between 82% and 96%. Which inequality correctly represents the relative order of the measures of central tendency for this distribution?
Mean < Median < Mode
Mode < Median < Mean
Median < Mean < Mode
Mean = Median = Mode
A biology class creates a scatter plot to analyze the relationship between daily sunlight exposure in hours (x) and the growth rate of tomato plants in centimeters per week (y) over a range of 2 to 8 hours of daily sunlight. The calculated line of best fit is y_hat = 1.4x + 3.2, with a strong correlation coefficient of r = 0.94. A student uses this model to predict the weekly growth of a plant subjected to 24 hours of continuous artificial light, obtaining y_hat = 1.4(24) + 3.2 = 36.8 cm. Which statistical principle explains why this prediction is unreliable?
The model is invalid because linear trend lines can only be used when the correlation coefficient r is exactly 1.0.
The calculation confused the explanatory variable with the response variable, so the student should have solved for x when y = 24.
The prediction is an interpolation error because plant growth cannot exceed the y-intercept value of 3.2 cm.
The prediction is an extrapolation far outside the domain of observed data (2 to 8 hours), where biological constraints may cause the linear relationship to break down or reverse.
A researcher finds a strong positive correlation (r = 0.88) between the monthly sales of ice cream in a metropolitan area and the number of reported swimming pool rescues during the same months. A municipal official proposes restricting ice cream sales near public pools to reduce drowning accidents. What fundamental statistical flaw undermines the official's proposal?
A correlation coefficient of 0.88 is statistically insignificant and indicates that the relationship between sales and rescues is purely random.
Scatter plots can only establish relationships between categorical variables; continuous variables require a contingency table to verify causality.
Correlation does not imply causation; the observed association between ice cream sales and pool rescues is confounded by a lurking variable, outdoor summer temperature, which drives both phenomena.
The linear model failed to include the negative slope produced by indoor swimming facilities during winter months.
Sections you finish are checked off in the contents.