10.1 Correlation & Scatter Plots
Key Takeaways
Scatter plots plot paired Cartesian coordinates (X_i, Y_i) to visually reveal the form, direction, and dispersion of bivariate relationships.
The Pearson correlation coefficient (r) quantifies the strength and direction of linear association between continuous variables, strictly ranging from -1.0 to +1.0.
A correlation coefficient near zero indicates the absence of a linear trend, but does not rule out strong non-linear or curvilinear patterns.
Statistical correlation proves numerical co-variation but never demonstrates cause-and-effect between process inputs and outputs.
Confounding or lurking variables frequently produce spurious correlations that vanish when true underlying root causes are experimentally controlled.
Correlation & Scatter Plots
Quick Answer: A scatter plot graphs paired coordinates on Cartesian axes to visually uncover relationships between process inputs () and outputs (). The Pearson correlation coefficient () quantifies linear association on a scale from (perfect negative) to (perfect positive), with indicating no linear link. Crucially, correlation demonstrates numerical association, never causation; unmeasured lurking variables frequently generate misleading statistical links that require controlled experimentation to verify. Independent CSSYB study guide by OpenExamPrep.
Bivariate Data Analysis in the Analyze Phase
In the DMAIC roadmap, the Analyze phase transitions an improvement team from describing process performance to discovering the fundamental root causes of variation and defects. While the Measure phase relies on univariate statistics (mean, standard deviation, histograms) to baseline individual metrics, univariate analysis cannot explain why an outcome fluctuates.
To isolate root causes, Six Sigma practitioners evaluate bivariate data—paired measurements of two characteristics collected from the same process units. Bivariate analysis investigates the Six Sigma transfer function:
Where is the Key Process Output Variable (KPOV) or customer CTQ metric, represents a candidate Key Process Input Variable (KPIV) or operating parameter, and represents common-cause random variation. By analyzing bivariate data, Yellow Belts separate vital process drivers from inconsequential noise variables.
Scatter Plots: Visualizing Bivariate Relationships
The primary graphical tool for exploring bivariate relationships is the scatter plot. A scatter plot plots paired observations on Cartesian coordinates, placing the independent input variable on the horizontal -axis and the dependent response on the vertical -axis.
Practitioners examine four core visual features on a scatter plot:
- Form: The overall geometric shape of the data points.
- Linear: Points cluster along a straight path where changes at a constant rate relative to .
- Curvilinear (Non-linear): Points trace a curved trajectory, such as a quadratic parabola or exponential curve. For example, chemical yield often rises with temperature up to an optimum point before dropping due to thermal degradation.
- Direction: The trajectory of the relationship.
- Positive Association: As increases, increases (upward slope from lower-left to upper-right).
- Negative Association: As increases, decreases (downward slope from upper-left to lower-right).
- Zero / No Association: Points form a diffuse, circular cloud with no discernible direction.
- Strength: The degree of dispersion around the central trend. Points tightly clustered around a line indicate a strong relationship, whereas widely dispersed points indicate a weak relationship.
- Outliers and Leverage: Scatter plots immediately expose isolated observations that depart from the general pattern. Extreme observations near the ends of the -axis exert high leverage and can distort statistical metrics.
Pearson Product-Moment Correlation Coefficient ()
Because visual interpretation of scatter plots can be subjective, Six Sigma uses the Pearson product-moment correlation coefficient () to quantify the strength and direction of linear association between two continuous variables:
Interpretation Scale and Bounds
The correlation coefficient is bounded strictly between and :
The algebraic sign reflects direction ( for positive slope, for negative slope), while the absolute magnitude () indicates strength:
| Correlation Range | Relationship Description | Process Interpretation |
|---|---|---|
| Perfect positive linear | Points fall exactly on an upward straight line | |
| Strong positive linear | Strong candidate root cause; tracks directly | |
| Moderate positive linear | Definite co-variation; influences with other factors | |
| Weak or no linear correlation | Little or no straight-line relationship | |
| Moderate negative linear | Inverse trend; increasing reduces | |
| Strong negative linear | Strong candidate root cause; increasing sharply reduces | |
| Perfect negative linear | Points fall exactly on a downward straight line |
Critical Assumptions and Boundaries
Yellow Belts must verify four key conditions when applying Pearson :
- Continuous Data: Both variables must be continuous interval or ratio measurements.
- Linearity: Pearson evaluates only linear association. For a strong U-shaped non-linear relationship, may equal zero because positive and negative slopes cancel out.
- Bivariate Normality: Variables should approximate normal distributions.
- Outlier Vulnerability: A single wild outlier can artificially inflate or suppress .
Critical Caveat: Correlation vs. Causation
The cardinal rule of bivariate analysis is that correlation does not equal causation. Calculating a high correlation coefficient () proves that two metrics move together mathematically, but does not prove that changing causes to change.
Misinterpreting correlation as causation frequently leads to flawed process changes due to:
1. Confounding and Lurking Variables
A lurking variable () is an unmeasured third factor that simultaneously influences both and .
Example: A manufacturing plant finds a strong positive correlation () between monthly electricity usage () and total product defects (). Cutting power will not eliminate defects: the lurking variable is monthly production volume (). High production volume increases electricity consumption and naturally yields more total defects.
2. Reverse Causality and Coincidence
Causality may flow in reverse ( causes ), or two unrelated metrics may correlate purely by chance or shared historical trends (spurious correlation).
Establishing True Causation
To prove causality in the Analyze phase, teams must:
- Validate Physical Mechanisms: Verify that a known engineering, physical, or operational law explains the relationship.
- Conduct Multi-Vari Studies: Stratify data across shifts, machines, and lots to verify the relationship persists across subgroups.
- Execute Controlled Experiments (DOE): Intentionally manipulate input under controlled conditions while holding all other factors constant to observe the response in .
When examining a scatter plot of furnace temperature (X) versus component tensile strength (Y), the plotted points form a distinct inverted U-shaped parabolic curve. The calculated Pearson correlation coefficient is r = 0.04. How should a Six Sigma Yellow Belt interpret this result?
A strong non-linear relationship exists between temperature and strength, but Pearson r cannot detect it because r measures only linear association.
There is no functional relationship between furnace temperature and component tensile strength because r is close to zero.
The furnace temperature is perfectly optimized across all ranges, indicating maximum process capability.
An arithmetic error must have occurred during computation, because any curved pattern always produces a negative correlation coefficient.
A project team analyzing customer service operations finds a strong positive correlation (r = 0.88) between the number of customer support representatives working a shift and the total count of customer complaints logged. What is the most appropriate statistical conclusion?
Hiring additional customer support representatives directly causes an increase in customer dissatisfaction and service errors.
The department manager should immediately reduce staffing levels to minimize the number of customer complaints received.
The calculated correlation coefficient is statistically invalid because Pearson r cannot exceed 0.50 in service environments.
The variables exhibit strong linear co-variation, but this does not prove causation; an unmeasured lurking variable, such as total customer transaction volume, may drive both.
Which of the following Pearson correlation coefficients indicates the strongest linear association between cutting speed and machined surface roughness?
+0.72
-0.89
+0.45
-0.15
Sections you finish are checked off in the contents.