10.1 Correlation & Scatter Plots

Key Takeaways

  • Scatter plots plot paired Cartesian coordinates (X_i, Y_i) to visually reveal the form, direction, and dispersion of bivariate relationships.

  • The Pearson correlation coefficient (r) quantifies the strength and direction of linear association between continuous variables, strictly ranging from -1.0 to +1.0.

  • A correlation coefficient near zero indicates the absence of a linear trend, but does not rule out strong non-linear or curvilinear patterns.

  • Statistical correlation proves numerical co-variation but never demonstrates cause-and-effect between process inputs and outputs.

  • Confounding or lurking variables frequently produce spurious correlations that vanish when true underlying root causes are experimentally controlled.

Last updated: September 2026

Correlation & Scatter Plots

Quick Answer: A scatter plot graphs paired coordinates (Xi,Yi)(X_i, Y_i) on Cartesian axes to visually uncover relationships between process inputs (XX) and outputs (YY). The Pearson correlation coefficient (rr) quantifies linear association on a scale from −1.0-1.0 (perfect negative) to +1.0+1.0 (perfect positive), with 0.00.0 indicating no linear link. Crucially, correlation demonstrates numerical association, never causation; unmeasured lurking variables frequently generate misleading statistical links that require controlled experimentation to verify. Independent CSSYB study guide by OpenExamPrep.

Bivariate Data Analysis in the Analyze Phase

In the DMAIC roadmap, the Analyze phase transitions an improvement team from describing process performance to discovering the fundamental root causes of variation and defects. While the Measure phase relies on univariate statistics (mean, standard deviation, histograms) to baseline individual metrics, univariate analysis cannot explain why an outcome fluctuates.

To isolate root causes, Six Sigma practitioners evaluate bivariate data—paired measurements of two characteristics collected from the same process units. Bivariate analysis investigates the Six Sigma transfer function:

Y=f(X)+ϵY = f(X) + \epsilon

Where YY is the Key Process Output Variable (KPOV) or customer CTQ metric, XX represents a candidate Key Process Input Variable (KPIV) or operating parameter, and ϵ\epsilon represents common-cause random variation. By analyzing bivariate data, Yellow Belts separate vital process drivers from inconsequential noise variables.


Scatter Plots: Visualizing Bivariate Relationships

The primary graphical tool for exploring bivariate relationships is the scatter plot. A scatter plot plots paired observations (Xi,Yi)(X_i, Y_i) on Cartesian coordinates, placing the independent input variable on the horizontal XX-axis and the dependent response on the vertical YY-axis.

Practitioners examine four core visual features on a scatter plot:

  1. Form: The overall geometric shape of the data points.
    • Linear: Points cluster along a straight path where YY changes at a constant rate relative to XX.
    • Curvilinear (Non-linear): Points trace a curved trajectory, such as a quadratic parabola or exponential curve. For example, chemical yield often rises with temperature up to an optimum point before dropping due to thermal degradation.
  2. Direction: The trajectory of the relationship.
    • Positive Association: As XX increases, YY increases (upward slope from lower-left to upper-right).
    • Negative Association: As XX increases, YY decreases (downward slope from upper-left to lower-right).
    • Zero / No Association: Points form a diffuse, circular cloud with no discernible direction.
  3. Strength: The degree of dispersion around the central trend. Points tightly clustered around a line indicate a strong relationship, whereas widely dispersed points indicate a weak relationship.
  4. Outliers and Leverage: Scatter plots immediately expose isolated observations that depart from the general pattern. Extreme observations near the ends of the XX-axis exert high leverage and can distort statistical metrics.

Pearson Product-Moment Correlation Coefficient (rr)

Because visual interpretation of scatter plots can be subjective, Six Sigma uses the Pearson product-moment correlation coefficient (rr) to quantify the strength and direction of linear association between two continuous variables:

r=∑i=1n(Xi−Xˉ)(Yi−Yˉ)∑i=1n(Xi−Xˉ)2∑i=1n(Yi−Yˉ)2r = \frac{\sum_{i=1}^{n} (X_i - \bar{X})(Y_i - \bar{Y})}{\sqrt{\sum_{i=1}^{n} (X_i - \bar{X})^2 \sum_{i=1}^{n} (Y_i - \bar{Y})^2}}

Interpretation Scale and Bounds

The correlation coefficient is bounded strictly between −1.0-1.0 and +1.0+1.0:

−1.0≤r≤+1.0-1.0 \le r \le +1.0

The algebraic sign reflects direction (++ for positive slope, −- for negative slope), while the absolute magnitude (∣r∣|r|) indicates strength:

Correlation RangeRelationship DescriptionProcess Interpretation
r=+1.0r = +1.0Perfect positive linearPoints fall exactly on an upward straight line
+0.70≤r<+1.00+0.70 \le r < +1.00Strong positive linearStrong candidate root cause; YY tracks XX directly
+0.30≤r<+0.70+0.30 \le r < +0.70Moderate positive linearDefinite co-variation; XX influences YY with other factors
−0.30<r<+0.30-0.30 < r < +0.30Weak or no linear correlationLittle or no straight-line relationship
−0.70<r≤−0.30-0.70 < r \le -0.30Moderate negative linearInverse trend; increasing XX reduces YY
−1.00<r≤−0.70-1.00 < r \le -0.70Strong negative linearStrong candidate root cause; increasing XX sharply reduces YY
r=−1.0r = -1.0Perfect negative linearPoints fall exactly on a downward straight line

Critical Assumptions and Boundaries

Yellow Belts must verify four key conditions when applying Pearson rr:

  • Continuous Data: Both variables must be continuous interval or ratio measurements.
  • Linearity: Pearson rr evaluates only linear association. For a strong U-shaped non-linear relationship, rr may equal zero because positive and negative slopes cancel out.
  • Bivariate Normality: Variables should approximate normal distributions.
  • Outlier Vulnerability: A single wild outlier can artificially inflate or suppress rr.

Critical Caveat: Correlation vs. Causation

The cardinal rule of bivariate analysis is that correlation does not equal causation. Calculating a high correlation coefficient (r=0.92r = 0.92) proves that two metrics move together mathematically, but does not prove that changing XX causes YY to change.

Misinterpreting correlation as causation frequently leads to flawed process changes due to:

1. Confounding and Lurking Variables

A lurking variable (ZZ) is an unmeasured third factor that simultaneously influences both XX and YY.

Example: A manufacturing plant finds a strong positive correlation (r=0.85r = 0.85) between monthly electricity usage (XX) and total product defects (YY). Cutting power will not eliminate defects: the lurking variable is monthly production volume (ZZ). High production volume increases electricity consumption and naturally yields more total defects.

2. Reverse Causality and Coincidence

Causality may flow in reverse (YY causes XX), or two unrelated metrics may correlate purely by chance or shared historical trends (spurious correlation).

Establishing True Causation

To prove causality in the Analyze phase, teams must:

  1. Validate Physical Mechanisms: Verify that a known engineering, physical, or operational law explains the relationship.
  2. Conduct Multi-Vari Studies: Stratify data across shifts, machines, and lots to verify the relationship persists across subgroups.
  3. Execute Controlled Experiments (DOE): Intentionally manipulate input XX under controlled conditions while holding all other factors constant to observe the response in YY.
Loading diagram...
Spurious Correlation vs. Verified Causation
Test Your Knowledge

When examining a scatter plot of furnace temperature (X) versus component tensile strength (Y), the plotted points form a distinct inverted U-shaped parabolic curve. The calculated Pearson correlation coefficient is r = 0.04. How should a Six Sigma Yellow Belt interpret this result?

A

A strong non-linear relationship exists between temperature and strength, but Pearson r cannot detect it because r measures only linear association.

B

There is no functional relationship between furnace temperature and component tensile strength because r is close to zero.

C

The furnace temperature is perfectly optimized across all ranges, indicating maximum process capability.

D

An arithmetic error must have occurred during computation, because any curved pattern always produces a negative correlation coefficient.

Test Your Knowledge

A project team analyzing customer service operations finds a strong positive correlation (r = 0.88) between the number of customer support representatives working a shift and the total count of customer complaints logged. What is the most appropriate statistical conclusion?

A

Hiring additional customer support representatives directly causes an increase in customer dissatisfaction and service errors.

B

The department manager should immediately reduce staffing levels to minimize the number of customer complaints received.

C

The calculated correlation coefficient is statistically invalid because Pearson r cannot exceed 0.50 in service environments.

D

The variables exhibit strong linear co-variation, but this does not prove causation; an unmeasured lurking variable, such as total customer transaction volume, may drive both.

Test Your Knowledge

Which of the following Pearson correlation coefficients indicates the strongest linear association between cutting speed and machined surface roughness?

A

+0.72

B

-0.89

C

+0.45

D

-0.15

Sections you finish are checked off in the contents.