10.2 Simple Linear Regression & Prediction

Key Takeaways

  • Simple linear regression mathematically models the functional relationship between a continuous predictor (X) and a continuous response (Y) via the equation Y_hat = b0 + b1*X.

  • The slope (b1) quantifies the estimated change in response Y for every one-unit increase in predictor X, while the Y-intercept (b0) represents the baseline value of Y_hat when X = 0.

  • The method of Ordinary Least Squares (OLS) calculates the regression line by minimizing the sum of squared vertical residuals between observed points and fitted values.

  • The coefficient of determination (R-squared) reflects the proportion of total variation in the response variable explained by the regression model, with R-squared = r^2 in simple linear regression.

  • While interpolation within the observed data range provides reliable predictions, extrapolating beyond the experimental boundaries introduces severe operational risk.

Last updated: September 2026

Simple Linear Regression & Prediction

Quick Answer: Simple linear regression develops a predictive mathematical equation—Y^=b0+b1X\hat{Y} = b_0 + b_1 X—quantifying how a continuous dependent response (YY) changes when an independent predictor (XX) is varied. Using Ordinary Least Squares (OLS), regression calculates the unique straight line that minimizes the sum of squared vertical residuals between data points and the line. The coefficient of determination (R2R^2) reflects the percentage of total variation explained by the model. Regression provides reliable predictions within the experimental data range, but extrapolating beyond observed boundaries creates severe operational risk. Independent CSSYB study guide by OpenExamPrep.

Purpose of Simple Linear Regression

Once scatter plots and correlation confirm a linear association between process variables, Six Sigma teams must quantify that link: How much does output YY change per unit change in input XX? What setting of XX achieves our target YY?

While correlation measures the degree of association, simple linear regression constructs a functional mathematical transfer function (Y=f(X)Y = f(X)). It establishes an explicit relationship between:

  • One Independent Predictor (XX): The controllable process setting, raw material parameter, or operational condition.
  • One Dependent Response (YY): The output quality characteristic, dimensional measurement, or cycle time metric targeted for improvement.

This predictive model enables teams to evaluate process sensitivity, establish operating tolerances, and design data-driven solutions.


The Linear Regression Model

Regression differentiates between the true population relationship and the sample model estimated from empirical data.

Theoretical Population Model

The theoretical relationship across the entire process population is expressed as:

Y=β0+β1X+ϵY = \beta_0 + \beta_1 X + \epsilon

Where β0\beta_0 is the population intercept, β1\beta_1 is the population slope, and ϵ\epsilon represents independent, normally distributed random error (ϵ∼N(0,σ2)\epsilon \sim N(0, \sigma^2)).

Fitted Sample Regression Equation

Using sample data, practitioners compute the fitted regression line:

Y^=b0+b1X\hat{Y} = b_0 + b_1 X

Where:

  • Y^\hat{Y} ("Y-hat"): The predicted mean response for a specified value of XX.
  • b1b_1 (Slope): The estimated change in response YY for each 1-unit increase in predictor XX. The slope represents the rate of change (ΔY/ΔX\Delta Y / \Delta X). A positive slope indicates YY increases as XX increases, whereas a negative slope indicates an inverse relationship.
  • b0b_0 (YY-Intercept): The estimated value of Y^\hat{Y} when X=0X = 0. If X=0X = 0 is outside operating reality (such as machine speed at 0 RPM during production), b0b_0 serves purely as a coordinate anchor without physical meaning.

The Method of Ordinary Least Squares (OLS)

To select the single best-fitting line among infinite candidates, statisticians use the Method of Ordinary Least Squares (OLS).

Residuals and Minimization

For every observation (Xi,Yi)(X_i, Y_i), the vertical distance between the actual observed value (YiY_i) and the fitted value (Y^i\hat{Y}_i) is the residual (eie_i):

ei=Yi−Y^i=Yi−(b0+b1Xi)e_i = Y_i - \hat{Y}_i = Y_i - (b_0 + b_1 X_i)

The sum of raw residuals always equals zero because positive and negative deviations cancel. OLS eliminates cancellation and penalizes large discrepancies by minimizing the Sum of Squared Errors (SSE):

SSE=∑i=1nei2=∑i=1n(Yi−Y^i)2\text{SSE} = \sum_{i=1}^{n} e_i^2 = \sum_{i=1}^{n} (Y_i - \hat{Y}_i)^2

Minimizing SSE produces the closed-form OLS formulas:

b1=∑(Xi−Xˉ)(Yi−Yˉ)∑(Xi−Xˉ)2=r(sYsX)b_1 = \frac{\sum (X_i - \bar{X})(Y_i - \bar{Y})}{\sum (X_i - \bar{X})^2} = r \left( \frac{s_Y}{s_X} \right) b0=Yˉ−b1Xˉb_0 = \bar{Y} - b_1 \bar{X}

The fitted OLS line always passes through the point of averages (Xˉ,Yˉ)(\bar{X}, \bar{Y}).


Coefficient of Determination (R2R^2)

To assess model adequacy, Six Sigma practitioners evaluate the Coefficient of Determination (R2R^2).

Regression partitions the total variation in the response variable into two additive components:

Total Sum of Squares (SST)=Regression Sum of Squares (SSR)+Error Sum of Squares (SSE)\text{Total Sum of Squares (SST)} = \text{Regression Sum of Squares (SSR)} + \text{Error Sum of Squares (SSE)}

Where SST measures total variation around Yˉ\bar{Y}, SSR measures variation explained by the regression line, and SSE measures unexplained residual noise. The coefficient of determination is:

R2=SSRSST=1−SSESSTR^2 = \frac{\text{SSR}}{\text{SST}} = 1 - \frac{\text{SSE}}{\text{SST}}

Key Properties and Interpretation

  • Range: 0.0≤R2≤1.00.0 \le R^2 \le 1.0 (or 0%0\% to 100%100\%).
  • Relation to Pearson rr: In simple linear regression, R2=(r)2R^2 = (r)^2. If r=0.90r = 0.90, then R2=0.81R^2 = 0.81.
  • Practical Interpretation: An R2R^2 of 0.850.85 means that 85%85\% of the total variation in response YY is explained by the linear relationship with predictor XX. The remaining 15%15\% represents unexplained common-cause variation or unmeasured variables.

Prediction: Interpolation vs. Extrapolation Risks

A validated regression equation allows practitioners to predict response values, but Yellow Belts must recognize the danger of extrapolation.

  • Interpolation (Statistically Valid): Predicting Y^\hat{Y} for predictor values within the observed data range (Xmin⁡≤X≤Xmax⁡X_{\min} \le X \le X_{\max}). Interpolated predictions are statistically valid and supported by empirical evidence.
  • Extrapolation (Severe Risk): Predicting Y^\hat{Y} for predictor values outside the observed data range (X<Xmin⁡X < X_{\min} or X>Xmax⁡X > X_{\max}).

Extrapolating assumes the linear relationship continues indefinitely. In reality, physical processes experience saturation, material degradation, or structural failure outside tested operating ranges, yielding misleading or impossible predictions.


Step-by-Step Worked Numerical Example

A Six Sigma team in a medical device facility evaluates sealing bar pressure in psi (XX) versus package seal strength in Newtons (YY).

1. Data and Summary Values

Five test samples yield:

  • Xˉ=50 psi\bar{X} = 50\text{ psi}, Yˉ=80 N\bar{Y} = 80\text{ N}
  • ∑(Xi−Xˉ)2=1,000\sum (X_i - \bar{X})^2 = 1,000
  • ∑(Xi−Xˉ)(Yi−Yˉ)=1,550\sum (X_i - \bar{X})(Y_i - \bar{Y}) = 1,550
  • Experimental range: 30 psi≤X≤70 psi30\text{ psi} \le X \le 70\text{ psi}

2. Computing Slope and Intercept

b1=1,5501,000=1.55 N/psib_1 = \frac{1,550}{1,000} = 1.55\text{ N/psi} b0=80−(1.55)(50)=80−77.5=2.5 Nb_0 = 80 - (1.55)(50) = 80 - 77.5 = 2.5\text{ N}

The fitted prediction model is:

Y^=2.5+1.55X\hat{Y} = 2.5 + 1.55 X

Interpretation: Each additional psi of sealing pressure increases predicted seal strength by 1.55 N1.55\text{ N}.

3. Model Evaluation and Prediction

Software output for the same five samples reports r=0.992r = 0.992, so R2=(0.992)2=0.984R^2 = (0.992)^2 = 0.984 (98.4%98.4\% of seal strength variation is explained by pressure).

  • Valid Interpolation: At X=55 psiX = 55\text{ psi} (within the 30–70 psi range):
Y^=2.5+1.55(55)=87.75 N\hat{Y} = 2.5 + 1.55(55) = 87.75\text{ N}
  • Extrapolation Danger: At X=200 psiX = 200\text{ psi}, the equation predicts Y^=312.5 N\hat{Y} = 312.5\text{ N}. In production, 200 psi would crush the blister pack, causing catastrophic seal failure.
Loading diagram...
Partitioning Variation in Linear Regression
Test Your Knowledge

A Six Sigma improvement team derives the following simple linear regression model to predict bearing operating temperature in Celsius (Y) based on machine shaft rotational speed in RPM (X): Y_hat = 22.5 + 0.045*X. If the machine is operated at 1,200 RPM, what is the predicted bearing temperature?

A

54.0°C

B

67.5°C

C

76.5°C

D

99.0°C

Test Your Knowledge

In an industrial regression study evaluating the impact of protective coating thickness (X) on metal corrosion resistance score (Y), the coefficient of determination is calculated as R^2 = 0.74. How should a Yellow Belt interpret this metric?

A

74% of the total variation observed in corrosion resistance scores is explained by the linear relationship with coating thickness, while 26% is unexplained variation.

B

The Pearson correlation coefficient between coating thickness and corrosion resistance equals 0.74.

C

For every additional micron of coating thickness applied, the corrosion resistance score increases by 0.74 units.

D

The probability of committing a Type I error during regression slope hypothesis testing is 26%.

Test Your Knowledge

A chemical plant models production yield (Y, in %) based on reactor operating temperature (X, in °C) across an experimental test range of 100°C to 160°C, yielding the equation Y_hat = 15 + 0.5*X. An engineer uses this equation to predict reactor yield at 350°C and calculates an expected yield of 190%. Why is this prediction statistically and operationally flawed?

A

The regression slope coefficient (b1 = 0.5) is too small to permit mathematical calculations beyond 200°C.

B

The regression model cannot be used because the Y-intercept (b0 = 15) is a positive number.

C

The Pearson correlation coefficient must be recalculated using non-parametric methods before calculating yield.

D

The prediction is an extrapolation far beyond the observed experimental range of 100°C to 160°C, where physical behavior and reaction dynamics may alter dramatically.

Sections you finish are checked off in the contents.