10.2 Simple Linear Regression & Prediction
Key Takeaways
Simple linear regression mathematically models the functional relationship between a continuous predictor (X) and a continuous response (Y) via the equation Y_hat = b0 + b1*X.
The slope (b1) quantifies the estimated change in response Y for every one-unit increase in predictor X, while the Y-intercept (b0) represents the baseline value of Y_hat when X = 0.
The method of Ordinary Least Squares (OLS) calculates the regression line by minimizing the sum of squared vertical residuals between observed points and fitted values.
The coefficient of determination (R-squared) reflects the proportion of total variation in the response variable explained by the regression model, with R-squared = r^2 in simple linear regression.
While interpolation within the observed data range provides reliable predictions, extrapolating beyond the experimental boundaries introduces severe operational risk.
Simple Linear Regression & Prediction
Quick Answer: Simple linear regression develops a predictive mathematical equation——quantifying how a continuous dependent response () changes when an independent predictor () is varied. Using Ordinary Least Squares (OLS), regression calculates the unique straight line that minimizes the sum of squared vertical residuals between data points and the line. The coefficient of determination () reflects the percentage of total variation explained by the model. Regression provides reliable predictions within the experimental data range, but extrapolating beyond observed boundaries creates severe operational risk. Independent CSSYB study guide by OpenExamPrep.
Purpose of Simple Linear Regression
Once scatter plots and correlation confirm a linear association between process variables, Six Sigma teams must quantify that link: How much does output change per unit change in input ? What setting of achieves our target ?
While correlation measures the degree of association, simple linear regression constructs a functional mathematical transfer function (). It establishes an explicit relationship between:
- One Independent Predictor (): The controllable process setting, raw material parameter, or operational condition.
- One Dependent Response (): The output quality characteristic, dimensional measurement, or cycle time metric targeted for improvement.
This predictive model enables teams to evaluate process sensitivity, establish operating tolerances, and design data-driven solutions.
The Linear Regression Model
Regression differentiates between the true population relationship and the sample model estimated from empirical data.
Theoretical Population Model
The theoretical relationship across the entire process population is expressed as:
Where is the population intercept, is the population slope, and represents independent, normally distributed random error ().
Fitted Sample Regression Equation
Using sample data, practitioners compute the fitted regression line:
Where:
- ("Y-hat"): The predicted mean response for a specified value of .
- (Slope): The estimated change in response for each 1-unit increase in predictor . The slope represents the rate of change (). A positive slope indicates increases as increases, whereas a negative slope indicates an inverse relationship.
- (-Intercept): The estimated value of when . If is outside operating reality (such as machine speed at 0 RPM during production), serves purely as a coordinate anchor without physical meaning.
The Method of Ordinary Least Squares (OLS)
To select the single best-fitting line among infinite candidates, statisticians use the Method of Ordinary Least Squares (OLS).
Residuals and Minimization
For every observation , the vertical distance between the actual observed value () and the fitted value () is the residual ():
The sum of raw residuals always equals zero because positive and negative deviations cancel. OLS eliminates cancellation and penalizes large discrepancies by minimizing the Sum of Squared Errors (SSE):
Minimizing SSE produces the closed-form OLS formulas:
The fitted OLS line always passes through the point of averages .
Coefficient of Determination ()
To assess model adequacy, Six Sigma practitioners evaluate the Coefficient of Determination ().
Regression partitions the total variation in the response variable into two additive components:
Where SST measures total variation around , SSR measures variation explained by the regression line, and SSE measures unexplained residual noise. The coefficient of determination is:
Key Properties and Interpretation
- Range: (or to ).
- Relation to Pearson : In simple linear regression, . If , then .
- Practical Interpretation: An of means that of the total variation in response is explained by the linear relationship with predictor . The remaining represents unexplained common-cause variation or unmeasured variables.
Prediction: Interpolation vs. Extrapolation Risks
A validated regression equation allows practitioners to predict response values, but Yellow Belts must recognize the danger of extrapolation.
- Interpolation (Statistically Valid): Predicting for predictor values within the observed data range (). Interpolated predictions are statistically valid and supported by empirical evidence.
- Extrapolation (Severe Risk): Predicting for predictor values outside the observed data range ( or ).
Extrapolating assumes the linear relationship continues indefinitely. In reality, physical processes experience saturation, material degradation, or structural failure outside tested operating ranges, yielding misleading or impossible predictions.
Step-by-Step Worked Numerical Example
A Six Sigma team in a medical device facility evaluates sealing bar pressure in psi () versus package seal strength in Newtons ().
1. Data and Summary Values
Five test samples yield:
- ,
- Experimental range:
2. Computing Slope and Intercept
The fitted prediction model is:
Interpretation: Each additional psi of sealing pressure increases predicted seal strength by .
3. Model Evaluation and Prediction
Software output for the same five samples reports , so ( of seal strength variation is explained by pressure).
- Valid Interpolation: At (within the 30–70 psi range):
- Extrapolation Danger: At , the equation predicts . In production, 200 psi would crush the blister pack, causing catastrophic seal failure.
A Six Sigma improvement team derives the following simple linear regression model to predict bearing operating temperature in Celsius (Y) based on machine shaft rotational speed in RPM (X): Y_hat = 22.5 + 0.045*X. If the machine is operated at 1,200 RPM, what is the predicted bearing temperature?
54.0°C
67.5°C
76.5°C
99.0°C
In an industrial regression study evaluating the impact of protective coating thickness (X) on metal corrosion resistance score (Y), the coefficient of determination is calculated as R^2 = 0.74. How should a Yellow Belt interpret this metric?
74% of the total variation observed in corrosion resistance scores is explained by the linear relationship with coating thickness, while 26% is unexplained variation.
The Pearson correlation coefficient between coating thickness and corrosion resistance equals 0.74.
For every additional micron of coating thickness applied, the corrosion resistance score increases by 0.74 units.
The probability of committing a Type I error during regression slope hypothesis testing is 26%.
A chemical plant models production yield (Y, in %) based on reactor operating temperature (X, in °C) across an experimental test range of 100°C to 160°C, yielding the equation Y_hat = 15 + 0.5*X. An engineer uses this equation to predict reactor yield at 350°C and calculates an expected yield of 190%. Why is this prediction statistically and operationally flawed?
The regression slope coefficient (b1 = 0.5) is too small to permit mathematical calculations beyond 200°C.
The regression model cannot be used because the Y-intercept (b0 = 15) is a positive number.
The Pearson correlation coefficient must be recalculated using non-parametric methods before calculating yield.
The prediction is an extrapolation far beyond the observed experimental range of 100°C to 160°C, where physical behavior and reaction dynamics may alter dramatically.
Sections you finish are checked off in the contents.