9.3 Measurement System Analysis: Bias, Linearity, and Stability
Key Takeaways
- Measurement System Analysis (MSA) partitions observed total process variation into actual part-to-part variation and measurement system error (sigma^2_total = sigma^2_part + sigma^2_gage), ensuring quality decisions are not compromised by gage error.
- Accuracy (trueness) encompasses the location components of measurement error—bias, linearity, and stability—while precision encompasses the spread components (repeatability and reproducibility).
- Bias represents the systematic difference between the observed measurement average and a certified reference master value, evaluated for statistical significance using a one-sample t-test at alpha = 0.05.
- Linearity measures how bias changes across the operating range of the gage, quantified by the slope (b) of the best-fit linear regression line of bias versus reference standard values.
- Stability tracks measurement system drift over calendar time by charting periodic measurements of a reference standard on an SPC control chart (X-bar & R or I-MR), where out-of-control signals detect calibration decay.
9.3 Measurement System Analysis: Bias, Linearity, and Stability
The Purpose of Measurement System Analysis (MSA) and AIAG Standards
In statistical quality control, data drives every critical decision: accepting or rejecting lots, computing capability indices ($C_p, C_{pk}$), tuning machine offsets, and declaring process stability. However, the numerical value registered by an inspector or automated gage is rarely the exact true value of the part. Every measurement contains an inherent degree of measurement error.
If measurement error is large relative to part tolerance or process variation, the manufacturing team operates under a false reality:
- Producer's Risk ($\alpha$ Error / False Alarm): Conforming parts near specification limits are misclassified as nonconforming due to gage error, leading to unnecessary scrapping, reworking, and downtime.
- Consumer's Risk ($\beta$ Error / Miss): Nonconforming parts are erroneously accepted and shipped to customers because gage inaccuracies mask out-of-spec dimensions.
To quantify and control this error, the Automotive Industry Action Group (AIAG) published the widely adopted Measurement System Analysis (MSA) Reference Manual (currently in its 4th Edition). Under MSA principles, the observed total process variation ($\sigma^2_{\text{Total}}$) is the mathematical sum of actual part-to-part manufacturing variation ($\sigma^2_{\text{Part}}$) and measurement system variation ($\sigma^2_{\text{MS}}$):
MSA evaluates whether the measurement system is capable of detecting process shifts and sorting good product from bad product before the data is used in quality decisions.
Accuracy (Trueness) versus Precision (Variation)
A cornerstone of the ASQ CQT body of knowledge is the rigorous distinction between accuracy (trueness) and precision (variation):
ACCURACY (LOCATION) VS. PRECISION (SPREAD):
Target 1: Accurate & Precise Target 2: Precise but BIASED (Inaccurate)
+-------+ +-------+
| * | | |
| *** | (Tight cluster on bullseye) | | *** (Tight cluster, offset)
| * | | | *** from bullseye!)
+-------+ +-------+
Target 3: Accurate but IMPRECISE Target 4: NEITHER Accurate NOR Precise
+-------+ +-------+
| * * | | * |
| * | (Centered average, but | * | (Scattered wide, off-center)
| * * | wide scattered spread!) | * |
+-------+ +-------+
- Accuracy (Trueness / Location Error): The closeness of agreement between the average of an infinite number of observed measurement results and a certified reference master value. Accuracy tells you where the measurement distribution is centered.
- The three location components of measurement error are Bias, Linearity, and Stability.
- Precision (Variation / Spread Error): The closeness of agreement between independent measurement results obtained under stipulated experimental conditions. Precision reflects the spread or dispersion of the measurement distribution, completely independent of whether the measurements are near the true target.
- The two precision components of measurement error are Repeatability and Reproducibility (detailed in Section 9.4).
| Measurement Attribute | Error Category | Core Question Answered | Error Components Evaluated |
|---|---|---|---|
| Accuracy (Trueness) | Location Error | Is the measurement centered on the true standard value? | Bias, Linearity, Stability |
| Precision (Variation) | Spread Error | Do repeated measurements cluster tightly together? | Repeatability (EV), Reproducibility (AV) |
Bias: Systematic Measurement Error and Statistical Significance
Bias is the systematic error or difference between the observed average of repeated measurements and the accepted, certified reference master value of the measured characteristic. Where $\bar{X}$ is the arithmetic mean of repeated measurements, and $X_{\text{ref}}$ is the certified reference master value (established by a calibration laboratory using NIST-traceable gage blocks or primary standards).
Bias Study Procedure (AIAG MSA Protocol)
- Sample Selection: Obtain one representative production part or standard whose critical characteristic has been independently measured in a controlled calibration laboratory to establish the reference master value ($X_{\text{ref}}$).
- Execution: Have a single appraiser measure the master part $n \ge 10$ to 15 times (conventionally $n = 15$ or $n = 20$) under normal shop-floor operating conditions. The trials must be conducted in random order or separated by intervening parts to prevent operator memory bias.
- Calculate Sample Mean ($\bar{X}$):
- Calculate Sample Standard Deviation ($s$):
- Calculate Bias:
- Calculate Percentage Bias:
Testing for Statistical Significance: The One-Sample $t$-Test
A critical question for quality technicians is: Is the calculated bias a true systematic error, or is it merely random sampling noise?
To answer this, MSA conducts a two-tailed one-sample $t$-test:
- Null Hypothesis ($H_0$): $\text{Bias} = 0$ (the measurement system has zero true bias, $\mu = X_{\text{ref}}$).
- Alternative Hypothesis ($H_1$): $\text{Bias} \ne 0$ (a statistically significant non-zero bias exists).
The $t$-Statistic Formula:
Where $s_{\bar{X}} = \frac{s}{\sqrt{n}}$ is the standard error of the mean, and the degrees of freedom is $df = n - 1$.
Decision Rule (at $\alpha = 0.05$):
- Look up the critical value $t_{\alpha/2, n-1}$ in a Student's $t$-distribution table (for example, for $n = 15, df = 14$, two-tailed $t_{0.025, 14} = 2.145$).
- If $|t_{\text{calc}}| > t_{\text{crit}}$ (or if the $p$-value $< 0.05$), reject the null hypothesis. The bias is statistically significant!
- Alternatively, compute the $95%$ Confidence Interval for Bias: If the confidence interval does not contain zero, the bias is statistically significant. The gage must be physically recalibrated, mechanically adjusted, or programmed with a software offset compensation factor.
Linearity: Evaluating Bias Across the Operating Range
A gage may exhibit zero bias when measuring a $1.000"$ standard, but develop an unacceptable $+0.003"$ bias when measuring a $5.000"$ part. This phenomenon is evaluated through Linearity.
[!IMPORTANT] Linearity: The difference in bias values across the expected operating range of the gage. It measures whether the instrument's accuracy remains constant or degrades as part sizes increase or decrease.
LINEARITY REGRESSION PLOT (BIAS VS. REFERENCE VALUE):
Bias (Y-axis)
^
| * (Part 5: Large positive bias)
| * (Part 4)
| * (Part 3: Bias ~ 0)
0 -+--------------*-------------------------------------> Reference Master Size (X-axis)
| * (Part 2) Slope b != 0 indicates Linearity Error!
| (Part 1: Negative bias)
|
Common Physical Causes of Linearity Error
- Lead Screw Pitch Wear: Progressive lead error in the precision screw of a micrometer or height gage.
- Optical / Capacitive Scale Distortion: Non-linear spacing or warping along the glass encoder scale of a digital caliper or CMM.
- Linkage Geometry Deflection: Wear or non-proportional lever movement in dial indicator gear trains.
- Electronic Amplifier Gain Error: Non-linear voltage amplification in LVDT electronic bore gages.
Linearity Study Protocol
- Select Reference Parts: Choose $g \ge 5$ parts (conventionally 5 parts) whose dimensions span the entire operating range of the gage (e.g., $1.000", 2.000", 3.000", 4.000", 5.000"$).
- Establish Reference Values: Measure each part in a certified calibration laboratory to determine exact master values: $X_{\text{ref}, 1}, X_{\text{ref}, 2}, \dots, X_{\text{ref}, 5}$.
- Conduct Measurements: Have one technician measure each part $m \ge 10$ times in completely randomized order.
- Calculate Average Bias per Part:
- Perform Simple Linear Regression: Fit a straight line modeling Bias ($Y$) as a function of the Reference Value ($X$): Where $b$ is the regression slope, and $a$ is the $Y$-intercept.
Calculating Regression Parameters:
Evaluating Linearity Metrics:
- Slope ($b$): Reflects the rate of change of bias per unit increase in measured length.
- Percentage Linearity:
- Coefficient of Determination ($R^2$): Measures the proportion of bias variation explained by the linear relationship with part size.
- Statistical Significance: A $t$-test is conducted on the slope ($H_0: b = 0$). If $p < 0.05$, the linearity error is statistically significant, indicating that the gage cannot be corrected by a single constant offset and requires mechanical reconditioning or multi-point calibration mapping.
Stability: Monitoring Measurement Drift Over Calendar Time
A measurement system may demonstrate excellent accuracy and precision during initial qualification. However, over weeks and months of shop-floor operation, environmental temperature swings, component aging, dirt accumulation, and mechanical wear cause the system to drift. This time-dependent variation is evaluated through Stability.
[!IMPORTANT] Stability (Consistency): The total variation in the measurements obtained with a measurement system on the same master standard or parts when measuring a single characteristic over an extended period of calendar time.
STABILITY CONTROL CHART (I-MR CHART ON REFERENCE MASTER):
Master Reading
^
| UCL = 2.0006" - - - - - - - - - - - - - - - - - - - - - - - - - -
| * * * (DRIFT!)
| CL = 2.0000" -------------*-------*-------*----*------------- (Shift detected)
| * *
| LCL = 1.9994" - - - - - - - - - - - - - - - - - - - - - - - - - -
+-------------------------------------------------------------------> Calendar Time
Week 1 Week 2 Week 3 Week 4
Stability Study Protocol
- Designate a Reference Standard: Dedicate a certified master part or gage block stack specifically for stability monitoring.
- Establish Sampling Frequency: Measure the reference standard periodically throughout production (e.g., once or twice per shift, daily, or weekly).
- Plot on SPC Control Charts:
- If subgroups of repeated measurements ($n = 3$ to 5) are taken: plot on an $\bar{X}-R$ control chart.
- If a single reading is taken at each time interval: plot on an Individuals and Moving Range ($I-MR$) control chart.
- Evaluate Out-of-Control Rules:
- Points beyond control limits ($\pm 3\sigma$): Signals an abrupt assignable cause (e.g., chipped contact stylus, dropped gage, cracked optical mirror).
- Trends (6 or 7 consecutive points continuously increasing or decreasing): Signals progressive wear, calibration drift, or gradual electronic degradation.
- Runs (8 consecutive points on one side of the centerline): Signals an uncorrected shift in the baseline zero datum or thermal normalization failure.
Shop-Floor Environmental Control and Drift Prevention
To preserve measurement stability, quality technicians must enforce strict shop-floor metrological controls:
- The Standard Reference Temperature (ISO 1): International standard reference temperature for dimensional metrology is $20^\circ\text{C}$ ($68^\circ\text{F}$). Inspection of precision parts with tight tolerances must occur in a temperature-controlled environment.
- Thermal Expansion Differential: Different materials expand at different rates: Where $\alpha$ is the coefficient of thermal expansion (aluminum $\approx 13 \times 10^{-6}/^\circ\text{F}$, tool steel $\approx 6.5 \times 10^{-6}/^\circ\text{F}$, granite $\approx 3.5 \times 10^{-6}/^\circ\text{F}$). If an aluminum aerospace housing is machined at $85^\circ\text{F}$ and inspected against steel gage blocks at $68^\circ\text{F}$, thermal shrinkage will cause catastrophic rejection unless temperature compensation is applied!
- Thermal Soaking Time: Workpieces and gage blocks brought from an unconditioned warehouse must be allowed to normalize (soak) in the inspection room—typically 1 hour per inch of cross-sectional thickness.
- Cleanliness and Air Quality: Air gages rely on clean, dried, oil-free compressed air. Oil mist or desiccant breakdown will clog pneumatic orifices, creating severe stability drift.
Step-by-Step Worked Numerical Examples
Worked Example 1: Full Bias Study with One-Sample $t$-Test
Scenario: A quality technician conducts a bias study on a digital micrometer used to inspect ground valve stems. A certified reference master gage block certified at $X_{\text{ref}} = 0.50000"$ is measured $n = 15$ times by one technician under production conditions. The valve stem blueprint tolerance is $0.5000" \pm 0.0020"$ (Total Tolerance $= 0.0040"$).
The 15 observed readings are:
$0.5004", 0.5003", 0.5005", 0.5002", 0.5004", 0.5006", 0.5003", 0.5004", 0.5005", 0.5003", 0.5004", 0.5002", 0.5005", 0.5004", 0.5006"$.
Step 1: Calculate Sample Mean ($\bar{X}$)
Step 2: Calculate Sample Standard Deviation ($s$)
Step 3: Calculate Bias and Percentage Bias
Step 4: Perform One-Sample $t$-Test for Statistical Significance
- Degrees of freedom: $df = n - 1 = 15 - 1 = 14$.
- Standard error of the mean:
- Compute $t_{\text{calc}}$:
- Look up critical $t$ at $\alpha = 0.05$ (two-tailed, $df = 14$): $t_{0.025, 14} = 2.145$.
Step 5: Interpretation and Confidence Interval
Since $|t_{\text{calc}}| = 12.96 > 2.145$, the bias is statistically significant ($p < 0.0001$).
$95%$ Confidence Interval for Bias:
Because the confidence interval does not contain zero, the micrometer exhibits an unadjusted systematic positive offset of $+0.0004"$. The technician must zero-reset the micrometer barrel before releasing it for production inspection.
Worked Example 2: Linearity Study Regression
Scenario: A linearity study is conducted on an electronic height gage across five calibrated reference standards spanning $1.000"$ to $5.000"$. The tolerance for parts measured by this gage is $0.0200"$. The average bias values obtained across 10 repeated trials per part are:
- Part 1 ($X_1 = 1.000"$): $\bar{B}_1 = +0.0002"$
- Part 2 ($X_2 = 2.000"$): $\bar{B}_2 = +0.0005"$
- Part 3 ($X_3 = 3.000"$): $\bar{B}_3 = +0.0008"$
- Part 4 ($X_4 = 4.000"$): $\bar{B}_4 = +0.0011"$
- Part 5 ($X_5 = 5.000"$): $\bar{B}_5 = +0.0014"$
Step 1: Calculate Summary Statistics
Step 2: Compute Regression Slope ($b$)
| Part $i$ | $X_i - \bar{X}$ | $(X_i - \bar{X})^2$ | $B_i - \bar{B}$ | $(X_i - \bar{X})(B_i - \bar{B})$ |
|---|---|---|---|---|
| 1 | $-2.0$ | $4.0$ | $-0.0006$ | $+0.0012$ |
| 2 | $-1.0$ | $1.0$ | $-0.0003$ | $+0.0003$ |
| 3 | $0.0$ | $0.0$ | $0.0000$ | $0.0000$ |
| 4 | $+1.0$ | $1.0$ | $+0.0003$ | $+0.0003$ |
| 5 | $+2.0$ | $4.0$ | $+0.0006$ | $+0.0012$ |
| Sum | $10.0$ | $+0.0030$ |
Step 3: Compute Intercept ($a$)
Step 4: Calculate Percentage Linearity
- Metrological Interpretation: The gage exhibits a progressive positive lead error of $+0.00030"$ per inch of travel. Because $R^2 = 1.00$ and the slope is non-zero, this linearity error indicates mechanical pitch wear along the vertical column rack, which cannot be cured by a simple zero adjustment.
Technician Inspection Scenarios & Common Exam Traps
Real-World Shop Scenario: Bore Gage Calibration Drift
A machining cell operator uses an electronic dial bore gage to inspect cylinder liners. At 7:00 AM, the gage is mastered inside the air-conditioned quality lab ($68^\circ\text{F}$) using a master ring gage. By 2:00 PM, the shop-floor ambient temperature reaches $92^\circ\text{F}$, and the bore gage indicates that cylinder liners are running undersize, prompting the operator to adjust the boring bar tool offset. At 4:00 PM, the quality auditor reinspects the parts in the lab and finds them oversized! Why did this happen?
- Root Cause Analysis: The technician mastered the bore gage at $68^\circ\text{F}$, but used it at $92^\circ\text{F}$. The aluminum gage extension rod expanded significantly in the shop heat, shifting the mechanical zero point (stability drift). Furthermore, the warm cylinder liners expanded during machining and contracted upon cooling in the lab. The technician failed to re-zero the gage on the shop floor and neglected thermal soaking rules.
Common Exam Traps for CQT Candidates
- Exam Trap 1: Confusing Bias with Repeatability: Bias is systematic error (distance from target center). Repeatability is random error (cluster width). Adjusting the zero screw fixes bias, but will never fix poor repeatability!
- Exam Trap 2: Misinterpreting the $t$-Test Result: If calculated $t$ is less than critical $t$ ($p > 0.05$), bias is statistically zero (inconsequential random noise). Do not attempt to adjust or calibrate a gage when the bias is statistically insignificant.
- Exam Trap 3: Linearity vs. Calibration Offset: If a gage has a constant bias of $+0.001"$ at $1"$, $2"$, and $3"$, the slope is $b = 0$. This gage has zero linearity error, but high bias. Linearity requires the bias to change across the measurement range.
- Exam Trap 4: Stability Sample Type: A stability study must always measure the exact same reference standard or master part. If an inspector measures different production parts over time, the chart reflects part manufacturing variation, not measurement system stability.
A quality technician performs a bias study on an electronic bore gage using a calibrated master setting ring certified at 2.00000 inches. The tolerance for the bore is ±0.0050 inches. Across 16 repeated measurements, the technician calculates a sample mean of 2.00045 inches and a sample standard deviation of s = 0.00040 inches. What is the calculated bias, what is the t-statistic, and is the bias statistically significant at alpha = 0.05 (critical t = 2.131)?
A linearity study is conducted on an outside micrometer across 5 gage block standards spanning its 0 to 4-inch operating range. The best-fit linear regression of average bias versus reference value yields: Bias = 0.0004 * (Reference) - 0.0002 inches, with an R-squared of 0.94 and a p-value for the slope of p = 0.008. How should the quality technician interpret these results?
A quality technician monitors the stability of an optical shaft measurement system by measuring a calibrated chrome master shaft once at the start of every 8-hour shift and plotting the values on an Individuals and Moving Range (I-MR) control chart. After two weeks of in-control operation, 7 consecutive individual measurements plot above the centerline, though none breach the upper control limit. What does this signal indicate, and what action should the technician take?