5.3 Measurement System Analysis (MSA) & Gage R&R

Key Takeaways

  • Total observed process variance is the additive sum of true part manufacturing variance and measurement system variance: σ²(Total) = σ²(Part) + σ²(Measurement System).
  • Repeatability (Equipment Variation, EV) quantifies within-appraiser variation under identical conditions, whereas Reproducibility (Appraiser Variation, AV) quantifies variation between different appraisers.
  • Under AIAG MSA guidelines, a Gage R&R result under 10% is acceptable, 10% to 30% is conditionally acceptable based on process criticality and risk, and over 30% is unacceptable.
  • The Number of Distinct Categories (ndc = 1.41 * PV / GRR) must equal or exceed 5 (ndc ≥ 5) for a measurement system to reliably discern process variation and calculate capability indices (Cp, Cpk).
  • Attribute measurement systems (Go/No-Go gages) are evaluated using cross-tabulation agreement and Cohen's Kappa statistic (κ), where κ ≥ 0.75 indicates good-to-excellent agreement between appraisers and standard truth.
Last updated: September 2026

5.3 Measurement System Analysis (MSA) & Gage R&R

In manufacturing quality control, inspectors frequently assume that their gages report absolute truth. When a digital height gage displays $2.5015\text{ in.}$, the shop floor assumes the part feature is exactly $2.5015\text{ in.}$ In reality, every observed measurement is a composite of two distinct sources of variation: the true variation of the manufacturing process, and the variation introduced by the measurement system itself. If the measurement system exhibits excessive variation, quality inspectors will reject perfectly conforming parts (producer's risk / Type I error / $\alpha$) or, far worse, accept nonconforming defectives (consumer's risk / Type II error / $\beta$). Measurement System Analysis (MSA) is the rigorous statistical discipline that qualifies measuring equipment, fixtures, software, and human operator procedures prior to conducting process capability studies or online statistical process control (SPC). For the ASQ Certified Quality Inspector, mastery of the AIAG MSA 4th Edition guidelines and Gage Repeatability and Reproducibility (Gage R&R) calculation is mandatory.


The Additive Law of Measurement Variance

The fundamental premise of MSA is governed by the additive law of variances:

σTotal Observed2=σTrue Part2+σMeasurement System2\sigma^2_{\text{Total Observed}} = \sigma^2_{\text{True Part}} + \sigma^2_{\text{Measurement System}}

Where:

  • $\sigma^2_{\text{Total Observed}}$ = The overall variance recorded in inspection records.
  • $\sigma^2_{\text{True Part}}$ = The actual manufacturing variation between parts produced by machine tools.
  • $\sigma^2_{\text{Measurement System}}$ (also called $\sigma^2_{MS}$ or $\sigma^2_{GRR}$) = The variance introduced by gages, fixtures, environment, and appraisers.
+-------------------------------------------------------------------------+
|                   TOTAL OBSERVED PROCESS VARIATION                      |
|                                                                         |
|  +---------------------------------+  +------------------------------+  |
|  |      TRUE PART VARIATION        |  |  MEASUREMENT SYSTEM VARIANCE |  |
|  |           (sigma^2_P)           |  |          (sigma^2_MS)        |  |
|  |  * Machine tool spindle runout  |  |  * Gage mechanical slop      |  |
|  |  * Material hardness variation  |  |  * Operator technique & feel |  |
|  |  * Tool wear over time          |  |  * Parallax and alignment    |  |
|  |  * Thermal expansion of stock   |  |  * Temperature gradients     |  |
|  +---------------------------------+  +------------------------------+  |
|                                   |                                     |
|                                   v                                     |
|               [ GAGE R&R = Repeatability + Reproducibility ]            |
+-------------------------------------------------------------------------+

If the measurement system consumes a large percentage of total variance, calculated process capability indices ($C_p$ and $C_{pk}$) become falsely depressed. A world-class manufacturing process capable of $C_p = 2.0$ will appear completely uncapable ($C_p < 1.0$) simply because the measuring tool is noisy and erratic!


The Five Components of Measurement System Error

Under AIAG MSA 4th Edition, measurement error is categorized into location (accuracy) errors and width (precision) errors:

+-------------------------------------------------------------------------+
|                 MEASUREMENT SYSTEM ERROR ARCHITECTURE                   |
|                                                                         |
|                  /                                    \                 |
|       LOCATION ERRORS (Accuracy)             WIDTH ERRORS (Precision)   |
|                 |                                        |              |
|       +---------+---------+                     +--------+--------+     |
|       |         |         |                     |                 |     |
|     Bias    Linearity  Stability          Repeatability    Reproducibility
|  (Offset)  (Over range) (Over time)      (Equipment / EV)  (Appraiser / AV)
+-------------------------------------------------------------------------+

1. Bias (Accuracy / Systematic Offset)

Bias is the difference between the average of multiple repeated measurements taken on a single part using the gage, and the accepted reference master value (calibrated ground truth):

Bias=XˉmeasuredXreference\text{Bias} = \bar{X}_{\text{measured}} - X_{\text{reference}}

  • Root Causes: Incorrect zero setting, worn contact points, uncompensated lead screw pitch error, or improper mastering.

2. Linearity (Consistency Across Operating Range)

Linearity assesses whether the gage's bias remains constant across its entire designated measuring range. A $0\text{ to } 6\text{-inch}$ caliper might have zero bias when measuring a $1.000\text{-inch}$ gage block, but exhibit a $-0.002\text{-inch}$ bias when measuring a $5.000\text{-inch}$ gage block.

  • Evaluation: Measure master standards across the full operational spectrum, plot Bias versus Reference Value, and calculate the linear regression slope: Linearity=Slope×Process Variation (or Tolerance)\text{Linearity} = |\text{Slope}| \times \text{Process Variation (or Tolerance)}

3. Stability (Consistency Over Time / Drift)

Stability (or drift) is the change in bias over extended periods of operation under standard environmental conditions.

  • Evaluation: An inspector measures a single certified master part 3 to 5 times at regular periodic intervals (e.g., once every shift or once per week) and plots the data on an SPC $\bar{X}$ and $R$ control chart. If points stay within statistical control limits with no trends or runs, the measurement system is stable.

4. Repeatability (Equipment Variation / EV)

Repeatability is the variation observed when one appraiser uses the same measuring instrument to measure the identical characteristic on the same physical part multiple times in the same environment.

  • Repeatability represents the internal, short-term capability of the gage hardware (internal friction, electronic sensor noise, structural compliance, mechanical backlash).

5. Reproducibility (Appraiser Variation / AV)

Reproducibility is the variation in the average of measurements made by different appraisers using the same instrument when measuring the identical characteristic on the same physical parts.

  • Reproducibility reflects human interaction: differing operator clamping forces, angle of sight (parallax), finger placement, part seating against datums, and interpretation of graduated lines.

Variable Gage R&R Study Design & Execution

A variable Gage R&R study isolates and quantifies Repeatability (EV) and Reproducibility (AV).

Standard AIAG Study Parameters

  • Number of Parts ($n$): Typically 10 parts. These parts must NOT be identical nominal parts; they must be selected to represent the full range of manufacturing process variation (including parts near the upper and lower specification limits).
  • Number of Appraisers ($k$): Typically 2 to 3 inspectors who normally perform this measurement task on the production line.
  • Number of Trials ($r$): Typically 2 to 3 repeated trials per part per appraiser.
  • Total Data Points: $10\text{ parts} \times 3\text{ appraisers} \times 3\text{ trials} = 90\text{ total readings}$.
  • Crucial Blinding Protocol: Parts must be marked with hidden serial numbers and completely randomized before being handed to the appraisers. Appraisers must never know which part number they are measuring, nor should they see their previous readings, eliminating human confirmation bias.

Calculation Methods: Average & Range vs. ANOVA

Inspectors must understand the mathematical mechanics of both calculation methods.

The Average & Range Method ($\bar{X}$ and $R$)

The Average & Range method calculates range statistics across appraisers and parts, applying empirical conversion factors ($K_1, K_2, K_3$ or $d_2^*$ factors):

  1. Equipment Variation (Repeatability / EV): Calculate the range of trials for each part for each appraiser, then find the grand average range ($\bar{\bar{R}}$): EV=Rˉˉ×K1=Rˉˉd2EV = \bar{\bar{R}} \times K_1 = \frac{\bar{\bar{R}}}{d_2^*} (where $K_1 = 1 / d_2^$ based on the number of trials $r$; for $r = 3$, $K_1 \approx 0.5908$)*.

  2. Appraiser Variation (Reproducibility / AV): Calculate the average measurement for each appraiser ($\bar{X}A, \bar{X}B, \bar{X}C$), find the difference between the maximum and minimum appraiser averages ($\bar{X}{\text{DIFF}} = \bar{X}{\text{max}} - \bar{X}{\text{min}}$), and subtract the repeatability contamination: AV=(XˉDIFF×K2)2EV2nrAV = \sqrt{(\bar{X}_{\text{DIFF}} \times K_2)^2 - \frac{EV^2}{n \cdot r}} (where $n$ is number of parts, $r$ is number of trials, and $K_2 = 1 / d_2^$ based on number of appraisers)*. Note: If the term under the radical is negative, $AV$ is assigned a value of 0.

  3. Combined Gage Repeatability & Reproducibility ($GRR$): GRR=EV2+AV2GRR = \sqrt{EV^2 + AV^2}

  4. Part Variation ($PV$): Calculate the range of average part dimensions ($R_p = \bar{X}{p\text{,max}} - \bar{X}{p\text{,min}}$): PV=Rp×K3PV = R_p \times K_3

  5. Total Variation ($TV$): TV=GRR2+PV2TV = \sqrt{GRR^2 + PV^2}

The ANOVA Method (Analysis of Variance)

The ANOVA method decomposes total variance using sum-of-squares mathematical models. It is statistically superior to the Average and Range method because:

  • It quantifies the Appraiser-by-Part Interaction variance ($AV \times PV$). An interaction occurs when Appraiser A consistently reads small parts oversized while Appraiser B reads large parts undersized.
  • It handles unbalanced study designs (e.g., missing data points or unequal trials).
  • It provides unbiased estimates of variance components with exact degrees of freedom.

Interpreting Results: %GRR & AIAG Acceptance Criteria

Gage R&R evaluates measurement error against two separate benchmarks: Process Variation and Engineering Tolerance.

1. %GRR Relative to Process Variation (%PV)

Evaluates the measurement system's ability to perform statistical process control and capability analysis ($C_p, C_{pk}$):

%GRR=(GRRTV)×100%\%GRR = \left( \frac{GRR}{TV} \right) \times 100\%

2. %GRR Relative to Tolerance (%Tolerance / Precision-to-Tolerance P/T Ratio)

Evaluates the measurement system's ability to make correct product acceptance decisions (accept/reject):

%P/T=(6σGRRUSLLSL)×100%\%P/T = \left( \frac{6 \cdot \sigma_{GRR}}{\text{USL} - \text{LSL}} \right) \times 100\%

(In AIAG MSA 4th Edition, a $6\sigma$ multiplier is standard; historical older standards utilized $5.15\sigma$, corresponding to $99%$ of the normal distribution).

AIAG MSA 4th Edition Benchmark Acceptance Criteria

%GRR / %ToleranceSystem StatusDecision & Action Mandate
Under 10% ($< 10%$)AcceptableThe measurement system is fully capable. Excellent for process control, tight tolerance verification, and capability studies.
10% to 30%Conditionally AcceptableMay be acceptable based upon the importance of the application, cost of gage hardware, or cost of repairing the measurement process. Requires customer or quality engineering approval.
Over 30% ($> 30%$)UnacceptableThe measurement system cannot reliably separate conforming parts from scrap. Measurement system must be quarantined and improved immediately.

Number of Distinct Categories ($ndc$)

The Number of Distinct Categories ($ndc$) metric quantifies the resolution and discrimination of the measurement system relative to actual process variation:

ndc=1.41×(σPartσGRR)=1.41×(PVGRR)ndc = 1.41 \times \left( \frac{\sigma_{\text{Part}}}{\sigma_{GRR}} \right) = 1.41 \times \left( \frac{PV}{GRR} \right)

(Always truncated to the nearest whole integer; never rounded up!)

NUMBER OF DISTINCT CATEGORIES (ndc) INTERPRETATION:

Process Spread: |<-------------------------------------------------------->|

1. ndc = 1 Category (UNACCEPTABLE):
   [======================================================================]
   Gage cannot distinguish any process variation; acts like a coarse light switch.

2. ndc = 2 to 4 Categories (MARGINAL / LOW):
   [============== High ==============][============== Low =============]
   Can only divide process into high/low or high/medium/low. Unusable for SPC!

3. ndc >= 5 Categories (ACCEPTABLE FOR SPC & CAPABILITY):
   [== 1 ==][== 2 ==][== 3 ==][== 4 ==][== 5 ==][== 6 ==][== 7 ==][== 8 ==]
   Gage divides process into 5+ distinct bins; capable of calculating Cp and Cpk.
  • $ndc \ge 5$ (Acceptable): The measurement system has sufficient discrimination to be used for Statistical Process Control (SPC) charting, process capability calculation ($C_{pk}$), and control plan execution.
  • $ndc = 2 \text{ to } 4$ (Marginal): The gage can only divide the process into coarse groupings (e.g., high, medium, low). It can detect massive process shifts, but cannot be used to estimate process capability.
  • $ndc < 2$ (Unacceptable): The measurement system has virtually zero discrimination. Measurement noise overwhelms true process variation.

Attribute Measurement System Studies (Go/No-Go Gages)

Many shop features are inspected using attribute gages (plug gages, thread ring gages, visual appearance standards, boundary samples) that output a binary pass/fail result.

Study Protocol

  • Sample Size: 50 parts selected such that: 20 parts are clearly conforming, 20 parts are clearly nonconforming, and 10 parts are borderline parts (close to upper and lower specification limits).
  • Appraisers & Trials: 3 appraisers evaluate all 50 parts twice ($r = 2$) in a completely randomized, blinded sequence.
  • All parts have a pre-established, certified true condition determined by precision variable laboratory measurement ("ground truth").

Cross-Tabulation & Cohen's Kappa Statistic ($\kappa$)

Attribute agreements are summarized in a contingency matrix comparing appraiser results against ground truth. To eliminate agreement occurring purely by random chance, Cohen's Kappa ($\kappa$) is calculated:

κ=PoPe1Pe\kappa = \frac{P_o - P_e}{1 - P_e}

Where:

  • $P_o$ = The observed proportion of pairwise agreement between appraisers (or between appraiser and reference standard).
  • $P_e$ = The hypothetical proportion of agreement expected by chance alone.

Interpreting Kappa ($\kappa$)

  • $\kappa > 0.75$: Excellent agreement between appraisers and standard (acceptable attribute gage).
  • $0.40 \le \kappa \le 0.75$: Moderate agreement; indicates need for improved visual standards, lighting, or inspector retraining.
  • $\kappa < 0.40$: Poor agreement; the attribute inspection system is invalid and untrustworthy.

Attribute Error Types

  • False Acceptance (Miss Rate / Consumer's Risk): Appraiser accepts a nonconforming part (shipping defects to customers). This is the most dangerous quality failure mode.
  • False Rejection (False Alarm / Producer's Risk): Appraiser rejects a conforming part (needlessly scrapping or reworking good parts).

Real Shop Inspection Scenarios & Troubleshooting MSA

[!WARNING] Exam Trap — High EV vs. High AV Root Causes: Quality exam questions frequently test whether an inspector can identify the root cause of a failing Gage R&R study:

  • If Repeatability (EV) is high: The problem is in the gage hardware or fixturing (loose dial clamp, worn indicator spindle, friction in slides, poor gage resolution, dirty anvils).
  • If Reproducibility (AV) is high: The problem is in the operators or standard operating procedure (differing torque on micrometer thimbles, visual parallax viewing the dial, lack of clear part locating instructions, differing feel on telescoping gages).
  • If Appraiser-by-Part Interaction is high: Different operators are holding, locating, or clamping specific part geometries differently.

Case Study: Resolving a 34% Gage R&R on CNC Turned Pins

  • Problem: An automotive plant conducts a Gage R&R study on a $\varnothing 0.5000\text{ in.} \pm 0.0004\text{ in.}$ ground pin using a 0–1 in. digital caliper. The study yields: $%GRR = 34.2%$ and $ndc = 3$. The system is rejected as unacceptable.
  • Troubleshooting & Corrective Action:
    1. Analysis of variance components reveals: $EV = 12.1%$ and $AV = 31.9%$. The overwhelming contributor to failure is Reproducibility ($AV$).
    2. Investigation reveals that three machinists were applying drastically different thumb pressures to the caliper slider, causing jaw tilt and massive Abbe error. Furthermore, one machinist was reading the part at the tip of the caliper jaws while another was reading near the beam.
    3. Corrective Action: The caliper is replaced with an outside micrometer equipped with a constant-force ratchet stop, mounted in a stationary benchtop stand. Machinists are trained to apply exactly two ratchet clicks.
    4. Re-study Results: The new study on the micrometer yields: $%GRR = 6.4%$ and $ndc = 11$. The measurement system is certified as fully acceptable.
Test Your Knowledge

A quality engineering technician conducts a variable Gage R&R study on a precision machined bore using 10 parts, 3 appraisers, and 3 trials per part. The resulting ANOVA analysis shows that Repeatability (Equipment Variation) is extremely high at 38% of total variation, while Reproducibility (Appraiser Variation) is only 4%. Which corrective action should the quality inspector recommend to address the primary source of measurement variation?

A
B
C
D
Test Your Knowledge

An automotive tier-one supplier conducts a Gage R&R study on a newly designed automated checking fixture used for a critical steering knuckle dimension. The study yields a %GRR relative to process tolerance of 22.4%, and the Number of Distinct Categories (ndc) is calculated as 7. According to AIAG MSA guidelines, how should this measurement system be classified?

A
B
C
D
Test Your Knowledge

In Measurement System Analysis, what is the minimum value required for the Number of Distinct Categories (ndc) for a variable measurement system to be considered capable of estimating process capability and supporting Statistical Process Control (SPC)?

A
B
C
D