9.4 Variable Gage R&R: Repeatability and Reproducibility
Key Takeaways
- Gage Repeatability (Equipment Variation, EV) quantifies the variation observed when one appraiser measures the same part multiple times with the same gage, representing short-term instrument noise.
- Gage Reproducibility (Appraiser Variation, AV) quantifies the variation observed in measurement averages across different appraisers measuring identical parts with the same gage, representing operator technique differences.
- Under AIAG guidelines, a measurement system is acceptable if %GRR < 10%, conditionally acceptable between 10% and 30% based on application criticality, and unacceptable if %GRR > 30%.
- The Number of Distinct Categories (ndc = 1.41 * PV / GRR) defines the gage's effective discrimination relative to process variation, strictly requiring ndc >= 5 for valid statistical process control.
- The ANOVA method is superior to the Average and Range method because it mathematically partitions and tests the Operator-by-Part interaction effect, revealing whether specific appraisers struggle with specific part geometries.
9.4 Variable Gage R&R: Repeatability and Reproducibility
Understanding Precision: Repeatability (EV) and Reproducibility (AV)
While accuracy (bias, linearity, stability) assesses whether a measurement system is centered on the true reference standard, precision assesses the spread or dispersion of repeated measurements. In industrial quality control, precision is evaluated through a Gage Repeatability and Reproducibility (Gage R&R) study.
Total measurement system precision error is mathematically partitioned into two distinct physical components:
PRECISION ERROR DECOMPOSITION:
TOTAL GAGE R&R VARIATION (sigma_GRR)
|
+----------------+----------------+
| |
REPEATABILITY REPRODUCIBILITY
Equipment Variation (EV) Appraiser Variation (AV)
| |
- Same appraiser - Different appraisers
- Same gage - Same gage
- Same part / characteristic - Same parts
- Short time intervals - Different techniques / clamping
- Inherent gage noise / play - Human interface differences
1. Repeatability (Equipment Variation, EV)
Repeatability is the variation observed when one appraiser uses the same gage to measure the same characteristic on the same identical part multiple times under identical operating conditions.
- What it reflects: The inherent mechanical, structural, and electrical noise of the measurement device itself.
- Root causes of high EV: Mechanical backlash in gear teeth, loose spindle bearings, sloppy slideways, electronic sensor drift, lack of gage rigidity, elastic fixture flexing, or friction in indicator dials.
2. Reproducibility (Appraiser Variation, AV)
Reproducibility is the variation observed in the average of measurements made by different appraisers using the same gage when measuring identical characteristics on the same set of parts.
- What it reflects: The variability introduced by the human beings operating the measurement equipment.
- Root causes of high AV: Inconsistent clamping force, differences in operator hand squeeze, varying sight angles causing parallax, improper part seating against datums, differing interpretations of zero, or poor operational training.
Combining EV and AV into Total Gage R&R ($GRR$)
Because repeatability and reproducibility are independent sources of variance, their variances sum directly:
Variable Gage R&R Study Architecture and Design Rules
To perform a statistically rigorous variable Gage R&R study, the quality technician must follow standardized AIAG design protocols:
The Standard Study Structure: 10-3-3 Architecture
- 10 Parts ($n = 10$): Representative parts selected from the manufacturing process.
- 3 Appraisers ($k = 3$): Operators who normally inspect the product on production shifts.
- 3 Trials ($r = 3$): Each operator measures each part three separate times.
- Total Data Points: $10 \text{ parts} \times 3 \text{ operators} \times 3 \text{ trials} = \mathbf{90 \text{ individual measurements}}$. (Note: For coarse screening or smaller operations, a $10 \times 2 \times 2 = 40$-point study is sometimes permitted, but the 10-3-3 structure is the industrial benchmark).
Three Inflexible Execution Rules
- Parts Must Span Process Variation: The 10 parts must not be chosen consecutively from one hour of production. They must be selected randomly across days or shifts to encompass the full range of historical manufacturing variation (from the low end to the high end of typical production). Selecting 10 nearly identical parts artificially deflates Part Variation ($PV$), causing calculated $%GRR$ to appear disastrously inflated!
- Blind Numbering: The 10 parts must be coded or labeled so that the appraisers cannot identify which part they are measuring. If an operator sees "Part 3" and remembers they measured it at $1.002"$ earlier, they will subconsciously bias their second reading to $1.002"$, destroying the study's validity.
- Completely Randomized Order: In Trial 1, Appraiser A measures all 10 parts in randomized order (e.g., 7, 2, 9, 1, ...). Then Appraiser B measures them in a newly randomized order, followed by Appraiser C. In Trials 2 and 3, the parts are re-randomized. Technicians must never allow an appraiser to measure the same part 3 times consecutively!
Mathematical Methodologies: Average and Range ($\bar{X}-R$) vs. ANOVA
Two mathematical techniques are recognized by AIAG for calculating Gage R&R:
1. The Average and Range ($\bar{X}-R$) Method
The Average and Range method is the traditional calculation technique historically favored on shop floors because it can be computed manually using range statistics and tabulations of $d_2^*$ constants.
Step-by-Step Computational Formulas:
- Range per Part Trial: For each appraiser on each part, calculate the range across trials: $R_{ij} = X_{\max} - X_{\min}$.
- Average Range per Appraiser: Calculate the average range for Appraiser A ($\bar{R}_A$), Appraiser B ($\bar{R}_B$), and Appraiser C ($\bar{R}_C$).
- Grand Average Range ($\bar{\bar{R}}$):
- Equipment Variation (Repeatability, $EV$): Where $K_1 = 1 / d_2^$. For $r = 3$ trials, $K_1 = 0.5908$ ($d_2^ \approx 1.693$).
- Appraiser Averages: Calculate the overall grand average measurement for each appraiser ($\bar{X}_A, \bar{X}_B, \bar{X}_C$).
- Appraiser Difference ($\bar{X}_{\text{diff}}$):
- Appraiser Variation (Reproducibility, $AV$): Where $n = 10$ parts, $r = 3$ trials, and $K_2 = 1 / d_2^*$ for $k$ appraisers (for $k = 3$ appraisers, $K_2 = 0.5231$). (Important rule: If the value under the square root is negative, set $AV = 0$).
- Total Gage R&R ($GRR$):
- Part Variation ($PV$): Calculate the average for each part across all appraisers and trials ($\bar{X}{P1}, \dots, \bar{X}{P10}$). Where $K_3 = 1 / d_2^*$ for $n$ parts (for $n = 10$ parts, $K_3 = 0.3146$).
- Total Variation ($TV$):
2. The ANOVA (Analysis of Variance) Method
The ANOVA method is the globally preferred modern standard for Gage R&R. While the range method only estimates main effects, ANOVA decomposes the total sum of squares into four statistical components:
- Parts Variance ($\sigma^2_{\text{Part}}$)
- Appraisers Variance ($\sigma^2_{\text{Appraiser}}$)
- Appraiser-by-Part Interaction Variance ($\sigma^2_{\text{Appraiser} \times \text{Part}}$)
- Equipment Error / Repeatability ($\sigma^2_{\text{Residual / Error}}$)
[!IMPORTANT] The Decisive Advantage of ANOVA: ANOVA explicitly measures the Operator-by-Part Interaction. If the interaction $p$-value is statistically significant ($p < 0.05$), it indicates that certain appraisers obtain systematically different readings on certain specific part geometries (for example, Appraiser A exerts excessive clamping force on thin-walled parts, but measures thick parts accurately). The Average and Range method cannot detect interaction effects!
AIAG Acceptance Benchmarks and Tolerance vs. Process Variation
Once total $GRR$ is calculated, it must be evaluated as a percentage against an engineering baseline. In modern MSA, two distinct baselines are evaluated:
1. $%GRR$ Relative to Total Variation ($\sigma_{\text{Process}}$)
Evaluates whether the measurement system is capable of monitoring process shifts for Statistical Process Control (SPC):
2. $%GRR$ Relative to Specification Tolerance ($\text{Tolerance}$)
Evaluates whether the measurement system is capable of inspecting and sorting parts against blueprint specifications: (Note: AIAG 4th Edition standardizes on $6\sigma$, representing $99.73%$ coverage, whereas older 3rd Edition manuals utilized $5.15\sigma$ for $99.0%$ coverage).
The AIAG Acceptance Criteria Matrix
| Percentage Range | Measurement System Capability Status | Required Engineering Action |
|---|---|---|
| $%GRR < 10%$ | Acceptable | Measurement system is fully capable. Recommended for all quality inspections and SPC charting. |
| $10% \le %GRR \le 30%$ | Conditionally Acceptable | May be acceptable based on application criticality, cost of replacement instrumentation, or customer concurrence. |
| $%GRR > 30%$ | Unacceptable | Measurement system is incapable. Cannot reliably distinguish good product from bad product. Requires immediate corrective action (redesign, re-fixturing, operator retraining, or gage replacement). |
Number of Distinct Categories ($ndc$) and Process Discrimination
A gage may possess high sensitivity, yet still be incapable of dividing process variation into meaningful sub-groups. To quantify practical measurement discrimination, the AIAG standard defines the Number of Distinct Categories ($ndc$):
[!IMPORTANT] Number of Distinct Categories ($ndc$): Represents the number of non-overlapping $97%$ confidence intervals that span the expected product variation. It defines the gage's ability to detect process variation. Note: Always truncate (round down) the calculated value to the nearest integer!
NUMBER OF DISTINCT CATEGORIES (ndc) INTERPRETATION:
ndc = 1: [=============================================] (UNACCEPTABLE: Gage sees parts as one chunk!)
Cannot distinguish parts; completely obscured by gage noise.
ndc = 3: [ Low Group ] [ Mid Group ] [ High Group ] (COARSE SCREENING ONLY)
Can only sort into coarse buckets; invalid for SPC control charts.
ndc >= 5: [1] [2] [3] [4] [5] [6] [7] [8] (ACCEPTABLE FOR SPC)
Gage cleanly separates natural process variation into distinct categories.
AIAG Criteria for $ndc$:
- $ndc \ge 5$: Acceptable. The measurement system possesses adequate discrimination to support Statistical Process Control (SPC) control charts and process capability analysis ($C_p, C_{pk}$).
- $2 \le ndc < 5$: Conditionally Acceptable for Coarse Sorting Only. The system can divide parts into high, medium, and low groups, but is too coarse to detect process trends or maintain variables control charts.
- $ndc < 2$ (i.e., $ndc = 1$): Unacceptable. Measurement noise completely swamps part-to-part variation. The system cannot distinguish one part from another.
Attribute Agreement Analysis: Fleiss' Kappa and Concordance
When quality inspections produce binary or nominal classifications (Pass/Fail, Conforming/Nonconforming, Good/Bad visual defect sorting) rather than continuous variables data, technicians must conduct an Attribute Agreement Analysis (Attribute Gage R&R).
Study Setup
- Select 30 to 50 parts with known reference master status (determined by engineering experts or laboratory metrology).
- The sample must include conforming parts, nonconforming parts, and critically: borderline / marginal parts.
- Have 2 to 3 appraisers inspect the parts across 2 to 3 trials in blind, randomized order.
Key Metrics Evaluated
- Within-Appraiser Agreement (Repeatability): How consistently does Operator A classify the same part across repeated trials?
- Between-Appraiser Agreement (Reproducibility): How frequently do Appraiser A, Appraiser B, and Appraiser C agree with each other on the same part?
- Agreement with Reference Standard (Accuracy / Effectiveness): How frequently do appraiser classifications match the true known master status?
Fleiss' Kappa ($\kappa$) Statistic
Raw percent agreement is deceptive because appraisers can agree purely by random guessing. To eliminate chance agreement, MSA calculates Cohen's Kappa (for 2 raters) or Fleiss' Kappa (for $> 2$ raters): Where $P_{\text{observed}}$ is the proportion of observed agreements, and $P_{\text{expected}}$ is the proportion of agreements expected purely by random chance.
| Kappa Value ($\kappa$) | Agreement Level | System Interpretation |
|---|---|---|
| $\kappa > 0.75$ | Excellent Agreement | Measurement system is highly capable for visual/attribute sorting. |
| $0.40 \le \kappa \le 0.75$ | Moderate / Acceptable Agreement | Conditionally acceptable; requires defect limit boundary samples and operator training. |
| $\kappa < 0.40$ | Poor Agreement | Unacceptable; attribute criteria are ambiguous and unreliable. |
Step-by-Step Worked Numerical Example: Full Gage R&R Calculation
Scenario Data
A quality technician conducts a variable Gage R&R study on a CNC lathe turning diameter. The study employs $n = 10$ parts, $k = 3$ operators, and $r = 3$ trials ($90$ total data points). The engineering print specification is $1.2500" \pm 0.0100"$ (Total Tolerance $= 0.0200"$).
The preliminary data analysis yields the following calculated summary statistics:
- Grand Average Range across all operators and parts: $\bar{\bar{R}} = 0.00085"$
- Operator Grand Averages: $\bar{X}_A = 1.25040"$, $\bar{X}_B = 1.25012"$, $\bar{X}_C = 1.24998"$
- Part Averages Range: $R_p = \bar{X}{P,\max} - \bar{X}{P,\min} = 1.25600" - 1.24350" = 0.01250"$
Constants from AIAG Tables:
- For $r = 3$ trials: $K_1 = 0.5908$
- For $k = 3$ operators: $K_2 = 0.5231$
- For $n = 10$ parts: $K_3 = 0.3146$
Step 1: Calculate Repeatability (Equipment Variation, $EV$)
Step 2: Calculate Reproducibility (Appraiser Variation, $AV$)
First, find the maximum difference between operator averages: Now compute $AV$:
Step 3: Calculate Total Gage R&R ($GRR$)
Step 4: Calculate Part Variation ($PV$)
Step 5: Calculate Total Variation ($TV$)
Step 6: Evaluate $%GRR$ Acceptance Criteria
- $%GRR$ Relative to Total Variation (Process):
- $%GRR$ Relative to Specification Tolerance:
- Evaluation: Both $%GRR_{TV}$ ($13.6%$) and $%GRR_{\text{Tol}}$ ($16.2%$) fall cleanly between $10%$ and $30%$. The measurement system is Conditionally Acceptable.
Step 7: Calculate Number of Distinct Categories ($ndc$)
- Truncating to the nearest integer: $ndc = 10$.
- Discrimination Evaluation: Because $ndc = 10 \ge 5$, the gage possesses excellent discrimination, capable of dividing part variation into 10 distinct categories. It is fully qualified for Statistical Process Control (SPC) charting.
Technician Inspection Scenarios & Common Exam Traps
Real-World Shop Scenario: Investigating High %GRR
A quality technician conducts a Gage R&R on an optical shaft measurement gage and finds $%GRR = 38%$ (unacceptable). The technician's manager immediately demands that all operators be sent to an off-site training course to improve their measurement technique. The technician reviews the data sheet and notes:
- $EV = 0.0018"$
- $AV = 0.0003"$
- Root Cause Diagnosis: The manager's decision is fundamentally flawed. $EV$ accounts for over $97%$ of the total measurement variance ($EV^2 / GRR^2 = (0.0018)^2 / [(0.0018)^2 + (0.0003)^2] = 0.973$). Appraiser variation ($AV$) is tiny. Training operators will not improve $EV$ by even one percent! The problem is mechanical equipment variation: worn spindle bearings, loose optical stage guide rails, or camera vibration. The technician must rebuild the mechanical fixture.
Common Exam Traps for CQT Candidates
- Exam Trap 1: Truncating $ndc$: Always round down (truncate) $ndc$ to the integer. If $ndc = 4.95$, it is 4, which fails the requirement of $ndc \ge 5$. Never round up to 5!
- Exam Trap 2: Variance Summation vs. Standard Deviations: Remember that variances add directly ($\sigma^2_{GRR} = \sigma^2_{EV} + \sigma^2_{AV}$), but standard deviations do not add linearly ($GRR \ne EV + AV$). You must take the square root of the sum of squares: $GRR = \sqrt{EV^2 + AV^2}$.
- Exam Trap 3: Handpicked "Golden" Parts in Study Design: If an exam question asks: "What happens if an engineer hand-selects 10 parts that are all machined near nominal?" Handpicking identical parts slashes Part Variation ($PV$). Because $%GRR = \frac{GRR}{\sqrt{GRR^2 + PV^2}}$, a tiny $PV$ forces the denominator to shrink, causing $%GRR$ to skyrocket falsely.
- Exam Trap 4: ANOVA Interaction Interpretation: If an ANOVA Gage R&R indicates a significant $p$-value for the Operator $\times$ Part interaction ($p < 0.05$), the root cause is that appraisers apply different techniques to different part sizes (e.g., operator deflection on thin parts, or inconsistent sight angles on tall parts).
During a variable Gage R&R study on a CNC lathe turning diameter, the technician calculates an Equipment Variation of EV = 0.0014 inches and an Appraiser Variation of AV = 0.0003 inches, resulting in a total %GRR of 34% (unacceptable). What is the primary source of the measurement error, and what corrective action should be taken?
A Gage R&R study conducted on a fuel injector nozzle seat bore produces a %GRR of 14.5% of process variation and 16.2% of specification tolerance. Under standard AIAG guidelines, how should the quality engineering team classify this measurement system?
A Gage R&R study yields a Part Variation of PV = 0.00480 inches and a total Gage R&R of GRR = 0.00150 inches. What is the Number of Distinct Categories (ndc), and does the measurement system have adequate discrimination to control the process?