6.4 Variable & Attribute Gage R&R Studies
Key Takeaways
- Precision is decomposed into Repeatability (Equipment Variation - EV: within-operator spread on identical parts) and Reproducibility (Appraiser Variation - AV: between-operator spread).
- A standard Variable Gage R&R study evaluates 10 representative parts, measured by 2 to 3 appraisers across 2 to 3 randomized blind trials, analyzed via either the Range method or the Two-Way ANOVA method.
- Acceptance criteria for Variable Gage R&R dictate that %GRR < 10% is acceptable, 10% to 30% is conditionally acceptable based on criticality, and > 30% is unacceptable; the Number of Distinct Categories (ndc) must be at least 5.
- The Number of Distinct Categories (ndc = 1.41 × σ_PV / σ_GRR) represents the number of non-overlapping confidence intervals the gage can distinguish; an ndc < 5 invalidates process capability studies regardless of %GRR.
- Attribute Agreement Analysis evaluates categorical inspection using within-appraiser repeatability, between-appraiser reproducibility, and accuracy against a master standard, assessed via Cohen's or Fleiss' Kappa (κ ≥ 0.75).
6.4 Variable & Attribute Gage R&R Studies
Quick Summary: While calibration addresses measurement accuracy (centering), Gage Repeatability and Reproducibility (Gage R&R) quantifies measurement precision (spread). In continuous systems, a Variable Gage R&R decomposes measurement error into Repeatability (Equipment Variation - EV) and Reproducibility (Appraiser Variation - AV). Under standard AIAG and CSSC guidelines, a measurement system is acceptable if %GRR is under 10% (or conditionally under 30%) and the Number of Distinct Categories ($ndc$) is 5 or greater. For discrete and visual inspections, an Attribute Agreement Analysis evaluates within-appraiser consistency, between-appraiser reproducibility, and accuracy against a master standard using Cohen's or Fleiss' Kappa ($\kappa \ge 0.75$).
Precision Decomposed: Repeatability vs. Reproducibility
Precision refers to the dispersion of repeated measurements taken on identical parts under specified operating conditions. Total measurement system variance ($\sigma^2_{\text{GRR}}$) is the sum of two distinct variance components:
Dissecting Gage R&R Precision
Gage R&R Variance (σ²_GRR)
│
┌─────────────────────┴─────────────────────┐
▼ ▼
Repeatability Reproducibility
(Equipment Variation - EV) (Appraiser Variation - AV)
• Same appraiser • Different appraisers
• Same gage • Same gage
• Same parts measured multiple times • Same parts measured multiple times
• Reflects gage mechanics, stiffness, • Reflects clamping pressure, angle,
bearing play, or electrical noise parallax, or procedural habits
1. Repeatability (Equipment Variation - EV)
Repeatability is the variation observed when one appraiser measures the same characteristic on the same parts multiple times using the same gage in the same physical environment.
- Root Causes of High Repeatability Error: Mechanical play in gage slides or bearings, loose fixtures, sensor noise, electrical interference, poor gage resolution, thermal contraction, or lack of instrument rigidity.
- Operational Fix: Repair, refurbish, or replace the physical measuring instrument, or stiffen the part clamping fixture.
2. Reproducibility (Appraiser Variation - AV)
Reproducibility is the variation observed in the average measurements made by different appraisers evaluating the same parts using the same gage.
- Root Causes of High Reproducibility Error: Differences in operator technique, variations in manual clamping torque, sightline parallax errors on analog scales, inconsistent gage positioning, or ambiguous procedural instructions in standard work.
- Operational Fix: Standardize operator work instructions, conduct training on gage seating, introduce mechanical torque-limiters (e.g., ratcheting thimbles), or automate the inspection station.
3. Operator-by-Part Interaction ($\sigma^2_{\text{Operator} \times \text{Part}}$)
In advanced studies, an interaction effect occurs when an appraiser measures certain parts differently than other parts (e.g., Appraiser A measures large parts with excessive clamping pressure, causing elastic deformation, but measures small parts normally). Interaction effects can only be isolated using the ANOVA method.
Variable Gage R&R Study Architecture
A Variable Gage R&R study evaluates continuous measurement systems. To generate statistically valid variance estimates, the study must follow a rigorous experimental protocol:
Standard Variable Gage R&R Matrix
10 Representative Parts (Spanning Full Process Variation Envelope)
┌───┬───┬───┬───┬───┬───┬───┬───┬───┬───┐
│ 1 │ 2 │ 3 │ 4 │ 5 │ 6 │ 7 │ 8 │ 9 │ 10│
└───┴───┴───┴───┴───┴───┴───┴───┴───┴───┘
│
┌────────────┼────────────┐
▼ ▼ ▼
Appraiser A Appraiser B Appraiser C (2 to 3 Frontline Operators)
│ │ │
┌────┴────┐ ┌────┴────┐ ┌────┴────┐
│ Trial 1 │ │ Trial 1 │ │ Trial 1 │ (2 to 3 Randomized, Blind Trials)
│ Trial 2 │ │ Trial 2 │ │ Trial 2 │
│ Trial 3 │ │ Trial 3 │ │ Trial 3 │
└─────────┘ └─────────┘ └─────────┘
Total Data Points = 10 Parts × 3 Appraisers × 3 Trials = 90 Observations
Experimental Protocol Rules
- Part Selection (10 Parts): Parts must be sampled to represent the entire natural range of process variation (including units from the low end, center, and high end of the distribution). Crucial Warning: Do not cherry-pick only nominal parts; doing so artificially suppresses Part-to-Part variation ($\sigma_{\text{PV}}$), causing the gage to appear artificially incapable!
- Appraiser Selection (2 to 3 Appraisers): Appraisers must be actual frontline operators or quality technicians who routinely perform the inspection, not metrology PhDs or laboratory supervisors.
- Randomized Blind Trials (2 to 3 Trials): The 10 parts must be re-coded (e.g., marked with hidden numbers on the reverse side) so appraisers cannot identify which part they are measuring. All parts must be measured in completely randomized order across trials to prevent recall bias.
Calculation Methods: Average & Range ($X$-bar & $R$) vs. ANOVA
Two mathematical approaches are used to analyze Variable Gage R&R data:
| Dimension | Average & Range ($X$-bar & $R$) Method | Two-Way ANOVA Method (Recommended) |
|---|---|---|
| Computational Basis | Uses range averages ($\bar{\bar{R}}$) and tabular constants ($K_1, K_2, K_3$) | Uses Sum of Squares ($SS$), Mean Squares ($MS$), and $F$-tests |
| Interaction Isolation | Cannot isolate Operator-by-Part interaction | Isolates and quantifies Operator $\times$ Part interaction |
| Unbalanced Data | Fails if trials or parts are missing | Readily handles unbalanced experimental designs |
| Variance Estimation | Biased for small sample sizes; less accurate | Unbiased variance component estimates; exact confidence intervals |
| Industry Status | Traditional manual calculation method | Standard method in all modern statistical software (Minitab, JMP) |
The Average and Range Method Formulas
Under the Range method, standard deviations are derived from range averages using standard AIAG/CSSC constants:
-
Equipment Variation (Repeatability): (Where $\bar{\bar{R}}$ is the average of all appraiser ranges, and $K_1 = 1/d_2^$ for trials $r$.)*
-
Appraiser Variation (Reproducibility): (Where $\bar{X}_{\text{diff}}$ is the difference between the highest and lowest appraiser means, $n = \text{parts}$, and $r = \text{trials}$.)
-
Gage R&R Standard Deviation:
-
Part-to-Part Variation: (Where $R_p$ is the range of part averages across all appraisers.)
-
Total Variation:
Variable Gage R&R Acceptance Criteria
A measurement system is evaluated across three primary benchmarks: %GRR of Total Variation (%TV), Precision-to-Tolerance Ratio (P/T Ratio), and the Number of Distinct Categories ($ndc$).
1. Percent of Total Variation (%TV / %GRR)
Evaluates measurement system variation relative to the total observed process variation:
| %GRR Range | Measurement System Capability | Practical Operational Action |
|---|---|---|
| < 10% | Acceptable | Measurement system is fully capable. Reliable for process control and capability analysis. |
| 10% to 30% | Conditionally Acceptable | May be acceptable based on application importance, gage replacement cost, or customer risk. |
| > 30% | Unacceptable | System cannot be trusted. Must be repaired, replaced, or recalibrated before proceeding. |
2. Precision-to-Tolerance Ratio (P/T Ratio / %Tolerance)
Evaluates measurement system variation relative to the customer's engineering tolerance band:
(Note: Historically, some automotive standards used $5.15\sigma$ representing a 99% interval; the modern CSSC and AIAG standard uses $6.0\sigma$ representing a 99.73% interval).
The identical threshold criteria apply: $< 10%$ is acceptable, $10% \text{ to } 30%$ is conditionally acceptable, and $> 30%$ is unacceptable.
3. Number of Distinct Categories ($ndc$)
Even if a gage exhibits an acceptable %GRR, it must possess sufficient resolution to divide process variation into distinct groups. The Number of Distinct Categories ($ndc$) represents the number of non-overlapping 97% confidence intervals across process spread:
[!IMPORTANT] The Truncation Rule: The calculated $ndc$ value is truncated (rounded down) to the nearest whole integer. For example, $ndc = 4.98$ truncates to 4.
$ndc$ Acceptance Benchmarks
- $ndc \ge 5$ (Acceptable): The measurement system can divide the process distribution into 5 or more distinct categories (e.g., low, medium-low, center, medium-high, high). Required for statistical process control (SPC) and capability analysis ($C_p / C_{pk}$).
- $ndc = 2$ to $4$ (Marginal / Coarse Screening Only): The gage can only classify parts into broad groups (e.g., high/low). Unacceptable for continuous SPC or capability estimation.
- $ndc = 1$ (Completely Unacceptable): The gage cannot distinguish part variation; all parts look identical to the measurement system.
Complete Worked Calculation: Variable Gage R&R Study
A medical implant manufacturer conducts a Variable Gage R&R study on an automated laser micrometer measuring knee joint femoral component thickness. The study design utilizes $n = 10$ parts, $k = 3$ appraisers, and $r = 2$ trials per appraiser.
Engineering specifications dictate: $\text{LSL} = 15.000\text{ mm}$, $\text{USL} = 16.000\text{ mm}$ (Tolerance Band $= 1.000\text{ mm}$).
From the two-way ANOVA analysis, the software extracts the following variance components:
- Repeatability Variance: $\sigma^2_{\text{EV}} = 0.000400\text{ mm}^2$
- Reproducibility Variance: $\sigma^2_{\text{AV}} = 0.000225\text{ mm}^2$
- Part-to-Part Variance: $\sigma^2_{\text{PV}} = 0.005625\text{ mm}^2$
Step 1: Compute Component Standard Deviations
Step 2: Calculate Gage R&R Variance and Standard Deviation
Step 3: Calculate Total Observed Variance and Standard Deviation
Step 4: Calculate %GRR (Percent of Total Variation)
Status: Unacceptable ($31.62% > 30%$ by total variation criteria).
Step 5: Calculate Precision-to-Tolerance (P/T Ratio)
Status: Conditionally Acceptable ($15.0%$ falls between $10%$ and $30%$ against tolerance).
Step 6: Calculate Number of Distinct Categories ($ndc$)
Truncating to the nearest integer yields:
Status: Unacceptable ($ndc = 4 < 5$).
Step 7: Practical Continuous Improvement Recommendation
Although the measurement system consumes only 15.0% of the customer's tolerance band, its %GRR of 31.62% and $ndc$ of 4 make it unacceptable for process control and capability analysis. The gage cannot divide the natural process spread into the minimum required 5 categories.
Looking at the variance contribution:
Because Repeatability (EV) accounts for 64% of total measurement noise, engineering should prioritize mechanical fixture rigidity and laser sensor cleaning before retraining appraisers.
Attribute Agreement Analysis (Attribute MSA)
When inspection data is discrete, binary, or qualitative (such as visual inspection of paint blemishes, document completeness audits, or customer service call compliance), continuous Gage R&R cannot be applied. Instead, the team executes an Attribute Agreement Analysis.
Four Dimensions of Attribute Agreement
1. Within-Appraiser (Repeatability) ▶ Does Inspector A agree with Inspector A
across Trial 1 and Trial 2?
2. Between-Appraiser (Reproducibility)▶ Does Inspector A agree with Inspector B
on identical samples?
3. Appraiser vs. Standard (Accuracy) ▶ Does Inspector A agree with the verified
Master Reference Standard?
4. Overall System vs. Standard ▶ Do ALL inspectors agree with each other
AND the Master Standard on ALL trials?
Study Architecture
- Sample Size: 30 to 50 samples selected from the process.
- Master Standard: Each sample must have a verified, pre-established master classification by an expert or committee.
- Crucial Part Mix: The sample set must include conforming parts, defective parts, and critically, borderline / marginal parts that challenge the visual defect threshold.
- Appraisers & Trials: 2 to 3 operational inspectors evaluate all parts in 2 independent, randomized, blind trials.
Statistical Evaluation: Fleiss' and Cohen's Kappa ($\kappa$)
Raw percentage agreement is a deceptive metric because appraisers can achieve high apparent agreement purely by chance (especially when the defect rate is low). Attribute MSA utilizes Kappa statistics (Cohen's Kappa for 2 appraisers; Fleiss' Kappa for multiple appraisers) to correct for chance agreement:
Where:
- $P_{\text{observed}}$ = Proportion of units where appraisers actually agree.
- $P_{\text{chance}}$ = Proportion of agreement expected purely due to chance.
Interpreting Kappa ($\kappa$)
- $\kappa > 0.75$ to $0.80$: Good to Excellent Agreement. The measurement system is acceptable for production inspection.
- $0.40 \le \kappa \le 0.75$: Moderate / Marginal Agreement. Measurement system needs improvement; requires clear boundary visual aids, refined operational definitions, and inspector retraining.
- $\kappa < 0.40$: Poor Agreement. The measurement system is unacceptable and statistically indistinguishable from coin flipping.
Misclassification Analysis: False Alarm vs. Miss Rate
In addition to Kappa, Attribute MSA evaluates two operational error rates:
- False Alarm Rate ($\alpha$ Risk / Type I Error): Proportion of true conforming parts rejected by appraisers (drives scrap and rework costs).
- Miss Rate ($\beta$ Risk / Type II Error): Proportion of true defective parts accepted by appraisers (drives customer defect escapes).
Critical Exam Traps to Avoid
- Trap 1: Directly Adding Standard Deviations Instead of Variances — Adding $EV$ and $AV$ directly to get $GRR$. Standard deviations never add directly: $\sigma^2_{\text{GRR}} = \sigma^2_{\text{EV}} + \sigma^2_{\text{AV}}$, so $GRR = \sqrt{EV^2 + AV^2}$.
- Trap 2: Confusing Repeatability with Reproducibility — Repeatability is equipment variation (same operator, same gage, same part). Reproducibility is appraiser variation (different operators, same gage, same part).
- Trap 3: Overlooking the Number of Distinct Categories ($ndc < 5$) — Approving a measurement system because %GRR is 22% (conditionally acceptable) while ignoring an $ndc$ of 3. If $ndc < 5$, the gage lacks sufficient resolution for SPC and capability studies.
- Trap 4: Assuming High Repeatability Proves an Attribute System is Capable — An appraiser can be 100% repeatable with themselves while being completely inaccurate against the master standard (e.g., consistently rejecting conforming parts). Appraiser vs. Standard agreement is mandatory.
A quality engineering team conducts a Variable Gage R&R study on an optical shaft micrometer. The analysis reveals that the percentage of total variation consumed by Gage R&R (%GRR) is 22.4%, and the software reports a Number of Distinct Categories (ndc) of 3. How should the Green Belt interpret these findings?
In a Variable Gage R&R study using 10 parts, 3 appraisers, and 2 trials, the ANOVA decomposition yields an Equipment Variation (Repeatability) standard deviation of EV = 0.030 mm and an Appraiser Variation (Reproducibility) standard deviation of AV = 0.040 mm. What is the combined Gage R&R standard deviation (GRR)?
A continuous improvement team conducts an Attribute Agreement Analysis on a visual surface inspection process across 3 inspectors evaluating 50 parts in 2 blind trials. The analysis reports that within-appraiser agreement is 96%, between-appraiser agreement is 92%, but the Fleiss' Kappa against the certified master standard is κ = 0.52. What does this outcome reveal about the inspection process, and what action is required?