7.2 Gauge Repeatability & Reproducibility (GR&R)
Key Takeaways
Repeatability (Equipment Variation or EV) measures the variation observed when one operator measures the same part repeatedly with the same gauge under identical conditions.
Reproducibility (Appraiser Variation or AV) measures the variation observed when different operators measure the same parts using the same gauge.
Total Gauge R&R variance combines equipment variation and appraiser variation: σ²GRR = σ²EV + σ²AV.
Under AIAG guidelines, a measurement system with %GR&R under 10% is acceptable, 10% to 30% is conditionally acceptable, and over 30% is unacceptable.
A capable measurement system must have a Number of Distinct Categories (ndc) of 5 or greater, while qualitative pass/fail inspection systems are validated using Attribute Agreement Analysis and Kappa statistics.
Gauge Repeatability & Reproducibility (GR&R)
Quick Answer: Gauge Repeatability and Reproducibility (GR&R) is the primary statistical study used in the Six Sigma Measure phase to evaluate measurement system precision for continuous data. Measurement variation is partitioned into Repeatability (Equipment Variation, EV), which captures variation from repeated measurements by one operator on the same part with the same gauge, and Reproducibility (Appraiser Variation, AV), which captures variation between different operators. Total Gauge R&R is . Industry benchmarks classify a system as acceptable if %GR&R is under 10%, conditionally acceptable between 10% and 30%, and unacceptable if over 30%. In addition, the Number of Distinct Categories () must be 5 or greater. Qualitative inspection is evaluated through Attribute Agreement Analysis and Kappa statistics. Independent CSSYB study guide by OpenExamPrep.
Components of Measurement System Precision: EV and AV
While calibration and bias studies address whether a gauge is centered on the true master value (accuracy), Gauge Repeatability and Reproducibility (GR&R) evaluates whether the system produces consistent readings upon repeated use (precision).
Total measurement system variation () is partitioned into two distinct components:
Repeatability: Equipment Variation (EV)
Repeatability represents the inherent variability of the measurement device. It is defined as the variation observed when one operator measures the same part multiple times using the same gauge in the same environment:
- Root Causes of High EV: Mechanical play, loose bearings, worn indicator dials, electrical sensor noise, or flexing fixtures.
- Corrective Actions for High EV: Rebuilding or servicing the gauge, improving part clamping rigidity, and reducing mechanical vibration.
Reproducibility: Appraiser Variation (AV)
Reproducibility represents the human and procedural variability between operators. It is defined as the variation observed among the average measurements made by different operators using the same gauge to evaluate the same parts:
- Root Causes of High AV: Ambiguous operational definitions, differing clamping force, improper probe angles, or inconsistent visual alignment.
- Corrective Actions for High AV: Developing clear Standard Operating Procedures (SOPs), conducting operator training, and implementing visual fixturing to standardize part positioning.
Designing a Variable Gauge R&R Study
A variable Gauge R&R study must be carefully structured to isolate equipment variation from appraiser variation across a representative production sample.
Standard Study Protocol
A typical study, following the AIAG Measurement Systems Analysis manual, uses:
- 10 Parts (): Selected from standard production to represent the entire process variation spread, including parts near nominal and near both specification limits.
- 2 to 3 Operators (): Trained technicians who routinely operate the equipment in daily production.
- 2 to 3 Trials (): Each operator measures every part 2 or 3 separate times.
- Blinded and Randomized Order: Parts must be secretly coded so operators cannot identify parts or recall previous readings. Measurement order must be fully randomized across trials.
Calculation Methods: Average and Range vs. ANOVA
Two mathematical approaches calculate GR&R components:
- Average and Range () Method: A traditional approach using range calculations and statistical correction factors () to estimate and . While simple for manual calculation, it assumes no interaction between operators and parts.
- Analysis of Variance (ANOVA) Method: The preferred modern method. ANOVA breaks down total variance into Part-to-Part, Operator (AV), Equipment (EV), and the Operator-by-Part Interaction. A significant interaction reveals that specific operators measure certain part geometries differently, highlighting targeted training or fixturing issues.
%GR&R Acceptance Criteria and Interpretation
Total measurement standard deviation () is expressed as a percentage of total process variation (%Study Variation) or specification tolerance (%Tolerance / P/T ratio):
AIAG Acceptance Thresholds
| %GR&R Range | System Status | Operational Action Required |
|---|---|---|
| Acceptable | The measurement system is fully capable. Measurement error accounts for a negligible fraction of process spread. | |
| Conditionally Acceptable | May be accepted based on application importance, gauge replacement cost, or customer agreement. | |
| Unacceptable | Incapable measurement system. Measurement error masks true process behavior. Root-cause remediation is mandatory before proceeding. |
Number of Distinct Categories ()
The Number of Distinct Categories () quantifies the measurement system's discrimination relative to process spread:
The resulting value is truncated to an integer:
- : Acceptable measurement system. The gauge has sufficient discrimination to monitor process variation and support control charting.
- : Conditionally acceptable for coarse grouping (e.g., low, medium, high), but inadequate for estimating process capability ().
- : Unacceptable. The gauge acts only as an attribute go/no-go checker and cannot distinguish part-to-part variation.
Attribute Measurement System Analysis (Attribute Agreement Analysis)
When inspection is qualitative (e.g., visual inspection for cosmetic flaws, solder joint pass/fail, or invoice auditing), Attribute Agreement Analysis assesses inspection consistency.
Study Design and Metrics
- Study Structure: 30 to 50 parts with known reference standards, including borderline conforming and defective units, evaluated by 2 to 3 inspectors in randomized order across 2 to 3 trials.
- Within-Appraiser Agreement (Repeatability): Consistency of an individual inspector across repeated evaluations of the same parts.
- Between-Appraiser Agreement (Reproducibility): Agreement among all inspectors evaluating the same parts.
- Appraiser vs. Standard Agreement (Accuracy): Percentage of inspector evaluations matching the certified master standard.
Cohen's and Fleiss' Kappa () Statistics
Because random chance produces some agreement, attribute systems rely on Kappa ():
- : Strong to excellent agreement beyond chance; the visual inspection system is acceptable.
- : Moderate agreement; requires boundary samples ("golden boards") and retraining.
- : Poor agreement; the inspection process is unreliable and requires immediate standard overhaul.
In a variable Gauge R&R study, a project team discovers that the majority of measurement error occurs because three different technicians obtain significantly different average readings while measuring the exact same parts using the same digital micrometer. Which component of measurement variation is responsible, and what is its standard abbreviation?
Equipment Variation (EV)
Part-to-Part Variation (PV)
Appraiser Variation (AV)
Process Linearity Error (LE)
A Yellow Belt conducts a variable Gauge R&R study on a critical CNC turning operation. The statistical output reveals a %GR&R of 34.5% and a Number of Distinct Categories (ndc) of 3. How should the team interpret this measurement system according to standard Six Sigma guidelines?
The measurement system is unacceptable; the gauge cannot reliably distinguish process variation, and measurement improvements must occur before completing the Measure phase
The measurement system is fully acceptable; an ndc value below 5 confirms that the measurement system has sufficient resolution for statistical process control
The measurement system is conditionally acceptable; any %GR&R between 30% and 50% is approved for baseline capability analysis without remediation
The gauge should be approved immediately because high %GR&R values indicate superior process capability (Cp > 2.0)
An assembly plant evaluates its visual inspection process for paint defects using Attribute Agreement Analysis with three inspectors and certified master panels. The calculated Fleiss' Kappa statistic is 0.84. What does this result indicate regarding the visual measurement system?
The inspection system has unacceptable bias and must be completely replaced by an automated optical sensor
Agreement among inspectors is driven entirely by random chance because Kappa exceeds 0.50
The repeatability of individual inspectors is poor, but between-appraiser reproducibility is high
The visual inspection system exhibits strong agreement beyond chance, confirming high consistency among inspectors and master standards
Sections you finish are checked off in the contents.