Proficiency Comparisons and Gage R&R
Key Takeaways
Normalized error compares a result difference with uncertainties under the stated correlation and coverage assumptions.
A satisfactory comparison score is evidence for that comparison, not automatic validation of every laboratory capability.
Gage R&R acceptance bands and distinct-category goals must be interpreted for the study and intended use.
Interlaboratory Comparisons (ILC) and Proficiency Testing (PT)
ISO/IEC 17025 clause 7.7.2 calls for comparison with other laboratories when available and appropriate, through proficiency testing and/or interlaboratory comparisons. Select suitable activities and interpret the results with their scope and uncertainty; one successful comparison does not validate every CMC or method.
In a typical PT scheme, a stable artifact (the transfer standard) is circulated among participating laboratories. Each laboratory measures the artifact and reports its measured value along with its expanded measurement uncertainty ().
The Normalized Error () Ratio
Performance is evaluated using the internationally standardized Normalized Error () ratio (ISO/IEC 17043):
where:
- = The measured value reported by the participant laboratory.
- = The assigned reference value established by the organizing reference laboratory (typically an NMI or accredited primary reference provider).
- = The expanded measurement uncertainty (, approximately confidence) declared by the participant laboratory.
- = The expanded measurement uncertainty () associated with the assigned reference value.
Interpretation and Decision Criteria
For independent laboratory and reference values with comparable expanded-uncertainty coverage, . Correlated results require the applicable covariance treatment. An absolute score at most one supports agreement for that comparison; it does not validate every uncertainty budget or the entire accredited scope. An unfavorable score requires investigation and risk-based action, not an automatic scope-wide suspension.
Measurement Systems Analysis (MSA) and Gage R&R Studies
While MAP monitors calibration laboratory standards, Measurement Systems Analysis (MSA) evaluates the capability of inspection equipment deployed on the production floor. Governed by the Automotive Industry Action Group (AIAG) guidelines, Gage Repeatability and Reproducibility (Gage R&R) studies evaluate total measurement system variation.
Partitioning Variance: Repeatability vs. Reproducibility
Total observed variation in an inspection process is decomposed into distinct variances:
- Repeatability (Equipment Variation / EV): The variation in measurements obtained with one measurement instrument when used several times by one appraiser while measuring the identical characteristic on the same parts. Reflects the inherent mechanical, electrical, and sensor repeatability of the tool.
- Reproducibility (Appraiser Variation / AV): The variation in the average of measurements made by different appraisers using the same instrument when measuring the identical characteristic on the same parts. Reflects human technique, clamping pressure, visual alignment, and ergonomic differences between operators.
AIAG Acceptance Criteria for %GRR
Gage capability is evaluated as a percentage of Total Variation (%TV) or as a percentage of product tolerance span (%Tolerance):
| Common guideline | Interpretation | Qualification |
|---|---|---|
| Below 10% | Often considered acceptable | Confirm the task and other performance requirements. |
| 10–30% | May be conditionally acceptable | Consider application, consequences, and policy. |
| Above 30% | Usually calls for improvement | Follow the agreed study and acceptance policy. |
These bands do not universally authorize or prohibit every use. State whether the denominator is tolerance or study variation, and use the agreed multiplier (for example six standard deviations).
Number of distinct categories
The usual study estimate is , with the result truncated to an integer under the relevant convention. It compares part variation with measurement variation. Values of five or more are a commonly used study goal, but suitability still depends on the intended task. A low value warns that the instrument may not resolve the part-to-part differences being studied. Do not treat the number as a universal guarantee for every control-chart application.
Control-chart signal rules belong to the check-standard lesson. Distinguish a time-sequence signal from a proficiency-comparison score or a gage-study acceptance criterion; they test different aspects of measurement performance.
Worked study interpretation
Assume the agreed study uses six standard deviations divided by tolerance. If measurement-system standard deviation is 0.004 mm and total tolerance is 0.200 mm, %GRR is 6 × 0.004 / 0.200 × 100 = 12%. Under a policy using the common bands, this calls for conditional assessment rather than automatic approval. The decision also considers bias, stability, linearity, consequences, and the intended inspection. If part standard deviation is 0.020 mm, ndc is 1.41 × 0.020 / 0.004 = 7.05, truncated to 7. A study denominator based on observed process variation can yield a different percentage than the tolerance denominator; state the basis so the interpretation is reproducible.
Designing an interpretable gage study
Select parts representing the intended process variation and operators representing actual use. Randomize or otherwise control the order so memory, drift, and fatigue do not masquerade as repeatability or operator differences. Repeatedly measuring one nearly identical master can evaluate a restricted repeatability condition, but cannot establish the full part-to-part variation needed for a representative distinct-category estimate.
Review operator-by-part interaction when the study design supports it. An operator may be consistent on smooth parts yet inconsistent on rough or awkward features. Such an interaction should not be hidden by averaging all observations into one number. If the study prompts a change to the fixture or procedure, repeat suitable measurements under the changed conditions and record the change. Compare like denominators when tracking improvement.
In a Proficiency Testing (PT) interlaboratory comparison, a participant calibration laboratory reports a measured value of 100.008 g with an expanded uncertainty (k=2) of 0.004 g. The accredited reference laboratory assigns a reference value of 100.002 g with an expanded uncertainty of 0.003 g. What is the normalized error (En) ratio, and how is the laboratory's performance evaluated? Assume uncorrelated results and comparable k = 2 coverage.
En = 0.86; satisfactory because it is less than 1.0
En = 1.20; unsatisfactory because |En| > 1.0
En = 0.60; satisfactory because |En| <= 1.0
En = 2.00; satisfactory because it falls within a 95% confidence interval
A plant’s approved gage-study policy prohibits acceptance inspection when %GRR exceeds 30% of tolerance. A bore-gauge study yields 34%. What follows?
Widen product tolerances without engineering review
Treat the gauge as acceptable because 34% is below 50%
Delete the least repeatable operator’s readings
Hold the gauge from that acceptance task and investigate measurement variation under the policy
Sections you finish are checked off in the contents.