Proficiency Comparisons and Gage R&R

Key Takeaways

  • Normalized error compares a result difference with uncertainties under the stated correlation and coverage assumptions.

  • A satisfactory comparison score is evidence for that comparison, not automatic validation of every laboratory capability.

  • Gage R&R acceptance bands and distinct-category goals must be interpreted for the study and intended use.

Last updated: October 2026

Interlaboratory Comparisons (ILC) and Proficiency Testing (PT)

ISO/IEC 17025 clause 7.7.2 calls for comparison with other laboratories when available and appropriate, through proficiency testing and/or interlaboratory comparisons. Select suitable activities and interpret the results with their scope and uncertainty; one successful comparison does not validate every CMC or method.

In a typical PT scheme, a stable artifact (the transfer standard) is circulated among participating laboratories. Each laboratory measures the artifact and reports its measured value along with its expanded measurement uncertainty (k=2k=2).

The Normalized Error (EnE_n) Ratio

Performance is evaluated using the internationally standardized Normalized Error (EnE_n) ratio (ISO/IEC 17043):

En=xlab−xrefUlab2+Uref2E_n = \frac{x_{\text{lab}} - x_{\text{ref}}}{\sqrt{U_{\text{lab}}^2 + U_{\text{ref}}^2}}

where:

  • xlabx_{\text{lab}} = The measured value reported by the participant laboratory.
  • xrefx_{\text{ref}} = The assigned reference value established by the organizing reference laboratory (typically an NMI or accredited primary reference provider).
  • UlabU_{\text{lab}} = The expanded measurement uncertainty (k=2k=2, approximately 95%95\% confidence) declared by the participant laboratory.
  • UrefU_{\text{ref}} = The expanded measurement uncertainty (k=2k=2) associated with the assigned reference value.

Interpretation and Decision Criteria

{∣En∣≤1.0  ⟹  Satisfactory Performance (Pass)∣En∣>1.0  ⟹  Unsatisfactory Performance (Fail / Action Required)\begin{cases} |E_n| \le 1.0 & \implies \textbf{Satisfactory Performance (Pass)} \\ |E_n| > 1.0 & \implies \textbf{Unsatisfactory Performance (Fail / Action Required)} \end{cases}

For independent laboratory and reference values with comparable expanded-uncertainty coverage, En=(xlab−xref)/Ulab2+Uref2E_n=(x_{lab}-x_{ref})/\sqrt{U_{lab}^2+U_{ref}^2}. Correlated results require the applicable covariance treatment. An absolute score at most one supports agreement for that comparison; it does not validate every uncertainty budget or the entire accredited scope. An unfavorable score requires investigation and risk-based action, not an automatic scope-wide suspension.


Measurement Systems Analysis (MSA) and Gage R&R Studies

While MAP monitors calibration laboratory standards, Measurement Systems Analysis (MSA) evaluates the capability of inspection equipment deployed on the production floor. Governed by the Automotive Industry Action Group (AIAG) guidelines, Gage Repeatability and Reproducibility (Gage R&R) studies evaluate total measurement system variation.

Partitioning Variance: Repeatability vs. Reproducibility

Total observed variation in an inspection process is decomposed into distinct variances:

σTotal2=σPart2+σGRR2\sigma_{\text{Total}}^2 = \sigma_{\text{Part}}^2 + \sigma_{\text{GRR}}^2 σGRR2=σRepeatability2+σReproducibility2=EV2+AV2\sigma_{\text{GRR}}^2 = \sigma_{\text{Repeatability}}^2 + \sigma_{\text{Reproducibility}}^2 = \text{EV}^2 + \text{AV}^2
  1. Repeatability (Equipment Variation / EV): The variation in measurements obtained with one measurement instrument when used several times by one appraiser while measuring the identical characteristic on the same parts. Reflects the inherent mechanical, electrical, and sensor repeatability of the tool.
  2. Reproducibility (Appraiser Variation / AV): The variation in the average of measurements made by different appraisers using the same instrument when measuring the identical characteristic on the same parts. Reflects human technique, clamping pressure, visual alignment, and ergonomic differences between operators.

AIAG Acceptance Criteria for %GRR

Gage capability is evaluated as a percentage of Total Variation (%TV) or as a percentage of product tolerance span (%Tolerance):

%GRR=(σGRRσTotal)×100%or%GRRTol=(6 σGRRUSL−LSL)×100%\%\text{GRR} = \left(\frac{\sigma_{\text{GRR}}}{\sigma_{\text{Total}}}\right) \times 100\% \quad \text{or} \quad \%\text{GRR}_{\text{Tol}} = \left(\frac{6\,\sigma_{\text{GRR}}}{\text{USL} - \text{LSL}}\right) \times 100\%
Common guidelineInterpretationQualification
Below 10%Often considered acceptableConfirm the task and other performance requirements.
10–30%May be conditionally acceptableConsider application, consequences, and policy.
Above 30%Usually calls for improvementFollow the agreed study and acceptance policy.

These bands do not universally authorize or prohibit every use. State whether the denominator is tolerance or study variation, and use the agreed multiplier (for example six standard deviations).

Number of distinct categories

The usual study estimate is ndc=1.41σPart/σGRRndc=1.41\sigma_{Part}/\sigma_{GRR}, with the result truncated to an integer under the relevant convention. It compares part variation with measurement variation. Values of five or more are a commonly used study goal, but suitability still depends on the intended task. A low value warns that the instrument may not resolve the part-to-part differences being studied. Do not treat the number as a universal guarantee for every control-chart application.

Control-chart signal rules belong to the check-standard lesson. Distinguish a time-sequence signal from a proficiency-comparison score or a gage-study acceptance criterion; they test different aspects of measurement performance.

Worked study interpretation

Assume the agreed study uses six standard deviations divided by tolerance. If measurement-system standard deviation is 0.004 mm and total tolerance is 0.200 mm, %GRR is 6 × 0.004 / 0.200 × 100 = 12%. Under a policy using the common bands, this calls for conditional assessment rather than automatic approval. The decision also considers bias, stability, linearity, consequences, and the intended inspection. If part standard deviation is 0.020 mm, ndc is 1.41 × 0.020 / 0.004 = 7.05, truncated to 7. A study denominator based on observed process variation can yield a different percentage than the tolerance denominator; state the basis so the interpretation is reproducible.

Designing an interpretable gage study

Select parts representing the intended process variation and operators representing actual use. Randomize or otherwise control the order so memory, drift, and fatigue do not masquerade as repeatability or operator differences. Repeatedly measuring one nearly identical master can evaluate a restricted repeatability condition, but cannot establish the full part-to-part variation needed for a representative distinct-category estimate.

Review operator-by-part interaction when the study design supports it. An operator may be consistent on smooth parts yet inconsistent on rough or awkward features. Such an interaction should not be hidden by averaging all observations into one number. If the study prompts a change to the fixture or procedure, repeat suitable measurements under the changed conditions and record the change. Compare like denominators when tracking improvement.

Test Your Knowledge

In a Proficiency Testing (PT) interlaboratory comparison, a participant calibration laboratory reports a measured value of 100.008 g with an expanded uncertainty (k=2) of 0.004 g. The accredited reference laboratory assigns a reference value of 100.002 g with an expanded uncertainty of 0.003 g. What is the normalized error (En) ratio, and how is the laboratory's performance evaluated? Assume uncorrelated results and comparable k = 2 coverage.

A

En = 0.86; satisfactory because it is less than 1.0

B

En = 1.20; unsatisfactory because |En| > 1.0

C

En = 0.60; satisfactory because |En| <= 1.0

D

En = 2.00; satisfactory because it falls within a 95% confidence interval

Test Your Knowledge

A plant’s approved gage-study policy prohibits acceptance inspection when %GRR exceeds 30% of tolerance. A bore-gauge study yields 34%. What follows?

A

Widen product tolerances without engineering review

B

Treat the gauge as acceptable because 34% is below 50%

C

Delete the least repeatable operator’s readings

D

Hold the gauge from that acceptance task and investigate measurement variation under the policy

Sections you finish are checked off in the contents.