4.5 Gage Correlation & Cross-Instrument Agreement

Key Takeaways

  • Calibration compares one instrument to a standard, gage R&R quantifies one measurement system’s own variation, and correlation compares two measurement systems to each other — they answer three different questions and cannot substitute for one another.
  • A defensible correlation study uses at least 10 parts whose values span the full tolerance, measured on both systems in a randomized, blind order, with repeats on each system.
  • The primary correlation output is the average signed difference (the offset or bias between systems), not a correlation coefficient; a Pearson r near 1.00 can coexist with a large constant offset.
  • A common acceptance criterion is that the between-system offset consume no more than about 10% of the feature tolerance, with the residual scatter small relative to the combined gage R&R.
  • When correlation fails, the resolution is to find the physical cause — different datums, filters, probe force, temperature, or evaluation algorithms — not to apply a fudge factor to make the numbers agree.
Last updated: September 2026

The Problem Correlation Solves

A receiving inspector rejects an incoming lot of housings for a bore that measures 1.2516 in on the shop's air gage against a 1.2500–1.2510 in requirement. The supplier's certificate of conformance reports 1.2505 in, measured on their coordinate measuring machine. Both instruments hold current calibration stickers. Both operators are trained. The parts have not changed between the two measurements.

Nothing in calibration or in a gage repeatability and reproducibility study answers this. Calibration proved each instrument agrees with its own standards. Gage R&R proved each instrument is repeatable within itself. Neither one asked the question that matters here: do these two measurement systems agree with each other on the same parts? That is what a gage correlation study determines, and Body of Knowledge topic II.C.4 requires you to be able to apply the method.

Three Studies, Three Questions

StudyQuestion it answersReferenceTypical output
CalibrationDoes this instrument agree with a traceable standard?A higher-echelon standardAs-found and as-left errors, uncertainty
Gage R&R (MSA)How much variation does this measurement system add?The parts themselves%GRR, ndc, appraiser and equipment variance
CorrelationDo two measurement systems agree with each other?The other systemAverage offset, scatter, agreement limits

Mixing these up is one of the most reliable exam traps in the metrology domain. A gage can be perfectly calibrated, have an excellent 6% GRR, and still disagree with a customer's CMM by 0.0008 in.


When a Correlation Study Is Required

  • Gage-to-gage. Two nominally identical instruments — two bench comparators, two hardness testers, two thread gages — are used on the same characteristic at different stations or on different shifts.
  • Manual-to-automated. The BoK calls this out by name. First article is measured by hand with micrometers and a height gage; production is monitored by an automated vision system or an in-line air gage. The two methods must be shown to agree before the automated method may be trusted for acceptance.
  • Customer-to-supplier. Part of a production part approval submission, or the standing agreement that decides who wins a dimensional dispute.
  • Site-to-site. The same part number inspected at two plants must be dispositioned the same way at both.
  • Before and after a change. New probe, new software version, new fixture, new operator population, relocation of the instrument.
  • After a correlation failure in the field. A customer rejection on a characteristic the shop passed is a mandatory trigger.

Designing the Study

A correlation study is a designed comparison, and the design choices determine whether the result means anything.

  1. Select parts that span the range. Use at least 10 parts — more is better — deliberately chosen so their true values cover the full tolerance band and ideally extend slightly beyond both limits. A study run on 10 parts clustered at nominal can only prove the systems agree at nominal, which is precisely where disagreement does not matter.
  2. Uniquely and permanently identify each part. Etch, engrave, or tag serial numbers. The single most common way a correlation study is invalidated is losing track of which part is which between the two measurement sessions.
  3. Define the measurement completely. Same feature, same datums, same location on the feature, same probing or contact points, same environmental soak. If the CMM evaluates a bore from 200 scanned points using a least-squares fit and the air gage reads a two-jet cross-section at one depth, you are not measuring the same thing and no amount of arithmetic will make the results agree.
  4. Randomize and blind. Present parts to each system in a different randomized order, and keep operators from seeing the other system's result. Knowing the expected number is a direct route to the confirmation bias covered in inspection-error training.
  5. Repeat. Measure each part two or three times on each system. Repeats separate genuine between-system offset from ordinary within-system noise.
  6. Record conditions. Temperature, humidity, operator, instrument serial number, software version, program revision, and date belong on the data sheet.

Analyzing the Data

Step 1: Average each system's readings per part

For each part, average the repeats within each system so each part contributes one value per system.

Step 2: Compute the per-part difference

di=Xi,System AXi,System Bd_i = X_{i,\text{System A}} - X_{i,\text{System B}}

Keep the sign. Discarding signs by taking absolute values destroys the offset information, which is the whole point of the study.

Step 3: Compute the average difference and its spread

The mean of the differences is the offset, also called the between-system bias: dˉ=1ni=1ndi\bar{d} = \frac{1}{n}\sum_{i=1}^{n} d_i

The standard deviation of the differences describes how consistent that offset is from part to part.

Step 4: Plot before you conclude

Two plots carry almost all of the information:

  • A scatter plot of System A versus System B with a 45-degree reference line. Points sitting parallel to but above the line reveal a constant offset. Points fanning away from the line as values increase reveal a proportional (slope) error, which usually means a scale or magnification difference.
  • A difference plot — each part's difference plotted against the average of the two systems. A flat cloud centered on zero means the systems agree. A tilted cloud means the offset changes with size. A widening cloud means agreement degrades at one end of the range.

Step 5: Judge against a criterion, not against a feeling

A widely used acceptance rule is that the between-system offset should consume no more than about 10% of the feature tolerance, and the scatter of the differences should be small compared with the combined repeatability of the two systems. On a 0.010 in tolerance, an offset of 0.0010 in is at the limit and an offset of 0.0025 in clearly fails.

The Correlation Coefficient Trap

Inspectors often report the Pearson correlation coefficient (r) and stop there. This is a serious analytical error and a favorite exam distractor. Correlation measures whether the two systems rank and track parts together, not whether they agree. If System A reads exactly 0.0030 in higher than System B on every single part, the two data sets have a perfect r = 1.000 while the systems disagree by an amount that would scrap the lot. High r with a large offset is a total agreement failure that looks like a success. Always report the average signed difference alongside any correlation statistic.


Worked Example

A shop correlates its bench comparator (System A) against the customer's CMM (System B) on a shaft diameter toleranced 0.7500–0.7520 in, a tolerance band of 0.0020 in.

PartComparator (A)CMM (B)Difference (A − B)
10.750420.75009+0.00033
20.750980.75067+0.00031
30.751310.75096+0.00035
40.751600.75131+0.00029
50.751860.75152+0.00034
60.752030.75171+0.00032

The differences average +0.00032 in and vary only from +0.00029 to +0.00035 in. The Pearson r for these two columns is essentially 1.000, which by itself looks excellent. But the offset consumes 0.00032 / 0.0020 = 16% of the tolerance, exceeding a 10% criterion, and it is consistently positive.

The consistency of the offset is diagnostic: a constant shift with no slope component points to a datum, setup, or mastering difference, not to random error. The investigation finds that the comparator was mastered with a gage block stack that had not been temperature-soaked, so it was reading a slightly expanded master. Re-mastering after a full soak collapses the offset to +0.00004 in. Note that the fix was physical — the study located a real cause. The wrong response would have been to program a −0.00032 in constant into the comparator to force agreement.


When Correlation Fails: Root Causes to Hunt

Symptom on the difference plotLikely physical cause
Constant offset, no slopeMastering error, zero-setting error, different datum origin, uncorrected temperature
Offset grows with size (slope)Scale factor, encoder or magnification error, thermal expansion coefficient wrong
Large scatter, no offsetOne system has poor repeatability; run gage R&R on each system separately
Offset only on some partsForm error interacting with different sampling strategies; two-point versus full-profile evaluation
Offset appears after software updateChanged fit algorithm, filter cutoff, or outlier-rejection setting

Two contact-method causes deserve specific mention because they generate exam questions. Probe or contact force differs between a hand micrometer, a bench comparator, and a CMM touch-trigger probe, and on compliant materials that force difference is measurable. Evaluation algorithm differs too: a bore evaluated by a least-squares circle, a minimum-circumscribed circle, and a two-point air-plug reading will return three different numbers on an out-of-round bore, and all three are correct for what they measure.


Round Robins and Interlaboratory Comparison

When more than two systems are involved — several plants, several suppliers, or an accredited lab network — the same logic scales up into a round robin or interlaboratory comparison. A single set of identified artifacts circulates on a fixed schedule; each participant measures and reports without seeing others' results; a coordinator compiles the data and computes each participant's deviation from the consensus or reference value. Laboratories operating to ISO/IEC 17025 participate in these proficiency-testing schemes as evidence of ongoing competence, and the same technique works perfectly well between a customer and its suppliers.


Documenting the Result

A correlation study is a quality record and must stand on its own years later. It should state the parts and their identification, both measurement systems including instrument serial numbers and software or program revisions, the operators, environmental conditions, the raw data, the computed offset and its spread, the acceptance criterion, the pass or fail conclusion, and any corrective action taken. In a customer-supplier context, both parties sign it, and it becomes the agreed method of record for that characteristic.


Common Exam Traps

  • Correlation is not calibration. Two instruments can correlate beautifully with each other and both be wrong relative to national standards. Correlation never establishes traceability.
  • Correlation is not gage R&R. Correlation compares systems; R&R characterizes one system's internal variation. A correlation study on a system with 45% GRR is meaningless because the noise swamps the offset you are trying to detect.
  • A high correlation coefficient does not mean agreement. Report the average signed difference.
  • Never "correct" a disagreement with an undocumented offset. Adjusting a gage to force agreement without finding the cause hides a real measurement problem and is a nonconformance in any audited quality system.
  • Same characteristic, same definition. If the two systems evaluate the feature differently — different datums, different fit algorithm, different number of points — they were never measuring the same quantity and correlation was doomed before the first part was measured.
Test Your Knowledge

A correlation study between a shop bench gage and a customer CMM on 12 shafts yields a Pearson correlation coefficient of 0.9998, and the shop gage reads an average of 0.0018 in higher than the CMM on every part. The feature tolerance is 0.004 in. How should the inspector interpret this result?

A
B
C
D
Test Your Knowledge

A quality engineer plans a manual-to-automated correlation study comparing hand micrometer readings against a new in-line laser scan micrometer. Which sampling approach produces the most defensible study?

A
B
C
D
Test Your Knowledge

Two hardness testers in the same lab both hold current calibration certificates traceable to NIST and each shows an acceptable gage R&R. A correlation study on the same set of test blocks nevertheless shows Tester 2 reading consistently 1.8 HRC higher than Tester 1. What does this result establish?

A
B
C
D