Bias, Linearity, Stability & P/T
Key Takeaways
- Bias is the average difference between measured values and a reference standard; large bias means accurate-looking precision can still be wrong.
- Linearity checks whether bias changes across the measurement range; stability checks whether measurement performance drifts over time.
- Percent agreement (and related attribute metrics) evaluate go/no-go or rating systems against themselves and against a known standard.
- %P/T (precision-to-tolerance) = (k·σ_measurement / tolerance)×100% with k often 6; under 10% is generally good, 10–30% marginal, over 30% unacceptable.
- Accept a measurement system only when bias, linearity, stability, R&R/%P/T (or attribute agreement), and business risk criteria all support the intended data use.
Bias, Linearity, Stability & P/T
Quick Answer: Beyond Gage R&R, CSSGB BoK III.E (Evaluate) requires judgment on bias, linearity, stability, percent agreement (attribute systems), and precision-to-tolerance (%P/T). Together these metrics answer: Is the system accurate, consistent over the range and over time, and precise enough relative to the specification for the decisions you will make?
Completing the MSA Picture
Gage R&R answers “How much noise does the measurement system add?” It does not fully answer:
- Are we measuring the true value on average? → Bias
- Does error stay constant from low to high readings? → Linearity
- Does performance hold this month the same as last month? → Stability
- For pass/fail gages, do people agree with each other and with truth? → Percent agreement
- Is precision small versus the tolerance? → %P/T
A system can pass R&R yet fail bias or P/T. Evaluate all relevant criteria for the CTQ and risk level.
Bias
Bias is the systematic difference between the average of measurements and a reference (master, calibrated standard, or accepted true value):
Often reported as a percent of reference or of tolerance:
Study approach: Measure a known standard many times (or several standards) under normal use conditions; compare the average to the certified value. Statistical tests can check whether bias differs from zero; for Green Belt practice, know the definition, the sign (high vs. low), and that calibration and method correction address bias, while R&R addresses precision.
| Bias finding | Typical action |
|---|---|
| Near zero vs. master | Accuracy acceptable for that point |
| Large positive bias | System reads high—recalibrate, correct procedure, or apply offset if justified |
| Large negative bias | System reads low—same corrective path |
| Bias changes with operator | Training / standardization issue (links to reproducibility) |
Precision vs. accuracy reminder: Tight clusters far from the master = precise but inaccurate (bias). Loose clusters around the master = accurate on average but imprecise (poor R&R).
Linearity
Linearity asks whether bias is constant across the operating range. A micrometer might be unbiased at mid-range but read high on small parts and low on large parts—disastrous if you use one correction factor everywhere.
Study approach:
- Select several reference standards spanning the full range of intended use (not only the center).
- Measure each multiple times; compute bias at each reference level.
- Plot bias vs. reference value. Fit a line (bias = $a + b \times$ reference).
Interpretation:
- Slope $b \approx 0$ and intercept $a \approx 0$ → good linearity and little fixed bias.
- Significant slope → bias depends on the magnitude of the measurement; fix gage design, fixtures, or use range-specific calibration.
- Scatter around the line → precision issues in addition to linearity.
Linearity problems often appear when a single %bias from one master is used to “clear” a wide product family.
Stability
Stability (measurement stability) is consistency of the measurement system over time. A system that passes R&R on Monday but drifts by Friday produces false trends and false special causes on control charts.
Study approach:
- Measure the same master(s) periodically (shift, day, week) under normal conditions.
- Plot averages (and ranges) on a control chart for the measurement system or track bias over time.
- Look for trends, shifts, cycles (temperature, warm-up, wear, battery, software updates).
Interpretation: Points beyond control limits or a clear drift mean the system is unstable—schedule calibration, environmental controls, preventive maintenance, or method locks. Stability is why MSA is not a one-time certificate; critical gages need ongoing checks in the Control phase.
Percent Agreement (Attribute MSA)
When the “measurement” is a classification (pass/fail, grade A/B/C), continuous R&R formulas do not apply directly. Attribute agreement analysis evaluates:
| Comparison | Question |
|---|---|
| Within appraiser | Does the same person repeat the same decision on the same unit? |
| Between appraisers | Do people agree with each other? |
| Vs. standard | Do decisions match the known correct classification? |
Percent agreement is the percentage of matched decisions in a given comparison (e.g., 45 of 50 trials match the master → 90% agreement).
Organizations set acceptance thresholds based on risk (often high agreement required for safety CTQs). Related statistics (kappa) adjust for agreement by chance; CSSGB items more often stress the concept of agreement with self, peers, and standard, plus the need for clear operational definitions and training when agreement is low.
Improve attribute systems by: better lighting/fixtures, clearer boundary samples (golden units), training on borderline cases, and reducing ambiguous categories.
Precision-to-Tolerance (P/T)
%P/T (also called the precision-to-tolerance ratio) compares measurement variation to the specification width:
- $\sigma_{\text{measurement}}$ is usually $\sigma_{\text{GRR}}$ from the R&R study.
- $k$ is commonly 6 (approximately 99.73% of a normal measurement-error distribution). Some older references use 5.15 (~99%); know that the constant multiplies σ, then divides by tolerance.
Worked example: Tolerance = USL − LSL = 2.0 units; $\sigma_{\text{GRR}} = 0.05$; use $k = 6$:
Common decision bands (aligned with AIAG-style teaching, same spirit as %GR&R):
| %P/T | Typical decision |
|---|---|
| < 10% | Measurement precision generally acceptable vs. tolerance |
| 10%–30% | Marginal — may accept based on criticality, cost, and use of data |
| > 30% | Unacceptable — measurement noise consumes too much of the tolerance |
%P/T vs. %GR&R:
- %GR&R compares measurement noise to observed process/part spread in the study.
- %P/T compares measurement noise to customer/engineering tolerance.
If the study parts barely vary, %GR&R can look terrible even when %P/T is fine. If the process is very wide but the tolerance is tight, %GR&R may look good while %P/T fails. Use both when a two-sided tolerance exists.
Decision Criteria: Accepting a Measurement System
A practical Evaluate-level checklist for CSSGB projects:
- Purpose fit — Will data be used for process control, capability claims, sorting, or engineering experiments? Higher risk → tighter criteria.
- Bias — Average error vs. reference acceptable (near zero; within internal calibration limits).
- Linearity — Bias does not change unacceptably across the range of use.
- Stability — No harmful drift over the relevant time horizon; control plan includes rechecks.
- Variables R&R — %GR&R preferably < 10% (or justified marginal 10–30%); ndc ≥ 5 for process analysis.
- %P/T — Preferably < 10% (or justified marginal) when tolerance is defined.
- Attribute — Percent agreement (and effectiveness vs. standard) meet risk-based thresholds; operational definitions clear.
- Business action — If any critical criterion fails: improve fixture/method, retrain, change gage technology, or change how the CTQ is measured—do not proceed as if the numbers were process truth.
| Scenario | Decision |
|---|---|
| %GR&R = 7%, %P/T = 8%, bias ≈ 0, stable, ndc = 8 | Accept for most process uses |
| %GR&R = 12%, %P/T = 28%, critical safety CTQ | Treat as inadequate; improve before capability claims |
| %GR&R = 9%, %P/T = 35% | Precision OK vs. process spread, not vs. tolerance—fix gage or widen understanding of tolerance risk |
| Attribute agreement vs. standard = 70% | Unacceptable for high-risk sorting; train and clarify standards |
| Excellent R&R last year, masters now drifting on charts | Stability failed—recalibrate and restore control |
Common Exam Traps
- Equating calibration sticker current with full MSA acceptance.
- Reporting only %GR&R when the question gives tolerance and asks about P/T.
- Confusing process stability (SPC on the product) with measurement stability (masters over time).
- Applying %P/T to attribute pass/fail data without a continuous σ.
- Ignoring linearity because one mid-range master showed low bias.
Bottom line: Bias, linearity, stability, percent agreement, and %P/T—together with Gage R&R—let a Green Belt evaluate whether the measurement system is acceptable. Only then should DMAIC treat the numbers as a trustworthy voice of the process.
A master standard certified at 10.000 mm is measured 25 times. The average of the readings is 10.040 mm. What is the bias, and what does it indicate?
A CTQ has USL − LSL = 1.20. From MSA, σ_GRR = 0.04. Using k = 6, what is %P/T and the usual classification?