Outlier Review and Worked Type A Analysis
Key Takeaways
A Grubbs test assumes an appropriate approximately normal model and does not alone authorize data deletion.
Preserve original observations and document the evidence and method supporting any exclusion or repeat run.
Use the spread of individual observations or the mean according to the actual reported measurement model.
Outlier Detection and Metrological Data Rejection Rules
During a series of repeated measurements, an occasional observation may appear widely disparate from the remaining data. Metrologists use formal statistical tests to determine whether an anomalous point is statistically significant.
Grubbs' Test
Grubbs' test (also known as the maximum normed residual test) is a commonly used method for detecting a single outlier in a sample that approximately follows a normal distribution.
Test Procedure:
Step 1. Calculate the sample mean and sample standard deviation including the suspected outlier. Step 2. Compute the Grubbs test statistic :
Step 3. Compare calculated with the critical value from the Grubbs table for sample size at significance level (or ):
- If , the suspected reading is statistically classified as an outlier at the selected significance level under the test assumptions.
| Sample Size () | Critical Value () | Critical Value () |
|---|---|---|
| 3 | 1.1543 | 1.1547 |
| 4 | 1.4812 | 1.4962 |
| 5 | 1.7150 | 1.7637 |
| 6 | 1.8871 | 1.9728 |
| 7 | 2.0200 | 2.1391 |
| 8 | 2.1266 | 2.2744 |
| 9 | 2.2150 | 2.3868 |
| 10 | 2.2900 | 2.4821 |
The table above is for the two-sided maximum-absolute-residual test. A preselected upper- or lower-tail test uses different critical values. Do not mix a one-sided table with the two-sided statistic.
Dixon's Q Test
Dixon's Q test is a quick ranking method used for small sample sizes ():
Step 1. Arrange the data in ascending order: . Step 2. Identify the suspect value (either or ) and compute the experimental ratio :
Step 3. Compare against tabulated . If , the reading is statistically anomalous.
ISO/IEC 17025 Compliance: Permitted vs. Prohibited Data Rejection
Caution
The Strict Metrological Rule on Data Rejection: A statistical test (Grubbs or Dixon) never gives a calibration technician the authority to discard a data point on its own! Statistical tests only establish that a data point is improbable under the assumed normal distribution.
Under ISO/IEC 17025 accreditation requirements and good metrological practice:
-
When Data Rejection Is Prohibited:
- A technician cannot delete a reading simply because it "ruins the uncertainty budget" or causes the test uncertainty ratio (TUR) to fall below .
- A technician cannot silently discard a reading merely because a statistic flags it; investigate the evidence and follow a justified, documented method.
- Undocumented "cherry-picking" to manufacture a favorable calibration outcome is data falsification.
-
When Data Rejection Is Permitted:
- A data point may be excluded under a justified documented procedure, for example when an assignable physical cause is verified and recorded on the calibration worksheet.
- Permissible assignable causes include: an observed electrical line voltage transient, an acoustic disturbance or floor vibration during an optical measurement, a confirmed speck of dust or lint on a gage block face, operator error in recording a digit, or an unseated connector.
- Remedial Action: Preserve original observations and investigate flagged points. If evidence establishes an invalid observation, document the exclusion and follow the approved replacement or repeat-run method. Statistical procedures can support documented treatment under their assumptions; a flag alone is not permission to discard an inconvenient result.
Step-by-Step Worked Calibration Example: Type A Analysis
A calibration technician calibrates a precision digital voltmeter by applying a nominal DC reference standard from a Josephson-traceable calibrator. The technician records repeated voltage indications under repeatability conditions:
Notice reading appears unusually high. Let us perform the complete Type A analysis and outlier investigation.
Step 1: Compute Sample Mean ()
Step 2: Compute Residuals and Sum of Squared Deviations
| Observation | Reading (V) | Residual (µV) | Squared residual (µV²) |
|---|---|---|---|
| 1 | 10.00012 | -20 | 400 |
| 2 | 10.00015 | 10 | 100 |
| 3 | 10.00009 | -50 | 2500 |
| 4 | 10.00014 | 0 | 0 |
| 5 | 10.00011 | -30 | 900 |
| 6 | 10.00016 | 20 | 400 |
| 7 | 10.00013 | -10 | 100 |
| 8 | 10.00010 | -40 | 1600 |
| 9 | 10.00014 | 0 | 0 |
| 10 | 10.00026 | 120 | 14400 |
| Sum | 100.00140 | 0 | 20400 |
The squared-residual sum is .
Step 3: Compute Sample Variance () and Sample Standard Deviation ()
Step 4: Perform Grubbs' Outlier Test on
Suspect value: . Residual magnitude: .
From the Grubbs critical value table for at :
Because , observation is statistically flagged as an outlier at significance level under the two-sided test.
- Hypothetical investigation: Assume synchronized instrument diagnostics and electrical monitoring establish that interference invalidated observation 10. A coincident HVAC start alone would not prove causation. Preserve original data and follow the approved documented exclusion and replacement or repeat-run method. The nine-reading calculation below illustrates arithmetic after justified exclusion, not permission to trim an inconvenient result.
Step 5: Recalculate Statistics for the Valid Sample ()
Sum of the remaining 9 readings: .
Step 6: Compute Standard Uncertainty of the Mean
Because the calibration certificate will report the mean of the 9 readings as the calibrated voltage:
Notice the metrological impact:
- The sample standard deviation represents the process repeatability (dispersion of single voltmeter readings).
- The standard uncertainty of the mean is the Type A standard uncertainty that enters the combined uncertainty budget for the reported mean value.
Official references (checked October 10, 2026): NIST Grubbs test definition and critical-value formula.
A Grubbs test flags one of ten mass-comparator readings under an assumed normal model. What is the appropriate initial response?
Automatically delete reading 7 from the dataset, recalculate the mean and variance for n = 9, and issue the calibration certificate
Replace reading 7 with the arithmetic average of the remaining 9 readings to preserve the degrees of freedom
Preserve the original observations, investigate the point and test assumptions, and document any justified exclusion or repeat-run decision.
Apply Dixon's Q test to override Grubbs' test and average the two resulting test statistics
Sections you finish are checked off in the contents.