Outlier Review and Worked Type A Analysis

Key Takeaways

  • A Grubbs test assumes an appropriate approximately normal model and does not alone authorize data deletion.

  • Preserve original observations and document the evidence and method supporting any exclusion or repeat run.

  • Use the spread of individual observations or the mean according to the actual reported measurement model.

Last updated: October 2026

Outlier Detection and Metrological Data Rejection Rules

During a series of repeated measurements, an occasional observation may appear widely disparate from the remaining data. Metrologists use formal statistical tests to determine whether an anomalous point is statistically significant.

Grubbs' Test

Grubbs' test (also known as the maximum normed residual test) is a commonly used method for detecting a single outlier in a sample that approximately follows a normal distribution.

Test Procedure:

Step 1. Calculate the sample mean xˉ\bar{x} and sample standard deviation ss including the suspected outlier. Step 2. Compute the Grubbs test statistic GG:

G=max⁡i=1,…,n∣xi−xˉ∣s=∣xsuspect−xˉ∣sG = \frac{\max_{i=1,\dots,n} |x_i - \bar{x}|}{s} = \frac{|x_{\text{suspect}} - \bar{x}|}{s}

Step 3. Compare calculated GG with the critical value GcritG_{\text{crit}} from the Grubbs table for sample size nn at significance level α=0.05\alpha = 0.05 (or α=0.01\alpha = 0.01):

  • If G>GcritG > G_{\text{crit}}, the suspected reading is statistically classified as an outlier at the selected significance level under the test assumptions.
Sample Size (nn)Critical Value GcritG_{\text{crit}} (α=0.05\alpha = 0.05)Critical Value GcritG_{\text{crit}} (α=0.01\alpha = 0.01)
31.15431.1547
41.48121.4962
51.71501.7637
61.88711.9728
72.02002.1391
82.12662.2744
92.21502.3868
102.29002.4821

The table above is for the two-sided maximum-absolute-residual test. A preselected upper- or lower-tail test uses different critical values. Do not mix a one-sided table with the two-sided statistic.

Dixon's Q Test

Dixon's Q test is a quick ranking method used for small sample sizes (3≤n≤103 \le n \le 10):

Step 1. Arrange the data in ascending order: x1≤x2≤⋯≤xnx_1 \le x_2 \le \dots \le x_n. Step 2. Identify the suspect value (either x1x_1 or xnx_n) and compute the experimental ratio QcalcQ_{\text{calc}}:

Qcalc=∣xsuspect−xnearest neighbor∣xmax⁡−xmin⁡=GapRangeQ_{\text{calc}} = \frac{|x_{\text{suspect}} - x_{\text{nearest neighbor}}|}{x_{\max} - x_{\min}} = \frac{\text{Gap}}{\text{Range}}

Step 3. Compare QcalcQ_{\text{calc}} against tabulated QcritQ_{\text{crit}}. If Qcalc>QcritQ_{\text{calc}} > Q_{\text{crit}}, the reading is statistically anomalous.

ISO/IEC 17025 Compliance: Permitted vs. Prohibited Data Rejection

Caution

The Strict Metrological Rule on Data Rejection: A statistical test (Grubbs or Dixon) never gives a calibration technician the authority to discard a data point on its own! Statistical tests only establish that a data point is improbable under the assumed normal distribution.

Under ISO/IEC 17025 accreditation requirements and good metrological practice:

  1. When Data Rejection Is Prohibited:

    • A technician cannot delete a reading simply because it "ruins the uncertainty budget" or causes the test uncertainty ratio (TUR) to fall below 4:14:1.
    • A technician cannot silently discard a reading merely because a statistic flags it; investigate the evidence and follow a justified, documented method.
    • Undocumented "cherry-picking" to manufacture a favorable calibration outcome is data falsification.
  2. When Data Rejection Is Permitted:

    • A data point may be excluded under a justified documented procedure, for example when an assignable physical cause is verified and recorded on the calibration worksheet.
    • Permissible assignable causes include: an observed electrical line voltage transient, an acoustic disturbance or floor vibration during an optical measurement, a confirmed speck of dust or lint on a gage block face, operator error in recording a digit, or an unseated connector.
    • Remedial Action: Preserve original observations and investigate flagged points. If evidence establishes an invalid observation, document the exclusion and follow the approved replacement or repeat-run method. Statistical procedures can support documented treatment under their assumptions; a flag alone is not permission to discard an inconvenient result.

Step-by-Step Worked Calibration Example: Type A Analysis

A calibration technician calibrates a precision digital voltmeter by applying a nominal 10.00000 V10.00000\text{ V} DC reference standard from a Josephson-traceable calibrator. The technician records n=10n = 10 repeated voltage indications under repeatability conditions:

x1=10.00012 Vx2=10.00015 Vx3=10.00009 Vx4=10.00014 Vx5=10.00011 Vx6=10.00016 Vx7=10.00013 Vx8=10.00010 Vx9=10.00014 Vx10=10.00026 V\begin{matrix} x_1 = 10.00012\text{ V} & x_2 = 10.00015\text{ V} & x_3 = 10.00009\text{ V} & x_4 = 10.00014\text{ V} & x_5 = 10.00011\text{ V} \\ x_6 = 10.00016\text{ V} & x_7 = 10.00013\text{ V} & x_8 = 10.00010\text{ V} & x_9 = 10.00014\text{ V} & x_{10} = 10.00026\text{ V} \end{matrix}

Notice reading x10=10.00026 Vx_{10} = 10.00026\text{ V} appears unusually high. Let us perform the complete Type A analysis and outlier investigation.

Step 1: Compute Sample Mean (xˉ\bar{x})

∑i=110xi=100.00140 V\sum_{i=1}^{10} x_i = 100.00140\text{ V} xˉ=100.00140 V10=10.000140 V\bar{x} = \frac{100.00140\text{ V}}{10} = 10.000140\text{ V}

Step 2: Compute Residuals and Sum of Squared Deviations

ObservationReading (V)Residual (µV)Squared residual (µV²)
110.00012-20400
210.0001510100
310.00009-502500
410.0001400
510.00011-30900
610.0001620400
710.00013-10100
810.00010-401600
910.0001400
1010.0002612014400
Sum100.00140020400

The squared-residual sum is 20400 (μV)2=2.040×10−8 V220400\ (\mu\mathrm{V})^2=2.040\times10^{-8}\ \mathrm{V}^2.

Step 3: Compute Sample Variance (s2s^2) and Sample Standard Deviation (ss)

ν=n−1=10−1=9\nu = n - 1 = 10 - 1 = 9 s2=∑(xi−xˉ)2n−1=2.040×10−8 V29=2.2667×10−9 V2s^2 = \frac{\sum (x_i - \bar{x})^2}{n - 1} = \frac{2.040 \times 10^{-8}\text{ V}^2}{9} = 2.2667 \times 10^{-9}\text{ V}^2 s=2.2667×10−9 V2=4.761×10−5 V=47.61 μVs = \sqrt{2.2667 \times 10^{-9}\text{ V}^2} = 4.761 \times 10^{-5}\text{ V} = 47.61\ \mu\text{V}

Step 4: Perform Grubbs' Outlier Test on x10x_{10}

Suspect value: x10=10.00026 Vx_{10} = 10.00026\text{ V}. Residual magnitude: ∣10.00026−10.000140∣=120 μV|10.00026 - 10.000140| = 120\ \mu\text{V}.

G=∣x10−xˉ∣s=120 μV47.61 μV=2.5205G = \frac{|x_{10} - \bar{x}|}{s} = \frac{120\ \mu\text{V}}{47.61\ \mu\text{V}} = 2.5205

From the Grubbs critical value table for n=10n = 10 at α=0.05\alpha = 0.05:

Gcrit(10,0.05)=2.2900G_{\text{crit}}(10, 0.05) = 2.2900

Because G=2.521>2.2900G = 2.521 > 2.2900, observation x10x_{10} is statistically flagged as an outlier at significance level α=0.05\alpha=0.05 under the two-sided test.

  • Hypothetical investigation: Assume synchronized instrument diagnostics and electrical monitoring establish that interference invalidated observation 10. A coincident HVAC start alone would not prove causation. Preserve original data and follow the approved documented exclusion and replacement or repeat-run method. The nine-reading calculation below illustrates arithmetic after justified exclusion, not permission to trim an inconvenient result.

Step 5: Recalculate Statistics for the Valid Sample (n=9n = 9)

Sum of the remaining 9 readings: ∑xi=90.00114 V\sum x_i = 90.00114\text{ V}.

xˉnew=90.00114 V9=10.0001267 V≈10.000127 V\bar{x}_{\text{new}} = \frac{90.00114\text{ V}}{9} = 10.0001267\text{ V} \approx 10.000127\text{ V} ∑i=19(xi−xˉnew)2=4.400×10−9 V2\sum_{i=1}^9 (x_i - \bar{x}_{\text{new}})^2 = 4.400 \times 10^{-9}\text{ V}^2 νnew=9−1=8\nu_{\text{new}} = 9 - 1 = 8 snew=4.400×10−98=2.3452×10−5 V=23.45 μVs_{\text{new}} = \sqrt{\frac{4.400 \times 10^{-9}}{8}} = 2.3452 \times 10^{-5}\text{ V} = 23.45\ \mu\text{V}

Step 6: Compute Standard Uncertainty of the Mean u(xˉ)u(\bar{x})

Because the calibration certificate will report the mean of the 9 readings as the calibrated voltage:

u(xˉ)=snewn=23.45 μV9=23.45 μV3=7.82 μVu(\bar{x}) = \frac{s_{\text{new}}}{\sqrt{n}} = \frac{23.45\ \mu\text{V}}{\sqrt{9}} = \frac{23.45\ \mu\text{V}}{3} = 7.82\ \mu\text{V}

Notice the metrological impact:

  • The sample standard deviation s=23.45 μVs = 23.45\ \mu\text{V} represents the process repeatability (dispersion of single voltmeter readings).
  • The standard uncertainty of the mean u(xˉ)=7.82 μVu(\bar{x}) = 7.82\ \mu\text{V} is the Type A standard uncertainty that enters the combined uncertainty budget for the reported mean value.

Official references (checked October 10, 2026): NIST Grubbs test definition and critical-value formula.

Test Your Knowledge

A Grubbs test flags one of ten mass-comparator readings under an assumed normal model. What is the appropriate initial response?

A

Automatically delete reading 7 from the dataset, recalculate the mean and variance for n = 9, and issue the calibration certificate

B

Replace reading 7 with the arithmetic average of the remaining 9 readings to preserve the degrees of freedom

C

Preserve the original observations, investigate the point and test assumptions, and document any justified exclusion or repeat-run decision.

D

Apply Dixon's Q test to override Grubbs' test and average the two resulting test statistics

Sections you finish are checked off in the contents.