8.4 Common Non-Normal Distributions: Binomial, Poisson & Chi-Square

Key Takeaways

  • Real-world process data frequently departs from normality due to physical zero-floors, mechanical wear-out patterns, or mixed operational streams, requiring specialized continuous or discrete distribution models.
  • The Binomial distribution models the number of defective units across n independent binary trials with constant probability p, serving as the mathematical foundation for p and np attribute control charts.
  • The Poisson distribution models discrete defect counts occurring across a continuous area of opportunity with a constant arrival rate λ, characterized by the unique mathematical property that its mean equals its variance (μ = σ² = λ).
  • The Chi-Square (χ²) distribution represents the sum of squared independent standard normal variables, is bounded at zero and skewed right, and governs hypothesis tests on population variance, goodness-of-fit evaluations, and contingency table tests of independence.
  • Normality is verified graphically using Normal Probability (Q-Q) plots and statistically via the Anderson-Darling test; when non-normality is confirmed (p <= 0.05), practitioners deploy Box-Cox power transformations, Johnson systems, or non-parametric alternatives.
Last updated: September 2026

8.4 Common Non-Normal Distributions: Binomial, Poisson & Chi-Square

Core Principle: In Six Sigma DMAIC projects, assuming data is normally distributed without empirical verification is a hazardous mistake. Real-world operational constraints—such as cycle times bounded at zero, mechanical wear-out rates, discrete defect occurrences, or contingency associations—routinely produce non-normal distributions (Binomial, Poisson, Chi-Square, Lognormal, Weibull, and Exponential). Quality practitioners must master the distinct mathematical properties and applications of these distributions, validate normality using Normal Probability Plots (Q-Q plots) and the Anderson-Darling test, and apply power transformations (Box-Cox, Johnson) or non-parametric alternatives when normality is rejected.


Why Real-World Operational Data Is Often Non-Normal

Introductory statistics courses often leave practitioners with the false impression that all industrial, healthcare, and transactional data naturally follows the Gaussian bell curve. In Six Sigma practice, this misunderstanding is known as the "Normality Myth." While the Central Limit Theorem guarantees that the distribution of sample means will approach normality as sample size increases, raw individual observations frequently violate normality due to inherent physical, operational, and structural constraints.

1. Physical and Boundary Constraints

  • Zero-Bound Variables (Natural Physical Floor): Many operational parameters cannot physically drop below zero ($X \ge 0$). Cycle times, transaction processing durations, machine setup times, customer waiting times, and hospital emergency room lengths-of-stay all have an absolute floor at zero but can extend upward indefinitely. This physical asymmetry creates heavy positive (right) skewness.
  • Geometric Form Tolerances: Dimensions such as flatness, concentricity, circularity, parallelism, and surface roughness ($Ra$) represent absolute deviations from perfection. A part cannot possess negative runout or negative flatness. Data clusters tightly near zero with an extended tail toward higher roughness or distortion.
  • Purity and Contaminant Levels: Chemical trace impurities, particulate counts in cleanrooms, and microbial counts in pharmaceutical water systems are bounded at zero and exhibit extreme positive skewness.

2. Operational and Process Constraints

  • Mixture of Sub-Processes: When data is collected from multiple machine spindles, distinct tool cavities, different operator shifts, or varying supplier batches without stratification, the resulting distribution can be bimodal, multimodal, or broad-shouldered (platykurtic).
  • Wear and Degradation Over Time: Tool wear, bearing fatigue, cutting insert degradation, and chemical catalyst decay do not occur uniformly. Failure rates often accelerate over time, generating skewed reliability distributions.
                    Operational Probability Distributions
                                      │
             ┌────────────────────────┴────────────────────────┐
             ▼                                                 ▼
     Discrete Distributions                           Continuous Distributions
   • Countable integer events                       • Unbroken continuous scale
   • Binomial: Defective units (pass/fail)          • Normal: Symmetrical bell curve
   • Poisson: Defect counts per opportunity         • Chi-Square: Sum of squared Z-scores
   • Hypergeometric: Finite sampling                • Lognormal: Skewed cycle times
   • Control Charts: p, np, c, u                    • Weibull: Life data & reliability
                                                    • Exponential: Constant failure rate

The Binomial Distribution (Discrete Defective Units)

The Binomial Distribution evaluates the number of "successes" (or defective units) observed across a fixed series of independent trials with binary outcomes. In quality engineering, it models the number of non-conforming items identified during inspection.

The Four Mandatory Conditions (BINS Acronym)

For a process to follow a binomial distribution, it must satisfy four rigid mathematical criteria:

  1. B — Binary Outcomes: Each trial has exactly two mutually exclusive outcomes: Conforming vs. Non-Conforming, Pass vs. Fail, Go vs. No-Go.
  2. I — Independent Trials: The outcome of any single trial does not affect or alter the probability of any subsequent trial.
  3. N — Number of Trials Fixed: The sample size or number of inspections ($n$) is fixed and predetermined prior to data collection.
  4. S — Same Probability: The probability of defect ($p$) remains constant from trial to trial across the entire sampling sequence.

Mathematical Formulation & Parameters

The probability of observing exactly $k$ non-conforming units in a sample of $n$ independent items is governed by the probability mass function (PMF):

P(X=k)=(nk)pk(1p)nk=n!k!(nk)!pk(1p)nkP(X = k) = \binom{n}{k} p^k (1 - p)^{n - k} = \frac{n!}{k!(n - k)!} p^k (1 - p)^{n - k}

Where:

  • $n$ is the total number of items inspected (sample size).
  • $k$ is the exact number of defective units observed ($k = 0, 1, 2, \dots, n$).
  • $p$ is the historical probability of a defective unit occurring on any single trial.
  • $(1 - p)$ is the probability of a conforming unit (often denoted as $q$).
  • $\binom{n}{k}$ is the binomial combination coefficient.

Key Moments of the Binomial Distribution

  • Mean (Expected Defective Count): μ=np\mu = n \cdot p
  • Variance: σ2=np(1p)\sigma^2 = n \cdot p \cdot (1 - p)
  • Standard Deviation: σ=np(1p)\sigma = \sqrt{n \cdot p \cdot (1 - p)}

Six Sigma Control Chart Connections

The Binomial distribution forms the theoretical foundation for two major Shewhart attribute control charts:

  • $p$-Chart (Fraction / Proportion Defective): Used when subgroup sample sizes ($n$) vary from period to period.
  • $np$-Chart (Number of Defectives): Used when subgroup sample sizes ($n$) remain strictly constant across all inspection intervals.

Normal Approximation to the Binomial

When sample size $n$ is large and $p$ is not near 0 or 1, calculating binomial factorials becomes computationally cumbersome. The Central Limit Theorem allows practitioners to approximate the binomial distribution using a normal distribution $N(\mu = np, \sigma^2 = np(1-p))$, provided both of the following threshold conditions are met:

np5andn(1p)5(conservative standard: 10)n \cdot p \ge 5 \quad \text{and} \quad n \cdot (1 - p) \ge 5 \quad \text{(conservative standard: } \ge 10\text{)}


The Poisson Distribution (Discrete Defect Counts)

While the Binomial distribution counts the number of defective units (a part is either good or bad), the Poisson Distribution models the number of discrete occurrences of a rare event (such as defects, blemishes, or imperfections) across a continuous, fixed area of opportunity.

An area of opportunity can be a defined unit of time (defects per hour), physical length (insulation pinholes per 1,000 meters of wire), surface area (paint bubbles per square meter of automotive hood), or volume (particulates per liter of intravenous fluid).

               Binomial vs. Poisson: Conceptual Difference

       Binomial Distribution                          Poisson Distribution
     "Counts Defective Units"                      "Counts Total Defects"
  ┌─────────────────────────────┐               ┌─────────────────────────────┐
  │ Part 1: PASS                │               │ Part 1: [ •  • ] (2 defects)│
  │ Part 2: FAIL (Scrapped)     │               │ Part 2: [      ] (0 defects)│
  │ Part 3: PASS                │               │ Part 3: [ •    ] (1 defect) │
  │ Part 4: FAIL (Scrapped)     │               │ Part 4: [ •••• ] (4 defects)│
  └─────────────────────────────┘               └─────────────────────────────┘
   Sample n = 4 units                            Total Defects = 7 defects
   Defective count k = 2 units                   Area of Opportunity = 4 units
   Metric: Proportion Defective (p)              Metric: Defects Per Unit (DPU = 1.75)

Mandatory Conditions for a Poisson Process

  1. Random and Independent: The occurrence of a defect at one location or instant does not alter the likelihood of another defect occurring elsewhere.
  2. Constant Arrival Rate: The average rate of occurrence ($\lambda$, "lambda") is constant throughout the entire area of opportunity.
  3. Infinitesimal Simplicity: Two events cannot occur at the exact same instantaneous point in time or space.
  4. No Theoretical Upper Bound: Unlike the Binomial distribution (where defects cannot exceed sample size $n$), a single unit can theoretically accumulate an infinite number of defects.

Mathematical Formulation & Parameters

The probability of observing exactly $x$ defects in a standard unit of opportunity is given by the probability mass function:

P(X=x)=eλλxx!for x=0,1,2,3,P(X = x) = \frac{e^{-\lambda} \lambda^x}{x!} \quad \text{for } x = 0, 1, 2, 3, \dots

Where:

  • $\lambda$ is the average number of defect occurrences per unit of opportunity.
  • $e \approx 2.71828$ is the base of the natural logarithm.
  • $x!$ is the factorial of the observed count $x$.

The Defining Property of the Poisson Distribution

A universally tested property of the Poisson distribution is that its mean and variance are mathematically identical:

μ=λσ2=λσ=λ\mu = \lambda \qquad \sigma^2 = \lambda \qquad \sigma = \sqrt{\lambda}

If an industrial printing process averages $\lambda = 4.0$ ink smudges per banner, the variance is identically $4.0\text{ smudges}^2$, and the standard deviation is $\sigma = \sqrt{4.0} = 2.0\text{ smudges}$.

Six Sigma Control Chart Connections

The Poisson distribution underpins two critical attribute control charts:

  • $c$-Chart (Count of Defects): Deployed when the area of opportunity (sample inspection unit) remains constant (e.g., exactly one aircraft wing panel, or exactly 100 square feet of carpet).
  • $u$-Chart (Defects Per Unit): Deployed when the inspection unit size varies (e.g., inspecting fabric bolts of differing lengths, or auditing daily customer calls with varying call volumes).

Poisson Approximation to the Binomial

When the number of binomial trials $n$ is very large ($n \ge 100$) and the probability of defect $p$ is very small ($p \le 0.05$), the binomial distribution converges directly to a Poisson distribution with parameter $\lambda = n \cdot p$. This allows rapid approximation of rare manufacturing defects without computing massive factorial combinations.


The Chi-Square ($\chi^2$) Distribution

The Chi-Square ($\chi^2$) Distribution is a fundamental continuous probability distribution widely utilized in Six Sigma for evaluating process variance, conducting goodness-of-fit assessments, and analyzing contingency tables.

Theoretical Derivation: Sum of Squared Normal Variables

The Chi-Square distribution is directly derived from the standard normal distribution. If $Z_1, Z_2, \dots, Z_k$ represent $k$ independent random variables drawn from a standard normal distribution ($Z_i \sim N(0, 1)$), the sum of their squared values follows a Chi-Square distribution with $k$ degrees of freedom:

Q=i=1kZi2=Z12+Z22++Zk2χ2(k)Q = \sum_{i=1}^k Z_i^2 = Z_1^2 + Z_2^2 + \dots + Z_k^2 \sim \chi^2(k)

                      The Chi-Square Distribution Architecture

     Probability Density f(χ²)
            ▲
            │   df = 2 (Monotonically decreasing)
            │───.
            │    \    df = 5 (Positively skewed)
            │     \   .--.
            │      \ /    \      df = 20 (Approaching Normal symmetry)
            │       │      \      _.-'"'-._
            │       /       \   .'         '.
            │      /         '.'             '--.
            └───┴─────────────────────────────────────► χ²
                0    2    4    6    8   10   12   14

Key Mathematical Properties of Chi-Square

  1. Strictly Non-Negative ($0 \le \chi^2 < \infty$): Because it is formed by summing squared numbers ($Z^2$), Chi-Square values can never be negative.
  2. Asymmetrical and Positively Skewed: The distribution starts at the origin and exhibits an extended right tail. The degree of skewness depends entirely on its degrees of freedom ($df = k$): Skewness=8k\text{Skewness} = \sqrt{\frac{8}{k}}
  3. Moments Defined Strictly by Degrees of Freedom ($k$):
    • Mean: $\mu = k$
    • Variance: $\sigma^2 = 2k$
    • Standard Deviation: $\sigma = \sqrt{2k}$
    • Peak (Mode): Occurs at $\chi^2 = k - 2$ (for $k \ge 2$).
  4. Convergence to Normality via Central Limit Theorem: As the degrees of freedom increase ($k > 30$), the skewness rapidly approaches zero, and the Chi-Square distribution converges toward a symmetrical Gaussian distribution $N(k, 2k)$.

The Three Major Six Sigma Applications of Chi-Square

1. Hypothesis Testing on a Single Population Variance ($\sigma^2$)

When a Green Belt needs to test whether a process improvement successfully reduced process variability below a target threshold $\sigma_0^2$, the test statistic follows a Chi-Square distribution with $df = n - 1$:

χ2=(n1)s2σ02\chi^2 = \frac{(n - 1)s^2}{\sigma_0^2}

Where $s^2$ is the sample variance calculated from a random sample of size $n$, and $\sigma_0^2$ is the hypothesized baseline variance.

2. Chi-Square Test of Independence (Contingency Tables)

In the Analyze phase, teams frequently investigate whether two categorical variables are statistically independent or associated. For example, does defect type depend on the manufacturing shift, operator, or production line?

Data is tabulated into a two-way contingency table with $r$ rows and $c$ columns. The expected frequency for each cell under the null hypothesis of independence is:

Eij=(Row i Total)×(Column j Total)Grand Total NE_{ij} = \frac{(\text{Row } i \text{ Total}) \times (\text{Column } j \text{ Total})}{\text{Grand Total } N}

The test statistic measures the cumulative standardized squared discrepancy between observed frequencies ($O_{ij}$) and expected frequencies ($E_{ij}$):

χ2=i=1rj=1c(OijEij)2Eij\chi^2 = \sum_{i=1}^r \sum_{j=1}^c \frac{(O_{ij} - E_{ij})^2}{E_{ij}}

With degrees of freedom: df=(r1)×(c1)df = (r - 1) \times (c - 1)

3. Chi-Square Goodness-of-Fit Test

Used to test whether an observed frequency distribution conforms to a theoretical probability distribution (such as a normal, Poisson, or binomial model):

χ2=k=1K(OkEk)2Ekdf=K1p\chi^2 = \sum_{k=1}^K \frac{(O_k - E_k)^2}{E_k} \qquad df = K - 1 - p

Where $K$ is the number of categories or bins, and $p$ is the number of distribution parameters estimated from the sample data.


Continuous Non-Normal Distributions in Reliability & Operations

In addition to the discrete models and the Chi-Square distribution, three continuous non-normal distribution families are ubiquitous in Lean Six Sigma engineering:

1. The Lognormal Distribution

A continuous random variable $X$ follows a Lognormal Distribution if the natural logarithm of $X$ is normally distributed:

Y=ln(X)N(μln,σln2)for X>0Y = \ln(X) \sim N(\mu_{\ln}, \sigma_{\ln}^2) \quad \text{for } X > 0

  • Characteristics: Strictly positive ($X > 0$), starts at zero, rises rapidly to an early peak, and features a long, heavy right tail.
  • Physical Genesis: Arises from multiplicative processes (where random shocks multiply together rather than add together).
  • Six Sigma Applications: Mean Time to Repair (MTTR), software bug resolution cycle times, financial claim sizes, and mechanical fatigue life under cyclical loading.

2. The Weibull Distribution

Developed by Swedish engineer Waloddi Weibull, the Weibull Distribution is the most versatile distribution in reliability engineering, product life data analysis, and survival modeling.

A Weibull distribution is defined by three parameters:

  • Shape Parameter ($\beta$ or "beta"): Governs the fundamental failure pattern.
  • Scale Parameter ($\eta$ or "eta" / characteristic life): Represents the 63.2nd percentile of life data.
  • Threshold / Location Parameter ($\gamma$ or "gamma"): The earliest time at which a failure can occur (often assumed to be zero).

The Bathtub Curve and the Shape Parameter ($\beta$)

The value of $\beta$ directly reflects the operational phase of the product lifecycle:

Shape ParameterFailure Rate TrendLifecycle PhaseSix Sigma Interpretation
$\beta < 1.0$DecreasingInfant MortalityEarly life failures caused by manufacturing defects, substandard soldering, or assembly flaws.
$\beta = 1.0$ConstantUseful LifePurely random failures independent of component age; identical to the Exponential distribution.
$\beta > 1.0$IncreasingWear-Out PhaseAge-dependent degradation, bearing fatigue, corrosion, and physical breakdown.
$\beta \approx 3.0 - 3.6$SymmetricalNormal ApproximationBell-shaped profile closely mimicking the standard normal distribution.

3. The Exponential Distribution

The Exponential Distribution models the continuous time between independent, random events occurring at a constant average rate $\lambda$:

f(x)=λeλxfor x0f(x) = \lambda e^{-\lambda x} \quad \text{for } x \ge 0

  • Characteristics: Strictly decreasing curve; peak density occurs at $x = 0$. Governed by a single parameter, the event rate $\lambda$ (where mean time between events $\theta = 1/\lambda$).
  • The Memoryless Property: The probability that an operating component survives an additional time $t$ is completely independent of how long it has already functioned: P(X>s+tX>s)=P(X>t)P(X > s + t \mid X > s) = P(X > t)
  • Six Sigma Applications: Mean Time Between Failures (MTBF) during stable "useful life," arrival intervals at call centers, and electrical component surge intervals.

Master Comparison Matrix: Key Six Sigma Distributions

DistributionData TypeRange of ValuesDefining ParametersMean ($\mu$)Variance ($\sigma^2$)Primary Six Sigma Application
NormalContinuous$-\infty < x < \infty$$\mu, \sigma$$\mu$$\sigma^2$Physical dimensions, weights, baseline capability ($C_p, C_{pk}$)
BinomialDiscrete$x \in {0, 1, \dots, n}$$n, p$$n \cdot p$$n \cdot p(1 - p)$Number of defective units (pass/fail), $p$ and $np$ control charts
PoissonDiscrete$x \in {0, 1, 2, \dots}$$\lambda$$\lambda$$\lambda$Number of defect occurrences per area of opportunity, $c$ and $u$ charts
Chi-SquareContinuous$0 \le x < \infty$$df = k$$k$$2k$Single variance hypothesis tests, contingency independence, goodness-of-fit
LognormalContinuous$x > 0$$\mu_{\ln}, \sigma_{\ln}$$e^{\mu + \sigma^2/2}$$e^{2\mu+\sigma^2}(e^{\sigma^2}-1)$Cycle times, repair times (MTTR), financial transaction sizes
WeibullContinuous$x \ge 0$$\beta$ (shape), $\eta$ (scale)Complex gammaComplex gammaReliability, life testing, accelerated life studies, wear-out modeling
ExponentialContinuous$x \ge 0$$\lambda$ (rate)$1 / \lambda$$1 / \lambda^2$Time between random events (MTBF), constant failure rate modeling

Testing for Normality: Graphical & Statistical Methodologies

Before executing parametric capability studies ($C_p, C_{pk}$) or standard hypothesis tests ($t$-tests, ANOVA), Green Belts must verify whether the normality assumption holds.

1. Graphical Evaluation: The Normal Probability Plot (Q-Q Plot)

A Normal Probability Plot (or Quantile-Quantile plot) plots ordered sample observations along the horizontal axis against theoretical standard normal quantiles along the vertical axis.

     Normal Data (Straight Line)              Heavy-Tailed / Skewed Data (Curved)

  Quantiles                                Quantiles
     ▲          /                             ▲          .---
     │         /                              │        .'
     │        /   Data falls on               │       /      S-curve indicates
     │       /    straight line               │      /       heavy tails
     │      /     ──► NORMAL                  │    .'        (Leptokurtic)
     │     /                                  │   /          ──► NON-NORMAL
     └───/──────────────► Sample              └───'──────────────► Sample
  • Interpretation Rule: If the process data is normally distributed, the data points form a tight, straight diagonal line along the 45-degree reference line, staying comfortably within the 95% confidence bands.
  • The Fat-Pencil Test: If a standard pencil laid over the line covers all data points, the normality assumption is generally acceptable for practical engineering.
  • Curved Pattern (Banana Shape): Indicates skewness (asymmetry).
  • S-Shaped Pattern: Indicates kurtosis departures (heavy or light tails).

2. Statistical Hypothesis Testing: The Anderson-Darling (AD) Test

While probability plots provide intuitive graphical assessment, the Anderson-Darling (AD) Test is the definitive statistical test used in Six Sigma software (e.g., Minitab). Unlike the Kolmogorov-Smirnov test, the Anderson-Darling test applies heavy mathematical weighting to the tails of the distribution, where non-normality most severely impacts defect rate and capability projections.

The Anderson-Darling Hypothesis Framework

H0 (Null Hypothesis): The sample data follows a normal distribution.\mathbf{H_0 \text{ (Null Hypothesis):}} \text{ The sample data follows a normal distribution.} Ha (Alternative Hypothesis): The sample data does NOT follow a normal distribution.\mathbf{H_a \text{ (Alternative Hypothesis):}} \text{ The sample data does NOT follow a normal distribution.}

The P-Value Decision Rule

Statistical software computes the Anderson-Darling test statistic ($A^2$) and an associated $p$-value. In Six Sigma, the standard significance threshold is $\alpha = 0.05$ (5%).

Test ResultStatistical DecisionNormality ConclusionSubsequent Operational Action
$p\text{-value} > 0.05$Fail to Reject $H_0$Data is normally distributed (assumption is tenable).Proceed with standard parametric capability indices ($C_p, C_{pk}$) and $Z$-score calculations.
$p\text{-value} \le 0.05$Reject $H_0$Data is statistically NON-NORMAL.Cannot use standard $C_p/C_{pk}$. Must transform data, fit non-normal curves, or use non-parametric tests.

[!TIP] Exam Memory Key: "If the $p$-value is low ($p \le 0.05$), the null must go!" Since the null hypothesis is that the data is normal, rejecting the null proves that the process data is non-normal.


Engineering Strategies for Handling Non-Normal Data

When the Anderson-Darling test confirms non-normality ($p \le 0.05$), applying standard normal formulas produces disastrously inaccurate capability estimates. Practitioners deploy four proven remediation strategies:

                     Remediation Strategies for Non-Normal Data
                                        │
         ┌──────────────────────────────┼──────────────────────────────┐
         ▼                              ▼                              ▼
    Investigate Special           Mathematical Data               Non-Parametric
    Causes & Stratify              Transformations                   Methods
  Eliminate mixed streams        Box-Cox Power (Y > 0)          Mood's Median, Mann-Whitney
  Isolate cavities / shifts      Johnson System (Any Y)         Distribution-Free Capability

1. Investigate Special Causes and Stratify

Before applying mathematical algorithms, investigate whether non-normality is caused by improper sampling. Combining parts from two different lathe spindles, mold cavities, or work shifts creates an artificial mixture distribution. Stratifying the data by machine or shift often reveals that each individual stream is perfectly normal.

2. Mathematical Data Transformation

Transforming data involves applying a mathematical function to every individual observation, compressing the elongated tail and restoring symmetrical bell-shaped normality.

A. Box-Cox Power Transformation

Developed by George Box and David Cox, the Box-Cox Transformation searches across an exponential family of power transformations governed by the parameter $\lambda$ (lambda):

Y={Yλ1λfor λ0ln(Y)for λ=0Y' = \begin{cases} \frac{Y^\lambda - 1}{\lambda} & \text{for } \lambda \ne 0 \\[6pt] \ln(Y) & \text{for } \lambda = 0 \end{cases}

  • Strict Mathematical Constraint: The Box-Cox transformation strictly requires all data values to be strictly positive ($Y > 0$). It cannot process zero or negative numbers. If zeros exist, a constant $c$ must be added to all values ($Y + c$).
  • Common $\lambda$ Values:
    • $\lambda = 1.0$: No transformation ($Y' = Y$).
    • $\lambda = 0.5$: Square root transformation ($Y' = \sqrt{Y}$), ideal for count-type data.
    • $\lambda = 0.0$: Natural log transformation ($Y' = \ln(Y)$), standard for lognormal cycle times.
    • $\lambda = -1.0$: Reciprocal transformation ($Y' = 1/Y$), common for rate/speed metrics.

B. Johnson Transformation

When data contains zeros, negative values, or severe multimodal skewness that Box-Cox cannot resolve, the Johnson Transformation selects among three functional curve families:

  • $S_B$ (Bounded): Accommodates data bounded by physical upper and lower walls.
  • $S_L$ (Lognormal): Accommodates unbounded positive data.
  • $S_U$ (Unbounded): Accommodates data spanning negative, zero, and positive values without requiring artificial offsets.

3. Non-Parametric (Distribution-Free) Statistical Methods

If transformations fail to normalize the dataset, practitioners utilize non-parametric tests that do not assume normality:

  • Replace 2-sample $t$-tests with the Mann-Whitney U-test or Mood's Median test.
  • Replace one-way ANOVA with the Kruskal-Wallis test.
  • Calculate non-parametric process capability using empirical percentiles (e.g., ISO 22514-2).

4. Direct Non-Normal Distribution Fitting

Rather than forcing non-normal data into a normal shape, modern statistical software can directly fit a specialized distribution (such as a 3-parameter Weibull model) and compute defect rates directly against the fitted distribution's cumulative probability tails.


Step-by-Step Worked Industrial Calculations

Worked Calculation 1: Chi-Square Test of Independence (Contingency Table)

Scenario: A Six Sigma Green Belt at an electronic assembly plant investigates whether the occurrence of two major defect categories (Surface Scratches vs. Solder Bridging) is independent of the manufacturing shift (Day Shift vs. Night Shift). A quality audit evaluates $N = 200$ defective assemblies.

Step 1: Formulate the Hypotheses

  • $H_0$: Defect category is independent of operating shift (no association).
  • $H_a$: Defect category is dependent on operating shift (statistically significant association).

Step 2: Construct the Observed Contingency Table ($O_{ij}$)

Operating ShiftSurface ScratchesSolder BridgingRow Totals
Day Shift301040
Night Shift7090160
Column Totals100100Grand Total $N = 200$

Step 3: Compute Expected Frequencies Under Independence ($E_{ij}$)

Using the formula $E_{ij} = \frac{(\text{Row Total}) \times (\text{Column Total})}{N}$:

  • Day Shift / Scratches: $E_{11} = \frac{40 \times 100}{200} = 20.0$
  • Day Shift / Bridging: $E_{12} = \frac{40 \times 100}{200} = 20.0$
  • Night Shift / Scratches: $E_{21} = \frac{160 \times 100}{200} = 80.0$
  • Night Shift / Bridging: $E_{22} = \frac{160 \times 100}{200} = 80.0$

Step 4: Compute the Chi-Square Test Statistic ($\chi^2$)

Calculate $\frac{(O - E)^2}{E}$ for each of the four cells:

(3020)220=10220=10020=5.00\frac{(30 - 20)^2}{20} = \frac{10^2}{20} = \frac{100}{20} = 5.00 (1020)220=(10)220=10020=5.00\frac{(10 - 20)^2}{20} = \frac{(-10)^2}{20} = \frac{100}{20} = 5.00 (7080)280=(10)280=10080=1.25\frac{(70 - 80)^2}{80} = \frac{(-10)^2}{80} = \frac{100}{80} = 1.25 (9080)280=10280=10080=1.25\frac{(90 - 80)^2}{80} = \frac{10^2}{80} = \frac{100}{80} = 1.25

Summing all four components: χ2=5.00+5.00+1.25+1.25=12.50\chi^2 = 5.00 + 5.00 + 1.25 + 1.25 = 12.50

Step 5: Determine Degrees of Freedom & Statistical Decision

df=(r1)×(c1)=(21)×(21)=1×1=1df = (r - 1) \times (c - 1) = (2 - 1) \times (2 - 1) = 1 \times 1 = 1

At a significance level $\alpha = 0.05$, the critical Chi-Square value for $df = 1$ is $\chi^2_{0.05, 1} = 3.841$.

Because the calculated test statistic ($\chi^2 = 12.50$) substantially exceeds the critical value ($3.841$), the associated $p$-value is $p < 0.001$. The team firmly rejects the null hypothesis and concludes that defect types are statistically dependent on shift. Solder bridging is disproportionately concentrated on the night shift, directing root-cause investigation to night-shift soldering oven temperatures or operator training.


Worked Calculation 2: Poisson Defect Probability

Scenario: A wave soldering machine produces printed circuit boards (PCBs). Process baseline data demonstrates that solder bridging defects follow a Poisson distribution with an average occurrence rate of $\lambda = 2.0$ defects per board.

What is the exact probability that a randomly selected PCB has zero defects ($x = 0$)?

P(X=0)=eλλxx!=e2.0(2.0)00!=e2.011=e2.00.1353P(X = 0) = \frac{e^{-\lambda} \lambda^x}{x!} = \frac{e^{-2.0} \cdot (2.0)^0}{0!} = \frac{e^{-2.0} \cdot 1}{1} = e^{-2.0} \approx 0.1353

There is a 13.53% probability that a board passes through wave soldering with zero solder defects, meaning 86.47% of boards will require touch-up rework. This provides the empirical baseline ($Y_0$) for the project charter.


Critical CSSC Exam Traps

  • Trap 1: Confounding Binomial (Defective Units) with Poisson (Defect Counts) — When an exam question involves inspecting continuous surfaces (e.g., paint bubbles on a hood, fabric tears per roll, surface flaws per glass panel), candidates mistakenly pick the Binomial distribution. Remember: if multiple defects can occur on a single unit, it is Poisson. If each unit is simply tagged pass or fail, it is Binomial.
  • Trap 2: Inverting the Anderson-Darling P-Value Decision — Assuming $p \le 0.05$ proves normality. The null hypothesis $H_0$ is that data is normal. Therefore, a $p$-value $\le 0.05$ rejects normality, proving the data is non-normal.
  • Trap 3: Attempting Box-Cox on Zero or Negative Values — Box-Cox formulas evaluate logarithms, which are mathematically undefined for zero and negative numbers. When data includes zeros or negative values, select the Johnson Transformation (or add an offset constant).
  • Trap 4: Miscalculating Chi-Square Degrees of Freedom for Contingency Tables — Dividing or computing degrees of freedom as $n - 1$. For a two-way contingency table, degrees of freedom is always $(r - 1)(c - 1)$.
  • Trap 5: Misinterpreting Weibull $\beta = 1.0$ as Normal — When a Weibull distribution has $\beta = 1.0$, it models a constant failure rate and collapses identically into an Exponential distribution. A Weibull distribution only approximates a normal distribution when $\beta \approx 3.0$ to $3.6$.
Loading diagram...
Six Sigma Probability Distribution Selection Architecture
Test Your Knowledge

A continuous improvement team conducts an Anderson-Darling normality test on 50 customer checkout cycle times during a Lean Six Sigma pilot. The statistical software reports an Anderson-Darling statistic A^2 = 1.482 and a corresponding p-value = 0.008. Using a standard significance level of alpha = 0.05, which conclusion and subsequent action must the Green Belt take?

A
B
C
D
Test Your Knowledge

A reliability engineer analyzes the time-to-failure data for a critical industrial hydraulic pump using a 2-parameter Weibull distribution. Software analysis indicates that the shape parameter beta is equal to 2.8. What does this shape parameter reveal regarding the failure pattern of the pump?

A
B
C
D
Test Your Knowledge

An inspection team monitors quality on an automotive sheet-metal stamping line. Technicians inspect stamped vehicle hood panels and count the total number of surface paint blemishes, dents, and scratch imperfections found across each 3.5-square-meter panel. Which probability distribution is most appropriate for modeling these imperfection counts?

A
B
C
D
Test Your Knowledge

A Lean Six Sigma project team investigates whether the occurrence of four distinct defect categories (Dimensional, Cosmetic, Functional, Packaging) is independent of three production shifts (Day, Swing, Night). The team constructs a contingency table with 3 rows and 4 columns. When performing a Chi-Square test of independence, how many degrees of freedom (df) should be used to determine the critical value?

A
B
C
D