9.3 Basic Distribution Types: Normal, Binomial & Poisson
Key Takeaways
The continuous Normal (Gaussian) distribution is symmetrical and bell-shaped, where the mean, median, and mode are identical.
The Empirical Rule establishes that approximately 68.27% of data lies within ±1σ, 95.45% within ±2σ, and 99.73% within ±3σ of the mean.
The standard normal score Z = (X - μ) / σ standardizes continuous observations to quantify defect probabilities against specification limits.
The Binomial distribution models discrete counts of defective units in independent pass/fail trials with constant probability p.
The Poisson distribution models discrete defect counts across a continuous area of opportunity, while skewed and bimodal shapes indicate process mixing or natural boundaries.
Basic Distribution Types: Normal, Binomial & Poisson
Quick Answer: Statistical analysis in Six Sigma relies on probability distributions to model process behavior. The continuous Normal (Gaussian) distribution is symmetrical (mean = median = mode) and governed by the Empirical Rule (68.27% within , 95.45% within , 99.73% within ), utilizing -scores to calculate process yields. Discrete distributions model count data: the Binomial distribution tracks binary pass/fail outcomes across independent trials, whereas the Poisson distribution tracks defect frequencies across an area of opportunity. Deviations from normality, such as skewness or bimodal clustering, provide immediate visual clues of underlying operational mixing or boundary effects. Independent CSSYB study guide by OpenExamPrep.
Continuous vs. Discrete Data Distributions in Six Sigma
In the Analyze phase of DMAIC, Six Sigma practitioners evaluate process performance by fitting collected baseline data to theoretical probability distributions. A probability distribution provides a mathematical model describing the likelihood of observing various values of a process output metric ().
The fundamental starting point for selecting an appropriate distribution model is distinguishing between continuous and discrete data:
- Continuous Data (Variables Data): Measurements taken on a continuous, uninterrupted numerical scale (such as dimensions in millimeters, cycle times in seconds, temperatures in Celsius, or tensile strength in megapascals). Continuous data is modeled using continuous distributions, predominantly the Normal (Gaussian) distribution.
- Discrete Data (Attribute Data): Categorical counts, classifications, or integer events (such as conforming versus nonconforming parts, customer billing complaints per month, or solder voids on a circuit board). Discrete data is modeled using discrete probability distributions, primarily the Binomial and Poisson distributions. The CSSYB BoK names the normal and binomial distributions; the Poisson distribution is included here because it is the usual model for counts of defects, and exam answer choices often contrast it with the binomial.
The Continuous Normal (Gaussian) Distribution
The Normal distribution is the mathematical cornerstone of statistical quality control and process capability modeling. Described by Carl Friedrich Gauss, it depicts the behavior of processes influenced by many independent, random common causes acting together.
Key Mathematical & Geometric Properties
- Symmetrical Bell Shape: The curve is perfectly symmetrical around the central arithmetic mean ().
- Coincidence of Central Tendencies: The mean, median, and mode are identical and located exactly at the center of the distribution.
- Asymptotic Tails: The tails of the curve extend infinitely in both positive and negative directions without ever touching the horizontal axis ().
- Total Probability: The total area under the probability density curve equals exactly (or probability).
The Empirical Rule (68-95-99.7 Rule)
For any data set that follows a normal distribution, the dispersion of observations relative to the population mean () and standard deviation () conforms to fixed, universal percentages:
- : Approximately 68.27% of all data points fall within one standard deviation of the mean (roughly 2 out of every 3 observations).
- : Approximately 95.45% of all data points fall within two standard deviations of the mean (roughly 19 out of every 20 observations).
- : Approximately 99.73% of all data points fall within three standard deviations of the mean (roughly 369 out of every 370 observations).
Only 0.27% of observations fall outside the control boundaries (0.135% in each tail), which forms the theoretical foundation for Shewhart three-sigma control limits.
The Standard Normal Distribution and the -Score
Because real-world processes operate with different units and scales, practitioners convert raw data into the Standard Normal Distribution (-distribution), which has a standardized mean of and standard deviation of .
The -score (standard score) measures the exact number of standard deviations an individual observation () lies above or below the population mean:
Where:
- = Specific value or customer specification limit (USL or LSL)
- = Process mean
- = Process standard deviation
Yellow Belt Calculation Example: A precision shaft grinding process has a mean diameter with . The customer's Upper Specification Limit (USL) is . The -score for the upper limit is:
Consulting a standard normal table reveals that leaves approximately of the distribution in the upper tail, indicating that about of shafts will exceed the customer's maximum allowable diameter.
Discrete Probability Distributions
When evaluating counts, classifications, and attribute pass/fail occurrences, Yellow Belts apply two core discrete distributions.
1. The Binomial Distribution
The Binomial distribution models discrete counts of successes or failures across a sequence of independent trials.
- Governing Conditions:
- The process consists of a fixed number of identical trials ().
- Each trial has only two mutually exclusive outcomes: success or failure (conforming or nonconforming).
- The probability of success () remains constant from trial to trial.
- All trials are statistically independent.
- Six Sigma Application: The Binomial distribution models the number of defective (nonconforming) units in an inspection lot. It provides the mathematical foundation for attribute -charts (proportion nonconforming) and -charts (number nonconforming).
2. The Poisson Distribution
The Poisson distribution models the frequency of discrete events or defects occurring within a specified, continuous area of opportunity (such as an interval of time, physical area, volume, or length).
- Governing Conditions:
- Events occur independently at a constant average rate ().
- The occurrence of an event in one interval does not influence occurrences in another.
- Multiple occurrences in an infinitesimally small interval are practically zero.
- Key Statistical Property: The mathematical variance of a Poisson distribution equals its mean ().
- Six Sigma Application: The Poisson distribution models the count of defects per unit—such as paint blemishes on an automobile hood, typing errors per standard contract page, or customer complaints logged per hour. It serves as the statistical engine for -charts (constant area of opportunity) and -charts (variable area of opportunity).
Critical Difference: Defective Units vs. Defects
Yellow Belts must never confuse defective units with defects:
- Defective Unit (Binomial): An entire product or transaction that fails to meet criteria (a cracked windshield). A unit is either good or bad.
- Defect (Poisson): An individual flaw or blemish on a unit (three small chips on a single windshield). A single defective unit may contain multiple distinct defects.
Diagnostic Interpretation of Distribution Shapes
When data is plotted on a histogram, the resulting geometric shape provides immediate diagnostic clues regarding underlying process dynamics.
| Distribution Shape | Geometric Characteristics | Mathematical Relationship | Diagnostic Root Cause Clues |
|---|---|---|---|
| Normal (Gaussian) | Symmetrical, single central bell-shaped peak | Stable process operating under natural, unconstrained common cause conditions. | |
| Positively Skewed (Right-Skewed) | Long tail extends toward higher positive values on the right | Natural lower physical boundary at zero (e.g., cycle times, customer queue times, rework hours, delivery transit delays). | |
| Negatively Skewed (Left-Skewed) | Long tail extends toward lower negative values on the left | Upper specification or physical ceiling boundary (e.g., raw material purity percentages, test scores on an easy exam, process yield rates). | |
| Bimodal Distribution | Two distinct peaks with a dip or valley between them | Two separate local modes | Mixture of two distinct populations (e.g., data combined from two different machines, two shifts, two vendor raw material lots, or two uncalibrated gauges). Requires stratification. |
When a Yellow Belt encounters a bimodal distribution during the Analyze phase, the mandatory investigative action is stratification—disaggregating the dataset by shift, machine, operator, or vendor to isolate and analyze the two underlying distributions separately.
A metal fabrication process produces steel pins with a normally distributed diameter having a mean (μ) of 50.0 mm and a standard deviation (σ) of 0.2 mm. According to the Empirical Rule, what percentage of pins will have a diameter falling between 49.6 mm and 50.4 mm?
50.00%
68.27%
95.45%
99.73%
A quality inspector at a textile mill inspects large 100-meter rolls of woven fabric and records the number of surface weave flaws, snags, and stains found on each roll. Which probability distribution is most appropriate for modeling these flaw counts?
Normal distribution
Poisson distribution
Binomial distribution
Uniform distribution
A Six Sigma team constructs a histogram of customer call wait times and observes a pronounced distribution where the bulk of the data clusters near zero on the left, but a long tail stretches far out to the right, causing the sample mean to be significantly greater than the median. What distribution shape does this represent, and what does it indicate?
Negatively skewed distribution indicating an artificial upper limit on call durations
Symmetrical normal distribution indicating an equal probability of short and long calls
Bimodal distribution indicating two competing telephone routing servers
Positively skewed distribution indicating a natural physical boundary at zero with occasional extreme delays
Sections you finish are checked off in the contents.