17.3 Common Cause vs Special Cause Variation
Key Takeaways
- Common cause variation is the stable, inherent variation of a process under its current design; improving it requires system change, not one-off blame.
- Special cause variation is assignable, intermittent, or unexpected; it should be investigated, contained, and removed so the process returns to (or establishes) a stable state.
- Control charts and pattern rules help distinguish the two; points beyond limits or non-random patterns often signal special causes.
- Audit CAPA that treats common-cause problems as special causes produces endless firefighting; treating special causes as “that’s just the process” leaves hazards unaddressed.
- Auditors evaluate whether management reacts correctly: special-cause reaction for assignable events, and fundamental process improvement for common-cause performance gaps.
17.3 Common Cause vs Special Cause Variation
/practice/cqaPractice questions with detailed explanations
Why this distinction is an audit skill
W. Edwards Deming and Walter Shewhart framed process improvement around a simple truth: how you react to variation determines whether quality improves or chaos grows. If managers treat every defect as a special event, they rewrite procedures daily, punish people for system noise, and never redesign the process. If they ignore true special causes, equipment failures, wrong materials, and unauthorized changes fester.
Quality auditors evaluate whether the QMS reaction matches the variation type. That evaluation shows up in findings about CAPA effectiveness, management review, SPC use, and process control.
Definitions
Common cause variation
- Present always when the process runs as currently designed
- Produced by many small, ordinary factors (materials lot-to-lot noise, ambient conditions within normal range, method tolerance, typical operator differences under standard work)
- Forms a stable, predictable distribution over time (in control)
- Reduced only by changing the system (equipment capability, method redesign, better materials, poka-yoke, automation, etc.)
Calling common-cause output “operator error” without evidence of deviation from standard work is a classic management mistake—and an audit observation when it drives ineffective CAPA.
Special cause variation
- Arises from specific, identifiable circumstances not part of the stable system
- Often intermittent, sudden, or linked to a change event
- Examples: broken tool, wrong batch of raw material, unapproved process change, power surge, new untrained temporary worker working outside procedure, software release bug
- Requires detection, containment, root-cause removal, and verification that the assignable cause will not recur
Special causes can make the process unstable (out of statistical control). Stability must often be restored before capability metrics are trusted (next section).
Comparison table
| Dimension | Common cause | Special cause |
|---|---|---|
| Nature | Inherent to current system | Assignable / exceptional |
| Predictability | Stable pattern over time | Unpredictable until cause is known |
| Typical signal | In-control chart; random scatter within limits | Point(s) beyond limits; runs, trends, cycles, mixtures |
| Management reaction | Improve the system / process redesign | Investigate, contain, eliminate cause |
| CAPA focus | Systemic controls, design, standards | Event-specific root cause + recurrence prevention |
| Wrong reaction | Firefight each point; blame individuals for system noise | Ignore spike as “normal variation”; no investigation |
Detecting special causes (practical signals)
Control charts are the classic tool, but auditors also use domain knowledge:
Statistical pattern signals (examples used in many SPC systems):
- Point beyond control limits
- Run of points on one side of the centerline
- Trend of steadily increasing or decreasing points
- Cyclic patterns suggesting shift/tooling/setup rhythm
- Sudden shift in level after a change point
Process/knowledge signals:
- Documented change (material, method, machine, measurement, environment, people)
- Maintenance event, scrap of a specific cavity, single supplier lot
- Correlation with a known incident log entry
Auditors should not invent Shewhart rules not used by the organization—but they should verify that the organization has a rational method to separate noise from signals and acts accordingly.
Worked mini-scenario — classification
Scenario A: Fill weight X-bar chart in control for 30 subgroups; Cpk is 0.85 because the process is centered but somewhat wide vs. specs. Defects occur at a steady low rate.
→ Variation dominating the defect rate is largely common cause. Fix: reduce variation or center better via system improvement (tooling, method, specs review)—not 30 separate CAPAs for each defective unit.
Scenario B: Same process in control for months; then three consecutive points above UCL after a new resin lot arrives; bulk density is out of COA.
→ Special cause (material). Fix: quarantine resin, contain product, correct supplier/receiving controls, verify return to stability.
Scenario C: Night shift mean is always higher than day shift for months, both “in control” when charted separately, but the combined chart looks unstable.
→ Likely two systems (stratification issue). Treat as systemic difference between shifts—training, setup, or equipment—not one-off “special” points every night.
CAPA implications (audit lens)
| Situation | Effective response | Ineffective response |
|---|---|---|
| Special cause event | Containment, root cause of the assignable factor, preventive controls for recurrence | “Retrain everyone” with no link to the actual cause |
| Common cause poor capability | Process improvement project, redesign, resource for variation reduction | Writing a CAPA for every scrap unit as if each were unique |
| Unstable process | Restore control first; then reassess capability | Publishing Cpk on out-of-control data as if it were a stable capability |
| Mixed streams | Stratify and improve each system | One plant-wide average hiding both problems |
Auditor questions that reveal maturity:
- How do you know this signal is special vs. common?
- What data (chart, log, change record) support that classification?
- Does the CAPA match the classification?
- After action, did the process stabilize and did the metric improve for the right reason?
Tampering and over-adjustment
A related concept (Deming): tampering—adjusting a stable process in reaction to common-cause noise—increases variation. Example: operators tweak a knobs every time a single unit is slightly high or low without evidence of special cause. Auditors who see constant micro-adjustments, “tribal” compensations, or undocumented tweaks should connect that behavior to increased dispersion and weak process control—often a system finding, not a single bad actor.
Audit evidence sources
- Control charts and reaction plans
- Nonconformance and CAPA databases (repeat causes vs. unique events)
- Change control and maintenance logs aligned to signal dates
- Management review actions (system projects vs. only firefighting KPIs)
- Interviews: Do supervisors distinguish “this is our normal noise” from “this is different”?
Exam anchors (Apply)
- Differentiate common vs. special with scenario data—not slogans
- Map reaction type to CAPA vs. system improvement
- Reject capability claims on unstable processes when the scenario shows special causes unaddressed
- Recognize stratification and mixed streams as system issues
- Avoid automatic blame of people for common-cause outcomes
Link to the rest of Domain V
A stable plating process shows an in-control chart for six months with Cpk = 0.9 and a steady small scrap rate. Management opens a separate CAPA for each scrap unit naming “operator error.” What is the best auditor assessment?
An X-bar chart has been in control. After a weekend maintenance event, two points exceed the upper control limit and thickness is high until a worn roller is replaced; then the chart returns to prior limits. Best classification?
Operators adjust a stable filling machine after every single bottle that is slightly above or below target, even when the process is in statistical control. Over months, variation increases. This behavior is best described as:
Why should an auditor challenge a published Cpk of 1.67 when the same period’s control chart shows multiple points beyond control limits and no investigation?