Process Capability Studies

Key Takeaways

  • A process capability study estimates how well a process can meet specifications for a defined CTQ, using representative data, known specs, and a valid variation model.
  • Before computing Cp/Cpk-type indices, confirm the process is in statistical control (stable) and that the data are approximately normal or handled with an appropriate method.
  • Characteristics must have clear operational definitions; specifications and tolerances must be the correct LSL/USL/target for that characteristic.
  • Rational sampling, independence assumptions, and measurement system adequacy (MSA) protect the study from garbage-in indices.
  • CSSGB BoK III.F.2 is Evaluate level: judge whether a capability study is valid and what its results mean—not just run software defaults.
Last updated: July 2026

Process Capability Studies (CSSGB BoK III.F.2 — Evaluate)

Quick Answer: A process capability study quantifies how well a stable process meets specifications for a chosen characteristic. Before trusting Cp, Cpk, Pp, or Ppk, confirm correct specs, a capable measurement system, statistical stability, and approximate normality (or a proper alternative). Evaluate-level skill means you can design the study and reject invalid “capability” claims.

What a Capability Study Is

A capability study answers: Given how this process currently runs, how does its variation and location compare with the allowed tolerance for this CTQ?

Typical study outputs:

  • Estimates of mean and standard deviation (within-subgroup and/or overall)
  • Capability indices (Cp, Cpk) and/or performance indices (Pp, Ppk), sometimes Cpm
  • Estimated percent nonconforming or DPMO under a stated model
  • Graphical evidence: control charts, histogram with specs, normal probability plot

Capability is a statement about a defined process and characteristic under the conditions sampled—not a permanent property of a machine nameplate.

Characteristics, Specifications, and Tolerances

Characteristic (Y): the measurable CTQ—diameter, cycle time, purity, pull strength, first-pass yield attribute rate, etc. Operationally define units, method, and time window.

Specifications: LSL, USL, and target T from the customer, drawing, regulation, or internal standard. Verify revision level; many “incapable” reports use obsolete limits.

Tolerance: for two-sided continuous specs, tolerance width = USL − LSL. Unilateral specs need one-sided analysis. Attribute requirements (for example “defectives ≤ 0.5%”) use different metrics (DPMO, yield)—do not force a variables Cp formula onto pure go/no-go data without a variables measurement.

Study inputEvaluate question
CharacteristicIs this the CTQ that drives customer pain?
Spec sourceCurrent, approved, and applicable to this product/process?
UnitsSame units for data and specs?
TargetMidpoint of specs, or a different customer-preferred T?
Subgroup / samplingDoes sampling match how the process generates variation?

Prerequisites: Stability First

Standard practice: estimate capability (Cp/Cpk using within-subgroup σ) only when the process is in statistical control. Out-of-control processes mix common-cause and special-cause variation; a single index then misrepresents future performance.

Stability checks before capability:

  1. Choose an appropriate control chart for the data type and subgrouping (X-bar and R/S, I-MR, etc.).
  2. Collect enough subgroups to see patterns (rules of thumb vary; many texts want ≥20–25 subgroups for a baseline study—know your organization’s standard).
  3. Resolve special causes or stratify (different shifts, tools, materials) so each studied stream is stable.
  4. Document the stable window used for the index calculation.

If leadership needs a number while the process is still chaotic, report process performance (Pp/Ppk from overall variation) as a snapshot of historical output, and clearly label that it is not a prediction of a controlled process. Do not market an out-of-control Ppk as “our capability.”

Prerequisites: Normality (and What If Not)

Classic Cp/Cpk interpretation (percent outside specs ≈ normal tail probabilities) assumes individuals are approximately normal (or that you work with a transformed metric that is).

Normality checks Green Belts use:

  • Histogram shape (skew, heavy tails, multimodality)
  • Normal probability plot (approx. straight line)
  • Formal tests (Anderson–Darling, Shapiro–Wilk) as supporting evidence—not sole decision makers on small n

If non-normal:

  • Investigate mixtures (two machines, two cavities)—stratify and study each stream
  • Consider transformation (for example Box–Cox) and compute capability on the transformed scale carefully, converting conclusions back to original units for communication
  • Use distribution-specific or nonparametric estimates of percent nonconforming
  • For strongly skewed cycle times, report percentiles and empirical % beyond specs alongside any index

Multimodality often means the “process” is not one process—capability of a blend of two setups is a management fiction.

Measurement System and Sampling Design

MSA: If gage R&R is poor, capability indices measure the measurement system + process muddle. Ensure discrimination and %GR&R are acceptable for the tolerance before a high-stakes capability claim.

Sampling for a study:

  • Rational subgroups for within-σ estimates (consecutive pieces from the same stream)
  • Cover the full range of common-cause conditions you claim the index represents (materials lots, shifts)—or explicitly narrow the claim
  • Avoid cherry-picking the “good week” after a firefight
  • Record time order so control charts are possible

Sample size: More data tighten estimates of μ and σ and of the indices. Very small n produces wide uncertainty; treat single-digit capability claims with skepticism.

Study Workflow (Evaluate-Ready)

  1. Define characteristic, CTQ link, specs/target, and population (product, line, period).
  2. Validate measurement (MSA / operational definition).
  3. Collect time-ordered data with planned subgrouping.
  4. Assess stability on control charts; act on special causes or stratify.
  5. Assess distribution shape; choose analysis path.
  6. Estimate within and overall σ as needed; compute indices and % nonconforming.
  7. Interpret with graphics: histogram + specs, capability plot, control chart appendix.
  8. Report boundaries: what process, what time window, short-term vs long-term, assumptions.

Interpreting Study Results (Without the Formula Deep Dive)

  • High Cp, low Cpk → enough potential spread vs. tolerance, but off-center—adjust location.
  • Low Cp → process variation too large for the tolerance even if centered—reduce σ or relax specs (specs only with customer authority).
  • Pp/Ppk much lower than Cp/Cpk → extra long-term / between-subgroup variation—instability or unmodeled shifts over time.
  • Stable + capable → good candidate for control-phase monitoring; still re-verify after changes.

Worked scenario: A coating thickness study uses 25 subgroups of 5. X-bar and R charts show one out-of-control point from a wrong viscosity batch; after removing that special-cause subgroup (and fixing the cause), charts are stable, histogram is roughly normal, LSL/USL confirmed. Only then does the team publish Cpk = 1.33. Publishing Cpk with the bad batch left in would understate true common-cause capability and overstate chaos as if it were permanent—or the reverse, depending on how the outlier sits. The Evaluate skill is knowing the index is conditional on the cleaned, stable story you documented.

Common Study Failures

  1. Specs wrong or one-sided treated as two-sided
  2. Capability computed during known process chaos
  3. Ignoring non-normality and multimodality
  4. Poor gage, rounded data, or inadequate discrimination
  5. Confusing a vendor’s short run with long-term plant performance
  6. Reporting four decimal places on Cpk when n is tiny

Bottom Line for III.F.2

A process capability study is a structured evaluation: right characteristic, right specs, adequate measurement, stable process, suitable distributional assumptions, then indices and nonconformance estimates. Skip the prerequisites and the number is theater. Master the study logic so Cp/Cpk/Pp/Ppk (next section) are computed only when they mean something.

Test Your Knowledge

A Green Belt calculates Cpk = 1.67 from last month’s production data. The X-bar chart for the same period shows multiple points beyond control limits, and two machines were blended in one dataset. What is the best evaluation of this study?

A
B
C
D
Test Your Knowledge

Before computing classical Cp and Cpk for a continuous CTQ, which pair of checks is most essential?

A
B
C
D