6.15 Capability Studies, Attributes Capability, Yield Metrics, and Long-Term Capability
Key Takeaways
- A capability study requires a stable process, a validated measurement system, a verified distribution, and a rational sampling plan that spans the sources of variation.
- Natural process limits are set by the process at the mean plus or minus three standard deviations; specification limits are set by the customer and are unrelated.
- Rolled throughput yield is the product of the first-pass yields of every step and is always lower than final yield.
- For attributes data, process sigma is obtained by converting the defect proportion to a Z value and adding the 1.5 sigma shift when reporting a short-term equivalent.
- Short-term capability uses within-subgroup variation and yields Cp and Cpk; long-term performance uses total variation and yields Pp and Ppk.
Designing and conducting a capability study
A capability index computed on unsuitable data is worse than no index, because it is believed. Four prerequisites, in order:
- Validated measurement system. If gage R&R consumes 30% of the tolerance, the observed spread is substantially measurement noise and the capability index is understated.
- Statistical stability. Capability is a prediction about future output, which is only defensible if the process is in control. Plot a control chart first; if special causes are present, remove them before computing capability.
- Verified distribution. Standard capability formulas assume normality. Test it, and if the data is non-normal, transform or use a percentile method.
- Rational sampling plan. The sample must span the sources of variation the study is meant to describe.
Sampling plan design
| Study purpose | Sampling approach | Typical size |
|---|---|---|
| Short-term (machine) capability | Consecutive pieces, single setup, single lot, single operator | 30-50 consecutive units |
| Long-term (process) performance | Subgroups spread across shifts, operators, lots, and setups | 25+ subgroups of 4-5, over 2-4 weeks |
| Setup verification | Small sample immediately after each setup | 5 units per setup, several setups |
The distinction is the whole point. A short-term study deliberately excludes between-setup and between-lot variation to estimate the process's inherent potential. A long-term study deliberately includes them to estimate what the customer will actually receive.
Report with the study: the sample size, the period covered, the stratification included, the normality result, the MSA result, and the control chart evidence. A capability number without that context cannot be evaluated.
Natural process limits versus specification limits
This distinction is tested repeatedly and confused constantly.
| Natural process limits | Specification limits | |
|---|---|---|
| Source | The process itself | The customer, the design, or the regulator |
| Formula | $\bar{x} \pm 3\hat{\sigma}$ | Given; not calculated from data |
| Also called | Voice of the process | Voice of the customer |
| Appear on | Control charts (as control limits for subgroup statistics) | Histograms and capability plots |
| Change when | The process changes | The requirement changes |
The two are completely independent. A process can be in perfect statistical control and produce 100% defective output, if its natural limits sit entirely outside the specification. Conversely, a process can be wildly out of control and still produce nothing outside specification if the tolerance is wide.
Two errors follow from confusing them, and both are serious:
- Putting specification limits on a control chart. Control limits describe what the process does; specification limits describe what is required. Charting subgroup means against specification limits is meaningless, because the distribution of means is narrower than the distribution of individuals by a factor of $\sqrt{n}$.
- Adjusting the process whenever a unit is outside specification. This is tampering: reacting to common-cause variation increases variation.
Yield and defect metrics
Definitions first. A unit is the item produced; a defect is any single non-conformance; a defective is a unit with one or more defects; an opportunity is a chance for a defect on a unit.
First pass yield and rolled throughput yield
First pass yield (FPY) for a step is the proportion of units passing that step with no rework. Rolled throughput yield (RTY) is the probability that a unit passes every step with no rework anywhere:
RTY is always lower than final yield, because final yield counts reworked units as good. The gap between them is the size of the hidden factory.
Worked example, four steps:
| Step | Units in | Passed first time | FPY |
|---|---|---|---|
| 1 Data entry | 1,000 | 950 | 0.950 |
| 2 Underwriting | 950 | 902 | 0.949 |
| 3 Verification | 902 | 875 | 0.970 |
| 4 Issuance | 875 | 866 | 0.990 |
So 86.6% of applications pass cleanly all the way through, even though final yield after rework may be 99%+. The normalized yield -- the average per-step yield -- is $\sqrt[4]{0.8658} = 0.9646$, useful for comparing processes with different numbers of steps.
When defects follow a Poisson model, yield relates to DPU as:
With $RTY = 0.8658$, total $DPU = -\ln(0.8658) = 0.1441$ defects per unit across the whole process.
Capability for attributes data
There is no Cp or Cpk for attributes data, because there is no continuous distribution and no specification width. Capability is expressed as a process sigma level derived from the defect rate.
Procedure:
- Compute DPMO (or the proportion defective, if working at the unit level).
- Convert the corresponding proportion to a $Z$ value from the standard normal table. This is $Z_{LT}$, the long-term sigma level.
- Add 1.5 to report the conventional short-term equivalent: $Z_{ST} = Z_{LT} + 1.5$.
Worked example. A process produces 6,210 defects per million opportunities.
- $DPO = 0.006210$, so the yield is 0.993790.
- The $Z$ value with 0.006210 in the upper tail is $Z_{LT} = 2.50$.
- $Z_{ST} = 2.50 + 1.5 = 4.0$, so the process is reported as a "4 sigma" process.
| DPMO | Long-term Z | Reported sigma level | Yield |
|---|---|---|---|
| 308,538 | 0.5 | 2.0 | 69.15% |
| 66,807 | 1.5 | 3.0 | 93.32% |
| 6,210 | 2.5 | 4.0 | 99.379% |
| 233 | 3.5 | 5.0 | 99.9767% |
| 3.4 | 4.5 | 6.0 | 99.99966% |
The step from 3 to 4 sigma removes roughly 90% of the defects; the step from 5 to 6 sigma removes another 98.5% of what remains. This is why capability improvement gets progressively harder and why the business case for the last increment must be made on consequence, not on defect count.
Short-term and long-term capability
| Short term | Long term | |
|---|---|---|
| Variation used | Within-subgroup, $\hat{\sigma}_{within} = \bar{R}/d_2$ or $\bar{s}/c_4$ | Total, $s$ computed over all individuals |
| Indices | $C_p$, $C_{pk}$ | $P_p$, $P_{pk}$ |
| Captures | Inherent process potential | Everything, including drift between subgroups |
| Data span | Short window, homogeneous conditions | Weeks or months, all conditions |
The 1.5 sigma shift is an empirical convention originating in Motorola's observation that process means drift over time. It is a default, not a law: where enough long-term data exists, compute the actual shift as $Z_{ST} - Z_{LT}$ rather than assuming 1.5.
Interpreting the gap between the two index sets is diagnostic:
| Observation | Interpretation | Action |
|---|---|---|
| $C_{pk} \approx P_{pk}$ | Process is stable; no meaningful drift between subgroups | Reduce inherent variation to improve |
| $C_{pk} \gg P_{pk}$ | Good inherent capability, poor control over time | Attack drift: setup, tool wear, material lots, shift differences |
| Both low | Inherent variation is too large for the tolerance | Fundamental process or design change; consider DFSS |
When only short-term data is available, the convention is to subtract 1.5 from the short-term $Z$ to estimate long-term performance. When only long-term data is available, add 1.5 to state the short-term equivalent. Say which convention you used; reporting a "6 sigma process" without stating whether it is short- or long-term is ambiguous by a factor of roughly 20,000 in defect rate.
A four-step process has first pass yields of 0.95, 0.98, 0.97, and 0.99. What is the rolled throughput yield, and why does it differ from the final yield the plant reports?
A process operates in statistical control but produces 12% of units outside the specification limits. What does this indicate about natural process limits and specification limits?
A capability study reports Cpk of 1.62 and Ppk of 0.94 on the same data set. What does the gap indicate and what should the team attack?