7.2 Failure Mode and Effects Analysis (FMEA) & RPN Calculation
Key Takeaways
- Failure Mode and Effects Analysis (FMEA) is a proactive, structured engineering methodology used to anticipate potential failure modes, evaluate their operational impact, and prioritize preventive countermeasures.
- FMEA is a dynamic, living document that must be initiated during Measure/Analyze, revised in Improve to evaluate countermeasure effectiveness, and institutionalized in Control as part of the Process Control Plan.
- The Risk Priority Number (RPN) is calculated as the mathematical product of Severity, Occurrence, and Detection (RPN = S * O * D), resulting in an absolute scale ranging from 1 to 1,000.
- The Detection rating scale is inverted: a rating of 1 represents almost certain detection (foolproof automated Poka-Yoke controls), whereas a rating of 10 represents complete inability to detect the defect before it reaches the customer.
- The High-Severity Rule mandates that any failure mode with a Severity rating of 9 or 10 (indicating safety hazards or regulatory non-compliance) must be remediated immediately, regardless of its overall RPN score.
7.2 Failure Mode and Effects Analysis (FMEA) & RPN Calculation
Core Principle: Failure Mode and Effects Analysis (FMEA) is a proactive, structured risk-assessment methodology designed to anticipate potential failures, evaluate their operational impact, and prioritize corrective actions before defects reach the customer. Unlike reactive troubleshooting, FMEA quantifies risk using three independent 1-to-10 rating scales: Severity ($S$), Occurrence ($O$), and Detection ($D$). Their mathematical product yields the Risk Priority Number ($\text{RPN} = S \times O \times D$), ranging from 1 to 1,000. Crucially, any failure mode with a Severity of 9 or 10 demands mandatory corrective action under the High-Severity Rule, regardless of overall RPN. FMEA is a living document that spans the entire DMAIC lifecycle.
Foundations & Evolution of FMEA
Originally developed in the late 1940s by the United States Military (formalized under standard MIL-P-1629) to assess the reliability of military equipment, FMEA was subsequently embraced by NASA for the Apollo space program and later standardized in the automotive industry by the Automotive Industry Action Group (AIAG). In Six Sigma, FMEA serves as the primary quantitative bridge between qualitative root-cause identification and targeted process improvement.
The FMEA Lifecycle Across DMAIC
MEASURE / ANALYZE IMPROVE PHASE CONTROL PHASE
┌─────────────────────┐ ┌─────────────────────┐ ┌─────────────────────┐
│ Baseline PFMEA │ │ Updated PFMEA │ │ Living Document │
│ • Map failure modes │ ────▶ │ • Action plans │ ────▶ │ • Institutionalized │
│ • Score initial RPN │ │ • Implement controls│ │ • Process control │
│ • Highlight S ≥ 9 │ │ • Recalculate RPN │ │ plan integration │
└─────────────────────┘ └─────────────────────┘ └─────────────────────┘
DFMEA vs. PFMEA
In modern operational excellence, practitioners distinguish between two primary branches of FMEA:
- Design FMEA (DFMEA): Focuses on product design, component geometry, material selection, physical tolerances, and engineering specifications. It evaluates potential failure modes that could cause product malfunction, shortened operational life, or safety hazards due to design weaknesses (e.g., a structural bracket cracking under dynamic load due to insufficient wall thickness).
- Process FMEA (PFMEA): Focuses on the manufacturing, assembly, or service delivery process. It assumes the product design is sound and examines how variations in tooling, machinery, operator handling, environmental factors, or standard operating procedures can induce defects (e.g., an operator failing to tighten a fastener to the specified torque due to an uncalibrated pneumatic gun).
| Feature | Design FMEA (DFMEA) | Process FMEA (PFMEA) |
|---|---|---|
| Core Focus | Product architecture, materials, tolerances | Manufacturing, assembly, transactional steps |
| Primary Objective | Maximize product reliability and end-user safety | Maximize process capability, yield, and consistency |
| Initiated When? | Conceptual design and prototype engineering | Tooling design, pilot production, DMAIC Analyze |
| Failure Modes | Material fatigue, thermal expansion, corrosion | Operator error, tooling wear, misplacement, contamination |
| Primary Solutions | Redesign dimensions, change materials, widen margins | Mistake-proofing (Poka-Yoke), automated sensors, SOPs |
Architecture of the FMEA Worksheet
A standardized Process FMEA worksheet follows an 11-column analytical sequence from left to right:
Standard PFMEA Worksheet Structure
┌──────────────┬──────────────┬──────────────┬───┬──────────────┬───┬──────────────┬───┬─────┬─────────────┬─────────┐
│ Process Step │ Failure Mode │ Failure │ S │ Cause of │ O │ Current │ D │ RPN │ Recommended │ Revised │
│ / Function │ (How it fails│ Effect │ │ Failure │ │ Controls │ │ │ Actions │ RPN │
├──────────────┼──────────────┼──────────────┼───┼──────────────┼───┼──────────────┼───┼─────┼─────────────┼─────────┤
│ Solder Wave │ Insufficient │ Open circuit;│ 8 │ Conveyor │ 4 │ Visual check │ 6 │ 192 │ Install auto│ S=8,O=2 │
│ Application │ solder depth │ board failure│ │ speed too │ │ 5 boards per │ │ │ optical gage│ D=2 │
│ │ │ in field │ │ fast │ │ shift │ │ │ & interlock │ RPN=32 │
└──────────────┴──────────────┴──────────────┴───┴──────────────┴───┴──────────────┴───┴─────┴─────────────┴─────────┘
- Process Step / Function: The specific operational stage or task being analyzed.
- Potential Failure Mode: The physical, procedural, or technical way in which the step can fail to meet requirements (the defect).
- Potential Effect(s) of Failure: The direct consequence of the failure mode on downstream operations, the final assembly, or the end customer.
- Severity ($S$): Numerical rating (1 to 10) evaluating the seriousness of the failure effect.
- Potential Cause(s): The specific operational mechanism or input variable ($X$) that triggers the failure mode.
- Occurrence ($O$): Numerical rating (1 to 10) evaluating the estimated frequency or likelihood that the specific cause will occur.
- Current Process Controls: The existing mechanisms (preventative procedures or detection inspections) currently in place to stop the cause or catch the defect before customer delivery.
- Detection ($D$): Numerical rating (1 to 10) evaluating the ability of the current controls to detect the failure mode or cause before downstream release.
- Risk Priority Number (RPN): Calculated summary index: $\text{RPN} = S \times O \times D$.
- Recommended Actions & Ownership: Specific countermeasures assigned to named individuals with strict target completion dates.
- Action Results: Recalculated Severity, Occurrence, Detection, and Revised RPN following implementation.
The 1-to-10 Rating Scales: Severity, Occurrence, and Detection
To ensure consistency across appraisers, organizations adopt standardized 1-to-10 rating rubrics. The table below synthesizes standard AIAG/CSSC scoring definitions:
| Rating | Severity ($S$): Impact of Failure Effect | Occurrence ($O$): Frequency of Cause | Detection ($D$): Ability to Detect Defect |
|---|---|---|---|
| 10 | Hazardous without warning: Endangers operator or customer; non-compliant with safety regulations. | Very High / Inevitable: Defect occurs $\ge 1$ in 2 cycles; $C_{pk} < 0.33$. | Absolute Uncertainty: No controls exist; defect is undetectable prior to customer receipt. |
| 9 | Hazardous with warning: Safety hazard or regulatory violation preceded by warning signs. | Very High: Defect occurs $\approx 1$ in 3 cycles; $C_{pk} \approx 0.33$. | Very Remote: Automated controls unlikely to detect; manual inspection unreliable ($<10%$ catch rate). |
| 8 | Very High: Primary product/service function inoperable; customer extremely dissatisfied. | High: Defect occurs $\approx 1$ in 8 cycles; $C_{pk} \approx 0.51$. | Remote: Manual visual or tactile inspection post-process ($10-30%$ catch rate). |
| 7 | High: Primary function degraded; product operates at reduced performance; major customer complaint. | High: Defect occurs $\approx 1$ in 20 cycles; $C_{pk} \approx 0.67$. | Very Low: Dual visual inspection or manual gauging post-process ($30-50%$ catch rate). |
| 6 | Moderate: Secondary function inoperable (e.g., air conditioning fails, audio display blank). | Moderate: Defect occurs $\approx 1$ in 80 cycles; $C_{pk} \approx 0.83$. | Low: Manual gauging on sample lots using calibrated attribute go/no-go gages. |
| 5 | Low: Secondary function degraded; customer experiences noticeable discomfort or annoyance. | Moderate: Defect occurs $\approx 1$ in 400 cycles; $C_{pk} \approx 1.00$. | Moderate: Calibrated variable gauging on a statistical sample at the workstation. |
| 4 | Very Low: Cosmetic blemish or minor defect noticed by $>75%$ of customers; scrap/rework incurred. | Occasional: Defect occurs $\approx 1$ in 2,000 cycles; $C_{pk} \approx 1.17$. | Moderately High: Calibrated variable gauging on 100% of finished units post-process. |
| 3 | Minor: Cosmetic blemish noticed by $50%$ of customers; minor rework on-line. | Low: Defect occurs $\approx 1$ in 15,000 cycles; $C_{pk} \approx 1.33$. | High: Automated inspection inline that flags defective parts and alerts operator. |
| 2 | Very Minor: Cosmetic imperfection noticed only by discriminating customers ($<25%$). | Very Low: Defect occurs $\approx 1$ in 150,000 cycles; $C_{pk} \approx 1.50$. | Very High: Automated inline sensor that rejects defective parts automatically. |
| 1 | None: No perceptible effect on product, process, or customer. | Remote: Failure is virtually impossible; $< 1$ in 1,500,000 cycles; $C_{pk} \ge 1.67$. | Almost Certain: Mistake-proofed (Poka-Yoke) design; defect cannot physically be produced. |
CRITICAL EXAM EMPHASIS — The Detection Scale Inversion: Notice that while high Severity (10) and high Occurrence (10) represent the worst conditions, Detection is inverted. A Detection rating of 1 is the best possible score (almost certain detection via Poka-Yoke), whereas a Detection score of 10 is the worst possible score (undetectable). Green Belts frequently make calculation or interpretation errors by assuming 10 represents "100% detection."
RPN Calculation & The High-Severity Rule
The Risk Priority Number (RPN) is calculated as:
Because each factor is rated on a 1-to-10 integer scale, the theoretical range of RPN is:
The Fatal Flaw of Pure RPN Ranking
Relying blindly on numerical RPN thresholds (e.g., establishing a corporate rule that "only failure modes with RPN $> 150$ require corrective action") is a dangerous operational error. Consider two failure modes from an automotive braking assembly:
-
Failure Mode A (Hydraulic Brake Line Severing):
- Severity: $S = 10$ (Total loss of braking; life-threatening crash without warning).
- Occurrence: $O = 2$ (Isolated incident; high-quality braided steel line).
- Detection: $D = 3$ (Automated high-pressure leak check at end-of-line).
- $\text{RPN}_A = 10 \times 2 \times 3 = 60$.
-
Failure Mode B (Aesthetic Paint Chip on Caliper Casting):
- Severity: $S = 3$ (Minor cosmetic blemish; no functional impact).
- Occurrence: $O = 7$ (Frequent chipped paint during bulk bin transfer).
- Detection: $D = 8$ (Visual check under poor lighting; high miss rate).
- $\text{RPN}_B = 3 \times 7 \times 8 = 168$.
If the team prioritizes corrective actions solely by RPN, they would expend engineering resources fixing cosmetic paint chips (RPN = 168) while ignoring catastrophic brake line failure (RPN = 60)!
The High-Severity Rule (Criticality Threshold)
To prevent this fatal failure mode, Six Sigma enforces the High-Severity Rule (also called Criticality Analysis):
The High-Severity Rule: Any potential failure mode with a Severity rating of 9 or 10 MUST receive immediate, mandatory corrective action, regardless of its Occurrence, Detection, or overall RPN score.
Failure modes with Severity 9 or 10 involve catastrophic safety hazards, physical injury, or violations of federal/statutory regulations. They cannot be tolerated simply because current detection is decent or occurrence is low.
Action Plans & Recalculating Revised RPN
Once failure modes are prioritized, the team develops targeted corrective actions. An essential Green Belt competency is understanding how different countermeasures impact the individual ratings:
How Countermeasures Alter FMEA Scores
TARGET RATING ENGINEERING MECHANISM REQUIRED
┌───────────────┐ ▶ Can generally ONLY be reduced by fundamentally changing the
│ Severity (S) │ design, chemistry, or process architecture (e.g., substituting
└───────────────┘ a non-toxic water solvent for an explosive hydrocarbon solvent).
┌───────────────┐ ▶ Reduced by removing root causes, improving process capability
│ Occurrence (O)│ (tightening Cp/Cpk), preventative maintenance, or redesigning
└───────────────┘ parts to make improper assembly physically impossible.
┌───────────────┐ ▶ Reduced by introducing automated inline optical sensors, 100%
│ Detection (D) │ electronic gauging, or Poka-Yoke error-proofing that halts the
└───────────────┘ process immediately upon detecting a non-conformance.
Worked Example: Tracking RPN Reduction
Consider a medical device manufacturer packaging sterile catheters:
-
Initial State:
- Process Step: Heat sealing Tyvek pouch.
- Failure Mode: Incomplete seal perimeter.
- Effect: Catheter non-sterile; risk of severe patient sepsis ($S = 9$).
- Cause: Heat element thermal drift ($O = 5$).
- Current Controls: Manual peel test on 2 pouches per shift ($D = 7$).
- Initial RPN: $9 \times 5 \times 7 = 315$.
-
Corrective Actions Implemented:
- Installed a dual-thermocouple automated temperature feedback loop that maintains sealing temperature within $\pm 1.0^\circ\text{C}$ (reduces Occurrence from $O = 5$ to $O = 2$).
- Installed an inline continuous optical laser vision scanner that inspects 100% of seal tracks and automatically diverts unsealed pouches (reduces Detection from $D = 7$ to $D = 2$).
- Note on Severity: Because an unsealed pouch still carries the catastrophic risk of sepsis if it ever reached a patient, the fundamental Severity remains $S = 9$.
-
Recalculated Revised RPN:
The project achieved an 88.6% reduction in overall risk priority while directly addressing a critical $S = 9$ condition.
Critical Exam Traps to Avoid
- Trap 1: The Detection Scale Inversion — Assuming that $D = 10$ means "100% detection probability." Remember: 1 is best (almost certain detection), and 10 is worst (undetectable).
- Trap 2: Prioritizing Solely by Total RPN Score — Forgetting the High-Severity Rule. If a question presents four failure modes and asks which to address first, immediately check for any $S \ge 9$. If one exists, it must be addressed regardless of whether another failure mode has a higher RPN.
- Trap 3: Believing Process Controls Can Easily Reduce Severity — Thinking that installing an automated camera reduces Severity. It does not! An inspection camera reduces Detection (and possibly Occurrence if linked to an automatic interlock). Severity reflects the customer's pain if the defect occurs and can only change if the product/process is redesigned.
- Trap 4: Viewing FMEA as a One-Time Static Document — Treating FMEA as a "check-the-box" milestone completed in Analyze and archived. On Green Belt exams, FMEA is always characterized as a dynamic, living document updated through Control.
A continuous improvement team at an aerospace valve manufacturer evaluates four potential failure modes during a Process FMEA session:
During a Process FMEA review at an automated bottling plant, the team investigates the failure mode 'Capping machine applies insufficient torque, resulting in leaking seals.' To mitigate this risk, the engineering team installs an automated inline load-cell sensor that tests the removal torque of 100% of sealed bottles on the conveyor, instantly rejecting non-conforming units. How will this corrective action alter the FMEA rating parameters?
A transactional Green Belt evaluates an insurance claims processing workflow. The team identifies the failure mode 'Incorrect claimant policy ID entered into billing ledger.' The effect is 'Claim rejected by payer and resubmitted, causing a 14-day reimbursement delay' (S = 6). The cause is 'Typing error by data entry clerk' (O = 5). The current control is 'Random spot-check audit of 2% of completed files' (D = 8). What is the initial RPN, and what is the most effective Lean Six Sigma action to reduce this risk?