8.2 Failure Mode and Effects Analysis (FMEA)

Key Takeaways

  • Failure Mode and Effects Analysis (FMEA) is a structured, proactive risk assessment methodology designed to identify, prioritize, and eliminate potential failures before they occur.

  • FMEA evaluates failure risks across three distinct 1-to-10 ordinal dimensions: Severity (S), Occurrence (O), and Detection (D).

  • The Risk Priority Number is calculated as RPN = Severity * Occurrence * Detection, yielding a potential composite score ranging from 1 to 1,000.

  • Detection scoring operates on an inverse scale, where 1 denotes near-certain detection by automated mistake-proofing and 10 denotes complete undetectability.

  • Any failure mode exhibiting a critical Severity score (9 or 10, indicating safety hazards or regulatory noncompliance) mandates corrective action regardless of the total RPN.

Last updated: September 2026

Failure Mode and Effects Analysis (FMEA)

Quick Answer: Failure Mode and Effects Analysis (FMEA) is a structured, proactive risk assessment tool used to identify, evaluate, and prioritize potential failure modes and their operational causes. Risks are scored across three 1-to-10 ordinal dimensions: Severity (SS, effect on customer), Occurrence (OO, probability of cause), and Detection (DD, capability of current controls to catch defects before escape). Multiplying these yields the composite Risk Priority Number (RPN=S×O×DRPN = S \times O \times D, range 1–1,000). Any item with a Severity score of 9 or 10 demands immediate corrective action regardless of the overall RPN. Independent CSSYB study guide by OpenExamPrep.

Purpose and Origins of FMEA

Organizations frequently struggle with reactive quality management: waiting for defects, equipment crashes, or customer complaints to occur before initiating root cause investigations. In contrast, Failure Mode and Effects Analysis (FMEA) provides a forward-looking, proactive methodology that identifies vulnerabilities during product or process development, allowing teams to engineer safeguards before failures manifest.

Originally developed in 1949 under U.S. military standard MIL-P-1629, FMEA gained widespread prominence during NASA's Apollo space missions, where mission-critical failures had to be prevented. In the 1970s, the automotive industry (led by Ford and codified by the Automotive Industry Action Group [AIAG]) established FMEA as a cornerstone of quality engineering. Today, FMEA is widely used across healthcare, aerospace, transactional operations, and manufacturing.

In Six Sigma DMAIC projects, FMEA bridges the Analyze and Improve phases. Teams leverage FMEA to dissect process maps, prioritize risk factors, identify critical input variables (XX's), and establish mistake-proofing priorities.

Types of FMEA

  • Design FMEA (DFMEA): Evaluates product designs, component geometry, material properties, tolerances, and interface connections before tooling release. It assumes the manufacturing process meets specifications and focuses on product robustness.
  • Process FMEA (PFMEA): Analyzes manufacturing, assembly, or transactional service workflows. It assumes the product design is sound and focuses on process step failures driven by operator technique, tooling wear, machine parameters, methods, and environmental conditions.

FMEA Structure and Terminology

An FMEA documents process vulnerabilities using a structured matrix containing key analytical elements:

  1. Process Step / Function: The specific process operation under review and its intended performance requirement.
  2. Potential Failure Mode: The physical, technical, or operational way in which the process step could fail to meet its intended function (e.g., incorrect fastener torque, missing data field, chemical overheating, contaminated fluid).
  3. Potential Effect of Failure: The downstream consequence or adverse impact experienced by subsequent internal operations, the final assembly, or the customer (e.g., vehicle stalls in traffic, structural fracture, billing dispute, regulatory sanction).
  4. Potential Cause: The specific operational root cause or trigger leading to the failure mode (e.g., operator fatigue, worn calibration spring, software buffer overflow, thermal sensor drift).
  5. Current Process Controls: The existing mechanisms deployed to mitigate risk:
    • Prevention Controls: Measures that prevent the failure cause from occurring (e.g., physical poka-yoke guide pins, automated shut-off limits, preventative maintenance schedules).
    • Detection Controls: Measures that detect the failure mode or cause after it occurs but before the deliverable escapes to the customer (e.g., manual visual inspection, optical barcode scanners, automated pressure testing).

The 1 to 10 Evaluation Criteria

Every identified failure mode is evaluated and scored on an ordinal scale from 1 to 10 across three critical criteria:

1. Severity (S)

Severity rates the seriousness of the failure effect on the customer or downstream operation:

  • 1: Minor or negligible effect; customer does not notice or experience performance degradation.
  • 4–6: Moderate disruption; customer experiences minor inconvenience or loss of secondary function.
  • 7–8: High severity; customer experiences primary system malfunction or major performance loss.
  • 9–10: Hazardous or catastrophic severity; poses severe safety risks to human life or violates governmental regulations, occurring with warning (9) or without warning (10).

Critical Rule on Severity: Severity can only be reduced by fundamentally redesigning the product or altering the core process architecture. Adding inspections, improving test gauges, or providing operator training does not alter how severe the failure is when it occurs; it only impacts detection or occurrence!

2. Occurrence (O)

Occurrence rates the likelihood or frequency that the specific failure cause will occur:

  • 1: Remote; prevention controls have effectively eliminated the cause, or it has essentially never been observed.
  • 4–6: Moderate; the cause produces occasional failures at a predictable rate.
  • 9–10: Very high; the cause produces failures frequently, and they are almost inevitable without new controls.
  • Each organization's rating table ties these ranks to specific failure rates (automotive tables, for example, express occurrence as failures per thousand). Use the table your organization or customer specifies, and keep it the same when you re-score after improvements.
  • Occurrence is reduced by eliminating causes, mistake-proofing (poka-yoke), and tightening process controls.

3. Detection (D)

Detection rates the ability of current controls to detect the failure mode or cause before it escapes to the customer:

  • Inverse Scaling: Detection uses an inverse scale where lower numerical scores represent superior detection capability!
  • 1: Almost certain detection; 100% automated mistake-proofing (poka-yoke) or interlocking sensors halt the line immediately upon an anomaly.
  • 4–6: Moderate detection; statistical sampling, manual gauge checks, or automated sorting systems catch defects with moderate reliability.
  • 9–10: Completely undetectable; no current controls exist, or controls rely on subjective human visual inspection with high escape probabilities.

Risk Priority Number (RPN) Calculation & Action Prioritization

The composite risk score is calculated by multiplying the three individual criteria:

Risk Priority Number (RPN)=Severity (S)×Occurrence (O)×Detection (D)\text{Risk Priority Number (RPN)} = \text{Severity (S)} \times \text{Occurrence (O)} \times \text{Detection (D)}

Because each factor ranges from 1 to 10, the total RPN ranges from 1 to 1,000.

Worked Example:

Consider an automated industrial bottle-capping operation in a beverage facility:

  • Process Step: Apply tamper-evident threaded closures to bottles.
  • Potential Failure Mode: Under-torqued cap (loose seal).
  • Potential Effect: Product leakage during transit, customer contamination risk (S=8S = 8).
  • Potential Cause: Drive belt slippage on the capping spindle (O=4O = 4, occurs occasionally).
  • Current Controls: Operator inspects 5 bottles manually every 2 hours (D=6D = 6, moderate escape risk).
  • RPN Calculation: RPN=8×4×6=192\text{RPN} = 8 \times 4 \times 6 = 192

The High-Severity Action Imperative

Many organizations historically applied arbitrary RPN thresholds (e.g., "any item with RPN >150> 150 requires action"). Modern quality engineering rejects relying exclusively on composite RPN. For example, the 2019 AIAG & VDA FMEA Handbook, widely used in automotive supply chains, replaced RPN ranking with Action Priority (AP) tables that weight severity first. The CSSYB Body of Knowledge still tests the RPN calculation, so know both the formula and its limitation.

Mandatory Priority Rule: Any failure mode with a Severity score of 9 or 10 must be addressed immediately with preventative actions, regardless of the total RPN score. Even if Occurrence is 1 and Detection is 1 (yielding an RPN of only 9 or 10), the catastrophic risk to human safety or regulatory compliance demands fail-safe preventative controls.

Post-Mitigation Revision

Once permanent corrective actions are implemented (such as installing a torque transducer with an automated reject arm), teams re-evaluate the process. If Occurrence drops to 1 and Detection drops to 1, the Revised RPN drops from 192 to:

Revised RPN=8×1×1=8\text{Revised RPN} = 8 \times 1 \times 1 = 8

This post-mitigation recalculation mathematically confirms significant risk reduction.

Loading diagram...
FMEA Risk Assessment & Mitigation Workflow
Test Your Knowledge

A project team analyzing a medical packaging process calculates the following scores for a heat-sealing failure mode: Severity = 9 (sterile barrier breach), Occurrence = 2 (rare occurrence), Detection = 2 (automated continuous thermal pressure sensor shuts down line upon deviation). What is the calculated Risk Priority Number (RPN), and what action is required?

A

RPN = 13; no action is required because the score is below the traditional 100-point threshold

B

RPN = 72; action can be deferred because the automated detection control is robust

C

RPN = 36; only routine monitoring is necessary because Occurrence is low

D

RPN = 36; immediate action is mandatory because Severity is 9, indicating a critical safety/regulatory hazard

Test Your Knowledge

Which parameter in a Failure Mode and Effects Analysis (FMEA) can ONLY be reduced by fundamentally redesigning the product or changing the core process architecture?

A

Severity

B

Occurrence

C

Detection

D

Risk Priority Number

Test Your Knowledge

In evaluating the Detection score within an FMEA, which of the following scenarios represents the lowest (most favorable) Detection numerical rating?

A

Periodic manual visual audits conducted by lead operators once per shift

B

An automated photoelectric sensor that immediately interlocks and shuts down the conveyor if a part is misaligned

C

End-of-line statistical quality control sampling inspecting 5 parts out of every batch of 100

D

Post-delivery customer reporting through an online warranty claims portal

Sections you finish are checked off in the contents.