5.6 Risk Management & Assessment

Key Takeaways

  • Risk Score is quantified as the product of Severity and Probability (RS = S x P), establishing boundaries for acceptable, ALARP, and unacceptable risk zones.
  • Failure Mode identifies how a component fails, Failure Effect describes the downstream customer impact, and Failure Cause pinpoints the underlying initiation driver.
  • Risk Mitigation encompasses four primary strategies: Avoidance (elimination), Reduction (probability/severity control), Transfer (liability shifting), and Acceptance (retaining residual risk).
  • The ALARP (As Low As Reasonably Practicable) principle dictates that risk in the yellow zone must be reduced unless the mitigation cost is grossly disproportionate to the risk reduction achieved.
Last updated: July 2026

5.6 Risk Management & Assessment

BoK placement note: Formal enterprise risk management is Domain VII. This section focuses on how continuous-improvement teams use risk scoring to prioritize preventive action, poka-yoke, and robust process changes before defects escape—bridging Domain V preventive tools with risk thinking.

Risk management is a proactive engineering discipline designed to identify, evaluate, prioritize, and mitigate vulnerabilities across product design and manufacturing lifecycles. By integrating quantitative risk assessment into quality engineering, organizations minimize financial losses, prevent safety hazards, and ensure compliance with international standards such as ISO 9001 and ISO 14971.


Risk Matrix: Probability vs. Severity Rating

A Risk Matrix is a qualitative or semi-quantitative tool used to categorize risk levels by evaluating the likelihood of a failure occurring (Probability / Occurrence) against the magnitude of its impact (Severity).

Constructing a $5 \times 5$ Risk Matrix

Risk is quantitatively evaluated as the product of Severity ($S$) and Probability ($P$):

Risk Score (RS)=Severity (S)×Probability (P)\text{Risk Score (RS)} = \text{Severity } (S) \times \text{Probability } (P)

Severity Level ($S$)DescriptionProbability Level ($P$)Quantitative Likelihood
5 - CatastrophicSystem loss, critical injury, regulatory shutdown5 - Frequent$> 10%$ failure rate ($> 1$ in 10)
4 - CriticalMajor system damage, severe process disruption4 - Probable$1%$ to $10%$ failure rate
3 - ModerateMinor injury, noticeable performance degradation3 - Occasional$0.1%$ to $1%$ failure rate
2 - MinorMinor process delay, slight customer annoyance2 - Remote$0.01%$ to $0.1%$ failure rate
1 - NegligibleImperceptible effect, no operational impact1 - Improbable$< 0.01%$ failure rate ($< 1$ in 10,000)
                       SEVERITY (S)
           1-Negligible 2-Minor 3-Moderate 4-Critical 5-Catastrophic
PROBABILITY (P)
5-Frequent       5        10       15         20           25
4-Probable       4         8       12         16           20
3-Occasional     3         6        9         12           15
2-Remote         2         4        6          8           10
1-Improbable     1         2        3          4            5

Legend: 1-6 = Low Risk (Green) | 8-12 = Medium Risk (Yellow) | 15-25 = High Risk (Red)

Risk Appetite and the ALARP Principle

Organizations establish risk boundaries to govern mitigation requirements:

  • Unacceptable Region (High Risk / Red, RS $\ge 15$): Risk must be mitigated immediately regardless of cost; operation cannot proceed.
  • ALARP Region (As Low As Reasonably Practicable / Yellow, $8 \le RS \le 12$): Risk is tolerated only if further reduction is economically or technologically impracticable. Cost-benefit analysis dictates action.
  • Acceptable Region (Low Risk / Green, RS $\le 6$): Residual risk is acceptable; standard process monitoring applies.

Failure Modes Identification

Systematic risk assessment requires clear differentiation between the failure elements within Design Failure Mode and Effects Analysis (DFMEA) and Process Failure Mode and Effects Analysis (PFMEA).

Key Terminology Definitions

  • Failure Mode: The specific manner or physical mechanism by which a component or process fails to satisfy its intended functional design requirements (e.g., fractured shaft, short circuit, incorrect torque).
  • Failure Effect: The downstream impact or consequence of the failure mode on higher-level assemblies, the end-user, or regulatory compliance (e.g., engine seizure, thermal runaway, brake loss).
  • Failure Cause: The specific design defect, process anomaly, or environmental factor that initiates the failure mode (e.g., hydrogen embrittlement, improper tool calibration, incorrect polymer fill rate).
[Failure Cause] ──(initiates)──> [Failure Mode] ──(results in)──> [Failure Effect]
 (e.g., Operator torque wrench   (e.g., Fastener under-torqued    (e.g., Structural joint separates 
      out of calibration)            at 15 Nm vs 35 Nm)             during vehicle operation)

Taxonomy of Failure Modes

Quality engineers categorize failure modes into five operational types:

  1. Total Loss of Function: Complete cessation of operation (e.g., broken drive belt).
  2. Partial Function: System operates below specified efficiency or output (e.g., pump operating at 60% rated flow rate).
  3. Intermittent Function: Operation toggles randomly between functional and non-functional states (e.g., loose solder joint causing intermittent signal).
  4. Degraded Function: Gradual decline in performance over time (e.g., sensor calibration drift due to thermal aging).
  5. Unintended / Inadvertent Function: System activates without command or in an unsafe state (e.g., airbag deployment without collision).

Risk Mitigation Strategies

Once risks are mapped and prioritized, quality engineers deploy four fundamental risk treatment strategies:

                         ┌─────────────────────────────────────────┐
                         │      Risk Treatment Decision Matrix     │
                         └────────────────────┬────────────────────┘
                                              │
         ┌──────────────────┬─────────────────┴──────────────────┬──────────────────┐
         ▼                  ▼                                    ▼                  ▼
  [1. Avoidance]     [2. Reduction]                       [3. Transfer]      [4. Acceptance]
  Eliminate risk     Lower Probability (Poka-Yoke)        Shift financial    Formal retention of
  source completely   or Severity (Protective Shield)      liability          residual low risk

1. Risk Avoidance

Risk avoidance involves altering the project plan, product design, or manufacturing process to completely eliminate the risk condition or risk driver.

  • Engineering Example: Eliminating a toxic solvent from a cleaning process by switching to aqueous ultrasonic cleaning, thereby avoiding chemical exposure hazards and hazardous waste regulations entirely.

2. Risk Reduction (Control & Mitigation)

Risk reduction applies engineering and administrative controls to decrease the probability of occurrence ($P$), reduce severity ($S$), or increase detection ($D$).

  • Reducing Probability: Implementing error-proofing (Poka-Yoke) fixtures that physically prevent incorrect part orientation.
  • Reducing Severity: Incorporating mechanical pressure-relief valves or structural crumple zones that absorb impact energy to lessen failure severity.

3. Risk Transfer (Sharing)

Risk transfer shifts the financial or operational burden of a risk to a third party. It does not eliminate the physical failure mode, but redistributes the consequence.

  • Engineering Example: Purchasing comprehensive warranty insurance, negotiating supplier indemnification clauses for defective component lots, or outsourcing highly hazardous chemical processing to certified specialized contractors.

4. Risk Acceptance (Retention)

Risk acceptance is an explicit managerial decision to retain residual risk without further expenditure, typically because the risk falls within defined tolerance levels or mitigation costs exceed potential loss impact.

  • Active Acceptance: Establishing a contingency fund or reserve inventory to absorb potential defect costs.
  • Passive Acceptance: Documenting minor cosmetic defects that require no corrective action or contingency planning.

Quantitative Mitigation Scenario

A manufacturer evaluates a wave soldering process defect:

  • Baseline: Severity $S = 4$, Occurrence $P = 4$. Initial Risk Score: $RS_{initial} = 4 \times 4 = 16$ (Unacceptable Red Zone).
  • Strategy Implemented: Installation of automated optical inspection (AOI) and closed-loop flux density control.
  • Post-Mitigation: Severity $S = 4$ (unaltered), Occurrence drops to $P = 2$. New Risk Score: $RS_{mitigated} = 4 \times 2 = 8$ (ALARP Yellow Zone).
  • Risk Reduction Index ($RRI$):

RRI=RSinitialRSmitigatedRSinitial×100%=16816×100%=50% risk reduction\text{RRI} = \frac{RS_{initial} - RS_{mitigated}}{RS_{initial}} \times 100\% = \frac{16 - 8}{16} \times 100\% = 50\% \text{ risk reduction}

Loading diagram...
Risk Treatment Decision Flowchart
Test Your Knowledge

Replacing a volatile chemical solvent with an aqueous ultrasonic cleaning process to eliminate workplace toxicity hazards is an example of which risk mitigation strategy?

A
B
C
D
Test Your Knowledge

In FMEA terminology, an operator using an uncalibrated torque wrench leading to a loose fastener joint that causes structural separation during operation represents which sequence?

A
B
C
D
Test Your Knowledge

An initial process risk has a Severity rating of 5 (Catastrophic) and an Occurrence rating of 4 (Probable), yielding a Risk Score of 20 (Unacceptable). Installing an error-proofing fixture reduces Occurrence to 1 (Improbable) while Severity remains 5. What is the Risk Reduction Index (RRI)?

A
B
C
D