8.3 Failure Modes and Effects Analysis (FMEA) in HTM

Key Takeaways

  • Joint Commission NPG.02.03.01 EP 7 requires hospitals to select at least one high-risk process every 18 months and conduct a proactive risk assessment.
  • A traditional FMEA maps a process, identifies failure modes, and scores each for Severity, Occurrence, and Detection on 1–10 scales.
  • The Risk Priority Number (RPN = S × O × D) ranges from 1 to 1,000, and Detection is scored inversely (1 = almost certain detection; 10 = undetectable).
  • High-severity failure modes (for example, S ≥ 8) deserve mitigation even when their RPN is low, because rare events can still kill patients.
  • The VA's Healthcare FMEA (HFMEA) uses five steps and a severity × probability hazard score with a decision tree instead of an RPN.
Last updated: September 2026

Failure Modes and Effects Analysis (FMEA) in HTM

While Root Cause Analysis (RCA) is inherently reactive—triggered only after an adverse event or sentinel injury has occurred—healthcare technology leaders must also champion proactive risk assessment. Waiting for a patient injury to identify a latent equipment hazard violates the principles of high-reliability healthcare organizations.


1. Proactive Risk Assessment vs. Reactive Incident Investigation

The Joint Commission Mandate for Proactive Risk Assessment

Joint Commission National Performance Goal NPG.02.03.01 EP 7 (effective January 2026; formerly LD.03.09.01 EP 7, and earlier LD.04.04.05) requires hospitals to select at least one high-risk process every 18 months and conduct a proactive risk assessment. EP 8 requires the hospital to use the results to improve safety. New technology deployments, complex handoffs, and high-hazard interventions are common choices.

Originally formulated in military and aerospace engineering (MIL-STD-1629A) and adapted by the Department of Veterans Affairs (VA) National Center for Patient Safety as Healthcare Failure Modes and Effects Analysis (HFMEA), FMEA provides a prospective, structured, team-based engineering methodology to identify where, how, and why a process or medical device might fail, allowing HTM leaders to implement safeguards before clinical deployment.

Two scoring styles. Traditional FMEA scores Severity, Occurrence, and Detection and multiplies them into a Risk Priority Number (RPN), as shown below. The VA's HFMEA uses five steps — define the topic, assemble the team, graphically describe the process, conduct the hazard analysis, and identify actions and outcome measures. It scores each failure mode's severity and probability on 4-point scales (hazard score = severity × probability, maximum 16) and then uses a decision tree that asks about single-point weaknesses, existing control measures, and detectability. Either method meets the proactive risk assessment requirement when applied rigorously.


2. An 8-Step FMEA Workflow for HTM

Executing an effective HTM FMEA requires a disciplined, step-by-step approach:

Step 1: Select High-Risk Technology Process
   ↓
Step 2: Assemble Multidisciplinary Expert Team
   ↓
Step 3: Process Flow Mapping & Failure Mode Identification
   ↓
Step 4: Determine Failure Causes & Clinical Effects
   ↓
Step 5: Assign Numeric Scores (Severity, Occurrence, Detection)
   ↓
Step 6: Calculate Risk Priority Number (RPN = S × O × D)
   ↓
Step 7: Prioritize Actionable Hazards (RPN > 100 or S ≥ 8)
   ↓
Step 8: Implement Engineering Controls & Recalculate Post-RPN

Step 1: Select a High-Risk Clinical Technology Process

The HTM department collaborates with Risk Management and the Quality Committee to identify a vulnerable technology-dependent workflow. Ideal candidates include deploying a new enterprise wireless telemetry network, transitioning to automated medication dispensing cabinets, introducing robotic surgical platforms, or integrating smart infusion pumps with electronic health record (EHR) auto-programming.

Step 2: Assemble a Cross-Functional Team

An FMEA cannot be performed by biomedical technicians in isolation. A robust team must represent all operational touchpoints:

  • Clinical Engineering / HTM leaders and biomedical technicians;
  • Frontline nursing staff and nurse educators;
  • Attending physicians or surgical department representatives;
  • Clinical Pharmacists (for drug delivery technology);
  • Healthcare Information Technology (HIT) and wireless network engineers;
  • Hospital Risk Management and Patient Safety officers.

Step 3: Map the Process Flow & Identify Failure Modes

The team deconstructs the selected process into sequential, granular operational steps. For each sub-step, the team asks: "In what ways could this step fail to perform its intended function?" These potential breakdowns are documented as Failure Modes.

Step 4: Determine Failure Causes & Clinical Effects

For every identified failure mode, the team determines:

  • Potential Causes: The root physical, technological, environmental, or human mechanisms that trigger the failure (e.g., component wear, battery drain, Wi-Fi blind spots, user misinterpretation).
  • Clinical Effects: The immediate and downstream clinical consequences to the patient (e.g., missed lethal arrhythmia alarm, drug under-dose, prolonged surgical time, hypoxic brain damage).

Step 5: Assign Numeric Ratings (Severity, Occurrence, Detection)

The multidisciplinary team scores each failure mode across three dimensions on standardized 1 to 10 ordinal scales:

  • Severity (S): Assesses the seriousness of the worst-case clinical effect on the patient.
  • Occurrence (O): Estimates the likelihood or frequency that the failure cause will occur during routine operation.
  • Detection (D): Rates the likelihood that the failure will be detected and intercepted by existing technical controls, alarms, or procedures before reaching the patient.

Step 6: Calculate the Risk Priority Number (RPN)

The team calculates the Risk Priority Number by multiplying the three scores: RPN=Severity×Occurrence×Detection\text{RPN} = \text{Severity} \times \text{Occurrence} \times \text{Detection} Because each factor ranges from 1 to 10, the total RPN spans a theoretical range from 1 to 1,000.

Step 7: Prioritize Critical Failure Modes for Action

The team analyzes the RPN distribution to establish intervention priorities. Healthcare organizations typically apply two critical thresholds:

  1. Aggregate RPN Threshold: Failure modes with an $\text{RPN} \ge 100$ (or within the top 20% of RPN values) require mandatory corrective action.
  2. Critical Severity Rule ($S \ge 8$): Any failure mode that carries a Severity rating of 8, 9, or 10 must be prioritized for risk mitigation regardless of its aggregate RPN. Even if an event has an occurrence of 1 (extremely rare) and high detection (D = 2), yielding an RPN of only 18, a potential consequence of patient death or irreversible injury cannot be dismissed as acceptable risk.

Step 8: Implement Risk Mitigation Actions & Recalculate RPN

The team designs and implements concrete engineering controls, system redesigns, and procedural interlocks to eliminate the failure cause or drastically improve detection. Once safeguards are active, the team re-evaluates the failure mode and calculates a Post-Mitigation RPN to prove measurable risk reduction.


3. Comprehensive FMEA Rating Scales & Scoring Criteria

Objective scoring requires rigorous calibration across team members. Subjective bias can artificially skew RPN scores.

Understanding the Detection Scale (The Inverse Logic)

A frequent point of confusion on certification exams and in clinical practice is the Detection (D) rating. Unlike Severity and Occurrence—where higher numbers represent worse outcomes and higher frequencies—Detection is an inverse rating:

  • A score of 1 represents Almost Certain Detection: The failure is immediately recognized and intercepted by an automated, fail-safe mechanism before it can impact the patient.
  • A score of 10 represents Absolute Undetectability: The failure is completely hidden, silent, and impossible to detect with existing controls, silently propagating to the patient.

Complete FMEA Rating Scale Matrix (1–10)

ScoreSeverity ($S$) — Clinical ImpactOccurrence ($O$) — FrequencyDetection ($D$) — Interception Probability
10Catastrophic: Patient death without warning; fatal regulatory violationExtremely High: Inevitable; occurs > 1 in 10 clinical usesUndetectable: Zero chance of detection; latent failure reaches patient silently
9Critical: Permanent life-altering disability or loss of organ functionVery High: Occurs frequently; approximately 1 in 20 usesExtremely Remote: Very unlikely to be detected prior to patient impact
8Major: Severe temporary harm requiring ICU admission or life supportHigh: Repeated occurrences; approximately 1 in 100 usesRemote: Low probability of detection; depends on rare visual cue
7Moderate-High: Significant harm requiring surgical or medical interventionModerately High: Occurs regularly; approximately 1 in 500 usesVery Low: Detected only through complex manual testing
6Moderate: Reversible injury requiring medical treatment or extended stayModerate: Occurs occasionally; approximately 1 in 1,000 usesLow: Detected through multi-step manual clinical checks
5Low-Moderate: Minor injury requiring outpatient care or observationLow-Moderate: Infrequent; approximately 1 in 5,000 usesModerate: 50% probability of detection during clinical setup
4Minor: Discomfort or minor delay; no medical intervention requiredLow: Relatively rare; approximately 1 in 10,000 usesModerately High: Likely to be detected during routine pre-use check
3Very Minor: Inconvenience to patient or staff; slight procedural delayVery Low: Rare; approximately 1 in 50,000 usesHigh: Automated alarm or prompt warns operator during use
2Negligible: Nuisance alert; zero patient discomfort or clinical impactRemote: Extremely rare; approximately 1 in 100,000 usesVery High: Immediate automated warning or lockout catches defect
1None: Completely unnoticeable; zero operational or clinical effectNearly Impossible: < 1 in 1,000,000 clinical usesAlmost Certain: Built-in hardware interlock prevents operation

4. Worked HTM FMEA Case Study: Enterprise Wireless Telemetry Monitoring System

To demonstrate the practical application of FMEA in clinical engineering, consider a worked analysis conducted during the hospital-wide replacement of an Enterprise Wireless Patient Telemetry Monitoring System.

Clinical Context & Scope

The hospital is replacing 250 ambulatory cardiac telemetry transmitters. Ambulatory cardiac patients wear battery-powered transmitters that broadcast real-time electrocardiogram (ECG) data over the hospital Wi-Fi network to a centralized monitoring station staffed by telemetry technicians.

Worked FMEA Worksheet

Process StepPotential Failure ModePotential Cause of FailureClinical Effect on PatientPre-$S$Pre-$O$Pre-$D$Pre-RPNCorrective / Preventive Action (Engineering Control)Post-$S$Post-$O$Post-$D$Post-RPNRPN Reduction
1. Signal TransmissionRF packet drop / blind spot during patient ambulationInadequate 5 GHz Wi-Fi access point coverage in diagnostic corridorsCentral monitor loses telemetry feed; undetected ventricular fibrillation; delayed resuscitation967378Conduct comprehensive RF site survey; deploy redundant medical-grade APs; configure transmitter to emit local audible alarm if out of network range for > 10 sec92236-342 (-90%)
2. Electrode Lead InterfaceECG lead wire detachment masquerading as flatline artifactElectroconductive gel desiccation or patient movement pulling snap leadTelemetry technician misinterprets as cardiac asystole, triggering false code blue, or ignores true arrest as artifact875280Standardize to active impedance-sensing snap leads; configure central station software with differentiated "Lead Off" acoustic chime distinct from lethal arrhythmia alarms83124-256 (-91%)
3. Power ManagementTransmitter battery depletion during active monitoringClinical staff failing to replace disposable alkaline batteries at scheduled shift intervalsAbrupt cessation of monitoring without warning; unobserved cardiac decompensation866288Replace disposable batteries with smart lithium rechargeable packs containing fuel-gauge microcontrollers; configure central station and mobile nurse pagers to alarm at 20% battery reserve82232-256 (-89%)
4. Alarm NotificationTelemetry technician alert fatigue causing delayed notificationExcessive non-actionable alarms (PVCs, high heart rate threshold set too low) masking critical eventsCritical delay in dispatching primary nurse to bedside during ventricular tachycardia886384Implement unit-specific evidence-based alarm defaults; establish delay timers for transient self-limiting tachycardia; integrate automated secondary escalation to nurse smartphones83248-336 (-88%)

Analysis of Case Study Outcomes

In this worked example, each of the four identified failure modes initially exhibited an RPN exceeding 250, driven by moderate-to-high occurrence rates and poor detection mechanisms. By implementing high-leverage engineering controls—such as redundant RF coverage, smart impedance-sensing hardware, intelligent battery fuel-gauging, and automated nurse escalation algorithms—the HTM team dramatically reduced Occurrence ($O$) and Detection ($D$) scores.

Critically, notice that Severity ($S$) remained unchanged at 8 or 9 across all four failure modes. In clinical engineering risk management, an engineered intervention rarely reduces the intrinsic severity of a clinical harm: if a patient in ventricular fibrillation is unmonitored, the physiologic severity remains catastrophic ($S = 9$). What the engineering controls accomplish is driving the probability of occurrence down to rare levels ($O = 2 \text{ or } 3$) and driving detection to near-certainty ($D = 1 \text{ or } 2$), reducing the overall Risk Priority Numbers by nearly 90% and protecting patient lives.

Loading diagram...
The 8-Step FMEA Workflow for HTM
Test Your Knowledge

An HTM multidisciplinary team is conducting a Failure Modes and Effects Analysis (FMEA), using Severity, Occurrence, and Detection scores, on the hospital's fleet of automated external defibrillators (AEDs) deployed in public ambulatory pavilions. One identified failure mode is "high-voltage capacitor failure resulting in inability to deliver shock during cardiac arrest." The team notes that the AED model features an automated daily internal diagnostic circuit that tests capacitor charging and flashes an external red LED status beacon while sounding an audible chirp if a fault occurs. How does this automated self-test feature influence the Failure Modes and Effects Analysis scoring?

A
B
C
D
Test Your Knowledge

A clinical engineering team performs an FMEA on a robotic surgical system prior to clinical go-live. The team calculates RPN scores for several failure modes:

  • Failure Mode A: Robotic arm servo motor stalls during organ retraction, causing sudden tissue laceration and severe hemorrhage (Severity = 10, Occurrence = 1, Detection = 3, RPN = 30).
  • Failure Mode B: Video console auxiliary touchscreen freezes, requiring a 60-second software reboot during non-critical surgical port placement (Severity = 3, Occurrence = 7, Detection = 5, RPN = 105).
  • Failure Mode C: Foot pedal coag switch experiences intermittent wire wear, requiring surgeon to depress pedal twice to activate (Severity = 4, Occurrence = 4, Detection = 4, RPN = 64). Based on standard healthcare risk assessment and FMEA prioritization rules, how should the HTM leadership prioritize corrective action?

A
B
C
D
Test Your Knowledge

A hospital Environment of Care (EOC) Committee is planning its annual risk management agenda. The HTM Director proposes conducting an FMEA on the hospital's central vacuum and medical gas pipeline distribution system, while a nursing administrator suggests that FMEAs are only legally required following an external Joint Commission survey citation. How should the HTM Director clarify the regulatory requirement and clinical justification for proactive risk assessments under Joint Commission standards?

A
B
C
D