4.2 Alarm Management & EEMUA 191 Guidelines
Key Takeaways
- EEMUA Publication 191 is the definitive international guidance document for alarm system specification, design, management, and continuous performance monitoring in process industries.
- The EEMUA 191 recommended target alarm priority distribution is approximately 80% Low Priority, 15% Medium Priority, and 5% High (or Urgent) Priority.
- An alarm flood is defined as a state where the control room operator receives more alarms than can be effectively processed, quantitatively defined as > 10 alarms per 10-minute period.
- Nuisance alarms—such as chattering, fleeting, standing, or bad-actor alarms—degrade operator situational awareness, induce alarm fatigue, and increase the risk of operators ignoring critical safety alarms.
- Alarm shelving (temporary manual suppression) must be strictly governed with mandatory authorization, time-outs, reason logging, and clear visual indications to prevent safety-critical barriers from being permanently disabled.
The Role and Purpose of Process Alarms
In modern computer-controlled process plants governed by Distributed Control Systems (DCS) and Programmable Logic Controllers (PLCs), the alarm system is the primary interface for alerting control room operators to abnormal process conditions.
According to EEMUA Publication 191 (Alarm Systems: A Guide to Design, Management and Procurement) and ANSI/ISA-18.2 / IEC 62682 (Management of Alarm Systems for the Process Industries), an alarm is explicitly defined as:
"An audible and visual means of informing the operator of an abnormal process condition or equipment malfunction that requires a timely human operator response."
+---------------------------------------------------------------+
| What Qualifies as an Alarm? |
+---------------------------------------------------------------+
| MUST REQUIRE A TIMELY OPERATOR ACTION! |
| If no operator response is possible or required: |
| --> It is NOT an alarm. It is an EVENT or STATUS MESSAGE. |
+---------------------------------------------------------------+
Critical Distinction: Alarm vs. Event vs. Safety Trip
- Alarm: Requires an immediate, defined human operator action to prevent an undesirable consequence (e.g., "High Reactor Temperature - Open Secondary Cooling Valve").
- Event / Status Message: An informational notification recording normal equipment state changes (e.g., "Pump P-101 Started" or "Filter Switch Complete"). Events should never trigger audible annunciators or flash on primary alarm screens.
- Safety Trip (SIS): An automated protective action executed independently by a logic solver to bring the process to a safe state without waiting for human intervention.
EEMUA Publication 191 Framework & Lifecycle
With the transition from physical panel-mounted annunciator tiles to software-based DCS consoles in the 1980s and 1990s, configuring an alarm shifted from a costly physical wiring task to a zero-cost software click. This led to massive "alarm inflation," where engineers configured alarms on virtually every instrument signal, overwhelming operators with thousands of useless notifications during process trips.
EEMUA 191 was established by the Engineering Equipment and Materials Users Association to restore rigor to alarm system design through a structured Alarm Management Lifecycle:
+--------------------------------+
| 1. Alarm Philosophy |
+--------------------------------+
|
v
+--------------------------------+
| 2. Alarm Identification |
+--------------------------------+
|
v
+--------------------------------+
| 3. Alarm Rationalization |
+--------------------------------+
|
v
+--------------------------------+
| 4. Detailed Design & Pri. |
+--------------------------------+
|
v
+--------------------------------+
| 5. Implementation |
+--------------------------------+
|
v
+--------------------------------+
| 6. Operation & Maintenance |
+--------------------------------+
|
v
+--------------------------------+
| 7. Monitoring & Audit (MOC) |
+--------------------------------+
- Alarm Philosophy: Setting company-wide standards for alarm definitions, priority rules, color coding, and performance targets.
- Alarm Identification: Identifying candidate alarms during HAZOP or risk assessment studies.
- Alarm Rationalization: Reviewing every candidate alarm against strict criteria: Is it documented? Does it have a unique cause? Is there a defined operator action? What is the consequence of inaction?
- Detailed Design: Assigning priority, deadbands, delay timers, and alarm limits based on rationalization data.
- Implementation & Operations: Commissioning configured alarms and training operators on alarm response procedures (ARP).
- Management of Change (MOC) & Audit: Ensuring no alarm threshold or priority is altered without formal engineering review and auditing system performance metrics continuously.
Alarm Prioritization & EEMUA 191 Target Distribution
Alarm priority determines the visual Annunciator color, audible tone, and screen positioning presented to the operator. Priority must be assigned strictly on the basis of severity of consequence and time available for operator response.
+--------------------------------+
| High / Urgent Priority | ~5% Target
| (Red / Flash / Heavy Tone) |
+--------------------------------+
| Medium Priority | ~15% Target
| (Amber / Yellow Tone) |
+--------------------------------+
| Low Priority | ~80% Target
| (Cyan / Blue Tone) |
+--------------------------------+
EEMUA 191 Recommended Priority Distribution Targets:
- Low Priority (~80% of total configured alarms): Minor deviations where operators have ample time (e.g., > 15-30 minutes) to take non-critical corrective steps.
- Medium Priority (~15% of total configured alarms): Moderate deviations threatening equipment damage or minor environmental release, requiring action within 5 to 15 minutes.
- High / Urgent Priority (~5% of total configured alarms): Severe safety, environmental, or major financial consequences requiring immediate operator intervention within minutes.
Warning: In un-rationalized plants, priority distribution is often inverted (e.g., 70% High, 20% Medium, 10% Low). When everything is flagged as High Priority, nothing is prioritized, leading directly to operator confusion during crises.
EEMUA 191 Performance Benchmark Metrics
EEMUA 191 provides precise numerical performance metrics to evaluate whether a plant's alarm system is healthy or dysfunctional.
| Alarm System Performance Metric | EEMUA 191 Target Benchmark | Manageable / Caution Level | Unacceptable / Dysfunctional Level |
|---|---|---|---|
| Average Alarm Rate (Normal Operation) | < 1 alarm per 10 minutes | 1 to 2 alarms per 10 minutes | > 2 alarms per 10 minutes (> 12/hr) |
| Peak Alarm Rate (10-minute window) | < 10 alarms per 10 minutes | 10 to 20 alarms per 10 minutes | > 10 alarms per 10 minutes |
| % of 10-min Periods with > 10 Alarms | < 1% of operational time | 1% to 5% of time | > 5% of operational time |
| Quantity of Standing Alarms at Shift Handover | < 5 to 10 standing alarms | 10 to 20 standing alarms | > 20 standing alarms |
| Top 10 "Bad Actor" Alarm Contribution | < 5% of total alarm count | 5% to 20% of total count | > 50% of total alarm count |
| Nuisance Alarms (Chattering / Fleeting) | Zero present | < 5 present | Continuous presence |
Alarm Floods and Nuisance Alarm Mechanisms
1. Alarm Floods
An alarm flood occurs when alarms enter the control room console at a rate significantly faster than an operator can read, comprehend, and respond to them. EEMUA 191 defines an alarm flood as any period where the alarm rate exceeds 10 alarms per 10 minutes.
During a major process trip (e.g., loss of main power or trip of a primary column feed pump), an un-rationalized DCS can generate 100 to 500 alarms within the first minute. Under such conditions:
- Critical safety alarms become buried in a sea of secondary process alarms.
- Operators suffer cognitive overload and visual paralysis.
- The risk of incorrect manual actions or complete loss of situational awareness skyrockets.
Technical Solutions for Alarm Floods:
- State-Based Alarm Suppression: Automatically suppressing lower-level alarms when equipment changes operational state (e.g., suppressing low-pressure and low-flow alarms on a pump when the pump stop signal is confirmed).
- Alarm Logic / First-Out Annunciation: Grouping cascading alarms so that only the root-cause initial trip ("First Out") is annunciated as High Priority, while downstream consequences are suppressed or downgraded.
Chattering Alarm Mechanism
Threshold ---------------------------------- (Alarm Line)
/\ /\ /\ /\ / / \ / \ / \ / \ / \ (Rapid Toggling)
/ \/ \/ \/ \/
With Deadband (Hysteresis)
Reset Line ---------------------------------
Alarm Line ---------------------------------
/-----------------------------\ (Stable Alarm)
/ ```
### 2. Nuisance Alarm Types and Engineering Remedies
Nuisance alarms account for up to 80% of all alarm traffic in poorly managed plants, directly creating **alarm fatigue** (where operators condition themselves to ignore or instantly acknowledge alarms without reading them).
- **Chattering Alarms:** Alarms that repeatedly transition into and out of the alarm state within seconds due to process noise near the threshold limit.
- *Engineering Remedy:* Implement a **deadband (hysteresis)**—a buffer zone (e.g., 2% of span) requiring the process variable to drop significantly below the alarm setpoint before resetting.
- **Fleeting Alarms:** Alarms that activate briefly but return to normal before an operator can respond.
- *Engineering Remedy:* Implement an **on-delay timer** requiring the signal to remain beyond the threshold for a sustained period (e.g., 5 seconds) before firing the alarm.
- **Standing Alarms:** Alarms that remain continuously in an active alarm state for hours, days, or months.
- *Engineering Remedy:* Perform maintenance on faulty instruments, re-rationalize setpoints, or formally shelve the alarm if maintenance is delayed.
- **Bad Actor Alarms:** A small subset of individual instruments (typically 10 or fewer) that generate the vast majority of total alarm volume due to mechanical wear or poor tuning.
- *Engineering Remedy:* Weekly monitoring of alarm logs to identify top 10 bad actors, followed by targeted instrument repair, re-calibration, or controller tuning.
---
## Alarm Shelving and Bypassing Governance
**Alarm shelving** refers to the manual temporary suppression of an alarm by a control room operator to prevent a known nuisance condition from distracting operations during maintenance or equipment outages.
While shelving is a necessary operational tool, unmanaged shelving destroys process safety protection by turning off critical barriers.
+---------------------------------------------------------------+
| Mandatory Alarm Shelving Controls |
+---------------------------------------------------------------+
| 1. Time-Limited Duration (Auto-un-shelve timer e.g. 8 hrs) |
| 2. Reason & Authorization Logging (Formal Shift Record) |
| 3. Continuous Visual Annunciator Display (Shelved List Tab) |
| 4. Dedicated Review at Every Shift Handover |
| 5. Prohibited for Safety Critical Alarms without MOC / Risk Ex |
+---------------------------------------------------------------+
### EEMUA 191 Best Practices for Alarm Shelving:
1. **Automated Timeout (Timer-Based Re-enabling):** Shelved alarms must automatically un-shelve after a pre-set duration (e.g., 4 or 8 hours) unless explicitly re-authorized by a shift manager.
2. **Explicit Operator Authorization & Reason Entry:** The DCS must force the operator to select a predefined reason (e.g., "Instrument Maintenance under PTW #1042") before shelving is accepted.
3. **Dedicated Console Visibility:** Shelved alarms must be displayed on a permanent, dedicated "Shelved Alarm Summary" tab on the operator console, ensuring incoming shift operators immediately see all suppressed alarms.
4. **Prohibition on Safety Critical Alarms:** Alarms designated as primary inputs to Safety Instrumented Functions (SIFs) or executive safety interlocks must never be shelved manually without a formal Management of Change (MOC) authorization and alternative compensatory safeguards.
According to EEMUA 191 guidelines, what is the recommended target priority distribution for alarm systems?
How does EEMUA 191 quantitatively define an 'alarm flood' condition in a control room?
Which engineering control is most effective at eliminating 'chattering alarms' caused by process signal noise near an alarm threshold?