8.2 Alarm Management

Key Takeaways

  • An alarm is an abnormal condition that requires a timely operator action. Run status, “batch complete,” and diagnostics without an action are not alarms.
  • ISA-18.2 and IEC 62682 are professional practice for 2027, not supplied exam PDFs. Only ISA-5.1 (2024) and IEC 61511-1 (2018) are supplied design standards.
  • Useful KPIs include average rate, peak/flood rate, standing count, chattering/fleeting, and stale alarms. Average rate can look fine while a flood is still occurring.
  • Shelving is operator-initiated and time-limited; suppression is designed state-based logic; out-of-service is authorized maintenance/MOC disable with return-to-service tracking.
  • After a compressor trip, first-out should capture the initiating cause. Consequence alarms (low flow after the machine is already stopped) should be suppressed, grouped, or lower priority.
Last updated: August 2026

Alarm is an action-required deviation, not a status bit

ISA-18.2 and IEC 62682 define an alarm as an indication of equipment malfunction, process deviation, or abnormal condition that requires a timely operator response. If no response is defined, it is not an alarm. Put it somewhere else: a status indication, an event log, a maintenance work order, or a diagnostic page.

Not alarms (exam traps):

  • Pump RUN / STOP bits (status).
  • “Batch complete,” “recipe downloaded,” “historian archive 80% full” unless a specific timed operator action exists.
  • Duplicate consequence bits that fire because the first alarm already stopped the machine.

Are alarms: high seal-oil differential pressure with a written action to start the standby pump within a stated time; high-high drum level with an action to cut inlet flow; first-out vibration trip that tells the operator why the compressor stopped.

For 2027, treat ISA-18.2 and IEC 62682 as the professional language of this topic. They are not on the NCEES supplied-standards list. The supplied PDFs remain ISA-5.1 (2024) and IEC 61511-1 (2018). You will not search ISA-18.2 clause numbers in the exam driver. You are still expected to know what a competent control-systems PE does with alarms, including how alarm response relates to a protection layer versus a true SIS function under IEC 61511-1.

KPIs: what the number is actually diagnosing

Published benchmark tables in ISA-18.2 (and the closely aligned IEC 62682 / EEMUA 191 practice) are order-of-magnitude operator-load targets, not laws of physics. Know what each metric means so you can pick the right fix.

Common professional targets (per operator console, three-priority systems):

  • Average annunciated rate — on the order of ~6/h “very likely acceptable,” ~12/h “maximum manageable” (about ~1 vs ~2 per 10 minutes).
  • Peak / flood — a 10-minute window with more than ~10 new alarms is a flood starting point; target ≤10 in the worst 10 minutes and ≪1% of time in flood.
  • Standing — alarms currently in the active list; a large standing list at shift change means the board starts already behind.
  • Chattering / fleeting — rapid return to normal and back (a common test is three or more activations in one minute); target zero with an action plan for any that exist.
  • Stale — remains in alarm a long time, often >24 h; target fewer than about five present on a given day, with an action plan.
  • Priority distribution of annunciated alarms — often cited near 80% low / 15% medium / 5% high. If everything is “emergency,” nothing is.
  • Top-10 contribution — a handful of tags should not be most of the load (practice targets often ~1–5% each as a prompt to hunt bad actors).

Average can lie. A console at 4 alarms/hour looks healthy until one 10-minute window dumps 47 new alarms. Always read average and peak together.

Priority and rationalization

Rationalization is the documented decision that an alarm deserves to exist: consequence of inaction, time available to act, operator action, setpoint, deadband, on/off delay, and priority. If you cannot name the action and the time, delete or demote it to a status.

Priority is a tie-breaker during peaks, not a filing cabinet for “safety vs quality vs environmental.” Assign it from severity × time to respond for the proximate consequence the operator can still affect. Inflating every alarm to “critical” because an unmitigated train wreck is imaginable destroys the priority system.

IEC 61511-1 still governs safety instrumented functions. An HMI alarm may be credited in a LOPA as an independent protection layer only if it meets the rules for that layer (independence, auditability, testing). Do not “make it SIL” by painting the alarm red.

Shelving versus suppression versus out-of-service

These three words are not synonyms:

  • Shelvingoperator-initiated, temporary, usually time-limited, with auto-return. Use it for a known short condition the philosophy allows (a filter swap with a documented max shelf time). Unauthorized or forgotten shelves are a KPI and an MOC leak.
  • Suppressiondesigned logic: when the pump is confirmed stopped, suppress the low-flow and low-discharge-pressure alarms that would otherwise chatter. State-based (mode-based) alarming is suppression done on purpose. It must be specified, tested, and visible (“suppressed by design,” not missing).
  • Out-of-service (OOS)authorized disable for maintenance or MOC, tracked like a bypass, with a return-to-service check. Taking a transmitter out for calibration is OOS, not “the operator deleted the limit.”

Exam trap: calling a standing stale alarm “shelved” when nobody authorized anything. Another trap: suppressing a first-out initiating alarm so the flood looks pretty and the cause is gone.

Worked example: flood after a compressor trip, and first-out

Given. Charge-gas centrifugal compressor C-301, 11,000 rpm class. Protective inputs include high shaft vibration, high discharge temperature, low lube-oil pressure, and low seal-oil differential. The anti-surge valve is on a fast loop. Downstream, a hydrotreater feed-flow alarm sits on the same operator console.

Event. Probe proximity vibration exceeds the trip. Within 20 seconds the alarm summary fills: high vibration, motor stop, low discharge flow, low discharge pressure, anti-surge deviation, downstream low feed flow, low lube-oil pressure (header sags after the shaft coasts), and a dozen analog “bad PV” alarms from transmitters that went to zero with the process.

What should have been first-out. High vibration is the initiating cause. Latch it as first-out in the compressor trip group so it remains identifiable after the list explodes. Consequence alarms that are expected once the machine is tripped (low discharge flow and pressure, surge-valve wide-open deviation, downstream low flow, motor-stop confirm) should be designed suppression or grouping when the trip is confirmed—not 15 independent high-priority horns. Lube-oil low pressure might be a true second initiating cause on another day; after a vibration trip it is often a follow-on. Rationalization decides whether it stays as a separate first-out group (lube system) or is suppressed when the compressor is confirmed stopped.

What the KPIs would show. Average rate for the shift might still look “OK.” Peak in that 10-minute window is a flood. If vibration chatters around the trip setpoint because deadband is zero, you also have a chatterer. If the downstream low-flow remains in alarm for 30 hours after the unit is shut down, it is stale—wrong state-based design, not a dedicated operator who “should just acknowledge harder.”

Human-factors link to 8.1. A P&ID-clone graphic that turns every consequence analog red at once is the visual form of the same flood. First-out plus suppression plus a calm overview analog of “compressor tripped — cause: vibration” is the HMI half of alarm management.

KPI versus what it diagnoses

KPITypical professional target (order of magnitude)What a bad number usually diagnosesFirst fix to consider
Average rate~6/h acceptable, ~12/h max manageable per consoleToo many configured alarms; status bits as alarms; bad actorsRationalize; remove non-action alarms
Peak / flood (10 min)≤10 new alarms; flood time ≪1%Cascading consequences; no first-out; no state-based suppressionFirst-out + designed suppression of expected follow-ons
Standing countSmall list at shift change (often cited ~<10 as a tight target)Uncleared conditions; idle-equipment alarms; ignored listState-based alarming; clear real problems
Chattering / fleetingZero (e.g. ≥3 times/min is a chatterer)No deadband/on-delay; noisy PV; limit in the noise bandDeadband, filter, delay; fix the measurement
Stale (>24 h class)Fewer than ~5 present in a day, with a planWrong operating-state design; latched hysteresis; using alarms as a work-order systemMode-based suppression; OOS/MOC; don’t alarm maintenance backlog
Priority mix~80/15/5 low/med/high annunciatedEverything defaulted to high; priority used as a category tagRe-rationalize on time-to-respond × proximate severity
Top-10 loadSmall share of total (often ~1–5% each as a prompt)Nuisance tags dominating the consoleFix the bad actors first
Test Your Knowledge

Which indication meets the professional definition of an alarm (ISA-18.2 / IEC 62682 practice) rather than a status or event bit?

A
B
C
D
Test Your Knowledge

A centrifugal compressor trips on high shaft vibration. Within 20 seconds the summary fills with low discharge flow, low discharge pressure, anti-surge deviation, downstream low feed flow, and motor-stop. Which design is the correct first-out / flood strategy?

A
B
C
D
Test Your Knowledge

A console averages 4 alarms per hour over a shift, but one 10-minute window produced 47 new alarms, and the same vibration tag activated 11 times in one minute. Which reading of the KPIs is correct?

A
B
C
D