2.3 Environmental, Health, and Safety (EHS) Risk Management

Key Takeaways

  • Reliability and EHS risk are connected because degraded equipment, emergency work, and transient operating states can increase exposure; the guide does not assign an unsupported universal percentage of incidents to those conditions.
  • Risk is formally defined as the product of Likelihood (Probability of Occurrence) and Severity (Consequence of Failure), establishing the mathematical basis for asset criticality and work prioritization.
  • A standardized 5x5 Risk Matrix establishes defensible thresholds across safety, environmental, and financial impacts, operationalizing the principle of As Low As Reasonably Practicable (ALARP).
  • OSHA 1910.119 Process Safety Management (PSM) mandates rigorous Mechanical Integrity programs, requiring documented inspections, testing, and compliance with Recognized And Generally Accepted Good Engineering Practices (RAGAGEP).
  • Mature safety cultures prioritize leading indicators—such as hazard identification work orders, proactive near-miss logs, and 100% safety-critical PM compliance—over historical lagging metrics like TRIR.
Last updated: September 2026

Environmental, Health, and Safety (EHS) Risk Management

Quick Answer: Reliability and EHS are connected disciplines. Risk control requires identifying hazards, evaluating likelihood and consequence under a defined method, complying with applicable requirements, maintaining process-safety-critical equipment, and tracking both preventive controls and outcomes.

In early industrial management, safety and maintenance were often treated as competing priorities. Plant leaders assumed that rigorous safety protocols inevitably slowed down maintenance execution, while maintenance was viewed simply as a repair function called upon after a mechanical failure had already introduced a hazard. Modern physical asset management has completely overturned this paradigm.

Under SMRP BoK Pillar 1, Function 1.6, certified maintenance and reliability professionals must integrate safety, health, and environmental risk directly into the asset management decision-making framework. A reliable plant is demonstrably a safe plant, a cost-effective plant, and an environmentally compliant plant. Disciplined reliability engineering is one of the most powerful risk mitigation strategies an industrial enterprise can deploy.


The Inextricable Link Between Reliability and Safety

Why Reactive Maintenance Breeds Catastrophic Risk

Equipment failures, emergency interventions, startups, shutdowns, and other non-routine conditions can expose workers and communities to hazards that are not present during stable operation. The magnitude varies by industry and event population, so do not attach an unsourced universal incident percentage. Use the facility's incident, near-miss, failure, and exposure data to identify the relevant risk.

When a machine fails unexpectedly, the operating environment instantly becomes volatile and hazardous:

  1. Compression of Time and High Stress: Technicians and supervisors face intense production pressure to restart equipment. Under this psychological stress, personnel take shortcuts: bypassing energy isolation steps, skipping formal Job Safety Analyses (JSAs), using makeshift rigging rather than engineered hoists, or entering equipment without verifying atmospheric testing.
  2. Loss of Containment and Process Instability: Mechanical breakdowns in fluid or gas processing systems often involve compromised primary containment—ruptured gaskets, blown pump mechanical seals, or split heat exchanger tubes. Technicians are forced to intervene in an environment containing toxic, flammable, corrosive, or high-pressure media.
  3. Unplanned Physical Exertion: Emergency repairs frequently demand awkward ergonomics, heavy manual lifting, and improvised access in unlit or cramped areas, leading directly to musculoskeletal strains, slips, trips, and falls.

The Safety Dividend of Planned and Scheduled Work

Conversely, when maintenance is planned and scheduled in advance:

  • Every hazard is systematically identified during the job planning phase.
  • Standard Operating Procedures (SOPs), specialized Personal Protective Equipment (PPE), engineered lifting plans, and detailed Lockout/Tagout (LOTO) procedures are incorporated into the printed job package.
  • Work is executed in a controlled, de-energized, cleaned, and decontaminated environment by technicians who are rested, properly equipped, and unburdened by artificial time pressure.

Risk Assessment Methodologies: Qualitative vs. Quantitative

To manage risk, reliability leaders must first quantify it. In asset management, Risk is formally defined as the product of two distinct variables:

Risk=Probability (Likelihood of Failure)×Severity (Consequence of Failure)\text{Risk} = \text{Probability (Likelihood of Failure)} \times \text{Severity (Consequence of Failure)}

Qualitative Risk Assessment

  • Definition: Evaluates risk using descriptive categorical ranking scales (e.g., Low, Moderate, High, Critical) based on the collective experience, engineering judgment, and consensus of a multidisciplinary team.
  • Application: Ideal for daily operational screening, job safety hazard evaluations, rapid work order prioritization, and initial plant-wide asset criticality ranking.
  • Strengths and Limitations: Highly cost-effective and rapid to deploy; however, it is inherently subjective and prone to individual cognitive bias or team consensus distortion.

Quantitative Risk Assessment (QRA)

  • Definition: Employs empirical statistical failure rate data, probabilistic mathematical modeling, and rigorous engineering simulations to calculate a discrete numerical probability of failure (e.g., $1.2 \times 10^{-4}$ events per operating year) and monetized or quantified consequence impacts.
  • Methodologies: Fault Tree Analysis (FTA), Event Tree Analysis (ETA), Layer of Protection Analysis (LOPA), and quantitative Failure Modes, Effects, and Criticality Analysis (FMECA).
  • Application: Mandatory for high-consequence industrial processes, nuclear power installations, offshore hydrocarbon extraction, and when designing Safety Instrumented Systems (SIS) to achieve specific Safety Integrity Levels (SIL) under IEC 61508 / 61511.

Designing and Applying a 5x5 Risk Matrix

A 5x5 Risk Matrix is the operational bridge that translates corporate risk tolerance into daily maintenance prioritization. It maps five discrete levels of failure likelihood against five discrete levels of failure consequence, producing twenty-five unique risk intersection cells.

   ▲  5 |  [M-5]   [H-10]  [C-15]  [C-20]  [C-25]    RISK BANDS:
 S │  4 |  [L-4]   [M-8]   [H-12]  [C-16]  [C-20]    [C] Critical (15-25)
 E │  3 |  [L-3]   [M-6]   [M-9]   [H-12]  [C-15]    [H] High (10-12)
 V │  2 |  [L-2]   [L-4]   [M-6]   [M-8]   [H-10]    [M] Medium (5-9)
   ▼  1 |  [L-1]   [L-2]   [L-3]   [L-4]   [M-5]     [L] Low (1-4)
        └─────────────────────────────────────────►
            1       2       3       4       5
         ◄──────────   LIKELIHOOD   ──────────►

Defining Probability and Consequence Scales

To prevent subjective arguing among team members, the matrix must be grounded in objective, documented definitions across each tier:

Likelihood (Probability) Tiers:

  1. Improbable / Rare: Highly unlikely to occur in the asset life cycle (< once per 25 years).
  2. Remote / Unlikely: Not expected to occur under normal operating conditions (once per 10 to 25 years).
  3. Credible / Possible: Might occur several times during the facility design life (once per 3 to 10 years).
  4. Probable / Frequent: Expected to occur periodically during normal operations (once per 1 to 3 years).
  5. Common / Continuous: Occurs repeatedly throughout an operating year (> once per year).

Severity (Consequence) Tiers:

  1. Negligible: Minor first-aid injury; zero environmental release; financial loss < $5,000.
  2. Moderate: Medical treatment injury (no lost workdays); contained on-site spill; financial loss $5,000–$50,000.
  3. Serious: Lost-time injury (reversible); reportable environmental spill with no off-site migration; financial loss $50,000–$250,000.
  4. Critical: Severe permanent disability; major environmental release with local off-site impact; financial loss $250,000–$1,000,000.
  5. Catastrophic: Single or multiple worker fatalities; massive uncontained toxic or flammable release with widespread community impact; financial loss > $1,000,000.

5x5 Risk Matrix Evaluation Table

Severity LevelLikelihood 1 (Rare)Likelihood 2 (Unlikely)Likelihood 3 (Possible)Likelihood 4 (Frequent)Likelihood 5 (Continuous)
Level 5: CatastrophicMedium (5)High (10)Critical (15)Critical (20)Critical (25)
Level 4: CriticalLow (4)Medium (8)High (12)Critical (16)Critical (20)
Level 3: SeriousLow (3)Medium (6)Medium (9)High (12)Critical (15)
Level 2: ModerateLow (2)Low (4)Medium (6)Medium (8)High (10)
Level 1: NegligibleLow (1)Low (2)Low (3)Low (4)Medium (5)

Risk Level Action Thresholds and the ALARP Principle

The calculated score ($1 \text{ to } 25$) dictates mandatory operational and maintenance actions governed by the principle of As Low As Reasonably Practicable (ALARP):

  • Low Risk (Score 1–4): Acceptable Zone. Risk is broadly acceptable. Managed through standard operating procedures, routine lubrication, and scheduled visual inspections. No special risk mitigation plans required.
  • Medium Risk (Score 5–9): ALARP / Tolerable Zone. Risk is tolerable only if further reduction is impracticable. Maintenance leadership must evaluate risk reduction measures (e.g., condition monitoring, improved PM tasks) and implement them unless the economic and operational cost of further reduction is grossly disproportionate to the risk benefit achieved.
  • High Risk (Score 10–14): Undesirable Zone. Requires formal, active risk mitigation. The asset cannot continue operating indefinitely without engineered safeguards, enhanced predictive technologies, or administrative controls. Requires written authorization from the Plant Manager to operate pending mitigation.
  • Critical Risk (Score 15–25): Unacceptable Zone. Operations must not commence or continue. Demands immediate equipment shutdown, emergency isolation, or immediate engineering redesign to eliminate the failure mode before normal production resumes.

Process Safety Management (PSM) & Mechanical Integrity (OSHA 1910.119)

In facilities handling hazardous, toxic, reactive, or highly flammable chemicals exceeding regulatory threshold quantities, asset management is legally governed by OSHA 29 CFR 1910.119: Process Safety Management of Highly Hazardous Chemicals. Enacted in 1992 following catastrophic chemical disasters, PSM contains fourteen interrelated elements designed to prevent catastrophic releases of toxic, reactive, or flammable liquids and gases.

Section (j) Mechanical Integrity: The Core Maintenance Mandate

While all fourteen elements impact the facility, Element (j): Mechanical Integrity (MI) is the exclusive domain and direct responsibility of the maintenance and reliability department. OSHA mandates that employers establish and implement written procedures to maintain the ongoing integrity of six critical process equipment categories:

  1. Pressure vessels and storage tanks
  2. Piping systems (including valves and piping components)
  3. Relief and vent systems and devices (e.g., pressure safety valves, rupture discs, flare stacks)
  4. Emergency shutdown systems (e.g., automated isolation valves, interlocks)
  5. Controls (including monitoring devices, sensors, alarms, and instrumentation)
  6. Pumps and rotating machinery

Recognized And Generally Accepted Good Engineering Practices (RAGAGEP)

Under OSHA PSM, maintenance intervals and inspection criteria cannot be arbitrarily established based on plant convenience or budget availability. Inspections and tests must strictly conform to Recognized And Generally Accepted Good Engineering Practices (RAGAGEP). RAGAGEP standards are established by consensus engineering organizations, including:

  • American Petroleum Institute (API): API 510 (Pressure Vessel Inspection Code), API 570 (Piping Inspection Code), API 653 (Tank Inspection, Repair, Alteration, and Reconstruction).
  • American Society of Mechanical Engineers (ASME): ASME Boiler and Pressure Vessel Code (Sections I, VIII), ASME B31.3 (Process Piping).
  • National Fire Protection Association (NFPA): Codes governing fire suppression pumps, flammable liquid storage, and explosion venting.

Inspection Documentation & Deficiencies

OSHA mandates that every inspection and test performed on PSM-covered equipment must be formally documented in the CMMS. The record must document:

  • Date of the inspection or test
  • Name of the person who performed the inspection or test
  • Serial number or unique equipment identifier
  • Specific description of the inspection or test performed
  • Quantitative results of the test
  • Documented verification of acceptable operating limits

Resolution of Equipment Deficiencies: If an inspection reveals that equipment wall thickness, relief valve set pressure, or vibration severity falls outside acceptable limits, the employer must correct the deficiency before further use, or implement documented interim controls verified by an engineering risk assessment demonstrating safe operation.

PSM MI MandateRegulatory Requirement (29 CFR 1910.119(j))M&R Operational ImplementationAudit Evidence Required
Written ProceduresEstablish and implement written maintenance procedures for covered equipment.Standard job plans, precision rebuild protocols, and OEM maintenance manuals stored in CMMS.Controlled document repository with revision histories and annual reviews.
Technician TrainingEnsure each employee maintaining process equipment is trained in hazards and procedures.Formal craft certification, precision maintenance training, and documented qualification matrices.Signed training records, competency test results, and annual refresher logs.
Inspection & TestingConduct inspections and tests following RAGAGEP codes and standards.Baseline thickness gauging, relief valve test bench recertifications, vibration and oil analysis routes.Non-destructive examination (NDE) inspection reports, API 510/570 inspection dossiers.
Equipment DeficienciesCorrect deficiencies outside acceptable limits before further use.Priority work orders triggered immediately upon out-of-spec inspection findings; formal engineering reviews.Completed corrective work orders with post-repair NDE reports and engineering sign-offs.
Quality AssuranceEnsure newly fabricated equipment and spare parts meet design specifications.Rigorous receipt inspection in MRO storeroom; verification of Material Test Reports (MTRs) for metallurgy.Positive Material Identification (PMI) logs and vendor Mill Test Certificates.

Critical Hazard Control Interfaces

Reliability professionals must master the interfaces between equipment maintenance execution and plant life-critical safety standards:

Control of Hazardous Energy: Lockout/Tagout (29 CFR 1910.147)

  • Zero Energy State: Maintenance cannot occur until all electrical, mechanical, hydraulic, pneumatic, chemical, thermal, and gravitational energies are isolated, locked, tagged, and verified by physical testing (try-step).
  • Complex Group Lockout: In major overhaul turnarounds involving dozens of craftspeople across multiple shifts, a Master Tag and Lockbox system must be administered by a designated Primary Authorized Employee. Each individual craftsperson must attach their personal lock to the group lockbox before performing work.

Permit-Required Confined Space Entry (29 CFR 1910.146)

  • Vessels, tanks, bins, and boilers represent acute life hazards (atmospheric toxicity, oxygen deficiency, engulfment, or internal mechanical hazards).
  • Reliable work execution demands strict permit protocols: atmospheric testing before entry (oxygen between 19.5% and 23.5%, flammability < 10% LEL, toxic contaminants like H2S < 10 ppm and CO < 25 ppm), continuous mechanical ventilation, a dedicated standby attendant stationed outside the portal, and non-entry emergency retrieval systems.

Management of Change (MOC - 29 CFR 1910.119(l))

  • Industrial history confirms that informal, unreviewed maintenance modifications are a leading cause of catastrophic disasters (e.g., the Flixborough explosion, caused by an unengineered temporary bypass pipe installed between chemical reactors).
  • Maintenance MOC Triggers: Any replacement that is not a true "Replacement in Kind" (RIK) requires a formal MOC. This includes altering piping metallurgy, installing a different mechanical seal type, modifying pump impeller diameters, changing lubrication formulations, adjusting relief valve set points, or modifying PLC control logic. Work must not proceed until an engineering safety review, hazard analysis, and P&ID documentation update are formally completed.

Safety Culture Maturity and Leading vs. Lagging Indicators

Hudson's Safety Culture Maturity Model

Industrial psychologist Patrick Hudson developed a widely adopted evolutionary model illustrating how an organization’s mindset regarding safety and reliability matures across five developmental stages:

[1. Pathological] ──> [2. Reactive] ──> [3. Calculative] ──> [4. Proactive] ──> [5. Generative]
  "Who cares?"         "Safety after       "We have rules        "We anticipate       "Safety is how
                       accidents"          & systems"            hazards"             we do business"
  1. Pathological: Safety is viewed as an annoying bureaucratic impediment. Leadership cares only about not getting caught by regulatory agencies.
  2. Reactive: Safety is taken seriously only after a serious accident occurs. The organization responds with punitive crackdowns and temporary enforcement, but lapses back into complacency.
  3. Calculative: The organization establishes formal safety manuals, procedures, audits, and checklists. Hazards are managed by strict adherence to written rules, but frontline ownership remains low.
  4. Proactive: Leadership and the workforce actively collaborate to anticipate hazards and equipment failures before they materialize. Frontline employees are empowered to stop work without fear of reprisal.
  5. Generative: Safety and reliability are fully synthesized into the fundamental corporate identity. The organization maintains chronic unease regarding operational risks, relentlessly seeks out subtle system weaknesses, and treats failure prevention as an ethical imperative.

The Danger of Exclusively Relying on Lagging Safety Metrics

Historically, industrial safety performance was judged entirely on Lagging Indicators:

  • Total Recordable Incident Rate (TRIR): $\frac{\text{Number of Recordable Injuries} \times 200,000}{\text{Total Hours Worked}}$
  • Days Away, Restricted, or Transferred (DART) Rate
  • Lost Time Incident Frequency (LTIF)

While lagging metrics are legally required, relying on them to manage plant safety is equivalent to driving an automobile forward while looking exclusively into the rearview mirror. A plant can achieve zero recordable injuries for an entire year simply through good fortune, even while critical pressure vessels are corroding, relief valves are stuck shut, and maintenance technicians are taking dangerous shortcuts on de-energization. When the catastrophic event occurs, lagging indicators provide zero warning.

Leading Indicators That Drive Proactive Safety and Reliability

Organizations can monitor leading indicators—measures of preventive activities, capacity, and control health—together with lagging outcomes:

  1. Preventive Maintenance Compliance on Safety-Critical Equipment (SCE): Mandatory 100% execution compliance within the 10% grace period for all testing on relief valves, toxic gas detectors, burner management interlocks, and emergency trip systems.
  2. Near-Miss and Hazard Identification Submission Rate: The volume of proactive safety observations and equipment hazard work orders submitted by technicians and operators.
  3. Percentage of Maintenance Work Planned and Scheduled: Track the site-defined planned-work share together with job-plan quality, emergency exposure, and safety outcomes; planning can reduce avoidable improvisation but a percentage alone does not prove risk control.
  4. Audit and Inspection Closure Rate: The speed and rigor with which documented safety and mechanical integrity deficiencies are permanently engineered out of the facility.
  5. Management of Change (MOC) Compliance: Percentage of equipment modifications executed with fully closed-out safety documentation and updated standard operating procedures.
Test Your Knowledge

Under OSHA 29 CFR 1910.119(j) Process Safety Management (PSM), what standard must an industrial facility adhere to when establishing inspection and testing frequencies, procedures, and acceptable operating tolerances for covered equipment such as pressure vessels, piping, and relief devices?

A
B
C
D
Test Your Knowledge

A facility applies the ALARP principle to a medium-risk scenario. What does ALARP require when assessing further controls?

A
B
C
D
Test Your Knowledge

Which of the following operational metrics serves as the most effective leading indicator of an organization's proactive maintenance safety performance, rather than a lagging indicator of past harm?

A
B
C
D