2.5 High Reliability Organizations and Resilience

Key Takeaways

  • HROs operate in complex, hazardous environments with error rates significantly lower than expected, typically achieving fewer than 1 catastrophic event per 100,000 or 1,000,000 operations.
  • The five core principles of High Reliability Organizations (HROs) developed by Weick and Sutcliffe include: Preoccupation with Failure, Reluctance to Simplify, Sensitivity to Operations, Commitment to Resilience, and Deference to Expertise.
  • Resilience engineering shifts patient safety from Safety-I (focusing on why things go wrong, which occurs in less than 1% to 10% of clinical encounters) to Safety-II (focusing on why things go right, representing over 90% to 99% of daily work).
Last updated: July 2026

Introduction to High Reliability Organizations (HROs)

High Reliability Organizations (HROs) are organizations that operate in complex, high-hazard domains for extended periods without serious accidents or catastrophic failures. Originally conceptualized through studying nuclear aircraft carriers, air traffic control systems, and nuclear power plants, HRO principles have been increasingly adopted in healthcare. While healthcare remains highly dynamic and prone to clinical variations, the HRO framework provides structured mechanisms to manage unexpected failures.

Traditional industries achieve high reliability through strict standardization, but HROs recognize that in complex systems, safety is not merely about enforcing compliance; it requires the capacity to adapt and manage the unexpected. Healthcare is highly variable because patients present with unique co-morbidities, and clinical protocols must constantly be tailored. Therefore, adopting an HRO mindset requires moving beyond blind rule-following to active, mindful organizing.

The Five Core Principles of HROs

Karl Weick and Kathleen Sutcliffe identified five core principles that define HROs. These are divided into two categories: anticipation (principles 1-3) and containment (principles 4-5).

1. Preoccupation with Failure

HROs do not view success as a permanent state. Instead, they actively search for potential vulnerabilities and treat any deviation or near miss as a symptom of a larger system flaw. Clinicians are encouraged to report near misses, which are analyzed with the same rigor as actual adverse events. A preoccupation with failure means that when something goes right, the organization does not become complacent; it asks why it went right and what potential failures were narrowly avoided. Every system anomaly is seen as a warning sign that requires investigation.

2. Reluctance to Simplify

In complex environments, simple explanations often obscure the systemic contributors to an error. HROs reject superficial root causes, such as blaming a single clinician's distraction, and instead investigate the underlying latencies, such as poor software interface design, staffing ratios, or conflicting policies. By avoiding simple explanations, HROs maintain a rich, nuanced understanding of their operational vulnerabilities. They recognize that clinical work is multi-dimensional and that errors are almost always the result of multiple aligned vulnerabilities rather than a single active failure.

3. Sensitivity to Operations

This principle stresses that true safety is determined by the actual work performed at the front line (work-as-done), not the idealized procedures documented in policy manuals (work-as-imagined). Leaders in HROs actively gather feedback from bedside nurses, technicians, and resident physicians to understand how systems function under real-world pressures. This requires direct observations, safety walk-rounds, and rapid feedback loops. A system is only as safe as its current operational state, and leaders must remain constantly aware of how frontline clinicians adapt when resources are constrained.

4. Commitment to Resilience

HROs recognize that despite all safety measures, unexpected events will occur. Resilience refers to the capability to anticipate trouble, absorb the impact of a disturbance, rapidly adapt, and recover. It requires maintaining margins of safety and building redundant pathways to prevent localized failures from cascading into catastrophic events. Resilient organizations do not panic when a crisis occurs; instead, they mobilize their resources, adjust their operations, and work collaboratively to restore safety.

5. Deference to Expertise

During normal operations, traditional hierarchy is respected. However, when an emergency or high-stress situation arises, decision-making authority migrates to the person with the most specific, relevant knowledge of the problem, regardless of their rank. For example, a senior surgeon deferring to an entry-level scrub nurse who points out a potential sterility breach is a direct application of this principle. Deference to expertise ensures that authority is flexible and aligned with situational knowledge rather than formal administrative status.

Resilience Engineering in Healthcare

Resilience engineering is a proactive safety paradigm that focuses on designing systems that can adapt and function safely in the face of complexity, uncertainty, and change. Rather than focusing solely on reducing errors, resilience engineering aims to enhance the system's capacity to adjust its operations before, during, or after changes and disturbances. This discipline defines four key capabilities of a resilient system:

  • Anticipating: Understanding what to expect, identifying potential future disturbances, and preparing for them (e.g., planning for surge capacity during a pandemic or seasonal flu).
  • Monitoring: Keeping track of what is happening in the current environment, including monitoring internal systems and external variables that could impact safety (e.g., tracking unit staffing ratios and patient acuity in real time).
  • Responding: Having the agility to react to disturbances by deploying adaptive strategies and adjusting clinical processes (e.g., implementing a rapid triage protocol when the emergency department experiences an influx of patients).
  • Learning: Extracting meaningful lessons from both successful adaptations and failures to continuously redesign the system (e.g., debriefing after a difficult clinical case to identify what enabled success or caused difficulty).

Safety-I vs. Safety-II Paradigm Shift

Traditionally, healthcare safety has operated under the Safety-I framework, which defines safety as the absence of negative events. Safety-I focuses on understanding why things go wrong and seeks to eliminate errors by enforcing compliance, restricting autonomy, and investigating accidents. However, Erik Hollnagel and colleagues introduced Safety-II, which defines safety as the ability to ensure that as many things as possible go right.

Safety-II recognizes that clinical work is inherently variable and that healthcare professionals must constantly adjust their performance to match the demands of the situation. Safety is not merely the absence of errors, but the presence of active adaptations that prevent harm. The table below highlights the key differences between these two paradigms:

DimensionSafety-ISafety-II
Definition of safetyAs few things as possible go wrong (absence of incidents).As many things as possible go right (presence of adaptations).
Safety focusAccidents, incidents, near-misses, and errors.Everyday clinical work, normal success, and adaptations.
Human roleA liability, hazard, or source of error to be controlled.A resource, adaptation engine, and source of system flexibility.
Management approachReactive (respond to failures) and constraint-based.Proactive (understand normal work) and enablement-based.
View of variabilityHarmful; something to be minimized through compliance.Necessary; allows the system to cope with unexpected disturbances.
Primary safety goalEliminate errors and enforce adherence to guidelines.Enhance the system's capacity to adjust and adapt dynamically.

Clinical Application and Exam Focus

For the CPPS exam, candidates must understand how HRO and resilience principles translate into clinical practice. For instance, rapid response teams (RRTs) allow any bedside clinician to bypass the traditional medical hierarchy and summon specialized critical care resources if they observe early signs of clinical deterioration. This embodies both deference to expertise and sensitivity to operations. Daily safety huddles, where multidisciplinary teams meet for 15 minutes to discuss potential safety concerns for the upcoming shift, demonstrate a preoccupation with failure.

Exam questions often test these concepts by asking candidates to identify which principle is demonstrated in a clinical vignette. If a nurse stops a procedure because of a gut feeling that something is wrong, and the team listens, that shows deference to expertise and preoccupation with failure. If an investigation looks beyond the individual nurse who gave the wrong dose to examine the pharmacist's workload and the design of the drug labels, that demonstrates a reluctance to simplify.

Loading diagram...
HRO Principles Organization
Test Your Knowledge

Under Weick and Sutcliffe's framework for High Reliability Organizations (HROs), which principle is specifically defined as treating any deviation or near miss as a symptom of a potentially larger, systemic vulnerability rather than an isolated incident?

A
B
C
D
Test Your Knowledge

Which of the following scenarios best demonstrates Weick and Sutcliffe's principle of 'Deference to Expertise' in a clinical setting?

A
B
C
D
Test Your Knowledge

How does the Safety-II paradigm, as defined in resilience engineering, differ fundamentally from the traditional Safety-I approach?

A
B
C
D