5.2 Evaluating Equipment Reliability & Gap Analysis
Key Takeaways
- The historical Nowlan and Heap airline study found that 89% of the items in its study population followed age-independent patterns; do not automatically assign that percentage to every modern industrial asset population.
- Age-based replacement is useful when a failure mode shows a defensible age relationship; condition-based, failure-finding, redesign, or run-to-failure policies may be better for other modes and consequences.
- Weibull analysis utilizes the shape parameter (beta) to diagnose failure physics: beta < 1 indicates infant mortality (burn-in defects), beta = 1 indicates constant random failure (exponential distribution), and beta > 1 reveals wear-out degradation.
- Crow-AMSAA tracking models cumulative failures versus cumulative operating time on log-log scales to determine reliability growth: a slope (alpha < 1) signals reliability growth, alpha = 1 indicates steady state, and alpha > 1 alerts to system deterioration.
- Reliability gap analysis establishes Best-Demonstrated Performance (BDP) as an internal operational ceiling, systematically comparing current Mean Time Between Failures against external industry benchmarks to identify production losses and prioritize engineering interventions.
Evaluating Equipment Reliability & Gap Analysis
Quick Answer: Reliability evaluation combines trustworthy event and exposure data, best-demonstrated internal performance, failure-pattern analysis, Weibull or other life-data methods where the data support them, and explicit comparison with required performance. The historic Nowlan and Heap findings are important evidence against assuming universal wear-out, not a universal percentage for all industrial equipment.
Measuring and Tracking Equipment Performance: Beyond Anecdotes
Traditional maintenance management often relies on anecdotal evidence, subjective technician opinions, or gross monthly repair costs to evaluate asset health. This qualitative approach fails to identify chronic failure modes, hides recurring micro-stoppages, and results in misallocated capital.
SMRP BoK Pillar 3 (Function 3.2) emphasizes disciplined, quantitative performance tracking. Reliability engineers monitor asset performance using standardized reliability metrics:
- Mean Time Between Failures (MTBF): $\text{MTBF} = \frac{\text{Operating Time}}{\text{Number of Unscheduled Failures}}$. Measures inherent operational runtime integrity.
- Failure Rate ($\lambda$): $\lambda = \frac{1}{\text{MTBF}}$. Represents the frequency with which an engineered system or component fails, expressed in failures per hour or per operating cycle.
- Chronic vs. Sporadic Failures: Sporadic failures are dramatic, sudden breakdowns (e.g., a catastrophic turbine shaft fracture) that trigger immediate root cause investigations. Chronic failures, by contrast, are recurring, smaller events (e.g., a pump packing leak requiring adjustment every three weeks). While sporadic failures capture management attention, chronic failures collectively account for 70% to 80% of total plant downtime losses and maintenance expense.
Establishing Best-Demonstrated Performance (BDP)
A powerful analytical tool in reliability engineering is establishing Best-Demonstrated Performance (BDP). BDP identifies the highest verified level of operational throughput, availability, and MTBF that an asset, production line, or facility has historically sustained over a meaningful operating window (e.g., a continuous 60-day or 90-day run without unpredicted stoppage).
- The Strategic Role of BDP: BDP proves what the existing asset configuration is physically capable of achieving under real-world plant conditions. If an automated packaging line demonstrated 96.5% availability and an MTBF of 420 hours over a consecutive 90-day period two years ago, but currently operates at 88.0% availability with an MTBF of 110 hours, the delta cannot be blamed on poor equipment design. The reliability gap must be attributed to operational changes, process drift, lubrication breakdown, operator turnover, or degraded maintenance execution discipline.
The Nowlan & Heap Failure Patterns: The Foundation of Modern Reliability
Until the late 20th century, industrial maintenance strategy was governed by a single universal assumption: as equipment ages and accumulates operating hours, its probability of failure inevitably increases. Consequently, plants scheduled periodic overhauls, tearing down healthy machinery on calendar cycles to "prevent wear-out."
In 1978, F. Stanley Nowlan and Howard F. Heap published their landmark study, Reliability-Centered Maintenance, commissioned by the U.S. Department of Defense and United Airlines. By analyzing decades of detailed commercial aircraft operating histories and component failures, they revolutionized physical asset management.
The Discovery: Six Universal Failure Patterns
Nowlan & Heap discovered that component failure behavior falls into six distinct curves, divided into two starkly different categories:
Category 1: Age-Related Failure Modes (Only 11% of Total)
- Pattern A (The Classic Bathtub Curve — 4%): High infant mortality, followed by a flat constant failure rate, terminating in a distinct, pronounced wear-out zone. Commonly found in simple mechanical components exposed to severe friction, erosion, or corrosion (e.g., pump impellers in acid service, furnace refractory linings).
- Pattern B (The Traditional Wear-Out Curve — 2%): Low initial failure probability, followed by a distinct, predictable wear-out threshold. Typical of components subject to direct sliding friction, tire tread wear, or brake pads.
- Pattern C (Gradual Aging Fatigue — 5%): Steadily rising failure probability with no distinct sharp wear-out knee. Typical of fatigue-prone structural steel, steam piping, or turbine blades subject to high-cycle thermal stress.
Category 2: Age-Independent Patterns in the Historical Study (89% of Its Items)
- Pattern D (Initial Low Failure with Sharp Rise — 7%): Low failure probability when brand new, followed by a rapid rise to a constant, age-independent random failure rate.
- Pattern E (Constant Random Failure — 14%): Completely flat, constant failure rate across the entire operating lifespan. Failure occurs randomly due to external process surges, electrical transients, foreign debris, or operator error. The age of the asset provides zero information about when it will fail.
- Pattern F (Infant Mortality Dominance — 68%): Very high initial failure probability (burn-in period), dropping sharply to a low, constant random failure rate for the remainder of its operational life. This single curve accounts for over two-thirds of all failure modes!
The Profound CMRP Takeaway
The study's finding that 89% of its item population followed patterns without a simple increasing-age wear-out relationship challenged the assumption that calendar overhaul is automatically effective. Tearing down complex rotating equipment, electrical panels, or hydraulic systems simply because 12 months have elapsed does not prevent failure. In fact, disassembling and rebuilding healthy machinery repeatedly re-exposes the equipment to Pattern F, introducing maintenance-induced defects (incorrect torque, misaligned couplings, seal damage, contamination, and unseated electrical connectors) that directly trigger infant mortality.
Nowlan & Heap 6 Failure Patterns Summary Table
| Pattern | Visual Profile | Population Share | Failure Rate Behavior $\lambda(t)$ | Dominant Physics / Causes | Optimal Maintenance Strategy |
|---|---|---|---|---|---|
| A (Bathtub) | High start, flat middle, steep rise | 4% | Decreasing $\to$ Constant $\to$ Rapidly Increasing | Burn-in assembly defects, followed by random events, ending in direct friction/corrosion wear-out. | Strict burn-in testing; run through flat zone; scheduled replacement before wear-out knee. |
| B (Wear-Out) | Flat baseline, distinct wear-out knee | 2% | Constant/Low $\to$ Sharp Increase | Sliding mechanical friction, direct abrasive wear, brake linings, elastomer seal hardening. | Time-directed scheduled replacement or overhaul based on operating hours or cycles. |
| C (Fatigue) | Steady upward slope throughout | 5% | Monotonically Increasing | Cyclic mechanical fatigue, thermal stress cycling, corrosion-fatigue, boiler tubing. | Predictive wall-thickness testing; nondestructive examination (NDE); time-directed replacement. |
| D (Early Rise) | Low initial, then constant plateau | 7% | Low $\to$ Sudden Rise $\to$ Constant Plateau | Complex mechanical/hydraulic assemblies that achieve initial run-in then experience constant random stress. | On-condition monitoring (vibration, oil analysis); do NOT schedule arbitrary teardown overhauls. |
| E (Random) | Completely horizontal straight line | 14% | Constant Failure Rate ($\lambda = \text{const}$) | Electrical transients, human error, external hydraulic shock, debris ingestion, lightning strikes. | Predictive condition-based maintenance (PdM); failsafe design; operator training; do NOT use time PM. |
| F (Infant Mortality) | Extreme high initial spike, dropping to low plateau | 68% | Sharp Initial Drop $\to$ Low Constant Plateau | Installation errors, improper alignment, incorrect torque, contamination, manufacturing defects. | Precision maintenance practices, commissioning standards, run-in testing, non-intrusive PdM. |
Weibull Analysis: Deciphering the Physics of Failure
While Nowlan & Heap provide the macro-level distribution of failure patterns, Weibull Analysis is the primary mathematical tool used by reliability engineers to diagnose the exact failure physics of a specific component population.
Developed by Swedish mathematician Waloddi Weibull, the two-parameter cumulative Weibull distribution is defined as:
Where:
- $t$ = Accumulated operating time, cycles, or mileage.
- $\beta$ (Beta - Shape Parameter): Dictates the slope of the Weibull probability plot and reveals the underlying failure physics.
- $\eta$ (Eta - Scale Parameter / Characteristic Life): The operational age at which 63.2% of the asset population will have failed, regardless of the value of $\beta$.
Interpreting the Shape Parameter ($\beta$)
The shape parameter $\beta$ is the most critical diagnostic number in life data analysis:
- $\beta < 1.0$ (Decreasing Failure Rate — Infant Mortality):
- Corresponds to Pattern F.
- Failure probability drops as operating time accumulates.
- Root Causes: Substandard parts, improper storage/preservation, poor lubrication during startup, pipe strain, misalignment, assembly errors, or manufacturing quality defects.
- Maintenance Action: Eliminate maintenance intrusion! Enforce precision maintenance standards (laser alignment, calibrated torque wrenches), conduct rigorous FAT/SAT commissioning tests, and reject time-based overhauls.
- $\beta = 1.0$ (Constant Failure Rate — Pure Random Failures):
- Corresponds to Pattern E (the Exponential Distribution).
- Failures occur randomly and independently of component age.
- Root Causes: Environmental shocks, operator errors, debris ingestion, sudden power surges, random external stresses.
- Maintenance Action: Time-directed overhauls are completely useless. Deploy condition monitoring (vibration, ultrasound, thermography) to detect early degradation signatures, or redesign the system for greater resilience.
- $\beta > 1.0$ (Increasing Failure Rate — Wear-Out Degradation):
- Corresponds to Patterns A, B, and C.
- Failure probability increases as the asset ages.
- Sub-classifications:
- $1.0 < \beta < 2.0$: Early fatigue or non-linear wear.
- $\beta \approx 3.5$: Approximates a symmetrical Gaussian Normal distribution (classic mechanical wear-out).
- $\beta > 4.0$: Rapid, steep wear-out (very narrow failure window with predictable end-of-life).
- Maintenance Action: Time-directed overhaul or scheduled component replacement is technically feasible and economically justified, provided the replacement interval is scheduled prior to the onset of the wear-out knee.
Weibull Beta Parameter Interpretation and Maintenance Tactics Table
| $\beta$ Value | Failure Rate Profile | Primary Failure Mechanisms | Prescribed Maintenance Tactic | Operational Example |
|---|---|---|---|---|
| $\beta = 0.5$ | Decreasing sharply | Assembly error, pipe strain, dirty lubricant at startup | Precision installation, commissioning audits, run-in tests | Newly installed mechanical seals failing within 72 hours of pump startup. |
| $\beta = 1.0$ | Constant (flat) | Lightning strikes, operator mishandling, debris ingestion | Condition-Based Monitoring (PdM), redesign, operator training | Electronic control boards, solenoid coils failing due to electrical line surges. |
| $\beta = 1.8$ | Moderately increasing | Mechanical fatigue, light cavitation, roller bearing spalling | On-condition monitoring (vibration analysis, oil debris analysis) | Overhung blower bearings operating under cyclic belt tension loads. |
| $\beta = 3.5$ | Increasing (Normal distribution) | General adhesive wear, friction degradation, seal lip wear | Predictive trending; planned component overhaul if PdM is not feasible | Conveyor pulley lag wear, slurry pump impellers, fleet vehicle brake pads. |
| $\beta = 6.0+$ | Increasing rapidly (Steep wear-out) | High-temperature oxidation, refractory spalling, UV decay | Hard-time scheduled replacement before the known failure threshold | Furnace oxygen sensor probes, aircraft turbine igniter plugs. |
Crow-AMSAA Reliability Growth Tracking
While Weibull analysis evaluates single failure modes on non-repairable parts, complex industrial systems (entire compressors, automated packaging cells, haul trucks) are repairable systems that experience multiple failures over time. To track whether overall system reliability is improving, deteriorating, or stagnating, reliability leaders deploy the Crow-AMSAA Model (developed by Dr. Larry Crow, based on Duane's postulate, under the Army Materiel Systems Analysis Activity).
The Mathematical Formulation
The cumulative number of failures $N(t)$ as a function of cumulative operating time $t$ is modeled by a power law non-homogeneous Poisson process:
Taking the natural logarithm of both sides yields a linear equation:
When cumulative failures are plotted against cumulative operating hours on log-log scale graph paper, the historical data points form a straight line whose slope is equal to $\alpha$ (alpha).
Interpreting the Crow-AMSAA Slope ($\alpha$)
The slope parameter $\alpha$ reveals the reliability trajectory of the repairable system:
- $\alpha < 1.0$ (Reliability Growth):
- The slope is less than 45 degrees.
- The failure rate is decreasing over cumulative operating time; MTBF is increasing.
- Operational Meaning: Proactive reliability initiatives, defect elimination, design modifications, and precision maintenance practices are working. Failures are becoming progressively less frequent.
- $\alpha = 1.0$ (Steady State):
- The slope is exactly 45 degrees.
- The system has a constant failure rate; MTBF is unchanging.
- Operational Meaning: The system is in a stable operating regime. Maintenance merely restores the asset to its historical baseline without altering inherent failure physics.
- $\alpha > 1.0$ (Reliability Deterioration):
- The slope is greater than 45 degrees.
- The failure rate is accelerating; MTBF is shrinking.
- Operational Meaning: System deterioration, aging wear-out, improper maintenance practices, chronic operator abuse, or escalating secondary damage cascades. Immediate management intervention is required.
Gap Analysis: Benchmarking Current Reliability against Industry Standards
A structured Reliability Gap Analysis compares current performance with required performance, internal Best-Demonstrated Performance (BDP), and genuinely comparable external data:
Quantifying the Financial Impact of the Gap
To translate an operational reliability gap into an executive business case, reliability managers quantify the Annual Opportunity Cost of Unreliability:
For example, if an automotive stamping plant currently runs at 84.0% availability against its approved target of 94.0%, the 10.0% gap across 6,000 planned annual operating hours represents 600 hours of lost production. At a contribution margin of $12,000 per hour, the unreliability gap inflicts $7.2 million in lost gross profit annually. Quantifying this gap provides the mathematical justification for deploying RCM, precision training, and advanced predictive maintenance technologies.
What is the most defensible use of the historic Nowlan and Heap finding that 89% of studied items followed age-independent patterns?
A fitted Weibull model for a defined fan population gives shape parameter beta = 0.65. Which interpretation is best supported, assuming the model and data are adequate?
When plotting cumulative failures against cumulative operating time on a log-log scale using the Crow-AMSAA reliability growth model, an engineer finds that the fitted slope parameter alpha equals 0.72. How should management interpret this result?