13.3 Equipment Maintenance Engineering, Availability (A), MTBF/MTTR & RCM
Key Takeaways
Maintenance strategies balance Corrective (run-to-failure), Preventive (time/usage intervals), and Predictive (condition monitoring via vibration, thermography, and oil analysis) approaches.
Maintainability models repair completion probability over duration , where repair rate .
System availability metrics distinguish Inherent Availability (inherent hardware design) from Operational Availability (including supply and administrative delays).
Reliability-Centered Maintenance (RCM) applies SAE JA1011 decision logic to classify failure consequences (hidden, safety, environmental, operational) and assign cost-effective maintenance tasks.
Overall Equipment Effectiveness benchmarks asset performance via , while preventive replacement is cost-justified only when hazard rate increases () and unplanned failure cost exceeds scheduled cost ().
13.3 Equipment Maintenance Engineering, Availability (A), MTBF/MTTR & RCM
Industrial maintenance engineering bridges the theoretical design reliability of equipment with its actual operational performance on the factory floor. Maximizing plant throughput, product quality, and worker safety requires structured maintenance management, rigorous availability metrics, and proactive failure-elimination methodologies.
1. Maintenance Strategies and Paradigms
Modern manufacturing and process industries categorize equipment maintenance into three primary operational paradigms:
+------------------------------+
| Maintenance Strategies |
+------------------------------+
|
+--------------------------+--------------------------+
| | |
+---------------+ +---------------+ +---------------+
| Corrective | | Preventive | | Predictive |
| Maintenance | | Maintenance | | Maintenance |
| (CM) | | (PM) | | (PdM) |
| Run-to-Failure| | Time/Usage | |Condition-Based|
+---------------+ +---------------+ +---------------+
1. Corrective Maintenance (CM / Reactive / Run-to-Failure)
- Operating Concept: Repair or replacement occurs exclusively following functional breakdown or obvious operational failure.
- Engineering Suitability: Cost-effective only for non-critical, redundant, or low-cost assets with negligible failure consequences, zero personnel safety hazard, zero environmental release risk, and minimal collateral machine damage (e.g., standard indicator bulbs, non-critical filter cartridges).
- Disadvantages: High secondary collateral damage, catastrophic unplanned outages, severe production schedule disruption, and excessive overtime labor costs.
2. Preventive Maintenance (PM / Scheduled / Time-Directed)
- Operating Concept: Component servicing, fluid replacement, recalibration, or overhaul occurs at fixed calendar intervals, run hours, or production cycle counts, regardless of apparent machine condition.
- Engineering Suitability: Highly effective for assets exhibiting predictable, aging-related wear-out mechanisms characterized by an Increasing Failure Rate ( in Weibull life models), such as mechanical seals, drive belts, and abrasive slurry impellers.
- Limitations: Over-maintenance; disassembling healthy machines induces infant mortality failures due to reassembly misalignments and contamination. Completely ineffective for random failure modes ().
3. Predictive Maintenance (PdM / Condition-Based Monitoring / CBM)
- Operating Concept: Asset health is continuously or periodically measured using non-intrusive diagnostic instrumentation. Maintenance intervention is scheduled only when physical degradation indicators exceed established warning thresholds.
- The P-F Interval (Potential Failure to Functional Failure):
- Point P represents the earliest moment at which progressive degradation is detectable via advanced instrumentation.
- Point F represents functional failure where the asset can no longer perform its specified operational duty.
- The P-F Interval represents the available time window for maintenance planners to order parts, stage tools, and schedule corrective repair during planned downtime before catastrophic failure occurs.
Condition
^
| Initial Good Condition
|---------\ Point P (Potential Failure Detected: Vibration / Oil)
| \
| \ <------ P-F Interval ------>
| \ \
| \ \ Point F (Functional Failure)
+----------------------------------------------------> Time (t)
Key Condition Monitoring Technologies
- Vibration Analysis (Spectral FFT): Detects unbalance (1X RPM), shaft misalignment (2X RPM), mechanical looseness (harmonics), and bearing rolling-element defects via characteristic fault frequencies: Ball Pass Frequency Outer Race (BPFO), Ball Pass Frequency Inner Race (BPFI), Ball Spin Frequency (BSF), and Fundamental Train Frequency (FTF).
- Lubricant & Ferrographic Analysis: Monitors lubricant viscosity, Total Acid Number (TAN), moisture content (Karl Fischer titration), and particulate wear debris morphology (cutting wear, fatigue spalling particles).
- Infrared Thermography: Detects thermal anomalies caused by excessive electrical contact resistance (loose terminal lugs), phase unbalance, refractory insulation breakdown, and localized bearing friction.
- Airborne and Structure-Borne Ultrasound: Pinpoints compressed gas leaks, high-pressure steam trap failures, and early subsurface bearing friction before vibration amplitudes rise.
- Motor Current Signature Analysis (MCSA): Identifies broken rotor bars, stator winding turn-to-turn shorts, and air-gap eccentricities in induction motors under operational load.
2. Maintainability Engineering and MTTR
Maintainability is the design characteristic of an item representing the probability that a failed system will be restored to full specified operating capability within a given duration , when maintenance is performed in accordance with prescribed procedures and resources.
Mathematical Formulation
Let the repair duration be a continuous random variable with probability density function . The Maintainability function is:
Exponential Repair Distribution
When repair tasks follow a Poisson restoration process with constant repair rate :
Mean Time To Repair (MTTR)
The expected duration required to perform corrective repair is:
Repair Time Percentiles
To determine the time window within which a specified percentage of repair actions will be completed:
Example: The time required to complete 90% of repairs ():
3. System Operational Availability Metrics
Availability represents the probability that a system is operating satisfactorily at time when called upon. In industrial engineering, availability metrics are strictly classified based on the operational scope and downtime categories included.
+-------------------------+
| System Availability |
+-------------------------+
|
+-----------------------------------+-----------------------------------+
| | |
+-----------------------+ +-----------------------+ +-----------------------+
| Inherent | | Achieved | | Operational |
| Availability (Ai) | | Availability (Aa) | | Availability (Ao) |
| Hardware Design Focus | | Preventive+Corrective | | Real-World Field Ops |
| Ideal Support Env. | | Excludes Supply Delay | | Total Field Downtime |
+-----------------------+ +-----------------------+ +-----------------------+
1. Inherent Availability ()
Inherent Availability is an intrinsic engineering design metric that reflects hardware maintainability under ideal operational conditions. It considers only corrective maintenance downtime, completely excluding scheduled preventive maintenance, spare parts supply delays, and administrative logistics wait time.
Where is the Mean Time Between Failures and is the Mean Time To Repair.
2. Achieved Availability ()
Achieved Availability expands the design metric to encompass both unscheduled corrective repairs and scheduled preventive maintenance shutdowns. It assumes ideal support conditions (tools, spare parts, and personnel are immediately available on site without logistics delay).
Where:
- is the Mean Time Between Maintenance (incorporating both corrective breakdowns and scheduled PM events):
- is the Mean Maintenance Downtime:
3. Operational Availability ()
Operational Availability is the true, real-world metric observed in field operations. It relates actual operational uptime to total calendar time, explicitly including all unscheduled repairs, scheduled PM, Logistics Delay Time (LDT) (waiting for parts delivery, specialized riggers, or rigging cranes), and Administrative Delay Time (ADT) (work permit authorization, lockout/tagout clearances, shift changeover delays).
Where Mean Downtime (MDT) represents the total average elapsed outage duration per maintenance event:
Because , the availability metrics always follow the strict hierarchical relationship:
4. Reliability-Centered Maintenance (RCM)
Developed in the late 1960s and early 1970s by Stanley Nowlan and Howard Heap for United Airlines and the U.S. Department of Defense, Reliability-Centered Maintenance (RCM) revolutionized equipment maintenance philosophy. Prior to RCM, industrial asset managers operated under the flawed premise that older components inherently exhibited higher failure probabilities, mandating frequent periodic strip-down overhauls. Nowlan & Heap's empirical data proved that only 11% of aircraft components followed wear-out curves, while 89% exhibited random or infant mortality failure profiles.
The Seven Essential RCM Questions (SAE JA1011 Standard)
Under the international standard SAE JA1011 (Evaluation Criteria for Reliability-Centered Maintenance (RCM) Processes), an authentic RCM process must address seven sequential questions for every physical asset:
- Functions: What are the functions and associated desired standards of performance of the asset in its present operating context?
- Functional Failures: In what ways can the asset fail to fulfill its specified functions?
- Failure Modes: What physical events, root causes, or degradation mechanisms cause each functional failure?
- Failure Effects: What specific physical evidence, alarms, secondary damage, or operational stoppages occur when each failure mode takes place?
- Failure Consequences: In what category does each failure matter?
- Hidden Failure Consequences: Failure of a non-evident safety device (e.g., pressure relief valve stuck closed, standby emergency generator failing to start).
- Safety Consequences: Failure modes causing injury or loss of human life.
- Environmental Consequences: Failure modes violating environmental legislation or operating permits.
- Operational Consequences: Direct production loss, scrap, repair cost, and financial downtime.
- Non-Operational Consequences: Direct repair cost only, where production is unaffected.
- Proactive Tasks: What proactive maintenance task (On-Condition / PdM, Scheduled Restoration, or Scheduled Discard) can be performed to predict or prevent the failure mode?
- Default Actions: If a technically feasible and cost-effective proactive task cannot be identified, what default action must be taken? Options include Failure-Finding Tasks (mandatory for hidden functions), Physical Engineering Redesign, or deliberate Run-to-Failure.
Failure Modes, Effects, and Criticality Analysis (FMECA)
RCM analyzes failure modes and their effects; many organizations pair it with an FMEA or FMECA that ranks failure risks with the Risk Priority Number (RPN):
Where each parameter is evaluated on a standard 1-to-10 scale:
- Severity (): Impact of failure effect (1 = unnoticeable; 10 = catastrophic hazard without warning).
- Occurrence (): Frequency or probability of occurrence based on failure rate (1 = virtually impossible; 10 = almost certain).
- Detection (): Probability that existing controls or diagnostic methods will detect the failure defect before functional failure occurs (1 = almost certain detection; 10 = absolute uncertainty of detection).
Critical Engineering Rule: Failure modes with very high severity (for example or , a safety or regulatory effect) must be addressed by redesign or fail-safe mitigation regardless of how low the composite RPN appears.
5. Total Productive Maintenance (TPM) & Overall Equipment Effectiveness (OEE)
Originating in Japan through the Japan Institute of Plant Maintenance (JIPM), Total Productive Maintenance (TPM) is an equipment-centric lean manufacturing system aimed at maximizing equipment effectiveness through the elimination of all equipment-related waste, breakdowns, and defects.
The Eight Pillars of TPM
- Autonomous Maintenance (Jishu Hozen): Empowers machine operators to conduct daily equipment cleaning, lubrication, bolting inspections, and basic adjustments, fostering operator ownership.
- Planned Maintenance: Professional maintenance engineers execute predictive condition monitoring and scheduled precision overhauls.
- Quality Maintenance: Zero defect manufacturing through machine condition management.
- Focused Improvement (Kobetsu Kaizen): Cross-functional teams eliminate the Six Big Losses.
- Early Equipment Management: Designing high maintainability and reliability into new machinery.
- Education and Training: Continuous skills development for operators and maintenance technicians.
- Safety, Health, and Environment: Achieving zero industrial accidents and environmental spills.
- TPM in Administration: Streamlining maintenance work orders, inventory, and procurement.
Overall Equipment Effectiveness (OEE)
Overall Equipment Effectiveness benchmarks the percentage of planned manufacturing time that is truly productive. It integrates three independent operating ratios:
The Six Big Losses Mapped to OEE Components
| OEE Metric | Category of Loss | Loss Mechanism Description | Root Cause / Engineering Solution |
|---|---|---|---|
| Availability | Loss 1: Equipment Breakdown | Unplanned catastrophic mechanical/electrical downtime (> 5 min) | FMECA, RCM, precision lubrication, vibration PdM |
| Availability | Loss 2: Setup & Adjustments | Die changes, tooling changeovers, warmup, calibration | Single-Minute Exchange of Die (SMED), standard work |
| Performance | Loss 3: Idling & Minor Stoppages | Brief stoppages (< 5 min), part jams, sensor misfires | De-bottlenecking chute feeds, sensor realignment |
| Performance | Loss 4: Reduced Operating Speed | Running machine below nameplate speed due to chatter or wear | Dynamic balancing, structural stiffening, operator training |
| Quality | Loss 5: Process Defects & Scrap | Off-spec parts produced during steady-state production | Statistical Process Control (SPC), Pokayoke, tooling wear |
| Quality | Loss 6: Startup & Yield Losses | Defective product produced during machine thermal warmup | Standardized warmup SOPs, automated pre-heating |
6. Mathematical Optimization of Preventive Replacement Scheduling
A central optimization task in industrial engineering is determining the cost-minimizing preventive replacement interval for equipment subject to wear-out degradation.
Age Replacement Policy Formulation
Under an Age Replacement Policy, a component is replaced whenever it reaches scheduled operating age , or immediately upon unannounced functional failure, whichever occurs first.
- Let represent the cost of a scheduled preventive replacement (performed during planned downtime with pre-staged parts).
- Let represent the cost of an unscheduled corrective replacement following breakdown (where , including secondary damage, emergency logistics, and lost throughput).
Cycle Cost and Length Expectations
A renewal cycle terminates upon either failure (at random time ) or scheduled replacement (at fixed time ):
- Probability of planned replacement at age :
- Probability of emergency breakdown replacement:
The expected cost per replacement cycle is:
The expected length of a replacement cycle is:
Applying the Renewal Reward Theorem, the long-run expected Total Cost per Unit Time is:
Optimality Condition
To find the optimal replacement interval , take the derivative and equate to zero. Using the Leibniz integral rule:
Setting the numerator of to zero yields the fundamental optimality condition:
Where is the instantaneous hazard rate at the optimal replacement age.
Crucial Engineering Insight
An optimal finite replacement interval can exist only if:
- The failure process exhibits an increasing hazard rate (, wear-out regime, ).
- The emergency failure cost is strictly greater than the planned replacement cost ().
If the asset follows an Exponential distribution (constant hazard rate ), the left side of the optimality equation simplifies to zero, proving mathematically that no finite replacement interval minimizes cost. Scheduled replacement of exponentially distributed assets always increases total operating costs.
7. Worked Numerical Examples: Availability, OEE & Maintenance Cost Optimization
Problem Formulation
Part A: Availability Analysis
A high-speed robotic packaging cell in a distribution facility has the following historical reliability parameters:
- Mean Time Between Failures: operating hours.
- Mean Time To Repair (corrective repair only): hours.
- Under actual field operating conditions, spare parts logistics introduce a Logistics Delay Time hours, while technician dispatch and permit sign-off introduce an Administrative Delay Time hours.
Calculate:
- The Inherent Availability ().
- The Operational Availability ().
Part B: Overall Equipment Effectiveness (OEE)
A manufacturing line operates an 8-hour shift (480 minutes). The operating schedule includes a scheduled 30-minute unpaid meal break (planned non-production time). During the operating shift, the line suffers 25 minutes of unplanned breakdown downtime and 15 minutes of tooling changeover adjustments. The ideal cycle time of the process is 0.50 minutes per finished part. During the shift, the cell produces 700 total parts, of which 28 parts fail dimensional inspection and are scrapped.
Calculate:
- Availability rate ().
- Performance rate ().
- Quality rate ().
- Overall Equipment Effectiveness ().
Step-by-Step Solutions
Solution to Part A
1. Inherent Availability ():
2. Operational Availability (): First, calculate the Mean Downtime ():
Now, calculate :
Engineering Assessment: Field support delays (logistics and administration) cause the cell to lose in operational availability compared to its inherent design potential ( vs ).
Solution to Part B
1. Planned Production Time:
2. Operating Time:
3. Availability Rate ():
4. Performance Rate ():
5. Quality Rate ():
6. Overall Equipment Effectiveness (OEE):
Benchmarking Insight: While an OEE of represents typical industrial manufacturing performance, world-class discrete manufacturing facilities target (, , ). Performance () is the primary loss driver in this cell, indicating micro-stoppages and running below design speed.
An automated packaging machine operates with a Mean Time Between Failures of MTBF = 380 operating hours and an average corrective repair duration of MTTR = 20 hours. When operating in an industrial plant environment, corrective repair actions require an average logistics delay time (LDT) of 15 hours for spare parts retrieval and an administrative delay time (ADT) of 5 hours for work permit clearance and technician dispatch. What are the Inherent Availability (Ai) and Operational Availability (Ao) of this packaging machine?
Inherent Availability Ai = 95.0%; Operational Availability Ao = 90.5%
Inherent Availability Ai = 90.5%; Operational Availability Ao = 95.0%
Inherent Availability Ai = 95.0%; Operational Availability Ao = 82.6%
Inherent Availability Ai = 97.4%; Operational Availability Ao = 92.7%
During a plant-wide reliability overhaul, an industrial reliability engineer evaluates maintenance policies according to the Reliability-Centered Maintenance (RCM) framework (SAE JA1011 / Nowlan & Heap). Which of the following statements correctly identifies the appropriate maintenance strategy based on failure characteristics and consequences?
Time-based periodic component overhaul is the optimal default strategy for all electronic control units because aging accelerates breakdown risk
Corrective run-to-failure maintenance should never be permitted under RCM guidelines, even when failure consequences are purely economic and non-critical
Condition-based on-condition tasks are appropriate only if a detectable potential failure (P) exists and the P-F interval is sufficiently long to take corrective action
Autonomous maintenance under TPM transfers the responsibility for major mechanical rebuilds and root-cause failure analysis from maintenance specialists to line operators
Sections you finish are checked off in the contents.
You've completed this section
Continue exploring other exams