6.1 Reliability in New Equipment Design & Life Cycle

Key Takeaways

  • Many reliability, maintainability, and lifecycle-cost drivers are determined during concept and design, when alternatives are easier and less costly to change; the exact percentage is project specific.
  • Life-cycle costing compares acquisition, installation, operation, maintenance, downtime, and end-of-life consequences using documented assumptions; acquisition is not a fixed universal fraction of TCO.
  • Life-Cycle Costing according to ISO 15663 and ISO 55000 evaluates Total Cost of Ownership (TCO)—including acquisition, installation, operating energy, maintenance, and decommissioning—demonstrating that initial purchase price typically represents only 10% to 20% of total lifetime expense.
  • Equipment reliability procurement standards must specify quantitative RAM targets (MTBF and MTTR), component standardization to control MRO inventory proliferation, and formal vendor verification data.
  • A rigorous commissioning framework progresses through Factory Acceptance Testing (FAT), Site Acceptance Testing (SAT), Cold Commissioning, and Hot Commissioning, culminating in comprehensive handover packages comprising asset hierarchies, BOMs, FMEA-based PM tasks, and baseline vibration/oil condition data.
Last updated: September 2026

Reliability in New Equipment Design & Life Cycle

Quick Answer: Integrate reliability and maintainability into concept, FEED, specification, procurement, testing, commissioning, and handover. Early decisions often have the greatest leverage because design alternatives remain open, while TCO, FAT, SAT, and acceptance criteria keep selection focused on required lifecycle performance.

The Capital Asset Life Cycle & The Cost of Change Curve

A fundamental reality of physical asset management is that an asset's ultimate operational reliability is largely predetermined long before the first maintenance technician turns a wrench. Historically, industrial organizations maintained a strict organizational silo between Capital Project Teams (tasked with delivering projects on time and under budget) and Operations & Maintenance (O&M) Teams (tasked with operating and repairing the plant for the next 20 to 40 years).

This disconnection creates the classic "lowest initial purchase price trap":

  • Capital project engineers select machinery based on the lowest initial tender bid to meet initial CapEx budget caps.
  • The selected equipment contains sub-grade seals, inferior bearings, inadequate lubrication systems, and unstandardized instrumentation.
  • Excess operating energy, downtime, recurring component failure, and excess spare holdings can follow, often outweighing the original purchase-price saving; quantify the case rather than applying a universal multiplier.

The Cost of Change vs. Ease of Influence Curve

The economic imperative for front-end reliability is illustrated by the Cost of Change Curve:

  • Conceptual & FEED Phases: The ability to influence ultimate equipment reliability is at its maximum (100%), while the financial cost of making design modifications is virtually zero (erasing lines on CAD drawings or editing equipment specification datasheets).
  • Detailed Engineering & Procurement: Equipment drawings are finalized, purchase orders placed, and long-lead forgings cast. Modifying design specifications now incurs engineering change orders and vendor penalty fees.
  • Construction & Installation: Equipment is mounted on concrete foundations and welded to process piping. Relocating a poorly positioned valve or modifying an undersized motor base now requires field cutting, re-welding, non-destructive testing (NDT), and structural rework.
  • Operations & Maintenance: Modifying equipment in service requires production line shutdowns, plant outages, Management of Change (MOC) hazard reviews, capital expenditure approvals, and costly physical retrofits. The cost of change is now 100 to 1,000 times greater than during the design phase.

Design for Reliability (DFR) & Design for Maintainability (DFM)

Reliability organizations should involve maintenance and reliability professionals in capital project design teams from the initial Front-End Engineering Design (FEED) phase through two core engineering disciplines:

1. Design for Reliability (DFR)

Design for Reliability is an engineering methodology aimed at eliminating potential failure modes during the conceptual and design phases of an asset's life cycle. DFR practices include:

  • Derating Components: Specifying electrical motors, mechanical drives, and structural shafts to operate comfortably below their maximum thermal, electrical, and mechanical ratings (e.g., loading electric motors to 75–80% of rated capacity to reduce winding insulation operating temperatures).
  • Failure Modes and Effects Analysis in Design (DFMEA): Conducting systematic reviews of preliminary equipment models to identify failure mechanisms (cavitation, thermal fatigue, stress corrosion cracking, erosion) and eliminating them through material selection (e.g., upgrading from 316 stainless steel to duplex stainless steel in high-chloride slurry service).
  • Redundancy & Fault Tolerance: Designing parallel or standby configurations for high-criticality assets ($N+1$ pump configurations with automatic auto-transfer switches) so that a single component failure does not cause a total process shutdown.
  • Environmental & Operational Robustness: Specifying heavy-duty sealing arrangements (magnetic bearing isolators instead of elastomer lip seals), IP66 / NEMA 4X electrical enclosures, and severe-duty coatings to withstand ambient temperature extremes, moisture, and chemical washdowns.

2. Design for Maintainability (DFM)

Design for Maintainability ensures that when maintenance, inspection, or repairs are required, they can be executed safely, rapidly, correctly, and cost-effectively. DFM focuses directly on reducing Mean Time To Repair (MTTR) through:

  • Physical Accessibility & Ergonomics: Providing permanent monorails, crane lifting lugs, unobstructed pull spaces for tube bundles, and clear technician access corridors around large machinery (eliminating the need to disassemble building walls or construct temporary scaffolding for routine pump overhauls).
  • Modularity & Quick-Change Subassemblies: Designing standardized "plug-and-play" modular cartridge assemblies (e.g., pre-assembled cartridge mechanical seals, drop-in replacement pump power ends, modular hydraulic power packs) that can be swapped in minutes in the field and rebuilt offline under clean workshop conditions.
  • Poka-Yoke (Error-Proofing): Designing physical components with asymmetrical bolt patterns, keyed shafts, or unique connector pinouts that make incorrect field reassembly physically impossible.
  • Integrated Condition Monitoring Facilities: Mandating pre-drilled and tapped flat-faced transducer mounting pads for vibration accelerometers, installed thermowells for temperature probes, permanently piped oil sampling ports with sample tubes positioned in active laminar flow zones, and full-port ball valves for ultrasonic leak testing.

Life-Cycle Costing (LCC) and Asset-Management Decisions

Life-Cycle Costing (LCC) is the formal economic methodology used to quantify the Total Cost of Ownership (TCO) of an asset across its entire operational lifespan—from initial concept to final disposal. ISO 15663 addresses life-cycle costing for petroleum, petrochemical, and natural-gas industries, while the ISO 55000 family addresses asset management. Use each within its scope and apply a documented lifecycle model suitable for the decision.

The Total Cost of Ownership (TCO) Iceberg

Purchase price is visible, but it may be only one of many material lifecycle costs. The relative shares of acquisition, energy, maintenance, downtime, support, and decommissioning depend on the asset and duty. Estimate them from evidence and show uncertainty.

Total Life-Cycle Cost (LCC)=Cacquisition+Cinstallation+Cenergy+Cmaintenance+Cunreliability+CdisposalS\text{Total Life-Cycle Cost (LCC)} = C_{\text{acquisition}} + C_{\text{installation}} + C_{\text{energy}} + C_{\text{maintenance}} + C_{\text{unreliability}} + C_{\text{disposal}} - S

Where:

  • $C_{\text{acquisition}}$ (Capital Acquisition Cost): Initial purchase price, engineering design fees, procurement administration, licensing, and freight.
  • $C_{\text{installation}}$ (Installation & Commissioning Cost): Civil foundations, structural steel, piping tie-ins, electrical hookups, FAT/SAT, alignment, and commissioning.
  • $C_{\text{energy}}$ (Operational Energy & Consumables): Cumulative electrical power, steam, natural gas, compressed air, and process water consumed over the asset's operating life (often the single largest cost component for continuous rotating equipment such as pumps, fans, and compressors).
  • $C_{\text{maintenance}}$ (Sustaining Maintenance Cost): Planned preventive maintenance, scheduled overhauls, predictive condition monitoring, replacement spare parts, and technician labor.
  • $C_{\text{unreliability}}$ (Cost of Unreliability): Lost production revenue, product quality scrap, expedited freight fees, and environmental/safety fines resulting from unscheduled equipment downtime.
  • $C_{\text{disposal}}$ (Decommissioning & Disposal Cost): Decontamination, environmental remediation, dismantling, hazardous waste disposal, and site restoration.
  • $S$ (Residual Salvage Value): Net financial recovery from asset resale or scrap metal recycling at end-of-life.

Life-Cycle Cost by Asset Phase

Life-cycle cost depends on asset type, duty, energy use, support model, analysis boundary, discounting, and consequence. There is no universal percentage split. Use phases to make costs and decisions visible:

PhaseRepresentative costsHigh-leverage reliability decisionsEvidence at the gate
Concept and front-end definitionstudies, alternatives, estimatingrequired function, duty, redundancy, maintainability, lifecycle horizonapproved requirements, assumptions, risk and LCC model
Design and procurementengineering, equipment, vendor servicesmaterials, operating envelope, accessibility, standardization, monitoring provisionsreviewed design and technical purchase specification
Fabrication and installationfabrication, civil, piping, electrical, quality controlworkmanship, cleanliness, foundations, piping strain, preservationinspection records, nonconformance closure, installation checks
Commissioning and handovertesting, training, spares, data loadingacceptance criteria, baseline condition, proof of function, ownershiptest evidence, punch-list status, controlled records
Operation and maintenanceenergy, labor, consumables, downtime, repair, supportoperating discipline, task strategy, defect elimination, obsolescenceperformance history, condition and work data, reviewed strategy
Decommissioningisolation, remediation, disposal, salvagesafe end state, environmental duties, information retentionapproved closure and disposal records

Discount future cash flows consistently and test uncertain drivers such as energy price, demand, failure rate, repair duration, and residual value. A low purchase price can be outweighed by energy, downtime, support, or disposal cost, but the model—not a slogan—must show it.

Developing Equipment Reliability Specifications

A technical purchase specification should convert the operating requirement into verifiable supplier obligations. Depending on the asset, it may include:

  1. required functions, capacity, duty cycle, operating envelope, environment, and design life;
  2. defined reliability, availability, and maintainability measures with the population, exposure, exclusions, and demonstration method stated;
  3. applicable legal, code, consensus-standard, owner, and cybersecurity requirements;
  4. materials, interfaces, guarding, lifting, access, modularity, isolation, and maintainability constraints;
  5. installation, preservation, cleanliness, alignment, balancing, lubrication, and quality requirements appropriate to the equipment;
  6. drawings, asset data, bills of material, software or firmware information, recommended spares, training, and lifecycle support;
  7. inspection, FAT, SAT, commissioning, acceptance, nonconformance, and warranty provisions.

Avoid inserting generic MTBF, MTTR, availability, vibration, alignment, or balance targets without a defined duty and verification basis. A contractual guarantee is useful only when the metric, test duration, operating conditions, exclusions, remedy, and responsible data source are unambiguous.

Commissioning & Acceptance Testing: FAT, SAT, and Phased Startup

A major contributor to infant mortality (Nowlan & Heap Pattern F, which accounts for 68% of industrial failure modes) is poor commissioning. Equipment that is improperly stored, unaligned, contaminated with construction debris, or operated with dry mechanical seals will fail within hours or days of plant startup. To mitigate infant mortality, capital projects execute a disciplined four-stage acceptance testing sequence:

1. Factory Acceptance Testing (FAT)

Conducted at the equipment vendor's manufacturing facility prior to crate shipment. FAT verifies that the equipment meets all mechanical, electrical, and performance specifications in the presence of the owner's reliability engineer:

  • Hydrostatic pressure testing of pressure-retaining casings.
  • Full-speed mechanical run tests verifying bearing housing vibration levels across all octaves.
  • Performance curve verification (flow, head, efficiency, Net Positive Suction Head Required [NPSHr], power draw).
  • Verification of emergency trip devices, safety interlocks, and control logic.
  • Resolution of all punch-list defects before the equipment leaves the factory floor.

2. Site Acceptance Testing (SAT)

Conducted after the asset arrives on-site, has been uncrated, inspected for transit damage, and permanently mounted on its plant foundation with field piping and electrical connections attached:

  • Static electrical insulation resistance (Megger) and motor circuit analysis.
  • Verification of foundation flatness, anchor bolt torque, and epoxy grout integrity.
  • Laser shaft alignment audits to verify zero residual pipe strain (uncoupled and coupled).
  • Instrument loop checks and automated safety shutoff testing.

3. Cold Commissioning (Dry / Static Testing)

Operating the machinery without process fluids or hazardous chemicals (often using clean water, dry air, or running uncoupled):

  • Rotational direction verification ("bump test" of motors to ensure proper phase rotation).
  • High-velocity oil flushing of hydraulic and lube oil consoles to verify target ISO 4406 cleanliness codes.
  • Instrument calibration verification and alarm setpoint validation.
  • Dynamic vibration testing under no-load or clean-water conditions.

4. Hot Commissioning (Live / Dynamic Testing)

Introducing live process fluids, raw materials, operational temperatures, and design operating pressures:

  • Establishing baseline operating thermal profiles via infrared thermography.
  • Capturing baseline vibration spectral signatures (FFT) across all bearing points at full operational speed and load.
  • Extracting initial baseline oil samples after 24 to 72 hours of continuous operation for laboratory viscosity, particle count, and spectrometric wear analysis.
  • Verifying mechanical seal face temperatures and barrier fluid flow rates.

Commissioning & Handover Gate Checklist Table

Project StageAcceptance GateRequired Verification & Quality CriteriaStakeholder Sign-OffFailure Consequences of Non-Compliance
ProcurementGate 1: Specification ReviewDFR/DFM compliance; pre-approved vendor list; MRO parts standardization; RAM targets included in RFP.Reliability Engineering, ProcurementProliferation of orphan spare parts; inaccessible machinery; unmaintainable components.
Factory FabricationGate 2: Factory Acceptance Test (FAT)Certified performance curve; full-load vibration within ISO Class 1; zero casing porosity or leak; control interlock test.Owner Reliability Rep, Vendor QA/QCReceiving substandard machinery; costly field machining retrofits; extended project delays.
Mechanical CompletionGate 3: Site Acceptance Test (SAT)Zero pipe strain (< 0.002 in dial deflection during flange bolt-up); soft foot < 0.002 in; anchor bolt torque verified.Mechanical Construction Lead, Maintenance LeadAccelerated bearing fatigue; chronic mechanical seal leakage; frame distortion; severe vibration.
Pre-StartupGate 4: Cold CommissioningMotor rotation direction verified; lube oil flushed to the approved equipment-specific ISO 4406 target; electrical Megger test > 100 Mohm; loop checks passed.Electrical & Instrumentation Lead, OperationsMotor running backward; immediate catastrophic bearing seizure from construction grit; electrical short.
Commercial OperationGate 5: Hot Handover & CloseoutFull-load baseline vibration logged in CMMS; baseline oil analysis recorded; complete BOMs loaded; PMs scheduled.Plant Operations Manager, Maintenance ManagerPremature infant mortality; inability to order replacement parts; missing PM schedules in CMMS.

Handover Documentation and CMMS Readiness

Handover should provide the information and capability needed to operate, maintain, troubleshoot, and modify the asset safely. The required package is risk- and asset-specific; commissioning should not be treated as complete merely because documents were transmitted.

A readiness gate may verify:

  1. Asset structure and identity: approved tags, parent-child relationships, nameplate attributes, locations, system boundaries, and serialized items where needed.
  2. Controlled technical information: approved drawings, manuals, data sheets, certificates, settings, software or firmware records, cause-and-effect information, and as-built revisions.
  3. Materials support: maintainable-item lists, illustrated parts information, critical-spare decisions, interchangeability, preservation requirements, and inventory links.
  4. Maintenance strategy: tasks justified by failure mode, consequence, regulation, warranty, or operating experience; frequencies, skills, tools, permits, measurements, and acceptance limits are explicit where applicable.
  5. Commissioning and baseline evidence: completed tests, resolved or accepted punch items, initial readings for the condition-monitoring techniques actually selected, and a clear reference to operating state and load.
  6. People and ownership: operating and maintenance training, demonstrated competence where required, warranty contacts, responsible record owners, and a process for resolving missing or incorrect data.

Not every asset needs vibration spectra, oil analysis, thermography, or motor-current baselines. Capture the modalities justified for the failure modes and monitoring plan. Likewise, an equipment bill of material should reach the maintainable level needed for planning and supply; demanding every possible vendor part can add noise without improving readiness.

After startup, compare actual performance with the defined requirements. Early failures, nuisance alarms, inaccessible tasks, inaccurate bills of material, and incorrect job assumptions should enter controlled defect and change processes. The handover package is a governed baseline that can be corrected as verified operating knowledge grows, not a static archive.

Test Your Knowledge

A chemical manufacturing company is designing a multi-million dollar plant expansion. During the Front-End Engineering Design (FEED) stage, the project team evaluates equipment layout and reliability standards. According to the Cost of Change principle in asset lifecycle management, why is reliability intervention most critical during this specific phase?

A
B
C
D
Test Your Knowledge

A procurement team is comparing slurry-pump bids. How should it use a documented life-cycle-cost model?

A
B
C
D
Test Your Knowledge

A commissioning plan labels a set of pre-process checks as cold commissioning. Which listed activity best fits that phase?

A
B
C
D