11.5 Building and Managing an Infrared Predictive Maintenance Program
Key Takeaways
- Infrared thermography sits in the predictive tier of the maintenance spectrum and is unique among condition monitoring technologies for its breadth, screening hundreds of energised assets per shift without contact.
- Inspection frequency is set by asset criticality - consequence of failure scored on safety, environmental, production loss, repair cost, and redundancy - never by convenience or route length.
- The two steps most often skipped are the two that keep a program funded: closing every exception to a verified, re-inspected repair, and reporting results as cost and downtime avoided rather than as thermal images.
- Programs must adopt a single severity criterion and state its reference (similar component versus ambient air), because mixing published delta-T tables inside one report makes the findings indefensible.
- Infrared identifies where to look and how urgently, while a second technology - ultrasound, vibration, oil analysis, or DLRO micro-ohm testing - usually determines the corrective action; single-technology shutdown recommendations destroy program credibility.
11.5 Building and Managing an Infrared Predictive Maintenance Program
A thermographer who can find a loose lug but cannot get it fixed has produced a photograph, not a maintenance outcome. The commercial value of infrared thermography is realised only inside a predictive maintenance (PdM) program: a documented, scheduled, criticality-driven system that finds exceptions, drives them to repair, verifies the repair, and can prove the money it saved. Level I curricula close with program implementation for exactly this reason — Level I personnel execute the routes, and a route that is badly designed cannot be rescued by good camera work.
1. From Reactive to Predictive
Maintenance strategies sit on a spectrum:
- Reactive (run-to-failure): repair after breakdown. Cheapest to schedule, most expensive to experience — unplanned outage, collateral damage, overtime, expedited parts.
- Preventive (time-based): service at fixed intervals regardless of condition. Predictable, but replaces healthy components and disturbs joints that were working.
- Predictive (condition-based): measure condition, act on evidence. Infrared thermography, vibration analysis, ultrasound, and oil analysis are the core measurement technologies.
- Proactive (root-cause): eliminate the failure mechanism — correct the load imbalance, fix the torque procedure, specify the right lubricant.
Infrared belongs to the third tier and feeds the fourth. Its unique strength is breadth: one thermographer can screen hundreds of assets per shift, non-contact and energised, which no other PdM technology can match.
2. Program Implementation Sequence
Course curricula — including the Infraspection Institute Level I module on infrared predictive maintenance programs — teach implementation as an ordered sequence rather than a checklist to be cherry-picked. Different programs number the steps differently, but the sequence below is the one industrial programs actually follow:
- Define objectives and scope. Write down what the program is for: fire-loss prevention, insurance compliance, uptime, energy, warranty verification, or all of these. Objectives determine which assets go on the route.
- Secure management commitment. Budget for equipment, training, and — critically — for the repairs the program will generate. A program that finds exceptions nobody funds dies in about two cycles.
- Build the asset register and rank criticality. List every candidate asset and score it on consequence of failure: safety, environmental, production loss, repair cost, redundancy. Criticality, not convenience, sets inspection frequency.
- Select equipment and software. Match detector resolution, NETD, lens set, and temperature range to the assets on the register. Buy the reporting and trending software at the same time, not later.
- Qualify personnel. Certify thermographers under an SNT-TC-1A written practice or an ISO 18436-7 central scheme, and define who may evaluate results and sign reports.
- Write procedures and adopt severity criteria. Specify the minimum load, the reference (similar component versus ambient), the standard whose ΔT table you are adopting, the environmental limits, and the required report fields. Adopt one criterion and apply it consistently.
- Design routes and frequencies. Sequence assets to minimise travel and PPE changes, and set frequency by criticality — for example quarterly for critical energised distribution, annually for general lighting panels.
- Capture baselines. Survey every asset under known, documented load and environmental conditions. Everything the program will ever claim about degradation is measured against these images.
- Execute surveys and issue exception reports. Follow the route, meet the load minimum, record the twelve mandatory report elements, and grade each exception against the adopted criteria.
- Close the loop: repair, verify, re-inspect. Track every exception to a work order, and re-image after repair under comparable load. An exception that is never re-inspected is an exception you cannot prove you fixed.
- Trend and report results. Plot ΔT history per asset, and report program KPIs and cost avoidance to management in their language, not in degrees Celsius.
- Audit and improve. Review missed findings, repeat exceptions, and false calls; feed them back into procedures, frequencies, and training.
Steps 10 and 11 are the ones most often skipped, and they are the two that keep the program funded.
3. Integration and Cross-Verification
No single technology sees every failure mode. Mature programs deliberately overlap them and require cross-verification before expensive action:
| Technology | Detects earliest | Confirms an infrared finding by |
|---|---|---|
| Airborne / contact ultrasound | Bearing friction, lubrication distress, air and steam leaks, arcing and corona | Distinguishing a blowing steam trap from flash steam; hearing friction before heat appears |
| Vibration analysis | Imbalance, misalignment, looseness, bearing defect frequencies | Identifying whether a hot bearing is a lubrication problem or a mechanical defect |
| Oil and grease analysis | Wear debris, contamination, additive depletion | Confirming lubricant condition behind an over- or under-lubrication signature |
| Motor circuit analysis | Winding insulation degradation, rotor bar defects | Explaining a hot stator that has no cooling restriction |
| Micro-ohm (DLRO) testing | Contact resistance, offline | Quantifying the joint resistance behind an electrical hot spot |
| Partial discharge / corona survey | Insulation breakdown in medium and high voltage | Explaining warm insulators and terminations at MV/HV |
The rule Level I thermographers should internalise: infrared identifies where to look and how urgently; a second technology usually determines what to do. Recommending a shutdown on a single uncorroborated thermal image is the fastest way to lose a program's credibility.
4. Frequency, Success Factors, and KPIs
Frequency follows criticality and rate of degradation. Electrical connection faults can progress from detectable to failed in weeks under changing load, so quarterly to annual intervals are typical for energised distribution; NFPA 70B has moved infrared inspection of electrical equipment from a recommendation to an expected maintenance activity, which has made an annual interval the practical floor for most facilities. Mechanical routes on critical rotating equipment often run monthly.
Success factors that separate programs that survive from programs that quietly stop:
- Repairs are funded and tracked to completion, then re-inspected.
- Surveys meet the minimum load requirement; low-load surveys generate false confidence.
- One severity criterion is adopted and applied consistently, with the reference stated.
- Baselines exist and are actually compared against.
- Results are reported in cost avoided and downtime avoided, not in thermal images.
KPIs worth tracking: route completion rate, exceptions found per survey, exception closure rate and mean time to repair, repeat-exception rate on previously repaired assets, false-call rate, and documented cost avoidance.
5. Worked Calculation: Proving the Program Pays
A plant runs an infrared PdM program over its electrical distribution system.
Annual program cost
- Thermographer labour: 240 hours at USD 85/hour = USD 20,400
- Camera depreciation and annual calibration = USD 6,500
- Training and recertification = USD 2,200
- Reporting software subscription = USD 1,400
- Total annual program cost = USD 30,500
Documented avoided losses this year
- One critical-tier busway joint corrected before failure. Historical unplanned outage on this bus: 9 hours at USD 14,000/hour of lost production = USD 126,000, plus USD 22,000 in equipment damage. Planned outage repair cost instead: USD 8,000. Net avoided = 126,000 + 22,000 - 8,000 = USD 140,000
- Two motor-feeder connections corrected during a scheduled outage rather than after failure: avoided motor rewind and downtime = USD 31,000
- Steam trap exceptions corrected, verified by ultrasound: USD 27,500 in recovered steam
Step 1 — total documented avoidance. 140,000 + 31,000 + 27,500 = USD 198,500
Step 2 — net benefit. 198,500 - 30,500 = USD 168,000
Step 3 — return on investment. ROI = net benefit / program cost = 168,000 / 30,500 = 5.5, i.e. 550%
Step 4 — payback period. Payback = program cost / avoidance rate = 30,500 / (198,500 / 12 months) = 1.84 months
Step 5 — the credibility rule. Every figure above must trace to a work order, a repair invoice, and a documented historical outage cost agreed with operations before the finding. An ROI built on hypothetical failures the thermographer imagined is worse than no ROI at all — it invites the first auditor who checks to shut the program down.
A facility has run an infrared route for two years. Exceptions are found and reported every quarter, but no repair work orders are tracked and no asset is ever re-imaged after repair. Which program failure does this represent, and why does it matter most?
A Level I thermographer finds an electric motor bearing housing running 18 °C above its identical sister machine. According to sound predictive maintenance program practice, what is the appropriate next action?
An infrared program costs USD 30,500 per year to run and documents USD 198,500 in avoided losses over the same year, every figure traceable to a work order and an agreed historical outage cost. What are the program's net benefit and simple return on investment?
You've completed this section
Continue exploring other exams