1.3 Measuring Performance & SMRP Standard Metrics
Key Takeaways
- Use both leading process indicators and lagging outcome indicators, and define every metric with an owner, scope, time basis, inclusions, exclusions, and data source.
- Maintenance cost as a percentage of replacement asset value divides annual maintenance cost by current replacement asset value; benchmark comparisons require consistent numerator and denominator definitions.
- MTBF measures operating time per failure and MTTR measures active repair time per failure; both require a clear population, failure definition, and observation period.
- The expression MTBF divided by MTBF plus MTTR is a two-state inherent-availability approximation under stated assumptions, not a universal operational-availability formula.
- Targets should come from asset requirements, business risk, historical capability, and genuinely comparable external data rather than from one universal world-class number.
1.3 Measuring Performance and SMRP Metrics
Quick Answer: A useful maintenance dashboard combines process measures that teams can act on now with outcome measures that show whether risk, cost, and production performance actually improved. Formula accuracy is necessary, but definition control is what makes a metric trustworthy.
Start with the decision
SMRP Body of Knowledge function 1.3 asks the professional to select key performance indicators, track them, and report them. Selection begins with a decision, not with a catalog of fashionable ratios. Ask:
- What business or asset objective does the measure support?
- Who can act when the result changes?
- What population, period, and event definition are used?
- Which source system supplies the numerator and denominator?
- What behavior might the measure accidentally encourage?
- Which companion measure prevents a misleading conclusion?
A KPI record should include a name, purpose, formula, unit, owner, data source, refresh frequency, inclusions, exclusions, target basis, and escalation rule. Version those definitions. A percentage calculated differently at two plants is not comparable merely because both reports use the same label.
Leading and lagging indicators
A leading indicator monitors execution of a process expected to influence a future outcome. Examples include completion of condition-monitoring routes, timely conversion of actionable findings into work, job-plan quality, material readiness, and PM completion within the facility's permitted window.
A lagging indicator describes a result that has already occurred. Examples include lost production, total maintenance cost, safety incidents, functional failures, and warranty claims.
The classification depends partly on the decision. MTBF is lagging when it reports last quarter's failures, but a downward MTBF trend can be an early warning for a larger future production loss. Avoid treating the labels as absolute. Pair measures in a cause-and-effect chain:
- verify the intended process was executed;
- verify that it found or prevented meaningful defects;
- verify that failure, risk, cost, or production performance changed.
High compliance alone does not prove task effectiveness. A plant can complete ineffective PM tasks on time.
Definition-control example
Suppose two sites report 90% schedule compliance. Site A divides completed scheduled work-order count by scheduled work-order count. Site B divides completed scheduled labor hours by scheduled labor hours and excludes approved shutdown cancellations. Both calculations can support management, but the values are not interchangeable. Reports must state the method.
The same issue appears in the word failure. A nuisance trip, loss of required function, degraded performance, and component replacement are different events. Before calculating MTBF, specify the asset boundary and event rule.
Maintenance cost as a percentage of replacement asset value
This measure compares the annual cost of maintaining assets with the current cost of replacing the defined asset base:
$\text{Maintenance cost as % RAV} = \frac{\text{annual maintenance cost}}{\text{replacement asset value}} \times 100$
Numerator control: Define whether maintenance labor and benefits, materials, contractors, rented tools, departmental overhead, and maintenance-related services are included. Separate expansion capital and other excluded costs consistently.
Denominator control: RAV is a current replacement estimate for the defined physical asset scope, not depreciated book value. State whether engineering, installation, commissioning, buildings, and site infrastructure are included. Exclude land or working capital if the governing definition excludes them.
Worked calculation
A facility defines annual maintenance cost as:
- labor and benefits: $2,100,000
- maintenance materials: $1,500,000
- maintenance contractors: $800,000
- rented tools and included overhead: $100,000
Total maintenance cost is $4,500,000. If the consistently defined RAV is $150,000,000:
The arithmetic is simple; interpretation is not. Published reference ranges vary by industry, asset age, duty, accounting practice, outsourcing, and RAV method. Compare only like definitions and investigate trends, risk, and service outcomes. A lower percentage can reflect efficiency, but it can also reflect deferred work or an inflated denominator.
MTBF and MTTR
For a repairable asset population:
$\text{MTBF} = \frac{\text{operating time in the observation period}}{\text{number of defined functional failures}}$
For active corrective maintenance time:
$\text{MTTR} = \frac{\text{total active corrective repair time}}{\text{number of defined functional failures}}$
Use consistent exposure units and asset boundaries. If ten identical pumps each operate 1,000 hours, the population exposure is 10,000 pump-hours. Calendar time is not automatically operating time.
MTTR is often used loosely. State whether it includes only hands-on diagnosis and repair, or also response, isolation, waiting for parts, testing, and administrative delay. Active repair time supports maintainability analysis; total restoration time supports an operational service decision. They answer different questions.
Two-state availability approximation
For a repairable system that alternates between operating and active corrective repair, with the assumptions represented by the selected MTBF and MTTR:
This is commonly called an inherent or two-state steady-state availability approximation. It excludes delays such as logistics, administration, planned maintenance, and supply constraints unless those times were built into the time definitions. Operational availability normally uses observed uptime divided by total required time or a model that includes all relevant downtime.
Example: During a 600-hour required period, a packaging system has 24 hours of active repair downtime across four failures and otherwise operates.
- operating time = 600 - 24 = 576 hours
- MTBF = 576 / 4 = 144 hours
- MTTR = 24 / 4 = 6 hours
- two-state estimate = 144 / (144 + 6) = 96.0%
Observed availability is also 576 / 600 = 96.0% in this simplified case because no other downtime category exists. If there were 20 hours waiting for parts, the active-repair approximation would no longer equal operational availability.
Cost of unreliability
Direct repair expense is only one consequence of failure. A business-impact model can include:
$\text{Failure impact} = \text{repair cost} + \text{lost contribution margin} + \text{quality loss} + \text{logistics or customer cost} + \text{safety and environmental consequence}$
Do not double count. Lost revenue and lost contribution margin are not the same measure, and a safety consequence should not be reduced to money when the organization uses a separate risk criterion. Document whether costs are actual, estimated, avoided, or probabilistic.
Building a balanced asset dashboard
A balanced dashboard may include:
| Decision perspective | Example measures |
|---|---|
| Business and financial | maintenance cost per unit, cost as % RAV, budget variance, risk exposure |
| Asset outcome | functional failures, availability, MTBF, quality loss, environmental events |
| Work process | schedule compliance, planned-work share, PM/PdM compliance, ready backlog |
| Capability | skill coverage, plan quality, failure-code completeness, corrective-action closure |
Set targets from asset performance requirements, regulatory duties, business risk tolerance, demonstrated process capability, and comparable reference data. Record whether a target is a mandatory limit, a forecast, an interim improvement goal, or a benchmark.
Common interpretation traps
- Gaming a single ratio: Reducing scheduled work can make schedule compliance rise without improving output.
- Changing scope silently: Removing contractors from the cost numerator creates a false improvement.
- Averaging unlike assets: Fleet MTBF can hide deterioration in one critical asset.
- Confusing correlation with cause: Compliance and reliability moving together does not prove one caused the other.
- Ignoring uncertainty: A small number of failures produces unstable rates; show event counts and exposure.
- Using universal targets: A number from another industry is a question prompt, not automatically a local standard.
The purpose of measurement is controlled learning and action. A good report states what changed, why it matters, what data limitations exist, and which owner will respond.
A plant's defined annual maintenance cost is $4,500,000 and the consistently scoped replacement asset value is $150,000,000. What is maintenance cost as a percentage of RAV?
Which measure is most directly a leading indicator of whether a preventive work process is being executed as designed?
A system is required for 600 hours, operates for 576 hours, and has four failures totaling 24 hours of active repair. With no other downtime, what are MTBF and the two-state availability estimate?