10.4 Failure Modes, Root-Cause Analysis, and Corrective Recommendations
Key Takeaways
- Level IV domain 4.3 Troubleshooting and Analysis carries 20 to 25 percent of the Level IV exam, the largest single analytical weighting in the program.
- Root-cause analysis distinguishes the failure mode, the failure mechanism, and the root cause, which are three different answers to three different questions.
- Most electrical equipment failures trace to insulation degradation from heat, moisture, contamination, or partial discharge, or to loose and corroded connections.
- Failure evidence must be preserved before repair, because the physical artifact is often the only record of what happened.
- A corrective recommendation that addresses only the failed component leaves the underlying cause in place to fail again.
Failure Modes, Root-Cause Analysis, and Corrective Recommendations
Quick Answer: Level IV domain 4.3 Troubleshooting and Analysis carries 20-25 % of the Level IV exam — the largest analytical weighting anywhere in the program. Its tasks are 4.3.1 Determine failure mode and cause, 4.3.2 Perform power quality analysis, and 4.3.4 Recommend corrective measures. Level III contributes 3.3.4 Identify and assess failures.
1. Three different questions
Candidates lose points by answering the wrong one.
| Question | Term | Example |
|---|---|---|
| What failed, observably? | Failure mode | Phase-to-ground fault in the B-phase winding |
| By what physical process? | Failure mechanism | Thermal degradation of turn insulation leading to turn-to-turn short and then to ground |
| Why did that process start? | Root cause | Blocked radiator reduced cooling; the blockage was never found because the inspection route omitted the rear of the unit |
Stopping at the mode produces "the winding shorted." Stopping at the mechanism produces "it overheated." Only the root cause tells you what to change so it does not happen again — and in this example the fix is a change to the inspection route, not a better winding.
A useful discipline is the five whys: the winding shorted; because the insulation degraded; because it overheated; because cooling was inadequate; because a radiator was blocked; because nothing in the maintenance program inspected it. The actionable answer is at the bottom, not the top.
2. The dominant failure mechanisms
Insulation failure — the largest single category across all apparatus:
| Mechanism | Signature |
|---|---|
| Thermal degradation | Brittle, darkened, embrittled insulation; the Montsinger rule of thumb holds that insulation life roughly halves for each 8-10 °C of sustained overtemperature |
| Moisture ingress | Low IR, low PI, elevated power factor, elevated water content in liquid |
| Contamination and tracking | Carbonized surface tracks, low IR that improves after cleaning |
| Partial discharge | Voids and treeing, ozone and nitric acid byproducts, ultrasonic and UV signature |
| Mechanical stress | Cracking from thermal cycling, vibration, or through-fault forces |
| Electrical overstress | Puncture path from a switching or lightning transient |
Connection failure — the second great category and the one thermography exists to find. The progression is a closed loop: a loose or corroded joint has elevated resistance, resistance produces I²R heat, heat causes differential expansion and further loosening plus accelerated oxidation, and both raise resistance again. Left alone it terminates in an open circuit or a fire.
Specific causes: inadequate or excessive torque, absence of Belleville washers where the design requires them, aluminium-to-copper interfaces without the correct connector and antioxidant compound, contamination in the joint, and thermal cycling from load variation.
Mechanical failure — binding and seized mechanisms from hardened or wrong lubricant, worn linkages, fatigue-fractured springs, and misalignment.
Contamination — moisture, dust, salt, industrial process residue, and animal or insect intrusion. Rodents, snakes, and birds cause a genuinely significant share of substation faults and the fix is a physical barrier, not an electrical one.
Through-fault damage — a transformer that survives repeated external faults accumulates mechanical winding displacement from the electromagnetic forces. It fails later, apparently spontaneously, and the real cause is a history of through-faults. Sweep frequency response analysis (SFRA) detects winding movement that no other test finds, which is why an SFRA signature taken at commissioning is valuable.
3. Preserving the evidence
This is the step most often destroyed by well-meaning haste, and it is worth stating first among the procedure.
Before anything is cleaned, cut, or replaced:
- Photograph everything, from wide context down to detail, including the surrounding equipment and the labels.
- Record positions — switch and breaker positions, relay targets, indicator flags, counter readings.
- Retrieve volatile data first. Relay event records, oscillography, sequence-of-events logs, and power quality recorder data may be overwritten or lost on power cycling. Download before restoring power.
- Take samples before disturbing — oil for DGA, SF₆, debris and residue.
- Do not clean the failure surface. The pattern of tracking, arcing, and deposition is the evidence.
- Preserve failed components. Bag and label the failed part; do not scrap it.
A repair crew that arrives and tidies up before the investigator does has destroyed the record permanently.
4. Structured investigation
- Make it safe. Isolate, ground, verify. Nothing else starts until this is done.
- Preserve evidence, per Section 3.
- Build the timeline. Use time-synchronized SOE and event records. What happened first? Did protection operate as designed? Was the operation correct for the condition present?
- Gather the history: maintenance records, previous test data, loading history, known events such as through-faults or lightning, and recent changes.
- Test the surviving equipment. Adjacent phases, adjacent units, and the upstream and downstream apparatus. A failure often has a cause outside the failed component.
- Analyze. Compare the physical evidence against the electrical evidence. Do they tell the same story? A winding failure with no protective operation, or a protective operation with no physical damage, are both important discrepancies.
- Identify the root cause — and continue until the answer is something that can be changed.
- Consider systemic implications. Do sister units share the condition? This is the question that turns one investigation into avoided failures.
Distinguish primary from secondary damage. A catastrophic failure destroys a great deal, and most of what an investigator sees was caused by the failure rather than causing it. Working outward from the least-damaged region toward the most-damaged usually locates the origin, because the initiating point is often not the worst-damaged point.
5. Common misdiagnoses
| Observed | Frequently blamed | Often actually |
|---|---|---|
| Repeated capacitor failures | Defective capacitors | Harmonic resonance with source inductance |
| Motor bearing failure on a VFD | Bearing quality | Shaft current / EDM fluting from common-mode voltage |
| Nuisance breaker tripping | Faulty trip unit | Load growth, harmonics with a non-true-RMS setting, or a coordination problem |
| Arrester failure | Defective arrester | MCOV wrongly selected for the system grounding method |
| Transformer failure after years of service | Age | Accumulated through-fault winding movement |
| Cable termination failure | Poor cable | Workmanship — semicon not removed, wrong stress relief, contamination during assembly |
| Relay misoperation | Relay defect | Settings, CT saturation, or wiring |
| Switchgear flashover | Insulation age | Contamination and moisture, or animal intrusion |
The pattern in every row: the component that failed is rarely the thing that needs to change.
6. Writing the corrective recommendation
A recommendation that only replaces the failed part guarantees recurrence. A complete recommendation has four layers:
- Immediate — make safe, isolate, provide temporary supply, protect personnel.
- Repair — restore the failed equipment to service, with the specific scope stated.
- Root-cause corrective — eliminate the cause. Change the setting, add the filter, correct the selection, fix the cooling, add the barrier, revise the procedure.
- Preventive / systemic — apply the lesson to sister equipment, revise the maintenance program, add the test that would have caught it, update the drawings or the specification.
Each item needs an owner, a priority, and a target date, or it is a suggestion rather than a corrective action.
The recommendation must also state what is not known. An investigation that cannot establish the root cause with confidence says so, states the candidate hypotheses, and recommends what would discriminate among them. A confident wrong root cause is worse than an honest open one, because it closes the investigation and leaves the real cause in place.
Exam trap: A question describes a transformer that failed with a shorted winding and asks for the root cause, offering "the winding insulation failed." That is the failure mode. The root cause is the answer to why the insulation degraded — inadequate cooling, sustained overload, moisture ingress, or a maintenance program that never inspected the relevant item.
A transformer fails with a shorted winding. Investigation finds the insulation was thermally degraded, cooling was inadequate because a radiator was blocked, and the blockage was never found because the inspection route omitted the rear of the unit. What is the root cause?
What must be done before a repair crew begins work at the site of an equipment failure?
A motor driven by a variable frequency drive suffers repeated bearing failures showing a fluting pattern on the race. What is the most likely cause?