12.4 Operational Troubleshooting
Key Takeaways
- A failed run is a bounded batch with an assignable local cause; a failed assay is a method that no longer performs across days, lots, and operators and must leave clinical service until re-optimized and revalidated.
- Instrument logs (volumes, errors, program names, temperatures) are primary evidence; read them for the failure timestamp before retitering the primary.
- When a control shift coincides with a new lot, quarantine that lot, restore a previously acceptable in-date lot if available, and do not report patients from affected runs.
- Repeating is restaining; investigating is explaining why the first attempt was untrustworthy and what will change; a blind repeat is not a corrective action.
- Quarantine physically segregates a suspect reagent; it is not discarding the laboratory's entire inventory, and it is not using the bottle for a STAT while the file is open.
12.4 Operational Troubleshooting
Quick Answer: Separate a failed run (one batch, assignable cause, controls or instrument out of spec) from a failed assay (the method no longer performs across days, lots, and operators). Open instrument logs and lot logs before you blindly restain. When a control shift coincides with a new lot, quarantine that lot. Repeating is not investigating. This section is operations; stain-appearance tables—blank slide, dirty background, wrong localization—belong with staining troubleshooting in Chapter 11.
This OpenExamPrep section is independent teaching on operational IHC troubleshooting covering published QIHC laboratory-operations topic areas. It is not an ASCP or CAP publication and does not claim Board or College approval. Chapter 11 asks what the stain looks like when chemistry fails. This section asks who stops the line, which records you open, and when a bottle leaves the bench.
Failed run versus failed assay
A failed run is a bounded event. Examples:
- On-slide controls fail on one rack while the previous rack the same morning passed with the same lots
- The autostainer aborts, a bulk bottle runs dry, a probe clogs, or a retrieval vessel boils dry
- One operator skips a blocking step on a manual batch
- A single control block is exhausted, dried out, cut from an old unstained spare (Section 12.1), or is the wrong tissue
Action: do not report patient results from the affected run. Document. Fix the assignable cause. Repeat the run with passing controls. The assay itself remains in service.
A failed assay is a method problem. Examples:
- Controls fail across days, instruments, and operators after obvious run defects are excluded
- Low-expressor positives disappear while 3+ controls still pass (sensitivity drift—the control lesson from Section 12.2 showing up as an operations failure)
- A new clone was treated as a lot change and the concordance file never existed (GPS 15 ignored)
- After a platform move or water-system change (GPS 14 territory), several markers misbehave despite fresh lots
Action: remove the assay from clinical service (or restrict it) until optimization and the appropriate revalidation or verification are done. Inform the medical director. Do not keep releasing "probably okay" patient slides because the send-out backlog is painful.
| Question | Failed run | Failed assay |
|---|---|---|
| How many events? | One batch or one instrument cycle | Recurrent after run defects are fixed |
| Controls | Fail with an assignable local cause | Fail, or pass deceptively (3+ only), across lots |
| Patient reporting | Hold that run; repeat after a defined fix | Hold the menu item |
| Typical first documents | That run's log, instrument error, that rack's lots | Trend of QC, multiple lot files, SOP versus bench |
| Fix | Repair cause; repeat the run | Re-optimize; revalidate or verify; possible clone or platform decision |
Mislabeling a failed assay as "another bad run" produces a month of silent under-calling. Mislabeling a dry bulk bottle as "the clone died" produces an unnecessary 20-plus-20 study while the instrument still has an empty tank.
This classification is not a DAB-versus-AEC appearance table. If you need to decide whether background is biotin, pigment, or drying artifact, go to Chapter 11. If you need to decide whether to quarantine a kit and stop the menu, stay here.
Instrument logs are part of the investigation
Automated stainers, retrieval modules, ovens, and refrigerators keep logs: temperatures, reagent volumes, program names, error codes, lid opens, abort times, maintenance due dates. When controls are pale, read the log for that timestamp before you retiter the primary.
Examples of log-detectable run failures:
- Bulk detection or wash ran out at slide 18 of 30
- Retrieval program was extended high-pH instead of the validated citrate-class program
- Instrument reported a dispense failure on positions 4–6
- Overnight refrigerator temperature excursion for the antibody carousel
- Stainer software ran protocol version 3 after the SOP moved to version 4 without a mapped update
If the log is empty because nobody enabled it, that is a documentation failure (Section 12.1), and you have lost the best evidence of a failed run. Do not invent a clone conspiracy when the log shows a dry bottle.
Manual benches still have logs: water-bath temperature, decloaker cycle printout, pH of working retrieval buffer, timer used. "We always do 20 minutes" is not a temperature record. A pale nuclear marker after someone used leftover microwave buffer that boiled down by half is an equipment-log story (evaporative concentration) as much as it is a stain-appearance story.
Review maintenance due dates while you are in the log. A neglected probe that clogs every Tuesday afternoon is a failed-run pattern with a scheduled cause. Cleaning the probe and documenting the maintenance is the fix; ordering a new clone because Tuesday looks weak is not.
Lot change coinciding with a shift
A classic operational pattern: Monday a new primary lot or detection lot is opened. Tuesday the low-expressor control is weak or the background jumps. Coincidence is a hypothesis, not a conclusion—but it is the first hypothesis to test.
Steps:
- Identify every reagent that changed (primary, detection, retrieval buffer, chromogen, even a new lot of charged slides).
- Quarantine the new suspect lot: physical separate shelf, labeled do-not-use, documented in the lot log and the corrective-action file.
- Return, if still in date and previously acceptable, to the prior lot and restain controls, including a low expressor, not only a 3+ core.
- If the prior lot restores the expected pattern, the new lot fails its 1-plus-1 (or kit equivalent) confirmation and stays out of service. Notify the vendor as the laboratory's policy requires; keep the bottle for the complaint.
- If both lots fail, you may be looking at an instrument, water, retrieval buffer, or assay problem rather than a single bottle.
A 1-plus-1 new-lot check that used only a screaming-positive tonsil can "pass" GPS 12 on paper and still miss a sensitivity drop. When a shift appears, re-challenge with low expressors even if the lot check was already checked off.
Do not blend leftover old lot into the new bottle to "average" them. That creates an unvalidated third reagent with no lot identity. Do not keep reporting Tuesday's patients because the 3+ control still looks like a postcard.
Repeating versus investigating
Repeating is staining again. Investigating is explaining why the first attempt was untrustworthy and what will be different.
Repeat-without-investigation looks like: control failed, restain the patient and the same control with the same lots, same program, same possibly clogged probe, until a slide looks pretty or the tissue is gone. That sequence can accidentally "work" if a dried pad is wet the second time, which then teaches the laboratory the wrong lesson: that IHC is a coin flip.
Investigation asks:
- Did this run fail or is the assay failing (table above)?
- What do the instrument log and lot log show for the same hour?
- Was the control tissue exhausted, overcut, stored as an old unstained slide, or the wrong block?
- Was the patient the problem (underfixed center, acid-decalcified bone, alcohol cytology on an NBF-validated assay)?
- Did anyone change dilution, retrieval time, or program name without 2-plus-2 or director review (GPS 13–15)?
- Are other markers on the same instrument failing (points to detection, water, platform) or only one primary (points to that vial or that retrieval)?
After you have a hypothesis, then repeat with a defined change: new lot, completed maintenance, recut from the block, on-slide low-expressor control, or send-out. Document the hypothesis and the result in the corrective-action file. A pathologist's "please repeat" is a request to investigate, not a request to burn the block. If the only remaining tissue is a 3 mm core, investigation before restain is how you avoid converting a failed run into a failed specimen.
When to quarantine a reagent
Quarantine means the reagent is physically segregated and barred from clinical use until a documented decision returns it, returns it with restriction, or discards or returns it to the vendor. It is not hiding the bottle in a drawer. It is not throwing away every antibody in the refrigerator.
Quarantine when:
- The new-lot 1-plus-1 (or 2-plus-2) confirmation fails
- A control shift is temporally tied to that bottle and the prior lot still works
- The reagent suffered a temperature excursion, a freeze-thaw the insert does not allow, or obvious contamination (precipitate, mold, diluted appearance)
- The vial is unlabeled, relabeled by hand without lot traceability, or past open-bottle dating
- The manufacturer issues a recall or the laboratory's incoming inspection fails (wrong clone in the box)
Do not automatically quarantine every antibody because one cytokeratin run was blank—that pattern says run or instrument until proven otherwise. Do not discard a quarantined predictive-kit lot before the medical director and, if needed, the vendor have seen it; you may need it as evidence. Do not use a quarantined lot "just for this STAT." A STAT patient is not a license to run a failed bottle.
Release from quarantine requires documented retesting that meets the laboratory's procedure (repeat 1-plus-1 with an adequate low expressor, or a director decision that the failure was a run defect unrelated to the bottle). Hope is not a release criterion. If release requires a method change (new retrieval time, new dilution), that is GPS 13 or 14 territory, not a quiet note on the fridge.
Operational decision path
When QC fails, walk this order:
- Stop reporting that run's patients.
- Confirm you are looking at the correct control tissue and a recut if the unstained control was old.
- Read instrument logs and lot logs for coincidences.
- Classify failed run versus failed assay.
- Quarantine implicated lots; restore a previously acceptable lot if available.
- Investigate preanalytic patient issues separately from reagent issues.
- Repeat only after a defined fix.
- If the method remains wrong, pull the assay and escalate to optimization and revalidation (Section 12.3).
That sequence is how laboratory operations protect patients. Pretty chromogen after a blind repeat is not the same as a controlled return to service. If the question on the exam is about what color the background is, Chapter 11. If the question is about whether to stop the line, this chapter.
Controls fail on a single autostainer rack. The instrument log shows the detection bulk empty after slide 12. Yesterday's racks with the same primary lot passed. How should this be classified and handled?
A new detection-kit lot is opened Monday. Tuesday the low-expressor HER2 control is weak while a 3-plus control is still strong. What is the best operational next step?