11.5 OT and ICS Incident Handling: Safety, Visibility, and Recovery
Key Takeaways
OT incident response begins with human and process safety, equipment protection, and reliable operation; confidentiality, integrity, and availability priorities are system- and hazard-specific rather than a universal reversed triad.
NIST SP 800-82 Rev. 3 advises extreme caution with active scanning on operational networks; approved device-aware scanning may be possible during planned outages, so “active scanning is always prohibited” is too absolute.
Passive network monitoring, controller and HMI logs, historian data, engineering-workstation evidence, change records, and controller logic comparisons provide complementary visibility.
Purdue levels and an industrial DMZ are useful reference patterns, not a mandate that every plant implements identical numbered zones or air gaps.
Containment and recovery must be approved with operations, control engineering, safety, and vendors; validate controller logic, recipes, setpoints, time sources, alarms, and safety functions before returning to normal control.
OT and ICS Incident Handling: Safety, Visibility, and Recovery
Operational Technology (OT) monitors or changes physical processes: power distribution, manufacturing, water treatment, building systems, transportation, and other cyber-physical environments. An incident can injure people, damage equipment, spoil product, or destabilize a process. The response team must therefore integrate cyber investigation with control engineering, operations, safety, maintenance, vendors, and business continuity.
NIST SP 800-82 Rev. 3 is the current final NIST guide as of this review; a Revision 4 initial public draft exists but is not final. The central principle is not a slogan that simply reverses the CIA triad. Safety, reliability, timing, integrity, and availability requirements depend on the process and hazard analysis. Confidentiality can also be critical for credentials, recipes, and engineering designs.
Architecture and Trust Boundaries
The Purdue-style hierarchy is a useful mental model:
- Levels 0 and 1 describe physical processes, sensors, actuators, programmable logic controllers (PLCs), remote terminal units (RTUs), and intelligent electronic devices.
- Level 2 contains supervisory control such as human-machine interfaces (HMIs) and local control servers.
- Level 3 supports site operations, engineering, historians, and manufacturing services.
- An industrial demilitarized zone, often called Level 3.5, can broker traffic between operations and enterprise services.
- Levels 4 and 5 represent enterprise and external services in common diagrams.
Real plants vary. Cloud-managed OT, safety networks, remote substations, vendor links, flat legacy segments, and multiple sites do not always fit identical levels. Treat the model as a way to identify zones and conduits, not proof that an IDMZ exists or that a safety system is physically air-gapped.
Build an incident map from actual diagrams, firewall and switch configurations, controller inventories, data flows, vendor remote-access paths, serial gateways, wireless links, and dependencies. Confirm the map with operators; an undocumented maintenance laptop or cellular modem may be the decisive path.
Detection and Safe Discovery
Prefer low-impact sources first:
- passive network taps or configured mirror ports, with attention to whether the sensor sees the relevant VLAN and protocol;
- firewall, VPN, jump-host, identity, DNS, and remote-access records;
- historian trends, alarm journals, sequence-of-events logs, HMI audit trails, controller diagnostics, and engineering change logs;
- engineering-workstation disk and memory evidence, removable-media records, project files, and vendor tool histories;
- controller logic, firmware, setpoints, recipes, and checksums compared with approved backups.
NIST warns that active scans may destabilize sensitive devices or alter process state. That does not mean every active query is universally forbidden. The asset owner may approve a device-aware method after passive discovery, vendor review, testing, backup verification, and safety analysis—preferably during a planned outage. Generic high-rate vulnerability sweeps and malformed probes should not be launched into production OT from an IT playbook.
Time correlation is difficult. Controllers may have unsynchronized clocks, limited logs, local time without a zone, or events overwritten quickly. Preserve raw records and document offsets rather than silently normalizing the originals.
Triage and Incident Scope
Determine whether the incident affects only the enterprise interface, an engineering workstation, supervisory control, basic control, or a safety function. Ask:
- Is the physical process stable, and are alarms and protective functions trustworthy?
- Has logic, firmware, a setpoint, recipe, calibration, or time source changed?
- Are commands arriving from an expected HMI, engineering station, remote vendor, or unknown path?
- Is the observed behavior malicious, a device fault, maintenance error, or process upset?
- What evidence will be overwritten if the controller reboots or the historian rolls over?
Cyber indicators must be interpreted with process data. A pump cycling rapidly could be malicious command injection, failed instrumentation, a control-loop problem, or legitimate operator action.
Safety-Governed Containment
Do not power down a PLC, block a control protocol, or isolate an HMI solely because a generic playbook says to disconnect the host. The incident commander and control authority choose actions within approved operating envelopes. Options include disabling a vendor remote-access path, filtering a known malicious address at a boundary, moving a compromised engineering workstation off the control network, failing over to a validated redundant controller, shifting to manual operation, or executing a controlled process shutdown.
Containment criteria should define who can authorize each action, required operator verification, rollback steps, communications, and conditions for emergency safe shutdown. Preserve safety instrumented system independence and verify that containment does not remove the last protective layer.
Evidence Collection
Capture configuration and volatile status through vendor-supported read-only functions where possible. Preserve controller project files, ladder logic or function blocks, firmware versions, run or program mode, fault registers, active connections, user and change logs, HMI projects, historian exports, switch tables, and packet captures. Hash exported files, record tool and cable versions, and note that reading a controller can itself create audit events or communication load.
An approved logic backup is valuable only if its provenance is known. Compare current logic with a version-controlled, signed or otherwise protected baseline and have a qualified engineer interpret differences. A checksum mismatch identifies change, not motive.
Eradication and Staged Recovery
Remove unauthorized remote tools and accounts, rotate credentials and certificates, patch supported components under change control, restore validated controller and HMI projects, and rebuild compromised Windows engineering stations from trusted media. Unsupported equipment may require compensating segmentation, protocol filtering, jump-host controls, or planned replacement rather than an untested patch.
Recovery proceeds in a safe sequence defined by operations. Verify sensors and actuators, controller logic, interlocks, alarm paths, historian collection, remote access, backups, and safety functions. Observe the physical process through a stabilization window before lifting heightened monitoring. The lessons-learned review updates network diagrams, asset inventory, vendor access, backup validation, and exercise scenarios.
Exam Scenario Method
For an OT question, reject absolute IT actions. Start with safety and process stability, involve the control authority, collect passive and vendor-supported evidence, and choose the least disruptive containment that stops harm. Active scanning is a risk-managed engineering decision, Purdue is a reference model, and restoration is complete only after both cyber integrity and physical process behavior are validated.
A responder wants to run a generic vulnerability scanner against operating PLCs. What is the best decision under NIST SP 800-82 Rev. 3 principles?
Run the scan immediately because availability is secondary during an incident
Active scanning is forbidden by law on every OT network
Use passive discovery first and permit only an owner-approved, device-aware method after safety, vendor, and outage considerations are addressed
Reboot each PLC before scanning to clear its network stack
Which statement about the Purdue model is most accurate during an OT investigation?
Every industrial network is legally required to implement identical Level 0 through Level 5 segments
A Level 3.5 label proves the safety system is physically air-gapped
Cloud-managed devices cannot be represented in zones and conduits
Purdue levels are a useful reference for trust boundaries, but responders must verify the plant's actual architecture, dependencies, and remote paths
An HMI is compromised but the physical process is stable. Which containment approach is most defensible?
Immediately remove power from every PLC in the cell
Have operations and control engineering approve isolation or failover of the HMI, preserve evidence, verify alternate control and alarms, and maintain a rollback path
Delete all historian records so the attacker cannot see them
Push an untested firmware update to every controller
Sections you finish are checked off in the contents.
You've completed this section
Continue exploring other exams