Playbook ID: PB-OT | Default severity: SEV-1 (the only downgrade path is a written engineering finding that no control system, safety function or process-network asset is in scope — downgrade to SEV-2, never lower, and record who signed it) | Owner: Incident Commander, paired with a named Engineering Authority who co-signs every OT action
Not for: IT-only ransomware at an organization with no physical process — use 14.1. An exploited perimeter appliance with no route to an operational site — use 14.12. A compromised badge or camera system is in scope here only if its failure has a physical consequence; otherwise treat it as ordinary IT.
Two adversary populations share this space and they want opposite things. The first is criminal and indiscriminate: 119 ransomware groups impacted 3,300 industrial organizations in 2025, up 49% from 80 groups the year before (Dragos). Those actors are usually not in your OT at all. They encrypt IT, and you shut the process down yourself because you cannot run it blind. The second population is patient and state-directed and is not trying to make money. CISA, NSA, FBI and Five Eyes partners documented Volt Typhoon pre-positioning on the IT networks of communications, energy, transportation and water utilities to enable disruption of OT functions, using living-off-the-land techniques with minimal malware and dwell times of at least five years in some victims (CISA AA24-038A).
Five years. Not a typo. That is a tenant, not an intruder.
What changed in 2026 is intent moving from access to understanding. Dragos named three new groups — AZURITE, PYROXENE, SYLVANITE — alongside continued ELECTRUM, KAMACITE, VOLTZITE and BAUXITE activity, and reported that KAMACITE systematically mapped control loops across US infrastructure through 2025 while ELECTRUM targeted distributed energy systems in Poland with deliberate attempts to affect operational assets. VOLTZITE compromised Sierra Wireless AirLink gateways to reach US midstream pipeline operations before pivoting to engineering workstations. AZURITE targets engineering workstations for operational data and long-term access (Dragos). Stolen control-loop documentation is not data theft. It is the design phase of an attack that has not run yet.
And here is the mistake teams make, every time, and it is a good-faith mistake made by competent people: an IT responder sees adversary traffic crossing into a process network and does what they have been trained to do for fifteen years — isolates the segment. In IT, a wrong containment call costs you an afternoon. In OT, you have just slammed a moving vehicle into park because you found malware in the infotainment system. Control loops lose their supervisory layer mid-sequence, operators lose view of a process that is still running, and a plant that was safe becomes a plant nobody can see. Actionable takeaway: the entry condition for every OT containment action in this playbook is a named Engineering Authority on the bridge who agrees the action is safe in the current process state. Not consulted afterwards. On the bridge, before.
| Role | Responsibility in PB-OT |
|---|---|
| Incident Commander | Owns the incident, the timeline and the IT-side response. Owns no OT action. Escalates the shutdown question to the operating authority rather than answering it. |
| Engineering Authority (control systems engineer, named per site) | Co-signs every action touching Level 3 and below. Determines what is safe in the current process state. Holds a veto, and the veto is final. |
| Process Safety Lead | Independent of both. Confirms safety functions remain available and unmodified. Owns the call to move to a safe state on safety grounds alone. |
| Operations Lead (IT) | IT-side containment, identity, evidence export, the boundary itself. |
| Control-System Vendor Liaison | Single channel to the OEM and integrator for validated patches, firmware verification and known-good logic. Vendors do not get ad-hoc remote access during an incident. |
| Communications Lead | Operator, site, customer and regulator messaging; coordinates with the site's existing safety and environmental notification process. |
| Scribe | UTC/ISO 8601 timeline, chain of custody, and — specific to this scenario — a log of every physical action taken in the field, by whom. |
| Legal Liaison | Legal hold, regulator engagement, sector reporting obligations. |
| Executive Sponsor | Approves anything that stops production or affects customers or the public. |
Markings used below: `EVIDENCE destroys or degrades evidence — capture first. TIP-OFF is visible to the adversary. (S) **requires the Engineering Authority's sign-off and may not be executed by an IT responder alone.** Where a step carries (S)`, an IT responder executing it unilaterally is a reportable safety event regardless of outcome.
| # | Action | Who | Done when | Evidence to capture |
|---|---|---|---|---|
| 1 | Declare T+0 at SEV-1 and open a joint bridge. No OT-side action is authorized until the Engineering Authority and Process Safety Lead have joined. If neither is reachable in 15 minutes, escalate to the site operating authority via the plant's own out-of-hours callout, not via IT's. | IC | Both roles present and named in the log | Declaration time (UTC/ISO 8601), triggering detection ID, names and join times |
| 2 | Ask the control room, not the tools: is the process behaving as expected? Distinguish loss of view (indications unreliable, control intact) from loss of control (commands not taking effect). These are different incidents with different urgencies. | Engineering Authority | Written operator statement recorded | Operator statement verbatim, shift log extract, time of first anomaly noticed |
| 3 | Verify at least three critical indications against independent physical instruments — local gauges, field readings, a second sensor on a different path. Do not proceed on HMI data alone. | Engineering Authority | Independent readings recorded and compared | Photographs of local instruments with timestamps, HMI screenshot for the same moment |
| 4 | Score the location of observed activity on the modified Purdue scale CISA uses in NCISS — 0 unsuccessful, 1 business DMZ, 2 business network, 3 business network management, 4 critical system DMZ, 5 critical system management, 6 critical systems, 7 safety systems (NCISS). This sets severity and, more importantly, who decides what happens next. | IC + Engineering Authority | Level assigned and recorded | Level, the specific asset that justified it, assessor names |
| 5 | Freeze OT change. Halt scheduled maintenance, planned configuration pushes, firmware updates and integrator work at every affected site. This is free, reversible, and it stops your own people from overwriting evidence in the next hour. | Engineering Authority | Change freeze acknowledged by every site and integrator | Freeze notice, acknowledgement list, work orders suspended |
| 6 | Inventory every active remote-access session into OT — vendor VPN, cellular and serial gateways, jump hosts, dial-in. Inventory only. Do not terminate yet — you need to know what legitimate operations depend on before you cut. | Ops Lead (IT) | Complete session list with owner per session | Session records, source IPs, accounts, start times, business owner per session |
| 7 | Start passive capture at the IT/OT boundary and at the process-network core, on a SPAN/mirror port or a passive tap. Passive is not a preference here — the ACSC/CISA logging guidance notes that excessive logging can adversely affect memory- and processor-constrained embedded OT devices, and that where OT devices cannot log you should log the traffic to and from them instead (Best Practices for Event Logging and Threat Detection). | Ops Lead (IT) | Capture running on both segments, writing to rotating files | pcap files with hashes, capture start time, interface and tap point, capturing host |
| 8 | Export the IT-side logs with the shortest retention first — RFC 3227 puts remote logging and monitoring data above configuration and archival media in the order of volatility, and in practice this tier is both the most useful and the first to age out (RFC 3227). | Ops Lead (IT) | Raw exports in the evidence store under legal hold | File hashes, query windows, exporting identity, source system names |
| 9 | Baseline the engineering workstations: hash every project and logic file, list last-modified times, and pull the engineering suite's own download/upload history to controllers. This is the asset AZURITE and VOLTZITE go for. | Engineering Authority + Ops Lead (IT) | Hash manifest produced for every EWS at the site | Hash manifest, EWS hostnames, tool versions, upload/download log export |
| 10 | Compare each controller's running program and configuration against the offline known-good master. Not against a copy stored on the network you are investigating. (S) | Engineering Authority | Every in-scope controller dispositioned as match / mismatch / unverifiable | Comparison output, master copy provenance and date, controller identifiers |
| 11 | Physical walk-down: record the position of every controller mode switch (RUN / PROGRAM / REMOTE), key switch and local/remote selector. A controller left in a writable mode is both a finding and an exposure. | Engineering Authority | Walk-down sheet complete and signed | Signed walk-down sheet, photographs, time of walk-down, walker's name |
| 12 | Do not run active discovery, vulnerability scanning or credentialed enumeration against Level 2 and below. If you need asset data, take it from the passive capture and from engineering's documentation. Record this constraint in the timeline so nobody re-litigates it at hour six. | IC | Constraint recorded and communicated to all responders | Timeline entry, distribution record |
# Passive capture at the IT/OT boundary. Run on a host attached to a SPAN/mirror
# port or a passive tap - never inline, and never on a control-system host.
# -i: the mirror interface. -s 0: full packets. -w/-C/-W: rotating 200MB files.
sudo tcpdump -i <mirror_iface> -s 0 -w /evidence/ot-boundary.pcap -C 200 -W 200
# Chain of custody: hash closed files only. Stop the capture and confirm no file
# is still being written before you run this - the file tcpdump had open will not
# hash the same way twice, and a hash that fails verification is worse than none.
# Record the output with the operator's name and the UTC time it was taken.
sha256sum /evidence/ot-boundary.pcap* > /evidence/ot-boundary.sha256# Engineering workstation project-file baseline. Read-only.
# Compare this manifest against the offline master hash list, not against a
# network copy - a network copy is inside the blast radius you are investigating.
Get-ChildItem -Path '<project_root>' -Recurse -File |
Get-FileHash -Algorithm SHA256 |
Export-Csv -Path 'C:\evidence\ews-project-hashes.csv' -NoTypeInformationThe sequence here is the reverse of your instincts. Establish a safe, known process state first; contain IT at full speed; break the boundary next; work inward last. Starting at the process end — pulling a switch, blocking a protocol, isolating a segment — removes the supervisory layer from a process that is still physically running, and the people who then have to manage that process are the ones standing next to it.
| # | Action | Who | Done when | Evidence to capture |
|---|---|---|---|---|
| 1 | Decide and record the target process state: continue normal, continue under local/manual control, controlled shutdown, or emergency shutdown. The Incident Commander does not make this call. | Operating authority, on Engineering Authority and Process Safety Lead recommendation | State selected, recorded, communicated to every operator | Decision record, decider's name and role, time, stated rationale |
| 2 | Contain in IT without restraint. Isolate hosts, revoke sessions and tokens, block C2 at the enterprise egress. The IT estate is yours and speed is a virtue there. `TIP-OFF` | Ops Lead (IT) | IT-side containment actions complete | Action log, isolated host list, revocation records |
| 3 | Break the IT/OT boundary at the single pre-agreed, previously tested break point, in the documented manner. If your organization has never tested this break, do not improvise it during an incident — put people in the control room and cut remote access instead (step 4). `TIP-OFF (S)` | Ops Lead (IT) + Engineering Authority | Boundary severed, control room confirms process still under control | Change record, before/after topology, confirmation from the control room with time |
| 4 | Disable remote access into OT: vendor accounts, integrator accounts, cellular and serial gateways, dial-in modems. Kill the account and the path — an account disabled at the IdP does not close a cellular modem someone can reach directly. `TIP-OFF (S)` | Ops Lead (IT) + Engineering Authority | Every session from step 6 of Phase 1 dispositioned | Per-session disposition, account disable records, physical confirmation for gateways |
| 5 | Where the process design supports it, move critical loops to local or manual control with operators physically present and briefed. This is the OT equivalent of degraded-mode operation, and it is what buys you the freedom to work upstream. (S) | Engineering Authority | Loops in local control, staffing confirmed | Loop list, staffing roster, time of transfer, operator acknowledgements |
| 6 | Do not use the safety instrumented system as a containment lever, and do not test, bypass, modify or reconfigure it during response. Confirm it is available and unmodified; that is the whole of your interaction with it. (S) | Process Safety Lead | SIS confirmed available and unmodified, in writing | SIS integrity check record, checker's name, method used |
| 7 | Quarantine compromised engineering workstations — but image them first and stand up a known-good replacement built from offline media. An EWS is the highest-value asset in the environment and the one most likely to be reimaged in a panic. `EVIDENCE TIP-OFF` | Ops Lead (IT) + Engineering Authority | Image captured and hashed; replacement in service | Disk image hash, acquisition tool and version, acquiring operator, replacement build provenance |
| 8 | Move the response onto out-of-band communications: phone bridge and a channel that does not traverse the compromised estate. CISA's guidance is to isolate in a coordinated manner using out-of-band methods such as phone calls, to avoid tipping off actors (CISA — I've Been Hit By Ransomware). | IC | Every responder on the out-of-band channel | Channel details, participant list, switchover time |
| 9 | Verify the historian and any data-diode or one-way path is still flowing in the intended direction only. A "read-only" historian link that has been reconfigured is a control path. (S) | Engineering Authority | Direction verified at the device, not from documentation | Device configuration export, verification method, verifier's name |
| 10 | Fall back to paper: manual logs, printed procedures, physical rounds on a defined interval. Do this before you need it, not after the HMIs go dark. | Engineering Authority | Paper procedures issued and rounds scheduled | Procedure versions issued, round schedule, first completed round sheet |
| # | Action | Who | Done when | Evidence to capture |
|---|---|---|---|---|
| 1 | Plan a single coordinated remediation event rather than removing findings as you discover them. Piecemeal containment tips your hand and the adversary abandons the burned infrastructure while keeping the access you have not found (Aldridge, Remediating Targeted-threat Intrusions). In OT this matters doubly, because your change windows are scarce and you get very few of them. | IC + Engineering Authority | Remediation event scoped, scheduled and briefed | Event plan, scope list, scheduled window, approvals |
| 2 | Rebuild engineering workstations from offline installation media and vendor-supplied images. Not from a backup that lived on the network you are cleaning. | Ops Lead (IT) + Vendor Liaison | Every EWS rebuilt and validated by engineering | Build provenance, media hashes, validation sign-off |
| 3 | Restore controller logic and configuration from the verified offline master, with engineering comparing checksums before and after the download. (S) | Engineering Authority | Every mismatched controller restored and re-verified | Pre/post checksums, master provenance, download records, engineer's signature |
| 4 | Verify firmware on affected devices against the vendor's published hashes through the Vendor Liaison. Verify — do not assume, and do not accept a firmware image someone downloaded during the incident from a general-purpose workstation. | Vendor Liaison + Engineering Authority | Firmware verified or replaced on every in-scope device | Vendor hash reference, verification output, device serials |
| 5 | Rotate credentials across the boundary: OT domain accounts, jump-host accounts, vendor accounts, gateway and appliance credentials, and any shared engineering account. `TIP-OFF` | Ops Lead (IT) | Rotation complete and old credentials confirmed rejected | Rotation records, first rejected-auth events |
| 6 | For credentials that genuinely cannot be rotated — hardcoded device passwords, a protocol with no authentication, a vendor account the OEM will not change outside a service call — record each one as an accepted risk with a named accepting executive, a compensating control, and a review date. Do not let it disappear into "remediated." | Engineering Authority + Executive Sponsor | Every non-rotatable credential entered in the risk register | Register entries, compensating control, accepting role, review date |
| 7 | Patch only what the OEM has validated for your configuration, in a window engineering owns. A CVSS score does not open a maintenance window; the process schedule does. Where patching is not possible, the answer is compensating controls, not a deferred ticket that quietly ages. | Vendor Liaison + Engineering Authority | Each in-scope vulnerability either patched, mitigated or registered | Vendor validation statement, change record, compensating control description |
| 8 | Re-scope before declaring eradication complete. If new adversary activity appears, contain it and return to analysis until the full scope and the initial vector are identified (CISA Federal Playbooks). | IC | No new activity across a defined observation window | Observation window definition, monitoring coverage, findings |
| # | Action | Who | Done when | Evidence to capture |
|---|---|---|---|---|
| 1 | Recover outward-in: identity and IT first, then the OT DMZ, then Level 3 supervisory, then Level 2, then controllers, and only then process restart. Reconnecting a cleaned process network to an uncleaned upper layer re-infects it in the order you just worked so hard to reverse. | IC + Engineering Authority | Each tier validated clean before the next reconnects | Per-tier validation record, reconnection times, validating role |
| 2 | Restore view before control. Operators get trustworthy indications, alarms and historian data back before anyone hands them a live command path. | Engineering Authority | Indications verified against field instruments again | Comparison record, operator acceptance, time |
| 3 | Prove control before load: loop checks, bump tests and alarm verification per the site's own commissioning procedure. (S) | Engineering Authority | Commissioning checklist complete and signed | Signed checklist, test results, tester names |
| 4 | Restart the process using the site's documented start-up procedure. This is an engineering procedure and it is not modified for the convenience of the incident. (S) | Operating authority | Process at target state and stable | Start-up log, deviations recorded and approved |
| 5 | Run an elevated-monitoring watch period with a defined duration and defined exit criteria, with the passive capture still running at the boundary. | Ops Lead (IT) | Watch period completed with no findings | Watch period definition, monitoring coverage, findings log |
| 6 | Re-baseline: fresh hash manifests for every EWS project file, fresh offline master copies of all controller logic, and a refreshed asset inventory reflecting what you actually found. | Engineering Authority | New masters stored offline and verified | New manifests, storage location, verification record |
| 7 | Formal return to normal, signed jointly by the Incident Commander, the Engineering Authority and the Process Safety Lead. Three signatures, because three different questions were answered. | IC | All three signatures recorded | Signed return-to-normal record with times |
| # | Action | Who | Done when | Evidence to capture |
|---|---|---|---|---|
| 1 | Joint hotwash with security, engineering, operations, safety and the OEM/integrator in the same room. If engineering was not in the room during the incident, that is finding number one. | IC | Hotwash held, findings owned and dated | Findings list with owner and due date per item |
| 2 | Review the IT/OT boundary specifically: what crossed it, why, and whether the break point worked as documented. | Engineering Authority + Ops Lead (IT) | Boundary review complete | Review record, remediation items |
| 3 | Close the gap in what you could not answer. If you could not tell whether a controller's logic had changed, the deliverable is offline golden copies and a scheduled comparison, not a monitoring product. | Engineering Authority | Gap register produced with owners | Gap register, owners, dates |
| 4 | Rehearse the boundary break and the fall-back-to-manual procedure in the next exercise cycle. If the break in Phase 2 step 3 could not be used because it had never been tested, it is now the top exercise objective. See Chapter 18. | IC | Exercise scheduled with the objective written in | Exercise plan, objective text, scheduled date |
| 5 | Share indicators through your sector ISAC and with CISA, and feed the control-loop and engineering-workstation observations back — this is exactly the telemetry that made the 2026 threat-group picture possible in the first place. | Legal Liaison + IC | Submission made | Submission record, recipient, content shared |
| 6 | Update this playbook and record the test date in its header. A playbook that survived a real incident unamended is a playbook nobody consulted. | IC | Playbook updated and version incremented | Version, changes, date, next review date |
Two clock families run in parallel here and they answer to different regulators. The cyber clock: if you are a TSA-designated pipeline or rail owner-operator, the Security Directives require reporting cybersecurity incidents to CISA within 24 hours of identification — a clock that is live today and shorter than CIRCIA's 72 hours will be when that rule is final (TSA ratification notice, 17 Jan 2025; CISA CIRCIA). Australian critical-infrastructure operators are commonly cited as facing a 12-hour SOCI Part 2B critical-incident clock; that figure was not confirmed against a primary Home Affairs or cyber.gov.au source for this book — verify it before encoding (see the verification note in Chapter 15). Every deadline, recipient and template is in Chapter 15; do not reconstruct them here and never quote one from memory.
The second family is the one IT teams forget entirely. A physical process incident may trigger safety and environmental reporting obligations that have nothing to do with cyber law, that are often faster, and that the plant already knows how to file. A release, an unplanned shutdown, an injury or a bypassed protective function has its own regulator and its own form. Your job is not to learn that regime during the incident — it is to make sure the Communications Lead is talking to the person at the site who owns it, in the first hour, before those two workstreams file inconsistent accounts of the same event.
Automate freely — everything read-only. IT-side detection enrichment. Starting the passive capture at the boundary on a defined trigger. Exporting and hashing short-retention logs. Re-running the engineering workstation project-file hash manifest and diffing it against the offline master on a schedule, so that "did the logic change?" is a query rather than a two-day investigation. Refreshing the asset inventory from passive data. Alerting on any OT protocol traffic sourced from a host that is not an engineering workstation or an HMI. None of these write to anything, and the last two are the highest-value pre-positioned detections you can build for the KAMACITE and VOLTZITE patterns.
Automate behind a human gate. Disabling a vendor remote-access account, blocking a source at the enterprise egress, and isolating an IT-side host that sits adjacent to the boundary. Reversible, scoped, and a false positive costs a phone call. Gate them on a named approver who is reachable out of hours, and have the automation attach the artefact that justified the action.
Never automate — and design the system so it cannot. Anything that writes to a controller, changes a setpoint, isolates a process-network segment, restarts an OT asset, or interacts with a safety instrumented system. Make this structural rather than procedural: no SOAR platform, agent or service account should hold a credential capable of writing to Level 2 or below. If the automation physically cannot reach the control layer, the 03:00 mistake becomes impossible instead of merely forbidden. That is the single highest-value architectural decision in this playbook, it costs nothing but discipline, and small operators can implement it as easily as large ones. Chapter 17 has the gate design in full.