The 2026 InfoSec Playbook · Scenario playbooks

#14.14 OT and ICS Incident

Playbook ID: PB-OT | Default severity: SEV-1 (the only downgrade path is a written engineering finding that no control system, safety function or process-network asset is in scope — downgrade to SEV-2, never lower, and record who signed it) | Owner: Incident Commander, paired with a named Engineering Authority who co-signs every OT action

#When to run this

Not for: IT-only ransomware at an organization with no physical process — use 14.1. An exploited perimeter appliance with no route to an operational site — use 14.12. A compromised badge or camera system is in scope here only if its failure has a physical consequence; otherwise treat it as ordinary IT.

#What you are dealing with

Two adversary populations share this space and they want opposite things. The first is criminal and indiscriminate: 119 ransomware groups impacted 3,300 industrial organizations in 2025, up 49% from 80 groups the year before (Dragos). Those actors are usually not in your OT at all. They encrypt IT, and you shut the process down yourself because you cannot run it blind. The second population is patient and state-directed and is not trying to make money. CISA, NSA, FBI and Five Eyes partners documented Volt Typhoon pre-positioning on the IT networks of communications, energy, transportation and water utilities to enable disruption of OT functions, using living-off-the-land techniques with minimal malware and dwell times of at least five years in some victims (CISA AA24-038A).

Five years. Not a typo. That is a tenant, not an intruder.

What changed in 2026 is intent moving from access to understanding. Dragos named three new groups — AZURITE, PYROXENE, SYLVANITE — alongside continued ELECTRUM, KAMACITE, VOLTZITE and BAUXITE activity, and reported that KAMACITE systematically mapped control loops across US infrastructure through 2025 while ELECTRUM targeted distributed energy systems in Poland with deliberate attempts to affect operational assets. VOLTZITE compromised Sierra Wireless AirLink gateways to reach US midstream pipeline operations before pivoting to engineering workstations. AZURITE targets engineering workstations for operational data and long-term access (Dragos). Stolen control-loop documentation is not data theft. It is the design phase of an attack that has not run yet.

And here is the mistake teams make, every time, and it is a good-faith mistake made by competent people: an IT responder sees adversary traffic crossing into a process network and does what they have been trained to do for fifteen years — isolates the segment. In IT, a wrong containment call costs you an afternoon. In OT, you have just slammed a moving vehicle into park because you found malware in the infotainment system. Control loops lose their supervisory layer mid-sequence, operators lose view of a process that is still running, and a plant that was safe becomes a plant nobody can see. Actionable takeaway: the entry condition for every OT containment action in this playbook is a named Engineering Authority on the bridge who agrees the action is safe in the current process state. Not consulted afterwards. On the bridge, before.

#Roles for this incident

RoleResponsibility in PB-OT
Incident CommanderOwns the incident, the timeline and the IT-side response. Owns no OT action. Escalates the shutdown question to the operating authority rather than answering it.
Engineering Authority (control systems engineer, named per site)Co-signs every action touching Level 3 and below. Determines what is safe in the current process state. Holds a veto, and the veto is final.
Process Safety LeadIndependent of both. Confirms safety functions remain available and unmodified. Owns the call to move to a safe state on safety grounds alone.
Operations Lead (IT)IT-side containment, identity, evidence export, the boundary itself.
Control-System Vendor LiaisonSingle channel to the OEM and integrator for validated patches, firmware verification and known-good logic. Vendors do not get ad-hoc remote access during an incident.
Communications LeadOperator, site, customer and regulator messaging; coordinates with the site's existing safety and environmental notification process.
ScribeUTC/ISO 8601 timeline, chain of custody, and — specific to this scenario — a log of every physical action taken in the field, by whom.
Legal LiaisonLegal hold, regulator engagement, sector reporting obligations.
Executive SponsorApproves anything that stops production or affects customers or the public.

Markings used below: `EVIDENCE destroys or degrades evidence — capture first. TIP-OFF is visible to the adversary. (S) **requires the Engineering Authority's sign-off and may not be executed by an IT responder alone.** Where a step carries (S)`, an IT responder executing it unilaterally is a reportable safety event regardless of outcome.

#Phase 1 — Detection and Triage

#ActionWhoDone whenEvidence to capture
1Declare T+0 at SEV-1 and open a joint bridge. No OT-side action is authorized until the Engineering Authority and Process Safety Lead have joined. If neither is reachable in 15 minutes, escalate to the site operating authority via the plant's own out-of-hours callout, not via IT's.ICBoth roles present and named in the logDeclaration time (UTC/ISO 8601), triggering detection ID, names and join times
2Ask the control room, not the tools: is the process behaving as expected? Distinguish loss of view (indications unreliable, control intact) from loss of control (commands not taking effect). These are different incidents with different urgencies.Engineering AuthorityWritten operator statement recordedOperator statement verbatim, shift log extract, time of first anomaly noticed
3Verify at least three critical indications against independent physical instruments — local gauges, field readings, a second sensor on a different path. Do not proceed on HMI data alone.Engineering AuthorityIndependent readings recorded and comparedPhotographs of local instruments with timestamps, HMI screenshot for the same moment
4Score the location of observed activity on the modified Purdue scale CISA uses in NCISS — 0 unsuccessful, 1 business DMZ, 2 business network, 3 business network management, 4 critical system DMZ, 5 critical system management, 6 critical systems, 7 safety systems (NCISS). This sets severity and, more importantly, who decides what happens next.IC + Engineering AuthorityLevel assigned and recordedLevel, the specific asset that justified it, assessor names
5Freeze OT change. Halt scheduled maintenance, planned configuration pushes, firmware updates and integrator work at every affected site. This is free, reversible, and it stops your own people from overwriting evidence in the next hour.Engineering AuthorityChange freeze acknowledged by every site and integratorFreeze notice, acknowledgement list, work orders suspended
6Inventory every active remote-access session into OT — vendor VPN, cellular and serial gateways, jump hosts, dial-in. Inventory only. Do not terminate yet — you need to know what legitimate operations depend on before you cut.Ops Lead (IT)Complete session list with owner per sessionSession records, source IPs, accounts, start times, business owner per session
7Start passive capture at the IT/OT boundary and at the process-network core, on a SPAN/mirror port or a passive tap. Passive is not a preference here — the ACSC/CISA logging guidance notes that excessive logging can adversely affect memory- and processor-constrained embedded OT devices, and that where OT devices cannot log you should log the traffic to and from them instead (Best Practices for Event Logging and Threat Detection).Ops Lead (IT)Capture running on both segments, writing to rotating filespcap files with hashes, capture start time, interface and tap point, capturing host
8Export the IT-side logs with the shortest retention first — RFC 3227 puts remote logging and monitoring data above configuration and archival media in the order of volatility, and in practice this tier is both the most useful and the first to age out (RFC 3227).Ops Lead (IT)Raw exports in the evidence store under legal holdFile hashes, query windows, exporting identity, source system names
9Baseline the engineering workstations: hash every project and logic file, list last-modified times, and pull the engineering suite's own download/upload history to controllers. This is the asset AZURITE and VOLTZITE go for.Engineering Authority + Ops Lead (IT)Hash manifest produced for every EWS at the siteHash manifest, EWS hostnames, tool versions, upload/download log export
10Compare each controller's running program and configuration against the offline known-good master. Not against a copy stored on the network you are investigating. (S)Engineering AuthorityEvery in-scope controller dispositioned as match / mismatch / unverifiableComparison output, master copy provenance and date, controller identifiers
11Physical walk-down: record the position of every controller mode switch (RUN / PROGRAM / REMOTE), key switch and local/remote selector. A controller left in a writable mode is both a finding and an exposure.Engineering AuthorityWalk-down sheet complete and signedSigned walk-down sheet, photographs, time of walk-down, walker's name
12Do not run active discovery, vulnerability scanning or credentialed enumeration against Level 2 and below. If you need asset data, take it from the passive capture and from engineering's documentation. Record this constraint in the timeline so nobody re-litigates it at hour six.ICConstraint recorded and communicated to all respondersTimeline entry, distribution record
shell
# Passive capture at the IT/OT boundary. Run on a host attached to a SPAN/mirror
# port or a passive tap - never inline, and never on a control-system host.
# -i: the mirror interface. -s 0: full packets. -w/-C/-W: rotating 200MB files.
sudo tcpdump -i <mirror_iface> -s 0 -w /evidence/ot-boundary.pcap -C 200 -W 200

# Chain of custody: hash closed files only. Stop the capture and confirm no file
# is still being written before you run this - the file tcpdump had open will not
# hash the same way twice, and a hash that fails verification is worse than none.
# Record the output with the operator's name and the UTC time it was taken.
sha256sum /evidence/ot-boundary.pcap* > /evidence/ot-boundary.sha256
PowerShell
# Engineering workstation project-file baseline. Read-only.
# Compare this manifest against the offline master hash list, not against a
# network copy - a network copy is inside the blast radius you are investigating.
Get-ChildItem -Path '<project_root>' -Recurse -File |
  Get-FileHash -Algorithm SHA256 |
  Export-Csv -Path 'C:\evidence\ews-project-hashes.csv' -NoTypeInformation

#Phase 2 — Containment

The sequence here is the reverse of your instincts. Establish a safe, known process state first; contain IT at full speed; break the boundary next; work inward last. Starting at the process end — pulling a switch, blocking a protocol, isolating a segment — removes the supervisory layer from a process that is still physically running, and the people who then have to manage that process are the ones standing next to it.

#ActionWhoDone whenEvidence to capture
1Decide and record the target process state: continue normal, continue under local/manual control, controlled shutdown, or emergency shutdown. The Incident Commander does not make this call.Operating authority, on Engineering Authority and Process Safety Lead recommendationState selected, recorded, communicated to every operatorDecision record, decider's name and role, time, stated rationale
2Contain in IT without restraint. Isolate hosts, revoke sessions and tokens, block C2 at the enterprise egress. The IT estate is yours and speed is a virtue there. `TIP-OFF`Ops Lead (IT)IT-side containment actions completeAction log, isolated host list, revocation records
3Break the IT/OT boundary at the single pre-agreed, previously tested break point, in the documented manner. If your organization has never tested this break, do not improvise it during an incident — put people in the control room and cut remote access instead (step 4). `TIP-OFF (S)`Ops Lead (IT) + Engineering AuthorityBoundary severed, control room confirms process still under controlChange record, before/after topology, confirmation from the control room with time
4Disable remote access into OT: vendor accounts, integrator accounts, cellular and serial gateways, dial-in modems. Kill the account and the path — an account disabled at the IdP does not close a cellular modem someone can reach directly. `TIP-OFF (S)`Ops Lead (IT) + Engineering AuthorityEvery session from step 6 of Phase 1 dispositionedPer-session disposition, account disable records, physical confirmation for gateways
5Where the process design supports it, move critical loops to local or manual control with operators physically present and briefed. This is the OT equivalent of degraded-mode operation, and it is what buys you the freedom to work upstream. (S)Engineering AuthorityLoops in local control, staffing confirmedLoop list, staffing roster, time of transfer, operator acknowledgements
6Do not use the safety instrumented system as a containment lever, and do not test, bypass, modify or reconfigure it during response. Confirm it is available and unmodified; that is the whole of your interaction with it. (S)Process Safety LeadSIS confirmed available and unmodified, in writingSIS integrity check record, checker's name, method used
7Quarantine compromised engineering workstations — but image them first and stand up a known-good replacement built from offline media. An EWS is the highest-value asset in the environment and the one most likely to be reimaged in a panic. `EVIDENCE TIP-OFF`Ops Lead (IT) + Engineering AuthorityImage captured and hashed; replacement in serviceDisk image hash, acquisition tool and version, acquiring operator, replacement build provenance
8Move the response onto out-of-band communications: phone bridge and a channel that does not traverse the compromised estate. CISA's guidance is to isolate in a coordinated manner using out-of-band methods such as phone calls, to avoid tipping off actors (CISA — I've Been Hit By Ransomware).ICEvery responder on the out-of-band channelChannel details, participant list, switchover time
9Verify the historian and any data-diode or one-way path is still flowing in the intended direction only. A "read-only" historian link that has been reconfigured is a control path. (S)Engineering AuthorityDirection verified at the device, not from documentationDevice configuration export, verification method, verifier's name
10Fall back to paper: manual logs, printed procedures, physical rounds on a defined interval. Do this before you need it, not after the HMIs go dark.Engineering AuthorityPaper procedures issued and rounds scheduledProcedure versions issued, round schedule, first completed round sheet

#Phase 3 — Eradication

#ActionWhoDone whenEvidence to capture
1Plan a single coordinated remediation event rather than removing findings as you discover them. Piecemeal containment tips your hand and the adversary abandons the burned infrastructure while keeping the access you have not found (Aldridge, Remediating Targeted-threat Intrusions). In OT this matters doubly, because your change windows are scarce and you get very few of them.IC + Engineering AuthorityRemediation event scoped, scheduled and briefedEvent plan, scope list, scheduled window, approvals
2Rebuild engineering workstations from offline installation media and vendor-supplied images. Not from a backup that lived on the network you are cleaning.Ops Lead (IT) + Vendor LiaisonEvery EWS rebuilt and validated by engineeringBuild provenance, media hashes, validation sign-off
3Restore controller logic and configuration from the verified offline master, with engineering comparing checksums before and after the download. (S)Engineering AuthorityEvery mismatched controller restored and re-verifiedPre/post checksums, master provenance, download records, engineer's signature
4Verify firmware on affected devices against the vendor's published hashes through the Vendor Liaison. Verify — do not assume, and do not accept a firmware image someone downloaded during the incident from a general-purpose workstation.Vendor Liaison + Engineering AuthorityFirmware verified or replaced on every in-scope deviceVendor hash reference, verification output, device serials
5Rotate credentials across the boundary: OT domain accounts, jump-host accounts, vendor accounts, gateway and appliance credentials, and any shared engineering account. `TIP-OFF`Ops Lead (IT)Rotation complete and old credentials confirmed rejectedRotation records, first rejected-auth events
6For credentials that genuinely cannot be rotated — hardcoded device passwords, a protocol with no authentication, a vendor account the OEM will not change outside a service call — record each one as an accepted risk with a named accepting executive, a compensating control, and a review date. Do not let it disappear into "remediated."Engineering Authority + Executive SponsorEvery non-rotatable credential entered in the risk registerRegister entries, compensating control, accepting role, review date
7Patch only what the OEM has validated for your configuration, in a window engineering owns. A CVSS score does not open a maintenance window; the process schedule does. Where patching is not possible, the answer is compensating controls, not a deferred ticket that quietly ages.Vendor Liaison + Engineering AuthorityEach in-scope vulnerability either patched, mitigated or registeredVendor validation statement, change record, compensating control description
8Re-scope before declaring eradication complete. If new adversary activity appears, contain it and return to analysis until the full scope and the initial vector are identified (CISA Federal Playbooks).ICNo new activity across a defined observation windowObservation window definition, monitoring coverage, findings

#Phase 4 — Recovery

#ActionWhoDone whenEvidence to capture
1Recover outward-in: identity and IT first, then the OT DMZ, then Level 3 supervisory, then Level 2, then controllers, and only then process restart. Reconnecting a cleaned process network to an uncleaned upper layer re-infects it in the order you just worked so hard to reverse.IC + Engineering AuthorityEach tier validated clean before the next reconnectsPer-tier validation record, reconnection times, validating role
2Restore view before control. Operators get trustworthy indications, alarms and historian data back before anyone hands them a live command path.Engineering AuthorityIndications verified against field instruments againComparison record, operator acceptance, time
3Prove control before load: loop checks, bump tests and alarm verification per the site's own commissioning procedure. (S)Engineering AuthorityCommissioning checklist complete and signedSigned checklist, test results, tester names
4Restart the process using the site's documented start-up procedure. This is an engineering procedure and it is not modified for the convenience of the incident. (S)Operating authorityProcess at target state and stableStart-up log, deviations recorded and approved
5Run an elevated-monitoring watch period with a defined duration and defined exit criteria, with the passive capture still running at the boundary.Ops Lead (IT)Watch period completed with no findingsWatch period definition, monitoring coverage, findings log
6Re-baseline: fresh hash manifests for every EWS project file, fresh offline master copies of all controller logic, and a refreshed asset inventory reflecting what you actually found.Engineering AuthorityNew masters stored offline and verifiedNew manifests, storage location, verification record
7Formal return to normal, signed jointly by the Incident Commander, the Engineering Authority and the Process Safety Lead. Three signatures, because three different questions were answered.ICAll three signatures recordedSigned return-to-normal record with times

#Phase 5 — Post-Incident

#ActionWhoDone whenEvidence to capture
1Joint hotwash with security, engineering, operations, safety and the OEM/integrator in the same room. If engineering was not in the room during the incident, that is finding number one.ICHotwash held, findings owned and datedFindings list with owner and due date per item
2Review the IT/OT boundary specifically: what crossed it, why, and whether the break point worked as documented.Engineering Authority + Ops Lead (IT)Boundary review completeReview record, remediation items
3Close the gap in what you could not answer. If you could not tell whether a controller's logic had changed, the deliverable is offline golden copies and a scheduled comparison, not a monitoring product.Engineering AuthorityGap register produced with ownersGap register, owners, dates
4Rehearse the boundary break and the fall-back-to-manual procedure in the next exercise cycle. If the break in Phase 2 step 3 could not be used because it had never been tested, it is now the top exercise objective. See Chapter 18.ICExercise scheduled with the objective written inExercise plan, objective text, scheduled date
5Share indicators through your sector ISAC and with CISA, and feed the control-loop and engineering-workstation observations back — this is exactly the telemetry that made the 2026 threat-group picture possible in the first place.Legal Liaison + ICSubmission madeSubmission record, recipient, content shared
6Update this playbook and record the test date in its header. A playbook that survived a real incident unamended is a playbook nobody consulted.ICPlaybook updated and version incrementedVersion, changes, date, next review date

#Decision points

#Communications and notification triggers

Two clock families run in parallel here and they answer to different regulators. The cyber clock: if you are a TSA-designated pipeline or rail owner-operator, the Security Directives require reporting cybersecurity incidents to CISA within 24 hours of identification — a clock that is live today and shorter than CIRCIA's 72 hours will be when that rule is final (TSA ratification notice, 17 Jan 2025; CISA CIRCIA). Australian critical-infrastructure operators are commonly cited as facing a 12-hour SOCI Part 2B critical-incident clock; that figure was not confirmed against a primary Home Affairs or cyber.gov.au source for this book — verify it before encoding (see the verification note in Chapter 15). Every deadline, recipient and template is in Chapter 15; do not reconstruct them here and never quote one from memory.

The second family is the one IT teams forget entirely. A physical process incident may trigger safety and environmental reporting obligations that have nothing to do with cyber law, that are often faster, and that the plant already knows how to file. A release, an unplanned shutdown, an injury or a bypassed protective function has its own regulator and its own form. Your job is not to learn that regime during the incident — it is to make sure the Communications Lead is talking to the person at the site who owns it, in the first hour, before those two workstreams file inconsistent accounts of the same event.

#Automation notes

Automate freely — everything read-only. IT-side detection enrichment. Starting the passive capture at the boundary on a defined trigger. Exporting and hashing short-retention logs. Re-running the engineering workstation project-file hash manifest and diffing it against the offline master on a schedule, so that "did the logic change?" is a query rather than a two-day investigation. Refreshing the asset inventory from passive data. Alerting on any OT protocol traffic sourced from a host that is not an engineering workstation or an HMI. None of these write to anything, and the last two are the highest-value pre-positioned detections you can build for the KAMACITE and VOLTZITE patterns.

Automate behind a human gate. Disabling a vendor remote-access account, blocking a source at the enterprise egress, and isolating an IT-side host that sits adjacent to the boundary. Reversible, scoped, and a false positive costs a phone call. Gate them on a named approver who is reachable out of hours, and have the automation attach the artefact that justified the action.

Never automate — and design the system so it cannot. Anything that writes to a controller, changes a setpoint, isolates a process-network segment, restarts an OT asset, or interacts with a safety instrumented system. Make this structural rather than procedural: no SOAR platform, agent or service account should hold a credential capable of writing to Level 2 or below. If the automation physically cannot reach the control layer, the 03:00 mistake becomes impossible instead of merely forbidden. That is the single highest-value architectural decision in this playbook, it costs nothing but discipline, and small operators can implement it as easily as large ones. Chapter 17 has the gate design in full.

#Pitfalls

This is one of the fourteen scenario playbooks in The 2026 InfoSec Playbook, a free field manual by Daniel Ramos. Written so somebody who has never read the book can pick it up mid-incident and run it. See all fourteen. Free, in full, no email wall.