How to design, run, score and close out the exercises that turn a plausible-looking playbook into one you know works — for the price of a conference room and three hours, not a seven-figure tool.
Who needs this: CISO, IR lead, SOC lead, detection engineers, exercise facilitators, executive sponsors, Legal, Comms, HR | Read time: 28 min | Maps to: CSF 2.0 IDENTIFY (ID.IM-01, ID.IM-02, ID.IM-03), GOVERN (GV.SC-08, GV.RR), PROTECT (PR.AT) | CIS v8.1 Controls 11, 14, 17, 18 | ISO/IEC 27001:2022 A.5.24, A.5.27, A.5.30, A.6.3
Welcome back, cyber-survivors. This is the chapter where we stop writing the plan and start finding out whether it is fiction.
Start with the failure that should be tattooed on the inside of every exercise planner's eyelids. When Equifax patched Apache Struts across the company, the notice telling people to do it went out to a distribution list — and, in GAO's words, "the recipient list for the notice was out-of-date and, as a result, the notice was not received by the individuals who would have been responsible for installing the necessary patch" (GAO-18-559). A stale mailing list. Not a zero-day, not a nation-state capability. A list nobody had ever sent a test message down.
That is the argument of this chapter. The controls that fail in real incidents are overwhelmingly the ones nobody exercised, and they fail in ways that are embarrassingly cheap to have discovered in advance. The British Library, website and intranet down, ended up running its response over social media and WhatsApp cascades (British Library, Learning Lessons from the Cyber-Attack). The Cyber Safety Review Board found Microsoft had stopped its manual rotation of consumer signing keys in 2021 after a cloud outage linked to the rotation process itself — a safety control abandoned precisely because using it hurt (CSRB). Every one is a testable proposition that went untested until an adversary tested it for free.
I have never once seen a tabletop exercise fail because it was too realistic. I have seen plenty fail because they were too comfortable — a scenario read from a slide, six people agreeing they would probably do the right thing, a document that changed nothing. This chapter is about the other kind, the one that ends with nine owned, dated findings and a playbook diff. Chapter 13 owns the lifecycle, the roles and the real-incident hotwash; Chapter 2 owns playbook metadata and the last_exercised field this chapter exists to populate; Appendix F carries the full scenario card set.
NIST SP 800-84 is still the canonical methodology, and the first thing it does is take a vocabulary away from you: "the term 'test' is reserved for testing systems or system components; it is not used to describe 'exercising' plans" (NIST SP 800-84).
That sounds like pedantry until you notice what it buys. A test produces a number: did the call-tree cascade complete inside the prescribed time limit, and in how many minutes. An exercise produces a judgement: did people make defensible decisions with the information they had. You need both, they cost wildly different amounts, and conflating them is how a program ends up with an annual tabletop, no measured restore time, and a contact list last verified in 2023.
SP 800-84 names four categories — tests, training, tabletop exercises, and functional exercises — and a four-phase event methodology every rung below inherits: Design → Develop → Conduct → Evaluate.
| Rung | Type | Typical cost | What it proves | What it cannot prove |
|---|---|---|---|---|
| 1 | Seminar | 1 hour, no prep | People know the plan exists and where it lives | Anything about behavior under pressure |
| 2 | Workshop | Half a day, moderate prep | The plan's gaps as authors see them; produces documents | That the plan survives contact |
| 3 | Tabletop (TTX) | 2–8 hours, 2–4 weeks prep | Decision quality, authority clarity, escalation paths, comms | That the tooling works or the timings are achievable |
| 4 | Drill / test | 1–4 hours, narrow scope | One capability, quantitatively — restore time, cascade time | Coordination across teams |
| 5 | Functional exercise | 1–2 days, 6–12 weeks prep | Real tools, real consoles, simulated events, measured timings | Physical/full-business disruption effects |
| 6 | Full-scale | Multi-day, months of prep | End-to-end organizational response including third parties | Nothing much more — but it costs accordingly |
Rungs 1–3 are discussion-based: people talk, nothing in production moves. Rungs 4–6 are operations-based: something happens, and something can break.
SP 800-84 adds two rules most programs violate immediately. First: separate senior-level and operational-level exercises before you combine them. "Senior-level teams and operational-level teams should participate in separate tabletop exercises initially because of their different levels of responsibility. Once these two groups have been exercised individually, both groups should participate in a combined exercise to validate coordination between the groups." Put the CFO in the same first session as the detection engineers and one of two things happens: the engineers stay quiet, or the CFO stops coming. Second: duration — 2–4 hours senior-level, 2–8 hours operational-level, with anything over four hours paired with a training session, because past that point you are teaching, not testing.
The off-the-shelf starting kit costs nothing. CISA's Tabletop Exercise Packages (CTEP) ship 100+ sample Situation Manuals organized by threat vector, plus an Exercise Brief deck, an Exercise Planner Handbook, and a Facilitator and Evaluator Handbook telling evaluators how to capture "strengths, areas for improvement, and recommendations" for the After-Action Report / Improvement Plan (CISA CTEP). CTEP delivers scenarios as time-phased injects rather than dumping the whole story at once, which is the most important design choice in the format. In the UK, NCSC's Exercise in a Box offers around 20 free exercises across a dozen-plus topics in micro, tabletop and simulation formats (NCSC), and NCSC now runs an assured-provider Cyber Incident Exercising scheme for those wanting a vetted external facilitator (NCSC).
The compliance floor, if you need one to get the invite accepted: SP 800-53 IR-3 requires testing the IR capability at an organization-defined frequency, noting that testing "includ[es] the use of checklists, walk-through or tabletop exercises, and simulations" (IR-3); IR-8 requires the plan be reviewed and approved at a defined frequency (IR-8); and SP 800-84 records 800-53's baseline expectation that federal agencies exercise or test contingency and IR capabilities at least annually. SP 800-61r3 rates ID.IM-02 — "Improvements are identified from security tests and exercises, including those done in coordination with suppliers and relevant third parties" — as High priority, tying supplier participation in through GV.SC-08 (NIST SP 800-61r3).
Actionable takeaway: Write the ladder into your plan with a named frequency per rung, then check which rungs you climbed in the last twelve months. Most programs find they did rung 3 once and rungs 4–6 never. The gap between "we do tabletops" and "we have measured our restore time" is where incidents live.
Almost everyone does this backwards. Someone reads a threat report, gets excited about a scenario, writes three pages of narrative, and then bolts on some objectives the scenario happens to touch. The result tests whatever the story author found interesting, which is not the same as whatever your program is weakest at.
SP 800-84's design sequence is the opposite order: determine the topic → determine the scope → identify objectives → identify participants → identify exercise staff → coordinate logistics. Objectives come third, before you have written a single word of story, and everything downstream is derived from them.
The second rule sounds wrong and is right. Keep the scenario short. SP 800-84: "A common misconception is that scenarios must be very detailed to be effective. Actually, it is often more effective to develop a short, concise scenario," because with long scenarios "participants often spend more time dissecting the scenario… than they spend on meeting the objectives." A detailed scenario invites the room to argue with your fiction. Every minute spent debating whether the EDR would really have missed that is a minute not spent finding out who is allowed to shut down production.
An objective is testable when a data collector who has never met your team can mark it true or false from what they observed. That means a role, an action, a threshold and a clock.
| Weak objective | Testable objective |
|---|---|
| "Test our incident response process" | The incident is declared at a stated severity by the on-call Incident Commander within 30 minutes of the first credible indicator |
| "Improve communication" | An approved holding statement exists within 60 scenario-minutes of the first external media enquiry, approved by the named authority in the plan |
| "Validate our backup strategy" | Backup, identity and hypervisor plane integrity is verified as part of triage, before any recovery decision is discussed |
| "Ensure leadership is engaged" | Every authority required by the playbook is exercised by its named holder or a named deputy; the number of decisions that stall awaiting an absent authority is zero |
Five to eight objectives is right for a three-hour operational tabletop. More than eight and your data collector cannot watch them all.
Once objectives exist, write the injects — the pre-scripted messages that force the decisions your objectives measure. SP 800-84 defines an inject as "a pre-scripted message that will be provided to participants during the course of an exercise," and specifies what each must carry: time, to whom, from whom, delivery means, and the message text. The chronological list of them is the Master Scenario Events List (MSEL) — "a chronologically sequenced outline of the simulated events and key event descriptions that participants will be asked to respond to," including expected actions and the objectives each event maps to. The MSEL is for exercise staff only; participants never see it.
Volume guidance from the same source is deliberately vague and deliberately correct: enough injects "to keep participants adequately occupied but… not be so many that participants will become overwhelmed." My working ratio, offered as judgement and not as NIST doctrine, is one inject per fifteen to twenty minutes of exercise clock, plus two spares in the facilitator's pocket for when the room solves something faster than expected.
Only now do you write the scenario narrative, and only enough to make the injects land. Then the artefact people skip, which is what turns scoring into vibes: evaluation criteria must be written before the exercise, "to ensure data collectors know what type of information to capture during the exercise."
Printed and in the room: the participant briefing; a Facilitator Guide (purpose, scope, objectives, scenario narrative, full question list, and a copy of the plan); a Participant Guide (the same, minus the question list); and the After-Action Report template, already populated with your objectives.
Actionable takeaway: Before your next exercise, write the objectives and the scoring sheet and get both signed off by the plan owner — with the scenario still unwritten. If you cannot get sign-off on the objectives, the scenario was never going to save you.
A complete operational-level tabletop you can run in three hours, grounded in what 2026 intrusions actually look like: identity-first entry, near-instant hand-off, deliberate targeting of the recovery path, extortion without encryption. Appendix F carries this and the rest of the scenario card set in printable form.
Profile assumed: roughly 400 staff, cloud identity provider with SSO, one on-premises file estate, an outsourced service desk, a cyber insurance policy, an IR retainer. Adjust the nouns, keep the decisions.
Scenario, in full — this is everything participants get up front: A routine quality review of service desk tickets has flagged a password reset and MFA re-enrolment completed six hours ago for a finance manager, requested by phone. The requester's identity was verified using employee number and date of birth. It is 09:40 on a Thursday.
Everything else arrives as an inject.
| ID | Objective (scored) |
|---|---|
| OBJ-1 | The incident is declared at a stated severity by the on-call Incident Commander within 30 minutes of inject 1, using the plan's declaration criteria |
| OBJ-2 | Identity containment is sequenced correctly — token and session revocation before credential reset, with OAuth grants enumerated — and the approving authority is named without debate |
| OBJ-3 | A documented isolate-or-observe decision is reached within 30 minutes of the scope changing, at the authority the plan specifies |
| OBJ-4 | Backup, identity and virtualization plane integrity is verified during triage, not deferred to recovery |
| OBJ-5 | An approved holding statement exists within 60 scenario-minutes of the first external enquiry |
| OBJ-6 | An out-of-band bridge, reachable without the primary identity provider, is established within 15 minutes of the identity plane being declared suspect |
| OBJ-7 | The extortion demand is routed into the payment-decision workflow with counsel and sanctions screening engaged, and no payment decision is taken on the response bridge |
| OBJ-8 | Every authority whose primary holder is unreachable has an accountable deputy identified within 20 minutes; stalled decisions are counted |
| # | Clock | Inject (to whom, by what means) | Decision it forces | Obj |
|---|---|---|---|---|
| 1 | 0:10 | Service desk QA report, emailed to SOC lead: password reset + MFA re-enrolment to a new device, completed by phone six hours ago | Is this an incident? Who declares, at what severity? | 1 |
| 2 | 0:25 | SIEM alert, to duty analyst: the same account authenticated from an unfamiliar network, created a mailbox rule, and granted OAuth consent to a third-party application with broad mail and file read scopes | Contain now or observe? Which containment action first? Who approves? | 2 |
| 3 | 0:40 | EDR alert, to Operations Lead: a legitimate remote-management tool was installed on a production file server 20 minutes after the token was first used | Is this still account compromise or is it an intrusion? Who can isolate a production server? | 3 |
| 4 | 1:00 | Phone call from the backup administrator to the IC: backup jobs failing since 03:00; a retention policy was modified by a service account overnight | Do we trust the recovery path? Do we isolate the backup network now? | 4 |
| 5 | 1:10 | Email to the general enquiries mailbox: a journalist asks for comment on "a security incident at your company" | Who speaks? What do we say? Does responding tip off the adversary? | 5 |
| 6 | 1:20 | Facilitator announcement: the identity provider is now considered suspect. Your incident bridge authenticates through it | Can you convene without SSO? Who has the out-of-band details? | 6 |
| 7 | 1:40 | Extortion note delivered to three executives' personal email: no encryption, 40 GB claimed including HR and contract data, 96-hour deadline, threats to notify your regulator and three named customers | Who owns the payment decision? What is the legal workflow? What clocks just started? | 7 |
| 8 | 1:55 | Insurer's breach response line: their panel requires an approved forensics provider. Your retained firm is not on the panel | Which contract governs? Who resolves it, and by when? | 7, 8 |
| 9 | 2:10 | Facilitator announcement: the only holder of break-glass credentials for the identity tenant is on annual leave, phone off. The CFO is airborne for four hours | Who deputises for each authority? How is that recorded? | 8 |
Two facilitator notes. Inject 6 produces the highest-value finding in almost every organization I have seen run something like it, and it is a pure announcement — no story required. Inject 8 exists because contract collisions are discovered at the worst possible moment and are trivially fixable in peacetime.
The threat model is not invented. Help desk impersonation to obtain password resets and MFA token transfers to attacker-controlled devices is documented TTP, and CISA explicitly notes that the presence of legitimate remote-management tools is not on its own malicious (CISA AA23-320A). The compressed timeline reflects Mandiant's finding that the median hand-off from initial-access broker to the operator who does the damage is now 22 seconds, down from over eight hours in 2022, and that operators increasingly target backup infrastructure, identity services and virtualization management planes — attacking your ability to recover rather than only your ability to operate (M-Trends 2026). Before the room debates payment in inject 7, it is worth knowing Coveware measured the Q2 2026 payment rate for data-exfiltration-only cases at 15% (Coveware by Veeam). The regulatory threat in that inject has precedent: ALPHV/BlackCat filed an SEC complaint against a victim for failing to disclose the breach ALPHV itself had caused.
Actionable takeaway: Run this as written next quarter, printed plan on the table, laptops closed. Then swap injects 7 and 8 for a supplier-breach pair and run it again the following quarter with your top vendor in the room — ID.IM-02 explicitly contemplates exercises "done in coordination with suppliers and relevant third parties."
A tabletop is a facilitated conversation with a scoring rubric attached. SP 800-84 specifies two staff roles as the minimum: a facilitator who leads the discussion and a data collector who records observations. Both must be thoroughly familiar with the plan and objectives, and both should meet beforehand and review previous exercises' results. One person cannot do both — facilitating takes all of your attention, and if you are also writing you record only what you already expected.
Open with the no-fault frame, out loud. Something close to: Nothing said in this room becomes a performance conversation. We are testing the plan, not the people. If the honest answer is "I have no idea," that is the most valuable thing you can say today, because it is a finding and I will write it down as one. CISA's version for real incidents applies identically: "Retrospectives must be blameless… Security incidents are rarely the result of one person's action. They are almost always the result of a failure of the overall system" (CISA IRP Basics).
Seat people away from their own teams. SP 800-84 is specific: participants are deliberately not seated with teammates, to encourage independent thinking and cross-exposure. It feels fussy for four minutes, then starts producing answers you would not otherwise have heard.
Keep the engineers from solving it. This failure mode is unique to security tabletops and it is not a discipline problem — it is what good engineers do. Someone starts designing the detection rule that would have caught inject 2, and the room follows, because that conversation is more comfortable than the one about who may call the CEO at 02:00. Two phrases handle most of it: "Assume it works — what do you do with the output?" for the person building the tool, and "Assume it doesn't — now what?" for the person whose plan depends on it. One rule resolves the rest: play the plan you have, not the plan you meant to write. When someone says "well, we'd obviously check the vault" and the plan does not say that, the data collector writes undocumented step relied upon and the exercise moves on.
Timekeeping. Hold the exercise clock visibly and give each inject a hard discussion budget. When the budget expires with no decision made, say so — "we are at time; the decision was not reached" — and let the data collector record it. A stalled decision is data; rescuing the room from a stall destroys it. Anything important but off-objective goes on a visible parking lot and gets an owner at the end.
The evaluator's job. Data collectors write against the criteria set in advance, capturing four things per inject: what was decided, who decided, how long it took from delivery, and what participants reached for — a document, a person, or a memory. That last one matters more than it looks. If four people reached for a colleague's memory rather than the playbook, you have a discoverability problem regardless of how correct the playbook's contents are.
Actionable takeaway: Name a facilitator and a separate data collector for every exercise, and have them meet a week beforehand with the objectives, the scoring sheet and the last exercise's findings in hand. If you cannot spare two people for three hours, you cannot spare the finding you were going to get.
Most organizations end an exercise with a warm feeling and a slide. Produce this instead.
Per-objective rating. Four levels, applied to the objective and never to a person:
| Rating | Meaning |
|---|---|
| Performed without challenges | The objective was met as written, within the stated threshold |
| Performed with minor challenges | Met, but late, or via an undocumented route, or only because one specific individual was present |
| Performed with major challenges | Partially met; the plan was materially wrong or unusable at this step |
| Unable to perform | Not met. No route existed |
"Only because one specific individual was present" is deliberately a minor challenge rather than a pass. Key-person dependency is the most common quiet finding in security exercises, and it never shows up unless you score for it.
Measured times. Record these regardless of rating, because they trend across exercises where ratings do not: time to declaration; time to assemble incident command; time to first containment action approved; time to out-of-band bridge established; time to first holding statement approved.
Stall count. The number of decisions that stopped awaiting an absent authority. This single integer is the most persuasive number you will take to an executive, because it converts "our escalation paths are unclear" into "on Thursday, four decisions stopped for an average of eleven minutes each, waiting for someone who was not reachable."
Then the artefacts. CISA's CTEP discipline is the After-Action Report paired with an Improvement Plan, and the pairing is what separates exercise value from exercise theatre, because the Improvement Plan is where every finding acquires an owner and a due date (CISA CTEP). SP 800-84 says the same: after the report, "the plan coordinator might assign action items to select personnel to update the IT plan" — and should then actually update it.
A finding record that survives contact with a busy quarter carries seven fields:
| Field | Example |
|---|---|
| ID | TTX-2026-Q4-F03 |
| Objective | OBJ-6 — out-of-band bridge |
| Observation | Bridge details existed only in the SSO-protected wiki; no participant could produce them offline |
| Owner | Named role (IR Lead), not a person's initials |
| Due date | A calendar date, not "next quarter" |
| Acceptance test | Three named responders produce dial-in details from a printed card with the tenant unreachable |
| Playbook change | PB-RANSOM §Comms — add out-of-band bridge to the printed contact card and to the header contact block |
The acceptance test field is the one people cut, and it is the one that makes the finding real. A finding without a written acceptance test closes when someone feels it is done.
Actionable takeaway: Score every objective, publish the stall count, and put exercise findings into the same tracking system as your vulnerability findings so they reach the same executive on the same report. Findings that live in a separate document die in it.
Exercises test whether the plan works. Purple teaming tests whether the detections and response actions the plan invokes actually fire. Different question, different budget, different failure mode.
The distinction from a penetration test is not snobbery. A pen test asks whether an attacker can get in, and is scored on findings — usually perimeter and application weaknesses. An adversary emulation asks whether, given an attacker already executing a specific known behavior on your estate, your telemetry sees it, your logic alerts on it, and your responders act on it. A clean pen test report alongside zero detection coverage is a very common combination, and the second condition is the one that determines how long an adversary lives in your network.
The material is free. MITRE's Center for Threat-Informed Defense publishes an Adversary Emulation Library of plans modeled on real threat actors' behaviors, in full emulation form (initial access through exfiltration) and micro emulation form (CTID). Micro emulations are the on-ramp for a small team: a single behavior, executed deliberately, checked against your SIEM, in an afternoon. The Purple Team Exercise Framework provides the open methodology for the collaborative CTI-plus-red-plus-blue version, with a named coordinator role and a flow from threat intelligence through attack planning, emulation, detection and response (PTEF). And RE&CT does for the response side what ATT&CK coverage mapping does for detection — coverage and gap analysis across response actions (RE&CT).
CISA builds emulation into post-incident activity, with a warning attached: "Advanced SOCs should consider emulating adversary TTPs to ensure recently implemented countermeasures are effective… This testing should be closely coordinated with a blue team to ensure that they are not mistaken for true adversary activity" (CISA Playbooks).
Report the coverage triple, never a percentage. For each prioritized technique, three separate values: do we have the telemetry (a visibility score, from a tool like DeTT&CT), do we have logic (a rule exists and is enabled), and has it fired on a validated test (a date). Green on all three is coverage; anything else is a named gap with a named owner. A mapped technique is not a validated detection, and a validated detection is not coverage (DeTT&CT / NVISO Labs). A technique with no telemetry is not a detection-engineering problem at all — it is an ingest and budget problem, and conflating the two is how teams burn a quarter writing rules that can never fire. Palantir's Alerting and Detection Strategy framework makes the point structurally: every documented detection carries a Validation section describing "the steps required to generate a representative true positive event which triggers this alert. This is similar to a unit test" (Palantir ADS).
One 2026-specific item belongs on every purple team's list this year. MITRE ATT&CK v19 split Defense Evasion into two tactics — TA0005, renamed Stealth, and a new TA0112 Defense Impairment — as of v19.2, current since 28 April 2026 (ATT&CK versions). Any coverage map, SIEM dashboard or purple-team report built on v18 or earlier now has a stale tactic axis. Chapter 9 owns detection engineering; the exercise-program obligation is narrower: re-baseline your coverage map against the pinned ATT&CK version once a year, and record which version each report was built on.
Actionable takeaway: Pick three techniques from your top scenario, run the micro emulations this month, and record the coverage triple for each. Three validated detections beat a spreadsheet claiming eighty percent coverage that nobody has ever fired a test through.
This is the part of the chapter with the best return per hour spent, and it needs no scenario, no facilitator and no budget. These are tests in SP 800-84's strict sense — quantifiable checks on whether a mechanism works — and every one has failed for real, in public, at an organization better resourced than yours.
| What to test | The test | Pass criterion | Where this failed for real |
|---|---|---|---|
| Call tree / notification list | Unannounced cascade; every recipient acknowledges | 100% acknowledgement within the plan's stated time | Equifax: the patch notice went to an out-of-date recipient list and never reached the people who would have installed it (GAO-18-559) |
| Out-of-band comms | Convene the bridge with corporate SSO treated as unavailable | Quorum present within 15 minutes, using details held offline | British Library: with website and intranet down, response ran over social media and WhatsApp cascades (British Library) |
| Backup restore | Restore one defined critical system to an isolated network; measure end to end | Restored, validated, and within the documented RTO | British Library lesson 8: "'Legacy' systems are not just hard to maintain and secure, they are extremely hard to restore" |
| Break-glass account | Use it in a change window; verify the alert fires and the audit record exists | Access succeeds, alert fires, use is reviewed | CSRB: Microsoft stopped manual key rotation in 2021 after an outage linked to the rotation process (CSRB) |
| After-hours escalation | Page the on-call chain at 02:00 on an unannounced weeknight | Human acknowledgement within the plan's threshold, at every tier | Sophos: 88% of ransomware encryption occurs outside business hours (Sophos) |
| Printed contact list | Ask three responders to physically produce their copy | Three copies produced, current version | CISA: "Print these documents and the associated contact list… During an incident, your internal email, chat, and document storage services may be down" (CISA IRP Basics) |
| IR retainer / insurer line | Call the number in the plan; time to reach a human; confirm contract currency and panel constraints | Human contact within SLA; no contract collision | Chapter 13's readiness table records "retainer expired" as a recurring finding |
| MFA exception register | Enumerate every internet-facing system without enforced phishing-resistant MFA | The list exists, is dated, and every entry has an owner and an end date | Change Healthcare: attackers used compromised credentials against a Citrix portal with no MFA, despite policy requiring it (Healthcare Dive); Colonial Pipeline: a legacy VPN profile "not intended to be in use," without MFA (Blount testimony); British Library lesson 3: MFA on all end-user technology "but not on certain supplier endpoints" |
| Detection validation currency | For each prioritized technique, the date it last fired on a test | No prioritized detection older than the documented validation interval | The coverage triple, above |
Look at the right-hand column and notice the pattern. Policy is universal; enforcement is not; and the exception is almost always at the seam with a third party or a legacy system. No playbook fixes that. A quarterly enumeration test surfaces it.
Do the call tree first. Not next quarter. This quarter. It costs one email and an hour of chasing acknowledgements, and it has already cost somebody else a great deal more.
Actionable takeaway: Put all nine rows on a recurring calendar with a named owner and a recorded result per run. None require a facilitator; most take under an hour. This is the highest-yield hour in the chapter.
Chapter 13 owns the post-incident review for real incidents. The exercise hotwash is the same instrument at lower stakes, and it happens immediately after the exercise, in the room, before anyone leaves. SP 800-84 gives the agenda as three questions the facilitator asks the participants: in which areas did they excel, where do they need training, and which parts of the plan should be updated. Fifteen minutes, verbal, no slides. The written report comes later; the hotwash catches what people will have rationalized away by Monday.
Blamelessness is not politeness, it is an information-gathering technique, and Chapter 13 sets out the evidence base for it. The exercise-specific consequence is narrower and worth saying plainly: the information you need lives in the head of the person who would look worst telling you. Blame is the mechanism by which you guarantee they do not.
Two current moves are worth importing into exercise reviews: the shift from "blameless" to "blame-aware" — everyone works within constraints, and some only become visible after the fact — and Calibrate, circulating draft findings before the review meeting so nobody is surprised in front of their peers (Howie: The Post-Incident Guide). Ambushing someone with a finding in a room full of colleagues buys you one finding and costs you a year of honest reporting.
Actionable takeaway: Run the verbal hotwash before anyone leaves the room, circulate draft findings for calibration within five business days, and hold the written review within ten. Momentum is the only thing that converts observations into changes.
Frequency is an organization-defined parameter under IR-3 and IR-8, which means you must choose and document it — "as needed" is not a frequency. Here is a defensible default set, and the tiers scale down honestly for a small team.
| Cadence | What | Rung | Minimum for a small org |
|---|---|---|---|
| Quarterly | Alert/notification/accountability cascade test | 4 | Same — it is an email and an hour |
| Quarterly | One 2–3 hour operational tabletop, rotating scenarios | 3 | One 90-minute micro-exercise from NCSC Exercise in a Box |
| Quarterly | Break-glass account use; restore of one defined critical system | 4 | Same, on your single most important system |
| Quarterly | Micro-emulation set against your top three techniques | 4 | Three atomic tests, checked in the SIEM |
| Semi-annually | Senior-level (executive) tabletop | 3 | Annual, 2 hours, with the leadership you have |
| Annually | Combined senior + operational exercise | 3 | Combine with the executive session |
| Annually | Functional exercise using real consoles and real timings | 5 | Substitute a full unannounced restore test |
| Annually | Joint exercise including at least one critical supplier (ID.IM-02, GV.SC-08) | 3 | A one-hour joint call walking the notification path |
| Annually | Re-baseline ATT&CK coverage against the pinned version | — | Same; the v19 tactic split makes this year non-optional |
| After every real incident | Blameless hotwash; then emulate the adversary's observed TTPs to verify the new countermeasures actually fire | 4 | The hotwash at minimum |
| On change | Any new system, supplier, regulation, or change of authority triggers a targeted review | — | Same |
That last row is not padding. NIST SP 800-61r3 enumerates where improvements come from and each is a trigger: evaluations and audits (ID.IM-01), tests and exercises (ID.IM-02, rated High), and the execution of operational processes (ID.IM-03, also High) (NIST SP 800-61r3). Calendar cadence alone produces a review that finds nothing, because the calendar does not know that you changed identity providers in March.
Actionable takeaway: Publish the exercise calendar twelve months out, with owners, and treat a missed exercise the way you treat a missed patch SLA — as a tracked exception with a named accepter. Exercises that float are exercises that slip.
Everything above is overhead unless the findings change the playbook. That is the whole point, and it is where most programs quietly stop.
The mechanism is described fully in Chapter 2, so here is only the exercise-side half. Every playbook header carries a last_exercised field; every exercise that touches a playbook updates it; and a CI check fails or flags any playbook whose date is older than your documented interval, flipping its status from Active to Draft. Not because someone noticed — because the pipeline noticed. A playbook nobody has rehearsed in a year is a hypothesis, and labeling it accurately is the cheapest honesty available to you.
Then the finding lifecycle: every after-action finding becomes an issue with an owner and a due date, carries a written acceptance test, produces an identified playbook change, and — the step everyone forgets — is re-tested at the next exercise touching the same objective. Findings that close on assertion reopen in production. A finding is verified when someone other than the owner has run the acceptance test.
Update the authorities every time, whether or not anything about them came up. CISA's hotwash objectives put "reviewing and updating roles, responsibilities, interfaces, and authority to ensure clarity" on the standing list, and it is there because unclear authority is a recurring real-world finding. Authorities rot faster than procedures. A reorganization does not send a notification to your playbooks.
One closing calibration, because this chapter has been enthusiastic and the enthusiasm has limits. Mandiant's conclusion from over 500,000 hours of 2025 incident response is that most intrusions still stem from human and systemic failures, not from novel adversary capability (M-Trends 2026). Exercises are how you find human and systemic failures before someone else monetises them. That is not a small claim, and it does not require a single new license.
Rehearse it, time it, write down what broke, and fix the thing before the calendar makes you do it again.
[IG1] [ID.IM-02] [CIS 17] [A.5.24][IG1] [ID.IM-02][IG2] [ID.IM-02][IG2] [ID.IM-02][IG2] [GV.RR] [A.6.3][IG2] [ID.IM-02][IG1] [ID.IM-02] [A.5.27][IG1] [ID.IM-02] [A.5.27][IG2] [ID.IM-02]last_exercised date, and an automated check flags or fails any playbook whose date exceeds the documented interval. [IG3] [ID.IM-02][IG1] [RS.CO] [CIS 17][IG1] [CIS 17] [A.5.29][IG1] [A.5.24][IG1] [CIS 11] [RC.RP] [A.5.30][IG2] [PR.AA][IG2] [RS.MA][IG1] [GV.SC-08] [CIS 15][IG2] [GV.SC-08] [ID.IM-02][IG3] [CIS 18] [DE.AE][IG3] [CIS 18][IG3] [DE.CM] [A.8.16][IG3] [DE.AE][IG3] [ID.IM-03] [DE.CM][IG1] [PR.AA] [CIS 6][IG2] [GV.RR] [ID.IM-01]