Companion volume · Edition 2026 · Intelligent Automation
The
Blast Radius
Incident response for managed service providers — what changes when the estate is not yours, your keys reach hundreds of companies at once, and your authority to use them stops at a contract line nobody has read since the day it was signed.
#The Blast Radius
Incident Response for Managed Service Providers
A standalone companion to The 2026 InfoSec Playbook
Daniel Ramos — CTO, Intelligent Automation, LLC Edition 2026.1 · September 2026
You hold the technical keys to hundreds of companies and the legal authority to use them at only some of them. Every other difference in MSP incident response follows from that gap — and from the fact that when you are patient zero, the incident is not one company's bad night. It is all of them, at once.
#A note before you start
This book is a companion to The 2026 InfoSec Playbook and does not repeat it. Where the main playbook covers something in depth — the incident response lifecycle, the fourteen scenario playbooks, the regulatory notification matrix — this book points at it rather than restating it. You do not need the main playbook open to use this one.
This is not legal advice. Parts 2, 4, 5 and Appendix B discuss contracts, liability and regulatory obligations because an MSP cannot respond to an incident without understanding them. All contract language here is illustrative — written to show you what a clause needs to achieve, not to be signed. Have your own counsel draft the real thing, and do it before the incident rather than during it.
Attribution and sources. Every incident, statute, deadline, command and statistic in this book traces to a primary or authoritative source, cited inline. Where something could not be confirmed it is marked with a Verify callout rather than asserted. Case studies are drawn from published post-incident reports, government advisories, court filings and regulator findings — they appear here to be learned from, not to be laughed at.
© 2026 Intelligent Automation, LLC.
#Contents
Part 1 — The Three Nights
- Scenario One — 23:45, Malware Is Spreading Across a Client Domain
- Scenario Two — 02:00, The Domain Controller Needs Isolating and the Client Is Unreachable
- Scenario Three — The Monday After, The Client's Auditor Wants the Evidence
Part 2 — Before the Night
Part 3 — The MSP-Side Playbooks
Part 4 — The Notification Chain
Part 5 — When the MSP Is the Regulated Entity
Part 6 — Templates and Tools
- Templates and Tools
- Appendix A — The Client Authority Matrix
- Appendix B — MSA and DPA clause library
- Appendix C — Per-tenant evidence pack manifest
- Appendix D — Six MSP tabletop scenarios
- Appendix E — The first ninety days
#Why This Book Exists
Hello, fellow keyholders. You are the people who get called when someone else's business stops, and you are about to spend three hours reading about the night it is your turn. Grab the coffee. This one is worth it.
#The Friday everyone remembers
It was 2 July 2021, the Friday afternoon before the US Independence Day weekend — the point in the calendar when the maximum number of on-call rotations are covered by the minimum number of humans, and everyone else is already on the road. A REvil affiliate exploited zero-days in internet-facing, on-premises Kaseya VSA servers: the remote monitoring and management platform whose customer base is overwhelmingly managed service providers (NCSC/ODNI factsheet).
Here is what that actually looked like from inside an MSP.
Nothing broke. The console was up. The agents checked in. What went out that evening was a routine agent-update procedure — the same mechanism used to push a hotfix on any ordinary Tuesday. Huntress's reconstruction of the chain is an authentication bypass in the VSA web interface, an authenticated session, a payload upload, and command execution by code injection (Tenable, quoting Huntress). Downstream, a PowerShell payload switched off Windows Defender real-time monitoring, IPS, IOAV protection and script scanning; a renamed copy of certutil.exe decoded an encrypted blob; the result side-loaded the ransomware into a legitimate, signed Windows Defender binary. Before any of that, the attackers deleted the IIS logs and the logs held in the application database (Truesec).
By 5 July, Kaseya's reported count was fewer than 60 direct clients and not more than 1,500 businesses supported by those clients — a roughly 25x amplification. REvil asked US$70 million for a universal decryptor. Individual firms reportedly paid between $40,000 and $220,000 before Kaseya announced on 23 July that it had obtained a universal key (NCSC/ODNI). The downstream reached a Swedish supermarket chain, kindergartens in New Zealand and public administration offices in Romania (The Record) — organizations that had never heard the name of the MSP whose server had just ended their week.
Now hold the number that matters. For those fewer than sixty providers, this was not a client incident they were helping with. There was no unaffected side of the business to run the response from, no other client to borrow a technician from, no console to work through — because the console was the delivery mechanism. It was every client, at once, delivered by their own tooling working exactly as designed. Nobody exploited the ransomware push feature. There is no ransomware push feature. There is a "deploy to all managed endpoints" feature, and it worked perfectly.
That is why generic incident response advice keeps missing. Almost every IR playbook ever written assumes one organization, one estate, one legal entity, one decision-maker, and a bad night that ends. Yours can assume none of those things.
Actionable takeaway: Before you read any further, answer one question out loud: if our RMM pushed something malicious to every managed endpoint tonight, what is the first thing we would do that does not require the RMM? If the answer takes longer than ten seconds, the rest of this book is for you.
#The thesis, in one sentence
Everything in this book follows from a single gap.
You hold the technical keys to hundreds of companies and the legal authority to use them at only some of them.
The keys are real. Your delegated admin relationships, your RMM agents, your backup consoles and your technicians' identities reach into estates you do not own and are contractually obliged to touch every day. The authority is thinner than almost anyone at an MSP believes. Published MSP master agreements are, overwhelmingly, scope, payment and liability documents. The CompassMSP master service agreement, for one public example, grants suspension rights only for non-payment, disclaims any guarantee that breaches can be prevented, caps liability at six months of fees, and contains no emergency authority clause, no incident response clause and no breach-notification-to-client timeline at all (CompassMSP MSA). If your MSA looks like that, you are in good company — most do — and Part 2 is where we fix it, with the drafting language in Appendix B.
NIST puts the same point in operational language. SP 800-61r3 says third-party responsibilities should be defined in a contract, including "authority to act on behalf of the organization" and "restrictions on what the service provider can do, such as ... making and implementing operational decisions (e.g., immediately deactivating certain services to contain an incident)" (NIST SP 800-61r3). Translation for 02:00: the credentials in your password manager answer the question can I. They do not answer may I, and only one of those two questions has a lawyer attached to it.
This is not a reason to freeze. It is a reason to decide the hard parts in daylight, when everyone is calm, caffeinated and not being paged. A technician who isolates a client's domain controller alone at 02:00 with no named approver is not a hero. They are an uninsured liability standing in a room where every possible outcome is somebody else's to argue about — and the whole point of what follows is that nobody at your firm should ever be put in that position again.
#The seven asymmetries
These are the seven ways MSP incident response differs structurally from the version in every other playbook. They recur throughout this book, so they get named once here.
1. You are the supply chain. Your compromise is everyone's compromise, simultaneously. Microsoft's own term for the strategy, describing NOBELIUM's campaign against CSPs and MSPs, is "compromise-one-to-compromise-many" — and Microsoft documented intrusions spanning four separate providers to reach a single final target (Microsoft Security Blog). Part 1 puts you inside three nights where this is true.
2. Authority is contractual, not technical. You can, therefore you may is false, and expensively so. US federal computer-crime law has a provision most MSPs have never been shown: 18 U.S.C. § 1030(a)(5)(A) reaches anyone who transmits a "command" and thereby "intentionally causes damage without authorization" — where "damage" is defined as any impairment to the integrity or availability of data, a program, a system, or information, and § 1030(g) allows a private civil action where loss aggregates to at least $5,000 in a year — a figure a few hours of production downtime clears without effort (18 U.S.C. § 1030). Your admin rights answer whether the access was authorized. That provision asks whether the damage was. No court has been found applying it to an IT provider doing incident response — the reading comes from the statutory text, not a case — which makes it an untested edge rather than a settled rule, and an untested edge is a poor place to be standing at 02:00. Part 2 turns this into clauses and a per-client authority matrix; Appendices A and B are the artifacts.
3. Triage is a portfolio problem. One indicator at one client is a question about two hundred — and the sweep answering it is incomplete by design. Defender advanced hunting looks back 30 days (Defender XDR data retention), multi-tenant views cap at 100 tenants (Defender multitenant management requirements), and the row budget is 50,000 divided by the number of tenants you selected (multitenant advanced hunting). Record which tenants returned "no data" and which returned "no result". They are different findings and only one of them is good news. Part 2 builds the sweep and pre-computes its limits.
4. Delegated admin is the attack path, and it is the same path you work through daily. When Microsoft moved partners from legacy DAP to GDAP, the default role set assigned to the Admin Agents group included Privileged Authentication Administrator — which can reset authentication methods for any user, including the customer's own Global Administrators — and Privileged Role Administrator, which can manage role assignments and PIM (Microsoft Learn, Microsoft-led GDAP transition). Those two roles together are a two-step path to full tenant control, sitting in relationships many MSPs have never audited because Microsoft created them automatically. Part 2 covers auditing and trimming them without breaking service delivery.
5. Evidence must be per-tenant, and commingling it is its own incident. Client A's evidence pack must contain zero Client B data, and your multi-tenant consoles are built to aggregate — a cross-tenant query exported to CSV is a commingled export by construction. Sentinel's documented MSSP best practice is one workspace per tenant precisely because it produces "fewer challenges regarding data ownership, data privacy and regulatory compliance" (Microsoft Learn). Segregation is the platform default; the engineering burden is on you not to break it. Part 2 covers scoped collection and the manifest that proves it.
6. Notification is a chain, and your contractual clock is usually tighter than any regulator's. The 72-hour GDPR clock, the 24-hour NIS2 clock and the four-business-day SEC clock all belong to your client. What binds you is your MSA — and MSP contracts routinely say "immediately", "within four hours", or "without undue delay upon becoming aware". Those are commercial terms with termination and indemnity consequences attached, and they are breached long before any statutory deadline is missed. They are also not uniform: one incident touching forty clients can involve a dozen different deadlines, formats and notice addresses. Part 4 has the chain, the clocks and the per-client fields to put in your PSA.
7. Patient zero might be you, and you must be able to rule that in or out fast. This is not hypothetical. In the DragonForce case documented by Sophos, the MSP's own SimpleHelp server was the entry point; the actors used the RMM's legitimate inventory features to gather "device names and configuration, users, and network connections" across multiple customer estates, then pushed a malicious installer through it (Sophos X-Ops). CISA now writes MSP advisories with an explicit instruction to contact your downstream customers (CISA AA25-163A). That email should already be drafted. Part 1's third scenario is this night; Part 2 builds the test that answers it.
Actionable takeaway: Print these seven. Put them on the wall next to the on-call phone. When an alert does not fit the playbook, it is almost always because one of these seven is the reason it does not fit.
#The roles, defined once
This book reuses the incident command roles from the main playbook, Chapter 13, and adds five MSP-side roles, used consistently from here on. They are roles, not people. At a twelve-person MSP one person will wear three of them, and that is fine, as long as they know which hat they are wearing for each call.
The six reused roles, in one line each, so you can work from this book alone. These restate the main playbook, Chapter 13; they do not replace it.
| Main-playbook role | Who this is | Owns |
|---|---|---|
| Incident Commander | The one person running the incident. At an MSP this is the MSP Incident Commander below, scoped across every affected tenant. | Decisions, delegation, tempo, and the running picture. Not the hands-on technical work. |
| Operations Lead | The senior hands-on responder. | Executing the technical response and reporting status back to the IC. |
| Communications Lead | Whoever speaks to anyone outside the response. | Client-facing and internal updates, on a stated schedule. |
| Scribe | A dedicated person who is neither the IC nor a responder. | The timeline: what happened, when, and what was decided, by whom, at what time. |
| Legal Liaison | Internal or external counsel. | Privilege, notification obligations, regulator and law-enforcement contact. |
| Executive Sponsor | The senior leader who can spend money and accept risk. At an MSP this is usually the Practice Owner. | Funding, external commitments, and decisions above the IC's authority. |
And the five MSP-side roles this book adds:
| Role | Who this is at a real MSP | Owns | Does NOT own |
|---|---|---|---|
| On-Call Technician | The person the alert wakes. | Detection, first triage, evidence preservation, escalation. | Declaring severity; any disruptive action beyond pre-authorized scope. |
| MSP Incident Commander | Service delivery manager or senior engineer. | The incident across all affected tenants; severity; blast-radius call; task assignment. | Contractual, insurance and regulatory decisions. |
| Client Authority Holder | The named person at the client who can approve disruptive action. Per client, per time-of-day, with a named deputy. | Approving or refusing action inside their own estate. | Anything at another client, or at the MSP. |
| Tenant Lead | The MSP engineer who owns the relationship and context for one client. | Client context, the authority-matrix lookup, client-facing continuity. | Portfolio-wide decisions. |
| Practice Owner | MSP owner or director. | Contractual, insurance and regulatory decisions; downstream notification sign-off; declaring a platform incident. | Running the technical response. |
The On-Call Technician is explicitly not the decision-maker, and that is a protection, not a demotion. Three reasons, all of them in the technician's favor. First, the legal exposure in asymmetry 2 attaches to whoever pressed the key; a documented instruction from a named role moves it to where it belongs. Second, most liability caps in MSP agreements are disapplied for gross negligence and willful misconduct — and a deliberate shutdown is by definition an intentional act, so the only thing separating "proportionate emergency response" from "willful misconduct" is the contemporaneous record of who decided and why. Third, and most practically: at 02:00, alone, with a phone ringing, nobody should be doing legal analysis. They should be doing triage, and escalating without shame.
#Blast radius: the severity dimension the main playbook does not have
Severity levels — SEV-1 through SEV-4, defined in the main playbook, Chapter 13 — measure how bad it is. Restated here in MSP terms, so this book works on its own. This does not replace the main playbook's definitions, and if your firm already runs its own schema, keep it and map it.
| Severity | What it means at an MSP |
|---|---|
| SEV-1 | Confirmed or strongly suspected compromise with active, spreading or unbounded impact: encryption in progress, an adversary holding administrative control, or any credible suspicion that MSP platform infrastructure is involved. Response runs continuously until it is over. |
| SEV-2 | Confirmed malicious activity whose scope is known and bounded, or a compromise whose scope is not yet established. Someone is awake and working it now. |
| SEV-3 | Suspicious activity that needs investigation but shows no evidence of adversary control: a single blocked detection, one credential to reset, an alert that does not reconcile. Worked in hours, not minutes. |
| SEV-4 | A security-relevant finding with no active adversary: configuration drift, a missed patch, an alert with a benign explanation. Ticketed and tracked. |
For an MSP that is half the question. The other half is how many legal entities does this reach, and through what. So every incident in this book carries two labels: a severity and a blast radius.
| Blast radius | What it means | Test |
|---|---|---|
| Single-tenant | Confined to one client's estate. No shared credential, shared tool or MSP-side system implicated. | The indicator has no provenance in anything the MSP operates. |
| Multi-tenant | The same actor, indicator or credential is confirmed at two or more clients that share nothing except you. | Cross-client correlation returns a hit, or a credential used at more than one client is implicated. |
| Platform | The MSP's own estate: RMM, PSA, backup console, documentation platform, partner tenant, or a technician identity. | Provenance touches anything on that list — confirmed or suspected. |
How the two dimensions interact:
| SEV-4 | SEV-3 | SEV-2 | SEV-1 | |
|---|---|---|---|---|
| Single-tenant | Ticket; Tenant Lead informed next business day. | On-Call Technician + Tenant Lead. | MSP IC paged. Client Authority Holder engaged. | MSP IC + Practice Owner. Full client-side IR. |
| Multi-tenant | Does not exist — floor is SEV-3. | MSP IC paged. Cross-client sweep opened. | MSP IC + Practice Owner. Parallel notifications begin. | Portfolio response. All affected Tenant Leads activated. |
| Platform | Does not exist. | Does not exist. | Does not exist. | All platform-level incidents are SEV-1. |
That bottom row is the rule this book will not negotiate on:
Any credible suspicion of a platform-level compromise is SEV-1 until disproven. Not until confirmed. Until disproven.
The reasoning is arithmetic, not drama. Fewer than sixty providers, up to 1,500 businesses. The cost of over-declaring a platform incident is a bad night and a postmortem that says you were cautious. The cost of under-declaring one is every client at once, reconstructed later from logs the attacker had time to delete. PagerDuty's public severity guidance puts the general form of the rule plainly: if you are unsure which level an incident is, treat it as the higher one, and reassess at the postmortem, never during (PagerDuty). For an MSP, that is not advice. It is the business model's insurance policy.
The cheap version, for a twelve-person MSP with no SOC: this is a laminated card by the on-call phone with those bullet points on it and one phone number. The expensive version detects the same things faster. It does not decide them better.
#How to use this book
Three doors. Take the one that matches tonight.
"It is happening now." → Part 1. Three clock-stamped scenarios written to be read while an incident runs: the first hour, the wrong instinct a competent technician will actually have and why it is wrong here, the decisions with decide-by times and named authorities, and the step tables. Skip everything else. Come back later.
"We are building the capability." → Part 2. The controls that turn those scenarios into ordinary Tuesdays: management-console hardening, delegated-admin least privilege, cross-tenant hunting within its real limits, technician identity, per-tenant evidence collection, and escalation design for an MSP that has one person awake.
"The owner needs to understand the exposure." → Parts 4 and 5. Contractual authority, liability in both directions, insurance, and the regulatory chain — including the ICO's first penalty against a processor (£3,076,320, against an IT supplier), which turned on gaps in MFA deployment, insufficient vulnerability scanning and inadequate patch management. The aggravating detail was a working MFA solution built but not rolled out, because of a perception that customers would resist it (Clifford Chance analysis of the penalty notice).
#A note on what this is not
This is a standalone companion to The 2026 InfoSec Playbook, not a chapter of it, and it deliberately does not repeat it. The main playbook covers incident command structure, severity definitions, identity containment sequencing, forensic fundamentals and the framework mappings — properly, at length, for a single organization. All of that still applies to you. Where this book needs it, it says "the main playbook, Chapter N" and moves on, because reprinting it would double the length and halve the usefulness. What is in here is only the delta: the seven asymmetries, and everything that follows from them. If a procedure below seems to skip a step you would expect, the step is in the main playbook and it has not been repealed.
One last thing before Part 1. The median time to attempted Active Directory compromise, across 661 incident response and MDR cases, has accelerated 70% year-over-year to 3.40 hours (Sophos 2026 Active Adversary Report). If your after-hours escalation path takes longer than that to reach someone who can decide, your architecture has already chosen the outcome — and no amount of talent on the night will change it. That is fixable this week, for free, with a phone tree and two names.
Know your keys, know your permissions, and never confuse the two — you are not managing two hundred networks, you are managing one with two hundred front doors.
#Sources
- ODNI/NCSC — Kaseya VSA Supply Chain Ransomware Attack factsheet
- The Record — Kaseya: more than 1,500 downstream businesses impacted
- Tenable — CVE-2021-30116: multiple zero-day vulnerabilities in Kaseya VSA
- Truesec — Kaseya supply chain attack targeting MSPs to deliver REvil ransomware
- CompassMSP — Master Service Agreement
- NIST SP 800-61r3 — Incident Response Recommendations and Considerations (PDF)
- 18 U.S.C. § 1030 — Cornell LII
- Microsoft Security Blog — NOBELIUM targeting delegated administrative privileges
- Microsoft Learn — Microsoft-led DAP to GDAP transition
- Microsoft Learn — Set up Microsoft Defender multitenant management
- Microsoft Learn — Advanced hunting in Microsoft Defender multitenant management
- Microsoft Learn — Data retention and data security in Microsoft Defender XDR
- Microsoft Learn — Prepare for multiple workspaces and tenants in Microsoft Sentinel
- Sophos X-Ops — DragonForce actors target SimpleHelp vulnerabilities to attack MSP, customers
- CISA AA25-163A — Ransomware actors exploit unpatched SimpleHelp RMM
- PagerDuty — Severity Levels
- Clifford Chance — ICO fines processor after inadequate security measures lead to widespread disruption
- Sophos — 2026 Active Adversary Report
#Part 1 — The Three Nights
#Scenario One — 23:45, Malware Is Spreading Across a Client Domain
Blast radius on arrival: unknown. That is the whole problem.
#23:45 — The page
You are the On-Call Technician. There are fourteen of you at this MSP and 180 clients on the book, and tonight the rotation is yours. Your phone goes off at 23:45 on a Tuesday — a detail worth noticing, since 88% of ransomware encryption happens outside business hours (Sophos, 2026 findings). Nobody schedules these for 10am.
The alerts are all from one tenant. Call it Client M — a 220-seat metal fabricator, two sites, one small server room, a night shift running the line. In the console you see:
- 23:41:07Z — credential-dumping behavior on
M-FS02, a file server: LSASS access from a non-standard parent process. - 23:42:19Z — the same on
M-APP01. - 23:43:55Z — remote service creation,
M-FS02→M-SQL01. - 23:44:30Z, 23:44:41Z, 23:45:02Z — three workstations in the same subnet, same signature.
Six hosts in four minutes, and the count is climbing while you read. Nobody at Client M is awake. Their IT manager's phone goes to voicemail, and the voicemail is full.
This is the night every MSP technician has rehearsed in their head and almost nobody has rehearsed on paper. Let us do it properly.
#23:47 — The wrong instinct, named
Here is what a competent technician does next, because it is obvious and because it feels like doing something: they click Isolate on M-FS02. Then Run antivirus scan. Then the next alert lands and they isolate that host too. They are working the queue, one host at a time, in the order the console hands them out.
That is not stupidity or laziness. It is the most natural response available to someone alone at midnight with a rising alert count, and it is what the console's interface is designed to encourage. It is also the thing that loses you this incident.
The failure has been documented for over a decade. Mandiant's Jim Aldridge described the chain exactly: responders remove the compromised systems they know about and feel accomplished; in doing so they "tip their hand to the attacker"; the attacker, still holding backdoors on systems the responders never found, "will take steps to ensure continued access," abandoning the burned malware, utilities and C2 and keeping everything else; and then "the responders will continue to be blind, and unaware; typically, this lasts until an outside party, e.g. law enforcement, notifies the organization again that they are compromised." (Aldridge, *Remediating Targeted-threat Intrusions*) CISA puts the same point in playbook language: "Some adversaries may actively monitor defensive response measures… Defenders should therefore develop as complete a picture as possible of the attacker's capabilities and potential reactions to avoid 'tipping off' the adversary." (CISA IR Playbooks)
Now read the alert stack again. Credential dumping came first. By the time you saw alert one, they already had material from LSASS. Isolating the host does not un-steal the credentials or close the accounts they open — it is changing the locks on the front door while the keys are already in someone's pocket, and doing it noisily enough that they hear you. That noise is the cost: it announces, at machine speed, that you are awake.
Three more reasons the piecemeal order fails specifically for an MSP:
- You have not asked the question only you can ask. If this indicator also exists at four other clients, the correct first action is not on this endpoint — it is on your own estate.
- You have not checked your authority. One endpoint isolation under a managed-EDR agreement is almost certainly inside it. Six, then a server, then the DC as the pattern climbs, is a slope with no landing. Nobody decides to isolate a domain controller without authority; they arrive there one defensible click at a time.
- The mechanics leak. Defender for Endpoint isolation is automatically lifted after seven days, and if the device is offline the service retries for up to three days then gives up (Microsoft, response actions on a device). Isolate nine hosts across ninety minutes and you have created nine expiry timestamps and an unknown number of silent failures, none written down anywhere.
Actionable takeaway: the first twenty minutes are for shape, not action. Know how big this is, whether it is only them, and whether it is you. Then contain everything you found, together, in one burst.
#23:45–23:55 — Phase 1: Establish it is real, and establish its shape
Two questions, ten minutes. Is it real? Behavioral credential-access detections have a known false-positive population — vulnerability scanners, backup agents, an engineer running a legitimate tool at an odd hour, an agent that updated last night. Your own change record is the fastest discriminator you own. What shape is it? One host with six alerts is a different incident from six hosts with one alert each.
| # | Action | Who | Done when | Evidence to capture |
|---|---|---|---|---|
| 1 | Open the incident in the tenant-scoped view, not the multi-tenant one. Record incident ID and earliest-alert UTC. | On-Call Technician | ID and first-alert UTC in the ticket | Incident export; ticket entry in UTC/ISO 8601 |
| 2 | Check change records and RMM job history for this tenant, last 72 hours. Identify any scan, patch, script or maintenance window running now. | On-Call Technician | Every concurrent automated activity named, or confirmed absent | RMM job history scoped to the site; PSA change tickets |
| 3 | Identify patient zero: earliest process ancestry across all alerted hosts, with the account it ran as. | On-Call Technician | A single earliest execution identified, or the ambiguity recorded | Device timeline per host; ancestry query text and results |
| 4 | Enumerate the accounts involved; mark each as domain admin, service account, MSP-managed or ordinary user. | On-Call Technician | Account list complete with privilege level | Account list; group membership export |
| 5 | Name the spread mechanism — remote service creation, WMI, scheduled task, SMB, RDP, admin tool. | On-Call Technician | Mechanism named for at least the first two hops | Query text and results; hashes of any observed binary |
| 6 | Check each affected host for a remote-access agent that is not yours. | On-Call Technician | Agent inventory compared against the deployed baseline | Installed-agent list per host; deviation noted |
| 7 | Count affected hosts, re-count ten minutes later, then assign provisional severity and blast radius. Default: SEV-2, blast radius unknown pending Phase 2. | On-Call Technician | Both counts and both classifications in the ticket | Host list with both timestamps and the delta; ticket entry with reasoning |
Row 7 matters more than it looks. Severity you can revise. Blast radius you must not guess. A single-tenant SEV-2 and a platform SEV-2 are different incidents with different commanders, and writing "single-tenant" at 23:52 because it feels true is how an MSP spends four hours solving the wrong problem.
Row 6 needs two cautions. Post-exploitation crews routinely install a second remote-access tool: ScreenConnect intruders deployed additional ScreenConnect clients and Atera trials (Trend Micro); N-central intruders persisted with Cloudflare tunnels, leaving a service named Cloudflared and svchost.exe in a user's Documents folder (Huntress). But presence alone proves nothing — "The use of these legitimate tools alone is not indicative of malicious activity" (FBI/CISA AA23-320A) — and your EDR may not even be looking, because "often RMM install paths are excluded from EDR inspection" (CISA, Guide to Securing Remote Access Software). Your own exclusion list is the adversary's safe harbour, and tonight is a bad night to find that out.
Takeaway: finish Phase 1 with four written facts — earliest execution, accounts, spread mechanism, host count. Not a theory. Four facts.
#23:55–00:05 — Phase 2: Is this only them? Then: is it us?
This is the phase that separates MSP incident response from ordinary incident response, and no single-tenant playbook contains it. One indicator at one client is a question about 180 companies.
#### The cross-tenant sweep
Fan the indicators out — and do the arithmetic before you press run, because the platform truncates silently and cheerfully.
- Defender multi-tenant advanced hunting covers a maximum of 100 target tenants in one view (Microsoft). At 180 clients you need at least two tenant groups.
- The row budget is brutal: "In multitenant environments, advanced hunting queries can return a maximum of 50,000 records in total. The result set from each individual tenant is capped at 50,000 divided by the number of tenants queried." (Microsoft) At 100 tenants that is 500 rows each. A tight
whereon a hash is fine. A broad process sweep will truncate and will not tell you. - Lookback is 30 days in advanced hunting, 180 days elsewhere in the portal (Microsoft).
- Older or non-Defender data needs Sentinel — capped at "up to 20 workspaces in a single query… we recommend including no more than 5" (Microsoft). And Sentinel is not reachable through GDAP at all; it needs B2B plus Azure Lighthouse (Microsoft). If that was never set up, tonight you have no Sentinel sweep. Record it as a gap; do not pretend otherwise.
- For a scripted fan-out,
POST /security/runHuntingQueryruns one call per tenant withThreatHunting.Read.All(Microsoft Graph).
| # | Action | Who | Done when | Evidence to capture |
|---|---|---|---|---|
| 1 | Normalize indicators: SHA-256 hashes, IPs, domains, and the behavior as a query predicate. | On-Call Technician | One row per indicator, typed | Indicator list, UTC-stamped |
| 2 | Split the estate into tenant groups of ≤100 and compute the per-tenant row budget before running. | On-Call Technician | Budget (50,000 ÷ N) written down | Group definitions; budget calculation in the ticket |
| 3 | Run the filtered sweep per group. Project few columns. Never use take. | On-Call Technician | All groups returned, row counts recorded per tenant | Query text, run time (UTC), per-tenant counts, TenantId retained |
| 4 | Classify every tenant: hit / no result / no data / not queried. | On-Call Technician | All 180 classified | Classification table appended to the ticket |
| 5 | For any tenant with a hit, open a separate ticket. Do not append to Client M's. | On-Call Technician | One ticket per affected tenant | Ticket IDs cross-referenced in the master timeline |
#### 00:00 — The harder half: is it us?
Now the question you must answer before you wake anyone.
This is not paranoia, it is the documented pattern. In the DragonForce case, the MSP's own SimpleHelp instance was the entry point; the actors then "gathered information on multiple customer estates managed by the MSP, including collecting device names and configuration, users, and network connections" and pushed a malicious installer through the RMM to client endpoints (Sophos X-Ops). Microsoft's name for the strategy is "compromise-one-to-compromise-many" (Microsoft).
| # | Action | Who | Done when | Evidence to capture |
|---|---|---|---|---|
| 1 | Delivery channel. Was the first malicious execution a child of your RMM agent, remote-control agent, or a script your platform pushed? | On-Call Technician | Parent process identified as MSP tooling or not | Ancestry query text and results; agent inventory for that host |
| 2 | Job history. Any script, job, patch or installer pushed to this tenant in 72 hours with no matching change ticket? Any push to many devices at once? | On-Call Technician | Every push matched to an approved change, or flagged | RMM job history scoped to the site; change-ticket cross-reference |
| 3 | Console integrity. Can every technician still log into the RMM and remote-control consoles? Any console admin account nobody created? Any gap in the console's own logs? | MSP Incident Commander | All three answered in writing | Console user list; console application logs for the window; note recording any log gap |
| 4 | Technician identity. Any console or PSA sign-in outside the on-call roster? Any remote-control session at an odd hour? Any failed authentication immediately following a password change? | On-Call Technician | Roster compared against 7 days of sign-in activity | Console and PSA sign-in export; on-call roster |
| 5 | Partner-tenant path. In the client tenant, any privileged change attributable to a partner identity with no corresponding interactive sign-in? | On-Call Technician | Entra audit and sign-in logs compared for the window | Entra audit export; Entra sign-in export; object-ID → technician mapping |
| 6 | Shared credentials. Is any account used in the lateral movement one your MSP also uses at other clients? | MSP Incident Commander | Every account classified client-only or MSP-shared | Credential inventory extract; vault access log |
Three of those need explaining, because the mechanics are counter-intuitive.
Row 4 is CISA's "intrusion canary." AA22-131A tells MSPs to review logs for "unexplained failed authentication attempts—failed authentication attempts directly following an account password change could indicate that the account had been compromised" (AA22-131A). One query, and the cheapest tell you own.
Row 5 is the delegated-admin trap. Partner access performed via PowerShell produces no sign-in record in the customer tenant — only the resulting modifications reach the customer's audit log, while portal and API access do produce sign-ins (AADInternals). So a privileged change attributable to a partner identity with no matching interactive sign-in is consistent with programmatic partner-side access: either your own automation, or somebody holding your partner credentials. Know which. And when you look, know what you are looking for — partner sign-ins render in the customer's logs as "{Governing tenant name} Technician", with a username of the form user_<object ID, dashes removed> (Microsoft). The customer-side log cannot name your individual technician. That mapping lives in your tenant, and only you can produce it.
Row 3 has history behind it. In the Kaseya intrusion the attackers deleted IIS logs and the logs held in the application database as a first stage (Truesec). In the ScreenConnect authentication-bypass campaign the exploit overwrote the internal user database, deleting all local users except a newly created admin (Huntress) — meaning "I can't log into my own console" is not an outage, it is an indicator of compromise. If your console logs have a hole at exactly the right time, treat yourself as patient zero until you can prove otherwise.
The answer changes everything downstream. If any check is positive or unanswerable: blast radius becomes platform; the Incident Commander becomes the Practice Owner; and you stop using the RMM as a collection tool, because CISA's instruction to a provider in this position is to isolate the management server or stop the process and contact downstream customers (CISA AA25-163A) — and you cannot both isolate it and collect through it.
Actionable takeaway: write these six checks as a saved query set and a one-page card, in daylight, reachable without the PSA. If answering "is it us" takes more than twenty minutes, that number is your real incident-response time — and every client inherits it.
#00:05–00:20 — Phase 3: Wake the right people, in the right order
Escalation order is content, not etiquette. Wrong order and you either burn twenty minutes briefing someone who cannot act, or you deliver a half-formed alarm to a client executive who calls their lawyer before you have a confirmed fact.
The order: On-Call Technician → MSP Incident Commander → Tenant Lead → Client Authority Holder.
| # | Action | Who | Done when | Evidence to capture |
|---|---|---|---|---|
| 1 | Page the MSP Incident Commander with four lines: what, how many hosts, blast-radius finding, what you have and have not done. | On-Call Technician | Acknowledged, with timestamp | Page record; the four-line brief, verbatim |
| 2 | IC assumes command, confirms or revises severity and blast radius, names a Scribe. Below three responders, the IC is the Scribe. | MSP Incident Commander | Handover recorded with UTC timestamp | Ticket entry naming IC, Scribe, severity, blast radius |
| 3 | Move to an out-of-band channel — a bridge or app not hosted in the affected tenant, and not in your own if blast radius is platform. | MSP Incident Commander | Bridge open, details distributed outside both estates | Bridge details; channel choice and reason logged |
| 4 | Wake the Tenant Lead. They bring context: what M-SQL01 does, whether the line depends on it, who really decides at midnight. | MSP Incident Commander | Tenant Lead on the bridge, or 10 minutes elapsed and the deputy paged | Page records with timestamps, every attempt |
| 5 | Read the tenant's authority tier and contact block before dialling: notification deadline, notice method and address, regulatory flags, insurer and panel counsel. | Tenant Lead | Tier and deadline read aloud on the bridge and recorded | Export of the tenant authority record as it stood tonight |
| 6 | Call the Client Authority Holder — primary, then alternate — on out-of-band numbers. Log number, time, outcome for every attempt. | Tenant Lead | Contact made, or the list exhausted and logged | Call log with UTC timestamps per attempt |
| 7 | Start the unreachable clock. At the pre-agreed elapsed time with no acknowledgment, the pre-authorized tier expands per contract. | MSP Incident Commander | Clock started, expiry time recorded | Ticket entry naming the clock, its length, and the clause it rests on |
Step 6 is the one that fails tonight. It is midnight, the primary's phone is on silent, and the alternate — if one is even named — left in March. That is the normal case, not the exception, and Scenario Two is entirely about it.
The asymmetry runs the other way too. In ACE American Insurance Co. v. Congruity 360 and Trustwave Holdings, an insurer that had paid its insured sued the technology vendors in subrogation — and the pleaded negligence against one was failure to properly notify appropriate parties, preventing timely proactive action and increasing damages (Hunton). Not failure to detect. Not failure to contain. Failure to tell someone. That is the On-Call Technician's specific exposure, and the only defense is a call log with timestamps.
#00:20 onward — Phase 4: Containment within your authority
Now you act — on everything you found at once, not in the order the console offers.
The governing distinction is not technical capability. You have the credentials for all of it. The distinction is authority, and it is contractual.
| Action | Typical tier under a managed-EDR agreement |
|---|---|
| EDR-isolate a single endpoint | Pre-authorized |
| Kill a process, quarantine a file, block a hash | Pre-authorized |
| Disable one compromised user and revoke its sessions | Pre-authorized |
| Block a specific IP, domain or URL at egress | Pre-authorized |
| Reset one privileged credential | Pre-authorized |
| Isolate a server or hypervisor host | Approval required |
| Isolate a domain controller | Approval required |
| Tenant-wide password reset or mass session revocation | Approval required |
| Take a production application or line offline | Executive approval |
| Restore from backup over live data | Executive approval |
That table is illustrative. The authoritative version is per-client, and Part 2's Client Authority Matrix is where each tenant's real tiers, approvers, deputies, out-of-band numbers and unreachable clock live. Nothing here overrides it.
One structural idea is worth stating plainly, because it is the best control in this whole book: make the RBAC role you hold in each tenant match the authority tier that tenant agreed to. Microsoft implements exactly this for its own managed service — grant Defender Experts Security Operator and "the experts can perform the required response actions on the incident on your behalf"; leave them at Security Reader and the same recommendations appear as pending actions awaiting the customer (Microsoft, Defender Experts for XDR). A tenant sitting at "approval required" should not have your standing account holding Global Administrator, because at 00:37 under pressure the technical capability wins the argument with the PDF. Every time.
| # | Action | Who | Done when | Evidence to capture |
|---|---|---|---|---|
| 1 | Trigger Collect investigation package on every affected host before isolating it. It is initiated remotely through the EDR service — no interactive session, no hands-on access — but it executes on the host, which is why the summary report records each command and its status. | On-Call Technician | Requested on every host; failures noted | Package ZIP per host, including CollectionSummaryReport.xls; API response codes |
| 2 | Isolate every confirmed host in one burst, not sequentially. Record submission time for each. | On-Call Technician | All confirmed hosts submitted inside one 5-minute window | Isolation records with UTC timestamps; Action center export |
| 3 | For servers, DCs, DNS and DHCP use selective isolation — isolation with exclusions, so the named services stay reachable — rather than full isolation, and only with approval. | MSP IC (approval); On-Call Technician (execution) | Approval recorded by named role, or action not taken | Approval record naming role, time, scope; the exclusion list used; action record |
| 4 | Contain identity in one atomic action per account: revoke sessions and reset the credential and remove attacker-created persistence (inbox rules, forwarding, added MFA methods, app registrations). | On-Call Technician | Every account processed; none left half-done | Per-account action log; Entra audit entries; rules/forwarding export taken before removal |
| 5 | Push the indicator as a block in the affected tenant only, with an expiry set. Not estate-wide yet. | On-Call Technician | Indicators created with expirationTime | Indicator export; per-tenant indicator count before and after |
| 6 | Verify by observation: no new token issuance, no new sign-ins, no new executions from contained principals. | On-Call Technician | Two clean checks, ten minutes apart | Verification query text and results, both intervals |
| 7 | Record every action, with the authority relied on, in the minute you take it. | Scribe | Every row in the action log names its authority | Contemporaneous action log with UTC timestamps and an authority column |
A note on step 3, because the two things get confused at exactly the wrong hour. Selective isolation is the action your technician takes — isolation with exclusions, supported on the server platforms you care about, and the one Microsoft points to where a web proxy would otherwise stop a device recovering from isolation (response actions on a device). Granular critical-asset containment is not. The narrow, port-and-direction containment that lets a domain controller, DNS or DHCP server keep answering is something the platform's own automatic containment may apply after it decides a device is a critical asset — without your responder choosing it and without an approval attached. Expect it, record it in the action log if it happens, and do not write it into a step table as a thing you execute.
Four mechanical facts decide whether this works, all documented, all of them things people learn at the worst moment:
- Rotating the password before revoking tokens fails. A refresh token is an independent bearer credential; Entra "can't directly revoke a session token issued by an application" (Microsoft), CAE propagation takes up to 15 minutes, and access tokens in CAE sessions stay valid for up to 28 hours (Microsoft, CAE). Reset-only leaves the adversary logged in and the user locked out — the worst of both.
- Your role may not be enough to save the client's own admin. Resetting the password of, or invalidating refresh tokens for, a privileged admin in a client tenant requires Privileged Authentication Administrator (Microsoft, GDAP roles by task). If your on-call holds only Help Desk Administrator, they cannot contain a compromised Global Admin at all. Find that out in daylight.
- Isolation has edges. Devices behind a full VPN tunnel cannot reach the EDR cloud service once isolated; proxies can prevent recovery from isolation; isolating a Hyper-V host blocks traffic to all its child VMs; unmanaged-device containment is recommended for no more than 100 devices at a time (Microsoft).
- Indicator blocking has a hard ceiling. "There's a limit of 15,000 indicators per tenant. Increases to this limit aren't supported"; CSV uploads cap at 500 per batch; CIDR is not supported (Microsoft). Set an expiry so incident indicators age out instead of permanently consuming a client's budget.
#Running in parallel from 23:50 — Evidence capture
Evidence collection is not a phase. It is a rail running the whole length of the night, starting before containment and never after it, because almost every artifact you need has a hard expiry and none of it can be created retroactively. The order rule, stated once: preserve, then contain. Holds are not retroactive, retention windows are short, and disabling an account before you have scoped it freezes the telemetry stream you were about to use.
| # | Action | Who | Done when | Evidence to capture |
|---|---|---|---|---|
| 1 | Start a contemporaneous action log at the first alert: UTC, ISO 8601, one row per action — time, action, reason, authority, operator. | Scribe / On-Call Technician | Running before the first containment action | The log itself; it is the primary artifact of the night |
| 2 | Open a per-tenant evidence container named by client and case ID. Never a shared forensics folder with subfolders. | On-Call Technician | Container exists, access granted per-case | Container path, access-control record, creation timestamp |
| 3 | Place the eDiscovery / legal hold in the client tenant before any containment touches mail or files. Only available where the roles were assigned at onboarding — see the note below. | Tenant Lead | Hold applied and timestamped, or recorded as unavailable and escalated | Hold record: case, locations, applied time — or the gap note and the escalation record |
| 4 | Export identity evidence first: Entra sign-ins, non-interactive sign-ins, service-principal sign-ins, audit logs for the window. | On-Call Technician | Exports complete, hashed, in the container | Query text, coverage window (UTC), row counts, SHA-256 per file |
| 5 | Request the EDR investigation package per host before isolating; preserve the command log for any live-response session. | On-Call Technician | One package per host, or a recorded failure reason | Package ZIP; CollectionSummaryReport.xls; live-response session ID and command log |
| 6 | Export your own MSP-side records for the same window: RMM activity log scoped to this site, remote-session audit, PSA tickets, on-call and paging record. | MSP Incident Commander | Exports complete with segment boundaries recorded | Segmented exports plus a segment index proving no gap |
| 7 | Hash every artifact at the moment of export; store the hash list separately. | On-Call Technician | Hash list complete and stored apart | Hash file with algorithm, tool and version |
| 8 | Record every gap: sources off, expired, throttled or truncated, with the reason and the date they would have started. | MSP Incident Commander | Gap list written | Gaps-and-limitations note |
Row 3 has a prerequisite, and it is not one you can satisfy tonight. The hold is placeable only where eDiscovery roles were already assigned in the client's own Purview portal at onboarding, to two named people per case. If nobody did that, there is no hold available to you at 23:50 — write it down as a gap with the time you discovered it, and escalate to the client's own administrator to place it. That is the honest version and it is survivable. Discovering it at 06:00, after containment has already touched mail, is not. Part 2 puts the role assignment where it belongs: on the onboarding checklist, not the incident one.
Why these expire, specifically. Entra sign-in and audit logs are retained 7 to 30 days by license, and diagnostic settings — the only durable route — capture only logs generated after they are configured (Microsoft). Defender advanced hunting sees 30 days, the portal 180, and on contract termination data is deleted no later than 180 days and is unrecoverable (Microsoft) — the sentence that matters when an unhappy client offboards after an incident. Search-UnifiedAuditLog returns a maximum of 50,000 records per run; if you get 50,000 back, records were dropped (Microsoft). Your own RMM log expires too: Datto RMM retains Activity Log data 180 days and caps CSV export at 500 rows per file (Datto RMM) — a completeness problem, not an inconvenience. And eDiscovery exports must be retrieved: the download window is 14 days, and export jobs auto-cancel at 7 days (Microsoft).
Two disciplines that cost nothing and save the case. One clock: UTC, ISO 8601, everywhere — including your PSA and technician workstations. The joint international logging guidance asks for "an accurate and trustworthy time source" used consistently across all systems, with UTC as "the preferred time standard" (Best Practices for Event Logging and Threat Detection). One custodian: the DOJ's position is that "as few employees as practicable should be assigned the responsibility of retaining custody of such information" (DOJ, *Best Practices for Victim Response*). At a fourteen-person MSP, that is one named person and one deputy, decided in advance.
One line about the ticket-writing itself: facts and timestamps only. No characterization of cause, no blame, no "we should have caught this." In Clorox v. Cognizant, the discovery target was the provider's own service-desk transcripts (The Register). Write every line as though a stranger will read it back to you in a room with a court reporter.
#06:30 — Preserving the answer to the client's first question
Some time between 07:00 and 08:30, Client M's owner will be told what happened. Their first question will not be about lateral movement or dwell time. It will be four words: "What did they get?"
You cannot answer that at 07:00 unless you took specific steps at 00:15. There is no retroactive route, and "we don't know" costs you the relationship even when it is honest.
| # | Action | Who | Done when | Evidence to capture |
|---|---|---|---|---|
| 1 | Hold before anything touches mail or files — the same hold as the evidence rail, and the load-bearing step. | Tenant Lead | Hold applied and timestamped | Hold record |
| 2 | Preserve mailbox, file, share and SharePoint/OneDrive access evidence for every implicated account, across the full dwell window — not just tonight — so you can later separate "accessed" from "could have accessed". | On-Call Technician | Export complete, coverage window recorded | MailItemsAccessed and file-access records from the Unified Audit Log (Microsoft Purview auditing); export segments plus segment index |
| 3 | Capture egress evidence — firewall, proxy, DNS — and record the appliance's configured retention as it stands tonight. | On-Call Technician | Logs pulled and retention config captured | Log export plus a dated capture of retention and quota settings |
| 4 | Enumerate what data classes exist on the affected hosts and in the affected accounts — the "what could they have got" side. | Tenant Lead | Data-class list with the query evidence behind each number | Scope-of-exposure note with supporting queries |
| 5 | Write the honest limits alongside the findings: what you can prove, what you can bound, what you cannot see and why. | MSP Incident Commander | Both documents exist before the client briefing | Scope-of-exposure note plus gaps-and-limitations note |
Row 3 is the one that gets skipped and the one that gets weaponized. If the firewall rolled its logs eight hours ago, that is a defensible fact if you can show the retention as configured at the time. If you cannot, the gap looks like concealment rather than a purchasing decision made three years ago by someone who has since left.
Row 5 protects everyone, including you. There is a real difference between "we have no evidence of exfiltration" and "we have evidence that no exfiltration occurred." The first is nearly always true. The second almost never is. Say the first, precisely, and put the second nowhere near an email.
#What should have been true at 23:44
Everything above is recoverable, and every failure point is a preparation task costing a morning and no license fee. Four of them.
1. The Client Authority Matrix existed, was current, and was readable in ten seconds. Per tenant: authority tier per action, primary and alternate approvers with out-of-band numbers, the unreachable clock and what it expires into, the contractual notification deadline and notice address, regulatory flags, insurer and panel counsel. Not a signed PDF in a shared drive — structured fields in the PSA, read at 00:05 rather than hunted for. And the RBAC you actually hold in that tenant matches the tier the client agreed to. Part 2 builds this, field by field.
2. The cross-tenant sweep was pre-written and its row budget pre-computed. Saved queries, tenant groups already defined under the 100-tenant ceiling, the 50,000-divided-by-N arithmetic already on the card, the Graph fan-out already tested, and "no data versus no result" already a column in the output. Nobody should meet the platform's caps for the first time at 00:02 while an adversary works. Part 2 covers the sweep and its honest limits.
3. The "is it us" check was six saved queries and a one-page card, not a research project. It is answerable in twenty minutes only if four preconditions were built beforehand: logs stored where your own RMM cannot reach or delete them; at least six months of retention, per CISA's MSP baseline (AA22-131A); customer data and credentials segregated so blast radius is bounded and attributable; and a mass-scripting safeguard, which is the best MSP-specific control any government body has published — "if an account attempts to push commands to 10 or more devices within an hour, retrigger security protocols, such as multifactor authentication (MFA), to ensure the source is legitimate" (CISA, Guide to Securing Remote Access Software). Part 2 §7 is the whole of this.
4. Evidence collection was already running before the incident started. Entra diagnostic settings streaming into each client's own workspace from onboarding day, because they capture nothing retroactively. Live response enabled per tenant at onboarding, because enabling it needs a permission your on-call does not hold. Monthly exports of delegated-access audit logs, RMM activity logs and the object-ID-to-technician mapping into your own store, because the consoles will have aged them out long before anyone asks. And a per-tenant evidence baseline saying which sources are on, at what retention, since when — which is what converts a gap into an explained gap. Part 6 has the manifest.
One more thing, and it should focus the mind. The ICO's first penalty against a data processor — £3,076,320 against an IT supplier — did not turn on sophistication. The findings were gaps in MFA deployment, insufficient vulnerability scanning and inadequate patch management, including a CVSS 10.0 vulnerability unpatched for two years. The provider had built a working MFA solution and not rolled it out, citing a perception that customers would be unwilling to implement it (ICO enforcement record; Clifford Chance; BleepingComputer).
If you are carrying an MFA exception right now for a client who pushed back, that email thread is exhibit A. Not eventually. Now.
None of the four items needs a SOC, a SIEM budget or headcount you cannot fund. They need one morning each and somebody's name against them. The technician who isolates a domain controller alone at 02:00 with no authority is not a hero — they are an uninsured liability, and the fault is not theirs. Our job as owners and delivery leads is to make sure nobody is ever standing there at 23:45 on a Tuesday holding a phone that nobody answers.
Get the matrix written, get the queries saved, get the logs out of your RMM's reach — and sleep a little better on the nights you are not on call.
#Sources
- Mandiant / Jim Aldridge — *Remediating Targeted-threat Intrusions*, Black Hat USA 2012
- CISA — Federal Government Cybersecurity Incident and Vulnerability Response Playbooks
- NIST SP 800-61r3 — Incident Response Recommendations and Considerations
- CISA AA22-131A — Protecting Against Cyber Threats to MSPs and their Customers
- CISA AA25-163A — Ransomware Actors Exploit Unpatched SimpleHelp RMM
- CISA and partners — Guide to Securing Remote Access Software
- FBI / CISA and partners AA23-320A — Scattered Spider, updated 29 July 2025
- ASD / CISA / FBI / NSA and partners — Best Practices for Event Logging and Threat Detection
- DOJ Cybersecurity Unit — Best Practices for Victim Response and Reporting of Cyber Incidents
- Sophos X-Ops — DragonForce actors target SimpleHelp vulnerabilities to attack MSP, customers
- Sophos — 2026 Active Adversary Report
- Help Net Security — Sophos identity-driven breaches findings
- Huntress — ScreenConnect authentication bypass
- Huntress — N-able N-central exploitation
- Trend Micro — Threat actor groups including Black Basta exploiting ScreenConnect
- Truesec — Kaseya supply-chain attack technical analysis
- Microsoft Security Blog — NOBELIUM targeting delegated administrative privileges
- Microsoft Learn — Defender multitenant management requirements
- Microsoft Learn — Advanced hunting in Defender multitenant management
- Microsoft Learn — Data retention and data security in Microsoft Defender XDR
- Microsoft Learn — Extend Microsoft Sentinel across workspaces and tenants
- Microsoft Graph — security: runHuntingQuery
- Microsoft Learn — Take response actions on a device
- Microsoft Learn — Overview of indicators in Defender for Endpoint
- Microsoft Learn — Revoke user access in Microsoft Entra ID
- Microsoft Learn — Continuous Access Evaluation
- Microsoft Learn — GDAP least-privileged roles by task
- Microsoft Learn — Microsoft Entra data retention
- Microsoft Learn — Monitor governing tenant admin activity in a governed tenant
- Microsoft Learn — Use a PowerShell script to search the audit log
- Microsoft Learn — Auditing solutions in Microsoft Purview
- Microsoft Learn — Limits in eDiscovery
- Microsoft Learn — Defender Experts for XDR managed detection and response
- AADInternals — Microsoft partners: The Good, The Bad, or The Ugly?
- Datto RMM — Activity Log
- EDPB — Guidelines 9/2022 on personal data breach notification, v2.0
- Art. 28 GDPR — Processor
- 18 U.S.C. § 1030 — Fraud and related activity in connection with computers
- Sophos — MDR Service Description
- Google Workspace — Premier terms of service
- Hunton — Cyber insurer sues policyholder's cyber pros (ACE American v. Congruity 360 and Trustwave)
- The Register — Clorox v. Cognizant
- ICO — Advanced Computer Software Group Limited enforcement
- Clifford Chance — ICO fines processor after inadequate security measures
- BleepingComputer — UK fines software provider £3.07 million for 2022 ransomware breach
#Scenario Two — 02:00, The Domain Controller Needs Isolating and the Client Is Unreachable
You can stop the encryption in ninety seconds. Nothing in the contract says you are allowed to.
Two hours ago this was Scenario One: credential dumping and lateral movement at a 220-seat manufacturer, alerts stacking, nobody at the client awake. If you are reading this cold you need only three facts. You are the On-Call Technician at a fourteen-person MSP with 180 clients. You hold Domain Admin in this environment because that is how the account was set up five years ago by someone who no longer works here. And the master services agreement says you provide "monitoring and management services," which is a description of a subscription, not a grant of emergency power.
Every MSP owner has thought about this night. Almost none have written it down. It is not a hard technical problem — isolating a domain controller is a two-minute job you have done in daylight without breaking a sweat. It is an authority problem, and authority problems have a nasty property: they get harder the longer you look at them, and you are looking at this one at two in the morning with a progress bar running in the other window.
A necessary note first. What follows is an operating framework built from published contracts, regulator guidance and statutory text. It is not legal advice, I am not a lawyer, and nothing here substitutes for counsel who has read your agreements. Where I describe contract language it is illustrative — the shape of a clause, not a clause.
#01:57 — What the console is telling you
Three alerts land within forty seconds, all from the same tenant. A shadow-copy deletion on DC01. A backup job failure on the Veeam server — not a timeout, an authentication failure, which is a different and much worse kind of failure. And a service-account logon from a workstation that has no business talking to a domain controller.
You do not need a threat-intelligence subscription to read that. CISA and the FBI's Akira advisory documents exactly this tradecraft: deleting Volume Shadow Copy Service copies, and going after the backup platform specifically with VeeamHax.exe, described there as a "plaintext credential leaking tool," and Veeam-Get-Creds.ps1 for "obtaining and decrypting accounts from Veeam servers" (CISA AA24-109A). Mandiant calls the wider pattern recovery denial — operators deliberately targeting backup infrastructure, identity services and virtualization management planes, attacking your ability to recover rather than only your ability to operate (M-Trends 2026).
Translation: someone is taking the fire exits off the building before they light the match.
The clock is not on your side. Secureworks' widely cited benchmark puts median time-to-encryption under 24 hours from initial access, with 10% inside five hours (Computer Weekly / Secureworks) — 2023 reporting with no directly comparable 2026 update, so treat it as a floor rather than a forecast. Sophos's 2026 dataset adds the part that makes this your problem: 88% of ransomware encryption happens outside business hours (Help Net Security on Sophos). Attackers pick 02:00 for the same reason you dread it. The people who can say yes are asleep.
At 02:00 you call the client's IT director. Voicemail. At 02:02 you call the CEO's mobile, the only other number in the record. Voicemail, and the greeting says he is on leave.
You now have every capability you need and no permission you can point to.
#The wrong instinct, named
Here is what a good technician does at 02:03, and I want to name it plainly because most of us would do it: you use the Domain Admin account, because you have it.
Not recklessly. You reason your way there in eight seconds — they are about to lose everything, I can stop it, waiting is obviously worse, and they gave me these credentials for exactly this. Every step feels correct. The conclusion is still a problem, and not for the reason you would guess. Isolating the DC may well be the right call. The problem is that "I had the credentials" is doing all the work in that reasoning, and it is the one part that carries no weight in the two places this decision gets reviewed: your client's boardroom, and — if it goes badly — a claims file.
There is a second instinct, the mirror image: you do nothing, because nobody said you could. You keep dialling, you write a long ticket note, you wait for 08:00. That one is worse, and on the current evidence it is the one more likely to get an MSP sued.
Takeaway: the question at 02:03 is not "can I?" and not "am I allowed?" It is "what makes this defensible either way, and what must I do in the next four minutes to make it so?"
#Why "I had the credentials" is not a defense
The law here is thinner than anyone would like, so take both directions and see the shape.
#### If you act
The contract. Two current, publicly posted MSP master agreements were read in full for this book and neither contains an emergency action clause. CompassMSP's MSA grants exactly two unilateral powers — service changes on thirty days' written notice, and suspension for non-payment — and states at §10.1 that services are rendered as an independent contractor and the agreement "does not create an employer-employee relationship" (CompassMSP MSA). Secure Data Technologies goes further in one direction, listing "Automated endpoint and identity isolation in the event of a detected threat event" as an included deliverable at §1.5.1 — but frames it as a service description, not a grant of authority, and then excludes "incident response beyond initial triage and alerting" as a separate billable engagement (Secure Data MSA, June 2025). If your MSA resembles either, it is a scope, payment and liability document. It is silent on tonight.
The agency argument, thinner than it sounds. Agency law does contain an emergency doctrine: where "unforeseen circumstances arise and it is impracticable to communicate with the principal," an agent "may do what is reasonably necessary in order to prevent substantial loss to his principal" (Saylor, *Law for Entrepreneurs*, restating Restatement agency doctrine). The catch: most MSAs expressly disclaim agency. Do not plan a night around emergency agency authority when your own contract says you are not an agent.
The statute nobody warns MSPs about. After Van Buren v. United States, 593 U.S. 374 (2021), the CFAA's "exceeds authorized access" route is largely closed to contract-based theories, and DOJ will not charge on the theory that authorization "was conditioned by a contract or policy" (Van Buren summary; DOJ Justice Manual 9-48.000). But 18 U.S.C. § 1030(a)(5)(A) is not an access provision. It reaches whoever "knowingly causes the transmission of a program, information, code, or command, and as a result of such conduct, intentionally causes damage without authorization." "Damage" at (e)(8) is "any impairment to the integrity or availability of data, a program, a system, or information"; "loss" at (e)(11) expressly includes "any revenue lost, cost incurred, or other consequential damages incurred because of interruption of service"; and § 1030(g) gives a private civil action where loss reaches $5,000 in a year (18 U.S.C. § 1030). Your admin credentials do not answer this, because (a)(5)(A) asks whether the damage was authorized, not whether the access was. No court has been found applying it to an IT provider doing incident response — the analysis comes from the statutory text, not a case — but it is the sharpest edge in the room.
For EU and UK clients, add Article 28(3)(a) GDPR, which requires a processor to act "only on documented instructions from the controller"; Article 82(2), which fixes processor liability where it "acted outside or contrary to lawful instructions of the controller"; and Article 28(10), under which a processor determining purposes and means "shall be considered to be a controller" (Art. 28 · Art. 82).
And note the quiet irony: your own limitation-of-liability clause is your main defense to a wrongful-shutdown claim, because business interruption and lost profits are exactly what it excludes. The danger is the carve-out. Caps are near-universally disapplied for gross negligence and willful misconduct, and a deliberate shutdown is an intentional act. Whether it is willful turns on whether it was reasonable and in good faith — which is decided by what you wrote down at 02:22.
#### If you do not act
ACE American Insurance Co. v. Congruity 360, LLC and Trustwave Holdings, Inc., No. 2:25-cv-15657 (D.N.J., filed 15 September 2025): a cyber insurer that paid $500,000 to its insured after a ransomware attack, then sued its insured's technology vendors in subrogation. The allegation against Trustwave, as pleaded, is that the vendor failed to properly notify appropriate parties of the incident, preventing timely proactive action and significantly increasing damages (Hunton · National Law Review). Nothing is adjudicated; these are allegations in a live complaint. But read what is pleaded: not failure to detect, not failure to contain. Failure to tell someone.
Mastagni Holstedt, A.P.C. v. LanTech, LLC — a Sacramento law firm suing its IT provider and backup vendor for negligence and breach of contract after a Black Basta attack, over $1 million sought, filed February 2024 (MSSP Alert). The industry body's own commentary focuses on the reported absence of a signed agreement (MSPAlliance). Docket number and status could not be confirmed, so take it as an illustration of what gets pleaded, not as a holding.
The practical asymmetry. Failure-to-act claims are being actively pleaded today, by subrogating insurers, with real dockets. Acting-without-authority claims against MSPs are largely theoretical and land inside a cap you drafted. That should inform your default — but it is not license to skip the paperwork, because the willful-misconduct carve-out and § 1030(a)(5)(A) are exactly where a bad night stops being insured.
#### The action-by-action reading
Analysis, not a quotation. Assumes a typical MSA with admin access granted, no emergency clause, and an independent-contractor clause.
| Action | Position under a silent MSA | Principal risk |
|---|---|---|
| Isolate a single endpoint via EDR | Strongest case. Commonly an included deliverable, and reversible | Low. Document and notify |
| Disable one compromised user account | Strong. Incidental to identity administration you already perform | Low |
| Block a specific IP or domain at egress | Defensible if narrow and reversible; not if it severs a production integration | Contract; § 1030(a)(5)(A) if it impairs availability |
| Tenant-wide password reset or mass session revocation | Weak. Tenant-wide user impact equals business impact | Breach of contract; business interruption |
| Isolate a domain controller | Weak. Authentication outage for the whole organization | Contract, tort, CFAA damage exposure |
| Shut down a production service or host | Weakest. This is the client's business decision | Contract, business interruption, willful-misconduct carve-out |
| Restore from backup over live data | Weakest of all. Destroys evidence and possibly current data | Spoliation, conversion, contract |
Actionable takeaway: reversible and single-asset, act and document. Tenant-wide or production-affecting, you need a name — a standing pre-authorization or a live approval from a listed approver. Isolating a DC sits on the wrong side of that line, which is why the rest of this section exists.
#The ladder — five rungs, in order, starting now
You have roughly twenty minutes of useful decision time. Spend it climbing, not agonizing. Each rung either resolves the problem or builds the record that makes the next rung defensible.
#### Rung 1 — 02:03 to 02:07: exhaust the authority you already have
Before you touch the DC, do everything you are unambiguously allowed to do. This is not stalling. It is the strongest fact in your later account: you took the least-disruptive effective action available, and escalated only when it was not enough.
| # | Action | Who | Done when | Evidence to capture |
|---|---|---|---|---|
| 1 | Isolate the workstation identified as the source of the DC logon, via EDR | On-Call Technician | Device shows isolated in the console | Action ID, UTC timestamp, operator identity, console name |
| 2 | Disable the compromised service account and revoke its sessions and tokens | On-Call Technician | Sign-ins fail; sessions terminated | Directory audit entries for disable and revoke, with UTC times |
| 3 | Block the identified C2 destinations at the client's egress point | On-Call Technician | Rule active and verified | Rule text, change record, config before and after |
| 4 | Collect the EDR investigation package from the DC and the source workstation | On-Call Technician | Packages downloaded and stored | Package plus CollectionSummaryReport.xls, which records the command used per data point and any error codes |
| 5 | Determine whether any immutable restore point predating the intrusion still exists | On-Call Technician | Restore-point inventory captured | Restore points with creation timestamps, immutability state, expiry |
| 6 | Confirm this is one tenant, and that it did not arrive through your RMM, PSA or a technician identity | On-Call Technician | Cross-tenant sweep complete | Query text, scope filter, run time, result counts |
Read the order of that table deliberately, because it inverts the rule Scenario One states. There the rule is preserve, then contain — holds are not retroactive and disabling an account before you have scoped it freezes the telemetry you were about to use. Here containment runs first, because destruction is in progress and the thing being destroyed is the evidence. That inversion is a choice, not an oversight, and it costs you something specific: once the service account is disabled you get no further sign-in records from that principal, so you lose the ability to watch which other footholds it pivots to. Write that down as a known gap at 02:07 rather than discovering it at 09:00. And step 4 is not optional or deferred — network isolation retains connectivity to the EDR service, so the investigation package can and must still be requested for every host touched in steps 1 to 3.
Step 6 is what separates MSP response from ordinary incident response; Scenario One covers it in full. If the answer is "it came through us," stop — that is a platform-level incident and a different playbook.
Step 5 changes the decision you are about to make. An immutable pre-intrusion restore point drops the cost of losing DC01 enormously, and with it the argument for aggressive unilateral action. No backups means the opposite. Know that answer before you decide, not after.
#### Rung 2 — 02:04 onward, in parallel: escalate inside your own MSP
Say this plainly, because somebody needs to: this decision is above the On-Call Technician's pay grade, and that is a feature of the design, not an insult to the technician.
NIST requires an incident response policy to state "roles, responsibilities, and authorities, such as which roles have the authority to confiscate, disconnect, or shut down technology assets" (NIST SP 800-61r3, §2.3). NCSC is blunter: decision-makers "must hold actual authority to approve major actions like taking systems offline," and deputies must be named for when primaries are unreachable (NCSC).
At a fourteen-person MSP those may be two people, one of whom is you wearing a second hat. Fine. The point is not headcount; it is that a named person other than the one holding the mouse says yes, and that yes is timestamped. The cheap version costs nothing: a second phone number in the on-call runbook and a rule that it always gets called.
#### Rung 3 — 02:05 to 02:20, in parallel: work the client contact tree properly
"I called and got voicemail" is not a contact tree. This is.
| # | Action | Who | Done when | Evidence to capture |
|---|---|---|---|---|
| 1 | Call the primary Client Authority Holder. If voicemail, leave the nature, the urgency, the action proposed and a call-back deadline | On-Call Technician | Call placed and logged | Time (UTC), duration, outcome, message content |
| 2 | Send the same content by SMS and by the out-of-band channel that does not depend on their M365 | On-Call Technician | Sent and logged | Message text, send timestamp, delivery status |
| 3 | Call the named deputy | On-Call Technician | Call placed and logged | As above |
| 4 | Call the documented out-of-hours path — operations manager, site number, answering service | On-Call Technician | Path exhausted | Each attempt logged separately with UTC times |
| 5 | Call executive escalation for a decision that stops business | MSP Incident Commander | Call placed and logged | As above |
| 6 | Declare the Authority Holder unreachable and start the unreachable clock | MSP Incident Commander | Declaration recorded | Ticket entry naming who declared it and on what basis |
Two disciplines inside that table. Use the phone, out of band. CISA is explicit: use out-of-band methods such as phone calls, because attackers "may monitor your organization's activity or communications to understand if their actions have been detected," and tipping them off "could cause actors to move laterally to preserve their access or deploy ransomware widely prior to networks being taken offline" (CISA — I've Been Hit By Ransomware). Emailing the IT director's compromised mailbox at 02:06 to say "we think you have ransomware" is a message to two audiences.
Log every attempt as you make it, not in the morning. ACE American v. Trustwave is, as pleaded, a failure-to-notify case. A call log with UTC timestamps is the artifact that answers it.
#### Rung 4 — 02:20: the imminent-harm reasoning, and its limits
If you have climbed three rungs and are still alone with it, this is the reasoning available to you, and vagueness here is what turns a defensible action into an indefensible one.
A specific, observed, imminent and substantial harm is in progress; the client cannot be reached despite documented and exhausted efforts; and the action proposed is the least disruptive one that will actually stop it. Every clause is load-bearing.
- Specific and observed. Shadow copies deleted, backup authentication failing, service-account logon from an unexpected host. Facts in your console with timestamps — not "it looked like ransomware."
- Imminent. The published time-to-encryption numbers are why this word applies at 02:20 and would not apply to the same indicators at 10:00 with the IT director three feet away.
- Substantial. Loss of the domain and the backups. If your honest assessment is "one file server, restorable from an immutable copy," it does not apply.
- Least disruptive effective. That is the next section. If a narrower action would have worked, you took the wrong one.
The limits, honestly:
- This is not a doctrine you can rely on. It is the structure a reviewer uses to decide whether you acted reasonably and in good faith — which is what keeps a deliberate shutdown out of the willful-misconduct carve-out to your own cap. It is not a permission.
- It does not survive an express prohibition. If the client's matrix says "DC isolation: prohibited without executive approval," that is the answer, and honouring it is the correct outcome.
- It does not extend to irreversible actions. Restoring over live data, wiping, rebuilding, and paying anything are outside it permanently.
- It weakens every hour you have known. Emergency reasoning is available to someone surprised at 02:00. It is not available to someone who saw the first indicator at 21:00, did not escalate, and reached for it at 02:00.
#### Rung 5 — if you act: what happens simultaneously
Not afterwards. Simultaneously. The action and the record are one step. Separate them and the record gets written by a tired person reconstructing from memory at 09:00 — exactly what DOJ guidance tells you to avoid: keep "a contemporaneous written record of all steps undertaken," because it "will minimize the need to rely solely on the recollections of personnel to reconstruct the order of events" (DOJ, *Best Practices for Victim Response and Reporting of Cyber Incidents*). The next two sections are that step.
#The proportionality ladder
"Isolate the DC" is not one action. It is at least six, with wildly different blast radii, and two of them destroy evidence you will need in about nine hours.
| Option | What it does | Business impact | Evidence impact | Tip-off risk | Reversible |
|---|---|---|---|---|---|
| Block specific egress destinations | Severs the identified C2 path only | Minimal | None | Low–moderate | Yes, instantly |
| Disable compromised accounts, revoke sessions | Removes the identity in use | Low, unless the account runs production services | None | Moderate | Yes |
| Selective isolation of the DC | Network isolation with defined exclusions, so the connectivity you named in advance survives; supported on Windows Server 2012 R2 and later | High, but narrower than full isolation, and lower the better your exclusions match what the site actually needs | None; host stays up | High | Yes |
| Full EDR network isolation of the DC | Disconnects the device from the network while retaining connectivity to the EDR service | High — authentication, DNS, DHCP for the site | None; host stays up, volatile state survives | High | Yes, and it auto-lifts after seven days |
| Physically or virtually disconnect | Pull the cable or remove the virtual adapter | High, as above | None, but you lose remote collection entirely | High | Yes, with hands on site or hypervisor access |
| Shut the DC down | Powers off the machine | High, plus restart risk | Destroys volatile evidence — memory, processes, network state, cached credentials | High | Only in the sense that you can power it back on |
Full isolation is not a shutdown, and that distinction is the whole game. Microsoft's device isolation "disconnects the compromised device from the network while retaining connectivity to the Defender for Endpoint service" — you keep investigative reach into the host while removing the attacker's (Take response actions on a device). RFC 3227 states the rule in one line: "Don't shutdown until you've completed evidence collection," and orders volatility from registers and memory down through disk to archival media (RFC 3227). Powering off a DC mid-incident buys the same business impact as isolation with strictly worse forensics. There is almost never a reason to do it.
Do not go looking for a critical-asset containment button, because there is not one. Microsoft does document granular critical-asset containment, and it does what the name suggests: it blocks only specific ports and directions, and it is built for domain controllers, DNS and DHCP servers so they keep running while the attacker's path is cut. But it is applied automatically. Attack disruption flags a malicious device, identifies its role, and applies a matching containment policy on its own. There is no operator procedure for it, and the analyst-invoked "contain device" action sitting next to it in the console applies to unmanaged devices — which your onboarded DC is not. Send a technician looking for it at 02:24 and they spend four minutes hunting a menu while shadow copies keep disappearing. Takeaway: the narrow option you can actually choose on a managed DC is selective isolation — network isolation with defined exclusions, supported on Windows Server 2012 R2 and later, so the connectivity you named survives while the rest of the path is cut. Set those exclusions in daylight, per tenant; at 02:24 you get whatever you configured in advance and nothing else.
Know the constraints before 02:24, because several will surprise you. Isolation auto-lifts after seven days. If the device is offline when you submit, the platform retries for up to three days and then you must reissue. Web proxies set by PAC, WPAD or static config can prevent a device recovering from isolation, which is why Microsoft points those environments at selective isolation instead. Devices behind a full VPN tunnel cannot reach the EDR cloud once isolated, so split tunnelling for EDR traffic is a prerequisite, not an optimization. On Linux, isolation is released if an administrator modifies or adds an iptables rule. And isolating a Hyper-V host blocks network traffic to all of its child VMs — which is how a technician intending to isolate one server takes down eleven. Automatic attack disruption will not cover this box either: its automatic isolation reaches workstations only, not servers (respond-machine-alerts).
Tipping off cuts differently tonight than in Scenario One. The correct default is Mandiant's: delay disruptive action until you can eradicate completely, because piecemeal containment makes responders "tip their hand," after which the attacker abandons the burned infrastructure and keeps the access you never found (Aldridge, *Remediating Targeted-threat Intrusions*, Black Hat USA 2012). Aldridge is equally explicit that this is not universal — whack-a-mole remains correct where harm is occurring in real time, his example being cash being stolen as you watch. Shadow copies being deleted is that case. Scenario One's discipline was observe before acting; this is the documented exception to it, and knowing which one you are in is the judgment the whole night turns on.
#The defensibility record
Print this one. Fill it as you go, in the ticket, in UTC, in ISO 8601, on a system that is not in the affected environment. Whether tonight was reasonable good-faith emergency action or willful misconduct is decided almost entirely by whether these ten fields were filled at 02:22 or reconstructed at 09:00.
| # | Field | What goes in it | Form it must take |
|---|---|---|---|
| 1 | Observed facts establishing imminent harm | The specific alerts, hosts, accounts and behaviors, each with its own timestamp and source console | Facts and IDs only. No characterization, no cause, no blame |
| 2 | Why waiting was not viable | The concrete harm expected and the basis for expecting it | One or two sentences, evidence-referenced |
| 3 | Contact attempts | Every call, SMS and out-of-band message to every listed contact | One row per attempt — name, role, channel, UTC time, outcome — including voicemails |
| 4 | Declaration of unreachable | Who declared it, when, on which exhausted path | Named role, UTC timestamp |
| 5 | Who at the MSP authorized the action | Named individual and role, time, and the channel it came through | Never blank. If nobody authorized it, that is the record |
| 6 | Action taken | Exact action, target, console, operator identity, platform action ID | Copy the action ID from the console; do not paraphrase |
| 7 | Why this was the least-disruptive effective action | The narrower options considered and why each was insufficient | Name the rejected options explicitly |
| 8 | Notification sent immediately afterwards | To whom, when, by which channel, with the text sent | Attach the sent message verbatim |
| 9 | Evidence preserved | Investigation packages, images, log exports | Per artifact: source, scope filter, operator, UTC export time, SHA-256 |
| 10 | Retention holds placed | Which automatic deletions you suspended, when, by whom | Include your own RMM and PSA purge jobs |
A paste-ready version for the ticket:
EMERGENCY ACTION RECORD — [tenant] — [incident ref]
All times UTC, ISO 8601.
OBSERVED: [alert / host / account / time / console]
IMMINENT HARM: [what is being lost, and on what evidence]
CONTACT ATTEMPTS: [name | role | channel | time | outcome] (one line each)
UNREACHABLE DECLARED: [role] at [time], basis: [path exhausted]
MSP AUTHORISATION: [name, role] at [time] via [channel]
ACTION: [action] on [target] via [console], action ID [id],
executed by [operator] at [time]
LEAST-DISRUPTIVE: [narrower options considered and why insufficient]
NOTIFIED: [recipient] at [time] via [channel] — text attached
EVIDENCE PRESERVED: [artefact | source | scope filter | export time | SHA-256]
HOLDS PLACED: [system | what was suspended | by whom | time]
Three disciplines make it worth having.
Everything in UTC, ISO 8601, from a synchronized source. The joint international event-logging guidance is direct: use Coordinated Universal Time "with the year listed first, followed by the month, day, hour, minutes, seconds, and milliseconds (e.g. 2024-07-25T20:54:59.649Z)," from an accurate and trustworthy time source used consistently across systems (Best Practices for Event Logging and Threat Detection). Your PSA renders local time; your EDR console may render the viewer's. Convert once, at entry, and record what you converted from.
Assume every line is read aloud. Start from the correct default: nothing you write is privileged. You are not the client's lawyer, and your tickets, alerts and internal chat are ordinary business records created for a business purpose. In Clorox v. Cognizant — a $380M claim filed July 2025 — the discovery target was the provider's own service-desk transcripts, showing agents resetting passwords and MFA without identity verification (The Register). Facts and timestamps only. No speculation about cause, no admissions, no "we should have."
Record who you are, in a form the client's logs cannot. Working through delegated admin, the client's own sign-in and audit logs show your technician as a display name of the form "{Governing tenant name} Technician," with a username of user_ followed by an object ID with the dashes removed — Microsoft confirms this is by design (GDAP FAQ · Monitor governing tenant admin activity). The mapping from object ID to human being lives in your tenant and nowhere else. Keep a standing exported object-ID-to-technician table, refreshed monthly. Without it nobody — including you — can prove which of your people did what on the client's DC at 02:22.
#What the contract should have said
Four clauses would have turned tonight from an ordeal into a procedure. All language below is illustrative shape only — what the clause must accomplish, not text to sign. Have counsel draft it. The full library is in Appendix B; the operational artifact behind it is the Client Authority Matrix in Part 2.
#06:40 — the morning after
At 06:40 the IT director calls back. His phone was on silent. He is going to be embarrassed about that, which matters more than you would think, because embarrassed people look for somewhere to put the feeling.
Run the conversation in this order.
- Current status first. "The environment is contained. Encryption was stopped. Here is what is up and what is down right now." Nobody hears anything else until they know whether the building is still on fire.
- What you did and when, as a log, not a narrative. Read the record. Times and actions, in order.
- The authority you acted under, without defensiveness and without apology. "At 02:19 we exhausted the contact path in your tenant record. At 02:21 our Practice Owner authorized network isolation of DC01 as the least disruptive action that would stop the destruction of your backups. Here is the record."
- What you did not do. You did not shut anything down. You did not restore over live data. You did not pay or contact anyone. Those were never yours to decide, and saying so establishes the boundary you respected.
- What you need from them now, as decisions with deadlines. Who is today's Authority Holder. Whether they are engaging counsel and their insurer. When the DC comes out of isolation, and who signs that off.
- The notification obligations you both now have, without guessing at their regulatory position. If you process personal data for them under EU or UK law, your duty is to notify them without undue delay, and the EDPB is explicit that the processor "does not need to first assess the likelihood of risk arising from a breach before notifying the controller" (EDPB Guidelines 9/2022 v2.0, ¶44). Their clock starts, in principle, when you tell them. Part 4 has the mechanics.
- No speculation about cause. Not the initial vector, not whose fault it is, not what should have been patched. You will know in two days. Say that.
When it went well, the conversation is short and the record does the work. Somebody wrote it all down at the time, a named person other than the technician authorized it, and the action was visibly the narrowest one that would have worked. Clients are far more reasonable about a disruptive night than folklore suggests, provided the disruption is explicable and the explanation is not being invented in front of them.
When it went badly, it goes badly in one of three shapes. If the record is thin — attempts unlogged, no MSP authoriser named — the conversation becomes about your process rather than their incident, and you will not get it back. If you overreached, shutting down more than you needed or touching production, the conversation becomes about the outage and the encryption becomes background. And if you under-reacted, waiting politely until 08:00 while the estate encrypted, the conversation is about why nobody woke anybody up — which is the specific allegation currently pleaded against a security vendor in a live subrogation suit.
#The uncomfortable truth
The technician who isolates a client's domain controller alone at 02:00, on their own judgment, with no framework and no named authoriser, is not a hero. They are exposed. They made a defensible call, probably the right one, in a way that leaves them personally holding a decision that belongs to the business — and if it turns out badly, the first question anyone asks will be about them rather than about the arrangement that left them there.
That is not a criticism of the technician. It is a criticism of every MSP that has not built the alternative, and I do not exempt anyone who has ever signed off an on-call rota and hoped, including me. The industry has been quietly running this risk on the goodwill of whoever is on the rota, and the goodwill has held up remarkably well. It should not have to.
The job is to make sure nobody is ever standing there alone. That needs no SOC, no platform and no budget. It needs an authority matrix with a second name on it, a second phone number that always gets called, an agreed number of minutes after which "unreachable" becomes a fact, and a ten-field record that takes two minutes to fill in. Four things, none of them expensive, all of them required to exist before 02:00 — because at 02:00 the only thing you can build is the record.
Actionable takeaway: pick one client this week — your largest, or the one whose downtime would hurt most — and get a named Authority Holder, a named deputy, and an agreed unreachable window into the tenant record. One client. This week. Then the next one at their renewal.
#What should have been true at 01:57
One — a completed Client Authority Matrix for this tenant, with one row per disruptive action marked pre-authorized, approval-required or prohibited; a named Authority Holder and deputy; an out-of-hours path with the date it was last tested; and an agreed number of minutes after which non-response becomes standing consent for the tier-two actions. That artifact turns tonight into a lookup instead of a judgment. Template in Part 2, worked example in Appendix A.
Two — an emergency action clause in the MSA, bound to the DPA as a documented instruction, with a minimum-extent-and-duration limiter, a post-hoc notice clock, a good-faith carve-in and a refusal shield. Part 2 covers what each clause does; Appendix B has the illustrative drafting language for counsel to work from.
Three — RBAC in the tenant that matches the authority tier the client agreed to. This is the idea worth stealing from the vendors: Microsoft implements the difference between "recommend" and "act" not in paperwork but in the role the provider holds — Security Reader surfaces recommended actions as pending, Security Operator lets the provider execute them (Defender Experts for XDR permissions). If a tenant is at "approval required," your standing account should not be sitting at Global Administrator, because at 02:00 under pressure the technical capability beats the contractual restriction every single time. Enforce the matrix in roles, not only in a PDF.
Stay contained, stay documented, and get the second phone number into the runbook before the next Tuesday night decides to test you.
#Sources
- CISA/FBI — #StopRansomware: Akira Ransomware (AA24-109A)
- Mandiant / Google Cloud — M-Trends 2026
- Computer Weekly / Secureworks — ransomware dwell times measured in hours
- Help Net Security on Sophos — identity-driven breaches report
- CompassMSP — Master Service Agreement
- Secure Data Technologies — Managed Services Agreement (June 2025)
- Saylor, *Law for Entrepreneurs* — Principal's Contract Liability
- Van Buren v. United States, 593 U.S. 374 (2021) — summary
- DOJ Justice Manual 9-48.000 — Computer Fraud
- 18 U.S.C. § 1030 — Cornell LII
- Article 28 GDPR
- Article 82 GDPR
- EDPB Guidelines 9/2022 v2.0 — personal data breach notification
- Hunton — cyber insurer sues policyholder's cyber pros (ACE American v. Congruity 360 & Trustwave)
- National Law Review — Cyber Insurer Sues Policyholder's Cyber Pros
- MSSP Alert — MSP sued by law firm over Black Basta ransomware attack
- MSPAlliance — When MSPs Get Sued
- NIST SP 800-61r3 — Incident Response Recommendations and Considerations
- NCSC — Cyber incident response processes
- Sophos — MDR Service Description
- CISA — I've Been Hit By Ransomware
- Microsoft Learn — Take response actions on a device (Defender for Endpoint)
- RFC 3227 — Guidelines for Evidence Collection and Archiving
- Mandiant / Aldridge — Remediating Targeted-threat Intrusions, Black Hat USA 2012
- DOJ — Best Practices for Victim Response and Reporting of Cyber Incidents
- ASD/CISA/FBI/NSA and partners — Best Practices for Event Logging and Threat Detection
- The Register — Clorox v. Cognizant
- Microsoft Learn — GDAP FAQ
- Microsoft Learn — Monitor governing tenant admin activity in a governed tenant
- AXIS Cyber Insurance Policy specimen (AXIS 1012561 0120).pdf)
- Google Workspace — Emergency Security Issue terms
- Blount — Senate HSGAC testimony, 8 June 2021 (Colonial Pipeline)
- Microsoft Learn — Defender Experts for XDR permissions
#Scenario Three — The Monday After, The Client's Auditor Wants the Evidence
#09:12 — three emails, one Monday
The incident was six weeks ago. By the standards of this trade it was a good one: business email compromise at Client A, a 140-seat professional services firm, caught on a Tuesday, contained by Wednesday, notified, closed. Everyone was gracious. Someone sent doughnuts.
Then Monday happens.
09:12. Client A's finance director forwards their external auditor's document request list. Twenty-two numbered items, two-week return date. Items 1 to 9 are about Client A's environment. Items 10 to 22 are about yours.
09:40. The client's cyber insurer's forensic accountant wants the recovery timeline, service by service, with restoration times, to calculate business interruption. They also want the date and time of discovery — meaning the exact moment your monitoring or your service desk first had the thing in hand.
11:20. An email from a law firm you have never heard of, on behalf of Client A, headed Preservation Notice. It lists systems, a date range, and custodians whose records must not be deleted. Two of the named custodians are your technicians.
Nobody has accused you of anything. Nobody needs to. Three proceedings have opened in parallel — the client's audit and regulatory track, the client's insurance claim, and a quiet assessment of whether the MSP caused or worsened it — and all three run on documents only you can produce. That is the turn this section is built on:
A lot of what they want is not evidence about the client. It is evidence about you.
Your technicians' access records. Your delegation history. Your patch reports. Your alerting configuration, and whether anyone ever switched it off. Your ticket timeline, the only durable record of when a human first looked at the alert. In Clorox v. Cognizant — $380M sought, filed 22 July 2025 — the centre of the pleading was the MSP's own service-desk transcripts, showing agents resetting passwords and MFA without verifying identity (The Register). The discovery target was the help desk. Not the malware.
#### The wrong instinct, and it is a good instinct
At 09:15 a competent service delivery manager opens the incident report written in week two — the tidy eight-page PDF with the executive summary, the timeline, the three remediation recommendations — and sends it to all three requesters with an offer to answer questions.
Generous, fast, and the worst move available. That PDF was written for a business purpose, as a client deliverable, by the party whose conduct is now in question. It characterises cause, its recommendations read as a list of things you should already have been doing, it carries no hashes, no scope filters and no chain of custody, and because you wrote it to be readable it summarises logs instead of producing them. Four audiences, four failures: the auditor cannot test a control from a narrative, the insurer cannot value a loss from a summary, the regulator reads your recommendations as an admission, and opposing counsel has a document you authored, characterizing cause, to put in front of a witness one sentence at a time.
The remedy is not to say less. It is to produce artifacts instead of assertions, scoped per requester, out of a pack you assembled deliberately.
#Four requesters, four different asks
They are not asking the same question in different fonts.
| Requester | What they actually want | What satisfies them |
|---|---|---|
| Client's auditor (SOC 2 / ISO 27001) | Evidence that named controls operated during the incident window — including yours, as a subservice organization | Your SOC 2 Type II with the period covered, a bridge letter for the gap since the report date, ISO certificate with scope statement and SoA, your CSOC list and your CUEC list, plus per-control operating evidence for the window |
| Client's cyber insurer | The discovery timestamp, the cost and effort record, service-by-service restoration times, and proof the controls declared on the application existed | Ticket timeline (first alert → first human → first client contact), restore session records, and a cost log with names, roles, hours and rates |
| Regulator | Broad, historical, mostly not incident-specific: policies, risk analyses, training records, contracts, incident logs | Documents that predate the incident and can be shown to predate it. OCR's stated position is that a short summary report "will likely not be sufficient" to show a risk analysis was accurate and thorough (HHS OCR) |
| Opposing counsel | Your ordinary-course business records — tickets, internal chat, RMM logs, change records — and a witness who will characterize cause | Nothing you volunteer. Produce through counsel, against a scoped request |
Two facts MSP owners routinely under-rate.
The regulator can come for you directly. In March 2025 the ICO fined Advanced Computer Software Group £3,076,320 — its first fine against a data processor — after attackers entered through a customer account without MFA and took data on 79,404 people (ICO). In January 2025 HHS OCR settled with Elgon Information Systems, a business associate, for $80,000 and a three-year corrective action plan over an intrusion through open firewall ports that took six days to detect (Nixon Peabody). And the FTC charged GoDaddy with, among other things, failing to adequately log and monitor security-related events in its hosting environment (FTC). Read that last one twice: inadequate logging is itself the violation, independent of whether it caused anything. Your evidence pack is not only how you prove what happened — its existence is a control you will be graded on.
Opposing counsel does not need your consent to reach your data. FRCP 34(a) reaches ESI in a party's "possession, custody, or control", and control — not location — decides it, so data you hold is generally discoverable from your client where the client has a contractual right to obtain it (Sedona Conference). Your MSA's data-access clause is an e-discovery clause whether you drafted it as one or not. Where you are not a party, Rule 45 subpoenas you directly.
Takeaway: build one pack, then produce subsets of it to each requester under a cover note stating exactly what is included and what is not. One pack, four productions. Never one PDF, four audiences.
#11:40 — the request list that is about you
Items 10 to 22 are the ones most MSPs have never assembled. Not because they are hard, but because nobody has asked before, and because half of them live in consoles that delete on a schedule you do not control.
The uncomfortable part: the joint CISA/NCSC-UK/ACSC/CCCS/NCSC-NZ/NSA/FBI advisory AA22-131A tells customers to demand exactly these records. It recommends storing the most important logs for at least six months, and that contracts require the MSP to "provide visibility — as specified in the contractual arrangement — to customers of logging activities, including provider's presence, activities, and connections" (CISA AA22-131A). Your client's auditor has read it. It is the source of half the list.
#### Technician access into the client tenant
Item 10, usually phrased as "a record of every access by provider personnel into our tenant during the incident window, identifying the individual." Learn how partner sign-ins actually render, because it is not what people expect (Monitor governing tenant admin activity; GDAP FAQ):
| Field | What the customer sees |
|---|---|
| Display name in sign-in and audit logs | {Governing tenant name} Technician — e.g. "Contoso IT Technician" |
| Username | user_{user object ID in the governing tenant, dashes removed} |
| Sign-in log filter | User contains Technician |
| Audit log filter | Initiated by (actor), startsWith the governing tenant name |
| Role needed in the customer tenant to read these | Reports Reader, Security Reader, Security Administrator, Global Reader or Global Administrator |
Microsoft confirms the user_ behavior is by design, and that is the whole point: the customer-side log cannot name your technician. The mapping from object ID to a human with employment dates exists only in your tenant. Without it, the record of who touched the client's environment is a list of hex strings — and an auditor will write that up exactly as unhelpfully as it sounds. Maintain an exported object ID → technician → employment-dates table, refreshed monthly. One script, and it answers a question nobody can answer retroactively once a technician has left.
#### Delegation history — and the one-year cliff
| Source | What it gives you | The expiry that bites |
|---|---|---|
| Partner Center GDAP activity log | Date-Time, Affected customer, Action, Performed by; CSV export | The export defaults to the most recent month — set From/To explicitly. Relationships Expired or Terminated for one year, or Approval pending for 90 days, are cleaned up automatically and become "no longer accessible or visible to the customer or the partner" |
| GDAP relationship records | Roles held, duration, expiry | Requests expire after 90 days unaccepted; maximum relationship duration two years; auto-extend adds 180 days at a time without customer consent, but a relationship carrying Global Administrator cannot be auto-extended |
| M365 Lighthouse audit logs | Create, edit, delete, assign and remote actions across Audit / Graph / Directory / Sign-in tabs, with the full request body; CSV export. Auditing "is enabled for all customers. It can't be disabled" | Time-range filters are last day / 7 days / 30 days only; new logs can take up to an hour to appear |
| Azure Lighthouse — customer Activity Log | "Event initiated by" name for provider actions | 90 days in the portal unless stored elsewhere; provider users and their role assignments do not appear in Access Control (IAM) or via role-assignment APIs |
(GDAP activity logs; Lighthouse audit logs; Monitor service provider activity)
Read the last row again, because it decides whether the audit goes well. When the auditor asks "who had access to our environment" and the client exports their own IAM, you are not in it. The complete answer exists only if you supply the GDAP and Lighthouse side — and an MSP that cannot produce it looks, from outside, exactly like an MSP that is hiding it.
Standing action: export GDAP activity and Lighthouse audit logs to your own store monthly. A 30-day filter and a one-year cleanup mean that by the time litigation starts, the console will not have it, and no support ticket brings it back.
#### The rest of the MSP file
| Record | Why they want it | Where it lives, and what kills it |
|---|---|---|
| Change records for the affected tenant | To test whether an MSP change caused or enabled the incident | PSA change tickets; Intune/Entra audit logs; Lighthouse deployment plans (apply/deploy/validate) |
| Patch and vulnerability evidence | Elgon was open firewall ports; GoDaddy was asset and update management | RMM patch reports; scanner history; exception register with named approver and expiry |
| Alerting configuration and its change history | "Was the alert that would have caught this ever enabled, and did someone turn it off?" | Sentinel analytics rules with version history; Defender custom detection rules (audited into the unified audit log); RMM monitor definitions |
| Ticket timeline: first alert → first human eyes → first client contact | The insurer's discovery timestamp; the regulator's "when did you know" | PSA plus alerting platform. Export with audit trails, time entries, attachments and internal notes — not the customer-facing summary |
| On-call and escalation record | Whether the technician escalated per policy | On-call roster, paging platform, call and SMS records |
| Technician joiners/movers/leavers with GDAP group membership dates | SOC 2 CC6.3 and CC6.8, and "who could have done this" | HR record plus Entra group membership audit |
| MFA and Conditional Access posture for your own admin identities | Advanced was fined over an entry point without MFA | Entra CA policy export plus sign-in logs |
| MSA, SLA, responsibility matrix, written scope changes | Establishes duty. Mastagni Holstedt v. LanTech is reported to have proceeded on a verbal agreement with no written contract (ChannelE2E) | The contract file |
| RMM and remote-control session audit | What the remote-access tool did in the tenant, and when | Datto RMM: 180-day retention, maximum 500 rows per CSV export (Datto). ConnectWise Control has basic and extended audit modes, extended including session video; the purge interval is whatever your database maintenance plan says (ConnectWise) |
#Segregation: Client A's pack contains zero Client B data
Every forensic source in Microsoft 365 is tenant-scoped by architecture. That is a gift — per-client segregation is the default. The entire burden is on you not to break it, using tools built to aggregate, because aggregating is what makes a small MSP viable.
Where commingling actually happens:
- Cross-tenant advanced hunting. Defender's multitenant management gives partners a consolidated view across tenants and hunting across multiple tenants at once (multitenant management). A cross-tenant KQL query exported to CSV is a commingled export by construction — and on the night, it was exactly the right query to run.
- Unfiltered console exports. Datto RMM's activity log filters by site, device and user. Export without the filter and you have exported the estate.
- Ticket and documentation exports. A technician's ticket history spans clients. So does the documentation tree holding diagrams and runbooks.
- Your own internal comms. A Teams channel or ticket thread covering several clients' incidents cannot be produced without redaction.
- The analyst's workstation. Live-response artifacts are pulled through one tenant's console and land on whichever laptop was to hand. That laptop is where folders get mixed.
Why this is its own incident. Put Client B's personal data in a pack handed to Client A's auditor and you have made an unauthorized disclosure of Client B's data. As processor you must notify Client B without undue delay, and you do not get to assess the risk first — the EDPB is explicit that the processor "does not need to first assess the likelihood of risk… it is the controller that must make this assessment" (EDPB Guidelines 9/2022). So: you emailed a spreadsheet, and now you are ringing a client who was having a perfectly good week to tell them they have a breach, caused by you, discovered while you were defending yourself over somebody else's. There is no version of that call that improves your position.
#### How to prove segregation — assert nothing, produce artifacts
| # | Control | What it produces for the pack |
|---|---|---|
| 1 | Every export runs scoped to one tenant, and the scope is captured | The exact query text — KQL, PowerShell or the console filter set — including the tenant/workspace/site filter, the console, the operator identity and the UTC run time. A query string is self-authenticating in a way a screenshot is not |
| 2 | Per-artifact tenant attestation in the manifest | Artifact → source system → scope filter → operator → export timestamp (UTC) → SHA-256 → reviewer who confirmed single-tenant content |
| 3 | Documented second-person review before anything leaves your custody | A named reviewer who checked each export for foreign-tenant identifiers. This is the control that catches the unfiltered CSV, and it is free |
| 4 | Architecture evidence | Workspace-per-tenant topology; the RBAC and security-group model showing which technicians could reach which tenant; GDAP group membership |
| 5 | Redaction log | What class of content was removed and why. Never redact by silently deleting rows |
Two facts make control 4 easy if you built it right and impossible if you did not. Microsoft's documented best practice for an MSSP is at least one Sentinel workspace per Entra tenant, with source logs staying in each spoke workspace and cross-tenant work done through Azure Lighthouse (prepare for multiple workspaces; manage multiple tenants as an MSSP) — a pooled single-workspace design is the design that fails this question. And Microsoft's delegation recommendation is to "assign a security group in your tenant to an approved role in the customer tenant… so that it includes only the technicians who help that customer" (cross-tenant delegated administration), which makes "who could have reached this tenant" a group membership export rather than an essay.
Actionable takeaway: carry the TenantId column into every artifact you save, and never run a query, export or collection for Client A from a console session authenticated to Client B. That one habit prevents most of what this section is about.
#Chain of custody when the evidence belongs to someone else
The forensic standards are written for a first-party responder. You are not one. The client owns the data, you have custody of a copy, and the authority under which you took it is contractual. That gap is where MSP packs fall over — not on the hashing, on the paperwork about who was allowed to press the button.
| Source | The operative requirement |
|---|---|
| RFC 3227 (BCP 55) | Chain of custody records where, when and by whom evidence was discovered and collected; where, when and by whom it was handled or examined; who had custody, for what period, and how it was stored. "The methods used to collect evidence should be transparent and reproducible." Keep dated notes; state whether a timestamp is local or UTC; proceed from volatile to less volatile; "don't shutdown until you've completed evidence collection" (RFC 3227) |
| NIST SP 800-86 | Log every custodian, their actions and times; store evidence securely when not in use; examine only a copy; verify integrity by computing and comparing message digests. Designate one evidence custodian with sole responsibility to document and label every item and record every action, by whom, where and when. Document the imaging software or hardware used — name, version, licensing. If it is unclear whether evidence needs preserving, preserve it (SP 800-86) |
| ISO/IEC 27037:2012 | Defines the Digital Evidence First Responder and Digital Evidence Specialist roles and four quality principles: auditability, repeatability, reproducibility, justifiability (ISO). ISO/IEC 27001:2022 A.5.28 Collection of evidence is the certification hook (ISMS.online) |
| ACPO / NPCC Guide v5 | Principle 3: "An audit trail or other record of all processes applied to digital evidence should be created and preserved. An independent third party should be able to examine those processes and achieve the same result" (NPCC) |
| NIST SP 800-61r3 | Under RS.AN-07, formal chain-of-custody procedures "might not be performed for every incident", "however, collected incident data is still considered evidence" (SP 800-61r3) |
That last line is the one to read aloud to a technician who thinks chain of custody is for people with evidence bags and gloves. You do not have to run a laboratory. You do have to be able to say who held it, when, and what they did to it.
The five deltas because you are the MSP:
- Ownership and custody are split. Every artifact carries a header naming the owning tenant, the collecting MSP entity and named technician, the authority relied on (MSA clause, GDAP relationship, client written instruction, or counsel engagement letter), and the destination store.
- One custodian per incident, not "whoever was on shift." The DOJ is blunt: "Ideally, as few employees as practicable should be assigned the responsibility of retaining custody of such information", because proper handling "can be useful in rebutting claims in subsequent legal proceedings… that electronic evidence has been tampered with or altered" (DOJ, *Best Practices for Victim Response*). Name a deputy — one person and no deputy is how a pack stalls for a fortnight in August.
- Your collection tools are in scope. RMM, EDR and remote-control agents are what an attacker targets, and in an MSP incident they may already be compromised. Record collector tool, version and license for every artifact. Where the RMM's integrity is in question, stop collecting through it: you cannot both isolate it and use it.
- Authority is documented before collection, not after. SP 800-86 puts the decision to collect and preserve in a legally usable way before collection begins.
- Hashing is the cheap part and it is the part people skip. Hash at the moment of export, record algorithm and tool, store the hash list separately from the evidence.
#The manifest
This is the deliverable. Create it empty as a template in your evidence store today, and fill it per tenant, per incident.
Every artifact carries the same nine fields: artifact path · source system · scope filter applied · export mechanism · operator · export timestamp (UTC, ISO 8601) · coverage window (UTC) · SHA-256 · completeness note. If an artifact cannot carry all nine, it does not go in the pack until it can.
The pack is one folder per tenant per incident, holding nine numbered groups: 00 cover and control, 10 identity, 20 Microsoft 365, 30 endpoint, 40 network, 50 backup, 60 MSP-side records, 70 assurance and legal, 80 analysis. The numbering is not decoration — it is what lets a stranger find the authority document without asking you, and what lets you hand a requester "groups 00, 60 and 70" as a sentence instead of a scavenger hunt.
The full directory tree and the 00_MANIFEST.csv schema are Appendix C. Copy them from there rather than retyping them, so every pack your firm ever produces has the same shape. What follows is the part Appendix C does not carry: what each group is actually for, and the thing that quietly destroys it.
#### What each group proves, and what kills it
00 — Cover and control. This group makes the other eighty per cent admissible rather than merely interesting. 00_AUTHORITY.pdf answers "on what basis did you take a copy of our client's mailbox?" 00_EVIDENCE_BASELINE.pdf is the quiet hero — which log sources were enabled, at what retention, since when, under which license — because it converts a hole in your data from an adverse inference into an explained gap.
10 — Identity. Entra sign-in and audit retention is license-dependent, between 7 and 30 days, and not retroactive. Diagnostic settings streaming to Log Analytics or Sentinel are the only durable route, they capture only what happens after configuration, and Microsoft warns it "might take up to three days for the logs to start appearing in the destination" (data retention; configure diagnostic settings). Produce the diagnostic-settings configuration and its creation date alongside the data — the config is the completeness proof. MSP trap: Log Analytics, Diagnostic settings, Workbooks and the Monitoring tab are listed as not supported via GDAP (supported workloads), so this is an onboarding task with an identity that can do it, not a 2am task.
20 — Microsoft 365. Two things to get right. Segmentation under the export caps, with the index as proof. And hashes: the eDiscovery metadata schema carries Native MD5 and Native SHA 256, both marked No for direct export and Yes for review-set export (metadata fields) — a direct export from search produces no native hashes, so route through a review set. The Summary.csv Total-versus-Actual delta plus the warnings-and-errors file is your completeness proof here (export from a review set). Mind the platform clocks too: search exports auto-cancel after 7 days and the download window is 14 days from creation (limits in eDiscovery). An export left unretrieved is simply gone.
30 — Endpoint and XDR. Defender data is portal-visible for 180 days, but advanced hunting reaches only 30 days unless streamed to Sentinel, and hunting exports cap at 100,000 rows / 64 MB with a 10-minute query timeout (data retention; advanced hunting). The investigation package ships with its own completeness artifact: CollectionSummaryReport.xls lists each data point, the command used, execution status and any error code (response actions on a device). Put that file in the pack, not just the folders. Live response produces a Command log and a unique Session ID used for auditing (live response) — that is your chain-of-custody artifact for everything you touched on the host.
40 — Network. Appliance storage rolls: FortiGate local disk logging is governed by diskfull, log-quota, maximum-log-age and roll-schedule (Fortinet). Capture the retention and quota values as they stood at the time of the incident and file them next to the logs. And keep NIST's caution in the analysis: analysts "should not assume that an activity is benign if security devices have not reported it as malicious."
50 — Backup. The insurer's group: restore-point inventory with immutability state and expiry, verification results, and a service-by-service recovery timeline feeding the business-interruption calculation. Watch your own retention — Veeam Backup Enterprise Manager's session history default is 13 weeks (Veeam), shorter than the interval between an incident and the average insurance dispute.
60 — MSP-side records. Everything in the section above. The group you cannot build retroactively, and the one the other side reads first.
70 — Assurance and legal. Your client's auditor treats you as a subservice organization. Under the carve-out method your controls are excluded from their report and disclosed as complementary subservice organization controls — a disclosure requirement, not a tested assertion — and their auditor tests only whether your client monitored you; under the inclusive method your controls are actually tested (Linford & Co). Read your own CUEC list before they do: it may say the client is responsible for something the client is convinced you do. And check your scope statement — a certificate whose ISO 27001 scope excludes the service line that failed will be read against you.
80 — Analysis. Three files carrying the pack's credibility. 80_master_timeline_UTC.csv, one row per event: UTC ISO 8601 timestamp · source system · source-rendered timestamp and zone · actor · action · target · artifact reference · confidence (observed / inferred). 80_scope_of_exposure.md: accounts, mailboxes, files, hosts, data classes, with the query evidence behind every number — because every number in that file ends up in a notification letter. The third one gets its own section.
#80_gaps_and_limitations.md — write it before they find it
The rule, and it is not negotiable: the gaps file is written by you and never discovered by the other side.
Every log source that was off, expired, throttled or truncated, with the reason and the date it would have started. Entra sign-ins that reach back only 30 days because the tenant is on P1 and diagnostic settings were configured on a stated date. The hunting query that returned exactly 100,000 rows. The firewall whose disk rolled at nine days. The mailbox export showing IsThrottled. The investigation package that failed because the laptop was on battery — Microsoft documents that collection "might fail if the target device has a low battery level or is on a metered connection," and a failed collection is itself a fact for the timeline.
Volunteering this is not a confession. It is the strongest single move in the pack.
It converts a hole into a finding you already own. "Sign-in logs begin 2025-11-04 because diagnostic settings were configured on that date under the tenant's P1 license" is an explained gap. Silence in the same place, found by someone else, is an unexplained gap — and unexplained gaps get characterized, by people whose job is to characterize them unfavourably.
It is the difference between a limitation and a spoliation argument. FRCP 37(e) does not create the duty to preserve; it supplies the sanctions when ESI that should have been preserved is lost because reasonable steps were not taken — measures no greater than necessary to cure prejudice under (e)(1), and only on a finding of intent to deprive, an adverse-inference instruction or dismissal under (e)(2) (Judicature). A contemporaneous, self-authored gaps file is direct evidence of reasonable steps and the opposite of intent.
It makes the rest of the pack credible. A production with no acknowledged limits reads as either a fabrication or an incomplete search. One that names its own edges reads as competent. Auditors relax visibly when they meet a limitations section, because it tells them the person who assembled the pack understood what they were assembling.
Actionable takeaway: open 80_gaps_and_limitations.md as the first file in the pack, not the last, and append every time a query truncates or a source comes back short. Written at the end, it is a memory test. Written as you go, it is a record.
#Timeline discipline
The joint Australian, US, UK, Canadian, New Zealand and partner guidance Best Practices for Event Logging and Threat Detection gives rules you can paste into an SOP (PDF):
- Establish an accurate and trustworthy time source used consistently across all systems, with the same date-time format everywhere, and multiple accurate time sources where possible in case the primary degrades.
- Synchronize and validate time servers throughout all environments, and capture significant events such as device boots and reboots.
- "Using Coordinated Universal Time (UTC) has the advantage of no time zones as well as no daylight savings, and is the preferred time standard. Implement ISO 8601 formatting… (e.g.
2024-07-25T20:54:59.649Z)." - Timesharing should be unidirectional: OT synchronises to IT, never the reverse.
ISO/IEC 27001:2022 A.8.17 Clock synchronization is the certification hook for the budget conversation (ISMS.online).
| MSP-specific hazard | Rule |
|---|---|
| Every console renders in a different zone — viewer local, tenant-configured, or UTC | The master timeline is UTC / ISO 8601 only. Each row records the source rendering alongside the normalized value. Convert once, at ingest, and record the conversion |
| Ingestion lag is not event time | Annotate known lags: Lighthouse audit logs up to an hour; Entra diagnostic settings up to three days to begin appearing at a new destination. A timeline treating ingest time as event time shows your response as later than it was |
| Your own clock is evidence | PSA timestamps, paging records and technician workstation clocks all feed "when did you know". Sync them to the same source and be able to prove it |
| Two tenants, one attacker | Each tenant's timeline stands alone. A separate internal master timeline correlates them. Do not produce the master to a single client |
| Collection times matter too | Per RFC 3227, record when each artifact was collected and whether the timestamp is local or UTC |
The confidence column is not optional. Every row is either observed — an artifact in this pack shows it — or inferred, meaning you reasoned to it. NIST SP 800-86 warns that file times may be wrong because the clock was never synchronized to an authoritative source, because seconds were omitted, or because an attacker altered them, and that OS-presented time can differ from the BIOS because of time-zone settings. An inference is legitimate. An inference presented as an observation is what destroys a witness, because you will be asked which artifact proves the row and there will not be one.
#Privilege, for someone who is not the client's lawyer
Start from the correct default: nothing you write is privileged.
Attorney-client privilege protects confidential communications between a client and its lawyer for the purpose of legal advice. You are not the client's lawyer and usually not the client's employee. Your tickets, alerts, runbooks, RMM logs, internal chat and the incident report you wrote as a client deliverable are business records created for a business purpose. They are discoverable. Every MSP that has typed "privileged and confidential" into a ticket subject line has achieved nothing except drawing attention.
The one route in. FRCP 26(b)(3)(A) protects documents prepared in anticipation of litigation or for trial "by or for another party or its representative (including the other party's attorney, consultant, surety, indemnitor, insurer, or agent)" (FRCP 26). An MSP retained by the client's outside counsel, under a separate written engagement scoped to legal advice or anticipated litigation, separately invoiced, not duplicating your BAU deliverables and not reused operationally, can sit inside "consultant… or agent."
Why it usually fails for you anyway. The Capital One / Clark Hill / Rutter's line holds that protection evaporates where the work would have been done anyway for business reasons, where the report is used for business purposes, or where it "only discussed facts and did not involve opinions and tactics." And the structural problem is worse for an MSP than for a DFIR firm: a consulting expert becomes a fact witness the moment it also acts as remediator or operator. You are the remediator. You rebuilt the servers, you reset the accounts. That work is operational, it is factual, and no protection attaches to it. Merely instructing an existing vendor to "report to counsel" is insufficient.
The four ways MSPs get pulled in: (1) the client's breach counsel asks for your logs and interviews — you are a witness, not a privileged participant, unless separately engaged; (2) counsel engages a DFIR firm and asks you to support it — your support work is generally not privileged, though the firm's report may be; (3) counsel engages you for a discrete analytical task — possible protection over that product if genuinely scoped and separated; (4) the conflict case.
Rules for the on-call technician, in ticket notes and chat. Facts and timestamps only. No characterization of cause, blame, legal exposure, or "we should have" — assume every line is read aloud in a deposition, as the SEC's SolarWinds complaint made of internal messages and presentations (SEC). Do not label ordinary operational work "privileged"; over-labeling invites a challenge that succeeds and taints the genuinely protected material with it. Keep counsel-directed work in a separate, access-controlled workspace with its own engagement reference, not in the PSA. Purview's Potentially privileged metadata field is a model output useful for screening — never produce it as a legal determination. And keep a contemporaneous decision record offline or on unaffected systems regardless of privilege, because you need it for regulators and the retrospective either way.
#Legal hold versus your deletion obligation
A genuine conflict, not a paperwork problem, with two horns.
Horn 1 — data protection law says delete. As processor, GDPR Article 28(3)(g) obliges you to delete or return personal data at the end of the service (Art. 28). The FTC's Blackbaud order goes further and makes indefinite retention itself the violation: publish and adhere to a retention schedule stating purpose, business need and "the set time frames for the deletion of that personal information (i.e., no indefinite retention)", and delete customer backup files no longer needed (FTC).
Horn 2 — the law of evidence says keep. GDPR Article 17(3)(e) disapplies the right to erasure where processing is necessary "for the establishment, exercise or defense of legal claims" — the recognized route for a legal hold, covering active litigation, reasonably anticipated litigation and regulatory investigations (Kennedys).
Resolve it in this order.
| # | Step | Why it sits here |
|---|---|---|
| 1 | Identify the trigger, record the date | The duty attaches when litigation is pending or reasonably foreseeable. For an MSP the trigger is earlier and quieter than "we got sued": the client declares an incident; their insurer opens a claim; a regulator makes contact; the client hints at a claim against you; your own insurer is notified; a client gives notice of termination during or shortly after an incident |
| 2 | Issue a written hold notice internally | Name the tenant, date range, custodians (technicians), systems in scope, and require acknowledgment. Preserve the notice and acknowledgements — they are the "reasonable steps" evidence under Rule 37(e) |
| 3 | Suspend the deletions you control, and record it | RMM/PSA purge jobs, Sentinel retention reductions, Purview retention edits, mailbox retention, backup expiry, your own log rotation. Sentinel gives a safety net: when you shorten a table's total retention, "Microsoft waits 30 days before removing the data, so you can revert the change" (manage data tiers) |
| 4 | Export before expiry where you cannot suspend | Vendor-side fixed retention does not care about your hold: Defender 180 days, Datto RMM 180 days, Entra 7–30 days. Export, hash, record |
| 5 | Narrow the hold, and write down the scoping reasoning | Article 17(3)(e) covers what is necessary for the claim, not "the entire backup estate". Scope by tenant, date range and data class |
| 6 | Segregate and lock the held set | Rather than freezing production. This also satisfies NIST's "store the evidence in a secure location when it is not being used" |
| 7 | Tell the client in writing what you hold and why | Expect them to want their own copy. A hold they learn about in discovery is a bad conversation |
| 8 | Log the release | A hold that is never released becomes a retention violation. Diary it |
Two caveats to carry. Whether Article 17(3)(e) extends to non-EU/UK proceedings is unresolved, so a hold existing only because of US litigation over EU data needs advice. And the Blackbaud order also requires an incident report to the FTC within 10 days of notifying any government entity of a breach — retention obligations and reporting obligations arrive together.
#What to capture continuously, ranked by how badly the absence hurts
A twelve-person MSP will not do all ten this quarter. Do them in this order.
| # | Capture | Why it ranks here | Cheap version |
|---|---|---|---|
| 1 | Entra diagnostic settings → Log Analytics/Sentinel, per tenant, today — AuditLogs, SignInLogs, NonInteractiveUserSignInLogs, ServicePrincipalSignInLogs, ProvisioningLogs, RiskyUsers, UserRiskEvents, MicrosoftGraphActivityLogs | Without it, a breach discovered on day 45 has no identity evidence at all, permanently, in every tenant | A configuration change, not a purchase. Do it at onboarding with an identity that can — it is not supported via GDAP |
| 2 | Purview Audit at the right tier, with retention policies actually created | 180 days is the floor, and the 10-year add-on policy is not retroactive | Standard's 180 days plus a scheduled monthly export to your own store |
| 3 | Monthly export of GDAP activity and Lighthouse audit logs | The 30-day Lighthouse filter and the one-year GDAP cleanup make this non-optional | A scheduled task and a folder. Genuinely free |
| 4 | Monthly export of RMM and remote-control audit logs, segmented under the caps | 180-day retention, 500 rows per CSV | Script the segmentation once, run it monthly |
| 5 | Sentinel per tenant, XDR tables extended past 30 days | Advanced hunting reaches only 30 days without it | 31–90 days of analytics retention is free storage on Sentinel-enabled workspaces; ingestion charges still apply |
| 6 | Standing object ID → technician → employment-dates mapping | The only way to turn user_a1b2c3… into a name | One CSV, monthly |
| 7 | Per-tenant evidence baseline — sources enabled, retention, since when, license | Converts a gap into an explained gap | A page per client in your documentation platform, refreshed on every license change |
| 8 | NTP / UTC / ISO 8601 everywhere, including PSA and technician workstations, with boot events captured | Your own clock is evidence | Configuration, not spend |
| 9 | Immutable, access-controlled evidence store with per-tenant partitions, separate from production | Access to delete or modify audit logs should be limited to personnel with a justified requirement | Immutable object storage with per-client containers beats a shared \\NAS\Forensics\ with subfolders, and costs less than the subfolder incident will |
| 10 | Contract language: logging visibility per AA22-131A, notification timing, the legal-hold carve-out, e-discovery data access, a named evidence custodian | Easier to agree at renewal than under a preservation notice | One MSA revision |
#Rehearse the pack in peacetime
The cheapest thing in this section: pick one real tenant and produce a complete evidence pack for a fictional incident, on the clock, with a stopwatch running.
Not a tabletop discussion. An actual production. Someone plays the auditor and issues a twenty-two-item request list. Someone assembles. A third person does the segregation review. You time it end to end and write down what broke.
What you will discover, roughly in this order: the RMM export stops at 500 rows and nobody knew; the audit search returned exactly 50,000 records, which means it returned an unknown number more; the direct eDiscovery export has no hashes in it; the technician who configured the client's diagnostic settings left in March and nobody can say what date they were set; the Lighthouse filter will not reach back far enough; and the one person who knows where the evidence store lives is on annual leave.
Every one of those is cheap to find in a drill and expensive to find in week six of a claim. Unlike most security exercises, this one leaves a durable artifact: a template pack, a scoped query library, and a segment-index script you keep.
#What should have been true
Three things, each of which would have made this Monday an afternoon's work.
1. Per-tenant log streaming, configured at onboarding. Entra diagnostic settings into each client's own workspace on day one, and Sentinel per tenant rather than one pooled workspace. It costs one configuration blade per client on day one, it is not available via GDAP so it has to be done deliberately, and it is not retroactive — which is the difference between "here is the identity evidence" and "the license only kept thirty days." See Part 2 on multi-tenant architecture and per-tenant telemetry.
2. The MSP-side records exported monthly, before anyone asked. GDAP activity, Lighthouse audit, RMM activity segmented under the 500-row cap, and the object-ID-to-technician map. None of it can be reconstructed once the console's window closes, all of it is a scheduled task, and it is the half of the request list that is about you. See Part 2 on evidence readiness.
3. An MSA that already answers the question. The legal-hold carve-out to the deletion clause, the logging-visibility commitment AA22-131A tells your clients to demand, a named evidence custodian, and a data-access clause written knowing it doubles as an e-discovery clause. Negotiated at renewal, calmly, by people who were not awake all night. See Part 2 on contractual authority.
The pack you assemble in week six is decided in week zero. Everything else is typing.
Keep the hashes, keep the clocks in UTC, and write your own gaps file before somebody else writes it for you.
#### Sources
- The Register — Clorox v. Cognizant
- ChannelE2E — MSP and backup vendor sued over cybersecurity breach
- HHS OCR — ransomware investigation settlements
- Nixon Peabody — HIPAA risk analysis failures and OCR settlements
- ICO — software provider fined £3m following 2022 ransomware attack
- FTC — action against GoDaddy over data security
- FTC — Blackbaud order on data retention
- [AXIS Cyber Insurance Policy specimen (AXIS 1012561 0120)](<https://www.axiscapital.com/docs/default-source/resources/axis-1012561-0120-(specimen).pdf>)
- Sedona Conference — Commentary on Possession, Custody, or Control
- CISA AA22-131A — Protecting Against Cyber Threats to MSPs and their Customers
- Microsoft — Monitor governing tenant admin activity in a governed tenant
- Microsoft — Cross-tenant delegated administration
- Microsoft — GDAP FAQ
- Microsoft — View and export GDAP activity logs
- Microsoft — Workloads supported by GDAP
- Microsoft — Review audit logs in Microsoft 365 Lighthouse
- Microsoft — Monitor service provider activity (Azure Lighthouse)
- Microsoft — Defender multitenant management
- Microsoft — Prepare for multiple workspaces and tenants (Sentinel)
- Microsoft — Manage multiple tenants as an MSSP (Sentinel)
- Microsoft — Manage data tiers and retention in Microsoft Sentinel
- Microsoft — Export, configure, and view audit log records
- Microsoft — Use a PowerShell script to search the audit log
- Microsoft — Document metadata fields in eDiscovery
- Microsoft — Export documents from a review set
- Microsoft — Limits in eDiscovery
- Microsoft — Data retention and data security in Microsoft Defender XDR
- Microsoft — Advanced hunting overview
- Microsoft — Take response actions on a device
- Microsoft — Investigate entities on devices using live response
- Microsoft — Microsoft Entra data retention reference
- Microsoft — How to configure diagnostic settings
- Datto RMM — Activity Log
- ConnectWise Control — Audit page
- Fortinet — log disk setting
- Veeam — configuring retention settings for index and history
- RFC 3227 — Guidelines for Evidence Collection and Archiving
- NIST SP 800-86 — Integrating Forensic Techniques into Incident Response
- NIST SP 800-61r3 — Incident Response Recommendations and Considerations
- ISO/IEC 27037:2012
- ISMS.online — ISO 27001 Annex A 5.28 Collection of evidence
- ISMS.online — ISO 27001 Annex A 8.17 Clock synchronisation
- NPCC — Good Practice Guide for Digital Evidence v5
- DOJ — Best Practices for Victim Response and Reporting of Cyber Incidents v2.0
- Joint guidance — Best Practices for Event Logging and Threat Detection
- EDPB Guidelines 9/2022 on personal data breach notification
- GDPR Article 28
- Kennedys — US litigation holds vs GDPR erasure obligations
- Judicature — Rule 37(e): The New Law of Electronic Spoliation
- FRCP 26 — Cornell LII
- Linford & Co — subservice organisations: carve-out vs inclusive
- SEC — press release 2023-227 (SolarWinds)
#Part 2 — Before the Night
#Before the Night: What Must Already Be True
Every scenario in Part 1 ended the same way: with two or three things that, had they existed on paper before the phone rang, would have turned a career-defining night into a Tuesday. This is where we build them.
None of it is exotic. There is no control here a twelve-person MSP cannot run, and where something needs a license tier you may not own, I have named the cheap version beside it. What the work does require is that it happen in daylight, on a calendar, with a name against each item — because every artifact below shares one property: nearly worthless if you start it during the incident, decisive if you finished it six months earlier. One number should shape all of it: median time to attempted Active Directory compromise is 3.40 hours, accelerating 70% year over year, in a case population where 84% of affected organizations had fewer than 1,000 employees (Sophos 2026 Active Adversary Report). If your escalation path is slower than your adversary's, the architecture has already picked the winner.
#1. The Client Authority Matrix
If you build nothing else from Part 2, build this. It answers one question per client, in advance: what may we do at 02:00 without asking, what must we ask about, and who exactly do we ask?
NIST puts the requirement on both sides. A client's IR policy must state "which roles have the authority to confiscate, disconnect, or shut down technology assets"; a provider contract must define "authority to act on behalf of the organization" and "restrictions on what the service provider can do, such as… immediately deactivating certain services to contain an incident" (NIST SP 800-61r3). CISA tells your clients the same, so they will eventually ask you for it (AA22-131A).
The best published model is commercial. Sophos's MDR service description offers Notify Only, Collaborate ("no Response Actions are taken without Customer/MSP's written consent") and Authorize (pre-authorization, notification afterwards) — plus the mode that solves the 02:00 problem, "Collaborate then Authorize," which authorises action "in the event Sophos does not receive acknowledgment from Customer/MSP after making reasonable attempts to contact all Customer defined contacts." It requires "at least one primary and one alternate authorized contact" and asks for acknowledgment "within sixty (60) minutes" (Sophos). Steal the structure.
#### 1.1 The template
Hold this as structured PSA fields, not a PDF. The on-call technician must read it in under a minute, from a phone, while the client's tenant is possibly the thing on fire.
Header block, per tenant.
| Field |
|---|
| Legal entity and contracting entity on the MSA |
| MSA signed? Emergency-authority clause present? Agency disclaimed? |
| DPA in place; containment recited as documented instructions (Y/N) |
| BAA in place; notification timeframe in days |
| PCI in scope; responsibility matrix version; who calls the acquirer |
| Other regimes: NIS2 / DORA / NYDFS / DFARS 7012 / FTC Safeguards / public company |
| Contractual notification deadline (hours); legal notice address and method |
| Client insurer, policy number, 24/7 claims line, panel breach counsel and DFIR firm |
| Our own carrier's claims line |
| Data residency constraints (EDR region, workspace region) |
| Authority tier as enforced in our RBAC |
Approver block, per tenant.
| Field |
|---|
| Client Authority Holder (primary): name, role, mobile, personal email |
| Client Authority Holder (deputy): name, role, mobile, personal email |
| Executive escalation, for shut-down-production decisions |
| Client legal/DPO contact; client communications lead |
| Out-of-band channel that survives their tenant being isolated |
| Acknowledgment SLA (minutes), and what happens on expiry |
| Unreachable window N (minutes) after which tier-2 actions become pre-authorized |
| Date the out-of-hours path was last tested, and by whom |
Action authority table. Mark each row Pre-authorised, Approval required or Prohibited, and name the approving role for every row that is not pre-authorized.
| # | Disruptive action | Tier | Approver (role) |
|---|---|---|---|
| A1 | Isolate a single endpoint (EDR containment) | ||
| A2 | Kill process / quarantine file / remove scheduled task | ||
| A3 | Disable one user; revoke sessions and tokens (tokens first, then password) | ||
| A4 | Block a specific IP, domain or URL at egress | ||
| A5 | Disable inbox rules or a mail-flow rule | ||
| A6 | Reset one privileged credential | ||
| A7 | Isolate a server or hypervisor host | ||
| A8 | Isolate a domain controller | ||
| A9 | Tenant-wide credential reset / mass session revocation | ||
| A10 | Block an egress class, or sever a site-to-site VPN | ||
| A11 | Suspend or take a production service offline | ||
| A12 | Restore from backup over live data | ||
| A13 | Engage third-party DFIR on the client's behalf | ||
| A14 | Notify a regulator on the client's behalf | ||
| A15 | Pay or negotiate a ransom | Prohibited | — |
Three rows carry constraints beyond preference. A13: check the client's carrier panel first — the wrong forensic firm can turn a covered claim into an argument. A14: a processor may notify on the controller's behalf "if the controller has given the processor the proper authorization and this is part of the contractual arrangements", but "the legal responsibility to notify remains with the controller" (EDPB Guidelines 9/2022). A15 is never yours: client, counsel, insurer.
#### 1.2 Make the RBAC match the paper
Microsoft implements this distinction in the role, not the paperwork: grant Defender Experts Security Operator and "the experts can perform the required response actions on the incident on your behalf"; leave them at Security Reader and the same actions sit under Pending actions, awaiting the customer (Microsoft Learn).
Do the same. If a tenant is "approval required" for domain-controller isolation, your standing account there should not be able to isolate a domain controller. At 02:00 technical capability beats contractual restriction every time, and that is not the technician's fault.
#### 1.3 Getting clients to fill it in
Ask at onboarding, and at renewal. Never mid-incident. Asking someone to nominate an out-of-hours decision-maker while their file server encrypts is asking a frightened person to sign under duress: bad answers, and any authority granted then gets re-litigated later by their counsel. Onboarding works because they are already answering questions about themselves; renewal works because you are already discussing what they get for their money. Three things make it land:
- Lead with the outage, not the breach. "If we see ransomware starting on your file server at 2am, do you want us to pull that server off the network immediately, or call you first? Both are legitimate. We need to know which, because whichever we guess, one of us is unhappy in the morning."
- Show them the buyer-side checklist they will measure you against. NCSC tells buyers to demand exactly this — severity-based notification timeframes "suitable for your company", with roles, responsibilities and liabilities clearly defined (NCSC-UK). A client who sees you volunteering the government's own checklist stops reading this as vendor paperwork.
- Price the tiers. A client wanting approval on everything is buying a slower response. Sophos says it plainly: "Notify Only can materially delay containment and disruption actions and may increase risk."
Then test the phone path — twice a year, unannounced, from the on-call phone, logged. Half of you will find a disconnected number. I have never met an MSP that ran this test and learned nothing.
Actionable takeaway: book ninety minutes this month and complete the matrix for your top ten clients. Then make "matrix agreed and phone path tested" a go-live blocker on onboarding and a standing renewal item.
#2. The MSA clauses you need
Everything below is illustrative — the shape of a clause and what it does, so you walk into a lawyer's office already knowing what you want. It is not legal advice and must not be pasted into a live agreement. Get your own counsel, in your own jurisdiction. I am a CTO with a research pack, not your solicitor.
The baseline is uncomfortable. Published MSP master agreements are overwhelmingly scope-payment-liability documents: the CompassMSP MSA grants unilateral powers only to change services on thirty days' notice and to suspend for non-payment — no emergency authority, no incident response clause, no breach-notification timeline to the client (CompassMSP MSA). Many also disclaim agency outright, quietly removing the emergency-agency argument you hoped to lean on.
Nine clauses close the gap. What each one is for, and what it costs you to be without it:
| Clause | What you need it to do | What breaks without it | Owner |
|---|---|---|---|
| 1. Emergency and exigent action | Define "Security Emergency" by trigger — confirmed or reasonably suspected unauthorized access, active malware propagation, credential compromise — grant power to take the matrix's pre-authorized actions, add the limiter (to the minimum extent and of the minimum duration required), and put post-hoc notice on a clock. That trigger → power → limiter → notice shape is Google's Emergency Security Issue structure (Google Workspace terms); the limiter is what makes it acceptable to the client's counsel | The US sharp edge is damage, not access. 18 U.S.C. § 1030(a)(5)(A) reaches anyone who "intentionally causes damage without authorization", damage being "any impairment to the… availability of data, a program, a system, or information", with a private right of action once loss reaches $5,000 (18 U.S.C. § 1030). Your admin credentials do not answer it: the statute asks whether the damage was authorized | Practice Owner with counsel; Tenant Lead owns the per-client schedule |
| 2. Right to suspend or disconnect | A security suspension right distinct from the non-payment right, on the same trigger, with the same limiter and notice — plus what happens if the client refuses reconnection conditions | Your only contractual power to disconnect anything is tied to an unpaid invoice. Published suspension clauses are near-universal for non-payment and far from universal for security (Law Insider) | Practice Owner |
| 3. Security incident cooperation | Mutual duties: the client makes people available, preserves systems on request, does not restore or reimage without agreement, keeps the Authority Holder current; you cooperate with their counsel, forensic firm and insurer | Read your existing clause before assuming it helps. Published "security incident" clauses impose duties on the provider and grant the provider nothing (Law Insider). A cooperation clause is not an authority clause, and without a mutual version your access to the client's own people during their incident is goodwill — scarce at 03:00 | Practice Owner drafts; MSP Incident Commander reviews for operational realism |
| 4. Evidence preservation and a legal-hold carve-out | Both parties preserve relevant logs, images and records on written notice of a suspected incident — plus an express carve-out to the deletion clause, because Article 28(3)(g) requires the processor to "delete or return all the personal data to the controller after the end of the provision of services" (Art. 28) | You are asked to delete the evidence of a dispute you are party to, and it arrives exactly when an unhappy client offboards after an incident. The platforms run their own clocks anyway: Defender data is "deleted from Microsoft's systems and is unrecoverable" no later than 180 days from contract termination or expiration (Defender XDR data retention), and expired delegated-admin relationships are cleaned up after one year (GDAP FAQ) | Practice Owner with counsel; the named MSP evidence custodian operates it |
| 5. Notification timing, both directions | Yours: a severity-banded clock, a legal notice address and method, a named recipient with out-of-hours contact, and no waiting for the investigation to finish. Theirs: tell you about any confirmed or suspected security event, credential compromise or regulator contact | For EU/UK clients you have no 72-hour budget of your own — "the processor does not need to first assess the likelihood of risk arising from a breach before notifying the controller", and "the controller should be considered as 'aware' once the processor has informed it" (EDPB Guidelines 9/2022); every hour spent confirming before telling comes out of their 72. Failure to escalate is also a live pleaded theory in a cyber insurer's subrogation action against its insured's technology vendors (Hunton). Nothing is adjudicated — but the shape of the claim points at the on-call technician | Practice Owner for the clause; MSP Incident Commander for making the number a PSA field |
| 6. Client responsibilities — the CUEC problem | What the client must do for your controls to work — maintain current authority holders, acknowledge within the SLA, keep an out-of-band channel, permit MFA everywhere, keep in-scope systems supported, report staff departures, not grant third parties admin access silently — aligned with the complementary user entity controls in your SOC 2 report | Your report says the client is responsible for things the client believes you do, and the gap surfaces during their audit or their incident (ScalePad; Secureframe). PCI is harder still: v4.0.1 requirement 12.9.2 obliges you to supply on request "Information about which PCI DSS requirements are the responsibility of the TPSP and which are the responsibility of the customer, including any shared responsibilities" (PCI DSS v4.0.1) | Practice Owner for the clause; your SOC 2 program owner reconciles it annually |
| 7. Limitation of liability and indemnification | Keep the cap and the exclusion of consequential loss — business interruption and lost profits are exactly what a wrongful-containment claim pleads. Pair the emergency clause with an express carve-in: good-faith action inside it is deemed authorized. And mirror the refusal shield — no liability for new or worsened malicious activity where the client declined a recommended action (Sophos) | Caps are almost always disapplied for gross negligence and willful misconduct, and a deliberate shutdown is an intentional act — whether it is willful turns entirely on your contemporaneous documentation. A cap also binds only the parties; a regulator, or the client's insurer subrogating, is not one | Practice Owner with counsel, checking every indemnity you give against your own E&O wording before signing — liabilities assumed by contract can fall inside your policy's contractual-liability exclusion |
| 8. Insurance requirements | The client carries cyber liability cover for data breaches, network security failures, business interruption and regulatory liability, produces proof on request, and — the pairing people miss — obtains a waiver of subrogation in your favor. Requiring the insurance without the waiver simply funds the party that will sue you. The waiver is a standard defense, not a wall: it has to align with the client's own policy wording, some policies restrict or override contractual waivers, and a waiver plus a liability cap reduces the exposure without eliminating it | Subrogation, above; and application misrepresentation. In Travelers v. International Control Services the insurer alleged the insured had represented MFA was required for email, remote network access and endpoint/server/directory access when investigation found it deployed only at the firewall; the parties agreed to an order rescinding the policy and declaring it void from inception (Insurance Journal). Read that as an MSP: you usually fill in that section of the client's application. Answer in writing, in your own words, with a dated evidence attachment | Practice Owner, with your own broker in the room |
| 9. The Data Processing Addendum | For every client whose processing falls under GDPR or UK GDPR: everything Article 28(3) requires, plus one MSP-specific move — recite the pre-authorized containment actions as documented instructions of the controller under Article 28(3)(a), which requires processing "only on documented instructions from the controller" (Art. 28). That recital puts your 02:00 isolation inside the instruction rather than outside it. Name your sub-processors too: Article 28(4) keeps you "fully liable" for them. For clients outside those regimes, the equivalent processing terms the applicable state or national law requires | Article 82(2) makes a processor liable where it "acted outside or contrary to lawful instructions" (Art. 82); containment outside documented instructions can strip the liability channelling that normally puts the controller in front of you and, at the extreme, re-characterize you as a controller under Article 28(10). A breach at your EDR or RMM vendor is your breach to notify to every controller in your book | Practice Owner with counsel; Tenant Lead keeps the sub-processor register current |
The illustrative drafting shape for each of these — what the clause does, the failure mode you will actually experience, and language to argue with — is Appendix B. Take that to counsel; do not take this table.
Actionable takeaway: run the review this quarter, in writing, per client: signed MSA at all? Emergency authority, or only non-payment suspension? Agency disclaimed? Notification number? where GDPR or UK GDPR applies, a DPA reciting containment as documented instructions? BAA number? PCI matrix naming who calls the acquirer? Named primary and deputy tested in 90 days? Client insurance plus subrogation waiver? RBAC matching the agreed tier? Any "no" is a finding; the ones that hurt at 02:00 are the second, the eighth and the tenth.
#3. Tenant isolation architecture for the MSP itself
The requirement, in one line: "segregating customer data sets (and services, where applicable) from each other — as well as from internal company networks — can limit the impact of a single vector of attack", alongside "Do not reuse admin credentials across multiple customers" (CISA Guide to Securing Remote Access Software; AA22-131A).
Four planes, which must not be one plane:
| # | Plane | Isolation requirement |
|---|---|---|
| 1 | MSP production tenant — staff email, files, finance, HR | Normal phishing exposure; must not hold the identities that administer clients |
| 2 | Management/identity plane — identities holding delegated access, RMM/PSA/backup admin | Tier 0 for every client you serve; administered only from a matching-tier workstation |
| 3 | Tooling plane — RMM and PSA consoles, documentation, secrets vault, backup console | Management interfaces behind VPN or IP allow-list, never internet-facing |
| 4 | Client tenants | Kept apart by default; the burden is on you not to break it |
Why a flat management plane is the specific failure. In a flat design, the identity that reads your company email is one escalation away from the identity holding Global Administrator in two hundred tenants. Microsoft states what workspace-per-tenant buys: "Ownership of data remains with each managed tenant… Ensures data isolation… Prevents data exfiltration from the managed tenants" (manage Sentinel workspaces at scale). Pooling inverts all three. And the direction of trust is not the one MSPs instinctively defend: CVE-2024-42448 in Veeam Service Provider Console (CVSS 9.9) allows, "From the VSPC management agent machine, under the condition that the management agent is authorized on the server… Remote Code Execution (RCE) on the VSPC server machine" (Veeam KB4679). One client's compromised agent reaching the console that manages everyone's backups. In a flat plane that is not lateral movement; it is portfolio movement.
The cheap version, for twelve people with no budget line.
- Separate identities, same tenant.
dan@reads email;dan.adm@holds delegated access and never receives mail or browses. Free, and the single biggest step. - One hardened machine per technician — or one shared — used only for administrative work. Microsoft's model blocks "email and web browsing — the most common phishing vector" (privileged access devices). A clean, current, EDR-covered laptop with mail and browsing restricted is a legitimate PAW at your scale.
- Move RMM and backup consoles off the public internet today. VPN or IP allow-list — configuration change, not a project.
- Kill shared credentials. Per-client accounts, repositories and keys; the ten largest and three most regulated first if you cannot do all at once.
- Break glass in the management plane. Two or more emergency accounts, cloud-only on
*.onmicrosoft.com, passkey or certificate-based, tied to no individual, excluded from lockout-capable Conditional Access, alerting on every sign-in, validated at least every 90 days and after IT staff changes (emergency access accounts). A lockout in your partner tenant is a simultaneous outage of administrative reach into every client you have.
Two asymmetries to disclose rather than have discovered. Lighthouse role assignments do not appear in the client's IAM blade or az role assignment list, and "you can onboard subscriptions and resource groups that have resource locks, [but] those locks don't prevent actions from being performed by users in the managing tenant" (cross-tenant management experiences). Both are documented and legitimate; both look like concealment if a client's auditor finds them first. Put them in the onboarding pack.
Takeaway: separate the identity that reads your email from the identity that administers your clients, this week. Everything else improves on that; nothing substitutes for it.
#4. Technician identity
Separate admin identities, no shared accounts. NCSC is explicit that MSPs should not use "generic shared management accounts" — actions must trace to "specific people's accounts" — and that administrative authentication should "only be performed from a privileged access workstation" (NCSC). This is not only hygiene: shared credentials make it impossible to distinguish "the attacker used our shared admin account" from "the client was breached separately," which is the exact question section 7 exists to answer.
Phishing-resistant MFA everywhere, and know who it protects. When a technician signs in to a client tenant via GDAP, "MFA is always required in the user's home tenant, and always trusted in the resource tenant", and cross-tenant MFA trust settings "aren't applied if an external user signs in using granular delegated admin privileges" (cross-tenant access settings). Read that plainly: the client cannot independently verify the strength of your technicians' MFA. Your partner tenant's authentication posture is their control. If your technicians use SMS codes, every client you serve is protected by SMS codes whether they know it or not; their only lever is binary — terminate the relationship.
The precedent is direct. The ICO fined a service provider £3,076,320 — the first UK penalty on a processor — after attackers entered through a system requiring only a username and password. The aggravating finding was not sophistication: the organization had developed a working MFA solution before the incident but had not rolled it out, citing a perception that customers would resist (BleepingComputer; Clifford Chance). If you carry an MFA exception because a client complained, that document is now exhibit A. Microsoft has made it contractual too: full MFA enforcement for Partner Center API access from 1 April 2026 — API calls made without MFA will be blocked (Partner Center announcements).
GDAP least privilege, done properly. The default is not least privilege. The Microsoft-led DAP-to-GDAP transition assigned nine roles to the Admin Agents group, two of which are Privileged Authentication Administrator — which can reset authentication methods for any user, including the customer's own Global Administrators — and Privileged Role Administrator, which manages role assignments and can therefore grant itself anything else (Microsoft-led GDAP transition). "GDAP by default" is a two-step path to full tenant control. Go and look at your MLT_* relationships today.
Because all roles in a relationship share one expiry and a Global-Administrator relationship cannot be auto-extended, use three relationships per client:
| Relationship | Duration | Roles | Access model |
|---|---|---|---|
| A — read/triage | Long (max 2 years) | Security Reader, Global Reader, Reports Reader, Service Support Administrator | Standing, all on-call staff |
| B — operational | Medium | Helpdesk / Password / User / Authentication / Intune / Exchange administrator | Standing, Tenant Lead and rota |
| C — incident elevation | Days | Global Administrator, or Privileged Authentication Administrator | JIT-eligible, approver-gated, MFA on activation; expiry does the revocation |
C is operational, not theoretical: an on-call holding only Helpdesk Administrator cannot revoke sessions or reset the password of a compromised Global Admin in the client tenant — that needs Privileged Authentication Administrator (least-privileged roles by task). Discovering that at 02:15 during a business email compromise is a bad way to learn it.
Just-in-time elevation. Lighthouse eligible authorizations activate for 30 minutes to 8 hours with MFA and up to ten approvers — and you must also create a permanent Reader authorization for the same principal or "the user can't elevate their role in the Azure portal" (create eligible authorizations). M365 Lighthouse JIT requires a request through My Access and approval from a JIT approver group (set up GDAP in Lighthouse).
Session recording — the honest position. There is none for cross-tenant admin sessions. What you have is an audit trail: M365 Lighthouse audit logs partner-side ("auditing is enabled for all customers. It can't be disabled", ~1 hour lag, filters reaching back 30 days), Purview and Entra audit logs client-side, and the Azure Activity Log, where "Event initiated by" names the user but "The tenant and role belonging to that user aren't shown", retained 90 days (Lighthouse audit logs; service provider activity). RMM-side recording is a product feature where you have it — ConnectWise Control documents extended auditing including session video, and tells administrators to check "how often data and extended auditing videos are deleted" by the database maintenance plan (Audit page). Find your purge interval before you need it.
And one artifact no vendor produces for you: in the client's logs your technicians appear as "{Governing tenant name} Technician" with a username of the form user_{object ID, dashes removed}, by design (monitor governing tenant activity). The customer-side log alone cannot name your technician. The object-ID-to-human mapping, with employment dates, lives in your tenant. Export it monthly.
Offboarding that actually removes delegated access. This is where the findings cluster: CC6.8 (timely removal of access for terminated users) and CC6.3 (periodic access review) are the most frequent SOC 2 exception sources (Scrut) — for an MSP, delegated-admin group membership persisting after a technician leaves, and stale access to clients you no longer serve. AA22-131A names the direction people forget too: "disabling MSP accounts can be overlooked when a contract terminates."
| # | Leaver action | Who | Evidence |
|---|---|---|---|
| 1 | Revoke sessions and refresh tokens on the admin identity, then disable it | Practice Owner | Identity audit entry, UTC |
| 2 | Remove from every partner-tenant group mapped to a delegated-admin relationship | Practice Owner | Before/after membership export |
| 3 | Remove from Lighthouse authorization groups and JIT-eligible assignments | Practice Owner | Delegations export |
| 4 | Disable RMM, PSA, backup, documentation and vault accounts | Tenant Lead | Per-console audit entry |
| 5 | Rotate every shared credential the leaver could reach | Tenant Lead | Vault audit log |
| 6 | Update the object-ID → technician → dates map | Practice Owner | Dated CSV in the evidence store |
| 7 | Revalidate break-glass accounts | Practice Owner | Test record, date, witness |
Step 7 is not optional after a departure: validate emergency access accounts "at least every 90 days" and additionally after IT staff changes.
Takeaway: run a quarterly access review that produces an artifact, not a feeling — group membership, delegated-admin relationships and Lighthouse delegations diffed against current staff and client lists, filed with a date and a reviewer name. A quarter of missing access reviews is the one finding no year-end scramble can fix.
#5. RMM and PSA hardening
Your RMM is a distribution mechanism. That is not a metaphor, it is the product's job description, and it is why six of the landmark incidents share one shape: an internet-facing management server belonging to the provider was authentication-bypassed, and the attacker then used the product's legitimate, documented, intended feature — push a script, push an installer, take control — to reach every endpoint downstream. No malware on the wire until the last step.
| What actually failed | The control that answers it |
|---|---|
| Kaseya VSA, July 2021 — internet-reachable on-prem server, agent channel trusted to run code, no gate between "a procedure exists" and "it runs on 3,000 machines." Fewer than 60 direct clients → up to 1,500 downstream businesses (NCSC/ODNI) | Console off the public internet; mass-deployment approval gate |
| Kaseya, 2021 — attackers deleted IIS logs and application-database logs as a first stage (Truesec) | Console audit logs shipped where the console cannot reach |
| ScreenConnect, Feb 2024 — CVE-2024-1709, authentication bypass, CVSS 10.0, versions ≤23.9.7, fixed in 23.9.8; it reached the setup wizard on a live instance and overwrote the user database, deleting all local users except a new admin. Over 8,200 publicly accessible servers counted on 21 February (Huntress; Sophos) | Same-day patch SLA; treat "I can't log into my console" as an indicator of compromise, not an outage |
| SimpleHelp, 2025 — an MSP's own instance compromised; the attacker used the RMM's inventory features to collect "device names and configuration, users, and network connections" across multiple customer estates, then pushed a malicious installer to client endpoints (Sophos X-Ops) | Alerting on mass-deployment and mass-inventory actions; per-technician device scoping |
| N-able N-central, Aug 2026 — CVE-2026-18577 was an incomplete patch for CVE-2026-18556: Hotfix 1 (2026.3.1.7, 2 Aug) did not close it; Hotfix 2 (2026.3.1.10, 6 Aug) added "additional hardening" for a "related attack path." CISA gave federal agencies until 6 August under BOD 26-04 (N-able; Rapid7) | Re-check every vendor advisory 72 hours after you patch |
| N-central post-exploitation — Take Control sessions against domain controllers, scripts and jobs pushed to many or all managed endpoints, unauthorized admin accounts, Cloudflare tunnels (Huntress) | Detect "unexpected use of my own console," not "the malware" |
The operating standard, in implementation order:
| # | Control | Grounding |
|---|---|---|
| 1 | Management interface off the public internet — VPN or IP allow-list only | "Place administrative interfaces of RMM behind a VPN or a firewall on a dedicated administrative network"; "allowlisting to limit communication with RMM capabilities to known IP address pairs" (CISA/FBI). Removes the precondition for four of the six incidents above |
| 2 | MFA on every console account, including in-product vendor support accounts; disable those unless actively needed | N-able's own post-exploitation advice; prefer FIDO/WebAuthn or PKI (AA23-320A) |
| 3 | Ship RMM and PSA audit logs off-platform | "Keep direct access to log servers — and the ability to delete or alter logs — out of reach of RMM tools" (CISA Guide). If the console is the compromised thing, its own log view is not evidence |
| 4 | Patch the console in hours on a KEV-class SLA, then re-read the advisory 72 hours later | 1,077 unique IPs were still on outdated N-central builds on 15 Aug 2025 (SecurityWeek); prioritize by KEV, not CVSS alone |
| 5 | Scope every technician to the clients they support; never "All Devices" | Datto RMM documents per-area security levels and agent policies that "prevent unauthorized job execution on sensitive devices" (Datto RMM); ScreenConnect documents role-based security with session groups (ConnectWise) |
| 6 | Script and mass-deployment approval workflow, with re-authentication on scale | "If an account attempts to push commands to 10 or more devices within an hour, retrigger security protocols, such as multifactor authentication (MFA), to ensure the source is legitimate" (CISA Guide) |
| 7 | Alert on mass-deploy actions, unexpected console accounts, unexplained password resets, support-account activity at odd hours | Exactly the 2026 N-central hunting list (Huntress) |
| 8 | Enable device approval; monitor agent encryption key changes | A key change "may indicate a legitimate reinstallation of the Agent or an attempt by an attacker to masquerade one device as another" (Datto RMM) |
| 9 | Allowlist your RMM; block installation and portable execution of every other; review EDR exclusions for RMM install paths | Attackers run RMM as "self-contained, portable executables" that "do not require administrator privileges", bypassing install-blocking; and "often RMM install paths are excluded from EDR inspection" (AA23-025A; CISA Guide) |
Item 9 has a free companion: LOLRMM, which CISA and partners name as "One open-source resource for identifying IOCs and Sigma rules associated with remote access tools" (LOLRMM). Once only one RMM is permitted to exist on your estate, "is this ours?" stops being a judgment call and becomes a lookup.
PSA, honestly. I have no vendor documentation for PSA-specific hardening and will not invent any. What is defensible: the PSA holds your client contact list, your notification SLAs, your ticket history and the insurer's "when did you know" timestamp. Apply the four controls that need no vendor guidance — MFA on every account, scoped API keys with an owner and an expiry, audit logs exported off-platform, and the assumption that the PSA may itself be down, which is why the hard-copy contact list exists.
Actionable takeaway: this week, confirm every management console — RMM, PSA, remote access, managed file transfer, backup — is on the current build, has MFA on every account including vendor support accounts, and is not reachable from the open internet. Then set a reminder to re-read each advisory 72 hours after patching. Not "when we next look." Seventy-two hours, on the calendar, with a name on it.
#6. Cross-tenant hunting capability
Build this in peacetime, when you can afford for it to be slow and wrong. Three parts: tooling that asks one question of many tenants, a query library written before the incident, and a written statement of what the sweep cannot see — the part that separates a professional answer from a reassuring one.
| Route | Reach | Hard limits |
|---|---|---|
| Defender multitenant advanced hunting | Endpoint, identity and mail telemetry where your role is assigned | 100 target tenants per view; 50,000 rows total, divided by the number of tenants queried; 30-day native lookback (mto-requirements; mto-advanced-hunting) |
Graph POST /security/runHuntingQuery, one call per tenant | Same data, scripted, scales past the console view | Needs ThreatHunting.Read.All and a multi-tenant app consented per client; "If a time filter is specified in both the query and the startTime parameter, the shorter time span is applied" (Graph) |
Sentinel cross-workspace union workspace("…") | Older than 30 days, or outside the Defender schema | "You can include up to 20 workspaces in a single query. However, for good performance, we recommend including no more than 5"; "can't scale above 100 workspaces" (extend Sentinel; workspace design) |
Do the arithmetic before you run: at 100 tenants in one view you get 500 rows per tenant. A tightly filtered query on one indicator is fine; a broad process-events sweep truncates silently and hands you a comforting, wrong answer.
Blocking is a separate operation with its own ceiling: 15,000 indicators per tenant, and "Increases to this limit aren't supported"; 500 per CSV batch; no CIDR notation for IP indicators (indicators overview; manage indicators). At 200 tenants that is 200 pushes: automate it, set an expiry on incident IOCs so they age out, and keep a per-tenant indicator inventory.
The query library lives in version control, one file per query, each carrying the question it answers in plain English, the indicator types it accepts, the expected row volume, the tenants it does not work in and why, and the date it last ran. At minimum: hash/IP/domain/URL presence; a second remote-access product on a managed endpoint; console account creation, password resets and privileged role assignment; partner-identity sign-ins outside working hours; directory changes by a partner identity with no corresponding interactive sign-in; and script or job pushes to more than ten endpoints in an hour.
The honest limits. A Microsoft-stack sweep answers exactly one question: has this indicator been seen in the last 30 days, on endpoints or mail, in tenants where my role is assigned and the license carries the telemetry. It does not cover tenants without hunting access or the right tier, Sentinel-only data (unreachable via GDAP at all — that needs B2B plus Lighthouse), anything past retention, non-Microsoft EDR estates, or firewall telemetry that never reached a workspace.
The cheap version. Outside the Microsoft stack, a scripted loop over your tenant list calling each vendor's API, writing per-tenant CSVs with a run timestamp and an explicit "query failed here, reason X" row, is a legitimate cross-tenant sweep. Slower and uglier; same two findings, and the gaps file is the part that matters.
Takeaway: run the sweep once this quarter against a harmless indicator, from a cold start, and time it. That number is your real detection fan-out time, and it belongs in your incident plan.
#7. The "is it us?" self-check
Patient zero might be you — it has been, for other MSPs, in at least six documented cases. This is a standing procedure, runnable by one technician in under thirty minutes, so the question gets a disciplined answer rather than an anxious one. Run it whenever an indicator appears at two or more clients that share nothing but you; anything anomalous touches your RMM, PSA, backup console, a technician identity or the partner tenant; a vendor discloses exploitation of a product in your management plane; or a client reports activity your change records cannot explain.
| # | Step | Time | Done when | Evidence |
|---|---|---|---|---|
| 1 | Freeze the question. Record the indicator, client, UTC timestamp, reporter. Begin no remediation | 0–2 min | Written in the incident log | Log entry, UTC |
| 2 | Correlate: is the indicator present at more than one client that shares nothing but us? Run the pre-written sweep from section 6 | 2–10 min | Sweep complete; hits and gaps both recorded | Query text, results, per-tenant hit/no-data table |
| 3 | Console record: does the RMM show script pushes, jobs, Take Control or remote sessions in the window that we cannot attribute to a named technician and a ticket? | 10–16 min | Every session and job attributed or flagged | Audit export, from the off-platform copy |
| 4 | Console integrity: new or unexpected console accounts, unexplained password resets, support-account activity at odd hours, inability to log in, new extensions or scripts nobody created? | 16–20 min | Account list reconciled against the staff list | Console user list with reconciliation notes |
| 5 | Log continuity: is there a gap in console or web-server logs at the relevant time? | 20–23 min | Continuity confirmed, or gap documented with start and end | Continuity note with timestamps |
| 6 | Delegated-admin path: does the client's audit log show a privileged change by a partner identity with no corresponding interactive sign-in? | 23–28 min | Each change matched to a sign-in, or flagged | Audit and sign-in exports filtered on Technician |
| 7 | Declare. Any unattributed result in steps 3–6 escalates to platform blast radius and wakes the MSP Incident Commander and Practice Owner | 28–30 min | Escalation made and acknowledged | Escalation timestamp and acknowledgment |
Why that order. Step 5 is not a formality — the Kaseya attackers deleted IIS and application-database logs as a first stage (Truesec); a hole at exactly the interesting time makes you patient zero until proven otherwise, and it only works if step 3 read the off-platform copy — which is why log export is item 3 in the hardening list, not item 9. Step 6 closes a specific forensic gap: partner access via PowerShell produces no customer-side sign-in record — only the resulting modifications appear in the customer's audit logs, whereas portal and API access do produce sign-ins (AADInternals). A privileged change by a partner identity with no matching sign-in is either your own automation or someone holding your credentials, and only your records can tell which. Step 1 forbids remediation because the instinct to clean the endpoint — a good instinct, held by good engineers — destroys the artifact that proves provenance and tips off an adversary still inside your management plane.
Actionable takeaway: print this and put it in the on-call bag beside the client contact list and the matrix summary. Drill it once against a benign indicator and time it. If it takes ninety minutes rather than thirty, that is your real number — fix the slowest step before you need it.
#8. The logging and evidence baseline, per tenant
Here is a sentence you want to say to a regulator, an insurer or a client's counsel without hesitating: "That log source was not enabled during that window, we know exactly why, and here is the record showing when it started." That is the difference between a gap and an explained gap. One is a limitation of the evidence; the other is read as adverse.
The artifact is a per-tenant baseline record, refreshed on every license change, one row per log source:
| Field | Note |
|---|---|
| Log source and destination | e.g. identity sign-in logs → client's own workspace |
| Enabled (Y/N) | |
| Retention | Portal retention by license, plus workspace retention |
| Since when | Date the setting was created — the field that does the work |
| License tier governing it | |
| Known limitation | e.g. not retroactive; up to three days before data appears |
| Last verified | Date and reviewer |
Cover identity sign-in and audit logs, the Unified Audit Log, EDR incidents and device timelines, SIEM tables and tiers, firewall and VPN logs with configured retention and quota, the RMM activity log, the PSA ticket audit trail, and backup job and restore-point history. The defaults you are baselining against:
- Identity activity logs: 7 to 30 days by license, non-retroactive. Diagnostic settings to a workspace are the only durable route, they capture only what is generated after configuration, and "it might take up to three days for the logs to start appearing in the destination" (Entra data retention; configure diagnostic settings).
- Unified Audit Log: 180 days by default; with the higher tier, identity, Exchange and SharePoint activity is retained a year.
Search-UnifiedAuditLogreturns a maximum of 50,000 records per run — exactly 50,000 back means records were dropped and you must split the range (audit log retention; search script). That cmdlet runs in Exchange Online PowerShell against a connected session, so confirm — per tenant, in daylight — that the module is installed on the on-call machine and the connection actually works. At 02:00 you want a row-cap problem, not an install-and-authenticate problem. - Defender: 180 days portal-visible, 30 days in advanced hunting, deleted no later than 180 days from contract termination (Defender XDR data retention).
- Data residency is set at provisioning and cannot be changed: "Once configured, you cannot change the location where your data is stored" (Defender for Endpoint data storage). Record each client's region — it answers "where was our data processed" in a breach questionnaire.
- RMM logs have their own caps. Datto RMM retains activity log data for 180 days and "A maximum number of 500 rows can be exported to a single CSV file" (Datto RMM Activity Log). Segment by date and site, and record the boundaries.
The baseline your clients will cite is six months — "store their most important logs for at least six months" (AA22-131A). Notice how many defaults above fall short of it without deliberate configuration.
Onboarding tasks that must never become incident tasks, because each needs a permission or a lead time the on-call will not have at 02:00: configure identity diagnostic settings to the client's workspace (not retroactive, three-day lead, and the relevant monitoring configuration is not supported through GDAP — supported workloads); enable Defender live response, which requires "Manage Portal Settings" and cannot be enabled and used in the same breath without it (live response); set the audit tier and create retention policies; assign eDiscovery roles in the client's own portal with two named people per case, because if the only member leaves "there's no way to access the data in the case" except via an eDiscovery Administrator (eDiscovery permissions); sign your IR collection scripts, since unsigned execution is a tenant-level toggle; and enforce NTP, UTC and ISO 8601 across your estate, PSA and workstations included, because cross-tenant timeline reconstruction is impossible without one trustworthy clock (joint logging guidance).
Standing monthly exports to your own store: delegated-admin activity logs and Lighthouse audit logs (that filter reaches back 30 days; expired relationships vanish after a year), RMM audit logs under the row cap with a segment index, and the object-ID-to-technician map. By the time anyone asks, the consoles will not have it.
Takeaway: the highest-leverage onboarding task an MSP can perform is streaming every client tenant's identity logs into that client's own workspace on day one. Without it, a breach discovered on day 45 has no identity evidence — in that tenant, permanently, with no remedy.
#9. Backups the MSP manages
The trust direction surprises people. CVE-2024-42448: "From the VSPC management agent machine, under the condition that the management agent is authorized on the server, it is possible to perform Remote Code Execution (RCE) on the VSPC server machine" — CVSS 9.9, paired with CVE-2024-42449, where an authorized agent can "leak an NTLM hash of the VSPC server service account and delete files on the VSPC server machine." Affected through 8.1.0.21377, fixed in 8.1.0.21999, "No mitigation available" other than upgrading (Veeam KB4679).
And the backup server is a credential store. CISA's Akira advisory documents VeeamHax.exe, a "plaintext credential leaking tool", and Veeam-Get-Creds.ps1 for "obtaining and decrypting accounts from Veeam servers", alongside deletion of Volume Shadow Copy Service copies (AA24-109A). Compromise the backup server, get the service account; in a shared-credential design, the service account gets every client. That is the whole argument for row 1.
| # | Control | Grounding |
|---|---|---|
| 1 | Per-client credentials, repository and encryption key | "Do not reuse admin credentials across multiple customers" (AA22-131A). One shared backup service account turns a single client incident into a portfolio event |
| 2 | Immutability the console's own administrator cannot revoke — object-lock or WORM at the storage layer, not a toggle in the console an attacker will own | "Ensure all backup data is encrypted, immutable (i.e., cannot be altered or deleted), and covers the entire organization's data infrastructure" (AA24-109A) |
| 3 | Offline copies with separate offline encryption keys | "Store backups separately and isolate from network connections"; "Maintain offline, encrypted backups with separate encryption keys" (AA22-131A) |
| 4 | KEV-class patch SLA for the multi-tenant console; interface off the internet | The VSPC CVEs had no mitigation other than upgrade |
| 5 | Alert on backup deletion, retention shortening, immutability disablement and job disablement | These precede encryption — and look identical to routine administration unless you alert on them |
| 6 | Test restores, not backup jobs | "Regularly update and test backups — including 'gold images'" (AA22-131A); NCSC tells buyers to demand "automated off-site backups with regular restore testing" (NCSC) |
| 7 | Treat the backup console as in scope for every client incident | If Client A is compromised and their agent is authorized on your shared console, the console is in scope until proven otherwise |
Capture restore-point inventory with immutability state and expiry as a standing monthly export. "What could we have recovered, and from when" then becomes a report rather than an archaeology project — and it is the same data the insurer needs for a business-interruption calculation.
Takeaway: if one backup admin credential covers every client you serve, it is the most valuable single object in your business and the attacker knows it. Split it, starting with the clients whose downtime you could not survive.
#10. Staffing and escalation that does not depend on heroics
The twelve-person MSP at 23:45, honestly: one technician on call, covering the whole book, with a phone, a laptop, a VPN and a multi-tenant console. Not a forensic analyst, not a lawyer, and with work in the morning. That is the norm, not a deficiency, and every design decision below assumes it.
- The on-call technician is a first responder, not a decision-maker. At 02:00 their job is: execute the pre-authorized actions, start the log, call the Client Authority Holder, run the "is it us?" check, escalate. Nothing there needs a judgment about contractual authority, because the matrix made those calls in daylight.
- Pre-position every permission the on-call will not have — live response enablement, diagnostic settings, eDiscovery roles. If they are 02:00 tasks, the evidence is already gone.
- Staff the elevation path or drop the gate. An approval-gated JIT path with a single sleeping approver fails closed at exactly the wrong moment.
- Two names per shift — a primary and a named escalation. At twelve people the second name is often the Practice Owner, and that is fine. A rota with one name and an assumption that someone will pick up is not.
- Make the portfolio trigger provenance, not severity. Any indicator touching the RMM, PSA, backup console, a technician identity or the partner tenant is a portfolio event from minute one. CISA's instruction to a compromised RMM operator is to isolate the server and contact downstream customers (AA25-163A) — a decision no single on-call technician should make alone, at any hour.
- Hard copies. "Maintain up-to-date hard copies of plans so responders can access them" (AA22-131A). The on-call bag holds the client contact list, the matrix summary, the self-check procedure, the notification numbers, your carrier's claims line and the break-glass envelope location. Your PSA may be the thing that is down.
One more thing that costs nothing and is skipped almost universally: decide in daylight who may declare an incident, who signs a client notification, when counsel and the insurer are engaged, and whether a technician may reimage before imaging. The default on the last is no. Restoring a client fast is your instinct and your job, and doing it before preservation destroys their evidence and your defense. Where a DFARS 252.204-7012 flow-down applies there is also a preservation duty running at least 90 days from submission of the cyber incident report to DoD, covering images of all known affected information systems and monitoring data (acquisition.gov) — which makes "reimage first" not merely unwise but non-compliant.
Actionable takeaway: publish an on-call rota with two names per shift, a documented handover, and a written statement of what the primary may do without calling anyone. Then run one unannounced call test per quarter — to your own rota, and to your ten largest clients' authority holders. The test costs an evening. The alternative costs a client base.
All of this is unglamorous, and all of it is cheaper than the night it prevents. The authority matrix is a spreadsheet. The clauses are a conversation with a lawyer you were having anyway. Isolation starts with a second username. Console hardening is mostly turning things off. The self-check is a printed page. Do the ten largest clients first, do them this quarter, and let the rest follow at renewal. Stay patched, stay documented, and never let a tired engineer be the only thing standing between a client and a decision nobody authorized.
#Checklist
- MSP-01A Client Authority Matrix exists for every managed client, naming a primary Client Authority Holder and a deputy with out-of-hours contact details.
[IG1][GV.RR-02] - MSP-02The matrix marks each disruptive action (isolate endpoint / server / DC, disable user, tenant-wide credential reset, block egress, suspend service, restore from backup, engage DFIR) as pre-authorized, approval-required or prohibited, with a named approving role.
[IG1] - MSP-03Each client's out-of-hours contact path has been dialled and verified within the last 180 days, with the date and tester recorded.
[IG1] - MSP-04The matrix defines an unreachable-escalation window in minutes, agreed in writing per client, after which specified actions become pre-authorized, with a post-hoc notification clock.
[IG2] - MSP-05The authority tier agreed with each client is enforced in RBAC — identity roles, EDR role assignments and RMM device scopes — not only in a document.
[IG2][PR.AA-05] - MSP-06Every managed client has a signed master agreement on file; an exception register lists any client without one, with a remediation date and owner.
[IG1] - MSP-07The master agreement contains an emergency action clause with a defined trigger, a minimum-extent-and-duration limiter, and a post-hoc notice obligation on a fixed clock.
[IG2] - MSP-08The master agreement contains a security suspension or disconnection right distinct from the non-payment suspension right.
[IG2] - MSP-09The master agreement contains an evidence-preservation obligation and an express legal-hold carve-out to the data-deletion and data-return clause.
[IG2][A.5.28] - MSP-10The notification deadline, legal notice address and named recipient for each client are structured PSA fields readable by the on-call technician, not a PDF in a document store.
[IG1][RS.CO-02] - MSP-11The complementary user entity controls in the MSP's SOC 2 report have been reconciled against client-responsibility clauses in each master agreement within the last 12 months.
[IG2][CC7.x] - MSP-12Every client contract requires the client to carry cyber liability insurance and to provide a waiver of subrogation in the MSP's favor, and the waiver has been checked against the client's own policy wording, since some policies restrict or override contractual waivers; exceptions are recorded and approved by the Practice Owner.
[IG2] - MSP-13A data processing addendum is in place for every client whose processing falls under GDPR or UK GDPR, reciting the pre-authorized containment actions as documented instructions of the controller under Article 28(3)(a) and naming the MSP's sub-processors; for clients outside those regimes, the equivalent processing terms required by the applicable state or national law.
[IG2] - MSP-14Answers given on a client's cyber insurance application on the client's behalf are provided in writing, dated, with supporting evidence attached, and retained.
[IG1] - MSP-15Administrative identities holding delegated access to client tenants are separate from those used for the MSP's own email and productivity work, with no shared or generic management accounts.
[IG1][PR.AA-01][CC6.3] - MSP-16Phishing-resistant MFA is enforced on every identity that can reach a client tenant or a management console, with no exception group; exceptions are documented, time-bound and approved by the Practice Owner.
[IG1][PR.AA-03] - MSP-17Administrative work on client tenants is performed only from a hardened, dedicated workstation with mail and general browsing restricted.
[IG2][PR.AA-05] - MSP-18Standing delegated access is read/triage only; write and privileged roles are obtained just-in-time through a documented activation path staffed whenever the on-call rota is live.
[IG2] - MSP-19Every delegated-admin relationship has been reviewed within the last 90 days, and Privileged Role Administrator and Privileged Authentication Administrator removed from any relationship that does not require them.
[IG2][PR.AA-05] - MSP-20At least two emergency access accounts exist in the MSP management plane — cloud-only, phishing-resistant, tied to no individual, excluded from lockout-capable conditional access — with alerting on every sign-in and validation at least every 90 days and after any IT staff change.
[IG2] - MSP-21Technician offboarding removes delegated access in a documented order and produces evidence: token revocation, delegated-admin group removal, Lighthouse and JIT removal, console account disablement, credential rotation, object-ID map update and break-glass revalidation.
[IG1][CC6.8] - MSP-22A quarterly access review is filed as an artifact — group membership, delegated-admin relationships and Lighthouse delegations diffed against current staff and client lists, with a reviewer name and date.
[IG2][CC6.3] - MSP-23No management console (RMM, PSA, remote access, managed file transfer, backup) exposes an administrative interface to the public internet; access is restricted by VPN or IP allow-list.
[IG1] - MSP-24Management-plane products are patched on a KEV-class SLA measured in hours, and every advisory acted on is re-checked 72 hours after patching for a superseding or incomplete fix.
[IG1] - MSP-25A script and mass-deployment approval workflow is in force, and an alert fires on any action pushing commands, scripts or installers to ten or more devices within an hour.
[IG2][DE.CM-09] - MSP-26RMM and PSA audit logs are shipped to a store the RMM and PSA cannot delete or alter, with at least six months of retention and documented export segmentation where row caps apply.
[IG2][A.8.15] - MSP-27A pre-written cross-tenant hunting query library exists in version control, has been executed end-to-end within the last 90 days, and records which tenants returned no matching records and which could not be queried, with the reason.
[IG2][DE.CM-01] - MSP-28A written "is it us?" self-check procedure exists, is runnable by one technician in under 30 minutes, is held in hard copy in the on-call bag, and has been drilled within the last 12 months.
[IG2][A.5.24] - MSP-29A per-tenant evidence baseline record states, for each log source, whether it is enabled, its destination, its retention, its license tier, the date it started and the date last verified.
[IG2][A.8.15][A.8.17] - MSP-30Backups are held with per-client credentials, repositories and encryption keys, with immutability enforced at the storage layer beyond the reach of backup-console administrators, and restore testing evidenced within the last 90 days.
[IG1][A1.x][RC.RP-01]
#Sources
- CISA and partners: AA22-131A — Protecting Against Cyber Threats to MSPs and their Customers · AA23-025A — Malicious Use of RMM Software · Guide to Securing Remote Access Software · AA23-320A — Scattered Spider · AA25-163A — SimpleHelp · AA24-109A — Akira · Kaseya guidance for affected MSPs · Best Practices for Event Logging and Threat Detection
- ODNI/NCSC — Kaseya VSA factsheet · NIST SP 800-61r3 · NCSC-UK — Choosing an MSP · NCSC-UK — Using MSPs to administer your cloud services
- Sophos MDR Service Description · DragonForce / SimpleHelp · ScreenConnect exploitation · 2026 Active Adversary Report
- Huntress — ScreenConnect authentication bypass · Huntress — N-central exploitation · N-able security update, 10 Aug 2026 · Rapid7 — CVE-2026-18577 · SecurityWeek · Truesec — Kaseya analysis · Mandiant — ScreenConnect hardening · Microsoft Security Blog — NOBELIUM and delegated admin · AADInternals — Microsoft partners
- Microsoft Learn: Microsoft-led DAP→GDAP transition · GDAP FAQ · Least-privileged roles by task · Workloads supported by GDAP · Partner Center October 2025 announcements · Cross-tenant access settings · Emergency access accounts · Entra data retention · Configure diagnostic settings · Monitor governing tenant admin activity
- Microsoft Learn: Azure Lighthouse cross-tenant management · Create eligible authorizations · Monitor service provider activity · Manage Sentinel workspaces at scale · Set up GDAP in M365 Lighthouse · Review Lighthouse audit logs
- Microsoft Learn: Defender multitenant requirements · Multitenant advanced hunting · Defender Experts managed response · Defender XDR data retention · Defender for Endpoint data storage · Live response · Indicators overview · Manage indicators · Extend Sentinel across workspaces · Workspace design · Audit log retention policies · Audit log search script · eDiscovery permissions · Graph runHuntingQuery · Privileged access devices
- Datto RMM — Security best practices · Datto RMM Activity Log · ConnectWise Control Security Best Practice Guide · ConnectWise Control Audit page · Veeam KB4679
- CompassMSP Master Service Agreement · Google Workspace terms · Law Insider — Security Incident · Law Insider — Suspension of Services · EDPB Guidelines 9/2022 · GDPR Article 28 · GDPR Article 82 · 18 U.S.C. § 1030 · DFARS 252.204-7012 · PCI DSS v4.0.1
- ICO — Advanced Computer Software Group enforcement · Clifford Chance · BleepingComputer · · Insurance Journal · Hunton — cyber insurer subrogation · Scrut — SOC 2 audit exceptions · ScalePad — SOC 2 MSP field guide · Secureframe on CUECs · LOLRMM
#Part 3 — The MSP-Side Playbooks
#The MSP-Side Playbooks
Every playbook in the main book assumes the responder is standing inside the affected organization. These five assume something worse: that the affected organization is you, and that between two and two hundred other companies are downstream of whatever you decide in the next hour.
That changes three things structurally. First, your tooling is in scope, so the console you would normally reach for is evidence, not equipment. Second, containment is a portfolio decision — killing your own RMM protects 180 clients and blinds you at all 180 simultaneously. Third, notification is not one email, it is N legally distinct notifications with client-specific content, several of them on contractual clocks tighter than any regulator's.
Each playbook below carries a default severity, a blast radius (single-tenant, multi-tenant, platform), entry criteria, five phases as step tables, decision points with a decide-by time and a named authority, notification triggers, and the pitfalls. Two markers appear in the step tables and they are not decoration:
Adjust the role names to your shop. Do not adjust the order.
These five carry no checklist codes of their own, because what makes them runnable is audited elsewhere in this book: MSP-19 and MSP-21 for the delegation and offboarding paths, MSP-23 through MSP-28 for management-plane exposure, patch SLA, mass-deployment approval, log shipping, the cross-tenant query library and the drilled self-check, MSP-42 for the offline contact list and the channel that does not traverse your RMM, and MSP-45 for the rehearsed notification chain. Green on those is what gives these playbooks somewhere to land.
#PB-MSP-PLATFORM — Your RMM or PSA is compromised
Default severity: SEV-1. Blast radius: platform. Runs: MSP Incident Commander, with Practice Owner as Executive Sponsor from minute one.
This is the one that ends companies. The landmark MSP incidents share a single shape: an internet-reachable management server belonging to the provider was authentication-bypassed, and the attacker then used the product's legitimate, documented, intended mass-deployment feature to reach every endpoint downstream. In the Kaseya VSA attack of July 2021 the ransomware arrived as an agent procedure named "Kaseya VSA Agent Hot-fix" — fewer than 60 direct MSP clients became up to 1,500 downstream businesses, roughly 25× amplification, on a holiday Friday evening chosen precisely because nobody would look at the console until Tuesday (ODNI/NCSC; Truesec). There was no malware on the wire until the very last step.
Entry criteria — any one of these:
- Console accounts you did not create, or console accounts that have vanished. The ScreenConnect CVE-2024-1709 authentication bypass (CVSS 10.0) reached the setup wizard on a configured production instance and overwrote the internal user database, deleting every existing local user except the attacker's new admin (Huntress). "I can't log into my console" is an indicator of compromise, not an outage.
- A script, job, agent procedure or Take Control session you cannot attribute to a technician and a ticket.
- A vendor advisory naming in-the-wild exploitation of a product you self-host and expose.
- A client reporting a second remote-access agent you did not deploy.
The wrong instinct, and it is a good instinct. The competent technician's first move is to log into the console and look. Do not. Logging in with your own credentials adds your session to the attacker's harvest, and every minute spent reading dashboards is a minute the deployment queue keeps draining. The Kaseya attackers' first action was deleting IIS logs and the logs held in the application database (Truesec). By the time you are reading the console, the console may already be a work of fiction. Your first move is to stop it deploying.
#### Phase 1 — Detection and triage
| # | Action | Who | Done when | Evidence to capture |
|---|---|---|---|---|
| 1 | Declare SEV-1 platform. Move all incident comms to the out-of-band channel (phone bridge + Signal group), not the PSA and not company email | On-Call Technician | Bridge open, Practice Owner joined | Declaration time (UTC), who was paged |
| 2 | Check whether logs were deleted. Compare console web logs, application-database audit records and your off-platform log copy for gaps at the same timestamp. A hole is the finding | MSP Incident Commander | Gap confirmed or ruled out, in writing | Log inventory with first/last event per source; hash the off-platform copy |
| 3 | Pull the deployment record: every script, job, agent procedure and remote session executed in the last 30 days, with initiating account and target count | Tenant Lead | Full export held off-platform | Query text, console, operator, UTC run time, SHA-256 |
| 4 | Identify unauthorized console accounts and password resets. N-central intruders created admin accounts with anomalous appended .invalid strings; hunt support-account activity at unusual times (Huntress) | MSP Incident Commander | Account list reconciled against HR roster | Account creation events, actor, source IP |
| 5 | Determine internet exposure and exact version; check the vendor advisory and the CISA KEV catalog | Operations Lead | Version and exposure documented | Authoritative version string, scan output |
| 6 | Assign the Scribe. Every decision, time and approver goes in a written log kept outside the affected platform | MSP Incident Commander | Log started | The log itself |
#### Phase 2 — Containment
Order matters more here than anywhere else in this book. Kill the deployment capability before you kill anything else. Every documented post-exploitation step in these incidents was a product feature — pushing scripts and jobs to many or all managed endpoints, launching Take Control sessions against downstream domain controllers, deploying additional remote-access clients as backup access (Huntress; Trend Micro). Disconnect the agents first and you have blinded yourself at every client at once, while whatever the attacker has already scheduled in the console is still sitting in the console. You have removed your view of the estate without reliably removing their reach into it. What your particular product does with queued work when an agent reconnects is a question to answer in daylight, in writing, as a named pre-incident task — not at 02:00 with the answer riding on it.
| # | Action | Who | Done when | Evidence to capture |
|---|---|---|---|---|
| 7 | Disable scripting, job scheduling, software deployment and remote-control initiation at the console — the capability, not the server [T] | Operations Lead | Deployment attempt fails in test | Config change record, UTC |
| 8 | Cancel or purge every pending and scheduled job. Verify the queue is empty, then verify again | Operations Lead | Queue empty on two checks | Pre- and post-state export |
| 9 | Isolate the management server from the internet, or stop the server process — CISA's explicit instruction to providers in this position (AA25-163A) [T] | MSP Incident Commander | No inbound reachability from an external host | Firewall rule, timestamp, who approved |
| 10 | Snapshot before you rebuild. Image the console host and preserve its logs, database and web logs before patching or reinstalling [E] if skipped | Operations Lead | Image hashed and stored per-tenant-neutral in the MSP evidence store | Imaging tool name, version, license; hash list stored separately |
| 11 | Force MFA re-enrollment and reset every console credential, including in-product vendor support accounts — disable those unless actively needed (N-able) | MSP Incident Commander | All accounts reset, support accounts disabled | Account-by-account record |
| 12 | Sweep client endpoints for the artifacts of this product's abuse: second remote-access agents, Cloudflared services, svchost.exe outside System32, unexpected installers | Tenant Lead | Sweep run across the estate, results per tenant | Per-tenant query, per-tenant result, "no data" versus "no result" recorded separately |
#### Phase 3 — Eradication
| # | Action | Who | Done when | Evidence to capture |
|---|---|---|---|---|
| 13 | Rebuild the management server from known-good media rather than patching in place. Patching is not remediation if you were already popped (Mandiant) | Operations Lead | New host up, old host preserved offline | Build record, source media hash |
| 14 | Re-check the vendor advisory 72 hours after you patch. N-central's Hotfix 1 did not close the hole; Hotfix 2 followed four days later (N-able) | Operations Lead | Advisory re-read, version current | Dated note of the re-check |
| 15 | Remove attacker persistence downstream: unauthorized accounts, tunnels, scheduled tasks, second RMM agents, and any EDR exclusion covering an RMM install path (CISA Guide) | Tenant Lead | Per-tenant clearance recorded | Per-tenant remediation log |
| 16 | Move the management interface behind a VPN or firewall on a dedicated administrative network; allowlist RMM communications to approved IP address pairs (CISA/FBI) | Operations Lead | External scan shows no exposed console | Scan output |
#### Phase 4 — Recovery and Phase 5 — Post-incident
Re-enable agents in waves, smallest and least critical client first, with a hold point between waves. Restore deployment capability last, with a two-person approval on any job targeting more than one client, and implement the mass-scripting safeguard while you are in there: re-trigger MFA when an account pushes commands to 10 or more devices within an hour (CISA Guide). Confirm the notification register is complete per client, not in aggregate.
Then produce a per-tenant evidence pack, not one master report, and reconcile your own diligence file — console patch history, MFA posture for every console account, the exception register with named approvers and expiry dates, alerting configuration and its change history. Answer in writing the question the ICO asked Advanced: was there a control you had built and not turned on? Advanced had a working MFA solution it had not rolled out, on a perception that customers would resist, and that perception cost £3,076,320 (Clifford Chance).
Actionable takeaway: today, write down how you would reach every one of your clients with your RMM and your PSA both dark — the list, the channel, and who holds the offline copy. If that takes you more than a minute to answer, that gap is what this playbook is about.
#PB-MSP-CROSS — Cross-tenant lateral movement
Default severity: SEV-2, escalating to SEV-1 on confirmation. Blast radius: multi-tenant. Runs: MSP Incident Commander, one Tenant Lead per affected client.
An indicator lands at one client. Three hours later you find it at a second client that shares no network, no supplier, no vertical and no staff with the first. The only thing those two companies have in common is you. That is not a coincidence to be explained; it is a finding to be acted on.
Entry criteria: the same file hash, IP, domain or attacker-created account observed at two or more clients with no shared infrastructure; or any indicator that touches the RMM, PSA, backup console, a technician identity or the partner tenant — which is a portfolio event from the first minute regardless of how small it looks.
The wrong instinct. The reflex is to work the second client's incident as a second incident, and it is the wrong shape entirely. Two single-tenant investigations run in parallel will each conclude "client-specific" and neither will ask the question that matters. The moment an indicator appears at unrelated tenants, stop opening tickets and start running the provenance test.
#### Phase 1 — Detection and triage
The three questions, in order, and each is answerable (synthesised from the documented cases in the pack):
| # | Action | Who | Done when | Evidence to capture |
|---|---|---|---|---|
| 1 | Q1. Is the indicator present at more than one client that shares nothing but us? Run the sweep before you theorize | MSP Incident Commander | Sweep complete across all tenants in the licensed, role-assigned set | Per-tenant query text, result, and "no data" recorded separately from "no result" |
| 2 | Q2. Did it arrive through our delivery channel? Check console deployment records, remote-session logs, and console web logs for the same window | Operations Lead | Console record reconciled against endpoint timeline | Console export, hashed |
| 3 | Q3 (Microsoft path). Is there a privileged change in the client tenant attributable to a partner identity with no matching interactive sign-in? Partner access via PowerShell produces no customer-side sign-in record — only the resulting modifications appear in the audit log (AADInternals) | Tenant Lead | Audit/sign-in correlation done per affected tenant | Both log exports, per tenant |
| 4 | Budget the sweep before you run it. Defender multitenant management caps at 100 target tenants and returns 50,000 rows total divided by the number of tenants queried — at 100 tenants, 500 rows each (Microsoft Learn). For anything older than the 30-day lookback, use Sentinel cross-workspace: 20 workspaces maximum per query, 5 recommended (Microsoft Learn) | Operations Lead | Query filtered to a specific indicator, few columns, no take; long-lookback batches run | Query text, batch list, results |
#### Phase 2 — Containment
| # | Action | Who | Done when | Evidence to capture |
|---|---|---|---|---|
| 5 | Place the legal/eDiscovery hold in each affected tenant before any containment action. Holds are not retroactive, and they cover mailbox, SharePoint, OneDrive and Teams content — that content only | Tenant Lead | Hold confirmed per tenant, two named people on each case | Hold record, custodian list |
| 6 | Export Entra sign-in and audit logs per tenant, now. No hold covers them. Retention is 7–30 days by license, and diagnostic settings streaming to a workspace are the only durable route — also not retroactive [E] if skipped | Tenant Lead | Export held off-platform for every tenant in scope | Export receipts per tenant; the diagnostic-settings configuration and its start date where one exists |
| 7 | Push the indicator as a block, per tenant. Limit is 15,000 indicators per tenant, not increasable; 500 per CSV batch; no CIDR notation (Microsoft Learn) | Operations Lead | Block confirmed in each tenant | Per-tenant push record with expirationTime set |
| 8 | Contain per tenant, at that tenant's authority tier. A pre-authorized action at Client A is not pre-authorized at Client B [T] | Tenant Lead | Actions taken match the per-tenant matrix | Approver name and time per action |
| 9 | If the MSP delivery channel is implicated, stop collecting through it and switch to out-of-band collection | MSP Incident Commander | Alternate collection path in use | Tool name, version, collector identity |
#### Phase 3 — Eradication and Phase 4 — Recovery
Eradicate per tenant and prove it per tenant: attacker accounts removed, persistence cleared, credentials rotated per client rather than through any shared admin account. Recovery is a per-tenant sign-off with the Client Authority Holder, and the portfolio is not recovered until the last tenant is. The structural fix is the one CISA has asked for since 2022: segregate customer datasets and services from each other and from your internal networks, and do not reuse admin credentials across multiple customers (AA22-131A). Shared credentials do not merely widen the blast radius; they make it impossible to distinguish "the attacker used our shared account" from "that client was breached separately" — which is the only question anyone will ask you.
#### Phase 5 — Post-incident: the notification problem, stated honestly
Every affected client must be told. They will ask you who else was affected. You cannot tell them. One client's incident is that client's confidential information, and disclosing it to another client — even a sympathetic one, even to be helpful, even when they have already guessed — is a second incident you have created yourself.
What you can say, and should say identically to everyone: that the incident touched more than one client, that each affected client has been notified individually, that you will not discuss any other client's environment with them and equally will not discuss theirs with anyone else. That last clause is the reassurance, and it is true.
Evidence segregation under pressure is where this playbook is actually lost. Multi-tenant consoles are built to aggregate, and the aggregation is the risk: a cross-tenant KQL query exported to CSV is a commingled export by construction. Three cheap controls. Every evidence-pack query runs scoped to one tenant, with the scope captured in the artifact header. Every export carries a per-artifact attestation — source system, scope filter, operator, UTC timestamp, SHA-256. And a named second person reviews every export for foreign-tenant identifiers before it leaves your custody. That third one catches the unfiltered CSV, and it costs ten minutes.
Actionable takeaway: pick one indicator from last month and run the provenance test on it now, in daylight, across your whole book. Time it. That number is how long the first hour of your worst night is going to take.
#PB-MSP-GDAP — Delegated admin relationship abuse
Default severity: SEV-1. Blast radius: platform. Runs: MSP Incident Commander with Practice Owner.
Microsoft has a name for this pattern, coined while watching NOBELIUM work through cloud service providers, MSPs and resellers to reach their downstream customers: "compromise-one-to-compromise-many." One documented chain crossed four distinct providers to reach a single final target — every provider in that chain was patient zero for the next (Microsoft). No product vulnerability was needed. Password spraying, token theft, API abuse and spear phishing were enough.
What a partner-tenant compromise actually grants downstream. Read the roles the Microsoft-led DAP-to-GDAP transition assigned to the Admin Agents group, because most MSPs never did. Alongside the readers and helpdesk roles it granted Privileged Authentication Administrator — which can reset authentication methods for any user, including the customer's own Global Administrators — and Privileged Role Administrator, which can manage role assignments and PIM, and therefore grant itself anything else (Microsoft Learn). "GDAP by default" is not least privilege; it is full tenant control in two steps, at every client with an active relationship. Legacy DAP was worse: an Admin Agent was Global Administrator in every DAP customer tenant, with no way to scope a partner user to a specific tenant (AADInternals).
Entry criteria: anomalous sign-in to the partner tenant or Partner Center; a GDAP relationship you did not request; privileged changes in a client tenant attributable to a partner identity with no matching interactive sign-in; a dormant partner account suddenly updated with MFA before login; Azure RunCommand paired with AOBO access; or a client reporting that "your technician" did something no technician did.
The wrong instinct. The reflex is to revoke every GDAP relationship immediately. It is understandable and it is a trap: revoking delegation also removes your ability to investigate and respond in the client tenants you just cut yourself off from, and each termination fires a notification to that client's Global admins before you have written a word of explanation. You will be explaining to 180 clients simultaneously, badly, on their terms.
#### Phase 1 — Detection and triage
| # | Action | Who | Done when | Evidence to capture |
|---|---|---|---|---|
| 1 | Enumerate every GDAP relationship, its roles, expiry and assigned security groups. Look for MLT_* names from the Microsoft-led transition, and for relationships nobody requested | MSP Incident Commander | Full inventory exported | Partner Center Delegated access export; set From/To dates explicitly — the activity-log export defaults to the most recent month |
| 2 | Export Partner Center activity logs and Microsoft 365 Lighthouse audit logs. Lighthouse auditing is on by default and cannot be disabled, but its filter is last day / 7 days / 30 days and new logs may take up to an hour to appear (Microsoft Learn) | Operations Lead | Exports held off-platform | CSV plus request bodies |
| 3 | Triage partner-tenant identities against Microsoft's own detection focus list: impossible travel, Entra modifications enabling persistence, activity targeting Global Administrators, dormant accounts updated with MFA before login (Microsoft) | MSP Incident Commander | Every partner identity classified | Sign-in log export |
| 4 | In each high-risk client tenant, filter customer sign-in logs on User contains Technician and audit logs on Initiated by (actor) matching your tenant name. Partner sign-ins render as "{Governing tenant name} Technician" with username user_{object ID, dashes removed} (Microsoft Learn) | Tenant Lead | Per-tenant correlation done | Both exports, per tenant |
| 5 | Verify partner-tenant break-glass accounts are intact and usable before you change anything | Practice Owner | Break-glass tested from the PAW | Test record |
#### Phase 2 — Containment
| # | Action | Who | Done when | Evidence to capture |
|---|---|---|---|---|
| 6 | Contain the partner-tenant identity itself: revoke sessions and reset credentials in one action, remove it from every GDAP security group, revoke its Lighthouse eligibility [T] | MSP Incident Commander | Identity confirmed inert | Action log with UTC times |
| 7 | Enforce MFA and Conditional Access for Partner Center access; audit for any non-Microsoft MFA that does not issue the expected claim — Entra and Partner Center check for the MFA claim, and what they cannot see does not count (Microsoft Learn) | Operations Lead | Every partner account challenged | CA policy export |
| 8 | Remove delegated administrative privileges that are no longer in use — Microsoft's own first mitigation for this campaign | MSP Incident Commander | Unused relationships terminated | Termination record per client |
| 9 | Strip Privileged Role Administrator and Privileged Authentication Administrator from any relationship that does not need them | MSP Incident Commander | Role sets reduced | Before/after role export |
| 10 | Notify each affected client before their Global admins receive the automated GDAP termination notification | Communications Lead | Notification sent per client at the contractual notice address | Acknowledgment register |
#### Phase 3 — Eradication
Rebuild the delegation model rather than restoring it. Three relationships per client: a standing long-duration read/triage set (Security Reader, Global Reader, Reports Reader, Service Support Administrator); a standing medium-duration operational set; and a short-duration relationship carrying Global Administrator only, requested per incident, where the expiry does the revocation for you. The mechanics constrain the design: maximum relationship duration is two years, a relationship carrying Global Administrator cannot be auto-extended, all roles in one relationship share a single expiry, and a name cannot be reused for 365 days after cleanup (GDAP FAQ). Make membership of the group behind the elevated relationship PIM-eligible, approver-gated and MFA-on-activation.
Then fix the partner tenant's own break-glass, because a lockout there costs you administrative reach into every client at once: two or more emergency accounts, cloud-only on *.onmicrosoft.com, not federated or synced, passwordless and phishing-resistant using methods your other admin accounts do not use, Global Administrator assigned active permanent in PIM rather than eligible, excluded from Conditional Access policies that block sign-in, stored in separate secure locations, used only from a privileged access workstation, alerted on every sign-in at severity 0, and validated at least every 90 days and after any IT staff change (Microsoft Learn).
#### Phase 4 — Recovery and Phase 5 — Post-incident
Restore delegation client by client with the Client Authority Holder's explicit agreement, and use the conversation to disclose two asymmetries before someone else finds them. First: the client has no way to independently verify the strength of your technicians' MFA — for GDAP sign-ins, MFA is always required in the technician's home tenant and always trusted in the resource tenant (Microsoft Learn). Your partner tenant's authentication posture is their control. Second: their risky-users report does not reflect external users, because risk is evaluated in the home directory — they cannot remediate a compromised technician identity, only you can. Both are documented; both look like concealment if the client finds them unaided.
Finally, export GDAP activity logs and Lighthouse audit logs monthly to your own store. Relationships in Expired or Terminated state are automatically cleaned up after one year and become invisible to partner and customer alike (GDAP FAQ). By the time anyone litigates, the console will not have it.
Actionable takeaway: open Partner Center this week and list every relationship still carrying Privileged Role Administrator or Privileged Authentication Administrator. Strip the ones nobody can justify, and put a name and a date against every one you keep.
#PB-MSP-TECH — Technician account compromise
Default severity: SEV-1. Blast radius: multi-tenant, potentially platform. Runs: MSP Incident Commander, with the Practice Owner and HR engaged from Phase 2.
One identity. Two hundred companies. That is the whole problem, and it is worth saying out loud that this is rarely a story about a careless technician. Scattered Spider's documented method against contracted IT help desks is layered social engineering across several calls: first to learn what a help desk requires for a password reset, then to gather that information for a targeted employee, then to convince help desk staff to reset passwords or transfer MFA tokens — enriched with PII from social media, OSINT, commercial intelligence tools and database leaks (AA23-320A). People who fall for that are not stupid. They are targeted by professionals with a script written for them.
Entry criteria: impossible-travel or anomalous sign-in on a technician identity; MFA method added or changed outside a known enrollment; a help-desk password reset or MFA transfer that cannot be tied to a verified requester; a technician reporting a strange call; console or client-tenant actions attributed to a technician who was asleep; or a lookalike domain of your brand in the yourname-helpdesk[.]com / yourname-sso[.]com pattern.
The wrong instinct. The reflex is to reset the technician's password immediately. That single action is wrong on three counts. It leaves the adversary logged in — a reset alone leaves access tokens valid for up to 28 hours in Continuous Access Evaluation sessions, application-issued session tokens that Entra cannot revoke at all, and any consented OAuth grant valid indefinitely. It tips off the adversary, because the technician is now locked out and phoning the help desk loudly. And it destroys nothing the attacker cares about while burning your only warning shot.
#### Phase 1 — Detection and triage
| # | Action | Who | Done when | Evidence to capture |
|---|---|---|---|---|
| 1 | Preserve first. Place holds and start log export before any containment action; Entra sign-in retention is 7–30 days by license and non-retroactive [E] if skipped | MSP Incident Commander | Holds placed in the partner tenant and each candidate client tenant | Hold record; export receipts |
| 2 | Scope what that identity touched, time-boxed to 30 minutes. Enumerate sessions, OAuth grants, registered devices, MFA methods, GDAP group memberships, Lighthouse eligibilities, RMM/PSA/backup console roles | MSP Incident Commander | Reach inventory complete | Group membership export with timestamps |
| 3 | Fan out across client tenants: filter customer sign-in logs on Technician and audit logs on your tenant as actor, for the full suspected window | Tenant Lead | Every tenant in the identity's reach checked | Per-tenant export |
| 4 | Pull remote-session and console records. Microsoft provides no session recording for cross-tenant admin sessions — your trail is Lighthouse audit logs partner-side, Purview Unified Audit Log and Entra audit logs client-side, and Azure Activity Log for Lighthouse actions, retained 90 days in the portal (Microsoft Learn). Review whatever session recording your RMM does provide against the roster, tickets and paging records | Operations Lead | Records exported before rotation; every session matched to a ticket or flagged | Exports with hashes; matched/unmatched session list |
| 5 | Talk to the technician. Early, privately, and as a witness rather than a suspect | MSP Incident Commander | First-hand account captured | Scribe's note: facts and times only |
#### Phase 2 — Containment
| # | Action | Who | Done when | Evidence to capture |
|---|---|---|---|---|
| 6 | Revoke sessions and reset the credential together; remove attacker-added MFA methods; revoke OAuth grants; disable registered devices [T] | MSP Incident Commander | No new tokens issued for the principal | Command output, UTC |
| 7 | Remove the identity from every GDAP security group, Lighthouse authorization, RMM scope and PSA role — group membership is the actual grant, not the account | Operations Lead | Membership zeroed | Before/after export |
| 8 | Where the identity acted inside a client tenant, contain per that tenant's authority tier and notify the Client Authority Holder | Tenant Lead | Per-tenant action taken with a named approver | Approver, time, action |
| 9 | Issue the technician a new identity with a new UPN rather than re-enabling the old one; keep the old object for evidence [E] if deleted | MSP Incident Commander | Technician working again | New object ID recorded in the mapping table |
| 10 | Check for the pattern, not just the account: monitor for yourbrand-helpdesk[.]com and yourbrand-sso[.]com permutations of your brand and your top clients' (AA23-320A) | Operations Lead | Monitoring live | Domain watchlist |
#### Phase 3 — Eradication, Phase 4 — Recovery
Eradication is architectural. Move technician identities to phishing-resistant FIDO2/WebAuthn or certificate-based MFA, which resists the push bombing and SIM swapping this actor uses. Put client-tenant administration on a privileged access workstation — Microsoft's model blocks email and web browsing on that device precisely because they are the phishing vector, and your client-tenant Global Admin identity is Tier 0 for hundreds of organizations, not one. Kill generic shared management accounts; NCSC's expectation is that actions trace to specific people's accounts (NCSC). And harden the help desk with a documented verification procedure for password and MFA-reset requests — that procedure is now a primary attack surface, not an administrative formality.
Recovery is per-client. Each affected Client Authority Holder gets the specific facts for their environment: what the identity could reach, what it did reach, what you have done, and what they should do.
#### Phase 5 — Post-incident: the HR dimension
Handle this part properly and you keep a good engineer. Handle it badly and you teach twelve people never to report anything again.
Three principles. Separate the access decision from the employment decision — suspending an identity is containment, it happens in minutes, and it is not a disciplinary finding; say that to the technician explicitly, in those words, at the moment you do it. Involve HR at Phase 2, not Phase 5, so the process is a process rather than an improvisation with an audience. And write facts, not characterizations — the ticket records times, actions and observations; it does not record blame, cause, or "we should have." Assume every line is read aloud in a deposition, because in a bad year it will be.
If the root cause was a social-engineering call, the finding is about your verification procedure, not the person who was called. I have never seen a help-desk verification procedure improve because someone was disciplined; I have seen several improve because someone felt safe describing exactly how the call went.
Actionable takeaway: take one technician who left in the last year and try to prove, from exports alone, that every delegated path they held is closed. Whatever you cannot prove is the size of your offboarding gap.
#PB-MSP-VENDOR — Your upstream vendor is compromised
Default severity: SEV-2 while the finding is vulnerable-only — you run an affected version and the Phase 1 hunt is clean. Blast radius: multi-tenant at that stage. Escalation: the moment exploitation here is confirmed, or cannot be ruled out, the blast radius is platform and the severity is SEV-1 — the rule from "Why This Book Exists" holds, and any credible suspicion of a platform compromise is SEV-1 until disproven — and this playbook hands over to
PB-MSP-PLATFORM. Runs: MSP Incident Commander; Practice Owner owns the disconnect decision.
Sooner or later a CVE drops in a product you have installed everywhere. Sometimes it is worse: on 28 May 2025 ConnectWise confirmed suspicious activity in its own ScreenConnect cloud environment, believed tied to a sophisticated nation-state actor (Qualys). "We moved to the vendor's cloud" removes your patching burden. It does not remove your blast radius. Your plan needs a branch for "our vendor was breached and we are downstream."
Entry criteria: a vendor security advisory or KEV addition for a product in your stack; credible in-the-wild exploitation reporting; a vendor disclosure of its own compromise; or a client or peer asking you about a headline before you had read it.
The wrong instinct. The reflex is to patch immediately and consider the matter closed. Three documented reasons that is not enough. Cleo's first fix did not fix it — 5.8.0.21 remained exploitable (BleepingComputer). N-able's Hotfix 1 addressed the original access point and a related attack path required Hotfix 2 four days later (N-able). And Rackspace applied Microsoft's mitigation rather than the patch, out of concern about operational disruption — and the attacker had a bypass for the mitigation (Dark Reading). Patch, then verify you are still patched. Not next sprint. This week.
#### Phase 1 — The first hour
| # | Action | Who | Done when | Evidence to capture |
|---|---|---|---|---|
| 1 | Answer three questions in this order: do we run it, at what exact version, and is the management interface reachable from the internet | On-Call Technician | All three answered from inventory, not memory | Version string, exposure scan result, UTC |
| 2 | Read the vendor advisory and check the CISA KEV catalog. KEV entries carry federal remediation deadlines that tell you how the government rates the urgency — CVE-2026-18577 was added on 3 August 2026 under BOD 26-04 with a deadline of 6 August, three days (CISA) | MSP Incident Commander | Advisory read in full, KEV status known | Advisory URL and retrieval time |
| 3 | Run the vendor's IOC list and any published detection script before you patch. Two questions, not one: does the vulnerability exist here, and was it already exploited here | Operations Lead | Hunt complete, results recorded | Query text, per-system result |
| 4 | Check the console's own logs for gaps and for unauthorized accounts, exactly as in PB-MSP-PLATFORM step 2 | MSP Incident Commander | Gap ruled in or out | Log inventory |
| 5 | Notify the carrier if a claim looks plausible, and confirm the panel before engaging any outside forensics firm — that is a consent event under most cyber and E&O policies | Practice Owner | Carrier notified, panel confirmed | Claim reference |
#### Phase 2 — Containment, Phase 3 — Eradication
Containment is the disconnect or the patch, plus MFA re-verification on every console account and disabling in-product vendor support accounts unless actively needed. Eradication is the re-check: read the advisory again 72 hours after you patched. Make that a calendar habit, not a good intention — an incomplete patch is a live zero-day, and "we patched on day one" was not a defense for anyone who patched N-central on 2 August 2026.
Then look sideways. Your backup console is the same class of target with the trust arrow reversed: CVE-2024-42448 in Veeam Service Provider Console (CVSS 9.9) allowed remote code execution on the VSPC server from an authorized management agent machine, with no mitigation available other than upgrading (Veeam KB4679). A single compromised client's backup agent is the attack surface for the console holding every client's backups. If a client is compromised and their agent is authorized on your shared console, that console is in scope until you prove otherwise.
#### Phase 4 — Recovery and client communication
Tell your clients before they read it in the news, and tell them even when nothing happened to them. Two authorities say so directly. AA22-131A expects MSPs to notify customers of confirmed or suspected security events on provider infrastructure — suspected, not confirmed (CISA). And NIS2 Art. 23(2) imposes a duty to communicate, without undue delay, to recipients potentially affected by a significant cyber threat — no incident required — telling them any measures or remedies they can take (Art. 23). An actively exploited, unpatched vulnerability in your RMM sits squarely inside that.
The message that works has four parts: what the vendor disclosed, whether you run it and in what configuration, what you have done and when, and what — if anything — the client needs to do. Send it per client at the contractual notice address, not as a broadcast to the day-to-day IT contact. An email to the helpdesk lead frequently is not contractual notice, and if the client later has to prove to a regulator when it became aware, your Teams message is a poor exhibit for both of you.
#### Phase 5 — Post-incident: the monitoring that makes this survivable
None of the above works if you find out from a client. The standing capability is unglamorous and cheap:
- A named owner for vendor advisory monitoring, with a named deputy, and an inbox that is read on weekends. Kaseya landed on a holiday Friday for a reason.
- A subscription to every vendor security bulletin in your stack — RMM, PSA, EDR, backup, MFT, firewall, documentation platform — plus the CISA KEV feed. Prioritize remediation by KEV, not by CVSS alone (AA22-131A).
- An asset inventory that answers "do we run it and at what version" in under five minutes. Include plugins, integrations and connectors — the 2019 GandCrab wave came through a plugin whose patch had been available since 2017 and which nobody owned (TechTarget).
- A designated security contact in Partner Center, with the CSP obligation to respond to security alerts within 24 hours or less.
- The 72-hour re-check, as a recurring task with an owner.
- A pre-drafted client advisory template you fill in rather than compose.
Every one of those is achievable by a twelve-person shop in an afternoon. The expensive version buys you a threat intelligence platform; the cheap version buys you a shared mailbox, a calendar reminder and a spreadsheet, and the cheap version has caught every incident in this section.
Actionable takeaway: pick the six items above, put a name against each, and set the first 72-hour re-check for the last thing you patched. If you cannot name the person who reads vendor advisories on a Saturday, that person is nobody, and nobody is on call this weekend.
Stay patched, stay segregated, and remember that in this business your worst night is a hundred other companies' worst night too.
#Sources
- CISA AA22-131A — Protecting Against Cyber Threats to MSPs and their Customers
- CISA AA25-163A — Ransomware Actors Exploit Unpatched SimpleHelp RMM
- FBI/CISA AA23-320A — Scattered Spider, updated 29 July 2025
- CISA et al. — Guide to Securing Remote Access Software
- CISA/FBI guidance for MSPs affected by the Kaseya VSA attack
- CISA — KEV addition, CVE-2026-18577, 3 August 2026 · CISA — I've Been Hit By Ransomware
- ODNI/NCSC — Kaseya VSA Supply Chain Ransomware Attack factsheet · Truesec — Kaseya technical analysis
- Huntress — ScreenConnect authentication bypass · Huntress — N-able N-central exploitation
- N-able — N-central security update, 10 August 2026 · Trend Micro — Black Basta and others exploiting ScreenConnect · Google Cloud/Mandiant — ScreenConnect remediation and hardening
- Sophos X-Ops — DragonForce actors target SimpleHelp to attack MSP customers · Sophos 2026 Active Adversary Report
- BleepingComputer — Clop claims Cleo data theft attacks · Dark Reading — Rackspace and the ProxyNotShell bypass · Qualys — CVE-2025-3935 added to KEV
- TechTarget — ConnectWise plugin flaw exploited in ransomware attacks on MSPs · Field Effect — Ingram Micro hit by SafePay ransomware · Veeam KB4679 — CVE-2024-42448
- Microsoft Security Blog — NOBELIUM targeting delegated administrative privileges · AADInternals — Microsoft partners
- Microsoft Learn: Microsoft-led DAP→GDAP transition · GDAP FAQ · GDAP least-privileged roles by task · Partner security requirements · Cross-tenant access settings · Manage emergency access admin accounts · Monitor governing tenant admin activity · Monitor service provider activity (Azure Lighthouse) · Review audit logs in Microsoft 365 Lighthouse · Advanced hunting in Defender multitenant management · Extend Microsoft Sentinel across workspaces and tenants · Overview of indicators in Defender for Endpoint
- EDPB Guidelines 9/2022 on personal data breach notification · NIS2 Article 23 · Commission Implementing Regulation (EU) 2024/2690
- Clifford Chance — ICO fines processor Advanced Computer Software Group
- NCSC-UK — Using MSPs to administer your cloud services
#Part 4 — The Notification Chain
#The Notification Chain
#The chain, drawn
You are the smoke detector. You are almost never the fire marshal.
In nearly every incident an MSP touches, the entity with the legal duty to tell a regulator is the client, not you. The entity holding the facts that make the duty assessable is you, not the client. And the entity operating under the shortest, hardest deadline in the whole sequence is — again — you, because your deadline came from a contract you signed rather than a statute somebody else has to comply with.
That is the whole problem in three sentences, and it explains why so many otherwise well-run MSPs have a bad time in the notification phase of an incident they handled competently everywhere else. The technical work was fine. The chain was where it went wrong.
A note before the templates. What follows summarises statute, regulator guidance and published contracts, and it is not legal advice — I am a CTO with a research pack, not your lawyer. The four templates are operational shapes, not drafted legal notices: the wording, and every clause reference inside them, belongs to your own counsel in your own jurisdiction, settled on a quiet afternoon rather than on the night. At 03:00 you want to be filling in blanks, not writing law.
Here is the chain, in the order it actually runs:
T+0 You discover. (Alert, ticket, vendor advisory, a client's phone call,
or a leak-site listing you did not want to read.)
T+? You classify: whose systems, whose data, which contracts, how many tenants.
T+? YOU NOTIFY EACH AFFECTED CLIENT.
Clock: CONTRACTUAL. Yours. Runs from YOUR discovery.
Often "immediately", "within four hours", "without undue delay".
T+? The client becomes AWARE.
For GDPR purposes the controller is, in principle, aware once you
have told them (EDPB Guidelines 9/2022, para 44).
T+? The client assesses risk / materiality. Their lawyers, their DPO,
their board, their insurer.
T+? The client notifies its regulator, its customers, its insurer.
Clocks: NIS2 24h early warning; GDPR 72h; NYDFS 72h; SEC 4 business
days from the materiality determination; HIPAA 60 days; FTC
Safeguards 30 days; state laws mostly 30 days.
Every clock on the bottom row belongs to the client, and each has its own trigger: NIS2's 24-hour early warning and 72-hour notification both run from the entity becoming aware of a significant incident (NIS2 Art. 23); NYDFS §500.17(a)'s 72 hours runs from determining that a cybersecurity incident occurred at the covered entity, its affiliates or a third-party service provider (23 NYCRR 500.17); the SEC's four business days runs from the materiality determination, not from discovery (SEC press release 2023-139); the FTC Safeguards Rule's 30 days runs from the financial institution's discovery of a notification event involving at least 500 consumers (16 CFR 314.4(j)). None of them is your deadline — unless you are yourself in scope, which an EU-established MSP under NIS2 is; see "Your own notifications" below. All of them are shortened by your delay.
Two facts govern everything downstream. Your clock starts before theirs and is shorter than theirs. And you owe it N times in parallel, with different content each time — the EDPB is explicit that where a processor serves multiple controllers affected by the same incident, "the processor will have to report details of the incident to each controller" (EDPB Guidelines 9/2022 v2.0, para 47). On a twelve-person MSP with one technician on call, that fan-out is the part that breaks, and it breaks because nobody pre-drafted anything.
Takeaway: the notification chain is a delivery problem before it is a legal problem. Build the delivery capability — per-client deadline, per-client recipient, per-client template — and the legal analysis becomes something your client's counsel does on their own time instead of something you improvise on yours.
#The clock that actually binds you
Ask an MSP owner what their notification deadline is and you will usually get a regulatory answer: 72 hours, GDPR. It is a good answer to a question nobody asked them. The 72 hours is the controller's, it runs from the controller's awareness, and it has nothing to do with you except that you are the reason it starts.
What binds you is your Master Services Agreement, your Data Processing Agreement and your Business Associate Agreement — shorter numbers, not uniform across your book, with commercial consequences that arrive whether or not a regulator ever hears about the incident. Miss a four-hour contractual notice and you have breached the contract. No risk threshold, no "no personal data was involved" defense, no supervisory authority exercising discretion. A clause and a timestamp.
There is a second, quieter version of this trap: many real MSP agreements contain no notification clock at all. The CompassMSP master agreement — a current, publicly posted example — has no incident response clause and no breach-notification-to-client timeline anywhere in it (CompassMSP MSA). Silence does not help you. It leaves you standing on whatever your DPA says, whatever the statute says, and whatever a court later decides "without undue delay" meant on the night in question.
The operational consequence is that the on-call technician needs one number, per client, visible in ten seconds. Not a signed PDF in a shared drive. A structured field in the PSA, next to the credentials and the after-hours phone list, that says: notify within N hours, by this method, to this address, to this named human. That is both the cheap version and the right version — one afternoon of contract reading and one custom field. There is no platform to buy.
#Article 33(2), and why "let me confirm first" spends the client's money
Every competent technician's instinct is to be sure before they worry anybody. It is a good instinct in almost every other part of this job. Here it is expensive, and the expense lands on someone else's balance sheet.
The EDPB settled the point, and the wording is worth reading twice: the processor "does not need to first assess the likelihood of risk arising from a breach before notifying the controller; it is the controller that must make this assessment on becoming aware of the breach. The processor just needs to establish whether a breach has occurred and then notify the controller." And then the sentence that defines the whole chain: "the controller should be considered as 'aware' once the processor has informed it of the breach" (EDPB Guidelines 9/2022 v2.0, para 44).
Translation. The 72 hours is a fuel tank that only starts emptying when you speak, and every hour you spend confirming or tidying is siphoned out of it before the client has driven a metre. Notify at hour nine because you wanted certainty and the client's counsel, DPO, forensics scoping, board briefing and regulator submission all have to happen in 63 instead of 72. They will not thank you for the certainty.
The guidance also anticipates how you are described in that submission: a controller "may find it useful to name its processor if it is at the root cause of a breach, particularly if this has led to an incident affecting the personal data records of many other controllers that use the same processor" (para 54). Assume the regulator learns your name on day one.
#What a notification must contain
A useful notification does five things and nothing else: says what is known, says what is not known, says what you have already done, says what you need the client to decide, and says when the next update arrives. NCSC's standard for incident communications is that they be "clear, consistent, authoritative, accessible and timely" (NCSC comms guidance) — worth pinning above the desk of whoever writes yours.
| # | Element | Why it is there | Failure if omitted |
|---|---|---|---|
| 1 | Client name, tenant/site identifier, notification reference | Proves a per-client notice, not a broadcast | Fails EDPB para 47; the client cannot file it |
| 2 | Discovery timestamp, UTC / ISO 8601 | Starts and evidences both clocks | Neither party can prove when awareness arose |
| 3 | What is known — observed facts only | The client's risk assessment runs on these | Client escalates back to you for basics |
| 4 | What is not yet known, listed explicitly | Prevents false reassurance and later retraction | Silence reads as "nothing else happened" |
| 5 | Actions already taken, with times | Evidences containment and good faith | Reads as inaction; weakens your own defense |
| 6 | Actions requested, each with a decision owner | Converts a warning into work | Client waits for you; you wait for the client |
| 7 | Named contact, out-of-hours number, bridge details | The client must reach a human | Escalation dies in a shared mailbox |
| 8 | Time of the next scheduled update | Stops the hourly "any news?" calls | Your Communications Lead spends the night on the phone |
| 9 | Statement that facts may change | Preserves your ability to correct | Every correction reads as a retraction |
Send it to the contractual notice address by the contractual method, and separately to the operational contact. A Teams message to the client's helpdesk lead frequently is not contractual notice, and it is a poor exhibit when the client later has to prove to a regulator exactly when it became aware. Keep the delivery evidence.
#### Template 1 — MSP-to-client initial notification
Variables in <ANGLE BRACKETS>. Plain text. No attachments on the first message. Illustrative shape only — have your own counsel settle the wording and the clause references before you ever need to send it.
Subject: Security incident notification — <CLIENT LEGAL ENTITY NAME> — ref <INCIDENT REF>
To: <CONTRACTUAL NOTICE ADDRESS>
Cc: <NAMED CLIENT AUTHORITY HOLDER>, <CLIENT SECURITY/IT CONTACT>
From: <MSP INCIDENT COMMANDER NAME AND ROLE>
This is a formal notification under <MSA CLAUSE REF / DPA CLAUSE REF / BAA
SECTION REF>.
1. WHAT WE KNOW
At <UTC TIMESTAMP, ISO 8601> we identified <FACTUAL DESCRIPTION OF THE
OBSERVED EVENT — e.g. "authentication to the <SYSTEM> administrative
console from an IP address not associated with any of our technicians">
affecting <NAMED SYSTEMS / TENANT / SITE>.
We discovered this at <UTC TIMESTAMP> via <DETECTION SOURCE>.
2. WHAT WE DO NOT YET KNOW
- Whether any data was accessed, copied or removed.
- The full list of affected systems and accounts.
- The root cause.
- Whether other clients of ours are affected.
We are investigating each of these and will report on them in the updates
below.
3. WHAT WE HAVE DONE
<UTC TIME> — <ACTION, e.g. "isolated <HOST> via endpoint protection">
<UTC TIME> — <ACTION, e.g. "revoked active sessions for <ACCOUNT>">
<UTC TIME> — <ACTION>
All actions taken so far fall within <PRE-AUTHORISED ACTIONS SCHEDULE REF>.
4. WHAT WE NEED FROM YOU
- <DECISION>, decided by <UTC TIME>, by <ROLE AT CLIENT>.
Options and consequences: <A> / <B>.
- Confirm the individual authorised to approve disruptive action tonight,
and a number that will be answered.
- Advise whether you require us to <SPECIFIC OPTIONAL ACTION>.
We will not take action outside the pre-authorised schedule without your
approval.
5. YOUR OWN OBLIGATIONS
You may have notification obligations of your own arising from this event.
We are not able to advise you on them. We are providing this notification
promptly so that you can take that advice with as much time as possible.
6. NEXT UPDATE
Our next written update will be sent by <UTC TIME>, and every
<INTERVAL> thereafter, whether or not there is new information.
Incident bridge: <DIAL-IN / LINK>
Direct contact: <NAME>, <ROLE>, <MOBILE>, available <HOURS>.
Facts in this notification are preliminary and may change as the
investigation continues.
Takeaway: the client's first question is never "what happened?" It is "what do I have to do, and by when?" Block 4 of the template above — WHAT WE NEED FROM YOU — is the part that earns the relationship. Write it first and let the rest hang off it.
#What never goes in it
Write every notification as though it will be read aloud, slowly, by opposing counsel, to a room that includes your client's insurer. In the incidents that go badly, it is. The SEC's action against SolarWinds and its CISO was built substantially on internal presentations, emails and instant messages (SEC press release 2023-227); the Clorox complaint against Cognizant put the vendor's own service-desk transcripts at the centre of the case (The Register). Your ticket notes are the same class of artifact.
The rules, drawn from NCSC's incident communications guidance (NCSC):
- No speculation about cause. "Avoid speculation or premature conclusions about the cause or extent of the incident, or who is behind it."
- Nothing you may have to retract. NCSC's own example is stating there is no known impact on personal data — a sentence that becomes a problem the moment your understanding changes. Say "we have not yet established whether…" instead.
- No hyperbole, in either direction. "Catastrophic" and "minor" are both characterizations, and both get quoted back.
- No attribution. Not the group name, not the country, not "this looks like the same crew that hit…".
- No unverified counts. Records, users, mailboxes — none of them until you can evidence them. A wrong number in an early notification propagates into a regulator submission and then into a correction.
- No admissions and no blame. Not "we should have patched that", not "this was the client's decision to defer MFA", not "our vendor let us down". Each may be true. None belongs in a notification.
- No characterization of legal exposure, ever, in any channel.
The discipline is the one the main playbook sets out for war-room hygiene: facts and timestamps go in; opinions, blame and speculation go nowhere. Distinguish "observed" from "assessed", and label which is which.
And the privilege point, because MSPs get this wrong in a specific way: nothing you write is privileged by default. You are not the client's lawyer, and your tickets, alerts, RMM logs and incident reports are business records created for a business purpose. The work-product route under FRCP 26(b)(3)(A) reaches material prepared in anticipation of litigation by a party's "consultant… or agent" (FRCP 26), but the Capital One, Clark Hill and Rutter's line narrowed it hard, and the mixed-role problem is usually fatal here — you are the remediator, which makes you a fact witness. Do not label ordinary operational work "privileged"; over-labeling invites a challenge that succeeds and taints whatever was genuinely protected. And when the client's counsel starts assessing whether you were at fault, your interests have diverged: get your own counsel then. Part 1's third scenario has the full treatment.
#The scheduled update, and the discipline of saying nothing
Set an update interval in the first notification, and then hit it. Every time. Including — especially — the updates where nothing has changed.
An update that says "no new findings since 03:00; the exports are still running; next update 09:00" costs ninety seconds and does three jobs. It proves the incident is being actively worked. It gives the client's own comms team something to say internally. And it builds a contemporaneous record of your attentiveness, which matters if anyone later argues you went quiet. The silence between hour four and hour eleven is where client trust dies, and it is also the gap a plaintiff's timeline slide is built around. Sophos's published MDR terms set a sixty-minute expectation on the customer to acknowledge outreach (Sophos MDR service description); hold yourself to a comparable standard in the other direction and put the number in the contract.
#### Template 2 — the scheduled update
Subject: UPDATE <N> — Security incident <INCIDENT REF> — <CLIENT NAME>
Time of this update: <UTC TIMESTAMP>
Status: <ACTIVE — CONTAINMENT / ACTIVE — INVESTIGATION / MONITORING /
CLOSED PENDING REPORT>
NEW SINCE LAST UPDATE
<FACTS ONLY, EACH WITH A UTC TIME. If nothing: "No new findings.">
CHANGED FROM PREVIOUS UPDATE
<ANY EARLIER STATEMENT NOW KNOWN TO BE INCOMPLETE OR INCORRECT, STATED
PLAINLY. If nothing: "Nothing previously reported has changed.">
STILL UNKNOWN
<CARRY FORWARD THE OPEN LIST; REMOVE ITEMS ONLY WHEN ANSWERED>
ACTIONS IN PROGRESS
<ACTION> — owner <NAME/ROLE> — expected <UTC TIME>
OPEN DECISIONS FOR YOU
<DECISION> — needed by <UTC TIME> — owner <ROLE AT CLIENT>
<If none: "No decisions are currently required from you.">
NEXT UPDATE: <UTC TIMESTAMP>
The "changed from previous update" block is the one people delete because it is uncomfortable. Keep it. A correction volunteered in the next scheduled update is a demonstration of rigour; the same correction discovered later by somebody else is a credibility event.
#Notifying forty clients at once
This is the section for the night the incident is yours: your RMM, your PSA, your partner tenant, your technician's identity. Kaseya in July 2021 ran from fewer than 60 direct MSP customers to as many as 1,500 downstream businesses (NCSC/ODNI factsheet). CTS, a UK legal-sector MSP, went down in November 2023 and took an estimated 80 to 200 law firms with it — no confirmed data theft at all, just a professional sector unable to trade (The Record). CISA now writes MSP advisories with "contact downstream customers" as an explicit provider action (CISA AA25-163A).
Five rules, and the order matters.
1. Tranche by contractual deadline, not by client size. Sort the notification list by shortest contractual clock first, then by regulatory exposure (public company, DIB flow-down, NYDFS, healthcare), then by everything else. The wrong order is by relationship value, and it fails specifically: your four-hour client sits at position nineteen because your biggest client sat at position one, and the record shows you chose. Sorting by clock is defensible. Sorting by revenue is a document you do not want produced.
2. Every notification carries that client's facts. One broadcast email fails EDPB para 47, fails the client's own reporting content needs, and lands identical wording in forty legal inboxes — a convenient way to help a coordinated claim form. Same template; change the tenant, the systems, the observed activity and the requested actions.
3. Never disclose one client's identity or status to another. Not the count, not the sector, not "we're seeing this at a couple of other clients too". A client asking "who else is affected?" gets: "I can confirm what we have observed in your environment. I can't discuss other clients' environments, and I'd give you the same answer if you were the one being asked about." Most clients understand that immediately, because they are on the other end of the same promise.
4. Assume clients talk to each other. Same trade bodies, same industry groups, quite often the same building. Two clients comparing your notifications is normal, and it is fine when the notifications are consistent in structure, timing and tone. It is not fine when one got a phone call at 02:00 and the other got an email at 11:00 with a different characterization of the same event. Consistency is not a nicety — it is the control that stops your notifications becoming evidence against each other.
5. Assume your own systems are untrustworthy. If the incident is in your PSA or RMM, that is both where your client contact list lives and where the notification would be sent from. Keep an offline copy of the list and a communications path that does not traverse the compromised platform — AA22-131A's blunt instruction is to maintain hard-copy plans accessible when networks are compromised (CISA AA22-131A).
#### Template 3 — the multi-client holding statement
For inbound questions from clients, press or partners while the per-client notifications are still going out. CISA says to rehearse this rather than improvise it: "If a reporter calls you, claiming to have data stolen from your file servers, what will you say? Having a good 'holding statement' will help" (CISA IRP Basics).
<MSP NAME> is responding to a security incident affecting <SYSTEM CLASS,
e.g. "one of our internal management platforms">, identified on
<DATE>. We began our response immediately and <FACTUAL CONTAINMENT
STATEMENT, e.g. "took the affected platform offline the same day">.
We are contacting affected clients directly with information specific to
their environments, and we are providing regular written updates to them.
We are not commenting on individual clients.
<IF ENGAGED:> We are working with external incident response specialists
and have notified <RELEVANT AUTHORITY, ONLY IF TRUE AND DECIDED>.
Our investigation is ongoing. We will not speculate about the cause or the
extent of the incident while that work continues, and we will correct
anything we report if our understanding changes.
Media enquiries: <NAME>, <EMAIL>, <PHONE>.
Clients: please contact <DEDICATED INCIDENT LINE / ADDRESS>.
Note what is not in it: no cause, no attribution, no reassurance about data, no client names, no numbers. Note what is: a dedicated client line. Notify forty clients simultaneously and your normal service desk will fall over, and the clients who cannot get through will assume the worst and tell each other so.
#### Template 4 — "this is our incident, not yours, but you are affected"
The hardest letter an MSP owner writes. Write it now, in daylight, so the version you send is not the version you drafted at 04:00 while your own business was on fire.
Subject: Security incident at <MSP NAME> affecting services provided to
<CLIENT LEGAL ENTITY NAME> — ref <INCIDENT REF>
To: <CONTRACTUAL NOTICE ADDRESS>
Cc: <NAMED CLIENT AUTHORITY HOLDER>
From: <PRACTICE OWNER NAME AND TITLE>
This is a formal notification under <MSA CLAUSE REF / DPA CLAUSE REF /
BAA SECTION REF>.
1. WHAT HAPPENED, AND WHERE
On <UTC TIMESTAMP> we identified a security incident affecting
<MSP SYSTEM — e.g. "the remote management platform we use to administer
client environments">. This is our system, not yours. Because that system
holds <ACCESS TO / DATA FROM> your environment, this incident affects you.
2. HOW IT AFFECTS YOUR ENVIRONMENT
<SPECIFIC TO THIS CLIENT: which of their systems the platform could reach,
which credentials it held, what activity has and has not been observed in
their tenant.>
3. WHAT WE HAVE OBSERVED IN YOUR ENVIRONMENT
Observed: <FACTS, WITH UTC TIMES>
Not observed to date: <SPECIFIC ITEMS CHECKED AND THE PERIOD COVERED>
Not yet checked: <SPECIFIC ITEMS AND WHEN WE EXPECT TO HAVE CHECKED THEM>
4. WHAT WE HAVE DONE
<UTC TIME> — <ACTION ON THE MSP PLATFORM>
<UTC TIME> — <ACTION IN THE CLIENT ENVIRONMENT, WITH THE AUTHORITY RELIED ON>
We have preserved the relevant logs and are holding them; we will not
delete anything relating to this period until you tell us to or a hold is
released.
5. WHAT WE ARE ASKING YOU TO DO
- <ACTION>, by <UTC TIME>, owner <ROLE AT CLIENT>.
- Consider whether you wish to engage your own incident response,
legal or insurance advisers. We will cooperate fully with anyone you
appoint and will provide evidence in the formats they ask for.
- Advise us of any notification obligations you believe apply to you, and
the information you need from us to meet them. We will prioritise those
requests.
6. WHAT WE ARE NOT SAYING
We are not yet able to tell you the cause. We are not able to tell you
whether data was removed. We will tell you as soon as we can evidence an
answer, and we will tell you if an earlier statement turns out to be wrong.
7. UPDATES
Written updates at <INTERVAL>, next at <UTC TIME>, whether or not there is
new information. A dedicated line for your team: <NUMBER>.
<NAME>
<TITLE>, <MSP LEGAL ENTITY NAME>
Two notes. First, no apology in section 1 and no promise in section 6 — an apology reads as an admission of fault to your insurer and to a court, and a promise you cannot keep gets quoted at you. Sympathy belongs on the phone call you make immediately after sending it, and you should make that call. Second, section 4's last sentence is doing real work: it tells the client you are preserving evidence rather than restoring over it, which is the first question their counsel will ask.
#Your own notifications: carriers first, and the notice trap
Your cyber policy and your Technology E&O policy do separate jobs — E&O is third-party cover for claims your clients bring against you; cyber is largely first-party cover for your own incident costs. Neither pays anything if you tell them late.
The published AXIS cyber specimen makes notice "a condition precedent to coverage": the insured must notify in writing as soon as practicable after discovery, and no later than 60 days after discovery or the end of the policy period, whichever is later. The same form allows incident-response expense without prior written consent for up to 72 hours from discovery, and only to retain a provider from the insurer's pre-approved panel (AXIS 1012561 0120 specimen.pdf)).
Read that as an operating instruction. Containment is not a "cost incurred" and never waits for a carrier. But the moment you are about to retain outside counsel, engage a DFIR firm or open any kind of negotiation — for yourself or on a client's behalf — that is a consent event, and reaching for your own preferred firm instead of their panel firm may be an uncovered cost you have just created for your client.
Your own regulator is a separate question from your clients' regulators, and it is genuinely yours. An MSP with an EU main establishment is named directly in NIS2 and owes a 24-hour early warning, a 72-hour notification and a one-month final report, each measured from your awareness of a significant incident, plus a duty to tell the recipients of your services about significant incidents and even significant cyber threats affecting them (NIS2 Art. 23). An MSP holding a DFARS 252.204-7012 flow-down reports to DoD within 72 hours of discovery, preserves images and monitoring data for at least 90 days, and gives the prime its incident report number (DFARS 252.204-7012). Part 5 covers all of it properly. The point here is narrower: notifying your clients does not discharge your own duties, and your own duties run on your own discovery.
#Law enforcement, and whose call it is
For an incident in a client's environment, the referral decision belongs to the client. It is their loss, their data, their relationship with an investigator who may ask them to delay other notifications, and their counsel's judgment about how that engagement interacts with everything else they owe. You advise; you do not dial.
What you owe instead is preparation: know the routes, offer them in the notification's "actions requested" section, and be ready to preserve what an investigator will want. In the United States the published routes are the FBI's IC3, a local FBI field office, or CISA at report@cisa.gov or 1-844-Say-CISA (CISA AA25-163A).
Two cautions. First, evidence before eradication: CISA's federal playbook direction is to coordinate with ICT service providers, vendors and law enforcement prior to initiating eradication, and a DFARS-scoped client carries a 90-day preservation duty that your instinct to reimage will destroy in ten minutes. Do not reimage before imaging. Not after the client has been reassured. Not once the ticket is closed. Before you touch it.
Second, a law-enforcement request to delay notification does not pause every other clock. Delay allowances differ by regime, and a client with securities obligations can find itself required to disclose publicly while an investigator asks it to stay quiet. That is a decision for the client's counsel, made with the facts you supplied — supply them early enough for the conversation to happen at all.
When the incident is yours, the referral decision is yours, and it belongs to the Practice Owner with counsel — not to whoever is awake.
Actionable takeaway: pick your three largest clients this week, open their agreements, and write the notification deadline, the notice method and the notice address into the PSA as structured fields. Then do the next three. Do not schedule a project. Do six contracts a week until the field is populated for your whole book, and put the same field on the onboarding checklist so it never falls behind again.
Stay patched, stay paranoid, and write the letter before you need it — the version you draft on a Tuesday afternoon is always better than the one you write at four in the morning.
#Checklist
- MSP-31Every client record carries the contractual notification deadline in hours, the contractual notice method, the legal notice address and a named recipient with an out-of-hours number, readable by the on-call technician in one screen without opening a contract.
[IG1][CIS 15][A.5.19] - MSP-32Every client record flags the regulatory regimes attaching to that client (GDPR controller, HIPAA covered entity, PCI, NYDFS, DFARS flow-down, public company, NIS2, DORA) and the categories of data held on their behalf.
[IG1][GV.SC] - MSP-33A written standing decision states that client notification is not delayed pending risk assessment, root-cause determination or confirmation of data impact, and names the role authorized to send.
[IG1][RS.CO] - MSP-34Four pre-approved notification templates exist and are reviewed at least annually: initial client notification, scheduled update, multi-client holding statement, and MSP-is-the-source notification.
[IG1][A.5.24] - MSP-35Notifications are issued to the contractual notice address by the contractual method, and delivery evidence (timestamp, recipient, transmission record) is retained with the incident file.
[IG1][A.5.28] - MSP-36A contemporaneous log records, in UTC/ISO 8601, the discovery time, every notification attempt, every contact reached or missed, and who authorized each action — maintained on a system unaffected by the incident.
[IG1][RS.MA][A.5.28] - MSP-37Every notification states the time of the next update, and updates are sent on schedule even when there is nothing new; a missed update is itself recorded as an incident-process exception.
[IG1][RS.CO] - MSP-38A written content standard prohibits cause speculation, attribution, fault admission, blame characterization, legal-exposure commentary and unverified record counts in any client-facing communication, and a second named person reviews every notification before it is sent.
[IG2][RS.CO] - MSP-39Every client file has been reviewed against a checklist confirming whether the MSA, DPA and BAA specify notification timing, method and notice address in both directions; each gap has a named owner (Practice Owner) and a remediation date tied to renewal.
[IG1][CIS 15][A.5.20] - MSP-40A multi-client notification runs from a tranche plan ordered by contractual deadline and then by regulatory exposure, never by account value, and the ordering rationale is recorded.
[IG2] - MSP-41Each affected client receives a notification containing that client's own facts; no single broadcast communication is relied on to discharge a per-client notification duty, and no notification discloses another client's identity, sector or status.
[IG2][RS.CO] - MSP-42An offline copy of the client notification contact list, and a communications path that does not traverse the RMM or PSA, are maintained and verified at least quarterly.
[IG1][A.5.30] - MSP-43The MSP's own cyber and Technology E&O carriers, claims lines, panel breach counsel and panel DFIR firms are recorded in the incident plan, and carrier notice sits on the incident checklist at the same tier as client notification.
[IG1][RS.CO] - MSP-44A standing rule prohibits engaging outside counsel, forensic firms or negotiators — for the MSP or on a client's behalf — before carrier notice and panel confirmation, with the sole exception of containment actions, which never wait.
[IG2] - MSP-45The notification chain is rehearsed at least annually against the largest realistic fan-out, measuring elapsed time from simulated discovery to the last client notification sent and to the first scheduled update, with the result reported to the Practice Owner.
[IG3][CIS 17][A.5.24]
#### Sources
- EDPB Guidelines 9/2022 on personal data breach notification, v2.0
- GDPR Article 33 · GDPR Article 28
- 45 CFR 164.410 · 45 CFR 164.314
- 16 CFR 314.4 (FTC Safeguards Rule)
- 23 NYCRR 500.17
- New York GBL §899-aa
- Privacy Rights Clearinghouse — Data Breach Notification Laws 50-State Survey, 2026
- NIS2 Directive Article 23
- gov.uk — Cyber Security and Resilience Bill: incident reporting factsheet
- DFARS 252.204-7012
- SEC press release 2023-139 — cybersecurity disclosure rules
- CISA AA22-131A — Protecting Against Cyber Threats to MSPs and their Customers
- CISA AA25-163A — SimpleHelp/ransomware advisory
- CISA — Incident Response Plan Basics
- NCSC-UK — Guidance on effective communications in a cyber incident
- NCSC-UK — Choosing a Managed Service Provider
- CompassMSP Master Service Agreement
- Sophos MDR Service Description
- AXIS Cyber Insurance Policy specimen (AXIS 1012561 0120).pdf)
- · Insurance Journal
- Hunton — ACE American v. Congruity 360 and Trustwave
- SEC press release 2023-227 — SolarWinds
- The Register — Clorox v. Cognizant
- FRCP 26
- NCSC/ODNI — Kaseya VSA supply chain ransomware attack factsheet
- The Record — cyberattack on MSP CTS impacts UK law firms
#Part 5 — When the MSP Is the Regulated Entity
#When the MSP Is the Regulated Entity
For fifteen years the compliance conversation at an MSP went like this: the client has the regulator, we have the tooling, our job is to help them pass. That arrangement is over. Somewhere between the EU naming "providers of managed services" in a directive book and the UK inventing a statutory category for us, the channel stopped being the plumbing behind regulated entities and became one.
This is the section the owner reads, and it is where I have deliberately turned the voice down. Where I could not confirm something, I say so rather than smooth it over — an owner making a certification decision on a half-remembered conference talk is exactly the failure this section exists to prevent.
#The three hats, worn at once
You are not in one regulatory relationship. You are in three, and one incident engages all three before breakfast.
| Hat | What it means | What puts you there |
|---|---|---|
| Regulated entity in your own right | The law names you. You register, you secure, you report on your own clock. | EU NIS2 (Annex I, sector 9); UK CSR Bill (pending); DORA only if designated a critical ICT third-party provider; CIRCIA if the final rule keeps the proposed criteria; CMMC where you hold a DoD contract or flow-down |
| Processor / service provider | Usually no direct regulator — HIPAA is the exception, where OCR can enforce against a business associate — but the law dictates what must be in your client's contract with you, and you feed their clock. | GDPR Art. 28 and 33(2); DORA Art. 30; FTC Safeguards §314.4(f); HIPAA business associate; NYDFS §500.11; PCI DSS 12.9 |
| Supply-chain risk your client must manage | Your incident is their reportable incident, and their regulator inspects how they diligenced you. | NIS2 Art. 21(2)(d); DORA Art. 28; NYDFS §500.11; FTC §314.4(f); SEC Item 1.05 third-party systems |
The hats do not take turns. Ransomware in your RMM makes you a regulated entity on a 24-hour clock, a processor owing individual notifications to forty controllers, and the subject of forty supply-chain post-mortems — all from the same minute, and all before anyone has confirmed whether data left the building. Actionable takeaway: put the three hats on one page, name which clients put you in which, and hand it to whoever answers the phone at 23:45. Nobody classifies well while a hypervisor is encrypting.
#EU NIS2 — this book has your name in it
MSPs and MSSPs are not swept in by implication. They are listed: Annex I, sector 9, "ICT service management (business-to-business)", as providers of managed services and of managed security services (scope guidance). Art. 6(39) covers services "related to the installation, management, operation or maintenance of ICT products, networks, infrastructure, applications or any other network and information systems… either on customers' premises or remotely" (Art. 6). It does not require you to hold client data, and it says "or remotely" out loud. A break/fix shop with a persistent RMM agent on a business customer's servers is inside that definition whether or not it has ever called itself an MSP.
Scope attaches at medium-sized enterprise under Art. 2(1), by reference to Recommendation 2003/361/EC (Art. 2; Recommendation).
| Size | Status |
|---|---|
| Under 50 staff and turnover/balance sheet at or under €10m | Out of scope by size, subject to the Art. 2(2) exceptions |
| 50 staff or more, or turnover and balance sheet both above €10m, to the medium ceiling | Important entity (Art. 3(2)) |
| 250 staff or more, or turnover above €50m and balance sheet above €43m | Essential entity (Art. 3(1)(a)) |
Most MSPs land as important entities (Art. 3) — ex-post supervision rather than routine inspection, and a lower fine floor. Being small is not a guarantee: Art. 2(2) applies the Directive regardless of size where the entity is the sole provider in a Member State of an essential service, where disruption "could have a significant impact on public safety, public security or public health," or where the entity is critical because of its specific importance at national or regional level. A twenty-five person MSP running IT for four regional hospitals is a live candidate. Ask your national authority in writing and file the answer; that email outranks a consultant's opinion when someone later asks why you never registered.
Which brings us to the duty most MSPs have already missed. Art. 27 requires managed service providers and managed security service providers, named explicitly, to give their competent authority entity name, sector, establishment addresses, current contacts, the Member States served, and their IP ranges; ENISA maintains the registry and the Directive's date was 17 January 2025 (Art. 27). It is a standing pre-incident obligation, and the one most commonly discovered during an incident. Jurisdiction at least is merciful: Art. 26 puts you under the Member State of your main establishment, and a non-EU MSP serving EU customers must designate a Union representative — without one, "any Member State in which the entity provides services may take legal actions against the entity" (Art. 26).
Two customer-facing duties sit in the same article and get missed because people stop reading at the deadlines. Art. 23(1) requires you to notify the recipients of your services, without undue delay, of significant incidents likely to adversely affect provision of those services. Art. 23(2) requires you to tell recipients potentially affected by a significant cyber threat — no incident needed — what measures they can take. An actively exploited, unpatched vulnerability in the RMM you run for two hundred companies is squarely inside 23(2). You owe them a communication, not a status page.
What "significant" means for an MSP specifically is set by Commission Implementing Regulation (EU) 2024/2690, which applies directly to MSPs and MSSPs and which almost nobody in the channel has read (EUR-Lex). Article 10 triggers include the service completely unavailable for more than 30 minutes; availability limited for more than 5% of Union users or more than one million, whichever is smaller, for over an hour; and data integrity or confidentiality compromised through suspected malicious action. Article 3(1) stacks general criteria on top, and criterion (e) is the one to read twice: "a successful, suspectedly malicious and unauthorized access to network and information systems occurred, which is capable of causing severe operational disruption." For an MSP, confirmed unauthorized access to the RMM, PSA or privileged-access tier is capable of that by definition. The 24-hour clock can start on a compromise that has produced no customer impact at all — and thirty minutes of total unavailability is a bar most of us clear a few times a year without formally classifying it. The failure here will be in classification, not detection.
Art. 20(1) requires management bodies to approve the risk-management measures and oversee implementation, and says they "can be held liable for infringements"; Art. 20(2) requires them to follow training (Art. 20). Where enforcement escalates, Art. 32(5) lets authorities suspend a certification covering part or all of the entity's services and ask a court to temporarily bar a person discharging managerial responsibilities at chief executive or legal representative level — powers that apply to essential entities (Art. 32). Fine floors under Art. 34: essential, at least €10m or 2% of global turnover, whichever is higher; important, at least €7m or 1.4% (Art. 34).
And the other direction — why your inbox is full of questionnaires. Art. 21(2)(d) requires in-scope entities to hold measures covering supply chain security, including the relationship with each direct supplier; Art. 21(3) requires them to weigh vulnerabilities specific to each supplier and the overall quality of that supplier's practices and secure development procedures (Art. 21). The recitals flag managed security service providers as warranting increased diligence given their operational integration and prior targeting. Every NIS2-regulated client must therefore diligence you and evidence that they did. The questionnaire is their compliance artifact, and your contract is what their regulator reads.
Transposition, honestly: due 17 October 2024 and still incomplete. The Commission opened infringements against 23 Member States in November 2024 and issued reasoned opinions to 19 in May 2025 (EC); on 8 July 2026 it referred Ireland, Spain, France and the Netherlands to the CJEU (ECSO tracker). Actionable takeaway: do not encode "NIS2" as one obligation. Encode a per-country matrix — threshold, registration portal, reporting portal, fine ceiling — for every Member State you serve. Where a state has not transposed, the Directive is not directly effective against you as a private entity, but the old NIS1 national law may still bite.
#UK Cyber Security and Resilience Bill — pending, and pointed at you
Status as of September 2026: in Parliament, not law. It cleared all Commons stages and had its Lords Second Reading on 14 July 2026 (Parliament). The Government's summary states the MSP measures arrive via secondary legislation, after an implementation consultation planned for 2026 (gov.uk). Nothing in it is enforceable against you today.
It creates a statutory category: Relevant Managed Service Providers — "a person who provides managed services in the UK (whether or not the person is established in the UK)" that is not a small or micro enterprise and not subject to public authority oversight (RMSP factsheet). The stated reason is unusually direct: MSPs "have unprecedented access to their customers' systems, making them an attractive target that cyber actors increasingly exploit." Duties: register with the regulator; implement appropriate and proportionate measures for the systems your managed services rely on; prevent and minimize incident impacts; notify significant incidents. The regulator is the Information Commission, successor to the ICO — so your cyber regulator will also be your data protection regulator.
Reporting shape (incident reporting factsheet): a light-touch notification within 24 hours of becoming aware an incident is occurring, sighting the NCSC; a full notification at 72 hours — the factsheet does not state whether that 72 hours runs from awareness or from the light-touch notice, and secondary legislation will settle it; NCSC informed at the same time as the regulator in a single submission; and a duty to notify affected customers with details, after the full regulator notification, so they can mitigate. On penalties, the Government describes a simplified two-band structure aligned with GDPR-style ceilings, with the higher band covering the security and notification duties and the standard band covering failure to register (enforcement factsheet). The amounts are in the [!VERIFY] box below, and they are not something to build a budget on yet.
Actionable takeaway: work out today whether you clear the headcount and turnover line — the definition is stable enough for that — then build one 24-hour notification capability, because the UK light-touch notice and the NIS2 early warning will be the same operational act with different recipients.
#EU DORA — you are not regulated, your contract is
DORA binds financial entities and designated critical ICT third-party providers. As an ordinary regional MSP you are not a DORA-regulated entity. What DORA does is rewrite your master agreement and make your financial-services client legally unable to sign the version you have used since 2019.
Art. 30(2) applies to every ICT service contract, regardless of criticality (Regulation (EU) 2022/2554): descriptions of services and the locations where they are provided and where data is processed; service levels with performance metrics; how availability, integrity, security and data protection are ensured; data recovery and return of data on insolvency or discontinuation; assistance during ICT incidents at no additional cost or at a pre-determined cost; cooperation with competent and resolution authorities; termination rights and notice periods; subcontracting conditions; and audit and inspection rights for the financial entity, its appointed third parties and the competent authority. Art. 30(3) adds more where the service supports a critical or important function: quantitative performance targets, notice periods for material developments, continuity plans, participation in the client's threat-led penetration testing, and exit strategies with mandatory transition periods. Art. 28(8) sets termination grounds that include "identified weaknesses in the provider's ICT risk management" — not a breach, just a weakness. Subcontracting is separately regulated by Commission Delegated Regulation (EU) 2025/532, in the Official Journal on 2 July 2025 (summary): your offshore NOC and backup vendor become your client's regulatory problem, and therefore yours.
Art. 28(3) requires financial entities to maintain a register of information on all ICT contractual arrangements, submitted annually to the national competent authority and forwarded to the ESAs; for the 2026 cycle Luxembourg's CSSF opened its portal on 11 February with submissions due 31 March, aligned to the ESAs' deadline (CSSF). At your service desk that arrives as a February spreadsheet request from every financial client at once: legal entity identifier, country of head office, country of data processing, service type, subcontracting chain, critical-or-important-function flag. Keep a standing, versioned data pack per client. Firms that improvise in March give inconsistent answers across clients, and inconsistency inside a supervisor's dataset is what draws the question.
Art. 31 lets the ESAs designate critical ICT third-party providers at group level, with Art. 31(11) allowing voluntary opt-in; first designations were made on 18 November 2025 (ESMA; EIOPA). For a designated provider, Art. 35 provides a periodic penalty payment of 1% of average daily worldwide turnover for up to six months — an exposure that never falls on the financial entity. Designation is out of reach for a regional MSP, but the criteria are systemic importance, substitutability and concentration, and an MSP that becomes the single operator for a cluster of small regulated firms is moving in the right direction on all three.
#GDPR — the processor layer, and where the 72 hours starts
Article 28 governs the DPA: process only on documented instructions; confidentiality commitments; Art. 32 security measures; sub-processor rules; assistance with data subject rights; assistance with Articles 32 to 36, which is where breach notification support lives; deletion or return at end of service; and making available all information necessary to demonstrate compliance, including contributing to audits. Art. 28(2) bars engaging another processor without prior specific or general written authorization and, under general authorization, requires you to inform the controller of intended additions so they can object. Art. 28(4) flows the same obligations down and leaves you fully liable to the controller for the sub-processor's performance (Art. 28). That last pair is the MSP-specific bear trap: your RMM SaaS, backup vendor, ticketing platform, offshore NOC and EDR vendor's cloud are all sub-processors, each needing authorization, an object-right mechanism and back-to-back terms, with your liability attached. Most MSP DPAs list three sub-processors. Most MSPs have fifteen.
Then the awareness question, which the EDPB settled in Guidelines 9/2022 v2.0 (PDF). Paragraph 44: the processor "does not need to first assess the likelihood of risk arising from a breach before notifying the controller"; it "just needs to establish whether a breach has occurred and then notify," and "the controller should be considered as 'aware' once the processor has informed it." Paragraph 47: where the processor serves multiple controllers affected by one incident, it "will have to report details of the incident to each controller." Paragraph 48: a processor may notify on a controller's behalf if authorized in the contract, but "the legal responsibility to notify remains with the controller."
Actionable takeaway: delete "let's confirm before we worry the client" from your incident vocabulary. Every hour you spend confirming comes off their 72, not yours, and the guidance forecloses the argument that you were being careful.
#CMMC — the bet-the-company decision, and Phase 2 is now
This is where MSPs are making their largest irreversible commercial decisions on the worst information, so I want to be precise. The 32 CFR Part 170 program rule took effect 16 December 2024; the 48 CFR acquisition rule took effect 10 November 2025, from which date DFARS 252.204-7021 appears in new DoD solicitations (Federal Register; DoD CIO). 32 CFR 170.3(e) sets four phases, each a calendar year apart from that date (32 CFR 170.3): Phase 1 from 10 Nov 2025, self-assessment at Level 1 or Level 2 as a condition of award; Phase 2 from 10 Nov 2026, adding Level 2 third-party (C3PAO) certification for applicable contracts; Phase 3 from 10 Nov 2027, Level 2 C3PAO for all applicable solicitations and Level 3 (DIBCAC) as a condition of award; Phase 4 from 10 Nov 2028, full implementation. Phase 2 begins in roughly two months, and the binding constraint is the assessment queue, not the technology.
Does the MSP itself need an assessment? Under 32 CFR 170.4 an External Service Provider is "external people, technology, or facilities that an organization utilizes for provision and management of IT and/or cybersecurity services on behalf of the organization" — with no requirement that it handle CUI. Security Protection Data is data stored or processed by assets providing security functions, expressly including "configuration data required to operate an SPA, log files generated by or ingested by an SPA, data related to the configuration or vulnerability status of in-scope assets, and passwords" (32 CFR 170.4). 32 CFR 170.19 then does the scoping (Cornell):
| You are… | Assessment consequence |
|---|---|
| An MSP with no DoD contract and no CUI, but holding Security Protection Data — their EDR, SIEM, patching, MFA, PAM, backups, credentials, logs | An ESP whose services are Security Protection Assets assessed inside the client's assessment. No certificate of your own; you are audited by proxy, against their System Security Plan. |
| An MSP that processes, stores or transmits the client's CUI and is not a CSP | Services assessed as part of the client's assessment. Still no separate certificate required. |
| An MSP offering a cloud service holding client CUI — multi-tenant file store, hosted VDI, backup cloud | You act as a CSP and must meet FedRAMP Moderate or DoD-recognized equivalency under DFARS 252.204-7012. The expensive path. |
| An MSP holding its own DoD prime contract or a flow-down involving CUI | An assessed organization in your own right, needing your own CMMC status at the level the contract requires. |
| Any of the above, certifying voluntarily | Permitted, and it reduces the client's assessment effort. The minimum assessment type is set by the client's contract. |
The trap is believing that "we don't touch CUI" puts you out of scope. It does not. Holding Security Protection Data — logs, configuration, vulnerability status, passwords — makes your platform a Security Protection Asset inside every DIB client's assessment. Your RMM, PSA, password vault and SIEM tenancy are all in that category, and they are all being assessed whether or not anyone sends you the report.
On equivalency: DFARS 252.204-7012(b)(2)(ii)(D) requires an external cloud provider handling covered defense information to meet requirements equivalent to the FedRAMP Moderate baseline (acquisition.gov). The DoD CIO memorandum of 21 December 2023 defines that as 100% of the FedRAMP Moderate baseline controls with no control-related POA&Ms, assessed by an accredited FedRAMP 3PAO, with a full body of evidence given to the contractor — and self-attestation is not permitted (memo; analysis). Note the asymmetry: equivalency is stricter than an actual FedRAMP authorization, because a real ATO tolerates open POA&Ms with risk acceptance and equivalency does not. If you are quietly running DIB client CUI on your own stack — the Hyper-V cluster sold as "our private cloud," the backup landing in your datacentre — you are a CSP under 7012 and probably not compliant. Three realistic paths: migrate that workload to an authorized environment, stop offering it to DIB clients, or fund a 3PAO equivalency assessment. That is a board paper, not a hallway conversation.
That preservation duty produces the most common MSP-caused evidentiary failure in this book: reimaging before imaging. Restoring the client fast is your instinct, your training and your contract. Do it in the wrong order on a 7012 client and you destroy their evidence and your own defense in one command.
#The rest of the map, at reference speed
| Regime | Applies to you? | Trigger | Clock, from what | To whom | Exposure |
|---|---|---|---|---|---|
| CIRCIA (CISA) | Potentially, in your own right — final rule unpublished, so nothing is required today | Covered cyber incident; separately, a ransom payment | 72h from reasonable belief the incident occurred; 24h from disbursement of a ransom payment | CISA | Not in force; do not plan around it |
| HIPAA business associate (45 CFR 164.410) | Yes, if you touch PHI for a covered entity | Discovery of a breach of unsecured PHI | Without unreasonable delay and no later than 60 calendar days from discovery — discovery being when you knew, or by exercising reasonable diligence would have known | The covered entity | Your BAA number, usually far shorter; and OCR can enforce against you directly as a business associate — see the Elgon settlement in Part 1 |
| PCI DSS 12.9 (v4.0.1) | Yes, as a third-party service provider | A customer request (12.9.2); a suspected compromise (12.10.1) | No standard clock; the brands' compromise programs expect immediate notification | The customer; the acquirer and brands per the responsibility matrix | Card-brand assessments and a forensic investigation flowed through your client |
| FTC Safeguards §314.4(j) (16 CFR 314.4) | Not directly — but the discovery is usually yours | Unauthorized acquisition of unencrypted customer information of 500+ consumers | As soon as possible, no later than 30 days from discovery | The FTC, by the client | A published filing naming a breach at a firm you administer |
| NYDFS §500.17 (23 NYCRR 500.17) | Not directly — but the rule names third-party service providers, so your incident is textually their reportable incident | The covered entity determines an incident occurred at it, an affiliate, or a third-party service provider; separately, an extortion payment | 72 hours from that determination; 24 hours after an extortion payment, plus a 30-day written explanation | The Superintendent, by the client | Contract termination and a §500.11 diligence failure |
| SEC Item 1.05 (Release 33-11216) | Not directly; may apply to both of you | The registrant determines an incident is material, including on third-party systems it uses | 4 business days from the materiality determination | The SEC, by the client | Your notification channel is their disclosure control |
Four of those rows deserve more than a cell.
HIPAA's discovery construction is the dangerous half. A breach is discovered on the first day it is known to the business associate "or, by exercising reasonable diligence, would have been known," imputed to any employee or agent other than the person who committed it. Your clock can start before anyone consciously registers a breach. An unreviewed alert queue is a running clock. Not a risk. A clock. And sixty days is a ceiling — your BAA has almost certainly shortened it to fifteen days or fewer, and the covered entity's own sixty-day budget for individuals and, at 500 or more, for HHS/OCR is consumed by whatever you take (HHS). Put your BAA number in the tenant record next to the after-hours phone list, not in a contracts folder.
PCI 12.9 is your requirement, not your client's. 12.9.1 requires you to give customers written agreements acknowledging your responsibility for the security of account data; 12.9.2 requires you to supply, on request, your compliance status and the split of which requirements are yours, which theirs, and which shared. 12.8.2 sets the scope trigger at providers "with which account data is shared or that could affect the security of the CDE" — you do not need to touch a card number, domain admin over an adjacent segment is enough. And 12.8.5's guidance addresses the MSP situation directly: where a primary provider contracts with secondary providers, "it is the responsibility of the primary TPSP to manage and monitor any secondary TPSPs." All of this is v4.0.1, the only active version, with the 51 future-dated requirements mandatory since 31 March 2025 (PCI SSC).
FTC Safeguards and NYDFS reach further than people expect. The Safeguards Rule binds financial institutions — a category covering mortgage lenders and brokers, finance companies, account servicers, collection agencies, credit counselors, tax preparation firms, investment advisors and finders — and defines a service provider as anyone "permitted access to customer information through its provision of services directly to a financial institution" (FTC). §314.4(f) makes your client select, contract with and periodically assess you; §314.4(h) makes them hold a written incident response plan, which means you are named in your clients' plans whether or not anyone told you. NYDFS §500.11(b) requires their contract with you to cover access controls including MFA under §500.12, encryption under §500.15, and notice of a cybersecurity event (23 NYCRR 500.11) — and since 1 November 2025 §500.12 requires MFA for any individual accessing any information system, which reaches your technicians' access into their tenant, not only their staff (NYDFS).
The SEC release matters if a client is public. The Commission is "not exempting registrants from providing disclosures regarding cybersecurity incidents on third-party systems they use," and materiality "is not contingent on where the relevant electronic systems reside or who owns them." Then the limit on how hard they must dig, which defines what your contract has to deliver: the rules "generally do not require that registrants conduct additional inquiries outside of their regular channels of communication with third-party service providers pursuant to those contracts" (Release 33-11216). Your contractual notification channel is your public client's disclosure control. If it is a shared mailbox nobody watches out of hours, their disclosure controls are defective and you built the defect.
#What this means commercially
Three things, and only one of them is a cost.
Questionnaires are now compliance artifacts. A NIS2-regulated client sending you a security questionnaire is not being difficult; Art. 21(3) obliges them to assess your practices and their regulator can ask to see the result. Same for a NYDFS covered entity under §500.11, an FTC-covered institution under §314.4(f), a DORA financial entity building its register. The wrong answer is to keep answering them one at a time, forever, at whatever quality the person on shift can manage. Build one maintained evidence set — SOC 2 Type II report with the period covered, a bridge letter for the gap since, ISO 27001 certificate with its scope statement and Statement of Applicability, your complementary subservice organization controls list, your complementary user entity controls list, and your subservice organization list with how you monitor them (Linford & Co; CompassITC) — and answer from it. Read your own complementary user entity controls before a client's auditor does; yours may say the client is responsible for something the client is certain you do.
Flow-down is arriving whether you renegotiate or not. DORA Art. 30 items, PCI 12.9.1 acknowledgements, GDPR Art. 28(3) terms, NYDFS §500.11(b) commitments, BAA clocks — these land in your agreement through your clients' obligations, not through your sales process. Your only choice is whether you author your position once, deliberately, with counsel, or accept forty versions drafted by forty client-side lawyers over three years.
Certification is the honest cost, and also the differentiator. The expensive paths are real: a 3PAO equivalency assessment if you host DIB client CUI, a C3PAO Level 2 assessment if you hold your own contract, SOC 2 Type II over a useful period. The cheap end is real too. NCSC-UK's Choosing a Managed Service Provider, published 24 November 2025 for the small businesses who are your prospects, names Cyber Essentials Plus as the government baseline with ISO 27001 and SOC 2 as further evidence (NCSC). That document is functionally the scorecard your prospects hold while they interview you: certifications, references, defined SLAs, notification timeframes by severity, least privilege on MSP access, and explicit documentation of accountability and liability for incidents. Every one of those questions is answerable by a twelve-person MSP that has done the work and unanswerable by a two-hundred-person MSP that has not. Regulatory standing does not correlate with headcount. It correlates with whether somebody sat down for two days and wrote the answers.
Actionable takeaway: pick the certification whose scope statement covers the service lines you actually sell, get it, and put the evidence set behind one link your salespeople can send. Then check the scope statement again. A certificate that excludes the service line that failed will be read against you — by the client's auditor, their insurer's subrogation counsel, and their regulator, in that order.
Regulatory standing used to be the thing you helped clients with. It is now the thing they check about you before they let you near the domain admin account. Register where you must, classify before you are forced to, and read your own scope statement before somebody else reads it for you.
#Checklist
- MSP-46A written scope determination exists for each regime — NIS2 Annex I sector 9 (with the Art. 2(1) size test and Art. 2(2) exceptions considered), UK RMSP definition, DORA counterparty status, CIRCIA proposed criteria, CMMC ESP/CSP status — reviewed annually and after any material change in headcount, turnover or client mix.
[IG1][GV.OC][A.5.36] - MSP-47Where the MSP has an EU main establishment, the NIS2 Art. 27 registration has been submitted (entity name, sector, establishment addresses, current contacts, Member States served, IP ranges) and the receipt retained; where the MSP is non-EU and serves EU clients, a Union representative has been designated in writing.
[IG1][GV.OC] - MSP-48A per-client regulatory register exists as structured PSA fields — contractual notification deadline in hours, legal notice address and method, named out-of-hours recipients, regime flags (GDPR / NIS2 / DORA / BAA / PCI / DFARS 7012 / NYDFS / public company), data categories held, and whether the MSP may notify a regulator on the client's behalf.
[IG1][ID.AM][CIS 15] - MSP-49One notification capability is built to the tightest clock the MSP faces (24 hours), with a pre-drafted T+0 pack per client, so only the recipient and form vary by regime.
[IG1][RS.CO][A.5.24] - MSP-50Client notifications are issued individually per affected client with that client's own facts, never as a single broadcast, and the template covers the content a controller needs for GDPR Art. 33(3).
[IG1][RS.CO] - MSP-51A written procedure states that the MSP notifies affected clients on establishing that a breach occurred, without first assessing likelihood of risk, and names the role authorized to send that notification out of hours.
[IG1][RS.CO][GV.RR] - MSP-52The sub-processor list in every client DPA reconciles to the actual delivery-path tooling inventory (RMM, PSA, backup, EDR cloud, ticketing, offshore NOC), with a working mechanism for notifying intended additions and allowing objection; reconciliation is performed at least quarterly.
[IG2][GV.SC][A.5.19][CIS 15] - MSP-53For every business associate relationship the BAA breach-reporting timeframe is recorded in the tenant record alongside the after-hours contacts, and the alert queue is reviewed on a defined daily cadence with the review evidenced, because HIPAA discovery is imputed on reasonable diligence.
[IG1][DE.AE][A.5.25] - MSP-54A PCI DSS responsibility matrix exists per service (and per client where the split differs), is producible on customer request under Requirement 12.9.2, and states explicitly which party notifies the acquirer and payment brands on suspected compromise.
[IG2][GV.SC][CIS 15] - MSP-55A versioned DORA data pack is maintained per financial-entity client — legal entity identifier, country of head office, countries of service provision and data processing, service type, subcontracting chain, critical-or-important-function flag — refreshed before the annual register-of-information cycle rather than during it.
[IG2][ID.AM][GV.SC] - MSP-56Agreements with financial-entity clients have been reviewed against DORA Art. 30(2) and, where the service supports a critical or important function, Art. 30(3): audit and inspection rights extending to the competent authority, incident assistance at no additional or pre-agreed cost, a documented exit plan with transition period, and TLPT cooperation.
[IG2][GV.SC][A.5.20] - MSP-57For each Defense Industrial Base client a written determination records whether the MSP handles CUI, handles Security Protection Data only, or operates a cloud service holding CUI — and, where it acts as a CSP, states the FedRAMP Moderate authorization or DoD-recognized equivalency position with its supporting body of evidence.
[IG2][GV.OC][GV.SC] - MSP-58Where a DFARS 252.204-7012 flow-down applies, the 72-hour reporting path to DoD and to the prime is documented with the report-number handover step, 90-day preservation of system images and monitoring data is a standing rule, and reimaging before imaging is prohibited by default with a named role able to authorize an exception.
[IG1][RS.CO][A.5.28] - MSP-59The management body has formally approved the cybersecurity risk-management measures and completed the required training, with the approval and training recorded in dated minutes retained as evidence.
[IG2][GV.OV][GV.RR] - MSP-60The MSP holds at least one third-party attestation or certification whose scope statement demonstrably covers the service lines it sells, and maintains a single evidence set — report with period covered, bridge letter, scope statement and Statement of Applicability, CSOC list, CUEC list, subservice organization list and monitoring evidence — ready to send without bespoke work.
[IG3][GV.SC][A.5.35][A.5.36]
#Sources
- NIS2 Directive, full text
- NIS2 Art. 2 (scope)
- NIS2 Art. 3 (essential and important entities)
- NIS2 Art. 6 (definitions)
- NIS2 Art. 20 (governance)
- NIS2 Art. 21 (risk-management measures, supply chain)
- NIS2 Art. 23 (reporting obligations)
- NIS2 Art. 26 (jurisdiction)
- NIS2 Art. 27 (registry)
- NIS2 Art. 32 (supervisory measures)
- NIS2 Art. 34 (administrative fines)
- Commission Implementing Regulation (EU) 2024/2690
- Commission Recommendation 2003/361/EC (SME definition)
- NIS2 MSP scope guidance
- European Commission, NIS2 transposition infringements
- ECSO NIS2 transposition tracker
- UK Parliament, CSR Bill Lords Second Reading
- gov.uk, CSR Bill summary factsheet
- gov.uk, Relevant Managed Service Providers factsheet
- gov.uk, incident reporting factsheet
- gov.uk, enforcement factsheet
- gov.uk, Research on UK managed service providers
- Regulation (EU) 2022/2554 (DORA)
- Commission Delegated Regulation (EU) 2025/532 (subcontracting RTS)
- CSSF, DORA register of information 2026 submission window
- ESMA, CTPP designations
- EIOPA, CTPP designations
- GDPR Art. 28
- GDPR Art. 33
- EDPB Guidelines 9/2022 on personal data breach notification, v2.0
- CISA, CIRCIA status
- CIRCIA NPRM, 89 FR (4 April 2024)
- CISA, CIRCIA NPRM informational overview_508c%20(locked).pdf)
- 48 CFR CMMC final rule (10 September 2025)
- DoD CIO, CMMC
- 32 CFR 170.3 (phased implementation)
- 32 CFR 170.4 (definitions)
- 32 CFR 170.19 (scoping)
- DFARS 252.204-7012
- DoD CIO, FedRAMP Moderate equivalency memorandum (21 December 2023)
- Secureframe, FedRAMP equivalency and CMMC
- 45 CFR 164.410
- HHS, Breach Notification Rule
- HHS, proposed Security Rule update
- PCI DSS v4.0.1 Requirements and Testing Procedures
- PCI SSC, future-dated requirements
- FTC, Safeguards Rule guidance
- 16 CFR 314.4
- 23 NYCRR 500.11
- 23 NYCRR 500.17
- NYDFS, cybersecurity resource centre
- NYDFS, how to report an extortion payment
- SEC Release 33-11216, cybersecurity disclosure final rule
- NCSC-UK, Choosing a Managed Service Provider
- Linford & Co, subservice carve-out vs inclusive audits
- CompassITC, subservice organizations in SOC reports
#Part 6 — Templates and Tools
#Templates and Tools
Everything from here on is meant to be torn out. If you read the rest of this book, nodded along, and then went back to the ticket queue, nothing changed. These five appendices are the parts that change something: a matrix you fill in, a clause library you hand to counsel, a manifest you copy into a folder structure, six exercises you can run on a Thursday afternoon with a whiteboard and no budget, and a ninety-day plan that starts with things that cost nothing.
Use them as drafts, not scripture. Your clients, your stack and your jurisdiction are not mine.
#Appendix A — The Client Authority Matrix
This is the single highest-value artifact in this book, and it is a spreadsheet. Not a platform, not a product, not a subscription. A spreadsheet with names and phone numbers in it, reviewed twice a year, that answers one question at 02:00: am I allowed to do this at this client, and if not, whose phone do I ring?
The structure below is synthesized from the requirements in Part 2 and Part 5 — NIST SP 800-61r3 on defining third-party authority and restrictions, CISA AA22-131A on contracts specifying who owns incident response, and the two published commercial implementations worth copying: the Sophos MDR service description's three authority modes with named primary and alternate approvers, and Microsoft Defender Experts' approach of encoding the authority tier in the RBAC role rather than in a PDF (NIST SP 800-61r3 · CISA AA22-131A · Sophos MDR service description · Microsoft Defender Experts for XDR).
#A.1 The blank template — header block
One record per client. Store it where the on-call technician can reach it in ten seconds without the PSA, because the PSA may be the thing that is down.
| Field | Value |
|---|---|
| Client legal entity name | |
| Contracting entity on the MSA (if different) | |
| Matrix version / effective date | |
| Signed by (client role) / signed by (MSP role) | |
| Next scheduled review date | |
| Contact path last tested (date) / result | |
| GDPR role (controller / processor / N.A.) | |
| BAA in place (Y/N) / notification window (days) | |
| PCI in scope (Y/N) / responsibility matrix version | |
| NIS2 / DORA / UK CSR status | |
| Contractual notification deadline to client (hours) | |
| Legal notice address and method | |
| Client cyber insurer / policy no. / 24-7 claims line | |
| Client panel breach counsel / panel DFIR firm | |
| MSP carrier claims line | |
| Data residency or jurisdictional constraints | |
| Acknowledgment SLA (minutes) | |
| Unreachable escalation window (minutes) and what unlocks |
#A.2 The blank template — approvers
| Role | Name | Mobile | Personal email | Out-of-hours OK | Escalate after |
|---|---|---|---|---|---|
| Primary Client Authority Holder | |||||
| Deputy Authority Holder | |||||
| Executive escalation (production-down decisions) | |||||
| Client legal / DPO | |||||
| Client communications lead | |||||
| Out-of-band channel (not their M365) |
The out-of-band row is not decoration. If your only route to the approver is the tenant you are about to isolate, you have no route to the approver.
#A.3 The blank template — action rows
Tiers: P = pre-authorized, act and notify · A = approval required from a listed approver · X = executive approval only · N = never the MSP's decision.
| # | Action | Tier | Named approver | Notify within | Notes / conditions |
|---|---|---|---|---|---|
| 1 | EDR-isolate a single endpoint | ||||
| 2 | Kill a process / quarantine a file | ||||
| 3 | Disable a single user account; revoke sessions and refresh tokens | ||||
| 4 | Block a specific IP, domain or URL at egress | ||||
| 5 | Disable inbox rules or a mail flow rule | ||||
| 6 | Reset one privileged credential | ||||
| 7 | Isolate a server or hypervisor host | ||||
| 8 | Isolate a domain controller | ||||
| 9 | Tenant-wide password reset / mass session revocation | ||||
| 10 | Block an egress class or sever a site-to-site VPN | ||||
| 11 | Take a production application offline | ||||
| 12 | Restore from backup over live data | ||||
| 13 | Engage third-party DFIR | ||||
| 14 | Notify a regulator on the client's behalf | ||||
| 15 | Pay or negotiate a ransom |
Rows 14 and 15 exist so that the answer is written down before anyone is tempted. Ransom decisions belong to the client with their counsel and their insurer. Regulatory notification on a controller's behalf requires express authorization and the legal responsibility stays with them regardless.
#A.4 Worked example — fictional client, illustrative only
EXAMPLE ONLY. Meridian Orthopaedic Group is invented for this appendix. The values are plausible, not recommended; a real matrix is negotiated per client.
| Field | Value |
|---|---|
| Client legal entity name | Meridian Orthopaedic Group LLC (240 staff, 6 sites) |
| Matrix version / effective date | v3.1 / 2026-04-01 |
| Signed by | COO (client) / Practice Owner (MSP) |
| Next scheduled review date | 2026-10-01 |
| Contact path last tested | 2026-07-14 — pass on primary, deputy reached on 2nd attempt (4 min) |
| GDPR role | N.A. (US only) |
| BAA in place / notification window | Yes / 10 calendar days from discovery |
| PCI in scope | Yes, SAQ-A-EP / responsibility matrix v2 (2026-02) |
| Contractual notification deadline to client | 4 hours from MSP discovery |
| Legal notice address | legal@ (per MSA §14) with copy to COO |
| Client cyber insurer / claims line | [carrier] / [24-7 number] / panel counsel and panel DFIR named in policy |
| Acknowledgment SLA | 60 minutes |
| Unreachable escalation | 45 minutes across all listed contacts unlocks tier A rows 7 and 8 |
| Role | Name | Out-of-hours | Escalate after |
|---|---|---|---|
| Primary Authority Holder | IT Director | Yes, mobile | 15 min |
| Deputy | Clinical Operations Manager | Yes, mobile | 15 min |
| Executive escalation | COO | Yes | 30 min |
| Out-of-band channel | Signal group "MOG-IR", 4 members, tested quarterly | — | — |
| # | Action | Tier | Approver | Notify within |
|---|---|---|---|---|
| 1 | EDR-isolate a single endpoint | P | — | 60 min |
| 3 | Disable a user; revoke sessions and tokens | P | — | 60 min |
| 4 | Block a specific IP/domain/URL | P | — | 4 h |
| 6 | Reset one privileged credential | P | — | 60 min |
| 7 | Isolate a server or hypervisor host | A | IT Director or deputy | at time of action |
| 8 | Isolate a domain controller | A | IT Director or deputy; after 45 min unreachable, unlocks on Practice Owner authorization (MSP Incident Commander recommends) | at time of action |
| 9 | Tenant-wide password reset | A | IT Director | at time of action |
| 11 | Take a production application offline | X | COO | at time of action |
| 11a | Exception: imaging PACS workstations | X | COO and Clinical Operations Manager — never unlocked by the unreachable clock | — |
| 12 | Restore from backup over live data | X | COO | — |
| 13 | Engage third-party DFIR | A | IT Director, after insurer panel check | — |
| 15 | Pay or negotiate a ransom | N | Client, counsel, insurer | — |
Read the unreachable clock precisely, because it is the row people misread at 02:00. It changes whose consent you are relying on. It does not remove the requirement for a named MSP authoriser. When it expires on rows 7 and 8, the action still needs the Practice Owner's yes, on the recommendation of the MSP Incident Commander, timestamped at the moment it was given — the rule set out in Part 1's second scenario. An expired clock is not a pre-authorization, and it is never the On-Call Technician's decision alone.
Row 11a is the point of doing this exercise with a real client. Clinical imaging is the one system this fictional client will not let anyone touch without a clinician in the loop, at any hour, for any reason — and the only way you find that out is by asking in daylight, not by discovering it at 02:00 with a scanner offline and a surgeon on the phone.
#A.5 Review cadence, and how to actually test the contact path
Three moments, and only three. At onboarding, as part of the standard build, before the first invoice. At renewal, alongside the commercial conversation, because that is when the client is already reading the paperwork. After any incident or near miss at that client, because that is when they care most and argue least. Never mid-incident — the matrix you write at 02:15 while an encryption run is going is not a matrix, it is a hostage note.
Between those, run a calendar reminder at six months. Anything untouched in twelve months is presumed wrong.
Testing the contact path is a separate exercise and takes about four minutes per client:
| # | Step | Who | Done when | Record |
|---|---|---|---|---|
| 1 | Pick a date and time deliberately outside business hours. Do not warn the contacts. | MSP Incident Commander | Scheduled | Date, time, tester |
| 2 | Call the primary Authority Holder on the listed mobile. State plainly that this is a scheduled contact test, not an incident. | Tenant Lead | Answered or voicemail | Answer time in minutes, or no answer |
| 3 | If no answer within the escalation window, call the deputy. | Tenant Lead | Answered or voicemail | Answer time, or no answer |
| 4 | Send a message on the out-of-band channel and ask for a reply. | Tenant Lead | Reply received | Reply time |
| 5 | Ask one question: "who else at your organization can approve taking a server offline tonight?" | Tenant Lead | Answered | Whether the answer matches the matrix |
| 6 | Update the matrix, record the test date and result, and raise a ticket for any number that failed. | Tenant Lead | Matrix saved | New version, gaps ticketed |
Step 5 is the one that finds the failure. Numbers go stale quietly; authority goes stale loudly, usually because the person on the form left in March.
#Appendix B — MSA and DPA clause library
Read this first. Everything in this appendix is illustrative drafting shape, written to show what a clause has to accomplish. It is not legal advice, it is not a quotation from any executed agreement, and it has not been drafted or reviewed by a lawyer for your jurisdiction. Take it to counsel who acts for you, in the jurisdiction of the contract, and let them write the language. The structures are modeled on published commercial models cited in each entry — chiefly the Google Workspace "Emergency Security Issue" suspension structure, the Sophos MDR authority-mode structure, and the contract-content checklist in DORA Article 30 (Google Workspace terms · Sophos MDR service description · DORA, Regulation (EU) 2022/2554).
Each entry gives you four things: what the clause does, what breaks without it, the failure mode you will actually experience, and who at your MSP owns getting it into the paper.
#B.1 Emergency security action
What it does. Defines a "Security Emergency" by trigger, grants the MSP power to take listed containment actions when one occurs, limits that power to the minimum extent and duration required, and imposes a post-hoc notice duty on a fixed clock. Copy the Google structure exactly: trigger definition → power → minimum-extent-and-duration limiter → post-hoc notice. The limiter is what makes a client's counsel sign it.
What breaks without it. Everything in Part 1's second scenario. Your authority to isolate a domain controller comes from an express clause, a documented standing pre-authorization, or a live approval from a named person. With none of the three, you are exposed for acting and exposed for not acting.
Failure mode. A technician with Domain Admin makes a defensible technical decision at 02:00 and an indefensible contractual one, and the argument afterwards is about willful misconduct — which is the carve-out to your own liability cap.
Owner: Practice Owner.
Illustrative shape. Define "Security Emergency" as a confirmed or reasonably suspected unauthorized access, active malware propagation, or credential compromise affecting the client environment. On a Security Emergency, the provider may take the actions listed at the pre-authorized tier of the Authority Matrix, to the minimum extent and for the minimum duration reasonably required to prevent or terminate the emergency; must notify the primary Authority Holder within a stated number of minutes of taking action and provide a written summary within a stated number of hours; and action taken in good faith under this clause is authorized, is not a breach, and any resulting unavailability is excluded from service credits.
#B.2 Right to suspend or disconnect
What it does. Gives you a security suspension right that is separate from your non-payment suspension right.
What breaks without it. Published MSP agreements near-universally carry suspension for non-payment and near-universally lack suspension for security. In the CompassMSP agreement, the only unilateral powers are 30-days'-notice service changes and suspension for non-payment (CompassMSP MSA). Thirty days' notice is not a containment control.
Failure mode. A client refuses to let you disconnect a compromised site from your management plane, and you are left choosing between contract breach and knowingly bridging an infected estate to the rest of your book.
Owner: Practice Owner.
#B.3 Security incident cooperation
What it does. Obliges both sides to cooperate — access, information, personnel, timelines — during an incident, at no additional cost or at a cost fixed in advance.
What breaks without it. Read your own service description carefully. One published MSP agreement excludes incident response beyond initial triage and alerting, making full-scale containment, forensics and remediation "a separate, billable engagement" (Secure Data Technologies MSA). That exclusion is a commercial choice you are entitled to make — but it must be a choice you made, not one you discover mid-incident. DORA Article 30(2) requires the opposite for EU financial-services clients: assistance in ICT incidents "at no additional cost or at a cost determined ex-ante".
Failure mode. At 02:40 someone has to decide whether the next eight hours are billable, and the person deciding is the one holding the keyboard.
Owner: Practice Owner, with the service delivery lead pricing it.
#B.4 Evidence preservation and the legal-hold carve-out
What it does. Requires both parties to preserve incident-related records, names an evidence custodian, and — critically — carves legal hold out of the contractual deletion obligation.
What breaks without it. GDPR Article 28(3)(g) requires the processor to delete or return all personal data at the end of the services (Art. 28 GDPR). Your standard offboarding clause almost certainly says the same. If litigation or a regulatory investigation is live, that clause and your preservation duty point in opposite directions, and you get to choose which one to breach.
Failure mode. A client terminates during an incident, invokes the deletion clause, and you delete the only complete copy of the evidence that would have shown you were not at fault. There is no version of that story that ends well for you.
Owner: Practice Owner with counsel; the evidence custodian role is named by the MSP Incident Commander.
Illustrative shape. The deletion and return obligations are expressly suspended in respect of any data subject to a legal hold, litigation, regulatory investigation or reasonably anticipated claim, until the hold is released in writing; the provider maintains a written record of what was held, when, by whom, and when released.
#B.5 Notification timing — MSP to client
What it does. States a number, states what starts it, and states the notice method and address.
What breaks without it. You are still bound. GDPR Article 33(2) obliges a processor to notify the controller "without undue delay" on becoming aware of a breach, and the EDPB is explicit that the processor does not first assess risk — establish that a breach occurred, then notify, because the controller's own 72-hour clock only starts when you tell them (EDPB Guidelines 9/2022). A HIPAA business associate has no later than 60 calendar days from discovery, where discovery is imputed on the day a reasonably diligent MSP would have known (45 CFR 164.410) — and your BAA has probably shortened that. The number is negotiated, not regulatory, and it is routinely much shorter than sixty days. Plan in days, not months, and read the actual number out of the actual agreement.
Failure mode. The two failures documented in Part 4 both live here: investigating before notifying, which spends the client's regulatory budget on your certainty; and notifying the day-to-day IT contact rather than the contractual notice address, which may not constitute notice at all.
Owner: Practice Owner sets the number; the service delivery lead makes sure it is a readable field in the PSA rather than a sentence in a PDF.
#B.6 Notification timing — client to MSP
What it does. Obliges the client to tell you about incidents in their own environment, to maintain current approver details, and to acknowledge your outreach within a stated window.
What breaks without it. You find out about your own client's ransomware from an EDR alert at 23:45, or from a journalist. Sophos's published service description carries the model on both halves: the customer must designate and keep current at least one primary and one alternate authorized contact with authority to approve response actions, and use commercially reasonable efforts to acknowledge outreach within 60 minutes (Sophos MDR service description).
Failure mode. The unreachable client of Part 1's second scenario, permanently. Without a contractual acknowledgment SLA you have no agreed definition of "unreachable", and therefore no defensible moment at which the unreachable escalation fires.
Owner: Practice Owner, enforced by the Tenant Lead at every review.
#B.7 Client responsibilities and CUECs
What it does. Lists what the client must do for your controls to work, and aligns that list with the complementary user entity controls in your SOC 2 report.
What breaks without it. Your SOC 2 says the client is responsible for reviewing your access, maintaining their own approvers and applying MFA to the accounts you do not manage. Your MSA says nothing about any of it. When those two documents disagree, the one without the auditor's name on it is the one that governs the relationship — and the gap between them is precisely where every "we thought you were doing that" conversation happens.
Failure mode. An MFA exception you carried for two years because a client's leadership objected. That specific pattern was an aggravating factor in the ICO's penalty against a processor: a working solution existed and was not rolled out on a perception that clients would resist (Clifford Chance analysis).
Owner: whoever owns your SOC 2 or ISO 27001 program, jointly with the Practice Owner.
#B.8 Limitation of liability, and the carve-in
What it does. Caps your exposure and excludes consequential loss — and then adds the part most MSPs miss, which is an express statement that good-faith action inside the emergency authority clause is authorized conduct.
What breaks without it. Your limitation clause is your primary defense against a wrongful-containment claim, because business interruption and lost profits are exactly the heads of loss it excludes. The danger is the carve-out: caps are almost always disapplied for gross negligence and willful misconduct, and a deliberate shutdown is by definition an intentional act. Your protection against re-characterization is not the clause — it is the contemporaneous record showing a good-faith, proportionate, minimum-extent decision.
Failure mode. You did the right thing at 02:00 and cannot prove it, and a cap you drafted stops protecting you at the exact moment you need it.
Owner: Practice Owner with counsel. Do not weaken the cap to win a deal without pricing what you just gave away.
Also worth asking counsel about: the refusal shield. Sophos publishes one — no liability for new or worsened malicious activity where the customer denied a requested authorization (Sophos MDR service description). That clause converts "the client said no" from an uninsured liability into a documented allocation of risk, and it is the piece most MSP agreements are missing.
#B.9 Insurance requirements
What it does. Requires each side to carry stated cover, to name the other where appropriate, to produce evidence on request, and — the row people forget — to waive subrogation against you.
What breaks without it. Your liability cap does not bind people who are not parties to your agreement. Regulators are not parties. Neither is your client's insurer exercising subrogation rights. A cyber insurer that paid out on a client's ransomware claim has sued that client's technology vendors directly, pleading negligence and breach of contract — including, against one vendor, a pleaded failure to properly notify appropriate parties of the incident, allegedly preventing timely action and increasing damages (Hunton · National Law Review). Nothing has been adjudicated. The pleading is the point: the alleged negligence is failure to escalate.
Failure mode. You survive the incident, the client's carrier pays the client, and then the carrier comes for you eighteen months later with your own ticket timeline as an exhibit.
Owner: Practice Owner. This one is not delegable.
#B.10 Sub-processor terms
What it does. Names your sub-processors, sets the authorization and change-notice mechanism, and flows your obligations down to them.
What breaks without it. GDPR Article 28(2) forbids engaging another processor without the controller's prior specific or general written authorization, and Article 28(4) makes you fully liable to the controller for that sub-processor's performance (Art. 28 GDPR). Read that as an MSP: a compromise at your RMM vendor is your breach to notify, to every affected client, on each of their clocks.
Failure mode. Your vendor is breached, and you cannot answer the first question every client asks — "which of our data touched them?" — because you have never maintained the list.
Owner: service delivery lead maintains the register; Practice Owner signs the change notices.
#B.11 Termination and offboarding, including delegated-access removal
What it does. Defines exit: transition assistance, data return and deletion, and the removal of every standing access path, with a checklist and an attestation.
What breaks without it. AA22-131A specifically calls out that disabling provider accounts at contract termination is commonly overlooked (CISA AA22-131A). For Microsoft-managed clients the paths are plural and none of them removes the others: GDAP relationships, Entra B2B guest objects, Azure Lighthouse delegations, RMM agents, backup console registrations, VPN accounts and shared credentials. Note two documented asymmetries you should disclose rather than let the client discover: Lighthouse role assignments do not appear in the client's own IAM blade or in az role assignment list, and client-side resource locks do not prevent actions by managing-tenant users (Azure Lighthouse cross-tenant management).
Failure mode. Standing privileged access into a former client's tenant, discovered by their new provider, in a security review, with your name on it. There is no innocent explanation that survives that meeting.
Owner: service delivery lead runs the checklist; Practice Owner signs the attestation and sends it.
#Appendix C — Per-tenant evidence pack manifest
Part 1's third scenario explains why this exists and what each artifact proves. This appendix is only the shape, so you can copy it into a folder and start filling it.
Two rules govern the whole pack. Nothing enters it without a hash and a recorded scope filter, because the test the pack has to pass is that the collection method was transparent and reproducible. And 80_gaps_and_limitations.md is written by you, not discovered by the other side. Volunteering the log source that was switched off is a credibility gain. Being caught by it is a credibility loss you do not recover in the same meeting.
#C.1 Per-artifact metadata
Every artifact carries these nine fields. This is the schema of 00_MANIFEST.csv.
| Column | Content |
|---|---|
artefact_path | Relative path within the pack |
source_system | Console, tenant and product that produced it |
scope_filter | The exact filter applied — tenant ID, date range, user set, host set |
export_mechanism | Portal export, PowerShell cmdlet, Graph call, API, manual |
operator | Named person who performed the export |
export_timestamp_utc | ISO 8601, UTC |
coverage_window_utc | Start and end of the data, ISO 8601, UTC |
sha256 | Hash of the file as exported |
completeness_note | Truncation, throttling, row caps, license limits, anything missing |
artefact_path,source_system,scope_filter,export_mechanism,operator,export_timestamp_utc,coverage_window_utc,sha256,completeness_note
#C.2 The directory structure
EVIDENCE_<CLIENT>_<INCIDENT-ID>/
├── 00_cover_and_control/
│ ├── 00_MANIFEST.csv # one row per artefact, schema above
│ ├── 00_CHAIN_OF_CUSTODY.pdf # discovered/collected/handled/transferred: who, when, where
│ ├── 00_AUTHORITY.pdf # MSA or DPA clause, GDAP record, written instruction, engagement letter
│ ├── 00_CUSTODIAN.txt # the single named MSP evidence custodian and deputy
│ ├── 00_HASHES.txt # algorithm, tool, version — stored separately from the artefacts
│ ├── 00_TOOLS.csv # every collection tool, version, licence
│ ├── 00_SEGREGATION_ATTESTATION.pdf # scope filters, reviewer name, redaction log
│ └── 00_EVIDENCE_BASELINE.pdf # sources enabled, retention, since when, licence tier
├── 10_identity/
│ ├── 10_entra_signins_<UTCrange>.json # incl. non-interactive and service principal
│ ├── 10_entra_auditlogs_<UTCrange>.json
│ ├── 10_conditional_access_policies_<date>.json # + CA policy change history
│ ├── 10_risky_users_risk_detections.csv
│ ├── 10_oauth_grants_serviceprincipals.csv
│ └── 10_privileged_role_assignments_<date>.csv
├── 20_m365_collaboration/
│ ├── 20_ual_export_<segment>.csv # + 20_ual_segment_index.csv proving segmentation under the cap
│ ├── 20_mailitemsaccessed_<user>.csv # retain the IsThrottled field
│ ├── 20_inboxrules_and_forwarding_<date>.csv
│ ├── 20_ediscovery_hold_record.pdf # case, locations, applied timestamp
│ └── 20_ediscovery_export/ # Summary.csv, load file with native hashes, warnings/errors, natives, text
├── 30_endpoint_xdr/
│ ├── 30_incident_<id>_export.json # + alert exports
│ ├── 30_device_timeline_<host>_<UTCrange>.csv
│ ├── 30_advanced_hunting/ # query text + results + row counts, flagging any truncation
│ ├── 30_investigation_package_<host>.zip # including CollectionSummaryReport.xls
│ ├── 30_liveresponse_<sessionid>_commandlog.txt
│ ├── 30_response_actions_audit.csv # isolation, scans, app restriction
│ └── 30_memory_image_<host>.raw # and/or disk image + acquisition log and hashes, where taken
├── 40_network_perimeter/
│ ├── 40_firewall_logs_<UTCrange>/ # + 40_firewall_retention_config_<date>.txt
│ ├── 40_vpn_remote_access_auth_<UTCrange>.csv
│ ├── 40_dns_proxy_egress_<UTCrange>/
│ └── 40_network_topology_<date>.pdf
├── 50_backup_recovery/
│ ├── 50_backup_job_history_<UTCrange>.csv
│ ├── 50_restore_points_inventory_<date>.csv # with immutability state and expiry
│ ├── 50_restore_sessions_and_verification.csv
│ └── 50_recovery_timeline.md # service-by-service restoration — feeds the insurer's BI calculation
├── 60_msp_side_records/
│ ├── 60_gdap_activity_log_<range>.csv # + 60_gdap_relationships_<date>.csv
│ ├── 60_lighthouse_audit_<range>.csv # all tabs
│ ├── 60_azure_lighthouse_activity_<range>.csv
│ ├── 60_technician_objectid_map_<date>.csv # object ID -> technician -> employment dates
│ ├── 60_rmm_activity_log_<site>_<range>.csv # + segment index where the export is row-capped
│ ├── 60_remote_session_audit_<range>.csv # + recordings where retained
│ ├── 60_psa_tickets_<tenant>_<range>/ # tickets, audit trails, time entries, attachments, internal notes
│ ├── 60_change_records_<range>.csv
│ ├── 60_patch_and_vuln_evidence_<range>/ # + exception register
│ ├── 60_alerting_config_and_history_<range>/
│ ├── 60_oncall_roster_and_paging_<range>.csv
│ └── 60_access_reviews_and_offboarding_<period>/
├── 70_assurance_and_legal/
│ ├── 70_soc2_typeII_<period>.pdf # + bridge letter
│ ├── 70_iso27001_cert_and_soa.pdf
│ ├── 70_cuec_responsibility_matrix_<client>.xlsx
│ ├── 70_legal_hold_notice_and_acknowledgements.pdf
│ ├── 70_retention_suspension_record.csv # what was suspended, when, by whom, released when
│ ├── 70_notification_record.pdf # regulator, insurer, client, law enforcement — with timestamps
│ └── 70_cost_and_effort_log.csv # persons, roles, time spent, approximate rates
└── 80_analysis/
├── 80_master_timeline_UTC.csv # schema below
├── 80_scope_of_exposure.md # accounts, mailboxes, files, hosts, data classes + the queries behind each number
└── 80_gaps_and_limitations.md # every source off, expired, throttled or truncated — with reason and date
#C.3 The master timeline schema
One row per event, UTC and ISO 8601 throughout, with the confidence column that keeps you honest:
timestamp_utc,source_system,source_rendered_timestamp,source_timezone,actor,action,target,artefact_ref,confidence
confidence takes exactly two values: observed or inferred. An auditor will forgive an inference. Nobody forgives an inference presented as an observation.
The 60_ directory is the one most MSPs have never assembled, and it is the one where your own diligence is on trial rather than your client's. Build it first.
#Appendix D — Six MSP tabletop scenarios
Six exercises, each self-contained, each runnable in ninety minutes. I have never once seen a tabletop fail because it was too realistic.
How to run one. You need two people who are not participants: a facilitator who leads the discussion and a data collector who writes down what happened. Keep the scenario short — the guidance in NIST SP 800-84 is blunt about this, warning that with long scenarios participants spend more time dissecting the scenario than meeting the objectives, and that a short, concise scenario is often more effective. Deliver injects in time sequence rather than handing out the whole story at once, which is how CISA's tabletop exercise packages are built. Write the evaluation criteria before the exercise. Finish with a hotwash, then an after-action report where every finding gets an owner and a due date (NIST SP 800-84 · CISA CTEP).
Score the plan, not the people. Four grades per objective — performed without challenges, with minor challenges, with major challenges, unable to perform — plus one number that matters more than all of them: how many decisions stalled waiting for an authority who was not in the room.
And one rule specific to this list. At least one inject in every exercise should have no good answer. The most useful sentence a team can learn to say out loud is "we do not know, and we cannot find out in time" — because that sentence, said at 03:00, is what triggers the right escalation instead of an hour of confident guessing. Injects marked [no-answer] below are there for exactly that.
#D.1 TTX-01 — 23:45, malware spreading across a client domain
Exercises: Part 1, Scenario One · Part 2 cross-tenant sweep · Appendix A. Participants: On-Call Technician, MSP Incident Commander, Tenant Lead. Objectives: (1) validate before remediating; (2) run the "is this only them, and is it us?" question inside the first twenty minutes; (3) act inside the authority tier and no further.
Opening. 23:45 on a Tuesday. EDR alerts are stacking from a 220-seat manufacturing client — credential dumping followed by lateral movement — and the host count is climbing. Nobody at the client is awake. You are alone.
| # | T+ | Inject | Decision it forces | What the facilitator listens for |
|---|---|---|---|---|
| 1 | 0 min | Four hosts alerting, then six, then nine. | Remediate hosts as they alert, or establish shape first? | Naming piecemeal containment as the wrong instinct — it tips off an adversary who still holds credentials and other footholds. |
| 2 | 10 min | The same file hash appears in a second, unrelated client's telemetry. | Client incident or portfolio incident? | Immediate escalation to portfolio. The one thing two unrelated clients share is you. |
| 3 | 20 min | Your cross-tenant hunt returns results for 62 of 180 tenants. The rest returned nothing. | Is "nothing" good news? | Distinguishing "no data" from "no result" — license gaps, missing role assignments and row caps all return silence. |
| 4 | 35 min | The client's authority matrix says endpoint isolation is pre-authorized; server isolation is not. Two servers are alerting. | Act, or wake someone? | Acting on the pre-authorized rows immediately, and starting the approver call tree for the rest — in parallel, not in sequence. |
| 5 | 50 min | An account manager asks whether the other 179 clients should be warned tonight. | Disclose, and how? | A blocking indicator with no client identity attached is fine. Naming the affected client to other clients is not. |
Success criteria. Portfolio escalation declared within 20 minutes of inject 2. Pre-authorized containment started before the call tree completes. No client's identity disclosed to another. A written log exists with timestamps.
#D.2 TTX-02 — 02:00, the DC needs isolating and the client is unreachable
Exercises: Part 1, Scenario Two · Appendix A · Appendix B. Participants: On-Call Technician, MSP Incident Commander, Practice Owner. Objectives: (1) exhaust the authority you already have before reaching for the authority you do not; (2) escalate above the technician; (3) produce a contemporaneous defensibility record while acting.
Opening. 02:00. Pre-encryption behavior on a client's domain controller — shadow copies deleting, backup jobs failing. You hold Domain Admin. You can network-isolate that DC in ninety seconds. The MSA says you provide "monitoring and management services" and is silent on emergency action.
| # | T+ | Inject | Decision it forces | What the facilitator listens for |
|---|---|---|---|---|
| 1 | 0 min | IT director's mobile goes to voicemail. Deputy field in the matrix is blank. | Proceed, or keep dialling? | Escalate to MSP Incident Commander and Practice Owner. This decision is above the technician's pay grade, and saying so protects them. |
| 2 | 8 min | CEO's mobile also voicemail. No out-of-band channel recorded. | Declare unreachable? | Whether "unreachable" has an agreed definition and clock. If not, that is the finding. |
| 3 | 15 min | Practice Owner asks: what is the least-disruptive effective action? | Isolate the DC, shut it down, disable accounts, or block egress? | Walking the proportionality ladder aloud, and naming which options destroy evidence and which tip off the adversary. |
| 4 | 20 min | Practice Owner authorises network isolation. | What do you record, right now? | Observed facts establishing imminent harm; every contact attempt with timestamps; who authorized; why this was least-disruptive; notification sent immediately after. |
| 5 | 06:40 | The IT director calls back, furious that authentication was down for four hours. | Who takes this call, and what is the first sentence? | The Practice Owner or MSP Incident Commander takes it, not the technician. First sentence states what was observed and what was done, in that order. |
| 6 | +2 days | Their insurer asks for your contemporaneous record of the decision. | Can you produce it? | Whether inject 4's record actually exists in a form someone else can read. |
Success criteria. Escalation above the technician within 10 minutes. A completed defensibility record. A named authoriser. Notification sent within the contractual window.
#D.3 TTX-03 — your RMM is compromised
Exercises: Part 3, PB-MSP-PLATFORM. Participants: everyone. This is the one where the whole firm sits in the room. Objectives: (1) kill mass-deployment capability before anything else; (2) communicate when your own PSA may be untrustworthy; (3) preserve console logs before they are gone.
Opening. 08:10. Three clients report identical ransom notes within eleven minutes. All three are managed through the same RMM instance. None of them share anything else.
| # | T+ | Inject | Decision it forces | What the facilitator listens for |
|---|---|---|---|---|
| 1 | 0 min | A technician says they cannot log into the RMM console. | Outage, or indicator? | Recognizing that inability to log in is itself an indicator — the ScreenConnect authentication bypass overwrote the user database and deleted all existing local users (Huntress). |
| 2 | 6 min | Someone proposes disconnecting all agents immediately. | Kill deployment first, or disconnect everything? | Killing the mass-deployment capability first, then deciding on agents — disconnecting agents also destroys your visibility. |
| 3 | 15 min | Console web logs have a two-hour hole. | What does the hole mean? | That console-side log deletion is an answer, not an obstacle. Kaseya's attackers deleted IIS and application-database logs as a first action (Truesec). If your logs have a hole at the right time, you are patient zero until proven otherwise. |
| 4 | 25 min | [no-answer] A client's counsel asks whether their data was accessed before the ransomware ran. Your console logs for that window are the ones that are gone. | Answer, hedge, or say you cannot know? | "We do not know, we cannot determine it from the available evidence, and here is what we are doing to narrow it." Anything more confident is a fabrication with your signature on it. |
| 5 | 40 min | You need to notify 180 clients and your PSA may be compromised. | What channel? | A pre-built out-of-band contact list, off-platform, hardcopy or in a separate service. AA22-131A asks for hardcopy plans accessible when networks are compromised (CISA AA22-131A). |
| 6 | 90 min | A journalist emails the Practice Owner. | Who speaks, and what is holding? | One named spokesperson, a holding statement that was written in peacetime, and no speculation about cause. |
Success criteria. Deployment capability disabled inside 15 minutes. Console logs preserved off-platform before remediation. Notification started via a channel that does not depend on the compromised platform. Nobody guessed at inject 4.
#D.4 TTX-04 — a technician account is compromised
Exercises: Part 3, PB-MSP-TECH · Part 2 technician identity. Participants: MSP Incident Commander, service delivery lead, Practice Owner, HR representative. Objectives: (1) scope what one identity touched across every tenant; (2) get the revocation order right; (3) handle the human dimension with dignity.
Opening. 14:20. An impossible-travel alert fires on a senior engineer's admin identity. They are on annual leave and their phone is off.
| # | T+ | Inject | Decision it forces | What the facilitator listens for |
|---|---|---|---|---|
| 1 | 0 min | Someone starts with a password reset. | Reset or revoke first? | Revoke sessions and refresh tokens first, then reset the password. Reversing the order leaves a valid token in the attacker's hands for the remainder of its lifetime. |
| 2 | 10 min | Which tenants did this identity reach? | Where do you look? | GDAP relationships and group membership, Lighthouse authorizations, RMM scoping, backup console rights. Four separate places, none of which shows the others. |
| 3 | 20 min | A privileged change appears in a client's audit log with no corresponding partner sign-in. | Attacker, or your own automation? | Knowing that partner access via PowerShell produces no customer-side sign-in record — only the resulting modifications appear (AADInternals) — so this is consistent with programmatic partner access, legitimate or not. |
| 4 | 35 min | The on-call only holds Help Desk administrator in the affected client. | Can they respond? | Recognizing that resetting a privileged admin's password or invalidating their refresh tokens needs Privileged Authentication Administrator (GDAP least-privileged roles by task). If that is a 14:20 discovery, it is an onboarding failure. |
| 5 | 50 min | The engineer calls in, upset, and asks whether they are being accused of something. | Who talks to them, and how? | A named person, a clear statement that the account is the subject and not the person, and no speculation in writing anywhere. |
| 6 | 70 min | Nine clients were reachable by that identity. | Do all nine get told? | Yes, separately, each with client-specific content. Not one broadcast email. |
Success criteria. Correct revocation order stated without prompting. All four access planes enumerated. Nine separate notifications drafted. The human conversation assigned to a named person.
#D.5 TTX-05 — the auditor's evidence request
Exercises: Part 1, Scenario Three · Appendix C. Participants: MSP Incident Commander, evidence custodian, service delivery lead, Practice Owner. Objectives: (1) produce a segregated per-tenant pack; (2) know your export caps before you hit them; (3) write the gaps file yourself.
Opening. Six weeks after an incident you handled well. The client's auditor, their cyber insurer and their outside counsel have all sent document requests. The requests overlap and none of them are identical.
| # | T+ | Inject | Decision it forces | What the facilitator listens for |
|---|---|---|---|---|
| 1 | 0 min | Can you answer all three with the same PDF? | One pack or three? | No. One evidence pack, three different cover letters and scopes. The auditor wants control operation, the insurer wants loss quantification, counsel wants a timeline. |
| 2 | 10 min | The request asks for technician access logs into their tenant for the six months before the incident. | Can you produce them? | Whether GDAP activity, Lighthouse audit, RMM session and PSA ticket exports are being taken monthly to your own store — or whether you are about to discover their retention windows the hard way. |
| 3 | 20 min | [no-answer] They want Entra sign-in logs from four months ago. Your diagnostic settings were configured two months ago and retention is not retroactive. | What do you say? | "That data does not exist and cannot be recovered." Then the date the export started and why. Retention changes are never retroactive (Entra diagnostic settings). |
| 4 | 30 min | A cross-tenant hunting export from the incident is in the evidence folder. | Is that admissible in this pack? | It is a commingled export by construction. It must be re-run scoped to this tenant, or excluded and explained. Commingling is its own reportable incident. |
| 5 | 45 min | Counsel asks whether your internal Teams messages about the incident are privileged. | Answer? | The correct default is that nothing you wrote is privileged. Say so, and get your own counsel before answering the next question. |
| 6 | 60 min | The client's counsel begins asking whether the MSP's patching contributed. | What changes? | The conflict moment. Your interests and the client's have separated. Stop, tell the Practice Owner, engage your own counsel and your carrier. |
Success criteria. A manifest with hashes and scope filters. A gaps file volunteered, not extracted. Zero other-client data in the pack. The conflict moment recognized at inject 6, not two weeks later.
#D.6 TTX-06 — a vendor zero-day in your remote-access tool
Exercises: Part 3, PB-MSP-VENDOR · Part 4. Participants: MSP Incident Commander, service delivery lead, Practice Owner, communications lead. Objectives: (1) decide patch-versus-disconnect in the first hour; (2) tell clients before they read it; (3) verify the fix actually fixed it.
Opening. 16:50 on a Friday. Your remote-access vendor publishes an advisory: pre-authentication remote code execution, CVSS 9.8, self-hosted instances affected, active exploitation observed. Your instance is self-hosted.
| # | T+ | Inject | Decision it forces | What the facilitator listens for |
|---|---|---|---|---|
| 1 | 0 min | Patch now, or take the server off the internet now? | Sequence. | Isolate from the internet or stop the service first, then upgrade. That is precisely the sequence CISA gave RMM operators for SimpleHelp (CISA AA25-163A). |
| 2 | 20 min | You patch. Is that remediation? | Patched equals safe? | No. Patching is not remediation if you were already compromised — the server needs forensic examination regardless. |
| 3 | 40 min | Do clients need telling, and what do you say? | Notify now with partial information, or wait for certainty? | Notify. CISA's MSP advisories carry an explicit "contact downstream customers" step, and the expectation is notification of suspected as well as confirmed events on provider infrastructure (CISA AA22-131A). |
| 4 | 3 days | The vendor issues a second advisory: the first hotfix was incomplete. | Do you notice? | Whether anyone re-checks vendor advisories after patching. This has happened — N-able's first N-central hotfix required a follow-up (N-able security update). A standing 72-hour re-check is the control. |
| 5 | 5 days | A client asks for written confirmation that their data was not accessed. | What can you honestly write? | What you looked at, what it showed, what you could not see, and the date your logging began. Never a bare assurance. |
Success criteria. Internet exposure removed within the first hour. Forensic examination scheduled independently of patching. Client notification sent the same day. A calendar entry for the 72-hour advisory re-check.
Actionable takeaway: Run TTX-01 and TTX-03 this quarter. They are ninety minutes each and they will find more gaps than a penetration test costing forty times as much. Then put every finding in a list with an owner and a due date, because an exercise without an improvement plan is a very expensive way to feel briefly worried.
#Appendix E — The first ninety days
For an MSP starting from nothing. Not a maturity model — a sequence, ordered by dependency, with the cost-nothing path named at every step.
The ordering matters and it is not arbitrary. You cannot run a credible cross-tenant sweep before you have a log baseline that tells you which tenants can even answer. You cannot rehearse an evidence pack before retention exists to fill it. You cannot do just-in-time elevation before technicians have separate admin identities to elevate. Doing these in the wrong order produces a lot of activity and no capability, which is the most expensive outcome available.
#E.1 Week one — the free things
Every item here costs time and nothing else. If you do only this section, you have removed the three failure modes that recur in every real MSP incident in Part 1's research: internet-exposed management planes, unenforced MFA on consoles, and nobody knowing who to call.
| # | Action | Who | Done when | Cost |
|---|---|---|---|---|
| 1 | Inventory every management plane: RMM, PSA, remote access, backup console, secrets vault, MFT, partner tenant. One row each: internet-reachable Y/N, MFA enforced Y/N, current build Y/N, who administers it. | Service delivery lead | The list exists and is complete | Free |
| 2 | Enforce MFA on every console account, including vendor support accounts inside the product. Disable those unless actively needed. | Service delivery lead | Zero accounts without MFA | Free |
| 3 | Get management interfaces off the public internet — VPN or IP allow-list. | Senior engineer | External scan shows nothing listening | Free (firewall rules) |
| 4 | Review EDR exclusions for RMM install paths. CISA names this as a known blind spot: RMM install paths are often excluded from EDR inspection (CISA Guide to Securing Remote Access Software). | Senior engineer | Exclusions documented and justified, or removed | Free |
| 5 | List every client with a signed MSA. Yes or no. No qualifiers. | Practice Owner | The list exists | Free |
| 6 | Enumerate every GDAP relationship and strip Privileged Role Administrator and Privileged Authentication Administrator from any that does not need them. | Service delivery lead | Export taken and trimmed | Free |
| 7 | Build the after-hours call tree — your people and, for your ten largest clients, theirs. Print it. | MSP Incident Commander | Hardcopy exists at the on-call desk | Free |
| 8 | Pick two names: who declares an incident, and who may authorize action at a client when the client cannot be reached. | Practice Owner | Written down and told to everyone | Free |
Item 8 is the one that matters most and takes ten minutes. Set against the median time to attempted Active Directory compromise cited in this book's opening, an escalation path that cannot reach a decision-maker inside a few hours has already chosen the outcome.
#E.2 Weeks two to four — visibility and the matrix
Depends on week one, because you cannot centralize logs you have not inventoried.
| # | Action | Depends on | No-budget path |
|---|---|---|---|
| 9 | Ship RMM and PSA audit logs off-platform, to somewhere the console itself cannot delete. Set six-month retention as the floor. | 1 | Scheduled CSV export to immutable object storage. Cheap storage beats no storage. |
| 10 | Per-tenant evidence baseline record: which log sources are on, at what retention, since when, at what license tier. | 1 | A spreadsheet. This is what later converts a gap into an explained gap. |
| 11 | Turn on the tenant-side logging you are entitled to at each client's existing license, and record the date you turned it on. | 10 | Free within licences you already pay for. Retention is never retroactive, so today is the cheapest day. |
| 12 | Author the Client Authority Matrix template and complete it for five clients — the five who would hurt most. | 5, 8 | Free. Appendix A is the template. |
| 13 | Script and mass-deployment approval as a two-person control, with alerting on any mass-deploy action. | 1 | Most consoles support scoped permissions natively. The CISA guidance offers a concrete threshold: re-trigger MFA when an account pushes commands to ten or more devices in an hour (CISA Guide). |
| 14 | Put the per-client notification deadline, notice address and approver names into the PSA as structured fields. | 5, 12 | Free. Custom fields. |
#E.3 Weeks five to eight — identity and isolation
Depends on weeks two to four, because scoping access requires knowing what access exists.
| # | Action | Depends on | No-budget path |
|---|---|---|---|
| 15 | Separate admin identities from daily-driver identities for everyone who touches a client tenant. | 6 | Free. New accounts, phishing-resistant MFA, no mailbox. |
| 16 | Restructure GDAP into layered relationships: standing read/triage, standing operational, and short-duration privileged requested per incident. | 6, 15 | Free. Note that all roles in one relationship share one expiry, which is why this has to be multiple relationships (GDAP FAQ). |
| 17 | Scope every technician to the clients they support, in RMM and in the identity plane. No "All Devices" by default. | 15 | Free. It is a permissions exercise, not a purchase. |
| 18 | Per-client backup credentials, per-client repository, per-client key. Immutability at the storage layer, not as a console setting. | 1 | Partly free. One shared backup service account across the estate converts one client's incident into a portfolio event. |
| 19 | Partner-tenant break-glass: two or more cloud-only emergency accounts, phishing-resistant, excluded from blocking Conditional Access, alerting on every sign-in, validated every 90 days (Microsoft emergency access accounts). | 15 | Free apart from two hardware keys. |
| 20 | Complete the authority matrix for the rest of the book, at renewal conversations. | 12 | Free, and it sells. Clients who are being asked NIS2 and DORA questions by their own customers understand exactly why you are asking. |
#E.4 Weeks nine to twelve — prove it works
| # | Action | Depends on | No-budget path |
|---|---|---|---|
| 21 | Write and pre-compute the cross-tenant sweep: the query, the tenant groups, the row budget per tenant. | 10, 11 | Free. Discovering at 02:40 that 100 tenants against a 50,000-row cap means 500 rows each is not a discovery you want to make live. |
| 22 | Produce a full per-tenant evidence pack for one real client, in peacetime, on the clock. | 10, 11 | Free, and it is the highest-yield day in this whole plan. You will find every export cap, missing hash and truncated CSV the cheap way. |
| 23 | Run TTX-01 and TTX-03 from Appendix D. Log findings with owners and due dates. | 7, 8, 12 | Free. Two afternoons. |
| 24 | Test the out-of-hours contact path for your top twenty clients using the procedure in A.5. | 12 | Free. Four minutes per client. |
| 25 | Contract review: for every client, answer in writing whether there is a signed MSA, whether it grants any emergency authority, whether it disclaims agency, what the notification clock is, whether a DPA or BAA exists and what number it carries, and whether your RBAC in that tenant matches the authority tier they agreed to. | 5, 12, 16 | Free to audit. Counsel costs money to fix — which is why you take them one prioritized list rather than 180 contracts. |
| 26 | Insurance: confirm your own cover, confirm the notice conditions, and confirm whether your carrier requires panel counsel or panel forensics. Do this before an incident, because engaging your own firm can be a consent event. | — | Free. One phone call to your broker. |
#E.5 The dependency picture, in one paragraph
Inventory before hardening, because you cannot secure what you have not listed. Logs before hunting, because a sweep across tenants with no telemetry returns silence that looks like safety. Identity separation before just-in-time elevation, because you cannot time-box an identity that also reads email. Authority matrix before tabletops, because an exercise that discovers you have no approver list teaches you one thing and then stops teaching. And evidence rehearsal last, because it tests everything upstream of it — which is exactly why it is the item most likely to get postponed, and exactly why it should not be.
Item 25 is the one that will surface the uncomfortable number: how many clients you serve without a signed agreement, or with one that gives you access and no authority. That number is almost never zero, at any MSP, of any size. Finding it out is not an indictment; it is Tuesday. Fixing it at renewal, one client at a time, with a lawyer who acts for you, is the whole job.
Stay patched, stay documented, and get the phone numbers tested before you need them — because the worst possible time to learn that a mobile number changed in March is at 02:00 in November.
#Sources
- NIST SP 800-61r3, Incident Response Recommendations and Considerations for Cybersecurity Risk Management — https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-61r3.pdf
- NIST SP 800-84, Guide to Test, Training, and Exercise Programs for IT Plans and Capabilities — https://nvlpubs.nist.gov/nistpubs/legacy/sp/nistspecialpublication800-84.pdf
- CISA AA22-131A, Protecting Against Cyber Threats to Managed Service Providers and their Customers — https://www.cisa.gov/news-events/cybersecurity-advisories/aa22-131a
- CISA AA25-163A, Ransomware Actors Exploit Unpatched SimpleHelp RMM — https://www.cisa.gov/news-events/cybersecurity-advisories/aa25-163a
- CISA/NSA/FBI/MS-ISAC/INCD, Guide to Securing Remote Access Software — https://www.cisa.gov/sites/default/files/2023-06/guide_to_securing_remote_access_software_final_508c_v3.pdf
- CISA Tabletop Exercise Packages (CTEP) — https://www.cisa.gov/resources-tools/services/cisa-tabletop-exercise-packages
- Sophos MDR Service Description — https://www.sophos.com/en-us/legal/mdr-description
- Microsoft Defender Experts for XDR — https://learn.microsoft.com/en-us/defender-xdr/managed-detection-and-response-xdr
- Google Workspace terms (Emergency Security Issue) — https://workspace.google.com/terms/2013/1/premier_terms.html
- CompassMSP Master Service Agreement — https://compassmsp.com/legal/master-service-agreement
- Secure Data Technologies Managed Services Agreement — https://www.securedatatech.com/wp-content/uploads/2025/06/Managed-Services-Agreement.pdf
- Regulation (EU) 2022/2554 (DORA) — https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CELEX:32022R2554
- GDPR Article 28 — https://gdpr-info.eu/art-28-gdpr/
- EDPB Guidelines 9/2022 on personal data breach notification, v2.0 — https://www.edpb.europa.eu/system/files/2023-04/edpb_guidelines_202209_personal_data_breach_notification_v2.0_en.pdf
- 45 CFR 164.410 — https://www.law.cornell.edu/cfr/text/45/164.410
- Hunton, cyber insurer subrogation suit against technology vendors — https://www.hunton.com/hunton-insurance-recovery-blog/policyholder-plot-twist-cyber-insurer-sues-policyholders-cyber-pros
- National Law Review, same matter — https://natlawreview.com/article/cyber-insurer-sues-policyholders-cyber-pros
- Clifford Chance, ICO penalty against a processor — https://www.cliffordchance.com/insights/resources/blogs/talking-tech/en/articles/2025/04/ico-fines-processor-after-inadequate-security-measures-lead-to-widespread-disruption.html
- Huntress, ScreenConnect authentication bypass analysis — https://www.huntress.com/blog/a-catastrophe-for-control-understanding-the-screenconnect-authentication-bypass
- Truesec, Kaseya supply-chain attack analysis — https://www.truesec.com/hub/blog/kaseya-supply-chain-attack-targeting-msps-to-deliver-revil-ransomware
- AADInternals, Microsoft partners: The Good, The Bad, or The Ugly? — https://aadinternals.com/post/partners/
- Microsoft Learn, GDAP least-privileged roles by task — https://learn.microsoft.com/en-us/partner-center/customers/gdap-least-privileged-roles-by-task
- Microsoft Learn, GDAP FAQ — https://learn.microsoft.com/en-us/partner-center/customers/gdap-faq
- Microsoft Learn, Azure Lighthouse cross-tenant management experiences — https://learn.microsoft.com/en-us/azure/lighthouse/concepts/cross-tenant-management-experience
- Microsoft Learn, manage emergency access admin accounts — https://learn.microsoft.com/en-us/entra/identity/role-based-access-control/security-emergency-access
- Microsoft Learn, configure Entra diagnostic settings — https://learn.microsoft.com/en-us/entra/identity/monitoring-health/howto-configure-diagnostic-settings
- N-able, N-central security update — https://www.n-able.com/blog/n-central-security-update-august-10-2026
#The MSP Readiness Checklist
Every control from every part of this book, in one place, ordered by implementation tier rather than by chapter. Work down the tiers, not across the parts — an MSP with every IG1 control in place is in materially better shape than one that has done half of Part 2 to IG3 and never tested a client contact tree.
| Tier | Who it is for |
|---|---|
IG1 | Essential. Every MSP needs this, whatever its size. |
IG2 | MSPs with someone whose actual job is security. |
IG3 | MSPs serving regulated clients, or facing adversaries who will spend real money. |
Readiness
Saved on this device#Tier 1 — Essential (30 controls)
- MSP-01A Client Authority Matrix exists for every managed client, naming a primary Client Authority Holder and a deputy with out-of-hours contact details.
- MSP-02The matrix marks each disruptive action (isolate endpoint / server / DC, disable user, tenant-wide credential reset, block egress, suspend service, restore from backup, engage DFIR) as pre-authorized, approval-required or prohibited, with a named approving role.
- MSP-03Each client's out-of-hours contact path has been dialled and verified within the last 180 days, with the date and tester recorded.
- MSP-06Every managed client has a signed master agreement on file; an exception register lists any client without one, with a remediation date and owner.
- MSP-10The notification deadline, legal notice address and named recipient for each client are structured PSA fields readable by the on-call technician, not a PDF in a document store.
- MSP-14Answers given on a client's cyber insurance application on the client's behalf are provided in writing, dated, with supporting evidence attached, and retained.
- MSP-15Administrative identities holding delegated access to client tenants are separate from those used for the MSP's own email and productivity work, with no shared or generic management accounts. [CC6.3]
- MSP-16Phishing-resistant MFA is enforced on every identity that can reach a client tenant or a management console, with no exception group; exceptions are documented, time-bound and approved by the Practice Owner.
- MSP-21Technician offboarding removes delegated access in a documented order and produces evidence: token revocation, delegated-admin group removal, Lighthouse and JIT removal, console account disablement, credential rotation, object-ID map update and break-glass revalidation. [CC6.8]
- MSP-23No management console (RMM, PSA, remote access, managed file transfer, backup) exposes an administrative interface to the public internet; access is restricted by VPN or IP allow-list.
- MSP-24Management-plane products are patched on a KEV-class SLA measured in hours, and every advisory acted on is re-checked 72 hours after patching for a superseding or incomplete fix.
- MSP-30Backups are held with per-client credentials, repositories and encryption keys, with immutability enforced at the storage layer beyond the reach of backup-console administrators, and restore testing evidenced within the last 90 days. [A1.x]
- MSP-31Every client record carries the contractual notification deadline in hours, the contractual notice method, the legal notice address and a named recipient with an out-of-hours number, readable by the on-call technician in one screen without opening a contract.
- MSP-32Every client record flags the regulatory regimes attaching to that client (GDPR controller, HIPAA covered entity, PCI, NYDFS, DFARS flow-down, public company, NIS2, DORA) and the categories of data held on their behalf.
- MSP-33A written standing decision states that client notification is not delayed pending risk assessment, root-cause determination or confirmation of data impact, and names the role authorized to send.
- MSP-34Four pre-approved notification templates exist and are reviewed at least annually: initial client notification, scheduled update, multi-client holding statement, and MSP-is-the-source notification.
- MSP-35Notifications are issued to the contractual notice address by the contractual method, and delivery evidence (timestamp, recipient, transmission record) is retained with the incident file.
- MSP-36A contemporaneous log records, in UTC/ISO 8601, the discovery time, every notification attempt, every contact reached or missed, and who authorized each action — maintained on a system unaffected by the incident.
- MSP-37Every notification states the time of the next update, and updates are sent on schedule even when there is nothing new; a missed update is itself recorded as an incident-process exception.
- MSP-39Every client file has been reviewed against a checklist confirming whether the MSA, DPA and BAA specify notification timing, method and notice address in both directions; each gap has a named owner (Practice Owner) and a remediation date tied to renewal.
- MSP-42An offline copy of the client notification contact list, and a communications path that does not traverse the RMM or PSA, are maintained and verified at least quarterly.
- MSP-43The MSP's own cyber and Technology E&O carriers, claims lines, panel breach counsel and panel DFIR firms are recorded in the incident plan, and carrier notice sits on the incident checklist at the same tier as client notification.
- MSP-46A written scope determination exists for each regime — NIS2 Annex I sector 9 (with the Art. 2(1) size test and Art. 2(2) exceptions considered), UK RMSP definition, DORA counterparty status, CIRCIA proposed criteria, CMMC ESP/CSP status — reviewed annually and after any material change in headcount, turnover or client mix.
- MSP-47Where the MSP has an EU main establishment, the NIS2 Art. 27 registration has been submitted (entity name, sector, establishment addresses, current contacts, Member States served, IP ranges) and the receipt retained; where the MSP is non-EU and serves EU clients, a Union representative has been designated in writing.
- MSP-48A per-client regulatory register exists as structured PSA fields — contractual notification deadline in hours, legal notice address and method, named out-of-hours recipients, regime flags (GDPR / NIS2 / DORA / BAA / PCI / DFARS 7012 / NYDFS / public company), data categories held, and whether the MSP may notify a regulator on the client's behalf.
- MSP-49One notification capability is built to the tightest clock the MSP faces (24 hours), with a pre-drafted T+0 pack per client, so only the recipient and form vary by regime.
- MSP-50Client notifications are issued individually per affected client with that client's own facts, never as a single broadcast, and the template covers the content a controller needs for GDPR Art. 33(3).
- MSP-51A written procedure states that the MSP notifies affected clients on establishing that a breach occurred, without first assessing likelihood of risk, and names the role authorized to send that notification out of hours.
- MSP-53For every business associate relationship the BAA breach-reporting timeframe is recorded in the tenant record alongside the after-hours contacts, and the alert queue is reviewed on a defined daily cadence with the review evidenced, because HIPAA discovery is imputed on reasonable diligence.
- MSP-58Where a DFARS 252.204-7012 flow-down applies, the 72-hour reporting path to DoD and to the prime is documented with the report-number handover step, 90-day preservation of system images and monitoring data is a standing rule, and reimaging before imaging is prohibited by default with a named role able to authorize an exception.
#Tier 2 — Dedicated security capability (28 controls)
- MSP-04The matrix defines an unreachable-escalation window in minutes, agreed in writing per client, after which specified actions become pre-authorized, with a post-hoc notification clock.
- MSP-05The authority tier agreed with each client is enforced in RBAC — identity roles, EDR role assignments and RMM device scopes — not only in a document.
- MSP-07The master agreement contains an emergency action clause with a defined trigger, a minimum-extent-and-duration limiter, and a post-hoc notice obligation on a fixed clock.
- MSP-08The master agreement contains a security suspension or disconnection right distinct from the non-payment suspension right.
- MSP-09The master agreement contains an evidence-preservation obligation and an express legal-hold carve-out to the data-deletion and data-return clause.
- MSP-11The complementary user entity controls in the MSP's SOC 2 report have been reconciled against client-responsibility clauses in each master agreement within the last 12 months. [CC7.x]
- MSP-12Every client contract requires the client to carry cyber liability insurance and to provide a waiver of subrogation in the MSP's favor, and the waiver has been checked against the client's own policy wording, since some policies restrict or override contractual waivers; exceptions are recorded and approved by the Practice Owner.
- MSP-13A data processing addendum is in place for every client whose processing falls under GDPR or UK GDPR, reciting the pre-authorized containment actions as documented instructions of the controller under Article 28(3)(a) and naming the MSP's sub-processors; for clients outside those regimes, the equivalent processing terms required by the applicable state or national law.
- MSP-17Administrative work on client tenants is performed only from a hardened, dedicated workstation with mail and general browsing restricted.
- MSP-18Standing delegated access is read/triage only; write and privileged roles are obtained just-in-time through a documented activation path staffed whenever the on-call rota is live.
- MSP-19Every delegated-admin relationship has been reviewed within the last 90 days, and Privileged Role Administrator and Privileged Authentication Administrator removed from any relationship that does not require them.
- MSP-20At least two emergency access accounts exist in the MSP management plane — cloud-only, phishing-resistant, tied to no individual, excluded from lockout-capable conditional access — with alerting on every sign-in and validation at least every 90 days and after any IT staff change.
- MSP-22A quarterly access review is filed as an artifact — group membership, delegated-admin relationships and Lighthouse delegations diffed against current staff and client lists, with a reviewer name and date. [CC6.3]
- MSP-25A script and mass-deployment approval workflow is in force, and an alert fires on any action pushing commands, scripts or installers to ten or more devices within an hour.
- MSP-26RMM and PSA audit logs are shipped to a store the RMM and PSA cannot delete or alter, with at least six months of retention and documented export segmentation where row caps apply.
- MSP-27A pre-written cross-tenant hunting query library exists in version control, has been executed end-to-end within the last 90 days, and records which tenants returned no matching records and which could not be queried, with the reason.
- MSP-28A written "is it us?" self-check procedure exists, is runnable by one technician in under 30 minutes, is held in hard copy in the on-call bag, and has been drilled within the last 12 months.
- MSP-29A per-tenant evidence baseline record states, for each log source, whether it is enabled, its destination, its retention, its license tier, the date it started and the date last verified.
- MSP-38A written content standard prohibits cause speculation, attribution, fault admission, blame characterization, legal-exposure commentary and unverified record counts in any client-facing communication, and a second named person reviews every notification before it is sent.
- MSP-40A multi-client notification runs from a tranche plan ordered by contractual deadline and then by regulatory exposure, never by account value, and the ordering rationale is recorded.
- MSP-41Each affected client receives a notification containing that client's own facts; no single broadcast communication is relied on to discharge a per-client notification duty, and no notification discloses another client's identity, sector or status.
- MSP-44A standing rule prohibits engaging outside counsel, forensic firms or negotiators — for the MSP or on a client's behalf — before carrier notice and panel confirmation, with the sole exception of containment actions, which never wait.
- MSP-52The sub-processor list in every client DPA reconciles to the actual delivery-path tooling inventory (RMM, PSA, backup, EDR cloud, ticketing, offshore NOC), with a working mechanism for notifying intended additions and allowing objection; reconciliation is performed at least quarterly.
- MSP-54A PCI DSS responsibility matrix exists per service (and per client where the split differs), is producible on customer request under Requirement 12.9.2, and states explicitly which party notifies the acquirer and payment brands on suspected compromise.
- MSP-55A versioned DORA data pack is maintained per financial-entity client — legal entity identifier, country of head office, countries of service provision and data processing, service type, subcontracting chain, critical-or-important-function flag — refreshed before the annual register-of-information cycle rather than during it.
- MSP-56Agreements with financial-entity clients have been reviewed against DORA Art. 30(2) and, where the service supports a critical or important function, Art. 30(3): audit and inspection rights extending to the competent authority, incident assistance at no additional or pre-agreed cost, a documented exit plan with transition period, and TLPT cooperation.
- MSP-57For each Defense Industrial Base client a written determination records whether the MSP handles CUI, handles Security Protection Data only, or operates a cloud service holding CUI — and, where it acts as a CSP, states the FedRAMP Moderate authorization or DoD-recognized equivalency position with its supporting body of evidence.
- MSP-59The management body has formally approved the cybersecurity risk-management measures and completed the required training, with the approval and training recorded in dated minutes retained as evidence.
#Tier 3 — Regulated and high-threat (2 controls)
- MSP-45The notification chain is rehearsed at least annually against the largest realistic fan-out, measuring elapsed time from simulated discovery to the last client notification sent and to the first scheduled update, with the result reported to the Practice Owner.
- MSP-60The MSP holds at least one third-party attestation or certification whose scope statement demonstrably covers the service lines it sells, and maintains a single evidence set — report with period covered, bridge letter, scope statement and Statement of Applicability, CSOC list, CUEC list, subservice organization list and monitoring evidence — ready to send without bespoke work.