Managed services put one organization's administrators inside another organization's estate, and this chapter tells both sides of that arrangement — the provider holding the keys and the customer who handed them over — what to build, what to demand, and what each is entitled to prove.
Who needs this: MSP and MSSP owners, service delivery leads, on-call technicians · CISOs, IT Directors and Heads of Procurement who buy managed services · Legal Liaison, Incident Commander, Executive Sponsor | Read time: 28 min | Maps to: CSF 2.0 GOVERN (GV.SC, GV.RR), IDENTIFY (ID.AM), PROTECT (PR.AA), DETECT (DE.CM), RESPOND (RS.CO, RS.MA) | CIS Controls 6, 15 | ISO/IEC 27001:2022 A.5.19–A.5.23, A.5.24, A.5.28, A.8.15
Welcome, fellow defenders. This is the chapter about the people who can log into your network without asking, and the chapter about being those people. Depending on your day job you are about to read one of two books. Read both.
On the Friday afternoon before the 2021 US Independence Day weekend, a supermarket chain in Sweden could not open its tills. Kindergartens in New Zealand lost their systems. Public administration offices in Romania went dark (The Record). None had bought a vulnerable product, misconfigured a firewall, or clicked anything. For a great many, the first genuinely useful fact anyone could offer was the name of a piece of software they had never purchased and could not have identified in a line-up.
A REvil affiliate had burned zero-days against Kaseya VSA servers that were on-premises and reachable from the internet — a remote monitoring and management product whose customer base is, overwhelmingly, managed service providers. As of 5 July the reported reach was fewer than 60 direct clients and no more than 1,500 businesses supported by those clients, against a US$70 million demand for a universal decryptor (NCSC/ODNI factsheet). Hold that arithmetic: roughly twenty-five downstream businesses per provider, none of whom had done anything wrong or could have seen it coming. Their exposure came from a decision that is sensible on every ordinary metric — hire a competent firm to run the IT so you can run the business.
Now turn it around and you are one of the fewer than sixty. Those firms spent the holiday weekend discovering that their disaster-recovery planning had quietly assumed a healthy management plane — and there was nothing left to recover from, and nothing left to recover with. Reconstruction of the chain puts an authentication bypass in the web interface first, then a session the attacker was treated as having earned, then an uploaded payload, then execution by injected code (Tenable, on Huntress's analysis) — and the first housekeeping act was wiping the web-server logs and the copies held in the application's database (Truesec). Note what did not happen: nothing in VSA misbehaved. An operator holding a session the product considered valid asked it to distribute a file to every agent it knew about, which is the single capability the product exists to sell.
That inversion — one incident, two entirely different experiences of it, joined by a standing administrative relationship — is the subject here. There is a 75,000-word companion volume to this book, The Blast Radius: Incident Response for Managed Service Providers, written entirely to the provider. This chapter is the gateway to it, and the half it does not contain: what the customer should demand, verify and write into the agreement.
Three questions. Answer them out loud.
Yes to 1 or 2 makes you a customer, and the second half of every section below is yours. Yes to 3 makes you a provider, and the first half is yours with the annex behind it.
Plenty of organizations answer yes to both and never notice — the accountancy practice that buys managed IT and also holds admin credentials in its clients' finance systems, the software vendor that outsources its help desk and also runs a platform its customers cannot administer without it. If that is you, everything below applies twice, in opposite directions.
Actionable takeaway: write down, this week, every organization that can administer something of yours and every organization you can administer. One list, two columns. Most of the work in this chapter begins with somebody being surprised by the second column.
Chapter 11 owns third-party and supply chain risk properly — the vendor register, tiering by access rather than spend, due diligence, SBOM and SLSA, OAuth grants, contract clauses, fourth parties, concentration risk. All of it applies here and none of it is repeated; read this chapter as the one tier of that program where the counterparty can log in and change things.
What makes managed services distinct is one structural fact. An ordinary vendor is something you integrate with; a managed service provider is something that administers you. The vendor holds an interface and some data. The provider holds a domain admin account, a delegated relationship carrying tenant-wide roles, an agent running as SYSTEM on every endpoint, and often the console governing your backups. That changes five things at once, each landing differently depending on which chair you sit in.
| The asymmetry | What it costs the provider | What it costs the customer |
|---|---|---|
| Administrative control is standing, not requested | One compromise resolves to every client at once, delivered by tooling behaving as designed. No client is unaffected while you work | Your exposure is set by the weakest control in a company you do not run. Your own MFA, patching and segmentation do not reach it |
| Capability and permission are separate, and only one was granted | You can isolate the domain controller. Whether you may is a contract question, at 02:00, with a lawyer attached | You granted the capability at onboarding and never granted or withheld the permission |
| Triage is portfolio-shaped, and the provider may be patient zero | One indicator at one client is a question about the whole book, and your own estate is the first thing to rule out | Your incident may be a symptom of someone else's, and only the provider can run the query that would show it |
| The record lives in two estates on two retention clocks | Much of what an auditor, insurer or opposing counsel wants after a client incident is about your operations | The logs proving who did what inside your tenant cannot name a human; the translation table sits in the provider's tenant |
| Notification is a relay and you do not run the first leg | Your contractual clock starts on your discovery and is routinely tighter than any statute your client works to | Your regulatory clock does not start until your provider speaks |
One number sets the tempo for all five. Across 661 incident response and MDR cases, median time to attempted Active Directory compromise is 3.40 hours, up 70% year over year, in a population where 84% of affected organizations had fewer than 1,000 employees — and 88% of ransomware encryption happens outside business hours (Sophos 2026 Active Adversary Report; Help Net Security). If the path from "the alert fires" to "somebody who can decide is awake" runs longer than that, the timing question is settled before the night begins and nobody on it gets a vote — true of the provider's escalation ladder and of the customer's phone tree alike.
Actionable takeaway: make managed services its own tier in your third-party program, with its own evidence demands and its own clauses. Grading your MSP on the questionnaire you send your payroll vendor is a category error with an outage attached.
One note before any of this. What follows is an operating framework assembled from published agreements, regulator guidance and statutory text. It is not legal advice, I am not a lawyer, and nothing here substitutes for counsel who has read your actual agreements. Where clause content is described it is illustrative — the shape of a clause and the job it has to do, never words to sign.
This is the most important idea in the annex and it compresses to one sentence: holding the credential answers "can I", and nothing else.
NIST puts the requirement on the contract plainly. Third-party responsibilities should be defined in the agreement, including "authority to act on behalf of the organization" and "restrictions on what the service provider can do, such as ... making and implementing operational decisions (e.g., immediately deactivating certain services to contain an incident)" (NIST SP 800-61r3). Real agreements rarely do either. A current, publicly posted master service agreement from a mid-market US provider grants unilateral powers only to change services on notice and to suspend for non-payment: no emergency authority clause, no incident response clause, no timeline for telling the client about a breach. Most look like that.
The exposure runs both ways. Act without authority and the sharp US edge is 18 U.S.C. § 1030(a)(5)(A), which is not an access provision: it reaches whoever transmits a command and thereby "intentionally causes damage without authorization", damage being "any impairment to the integrity or availability of data, a program, a system, or information", with a private civil action once loss reaches $5,000 in a year (18 U.S.C. § 1030). Notice what (a)(5)(A) does not ask. It does not ask whether you were let in. It asks whether the impairment you caused was authorized, and a credential — however legitimately issued, however routinely used — is silent on that. To my knowledge the provision has never been turned on a provider acting inside an incident, so this is a reading of the text rather than a rule anybody has tested, and an untested reading is thin ground to stand on while a progress bar runs. Customers should notice the corollary: if nobody ever wrote down which damage was authorized, neither side has an answer to give. Fail to act and the claims are already on dockets: in ACE American Insurance Co. v. Congruity 360, LLC and Trustwave Holdings, Inc., No. 2:25-cv-15657 (D.N.J., filed 15 September 2025), a cyber insurer sued its insured's technology vendors in subrogation, and the pleaded negligence against one is failure to properly notify appropriate parties (Hunton). None of that is proven; a complaint is argument, not a finding. What is worth reading is the shape of the argument. The insurer did not plead that the vendor missed the intrusion or fumbled the containment. It pleaded silence. For a buyer, that reframes what a notification clause is actually for — it is not paperwork, it is the allegation you are pre-empting.
The annex resolves this with a per-client Client Authority Matrix. Compressed, it answers three questions for one relationship: what may be done without asking, what must be asked about, and whose phone rings. A workable default set of tiers: pre-authorized for reversible single-asset work — isolate one endpoint, kill a process, quarantine a file, block a specific address, disable one compromised account and revoke its sessions; approval required from a named customer authority holder for anything server-level or tenant-wide — isolating a server, hypervisor host or domain controller, mass credential reset or session revocation, and engaging third-party forensics or notifying a regulator on the customer's behalf; executive approval for taking a production service offline or restoring over live data; and prohibited, always, for paying or negotiating a ransom, which is never the provider's decision to make.
A published commercial model already does this well: Sophos grades authority from Notify Only through Collaborate to Authorize, and adds "Collaborate then Authorize," which authorises action "in the event Sophos does not receive acknowledgment from Customer/MSP after making reasonable attempts to contact all Customer defined contacts" (Sophos) — bounded by the limiter a customer's counsel will want, action "to the minimum extent and of the minimum duration required" (Google Workspace terms).
Everything above was written for the provider. Now flip it, because this artefact belongs to you at least as much as to them. You granted a third party the ability to disconnect your domain controller and have almost certainly never told them whether they may. That silence is not neutrality; it is a decision, made by default, that whoever is awake at 02:00 chooses.
Several landmark incidents share one shape. A management server belonging to the provider was reachable from the internet, was authentication-bypassed, and was then used to do exactly what it is built to do: push a script, push an installer, take control of a session. Nothing exotic crosses the wire until the last step, because the product's own distribution feature carries the payload.
The identity plane has its own version. Two of the roles Microsoft handed the Admin Agents group by default at the GDAP cutover are worth knowing by name (Microsoft). One is Privileged Authentication Administrator: it resets any user's authentication factors, Global Administrators included. The other is Privileged Role Administrator: it hands out role assignments, which means it can hand itself whatever it is missing. Chain them and a partner reaches the whole tenant in two moves. Nobody chose that configuration — the platform wrote it during the migration, which is exactly why so few of those relationships have ever been read. This is Chapter 4's access review pointed outward. The review that matters is not of your own administrator groups but of the standing roles another company holds into your tenant, and everything Chapter 4 argues about privileged access management and just-in-time elevation applies to a partner's roles exactly as it does to your own — with the difference that you cannot run the elevation workflow, so you have to buy it. Microsoft's term for the strategy exploiting it is "compromise-one-to-compromise-many", documented across intrusions spanning four providers to reach one final target (Microsoft Security Blog). Backups are no safer: CVE-2024-42448 in Veeam Service Provider Console (CVSS 9.9) allows remote code execution on the console server from an authorized management agent machine (Veeam KB4679).
Provider side, the standard is short and mostly free. Get every management interface off the public internet behind a VPN or IP allow-list, which is what CISA and the FBI told providers after Kaseya (CISA/FBI). Enforce phishing-resistant MFA on every console account, vendor support accounts included. Ship console audit logs somewhere the console cannot delete them. Patch on the KEV-driven clock Chapter 10 sets, and put an internet-facing management console at the top of that queue rather than in it — Chapter 10's prioritization assumes a vulnerability reaches one estate, and this one reaches every customer you have. Then re-read the advisory 72 hours later. Scope technicians to the customers they support. And add the one mass-deployment brake in the public guidance that was written specifically for this business model: "if an account attempts to push commands to 10 or more devices within an hour, retrigger security protocols, such as multifactor authentication (MFA), to ensure the source is legitimate" (CISA, *Guide to Securing Remote Access Software*). The same guide names the blind spot underneath it all: RMM install paths are often excluded from EDR inspection, which makes your own exclusion list the adversary's quiet room.
Customer side, this is why "do you have MFA?" is the wrong question. It is a yes/no with a yes attached, and you cannot verify it: when a technician signs into your tenant through a delegated relationship, MFA is required and evaluated in their home tenant and trusted in yours, and cross-tenant MFA trust settings are not applied to granular delegated admin sign-ins at all (Microsoft). Your provider's authentication posture is silently your control, and your only lever is the contract. Ask for artefacts instead — Chapter 11's diligence pack is the container, and these are the six things you can only ask of a company that administers you: an export of every identity that can reach your environment, with the delegated relationships and roles held and when they expire; the factor type and policy proving phishing-resistant MFA on those identities rather than on staff email; confirmation that no management console is reachable from the public internet; the patch SLA in hours for remote-access and RMM products, with the re-check step after the fix; what fires when one of their accounts pushes a script to ten or more of your devices; and a straight answer on whether any administrative credential, repository or backup key is reused across customers, which the joint advisory tells them not to do.
Then do the four things on your own side of the boundary that need nobody's cooperation: keep provider accounts out of your internal administrator groups; scope them to the systems the provider actually manages; audit that they serve the purpose they were created for; and disable them when the contract ends, which the joint advisory notes is commonly overlooked at termination (CISA AA22-131A).
When an indicator lands, an internal team asks how far it has spread inside one estate. A provider must answer two further questions first, and no single-organization playbook contains either.
Is it only them? Two customers in different sectors, different countries, different payroll systems, doing business with none of the same people — and the same attacker-created account in both. Correlation that thin has exactly one plausible bridge, and the provider is standing on it. Treat that as an answer, not a curiosity. The sweep answering it must be pre-built, because the platforms cap it in ways nobody wants to discover mid-incident: one view of Defender's multi-tenant hunting reaches at most 100 tenants; the result ceiling is 50,000 rows shared out across however many tenants you selected, which is 500 apiece at the maximum; and it sees back 30 days natively (Microsoft). A tightly filtered hash query is fine. A broad process sweep truncates silently and hands back a comforting, wrong answer.
Is it us? The annex makes this a standing, timed procedure one technician can run in under thirty minutes across delivery channel, job history, console integrity, technician identity, the delegated-access path and shared credentials. Two steps are counter-intuitive enough to repeat. A gap in the management console's own logs at exactly the interesting moment is a finding, not a glitch — log deletion was the opening move at Kaseya. And privileged changes in a customer tenant attributable to a partner identity with no corresponding interactive sign-in are consistent with programmatic partner-side access, because a partner working through PowerShell leaves the customer's sign-in log empty while the portal and the API both write to it (AADInternals). That is either your own automation or somebody holding your partner credentials, and only your records can tell which.
Import the annex's rule whole: any credible suspicion that the provider's own platform is involved is the highest severity until disproven — not until confirmed. Highest severity means SEV-1 as Chapter 13 defines it, declared here on suspicion rather than on confirmation, which is the one place a provider departs from that scheme. The downside of calling it too loudly is one wasted night. The downside of calling it too quietly is the whole book at once — and by the time anybody reconstructs the sequence, the evidence has been curated by the intruder.
Customer side, this makes a genuinely useful procurement question: how long does it take you to check every customer you serve for one indicator, and when did you last do it? Ask for the number and the date. Some providers have a query library, a tested fan-out and a documented list of tenants the sweep cannot reach, and will answer in minutes with caveats. Some will say they would "have a look" — which means the sweep is a research project run under pressure by whoever happens to be awake, and that is a real answer worth hearing honestly. It is also the gap to fix together at the next service review, not a reason to leave. The joint advisory expects a provider to raise the alarm with its customers on events touching provider infrastructure that are suspected as well as confirmed (AA22-131A). Suspected. Read that word twice.
Actionable takeaway for providers: this quarter, pick an indicator that means nothing, start from a blank console, and time yourself to a complete answer. Whatever the stopwatch says is your true portfolio triage speed. Put it in the incident plan, and — genuinely — in the sales deck.
Six weeks of quiet after an incident closes is not the same as safety. What follows is three separate readings of the same event — the compliance one, the money one, and the fault one. None of them arrives as an accusation and none of them needs to, because all three are settled by paper, and a great deal of the relevant paper describes the provider rather than the customer.
Here is the turn that surprises providers and that customers should understand before they need it. Much of what gets requested is not evidence about the customer's environment. It is evidence about the provider's operations. Technician access records. Delegation history. Patch reports. Alerting configuration and its change history — "was the detection that would have caught this ever enabled, and did somebody switch it off?" The service-desk timeline, which is usually the only surviving proof of the moment a person actually opened the alert. The on-call and escalation record. In Clorox v. Cognizant, a $380 million claim filed on 22 July 2025, what sat at the heart of the pleading was chat logs from the provider's own support queue, in which agents were said to have reset credentials and multi-factor enrolments for callers whose identity nobody had established (The Register). The discovery target was the provider's own service-desk call handling — ticket notes, call recordings, verification steps — not forensic artefacts. Help-desk identity verification is Chapter 4's control, and that pleading is what it looks like when it is missing. The difference worth sitting with is that the help desk in question was not the customer's. They could not watch the verification step being skipped, could not audit it, and could only ever have contracted for it.
Two structural facts make this harder than it looks. First, the customer-side log physically cannot name the provider's technician: through a delegated relationship, sign-in and audit entries render with a display name of the form "{Governing tenant name} Technician" and a username of user_ plus an object ID with the dashes removed, by design (Microsoft). The translation to a human being with employment dates exists in one place only — the provider's own tenant — and no support ticket recovers it once a technician has left. Second, everything expires on schedules neither party chose: directory sign-in and audit history survives somewhere between a week and a month depending on what the customer pays for, and turning on log streaming buys you nothing that happened before you turned it on (Entra); endpoint telemetry reaches back 30 days in advanced hunting and is deleted no later than 180 days from contract termination, unrecoverable (Defender XDR). Retention is tied to the contract, and the contract is the first thing an angry customer ends — so the deletion timer and the litigation risk start on the same day, running in opposite directions. Export at offboarding, not after it. Chapter 9 argues retention as a detection question — can you still see the thing that happened. In this relationship it is also an evidence question, and the answer is set by a license tier chosen in somebody else's tenant, which is why it belongs in the agreement rather than in the logging strategy.
Provider side. Build one pack and produce subsets to each requester under a cover note saying what is in and what is not. Chapter 13's chain of custody supplies most of what each artefact has to answer on its own — where it came from and under what filter, who pulled it by what mechanism at what UTC timestamp in ISO 8601, what window it genuinely covers, and a hash showing it has not moved since. Add the field a single-organization responder never needs: which tenant, and under which delegated relationship. Nine fields, filled at export, never afterwards. Tenants come apart by default, which means every commingling failure is something the provider actively built: a cross-tenant query written into a single CSV has already mixed the estates, and a pack assembled from it puts one customer's data in another customer's hands — a notifiable breach of its own, arriving on top of the one you were documenting. Then write the gaps file yourself, first rather than last: each source that was disabled, aged out, rate-limited or cut short, why, and the earliest date it does cover. "These logs begin on this date because the setting was created then, under this license" is an explained limitation. The same hole discovered by the other side, with nothing beside it, invites the least flattering reading available — and the people finding it are paid to find the least flattering reading.
Customer side — the half written from your chair. You are entitled to this material, and the moment to establish that is contract signature rather than week six of a claim. The joint international advisory tells you what to ask for: contracts should require the provider to "provide visibility — as specified in the contractual arrangement — to customers of logging activities, including provider's presence, activities, and connections", and organizations should store their most important logs for at least six months (CISA AA22-131A). NCSC-UK adds that you should verify the provider maintains security logs, understand the retention periods, and confirm your own access rights to logging data during an incident (NCSC-UK). Four clauses for your counsel: a logging and visibility obligation naming what the provider records about its own presence in your estate, at what retention, and your right to receive it within a stated number of business days, during and after an incident and after termination; the object-ID-to-technician mapping exported monthly; a legal-hold carve-out to the deletion clause, so nobody must choose between breaching the contract and destroying the record in the week a preservation notice lands; and evidence preservation on written notice both ways, including a prohibition on reimaging before imaging — restoring you quickly is your provider's instinct, training and contract, and in the wrong order it destroys your evidence and their defense in one command.
Actionable takeaway for both: rehearse the production in peacetime, against a stopwatch. A drill turns up the same three discoveries almost every time: an export limit nobody had ever hit before, a configuration whose author has since left and whose start date nobody can now evidence, and a single point of human failure standing between you and the archive. All three are trivial in a rehearsal and disqualifying in a production request.
Drawn properly, most of the confusion evaporates.
PROVIDER LANE CUSTOMER LANE
─────────────────────────────── ───────────────────────────────
Discovers. Alert, ticket, vendor
advisory, or a customer ringing in.
│
Classifies. Whose systems, whose
data, which agreements, how many
customers are inside it.
│
Sends the notice. ──────────────► "Nobody has told us yet" ends here.
Clock: contractual, and it has AWARENESS TRANSFERS ON THIS ARROW.
been running since discovery. Every downstream clock starts now.
│
Decides materiality. Counsel, DPO,
board, insurer.
│
Files. Regulators, its own customers,
its insurer.
Every clock in the right-hand lane belongs to the customer, and Chapter 15 works them properly — the notification decision tree, the materiality call, every regulatory clock with its trigger — with the matrix in Appendix C as the lookup. Do not run this section as your regulatory reference. What belongs here is the one thing all of those clocks have in common: not one of them starts until the provider speaks. NYDFS §500.17(a) is the rule that says so on its face, running its 72 hours from determining that an incident occurred at the covered entity, its affiliates or a third-party service provider (23 NYCRR 500.17).
The left-hand lane carries statutory duties as well as contractual ones, and HIPAA is the one providers most often misplace. As a business associate the provider has no later than 60 calendar days from its own discovery to notify the covered entity, where discovery includes the first day the breach would have been known by exercising reasonable diligence (45 CFR 164.410) — so an unreviewed alert queue is a running clock, and it is running on the provider's side of the arrow, not the customer's. The number actually negotiated into a BAA is routinely far shorter than 60 days, and whatever it is, it eats into the covered entity's own 60.
Provider side. The instinct to be certain before worrying anybody is a good instinct almost everywhere else in this job. Here it spends money that is not yours: every hour of confirming comes out of the customer's 72. Your binding number is usually neither figure — it is whatever the master agreement says, commonly "immediately", four hours or twenty-four, breached long before any statutory deadline is missed, and not uniform across your book. Engineer to the harshest deadline you carry, once. One notification pipeline, built to fire inside a day of the moment you first knew, running somewhere your own estate going dark cannot silence it. After that the work is clerical — swap the addressee, swap the template, send.
Customer side — take this sentence to your next renewal: your provider's notification SLA is your regulatory exposure. If the deadline is not in the contract in hours, you do not have one; you have a phrase somebody else gets to interpret years afterwards, and nobody named to receive the notice, which lets an argument start about whether you were properly notified at all. Fix four things: a period in hours banded by severity, with a legal notice address and a named operational recipient with an out-of-hours number, which is what both AA22-131A and NCSC-UK tell buyers to demand; a written standing instruction that notification is not delayed pending root cause, risk assessment or confirmation of data impact; a duty to notify you of confirmed or suspected events on the provider's own infrastructure, not only inside your tenant; and a scheduled update cadence hit even when nothing has changed. If you are a public company, add one line: the SEC declined to exempt registrants from disclosing incidents on third-party systems they use, while noting the rules "generally do not require that registrants conduct additional inquiries outside of their regular channels of communication with third-party service providers pursuant to those contracts" (Release 33-11216). The Commission is telling registrants they may rely on the contractual channel — which makes that channel part of your disclosure controls, and its weaknesses yours to explain. Name a person, not a function. Give the clause an out-of-hours number. A notice route that only works during office hours is a control that only works when nothing is happening.
For fifteen years the compliance conversation ran one way: the customer had the regulator, the provider had the tooling, and the provider's job was to help the customer pass. That is coming apart, and the direction of travel is consistent even where the detail is still moving.
EU NIS2 does not sweep providers in by implication; it lists them, in Annex I, sector 9, "ICT service management (business-to-business)", with Art. 6(39) reaching services related to the installation, management, operation or maintenance of ICT products, networks, infrastructure and applications, "either on customers' premises or remotely" (Art. 6). Holding customer data is not required. Size is the gate, though: scope attaches at medium-sized enterprise under Art. 2(1), so a provider under 50 staff with turnover and balance sheet at or below €10m is out on size — unless Art. 2(2) pulls it back in, which it can where the provider is the sole or critical operator for services that matter at national or regional level (Art. 2). A twenty-five-person shop running the IT for four regional hospitals is a live candidate, and that call belongs to the Member State, not to you. Providers that are in mostly land as important entities under Art. 3(2) (Art. 3) and report on the Art. 23 ladder: early warning within 24 hours of awareness of a significant incident, notification within 72 hours of awareness, final report within one month of that notification (Art. 23). Art. 27 adds a standing pre-incident duty to register with the competent authority, dated 17 January 2025 and most often discovered mid-incident (Art. 27); transposition remains incomplete, with four Member States referred to the Court of Justice on 8 July 2026 (ECSO tracker). The customer-side consequence explains your questionnaire inbox: Art. 21 requires in-scope entities to hold supply chain measures covering each direct supplier and to weigh that supplier's specific vulnerabilities and the quality of its practices (Art. 21). Diligencing your provider, and evidencing it, is your own compliance artefact now.
The UK Cyber Security and Resilience Bill invents a statutory category, Relevant Managed Service Provider — "a person who provides managed services in the UK (whether or not the person is established in the UK)" that is not a small or micro enterprise — with duties to register, secure and notify significant incidents to the Information Commission, on a 24-hour light-touch notice to the regulator sighting the NCSC, measured from becoming aware that an incident is occurring, a fuller notification at 72 hours from the same trigger, and a duty to tell affected customers after the regulator notification has gone in (RMSP factsheet; incident reporting). The stated reason is unusually direct: providers "have unprecedented access to their customers' systems, making them an attractive target that cyber actors increasingly exploit." Status matters more than content — it is in Parliament and it is not law, Lords Second Reading was on 14 July 2026 (Parliament), and the provider duties arrive through secondary legislation after an implementation consultation (gov.uk).
EU DORA does not regulate an ordinary regional provider. It rewrites the master agreement instead, and makes a financial-services customer legally unable to sign the 2019 version. Art. 30(2) binds every ICT service contract, critical or not, and two of its requirements end long-standing channel practice: incident assistance at no additional cost or at a pre-determined cost, which closes the familiar carve-out treating real incident work as a separate chargeable project, precisely when a provider would reach for it; and audit and inspection rights extending to the financial entity, its appointed third parties and the competent authority (Regulation (EU) 2022/2554). An audit-rights carve-out is not a negotiating position; it puts the customer out of compliance, so the deal stalls. Art. 28(3)'s annual register of information arrives at a service desk as a February spreadsheet request from every financial customer at once (CSSF).
CMMC is a live commercial decision for providers serving defense customers. Phase 2 begins 10 November 2026, adding Level 2 third-party certification for applicable contracts (32 CFR 170.3; DoD CIO). The trap is believing that not touching controlled unclassified information puts you out of scope. An External Service Provider is defined without any requirement to handle CUI, and Security Protection Data expressly includes configuration data, log files, vulnerability status and passwords (32 CFR 170.4) — which puts an RMM, a password vault and a SIEM tenancy inside every defense customer's assessment whether or not anyone sends the provider the report. A provider actually hosting customer CUI acts as a cloud service provider under DFARS 252.204-7012 and must meet FedRAMP Moderate or DoD-recognized equivalency (DoD CIO memorandum). Nothing about that decision fits in a corridor — it is a capital question with a compliance deadline attached.
Direct enforcement against providers already exists. In March 2025 the ICO issued its first penalty against a data processor — £3,076,320 against an IT supplier, after attackers entered through a customer account protected by nothing but a username and password, with data taken on 79,404 people. Nothing in the findings involved sophistication — incomplete MFA rollout, thin vulnerability scanning, patching that had fallen behind. What made it worse was that a functioning MFA capability had been developed and then left on the shelf, on the assumption that customers would object to it (ICO; Clifford Chance). Somewhere in a mailbox there is a thread where a customer said MFA was too much friction and somebody at the provider agreed to leave them out of the rollout. In the ICO's findings, a thread like that is not context. It is the finding. Go and read yours this week, not after the next assessment. Two more make the point elsewhere. A business associate reached an $80,000 settlement with HHS OCR, plus three years of corrective action, over an intrusion that came in through firewall ports left open and ran six days before anyone spotted it (Nixon Peabody). And one of the FTC's charges against a hosting provider was simply that it did not adequately log and monitor security events (FTC) — the logging gap standing as the violation on its own, regardless of what it did or did not cause.
Actionable takeaway for providers: hold the attestation, and hold it at the right size — a certificate scoped to the corporate network is not a certificate covering the RMM. Publish the evidence set at a stable URL so nobody is rebuilding it deal by deal. Then reread the scope paragraph the way an adversary would, because the gap between what you sell and what the boundary covers is never discovered by you. Actionable takeaway for customers: ask which regimes name your provider directly, whether they have registered where registration is required, and which attestation covers the service you actually buy. Then ask for three things alongside the report: the period it genuinely covers, a bridge letter for the months since that period ended, and the complementary user entity controls list — the controls the provider's own report says you are responsible for. Read that last one slowly. It routinely assigns you something you assumed they ran. None of this scales with headcount. It scales with whether one person has ever blocked out two days and written the answers down.
The annex is built around three nights. Compressed here, they show a provider what the volume contains and a customer what a competent provider has already thought through. If any of them prompts "we could not do that tonight", that is exactly the diagnostic they exist for.
You have met their single-organization cousins already. The first night is Chapter 14's account takeover and identity provider playbooks arriving together; the second is the opening of its ransomware playbook, read before anyone has typed the word. The annex's PB-MSP- playbooks sit alongside those fourteen rather than replacing them, and the distinction is the whole point of this chapter: Chapter 14 is written for an organization defending itself, and these are written for a provider defending forty organizations at once.
Credential dumping on a file server. The same on an application server. Remote service creation to a third host. Then three workstations in one subnet, all inside four minutes, the count still climbing while the on-call technician reads. Nobody at the customer is awake and the IT manager's voicemail is full.
Working the console top to bottom is what the interface is built to encourage, and it loses the incident twice over. It is documented to fail: remove the compromised systems you know about and the responders "tip their hand" to the attacker, who abandons the burned tooling, keeps the access nobody found, and goes quiet until an outside party tells the organization it is still compromised (Aldridge, Black Hat USA 2012; CISA IR Playbooks). It announces at machine speed that somebody is awake, which buys the intruder a reason to drop the tooling you can see and keep the access you cannot. And it arrives late regardless: the first alert in the stack was credential theft, and no amount of isolation puts those back.
The two questions belonging to a provider and nobody else are what this night turns on. Is it only them — the pre-written cross-tenant sweep, its row budget computed in advance and its unreachable tenants recorded by name. And is it us — delivery channel, unexplained job pushes, console account integrity, technician sign-ins outside the rota, privileged changes with no matching interactive sign-in. Both answered before anyone is woken, because the answer decides who commands the incident and whether the RMM is an instrument or a suspect. A suspect is not a collection tool. CISA tells a provider in that position to take the management server off the network and warn the customers downstream of it (AA25-163A) — and a server you have just disconnected cannot also be your instrument. Then act on the whole footprint at once — every confirmed host and identity inside a few minutes of each other, with collection already queued — rather than trickling containment across an hour and a half and teaching the intruder your tempo. Preserve, then contain.
Three alerts, forty seconds apart. Volume shadow copies gone. The nightly backup did not time out — it was refused credentials, which is the version of that alert to be afraid of. And a service account signing in from a desktop that has never in its life had cause to speak to a domain controller. That tradecraft is documented — deletion of Volume Shadow Copy Service copies alongside tooling built to extract credentials from the backup platform (CISA AA24-109A). Somebody is removing the fire escapes before lighting the match.
The technician has every capability required and no permission they can point at. The ladder, in miniature, is four rungs climbed in parallel. Exhaust the authority you already hold — isolate the source workstation, revoke the compromised account's sessions and then disable it, block the identified egress, collect the investigation package, and establish whether an immutable pre-intrusion restore point exists, because that single fact changes the cost of losing the machine. The ordering rule in the middle of that list is Chapter 4's, and it matters more here than anywhere else in the book, because the identity you are racing may hold roles in several tenants and every minute of a live token spends them all. Whether the restore point is a fact you can confirm in four minutes or a hope you test at 03:00 was settled months ago by the immutability and restore-testing work in Chapter 12. Escalate inside your own organization, because a named person other than the one holding the mouse must say yes. Work the contact tree and log every attempt as you make it, in UTC, because the pleaded claim in the live subrogation case is failure to notify. Then take the narrowest containment that actually works.
Whatever is decided, the action and the record are one step rather than two. The annex calls the artefact a defensibility record: ten fields filled at the time, in UTC, on a system outside the affected environment — the facts establishing imminent harm, why waiting was not viable, every contact attempt as its own row, who declared the customer unreachable, who at the provider authorized the action, the exact action with its console action ID, why it was the least disruptive option that would have worked, the notification sent afterwards, the evidence preserved, the holds placed. One of those fields has no permissible empty state: the name of the person who authorized it. An empty box there is not a missing field, it is an answer. And the difference between an emergency action and an unauthorized one is largely chronological — a record made while the console is still open reads as a decision, and the same record typed after breakfast reads as a justification.
Customer half of this night: all of it goes faster when you have named a deputy who answers the phone. And you will know inside ninety seconds whether the call you get afterwards was rehearsed. A rehearsed one opens with the state of your business rather than with reassurance. It reads times off a record instead of recalling them. It names the person who said yes and does not apologize for them. And it volunteers the ceiling — the things they deliberately stopped short of, and why. If you have to ask for that last part, ask; the answer tells you whether anybody was watching the boundary while you were asleep.
The incident was handled well. Somebody sent doughnuts. Then a Monday arrives with three emails: the auditor's twenty-two-item document request, the insurer's forensic accountant wanting the discovery timestamp and a service-by-service recovery timeline, and a preservation notice from a law firm nobody has heard of naming two of the provider's technicians as custodians.
The generous instinct is to send all three the tidy incident report written in week two. It is the worst move available. That report is a persuasive document produced by an interested party, which is precisely what none of these three wanted. Auditors and insurers both need the underlying thing — a control record, a cost line — and a narrative summary hands over neither. Worse, the report carries a theory of cause in it and a byline on it, and in the one proceeding where that matters, both are now the other side's property. Produce artefacts instead, scoped per requester, from one deliberately assembled pack. And note what nobody in the room finds obvious: opposing counsel does not need the provider's consent to reach its data, because control rather than location decides discoverability (Sedona Conference). Nobody signs a data-access clause thinking about discovery, and it becomes one anyway — your right to obtain their records is exactly the hook that makes those records reachable through you. Customer half: whatever arrives in your inbox in week six was determined by what you insisted on at signature. The intervening work is administration.
This chapter is where the argument lives. The Blast Radius is where the procedures live — a standalone companion written entirely from the provider's chair, which points at this book rather than repeating it.
| Part | What is in it |
|---|---|
| Why the annex exists | The seven asymmetries in full, the MSP-side incident command roles, and blast radius as a second severity dimension alongside SEV-1 to SEV-4 |
| Part 1 — The Three Nights | The three scenarios above at full length: clock-stamped phases, step tables with evidence columns, decisions with decide-by times and named authorities |
| Part 2 — Before the Night | The ten things that must already be true — authority matrix, clauses, tenant isolation architecture, technician identity and delegated-access least privilege, RMM and PSA hardening, cross-tenant hunting, the self-check, the per-tenant logging baseline, provider-managed backups, escalation design |
| Part 3 — The MSP-side playbooks | PB-MSP-PLATFORM (your RMM or PSA is compromised), PB-MSP-CROSS (cross-tenant lateral movement), PB-MSP-GDAP (delegated admin relationship abuse), PB-MSP-TECH (technician account compromise), PB-MSP-VENDOR (your upstream vendor is compromised) |
| Part 4 — The notification chain | The clocks, the content standard, four notification templates, and how to notify forty customers at once without breaching anyone's confidence |
| Part 5 — When the provider is the regulated entity | NIS2, the UK Bill, DORA, GDPR's processor layer, CMMC, and the rest of the map at reference speed |
| Part 6 — Templates and tools | The Client Authority Matrix, blank and worked; an MSA and DPA clause library; a per-tenant evidence pack manifest; six tabletop scenarios; and a first-ninety-days plan that opens with the things that cost nothing |
| The MSP Readiness Checklist | Sixty MSP-xx controls, tiered IG1 to IG3 |
Those sixty controls are a separate, deeper set from the PROV codes below. PROV is this chapter's own — a joint standard for the relationship, roughly half of it addressed to the customer. Green on PROV is where a serious relationship starts; green on MSP-xx is what a provider owes its whole book.
If you take one thing into work tomorrow, take the smallest one. Providers: choose the account you would least like to explain to a journalist. Before the week is out, its record should carry a person who can authorize an outage, somebody who can do it when the first one is unreachable, and the number of minutes you are permitted to wait. Then dial the number and see who answers. Customers: send your provider three questions — how long does it take you to sweep every customer for one indicator, what may you do to us at 02:00 without asking, and may I have the list of your people who can reach my tenant. Not next quarter. This week.
Know exactly who can turn your keys, write down what you have permitted them to unlock, and check the phone number before the night that needs it.
Domain code PROV. Every item carries an audience tag so it can be filtered on sight. These are the relationship-level controls; the sixty MSP-xx controls in The Blast Radius are the provider's deeper operating standard and are not duplicated here. PROV assembles into Appendix A alongside every other chapter's controls, where the audience tag is what lets a buyer and a provider run the same list from opposite sides of the same door.
[IG1] [GV.SC] [ID.AM] [CIS 15] [Both][IG1] [GV.RR-02] [A.5.24] [Both][IG1] [GV.RR-02] [Customer][IG2] [Both][IG1] [Both][IG2] [PR.AA-05] [Both][IG1] [PR.AA-05] [CIS 6] [Customer][IG1] [PR.AA-01] [A.5.20] [Customer][IG2] [PR.AA-05] [Customer][IG1] [GV.SC] [Both][IG1] [Provider][IG1] [PR.AA-03] [Provider][IG2] [DE.CM-09] [Provider][IG1] [PR.AA-01] [RC.RP-01] [Provider][IG2] [DE.CM-01] [Provider][IG2] [A.5.24] [Provider][IG2] [GV.SC] [Customer][IG1] [RS.CO-02] [A.5.20] [Both][IG1] [RS.CO-02] [Both][IG2] [GV.SC] [RS.CO-02] [Both][IG2] [A.5.28] [A.8.15] [Customer][IG2] [A.8.15] [A.5.28] [Provider][IG2] [A.5.28] [Both][IG2] [GV.SC] [Customer][IG3] [A.5.24] [CIS 17] [Both]