In this article:
Want us to find IT vendors for you?
Share your vendor requirements with one of our account managers, then we build a vetted shortlist and arrange introductory calls with each vendor.
Book a call

How to test an MDR provider before you sign: the seven drills that expose the gaps

A standard MDR trial tests the platform, not the service. Seven drills reveal response times, containment authority, and coverage before you sign.

Author:
Date

Three providers are on your shortlist. All three detect the same techniques, cite the same frameworks, and produce reference customers who sound genuinely happy. Each one offers a 30-day trial.

That trial will not tell you what you need to know, and the reason sits in the providers' own contracts. Expel excludes the first 60 days of onboarding from its service level agreement.

Sophos excludes 60 days from purchase for new customers. A standard 30-day pilot therefore runs entirely inside the window where the provider carries no contractual obligation to perform at all.

So you spend a month watching a service that is not yet on the hook, then sign a three-year agreement on the strength of it. The way out is to stop treating the pilot as a demonstration and start running it as a rehearsal.

Why a standard MDR trial tests the platform and not the service

A typical 30-day trial measures four things:

  • How quickly sensors deploy across your estate.
  • How the console is laid out.
  • How many alerts arrive per day.
  • Whether the service catches a commodity test file.

Every serious provider passes all four. Those are properties of software, and the software is mature across this category.

What you are actually contracting for is a group of people, a set of permissions, and a documentation trail. Someone triages your alert at three in the morning. That person reaches a conclusion, decides whether to wake you, and is either permitted to isolate a host or required to ask first. Afterwards, a record either exists or it does not.

None of that appears on a dashboard. Deployment speed tells you nothing about what happens in the forty minutes after a detection fires.

You cannot establish it by asking, either. Every provider answers yes to whether they can contain automatically, and the answer lives in the conditions attached to the yes.

Sophos, for example, publishes three separate authority modes, and in one of them no response action is taken without your written consent. Both the yes and the caveat are accurate.

Confusing the platform for the service has a specific consequence. You sign on software quality and then live with service quality for the length of the term, and by the time the difference becomes visible you are renewing rather than choosing.

What a provider means by response time, and why the number is not comparable

Before you can set a pass threshold for anything, you need to know which interval you are measuring. Providers publish response commitments that look comparable and measure different things.

Expel's contractual clock runs from the moment an alert fires to the moment an analyst acknowledges it and begins triage, at 15 minutes for critical severity and 30 for high, with service credits attached.

Sophos MDR Plus measures from its own identification of a priority investigation to either customer outreach or the initiation of a response action, at 60 minutes for 90% of priority investigations.

Secureworks measures from the point it initiates a case to the point you are notified of its analysis, also at 60 minutes, also credit-backed. Critical Start publishes a 10-minute notification for critical alerts and a 60-minute time to resolve across all priorities.

Two of those are 60 minutes. They are not the same 60 minutes. One starts when the provider decides a case exists, the other starts when a case is initiated for analysis, and neither starts when the event actually happened on your network.

Sophos is explicit that its base-tier figures are targets rather than a service level, and that they may vary by integration type and data availability. That last clause matters more than the number in front of it.

Provider Published clock Where the clock starts Remedy attached
Expel Mean time to triage, 15 min critical / 30 min high Alert fires Service credit
Sophos MDR Plus Time to respond, 60 min for 90% of priority investigations Sophos identifies a priority investigation Service credit
Sophos MDR (base) Case creation and initial action targets Detection ingestion, varies by integration Target only
Secureworks Taegis Threat investigation, 60 min to notification Secureworks initiates the case Service credit
Critical Start 10 min critical notification, 60 min time to resolve Alert classification Service credit
CrowdStrike 1-10-60 benchmark, median time to contain reported Initial detection Benchmark, not a term
SentinelOne Not publicly documented Not publicly documented Not publicly documented
Arctic Wolf 2 hr notification (2023 reseller service guide) Arctic Wolf's discovery of the incident Not publicly documented

There is an awkward pattern in that table. The providers who publish hard, credit-backed definitions are Expel, Sophos, Secureworks and Critical Start. The providers most often reaching a three-way shortlist are the ones where the commitment cannot be read off a public document.

CrowdStrike's 1-10-60 framing is a recommended benchmark rather than a contractual term. SentinelOne publishes no numeric response commitment I could locate. Arctic Wolf's own current figure is not published; the only dated public number I found is a two-hour notification from discovery, in a 2023 reseller service guide, and it conflicts with other figures circulating in third-party summaries.

That is the argument for drilling rather than asking. When the commitment is not documented, generating the condition and timing it yourself is the only way to establish what you are buying.

What to fix before the pilot starts, or it proves nothing

A pilot run on incomplete telemetry measures your gaps rather than their service. Providers document that this is your responsibility, in language that survives into the contract.

Sophos requires its service software on at least 80% of licensed volume, described as necessary for sufficient visibility, and states that unremediated health conditions may reduce service quality or limit its ability to investigate and take response actions.

Secureworks is blunter, warning that customer noncompliance can lead to reduced capability, suspension of managed components or service levels, or a shift to monitor-only.

Rapid7 requires you to supply an asset exclusion list during setup. Your containment boundary is something you define, and if you skip that conversation you inherit a default you never chose.

One technical point worth settling per integration. Identity and cloud control plane events reach these services through connector APIs and log streaming, and no provider in scope publishes ingestion latency or audit API polling intervals. Confirm the path is live and confirm how fresh the data is, because a detection that depends on an identity feed cannot fire faster than that feed arrives.

Readiness item Why it matters How to verify it What the pilot measures instead if it is missing
Asset inventory completeness Coverage percentage is a contractual precondition, not a courtesy Reconcile the agent count against DHCP, DNS and switch port data, not against the asset register Detection on the subset you already knew about
Endpoint and server agent coverage Sophos requires 80% of licensed volume for sufficient visibility Request the deployment percentage against licensed volume in writing A service operating below its own documented visibility floor
Linux and server telemetry Server-side compromise is the scenario most pilots never test Confirm the tier purchased includes Linux, and that agents report Workstation detection quality, extrapolated to servers
Identity event forwarding Credential attacks leave no endpoint signal to detect Generate a benign sign-in anomaly and confirm it appears in the provider console Nothing. The drill returns a false negative you will misread
Cloud control plane audit logs Control plane abuse is audit-log-only activity Confirm the connector is authorised and events are landing, per cloud account Your cloud posture, unmonitored, for the length of the term
Unmanaged devices An uninstrumented asset is invisible by definition Passive network discovery against the agent inventory A coverage gap you will discover during a real incident
Named on-call contacts and escalation tree Providers disclaim delay caused by unreachable contacts Hand over a written tree with out-of-hours numbers and test one Your own phone tree, and an excluded SLA breach
Baseline time to notice Improvement is unprovable without a starting number Record your current detection-to-action interval for 30 days prior An absolute number with nothing to compare it against

If several rows in that table look uncertain, fix the telemetry before you start the clock on a trial. Device logging is the usual culprit, and an endpoint stack audit will surface the rest. Providers know that a pilot on thin telemetry reflects well on them.

The seven drills to run in an MDR pilot

Two rules apply across all seven.

First, on method, point at recognised frameworks rather than techniques. MITRE ATT&CK and its managed services evaluation methodology cover detection coverage. NIST SP 800-115 covers technical test planning. NIST SP 800-84 covers the human-response drills, which are functional exercises rather than technical tests. Use the provider's own supported validation tooling where it exists.

Second, record every drill as four wall-clock timestamps: event generation, detection, notification to you, and containment. You compute the interval. Do not accept the provider's number, because their clock starts somewhere you did not choose.

Drill 1: a commodity detection during business hours

Generate a recognised benign test artefact through a supported validation tool, on a covered workstation, mid-morning.

Tell them nothing beyond the standing testing authorisation already in place.

Observe all four timestamps, the severity assigned, whether a case was created, and which channel the notification arrived on.

Pass requires notification through the agreed channel with a case reference and a named owner.

A weak result here means the notification path itself is broken, and every subsequent drill measures plumbing rather than judgement. Fix it and rerun before continuing.

Drill 2: the same event at 2am on a Sunday

This is the only drill that matters, and almost nobody runs it, because scheduling it is inconvenient for everyone including you. Run it anyway.

Generate the identical artefact from drill one, on a comparable asset, at the least convenient hour available.

Tell them nothing extra. Same authorisation, same channels.

Observe the same four timestamps, plus who responded and where they were.

Pass requires the human interval to hold within a reasonable margin of the business-hours result.

A weak result shows as an alert that fires on time and a human who arrives hours later. That gap is the service you are buying, and it will not improve after signature.

Drill 3: credential-based activity with no malware present

Generate benign anomalous authentication activity through your identity provider, using its own test facilities or a supported validation tool.

Tell them nothing. This drill only works unprompted.

Observe whether a case opens from identity events alone, and how long the identity feed takes to surface.

Pass requires a case raised from identity telemetry with no endpoint signal involved.

A weak result means identity is a slide in the deck rather than a monitored source. Check the connector before blaming the analyst, because if the audit path is not live no tier can detect this.

Drill 4: a containment request

Generate nothing. Ask the provider to isolate a specific host.

Tell them exactly what you want contained, through the normal channel, and record who you spoke to.

Observe elapsed time from request to enforcement, who authorised it, and what the action actually operated on.

Pass requires enforcement inside the agreed window with an audit record naming the authoriser.

A weak result is worth reading carefully. Agent-level network containment is not upstream firewall or NAC quarantine. Session revocation is not a password reset, and Huntress documents the former without the latter. Sophos limits agentless third-party endpoints to isolation and un-isolation only. Critical Start applies a two-person review to every action, which is a design choice rather than a delay.

Drill 5: a false positive spike, and who owns the filtering decision

Whether tuning is billable is not publicly documented by any provider in scope, so ask a better question. Who decides what stops being an alert, and can you see what they filtered?

Generate a benign, repetitive pattern that resembles an alertable condition, through a supported tool.

Tell them nothing until the volume registers.

Observe who notices first, how quickly a suppression is applied, and whether you receive a record of what was suppressed.

Pass requires the provider to notice before you do and to document the change.

A weak result connects to contract language. Sophos filters activity in its reasonable judgment, states no obligation to retroactively re-investigate filtered activity, and excludes data generated by detection content you authored. Establish during the pilot whose detection content governs, and whether you can audit the filter list.

Drill 6: escalation to a named human analyst

Generate nothing. Call the out-of-hours number during an open case from an earlier drill.

Tell them you want the analyst currently working your case.

Observe whether you reach a person with your case in front of them, what they can see, and what they are permitted to decide.

Pass requires a named analyst with case context and stated authority.

A weak result distinguishes staffing models. A named team, as with Arctic Wolf's concierge structure or Rapid7's assigned advisor, behaves differently from an undifferentiated pool. Both are legitimate. Only one of them knows your environment at 3am.

Drill 7: a server-side, Linux or unmanaged asset event

Generate a benign detection on a Linux server or a comparable non-workstation asset, using supported tooling.

Tell them nothing.

Observe whether the asset is covered at all, whether severity handling differs, and whether containment authority extends to it.

Pass requires detection and the same response authority you verified in drill four.

A weak result is usually a coverage gap rather than a detection failure, and that distinction matters. This scenario is established public methodology; MITRE's 2024 managed services round emulated ransomware deployment to Windows and Linux ESXi servers.

If your servers sit outside the purchased tier, the drill has told you something about your order form rather than their analysts. Telemetry differences between endpoint platforms surface here more often than anywhere else.

How to run the drills without voiding your trial terms

Running these unannounced feels appealing and collides with what providers actually document.

Sophos excludes incident response for penetration testing, red team engagements and adversary simulation. Critical Start excludes red-team and pen-test activity from SLA scope during defined windows. An unannounced drill can therefore fall outside the commitment you are trying to measure, which defeats the exercise entirely.

So get written testing permission into the trial agreement before onboarding starts. It should specify the scope, the timing windows, the tooling categories permitted, and confirmation that drill activity remains inside SLA measurement. That last clause is the one providers resist, and their response tells you something either way.

A provider who declines outright is telling you the service performs best unobserved. That is a finding, not an obstacle.

Draw your rules of engagement from published guidance rather than inventing them. SP 800-115 covers scope definition, timing, data handling, escalation paths and the conditions under which testing stops. Add an abort signal and a named coordinator on each side.

Cloud activity needs separate clearance. AWS permits assessment of listed services without prior approval, and requires a simulated events form submitted at least two weeks ahead for red team exercises, simulated phishing and comparable activity.

Microsoft no longer requires notification for testing your own tenant, and states it may interrupt attacks in progress whether or not they form part of a valid test.

Internally, tell whoever would otherwise respond. Your own team, your service desk, and any incumbent provider. Where responsibilities are already split across an MSP or co-managed arrangement, agree in advance who stands down during a drill window, or you will spend a Sunday managing an incident you created.

What to record, and why the record is your negotiation lever

The evidence file is the deliverable. For each drill: four timestamps, the analyst identity, the verbatim notification text, the actions taken, and who authorised them.

Verbatim matters more than it sounds. A notification reading escalated for review and one reading host isolated, awaiting your confirmation describe different services, and paraphrase erases the difference.

The conversion step is where this pays. Because providers anchor their clocks at different points, your own four timestamps let you restate observed performance in whichever definition you want written into the agreement.

You can express the same drill as time to notify, time to contain, or detection to enforcement, and propose the one that reflects what you care about.

Where a provider publishes no numeric commitment, the observed interval becomes the proposed term. That is a materially stronger position than accepting silence and hoping.

Record one more thing: whether an incident record was produced at all, and in what form. Documentation deliverables and after-action cadence are not publicly documented for most providers in scope, which makes the pilot the only place to establish them.

Scoring the pilot

Weight the drills by consequence rather than by convenience.

Drill Weight What a pass requires
Out-of-hours detection 25% Human interval within a reasonable margin of business hours
Containment request 25% Enforcement inside the agreed window, with a named authoriser on record
Server, Linux or unmanaged asset 15% Detection plus the same response authority as a workstation
Credential activity, no malware 15% Case raised from identity telemetry alone
Escalation to a named analyst 10% A reachable human with case context and stated authority
False positive spike 5% Provider notices first and documents the suppression
Business-hours commodity detection 5% Notification path functions end to end

Fold these weights into your existing vendor scorecard rather than scoring MDR in isolation. The general method for structuring the evaluation sits in the vendor proof of concept guide; everything above is the category-specific layer on top of it.

The five pilot results that should stop the deal

Notification arrived and nobody was permitted to act. You have bought monitoring at response pricing. The alert was never the hard part.

Out-of-hours performance materially worse than in-hours. Attacks are timed for exactly that window. A service that performs on Tuesday afternoon is protecting you during the hours you were already covered.

Identity or cloud control plane events invisible. Credential-based intrusion leaves no endpoint signal, so this is a blind spot rather than a limitation. It will not be fixed by tuning.

Containment required a ticket and a business-day wait. Ransomware propagation does not queue. If enforcement depends on a human being at a desk, the containment claim describes a capability rather than a service.

No incident record produced afterwards. Without a record you cannot demonstrate the control to an auditor, support a claim, or reconstruct what happened. You are also unable to hold the provider to anything, because nothing was written down.

Each of these is disqualifying rather than negotiable. They are properties of how the service is built and staffed, and no commercial concession changes them.

Get the shortlist before you run the drills

Filter with your environment, telemetry scope and out-of-hours cover and discover from a catalog of 1200+ curated vendors. No vendor sees your details until you decide to talk. It's private and free.

Find MDR vendors

FAQ

How long should an MDR pilot run?

Longer than the 60-day SLA exclusion window if you can negotiate it, because Expel and Sophos both disclaim service levels during onboarding. A 30-day trial sits entirely inside that period, so treat it as a readiness test and push the response drills as late as possible.

Can I run tests during a trial without permission?

No. Sophos excludes incident response for adversary simulation and penetration testing, and Critical Start excludes red-team activity from SLA scope during defined windows. Testing without written permission can place the drill outside the commitments you are measuring.

Should the pilot run on production?

Production gives the only representative answer, which is why DORA requires threat-led penetration testing on live production systems for in-scope financial entities, alongside controls to manage the risk. Apply the same logic: production, with an abort signal and change control.

Do providers know when they are being tested?

Often, and it matters less than you would think. Written authorisation makes the activity visible anyway, and what you are measuring is elapsed time and authority rather than whether you fooled anyone.

Can I pilot two providers at once?

Yes, and it is the fastest way to normalise for your own environment. Run identical drills against both, on comparable assets, in the same windows. Overlapping agents need coordination first.

What if the provider will not allow testing?

Ask what they will allow, in writing. A provider offering a supported validation path is behaving reasonably. One refusing any observation of response behaviour has answered the question you were asking.

Is 30 days long enough to judge tuning?

For volume, yes. For ownership, yes, because the filtering question is answerable in a week. For steady-state alert quality, no; that emerges over a quarter, which is why the suppression record matters more than the alert count.

What is the difference between an MDR pilot and a purple team exercise?

A purple team exercise validates detection coverage collaboratively against known techniques. An MDR pilot validates the human service around that coverage. Worth noting that MITRE's managed services evaluation assessed how providers reported on adversary behaviour, and did not assess their ability to block or remediate.