How to measure a red team engagement
Do not count findings. A red team stops looking once it has a route, so the number of issues in the report is a measure of how the engagement was run rather than of your security. The measures that survive scrutiny are all timings and ratios drawn from two timelines placed side by side: what the team did, and what your defenders saw. Everything useful comes out of the gap between them, which is why the recording arrangements matter more than the choice of metric.
Why is the findings count the wrong measure?
Because the engagement is not trying to be exhaustive. A red team pursuing an objective abandons exploitable weaknesses that it does not need, so a report with four findings may reflect an efficient path rather than a healthy estate, and a report with forty may reflect a team that ran out of stealth and started enumerating. The count varies with method, not with risk.
Counting also creates the wrong incentive on both sides. Suppliers paid against a perception of value learn to pad, defenders learn to argue individual findings down, and the conversation moves to severity ratings and away from the question of why nobody noticed a fortnight of activity. That argument consumes the debrief that should have produced the plan.
Severity ratings have the same problem in miniature. A rating is a property of a finding in the abstract, and the same misconfiguration is critical in one estate and irrelevant in another depending on what sits next to it. Argue about the path instead, because the path is specific to you and cannot be scored by anyone who has not seen your environment.
The test for whether your measurement is sound: if the same engagement were run by a quieter team who reached the objective in three steps, would your score improve or worsen? Under a findings count it improves, which is backwards, because a quieter attacker is worse news for you rather than better.
What should you measure instead?
Time and coverage, per objective. Time to first telemetry, time to first human looking, time to correct attribution, time to containment. Then coverage as a ratio: of the techniques the team executed, how many produced a record anywhere, how many produced an alert, and how many produced an alert somebody acted on. Those four numbers describe a pipeline, and each drop between them points at a different team and a different fix.
Add two outcome measures. How many objectives were achieved, and how many were achieved without detection at any stage, because an objective reached under observation is a different result from one reached invisibly. And record whether the response actually worked when it happened: containment attempted is not containment achieved, and a team that isolated the wrong host has a finding worth more than any vulnerability in the report.
One measure worth adding is whether the detection led to the right conclusion. Teams frequently notice something, classify it as a false positive or as routine administrative activity, and close it, which registers as a detection under any counting scheme and as a miss in reality. Record classification accuracy separately, because the remedy is analyst context rather than another rule.
None of these are comparable to an industry benchmark, and you should distrust any supplier who offers one. They are comparable to your own previous engagement with the same starting position and the same objective, which is the only honest baseline available.
Which measures are worth capturing?
The set below covers the pipeline from telemetry to containment. The last column is the point: each measure implies a specific remedial owner, which is what stops the metric becoming decoration.
| Measure | How to capture it | A poor result looks like | What it points at |
|---|---|---|---|
| Time to first telemetry | Red team timestamps against log search after the fact | No record of an action exists anywhere | Log pipeline and sensor coverage |
| Time to first human review | Alert queue and case timestamps | Alert raised on day one, opened on day nine | Triage capacity or alert volume |
| Techniques recorded versus executed | Reconcile the action list with searchable logs | Half the actions leave no trace | Instrumentation, usually cloud or container |
| Techniques alerted versus recorded | Compare alerts fired with records found | Everything recorded, almost nothing alerted | Detection engineering backlog |
| Time from detection to containment | Case notes plus the change or isolation record | Detected in hours, contained in days | Authority to act, and out-of-hours process |
| Objectives reached undetected | Objective list annotated at the joint debrief | All objectives reached with no detection | The programme, not a single control |
| Deconfliction requests logged | The trusted contact's request log | Zero requests over a long engagement | Nobody was looking, or nobody escalated |
Where do these numbers mislead you?
Fast detection on a loud technique tells you little. If the team was detected while running a scan, you have learned that your tooling catches scanning, which was never in doubt. Weight the coverage ratio towards the quiet techniques, particularly the ones using legitimate credentials and built-in tooling, because those are the ones a competent intruder will use and the ones your baseline cannot separate from normal work.
Detection caused by the exercise itself is the second trap. Once a deconfliction request has been made, the team monitoring is primed, and every subsequent detection is contaminated. Record which detections happened before the first deconfliction call and treat later ones separately. The same applies to announced exercises, where the whole window is primed by definition.
Sample size is the third trap. Every one of these measures comes from a handful of events in one engagement, so a difference of a few hours between years is noise rather than progress. Treat large movements as signal and small ones as nothing, and lean on continuous purple team measures for anything you intend to trend monthly.
Finally, do not compare engagements with different starting positions. An assumed-breach engagement that begins inside your cluster will show worse detection numbers than a perimeter engagement, not because you got worse but because the cluster is less instrumented than the edge. Comparability requires the same foothold, the same objective and ideally the same supplier, which is an argument for planning the retest at the time you commission the first exercise.
What should the report timeline look like?
Two columns, one page, same clock. On the left, every action the team took with a timestamp and the technique it corresponds to. On the right, what your side recorded: log events found afterwards, alerts fired, cases opened, decisions taken, and deconfliction calls. Insist on this in the statement of work, because a supplier who did not keep action-level timestamps cannot produce it retrospectively.
Read it for three patterns. Long silences, where the team operated for days with nothing on the right hand side, which are instrumentation problems. Rows where the right hand column has a log event but no alert, which are rule-writing work. And rows where an alert exists but the next entry is hours or days later, which are process and staffing problems and usually the cheapest of the three to fix.
Then agree the narrative in the room rather than by email. Defenders routinely find records the red team did not know existed, which improves your numbers, and occasionally the red team describes an action nobody can find any trace of, which is the most valuable line in the document.
How do you report progress without a score?
Repeat the exercise with the same starting position and the same objective, and report the change in the timings. That is a defensible statement to a board: last year the objective was reached undetected in nine days, this year the activity was attributed on day two and contained the same afternoon. No index, no maturity level, no comparison with unnamed peers.
Between engagements, report the pipeline measures from your own purple team work, because they move monthly rather than annually. Techniques with no telemetry, techniques recorded but unalerted, and detections that have been re-tested since the last platform change are three counts you control and can improve without buying another engagement.
Keep one qualitative statement in the pack as well, in the defenders' own words: what they saw, what they thought it was, and what stopped them acting sooner. It is not a metric and it is usually the most persuasive paragraph in the document, because it describes a decision rather than a number, and decisions are what the funding changes.
Resist converting any of this into a single percentage. A composite score hides exactly the distinction that matters, between a missing log source that costs a project and a missing rule that costs an afternoon, and it is the mechanism by which a programme stops being about defensible decisions and starts being about the number going up.
Common questions
- How do you measure the success of a red team engagement?
- By timings and coverage ratios rather than findings. Record time to first telemetry, time to first human review, time to correct attribution and time to containment, then the ratio of techniques executed to techniques recorded, alerted and acted upon. Add objectives achieved and objectives achieved without any detection. Each measure points at a specific owner and a specific fix.
- Why is the number of findings a bad red team metric?
- Because the engagement is not exhaustive by design. A team pursuing an objective ignores weaknesses it does not need, so a short findings list can mean an efficient path rather than a healthy estate. Worse, the count improves when a quieter attacker reaches the objective in fewer steps, which is a worse outcome for you, so the metric moves in the wrong direction.
- What is a good time to detect a red team?
- There is no defensible industry figure, and any supplier offering one is selling. The only honest baseline is your own previous engagement with the same starting position and the same objective. Judge progress by movement in your own timings, and treat detection of loud techniques such as scanning as near worthless compared with detection of activity using legitimate credentials.
- What should a red team report contain?
- A two-column timeline on a single clock: every action taken with a timestamp and its technique on one side, and what your side recorded, alerted, opened as a case and decided on the other, including deconfliction calls. Require it in the statement of work, because a supplier who did not keep action-level timestamps cannot reconstruct it afterwards and the detection assessment becomes impossible.
- How do you report red team results to a board?
- As a change in timings against a repeated exercise with the same foothold and objective, stated plainly. For example, the objective was previously reached undetected over several days and was attributed within hours this time. Avoid composite scores, because they merge a missing log source that needs a project with a missing rule that needs an afternoon, which is the distinction the board is funding.