Red teaming guide

How to measure a red team engagement

Do not count findings. A red team stops looking once it has a route, so the number of issues in the report is a measure of how the engagement was run rather than of your security. The measures that survive scrutiny are all timings and ratios drawn from two timelines placed side by side: what the team did, and what your defenders saw. Everything useful comes out of the gap between them, which is why the recording arrangements matter more than the choice of metric.

Why is the findings count the wrong measure?

Because the engagement is not trying to be exhaustive. A red team pursuing an objective abandons exploitable weaknesses that it does not need, so a report with four findings may reflect an efficient path rather than a healthy estate, and a report with forty may reflect a team that ran out of stealth and started enumerating. The count varies with method, not with risk.

Counting also creates the wrong incentive on both sides. Suppliers paid against a perception of value learn to pad, defenders learn to argue individual findings down, and the conversation moves to severity ratings and away from the question of why nobody noticed a fortnight of activity. That argument consumes the debrief that should have produced the plan.

Severity ratings have the same problem in miniature. A rating is a property of a finding in the abstract, and the same misconfiguration is critical in one estate and irrelevant in another depending on what sits next to it. Argue about the path instead, because the path is specific to you and cannot be scored by anyone who has not seen your environment.

The test for whether your measurement is sound: if the same engagement were run by a quieter team who reached the objective in three steps, would your score improve or worsen? Under a findings count it improves, which is backwards, because a quieter attacker is worse news for you rather than better.

What should you measure instead?

Time and coverage, per objective. Time to first telemetry, time to first human looking, time to correct attribution, time to containment. Then coverage as a ratio: of the techniques the team executed, how many produced a record anywhere, how many produced an alert, and how many produced an alert somebody acted on. Those four numbers describe a pipeline, and each drop between them points at a different team and a different fix.

Add two outcome measures. How many objectives were achieved, and how many were achieved without detection at any stage, because an objective reached under observation is a different result from one reached invisibly. And record whether the response actually worked when it happened: containment attempted is not containment achieved, and a team that isolated the wrong host has a finding worth more than any vulnerability in the report.

One measure worth adding is whether the detection led to the right conclusion. Teams frequently notice something, classify it as a false positive or as routine administrative activity, and close it, which registers as a detection under any counting scheme and as a miss in reality. Record classification accuracy separately, because the remedy is analyst context rather than another rule.

None of these are comparable to an industry benchmark, and you should distrust any supplier who offers one. They are comparable to your own previous engagement with the same starting position and the same objective, which is the only honest baseline available.

Which measures are worth capturing?

The set below covers the pipeline from telemetry to containment. The last column is the point: each measure implies a specific remedial owner, which is what stops the metric becoming decoration.

MeasureHow to capture itA poor result looks likeWhat it points at
Time to first telemetryRed team timestamps against log search after the factNo record of an action exists anywhereLog pipeline and sensor coverage
Time to first human reviewAlert queue and case timestampsAlert raised on day one, opened on day nineTriage capacity or alert volume
Techniques recorded versus executedReconcile the action list with searchable logsHalf the actions leave no traceInstrumentation, usually cloud or container
Techniques alerted versus recordedCompare alerts fired with records foundEverything recorded, almost nothing alertedDetection engineering backlog
Time from detection to containmentCase notes plus the change or isolation recordDetected in hours, contained in daysAuthority to act, and out-of-hours process
Objectives reached undetectedObjective list annotated at the joint debriefAll objectives reached with no detectionThe programme, not a single control
Deconfliction requests loggedThe trusted contact's request logZero requests over a long engagementNobody was looking, or nobody escalated

Where do these numbers mislead you?

Fast detection on a loud technique tells you little. If the team was detected while running a scan, you have learned that your tooling catches scanning, which was never in doubt. Weight the coverage ratio towards the quiet techniques, particularly the ones using legitimate credentials and built-in tooling, because those are the ones a competent intruder will use and the ones your baseline cannot separate from normal work.

Detection caused by the exercise itself is the second trap. Once a deconfliction request has been made, the team monitoring is primed, and every subsequent detection is contaminated. Record which detections happened before the first deconfliction call and treat later ones separately. The same applies to announced exercises, where the whole window is primed by definition.

Sample size is the third trap. Every one of these measures comes from a handful of events in one engagement, so a difference of a few hours between years is noise rather than progress. Treat large movements as signal and small ones as nothing, and lean on continuous purple team measures for anything you intend to trend monthly.

Finally, do not compare engagements with different starting positions. An assumed-breach engagement that begins inside your cluster will show worse detection numbers than a perimeter engagement, not because you got worse but because the cluster is less instrumented than the edge. Comparability requires the same foothold, the same objective and ideally the same supplier, which is an argument for planning the retest at the time you commission the first exercise.

What should the report timeline look like?

Two columns, one page, same clock. On the left, every action the team took with a timestamp and the technique it corresponds to. On the right, what your side recorded: log events found afterwards, alerts fired, cases opened, decisions taken, and deconfliction calls. Insist on this in the statement of work, because a supplier who did not keep action-level timestamps cannot produce it retrospectively.

Read it for three patterns. Long silences, where the team operated for days with nothing on the right hand side, which are instrumentation problems. Rows where the right hand column has a log event but no alert, which are rule-writing work. And rows where an alert exists but the next entry is hours or days later, which are process and staffing problems and usually the cheapest of the three to fix.

Then agree the narrative in the room rather than by email. Defenders routinely find records the red team did not know existed, which improves your numbers, and occasionally the red team describes an action nobody can find any trace of, which is the most valuable line in the document.

How do you report progress without a score?

Repeat the exercise with the same starting position and the same objective, and report the change in the timings. That is a defensible statement to a board: last year the objective was reached undetected in nine days, this year the activity was attributed on day two and contained the same afternoon. No index, no maturity level, no comparison with unnamed peers.

Between engagements, report the pipeline measures from your own purple team work, because they move monthly rather than annually. Techniques with no telemetry, techniques recorded but unalerted, and detections that have been re-tested since the last platform change are three counts you control and can improve without buying another engagement.

Keep one qualitative statement in the pack as well, in the defenders' own words: what they saw, what they thought it was, and what stopped them acting sooner. It is not a metric and it is usually the most persuasive paragraph in the document, because it describes a decision rather than a number, and decisions are what the funding changes.

Resist converting any of this into a single percentage. A composite score hides exactly the distinction that matters, between a missing log source that costs a project and a missing rule that costs an afternoon, and it is the mechanism by which a programme stops being about defensible decisions and starts being about the number going up.

Common questions

How do you measure the success of a red team engagement?
By timings and coverage ratios rather than findings. Record time to first telemetry, time to first human review, time to correct attribution and time to containment, then the ratio of techniques executed to techniques recorded, alerted and acted upon. Add objectives achieved and objectives achieved without any detection. Each measure points at a specific owner and a specific fix.
Why is the number of findings a bad red team metric?
Because the engagement is not exhaustive by design. A team pursuing an objective ignores weaknesses it does not need, so a short findings list can mean an efficient path rather than a healthy estate. Worse, the count improves when a quieter attacker reaches the objective in fewer steps, which is a worse outcome for you, so the metric moves in the wrong direction.
What is a good time to detect a red team?
There is no defensible industry figure, and any supplier offering one is selling. The only honest baseline is your own previous engagement with the same starting position and the same objective. Judge progress by movement in your own timings, and treat detection of loud techniques such as scanning as near worthless compared with detection of activity using legitimate credentials.
What should a red team report contain?
A two-column timeline on a single clock: every action taken with a timestamp and its technique on one side, and what your side recorded, alerted, opened as a case and decided on the other, including deconfliction calls. Require it in the statement of work, because a supplier who did not keep action-level timestamps cannot reconstruct it afterwards and the detection assessment becomes impossible.
How do you report red team results to a board?
As a change in timings against a repeated exercise with the same foothold and objective, stated plainly. For example, the objective was previously reached undetected over several days and was attributed within hours this time. Avoid composite scores, because they merge a missing log source that needs a project with a missing rule that needs an afternoon, which is the distinction the board is funding.

More on Red teaming

Let’s create something out of this world together.

Have a project in mind? Contact us for expert design and development solutions. Let’s discuss how we can help grow your business.

Azaadi Offer

Claim a free security assessment

Until 31 August we're covering the cost of a full vulnerability assessment and penetration test. Mention it in your message and we'll scope it with you.

  • Web application testing, authenticated and unauthenticated
  • Mobile application testing across iOS and Android
  • External network and infrastructure assessment
  • Manual exploitation by engineers, not scanner output

Testing and the report are free. Fixing what we find is quoted separately, with no obligation to accept.

Read the full offer

Tell us what you are trying to build and we will tell you plainly whether we are the right people for it. Book a call with an expert to work through the detail, or ask for a fixed quote if the scope is already clear. No obligation either way.

Four fields is all we need to get started.

Fastnexa Logo

© 2026 fastnexa. All rights reserved.