Measure purple team value: detections built per technique tested, gaps closed, time to new detection and coverage gained.
A purple team is only worth its cost if findings become durable detections, so gap closure and re-validation carry more than half the weight. Speed matters because a gap known for three months is an unmitigated risk, and quality is scored explicitly because shipping noisy rules to close a gap on paper makes the SOC worse. The initial detection rate is reported separately: it measures the estate you started with, not the team's work. Red teams produce reports and blue teams produce alerts; the purple metric is the only one that tracks whether a finding actually changed what the SOC can see.
Purple Team Effectiveness
effectiveness = 0.30×gapClosure + 0.25×validation + 0.20×speed + 0.15×quality + 0.10×coverageGain, where gapClosure = detections shipped ÷ gaps found and speed = 100 − 3 × days to detection.
effectiveness = 0.30×gapClosure + 0.25×validation + 0.20×speed + 0.15×quality + 0.10×coverageGain, where gapClosure = detections shipped ÷ gaps found and speed = 100 − 3 × days to detection. A purple team is only worth its cost if findings become durable detections, so gap closure and re-validation carry more than half the weight. Speed matters because a gap known for three months is an unmitigated risk, and quality is scored explicitly because shipping noisy rules to close a gap on paper makes the SOC worse. The initial detection rate is reported separately: it measures the estate you started with, not the team's work.
Red teams produce reports and blue teams produce alerts; the purple metric is the only one that tracks whether a finding actually changed what the SOC can see.
This calculator takes 8 inputs: Techniques exercised in the cycle, Detected on the first attempt, Gaps identified, New or improved detections shipped, Detections re-tested and confirmed, Mean days from gap to shipped detection, Cycle length, Detections that generated unacceptable noise. The pre-filled defaults are a realistic starting point — replace them with figures from your own environment for a result you can act on.
No — it is the reason the exercise was worth running. A 37% first-run rate on realistic techniques is normal; what matters is whether the other 63% turned into validated detections before the next cycle.
Because a rule that fires two hundred times a day is switched off or ignored within a fortnight, which means the gap is still open while the metrics say it closed. Counting noise against the score keeps the number honest.