Score data classification programme coverage and derive the control baseline each tier requires across a data estate.
Classification coverage has to be volume-weighted, because the unstructured estate is usually the majority of the data and the minority of the labelling effort. Accuracy is applied as a multiplier since a wrong label is actively harmful — it grants access or relaxes encryption on the basis of a mistake. The largest single distinction is whether labels drive enforcement: labels that only inform behaviour are documentation, not control. Classification is the dependency for least privilege, DLP and retention, which is why a stalled classification programme quietly blocks three other programmes.
Data Classification
coverage = structuredShare × structuredClassified + unstructuredShare × unstructuredClassified; programme = 0.55×(coverage × labelAccuracy) + 0.35×controlsEnforced + 10 for automated discovery.
coverage = structuredShare × structuredClassified + unstructuredShare × unstructuredClassified; programme = 0.55×(coverage × labelAccuracy) + 0.35×controlsEnforced + 10 for automated discovery. Classification coverage has to be volume-weighted, because the unstructured estate is usually the majority of the data and the minority of the labelling effort. Accuracy is applied as a multiplier since a wrong label is actively harmful — it grants access or relaxes encryption on the basis of a mistake. The largest single distinction is whether labels drive enforcement: labels that only inform behaviour are documentation, not control.
Classification is the dependency for least privilege, DLP and retention, which is why a stalled classification programme quietly blocks three other programmes.
This calculator takes 9 inputs: Data estate size, Share held in structured stores, Structured data classified, Unstructured data classified, Estate at the highest tier, Estate at the confidential tier, Automated discovery and labelling in place, Sampled label accuracy, Labels that drive enforced controls. The pre-filled defaults are a realistic starting point — replace them with figures from your own environment for a result you can act on.
With discovery of the highest tier. Find and label the restricted data first, apply enforced controls to it, then widen. Labelling the entire estate uniformly before any control is wired up produces metadata and no risk reduction.
Not at scale, and not consistently. User labelling works as a supplement to automated classification with a small number of tiers — three or four. Six-tier schemes with mandatory user selection reliably produce guesses.