Data classification: finding your sensitive data before you protect it
Data classification is the practice of labelling information by how sensitive it is, so you can protect it in proportion. It sounds like bureaucracy, and done badly it is. Done well, it is the foundation the other controls stand on.
The reason is simple. You cannot apply least privilege to data you cannot identify. You cannot tell DLP what to watch for if you have not defined what "sensitive" means. You cannot meet GDPR's obligations over personal data without knowing where that data sits. And you cannot scope a breach in 72 hours if you do not know what the affected systems held. Classification is what turns "protect our data" from a slogan into something specific. The mistake most companies make is over-engineering it: a six-tier taxonomy with elaborate handling rules that nobody follows. A scheme people actually use beats a perfect one they ignore.
- Data classification is labelling information by how sensitive it is, so you can protect it in proportion. Done well, it is the foundation the other controls stand on.
- You cannot apply least privilege to data you cannot identify, tell DLP what to watch without defining "sensitive," or scope a breach in 72 hours if you do not know what the affected systems held.
- The mistake most companies make is over-engineering: a six-tier taxonomy with elaborate handling rules that nobody follows. A scheme people actually use beats a perfect one they ignore.
- Sensitive data does not stay where you expect. It spreads into SaaS, files, email, endpoints, and personal accounts, and that is where most companies underestimate what they hold.
- A classification scheme is only worth the effort if it drives action: access, DLP, retention, and incident response should all change based on the label.
a classification scheme that gets used
Keep it to a small number of tiers with clear, plain-language definitions. A workable default is four:
- Public. Information you would happily put on your website. Marketing material, public docs. No restriction.
- Internal. Ordinary business information that should not be public but is not sensitive. Most day-to-day work. The default tier.
- Confidential. Information that would cause real harm if exposed: customer data, personal data, contracts, financials, source code, credentials. Access restricted to those who need it.
- Restricted. The most sensitive subset: large volumes of personal data, regulated data, the crown jewels. Tightest access, strongest controls, often extra monitoring.
Two principles keep it usable. First, default everything to Internal, so people only have to make a decision for the things that are more sensitive. Second, tie each tier to concrete handling rules, especially access: who can reach Confidential and Restricted data, and how. A label with no consequence is just decoration.
finding where the sensitive data actually lives
Classification on paper is easy. The hard part is that sensitive data does not stay where you expect. It spreads. Finding it means looking in the places it accumulates:
Structured systems. Databases, the CRM, the HR system, the billing platform, the data warehouse. The obvious homes, and usually the easiest to reason about.
SaaS applications. Sensitive data ends up in support desks, project tools, shared drives, and the long tail of SaaS that teams adopted themselves. This is where it hides, and where most companies underestimate what they hold.
Files and collaboration tools. Spreadsheets of customer data, exported reports, contracts in shared folders. Unstructured data is the hardest to track and often the most exposed.
Email and messaging. Personal data and confidential attachments flow through mailboxes and chat constantly, and tend to stay there.
Endpoints and personal accounts. Data copied to laptops, personal cloud drives, and personal AI tools, the hardest to see and the reason browser security and shadow IT discovery matter.
You do not need to map every file on day one. Start with the systems most likely to hold Confidential and Restricted data, and the access to them, then widen.
from classification to control
A classification scheme is only worth the effort if it drives action. Once you know what you have and where it lives, the labels should change how the data is treated:
Access. This is the primary control. Confidential and Restricted data should be reachable only by those who need it, with least privilege enforced and access reviewed. Classification is what tells you which data deserves the tightest access.
Data-loss prevention. Now you can tell DLP what to watch: classification defines what "sensitive" means so the tooling can flag it leaving.
Retention and minimisation. Data you do not hold cannot leak. Classification surfaces what you are keeping that you do not need, which is both a risk reduction and a GDPR expectation.
Incident response. When something happens, knowing a system's classification tells you instantly how serious it is and what obligations apply.
keeping it alive
Classification decays as new data and systems appear, so build it into the flow rather than treating it as a project. Set a default tier so most data is handled without a decision. Fold a "what data does this hold" question into how new tools get adopted. And revisit the high-sensitivity systems on a cadence, because that is where the risk concentrates. The aim is a living picture you maintain, not a one-off audit you repeat from scratch.
let's start with a conversation
Most first conversations start with not quite knowing what you have or where to begin. That's normal, and it's exactly where we're useful.
Tell us what prompted this. An upcoming audit, an incident, a client's security questionnaire, or just a sense that things have gotten messy.
We'll take it from there

+48 783 762 997
julian@unshadowit.com

