Skip to content

Ask ten vendors about data sovereignty, and you will get ten answers about data residency. The two are not the same, and the gap between them is where most regulated enterprises get exposed. Most RFP templates ask whether data residency is satisfied. Almost none asks: who else can read the raw data? Ask the wrong question, and you get the wrong answer. Worse, most organizations discover the difference only after an incident forces the issue, which is the worst possible time to learn it.

Three words that are not the same

Let’s clear the air around the definitions first, because the confusion lies in treating them as one.

Residency is where your data physically sits. Ownership is who holds legal title to it on paper. Sovereignty is who can actually read it, process it, or be compelled to hand it over. Residency and ownership matter, and no one is arguing otherwise. But sovereignty is the one about control.

“The number of parties who can be compelled to disclose your data is exactly the number who can read it.”

Which raises a harder question: Who is actually reading your data? That link between reading and legal exposure is exactly where most DSPM tools break.

Where do most DSPMs fail?

To classify sensitive data, a DSPM has to read it first. Most tools send a readable copy to a frontier AI model, often running in another jurisdiction.

Picture a hospital in Abu Dhabi. Its patient records live in a UAE data center, exactly as the regulator requires. To tag what is sensitive, the DSPM ships those records to an AI service abroad. The records at rest never left the country, but a readable copy did. Check the compliance dashboard, and it reads all green.

“The exposure is never in the storage layer. It lives in what happens to the data in transit and in use. Residency holds. Ownership holds. But sovereignty fails.”

Three scenarios “protected on paper” has already failed

The hospital case is not a one-off. There is a name for this pattern, “protected on paper,” where every box is checked, and the data is exposed anyway. Here are three significant examples:

Samsung. Within about 20 days of permitting ChatGPT, Samsung engineers pasted semiconductor source code and a confidential meeting transcript into it. Good employees, no bad intent, just saving time. Once submitted, the data could not be recalled; its IP information was in a public LLM’s ecosystem.

The US CLOUD Act. Under US jurisdiction, an organization can be compelled to produce data regardless of where it is stored. Access follows corporate control, not the location of the data. That is a real problem, and it is rarely the question anyone asks while evaluating or buying a DSPM solution.

The recent Anthropic’s Fable and Mythos model rollback. Anthropic, under the US government’s directive, suspended two frontier models for an entire class of users based on their nationality. To stay compliant, the provider (Anthropic) had to disable those models the same day. For everyone who was using them, whether they thought they had the right model or not: sovereign cloud, in-region data, access literally vanished overnight.

These are some of the things we should be thinking about: What models are being leveraged in the AI-powered DSPMs? Every customer should be asking this of the technology providers they’re working with, so they understand the nuances and some of the challenges involved in acquiring these technologies.

None of these was a breach. Each was a routine, sanctioned workflow that neither residency nor ownership did anything to stop. This means the problem is architectural, and hence the solution must be as well.

The fix is architectural

Let’s give the traditional DSPM its due for what it’s worth. The sensitive data sits in the customer’s own environment, on file shares, in a database, or in a document repository. To classify it, the tool ingests that data into a cloud, usually a vendor’s multi-tenant cloud, and often across a border, simply because that is where the classification service runs. The moment it reaches a frontier model in another geography, the customer is exposed to whatever jurisdiction that infrastructure sits in, and to every sub-processor in the vendor’s stack they may never have been told about.

So every DSPM built this way shares the same structural stack. This is architectural, not a settings problem. And because the problem is architectural, the fix has to be too. You change the design so the reading step happens somewhere the organization controls. You flip the order and protect the data before anything ever reads it.

You might object that your DSPM does not do this. It uses regex or pattern matching instead. That works for the easy data, but it inherits the same limits DLP hit years ago. Ask it to accurately tell you what the data is, read the context, judge the sensitivity, and categorize it correctly, and the old methods fall short. Tools/DSPMs built on them are being phased out.

The hard data is the reason. If you want to find your IP, the source code, the semiconductor designs, and classify them correctly, you have to use these models. Most DSPMs are already moving down this path. The point is not to avoid AI. It is to understand what your DSPM does with your data once it gets there, and how that squares with residency, ownership, and sovereignty.

So, how do you protect the data that the DSPM has to read? Two approaches, and neither asks you to give up AI.

One: change what gets read

Transform the sensitive values before any classifier sees them. Mask and tokenize a record, a name, a national ID, a balance, so the model reads structure and patterns, never the raw data. The tokenization is contextual, so the model still knows a token is a national ID or a medical record number and classifies correctly. So that a legal order served on the vendor yields only meaningless tokens.

Two: keep the model in-country

Build on an open-weight model, one you can run in any geography you control. Think of an open-source equivalent to a frontier model, deployed inside your own boundary. It classifies the data in place, so raw values never leave the jurisdiction, and no foreign entity ever receives them. The data stays resident, ownership intact, and sovereign, because there is no outside destination at all.

Both approaches keep AI in play and keep raw data out of reach. Either way, you achieve the accuracy you need without trading away residency, ownership, or sovereignty. And you should not take either on faith. Ask vendors to prove it live, in real time, not on a slide.

Then make the protection travel with the data

Classifying data without exposing it is only half the job. Once a DSPM has found and tagged what matters, the real question is what happens to that protection when the file moves.

Picture a due-diligence data room in a cross-border merger, where classified files are shared with an outside advisory firm in another country.

Perimeter controls hold only while the file sits still. The moment someone downloads it or emails it out, the protection is gone. Bind the controls to the file itself instead, and they travel with it across clouds, borders, and partners.

“Sovereignty stops being a property of an address and becomes a property of the data.”

Sovereign by design

If you carry one idea out of all this, make it this one.

Look for technologies that are sovereign by design, not shoehorned into it after the fact. Trust baked into the architecture holds. Trust bolted on with a setting does not. So, put one question to any vendor you work with. Can you read my data, or be compelled to disclose it? Residency is not sovereignty. Do not conflate the two.

Before your next DSPM evaluation, take these five questions with you.