Trust & Privacy

Why We Describe De-Identification as a Tool, Not a Promise

· By Ian Vardy, CEO, Soma Health

Two things get conflated constantly: a model holding context while it works, and a vendor training on your content. They are different, and only one of them persists. Here is the distinction, plus the part that matters most — the strongest de-identification happens on your own machine, before anything is sent anywhere.

De-identification is a set of techniques that lower risk. It is not a guarantee that content can never be connected back to a person, and we don't describe it as one. The strongest version of it happens locally, on the clinician's own machine, before anything is transmitted at all — because content that never leaves can't be exposed by anything downstream.

I build this software, so this is our own posture rather than neutral commentary. I'm writing it out because the vocabulary around it is the muddiest in our industry, and because two genuinely different things — a model holding context, and a vendor training on your content — get treated as the same worry.

A brass padlock centered on a blue-striped wall

The strongest protection is content that never leaves the machine in the first place.

Memory and training are not the same thing

This distinction comes up in almost every conversation I have, and untangling it is the single most useful thing in this article.

Memory, in the sense that matters here, is working context. When a model produces text, it has the material it's currently working with in front of it. That's how it can refer back to something from earlier in the same document. It exists for the duration of the work and then it's gone — the same way a calculator holds a number while you're using it. It doesn't accumulate, and it doesn't carry into anyone else's work.

Training is different in kind. Training means content is used to adjust the model itself, so that its future behaviour is shaped by it. That's persistent, it's cumulative, and it's the thing people are actually worried about when they ask "does it remember my clients?"

Conflating them produces two opposite mistakes. Some people assume any AI tool is quietly learning from their files, which usually isn't true. Others hear "we don't train on your data" and conclude nothing is processed at all, which definitely isn't true. The honest position is that content is processed while the work happens, and whether it's retained or learned from afterward is a separate question with a separate answer — one you can find in a vendor's terms, as I've written about in does an AI tool train on your client notes.

For us: content used to improve the product is off by default, and enabling it takes both a clinician's opt-in and each individual client's separate approval. Neither alone is sufficient.

Why is local de-identification the strongest version?

Because it changes what exists to protect rather than how well it's protected.

Every other safeguard — encryption, access control, retention limits, contractual restrictions — is a control applied to content that has already been transmitted. They're all worth having, and they all depend on somebody's systems and somebody's promises continuing to work.

Content that was reduced or stripped before it was ever sent doesn't depend on any of that. There's no version of a downstream failure that exposes something which never left the machine. That's a categorically stronger position than a well-secured copy of the full thing, and it's why I'd rather be judged on where the reduction happens than on any badge.

It's also the question I'd put to any vendor in this category: what happens on my device before anything is transmitted, and can you describe it specifically? A vendor who has thought about this will have a concrete answer. A vendor who hasn't will describe their encryption instead — which is answering a different question.

A laptop showing a padlock and a secured connection

Every other safeguard protects content that has already been sent. This one changes what gets sent.

So why isn't de-identification a promise?

Because the research doesn't support the stronger claim, and I'd rather say so than repeat a comfortable phrase.

The privacy literature has been direct about the limits of anonymisation, particularly as the ability to link datasets improves. Removing direct identifiers meaningfully reduces risk. It does not produce a mathematical guarantee that content can never be reconnected to a person, and anybody claiming that guarantee is overselling something they can't underwrite.

So we describe it as what it is: a control that lowers risk, layered with others, applied as early in the pipeline as we can manage. Not a magic word that makes an obligation disappear — and the obligation genuinely doesn't disappear, since responsibility for client information stays with the custodian regardless of what a vendor's page says.

What should you take from this?

Three things.

Ask which one a vendor is describing. "We don't train on your data" and "nothing is retained" and "it's de-identified" are three separate claims. A vendor should be able to make each one independently and tell you which they're making.

Ask what happens before transmission. It's the question that separates a design decision from a marketing position.

Treat de-identification as one layer. It works alongside encryption, retention limits and access control. On its own it isn't a guarantee, and a vendor presenting it as one has told you something about how carefully they read.

If you want to run those questions at us, start here and then ask them. And the general framework for interrogating a processing pipeline — including ours — is in what to ask about where your data is processed.

— Ian

Ian Vardy
Ian Vardy
Founder & CEO, Soma Health

Ian is building Soma — AI tools that give clinicians their time back by drafting documentation, so therapists and psychologists can focus on their clients. He writes about clinical reporting, AI, and running a clinician-first software company.

See how Soma drafts reports →