Trust & Privacy

Does an AI Tool Train on Your Client Notes? Where the Answer Is Actually Written

· By Ian Vardy, CEO, Soma Health

The answer is almost never under a heading that says 'training'. It sits in the terms of service under service improvement, product development, or aggregated and de-identified data. Here is exactly where to look, which four phrases to search for, and what each one actually permits.

If you want to know whether an AI tool trains on your client notes, don't look at the marketing page — look at the terms of service, and don't search for the word "training". The permission is nearly always written under a different heading: service improvement, product development, or aggregated and de-identified data. That's where the real answer lives.

I build one of these tools, so I've read a great many of these documents, including ours. I'm not your lawyer and this isn't legal advice. What I can do is show you where the sentence usually hides and what the standard phrasings actually permit, because "do you train on my data?" is the question I get asked more than any other and it deserves a better answer than a reassuring yes-or-no on a sales call.

A laptop showing a padlock and a secured connection

The answer to this question is a sentence in a document, not a badge on a homepage.

Where is the answer actually written?

Three places, in descending order of how much they'll tell you.

The terms of service or data processing agreement. This is the binding one. Everything else is commentary. If a vendor has a separate DPA, ask for it even if you're a solo practice — it's usually far more specific than the public privacy policy.

The privacy policy. Broader and vaguer, but it often names the categories of use, which is enough to tell you what questions to ask next.

The sub-processor list. Frequently overlooked and often the most revealing document a vendor publishes. It names the other companies involved in delivering the service, which tells you who else is touching the data and under whose terms. If a vendor doesn't publish one, that's worth asking about directly.

What you're looking for in all three is not a promise. It's a permission — the sentence that describes what the vendor is allowed to do, which is a different thing from what they currently do.

What words should I search for?

Open the document and use your browser's find function on these, in this order:

  • "improve" — as in improve our services. This is the single most common wrapper for a training permission.
  • "aggregate" and "de-identif" — usually paired, and usually the qualifier that makes a broad permission sound narrow.
  • "machine learning" or "models" — sometimes stated plainly, more often in a carve-out.
  • "third part" — catches both third party and third parties, and finds the clause about who else gets access.
  • "retain" and "delete" — how long anything persists, and what deletion actually removes.

Five searches, maybe ten minutes. It will tell you more than an hour on a demo call.

What does "aggregated and de-identified" actually change?

This phrase does more work in our industry than any other, and it's worth understanding what it does and doesn't guarantee.

De-identification is a set of techniques for reducing how identifiable data is — stripping names, dates, locations and other direct identifiers. It genuinely reduces risk and it's a control worth having. But the privacy literature has been direct about the limits of anonymisation, particularly as linkage and modelling improve. It is a control, not a guarantee of non-identifiability.

Which matters here because "we only use aggregated, de-identified data to improve our services" is a sentence that permits training. It's just training on material the vendor has processed first. That may be perfectly acceptable to you — but you should know you're agreeing to it rather than discovering it later.

The follow-up question I'd ask: who decides what counts as de-identified, and is that process described anywhere? A vendor with a real answer will describe the method. A vendor without one will repeat the phrase.

A professional reviewing a thick stack of documents at a glass desk

Ten minutes with the terms of service beats an hour on a sales call.

Is "off by default" the same as "never"?

No, and the difference matters.

Off by default means the permission exists in the contract and the switch is currently set to off. That's a meaningful protection — it means nothing happens without a deliberate act — but the permission is still there, and defaults can be changed at renewal or in a new version of the terms.

Never would mean the permission doesn't exist at all. Very few products in this category can say that honestly, and you should be a little suspicious of one that claims it without pointing at the clause.

So the useful question isn't "do you train on my data?" — to which almost everyone says no. It's "what would have to happen for my data to be used to improve your product, and who has to agree?" That question has a specific answer, and the specificity of the reply tells you what you need to know.

There's a related question underneath this one about how your obligation works when data goes to a vendor at all, which I've written about separately in the plain guide to what PIPEDA actually says.

What does Soma do?

Since I'm telling you to interrogate vendors, here's ours, plainly.

Data used to improve the product is off by default. Turning it on requires two separate acts: the clinician opts in, and then each individual client separately approves. Neither alone is enough. Clients are aliased in the first place, so we're not holding direct identifiers. Content is encrypted in transit and at rest, and our documentation processing runs on edge infrastructure in Montreal, with a single step in report drafting where content transits a large-model provider outside Canada — I'd rather write that down than let a map imply otherwise.

You should not take that on my word. Ask me for the clause, the same way you'd ask anyone else. The full set of criteria I'd apply to any vendor in this category, including us, is in how to choose documentation AI for a clinical practice. If you want to see what the drafting itself does with the material you give it, that's here.

Ten minutes and five search terms. That's the whole method, and it works on every vendor you'll evaluate.

— Ian

Ian Vardy
Ian Vardy
Founder & CEO, Soma Health

Ian is building Soma — AI tools that give clinicians their time back by drafting documentation, so therapists and psychologists can focus on their clients. He writes about clinical reporting, AI, and running a clinician-first software company.

See how Soma drafts reports →