AI in Practice

Trialling an AI Report Tool Without Risking a Client File — What I'd Do

· By Ian Vardy, CEO, Soma Health

You can run a genuinely informative trial of a report-drafting tool without putting a single real client file through it — using a synthetic case you build once, a fixed set of measurements, and a clear stopping rule. Here is the protocol I'd use, including on us.

You can run a real trial of an AI report tool without a single real client file going through it. Build one synthetic case, run it through every tool you're considering, measure the same four things each time, and decide against a threshold you set beforehand. It takes an afternoon to set up and it will tell you more than any demo.

I build one of these tools, so read this knowing I have an interest in you trialling carefully rather than being persuaded quickly. That's genuinely my preference — a customer who tested is a customer who stays.

Hands typing on a laptop showing a nearly blank document on a desk

The trial that tells you something starts before any real client information exists.

Why not just trial it on a real case?

Because of the order of operations. Evaluating a vendor means deciding whether you trust them with client information — and running the trial on client information means you've already answered that question in order to ask it.

There's also a consent problem. Putting a real file through a tool you haven't adopted, for the purpose of deciding whether to adopt it, is difficult to describe cleanly in a consent conversation. Clinicians raised this with me before I'd thought it through properly.

And practically, a real case makes a poor test. You can only run it once, you can't compare tools on identical input, and you'll be judging output you're emotionally invested in.

What should the test material be?

A synthetic case you construct yourself. Invent the person. Invent the history. Make up the scores. Nothing traceable to anyone, including composites of real clients — a composite is still derived from real files.

Make it representative rather than easy. Specifically:

  • Full length. If your real reports run twenty pages, a two-page test tells you nothing about the tool at the length where it matters.
  • An uneven profile. A flat, tidy case flatters every tool. A scattered one is where they separate.
  • Something missing. Omit a piece of information a real file would sometimes lack, and see whether the tool flags the gap or fills it in confidently.
  • Your own structure. Your section order, your headings, your typical length.

Build it once, save it, and reuse it for every vendor. Identical input is what makes the comparison mean anything.

What do you actually measure?

Four things, written down before you start.

Time from draft to signed. Not "did it produce something" — how long from receiving the draft to a document you'd put your name on, against your own current baseline for the same work. If you don't know your baseline, measure that first; there's a method in how I'd measure whether an AI tool actually saves you time.

Edit distance in the sections that matter. How much rewriting did the integration section need, versus the procedures section? Heavy editing of boilerplate is fine. Heavy editing of the analytic sections means the tool isn't doing the part you needed.

Fabrication count. Did anything appear in the draft that wasn't in what you supplied? This is the one to be strict about. A single invented finding in a test case is a disqualifying result, not a rough edge — and the gap you deliberately left is how you'll catch it.

Voice match. Read it aloud. Does it sound like you, or like a competent stranger? Rewriting someone else's prose into your voice is often slower than writing your own.

Hands typing on a laptop with handwritten notes beside it on a desk

Identical synthetic input across every vendor is what makes the comparison mean anything.

What should a vendor let you do?

The trial itself is a test of the vendor, not just the product. What I'd expect:

  • Your own material, not their sample. A demo document is chosen to flatter the demo. If you can't run your own structure through it before paying, that's an answer.
  • A trial without a card, or with a clean exit. And a plain statement of what happens to your test content afterward.
  • A straight answer on what the trial version does differently. Some trials run a smaller model or a shorter length limit. Fine — but you need to know, or you're evaluating a different product from the one you'd buy.
  • Someone technical to talk to. Ask a specific question and see who answers.

What did clinicians say went wrong in their trials?

Three things, all recognisable.

They trialled during a quiet week. Which means they measured the tool under conditions that don't resemble the week they actually needed help in. Trial it during a busy stretch or the result won't transfer.

They didn't set a threshold beforehand. Without a number decided in advance, the decision drifts into how the demo felt. Write down what "good enough to adopt" means before you see any output.

Only one person tried it. In a group practice this is a real trap — a tool that suits one clinician's writing style can be useless to a colleague. If more than one person will use it, more than one person runs the case.

What would I say about trialling us?

Run exactly this protocol on us, with the same synthetic case you use on everyone else, and hold us to the same fabrication threshold. If our draft doesn't beat your blank-page baseline on a full-length uneven case, we haven't earned the seat.

The wider set of criteria worth scoring every vendor against — data path, authorship, pricing at your real volume, exit terms — is in how to choose documentation AI for a clinical practice. If you want to see the drafting before you build a test case, it's here.

One synthetic case, four measurements, a threshold set in advance. That's the whole protocol.

— Ian

Ian Vardy
Ian Vardy
Founder & CEO, Soma Health

Ian is building Soma — AI tools that give clinicians their time back by drafting documentation, so therapists and psychologists can focus on their clients. He writes about clinical reporting, AI, and running a clinician-first software company.

See how Soma drafts reports →