AI in Practice

Can I Use ChatGPT for Psychological Reports? What I'd Check First

· By Ian Vardy, CEO, Soma Health

The most common question I get. A general-purpose assistant will produce something that reads like a report — the problems are what it does with your client information, what it invents when material is missing, and whether you can verify any of it. Here is what I would check before trying it.

You can put material into a general-purpose AI assistant and get back something that reads like a psychological report. Whether you should comes down to three things: what happens to the client information you paste in, what the tool does when your material is incomplete, and whether you can trace any sentence back to something you actually supplied.

I build a competing product, so weigh that. I'm also not going to pretend nobody does this — a lot of clinicians have tried it, several have told me about it, and the honest answer is more useful than a warning.

Hands typing on a laptop showing a nearly blank document on a desk

It will produce something. The question is what it did with your material to get there.

What happens to what you paste in?

This is the first question and it's answerable without trusting anyone's marketing.

General-purpose assistants come in consumer and business tiers with materially different terms. On some consumer tiers, content may be used to improve the service by default; on business and enterprise tiers it typically isn't. The setting exists, the default varies, and it changes without much fanfare.

So: find the data-controls setting in the account you're actually using, check what it's set to, and check whether your tier's terms permit training. Don't rely on remembering what was true last year. The method for finding this in any vendor's terms — which words to search for — is in does an AI tool train on your client notes.

The related question is where processing happens geographically, which for a Canadian practice may matter to your own privacy assessment. General-purpose assistants generally don't let you choose.

What does it do when your material is incomplete?

This is the failure mode I'd worry about most, and it's specific to how these systems work.

A general-purpose assistant is built to produce a fluent, complete-looking response. If your input is missing something a report would normally contain, the strong tendency is to produce plausible text to fill the gap rather than to flag the absence. In ordinary use that's helpful. In a clinical document it's the worst possible behaviour, because the invented sentence reads exactly like the true ones.

There's a simple test. Deliberately leave something out of your material — a section of history, one set of results — and see whether the output names the gap or quietly papers over it. Run that before you run anything real. More on why fluency hides this in AI report writing accuracy: what to realistically expect.

Can you verify what comes back?

A review is only meaningful if you can check a claim against what you supplied.

With a purpose-built drafting tool, the draft is generated from structured material you provided, so tracing a sentence back is possible. With a long freeform conversation, the provenance of any particular sentence is much harder to establish — you'd have to re-derive it, which is often slower than writing it yourself and definitely riskier with your name on the document.

That's the difference that matters, and it isn't about model quality. It's about whether the workflow makes your signature mean something. I've written about why that line matters in why the clinician stays the author.

A stack of printed report forms with reading glasses and a pen on a desk

A review you can't check against source material isn't a review.

What about the format?

The practical friction people describe most.

A general-purpose assistant doesn't know your section order, your headings, your length, or the phrasings you always use. You can tell it, every time, in a long prompt — and people do, and it drifts anyway across a long document. Rewriting output into your own structure is frequently slower than starting from your own template. There's more on that test in psychological assessment report format: what any AI draft has to match.

So what would I actually check first?

Five things, in this order, before any real client material is involved:

  1. The data-control setting on the exact account you'd use, and what your tier's terms permit.
  2. The gap test — leave something out, see whether it flags or fabricates.
  3. Traceability — pick a sentence from the output and try to point at where it came from.
  4. Format fidelity — does it hold your structure across twenty pages, not two.
  5. Your college's position, which is the one I genuinely can't answer for you.

If it clears all five for how you work, that's a legitimate answer and I'd rather you knew it than bought something you didn't need. If it doesn't, the gaps are exactly what a purpose-built tool is for — ours is here, and I'd apply the same five checks to us.

— Ian

Ian Vardy
Ian Vardy
Founder & CEO, Soma Health

Ian is building Soma — AI tools that give clinicians their time back by drafting documentation, so therapists and psychologists can focus on their clients. He writes about clinical reporting, AI, and running a clinician-first software company.

See how Soma drafts reports →