If you're searching for reviews of AI scribes to work out what to use for assessment reports, the reviews are measuring the wrong thing. They test transcription accuracy, speaker separation and how quickly a note appears after a session. A psychological report isn't a record of a conversation, so none of those numbers predicts anything about it.
I build a report-drafting tool, so I have a stake in this distinction being understood. It's also just true, and it's the single most common category confusion I encounter.

Scribe reviews measure a session-note product. That's a different product.
What scribe reviews actually measure
Look at any of them and the criteria are consistent:
- Transcription accuracy — did it hear the words correctly
- Speaker separation — did it know who said what
- Latency — how fast the note appeared afterward
- Note format fit — does the output match a standard note structure
- Ambient capture quality — how it handles a real room
Every one of those is a fair test of a session-note product. Not one of them tells you how a tool handles a document synthesised from scores, records, observations and professional reasoning.
Why a report isn't a scribe problem
A scribe's input is speech. Its job is to lose as little as possible on the way to text.
An assessment report's input is a set of results and observations you've already gathered, most of which was never spoken aloud. There's no audio. There's no conversation to transcribe. The work is organising and describing around reasoning you've already done — the opposite direction of travel from speech-to-text. The reasoning itself stays yours in either category; no tool's job description includes it.
Which means a tool can score excellently on every scribe criterion and be useless to you, because it was built to solve a problem you don't have. The four categories are separated out in AI scribe vs dictation vs ambient vs transcription.
The criteria that would actually predict it
If someone were reviewing tools for report writing rather than note-taking, they'd measure:
Full-length output. Does it carry every section of a twenty-page document — the descriptive sections drafted from your material, the interpretive ones scaffolded for the judgment only you can supply — or a tidy paragraph.
Structural fidelity. Does the draft arrive in your section order and length, or a house style you'll rewrite.
Behaviour on missing input. Does it flag a gap or fill it confidently. This is the one that matters most and no scribe review tests it.
Traceability. Can you check a sentence against what you supplied.
Performance on an uneven case. Tidy inputs flatter everything. A scattered profile is where tools separate.
Draft-to-signed time against your own baseline — not generation time, which is the number that gets quoted.
Six criteria, none of which appear in a scribe review. If you want them applied to us as well, that's the best AI for psychological report writing? five things I'd judge it on.

The tests that predict report performance aren't in any scribe review.
Why the confusion persists
Partly vocabulary — "AI scribe", "ambient AI", "clinical documentation AI" and "report writing AI" get used interchangeably, sometimes by vendors who'd rather you didn't distinguish.
Partly market shape. Session notes are a far bigger market than assessment reports, so most products, most reviews and most search results are about notes. If you write assessments, you're reading content aimed at somebody else's job.
And partly because the products genuinely do overlap at the edges. Some vendors offer both. The useful question isn't which label a tool wears — it's what goes in and what comes out. Speech in, note out, is one product. Structured material in, long document out, is another. In either one, the clinician stays the author: you review, edit, and sign every word, and the clinical reasoning in the document is yours, not the tool's.
What I'd do instead of reading reviews
Build one synthetic case at full length with an uneven profile, run it through everything you're considering, and measure the six criteria yourself. It takes an afternoon and it beats any review, including a favourable one about us.
The protocol, without putting a real client file anywhere, is in trialling an AI report tool without risking a client file. If you'd like to run it on ours, start here.
— Ian
