Transcription turns speech into text. Dictation turns your speech into text you intended to write. An ambient tool listens to a conversation without anyone addressing it. An AI scribe takes any of those and produces a structured document. They're four different products, and vendors use the words almost interchangeably.
I build in this category, so I've watched the vocabulary blur — sometimes carelessly, occasionally on purpose. Knowing which one you're being shown changes what questions you should ask, because the privacy implications and the failure modes are genuinely different.

Four words, four products. The differences matter more than the marketing suggests.
What does each one actually do?
Transcription converts speech to text and stops. The output is words in order, attributed to speakers if you're lucky. It makes no decisions about structure or relevance. Everything downstream is still your job.
Dictation is transcription of speech you're producing deliberately, for the document. You compose out loud; the tool types. The intelligence is entirely yours — the tool's job is accuracy and not making you fight it.
Ambient describes how audio is captured rather than what happens next: a tool listening to a naturally occurring conversation nobody is directing at it. Ambient capture usually feeds transcription, which usually feeds something else. It's a property of the front end, not a product category, which is exactly why it gets used loosely.
AI scribe is the one doing the most work and carrying the most ambiguity. It takes source material and produces a structured document — selecting, organising and phrasing. That's a different order of operation from the first three, because it makes editorial decisions.
Why does the distinction matter for privacy?
Because the questions you'd ask are different for each.
For anything with ambient capture, the first question is about the room: who else is present, what they've been told, and how someone declines. That's a consent conversation rather than a software feature, and it's the reason clinicians tell me ambient tools get the most scrutiny from their colleges and their clients.
For transcription and dictation, the questions are about the pipeline: where the audio goes, whether it's retained, and what happens to the text afterward.
For an AI scribe, add a third: what did the tool decide, and can you verify it. A structured document is the product of choices, and you're signing those choices.
There's a general framework for the pipeline questions in what to ask about where your data is processed.
What are the different failure modes?
This is the practical reason to know which one you have.
Transcription fails by mishearing. Wrong word, wrong speaker, dropped phrase. Annoying, usually visible, and correctable because you can compare against what was said.
Dictation fails by mistyping you. Same category of error, and you'll normally catch it because you know what you meant.
Ambient fails by capturing the wrong thing. Side conversations, a colleague in the corridor, material nobody intended to be part of the record.
An AI scribe fails by writing something confident and wrong. This is the failure mode that matters, because a well-formed sentence that nobody said is much harder to notice than a garbled one. The output reads correctly, so the error hides inside fluency.
That difference should change how you review. Checking a transcript is proofreading. Checking a structured document means verifying claims against source material — a different task, and a slower one.

A garbled transcript announces its error. A fluent invented sentence doesn't.
Which one solves an assessment write-up?
None of the first three, which is the point worth making.
A full assessment report isn't a record of a conversation — it's a synthesis of scores, history, observations and professional reasoning, most of which was never spoken aloud. Transcription has nothing to transcribe. Ambient has no conversation that contains the document. Dictation helps only if you were going to compose the whole thing out loud, which almost nobody does for a twenty-page report.
That's why session-note tools and assessment-report tools are genuinely different products even when the marketing is identical, and why I'd rather lose a therapist writing progress notes to something built for notes than pretend one tool serves both.
What should you actually ask a vendor?
Three questions that cut through the vocabulary:
"What is the input, and what is the output?" Audio to text is one product. Structured source material to finished document is another.
"What decisions does it make?" If the answer is none, it's transcription with better branding. If it selects or organises, ask how you verify those choices.
"Can I trace any sentence back to what I supplied?" This is the one I'd press hardest, because it decides whether reviewing the output is a real check or a formality.
The wider criteria are in how to choose documentation AI for a clinical practice, and the protocol for testing any of them safely is in trialling an AI report tool without risking a client file.
For the record, ours is a drafting tool — structured source material in, a first draft out, which you correct and sign. That's what it does, and it's deliberately not the other three.
— Ian
