To find out whether an AI tool saves you time, measure your own baseline for two weeks first, then measure the same work with the tool for two weeks, and compare only the stage the tool actually touches. A vendor's headline number — mine included — is a claim about somebody else's practice.
I build one of these tools and I'd rather you held me to a number you produced than one I did. This is the method I'd use, and it's deliberately simple enough that you'll actually finish it.

The only number that means anything is the change against your own baseline.
Why is a vendor's time-saving number close to useless?
Not because vendors are lying — mostly they aren't — but because the number is unanchored.
A time saving depends on the length of your documents, your section structure, how much you edit, how fast you write already, and what stage of the work the tool touches. Change any of those and the figure moves. A saving measured on short session notes tells you nothing about a twenty-page assessment report.
There's also a reporting asymmetry worth knowing about: comparisons tend to be drawn against a slow baseline, because that's the version that makes the difference visible. Your baseline is whatever you actually do now, which is usually faster than the one in the case study.
So the number to want isn't "how much time does this save?" It's "how much time does this save me, on my documents?" That has an answer, and you can get it in a month.
What's the baseline, and how do you get one?
Two weeks of honest timing before you change anything.
Time only the stage you're evaluating. If you're assessing a report-drafting tool, time from "I have everything I need and I'm starting to write" to "this is signed". Don't include testing, scoring or the feedback session — the tool doesn't touch them, so including them dilutes the measurement toward zero.
Record three things per document: the stage time, the document type, and roughly how long the finished piece is. A note on your phone is enough. What you want at the end is a small table, not a precise average — five or six reports gives you a usable picture of your normal range.
Most people find the baseline itself useful regardless of what they buy. Several clinicians told me the exercise was the first time they'd seen the write-up separated from the rest of the assessment, which is also the thing that makes the bottleneck arguable in a staffing conversation. The broader arithmetic is in what an assessment report actually costs a practice to produce.
What do you time once the tool is in?
Exactly the same interval, on the same kind of documents.
The measurement is draft to signed, and it starts when the draft arrives — not when the tool finishes generating. Time spent reading, correcting, restructuring and rewriting is all part of the cost. A tool that produces a draft in ninety seconds and then takes you two hours to fix has not saved you two hours.
Keep timing for at least two weeks, and keep going past the first few. Nearly everyone gets faster with a new tool over the first several documents as they learn what to hand it and what to expect back. If you stop after three, you'll measure the learning curve rather than the tool.

Draft to signed. The clock starts when the draft arrives, not when the tool stops generating.
What will make the number lie to you?
Three things, and all three are common.
Timing a quiet fortnight. A tool trialled during a light week is measured under conditions that don't resemble the week you needed it. Run the comparison across a normal or busy stretch or the result won't transfer.
Changing two things at once. New tool plus a new template plus a new protected writing block, and you'll never know which one moved the number. Change one thing.
Counting the wrong stage. If the tool touches drafting and you measure the whole assessment, a real 40 percent improvement in the write-up shows up as a modest overall change and looks like noise.
There's a subtler one too: quality drift. If the "signed" document is quietly worse than your baseline document, the time saving isn't real, it's borrowed. Which leads to the check that matters most.
What does a real result look like?
Two numbers and one judgment.
The time delta on the stage the tool touches, expressed against your baseline range rather than as a single figure — "usually three to four hours, now usually two" is a more truthful result than "38 percent faster".
Consistency. Did every document improve, or one dramatically while the rest stayed flat? A single spectacular case is usually a case that suited the tool, not a pattern.
And the judgment: would you have signed these anyway? Pull two of the tool-assisted documents and read them cold against two baseline ones. If the assisted versions are thinner in the analytic sections, the tool moved work rather than removing it — and you've traded quality for hours without deciding to.
If the delta is real, consistent, and the documents hold up, that's a genuine result. If it isn't, you've spent a month and learned something true about your own practice, which is worth having either way.
Run this on us with everything else. The protocol for the trial that feeds it — synthetic case, no real client file — is in trialling an AI report tool without risking a client file, and the drafting itself is here.
— Ian
