Somebody has shown you a slide with hours saved on it, and you did not believe it. You were right not to. That number is almost always a usage count multiplied by an assumed minutes saved per task, and neither figure was observed. It was supplied, usually by the party with an interest in the answer.
This does not tell you the AI is failing. It tells you nobody has measured it. Those are different problems, and only one of them is expensive to fix.
Most AI measurement is theatre, and the mechanism is simple
Look at how that number is built. The usage count is observed. The minutes saved per task are assumed, the loaded hourly rate is assumed, and the product of the two is presented to people with no way to audit any step of it. Usage measures the system, not the business.
The second failure is quieter. The metric is chosen after the results are in, so whichever number moved becomes the headline. If the reported measure changes between reviews, you are looking at a claim with a chart around it.
Four questions get confused into one
Separating them is most of the work. They form a ladder, and each level is necessary for the one above it while proving nothing about it on its own. Almost every AI report jumps from the first question straight to the fourth.
1. Did the system operate?
Availability, error rate, failed and retried tool calls, queue depth and queue age, latency at the tail rather than the average, cost per successful run against budget, and how often a run needed a human to intervene before it could finish.
The failures that matter here are the silent ones. A run that completes and returns something useless is recorded as a success in most logs. Unless somebody instruments the difference between finished and correct, these metrics look healthy while the workflow rots underneath them.
2. Did people use it?
Active users measured against the population who could be using it, per workflow rather than per licence. Suggestions accepted, edited and rejected, counted as three outcomes rather than two. What the edits consistently change. Sessions started and abandoned, which nobody logs and everybody should.
The rejection rate is the most useful signal an AI system produces, and it is routinely misread. A step people keep rejecting is telling you the input context is wrong, the output arrives at the wrong moment in their day, or the step should not have been automated at all. Treating that as a training or attitude problem is how a system dies quietly while its dashboard stays green.
A system that is used but always edited before it goes out is a drafting tool. That is a legitimate permanent steady state for many workflows, and calling it a failed automation is a category error.
Adoption is the leading indicator. It moves before workflow metrics and long before financial ones, and it is the one question whose answer is decisive on its own: a system nobody uses cannot have worked, whatever the operating metrics say.
3. Did the workflow improve?
Cycle time from a start event to an end event that a system records, expressed as a distribution rather than an average, because the tail is where the cost and the complaints live. Backlog size and backlog age. Rework, exception and escalation rates, each defined in writing before anyone reads the numbers. Quality, scored by named humans against a rubric fixed in advance. Volume, so that throughput per unit of incoming work is comparable when demand moves.
4. Did the investment create value?
There are only five honest categories: revenue won that would not otherwise have been won, revenue retained, capacity created and then redeployed, cost avoided, and risk reduced. Each one requires a named owner inside the business rather than at the supplier, a stated causal assumption, and a decision that actually changed. Where the assumption does not hold, the correct thing to report is that the category is not available.
Saved time is not cash
Released time becomes money in exactly two ways. The freed capacity is redeployed onto work that earns or saves, or the headcount changes. In most businesses neither happens. The time is absorbed, which is a real benefit and belongs in the workflow section, not a financial one.
So the test for a value claim is not how much time was released. It is which decision changed as a result. A budgeted hire not made, a contract not renewed, agency work brought back in-house, overtime not paid, a service level met without adding a shift: those are decisions, and a decision leaves evidence in the accounts.
The one AI outcome we publish has that shape. We have built an order-entry agent for a UK builders' merchant on OpenAI's GPT-5.6, now in staging, that reads incoming orders and writes them into their system. The hire they had budgeted at £20,000 a year was not made. What makes that a value claim is not the hours. It is the vacancy that was closed.
Pipeline is not revenue either, so an AI system credited with creating opportunities has to be followed through to won and then to collected before anything is claimed. And revenue is not cash. Where the workflow touches billing or collection, the number that counts is what arrived in the bank, not what was invoiced.
A baseline captured after the build is a memory
People reconstruct the before-state once they already know the after-state, and the reconstruction bends towards the answer. An estimate produced after the fact is not a weaker measurement. It is a different kind of object, and presenting it as one is where most AI business cases lose a finance team.
If you did not measure before, you cannot claim improvement. Say that plainly, keep whatever operating and adoption evidence you do have, and start the baseline now so the next change is measurable. That answer survives a board meeting. An invented number does not survive the first person who asks how it was derived.
What to capture before a build starts:
- The workflow named, with a start event and an end event that some system already records.
- Volume across a period long enough to contain a peak and a trough.
- Cycle time as a distribution, taken from timestamps rather than from asking people how long things take.
- Rework, exception and escalation rates, with each definition agreed in writing.
- A quality sample, scored by named humans against a rubric that is fixed and kept, so the same rubric can be applied afterwards.
- The cost basis: who does this work today, at what loaded cost, and which budget line it sits on.
- The decision that would change if the number moved. If no decision would change, the workflow is not ready to be built.
The part clients want to skip is the part that cannot be recovered later. On most workflows the start and end events are not recorded anywhere, because the work happens in inboxes, spreadsheets and conversations. Establishing a baseline therefore means instrumenting the manual process first and living with it for a few weeks before a line of code is written. That is why our 12-month AI programme baselines a department before anything is designed or built, and why the final stage of every wave is improve or stop rather than a launch announcement.
Attribution is genuinely hard, and a clean number is a sales tactic
An AI system is almost never the only thing that changed. A quarter contains seasonality, pricing moves, staff joining and leaving, a competitor's behaviour, and usually a process change shipped by someone else in the same weeks. People also work differently when they know a process is being watched.
Anyone who hands you a single clean percentage has resolved all of that by ignoring it. What is actually available to you:
- A staged rollout with a holdout group, where volume is high enough for the comparison to mean anything.
- Deliberate on and off periods on the same team, which is cruder but works at low volume.
- Cohort comparison between teams doing comparable work, accepting that teams differ.
- Pre-registering the metric, the direction and the threshold before the build, so the measure cannot be selected afterwards.
Where none of those is possible, the evidence is directional. Label it that way. A directional finding you can defend is worth more than a precise one you cannot.
What a monthly review should contain
Run it in the order of the four questions, so nobody argues about value before adoption is established. Beyond the numbers, a review worth the hour contains:
- Each number with its derivation attached, and a label saying whether it was observed or assumed.
- One named owner per number, inside the business rather than at the supplier.
- The decisions taken since the last review, and what evidence they rested on.
- What was stopped, turned off or reduced in scope, and why.
- What remains unproven, stated as unproven rather than omitted.
- Any change in autonomy: which specific actions moved along the ladder from read to draft to recommend to act with approval, and what operating evidence justified the move.
A review that reports improvement every month is not a review. It is a status update with numbers on it, and the first genuine problem will arrive as a surprise.
Telling an improvement from a good quarter
Five tests, and a real improvement passes most of them:
- It persists across a period longer than the workflow's own cycle, so a seasonal effect has time to reverse.
- It survives a personnel change, a holiday period and a demand spike.
- It is visible in the distribution, particularly the slow tail, and not only in the average.
- It holds when volume moves, which is why throughput per unit of incoming work matters more than a raw count.
- It is corroborated by a second measure that moves for the same reason, ideally one nobody was optimising.
The final check is the one that settles most arguments. You should be able to name the step that got faster or more accurate and explain the mechanism at that step. If nobody can, the number is a good quarter, and good quarters end. Further reading on how capability builds in stages: the AI transformation ladder.
Where to start if AI is already running
If something is live and you cannot tell whether it is worth anything, do not begin by measuring it harder. Begin by working out which of the four questions you can still answer with evidence, which baselines have already been lost, and which workflows are worth instrumenting before the next build.
That is what our diagnostic assessment covers. It is short, fixed in scope, and carries no obligation. Measurement is far easier when the scope is a single department, which is the case for a department-sized rollout.
Stay Updated with Our Latest Insights
Get expert HubSpot tips and integration strategies delivered to your inbox.




