Skip to content
elementai
Consultation
GUIDEMETRICS

How to measure the return on an AI implementation so the numbers mean something

Marcin PrzybyłFounder, element aiJune 16, 20267 min read

The most common mistake in measuring AI implementations is not picking the wrong metrics — it is starting to measure after go-live. At that point every number is useless, because there is nothing to set it against, and the discussion about results turns into an exchange of impressions. Below are four indicators that make sense in service processes, three that mislead with grim regularity, and one rule you have to settle before day one of the pilot.

IN SHORT
  • Measuring before the implementation matters more than measuring after it — with no baseline there is nothing to compare against.
  • Four metrics are enough: share of cases handled without a human, time to response, cost per case, and number of corrections.
  • A "percentage of correct answers" with no definition of correct is a metric that always comes out well.

Rule one: the baseline comes first

Before anything is implemented you need three numbers for the chosen process: how many cases a month, how many minutes one takes, and the hourly rate of the people handling them. That is the only moment when they can be measured honestly — after go-live nobody will reconstruct how long a case really took "before". We collect those three numbers during the audit and record them together with their definitions, because here the definition matters more than the value.

Definition means: is a "case" one email or a whole thread. Do we count handling time from arrival or from opening. Do we include abandoned threads. If someone counts it differently after go-live, the comparison stops meaning anything — and definitions drifting apart causes far more arguments about results than an actual absence of results.

Laptop screen showing charts and analytics data
Dashboard indicators are worth exactly as much as the definitions they were counted on — and as much as the measurement taken before go-live. Photo: Luke Chesser · Unsplash

Four metrics worth counting

01

Share of cases handled without a human. The single most important number of the whole implementation. Counted across all cases, not the ones the agent picked — otherwise it climbs on its own as the scope narrows.

02

Time from enquiry to response. Measured across the full volume, escalations included. The average can mislead when the tail is long, so keep the median and the 90th-percentile value next to it.

03

Cost of handling one case. Labour cost plus system cost, divided by the number of cases. The only metric that puts time saved and money spent on the implementation into a single number.

04

Number of corrections after the agent. How many answers the team had to change before sending or correct afterwards. It is a quality indicator and, at the same time, an early signal that the rules need adjusting.

Three metrics that sound good and mislead

First: the "percentage of correct answers" with no definition of correct. It always comes out high, because everyone judges differently and nobody checks a random sample. If you are going to use it, use it on a random sample, scored against written criteria, by someone who did not build the system.

Second: the number of enquiries handled. It goes up when the agent works, but it also goes up when customers write more often because they get faster answers. On its own it says nothing about savings. Third: customer satisfaction measured right after go-live. The novelty effect and the change of channel distort the result enough that a sensible comparison is only possible a quarter later.

"it works well"four numbers with definitions

The difference between a discussion of impressions and a conversation after which you can decide to extend the implementation to the next process — or to pull it.

What those numbers will not show

Two things escape measurement and can matter more than the savings. First: what the team does with the time it gets back. If 160 hours a month go back into substantive work, the implementation earns a second time — but you will not see that in the cost per case. Second: predictability. An answer within seconds, in the evening and at the weekend too, changes the customer experience in a way none of the four metrics above will capture.

So when we wrap up an implementation we report the numbers and describe those two effects separately, instead of trying to force them into an indicator. An indicator that covers everything usually measures nothing.

Frequently asked questions

How long before the return shows?

With a well-chosen process — one that takes several hundred hours a year and has clear rules — the cost of implementation usually pays back within a few months. With a process carrying a high share of unusual cases the period stretches noticeably, and that is one of the reasons we choose the process during the audit.

Who should count these metrics?

Someone on your side, on data from your systems. We provide the definitions, the logs and the counting method, but a result you cannot reproduce without us is worth little when you are deciding on the next processes.

Can the effect be measured without a pilot?

It can be estimated — from the three numbers in the audit plus an assumption about the share of unusual cases. That is an estimate, not a measurement, and we describe it that way. The pilot turns it into a number counted on real traffic.

All articles

Find out which process to hand over first

Free consultation: 30 minutes, one process and a first estimate of the time and money you'll win back. We reply within 24 hours.

Free · 30 minutes · no commitment