Artificial intelligence
A voice agent in service: what you will not hear unless you listen
A voice agent dashboard is filled in by the agent itself. The only reliable instrument is an hour of listening a week.
A voice agent put into service produces reassuring figures very quickly: so many calls handled, such an average duration, such a percentage resolved. Those figures are sincere and they prove nothing, for a reason the first section sets out: the agent writes them.
Meanwhile things happen on the line that nobody sees. One caller repeats the same sentence three times. Another asks for a person and does not get one. A third hangs up midway, which is recorded as a completed call.
This article describes the only practice that gives access to that: listening, regularly, to a small number of calls. It says how many, at what rhythm, what to listen for, what to change afterwards and what to leave strictly alone. The companion article covers the decision to deploy — the thirty-second test, what never to entrust to a machine — and none of it is repeated here.
It carries no chart. The first section explains why the figures available on this subject cannot serve as evidence, for a reason different from every one our other articles give.
The dashboard is filled in by the party it grades
A voice agent produces its own statistics. It decides a call was resolved, it decides an intent was understood, it counts the transfers it chose to make. The system being measured and the measuring instrument are the same thing.
That is not an accusation of cheating, and it is more awkward than one. The agent classifies correctly by its own rules: a call where it said something and the caller hung up without complaining meets exactly the definition of resolved it was given. The fault is in the definition, not in the counting.
This is why we publish no resolution rate, neither ours nor anybody else, and why you should be wary of the ones you will be shown. A figure produced by the party it grades is not a weak measurement, it is one that cannot be wrong — it is true by construction, which is precisely the problem.
The only outside instrument available is the ear of somebody with no stake in the result. It is artisanal, it produces no curve, and it is the only thing that can say a call classified as resolved ended with an irritated caller hanging up.
The listening session: ten calls, one hour, every week
The practice fits into three numbers and beats any report. Ten calls, drawn at random across the week, listened to in full, by somebody who knows the trade. About an hour, once a week.
The draw has to be random and that is the easiest point to get wrong. Listening to the shortest calls, or to the ones the system flagged, means trusting the classification the previous section just said is worthless. Take the first call of each hour, or one in ten in order: any mechanical rule will do except the one that chooses.
The person matters more than the method. You need somebody who knows what a caller is actually asking for — from reception, the counter, after-sales — because half the defects are only audible if you know what the right answer would have been. A technical supplier hears a correct transcription where your salesperson hears a useless reply.
Ten calls a week sounds derisory against several hundred. It is not a statistical sample and does not claim to be: a voice agent defects are repetitive, so a real problem shows up in the ten. A problem that never shows up in the ten is not your priority.
The four ways a call ends
During the session every call falls into one of four boxes, and the sorting takes a second. That is the whole grid, and it replaces every indicator on the dashboard.
Handled: the caller got what they wanted with no human involved. Transferred: the agent passed it on, which is a success rather than a failure — an agent that never transfers is more worrying than one that transfers often.
Abandoned: the caller hung up before getting anything. That is the box that matters, and it is the one the dashboard nearly always files somewhere else.
And called back: the caller hung up and rang again shortly after, which is the most expensive and the most invisible case. It appears nowhere because it presents as two independent calls, and the only way to spot it is recognising a number while listening.
The transfer that did not happen
The most frequent defect of a voice agent in service is not answering badly, it is carrying on answering when it should have passed the call on. It is built to help, and it helps to the end, including when the right help was somebody else.
On the recording this is very audible and always the same: the caller rephrases. Two rephrasings are a sign, three are a certainty. The content can be perfectly correct at every turn; what is missing is any recognition that the conversation is not advancing.
The fix is not better answers but a giving-up rule: past two turns with no progress, hand over. That is counter-intuitive for somebody who paid for an agent — you are asking it to do less — and it is nearly always the change that most improves how the service is perceived.
You also need one explicit exit sentence: I am putting you through to somebody. Holding phrases — let me see what I can do — extend the call without promising anything, and that is exactly where callers hang up.
What the caller silence means
Silences are half the information in a call and they appear in no report. There are three kinds, and they mean nothing like the same thing.
A short silence after a question from the agent is normal: the person is thinking. A long silence after an answer from the agent is not: that is somebody who does not know what to do with what they just heard, and who will either rephrase or hang up.
The third is the most revealing: silence while the agent is still speaking. It indicates the person has stopped listening, usually because the answer is too long. That is a length defect rather than a content one, and it is corrected by cutting rather than adding.
This category explains why a transcript does not replace listening. The text of a call where the caller waits six seconds before replying is identical to the text of a fluent one; the difference is audible and does not write down.
The phrases it failed to understand twice
Every listening session produces a small list: the wordings the agent did not understand. It is the raw material for everything that follows, and it is worth keeping in one cumulative file rather than dealt with and forgotten.
The sorting rule is counting rather than irritation: a wording heard once is a curiosity, heard twice it is a gap. Without the cumulative file the second occurrence arrives three weeks after the first and nobody connects them — each stays a curiosity forever.
The gaps nearly always fall into three families, and naming them helps because they are not fixed in the same place. Local vocabulary, the most frequent: the way your customers actually name your products, which is not the way your catalogue does. Compound requests, where two questions arrive in one sentence. And requests with no relation to what you sell, whose only right answer is a transfer.
The first family is what makes a voice agent usable or not, and it is word work rather than technical work. It fills up by listening, never by thinking in advance about what people might say.
The call back: what happens after hanging up
A caller who rings again within the hour is the strongest signal you have, and no tool produces it. It says the first call failed, completely, regardless of how it was classified.
Spotting it takes little: in the call log, look for numbers appearing twice on the same day. It is five minutes of work and it gives a short list, which is the list of calls to listen to first the following week.
The cause is nearly always one of two, and listening settles it. Either the caller understood they would get nothing and rang back hoping to reach a person — in which case it is a transfer problem, section three. Or they got an answer and then found it wrong or incomplete, which is worse because they went away for a while believing they knew.
A call-back rate falling while the dashboard resolution rate stays flat is the one reliable indication that your agent is improving. It is also the only measurement on this page that resembles a figure, and it is usable precisely because your telephone system produces it rather than the agent.
What to change weekly, and what to leave alone
The listening session produces urges to modify, and the discipline is to satisfy only two or three a week. An agent modified on ten points is an agent nobody can explain any more.
What can be changed freely and safely: vocabulary, meaning the wordings the agent has to recognise; the length of answers, nearly always downwards; and the transfer rules, nearly always towards more transfers.
What must not be touched weekly: the structure of the conversation, and the scope of what the agent handles. Those two are design decisions, taken quarterly, and altering them at the rhythm of listening produces a shapeless system where every part answers the memory of one call.
Note what you change and the date, in the same file as the list of misunderstandings. Without that, in a month nobody will know whether an improvement came from a setting or a quieter week, and that is how teams end up believing false things about their own tool.
The hours when it should not answer
A voice agent does not need to cover the whole day, and making it cover the whole day is often what makes it unpopular. The calculation is not technical: it is about what happens when it does not answer.
During working hours the question is who would answer otherwise. If somebody would and they are available, the agent on overflow — after several rings — is nearly always better than the agent in front, and that is the use the neighbouring article recommends for a start.
Nights and weekends are the reverse: there is nobody, and an agent taking a message and reporting an intent is infinitely better than a phone ringing into an empty room. That is where it earns most, and it is where it gets installed last, because the volume there is low and the temptation is to put it where volume is high.
Finally there are moments when it is better that it does not answer at all: periods when you know the calls will all be of one kind and all difficult — an announced stockout, an incident, a promotion that went wrong. Switching it off for a day is an option, and it beats three hundred people irritated by a machine repeating that it does not understand.
When to narrow it rather than improve it
The reflex in front of a disappointing agent is to extend it: more cases, more vocabulary, more scenarios. That is nearly always the wrong direction, and listening shows it faster than any reasoning.
The signal is simple: if the well-handled calls all belong to two or three categories and the rest go to transfer or abandonment, then your agent is good at two or three things. Narrowing it to those, and transferring everything else immediately, improves everybody experience overnight.
An agent that clearly announces a narrow domain is perceived as competent; the same agent accepting everything and half failing is perceived as an obstacle. The difference in perception is enormous and the work to get it is a subtraction, which makes it fast and cheap.
It also answers the owner who asks why it is not made to do more. The right answer is to play them, on the recording, three calls from the category they wanted added: either those are plainly handleable, or the question closes itself.
Stopping it cleanly
Stopping has to be planned before it is needed, because a voice agent unplugged badly leaves a silent line, and a silent line is worse than anything this page describes.
The minimum is knowing, in advance and in writing, how the line returns to ringing normally, who is allowed to do it, and how long it takes. That information lives with your operator or in your telephone installation rather than with the agent supplier, and at least two people must hold it.
Temporary stopping is the common case: a hard day, a promotion, an incident. It has to be decidable in five minutes by somebody on your side, without calling a supplier. If that is not true today, it is the first thing to fix, before any vocabulary tuning.
A permanent stop deserves the same soberness as a migration: cut it, check the line rings, listen to two real calls the same day. And keep the recordings and the cumulative list of misunderstandings, which stay useful with no agent at all — they are the only written description of what your customers actually ask for on the phone.
What we do, and what we refuse
Our share is the listening apparatus and what it feeds: the random draw of calls, the four-box grid, the cumulative misunderstanding file, the giving-up rule after two turns, and spotting call-backs in the telephone log. We run the first sessions with you and then you hold them alone, because that is where they have value.
We refuse to publish a resolution rate, including our own, for a precise reason: the figure is produced by the agent, which is the party it grades. It is not approximate, it is true by construction. A supplier showing you one is showing you the definition they wrote themselves, and the question to ask is what they count as resolved rather than how many.
We also refuse to run the listening sessions in your place over time. Half the defects are only audible to somebody who knows what the caller actually wanted, and we do not know that as well as your reception does. A supplier selling that service is selling an hour of their time where an hour of yours is worth five times more.
Finally, two things on this page can be done today without us: listen to ten random calls this week, and look in the telephone log for numbers appearing twice on the same day. In an hour you will know whether your agent is doing any good, and the dashboard will not have told you.
Frequently asked questions
Why not trust the resolution rate on display?
Because it is produced by the agent itself, which is the party it grades. It is not an approximate measurement, it is one that is true by construction: a call where the agent said something and the caller hung up without complaining meets exactly the definition of resolved it was given. The question to put to a supplier is what they count as resolved, not how many.
How many calls should be listened to, and how often?
Ten a week, drawn at random, listened to in full, by somebody who knows the trade. It is not a statistical sample and does not claim to be: a voice agent defects are repetitive, so a real problem shows up in the ten. The draw must be mechanical — one in ten in order, or the first of each hour — because listening to what the system flagged means trusting the classification you are wary of.
My agent transfers a lot. Is that a bad sign?
No, the opposite. An agent that never transfers is more worrying: it is built to help and it helps to the end, including when the right help was somebody else. On the recording you hear it in the caller rephrasing — twice is a sign, three times a certainty. The useful fix is a giving-up rule after two turns with no progress, not better answers.
How do you spot the calls that really failed?
In the telephone log, look for numbers appearing twice on the same day. A caller ringing back within the hour says the first call failed, however it was classified. Five minutes of work, it gives the list of calls to listen to first, and it is reliable because your telephone system produces it rather than the agent.
Should the agent be extended when it disappoints?
Almost never. If the well-handled calls all belong to two or three categories, your agent is good at two or three things: narrow it to those and transfer the rest immediately. An agent announcing a narrow domain is perceived as competent; the same one accepting everything and half failing is perceived as an obstacle. The work is a subtraction, so it is quick.
How do you switch the agent off in a crisis?
It has to be decidable in five minutes by somebody on your side, without calling a supplier, and at least two people must know how. The information lives with your operator or in your telephone installation rather than with the agent supplier. If that is not true today, fix it before any vocabulary tuning: an agent unplugged badly leaves a silent line.
Where we come in
Ten calls drawn at random and listened to in full with somebody from reception are worth more than any report. Half the faults are audible and unmeasurable.
- We set up the random draw and the grid, then leave them with you.
- We pull the numbers that ring back more than once in a day.
- We attend the first sessions, not the ones after.
We will not hold these sessions for you over time: what is learned there is only worth anything to the people who answer the phone.
Read next
Measuring darija transcription: the protocol, and the figure we do not publish
Nobody publishes an error rate for Algerian darija, us included. Here is the protocol that would produce one, and what to know before reading it.The cost of a call, minute by minute
Four meters run during a voice call and none of them counts the same thing. Confusing them gives an estimate wrong by a factor, not by a margin.The words transcription does not know
The voice agent understands the sentence and gets the name wrong. That is the opposite of what people fear, and it is what makes an appointment unusable.
Let us talk about your project
A free audit, no commitment: we look at your online presence and tell you what is holding it back.