AI receptionist guide

Measuring whether an AI receptionist works

Containment rate is the number every vendor leads with and the one you should trust least. It counts calls that ended without reaching a human, which means a caller who lost patience and hung up is scored identically to a caller who got what they needed. A metric that cannot distinguish success from abandonment is not a quality measure, it is a reassurance. Here is a set that can fail, and how to find the failures that leave no trace in any dashboard.

Why does containment rate mislead?

Because it measures the absence of a transfer rather than the presence of an outcome. Three very different calls all count as contained: the caller who booked an appointment, the caller who was told something wrong and believed it, and the caller who hung up at the second misheard question. Reported as one figure, they average into a number that goes up when the system gets worse at handing over.

That perverse incentive is the deeper problem. Optimising for containment means discouraging handovers, and handovers are the safety mechanism. A configuration change that makes the agent more persistent will raise containment and lower customer satisfaction at the same time, and nothing in the metric shows the second half.

If you keep containment at all, keep it split three ways: resolved, abandoned by the caller, and ended by the agent. The first is the only one worth celebrating, and the second is usually larger than anyone expects when it is first separated out.

Splitting it is easier than it sounds. An abandoned call has a signature: it ends during or just after an agent turn, no outcome was written to any other system, and there is often a second call from the same number soon afterwards. A resolved call ends with the caller closing the conversation and a record existing elsewhere. Those two patterns can be separated with a query rather than a project.

What should you measure instead?

Four things, and one of them is not about the agent at all. Answer rate, which is the share of inbound calls that got a live response of any kind, including the ones your team took. Outcome rate, which is the share of agent-handled calls that produced a booking, a qualified enquiry with a usable number, or a factually correct answer. Handover completion, which is the share of attempted transfers where a person actually spoke to the caller. And promise keeping, which is the share of committed callbacks that happened.

Answer rate is the number that justifies the project, and it needs a baseline taken before you switch anything on. Your phone provider can report unanswered and out-of-hours calls, and without that figure you will be comparing the agent against an impression.

Promise keeping is the one that will embarrass you, which is why it belongs on the list. If the agent tells forty callers a week that someone will ring before ten, and thirty of those calls happen, the agent is working and the business is not, and the customer cannot tell the difference.

MetricWhat it tells youHow it misleadsPair it with
Containment rateCalls that avoided a transferCounts abandonment as successSplit by resolved, abandoned, ended
Answer rateCalls that got a live responseSays nothing about qualityOutcome rate on the same calls
Outcome rateBookings and usable enquiriesNeeds a definition you can auditA sample checked by hand
Average handle timeCost per call, and frictionShort can mean efficient or abandonedTurn count and abandonment point
Handover rateHow often it steps backLow looks good, often means overreachHandover completion rate
Caller satisfaction surveyStated experienceOnly the extremes respondRepeat call rate within a day
Promise keepingWhether commitments were honouredNobody owns it, so nobody reports itThe overnight capture list

How do you tell a completed call from a good one?

By defining outcomes in advance as things that leave evidence in another system. A booking exists in the calendar. An enquiry exists in the CRM with a number that dials. A question was answered from a source you can point at. If the only evidence is the agent's own note that the call went well, you are measuring the agent's opinion of itself.

Then verify by sampling, because automated outcome tagging drifts. Pull twenty calls a week that the system marked successful and check them against the evidence. Errors cluster in a predictable place: calls where something was captured but captured wrongly, particularly numbers, which look like clean successes until someone tries to ring back.

Also track false completion separately once you find it. A call that ends with a booking at the wrong time is worse than a call that ended with a message, and if both are counted as outcomes you have hidden your most expensive failure inside your best metric.

What does a healthy handover rate look like?

There is no single right figure, and anyone quoting one is not accounting for scope. An agent restricted to bookings and messages will hand over often and correctly; an agent given a wide brief will hand over less because it is attempting more, which is not the same as helping more. What you can read is the direction and the mix.

Direction first. A handover rate that falls after a configuration change should prompt the question of what the agent is now attempting that it used to pass on, and a sample of those calls will answer it. A rate that rises usually means either a new type of caller or something in the business changed that the agent was never told about, such as a new service or a price.

Then the mix of reasons. Handovers caused by the caller asking are healthy and should be the largest group. Handovers caused by repeated recognition failure on the same question are a defect, and the specific question tells you what to fix. Grouping handovers by cause turns a vague number into a work list.

How do you find the calls that failed silently?

Look for repeat calls from the same number within a short window. Two calls from one number inside an hour is the clearest signal of a failed call that every standard metric records as two successes. Build that report first; it takes one query against your call log and it will change what you think is happening.

Then look at where callers stop. A hang-up immediately after an agent turn tells you which sentence loses people, and it is usually the same three or four sentences: an over-long greeting, a question the caller does not see the point of, or a readback of something misheard. Reading the last two turns of every abandoned call is the highest-yield half hour available.

Finally, search transcripts for the phrases callers use when the system is failing them. No, I said. That is not what I asked. Can I speak to someone. Hello, hello. Each is a marker for a defect that produced no error and no alert, and counting them weekly gives you a quality trend that does not depend on anyone's dashboard.

What should you review every week?

Thirty minutes, four numbers and fifteen calls. The numbers: answer rate against your pre-launch baseline, outcome rate, handover completion, and callbacks promised against callbacks made. The calls: five picked at random, five that ended soonest after an agent turn, and five from numbers that rang more than once that week.

Keep it a written note rather than a dashboard glance, because the value is in the pattern between weeks. Two consecutive weeks of the same misheard street name is a vocabulary fix. Two consecutive weeks of the same unanswerable question is a script or a knowledge fix. Neither is visible in a single week's figures.

And review one thing that is not about the agent: what your team did with what it captured. Most voice agent disappointments turn out to be handover and follow-up failures rather than conversation failures, and the review that only examines the AI will keep finding the AI innocent while customers keep leaving.

Give the review an owner and a fixed slot, because it is the first thing dropped when a week gets busy, and the system degrades quietly rather than breaking. A month without transcript review typically ends with a vocabulary problem nobody caught, a promise nobody kept, and a dashboard that still reads well. That half hour is the maintenance schedule for the whole thing, not an optional extra for the curious.

Common questions

What is containment rate, and why is it a poor metric?
Containment rate is the share of calls that ended without reaching a human. It treats a caller who booked an appointment, a caller who was given a wrong answer, and a caller who hung up in frustration as the same result. It also rewards discouraging handovers, which are the safety mechanism. If you keep it, split it into resolved, abandoned by the caller, and ended by the agent.
What metrics actually show whether a voice agent is working?
Four. Answer rate, the share of inbound calls that got a live response, measured against a baseline taken before launch. Outcome rate, the share of handled calls producing a booking, a usable enquiry or a correct answer. Handover completion, the share of attempted transfers where a person actually spoke. And promise keeping, the share of committed callbacks that happened.
What is a good handover rate for an AI receptionist?
There is no single figure, because it depends entirely on the scope you gave the agent. Read the direction and the reason mix instead. Handovers because the caller asked are healthy and should dominate. Handovers caused by repeated recognition failure on the same question are defects, and the question involved tells you what to fix.
How do we find calls that went wrong without anyone noticing?
Three reports. Repeat calls from the same number within an hour, which every standard metric records as two successes. Hang-ups immediately after an agent turn, where the last two turns show which sentence loses people. And a transcript search for phrases such as no, I said, that is not what I asked, and can I speak to someone.
Is average handle time a quality metric?
It is a cost metric that sometimes reveals quality, and it points both ways. A falling average can mean the agent is more efficient or that more callers are giving up early, and the two look identical in the number. Read it alongside turn count and the point at which calls end, and treat any sudden drop as a possible abandonment problem.
How often should we review call transcripts?
Weekly, for about thirty minutes. Read five random calls, five that ended soonest after an agent turn, and five from numbers that rang more than once that week. Write the findings down rather than glancing at a dashboard, because the value is in the pattern across weeks: the same misheard street name twice is a vocabulary fix.

More on AI receptionist

Let’s create something out of this world together.

Have a project in mind? Contact us for expert design and development solutions. Let’s discuss how we can help grow your business.

Azaadi Offer

Claim a free security assessment

Until 31 August we're covering the cost of a full vulnerability assessment and penetration test. Mention it in your message and we'll scope it with you.

  • Web application testing, authenticated and unauthenticated
  • Mobile application testing across iOS and Android
  • External network and infrastructure assessment
  • Manual exploitation by engineers, not scanner output

Testing and the report are free. Fixing what we find is quoted separately, with no obligation to accept.

Read the full offer

Tell us what you are trying to build and we will tell you plainly whether we are the right people for it. Book a call with an expert to work through the detail, or ask for a fixed quote if the scope is already clear. No obligation either way.

Four fields is all we need to get started.

Fastnexa Logo

© 2026 fastnexa. All rights reserved.