Your voice agent is containing 78% of calls. That number is lying to you.
Containment is the metric every voice AI company reports, because it's the easiest one to compute and the hardest one for a buyer to disprove. It also has nothing to do with whether your customer got helped

9 min read time
Somewhere in your building, both of these are true
There's a slide that says the AI contained 78% of calls last month. Everyone in that meeting nods.
Two floors down, in a queue nobody presents, a customer is calling for the third time about the same broken meter.
Neither number knows the other exists. That's not a reporting oversight — it's the metric working exactly as designed.
What containment actually measures
The formula is simple enough to fit in a sentence: the percentage of calls that were handled without transferring to a human.
Read it again and notice what isn't in there. Nothing about whether the caller got an answer. Nothing about whether the answer was correct. Nothing about whether the right ticket was raised, whether the promise made on that call is one your business can keep, or whether the same person called back forty minutes later on WhatsApp.
A call where the agent said "I'm not able to help with that" and the customer hung up in frustration is a contained call. A call where the bot cheerfully raised a billing ticket for what was actually a supply fault is a contained call. A call that dropped at second nineteen because a carrier leg degraded is, in most dashboards, a contained call.
Containment measures one thing honestly: how little your human team was bothered. That's a cost metric wearing the costume of a quality metric.
THE UNCOMFORTABLE TRUTH
Every one of the four most common failures in production makes containment go up while the operation gets worse. The metric doesn't just miss the problem. It rewards it.
The four ways containment inflates
They're the same four everywhere — utility, EV, real estate, education. Different floors, identical failure.
Silent abandonment. The caller gives up. No transfer request, no escalation, no complaint — they just leave. Invisible unless you're instrumenting for it, because the system did precisely what it was told: it didn't transfer the call. Human floors have tracked abandonment for thirty years. It quietly disappeared the moment the agent stopped being human.
False resolution. The agent answers confidently and wrongly. The customer has no reason to doubt it, ends the call satisfied, and finds out three days later that the connection was never restored. This is the expensive one — it takes containment credit and generates a second, angrier contact, usually on a channel that reports to a different manager.
Disposition drift. "It stopped working right after the last bill" is a fault report. Tag it as a billing query, close it, and the numbers look immaculate: contained, dispositioned, closed same-day. The field engineer never gets dispatched. We've walked into live floors running north of 30% incorrect service-request categorisation, every one of those calls counting as a win.
Channel leakage. The customer gives up on the phone and opens WhatsApp. Because voice and digital live in separate systems with separate owners, the failed call and the follow-up are never joined. Voice reports success. Digital reports growth. Nobody reports the truth.
The reason nobody catches it has nothing to do with AI
You cannot audit a number you sample.
The standard on most floors is manual review of 1–2% of calls, days after the fact, by a small QA team working a spreadsheet. On 50,000 calls a month that's a few hundred conversations — chosen, in practice, by whichever queue was convenient on Thursday.
At that sample size a fifteen-point gap between containment and real resolution is statistically invisible. You'd need the failure to be both frequent and evenly distributed to reliably catch it, and production failures are neither. They cluster. In one language. One campaign. One intent. One node that started degrading on a Tuesday.
So the gap survives — not because anyone's dishonest, but because the measurement habit was built for an era when listening to every call was physically impossible. That constraint expired years ago. The habit didn't.
Which explains a set of numbers the industry keeps repeating without connecting: Gartner puts the share of AI agent pilots that never reach production somewhere near 89%. S&P Global found 42% of companies abandoning most of their AI initiatives in 2025, up from 17% a year earlier. MIT's Project NANDA found the overwhelming majority of generative AI pilots showing no measurable P&L return at all.
Put those alongside the containment rates being reported in the same period and only one explanation fits. A lot of pilots were succeeding on the dashboard and failing in the building.
So replace it. Here's the number we report instead.
Call it verified resolution. A call counts only if all five of these hold.
01. The intent was captured correctly.
Judged against the transcript, not against the disposition code the system assigned itself. This is the one that catches "it stopped working after the last bill" being filed as a billing query.
02. The action was executed in your system of record.
The ticket exists. The slot is booked. The CRM reflects it. Not promised on the call — executed during it. A confident closing line is not an outcome.
03. The disposition matches the intent, and no repeat contact follows within seven days — on any channel.
Voice, WhatsApp, SMS, email, walk-in. One customer, one issue, one identity. If you can't join a callback to its original call, your resolution rate isn't low or high. It's unknowable.
04 The call scored above threshold on your rubric.
Verification performed, disclosures made, nothing promised your business can't deliver, tone appropriate to the situation. Your rubric — your categories, your weights, your definition of resolved.
And the denominator is calls offered, not calls handled. If a call never connected, that's the operation's failure, not an exclusion from its statistics.
Three things happen when you adopt this.
Your number goes down. Ours did. Every honest number does.
It becomes falsifiable — every criterion can be checked by someone who doesn't work for the vendor. That's the entire point. A metric your vendor defines, computes and reports to you isn't measurement. It's collateral.
And it starts pointing at fixes. Containment failing tells you nothing you can act on. Verified resolution failing tells you which criterion broke — and a disposition failure is a knowledge base problem, a repeat-contact failure is a handoff problem, a rubric failure is a script problem. Three different fixes, three different owners, three different weeks.
You can't compute it from a sample
The criteria are per-call and the failures cluster, so the only sampling rate that works is 100%.
That's a change in kind, not degree. Every conversation — ours and your team's — scored against the same rubric within minutes of ending, not days. Categories and weights you own and can re-weight, because a missed disclosure means something different in a utility than it does in retail. And any score overridable by a manager who listened and disagrees, because a rubric nobody can argue with is a rubric nobody trusts.
The by-product turns out to be worth as much as the metric. Score everything and you stop having a call archive and start having a searchable record of every conversation your business had that day — which intents are rising, which promise is being made that shouldn't be, which language is underperforming, which new joiner needs coaching on verification and which needs it on closure.
That's why Auto QC is a product here and not a reporting tab. It's also the only reason we're willing to publish a resolution figure at all: across our production deployments 83% of calls resolve end to end, 100% of calls are scored, and we can show the working on every single one. We'd rather defend 83% line by line than quote 95% that evaporates the first time someone joins a callback to its original call.
Six questions to take into your next demo
They take four minutes and tell you more than the demo will.
01 How do you define a resolved call? If the answer is "not transferred," stop there.
02 What percentage of calls do you score, and how long after the call? Below 100%, or measured in days, means nobody knows what happened on your floor yesterday.
03 Is abandonment inside the AI session reported separately? If it isn't instrumented, it's being counted as a success.
04 How do you detect a repeat contact that arrives on a different channel?
05 Who owns the rubric, and can I re-weight it and override a score? If the vendor owns the scoring, the vendor is grading their own homework.
06 When a node in the call path degrades, what happens to the calls in flight — and does the dashboard show it?
Anyone selling you a toolkit will struggle with at least four of these, because the honest answer to most of them is you'll build that.
Why this number, out of all of them?
Containment became the headline metric because it was cheap to compute back when reviewing every call was impossible. It isn't impossible any more.
Every call can be scored. Every score can be explained. Every claim a vendor makes about your operation can be checked against your own recordings, in your own languages, on your own rubric.
Which means the industry has run out of excuses for reporting a number that describes how little the customer bothered us — and calling it service.
We don't sell software. We run the operation, and we answer for the number at month end. If you want to know what your real resolution rate is, don't argue about definitions. Score a week of your own calls and look.
Send us a week of your calls. We'll score them against your rubric and show you the gap.
Human-AI Synergy
We're teaming up AI and humans to bring the personal touch back to customer interactions—at scale.
© 2025, Lares Pvt. Ltd. - All rights reserved.
