AR · July 2026 · 7 min read
Deflection is the number your vendor likes. Here is the one you need.
Four metrics worth tracking, one worth ignoring, and how to tell a working agent from customers giving up.
If your dashboard shows deflection rising and nothing else, you cannot tell whether the AI is working or whether your customers have quietly stopped trying.
Why the default metric misleads
Deflection counts any conversation that did not reach a human. A customer who got a correct answer counts. So does one who gave up, and one who churned. The metric cannot distinguish them, and it is the number nearly every vendor leads with.
Track these four
| Metric | What it tells you | How to get it |
|---|---|---|
| Confirmed resolution | Problem actually solved | Post-conversation confirmation, not inference |
| Escalation context quality | Whether handoff works | Sample escalated tickets, check for repetition |
| Hallucination count | Confident wrong answers | Manual review of a weekly sample |
| Actual invoice | What it really costs | Your billing page, not the vendor's calculator |
Confirmed resolution is the one that matters
Ask your platform for resolution confirmed by the customer rather than inferred from silence. Vendors who have the number will give it to you. Vendors who redirect to deflection have told you something.
The gap between deflection and confirmed resolution is your real error rate. If deflection is 70% and confirmed resolution is 40%, roughly 30% of your customers went away unhelped and uncounted.
Sample manually. There is no substitute.
Read twenty AI conversations a week. Not a dashboard, the actual transcripts. Hallucinations and tone failures do not appear in aggregate metrics and are obvious within two minutes of reading.
Watch satisfaction, separately, by channel
If CSAT falls while deflection rises, customers are being deflected rather than helped. A Reddit thread on this notes roughly one in five customers report zero benefit from AI support, and that people are less forgiving of an AI mistake than a human one.
Track cost per resolved ticket, not cost per seat
On per-resolution pricing this rises as the AI improves. On per-seat it falls with volume. Knowing which direction yours moves tells you whether scaling helps or hurts.
Metric definitions here are read from vendor material and public buyer guides. We do not run performance tests.
Frequently asked
What metrics should I track for AI customer service?
Confirmed resolution rate, escalation context quality, hallucination count from manual sampling, and the actual invoice. Deflection alone tells you very little.
What is a good deflection rate?
The question is not answerable without knowing what triggers it. Ask for confirmed resolution instead.
How do I know if customers are giving up?
Compare deflection against confirmed resolution. The gap is customers who neither reached a human nor got helped.
How often should I review AI conversations manually?
Weekly, around twenty transcripts. Hallucinations and tone failures are invisible in aggregate metrics and obvious in transcripts.
What is a hallucination in customer support?
A confidently stated wrong answer, usually invented when documentation does not cover the question. These are the expensive failures because customers act on them.
Should CSAT change after deploying AI?
If it falls while deflection rises, customers are being deflected rather than helped. Track them together or the trade is invisible.
How do I measure escalation quality?
Sample escalated tickets and check whether the customer repeated themselves. If they did, the handoff is losing context.
What does the vendor dashboard not show me?
Usually confirmed resolution, hallucination rate and the deflection-resolution gap. All three require asking or sampling.
How do I calculate cost per resolved ticket?
Total monthly spend including seats divided by confirmed resolutions. Not by deflections, which inflates the denominator.
When should I conclude it is not working?
When confirmed resolution stays low after documentation is fixed, or when cost per resolved ticket exceeds what the human handling cost.
Tools mentioned
Full reviews, pricing tiers and where each one breaks.
Zendesk QA
ZendeskFormerly Klaus. Auto-scores every conversation and now grades AI agents alongside humans.
Fin by Intercom
The most polished autonomous agent on the market, attached to the pricing model buyers complain about most.
Freshdesk with Freddy AI
Cheaper than Zendesk with a comparable feature list. The trade shows up in depth rather than breadth.
eesel AI
Trains on your existing docs and tickets and works inside the help desk you already run, instead of replacing it.
You can also look into
How to Reduce Support Ticket Volume with AI - Proven Strategies (2026)
Which ticket types actually automate, how much to expect, and the metric that hides whether it worked.
Somebody claims 93% autonomous resolution. Here is what that number can hide.
Vendors publish autonomy rates between 30% and 93%. The range is that wide because they are not measuring the same thing.
Best practices for AI customer support, written as things that go wrong
Best-practice lists are usually a vendor describing its own feature set. These are the failure modes teams actually report, and what prevents each one.
How to reduce customer support costs with AI (and the pricing model that undoes it)
Every guide on this topic promises 30-40% savings. None of them mention that the most common AI pricing model charges you more the better the AI works.
WRITTEN BY AR · UPDATED 2026-07-29
I read the fine print. Vendor pricing pages, billing definitions, terms, funding filings and acquisition notices — then I do the arithmetic nobody publishes: what a platform actually costs at your volume, what its headline metric is really counting, and who owns it now. I do not run benchmarks, and no page here pretends otherwise.