guide
AI Resolution Rate Claims Explained - Why 93% and 30% Both Mean Nothing (2026)
One vendor says 93%. Another says 30%. Both are telling the truth.
Search for AI customer support and you will be told, on the same page of results, that AI resolves 93% of tickets and that it resolves 30% of them. Both numbers are published by credible parties. Neither is lying. They are measuring different things, and the difference is where your money goes.
OpenAI's case study on MavenAGI cites 93% of customer support questions answered autonomously and a 60% reduction in time to resolve. Intercom has publicly described seeing 30-50% of conversations resolved by AI. That is a threefold gap between vendors selling into the same buyer.
Three different things get called the same word
The gap is not performance. It is definitional. There are at least three distinct events that get reported as one metric:
- Answered — the AI produced a reply. Says nothing about whether it was right.
- Deflected — the conversation did not reach a human. Includes the customer giving up and leaving.
- Resolved — the customer's problem was actually solved, ideally confirmed by the customer.
A platform reporting 93% on the first definition and one reporting 35% on the third could be performing identically. You cannot tell from the marketing page, and in most cases you cannot tell from the sales deck either.
Why deflection is the number vendors prefer
Deflection is larger, easier to measure, and does not require asking the customer anything. It also counts your worst outcomes as wins. A customer who asks a question, gets a useless answer, gives up and never contacts you again is a deflection. So is a customer who churns quietly. Nothing in the metric distinguishes them from a genuine success.
“Deflection rate? Most platforms optimize for this. Resolution rate...”— Text Inc, SaaS buyer's guide 2026
The industry's own buyer guides say this out loud. It still shows up as the headline number on nearly every vendor site.
This is not only a measurement problem. It is a billing problem.
Where a platform bills per resolution, the definition sets your invoice. Intercom's Fin publishes $0.99 per resolution. Fini publishes $0.69. Sierra is reported at roughly $1.50, negotiated rather than published.
| Platform | Published rate | Model |
|---|---|---|
| Fini | $0.69 | Per resolution |
| Fin by Intercom | $0.99 | Per resolution |
| Sierra | ~$1.50 (reported) | Per resolved interaction |
| Zendesk | ~$55/seat + add-ons | Per seat |
Ranked on rate alone, Fini wins. But a platform that bills when the customer stops replying will charge you for abandonment, and a platform that bills only on confirmed resolution will not. The cheaper rate can produce the larger bill. Comparing the numbers without the definitions behind them is not a comparison.
What to ask instead
Four questions, in this order. They take about ten minutes on a sales call and they are worth more than any benchmark you will be shown.
- What specific event triggers your headline autonomy number? Get it in writing.
- What event triggers a billable resolution? These are frequently not the same event.
- What is your resolution rate confirmed by the customer, not inferred? Vendors who have this number will give it to you. Vendors who deflect the question have answered it.
- What happens to the metric when the customer abandons mid-conversation? Success, failure, or excluded?
The fourth question is the one that produces the most revealing silences.
We are reading every vendor's published billing and metric definitions and will publish them side by side. Until that is finished, this post deliberately does not summarise any individual vendor's definition second-hand — doing so would be exactly the unsourced claim it is arguing against.
Six claims, six definitions
Collected in one place, because the spread is the point rather than any individual figure.
| Vendor | Claim | Stated basis |
|---|---|---|
| Gorgias | Up to 60% of repetitive tasks | Scoped to repetitive, which is honest |
| Intercom / Fin | Up to 59%, elsewhere 30-50% | Two figures from one vendor |
| Decagon | ~70% average, 80%+ some customers | Enterprise, vendor-led implementation |
| Klarna | 67% of chats, month one | Peak, before the 2025 walkback |
| Tidio | 67% of common questions | Scoped to common, which is a large caveat |
| MavenAGI | Very high headline | Basis not published |
Notice how much work the qualifiers do. Up to 60% of repetitive tasks and 60% of tickets are entirely different claims, and only one of them is being made.
The three questions that decode any of them
- What is in the denominator? Removing out-of-scope queries is how 80% figures get built.
- Over what window? A conversation counted resolved on Monday that reopens Thursday is only resolved if nobody checks.
- Was this measured on all traffic or a selected deployment? Vendor-led enterprise implementations are not a benchmark for your self-serve rollout.
Against Gartner's published picture — only 14% of service issues fully resolved in self-service across 5,728 customers surveyed in December 2023, and only 36% even for issues customers themselves called very simple — those headline numbers describe a mature ceiling, not a starting point.
The same vendor, two different numbers
The clearest demonstration that these figures are snapshots rather than properties: one vendor publishing its own average resolution rate twice, with different answers.
| Source | Claim | Sample |
|---|---|---|
| Vendor's own ROI page | 76% average, top performers 80-84% | 12,000 customers |
| Cited in third-party coverage | 66% average | 6,000 customers |
Ten percentage points and half the customer base. Both are probably true of the moment they were measured, which is exactly the point — a resolution rate describes a changing population of deployments, not the software.
It also matters commercially. Sixty-six and seventy-six percent are a difference of roughly one agent in ten when you plan headcount from them, and neither figure comes with a date attached in most places it gets quoted.
What the qualifiers are doing
Read the claims again with attention to their scope and most of the spread disappears.
| Claim as written | What it actually says |
|---|---|
| Up to 60% of repetitive tasks | A ceiling, on a subset, of tasks not tickets |
| 67% of common questions | A ceiling, on the easiest slice of volume |
| Up to 59% | A ceiling, unqualified denominator |
| ~70% average across customers | A mean, on vendor-led enterprise implementations |
Three of those four are ceilings rather than averages, and two are scoped to a subset chosen because it automates well. None of them is a prediction about your queue, and none is dishonest — the qualifiers are right there in the sentence, doing work most readers skip past.
Where that leaves the number
A 93% autonomy figure is not necessarily inflated and a 35% figure is not necessarily worse. Without the definition attached, neither number carries information. Treat any autonomy claim without a stated measurement event as marketing, including the flattering ones, and especially the ones that appear alongside a per-resolution price.
SOURCES
WRITTEN BY AR · UPDATED 2026-07-28
I read the fine print. Vendor pricing pages, billing definitions, terms, funding filings and acquisition notices — then I do the arithmetic nobody publishes: what a platform actually costs at your volume, what its headline metric is really counting, and who owns it now. I do not run benchmarks, and no page here pretends otherwise.