What 'automation rate' means, and how AI support vendors inflate it
Every AI support vendor quotes one number before any other: the automation rate. Siena’s case studies say 79% for Simple Modern. Intercom markets an average around 67% for Fin. Our own landing page says roughly 30% of a store’s volume in the first weeks. All of these are automation rates, and no two of them mean the same thing.
That is not always a lie. It is worse in a way: it is a metric with no standard definition, quoted without its denominator, measured by the party who benefits from the bigger number. This page explains what the words can mean, where the number gets padded, and the four questions that turn a marketing claim into a comparable figure.
Three different metrics, one name
When a vendor says “automation rate,” it means one of three things. They are not interchangeable, and the gap between them is where most of the inflation lives.
- Deflection. The customer got an automated answer and never contacted a human. Cheapest to claim: silence counts as success. If the answer was wrong and the customer gave up, deflection still scores it.
- Resolution. The AI handled the conversation end to end and the customer went away satisfied. Now satisfaction needs a definition, so vendors measure it with a post-conversation rating, a no-reopen window, or a human sampling conversations. Each choice moves the number.
- Containment. The conversation never left the AI’s queue. Close to deflection, but usually measured inside one channel, which quietly drops everyone who gave up in the chat and emailed instead.
Deflection is always the biggest number. Resolution is the honest one. Containment is whatever the vendor’s funnel says it is. When you see a rate quoted with no label, assume deflection and negotiate from there.
Where the number gets padded
Five mechanics do most of the work. None of them require anyone to lie.
1. The eligible-subset trick. “Our AI resolves 90% of tickets” often means 90% of the tickets the AI is allowed to touch. If only 40% of your volume is AI-eligible, the honest org-wide math is 40% times 90%, about 36% of total volume, before platform costs. Industry analysis puts the realistic first-year net reduction for a whole support operation at 20 to 35%, against the 85 to 95% per-ticket reduction quoted in headlines. Both numbers are true. They describe different denominators.
2. The cherry-picked cohort. Case-study customers are the vendor’s best fits: clean policies, simple catalogues, motivated teams. Siena publishes 36 named customers with granular numbers; those are stores where the deployment worked. The honest question for a case study is not “did it work there” but “am I shaped like there.”
3. The flexible definition of resolved. Gorgias bills its AI per resolved conversation and rebills as a human ticket if a person joins within 72 hours. That is a billing definition, and a reasonable one. But “resolved” can also mean the customer stopped replying, the ticket auto-closed after N days, or the AI sent a link. When the same word sets the invoice and the success metric, ask which definition is attached to which number.
4. Ticket splitting. One angry customer with a missing order can become three conversations across email and chat. If each automated exchange counts as a resolution, the rate rises while the customer’s problem is once and unsolved. Per-conversation billing has the same arithmetic working on your invoice.
5. The satisfaction filter. CSAT is often collected only on AI-handled conversations. Customers who lose patience and escalate to a human exit the AI’s denominator before they can vote. The AI’s score stays high partly because the angriest people left the room.
Claimed vs measured
The checkable version of this gap exists in the wild. Intercom markets an average resolution rate around 67% for Fin, aggregated from customer reviews. Independent reviews and field reports typically land real-world deflection in the 50 to 80% range, tracking almost entirely with how complete the knowledge base is, and at least one independent test measured Fin under 40%. Nothing here is fabricated. One number is a ceiling under good conditions, the other is a median including messy ones. A vendor quoting the ceiling as the expectation is not lying to you, but they are letting you lie to yourself.
The same discipline applied to our own deployment: our first store runs at roughly 30% of support volume handled, measured as cases the AI closed without human involvement and the monthly audit did not overturn. We quote it because it is what a real first-month deployment did, not because 30% is a impressive number. It happens to be inside the honest 20 to 35% net range the industry analysis predicts.
The four questions that fix the number
Ask any vendor, including us, these four, and write down the answers:
- What counts as automated? Deflection, resolution or containment? What closes the gap between their definition and a customer who stopped replying?
- Of which denominator? All inbound conversations, or the AI-eligible subset? What share of a store like yours is typically eligible?
- Measured by whom, over what period? Vendor-reported or independently audited? Does the period include peak season, when policies strain and volume triples?
- What happens to the misses? How is an AI error detected, how fast is it corrected, and who is accountable? An automation rate with no error story is a revenue number, not an operations number.
A vendor who answers all four in plain language has earned a trial. A vendor who answers with another case study has answered a different question.
What to do with your own number
Whatever you deploy, measure it yourself, with a definition you would defend to your accountant:
- Baseline first. Cases per hundred orders, for a full month, before any AI touches anything.
- Define resolved your way. The customer did not reopen, re-ask in another channel, or refund within seven days. Harsh, honest, comparable.
- Read the misses monthly. Not the rate: the transcripts of what it got wrong. The rate tells you what to bill; the misses tell you what to fix.
Automation rate is a fine metric. It is just a privately defined one. Hold it to a public definition and the market gets comparable fast.
We build managed AI teammates for online stores, and the number we report to our stores every month is theirs, not ours: what it handled, what it missed, what it should learn next. If you want to see what that reporting looks like on your own queue, the conversation starts with Anna.
Sources: Intercom Fin resolution claims (G2 review aggregation and Intercom marketing, 2026); independent field reviews (getmacha, superframeworks, 2026); Lorikeet and Digital Applied cost benchmarks (2026); Gorgias AI billing terms (eesel AI and Featurebase teardowns, September 2026); Siena published case studies. Loqum deployment figures are our own, stated with their measurement method.