· Prakash Natarajan · Reliability · 16 min read

Deflection Rate, Explained: Formula and Benchmarks

Deflection rate can look great and still hide a broken support product. The formula, real benchmarks, and the cost-per-resolution number that matters.

Deflection rate can look great and still hide a broken support product. The formula, real benchmarks, and the cost-per-resolution number that matters.

Deflection rate measures the share of customer contacts an AI agent resolves without ever routing to a human, calculated as deflected contacts divided by total contact attempts. Published vendor benchmarks put a workable range anywhere from 25 percent in regulated industries to 75 percent in ecommerce, but that single number says almost nothing about whether the conversations behind it were actually solved. For a team building an AI SaaS product, the figure worth reporting pairs deflection with a resolution check and a real cost-per-resolution number your customer can verify against their own support bill.

What is deflection rate, and how do you calculate it?

Deflection rate is the percentage of inbound support contacts that never reach a human agent, whether because a bot answered them, a help article satisfied the customer, or the conversation simply ended on its own.

deflection rate formula calculation

The customer support platform Decagon publishes the formula most vendors use: deflected contacts divided by total contact attempts, multiplied by 100. Its own worked example takes 6,500 deflected contacts out of 10,000 total attempts for a deflection rate of 65 percent. You can run the same math on your own numbers. Say your agent handles 4,000 conversations in a month and 2,600 of them close without a human ever getting looped in: that’s 2,600 divided by 4,000, or 65 percent as well. The formula itself stays simple on purpose. Deciding what counts as “deflected” in the first place is the real work, since a customer who got a real answer and a customer who gave up in frustration both count the same way under this definition unless you build a separate check for which is which.

Most teams pull the total-attempts number from whatever system routes conversations, a help desk, a live chat widget, or the AI agent’s own session log, and the deflected count from however many of those sessions closed without a human touch. If your product serves multiple customers, run this calculation per customer rather than as one blended company-wide figure. A single aggregate number hides which of your customers the agent is actually working for and which ones are dragging the average down.

What counts as a good deflection rate for an AI agent?

There is no single “good” deflection rate. What counts as strong varies by industry, by how complex the underlying requests are, and by how long the agent has been live.

deflection rate benchmarks by industry comparison chart

Decagon’s published benchmarks put ecommerce support in the 55 to 75 percent range, SaaS support in the 40 to 60 percent range, and financial services or healthcare support in the 25 to 45 percent range, reflecting how much more caution and verification those regulated categories require before an agent can close something on its own. The same source describes a maturity curve where a newly launched agent starts around 30 to 40 percent and climbs toward 60 to 75 percent over three to six months, as the team feeds it more edge cases and expands what it’s trusted to handle. If you are building an AI SaaS product for other companies to use, that range should shape what you promise a prospective customer rather than a flat target. A vertical SaaS tool selling into healthcare should never be judged against an ecommerce benchmark, and an agent that just launched three weeks ago should never be compared to one that has been tuned for six months.

Treat these ranges as a sanity check rather than a scoreboard. A number that lands wildly outside the published range for a given industry, in either direction, is worth investigating before you report it anywhere. A number that lands right in the middle of the range isn’t automatically healthy either, which is exactly the trap the next section gets into.

How is deflection rate different from resolution rate and containment rate?

These three terms get used interchangeably in vendor marketing, and that’s exactly the problem: they measure different things, and only one of them tells you whether the customer’s actual problem got fixed.

comparing deflection rate containment rate and resolution rate metrics

The AI support platform Fin lays out a useful hierarchy, ranked by how much it actually proves. Deflection rate just counts whether a query ever reached a human, so a customer who clicked away from an unhelpful FAQ page still counts as deflected. Containment rate narrows that to AI conversations specifically, measuring whether the chat ended without an escalation, which filters out passive self-service but still says nothing about whether the answer was right. Automated resolution rate is the first metric that tries to check whether the AI actually addressed what the customer needed, verified through a survey, a completed action like a refund or an account change, or an explicit confirmation from the customer. The table below lays out what each one actually measures and where it falls short on its own.

MetricWhat it actually measuresWhat it misses
Deflection rateAny contact that never reaches a humanWhether the issue was solved, or the customer just gave up
Containment rateAI conversations that end without escalationWhether the AI’s answer was correct or the customer just stopped replying
Automated resolution rateWhether the AI’s action or answer matched what the customer neededHow resolution gets defined varies by vendor, and self-reported numbers aren’t independently audited

Fin’s own published numbers illustrate the gap between these tiers: across roughly 12,000 customers the company reports an average resolution rate near 76 percent, with top performers between 80 and 84 percent, while a deflection or containment number for the same conversations would read higher because it counts cases that never technically escalated but also never got properly solved. If you’re reporting one number to a customer, resolution rate is the one worth defending. Deflection rate on its own is a volume metric, not a quality metric, and treating it as proof of a working product is the single most common way these dashboards get misread.

Why can a high deflection rate still mean your agent is failing?

A high deflection rate can climb for the worst possible reason: customers giving up rather than escalating, because the self-service flow in front of them was confusing, slow, or clearly unhelpful.

chat conversation abandoned before escalation

Decagon’s own glossary carries a direct warning about this: “an operation can post high deflection while resolution quietly declines.” Forethought makes a related point from the customer experience side, noting that pushing deflection rate up through poorly built self-service tools can drag customer satisfaction down at the same time, especially when the self-serve error rate, the share of automated attempts that fail outright, climbs alongside it. Neither of those effects shows up in the deflection number itself. A dashboard reporting 65 percent deflection looks identical whether that 65 percent represents genuinely solved problems or a wall of customers who clicked away because nothing on screen helped them.

The sharper signal, according to Fin’s writeup, is the reopen rate: the share of “resolved” conversations where the same customer comes back with the same issue within 24 to 48 hours. A conversation that gets marked resolved and reopens two days later almost certainly wasn’t resolved at all, it just got closed. Tracking reopen rate next to deflection rate catches the failure mode that deflection rate alone is structurally blind to. This is also where trace-level visibility into what the agent actually did, not just whether the conversation ended, starts to matter: testing an AI agent before it ships and watching its live traces after it ships are both ways of catching a false resolution before it inflates a number you’re about to hand to a customer.

How do you turn deflection rate into a cost-per-resolution number customers will trust?

Deflection rate on its own doesn’t produce a dollar figure. To get a cost-per-resolution number a customer will actually believe, you need to combine deflection with a resolution check and the real cost difference between an automated and a human-handled contact.

cost per resolution calculation dashboard comparing automated and human support costs

Decagon publishes a cost range for this: automated contacts running $0.10 to $1.00 each versus $8 to $15 for a human-handled contact, and estimates that each additional percentage point of deflection across 100,000 monthly contacts saves roughly $70,000 to $140,000. Those figures come from one vendor’s own published estimates rather than an independently audited study, so treat them as a starting benchmark to sanity-check your own numbers against, not a number to repeat as fact about your specific product. Run the same math with your own cost data: take your actual per-contact cost for automated and human-handled conversations, multiply the difference by however many contacts your agent genuinely resolved, not just deflected, and you get a defensible savings figure instead of a borrowed one.

The word “genuinely” is doing real work in that sentence. A raw deflection count multiplied by a cost difference overstates the number, because it includes every customer who gave up along with every customer who got helped. A cleaner cost-per-resolution figure starts from your automated resolution rate, discounts out anything your reopen-rate check flags as a false resolution, and only then multiplies by the cost gap. If 2,600 conversations were deflected out of 4,000 total, but your reopen-rate check shows 200 of those came back within 48 hours with the same problem, your genuinely resolved count is 2,400, not 2,600, and that’s the number that should drive the dollar figure you show a customer. Showing the inflated number once and then having a customer’s own support logs contradict it later costs you more credibility than the gap between the two numbers was ever worth.

What do you need to measure this correctly across tenants?

If your product serves more than one customer, deflection rate and cost-per-resolution have to be calculated and reported per tenant, never as one blended number across your whole customer base.

multi-tenant dashboard showing separate isolated customer metrics panels

A blended average hides exactly the information a customer is paying to see. If one customer’s agent deflects 70 percent of contacts and another’s deflects 35 percent, a combined 52 percent average tells neither customer anything true about their own results, and it makes the underperforming customer’s account look better than it actually is right up until they compare notes with someone else. Calculating this correctly means every conversation, every tool call, and every reopen check needs to be tagged with which customer’s agent produced it, from the moment the agent runs, not reconstructed after the fact from a shared log. That tagging also has to hold up as an access boundary: a customer viewing their own deflection dashboard should never be able to see, even accidentally, the traces or numbers belonging to another customer on the same platform.

This is the specific engineering problem an embeddable, white-label reporting layer has to solve that a general-purpose observability tool usually doesn’t, because most observability tooling was built for one team watching its own agent, not for a platform reporting isolated numbers to dozens of separate paying customers at once. AiAgRe traces every tool call an agent makes with per-tenant attribution built in from the start, so the deflection and cost-per-resolution numbers on a customer’s dashboard are scoped to that customer alone, and the underlying trace data backing each number is the same data an AI agent dashboard shows them, not a separately reconstructed estimate. See multi-tenant analytics for what that tagging and access-boundary work actually involves.

Reporting a deflection rate you can actually stand behind

Deflection rate is a real, useful number, but only when it travels with the two things this piece has walked through: a resolution check that catches the customers who gave up instead of getting helped, and a cost calculation built on your genuinely resolved count rather than your raw deflected count. Report deflection alone and you’re reporting volume. Pair it with resolution rate, a reopen-rate check, and a per-tenant cost figure, and you’re reporting something a customer can hold you to.

Start by pulling your own numbers this week: total contacts, deflected contacts, and however many of those came back within 48 hours with the same issue. That reopen count alone will tell you more about whether your deflection rate is real than the deflection rate itself does. If you’re building the kind of white-label AI SaaS product where these numbers eventually show up on a customer’s own dashboard, AI agent monitoring is what keeps that reopen-rate check running continuously in production instead of as a one-time audit, so the number your customer sees stays accurate as their agent keeps handling new conversations. See pricing for how tracing and per-tenant dashboards fit together in practice.

Frequently asked questions

What is deflection rate in customer support?

Deflection rate is the percentage of customer contacts that get resolved without ever reaching a human agent, calculated as deflected contacts divided by total contact attempts and multiplied by 100. It counts anything that avoided a human handoff, whether the underlying issue was actually solved or the customer simply gave up.

What is a good deflection rate for an AI agent?

Published benchmarks put ecommerce support around 55 to 75 percent, SaaS support around 40 to 60 percent, and regulated categories like financial services or healthcare around 25 to 45 percent, with newly launched agents typically starting near 30 to 40 percent and climbing over three to six months. Treat these as rough sanity checks against your own industry rather than a universal target.

What is the difference between deflection rate and containment rate?

Deflection rate counts any contact that never reaches a human, including self-service clicks and abandoned chats. Containment rate narrows that specifically to AI chat conversations that end without an escalation. Neither one confirms the customer’s problem actually got solved, which is what automated resolution rate tries to verify separately.

How do you calculate cost per resolution for an AI agent?

Take the cost difference between an automated contact and a human-handled contact, then multiply it by your genuinely resolved contact count, not your raw deflected count. Subtract out any conversations your reopen-rate check flags as false resolutions first, since those never actually saved the cost of a human contact if the customer had to come back anyway.

Can deflection rate be gamed or inflated?

Yes. Because deflection rate counts any contact that avoids a human, a confusing self-service flow that causes customers to give up rather than escalate will still show up as a high deflection rate. That’s why a reopen-rate check, tracking how often a “resolved” conversation comes back with the same issue within 24 to 48 hours, matters as much as the deflection number itself. A related but sneakier version of the same problem is a hallucinated completion, where the agent claims it resolved something no tool call actually backed up; see AI agent hallucination for how that specific pattern inflates a deflection number without ever showing up in a reopen check.

How do you measure deflection rate for a multi-tenant AI product?

Calculate and report it separately for each customer rather than blending your whole customer base into one average, since a blended number hides which customers the agent is actually working for. Every conversation and tool call needs to be tagged with which customer’s agent produced it from the moment it happens, and one customer should never be able to see another customer’s traces or numbers.

Related reading: AI agent testing covers how to catch a false resolution before it ever reaches a customer-facing number, AI agent hallucination covers the specific false-completion pattern that inflates a deflection number quietly, AI agent monitoring is what keeps a reopen-rate check running continuously once the agent is live, an AI agent dashboard is where the deflection and cost-per-resolution numbers this piece walks through actually get shown to your customers, and escalation rate for AI agents covers the confidence-threshold side of the same reporting problem. See pricing for how AiAgRe’s tracing and white-label dashboards fit together.

Back to Blog