· Prakash Natarajan · Reliability · 16 min read

AI Agent Testing: What to Check Before Customers Do

AI agent testing catches tool failures and bad outputs before they land in the deflection and cost numbers you report to customers. Here is what to test first, with real failure examples.

AI agent testing catches tool failures and bad outputs before they land in the deflection and cost numbers you report to customers. Here is what to test first, with real failure examples.

AI agent testing means running an agent against a set of real and adversarial scenarios before it ever reaches a customer, checking not just whether it produces the right final answer but whether every tool call, handoff, and piece of reasoning behind that answer was actually correct. For a support copilot or a sales assistant, that distinction matters more than it sounds, because an agent can complete a task the wrong way and still read as flawless in the transcript. That means the deflection rate or resolution rate you show your own customers can be built on top of failures nobody ever caught. This piece covers the failure modes worth testing for, how to build a small dataset that actually catches them, when to reach for a code-based check versus a model acting as the judge, and the extra question multi-tenant products cannot skip: making sure a failure in one customer’s agent run never touches another customer’s data.

What actually breaks when an AI agent runs in production?

Most agent failures never show up as an error message. They show up as a fluent, confident response that quietly did the wrong thing, which is exactly why they are so easy to miss in a quick read of the transcript.

ai agent testing failure modes in a live trace

The clearest example is an agent that claims an action happened without ever calling the tool that would have made it happen: it tells the user their subscription was cancelled, or their ticket was escalated, and the polished, confident tone of the reply hides the fact that no API call was ever made. You only catch this by reading the actual trace, the ordered record of every tool call and tool response behind the reply, not by scoring the final message on its own. A close cousin is the tool call that fires but with the wrong parameters, so it technically succeeds while doing something slightly different from what the user asked for, and the agent narrates the result as if it went exactly to plan.

Two other patterns show up constantly once you start looking. The first is the interrogation loop, where the agent asks the user for information they already gave two messages earlier, because the agent lost track of context rather than because the user was unclear. It reads as politeness in the transcript and as a broken memory when you look closer. The second is the budget burner: a task that finishes correctly but only after a dozen extra reasoning steps or tool calls that never needed to happen, quietly multiplying your cost per resolution without changing the final answer at all. None of these show up if the only thing you evaluate is whether the last message sounds right, which is the single biggest gap in how most teams start testing an agent.

How do you build a test set that actually catches those failures?

You build it by writing a small number of real examples yourself before you automate anything, because a test set assembled entirely from unreviewed synthetic data teaches you very little about how your specific agent actually fails.

building a golden dataset for ai agent evaluation

Start with five to ten hand-written examples for each capability your agent has, covering the ordinary happy path, one clearly out-of-scope request, and one adversarial attempt to get the agent to say or do something it shouldn’t. Write these from real conversations if you have any, or from your own honest guess at how a first-time user would actually phrase things, not from a list of tidy textbook questions. Once you’re running this small set regularly and you’ve watched it catch a handful of real regressions, grow it to twenty or thirty cases per capability, adding any production conversation that broke the agent in a way your existing cases didn’t cover. That growth path matters more than the exact number you start with: a small, honestly curated set that gets reviewed weekly beats a large set generated once and never revisited.

Keep every example paired with an expected outcome you can actually check, not just an expected reply. For a tool-calling agent that means recording which tool should be called, with which arguments, and what the end state should look like afterward, so a passing test means the agent did the right thing and not just that it said something plausible. Synthetic data generation skips this step most often, and skipping it is exactly why a dataset built that way misses the failures described in the previous section.

Should you use code-based checks or a model as the judge?

Use both, because they catch different things. Code-based checks are cheap, fast, and completely reliable for anything you can define exactly, while a model grading the output is the only practical way to score something as subjective as tone or a partially correct answer. If you’d rather not hand-roll either, the established AI agent evaluation tools cover both approaches out of the box.

comparing code assertions and llm judge scores for ai agent testing

Reach for a code-based assertion whenever the correct outcome has a hard, checkable shape: the right tool was called, with the right arguments, a forbidden phrase never appears in a compliance-sensitive reply, or the output matches a schema your downstream system expects. These checks run in milliseconds and never disagree with themselves, so put as much of your suite here as the task allows. Reach for a model as the judge only once you’ve exhausted what a deterministic check can cover: whether the agent addressed the user’s real question, whether a claimed resolution is genuinely supported by the trace, or whether tone stayed appropriate through a frustrated exchange.

A judge model is only trustworthy once you’ve calibrated it. Take ten to twenty transcripts, grade them yourself first, then run the judge on the same set and check whether its scores agree with your own read closely enough to trust. If they don’t, the problem is almost always a vague grading prompt rather than a bad model, so tighten the rubric until agreement holds. Recalibrate any time you change the underlying agent model or materially rewrite its instructions, because a judge tuned against one model’s phrasing habits can quietly drift out of sync with a different one.

How does agent testing protect the ROI numbers you show customers?

Every failure mode above turns into a wrong number somewhere in your reporting, not just a bad user experience, because a deflection rate or resolution rate is only as honest as the test coverage behind the agent producing it.

ai agent testing tied to customer facing roi metrics

Walk the chain through: an agent that claims a task completed without ever calling the tool inflates your resolution count directly, because whatever logic marks a conversation “resolved” almost certainly trusted the agent’s own closing message. A tool call fired with slightly wrong arguments still returns a plausible-sounding answer, so it gets logged as a successful deflection even though it solved the wrong problem. The cost side moves just as quietly: a budget-burning run with a dozen unnecessary extra steps doesn’t fail, so nothing flags it, but it raises your true cost per resolution every time it happens, unnoticed until the bill arrives. None of this shows up if your evals only score whether a reply sounds right, which is why the testing work above has to connect to the reporting layer, not live in a separate QA process.

The fix is to tie a handful of your evals directly to how a conversation actually gets classified for reporting, not only to whether the final reply reads well. If your product marks a conversation “resolved” based on a closing-intent classifier or an explicit tool call, write a test that checks the classifier only fires when the underlying trace genuinely supports it, and run that test on the same cadence as everything else. This is the layer AiAgRe’s tracing sits underneath: because every event is captured at the trace level and tied to the same deflection, resolution, and cost-per-resolution definitions your customer-facing dashboard renders, a test that catches a ghost action in the trace is the same test that keeps your reported ROI number honest. Testing an agent and proving its ROI to your own customers stop being two separate projects once the eval suite is reading from the same trace your dashboard reads from.

What do you need to test before a white-label or multi-tenant rollout?

You need to test that a failure in one customer’s agent run can never leak into another customer’s traces, dashboard, or aggregate metrics, which is a distinct question from whether the agent itself works.

testing tenant isolation before a multi tenant ai agent rollout

Run your entire eval suite under two separate fake tenant identities and confirm, explicitly, that neither tenant’s results, traces, or dataset entries ever appear in the other’s view. This sounds obvious until you’ve watched a filtered-after-the-fact dashboard leak a stray row because the filter lived in the display layer instead of at data capture, a common mistake in systems that added multi-tenancy after the fact. Test the auth boundary the same way: a token scoped to one customer’s data should fail cleanly against another customer’s records, never silently return them.

The second thing worth testing, and the one teams skip most often, is whether your metric definitions stay identical no matter which customer’s dashboard is rendering them. If “resolved” means something subtly different depending on which tenant’s configuration is active, the number stops being comparable across your own customer base. Write a test that runs the same conversation through two different tenant configurations and checks that a resolution gets classified the same way both times, unless the tenant genuinely customized that logic on purpose. Getting this right before a white-label rollout is far cheaper than explaining to two customers why their numbers were computed differently, and it’s the same isolation problem a multi-tenant analytics layer has to solve by default, not patch in after a leak.

How often should evals run once the agent is live?

Run the full suite on every code change inside CI, and run a smaller, sampled slice continuously against live production traffic, because pre-release testing alone will always miss the failures your actual users eventually find.

sampling live production traffic for continuous ai agent evaluation

In CI, every pull request that touches the agent’s prompt, tools, or model should trigger your full code-based suite plus your calibrated judge checks, so a regression gets caught before it ships rather than after a customer notices. In production, running the judge model against every conversation gets expensive fast and rarely buys much extra signal, so a common starting point is sampling five to ten percent of live conversations, weighted toward longer conversations, ones a user rated poorly, or ones that used a recently shipped tool. Alert on the aggregate trend, not any single flagged conversation: one bad sample is noise, a sustained drop across a day of sampled traffic is signal.

Close the loop weekly. Pull the production conversations your sampled evals flagged, review the ones that genuinely represent a new failure mode, and add the clearest two or three straight into your hand-curated dataset from the second section. This is the habit that separates a team that tests once before launch from one that keeps its ROI numbers trustworthy for as long as the product is live.

What to test this week if you have not started

Write five to ten examples for your agent’s most-used capability, covering the happy path, one out-of-scope request, and one adversarial attempt, and pair each with a checkable expected outcome rather than just an expected reply. Add two or three code-based assertions on top of that set for anything with a hard, checkable shape, like the right tool getting called with the right arguments. Add one calibrated judge check for whatever is genuinely subjective in your agent’s output, grading ten transcripts yourself first so you know the judge agrees with your own read before you trust it. Then take the single most important step most teams skip entirely: tie at least one of those checks directly to the logic that marks a conversation resolved or deflected in your reporting, so the eval suite and the ROI number you show customers are reading from the same evidence. If your product is multi-tenant or white-label, run the whole thing under two separate fake tenants before any of it ships customer-facing, and confirm neither one can see the other’s data.

None of this needs a large team or an enterprise QA budget. It needs a handful of honest examples, a couple of hours calibrating one judge, and a habit of feeding real production failures back into the set every week. If you’re building the kind of AI SaaS product where your own customers will eventually see the deflection rate and cost-per-resolution numbers your agent produces, AiAgRe ties that same trace data straight into a white-label dashboard, so the number your test suite is protecting is the same number your customer ends up looking at.

Frequently asked questions

What is AI agent testing?

AI agent testing is the practice of running an AI agent against a curated set of real and adversarial scenarios to check that its tool calls, reasoning steps, and final outputs are correct, not just that the final reply reads well. It covers both automated checks run in CI before every release and continuous sampling of live production conversations after the agent ships.

How many test cases do you need before shipping an AI agent?

Start with five to ten hand-written examples per capability, covering the happy path, an out-of-scope request, and an adversarial attempt. Grow that to twenty or thirty cases per capability once the smaller set has caught a few real regressions and you’ve started folding in genuine production failures, rather than trying to write a large dataset up front.

Can AI agent testing be fully automated?

Most of it can. Code-based assertions and a calibrated LLM-as-judge check can run unattended in CI and against sampled production traffic. What cannot be fully automated is calibration itself: a human still needs to grade a small set of transcripts periodically to confirm the judge model’s scores still agree with real human judgment, especially after changing the underlying agent model.

What is the difference between an eval and a test for an AI agent?

In practice the terms overlap heavily and most teams use them interchangeably. A “test” more often implies a pass-or-fail check run in CI before a release, while an “eval” more often implies a broader score, sometimes graded by a model, that gets tracked over time. AI agent testing in this piece covers both, because a healthy suite needs the hard pass-or-fail checks and the graded, trend-tracked evals together.

How do you test a multi-tenant AI agent without mixing up customer data?

Run your full eval suite under at least two separate fake tenant identities and confirm neither tenant’s traces, results, or dataset entries ever surface in the other’s view. Also test the access boundary directly, by deliberately requesting one tenant’s data with a token scoped to another tenant and confirming it fails cleanly instead of silently returning the wrong records.

How often should you re-run AI agent evals after shipping?

Run the full suite on every code change that touches the agent, inside CI. Once live, sample somewhere between five and ten percent of production conversations continuously, weighted toward longer conversations and recently changed tools, and review the flagged results on a weekly cadence to pull genuine new failure modes back into your hand-curated dataset.

Related reading: once your agent’s behavior is covered by tests, AI agent monitoring is what watches those same traces continuously in production, and an AI agent dashboard is where the deflection, resolution, and cost numbers this testing protects actually get shown to your customers. AI agent benchmarks covers the gap between a lab score like AgentBench or tau-bench and the production failures this piece’s test set is built to catch. See pricing for how AiAgRe’s tracing and white-label dashboards fit into your stack.

Back to Blog