· Prakash Natarajan · Reliability · 17 min read

Why Multi-Agent LLM Systems Fail in Production

Multi-agent LLM systems fail for one of three root reasons, not a random assortment of bugs. The MAST taxonomy, what each looks like in a real trace, and where it shows up in a customer's numbers.

Multi-agent LLM systems fail for one of three root reasons, not a random assortment of bugs. The MAST taxonomy, what each looks like in a real trace, and where it shows up in a customer's numbers.

Multi-agent LLM systems fail for one of three reasons, according to the largest public study of the problem so far: a flaw in how the task was specified before anything ran, a breakdown in how the agents coordinated with each other mid-task, or a missing check that should have caught the wrong answer before it shipped. That taxonomy comes from MAST, the Multi-Agent System Failure Taxonomy, built from more than 1,600 annotated execution traces collected across seven popular open-source frameworks. Each of the three failure types leaves a different fingerprint in a live trace, and in a product that reports its own metrics to paying customers, each one eventually surfaces as a specific number: a deflection that should have happened and didn’t, an escalation that fired for the wrong reason, or a resolution that took three times longer than it should have. Naming which of the three you’re actually looking at, before reaching for a fourth agent to patch it, is the difference between a real fix and a more complicated version of the same failure.

Why Do Multi-Agent LLM Systems Fail More Than a Single Well-Built Agent?

A multi-agent system fails more often than a single well-built agent because splitting a task across specialized agents adds an entire new failure surface, coordination between agents, on top of every failure mode a single agent already had, without removing any of them.

why multi-agent llm systems fail

The pitch behind multi-agent design sounds reasonable on paper: give one agent a narrow job (plan the task), give another a narrow job (call the right tools), give a third a narrow job (write the final reply), and each one should be more reliable at its slice than a single generalist agent trying to do all three at once. In practice, that narrowing buys you specialization but costs you a hand-off, and a hand-off is a new place for information to get lost, assumed, or duplicated. A team named Cemri, Pan, Yang, and a group of co-authors (several with UC Berkeley affiliations, including Ion Stoica, Matei Zaharia, and Joseph Gonzalez) set out to measure exactly how often that trade goes wrong, and built MAST-Data, a dataset of over 1,600 execution traces gathered across seven widely used multi-agent frameworks, to do it. Rather than guessing at failure causes from a handful of anecdotes, they had human annotators manually code 150 of those traces by hand, reaching a Cohen’s kappa of 0.88 between annotators, a strong agreement score that means the categories they landed on weren’t one researcher’s opinion but something two independent people, reading the same trace, consistently agreed had gone wrong. That manual pass is what became MAST, and an LLM-as-judge pipeline validated against that same human baseline is what let the team apply it across the full 1,600-trace dataset instead of stopping at 150. The same idea, a small hand-labeled set validating a larger automated one, is the backbone of building a reliable golden dataset for your own agent, and it’s worth noticing that the researchers who study why these systems fail used the exact discipline a team should use to catch failures before shipping.

What Are the Three MAST Failure Categories, and What Causes Each One?

MAST groups every failure into three categories: system design issues, where the task or the agents’ roles were specified badly before anything ran; inter-agent misalignment, where agents miscommunicate or duplicate work mid-task; and task verification gaps, where nobody, human or agent, checked the output before it shipped.

mast taxonomy three failure categories

Those three categories roll up 14 distinct failure modes the researchers identified during their manual coding pass, everything from an agent given an ambiguous role with no clear boundary against a teammate’s role, to a conversation that terminates early because one agent assumed a condition was met that never actually was, to a final answer that passes review simply because nothing in the pipeline was built to review it. System design issues happen before a single agent runs: someone wrote a prompt or a role definition that left real ambiguity in it, and every agent that reads that ambiguous instruction fills the gap with its own assumption, and those assumptions don’t have to match. Inter-agent misalignment happens while the system runs: one agent’s output becomes another agent’s input, and if the receiving agent trusts that input without checking it, or the two disagree about what state the task is actually in, the conversation drifts without either agent noticing. Task verification gaps happen at the end: a multi-agent pipeline produces a final answer, and unless something, a rule-based check, a second model acting as a judge, or a human, actually looks at that answer before it goes out, a wrong one ships exactly as smoothly as a right one. The paper doesn’t publish a clean percentage split of how often each category shows up (treat any specific split you see quoted elsewhere as an unverified secondhand number), but knowing the three buckets exist is enough to start sorting your own incidents into one of them instead of filing every multi-agent bug under the same vague “the agents got confused” label.

What Does a System-Design Failure Actually Look Like in a Trace?

A system-design failure shows up in a trace as two agents both operating on the same assumption about who owns a step, with neither one’s trace actually showing that step executed.

system design failure in an agent trace

Picture a support copilot built with a triage agent and a resolution agent. The triage agent’s role is written as “route the ticket and confirm the account is in good standing,” and the resolution agent’s role is written as “resolve the ticket using the account’s current status.” Read separately, both roles sound complete. Read together, neither one explicitly says which agent is responsible for the actual account-status lookup call, so each one assumes the other already made it. Months later, a trace review turns up a resolved ticket where the final reply confidently states the account is current, but the account-lookup tool was never called by either agent anywhere in that trace. That’s a system-design failure, and it’s specifically the kind that only shows up if you’re logging a real span for every tool call an agent makes, not just the final merged reply the customer sees. What to trace first for a single agent still applies here, but a multi-agent system adds one more requirement on top of it: the trace needs to show which agent owned each step, not just that the step happened somewhere in the pipeline.

How Does Inter-Agent Misalignment Show Up in a Deflection Rate?

Inter-agent misalignment shows up in a deflection rate as a conversation that looks resolved in the transcript but quietly reopens later, because one agent closed the ticket based on a state a second agent hadn’t actually reached yet.

inter-agent misalignment and deflection rate

Take a two-agent setup where one agent drafts the customer-facing reply and a second agent, running slightly behind it, is still fetching the order details the reply is supposed to reference. If the drafting agent doesn’t wait on a confirmed signal from the fetching agent, and instead proceeds once a timeout or a partial response arrives, it can ship a reply that references order details that were never actually confirmed correct. The conversation gets marked as deflected the moment that reply goes out, because from the deflection rate formula’s point of view, a contact that never reached a human counts as deflected regardless of whether the answer inside it was right. The customer reads a wrong detail, opens the ticket again the next day, and now the same contact counts twice against a metric that was supposed to measure success the first time. This is exactly why a deflection number on its own, without a resolution check sitting behind it, can quietly reward a race condition between two agents instead of catching it. A trace that timestamps each agent’s individual completion, not just the moment the combined reply went out, is what turns that silent reopen into a visible pattern instead of a mystery a support lead notices three weeks later in the reopen-rate report.

What Happens When Nobody Verifies the Final Output?

When no independent step verifies a multi-agent system’s final output, a wrong or fabricated answer ships with exactly the same confidence as a correct one, because nothing in the pipeline was ever built to tell the two apart.

task verification gap in multi-agent output

This is the same confident-but-wrong shape covered in more depth in AI agent hallucination, and a multi-agent pipeline makes it easier to fall into rather than harder, because each individual agent in the chain can look reasonable in isolation while the combination still produces a wrong final answer. A planning agent picks a plausible-sounding approach, a tool-calling agent executes it without flagging that the approach didn’t quite fit the request, and a writing agent turns the result into fluent prose, and at no point does anything in that chain ask “is this actually right.” The cost of that gap rarely shows up as one dramatic incident. It shows up as a slow drift upward in how often conversations bounce back for a second attempt, which is exactly the pattern covered in escalation rate for AI agents: a system with no verification step doesn’t necessarily escalate more, it resolves wrong more, and each of those wrong resolutions that eventually gets caught, whether by a frustrated customer or a downstream process, adds a second and sometimes third pass onto a ticket that should have taken one. Adding a verification step doesn’t have to mean a full second LLM judging every response. A deterministic assertion (did the tool call that was supposed to run actually run, does the final answer reference a value that exists in a verified tool result) catches a meaningful share of this category for a fraction of the cost of a model-based check, and it’s worth building that cheap check before reaching for a more expensive one.

Why Does One Tenant’s Coordination Failure Become Another Tenant’s Incident?

A coordination failure becomes another tenant’s incident when the agents involved share infrastructure across tenants, a queue, a memory store, a shared rate limit, so one tenant’s stuck loop or malformed hand-off consumes capacity or leaks context meant for someone else.

multi-tenant blast radius of a coordination failure

This is the failure category that a single-tenant discussion of MAST never has to consider, and it’s the one a B2B2C AI SaaS product can’t afford to ignore. If a coordinator agent pulls the next task off a shared processing queue without a tenant boundary enforced at that queue, one tenant’s misaligned agents retrying the same failed step in a loop can starve every other tenant’s requests behind it in line. If two agents share a memory store scoped by session rather than by tenant, a coordination bug that causes one agent to read the wrong context can, in the worst case, pull in a fragment that belongs to a different customer’s conversation entirely. Neither of these is a hypothetical edge case once you’re running the same multi-agent pipeline for more than one paying customer at a time, and neither shows up in a taxonomy built from single-tenant academic benchmarks, because academic benchmarks don’t have tenants. Tenant isolation for AI agents covers the database-level version of this problem in depth, and the multi-agent version adds one more layer on top: the isolation boundary has to hold not just for stored data, but for whatever shared queue, cache, or rate limit sits between your agents while they’re actively coordinating on someone’s live conversation.

How Do You Catch These Three Failure Types Before a Customer Does?

You catch each of the three failure types with a different check, because a single test suite tuned for one category will walk right past the other two: a specification review before anything runs, a per-hand-off trace while agents are coordinating, and an independent verification step before a reply ships.

catching multi-agent failures before shipping

For system design, the check happens before deployment: read every agent’s role definition side by side, the way you’d review a set of API contracts, and look specifically for a responsibility that two roles both seem to claim, or that neither one clearly claims. For inter-agent misalignment, the check happens in the trace itself: log each agent’s hand-off as its own span with its own timestamp and its own completion signal, so a race condition between two agents is something you can point to in a trace viewer after the fact, not something you have to reconstruct from a confused customer’s description of what happened. For task verification, the check happens at the exit: a deterministic assertion or a separate judge call sitting between the last agent’s output and whatever counts that conversation as resolved, so a wrong answer has to clear one more gate before it reaches a customer-facing number. None of these three checks replace the others, and running only one of them will catch exactly one category of failure while leaving the other two invisible. Once you’ve caught a handful of real failures from each category, turning them into a golden dataset of known multi-agent failure traces, tagged by which MAST category they belong to, gives you something to regression-test the next version of your pipeline against, instead of finding out the same coordination bug came back the hard way.

Naming the Failure Type Is the Fix, Not Adding a Fourth Agent

The instinctive first response to a multi-agent bug is usually to add another agent: a supervisor to watch the others, a critic to double-check the output, a router to catch a bad hand-off before it lands. Sometimes that’s the right move, but often it just adds a fourth agent’s worth of new coordination surface on top of a problem that was never a coordination problem in the first place. A system-design failure needs a clearer specification, not a supervisor watching agents follow an ambiguous one. A verification gap needs one deterministic check, not a second LLM agent that inherits the same blind spots as the first. Sorting the failure into one of MAST’s three categories before reaching for a fix is what tells you which of those two situations you’re actually in. AiAgRe traces every agent-to-agent hand-off as its own span, tagged with tenant identity from the first request, so when a multi-agent conversation breaks, the trace shows exactly which boundary it broke at, a bad spec, a missed hand-off, or a missing check, instead of leaving you to guess which of the three it was from a transcript alone.

Frequently asked questions

What is the MAST taxonomy?

MAST, the Multi-Agent System Failure Taxonomy, is a classification system built from over 1,600 annotated execution traces across seven multi-agent frameworks, sorting every failure into one of three root categories: system design issues, inter-agent misalignment, and task verification gaps.

How many distinct failure modes does MAST identify?

The researchers identified 14 distinct failure modes during their manual coding pass, which roll up into the three broader categories. The 150-trace manual analysis that produced them reached a Cohen’s kappa of 0.88 between human annotators, a strong agreement score for a qualitative coding task.

Does adding more agents to a pipeline make it more reliable?

Not by itself, since splitting a task across more agents adds specialization, but it also adds hand-offs, and each hand-off is a new place for a system-design gap or a coordination failure to happen. More agents without a clearer specification and a verification step usually just adds more surface for the same three failure categories to show up on.

What’s the difference between a system-design failure and an inter-agent misalignment failure?

A system-design failure exists before anything runs, baked into an ambiguous role or task specification that every agent interprets differently. An inter-agent misalignment failure happens while the system is running, when one agent’s output becomes another agent’s input and the two lose sync about what state the task is actually in.

Can a single verification step catch all three MAST failure categories?

No, because a verification step at the end catches a wrong final answer, but it won’t catch a system-design ambiguity that never surfaces as an obviously wrong answer, and it won’t catch a race condition between two agents that happened to still produce a plausible-looking reply. Each category needs its own check.

How do you test for multi-agent failures before shipping a new version?

Build a small dataset of real traces where each of the three failure types actually happened, tagged by category, and run your next version’s traces against it before rollout. That’s the same golden-dataset discipline the MAST researchers used to validate their own automated classifier against a human-coded baseline.

Related reading: multi-agent system architecture covers the orchestrator-worker and hierarchical patterns whose hand-off points are exactly where these failures start, AI agent observability covers what to trace first for a single agent before you add the multi-agent hand-off layer on top, AI agent hallucination covers the confident-but-wrong failure shape a missing verification step lets through, deflection rate and escalation rate cover the two customer-facing numbers a coordination failure quietly distorts, tenant isolation for AI agents covers the shared-infrastructure boundary a multi-tenant coordination failure can cross, and how to build a golden dataset covers turning real failure traces into a regression check for the next release.

Back to Blog