· Prakash Natarajan · Reliability · 16 min read

Multi-Agent System Architecture: Patterns That Ship

Four multi-agent architecture patterns show up in production again and again. What each one costs in tokens, when a single agent is still the right call, and what breaks first once a second tenant shows up.

Four multi-agent architecture patterns show up in production again and again. What each one costs in tokens, when a single agent is still the right call, and what breaks first once a second tenant shows up.

Multi-agent system architecture is the design decision for how you split one job across more than one LLM agent and coordinate what each one does: a central orchestrator that delegates to subagents, a skills-based agent that loads a specialized prompt on demand, a handoff chain where control passes between agents as state changes, or a router that fans a request out to several agents in parallel and merges their answers. Each pattern trades a different amount of latency, token cost, and coordination risk for a different kind of reliability, and the pattern that wins for a research assistant answering one-off questions is not the pattern that wins for a support copilot handling the same conversation for months. The teams that get this right pick the architecture the failure data justifies for their specific workload, not the one that reads best in a conference talk, and they build in a tenant boundary from the first line of orchestration code rather than retrofitting one after the first customer complaint.

What Are the Main Multi-Agent System Architecture Patterns?

Four patterns cover almost every production multi-agent system: a centralized orchestrator delegating to subagents, a skills-based agent that loads specialized prompts on demand instead of spinning up separate agent instances, a handoff chain where the active agent changes as the conversation’s state changes, and a router that dispatches a request to several agents in parallel and synthesizes their results.

multi agent system architecture patterns

LangChain’s engineering team laid these four out with a useful side-by-side comparison, and the numbers behind each one are worth sitting with before picking one. In the centralized orchestration pattern, subagents stay stateless and run behind strong context isolation, but a single request still needs four model calls where the other three patterns need three, because every result has to flow back through the orchestrator before it reaches the user. On a repeat request, the stateful patterns, skills and handoffs, save 40 to 50 percent of that cost because they carry context forward instead of rebuilding it from scratch each time. On a multi-domain query that spans more than one specialized area, the subagent and router patterns process 67 percent fewer tokens than the skills pattern, because parallel dispatch avoids loading every specialized prompt into one growing context window. None of the four is a universal winner. A support copilot handling a long-running account issue benefits from the stateful savings of handoffs; a one-shot research query that touches several unrelated domains benefits from the router’s parallel fan-out; and a workload where you need a hard audit trail of who decided what benefits from the orchestrator pattern’s extra call, because that call is also the place a verification check naturally lives.

When Should You Move From a Single Agent to a Multi-Agent System?

Move from a single agent to a multi-agent system only once a single agent, with tools added first, has genuinely hit a limit it can’t solve by better prompting or better tool descriptions, because a multi-agent system costs meaningfully more to run and coordinate than the single-agent baseline it replaces.

single agent versus multi agent system

Anthropic’s own engineering write-up on the multi-agent research system they built is unusually direct about the tradeoff: a lead agent with several subagents outperformed a single well-tuned agent by 90.2 percent on their internal research evaluation, but that same multi-agent setup used about 15 times more tokens than a plain chat interaction, and in their own testing, token usage alone explained roughly 80 percent of the variance in how well the system performed on a public research benchmark. That is a real performance gain bought with a real cost increase, and the team’s own recommendation, before reaching for a second agent, is to start with a single agent, add tools to it, and only split it into a multi-agent system once you’ve hit a limit tools alone can’t fix. For a support copilot or a sales assistant built on top of a fixed per-conversation budget, that 15x multiplier is not an abstract number, it is the difference between a cost-per-resolution figure your customer will accept and one that erases the margin the whole product depends on. The right test before adding a second agent is concrete: can you name the specific task type a single agent keeps failing at, and can you show that better tool descriptions or a longer system prompt didn’t fix it? If you can’t answer both, the fix is probably not architecture yet.

How Does the Orchestrator-Worker Pattern Actually Work in Production?

The orchestrator-worker pattern works by giving one lead agent the job of reading the request, breaking it into a small number of parallel subtasks, spawning a subagent for each one, and then synthesizing the subagents’ individual results into one final answer, with the lead agent as the only point that talks to the user directly.

orchestrator worker pattern in production

Anthropic’s write-up describes scaling this deliberately by query complexity: a simple factual question gets one agent making three to ten tool calls, while a genuinely complex comparison might spin up ten or more subagents each working a different angle in parallel. Two things separate a production version of this pattern from a demo. First, tool descriptions matter more than most teams expect: the same engineering team found that a tool-testing agent that rewrote its own tool descriptions after diagnosing failures produced a 40 percent decrease in task completion time, which means a vague tool description is not a minor annoyance, it actively sends subagents down the wrong path. Second, the orchestrator has to guard against a specific failure this pattern invites, subagents that keep working past the point of diminishing returns, spawning more of themselves for a query simple enough to need one, or grinding through search queries long after they had enough to answer. The fix isn’t a supervisor watching the watchers, it’s an explicit stopping condition in the orchestrator’s own prompt: a stated budget of calls or subagents per complexity tier, checked against actual usage before the orchestrator lets a task keep running. Get the stopping condition right and the pattern scales cleanly with query difficulty. Skip it and the same architecture that handled a hard query efficiently will burn the same budget on an easy one nobody needed spent.

What Does a Handoff-Based Architecture Look Like for a Support Copilot?

A handoff-based architecture changes which agent is active as the conversation’s state changes, passing control from one specialized agent to the next through a tool call that also carries the accumulated state forward, so a triage agent can hand a ticket to a billing agent without either one rebuilding context the other already gathered.

handoff based architecture for a support copilot

This is the pattern that maps most directly onto a real support product, because a support conversation is naturally staged: intake and classification, then a specialist step (billing, technical, account access), then a close-out step that confirms resolution and logs the outcome. LangChain’s comparison of the four patterns found handoffs save 40 to 50 percent of the token cost on a repeat interaction precisely because the receiving agent inherits state instead of re-deriving it, which matters a great deal on a support workload where the same customer often comes back within the same session. The design risk specific to this pattern is the handoff boundary itself: if the triage agent’s role definition and the billing agent’s role definition both implicitly assume the other one already confirmed the customer’s account status, neither one actually does it, and the conversation resolves confidently on an unconfirmed assumption. That exact failure shape, and how to catch it in a trace before a customer does, is covered in more depth in why multi-agent LLM systems fail in production. The practical fix here is narrow: every handoff should carry an explicit, named payload of what the receiving agent can assume is already true, not an implicit “the previous agent probably handled it.”

How Do You Keep a Multi-Agent Architecture Safe Across Tenants?

You keep a multi-agent architecture safe across tenants by tagging every agent, every handoff, and every shared resource, queue, cache, or rate limit, with the tenant identity from the first request, so a coordination bug in one tenant’s conversation can’t consume capacity or read context that belongs to a different customer.

tenant safe multi agent architecture

Neither the LangChain comparison nor Anthropic’s research-system write-up has to think about this problem, because both describe a single team running the architecture for its own internal use. A B2B2C AI SaaS product doesn’t get that luxury: the same orchestrator, handoff chain, or router pattern is running the same code path for dozens or hundreds of paying customers at once, and the architecture decision that looked purely about latency and token cost in isolation becomes a security boundary the moment a second tenant shares any part of it. A shared task queue without a tenant partition lets one tenant’s misbehaving agent, stuck retrying a failed step in a loop, starve every other tenant’s requests behind it. A memory store scoped by session instead of by tenant can let a coordination bug pull a fragment of one customer’s conversation into another customer’s context window. Tenant isolation for AI agents covers the database and memory version of this boundary in depth, and the architecture layer adds one more requirement on top: whichever of the four patterns you pick, the orchestrator’s own state, the handoff payload, and the router’s dispatch queue all need the tenant identity baked in from the first call, not added once the first cross-tenant leak gets reported. AiAgRe’s SDK tags every span with tenant identity at ingestion for exactly this reason, so the architecture choice and the tenant boundary are the same decision instead of two separate projects.

How Do You Know Whether the Multi-Agent Architecture Is Actually Working?

You know a multi-agent architecture is actually working when its cost-per-resolution and deflection numbers hold steady as conversation volume grows, not just when a demo query returns a good-looking answer, because the architecture’s real cost shows up in aggregate token spend and coordination failures that a single test case never surfaces.

measuring multi agent architecture performance

Anthropic’s evaluation approach for their own research system is a useful floor: start with a small sample, around twenty queries, for fast iteration, use an LLM-as-judge rubric scored against factual accuracy, citation accuracy, completeness, source quality, and tool efficiency, and treat a single consistent judge call as more reliable than an ensemble for this kind of scoring, while keeping a human review pass in place because it catches failure modes an automated judge misses (their own example was a judge favoring high-volume content over more authoritative sources, a bias a human reviewer caught immediately). That’s a solid starting point for correctness, but it stops short of the number that actually matters to a B2B2C product: what a resolved conversation cost to produce, and whether that cost holds up once the same architecture runs for a real tenant’s real volume instead of a twenty-query eval set. Cost-per-resolution and deflection rate is the pairing that catches what a pure correctness eval misses, a coordination inefficiency that doesn’t produce a wrong answer, just a needlessly expensive right one, four model calls where three would do, or a router pattern re-loading context a handoff pattern would have carried forward for free. Tie the architecture choice to that number, tracked per tenant so one customer’s heavier workload doesn’t quietly distort another’s, and the pattern that looked cheapest in an internal benchmark either holds up under real traffic or gets replaced by one that does.

Pick the Architecture the Failure Data Justifies

The instinct once a single agent starts struggling is to reach for the most sophisticated pattern available, a full orchestrator with several specialized subagents, because it sounds like the serious answer. Often the handoff pattern, or even a single agent with a better tool description, solves the same problem for a fraction of the token cost, and the only way to know which is true for your workload is to measure the failure you’re actually trying to fix before picking the architecture meant to fix it. A support copilot with a staged workflow usually wants handoffs. A research or discovery task that genuinely benefits from parallel exploration usually wants an orchestrator with subagents. A workload that mostly needs a different specialized prompt per request, without the overhead of a second agent instance, usually wants the skills pattern. None of the four patterns is free, and none of them is safe by default the moment more than one tenant shares the system running it. AiAgRe traces every agent, every handoff, and every dispatch with tenant identity attached from the first request, so whichever of these patterns you choose, the cost-per-resolution number and the coordination trace behind it are tied to the same tenant-scoped record from day one instead of two systems you have to reconcile after the fact.

Frequently asked questions

What is multi-agent system architecture?

Multi-agent system architecture is the design of how a task gets split across more than one LLM agent and how those agents coordinate, commonly one of four patterns: centralized orchestration with subagents, a skills-based single agent that loads specialized prompts on demand, a handoff chain that passes control as conversation state changes, or a router that dispatches to several agents in parallel and merges the results.

Is a multi-agent system always better than a single agent?

No. Anthropic’s own engineering team recommends starting with a single agent and adding tools before splitting into multiple agents, since a multi-agent setup can use roughly 15 times more tokens than a single-agent chat interaction for the same task. The performance gain has to be real and measured, not assumed, before the added cost is worth it.

What is the orchestrator-worker pattern?

The orchestrator-worker pattern uses one lead agent to break a request into subtasks, spawn a subagent for each one, and synthesize the subagents’ results into a single final answer. It scales well with query complexity when the orchestrator enforces an explicit budget of calls or subagents per task, and scales poorly when subagents are left to keep working past the point of diminishing returns.

What is a handoff-based multi-agent architecture?

A handoff-based architecture passes control between specialized agents as a conversation’s state changes, with each handoff carrying accumulated state forward so the receiving agent does not have to rebuild context from scratch. It saves 40 to 50 percent of token cost on repeat interactions compared to a stateless pattern, which makes it a strong fit for staged workflows like support triage followed by a specialist step.

How do you keep a multi-agent architecture safe when it serves more than one customer?

Tag every agent, handoff, and shared resource, queues, caches, rate limits, with the tenant identity from the first request. Without that boundary, one tenant’s coordination bug in a shared queue or memory store can consume capacity or expose context that belongs to a different tenant entirely.

How do you measure whether a multi-agent architecture is worth its added cost?

Track cost-per-resolution and deflection rate per tenant as real traffic scales, not just correctness on a small evaluation set. An architecture that scores well on twenty test queries can still lose money in production if it takes more model calls per resolution than a simpler pattern would need for the same outcome.

Related reading: why multi-agent LLM systems fail in production covers the specific coordination failures each of these patterns can introduce once they’re live, tenant isolation for AI agents covers the database and memory boundary the tenant-safety section here builds on, deflection rate covers the cost-per-resolution formula referenced in the measurement section, AI agent observability covers what to trace first regardless of which architecture pattern you pick, golden dataset for AI agents covers building the evaluation set this piece’s measurement section depends on, and AI agent multi-tenant analytics is the product page for the tenant-tagged tracing this piece describes.

Back to Blog