· Prakash Natarajan · Reliability · 16 min read
LLM Agent Frameworks: The Three Axes Nobody Scores
Every LLM agent framework roundup scores developer experience and integration count. Here is what six of them actually give you for tracing, eval, and multi-tenant isolation, with the API-level specifics the lists leave out.

LLM agent frameworks differ less on whether they can build an agent, most of them can, and more on what happens after you ship one: whether it hands you a way to trace a single customer’s calls, whether it gives you any evaluation hooks at all, and how much of your code you rewrite the day you decide to switch. LangChain, CrewAI, LlamaIndex Workflows, the OpenAI Agents SDK, Google’s ADK, and Microsoft’s Agent Framework each answer those three questions differently, and none of the roundups ranking them right now put a number on any of the three. This piece does, with the specific API and configuration details behind each answer, then walks through what changes once the agent you build with any of them has to prove its value to a paying customer rather than just work in a demo.
What Should You Actually Score an LLM Agent Framework On?
Score an LLM agent framework on three things that show up only after launch: whether it gives you per-conversation tracing you can filter by customer, whether it ships an evaluation surface or leaves you to build one from scratch, and how tightly your application code is coupled to that framework’s own abstractions.

Every framework in this piece can call a model, call a tool, and hand control between two agents, that part of the comparison is close to solved and mostly a matter of taste in API design. Tracing is where the real gap opens up. Some frameworks record a full span tree of every model call, tool call, and handoff by default; others log only what you explicitly instrument, which means the first time you need to answer “what did this agent do for customer 4471 last Tuesday” you find out which category yours fell into. Evaluation support ranges from a built-in test suite with simulated users to nothing beyond “bring your own eval library.” And lock-in is the axis nobody talks about because it only costs you something once, the day a framework’s roadmap stalls or a better primitive ships elsewhere and rewriting your orchestration logic turns out to touch half the codebase. None of these three axes shows up in a GitHub star count or an integration-count table, which is exactly why the next six sections walk through each framework against all three, not just the first one.
How Does LangChain’s LangGraph Handle State and Tracing?
LangGraph, the orchestration layer built on top of LangChain, handles state through built-in checkpointing that persists conversation history and context across sessions, and it hands tracing off entirely to LangSmith, the paired observability platform from the same team.

LangChain’s own comparison of its ecosystem describes LangGraph’s state handling as memory that “stores conversation histories and maintains context over time,” paired with native token-by-token streaming that surfaces an agent’s reasoning and actions as they happen rather than only at the end of a run. The trace side lives in LangSmith, which the same page describes as the tool that lets a team “debug every agent decision, eval changes, and deploy in one click,” and that pairing is also where the lock-in shows up: LangGraph’s state model and LangSmith’s tracing schema are designed around each other, so getting the full picture means adopting both pieces, not just the orchestration layer. The upside of that coupling is real. LangChain reports over 1,000 integrations connecting models to data systems and tools, the deepest ecosystem of any framework in this list, and the abstraction layer is deliberately built so your application logic does not have to change when the underlying model does. The cost shows up in debugging: those same abstraction layers, several layers deep in a complex graph, can make an edge-case failure harder to trace back to the exact node that caused it, which is a fair trade for teams that value the ecosystem breadth over a thinner, more inspectable stack. Anyone deciding between LangGraph and a lighter-weight alternative for a specific use case should also read this site’s own CrewAI vs LangChain comparison, which covers where the two diverge in practice.
What Makes CrewAI Different From a Graph-Based Orchestrator?
CrewAI organizes agents around roles, not a graph, with each agent assigned a job title, a goal, and a set of tools, and a “crew” that coordinates how those role-based agents hand work to each other, which makes it faster to prototype than a graph you have to draw by hand.

That role-based mental model is the entire pitch: instead of wiring nodes and edges, a team names a researcher, a writer, and an editor, gives each one a goal, and lets CrewAI’s runtime figure out the handoff sequence, which is a genuinely faster way to get a first working multi-agent setup on the screen. CrewAI also ships Model Context Protocol support out of the box and, per its own materials, integrates with more than 700 external applications, plus a monitoring dashboard and built-in training tools for iterating on agent behavior. The tracing gap is the specific weakness LangChain’s own comparison calls out: CrewAI’s “action traces don’t reflect actual execution” in some configurations, alongside gaps in asynchronous execution and streaming, and friction when the model provider is not OpenAI. That matters more than a generic “less mature” label, because a trace that does not reflect what actually ran is worse than no trace at all, it tells you the agent behaved correctly when it might not have. For a team weighing rapid prototyping against a trace you can actually trust in production, that trade-off is the one to test directly against your own workload rather than take on either vendor’s word.
Is LlamaIndex Workflows Built for Production Agents or Just RAG?
LlamaIndex Workflows extends the framework’s original retrieval-augmented generation roots into event-driven agent orchestration, where steps fire in response to events rather than following a fixed sequence, and it is production-capable but still carries the retrieval-first design of the library it grew out of.

The event-driven model is genuinely useful for agents whose next step depends on what a retrieval or tool call just returned, since a workflow reacts to the event rather than requiring you to pre-plan every branch, and LlamaIndex’s deep ecosystem of data connectors means a retrieval-heavy agent gets that piece essentially for free. The friction shows up in two places LangChain’s own review of the framework names directly: agent handoff failures where a receiving agent stops responding under certain conditions, and tracing problems specifically under concurrent execution, exactly the scenario a multi-tenant product runs constantly once more than one customer’s requests are in flight at the same time. The same review also flags the framework as more boilerplate-heavy than some alternatives for teams building agents that are not primarily retrieval-driven. None of that makes LlamaIndex Workflows a poor choice, it makes it a strong one specifically for retrieval-heavy agents and a weaker fit for a general-purpose multi-agent system, which is the same distinction this site draws out fully in LangChain vs LlamaIndex.
What Does the OpenAI Agents SDK Give You That the Others Don’t?
The OpenAI Agents SDK ships tracing turned on by default, capturing every model generation, tool call, handoff, guardrail check, and custom event as a distinct span in a single run’s trace tree, with no separate observability product required to see it.

That default-on tracing is the SDK’s clearest advantage over the field: generation spans, function-call spans, handoff spans, guardrail spans, even audio operations for voice agents, all get wrapped automatically, and a team can disable it only by explicitly setting an environment variable or calling a dedicated function, which is the opposite default from frameworks that require you to add instrumentation yourself. Traces route to OpenAI’s own backend by default through a batch processor, but the SDK exposes add_trace_processor() to send the same span data to a second destination and set_trace_processors() to replace the default entirely, with more than thirty documented integrations including Weights and Biases, Arize Phoenix, MLflow, Braintrust, LangSmith, and Datadog. The API surface itself is deliberately minimal, with clean handoff primitives between agents that keep a multi-agent setup readable, and Model Context Protocol support comes built in. The gap is durability: the SDK has no native mechanism for a workflow to survive a process crash mid-run and pick back up where it left off, so a team that needs that guarantee has to bolt on an external system like Temporal or DBOS, adding an operational dependency the SDK itself does not manage. For a team that already lives in the OpenAI ecosystem and wants tracing without standing up a separate tool first, that trade is an easy one to accept.
How Does Google’s ADK Handle Evaluation and Enterprise Deployment?
Google’s Agent Development Kit ships built-in evaluation tooling, criteria-based scoring, simulated users, simulated environments, and custom metric definitions, which puts real eval tooling in the framework itself rather than leaving it to a separate product, and it deploys across five languages: Python, TypeScript, Go, Java, and Kotlin.

The evaluation tooling is the standout feature here, since none of the other five frameworks in this piece ship simulated-user testing as a first-party capability, and ADK pairs it with session management that supports rewinding a conversation to an earlier state and migrating sessions between environments, useful for reproducing a bug a customer reported without needing their exact live session. Observability runs on three layers, logging, metrics, and distributed traces, with the trace layer integrating natively with Google Cloud Trace once an agent is deployed to Google Cloud specifically. That last clause is also the catch: the framework’s benefits concentrate heavily on GCP, and LangChain’s own review of ADK notes that non-Cloud environments lose much of what makes the framework distinctive, plus a specific operational issue where session state gets lost if a Cloud Run container restarts mid-conversation, a failure mode worth testing directly before it reaches a customer’s session in production. For a team already committed to Google Cloud, ADK’s evaluation tooling and multi-language support are hard to match elsewhere. For a team on a different cloud, the same framework loses a meaningful share of its advantage.
Where Does Microsoft’s Agent Framework Fit After Merging AutoGen and Semantic Kernel?
Microsoft’s Agent Framework is the unified successor to both AutoGen and Semantic Kernel, built around graph-based workflows with Python and .NET reaching feature parity at version 1.0, and it ships migration assistants specifically for teams moving off either of the two frameworks it replaces.

The consolidation itself is the headline: instead of choosing between AutoGen’s research-oriented multi-agent conversation model and Semantic Kernel’s enterprise orchestration primitives, a team now targets one framework that absorbed both, and Azure AI Foundry’s responsible AI tooling plugs in directly for teams that need policy and safety controls as part of the deployment pipeline rather than bolted on afterward. That consolidation is genuinely useful for any team already invested in Semantic Kernel, and this site’s own Semantic Kernel vs LangChain piece is the place to start if that migration path is the open question. The framework’s own review from LangChain flags two specific issues worth testing before committing: sequential context handling that does not always behave the way a team coming from AutoGen’s async model expects, and function-approval scoping gaps, plus provider adapter coverage that thins out noticeably once you step outside Azure OpenAI as the model provider. For a team already on Azure with Azure OpenAI as the primary model, those gaps mostly disappear. For a team on a mixed-provider stack, they are worth a proof of concept before the migration, not after.
Which Framework Should You Actually Pick for a Multi-Tenant Product?
Pick the framework whose tracing and evaluation model matches how your product actually runs, single global agent or one instance per paying customer, because every framework in this piece was built to help the team that wrote it debug an agent, not to help that team’s own customers see proof their agent is working.
That distinction matters more than any single feature row above. LangGraph and LangSmith give a team deep tracing paired to a wide ecosystem. The OpenAI Agents SDK gives default-on span tracing with minimal setup. Google’s ADK gives genuine evaluation tooling with simulated users. CrewAI gives the fastest path to a first working prototype. But none of the six frameworks compared here were built with a second layer of customers in mind, the paying users of the product a team builds on top of the agent, who need to see their own deflection rate, their own cost per resolution, and their own trace history without seeing any other tenant’s data mixed in. That is a different problem from the one these frameworks solve, closer to what this site covers in tenant isolation for AI agents, and it is the reason a team should treat “which framework” and “how do I prove this agent works to the people paying for it” as two separate decisions rather than one. AiAgRe sits on top of whichever framework a team already picked, tracing what the agent does per tenant and turning that into a dashboard the team’s own customers see, so the framework choice above stays about building the agent, not about building the proof layer on top of it. See pricing for how that layer fits on top of an existing LangGraph, CrewAI, or Agents SDK build.
Frequently asked questions
What is the best LLM agent framework overall?
There isn’t one. LangGraph fits teams that want the deepest ecosystem and are willing to adopt LangSmith for tracing. The OpenAI Agents SDK fits teams that want tracing on by default with minimal setup. Google’s ADK fits teams on GCP that want built-in evaluation tooling. CrewAI fits teams that want the fastest path to a first working prototype. The right one depends on which of the three axes, tracing, evaluation, and lock-in, matters most for the specific product being built.
Does LangGraph require LangSmith to work?
No, LangGraph runs without LangSmith, but the deep tracing and evaluation features come specifically from the LangSmith integration, so a team that skips LangSmith gets the orchestration layer without the observability layer that makes debugging a complex graph practical.
Is the OpenAI Agents SDK tied to OpenAI’s own models?
The SDK’s tracing and handoff primitives are model-agnostic in design, but its tightest integration and default configuration assume OpenAI models, and teams running other providers should confirm current provider support before committing, since that coverage changes as the SDK matures.
Can you switch LLM agent frameworks later without a full rewrite?
Rarely without real cost. The more a team leans on a framework’s specific abstractions, LangGraph’s graph state, CrewAI’s role objects, ADK’s session model, the more application code has to change on a switch, which is why lock-in belongs in the decision up front rather than getting discovered the day a switch becomes necessary.
Do any of these frameworks handle multi-tenant isolation natively?
No. Tenant isolation, keeping one customer’s agent conversations, tool access, and data fully separated from another’s, is an application-layer responsibility in every framework covered here, not something any of the six ship as a built-in guarantee.
How does CrewAI’s tracing compare to LangGraph’s or the OpenAI Agents SDK’s?
CrewAI’s tracing is the least reliable of the three in independent reviews, with reported cases where recorded action traces do not reflect what the agent actually executed, particularly under async or streaming configurations, which is worth testing directly against your own workload before trusting it for production debugging.
Related reading: AI agent vs LLM covers the decision one level up, whether the loop these frameworks structure is even the right call before picking one of them, why multi-agent LLM systems fail in production covers what breaks once a framework’s delegation or hand-off pattern runs against real traffic, and AI agent testing covers evaluating whichever framework you land on before it ships.