CrewAI Reliability

CrewAI observability for crews that have to be right

CrewAI makes multi-agent orchestration easy to write and hard to verify. A crew that delegates in a loop, truncates at max_iter, or hands back a task output in the wrong shape still returns a confident string and exits zero. Getting a crew to production means tracing the run structure, every agent, task, delegation, and tool call, and then turning those traces into numbers your team and your customers can act on.

The problem

A crew that fails still returns a string

Single-agent failures are relatively easy to spot: one prompt, one answer, and the answer is wrong. Crews fail differently. The output is assembled from several agents, each with its own role, tools, and iteration budget, and a defect in any one of them gets smoothed over by the agents downstream. The final answer reads fine. The failure is upstream and invisible.

Four failure modes account for most of what teams hit once a crew leaves the demo:

Delegation cycles. With allow_delegation enabled, agent A asks agent B, which asks A again. Nothing crashes. Tokens and wall-clock time climb until an iteration cap or a rate limit ends the run somewhere arbitrary.

Silent truncation. An agent that exhausts max_iter doesn't raise; it answers with whatever context it accumulated. That is a partial result presented as a complete one, and it is indistinguishable from success unless you recorded the iteration count.

Output contract drift. A task declares an expected_output, or an output_pydantic model, and the model returns prose that almost fits. Parsing succeeds often enough that the failures look random rather than systematic.

Swallowed tool errors. When a tool raises, CrewAI feeds the error back to the LLM so it can recover. Often it does. Sometimes it invents a plausible value instead, and the fabricated value flows into the next task as if it were real data.

What you already have

CrewAI's own tracing already exists. It stops at your engineers.

CrewAI ships tracing under the name AMP. Run crewai login once, set tracing=True on your Crew or Flow (or the equivalent CREWAI_TRACING_ENABLED=true environment variable), and every agent decision, task execution, tool call, and LLM request shows up in a dashboard at app.crewai.com with no other code change. For a developer chasing down why one run went wrong, that's genuinely the right first step, and it costs nothing to turn on.

What AMP doesn't do is turn those traces into a number anyone outside engineering can act on. It has no concept of a resolution, a deflection, or a cost you could defend to a paying customer, and no notion of a customer at all, so there's no way to scope a view so customer A sees only customer A's crew runs. That's not a gap in AMP; it just answers a different question. AMP tells your team what a crew did. It was never built to tell your customer what their agent is doing for them, and it has no multi-tenant model to do it safely even if it tried.

Most teams running CrewAI in production end up with both layers: AMP (or an OTel exporter pointed at any tool that reads GenAI semantic conventions) for engineering debugging, and something downstream of the same traces for the outcome metrics and the customer-facing dashboard. The rest of this page is about that second layer.

What to capture

Six things worth tracing in a CrewAI run

Instrument these and a bad crew run stops being a mystery you reproduce by hand.

The crew execution graph

Which agent ran which task, in what order, and what each task returned. In a hierarchical process that graph is decided at runtime by the manager LLM, so it has to be recorded rather than inferred from your code.

Delegation edges and depth

Every delegation as an edge between two agents inside one trace. Depth and repeat count are what turn "the crew was slow" into "agents A and B ping-ponged eleven times".

Iteration and rate-limit budgets

Iterations used against max_iter, and requests against max_rpm. A run that finishes at its ceiling is a truncated run, and it deserves a different label than one that finished because it was done.

Tool calls with real inputs and outputs

The arguments the agent passed, what came back, and whether the tool raised. Tool errors that the LLM recovered from are the ones that go missing, and they are the ones that produce confident wrong answers.

Task output validation

Whether each task's result actually matched its expected_output or output schema, recorded per task instead of judged once at the end of the crew.

Cost attributed per agent

Tokens and dollars broken out by agent and task, not just per crew run. One over-eager researcher agent is usually responsible for most of a crew's bill, and you cannot see that from a single total.

Getting data in

Start with the exporter you already have

CrewAI apps instrumented for OpenTelemetry can ship traces without touching crew code. The ingest gateway reads the OTel GenAI semantic conventions and scopes every span to an org and a customer on arrival.

bash
# Point any OpenTelemetry-instrumented CrewAI app at the ingest gateway.
# No change to your crew, agent, or task definitions.

export OTEL_EXPORTER_OTLP_ENDPOINT="https://ingest.aiagre.com"
export OTEL_EXPORTER_OTLP_HEADERS="authorization=Bearer aig_your_key,\
x-aiagre-customer-id=cus_acme,\
x-aiagre-agent-id=support-crew"
export OTEL_SERVICE_NAME="support-crew"

python run_crew.py

Spans arrive as structured runs, not log lines: prompts, tool calls, tokens, model, latency, and status, keyed by trace ID.

Then record what the crew was actually for

Traces answer what happened inside a run. Outcomes answer whether the run did its job: the input to deflection rate, resolution rate, and cost per resolution.

typescript
// Traces show what the crew did. Outcomes show whether it was worth it.
import { init, withRun, traceChain, trackOutcome } from "@aiagre/node";

init({
  apiKey: process.env.AIAGRE_API_KEY!,
  customerId: "cus_acme",   // your end customer
  agentId: "support-crew",  // the crew this run belongs to
});

await withRun({ sessionId }, async () => {
  const result = await traceChain({ name: "crew.kickoff" }, () =>
    kickoffCrew(ticket)
  );

  await trackOutcome({
    name: "ticket.handled",
    outcome: result.escalated ? "handoff_to_human" : "resolved",
    metadata: { queue: ticket.queue, taskCount: result.tasks.length },
  });
});

The Node SDK covers TypeScript services and JavaScript agent stacks; Python crews use the OTLP path above. See AI agent reliability for how traces and outcomes fit together.

The part tools skip

Your customers do not want a trace viewer

Everything above serves your engineers. If your crew runs inside a product other companies pay for, a second audience shows up with a different question: is this agent saving us money? A span waterfall does not answer that, and it was never meant to.

Answering it means deriving a small set of numbers from the same traces: deflection rate, cost per resolution, resolution rate, and being explicit about how each is defined, because none of them are standardized across the industry. The honest move is to publish your formula alongside the number.

The harder half is architectural. Showing customer A a slice of your crew data, with no path by which A can see customer B's runs, means every span carries an org identity and a customer identity from ingestion, and every read is scoped the same way. A filter bolted onto a single-tenant dashboard is one query bug away from a tenant leak.

AiAgRe scopes every ingested event to an org and a customer, then issues short-lived embed tokens so each customer's dashboard reads only their own crew runs, styled to match your product, not ours.

The full loop

Observe, prove, remediate

Tracing a crew is the first third of the problem. The other two thirds are what keep a crew reliable after the first incident.

Observe

Full crew traces: agents, tasks, delegations, tool calls, tokens, and status, tied to a trace ID you can replay instead of reconstruct.

Prove

Deflection rate, cost per resolution, and resolution rate computed from those traces, scoped per customer and defensible when a customer asks how the number was derived.

Remediate

Cluster recurring crew failures by type, tighten the guardrails and limits that matter, then confirm the numbers actually moved.

Framework-agnostic on purpose:The same ingest path covers LangGraph, LangChain, LlamaIndex, and anything else that speaks OTLP, so swapping orchestrators does not reset your metrics history.

FAQs

CrewAI observability: frequently asked questions

Common questions from teams running CrewAI crews in production.

What is CrewAI observability?

CrewAI observability means capturing the structure of a crew run, not just its final output: which agent took which task, every tool call and tool response, the delegations between agents, tokens and cost per agent, and whether each task's output matched the expected_output it declared. Without that structure, a crew that returns a plausible but wrong answer looks identical to one that worked.

Why do CrewAI crews fail silently?

Because most CrewAI failure modes end in a successful-looking return value. An agent that exhausts max_iter stops reasoning and answers with whatever it has. A tool that raises gets its error fed back to the LLM, which often writes around it instead of surfacing it. A hierarchical manager can bounce a task between agents until the context window fills. In every case kickoff() returns a string, the process exits zero, and nothing in your application logs marks the run as degraded.

How do I monitor a CrewAI crew in production?

Two paths. If your crew already emits OpenTelemetry through an instrumentation library, point the OTLP exporter at an ingest endpoint that understands the GenAI semantic conventions and you get traces with no code change. If you want business outcomes as well, wrap kickoff() and record what the crew was for: resolved, handed off to a human, abandoned. Traces tell you what happened inside the run; outcomes tell you whether the run was worth its cost.

How do you catch a CrewAI delegation loop?

Track delegation as an edge between agents inside a single trace, then alert on depth and repetition rather than on total runtime. A crew that delegates A to B to A to B is burning tokens on a cycle long before it hits a timeout. CrewAI gives you the levers to bound it, allow_delegation, max_iter, max_rpm, but you need the trace to know which crew needs which limit tightened.

Does CrewAI observability work with the hierarchical process?

It matters more there. In a sequential process the task order is fixed and easy to reason about from the outside. In a hierarchical process the manager LLM decides assignment at runtime, so the execution graph is different on every run and the only way to know what actually happened is to record it. Manager decisions should be first-class spans, not hidden inside one opaque crew span.

Can I show CrewAI metrics to my own customers?

Only safely if every span carries a customer identity from the moment it is written. AiAgRe tags each ingested event with both an org identity (you) and a customer identity (your end customer), and issues short-lived embed tokens so a customer's dashboard can only ever read their own crew runs, cost, and resolution numbers, rendered inside your product under your branding.

Do I still need CrewAI's built-in AMP tracing if I use AiAgRe?

Most teams keep both; they cover different layers. AMP is CrewAI's own trace viewer: fast to turn on with crewai login and tracing=True, and scoped to your engineering team inside app.crewai.com. It doesn't carry over if you later add a crew built on LangGraph or a raw API, and it has no outcome or customer-facing layer. AiAgRe reads the same run data over OTLP, works the same way across frameworks, and turns it into the deflection rate, cost per resolution, and per-customer dashboard that AMP was never built to expose. If all you need is why one run failed, AMP alone covers it. If you need to show a paying customer what their agent is doing for them, that's the layer AMP stops at.

Ship a crew you can prove works

Request access and we'll help you instrument your first crew and scope a customer-facing dashboard.