All guides

LLM observability for AI agents

Compare agent observability tools by model calls, tool execution, retries and task outcomes.

$20 per user / month. Payment required to activate your workspace.

Lunary trace showing nested model and tool calls
Agent tracing in Lunary

Follow the whole task, from the first model call to the result your user receives.

Complete traces

Connect model, retrieval and tool operations to a parent task.

Useful metrics

Measure success, latency and cost per task.

Lunary

Inspect application runs alongside customer conversations.

At a glance

LLM observability for AI agents: at a glance
PlatformReason to include itAgent-specific test
LunaryApplication runs plus conversation contextFollow a user issue through nested tool calls
LangfuseTracing and evaluation workflowsTurn a failed observation into a repeatable test
LangSmithTracing, experiments, and prompt workInspect a nested failure and compare a candidate
HeliconeGateway plus request observabilityVerify provider behavior and surrounding tool context

Capture what actually executed

A model’s tool-call request is different from the tool’s execution. Record both, including retries, errors and the final user-visible result.

  • Stable task and parent IDs.
  • Model, prompt and release context.
  • Tool results, duration and errors.

Match the tool to your architecture

Compare Lunary for conversation context, Langfuse or LangSmith for tracing and evaluations, and Helicone when you also need a provider gateway.

Start with one agent

Use Lunary’s SDK or OpenTelemetry integration. Verify that a tool failure appears under the right parent run and that the trace explains the final response.

Questions & answers

Is tracing model calls enough?

Usually not. Include tools, retrieval, retries and the final task outcome.

What should I measure?

Task success alongside latency and cost. An HTTP 200 does not establish a successful user outcome.

By Lunary · Documentation reviewed Sep 16, 2026

Sources & methodology (10)

Based on vendor documentation, not an independent benchmark or hands-on rating. Features and plans can change.

Build with Lunary.

Trace, evaluate, and improve your AI applications.

$20 per user / month. Payment required to activate your workspace.