LLM observability for AI agents
Compare agent observability tools by model calls, tool execution, retries and task outcomes.
$20 per user / month. Payment required to activate your workspace.

Follow the whole task, from the first model call to the result your user receives.
Complete traces
Connect model, retrieval and tool operations to a parent task.
Useful metrics
Measure success, latency and cost per task.
Lunary
Inspect application runs alongside customer conversations.
At a glance
| Platform | Reason to include it | Agent-specific test |
|---|---|---|
| Lunary | Application runs plus conversation context | Follow a user issue through nested tool calls |
| Langfuse | Tracing and evaluation workflows | Turn a failed observation into a repeatable test |
| LangSmith | Tracing, experiments, and prompt work | Inspect a nested failure and compare a candidate |
| Helicone | Gateway plus request observability | Verify provider behavior and surrounding tool context |
Capture what actually executed
A model’s tool-call request is different from the tool’s execution. Record both, including retries, errors and the final user-visible result.
- Stable task and parent IDs.
- Model, prompt and release context.
- Tool results, duration and errors.
Match the tool to your architecture
Compare Lunary for conversation context, Langfuse or LangSmith for tracing and evaluations, and Helicone when you also need a provider gateway.
Start with one agent
Use Lunary’s SDK or OpenTelemetry integration. Verify that a tool failure appears under the right parent run and that the trace explains the final response.
Questions & answers
Is tracing model calls enough?
Usually not. Include tools, retrieval, retries and the final task outcome.
What should I measure?
Task success alongside latency and cost. An HTTP 200 does not establish a successful user outcome.
By Lunary · Documentation reviewed Sep 16, 2026
Sources & methodology (10)
Based on vendor documentation, not an independent benchmark or hands-on rating. Features and plans can change.
Build with Lunary.
Trace, evaluate, and improve your AI applications.
$20 per user / month. Payment required to activate your workspace.