Skip to main content
When an agent misbehaves in production, diagnosing the cause is different from debugging a traditional application. You can’t pause a language model mid-thought with a breakpoint. Instead, you must build observability: log every step, measure the right metrics, and capture precise timing. The goal is simple: know what the agent did, what each tool returned, and how long each step took.
  • Log incoming user messages.
  • Log model decisions to call tools, including tool name and arguments.
  • Log tool results and token counts.
  • Log the model’s final response.
  • At request completion, log totals: elapsed time, total tokens, tool call counts.
A neon-style infographic titled "LOG EVERYTHING" outlining to log, measure, and time user messages, tool calls, and LLM responses. A "Capture Every Step" section shows colored tags for User Message, Tool Call, Tool Result, LLM Response, and Total Summary.

Example: fully logged flight booking request

  • 12:01:00 — User: “Book me a flight to NYC next Friday under $300”
  • The model selects Search Flights (with date and price cap).
  • 12:01:01 — Tool call: search_flights
  • 12:01:02 — Tool result: 4 flights found (1,200 tokens)
  • 12:01:02 — Tool call: check_calendar
  • 12:01:03 — Tool result: calendar loaded (800 tokens)
  • 12:01:04 — Final model text response
  • Total: 3 LLM calls, 2 tool calls, 2,200 tokens, 4.0s
The corresponding log trace:
That trace shows exactly what happened. If the answer is wrong, inspect the tool results. If the request is slow, see which step consumed the most time.
Log everything, but avoid storing sensitive user data in plain text. Implement redaction, anonymization, or tokenization for personally identifiable information (PII) before long-term storage.

Key metrics to capture for every request

Track these six metrics consistently. They help you spot regressions, prevent runaway costs, and identify reliability problems.
A retro-styled infographic titled "KEY METRICS." It lists six metrics to track on every request — Tokens/Request, Tool Calls/Request, Latency, Error Rate, Loop Iterations, and Context Usage — each with a short description.

Common failure modes and debugging steps

  • Agent returns wrong answers
    • Inspect logs for tool results and timestamps. Often the reasoning is correct but the data is stale or erroneous. Fix the tool or data source rather than the model.
  • Agent calls the wrong tool
    • Make tool descriptions explicit. If two tools look similar to the model (e.g., “Search Web” vs “Search Flights”), clarify usage conditions in the tool description and system prompt.
  • Agent loops without progress
    • Usually the model cannot interpret a tool’s output. Improve the result format (structured JSON, clear keys), add instructions on how to handle that output, or include success/failure flags in tool responses.
  • Responses are too slow
    • Break down latency into LLM vs. external API time. Large context windows increase LLM latency; slow third-party APIs increase tool latency. Optimize by compressing context and caching tool results.
  • Context window overflow
    • Summarize or compress old turns, cap included history, or implement retrieval strategies to avoid reaching ~80% of the model’s context limit.

Real-time event streaming and observability (OpenClaw)

OpenClaw streams agent events in real time via a Pub/Sub event system: text generation, tool calls, tool results, errors, retries, and context compression are emitted as events. Streaming these events enables:
  • Live debugging and replay of agent behavior.
  • Building progress indicators and structured UIs that reflect each step.
  • Postmortem analysis and searchable transcripts.
Subscribe to agent events (example):
A retro-style slide titled "OpenClaw Observability" and "The Debugging Mindset" showing contrasting debugging approaches ("NOT stepping through code" vs "READING the transcript"). It highlights that conversation history is the primary debugging tool and shows buttons labeled "LOG IT," "SEARCH IT," and "STORE IT."

The debugging mindset for agents

You are not stepping through functions — you are reading a transcript. Every tool call, every result, and the model’s reasoning is your primary debugging interface. Make logs searchable, structured, and retained long enough for analysis and postmortems.
Beware of cost and privacy: verbose logging increases token usage and storage. Balance fidelity with cost and compliance by sampling, redacting PII, and aggregating metrics.

Checklist: build observability from day one

  • Design logs as structured events (timestamp, type, payload, tokens, duration).
  • Emit both high-level summaries and raw traces for deep debugging.
  • Track the six key metrics per request and surface alerts on outliers.
  • Stream events for live UIs and easier replay.
  • Ensure PII is handled (masking/redaction) and log retention complies with policy.
  • Use tools like OpenTelemetry and Prometheus for metrics and tracing integrations.
Build monitoring and observability into your agent system from the start — readable transcripts and structured metrics are the fastest path from incident to resolution.

Watch Video

Practice Lab