> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Monitoring and Debugging

> Guide to monitoring and debugging AI agents by logging every step, tracking key metrics, streaming events, and enforcing PII protection for observability and incident resolution.

When an agent misbehaves in production, diagnosing the cause is different from debugging a traditional application. You can't pause a language model mid-thought with a breakpoint. Instead, you must build observability: log every step, measure the right metrics, and capture precise timing. The goal is simple: know what the agent did, what each tool returned, and how long each step took.

* Log incoming user messages.
* Log model decisions to call tools, including tool name and arguments.
* Log tool results and token counts.
* Log the model's final response.
* At request completion, log totals: elapsed time, total tokens, tool call counts.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/z7NmHsFQN9LCEiD0/images/AI-Agents-for-Beginners-OpenClaw-Case-Study/Production-OpenClaw/Monitoring-and-Debugging/log-everything-user-tool-llm-steps.jpg?fit=max&auto=format&n=z7NmHsFQN9LCEiD0&q=85&s=e917ef345cb62f275e243291794860dd" alt="A neon-style infographic titled &#x22;LOG EVERYTHING&#x22; outlining to log, measure, and time user messages, tool calls, and LLM responses. A &#x22;Capture Every Step&#x22; section shows colored tags for User Message, Tool Call, Tool Result, LLM Response, and Total Summary." width="1920" height="1080" data-path="images/AI-Agents-for-Beginners-OpenClaw-Case-Study/Production-OpenClaw/Monitoring-and-Debugging/log-everything-user-tool-llm-steps.jpg" />
</Frame>

## Example: fully logged flight booking request

* 12:01:00 — User: "Book me a flight to NYC next Friday under \$300"
* The model selects `Search Flights` (with date and price cap).
* 12:01:01 — Tool call: `search_flights`
* 12:01:02 — Tool result: 4 flights found (1,200 tokens)
* 12:01:02 — Tool call: `check_calendar`
* 12:01:03 — Tool result: calendar loaded (800 tokens)
* 12:01:04 — Final model text response
* Total: 3 LLM calls, 2 tool calls, 2,200 tokens, 4.0s

The corresponding log trace:

```text theme={null}
agent.log

[12:01:00] User: "Book me a flight to NYC next Friday under $300"
[12:01:00] Model: claude-sonnet-4-20250514  Tools: [search_flights, check_calendar, book_flight]

[12:01:01] LLM → tool_call → search_flights(dest="NYC", date="2026-03-13", max_price=300)
[12:01:02] Result: 4 flights found (1,200 tokens)

[12:01:02] LLM → tool_call → check_calendar(date="2026-03-13")
[12:01:03] Result: calendar loaded (800 tokens)

[12:01:04] LLM → text → "I found a $280 flight at 2:30pm..."

[12:01:04] Total: 3 LLM calls 2 tool calls 2,200 tokens 4.0s
```

That trace shows exactly what happened. If the answer is wrong, inspect the tool results. If the request is slow, see which step consumed the most time.

<Callout icon="lightbulb" color="#1CB2FE">
  Log everything, but avoid storing sensitive user data in plain text. Implement redaction, anonymization, or tokenization for personally identifiable information (PII) before long-term storage.
</Callout>

## Key metrics to capture for every request

Track these six metrics consistently. They help you spot regressions, prevent runaway costs, and identify reliability problems.

| Metric | Why it matters | How to measure |
| - | -: | - |
| Tokens per request | Primary cost signal. Spikes mean higher API costs or inefficient prompts. | Sum tokens consumed by LLM & tools per request. |
| Tool calls per request | High counts can indicate inefficient chains or looping. | Count unique tool invocations per session/turn. |
| Latency | End-to-end time; split into LLM time and tool execution time to find bottlenecks. | Measure wall-clock time per LLM call and per tool call. |
| Error rate | Tool or API failures degrade UX and loop behavior. | % of failed tool calls / total tool calls. |
| Loop iterations | Excessive loops indicate the model isn't progressing. | Count agent loop iterations per request. |
| Context window usage | Hitting the context limit causes truncation or degraded reasoning. | Track tokens in history vs. model context window (e.g., % of limit). |

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/z7NmHsFQN9LCEiD0/images/AI-Agents-for-Beginners-OpenClaw-Case-Study/Production-OpenClaw/Monitoring-and-Debugging/retro-key-metrics-every-request.jpg?fit=max&auto=format&n=z7NmHsFQN9LCEiD0&q=85&s=d5225829ec2ddd43efaaca31ab3cc4b8" alt="A retro-styled infographic titled &#x22;KEY METRICS.&#x22; It lists six metrics to track on every request — Tokens/Request, Tool Calls/Request, Latency, Error Rate, Loop Iterations, and Context Usage — each with a short description." width="1920" height="1080" data-path="images/AI-Agents-for-Beginners-OpenClaw-Case-Study/Production-OpenClaw/Monitoring-and-Debugging/retro-key-metrics-every-request.jpg" />
</Frame>

## Common failure modes and debugging steps

* Agent returns wrong answers
  * Inspect logs for tool results and timestamps. Often the reasoning is correct but the data is stale or erroneous. Fix the tool or data source rather than the model.

* Agent calls the wrong tool
  * Make tool descriptions explicit. If two tools look similar to the model (e.g., "Search Web" vs "Search Flights"), clarify usage conditions in the tool description and system prompt.

* Agent loops without progress
  * Usually the model cannot interpret a tool's output. Improve the result format (structured JSON, clear keys), add instructions on how to handle that output, or include success/failure flags in tool responses.

* Responses are too slow
  * Break down latency into LLM vs. external API time. Large context windows increase LLM latency; slow third-party APIs increase tool latency. Optimize by compressing context and caching tool results.

* Context window overflow
  * Summarize or compress old turns, cap included history, or implement retrieval strategies to avoid reaching \~80% of the model’s context limit.

## Real-time event streaming and observability (OpenClaw)

OpenClaw streams agent events in real time via a Pub/Sub event system: text generation, tool calls, tool results, errors, retries, and context compression are emitted as events. Streaming these events enables:

* Live debugging and replay of agent behavior.
* Building progress indicators and structured UIs that reflect each step.
* Postmortem analysis and searchable transcripts.

Subscribe to agent events (example):

```javascript theme={null}
// Subscribe to agent events via pi-embedded-subscribe
const events = await client.piEmbeddedSubscribe({ sessionId });

for await (const event of events) {
  switch (event.type) {
    case "text_generation":
      stream(event.text);
      break;
    case "tool_call":
      log(event.name, event.args);
      break;
    case "tool_result":
      log(event.result, event.tokens);
      break;
    case "error":
      alert(event.message);
      break;
    case "context_compression":
      log("compressing...");
      break;
  }
}
```

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/z7NmHsFQN9LCEiD0/images/AI-Agents-for-Beginners-OpenClaw-Case-Study/Production-OpenClaw/Monitoring-and-Debugging/openclaw-observability-debugging-mindset-transcript-buttons.jpg?fit=max&auto=format&n=z7NmHsFQN9LCEiD0&q=85&s=74a849db2dcc0633499cb588a3a29a79" alt="A retro-style slide titled &#x22;OpenClaw Observability&#x22; and &#x22;The Debugging Mindset&#x22; showing contrasting debugging approaches (&#x22;NOT stepping through code&#x22; vs &#x22;READING the transcript&#x22;). It highlights that conversation history is the primary debugging tool and shows buttons labeled &#x22;LOG IT,&#x22; &#x22;SEARCH IT,&#x22; and &#x22;STORE IT.&#x22;" width="1920" height="1080" data-path="images/AI-Agents-for-Beginners-OpenClaw-Case-Study/Production-OpenClaw/Monitoring-and-Debugging/openclaw-observability-debugging-mindset-transcript-buttons.jpg" />
</Frame>

## The debugging mindset for agents

You are not stepping through functions — you are reading a transcript. Every tool call, every result, and the model’s reasoning is your primary debugging interface. Make logs searchable, structured, and retained long enough for analysis and postmortems.

<Callout icon="warning" color="#FF6B6B">
  Beware of cost and privacy: verbose logging increases token usage and storage. Balance fidelity with cost and compliance by sampling, redacting PII, and aggregating metrics.
</Callout>

## Checklist: build observability from day one

* Design logs as structured events (timestamp, type, payload, tokens, duration).
* Emit both high-level summaries and raw traces for deep debugging.
* Track the six key metrics per request and surface alerts on outliers.
* Stream events for live UIs and easier replay.
* Ensure PII is handled (masking/redaction) and log retention complies with policy.
* Use tools like OpenTelemetry and Prometheus for metrics and tracing integrations.

## Links and references

* OpenTelemetry: [https://opentelemetry.io/](https://opentelemetry.io/)
* Prometheus: [https://prometheus.io/](https://prometheus.io/)
* Guidance on logging and PII: [https://www.privacyguidance.example/](https://www.privacyguidance.example/) (replace with your org policy)

Build monitoring and observability into your agent system from the start — readable transcripts and structured metrics are the fastest path from incident to resolution.

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/ai-agents-for-beginner-openclaw-case-study/module/b8b38b25-c4eb-425f-a093-cec426365977/lesson/882e05ae-3f04-4c2a-b1b0-08668bd689cc" />

  <Card title="Practice Lab" icon="flask-conical" cta="Learn more" href="https://learn.kodekloud.com/user/courses/ai-agents-for-beginner-openclaw-case-study/module/b8b38b25-c4eb-425f-a093-cec426365977/lesson/ac796dee-402f-402a-b421-4ac5c2dac327" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.