> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Lab Walkthrough Add Monitoring

> Explains adding lightweight instrumentation to a safe agent to log tool calls, durations, token usage, iteration warnings, and an execution summary for debugging and profiling

Your safe agent currently enforces iteration limits and catches tool errors, but it lacks visibility into what happened during a run: which tools were called, how long they took, and how many tokens the model consumed. In this lesson we add simple, low-overhead instrumentation so you can debug, profile, and optimize agent executions.

Overview

* Environment: Python, virtualenv, and the OpenAI SDK are preinstalled.
* Working directory: `/root/code`.
* Goal: start from the safe agent and add logging for tool calls, timing, token usage, and iteration warnings.

Why instrument an agent?

* Faster debugging of tool failures.
* Identify slow tools and optimize or cache them.
* Track token usage to control costs.
* Get advance warning before hitting the iteration cap.

Baseline (what you start from)

* The baseline already enforces a max iteration cap of 10 and includes try/except handling for tool execution.
* Example baseline snippet (trimmed for clarity):

```python theme={null}
from openai import OpenAI
import os
import json
import time

client = OpenAI()
model = "openai/gpt-4.1-mini"
MAX_ITERATIONS = 10

def execute_tool(name, args):
    try:
        if name == "check_calendar":
            result = check_calendar(**args)
        # other tool branches...
    except Exception as e:
        result = f"Error: {str(e)}"
    return result
```

What we'll add

1. `run_log`: a list of dictionaries capturing each tool call: tool name, args, truncated result, and duration\_ms.
2. Timing inside `execute_tool`: measure start/end, compute duration in milliseconds, and append a compact trace to `run_log`.
3. Token tracking: maintain running totals of prompt and completion tokens after each API call.
4. Iteration tracking: maintain `iteration_count` and print a warning when within the last two iterations.
5. Execution summary: when the agent finishes, print iterations used, total tokens, number of tool calls, each call's timing, and total wall-clock time.

Quick links and references

* [OpenAI Python SDK](https://platform.openai.com/docs/api-reference)
* [Python logging best practices](https://docs.python.org/3/howto/logging.html)
* [Virtual environments (venv)](https://docs.python.org/3/library/venv.html)

Instrumentation: add `run_log` and an instrumented `execute_tool`

* Initialize `run_log` and token counters immediately before your agent loop.
* Replace the old `execute_tool` with the instrumented version below. It measures duration and stores a truncated string result so logs remain compact.

```python theme={null}
def execute_tool(name, args, run_log):
    start_time = time.time()
    try:
        if name == "check_calendar":
            result = check_calendar(**args)
        elif name == "some_other_tool":
            result = some_other_tool(**args)
        else:
            result = f"Unknown tool: {name}"
    except Exception as e:
        result = f"Error: {str(e)}"
    duration = time.time() - start_time
    # Ensure we store a string (truncated) so logs stay compact
    run_log.append({
        "tool": name,
        "args": args,
        "result": str(result)[:100],
        "duration_ms": round(duration * 1000)
    })
    return result
```

Track tokens after each API call

* Maintain running counters for prompt and completion tokens and update them after each `client.chat.completions.create` call.
* Different SDK versions may return `usage` either as an attribute or a dict — handle both forms.

```python theme={null}
total_prompt_tokens = 0
total_completion_tokens = 0

# Example call to the Chat Completions endpoint using the OpenAI client
response = client.chat.completions.create(model=model, messages=messages, tools=tools)

# Update token counters if usage is present (handle both attribute and dict forms)
usage = None
if getattr(response, "usage", None) is not None:
    usage = response.usage
elif isinstance(response, dict) and response.get("usage") is not None:
    usage = response.get("usage")

if usage is not None:
    # prompt_tokens may be an attribute or a key depending on the SDK
    if isinstance(usage, dict):
        prompt_tokens = usage.get("prompt_tokens")
        completion_tokens = usage.get("completion_tokens")
    else:
        prompt_tokens = getattr(usage, "prompt_tokens", None)
        completion_tokens = getattr(usage, "completion_tokens", None)

    if prompt_tokens is not None:
        total_prompt_tokens += int(prompt_tokens)
    if completion_tokens is not None:
        total_completion_tokens += int(completion_tokens)
```

Iteration tracking and advance warning

* Keep `iteration_count` and increment it at the start of each loop iteration.
* Print an advance warning when you are within the last two iterations so you can take corrective action or gracefully exit.

```python theme={null}
iteration_count = 0

for iteration in range(MAX_ITERATIONS):
    iteration_count += 1
    # Warn when within the last two iterations (e.g., for MAX=10, warn on 9 and 10)
    if iteration_count >= MAX_ITERATIONS - 1:
        print(f"WARNING: approaching iteration limit ({iteration_count}/{MAX_ITERATIONS})")

    # Call the model (example)
    response = client.chat.completions.create(model=model, messages=messages, tools=tools)

    # (Handle response, update tokens, call tools via execute_tool, etc.)
```

Record wall-clock time and print an execution summary

* Record a wall-clock start time before the loop and compute elapsed time after it completes.
* Print a concise summary with iterations used, tokens, tool call counts and timings, and total elapsed time.

```python theme={null}
wall_start = time.time()

# ... agent loop runs here ...
elapsed = round(time.time() - wall_start, 2)
print("\n=== Execution Summary ===")
print(f"Iterations used: {iteration_count}/{MAX_ITERATIONS}")
print(f"Total tokens: {total_prompt_tokens + total_completion_tokens}")
print(f" - Prompt tokens: {total_prompt_tokens}")
print(f" - Completion tokens: {total_completion_tokens}")
print(f"Tools called: {len(run_log)}")
for entry in run_log:
    print(f" - {entry['tool']}: {entry['duration_ms']} ms, args={entry['args']}, result={entry['result']}")
print(f"Total elapsed time (s): {elapsed}")
```

Instrumentation summary table

| Metric | What it records | Example |
| - | - | - |
| `run_log` | One entry per tool call: `tool`, `args`, truncated `result`, `duration_ms` | `{"tool":"check_calendar","args":{...},"result":"OK","duration_ms":120}` |
| Tokens | Aggregated `prompt` and `completion` tokens from `response.usage` | `prompt_tokens: 45`, `completion_tokens: 12` |
| Iterations | `iteration_count` and warnings when close to `MAX_ITERATIONS` | `WARNING: approaching iteration limit (9/10)` |
| Wall-clock | Total elapsed seconds for the whole run | `Total elapsed time (s): 3.42` |

Full integration notes

* Initialize `run_log`, `total_prompt_tokens`, `total_completion_tokens`, and `iteration_count` immediately before entering your agent loop.
* Replace your existing `execute_tool` with the instrumented version and call it as `execute_tool(name, args, run_log)`.
* After every call to `client.chat.completions.create`, update token counters from `response.usage`.
* Preserve try/except handling inside `execute_tool` so tool failures continue to produce descriptive results that are logged.
* Keep the safety cap `MAX_ITERATIONS` unchanged — instrumentation should not alter agent control flow.

<Callout icon="lightbulb" color="#1CB2FE">
  This instrumentation records what happened (tools called, durations, token usage, and iteration counts) without changing agent decision logic or safety limits. Use these logs to pinpoint slow tools, unexpected exceptions, or high token usage.
</Callout>

Security and privacy note

<Callout icon="warning" color="#FF6B6B">
  Be careful what you log. Logs may contain sensitive information from tool arguments or results. Mask or redact secrets before writing logs to persistent storage or sharing them.
</Callout>

Run it

* Run the monitored agent. It will:
  * Call tools (for example, `check_calendar`) via the instrumented `execute_tool`.
  * Append compact entries to `run_log` for each tool call.
  * Print a one-line answer (your agent's response) followed by the Execution Summary block showing iterations, tokens, per-tool durations, and total elapsed time.

Example output snapshot (illustrative)

```text theme={null}
Agent answer: "I found a free slot on Tuesday at 3pm."
=== Execution Summary ===
Iterations used: 3/10
Total tokens: 120
 - Prompt tokens: 75
 - Completion tokens: 45
Tools called: 2
 - check_calendar: 150 ms, args={'date':'2026-10-07'}, result=OK
 - send_email: 320 ms, args={'to':'user@example.com'}, result=Email sent
Total elapsed time (s): 1.45
```

Next steps

* Persist `run_log` to a JSON file for later analysis or attach it to a debugging UI.
* Add contextual IDs to log entries (request\_id, session\_id) to correlate logs from multiple runs.
* If costs are a concern, use the token counters to create alerts when usage exceeds thresholds.

By adding these lightweight instrumentation primitives you can quickly surface where time and tokens are spent and iterate on performance or correctness with much more confidence.

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/ai-agents-for-beginner-openclaw-case-study/module/b8b38b25-c4eb-425f-a093-cec426365977/lesson/ea214870-02c6-4d48-ad76-68e96f3154e3" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.