Skip to main content
Your safe agent currently enforces iteration limits and catches tool errors, but it lacks visibility into what happened during a run: which tools were called, how long they took, and how many tokens the model consumed. In this lesson we add simple, low-overhead instrumentation so you can debug, profile, and optimize agent executions. Overview
  • Environment: Python, virtualenv, and the OpenAI SDK are preinstalled.
  • Working directory: /root/code.
  • Goal: start from the safe agent and add logging for tool calls, timing, token usage, and iteration warnings.
Why instrument an agent?
  • Faster debugging of tool failures.
  • Identify slow tools and optimize or cache them.
  • Track token usage to control costs.
  • Get advance warning before hitting the iteration cap.
Baseline (what you start from)
  • The baseline already enforces a max iteration cap of 10 and includes try/except handling for tool execution.
  • Example baseline snippet (trimmed for clarity):
What we’ll add
  1. run_log: a list of dictionaries capturing each tool call: tool name, args, truncated result, and duration_ms.
  2. Timing inside execute_tool: measure start/end, compute duration in milliseconds, and append a compact trace to run_log.
  3. Token tracking: maintain running totals of prompt and completion tokens after each API call.
  4. Iteration tracking: maintain iteration_count and print a warning when within the last two iterations.
  5. Execution summary: when the agent finishes, print iterations used, total tokens, number of tool calls, each call’s timing, and total wall-clock time.
Quick links and references Instrumentation: add run_log and an instrumented execute_tool
  • Initialize run_log and token counters immediately before your agent loop.
  • Replace the old execute_tool with the instrumented version below. It measures duration and stores a truncated string result so logs remain compact.
Track tokens after each API call
  • Maintain running counters for prompt and completion tokens and update them after each client.chat.completions.create call.
  • Different SDK versions may return usage either as an attribute or a dict — handle both forms.
Iteration tracking and advance warning
  • Keep iteration_count and increment it at the start of each loop iteration.
  • Print an advance warning when you are within the last two iterations so you can take corrective action or gracefully exit.
Record wall-clock time and print an execution summary
  • Record a wall-clock start time before the loop and compute elapsed time after it completes.
  • Print a concise summary with iterations used, tokens, tool call counts and timings, and total elapsed time.
Instrumentation summary table Full integration notes
  • Initialize run_log, total_prompt_tokens, total_completion_tokens, and iteration_count immediately before entering your agent loop.
  • Replace your existing execute_tool with the instrumented version and call it as execute_tool(name, args, run_log).
  • After every call to client.chat.completions.create, update token counters from response.usage.
  • Preserve try/except handling inside execute_tool so tool failures continue to produce descriptive results that are logged.
  • Keep the safety cap MAX_ITERATIONS unchanged — instrumentation should not alter agent control flow.
This instrumentation records what happened (tools called, durations, token usage, and iteration counts) without changing agent decision logic or safety limits. Use these logs to pinpoint slow tools, unexpected exceptions, or high token usage.
Security and privacy note
Be careful what you log. Logs may contain sensitive information from tool arguments or results. Mask or redact secrets before writing logs to persistent storage or sharing them.
Run it
  • Run the monitored agent. It will:
    • Call tools (for example, check_calendar) via the instrumented execute_tool.
    • Append compact entries to run_log for each tool call.
    • Print a one-line answer (your agent’s response) followed by the Execution Summary block showing iterations, tokens, per-tool durations, and total elapsed time.
Example output snapshot (illustrative)
Next steps
  • Persist run_log to a JSON file for later analysis or attach it to a debugging UI.
  • Add contextual IDs to log entries (request_id, session_id) to correlate logs from multiple runs.
  • If costs are a concern, use the token counters to create alerts when usage exceeds thresholds.
By adding these lightweight instrumentation primitives you can quickly surface where time and tokens are spent and iterate on performance or correctness with much more confidence.

Watch Video