Skip to main content
Time for another hands-on lab. You already have a tool-calling agent that works on the happy path, but it has two critical failure modes that will hurt you in production:
  • tools that crash (uncaught exceptions), and
  • an unbounded loop that can retry forever (burning API credits).
This walkthrough shows how to harden a simple Python agent (single-file: safe_agent.py) so it behaves safely in production: tools return errors as results, and the main loop is bounded with a graceful final attempt. Environment
  • Working directory: /root/code
  • File: safe_agent.py
  • Prereqs: Python, virtual environment, OpenAI SDK (pre-installed)
A retro-style graphic with neon pixelated text reading "HANDS-ON LAB" and "BUILD A SAFE AGENT." Below it is the subtitle: "Your tool-calling agent, hardened for production," all on a dark grid background.
Overview of the baseline (unsafeguarded)
  • File to edit: safe_agent.py
  • Start with this minimal agent: OpenAI client, a check_calendar tool, execute_tool dispatcher, and a while True loop. This is intentionally vulnerable so you can see the failure modes.
Baseline (unsafeguarded) example
Failure modes, fixes, and benefits Failure mode #1 — tool crashes
  • Problem: if a tool raises an exception, the entire script crashes, the user sees a traceback, and the conversation ends.
  • Fix: catch exceptions in execute_tool and return a descriptive error string. The model will receive "Error: ..." as the tool output and can adjust behavior (try another tool, explain the problem, notify the user).
Add a flaky tool to test this behavior (it always raises), and update execute_tool to return an error string on exceptions. Patch for safe tool execution
Wire flaky_tool into the tool list you provide the model (for example, in the system or tool description messages) and ask the agent to use it. The script should not crash — instead the model will receive a result beginning with "Error" and can recover.
Return errors to the agent (as strings) rather than raising them from execute_tool. This lets the model handle failures gracefully and maintain conversation continuity.
Failure mode #2 — infinite loops
  • Problem: while True exits only when the model signals finish (finish_reason == “stop”). If tools are flaky or the model never returns a final finish, the agent can retry forever and burn credits.
  • Fix: replace while True with a bounded loop (e.g., for iteration in range(MAX_ITERATIONS)). Use a for-else clause: the else block runs only when the loop exhausts iterations without a break. In this case, append a prompt asking for a best-effort answer and make one final completion call.
Main loop with iteration logging and graceful exhaustion
Do not use unbounded retry loops in production agents. Always cap retries (e.g., MAX_ITERATIONS) and provide a clear final attempt so users get a best-effort response instead of the process running indefinitely.
Putting it all together — a safer safe_agent.py
  • The minimal improvements to apply:
    1. Add flaky_tool for testing.
    2. Make execute_tool return errors (string) instead of raising.
    3. Replace while True with a bounded loop and use for-else to perform a final completion if the agent never finishes.
Example complete script (combine the patches above):
Quick test suggestions
  • Ask the agent to use flaky_tool. You should see:
    • iteration logs increment,
    • execute_tool returning "Error: Service unavailable. Try a different approach.",
    • the model receives the error string and should either try a different tool or produce a best-effort answer when iterations are exhausted.
  • Verify that the process never crashes with a stack trace from flaky_tool.
Why these changes matter (summary)
  • Returning errors as tool results keeps the conversation alive and allows the agent to recover.
  • Bounded loops prevent runaway API usage and give you a chance to provide a final best-effort response.
  • Together they make your tool-calling agent resilient and production-ready.
References and further reading Good luck! Run the script again after applying the patches — you should see iteration counters, handled tool errors, and graceful termination instead of crashes or infinite retries.

Watch Video