Skip to main content
Zyppy, Savvy, and Meshy mostly handle text: questions, research, and memories. Tasks that require actual code execution — running scripts, querying databases, or automating system workflows — belong to Codey. Codey is the code and automation specialist: he runs Python scripts, executes system commands, and queries databases. Because code execution is inherently unpredictable, Codey needs the strongest error handling of all agents. In this article we explain how agents fail, and show practical approaches to make them resilient in production.
A retro-styled interface showing "Text Agents" with small colorful robot icons (Zippy, Savvy, Meshy) and a central green robot named CODY labeled "Code & Automation Specialist." Below are three neon buttons labeled "PYTHON SCRIPTS," "DATABASES," and "SYSTEM COMMANDS."
Codey operates in the real world: network errors, runtime exceptions, timeouts, and side effects happen frequently. Design agents to detect and react to these failures instead of hiding them.
Agents fail, tools crash, APIs time out, and models can hallucinate. A production-ready agent must anticipate these failure modes and recover without breaking the user experience.
A black background with large, colorful pixelated text reading: "AGENTS FAIL. TOOLS CRASH. APIS TIME OUT. MODELS HALLUCINATE." A small green robot icon appears in the top-right and a yellow button near the bottom says "HANDLE FAILURES • WITHOUT BREAKING."

Overview — common failure modes

There are six common failure modes to design for: Below we walk through practical strategies for each failure mode.

Tool failures

When a tool call raises an exception, do not swallow it. Instead, make the failure visible to the agent so the LLM can retry, choose a different tool, or ask for clarification. Wrap external tool calls in try/except, capture diagnostic information (exception type and traceback), and present that back to the agent as a tool result. Example pattern:
Give the LLM the error message as part of the conversation so it can select a next step (retry, switch tools, or ask for clarification). Also implement exponential backoff for transient network errors.

Bad tool selection

Bad tool selection is often a design problem rather than a runtime exception. If the agent consistently picks the wrong tool:
  • Improve tool descriptions and add concrete examples.
  • Include explicit negative examples: show when not to use the tool.
  • Constrain tool usage with lightweight validators or type checks before execution.
  • Use server-side validation to reject tool calls with invalid argument shapes.
Better metadata and clearer examples significantly reduce hallucinated or incorrect tool choices.

Hallucinated tool calls

Agents sometimes invent a tool name or pass invalid arguments. Always validate tool calls on the platform side. If a tool name is unknown or arguments are malformed, return a structured error back to the agent:
  • Return machine-readable error codes plus human-readable diagnosis.
  • Suggest the correct tool or required argument schema where possible.
  • Log the hallucinated call for monitoring and retraining.

Infinite loops

Unbounded iteration is costly and can escalate bills or resource usage. Always set hard iteration limits, time budgets, and inspect the model’s finish reason when available. Example iteration guard:
Other useful guards: a total time budget, token budget, and per-tool retry limits.
Never allow unbounded execution of user-defined scripts or arbitrary shell commands. Always sandbox and constrain resource usage to avoid runaway processes or security risks.

Context overflow

When a conversation grows past the model’s context window, compress older content. Keep recent exchanges in full fidelity and summarize or embed earlier ones to preserve the thread while freeing tokens for current reasoning. Common approaches:
  • Periodically summarize earlier messages and replace them with compressed summaries.
  • Use embeddings to store long-term memory and retrieve only relevant snippets.
  • Detect when the context is nearing capacity and trigger compression automatically.
The case study’s implementation detects context pressure and compresses earlier messages to make room for new content.
A stylized UI graphic titled "CONTEXT OVERFLOW" showing a context window split into "old · summarized" and "recent · full fidelity" with a highlighted "COMPRESS OLD CONTENT" button. Below it is a panel labeled "OPENCLAW IMPLEMENTATION" with the subtitle "Detects context filling up."

Model errors and provider fallback

Model APIs can fail with rate limits, outages, or malformed responses. Implement provider fallback: try a primary provider, and if it fails or returns an invalid response, retry with a secondary provider. Normalize request and response formats so switching providers does not require changing higher-level logic. Example fallback sketch:
Keep responses normalized (tokens, finish reasons, metadata) so the agent logic can remain provider-agnostic.

Five core principles for robust error handling

  1. Expect failure. Design for tool and model failures from day one.
  2. Surface errors to the agent. Don’t silently swallow failures — let the LLM adapt.
  3. Set hard limits. Constrain iterations, tokens, and time to prevent runaway costs.
  4. Log everything. Detailed logs (including tool inputs/outputs and tracebacks) make diagnosis and reproduction possible.
  5. Fail gracefully to users. Return clear, actionable messages instead of cryptic traces.

Implementation summary

This implementation bundles the patterns above:
  • Model fallback across providers.
  • Tool error recovery that passes errors back to the agent as structured results.
  • Context compression for long conversations.
  • Sandboxed execution for untrusted code and system commands.
  • Rate limits and iteration guards to prevent runaway loops.
These layers work together to make production agents resilient, observable, and safer to run at scale. Error handling is essential for production agents. Know the failure modes, build recovery strategies, surface errors to the agent so it can adapt, set hard limits to control costs, and log everything for debugging. When things go wrong, fail gracefully to the user.
A neon-styled slide titled "OPENCLAW PATTERNS" listing resilience patterns like Model Fallback, Error Recovery, CTX Compression, Sandbox Exec, Rate Limiting, and Resilient in Production. Below is a "THE TAKEAWAY" section with tips such as "Know your failure modes," "Fail gracefully," "Expect • Design • Recover," and "Set hard limits."

Further reading and references

Use these references to learn provider-specific retry semantics, rate-limiting guidance, and SDK best practices.

Watch Video