> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Error Handling

> Strategies for detecting, handling, and recovering from failures in AI agents and automation

Zyppy, Savvy, and Meshy mostly handle text: questions, research, and memories. Tasks that require actual code execution — running scripts, querying databases, or automating system workflows — belong to Codey.

Codey is the code and automation specialist: he runs Python scripts, executes system commands, and queries databases. Because code execution is inherently unpredictable, Codey needs the strongest error handling of all agents. In this article we explain how agents fail, and show practical approaches to make them resilient in production.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/WUKBeXdogksN49S5/images/AI-Agents-for-Beginners-OpenClaw-Case-Study/Building-AI-Agents/Error-Handling/retro-text-agents-cody-code-automation.jpg?fit=max&auto=format&n=WUKBeXdogksN49S5&q=85&s=96b6942a7da0647bc8aaaec44687b52a" alt="A retro-styled interface showing &#x22;Text Agents&#x22; with small colorful robot icons (Zippy, Savvy, Meshy) and a central green robot named CODY labeled &#x22;Code & Automation Specialist.&#x22; Below are three neon buttons labeled &#x22;PYTHON SCRIPTS,&#x22; &#x22;DATABASES,&#x22; and &#x22;SYSTEM COMMANDS.&#x22;" width="1920" height="1080" data-path="images/AI-Agents-for-Beginners-OpenClaw-Case-Study/Building-AI-Agents/Error-Handling/retro-text-agents-cody-code-automation.jpg" />
</Frame>

<Callout icon="lightbulb" color="#1CB2FE">
  Codey operates in the real world: network errors, runtime exceptions, timeouts, and side effects happen frequently. Design agents to detect and react to these failures instead of hiding them.
</Callout>

Agents fail, tools crash, APIs time out, and models can hallucinate. A production-ready agent must anticipate these failure modes and recover without breaking the user experience.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/WUKBeXdogksN49S5/images/AI-Agents-for-Beginners-OpenClaw-Case-Study/Building-AI-Agents/Error-Handling/agents-tools-apis-models-failures-robot.jpg?fit=max&auto=format&n=WUKBeXdogksN49S5&q=85&s=05b9b3260f85d5eefa48f2b73834f1a5" alt="A black background with large, colorful pixelated text reading: &#x22;AGENTS FAIL. TOOLS CRASH. APIS TIME OUT. MODELS HALLUCINATE.&#x22; A small green robot icon appears in the top-right and a yellow button near the bottom says &#x22;HANDLE FAILURES • WITHOUT BREAKING.&#x22;" width="1920" height="1080" data-path="images/AI-Agents-for-Beginners-OpenClaw-Case-Study/Building-AI-Agents/Error-Handling/agents-tools-apis-models-failures-robot.jpg" />
</Frame>

## Overview — common failure modes

There are six common failure modes to design for:

| Failure mode | What happens | Mitigation |
| - | - | - |
| Tool failures | Tool call errors (API down, bad input, timeout) | Surface the error to the agent, log diagnostics, retry/backoff, use validation |
| Bad tool selection | Agent picks an unsuitable tool for the task | Improve tool metadata, examples, and add negative examples and validators |
| Hallucinated tool calls | Agent invents a tool name or passes invalid args | Validate tool calls server-side, reject unknown tools, return clear error results |
| Infinite loops | Agent repeats steps without progress | Enforce hard iteration/time limits and inspect model finish reasons |
| Context overflow | Conversation exceeds model context window | Compress or summarize older messages; keep recent messages in full fidelity |
| Model/provider errors | Rate limits, outages, malformed responses | Model fallback, request normalization, and consistent validation |

Below we walk through practical strategies for each failure mode.

## Tool failures

When a tool call raises an exception, do not swallow it. Instead, make the failure visible to the agent so the LLM can retry, choose a different tool, or ask for clarification. Wrap external tool calls in try/except, capture diagnostic information (exception type and traceback), and present that back to the agent as a tool result.

Example pattern:

```python theme={null}
import traceback

def call_tool_and_report(tool, *args, **kwargs):
    try:
        result = tool.run(*args, **kwargs)
        return {"ok": True, "result": result}
    except Exception as e:
        tb = traceback.format_exc()
        error_message = f"TOOL ERROR: {type(e).__name__}: {e}\n{tb}"
        # send_error makes the error visible to the agent (e.g., as a tool result)
        send_error(error_message)
        return {"ok": False, "error": error_message}
```

Give the LLM the error message as part of the conversation so it can select a next step (retry, switch tools, or ask for clarification). Also implement exponential backoff for transient network errors.

## Bad tool selection

Bad tool selection is often a design problem rather than a runtime exception. If the agent consistently picks the wrong tool:

* Improve tool descriptions and add concrete examples.
* Include explicit negative examples: show when *not* to use the tool.
* Constrain tool usage with lightweight validators or type checks before execution.
* Use server-side validation to reject tool calls with invalid argument shapes.

Better metadata and clearer examples significantly reduce hallucinated or incorrect tool choices.

## Hallucinated tool calls

Agents sometimes invent a tool name or pass invalid arguments. Always validate tool calls on the platform side. If a tool name is unknown or arguments are malformed, return a structured error back to the agent:

* Return machine-readable error codes plus human-readable diagnosis.
* Suggest the correct tool or required argument schema where possible.
* Log the hallucinated call for monitoring and retraining.

## Infinite loops

Unbounded iteration is costly and can escalate bills or resource usage. Always set hard iteration limits, time budgets, and inspect the model's finish reason when available.

Example iteration guard:

```python theme={null}
MAX_ITER = 10

for i in range(MAX_ITER):
    response = call_model(current_context)

    # Many model APIs include a finish reason; check it when available
    finish_reason = response.get("finish_reason")  # e.g., "stop", "length", "max_tokens"
    if finish_reason == "stop":
        # model indicates it's done
        break

# If we hit the iteration limit, force a final response
if i + 1 >= MAX_ITER:
    force_response("You have reached the maximum number of reasoning steps. Provide the best final answer you can with the information available.")
```

Other useful guards: a total time budget, token budget, and per-tool retry limits.

<Callout icon="warning" color="#FF6B6B">
  Never allow unbounded execution of user-defined scripts or arbitrary shell commands. Always sandbox and constrain resource usage to avoid runaway processes or security risks.
</Callout>

## Context overflow

When a conversation grows past the model's context window, compress older content. Keep recent exchanges in full fidelity and summarize or embed earlier ones to preserve the thread while freeing tokens for current reasoning.

Common approaches:

* Periodically summarize earlier messages and replace them with compressed summaries.
* Use embeddings to store long-term memory and retrieve only relevant snippets.
* Detect when the context is nearing capacity and trigger compression automatically.

The case study's implementation detects context pressure and compresses earlier messages to make room for new content.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/WUKBeXdogksN49S5/images/AI-Agents-for-Beginners-OpenClaw-Case-Study/Building-AI-Agents/Error-Handling/context-overflow-compress-old-content-openclaw.jpg?fit=max&auto=format&n=WUKBeXdogksN49S5&q=85&s=a00084a2594417ec96244eb7e89015f1" alt="A stylized UI graphic titled &#x22;CONTEXT OVERFLOW&#x22; showing a context window split into &#x22;old · summarized&#x22; and &#x22;recent · full fidelity&#x22; with a highlighted &#x22;COMPRESS OLD CONTENT&#x22; button. Below it is a panel labeled &#x22;OPENCLAW IMPLEMENTATION&#x22; with the subtitle &#x22;Detects context filling up.&#x22;" width="1920" height="1080" data-path="images/AI-Agents-for-Beginners-OpenClaw-Case-Study/Building-AI-Agents/Error-Handling/context-overflow-compress-old-content-openclaw.jpg" />
</Frame>

## Model errors and provider fallback

Model APIs can fail with rate limits, outages, or malformed responses. Implement provider fallback: try a primary provider, and if it fails or returns an invalid response, retry with a secondary provider. Normalize request and response formats so switching providers does not require changing higher-level logic.

Example fallback sketch:

```python theme={null}
providers = [claude_client, gpt_client, gemini_client]

for provider in providers:
    try:
        response = provider.call_model(prompt, **params)
        if valid_response(response):
            return response
    except ProviderError as e:
        log.warning(f"{provider.name} failed: {e}")
# If all providers fail, return a user-facing error message
return {"error": "All model providers are currently unavailable. Please try again later."}
```

Keep responses normalized (tokens, finish reasons, metadata) so the agent logic can remain provider-agnostic.

## Five core principles for robust error handling

1. Expect failure. Design for tool and model failures from day one.
2. Surface errors to the agent. Don’t silently swallow failures — let the LLM adapt.
3. Set hard limits. Constrain iterations, tokens, and time to prevent runaway costs.
4. Log everything. Detailed logs (including tool inputs/outputs and tracebacks) make diagnosis and reproduction possible.
5. Fail gracefully to users. Return clear, actionable messages instead of cryptic traces.

## Implementation summary

This implementation bundles the patterns above:

* Model fallback across providers.
* Tool error recovery that passes errors back to the agent as structured results.
* Context compression for long conversations.
* Sandboxed execution for untrusted code and system commands.
* Rate limits and iteration guards to prevent runaway loops.

These layers work together to make production agents resilient, observable, and safer to run at scale.

Error handling is essential for production agents. Know the failure modes, build recovery strategies, surface errors to the agent so it can adapt, set hard limits to control costs, and log everything for debugging. When things go wrong, fail gracefully to the user.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/WUKBeXdogksN49S5/images/AI-Agents-for-Beginners-OpenClaw-Case-Study/Building-AI-Agents/Error-Handling/openclaw-resilience-patterns-takeaway.jpg?fit=max&auto=format&n=WUKBeXdogksN49S5&q=85&s=e8eda4004840248c729e5ea85e69eab0" alt="A neon-styled slide titled &#x22;OPENCLAW PATTERNS&#x22; listing resilience patterns like Model Fallback, Error Recovery, CTX Compression, Sandbox Exec, Rate Limiting, and Resilient in Production. Below is a &#x22;THE TAKEAWAY&#x22; section with tips such as &#x22;Know your failure modes,&#x22; &#x22;Fail gracefully,&#x22; &#x22;Expect • Design • Recover,&#x22; and &#x22;Set hard limits.&#x22;" width="1920" height="1080" data-path="images/AI-Agents-for-Beginners-OpenClaw-Case-Study/Building-AI-Agents/Error-Handling/openclaw-resilience-patterns-takeaway.jpg" />
</Frame>

## Further reading and references

* OpenAI API — [https://platform.openai.com/docs](https://platform.openai.com/docs)
* Anthropic/Claude docs — [https://www.anthropic.com/](https://www.anthropic.com/)
* Google Cloud AI (Gemini) docs — [https://cloud.google.com/ai-platform](https://cloud.google.com/ai-platform)

Use these references to learn provider-specific retry semantics, rate-limiting guidance, and SDK best practices.

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/ai-agents-for-beginner-openclaw-case-study/module/d77598d4-d1d3-4768-97da-03ead60bf984/lesson/20c1b6fc-5c33-4f42-9751-2a729a592bf6" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.