
Codey operates in the real world: network errors, runtime exceptions, timeouts, and side effects happen frequently. Design agents to detect and react to these failures instead of hiding them.

Overview — common failure modes
There are six common failure modes to design for:
Below we walk through practical strategies for each failure mode.
Tool failures
When a tool call raises an exception, do not swallow it. Instead, make the failure visible to the agent so the LLM can retry, choose a different tool, or ask for clarification. Wrap external tool calls in try/except, capture diagnostic information (exception type and traceback), and present that back to the agent as a tool result. Example pattern:Bad tool selection
Bad tool selection is often a design problem rather than a runtime exception. If the agent consistently picks the wrong tool:- Improve tool descriptions and add concrete examples.
- Include explicit negative examples: show when not to use the tool.
- Constrain tool usage with lightweight validators or type checks before execution.
- Use server-side validation to reject tool calls with invalid argument shapes.
Hallucinated tool calls
Agents sometimes invent a tool name or pass invalid arguments. Always validate tool calls on the platform side. If a tool name is unknown or arguments are malformed, return a structured error back to the agent:- Return machine-readable error codes plus human-readable diagnosis.
- Suggest the correct tool or required argument schema where possible.
- Log the hallucinated call for monitoring and retraining.
Infinite loops
Unbounded iteration is costly and can escalate bills or resource usage. Always set hard iteration limits, time budgets, and inspect the model’s finish reason when available. Example iteration guard:Never allow unbounded execution of user-defined scripts or arbitrary shell commands. Always sandbox and constrain resource usage to avoid runaway processes or security risks.
Context overflow
When a conversation grows past the model’s context window, compress older content. Keep recent exchanges in full fidelity and summarize or embed earlier ones to preserve the thread while freeing tokens for current reasoning. Common approaches:- Periodically summarize earlier messages and replace them with compressed summaries.
- Use embeddings to store long-term memory and retrieve only relevant snippets.
- Detect when the context is nearing capacity and trigger compression automatically.

Model errors and provider fallback
Model APIs can fail with rate limits, outages, or malformed responses. Implement provider fallback: try a primary provider, and if it fails or returns an invalid response, retry with a secondary provider. Normalize request and response formats so switching providers does not require changing higher-level logic. Example fallback sketch:Five core principles for robust error handling
- Expect failure. Design for tool and model failures from day one.
- Surface errors to the agent. Don’t silently swallow failures — let the LLM adapt.
- Set hard limits. Constrain iterations, tokens, and time to prevent runaway costs.
- Log everything. Detailed logs (including tool inputs/outputs and tracebacks) make diagnosis and reproduction possible.
- Fail gracefully to users. Return clear, actionable messages instead of cryptic traces.
Implementation summary
This implementation bundles the patterns above:- Model fallback across providers.
- Tool error recovery that passes errors back to the agent as structured results.
- Context compression for long conversations.
- Sandboxed execution for untrusted code and system commands.
- Rate limits and iteration guards to prevent runaway loops.

Further reading and references
- OpenAI API — https://platform.openai.com/docs
- Anthropic/Claude docs — https://www.anthropic.com/
- Google Cloud AI (Gemini) docs — https://cloud.google.com/ai-platform