
Where failures typically occur
When developing a Bedrock-based application, most problems fall into three categories:- Model output
- Infrastructure
- Inputs and context

Quick reference: failure types and mitigations
Solution overview
To build resilience, adopt these core practices:- Implement error handling and retry logic in application code.
- Validate and sanitize model outputs — never treat them as production-ready by default.
- Apply guardrails to filter unsafe or disallowed prompts and outputs.
- Validate inputs to enforce required formats, prompt length, language, and schema.
- Add structured logging and observability to capture failure patterns and frequency.
Guardrails and safety (important)
Implement input and output guardrails that detect disallowed or unsafe content (for example, instructions to commit fraud or other illegal acts). Block or sanitize such prompts before sending them to the model and log these events for review.
Workflow considerations
A robust workflow detects unexpected conditions, decides whether to retry or fail, and provides a clear fallback path. Typical branching logic:- Detect error (API error, timeout, malformed response).
- Classify the error: retryable (transient network or throttling) vs non-retryable (invalid input, logic error, guardrail violation).
- If retryable: perform limited retries with exponential backoff. If still failing, fall back.
- If non-retryable: return a helpful error or ask the user to rephrase/provide missing context.
- For malformed responses: consider modifying the prompt (e.g., stricter instructions, schema enforcement) instead of blindly resubmitting the same request.
- Log all events and metrics: error counts, latencies, retry attempts, validation failures.

Example: handling an API call failure in Python
Below is an improved Python example showing a Bedrock Runtime call with classification of errors, limited retries using exponential backoff, and basic response validation. In production, replaceprint statements with structured logging and push metrics to your observability backend.
- Separate transient vs non-transient errors and only retry appropriate cases.
- Use exponential backoff with a retry cap to avoid thundering herds and excessive costs.
- Validate response content (JSON/schema). If the model returns malformed content, consider a fix-and-retry with stricter prompt instructions.
- In production, replace prints with structured logs and emit metrics for retries, failures, and latencies.
Best practices recap for code-level handling
- Distinguish error classes: network/timeouts (retryable), 5xx (usually retryable), 4xx (usually non-retryable), malformed content (repair with re-prompt).
- Limit retries and implement jitter to avoid synchronized retries.
- Validate output against a schema when expecting structured data (e.g., JSON schema).
- Log both application-level and model-level anomalies to help refine prompts and guardrails over time.
Always validate and sanitize model outputs before consuming them downstream. Logging and metrics are essential for diagnosing recurring failures and improving prompts and workflows.
Summary / Key takeaways
- Failures come from model outputs, infrastructure, or inputs/context — identify which class you’re dealing with.
- Implement structured error handling and retries for transient failures; use clear fallback behavior for persistent failures.
- Sanitize and validate every model response; enforce JSON/schema constraints if you rely on structured output.
- Add logging and observability so you can measure error rates and iterate on prompts, guardrails, and system design.
Links and references
- Amazon Bedrock documentation
- boto3 documentation
- Botocore exceptions guide
- Article on retry and exponential backoff patterns: https://aws.amazon.com/blogs/architecture/exponential-backoff-and-jitter/