> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Handling Errors Troubleshooting and Edge Cases Part 1

> Guidance on error handling, retries, input and output validation, guardrails, and observability for building resilient Amazon Bedrock GenAI applications, with a Python retry example

In this lesson we cover error handling, troubleshooting, and edge cases for GenAI applications built on Amazon Bedrock. The goal is to design resilient, predictable workflows that detect failures, recover when possible, and fail gracefully when necessary. You’ll see practical guidance, mitigation patterns, and a Python example demonstrating API-level error handling and retries.

Below is the problem → solution → workflow → results flow that guides this lesson.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/BVCvDn4rl3j0TCQq/images/Introduction-to-Amazon-Bedrock/Best-Practices-and-Optimization/Handling-Errors-Troubleshooting-and-Edge-Cases-Part-1/lecture-flow-resilient-genai-error-handling.jpg?fit=max&auto=format&n=BVCvDn4rl3j0TCQq&q=85&s=e12a49ecda55a64e60ceb21815704b39" alt="A slide titled &#x22;Lecture Flow&#x22; showing a flowchart of boxes (Problem → Solution → Workflow → Results → Key Takeaway → What's Next) about designing resilient GenAI systems and implementing error-handling. It notes that GenAI can fail unpredictably and recommends building controlled, failure-tolerant workflows for more reliable AI." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Best-Practices-and-Optimization/Handling-Errors-Troubleshooting-and-Edge-Cases-Part-1/lecture-flow-resilient-genai-error-handling.jpg" />
</Frame>

## Where failures typically occur

When developing a Bedrock-based application, most problems fall into three categories:

* Model output
* Infrastructure
* Inputs and context

Each category has distinct symptoms and mitigation strategies.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/BVCvDn4rl3j0TCQq/images/Introduction-to-Amazon-Bedrock/Best-Practices-and-Optimization/Handling-Errors-Troubleshooting-and-Edge-Cases-Part-1/problem-failure-sources-model-infra-inputs.jpg?fit=max&auto=format&n=BVCvDn4rl3j0TCQq&q=85&s=cca69a301755cbb456fa0a13818286d4" alt="A dark-themed slide titled &#x22;Problem&#x22; with three panels. It lists failure sources: Model Output (incomplete/malformed responses, hallucinations, token limits, guardrail violations), Infrastructure (API errors/timeouts, throttling, dependency failures), and Inputs & Context (empty retrievals, prompt edge cases, invalid input formats)." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Best-Practices-and-Optimization/Handling-Errors-Troubleshooting-and-Edge-Cases-Part-1/problem-failure-sources-model-infra-inputs.jpg" />
</Frame>

### Quick reference: failure types and mitigations

| Failure type | Common symptoms | Suggested mitigations |
| - | - | - |
| Model output | Incomplete/malformed responses, hallucinations, irrelevant content, token-limit truncation, guardrail violations | Enforce output schema, sanitize & validate outputs, re-prompt with clarifying instructions, apply safety filters and guardrails |
| Infrastructure | API errors, timeouts, throttling, downstream service failures (e.g., vector store, search) | Add retries with exponential backoff for transient errors, circuit breakers, timeouts, graceful fallbacks, monitoring/alerts |
| Inputs & context | Empty retrieval results, overly long or empty prompts, unsupported languages, invalid formats | Validate inputs, enforce required schemas and prompt lengths, provide user guidance and fallback responses |

## Solution overview

To build resilience, adopt these core practices:

* Implement error handling and retry logic in application code.
* Validate and sanitize model outputs — never treat them as production-ready by default.
* Apply guardrails to filter unsafe or disallowed prompts and outputs.
* Validate inputs to enforce required formats, prompt length, language, and schema.
* Add structured logging and observability to capture failure patterns and frequency.

### Guardrails and safety (important)

<Callout icon="warning" color="#FF6B6B">
  Implement input and output guardrails that detect disallowed or unsafe content (for example, instructions to commit fraud or other illegal acts). Block or sanitize such prompts before sending them to the model and log these events for review.
</Callout>

## Workflow considerations

A robust workflow detects unexpected conditions, decides whether to retry or fail, and provides a clear fallback path. Typical branching logic:

* Detect error (API error, timeout, malformed response).
* Classify the error: retryable (transient network or throttling) vs non-retryable (invalid input, logic error, guardrail violation).
* If retryable: perform limited retries with exponential backoff. If still failing, fall back.
* If non-retryable: return a helpful error or ask the user to rephrase/provide missing context.
* For malformed responses: consider modifying the prompt (e.g., stricter instructions, schema enforcement) instead of blindly resubmitting the same request.
* Log all events and metrics: error counts, latencies, retry attempts, validation failures.

If you sketch the flow:

User request → call Bedrock model → on success: validate & return → on failure: retry/repair or return fallback.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/BVCvDn4rl3j0TCQq/images/Introduction-to-Amazon-Bedrock/Best-Practices-and-Optimization/Handling-Errors-Troubleshooting-and-Edge-Cases-Part-1/bedrock-error-handling-retry-fallback-flow.jpg?fit=max&auto=format&n=BVCvDn4rl3j0TCQq&q=85&s=7f2b9f9fa17d323883318d81f24f1e8c" alt="A dark-themed flowchart titled &#x22;Workflow: Error-Handling Workflow.&#x22; It diagrams steps from User Request to calling a Bedrock model, then branching on success (validate/return response) or failure (retry, handle/fix, or return a fallback)." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Best-Practices-and-Optimization/Handling-Errors-Troubleshooting-and-Edge-Cases-Part-1/bedrock-error-handling-retry-fallback-flow.jpg" />
</Frame>

You may also involve the user: prompt for clarification or ask them to rephrase when automatic corrections are inappropriate.

## Example: handling an API call failure in Python

Below is an improved Python example showing a Bedrock Runtime call with classification of errors, limited retries using exponential backoff, and basic response validation. In production, replace `print` statements with structured logging and push metrics to your observability backend.

```python theme={null}
import json
import time
import logging
import boto3
from botocore.exceptions import ClientError, ReadTimeoutError, EndpointConnectionError

logger = logging.getLogger("bedrock-client")
logger.setLevel(logging.INFO)

bedrock = boto3.client("bedrock-runtime", region_name="us-east-1")

def invoke_with_retries(model_id, payload, max_retries=3, base_backoff=1.0):
    attempt = 0
    while True:
        attempt += 1
        try:
            response = bedrock.invoke_model(
                modelId=model_id,
                body=json.dumps(payload),
                contentType="application/json",
                accept="application/json"
            )
            # Example: read the response body (may be streaming depending on SDK)
            body = response.get("body")
            if hasattr(body, "read"):
                text = body.read().decode("utf-8")
            else:
                text = body

            # Basic validation: ensure we have JSON (enforce schema in production)
            try:
                parsed = json.loads(text)
            except json.JSONDecodeError:
                logger.warning("Model returned non-JSON; will attempt a single repair retry.")
                raise ValueError("Invalid JSON in model response")

            # TODO: run JSON schema validation here if you depend on structured output
            return parsed

        except (ReadTimeoutError, EndpointConnectionError) as e:
            # Transient network issues: retry with exponential backoff
            if attempt <= max_retries:
                backoff = base_backoff * (2 ** (attempt - 1))
                logger.warning("Transient error: %s. Retrying in %.1fs (attempt %d/%d)", e, backoff, attempt, max_retries)
                time.sleep(backoff)
                continue
            else:
                logger.error("Exceeded retries for transient error: %s", e)
                raise

        except ClientError as e:
            # ClientError can include 4xx/5xx errors. Inspect status code/message to decide action.
            code = e.response.get("Error", {}).get("Code", "")
            msg = e.response.get("Error", {}).get("Message", "")
            logger.error("Bedrock ClientError: %s - %s", code, msg)

            # Example logic: do not retry on 4xx; consider retry on 5xx or throttling
            status_code = e.response.get("ResponseMetadata", {}).get("HTTPStatusCode")
            if status_code and 500 <= status_code < 600 and attempt <= max_retries:
                backoff = base_backoff * (2 ** (attempt - 1))
                logger.warning("Server error (%d). Retrying in %.1fs...", status_code, backoff)
                time.sleep(backoff)
                continue
            else:
                raise

        except ValueError as e:
            # Non-JSON or malformed model output: attempt a single re-prompt or fallback
            logger.info("Malformed model output: %s", e)
            # Optionally: modify the prompt (e.g., add "Return valid JSON matching schema X") and retry once
            if attempt <= 1:
                payload["inputText"] = "Please respond with valid JSON only. " + payload.get("inputText", "")
                continue
            else:
                raise

# Usage
payload = {"inputText": "Summarize this review"}
try:
    result = invoke_with_retries("amazon.nova-lite-v1:0", payload)
    # Process result: validate schema, sanitize content, then return to user
    print("Parsed model result:", result)
except Exception as exc:
    # Provide a graceful fallback or user-facing error
    logger.exception("Failed to get a valid response from Bedrock: %s", exc)
    print("Unable to provide an answer at this time. Please try again or rephrase your request.")
```

Notes on this example:

* Separate transient vs non-transient errors and only retry appropriate cases.
* Use exponential backoff with a retry cap to avoid thundering herds and excessive costs.
* Validate response content (JSON/schema). If the model returns malformed content, consider a fix-and-retry with stricter prompt instructions.
* In production, replace prints with structured logs and emit metrics for retries, failures, and latencies.

### Best practices recap for code-level handling

* Distinguish error classes: network/timeouts (retryable), 5xx (usually retryable), 4xx (usually non-retryable), malformed content (repair with re-prompt).
* Limit retries and implement jitter to avoid synchronized retries.
* Validate output against a schema when expecting structured data (e.g., JSON schema).
* Log both application-level and model-level anomalies to help refine prompts and guardrails over time.

<Callout icon="lightbulb" color="#1CB2FE">
  Always validate and sanitize model outputs before consuming them downstream. Logging and metrics are essential for diagnosing recurring failures and improving prompts and workflows.
</Callout>

## Summary / Key takeaways

* Failures come from model outputs, infrastructure, or inputs/context — identify which class you're dealing with.
* Implement structured error handling and retries for transient failures; use clear fallback behavior for persistent failures.
* Sanitize and validate every model response; enforce JSON/schema constraints if you rely on structured output.
* Add logging and observability so you can measure error rates and iterate on prompts, guardrails, and system design.

## Links and references

* [Amazon Bedrock documentation](https://docs.aws.amazon.com/bedrock/)
* [boto3 documentation](https://boto3.amazonaws.com/v1/documentation/api/latest/index.html)
* [Botocore exceptions guide](https://botocore.amazonaws.com/v1/documentation/api/latest/reference/exceptions.html)
* Article on retry and exponential backoff patterns: [https://aws.amazon.com/blogs/architecture/exponential-backoff-and-jitter/](https://aws.amazon.com/blogs/architecture/exponential-backoff-and-jitter/)

Further guidance on guardrails and enforcing safety and output constraints in Bedrock integrations is available in the Bedrock documentation and security best-practices.

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/introduction-to-amazon-bedrock/module/1a696c4d-73f8-4ae4-bcc4-cfbe9c6f03ff/lesson/3b4f1c0d-7468-44ce-bfe2-295e00245c56" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.