Skip to main content
This article covers practical patterns for making Bedrock runtime calls robust: retries with exponential backoff, validating model outputs, handling RAG (retrieval-augmented generation) edge cases, managing token limits, and dealing with safety/guardrails. Each section includes code examples and actionable best practices you can adopt.

Retry loop with exponential backoff (boto3 example)

Wrap calls to the Bedrock runtime client with an explicit retry loop and exponential backoff to handle transient errors such as throttling. The example below shows:
  • Creating a boto3 Bedrock Runtime client.
  • Iterating attempts from 1..N for human-friendly logging.
  • Retrying only on specific error codes (e.g., ThrottlingException).
  • Breaking early on success and logging failures when unretryable or final attempt fails.
Best-practice notes:
  • Use range(1, max_retries + 1) so logs show attempts starting at 1.
  • 2 ** attempt implements exponential backoff. Consider adding jitter (randomized sleep) to avoid synchronized retries across clients.
  • Retry selectively — limit retries to transient/throttling errors. For other errors, log and stop.
  • Record metrics (counts of throttles, retries, latency) so you can monitor trends and set alarms.
Retry decision matrix (quick reference):

Validate model output (guard against malformed responses)

Always validate and parse model responses before using them. Model output may be malformed JSON, missing keys, or unexpected types. Handle parsing errors and validation issues explicitly so you can decide whether to sanitize, retry, or return a controlled error.
Validation tips:
  • Catch json.JSONDecodeError for invalid JSON.
  • Use explicit validation: check required fields and types; raise ValueError or a custom exception to unify handling.
  • Decide a policy per validation failure: retry the model call, sanitize the response, ask the model to reformat, or return a safe fallback to the user.
  • Log schema failures with examples (scrub any PII) to improve prompts or post-processing.

Retrieval-Augmented Generation (RAG) — when no context is found

When your retrieval step returns an empty set (no relevant chunks), define how your application should respond to avoid hallucination and unnecessary cost. Typical strategies:
  • Proceed and tell the model explicitly that no supporting context was found.
  • Refuse to answer and prompt the user to rephrase or provide more detail.
  • Return a safe fallback message that avoids speculation.
Example:
When retrieval returns no context, make that explicit to the user (e.g., “No supporting documents found”). This reduces the risk of the model producing plausible but ungrounded (hallucinated) answers.

Handling token limits and context constraints

Models enforce a fixed context window (input tokens + output tokens). Exceeding the model’s limit can cause request failures or truncated output. Use these strategies to manage token constraints:
  • Summarize or truncate long inputs before sending.
  • Limit retrieved chunks in RAG (e.g., top 3 instead of top 10).
  • Set an explicit max tokens (or equivalent) to cap the output size.
  • Preprocess or compress retrieved documents to extract the most relevant spans.
  • Break tasks into multiple calls or multi-step workflows when appropriate.
A dark-themed infographic titled "Workflow: Handling Token Limits and Context Constraints" showing input and output progress bars with red "MAX" markers and an orange warning about exceeding limits. On the right it lists strategies like truncating or summarizing input, limiting retrieved chunks (RAG), and constraining output length.
Example prompt guidance:
  • Instruct the model with explicit output constraints, for example: “Summarize this document in one paragraph of no more than 200 words.”
  • Use server-side limits in your API call to enforce a maximum number of output tokens.

Handling safety responses and guardrails

When input violates safety policies (illegal, unsafe, or otherwise disallowed), either block the request client-side or rely on managed guardrails. Provide clear, user-friendly feedback and do not surface internal implementation details. Guidelines:
  • Return a short, clear message such as: “This prompt cannot be processed because it violates our safety policy.”
  • Avoid exposing internal guardrail names, policy internals, or implementation details.
  • Degrade the UI gracefully (do not crash the app).
  • Log safety events (without sensitive user content) so you can audit trends and tune safeguards.
A slide titled "Workflow: Handling Safety Responses (Guardrails)" listing four guidelines for blocked content: provide clear user-friendly feedback, do not expose internal details, handle the response gracefully in the app, and log and monitor safety events.
Always log safety-related interceptions (without exposing sensitive user data) so you can audit and improve your detection logic and understand trends in abusive or dangerous inputs.

Quick checklist before deploying to production

  • Implement selective retries with exponential backoff and jitter.
  • Validate model outputs and fail early when schema expectations are unmet.
  • Define clear RAG fallback behavior for empty retrievals.
  • Enforce token limits via input compression and API-level output caps.
  • Display user-friendly messages for safety blocks and log events securely.
  • Instrument metrics for throttling, retries, validation failures, and safety interceptions.

Watch Video