Retry loop with exponential backoff (boto3 example)
Wrap calls to the Bedrock runtime client with an explicit retry loop and exponential backoff to handle transient errors such as throttling. The example below shows:- Creating a boto3 Bedrock Runtime client.
- Iterating attempts from 1..N for human-friendly logging.
- Retrying only on specific error codes (e.g.,
ThrottlingException). - Breaking early on success and logging failures when unretryable or final attempt fails.
- Use
range(1, max_retries + 1)so logs show attempts starting at 1. 2 ** attemptimplements exponential backoff. Consider adding jitter (randomized sleep) to avoid synchronized retries across clients.- Retry selectively — limit retries to transient/throttling errors. For other errors, log and stop.
- Record metrics (counts of throttles, retries, latency) so you can monitor trends and set alarms.
Validate model output (guard against malformed responses)
Always validate and parse model responses before using them. Model output may be malformed JSON, missing keys, or unexpected types. Handle parsing errors and validation issues explicitly so you can decide whether to sanitize, retry, or return a controlled error.- Catch
json.JSONDecodeErrorfor invalid JSON. - Use explicit validation: check required fields and types; raise
ValueErroror a custom exception to unify handling. - Decide a policy per validation failure: retry the model call, sanitize the response, ask the model to reformat, or return a safe fallback to the user.
- Log schema failures with examples (scrub any PII) to improve prompts or post-processing.
Retrieval-Augmented Generation (RAG) — when no context is found
When your retrieval step returns an empty set (no relevant chunks), define how your application should respond to avoid hallucination and unnecessary cost. Typical strategies:- Proceed and tell the model explicitly that no supporting context was found.
- Refuse to answer and prompt the user to rephrase or provide more detail.
- Return a safe fallback message that avoids speculation.
When retrieval returns no context, make that explicit to the user (e.g., “No supporting documents found”). This reduces the risk of the model producing plausible but ungrounded (hallucinated) answers.
Handling token limits and context constraints
Models enforce a fixed context window (input tokens + output tokens). Exceeding the model’s limit can cause request failures or truncated output. Use these strategies to manage token constraints:- Summarize or truncate long inputs before sending.
- Limit retrieved chunks in RAG (e.g., top 3 instead of top 10).
- Set an explicit
max tokens(or equivalent) to cap the output size. - Preprocess or compress retrieved documents to extract the most relevant spans.
- Break tasks into multiple calls or multi-step workflows when appropriate.

- Instruct the model with explicit output constraints, for example: “Summarize this document in one paragraph of no more than 200 words.”
- Use server-side limits in your API call to enforce a maximum number of output tokens.
Handling safety responses and guardrails
When input violates safety policies (illegal, unsafe, or otherwise disallowed), either block the request client-side or rely on managed guardrails. Provide clear, user-friendly feedback and do not surface internal implementation details. Guidelines:- Return a short, clear message such as: “This prompt cannot be processed because it violates our safety policy.”
- Avoid exposing internal guardrail names, policy internals, or implementation details.
- Degrade the UI gracefully (do not crash the app).
- Log safety events (without sensitive user content) so you can audit trends and tune safeguards.

Always log safety-related interceptions (without exposing sensitive user data) so you can audit and improve your detection logic and understand trends in abusive or dangerous inputs.
Quick checklist before deploying to production
- Implement selective retries with exponential backoff and jitter.
- Validate model outputs and fail early when schema expectations are unmet.
- Define clear RAG fallback behavior for empty retrievals.
- Enforce token limits via input compression and API-level output caps.
- Display user-friendly messages for safety blocks and log events securely.
- Instrument metrics for throttling, retries, validation failures, and safety interceptions.
Links and references
- Amazon Bedrock documentation
- boto3 — AWS SDK for Python
- botocore exceptions reference
- Best practices for retriable API requests (exponential backoff and jitter) — see many cloud SDK guides for reference.