> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Release Management With Shadow Release

> Explains using shadow releases to evaluate new generative AI models on real traffic safely by routing requests to primary and shadow models and analyzing offline metrics.

In this lesson you’ll learn how to manage safe rollouts for Generative AI (GenAI) applications using a shadow release pattern. Shadow releases allow you to validate new models or prompt changes against real traffic without exposing users to unvetted behavior.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/BVCvDn4rl3j0TCQq/images/Introduction-to-Amazon-Bedrock/Advanced-Topics-Optional/Release-Management-With-Shadow-Release/shadow-release-request-routing-bedrock-agentcore.jpg?fit=max&auto=format&n=BVCvDn4rl3j0TCQq&q=85&s=ade3385ec0e775ed374d02ac1bcef840" alt="A slide showing a lecture flow diagram with blue gradient boxes connected by arrows. It outlines steps from &#x22;Problem: Changes to AI systems are unpredictable&#x22; to &#x22;Solution: Shadow Release,&#x22; then &#x22;Workflow: Request Routing,&#x22; leading to &#x22;Results: Safer release management,&#x22; &#x22;Key Takeaway: careful release strategies,&#x22; and &#x22;What's Next: Introduction to Bedrock AgentCore.&#x22;" width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Advanced-Topics-Optional/Release-Management-With-Shadow-Release/shadow-release-request-routing-bedrock-agentcore.jpg" />
</Frame>

Overview: GenAI systems are non-deterministic — a model upgrade, different multi-turn prompt logic, or switching to a new model family can change outputs in ways that are difficult to predict. Shadow releases reduce risk by letting you evaluate candidate models using production traffic while returning only the primary model’s result to users.

<Callout icon="lightbulb" color="#1CB2FE">
  Shadow releases let you "shadow" production requests to candidate models for offline evaluation. This enables evidence-driven decisions for promoting new models or prompt changes without impacting users.
</Callout>

## The problem: unpredictable behavior after release

GenAI models do not always behave like deterministic software. A change that passed tests can still produce unexpected outputs in production because:

* Model versions differ in reasoning, response style, or hallucination behavior.
* Prompt edits can alter answers in subtle ways.
* Tokenization and response length can change cost and latency characteristics.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/BVCvDn4rl3j0TCQq/images/Introduction-to-Amazon-Bedrock/Advanced-Topics-Optional/Release-Management-With-Shadow-Release/genai-unknown-behavior-timeline.jpg?fit=max&auto=format&n=BVCvDn4rl3j0TCQq&q=85&s=a258d08c31ed83254477bc2936aa99fb" alt="A slide titled &#x22;Problem: GenAI Doesn't Behave Like Normal Software&#x22; showing a timeline from &#x22;Current model&#x22; to &#x22;Release&#x22; and &#x22;Ship change&#x22; that ends at &#x22;Unknown behavior&#x22; marked by a red alert icon. It illustrates that generative AI can change after release and produce unexpected behavior." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Advanced-Topics-Optional/Release-Management-With-Shadow-Release/genai-unknown-behavior-timeline.jpg" />
</Frame>

These uncertainties create real user-experience risk after a release. Shadow release is a practical pattern to mitigate that risk.

## Shadow release pattern (high level)

How it works:

* User submits a request to your application.
* The app routes the request to the primary model and returns that response to the user.
* In parallel, the identical request is sent to one or more shadow (candidate) models.
* Shadow responses are recorded (logs, S3, analytics pipeline) and never shown to users.
* Offline analysis compares primary vs. shadow across accuracy, tone, latency, cost, and other metrics to determine promotion decisions.

Typical architecture: API Gateway → Lambda (request router) → Bedrock inference calls to primary and shadow models. The router returns the primary output immediately while storing shadow outputs for later analysis.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/BVCvDn4rl3j0TCQq/images/Introduction-to-Amazon-Bedrock/Advanced-Topics-Optional/Release-Management-With-Shadow-Release/api-gateway-lambda-primary-shadow-routing.jpg?fit=max&auto=format&n=BVCvDn4rl3j0TCQq&q=85&s=0ae348e3bddc6cf9a786a0185c85e2ae" alt="A flow diagram showing an API Gateway sending user requests to a Lambda &#x22;Request Router&#x22; that routes each request to a primary model (returning the primary response to the user) and to multiple shadow models whose responses are stored and compared. The title reads &#x22;Solution: Route Every Request to Both Models.&#x22;" width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Advanced-Topics-Optional/Release-Management-With-Shadow-Release/api-gateway-lambda-primary-shadow-routing.jpg" />
</Frame>

## Implementation notes and options

Where to implement the request router:

* AWS Lambda (simple, serverless).
* Containerized service (ECS, Fargate, EKS) for higher throughput.
* Embedded in a Java/Node.js service if you control the app runtime.
* AWS Step Functions for orchestrating parallel calls and aggregating results with visual workflows.

Design considerations:

* Shadow responses must never affect the primary response path.
* Persist shadow outputs to an analytics store (S3, CloudWatch Logs, or a dedicated pipeline) for comparison.
* Make shadow calls concurrently to avoid adding latency to the primary path (or isolate latency by not waiting for shadow calls in the user-facing response).
* Track token usage and latency per model to estimate cost and performance impact at scale.

<Callout icon="warning" color="#FF6B6B">
  Shadow responses must never be surfaced to users or change the user-facing result. Ensure logging and persistence are isolated from the primary response path.
</Callout>

## Example: Lambda request router (Python)

This example demonstrates the end-to-end flow in a Lambda handler:

* Read incoming request
* Call the primary model and return its result
* Call the shadow model(s) with the same prompt
* Log or persist shadow outputs for offline analysis

Adjust model IDs, retry/backoff logic, and logging sinks for your environment.

```python theme={null}
import json
import boto3
import concurrent.futures

bedrock = boto3.client("bedrock-runtime")

PRIMARY_MODEL = "amazon.nova-lite-v1:0"
SHADOW_MODELS = ["amazon.nova-pro-v1:0"]  # add more candidates as needed

def call_model(model_id, messages):
    """Invoke Bedrock and return the plain text response and metadata."""
    response = bedrock.converse(modelId=model_id, messages=messages)
    text = response["output"]["message"]["content"][0]["text"]
    # Optionally capture tokens, latency, and raw response for analysis
    return {"model": model_id, "text": text, "raw": response}

def lambda_handler(event, context):
    body = json.loads(event.get("body", "{}"))
    user_input = body.get("message", "")

    messages = [{"role": "user", "content": [{"text": user_input}]}]

    # Call primary model synchronously (returned to user)
    primary_result = call_model(PRIMARY_MODEL, messages)
    primary_text = primary_result["text"]

    # Call shadow models concurrently, do not await for response to return to user
    with concurrent.futures.ThreadPoolExecutor(max_workers=len(SHADOW_MODELS)) as executor:
        futures = {
            executor.submit(call_model, model, messages): model for model in SHADOW_MODELS
        }
        for fut in concurrent.futures.as_completed(futures):
            try:
                shadow_result = fut.result()
                # Persist the shadow result for offline analysis (CloudWatch, S3, DB, etc.)
                print("SHADOW:", shadow_result["model"], shadow_result["text"])
                # Optionally store `shadow_result["raw"]` for metric extraction
            except Exception as e:
                # Handle and log individual shadow call errors without impacting the primary response
                print("Shadow call error:", e)

    # Return the primary model's response to the user
    return {"statusCode": 200, "body": json.dumps({"response": primary_text})}
```

## Metrics to collect and compare

Collect multiple dimensions so promotion decisions are evidence-driven:

| Metric | Why it matters | Example storage |
| - | - | - |
| Accuracy / correctness | Ensures the candidate produces correct task outputs | `S3` or analytics DB |
| Tone & alignment | Aligns with product voice, guardrails, safety | `CloudWatch Logs`, review tool |
| Latency | Affects UX and concurrency needs | `CloudWatch Metrics` |
| Cost (tokens) | Different tokenization => different cost at scale | `S3` + token counters |
| Failure/retry rates | Model or network instability | `CloudWatch` |

Tokenization differences can materially affect cost when operating at scale — measure tokens per-request per-model.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/BVCvDn4rl3j0TCQq/images/Introduction-to-Amazon-Bedrock/Advanced-Topics-Optional/Release-Management-With-Shadow-Release/compare-candidate-models-workflow.jpg?fit=max&auto=format&n=BVCvDn4rl3j0TCQq&q=85&s=90ef59e80f1b0ee53ddf9178f8d2ab98" alt="A dark-themed slide titled &#x22;Workflow: Compare Candidate Models&#x22; with four numbered cards showing icons and labels: Accuracy, Tone and consistency, Latency, and Cost (tokens). Each card highlights a criterion for comparing models." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Advanced-Topics-Optional/Release-Management-With-Shadow-Release/compare-candidate-models-workflow.jpg" />
</Frame>

## When to use a shadow release

Use shadow releases for scenarios where behavior may change and you want low-risk evaluation:

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/BVCvDn4rl3j0TCQq/images/Introduction-to-Amazon-Bedrock/Advanced-Topics-Optional/Release-Management-With-Shadow-Release/when-to-use-shadow-release.jpg?fit=max&auto=format&n=BVCvDn4rl3j0TCQq&q=85&s=daf88a5992e5b017b5de510b60d02f8c" alt="A dark-blue slide titled &#x22;Workflow: When to Use&#x22; that lists four reasons to use a shadow release: 1) Model upgrades, 2) Prompt changes, 3) New features (e.g., agents), and 4) Performance improvements." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Advanced-Topics-Optional/Release-Management-With-Shadow-Release/when-to-use-shadow-release.jpg" />
</Frame>

* Model upgrades (new foundation model versions)
* Prompt or system-message changes
* New product features (agents, decision logic, guardrails)
* Performance or cost optimization experiments

You can also use shadow releases alongside other strategies:

* Canary releases: expose a small percentage of live traffic to a candidate model.
* A/B testing: return different content to different users for direct comparison (user-visible).
  Shadow release differs because it uses real traffic while keeping user experience unchanged.

## Expected outcomes and benefits

* Safer production releases: observe candidate behavior on real requests before promotion.
* Evidence-driven promotion: promote models based on measured metrics rather than intuition.
* Continuous improvement: iterate on prompts, guardrails, and model selection with minimal user risk.
* Operational visibility: capture behavior changes (latency, token usage, hallucinations) early.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/BVCvDn4rl3j0TCQq/images/Introduction-to-Amazon-Bedrock/Advanced-Topics-Optional/Release-Management-With-Shadow-Release/genai-release-strategy-shadow-releases-risk.jpg?fit=max&auto=format&n=BVCvDn4rl3j0TCQq&q=85&s=bd923dce684a2161bf322ef6012c0a10" alt="A presentation slide titled &#x22;Key Takeaways&#x22; with three numbered points. The points state: GenAI systems need careful release strategies; shadow releases safely evaluate real-world changes; and these practices improve systems while minimizing risk." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Advanced-Topics-Optional/Release-Management-With-Shadow-Release/genai-release-strategy-shadow-releases-risk.jpg" />
</Frame>

## Key takeaways

* GenAI systems require careful release strategies because model/version changes can produce very different results.
* Shadow releases let you evaluate real-world changes safely without exposing users to unvetted behavior.
* Combine metrics (accuracy, tone, latency, and cost) and automated pipelines to make promotion decisions reliable and repeatable.

This concludes the short lesson on release management with shadow releases. The course also includes an introduction to Bedrock AgentCore and orchestration techniques for more advanced release workflows.

## Links and references

* Amazon Bedrock documentation: [https://docs.aws.amazon.com/bedrock/](https://docs.aws.amazon.com/bedrock/)
* AWS Lambda: [https://docs.aws.amazon.com/lambda/](https://docs.aws.amazon.com/lambda/)
* AWS Step Functions: [https://docs.aws.amazon.com/step-functions/](https://docs.aws.amazon.com/step-functions/)
* Best practices for canary and A/B testing: [https://aws.amazon.com/what-is/continuous-delivery/](https://aws.amazon.com/what-is/continuous-delivery/)

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/introduction-to-amazon-bedrock/module/7af9f623-7d4e-447a-8b21-6e635dfaccfa/lesson/b1c120e7-dc69-4f96-81a7-4c265c5c9698" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.