> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Multimodal Capabilities Text and Images Part 2

> Explains streaming multimodal responses and image generation workflows, including handling image bytes, Base64 decoding, and choosing appropriate API endpoints for low latency and image creation.

In this lesson we cover streaming responses from a multimodal model and the typical image-generation flow. Streaming changes only how the model delivers output — not how you structure the request. You still provide messages, image bytes, and other multimodal inputs the same way; the SDK method is what differs (for example, `ConverseStream` vs `Converse`).

## Streaming multimodal responses

Streaming delivers generated content incrementally as the model produces it. This is useful for low-latency UI updates (progressive rendering, partial transcript display) or when you want to display text as it arrives rather than waiting for the full response.

Below is a concise Python example that reads an image from disk, sends it to a streaming Converse API, and prints incoming text deltas in real time.

```python theme={null}
import boto3

client = boto3.client("bedrock-runtime", region_name="us-east-1")

with open("diagram.png", "rb") as f:
    image_bytes = f.read()

response = client.converse_stream(
    modelId="amazon.nova-lite-v1:0",
    messages=[
        {
            "role": "user",
            "content": [
                {"text": "Explain, step by step, how to analyze a process diagram image."},
                {
                    "image": {
                        "format": "png",
                        "source": {"bytes": image_bytes}
                    }
                }
            ]
        }
    ]
)

# Iterate over streaming events and print text deltas as they arrive
for event in response.get("stream", []):
    if "contentBlockDelta" in event:
        delta = event["contentBlockDelta"]["delta"]
        if "text" in delta:
            print(delta["text"], end="", flush=True)
```

What to look for in the streaming loop:

* Each streaming event may contain a `contentBlockDelta`.
* The `delta` object holds incremental content chunks.
* In the example we extract `delta["text"]` and print it to standard output, appending chunks together in real time.

## Image generation (model invoke vs conversation)

When the goal is to generate images, many foundation models expose an `invoke_model`-style endpoint rather than a conversational `Converse` endpoint. Typical behavior:

* You send a text prompt (and optionally image parameters such as size, count).
* The model returns image data, usually Base64-encoded strings.
* Decode the Base64 before saving or rendering.

<Callout icon="lightbulb" color="#1CB2FE">
  When a model returns generated images, they are commonly Base64-encoded in the response payload. Decode them (for example with Python's `base64.b64decode`) before writing to disk or otherwise using the bytes.
</Callout>

Example: request an image from a model and write the first returned image to a PNG file.

```python theme={null}
import boto3
import json
import base64

client = boto3.client("bedrock-runtime", region_name="us-east-1")

# Build the request body for image generation
body = {
    "text": [
        {
            "role": "user",
            "content": [
                {
                    "text": "A futuristic classroom with holographic displays and students learning AI"
                }
            ]
        }
    ],
    "image": {
        "size": {"width": 1024, "height": 1024},
        "count": 1
    }
}

# Call the model that supports image generation (e.g., Nova Canvas)
response = client.invoke_model(
    modelId="amazon.nova-canvas-1:latest",
    body=json.dumps(body),
    contentType="application/json"
)

# Read and parse the response body
response_body = response["body"].read().decode("utf-8")
result = json.loads(response_body)

# Extract the first image (Base64) and write to file
image_b64 = result["images"][0]
with open("generated.png", "wb") as out_f:
    out_f.write(base64.b64decode(image_b64))
```

Notes about this image-generation flow:

* Request/response schema for `invoke_model` can vary by model. Adapt `body` to the model's expected structure.
* The `images` element often contains Base64 strings; some models may wrap these strings in metadata objects — adjust extraction accordingly.
* Instead of writing to disk, you can stream decoded bytes to S3, serve them directly to a UI, or embed them in an HTTP response.

## Quick comparison

| Use case | Endpoint / method | Common response format | When to use |
| - | -: | - | - |
| Streaming conversational output (text) | `ConverseStream` | Incremental `contentBlockDelta` events | Low-latency partial text display |
| Non-stream conversational output | `Converse` | Full generated text in single response | Simpler flow, small responses |
| Image generation | `invoke_model` (or model-specific) | Base64 `images` array (often) | Generate and store/serve images |

## Why multimodal matters

Multimodal models let applications accept real-world inputs (images, documents, scans) and convert them into structured, actionable information. For example:

* A model can extract order numbers, line items, and totals from a scanned purchase order.
* Extracted data can trigger downstream workflows (ERP, billing, fulfillment) without manual review.
* This leads to richer UX, more natural interactions, and closer integration with enterprise systems.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/BVCvDn4rl3j0TCQq/images/Introduction-to-Amazon-Bedrock/Advanced-Topics-Optional/Multimodal-Capabilities-Text-and-Images-Part-2/results-slide-five-benefits.jpg?fit=max&auto=format&n=BVCvDn4rl3j0TCQq&q=85&s=10e4d42f7934f78d540757accf4751fc" alt="A slide titled &#x22;Results&#x22; showing five numbered cards with colorful circular icons and short captions. The captions list benefits like richer UX accepting images/documents, better enterprise document workflows, more natural app interactions, broader app capabilities, and better model selection." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Advanced-Topics-Optional/Multimodal-Capabilities-Text-and-Images-Part-2/results-slide-five-benefits.jpg" />
</Frame>

## Summary

* Streaming changes delivery, not request structure — send the same messages and image bytes but use a streaming API to receive deltas.
* For image generation, decode Base64-encoded outputs before saving or serving.
* Choose the right foundation model and adapt the request/response handling to the model’s expected schema.
* Store or render artifacts according to your application needs (file system, S3, or direct UI streaming).

## Links and references

* [Amazon Bedrock documentation](https://docs.aws.amazon.com/bedrock/)
* Boto3 client reference: [https://boto3.amazonaws.com/v1/documentation/api/latest/reference/services/bedrock-runtime.html](https://boto3.amazonaws.com/v1/documentation/api/latest/reference/services/bedrock-runtime.html)
* Python base64 decoding: [https://docs.python.org/3/library/base64.html](https://docs.python.org/3/library/base64.html)

This lesson wraps up multimodal streaming and image generation basics. Next logical topics include managing conversation context, stateful chat flows, and advanced grounding techniques for reliable multimodal pipelines.

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/introduction-to-amazon-bedrock/module/7af9f623-7d4e-447a-8b21-6e635dfaccfa/lesson/b2a114f1-bd2e-4f98-8050-0c117a530ace" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.