ConverseStream vs Converse).
Streaming multimodal responses
Streaming delivers generated content incrementally as the model produces it. This is useful for low-latency UI updates (progressive rendering, partial transcript display) or when you want to display text as it arrives rather than waiting for the full response. Below is a concise Python example that reads an image from disk, sends it to a streaming Converse API, and prints incoming text deltas in real time.- Each streaming event may contain a
contentBlockDelta. - The
deltaobject holds incremental content chunks. - In the example we extract
delta["text"]and print it to standard output, appending chunks together in real time.
Image generation (model invoke vs conversation)
When the goal is to generate images, many foundation models expose aninvoke_model-style endpoint rather than a conversational Converse endpoint. Typical behavior:
- You send a text prompt (and optionally image parameters such as size, count).
- The model returns image data, usually Base64-encoded strings.
- Decode the Base64 before saving or rendering.
When a model returns generated images, they are commonly Base64-encoded in the response payload. Decode them (for example with Python’s
base64.b64decode) before writing to disk or otherwise using the bytes.- Request/response schema for
invoke_modelcan vary by model. Adaptbodyto the model’s expected structure. - The
imageselement often contains Base64 strings; some models may wrap these strings in metadata objects — adjust extraction accordingly. - Instead of writing to disk, you can stream decoded bytes to S3, serve them directly to a UI, or embed them in an HTTP response.
Quick comparison
Why multimodal matters
Multimodal models let applications accept real-world inputs (images, documents, scans) and convert them into structured, actionable information. For example:- A model can extract order numbers, line items, and totals from a scanned purchase order.
- Extracted data can trigger downstream workflows (ERP, billing, fulfillment) without manual review.
- This leads to richer UX, more natural interactions, and closer integration with enterprise systems.

Summary
- Streaming changes delivery, not request structure — send the same messages and image bytes but use a streaming API to receive deltas.
- For image generation, decode Base64-encoded outputs before saving or serving.
- Choose the right foundation model and adapt the request/response handling to the model’s expected schema.
- Store or render artifacts according to your application needs (file system, S3, or direct UI streaming).
Links and references
- Amazon Bedrock documentation
- Boto3 client reference: https://boto3.amazonaws.com/v1/documentation/api/latest/reference/services/bedrock-runtime.html
- Python base64 decoding: https://docs.python.org/3/library/base64.html