Skip to main content
In this lesson we cover streaming responses from a multimodal model and the typical image-generation flow. Streaming changes only how the model delivers output — not how you structure the request. You still provide messages, image bytes, and other multimodal inputs the same way; the SDK method is what differs (for example, ConverseStream vs Converse).

Streaming multimodal responses

Streaming delivers generated content incrementally as the model produces it. This is useful for low-latency UI updates (progressive rendering, partial transcript display) or when you want to display text as it arrives rather than waiting for the full response. Below is a concise Python example that reads an image from disk, sends it to a streaming Converse API, and prints incoming text deltas in real time.
What to look for in the streaming loop:
  • Each streaming event may contain a contentBlockDelta.
  • The delta object holds incremental content chunks.
  • In the example we extract delta["text"] and print it to standard output, appending chunks together in real time.

Image generation (model invoke vs conversation)

When the goal is to generate images, many foundation models expose an invoke_model-style endpoint rather than a conversational Converse endpoint. Typical behavior:
  • You send a text prompt (and optionally image parameters such as size, count).
  • The model returns image data, usually Base64-encoded strings.
  • Decode the Base64 before saving or rendering.
When a model returns generated images, they are commonly Base64-encoded in the response payload. Decode them (for example with Python’s base64.b64decode) before writing to disk or otherwise using the bytes.
Example: request an image from a model and write the first returned image to a PNG file.
Notes about this image-generation flow:
  • Request/response schema for invoke_model can vary by model. Adapt body to the model’s expected structure.
  • The images element often contains Base64 strings; some models may wrap these strings in metadata objects — adjust extraction accordingly.
  • Instead of writing to disk, you can stream decoded bytes to S3, serve them directly to a UI, or embed them in an HTTP response.

Quick comparison

Why multimodal matters

Multimodal models let applications accept real-world inputs (images, documents, scans) and convert them into structured, actionable information. For example:
  • A model can extract order numbers, line items, and totals from a scanned purchase order.
  • Extracted data can trigger downstream workflows (ERP, billing, fulfillment) without manual review.
  • This leads to richer UX, more natural interactions, and closer integration with enterprise systems.
A slide titled "Results" showing five numbered cards with colorful circular icons and short captions. The captions list benefits like richer UX accepting images/documents, better enterprise document workflows, more natural app interactions, broader app capabilities, and better model selection.

Summary

  • Streaming changes delivery, not request structure — send the same messages and image bytes but use a streaming API to receive deltas.
  • For image generation, decode Base64-encoded outputs before saving or serving.
  • Choose the right foundation model and adapt the request/response handling to the model’s expected schema.
  • Store or render artifacts according to your application needs (file system, S3, or direct UI streaming).
This lesson wraps up multimodal streaming and image generation basics. Next logical topics include managing conversation context, stateful chat flows, and advanced grounding techniques for reliable multimodal pipelines.

Watch Video