Skip to main content
In this lesson we’ll explore multimodal capabilities of foundation models: how to send and receive more than just text (images, documents, video, and other media). Many modern applications require the model to ingest binary content or to produce media as part of the response—so understanding the expected input formats, request shapes, and response parsing is essential. Common multimodal use cases:
  • Ask the model “What’s in this image?” and get a concise description.
  • Supply a PDF and ask targeted questions about the content.
  • Provide a short video clip for summarization or to extract timestamps and captions.
  • Request image or video generation as part of a creative workflow.
A presentation slide titled "Problem: Modern GenAI Apps don't just take text" that shows four turquoise icons and examples of inputs/outputs. The examples read: "Text + Image for analysis," "PDF or document for question answering," "A video reference for summarization," and "Image or video output for creative generation."
Before integrating a model, verify what modalities it supports. Many models remain text-in/text-out only; others accept images, audio, or documents. Check the model catalog for supported input/output types and any required formats.
Always consult the model catalog to confirm input types (image, document, audio, video) and output types, plus any required metadata or formatting (for example: format, name, or source fields).
A presentation slide titled "Problem: Multimodal Apps need more than a Prompt String" showing a schematic "multimodal request anatomy" icon and a list of challenges (different input formats, limited model modality support, varied output types, and the need for correct packaging/parsing).
Consider how a model returns outputs:
  • Single-response models return the full result in one response.
  • Streaming-capable models return incremental chunks (text or media parts). Your application must handle chunked responses when using streaming endpoints or methods (for example, a ConverseStream or streaming variant of an invoke API).
Not all SDK methods support every workflow. Some models expose streaming APIs for incremental output while others only support non-streaming invoke-style calls. Plan routing, timeouts, and error handling based on the model’s capabilities.
Recommended approach
  • Use a small set of repeatable patterns to package inputs and parse outputs for each modality.
  • Route requests in your app by modality and by the selected model’s capabilities.
  • Use a conversational API (Converse) for multi-turn, contextful dialogs, or a one-shot InvokeModel-style call for single requests.
  • Keep request/response handling modular to enable swapping models without large refactors.
A dark-themed flowchart titled "Workflow: Building a Multimodal Request" showing linked steps from "Pick a multimodal-capable model" to "Build the request body," "Send binary content," then "Parse text/stream/media," "Present or store result," and "Add validation and fallback logic in the app."
Workflow summary (quick reference)

Examples

Example 1 — Image in, Text out (local file) This example sends a local image file alongside a text prompt and uses a Converse call to ask the model to describe the image. The message content mixes text and a binary image object so you can swap models with minimal changes.
Notes:
  • Ensure the image file is accessible to the runtime environment (local filesystem, mounted volume, or container).
  • Mixing text and an image object in a single message simplifies changing models or switching from local files to remote object stores.
Example 1 (variant) — Image in from S3, Text out Images often live in S3 (or another object store). Retrieve the image bytes from S3 and forward the bytes in the same message structure.
A stylized workflow diagram titled "Workflow: Example 1 Variant – Image From S3, Text Out" showing a cloud/browser icon in the center. It depicts an S3 bucket feeding a "Send" box (text question + image content) into a service and returning a "Receive" box with a text answer, with S3 and Bedrock SDK clients shown.
Example 2 — Document in (PDF), Text out Supplying a document (PDF) plus a question lets the model answer based on the document context. The example below sends a PDF’s bytes in the message content.
A stylized workflow diagram titled "Workflow: Example 2 – Document In, Text Out" showing two colored rounded panels labeled "Send" and "Receive" connected through a central speech-bubble icon. The left panel includes a small box labeled "Question / Document content."
Best practices and key takeaways
  • Confirm model capabilities in the model catalog (input/output modalities, required metadata).
  • Package binary content as the model expects (specify format, name, and provide source.bytes).
  • Implement robust parsing for both single-response and streaming outputs; handle partial or chunked data.
  • Add validation and fallback logic (e.g., if the model returns no text, fallback to retry or a simpler model).
  • Keep request/response handling modular to enable model or workflow changes without large refactors.
Additional resources

Watch Video