> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Multimodal Capabilities Text and Images Part 1

> Overview of multimodal foundation model usage, covering supported modalities, request and response formats, streaming, code examples, and integration best practices

In this lesson we'll explore multimodal capabilities of foundation models: how to send and receive more than just text (images, documents, video, and other media). Many modern applications require the model to ingest binary content or to produce media as part of the response—so understanding the expected input formats, request shapes, and response parsing is essential.

Common multimodal use cases:

* Ask the model “What’s in this image?” and get a concise description.
* Supply a PDF and ask targeted questions about the content.
* Provide a short video clip for summarization or to extract timestamps and captions.
* Request image or video generation as part of a creative workflow.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/BVCvDn4rl3j0TCQq/images/Introduction-to-Amazon-Bedrock/Advanced-Topics-Optional/Multimodal-Capabilities-Text-and-Images-Part-1/genai-multimodal-inputs-outputs-slide.jpg?fit=max&auto=format&n=BVCvDn4rl3j0TCQq&q=85&s=0bd1f84bbff7f045a424f3f8447a62b4" alt="A presentation slide titled &#x22;Problem: Modern GenAI Apps don't just take text&#x22; that shows four turquoise icons and examples of inputs/outputs. The examples read: &#x22;Text + Image for analysis,&#x22; &#x22;PDF or document for question answering,&#x22; &#x22;A video reference for summarization,&#x22; and &#x22;Image or video output for creative generation.&#x22;" width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Advanced-Topics-Optional/Multimodal-Capabilities-Text-and-Images-Part-1/genai-multimodal-inputs-outputs-slide.jpg" />
</Frame>

Before integrating a model, verify what modalities it supports. Many models remain text-in/text-out only; others accept images, audio, or documents. Check the model catalog for supported input/output types and any required formats.

<Callout icon="lightbulb" color="#1CB2FE">
  Always consult the model catalog to confirm input types (image, document, audio, video) and output types, plus any required metadata or formatting (for example: `format`, `name`, or `source` fields).
</Callout>

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/BVCvDn4rl3j0TCQq/images/Introduction-to-Amazon-Bedrock/Advanced-Topics-Optional/Multimodal-Capabilities-Text-and-Images-Part-1/multimodal-request-anatomy-challenges.jpg?fit=max&auto=format&n=BVCvDn4rl3j0TCQq&q=85&s=3ec335782f1ee7400a76da9e97b3772b" alt="A presentation slide titled &#x22;Problem: Multimodal Apps need more than a Prompt String&#x22; showing a schematic &#x22;multimodal request anatomy&#x22; icon and a list of challenges (different input formats, limited model modality support, varied output types, and the need for correct packaging/parsing)." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Advanced-Topics-Optional/Multimodal-Capabilities-Text-and-Images-Part-1/multimodal-request-anatomy-challenges.jpg" />
</Frame>

Consider how a model returns outputs:

* Single-response models return the full result in one response.
* Streaming-capable models return incremental chunks (text or media parts). Your application must handle chunked responses when using streaming endpoints or methods (for example, a `ConverseStream` or streaming variant of an invoke API).

<Callout icon="warning" color="#FF6B6B">
  Not all SDK methods support every workflow. Some models expose streaming APIs for incremental output while others only support non-streaming invoke-style calls. Plan routing, timeouts, and error handling based on the model's capabilities.
</Callout>

Recommended approach

* Use a small set of repeatable patterns to package inputs and parse outputs for each modality.
* Route requests in your app by modality and by the selected model’s capabilities.
* Use a conversational API (Converse) for multi-turn, contextful dialogs, or a one-shot InvokeModel-style call for single requests.
* Keep request/response handling modular to enable swapping models without large refactors.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/BVCvDn4rl3j0TCQq/images/Introduction-to-Amazon-Bedrock/Advanced-Topics-Optional/Multimodal-Capabilities-Text-and-Images-Part-1/multimodal-workflow-build-send-parse-validate.jpg?fit=max&auto=format&n=BVCvDn4rl3j0TCQq&q=85&s=f3ed729f4da73446f09690c8f4e95701" alt="A dark-themed flowchart titled &#x22;Workflow: Building a Multimodal Request&#x22; showing linked steps from &#x22;Pick a multimodal-capable model&#x22; to &#x22;Build the request body,&#x22; &#x22;Send binary content,&#x22; then &#x22;Parse text/stream/media,&#x22; &#x22;Present or store result,&#x22; and &#x22;Add validation and fallback logic in the app.&#x22;" width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Advanced-Topics-Optional/Multimodal-Capabilities-Text-and-Images-Part-1/multimodal-workflow-build-send-parse-validate.jpg" />
</Frame>

Workflow summary (quick reference)

| Step | Action |
| - | - |
| 1 | Pick a multimodal-capable model (consult the model catalog). |
| 2 | Build the request body and include any binary content (images, PDFs, video) in the supported format (`format`, `name`, `source.bytes`). |
| 3 | Send the request to the runtime (use the streaming endpoint if the model supports streaming). |
| 4 | Parse the response (single-response body, streaming text chunks, or binary media bytes). |
| 5 | Present or store the result; add validation and fallback logic to handle unexpected outputs. |

## Examples

Example 1 — Image in, Text out (local file)
This example sends a local image file alongside a text prompt and uses a Converse call to ask the model to describe the image. The message content mixes text and a binary image object so you can swap models with minimal changes.

```python theme={null}
import boto3

client = boto3.client("bedrock-runtime", region_name="us-east-1")

with open("invoice.png", "rb") as f:
    image_bytes = f.read()

response = client.converse(
    modelId="amazon.nova-lite-v1:0",
    messages=[
        {
            "role": "user",
            "content": [
                {"text": "What is this image showing? Summarize it in 3 bullet points."},
                {
                    "image": {
                        "format": "png",
                        "source": {"bytes": image_bytes}
                    }
                }
            ]
        }
    ]
)

# Safely extract text result (check response shape in case of model differences)
answer = response.get("output", {}).get("message", {}).get("content", [])
if answer:
    text_parts = [c.get("text") for c in answer if "text" in c]
    print("\n".join([p for p in text_parts if p]))
else:
    print("No text output returned; inspect the full response:", response)
```

Notes:

* Ensure the image file is accessible to the runtime environment (local filesystem, mounted volume, or container).
* Mixing text and an image object in a single message simplifies changing models or switching from local files to remote object stores.

Example 1 (variant) — Image in from S3, Text out
Images often live in S3 (or another object store). Retrieve the image bytes from S3 and forward the bytes in the same message structure.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/BVCvDn4rl3j0TCQq/images/Introduction-to-Amazon-Bedrock/Advanced-Topics-Optional/Multimodal-Capabilities-Text-and-Images-Part-1/workflow-image-s3-text-out.jpg?fit=max&auto=format&n=BVCvDn4rl3j0TCQq&q=85&s=817af7250a42eb14026ed98e2c622d18" alt="A stylized workflow diagram titled &#x22;Workflow: Example 1 Variant – Image From S3, Text Out&#x22; showing a cloud/browser icon in the center. It depicts an S3 bucket feeding a &#x22;Send&#x22; box (text question + image content) into a service and returning a &#x22;Receive&#x22; box with a text answer, with S3 and Bedrock SDK clients shown." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Advanced-Topics-Optional/Multimodal-Capabilities-Text-and-Images-Part-1/workflow-image-s3-text-out.jpg" />
</Frame>

```python theme={null}
import boto3

# Clients
s3 = boto3.client("s3")
bedrock = boto3.client("bedrock-runtime", region_name="us-east-1")

# S3 location
bucket = "my-image-bucket"
key = "images/invoice.png"

# Read image from S3
s3_response = s3.get_object(Bucket=bucket, Key=key)
image_bytes = s3_response["Body"].read()

# Send to Bedrock Runtime via Converse
response = bedrock.converse(
    modelId="amazon.nova-lite-v1:0",
    messages=[
        {
            "role": "user",
            "content": [
                {"text": "What is this image showing? Summarize in 3 bullet points."},
                {
                    "image": {
                        "format": "png",
                        "source": {"bytes": image_bytes}
                    }
                }
            ]
        }
    ]
)

answer = response.get("output", {}).get("message", {}).get("content", [])
if answer:
    print(answer[0].get("text"))
else:
    print("No response text found; inspect full response:", response)
```

Example 2 — Document in (PDF), Text out
Supplying a document (PDF) plus a question lets the model answer based on the document context. The example below sends a PDF's bytes in the message content.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/BVCvDn4rl3j0TCQq/images/Introduction-to-Amazon-Bedrock/Advanced-Topics-Optional/Multimodal-Capabilities-Text-and-Images-Part-1/workflow-example2-document-text-send-receive.jpg?fit=max&auto=format&n=BVCvDn4rl3j0TCQq&q=85&s=7a2d69916aa74544bef07dff68b6febc" alt="A stylized workflow diagram titled &#x22;Workflow: Example 2 – Document In, Text Out&#x22; showing two colored rounded panels labeled &#x22;Send&#x22; and &#x22;Receive&#x22; connected through a central speech-bubble icon. The left panel includes a small box labeled &#x22;Question / Document content.&#x22;" width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Advanced-Topics-Optional/Multimodal-Capabilities-Text-and-Images-Part-1/workflow-example2-document-text-send-receive.jpg" />
</Frame>

```python theme={null}
import boto3

client = boto3.client("bedrock-runtime", region_name="us-east-1")

with open("policy.pdf", "rb") as f:
    pdf_bytes = f.read()

response = client.converse(
    modelId="amazon.nova-lite-v1:0",
    messages=[
        {
            "role": "user",
            "content": [
                {"text": "Summarize the return policy in plain English."},
                {
                    "document": {
                        "format": "pdf",
                        "name": "policy",
                        "source": {"bytes": pdf_bytes}
                    }
                }
            ]
        }
    ]
)

answer = response.get("output", {}).get("message", {}).get("content", [])
if answer:
    # Models may return multiple content items; extract text parts.
    text_parts = [c.get("text") for c in answer if "text" in c]
    print("\n".join([p for p in text_parts if p]))
else:
    print("No text output returned; inspect the full response:", response)
```

Best practices and key takeaways

* Confirm model capabilities in the model catalog (input/output modalities, required metadata).
* Package binary content as the model expects (specify `format`, `name`, and provide `source.bytes`).
* Implement robust parsing for both single-response and streaming outputs; handle partial or chunked data.
* Add validation and fallback logic (e.g., if the model returns no text, fallback to retry or a simpler model).
* Keep request/response handling modular to enable model or workflow changes without large refactors.

Additional resources

* [Amazon Bedrock Runtime Documentation](https://docs.aws.amazon.com/bedrock/latest/userguide/what-is-bedrock.html)
* [Amazon S3 Developer Guide](https://docs.aws.amazon.com/AmazonS3/latest/userguide/Welcome.html)
* Model catalogs and SDK docs (check your platform/provider for the latest model capability listings and API details).

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/introduction-to-amazon-bedrock/module/7af9f623-7d4e-447a-8b21-6e635dfaccfa/lesson/fa81c4de-3188-402b-97f9-343ccea5166a" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.