> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Integrating Bedrock With Amazon S3 Part 1

> Using Amazon Bedrock with S3 to automate and scale generative AI processing of files, including code examples, workflow patterns, and production considerations.

In this lesson we cover how to integrate Amazon Bedrock with Amazon S3 to apply generative AI at scale to files stored in the cloud. We'll follow this sequence:

* Define the problem (difficulty processing existing data)
* Describe the solution pattern (S3 as input/output, Bedrock for inference)
* Show example code using the AWS SDK (boto3)
* Walk through a short demo (console view)
* Summarize next steps and expected outcomes

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/4OlDw81IoiRnJTCQ/images/Introduction-to-Amazon-Bedrock/Integrating-Amazon-Bedrock-With-Other-AWS-Services/Integrating-Bedrock-With-Amazon-S3-Part-1/lecture-flow-s3-sdk-ai.jpg?fit=max&auto=format&n=4OlDw81IoiRnJTCQ&q=85&s=6169dfb1fbf69e8c696bdc170f5d3372" alt="A slide titled &#x22;Lecture Flow&#x22; showing a blue flowchart of rounded boxes: Problem → Solution → Workflow → Demonstration → Results → Key Takeaway → What's Next. The boxes include notes about using Amazon S3, accessing S3 with the SDK, and applying AI to data in S3." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Integrating-Amazon-Bedrock-With-Other-AWS-Services/Integrating-Bedrock-With-Amazon-S3-Part-1/lecture-flow-s3-sdk-ai.jpg" />
</Frame>

Problem

Most organizations have large volumes of text data—documents, reports, logs, call transcripts, and customer feedback—scattered across many files. Manually reading and summarizing this content is slow and doesn't scale: what works for 100 documents won't work for 100,000.

We need a reliable, scalable pattern to apply generative AI (summarization, transformation, sentiment analysis, classification, etc.) directly to data stored in S3 so processing can be automated and parallelized.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/4OlDw81IoiRnJTCQ/images/Introduction-to-Amazon-Bedrock/Integrating-Amazon-Bedrock-With-Other-AWS-Services/Integrating-Bedrock-With-Amazon-S3-Part-1/difficult-to-process-stored-text-data.jpg?fit=max&auto=format&n=4OlDw81IoiRnJTCQ&q=85&s=9fabbdf8e08aae1274d4fad105422bbb" alt="A slide titled &#x22;Problem: Difficult to Process Existing Data&#x22; that lists three types of stored text data—documents and reports, logs and transcripts, and customer feedback files—and briefly describes each. A highlighted footer notes the key need for organizations to apply AI to summarize, analyze, or transform this data without manual effort." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Integrating-Amazon-Bedrock-With-Other-AWS-Services/Integrating-Bedrock-With-Amazon-S3-Part-1/difficult-to-process-stored-text-data.jpg" />
</Frame>

Solution: Use Amazon S3 with Bedrock

A common, scalable pattern is:

* Store source files in S3 (use buckets or prefixes to organize by team/use-case).
* Read objects from S3 in your application or serverless function.
* Send the file contents to Amazon Bedrock (invoke the chosen foundation model) for inference.
* Write generated outputs back to S3 (use a separate prefix or bucket for outputs).
* Optionally trigger the pipeline automatically using S3 Event Notifications (to Lambda, SQS, or SNS) for event-driven processing.

This pattern treats S3 as the durable storage layer (input + output), with the application layer responsible for calling Bedrock and handling orchestration.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/4OlDw81IoiRnJTCQ/images/Introduction-to-Amazon-Bedrock/Integrating-Amazon-Bedrock-With-Other-AWS-Services/Integrating-Bedrock-With-Amazon-S3-Part-1/amazon-s3-integration-slide.jpg?fit=max&auto=format&n=4OlDw81IoiRnJTCQ&q=85&s=004df9199772e1c7015464940f58b31e" alt="A presentation slide titled &#x22;Solution: Use Amazon S3 Integration&#x22; with a large S3 bucket icon at the top. Below it are four circular icons and labels: &#x22;Store input documents,&#x22; &#x22;Store output results,&#x22; &#x22;Trigger processing via events,&#x22; and &#x22;Integrate with Retrieval Augmentation.&#x22;" width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Integrating-Amazon-Bedrock-With-Other-AWS-Services/Integrating-Bedrock-With-Amazon-S3-Part-1/amazon-s3-integration-slide.jpg" />
</Frame>

S3 pattern at a glance

| Resource / Prefix | Purpose | Example |
| - | -: | - |
| `s3://my-bucket/input/` | Source documents to process | `s3://my-bucket/input/reviews.txt` |
| `s3://my-bucket/output/` | Generated results (summaries, analyses) | `s3://my-bucket/output/reviews_summary.txt` |
| S3 Event Notification | Trigger processing when new objects are uploaded | Invoke Lambda or send to SQS |
| Bedrock runtime | Model inference endpoint | `invoke_model` via `boto3.client("bedrock-runtime")` |

Code example (single-file, single-object)

This compact Python example uses boto3 to:

1. Read a text object from S3
2. Call the Bedrock runtime (invoke\_model) with a chat-style messages payload
3. Parse the model response (with flexible extraction to handle different model response shapes)
4. Write the generated summary back to S3

Replace bucket names, keys, and the `modelId` with values appropriate for your environment. Ensure your AWS credentials and region are configured (via environment variables, shared credentials file, or an IAM role).

```python theme={null}
import boto3
import json

# Initialize clients (set region as needed)
s3 = boto3.client("s3")
brt = boto3.client("bedrock-runtime", region_name="us-east-1")

# Read input text from S3
text = s3.get_object(Bucket="my-bucket", Key="input.txt")["Body"].read().decode("utf-8")

# Prepare a chat-style payload for Bedrock
payload = {
    "messages": [
        {
            "role": "user",
            "content": [
                {"type": "text", "text": f"Summarize this text in 3 bullet points:\n\n{text}"}
            ]
        }
    ],
    "max_tokens": 200,
    "temperature": 0.5
}

# Call Bedrock runtime (invoke_model)
response = brt.invoke_model(
    modelId="us.meta.llama3-3-70b-instruct-v1:0",
    contentType="application/json",
    accept="application/json",
    body=json.dumps(payload)
)

# Extract text from the response body (handle common shapes)
resp_body = response["body"].read().decode("utf-8")
try:
    resp_json = json.loads(resp_body)
    if "output" in resp_json and "message" in resp_json["output"]:
        summary = resp_json["output"]["message"]["content"][0].get("text")
    elif "results" in resp_json and isinstance(resp_json["results"], list):
        summary = resp_json["results"][0].get("output_text") or resp_json["results"][0].get("content")
    else:
        # Fallback to raw body if the JSON shape is unexpected
        summary = resp_body
except Exception:
    # If parsing fails, use raw body
    summary = resp_body

# Write the summary back to S3 (use a separate output key or prefix)
s3.put_object(Bucket="my-bucket", Key="summary.txt", Body=summary)
```

How this works (step-by-step)

1. s3.get\_object(...) reads the object from the specified S3 bucket/key and returns a streaming Body; the example reads and decodes it into a string.
2. brt.invoke\_model(...) sends a JSON payload to the Bedrock runtime. This includes chat-style `messages` and inference parameters such as `max_tokens` and `temperature`.
3. Models can return responses in different JSON shapes. The example includes multiple checks to extract the generated text; adjust extraction to match your model's response format.
4. s3.put\_object(...) writes the generated summary back to S3. Use a dedicated output prefix (for example, `output/summary.txt`) to keep inputs and outputs organized.

<Callout icon="lightbulb" color="#1CB2FE">
  S3 terminology tip: files in S3 are called "objects" and are identified by object keys. Use logical prefixes such as `input/` and `output/` (or separate buckets) to separate source data from generated results and simplify processing.
</Callout>

Scaling to multiple files

The same pattern scales to many objects. Below is a simple iteration pattern: list objects under an `input/` prefix, process each file, and write results under an `output/` prefix. In production, add proper error handling, pagination, retries, rate limiting, and parallelization.

```python theme={null}
import os
import boto3
import json

s3 = boto3.client("s3")
brt = boto3.client("bedrock-runtime", region_name="us-east-1")

bucket = "my-bucket"
input_prefix = "input/"
output_prefix = "output/"

# List objects under the input prefix
resp = s3.list_objects_v2(Bucket=bucket, Prefix=input_prefix)
for obj in resp.get("Contents", []):
    key = obj["Key"]
    if key.endswith("/"):  # skip folder placeholders
        continue

    # Read each input object
    text = s3.get_object(Bucket=bucket, Key=key)["Body"].read().decode("utf-8")

    # Prepare the payload for Bedrock
    payload = {
        "messages": [
            {
                "role": "user",
                "content": [
                    {"type": "text", "text": f"Summarize this text in 3 bullet points:\n\n{text}"}
                ]
            }
        ],
        "max_tokens": 200,
        "temperature": 0.5
    }

    # Invoke Bedrock
    response = brt.invoke_model(
        modelId="us.meta.llama3-3-70b-instruct-v1:0",
        contentType="application/json",
        accept="application/json",
        body=json.dumps(payload)
    )

    # Extract model output (robust to multiple shapes)
    resp_body = response["body"].read().decode("utf-8")
    try:
        resp_json = json.loads(resp_body)
        if "output" in resp_json and "message" in resp_json["output"]:
            summary = resp_json["output"]["message"]["content"][0].get("text")
        elif "results" in resp_json and isinstance(resp_json["results"], list):
            summary = resp_json["results"][0].get("output_text") or resp_json["results"][0].get("content")
        else:
            summary = resp_body
    except Exception:
        summary = resp_body

    # Preserve filename and write to the output prefix
    filename = os.path.basename(key)
    output_key = f"{output_prefix}{filename.rsplit('.', 1)[0]}_summary.txt"
    s3.put_object(Bucket=bucket, Key=output_key, Body=summary)
```

<Callout icon="warning" color="#FF6B6B">
  Production considerations: monitor and control model usage to avoid unexpected costs, add retries/backoff for transient errors, paginate `list_objects_v2` results, and ensure the IAM role or credentials used have least-privilege access to the relevant S3 buckets and Bedrock APIs.
</Callout>

Demo walkthrough (console)

In the demo bucket shown in the console, there are two prefixes: Input and Output. The Input folder contains a file with ten short customer reviews that we can use as sample input. The processing code reads that file from S3, sends it to Bedrock for summarization (or sentiment analysis), and writes the generated results back to the Output prefix.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/4OlDw81IoiRnJTCQ/images/Introduction-to-Amazon-Bedrock/Integrating-Amazon-Bedrock-With-Other-AWS-Services/Integrating-Bedrock-With-Amazon-S3-Part-1/dark-screen-10-customer-reviews-cursor.jpg?fit=max&auto=format&n=4OlDw81IoiRnJTCQ&q=85&s=5dc19fe365b591c05cfb607cfcdc6925" alt="A dark-themed computer screen showing a text file of ten numbered customer reviews, with a large white mouse cursor visible near the left. The reviews are short praise/complaint lines about a product." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Integrating-Amazon-Bedrock-With-Other-AWS-Services/Integrating-Bedrock-With-Amazon-S3-Part-1/dark-screen-10-customer-reviews-cursor.jpg" />
</Frame>

At scale, replace the simple listing loop with an event-driven approach using S3 Event Notifications:

* New object uploaded -> S3 Event Notification -> Lambda / SQS -> worker consumes message and calls Bedrock
* Or use a batch worker that polls SQS to process many files in parallel

Next steps

Consider expanding this S3 + Bedrock pattern with:

* Retrieval-augmented generation (RAG) and semantic search (index embeddings for fast lookup)
* Preprocessing for binary formats (extract text from PDFs, Word documents, images with OCR)
* Robust orchestration (Step Functions, SQS, or containerized workers) for higher throughput and retry semantics
* Monitoring, billing alerts, and logging for model usage

Links and references

* Amazon Bedrock overview: [https://aws.amazon.com/bedrock/](https://aws.amazon.com/bedrock/)
* Amazon S3 documentation: [https://docs.aws.amazon.com/s3/](https://docs.aws.amazon.com/s3/)
* Boto3 documentation: [https://boto3.amazonaws.com/v1/documentation/api/latest/index.html](https://boto3.amazonaws.com/v1/documentation/api/latest/index.html)
* S3 Event Notifications: [https://docs.aws.amazon.com/AmazonS3/latest/userguide/notification-how-to.html](https://docs.aws.amazon.com/AmazonS3/latest/userguide/notification-how-to.html)
* AWS IAM best practices: [https://docs.aws.amazon.com/IAM/latest/UserGuide/best-practices.html](https://docs.aws.amazon.com/IAM/latest/UserGuide/best-practices.html)

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/introduction-to-amazon-bedrock/module/6a77d10e-3172-4684-96bf-8e168372fae5/lesson/01a0d873-93dd-4a5c-b76f-6d63e6790ed2" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.