> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Managing Costs and Optimizing Performance Part 4

> Guidance on reducing costs and boosting performance for Amazon Bedrock knowledge bases through incremental syncs, limited retrieval, caching, batching, precomputed embeddings, and appropriate model choice

When managing knowledge bases in Amazon Bedrock, plan for re-ingestion from the start. If you initially ingest 1,000 documents, those documents will change over time: new documents arrive, and existing documents are updated. Re-ingesting the entire corpus on every change is wasteful — instead, apply incremental updates only to changed items.

Amazon Bedrock provides a synchronize (sync) operation for knowledge bases that performs incremental updates: it applies only changes rather than re-ingesting the whole corpus. Sync does not run automatically — you must decide how and when to trigger it (for example, event-driven when new data arrives, or on a schedule such as every six hours using Amazon EventBridge).

<Callout icon="lightbulb" color="#1CB2FE">
  Use incremental sync to update only changed documents. Trigger sync either by events (when data arrives) or on a schedule ([Amazon EventBridge](https://docs.aws.amazon.com/eventbridge/latest/userguide/what-is-amazon-eventbridge.html)) depending on your ingestion pattern and SLA.
</Callout>

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/BVCvDn4rl3j0TCQq/images/Introduction-to-Amazon-Bedrock/Best-Practices-and-Optimization/Managing-Costs-and-Optimizing-Performance-Part-4/workflow-manage-kb-costs-reingest-changed.jpg?fit=max&auto=format&n=BVCvDn4rl3j0TCQq&q=85&s=8d149c110b239018a6fd07c38f797e6d" alt="A slide titled &#x22;Workflow: Manage Knowledge Base Costs&#x22; showing three document icons with only the center one highlighted as &#x22;Changed&#x22; while the other two are labeled &#x22;Unchanged&#x22; to illustrate avoiding re‑ingesting unchanged data. A callout says &#x22;Re‑ingest only the changed file; skip the other two&#x22; and suggests using versioning or timestamps to detect updates." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Best-Practices-and-Optimization/Managing-Costs-and-Optimizing-Performance-Part-4/workflow-manage-kb-costs-reingest-changed.jpg" />
</Frame>

## Detecting changed documents

Detect changes using one or more of the following methods and feed the changed-document list to Bedrock's sync operation:

* Versioning: increment a version number when a file changes.
* Timestamps: compare last-modified timestamps.
* Checksums/hashes: compute a checksum and compare to the stored value.
* Event-driven notifications: publish an event when content changes and trigger a sync for that document.

If real-time eventing is not available or needed, schedule periodic syncs using Amazon EventBridge (hourly, every few hours, daily) to keep the knowledge base reasonably fresh without re-ingesting everything.

<Callout icon="warning" color="#FF6B6B">
  Avoid re-ingesting the entire knowledge base on every change. That increases token usage, processing cost, and embedding recomputation. Incremental syncs are the cost-efficient approach.
</Callout>

## Example: Limit retrieved chunks in a RetrieveAndGenerate call

When using retrieval-augmented generation (RAG) with Bedrock, limit the number of retrieved chunks to reduce context injected into the model. Fewer chunks means fewer input tokens and lower cost.

In the example below, the vector search configuration returns only the top three chunks by specifying `numberOfResults: 3`.

```python theme={null}
response = bedrock_agent_runtime.retrieve_and_generate(
    input={"text": "What is the returns policy?"},
    retrieveAndGenerateConfiguration={
        "knowledgeBaseConfiguration": {
            "knowledgeBaseId": "kb-123",
            "modelArn": "amazon.nova-lite-v1:0",
            "retrievalConfiguration": {
                "vectorSearchConfiguration": {
                    "numberOfResults": 3  # limit retrieved chunks
                }
            }
        }
    }
)
```

By restricting retrieved chunks you control the size of the augmented prompt (fewer input tokens). This is one of the most effective levers for cost and performance control in RAG workflows.

## Operational recommendations (quick reference)

| Recommendation | Why it matters | Example / How-to |
| - | -: | - |
| Start with smaller models | Lower cost and often sufficient for many tasks | Use `amazon.nova-lite-*` or similar; move to larger models only if metrics justify it |
| Track token usage | Understand cost drivers and tune parameters | Log request/response sizes and `max_tokens` |
| Use incremental sync | Avoid unnecessary embeddings and reprocessing | Detect changes via version/timestamp/checksum and call Bedrock sync only for changed items |
| Cache repeated prompts | Save tokens and latency for identical queries | Serve cached responses for repeat queries |
| Batch and pre-compute | Reduce call volume and reprocessing | Pre-compute embeddings; batch similar requests |
| Monitor billing | Detect cost anomalies and set alerts | Use [AWS Billing Console](https://console.aws.amazon.com/billing/home) and [AWS Cost Explorer](https://aws.amazon.com/aws-cost-management/aws-cost-explorer/) |

## Simple load-test example

A minimal timing loop helps estimate baseline response times and detect performance regressions. Replace the print statements with real model calls or HTTP requests to Bedrock for an actual load test.

```python theme={null}
import time

queries = ["What is the returns policy?"] * 10  # simulate 10 identical queries

start = time.time()

for q in queries:
    # simulate model call; replace with actual Bedrock call in real tests
    print(f"Processing: {q}")

end = time.time()

print(f"Total time: {end - start:.2f} seconds")
```

From these simple measurements you can derive practical actions: refine prompts to reduce token usage, add caching layers to avoid repeated calls, batch requests where appropriate, and provision throughput (or scale horizontally) when necessary.

## Expected outcomes

Adopting these strategies will typically yield:

* Lower operational costs by reducing unnecessary token usage and choosing appropriate models.
* Faster responses because less irrelevant data is sent to the model.
* Better scalability — moving from a handful of users to hundreds or thousands without runaway costs.
* More predictable usage and billing.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/BVCvDn4rl3j0TCQq/images/Introduction-to-Amazon-Bedrock/Best-Practices-and-Optimization/Managing-Costs-and-Optimizing-Performance-Part-4/genai-cost-tips-key-takeaways-slide.jpg?fit=max&auto=format&n=BVCvDn4rl3j0TCQq&q=85&s=ee921b4044985f1b5975d691613bca16" alt="A presentation slide titled &#x22;Key Takeaways&#x22; that lists numbered recommendations for cost-efficient GenAI—choose the right model, use RAG, reuse embeddings, and use caching/batching. The slide has a dark left panel with blue numbered markers and the text on a light background." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Best-Practices-and-Optimization/Managing-Costs-and-Optimizing-Performance-Part-4/genai-cost-tips-key-takeaways-slide.jpg" />
</Frame>

## Key practical takeaways

1. Choose the right model for the task; don’t default to the largest model for everything.
2. Use RAG (retrieval-augmented generation) and Bedrock knowledge-base syncs so only relevant chunks are injected into prompts.
3. Reuse pre-computed data (embeddings) rather than recomputing them repeatedly.
4. Use caching and batching to avoid unnecessary Bedrock calls and lower overall token usage.
5. Prefer serving from cache or pre-computation for repeated identical prompts; use batching where real-time responses are not required.

Later sections will cover error handling, troubleshooting, edge cases, and defensive coding practices for reliable production deployments.

## Links and references

* Amazon EventBridge: [https://docs.aws.amazon.com/eventbridge/latest/userguide/what-is-amazon-eventbridge.html](https://docs.aws.amazon.com/eventbridge/latest/userguide/what-is-amazon-eventbridge.html)
* AWS Billing Console: [https://console.aws.amazon.com/billing/home](https://console.aws.amazon.com/billing/home)
* AWS Cost Explorer: [https://aws.amazon.com/aws-cost-management/aws-cost-explorer/](https://aws.amazon.com/aws-cost-management/aws-cost-explorer/)

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/introduction-to-amazon-bedrock/module/1a696c4d-73f8-4ae4-bcc4-cfbe9c6f03ff/lesson/315babf3-2fdb-472a-a083-7d2d6eb00c5c" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.