> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Managing Costs and Optimizing Performance Part 3

> Guidance on reducing cost and latency by choosing appropriate model sizes, constraining outputs, optimizing prompts and RAG chunking for efficient inference and knowledge base costs.

Choosing the right model and constraining its outputs are two of the most effective levers for controlling cost and latency. Use smaller models for classification, summarization, and other high-volume, low-complexity requests; reserve larger models for tasks that require deep reasoning, multi-step workflows, or external tool calls (for example, when the model must call APIs, query databases, or orchestrate multiple steps).

## Model selection: match size to task

* Small models: fast, low-cost, ideal for classification, sentiment analysis, short summaries, and deterministic transformations.
* Large models: higher latency and cost, best for complex reasoning, multi-step tasks, and situations requiring broad context or advanced synthesis.
* When unsure, benchmark both representative smaller and larger models on your workload and measure accuracy, latency, and cost.

## Control output size and format

Explicitly constraining output length and format reduces token usage and prevents verbose, costly responses. When invoking a model, set a hard token limit and provide a strict output template.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/BVCvDn4rl3j0TCQq/images/Introduction-to-Amazon-Bedrock/Best-Practices-and-Optimization/Managing-Costs-and-Optimizing-Performance-Part-3/workflow-control-output-size-tips.jpg?fit=max&auto=format&n=BVCvDn4rl3j0TCQq&q=85&s=21778ede2d9dfc7af80a31640d791e53" alt="A presentation slide titled &#x22;Workflow: Control Output Size&#x22; showing three numbered tips: 01 cap with maxTokens, 02 use clear instructions, and 03 use constrained formats (e.g., JSON, bullets). Each tip is displayed in a colored boxed panel with brief explanations." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Best-Practices-and-Optimization/Managing-Costs-and-Optimizing-Performance-Part-3/workflow-control-output-size-tips.jpg" />
</Frame>

Example: limit tokens and ask for a concise format using the Bedrock runtime (boto3):

```python theme={null}
import boto3
import json

bedrock = boto3.client("bedrock-runtime", region_name="us-east-1")

prompt = "Summarize the benefits of cloud computing in 3 bullet points."

response = bedrock.invoke_model(
    modelId="amazon.nova-lite-v1:0",
    body=json.dumps({
        "inputText": prompt,
        "textGenerationConfig": {
            "maxTokenCount": 100,    # limit output size
            "temperature": 0.5
        }
    }),
    contentType="application/json",
    accept="application/json"
)

result = json.loads(response["body"].read())
print(result["results"][0]["outputText"])
```

In the example above, `maxTokenCount` is set to 100. Coupled with an instruction that the output must be exactly three bullet points, this prevents rambling and reduces inference cost.

## Prompt design: reduce input token cost

The input context also contributes to cost. Avoid repeatedly sending large or irrelevant documents. In retrieval-augmented generation (RAG), only pass the most relevant chunks into the prompt to keep both input and output token usage low.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/BVCvDn4rl3j0TCQq/images/Introduction-to-Amazon-Bedrock/Best-Practices-and-Optimization/Managing-Costs-and-Optimizing-Performance-Part-3/workflow-prompt-design-costly-mistakes.jpg?fit=max&auto=format&n=BVCvDn4rl3j0TCQq&q=85&s=afbdcca83ff30558253b313948bf4cd1" alt="A slide titled &#x22;Workflow: Prompt Design Impacts Cost&#x22; listing three costly mistakes: repeating large context unnecessarily, injecting irrelevant documents in RAG, and sending entire documents instead of chunks. Each item appears in an orange-outlined box with an X icon on a dark blue background." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Best-Practices-and-Optimization/Managing-Costs-and-Optimizing-Performance-Part-3/workflow-prompt-design-costly-mistakes.jpg" />
</Frame>

Naive approach — embedding an entire document in the prompt (costly):

```python theme={null}
full_document = """
Returns are accepted within 30 days...
Shipping takes 5-7 days...
Warranty lasts 1 year...
[...very long document...]
"""

prompt = f"""
Use the following document to answer the question:

{full_document}

Question: What is the returns policy?
"""
```

Cost-effective approach — retrieve only relevant chunks and pass a concise context:

```python theme={null}
full_document = """
Returns are accepted within 30 days...
Shipping takes 5-7 days...
Warranty lasts 1 year...
[...very long document...]
"""

# Assume retrieval returns only the relevant passages:
retrieved_chunks = [
    "Items can be returned within 30 days of delivery.",
    "Refunds are issued to the original payment method."
]

context = "\n".join(retrieved_chunks)

prompt = f"""
Use the following context to answer the question:

{context}

Question: What is the returns policy?
"""
```

By retrieving and sending only the relevant paragraphs you reduce input token size and typically improve grounding and accuracy. For more on RAG concepts, see [Retrieval-augmented generation](https://en.wikipedia.org/wiki/Retrieval-augmented_generation) and the [Amazon Bedrock documentation](https://aws.amazon.com/bedrock/).

<Callout icon="lightbulb" color="#1CB2FE">
  When you require a strict structured output (for example JSON), include the exact schema or a template in the prompt. This reduces malformed outputs and prevents extraneous text that consumes tokens. For example, supply a JSON schema and instruct: “Return only valid JSON matching this schema.”
</Callout>

## Knowledge base (RAG) billing and chunk-size trade-offs

When you use managed RAG services or Bedrock Knowledge Bases, costs arrive from multiple places. Monitor and optimize across the whole pipeline:

| Billing dimension | What it covers | Optimization tips |
| - | - | - |
| Ingestion | Embedding calls, chunking, parsing | Choose embedding model and chunk size wisely; batch embeddings to reduce overhead |
| Retrieval | Vector store query and storage costs | Select the appropriate vector store for your scale and query patterns |
| Inference | Model compute for the augmented prompt | Minimize augmented prompt size and pick the right foundation model for the task |

Chunk size directly affects both ingestion and runtime costs:

* Larger chunks → fewer embeddings (lower ingestion cost) but coarser retrieval. Retrieved chunks may include irrelevant information, increasing augmented prompt size and runtime cost.
* Smaller chunks → more embeddings (higher ingestion cost) but more precise retrieval. Precise retrieval injects only tightly relevant content into the prompt and can lower runtime cost.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/BVCvDn4rl3j0TCQq/images/Introduction-to-Amazon-Bedrock/Best-Practices-and-Optimization/Managing-Costs-and-Optimizing-Performance-Part-3/knowledge-base-costs-chunk-size-comparison.jpg?fit=max&auto=format&n=BVCvDn4rl3j0TCQq&q=85&s=126409b6bf1fe7c9392da4a1abff3837" alt="A slide titled &#x22;Workflow: Manage Knowledge Base Costs&#x22; comparing larger chunks (fewer embeddings, coarse retrieval, higher runtime cost) with smaller chunks (more embeddings, precise retrieval, lower runtime cost)." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Best-Practices-and-Optimization/Managing-Costs-and-Optimizing-Performance-Part-3/knowledge-base-costs-chunk-size-comparison.jpg" />
</Frame>

<Callout icon="warning" color="#FF6B6B">
  Evaluate trade-offs carefully: choose chunk sizes and embedding models that balance ingestion vs. runtime costs for your workload. Continuously monitor costs across ingestion, retrieval, and inference and adjust parameters based on observed accuracy and spend.
</Callout>

## Quick reference

| Topic | Recommendation |
| - | - |
| Use smaller models when possible | For high-volume, low-complexity tasks (classification, short summaries) |
| Constrain outputs | Use `maxTokenCount`, templates, and strict format instructions |
| Prompt design | Retrieve and send only relevant context; avoid duplicating large documents |
| RAG pipeline | Monitor ingestion, retrieval, and inference costs and tune chunk sizes and embedding models |

Summary: favor smaller models for simple tasks, reserve large models for complex reasoning, constrain outputs (token limits and strict formats), design precise prompts, and use RAG intelligently—retrieve only necessary chunks to reduce token usage and cost.

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/introduction-to-amazon-bedrock/module/1a696c4d-73f8-4ae4-bcc4-cfbe9c6f03ff/lesson/6d04f12c-d448-4dd2-a87d-eb62f3b0aebc" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.