> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Managing Costs and Optimizing Performance Part 1

> Strategies to reduce costs and latency for GenAI on Amazon Bedrock using model selection, prompt trimming, caching, retrieval, hosting choices, and monitoring

In this lesson you’ll learn practical strategies to control costs and optimize latency for GenAI applications built on Amazon Bedrock. The objective is to scale predictably—keeping token spend and response times manageable—by addressing common cost drivers across model choice, prompt engineering, caching, and hosting.

We’ll cover:

* Typical cost and latency drivers you’ll run into when moving from prototype to production.
* Targeted optimizations you can apply immediately.
* Hosting trade-offs and monitoring approaches to keep spend predictable.

The problem

Many teams prototype GenAI solutions on Bedrock and then find costs and token usage growing unpredictably as traffic scales. Typical causes include:

* High token consumption from large or redundant prompts.
* Always choosing the largest (and most expensive) model for convenience.
* No caching or reuse of previously generated outputs.
* Sending entire documents into the model instead of only relevant context.

These issues often become obvious when a project moves from a small pilot to production workloads where token-volume multiplies.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/BVCvDn4rl3j0TCQq/images/Introduction-to-Amazon-Bedrock/Best-Practices-and-Optimization/Managing-Costs-and-Optimizing-Performance-Part-1/problem-slide-token-usage-large-prompts.jpg?fit=max&auto=format&n=BVCvDn4rl3j0TCQq&q=85&s=b0019e4950c7568e543928724911141e" alt="A dark-themed presentation slide titled &#x22;Problem&#x22; showing five numbered boxes listing issues: 01 High token usage, 02 Large prompts, 03 Inefficient model selection, 04 Poor caching or reuse, and 05 Scaling usage." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Best-Practices-and-Optimization/Managing-Costs-and-Optimizing-Performance-Part-1/problem-slide-token-usage-large-prompts.jpg" />
</Frame>

If left unaddressed, unoptimized GenAI workloads on Bedrock can become prohibitively expensive and jeopardize the ROI of your product.

Solution overview

Control cost and latency at multiple layers:

1. Model and hosting selection — pick the right model and runtime configuration.
2. Prompt engineering — trim and format prompts to reduce tokens.
3. Caching and deduplication — avoid repeated calls for identical requests.
4. Retrieval (RAG) — send only relevant document chunks, not entire documents.
5. Monitoring and usage controls — measure and alert on the right metrics.

Choose the right model and hosting

* Use smaller, cheaper models for straightforward tasks (e.g., email summarization, format conversion). These typically deliver sufficient quality at a fraction of the token and latency cost compared to 70B models.
* Reserve larger models for tasks that truly need advanced reasoning or creativity; they are more expensive per input/output token and usually increase latency.

Trim your prompts

* Every input token counts: shorter, focused prompts reduce both billing and compute time.
* Remove irrelevant context and avoid repeated or duplicated instructions across requests.

Example: a bloated prompt consuming 32 tokens will cost more and increase latency compared to a concise 12-token prompt that contains only the necessary context. Keep prompts concise, explicit, and structured.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/BVCvDn4rl3j0TCQq/images/Introduction-to-Amazon-Bedrock/Best-Practices-and-Optimization/Managing-Costs-and-Optimizing-Performance-Part-1/trim-prompts-32-to-12.jpg?fit=max&auto=format&n=BVCvDn4rl3j0TCQq&q=85&s=ae2bf03b5e7b105ed51c8c84cb939bff" alt="A presentation slide titled &#x22;Solution&#x22; recommending reducing input and output token usage. It compares a bloated prompt of 32 tokens (higher cost & latency) with a trimmed prompt of 12 tokens (lower cost & faster), shown with token block graphics." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Best-Practices-and-Optimization/Managing-Costs-and-Optimizing-Performance-Part-1/trim-prompts-32-to-12.jpg" />
</Frame>

Cache responses and prompt context

* Implement an application-level cache using technologies such as Redis, Memcached, or Amazon ElastiCache to store model outputs and frequently reused prompt context.
* Compute a deterministic cache key from a normalized user query and any relevant parameters (temperature, system prompt version, retrieval results). On cache hit, return the cached output; on miss, call Bedrock and store the result.

Example pseudocode for a simple cache lookup:

```python theme={null}
# Pseudocode
key = sha256(normalize(user_query) + "|" + model_name + "|" + params)
cached = cache.get(key)
if cached:
    return cached
result = call_bedrock(model=model_name, prompt=prompt, params=params)
cache.set(key, result, ttl=3600)
return result
```

<Callout icon="lightbulb" color="#1CB2FE">
  Caching is your responsibility: Bedrock does not automatically cache generated outputs for you. Prompt-level caching tools can prevent resending identical context, but they do not replace application-level caching of generated outputs.
</Callout>

Apply RAG (Retrieval-Augmented Generation) for large-context tasks

* When a task requires knowledge from large documents, index the documents and retrieve only relevant chunks for the query before calling the model.
* RAG lowers token usage and improves answer relevance because you provide only focused context instead of the entire document.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/BVCvDn4rl3j0TCQq/images/Introduction-to-Amazon-Bedrock/Best-Practices-and-Optimization/Managing-Costs-and-Optimizing-Performance-Part-1/apply-rag-avoid-large-prompts.jpg?fit=max&auto=format&n=BVCvDn4rl3j0TCQq&q=85&s=9bde7f76f36659b20eddb918d49f1726" alt="A slide titled &#x22;Solution: Apply RAG to avoid large prompts&#x22; showing a flow diagram: a full document goes into a RAG component, which produces relevant chunks that are then passed to a model (brain icon)." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Best-Practices-and-Optimization/Managing-Costs-and-Optimizing-Performance-Part-1/apply-rag-avoid-large-prompts.jpg" />
</Frame>

Monitor and understand costs

* Use the AWS Billing console to review invoices and drill into cost drivers: [https://console.aws.amazon.com/billing/home](https://console.aws.amazon.com/billing/home)
* Use AWS Cost Explorer for visual analysis and to identify the services, APIs, or models consuming the most budget: [https://console.aws.amazon.com/cost-management/home](https://console.aws.amazon.com/cost-management/home)
* Tag workloads and Bedrock requests where possible so you can attribute spend to features, teams, or customers.

Key metrics to track

| Metric | Why it matters | Example tool |
| - | - | - |
| Input tokens per request | Primary driver of model cost | Bedrock metrics / app logs |
| Output tokens per request | Directly affects spend and latency | Bedrock metrics / app logs |
| Requests per second (RPS) | Capacity planning, throttling, provisioning decisions | CloudWatch |
| Cache hit rate | Higher hit rate reduces model calls & cost | Application metrics / Redis |
| Average latency | User experience & SLA planning | CloudWatch / APM tools |

Be deliberate about Bedrock hosting modes

There are three common Bedrock hosting approaches—each has trade-offs between predictability, cost, and control. Use the table below to choose the best fit for your workload.

| Hosting mode | Billing model | Predictability | Best for |
| - | - | - | - |
| Serverless On-Demand | Pay per token / per request | Lower predictability (multi-tenant) | Variable or low-volume workloads; minimal ops |
| Serverless + Provisioned Throughput | Reservation + per-token | More predictable latency and capacity | Consistent steady-state traffic where reservations are justified |
| Marketplace / Dedicated Instance | Instance-hours (dedicated VM) | High predictability, you control instance size | High-volume, latency-sensitive workloads that require dedicated resources |

* Serverless On-Demand
  * Fully managed; billed per token.
  * Good for bursty or unpredictable traffic where you want no long-term commitments.
  * Response times can be variable due to shared infrastructure.

* Serverless + Provisioned Throughput
  * Reserve throughput for a model for a term (monthly/6/12 mo) to get capacity guarantees and more stable latency.
  * Reservation costs apply even if you under-utilize capacity.

<Callout icon="warning" color="#FF6B6B">
  Provisioned throughput is a reservation: you pay for the capacity you reserve even if you don’t use it. Only choose this option if you can reasonably commit to consistent usage to justify the cost.
</Callout>

* Marketplace / Dedicated Instance
  * AWS provides a dedicated VM instance (pick instance type, e.g., p5en.48xlarge) running a model for your account only.
  * Predictable performance and full control, but you pay for instance hours while it runs.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/BVCvDn4rl3j0TCQq/images/Introduction-to-Amazon-Bedrock/Best-Practices-and-Optimization/Managing-Costs-and-Optimizing-Performance-Part-1/choose-appropriate-hosting-comparison.jpg?fit=max&auto=format&n=BVCvDn4rl3j0TCQq&q=85&s=d3c691b5dda50eb9f2e1b9791c6a0d67" alt="A slide titled &#x22;Workflow: Choose Appropriate Hosting&#x22; showing a comparison table of hosting options (Serverless On‑Demand, Serverless + Provisioned Throughput, Marketplace). The table lists features like &#x22;Fully managed by AWS,&#x22; &#x22;Dedicated instance,&#x22; &#x22;Pick your instance type,&#x22; &#x22;Predictable capacity,&#x22; and &#x22;Pay per token&#x22; with yes/no values." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Best-Practices-and-Optimization/Managing-Costs-and-Optimizing-Performance-Part-1/choose-appropriate-hosting-comparison.jpg" />
</Frame>

Operational tips and quick wins

* Implement a rate limiter and soft quota per user to prevent accidental cost spikes.
* Add telemetry for token counts per request so you can create alerts when average tokens or spend exceed thresholds.
* Use smaller models for pre-processing tasks (classification, extraction) and reserve larger models for final generation steps when necessary.
* Batch requests where possible (e.g., multiple small prompts in one call) to reduce overhead and per-request latency.
* Cache common completions (e.g., templated responses) and purge or version cache keys when system prompts change.

Next steps

* Start with instrumentation: log per-request input/output tokens, model used, and latency.
* Add caching and create an A/B test to measure cost savings and latency improvements.
* Evaluate provisioning vs. serverless based on observed steady-state traffic and latency SLA needs.
* Use RAG to limit the size of prompts for document-heavy features.

References

* Amazon Bedrock: [https://aws.amazon.com/bedrock/](https://aws.amazon.com/bedrock/)
* AWS Billing Console: [https://console.aws.amazon.com/billing/home](https://console.aws.amazon.com/billing/home)
* AWS Cost Explorer: [https://console.aws.amazon.com/cost-management/home](https://console.aws.amazon.com/cost-management/home)
* Redis: [https://redis.io/](https://redis.io/)
* Memcached: [https://memcached.org/](https://memcached.org/)
* Amazon ElastiCache: [https://aws.amazon.com/elasticache/](https://aws.amazon.com/elasticache/)

Summary

* Optimize across model selection, prompt engineering, caching, and retrieval to reduce Bedrock costs and improve latency.
* Monitor key metrics (tokens, RPS, cache hit rate) and choose hosting based on workload predictability and latency requirements.
* Implement quick wins (caching, trimming prompts, RAG) to materially reduce spend while preserving user experience.

Code patterns and full end-to-end implementations are outside the scope of this article, but the guidance here helps you identify where to invest engineering effort to achieve predictable, cost-efficient GenAI production deployments.

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/introduction-to-amazon-bedrock/module/1a696c4d-73f8-4ae4-bcc4-cfbe9c6f03ff/lesson/f4c67de8-7371-4431-8d31-b9ba9707b022" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.