Skip to main content
In this lesson you’ll learn practical strategies to control costs and optimize latency for GenAI applications built on Amazon Bedrock. The objective is to scale predictably—keeping token spend and response times manageable—by addressing common cost drivers across model choice, prompt engineering, caching, and hosting. We’ll cover:
  • Typical cost and latency drivers you’ll run into when moving from prototype to production.
  • Targeted optimizations you can apply immediately.
  • Hosting trade-offs and monitoring approaches to keep spend predictable.
The problem Many teams prototype GenAI solutions on Bedrock and then find costs and token usage growing unpredictably as traffic scales. Typical causes include:
  • High token consumption from large or redundant prompts.
  • Always choosing the largest (and most expensive) model for convenience.
  • No caching or reuse of previously generated outputs.
  • Sending entire documents into the model instead of only relevant context.
These issues often become obvious when a project moves from a small pilot to production workloads where token-volume multiplies.
A dark-themed presentation slide titled "Problem" showing five numbered boxes listing issues: 01 High token usage, 02 Large prompts, 03 Inefficient model selection, 04 Poor caching or reuse, and 05 Scaling usage.
If left unaddressed, unoptimized GenAI workloads on Bedrock can become prohibitively expensive and jeopardize the ROI of your product. Solution overview Control cost and latency at multiple layers:
  1. Model and hosting selection — pick the right model and runtime configuration.
  2. Prompt engineering — trim and format prompts to reduce tokens.
  3. Caching and deduplication — avoid repeated calls for identical requests.
  4. Retrieval (RAG) — send only relevant document chunks, not entire documents.
  5. Monitoring and usage controls — measure and alert on the right metrics.
Choose the right model and hosting
  • Use smaller, cheaper models for straightforward tasks (e.g., email summarization, format conversion). These typically deliver sufficient quality at a fraction of the token and latency cost compared to 70B models.
  • Reserve larger models for tasks that truly need advanced reasoning or creativity; they are more expensive per input/output token and usually increase latency.
Trim your prompts
  • Every input token counts: shorter, focused prompts reduce both billing and compute time.
  • Remove irrelevant context and avoid repeated or duplicated instructions across requests.
Example: a bloated prompt consuming 32 tokens will cost more and increase latency compared to a concise 12-token prompt that contains only the necessary context. Keep prompts concise, explicit, and structured.
A presentation slide titled "Solution" recommending reducing input and output token usage. It compares a bloated prompt of 32 tokens (higher cost & latency) with a trimmed prompt of 12 tokens (lower cost & faster), shown with token block graphics.
Cache responses and prompt context
  • Implement an application-level cache using technologies such as Redis, Memcached, or Amazon ElastiCache to store model outputs and frequently reused prompt context.
  • Compute a deterministic cache key from a normalized user query and any relevant parameters (temperature, system prompt version, retrieval results). On cache hit, return the cached output; on miss, call Bedrock and store the result.
Example pseudocode for a simple cache lookup:
Caching is your responsibility: Bedrock does not automatically cache generated outputs for you. Prompt-level caching tools can prevent resending identical context, but they do not replace application-level caching of generated outputs.
Apply RAG (Retrieval-Augmented Generation) for large-context tasks
  • When a task requires knowledge from large documents, index the documents and retrieve only relevant chunks for the query before calling the model.
  • RAG lowers token usage and improves answer relevance because you provide only focused context instead of the entire document.
A slide titled "Solution: Apply RAG to avoid large prompts" showing a flow diagram: a full document goes into a RAG component, which produces relevant chunks that are then passed to a model (brain icon).
Monitor and understand costs Key metrics to track Be deliberate about Bedrock hosting modes There are three common Bedrock hosting approaches—each has trade-offs between predictability, cost, and control. Use the table below to choose the best fit for your workload.
  • Serverless On-Demand
    • Fully managed; billed per token.
    • Good for bursty or unpredictable traffic where you want no long-term commitments.
    • Response times can be variable due to shared infrastructure.
  • Serverless + Provisioned Throughput
    • Reserve throughput for a model for a term (monthly/6/12 mo) to get capacity guarantees and more stable latency.
    • Reservation costs apply even if you under-utilize capacity.
Provisioned throughput is a reservation: you pay for the capacity you reserve even if you don’t use it. Only choose this option if you can reasonably commit to consistent usage to justify the cost.
  • Marketplace / Dedicated Instance
    • AWS provides a dedicated VM instance (pick instance type, e.g., p5en.48xlarge) running a model for your account only.
    • Predictable performance and full control, but you pay for instance hours while it runs.
A slide titled "Workflow: Choose Appropriate Hosting" showing a comparison table of hosting options (Serverless On‑Demand, Serverless + Provisioned Throughput, Marketplace). The table lists features like "Fully managed by AWS," "Dedicated instance," "Pick your instance type," "Predictable capacity," and "Pay per token" with yes/no values.
Operational tips and quick wins
  • Implement a rate limiter and soft quota per user to prevent accidental cost spikes.
  • Add telemetry for token counts per request so you can create alerts when average tokens or spend exceed thresholds.
  • Use smaller models for pre-processing tasks (classification, extraction) and reserve larger models for final generation steps when necessary.
  • Batch requests where possible (e.g., multiple small prompts in one call) to reduce overhead and per-request latency.
  • Cache common completions (e.g., templated responses) and purge or version cache keys when system prompts change.
Next steps
  • Start with instrumentation: log per-request input/output tokens, model used, and latency.
  • Add caching and create an A/B test to measure cost savings and latency improvements.
  • Evaluate provisioning vs. serverless based on observed steady-state traffic and latency SLA needs.
  • Use RAG to limit the size of prompts for document-heavy features.
References Summary
  • Optimize across model selection, prompt engineering, caching, and retrieval to reduce Bedrock costs and improve latency.
  • Monitor key metrics (tokens, RPS, cache hit rate) and choose hosting based on workload predictability and latency requirements.
  • Implement quick wins (caching, trimming prompts, RAG) to materially reduce spend while preserving user experience.
Code patterns and full end-to-end implementations are outside the scope of this article, but the guidance here helps you identify where to invest engineering effort to achieve predictable, cost-efficient GenAI production deployments.

Watch Video