Skip to main content
When choosing hosting for an application backed by Amazon Bedrock, think carefully about how costs scale as traffic grows. During early development or low-traffic phases, a serverless on-demand model (billed per request or per token depending on the provider/pricing) can be extremely affordable — often just a few dollars per month. However, once your application receives significant concurrent traffic (for example, hundreds of concurrent users), on-demand token billing can rapidly increase costs because you pay for tokens consumed by every request. If you expect steady, predictable load, plan ahead: provisioned throughput or a Marketplace (dedicated instance) can reduce costs and provide predictable performance. If traffic is unpredictable with spikes, on-demand remains attractive. If traffic is steady and growing, on-demand costs will scale roughly linearly with usage — consider provisioned throughput or Marketplace/dedicated instances to control cost and capacity.
Provisioned throughput trades variable cost for a fixed monthly (or periodic) fee. You pay for the reserved capacity whether you use it or not, so this is best when demand is predictable and sustained.
A dark-themed slide titled "Workflow: Choose Appropriate Hosting" showing a comparison table of hosting options (Serverless On-Demand, Serverless + Provisioned Throughput, Marketplace). Rows list features like "Fully managed by AWS," "Dedicated instance," "Pick your instance type," "Predictable/Guaranteed capacity," and "Pay per token" with answers for each option.
Marketplace (dedicated instances) provides the most control over instance-level resources — CPU, memory, disk, network throughput — and lets you choose the specific model and instance sizing. Marketplace/dedicated-instance billing is typically hourly for the instance(s) you run, not per-token. Example scenarios (illustrating trade-offs)
  • On-demand billing example: If each request costs £0.002, serving 10,000 requests/day costs £20/day. If traffic falls to 1,000 requests/day, cost drops to £2/day. On-demand keeps cost proportional to usage.
  • Provisioned throughput example: If you reserve capacity for 10,000 requests/day at a fixed cost of £15/day, the effective cost per request is lower at full utilization. But if you only use 2,000 requests/day, you still pay £15/day, raising the effective cost per actual request.
Which option to choose:
  • Use on-demand when traffic is low or unpredictable.
  • Use provisioned throughput when demand is high and consistent.
  • Use Marketplace/dedicated instances when you need instance-level control and predictable performance tied to instance sizing.
An infographic comparing "On-Demand (Pay per Use)" versus "Provisioned Throughput (Reserved)" pricing, showing example calculations for 10,000 and reduced request volumes and resulting cost-per-request. It highlights daily costs, how costs change when traffic drops, and notes about cost consistency vs savings at high usage.
What drives cost in Amazon Bedrock? Your application controls these costs by design:
  • Reduce prompt and context size where possible.
  • Select an appropriate model for each task (see model routing below).
  • Add caching to avoid repeated calls for the same query or response.
Caching options to reduce repeated token usage:
  • Redis — an in-memory cache for very low latency lookups.
  • DynamoDB — a durable key-value store for longer-lived results and larger datasets.
A dark-themed flowchart titled "Workflow: What Drives Amazon Bedrock Cost?" showing an app sending small and large prompts through a cache to AWS-managed Bedrock, which routes requests to different models. The diagram labels the cache (Redis/DynamoDB) and shows Model A with a single dollar sign and Model B with three dollar signs to indicate cost differences.
Model selection matters Choose a model that matches the task in capability, latency, and cost. Lightweight models are faster and cheaper but less capable; large models are slower and more expensive but better for complex reasoning. Example routing strategy:
  • Route high-quality, complex tasks to larger models even when the prompt is small (for better reasoning or accuracy).
  • Route high-volume, simple tasks (classification, short summaries) to smaller, cheaper models to reduce costs.
Model examples:
  • Amazon Nova Micro — small, fast, and low-cost; suitable for text summarization, simple completions, or high-volume lightweight tasks.
  • Meta Llama 3.1 70B Instruct — large, slower, costlier; suitable for complex reasoning and high-quality generation.
A side-by-side comparison chart titled "Workflow: Model Selection Matters" contrasting a lightweight model and a large capable model. The left column lists Amazon Nova Micro as small, fast, and low-cost for simple tasks; the right column lists Meta Llama 3.1 70B Instruct as large, slower, and more expensive.
Best practices summary
  • Start with on-demand during development or when traffic is unpredictable.
  • Add caching (Redis, DynamoDB) to reduce duplicate model calls and token use.
  • Route different prompts to different models based on cost vs. quality requirements.
  • Move to provisioned throughput or Marketplace instances when you need predictable performance and cost for sustained traffic.
  • Monitor both input and output token usage and track model-specific costs — these are your primary cost drivers.
Design your application to control prompt size, model selection, and caching. These levers give you the most direct control over Amazon Bedrock costs while preserving required performance and response quality.
Links and references

Watch Video