Provisioned throughput trades variable cost for a fixed monthly (or periodic) fee. You pay for the reserved capacity whether you use it or not, so this is best when demand is predictable and sustained.

- On-demand billing example: If each request costs £0.002, serving 10,000 requests/day costs £20/day. If traffic falls to 1,000 requests/day, cost drops to £2/day. On-demand keeps cost proportional to usage.
- Provisioned throughput example: If you reserve capacity for 10,000 requests/day at a fixed cost of £15/day, the effective cost per request is lower at full utilization. But if you only use 2,000 requests/day, you still pay £15/day, raising the effective cost per actual request.
- Use on-demand when traffic is low or unpredictable.
- Use provisioned throughput when demand is high and consistent.
- Use Marketplace/dedicated instances when you need instance-level control and predictable performance tied to instance sizing.

Your application controls these costs by design:
- Reduce prompt and context size where possible.
- Select an appropriate model for each task (see model routing below).
- Add caching to avoid repeated calls for the same query or response.
- Redis — an in-memory cache for very low latency lookups.
- DynamoDB — a durable key-value store for longer-lived results and larger datasets.

- Route high-quality, complex tasks to larger models even when the prompt is small (for better reasoning or accuracy).
- Route high-volume, simple tasks (classification, short summaries) to smaller, cheaper models to reduce costs.
- Amazon Nova Micro — small, fast, and low-cost; suitable for text summarization, simple completions, or high-volume lightweight tasks.
- Meta Llama 3.1 70B Instruct — large, slower, costlier; suitable for complex reasoning and high-quality generation.

- Start with on-demand during development or when traffic is unpredictable.
- Add caching (Redis, DynamoDB) to reduce duplicate model calls and token use.
- Route different prompts to different models based on cost vs. quality requirements.
- Move to provisioned throughput or Marketplace instances when you need predictable performance and cost for sustained traffic.
- Monitor both input and output token usage and track model-specific costs — these are your primary cost drivers.
Design your application to control prompt size, model selection, and caching. These levers give you the most direct control over Amazon Bedrock costs while preserving required performance and response quality.
- Amazon Bedrock overview: https://learn.kodekloud.com/user/courses/introduction-to-amazon-bedrock
- Redis (in-memory caching): https://redis.io/
- Amazon DynamoDB (durable key-value store): https://aws.amazon.com/dynamodb/
- AWS pricing and instance guidance: https://aws.amazon.com/pricing/