AWS Budgets: track and alert on spend
AWS Budgets lets you create departmental or cost-center budgets and measure consumption against them. Most organizations implement budgets by tagging resources—virtual machines, containers, Lambda functions, and other services—with cost-allocation tags. Tagging enables you to correlate resource usage to the right budget and generate department-level reports.Budgets are for tracking and alerting only — they do not enforce or cap spend. For example, a “VM budget” of 10,000. Budgets can predict overruns and send alerts, but they don’t block consumption.
Best practice: create budgets per product, team, or cost center and make cost-allocation tags mandatory in CI/CD or provisioning pipelines. Combine budgets with alerts to drive early investigation before month-end surprises.
Monitor Bedrock usage with CloudWatch
Amazon CloudWatch collects metrics for many AWS services, including Bedrock. When you know the cost per token for a given model, CloudWatch metrics such asInputTokenCount and OutputTokenCount allow you to estimate spend at granular intervals (5 minutes, 10 minutes, or hourly). Combine token counts with model pricing to compute estimated spend for each interval and to identify spikes in average token usage.

Diagnostic scenarios and step-by-step guidance
Use the following scenarios to triage unusual or unexpected Bedrock costs. The images below illustrate the investigative workflows. Scenario 1 — Costs doubled over the last week with no traffic increase, no model change, and no deployment changes:- Inspect CloudWatch metrics to check token counts per request.
- If per-request token counts increased, query CloudWatch Logs for model invocation details and token counts.
- Common root cause: prompts have grown over time, increasing token usage per request.
- Fixes: reduce prompt size, supply context via Bedrock knowledge bases (document chunking), or tighten output limits to lower generated tokens and preserve budget for necessary inputs.

- Use CloudWatch metrics and Cost Explorer to drill into dimensions such as
modelId,region, andusageType. - You may discover the application is using a more expensive model than necessary for routine tasks.
- Fix: match the model to the task — adopt smaller or more cost-effective models for routine workloads and reserve larger models for tasks that require higher-quality outputs or deeper reasoning.

Quick reference — investigation tools
Tip: instrument your application to log model name, prompt length (tokens), output length (tokens), and requestId. This makes Correlation between metrics and logs straightforward.
Expected outcomes from cost-management practices
By implementing these monitoring and control practices you should achieve:- Predictable spend as usage grows, reducing surprise end-of-month bills.
- Engineering teams making conscious model choices (start with smaller models and scale up when necessary).
- Lower token consumption by trimming prompts and constraining outputs (
maxTokensand more focused prompts). - Clear visibility into linear costs as user counts increase, enabling better capacity and budget planning.

Next steps and references
- Instrument your application logs to include token counts and model identifiers.
- Create budgets with cost-allocation tags and set alerts to drive early investigation.
- Use CloudWatch dashboards and scheduled reports to monitor trends and runbooks for cost incidents.
- AWS Budgets: https://docs.aws.amazon.com/cost-management/latest/userguide/budgets-managing-costs.html
- Amazon CloudWatch: https://docs.aws.amazon.com/cloudwatch/
- CloudWatch Logs Insights: https://docs.aws.amazon.com/AmazonCloudWatch/latest/logs/AnalyzingLogData.html
- AWS Cost Explorer: https://docs.aws.amazon.com/cost-management/latest/userguide/what-is-cost-explorer.html
- Amazon Bedrock: https://aws.amazon.com/bedrock/