- Common cost drivers for GenAI applications on Bedrock.
- How to combine monitoring (CloudWatch) and cost tools (Cost Explorer, Budgets) to find root causes.
- A step-by-step troubleshooting workflow.
- Practical optimizations and expected impact.

- Large input prompts increase token usage and directly raise costs.
- More powerful (larger) models cost more per request or per token.
- High request volume multiplies per-request charges.
- Lack of instrumentation and visibility makes it difficult to identify where spend is coming from.
- Are unusually large prompts coming from a particular part of the app (e.g., ingestion pipelines, interactive UI, background jobs)?
- Which models are being targeted for those prompts?
- Are large prompts being routed to multi-billion-parameter models unnecessarily?
- Emit metrics from your application that capture model usage: token counts (input/output), model name/version, and invocation counts, plus latency and error rates.
- Visualize these metrics in Amazon CloudWatch dashboards and create alarms for anomalies or thresholds.
- Centralize application logs and model-invocation details in CloudWatch Logs to analyze request patterns and trace sources of large prompts or volume spikes.
- Use AWS Cost Explorer and AWS Budgets to view and alert on spend by service, account, tags, and usage type.
- Metrics: emit
tokens_in,tokens_out,model_invocations, andlatency_mswith relevant dimensions (model_name,service,environment,feature_flag). - Dashboards: create views that show token usage by model, invocation rate, and cost by service over time.
- Alarms: add alerts for sudden token spikes, unexpected model usage, or cost thresholds.

- Identify likely cost drivers (large prompts, model choice, request volume, lack of caching).
- Look for supporting evidence in CloudWatch Metrics (token counts, invocation spikes, latency).
- Analyze CloudWatch Logs to find where requests originate and to inspect payload sizes.
- Consult Cost Explorer and Budgets to map usage to dollars and to detect daily/weekly trends.
- Implement optimizations and re-measure to validate impact.

-
Model selection
- Use smaller, more efficient models for routine tasks (formatting, simple summarization, classification).
- Reserve larger models for tasks that truly require higher reasoning or creativity.
- Route requests dynamically: choose model by intent or task complexity.
-
Prompt optimization
- Make prompts concise to reduce input tokens.
- Provide precise instructions and enforce output schemas (e.g., JSON-only responses) to control verbosity.
- Set a maximum output token limit where appropriate.
-
Call reduction
- Avoid duplicate or unnecessary repeated calls; batch where possible.
- Debounce UI-driven calls to reduce accidental bursts.
-
Caching and alternative architectures
- Cache identical or similar responses at the application level to avoid repeated model invocations.
- Use RAG (retrieval-augmented generation) with a vector DB: store large documents in the vector store and include only relevant chunks via similarity search to keep prompts small.
Bedrock does not deduplicate or cache responses for you. If your application has repeated identical queries, consider a caching layer (Redis/Memcached via Amazon ElastiCache or another store) to store and serve repeat responses within an acceptable TTL.
-
Application caching (Redis / Memcached)
- Workflow: check cache → serve if valid → otherwise call Bedrock → store result in cache.
- Useful for idempotent queries and expensive repeated computations.
-
Retrieval-augmented generation (RAG) with a vector database
- Store documents or long contexts as embeddings.
- For each query, perform a similarity search to fetch the most relevant chunks and send only those chunks as prompt context.
- This reduces prompt size and keeps relevant information available to the model without sending entire documents.
- CloudWatch: technical telemetry (tokens, invocations, latency).
- Cost Explorer: spend visualizations and grouping by service, region, tag, and usage type.
- AWS Budgets: create alerts for specific spend thresholds or forecasted overages.

- Location: AWS Console → Billing and Cost Management → Cost Explorer.
- Capabilities:
- Visualize spend as stacked bar charts, line graphs, or scatter plots.
- Set date ranges and granularity (hourly, daily, monthly).
- Group and filter by service, Region, linked accounts, tags, and usage type.

- Adjust granularity to reveal hourly bursts versus long-term trends.
- Use grouping and tagging to attribute cost to teams, features, or environments.
- Exclude credits or specific accounts if you need to model post-credit spend.
- Leverage forecasting and AWS Cost Anomaly Detection to catch unexpected changes earlier.
- Instrument your app to emit token counts, model invocations, and latency to CloudWatch.
- Use CloudWatch Logs to trace request patterns and identify sources of large prompts or high-volume usage.
- Use AWS Cost Explorer and AWS Budgets to translate those technical signals into dollars and monitor spend over time.
- Apply targeted optimizations—model selection, prompt tuning, caching, RAG—and measure the impact iteratively.