- Problem: lack of visibility into Bedrock applications
- Solution: use CloudWatch for observability
- Workflow: CloudWatch Metrics and CloudWatch Logs
- Results: clear visibility and faster troubleshooting
- Summary and next steps (including a lab exercise)

- Inputs (prompts, request size)
- Model responses (results, error codes)
- The sequence of events that produced the metric behavior

- CloudWatch Metrics: aggregated, time-series data such as invocation counts, token counts, latencies, and errors.
- CloudWatch Logs: detailed request/response capture (invocation logging) plus application logs you emit from your code.

- Bedrock throttling
- Upstream network timeouts
- Large prompts or unexpected input patterns that increase token usage and cost
Invocation logging captures complete request and response data. If you enable it, be mindful of potentially sensitive information (PII) and the storage/cost implications of logging full payloads.
- When a Bedrock call is initiated
- The model and parameters used for the request
- Response status or error details
- Business-level context to help correlate with metrics
- Open the CloudWatch console and go to Metrics.
- Select the Bedrock namespace.
- Choose metrics and filter by dimensions (model, region, etc.).
- Add them to a chart or custom dashboard and set the desired time range.
Token-based metrics are especially important because Bedrock pricing is often token-based. Monitoring
InputTokenCount and OutputTokenCount helps you detect cost trends and decide on mitigations (prompt engineering, rate limits, provisioned throughput).
Here’s an example from the CloudWatch console showing the Bedrock namespace and several Bedrock-specific metrics. Note that metrics only appear once you’ve generated Bedrock usage in the account.

- Use namespaces and dimensions to scope metrics to the relevant model, region, or environment.
- Build focused dashboards (per application, team, or business unit) that surface only actionable metrics.
- Adjust time ranges for near real-time monitoring or retrospective troubleshooting.
- Create CloudWatch Alarms on critical signals (high latency, increased error rate, token spikes) to trigger notifications or automated remediation.
- Remember that CloudWatch metrics are aggregated before visualization; the charts refresh as new data points arrive rather than continuously streaming.
- Alert fires on a metric (e.g., latency or token spike).
- Open the associated dashboard and check dimensions to narrow affected models/regions.
- Inspect correlated logs (invocation logs + application logs) to find the root cause: throttling, upstream failures, or large request payloads.
- Apply remediation: tweak prompts, increase capacity, add retries/backoff, or redact PII in logs.
- Combine CloudWatch Metrics and CloudWatch Logs to see both the “what” and the “why.”
- Enable invocation logging only when needed; consider privacy and cost implications.
- Instrument your app to emit contextual logs that can be correlated with Bedrock invocations.
- Build dashboards and alarms to detect and respond to issues quickly.