- Serverless (On-Demand): the simplest option — call the model via the AWS SDK by name and Bedrock handles compute and autoscaling for you.
- Serverless + Provisioned Throughput: still serverless operationally, but you reserve throughput for predictable latency and capacity.
- Bedrock Marketplace (Dedicated Hosting via SageMaker): used when models need dedicated VMs, custom containers, specific GPUs, or special licensing — Bedrock provisions a SageMaker-managed VM for those cases.
Serverless (On-Demand)
Serverless on-demand is the fastest path to get started. You invoke a named model through Bedrock using the AWS SDK (or APIs) and Bedrock transparently allocates the compute required. Key points:- No infrastructure to manage — Bedrock is fully managed.
- Billing is typically per token consumed (input + output).
- Automatic scaling makes it ideal for spiky or unpredictable traffic.
- Multi-tenant hosting may cause variable latency during region-wide demand.
Serverless + Provisioned Throughput
If your workload has predictable, steady traffic and you require consistent latency, provisioned throughput gives you a performance guarantee while retaining serverless operations. In the Bedrock console you specify:- Which model to reserve throughput for.
- The throughput capacity to purchase.
- A term commitment (for example, 1-year or 3-year).
- Predictable response times and reduced throttling risk.
- Still managed by Bedrock (no instance management).
- Best for forecastable, steady-state traffic requiring low-latency SLAs.

Bedrock Marketplace (Dedicated Hosting via SageMaker)
Marketplace-hosted models are still delivered through Bedrock, but Bedrock uses SageMaker hosting behind the scenes to provision a dedicated VM for your model. This gives you an isolated runtime — for example, anml.p5en.48xlarge — that does not share CPU/GPU or memory with other tenants.
Common reasons to choose Marketplace hosting:
- Custom containers for non-standard runtimes or model server code.
- Dedicated GPU or VM resources for high-performance inference.
- Licensing or compliance requirements that mandate isolated compute.

Key Differences and Billing Model
Below is a compact comparison to help you decide quickly which option fits your needs.
A token is the unit used to measure model input and output size (e.g., tokens consumed by the prompt and by the generated response). In serverless billing, your cost scales with the number of tokens consumed; with provisioned throughput you purchase capacity in advance; with Marketplace you pay for the dedicated managed instance.
When to Use Each Option
- Serverless (On-Demand)
- Use when you want simple integration and minimal operational overhead.
- Best for spiky or unpredictable workloads where per-token billing and autoscaling are acceptable.
- Serverless + Provisioned Throughput
- Use when you can forecast traffic and require predictable latency and throughput SLAs.
- Ideal if you want reservation-based performance but still want Bedrock to manage the environment.
- Marketplace (Dedicated)
- Use when the model needs a custom container, dedicated GPUs/VMs, or special licensing and isolation.

What Bedrock Enables
Amazon Bedrock offers a managed platform to build GenAI applications with options to fit operational and performance needs:- Serverless access to multiple foundation models from various providers.
- Automatic scaling with serverless on-demand or reserved capacity via provisioned throughput for predictable performance.
- A model catalog with Marketplace-hosted options when customization or isolation is required.
- Production-ready features such as guardrails, content filtering, semantic search for knowledge integration, and agent tooling to allow models to call APIs or databases.

Links and references
- Amazon Bedrock Overview
- Amazon SageMaker Hosting
- AWS Pricing
- Kubernetes Basics (for general compute orchestration concepts)