Skip to main content
When you evaluate performance and hosting for a foundation model in Amazon Bedrock, you typically choose one of three provisioning approaches:
  • Serverless (On-Demand): the simplest option — call the model via the AWS SDK by name and Bedrock handles compute and autoscaling for you.
  • Serverless + Provisioned Throughput: still serverless operationally, but you reserve throughput for predictable latency and capacity.
  • Bedrock Marketplace (Dedicated Hosting via SageMaker): used when models need dedicated VMs, custom containers, specific GPUs, or special licensing — Bedrock provisions a SageMaker-managed VM for those cases.
Below we walk through each option, when to use it, and the trade-offs to help you pick the right hosting model for production GenAI workloads.

Serverless (On-Demand)

Serverless on-demand is the fastest path to get started. You invoke a named model through Bedrock using the AWS SDK (or APIs) and Bedrock transparently allocates the compute required. Key points:
  • No infrastructure to manage — Bedrock is fully managed.
  • Billing is typically per token consumed (input + output).
  • Automatic scaling makes it ideal for spiky or unpredictable traffic.
  • Multi-tenant hosting may cause variable latency during region-wide demand.
Use serverless on-demand when simplicity and elastic scaling are more important than absolute, predictable latency.

Serverless + Provisioned Throughput

If your workload has predictable, steady traffic and you require consistent latency, provisioned throughput gives you a performance guarantee while retaining serverless operations. In the Bedrock console you specify:
  • Which model to reserve throughput for.
  • The throughput capacity to purchase.
  • A term commitment (for example, 1-year or 3-year).
This model is analogous to buying provisioned IOPS for databases — you purchase reserved capacity that is available regardless of other tenants. Benefits:
  • Predictable response times and reduced throttling risk.
  • Still managed by Bedrock (no instance management).
  • Best for forecastable, steady-state traffic requiring low-latency SLAs.
An infographic titled "Workflow: Serverless + Provisioned Throughput" showing four numbered cards with icons. The cards note: reserve capacity for performance, pay for provisioned capacity, best for steady workloads (consistent latency & throughput), and avoid throttling under load.
If you observe throttling or unpredictable latency under multi-tenant load, moving from pure serverless to provisioned throughput is a common remediation.

Bedrock Marketplace (Dedicated Hosting via SageMaker)

Marketplace-hosted models are still delivered through Bedrock, but Bedrock uses SageMaker hosting behind the scenes to provision a dedicated VM for your model. This gives you an isolated runtime — for example, an ml.p5en.48xlarge — that does not share CPU/GPU or memory with other tenants. Common reasons to choose Marketplace hosting:
  • Custom containers for non-standard runtimes or model server code.
  • Dedicated GPU or VM resources for high-performance inference.
  • Licensing or compliance requirements that mandate isolated compute.
Because a dedicated instance is provisioned for your model, Marketplace is the right choice when isolation, specific instance shapes, or custom packaging are required.
A dark blue slide titled "Workflow: Bedrock Marketplace" with the subheading "Some models require:" and three outlined boxes labeled "Custom containers," "Dedicated GPUs," and "Special licensing constraints." Each box includes a small light-blue icon illustrating its requirement.

Key Differences and Billing Model

Below is a compact comparison to help you decide quickly which option fits your needs.
A slide showing a comparison table titled "Workflow: Serverless vs Provisioned vs Marketplace" that contrasts features (like fully managed by AWS, dedicated instance, instance type, predictable capacity, and pay-per-token) across three deployment options. The table uses three columns labeled Serverless On-Demand, Serverless + Provisioned Throughput, and Marketplace with yes/no entries.
A token is the unit used to measure model input and output size (e.g., tokens consumed by the prompt and by the generated response). In serverless billing, your cost scales with the number of tokens consumed; with provisioned throughput you purchase capacity in advance; with Marketplace you pay for the dedicated managed instance.

When to Use Each Option

  • Serverless (On-Demand)
    • Use when you want simple integration and minimal operational overhead.
    • Best for spiky or unpredictable workloads where per-token billing and autoscaling are acceptable.
  • Serverless + Provisioned Throughput
    • Use when you can forecast traffic and require predictable latency and throughput SLAs.
    • Ideal if you want reservation-based performance but still want Bedrock to manage the environment.
  • Marketplace (Dedicated)
    • Use when the model needs a custom container, dedicated GPUs/VMs, or special licensing and isolation.
A slide titled "Workflow: When to Use Provisioned vs Marketplace?" with two panels comparing use cases. The Provisioned panel lists consistent latency, predictable traffic, and staying serverless; the Marketplace panel lists needing a specific third‑party model, customization, and isolation.

What Bedrock Enables

Amazon Bedrock offers a managed platform to build GenAI applications with options to fit operational and performance needs:
  • Serverless access to multiple foundation models from various providers.
  • Automatic scaling with serverless on-demand or reserved capacity via provisioned throughput for predictable performance.
  • A model catalog with Marketplace-hosted options when customization or isolation is required.
  • Production-ready features such as guardrails, content filtering, semantic search for knowledge integration, and agent tooling to allow models to call APIs or databases.
A dark-blue presentation slide titled "Results: Bedrock Enables Orgs to Build Advanced GenAI Applications Today" with four numbered cards. Each card highlights a feature: serverless access to foundation models, scalable or provisioned throughput, choice of multiple model providers, and building with agents, knowledge bases, and guardrails.
If you take one thing away: Bedrock provides an on-demand, scalable, fully managed architecture that lets you invoke foundation models through AWS APIs and choose the hosting option that matches your performance, isolation, and customization requirements. This concludes the lesson on Bedrock architecture. Supported foundation models and provider-specific guidance are covered in a separate article.

Watch Video