> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Inference Profiles

> Explains Amazon Bedrock inference profiles, why some model IDs need prefixes like us., how to add them in code, and routing and capacity tradeoffs.

In this article we explain inference profiles in Amazon Bedrock: what they are, why some model IDs include prefixes like `us.`, how to use them in code, and the trade-offs to consider. You'll learn how inference profiles affect routing and capacity so you can reliably invoke Bedrock models from your application.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/tJmUiudNjsCWp_bm/images/Introduction-to-Amazon-Bedrock/Getting-Started-With-Amazon-Bedrock/Inference-Profiles/lecture-flow-reliable-model-invocation.jpg?fit=max&auto=format&n=tJmUiudNjsCWp_bm&q=85&s=c4dd9026144e94d07a1d1927542720f1" alt="A slide titled &#x22;Lecture Flow&#x22; showing a horizontal flowchart of six blue rounded boxes (Problem → Solution → Workflow → Demonstration → Results → Key Takeaway) connected by arrows. Each box has short notes about model IDs, inference profiles, using prefixes, and reliable model invocation." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Getting-Started-With-Amazon-Bedrock/Inference-Profiles/lecture-flow-reliable-model-invocation.jpg" />
</Frame>

## Problem

When selecting a model from the Bedrock model catalog, many developers copy the catalog's programmatic model ID into their CLI or application code. Sometimes that exact identifier—copied directly from the catalog—fails when you try to invoke the model. This is confusing because the catalog ID seems correct, yet some invocations succeed and others fail.

The mismatch is frequently caused by inference profiles: Bedrock's runtime routing configuration that may require a system-defined prefix on the model ID.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/tJmUiudNjsCWp_bm/images/Introduction-to-Amazon-Bedrock/Getting-Started-With-Amazon-Bedrock/Inference-Profiles/model-id-mismatch-catalog-confusion.jpg?fit=max&auto=format&n=tJmUiudNjsCWp_bm&q=85&s=993e645fc929cdfefcc403e49a21dbee" alt="A presentation slide titled &#x22;Problem: Some Model IDs Don't Work Even With the Right Model ID.&#x22; It shows four numbered cards listing issues: developers invoking models, incorrect model IDs causing failures, model catalog format differences, and teams needing to know which identifier to use." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Getting-Started-With-Amazon-Bedrock/Inference-Profiles/model-id-mismatch-catalog-confusion.jpg" />
</Frame>

## What are inference profiles?

Inference profiles are platform-provided deployment/routing configurations for certain Bedrock foundation models. An inference profile lets Bedrock route your request across a specified geography (for example, within U.S. regions) rather than binding it to a single regional endpoint. This increases the chance Bedrock finds available capacity for your request.

Key characteristics:

* Platform-level routing mechanism provided by AWS/Bedrock.
* Do not change the model architecture or user-facing APIs.
* Often system-defined (for example, `us`); users generally cannot create these profiles.
* Adding a profile prefix to the model ID (for example `us.`) tells Bedrock to route the request within that profile scope.

Table: quick reference

| Topic | Summary | Example |
| - | - | - |
| Purpose | Route requests across a geographic profile to find capacity | `us.meta-llama-3-1-8b-instruct-v1` |
| Scope | System-defined routing (not user-created) | `us.` indicates U.S. inference profile |
| Effect on model | No change to model behavior or architecture | Same prompts and parameters apply |

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/tJmUiudNjsCWp_bm/images/Introduction-to-Amazon-Bedrock/Getting-Started-With-Amazon-Bedrock/Inference-Profiles/inference-profile-deployment-config-slide.jpg?fit=max&auto=format&n=tJmUiudNjsCWp_bm&q=85&s=2833437d67791a208f93199e566574d6" alt="A presentation slide titled &#x22;Solution: Use an Inference Profile&#x22; that defines an inference profile as &#x22;a deployment configuration for a foundation model.&#x22; Below it lists what it does: routes requests within a region, manages capacity, and controls model access." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Getting-Started-With-Amazon-Bedrock/Inference-Profiles/inference-profile-deployment-config-slide.jpg" />
</Frame>

## How it looks in code

You typically only change the `modelId` string to include the inference profile prefix; your application logic remains the same. Below is a concise Python example using the Bedrock runtime client to illustrate the difference. The key change is adding the `us.` prefix to the `modelId` to select the system-defined U.S. inference profile.

```python theme={null}
# python
import boto3
import json

# Create a Bedrock runtime client (region can remain your default)
client = boto3.client("bedrock-runtime")

# Model ID without inference profile (may fail for some models)
model_id_plain = "meta-llama-3-1-8b-instruct-v1"

# Model ID with system inference profile (routes within US regions)
model_id_with_profile = "us.meta-llama-3-1-8b-instruct-v1"

# Simple payload: ask for 2 short bullet points
payload = {
    "input": "Explain Bedrock in two short bullet points.",
    # Runtime-specific parameters such as max tokens may be placed in the request body
    # or passed using provider-specific fields. This is a simplified example.
    "parameters": {"max_tokens": 200}
}

response = client.invoke_model(
    modelId=model_id_with_profile,
    contentType="application/json",
    accept="application/json",
    body=json.dumps(payload).encode("utf-8")
)

# The response body is returned as bytes; decode and print
body = response["body"].read().decode("utf-8")
print(body)
```

Notes on the code:

* This example demonstrates how the `modelId` is formed. SDKs and API versions may differ in exact parameter names and locations for runtime options.
* For production use, consult the AWS Bedrock developer guide and the specific SDK docs for your language:
  * AWS Bedrock Developer Guide: [https://docs.aws.amazon.com/bedrock/latest/devguide/](https://docs.aws.amazon.com/bedrock/latest/devguide/)

<Callout icon="lightbulb" color="#1CB2FE">
  Some Bedrock catalog entries require a system-defined inference profile prefix (for example, `us.`). If an invocation fails using the catalog ID, retry with the relevant inference profile prefix before changing other parts of your code.
</Callout>

## Why use inference profiles?

Benefits:

* Reliability: Fewer invocation failures because Bedrock can route to any region inside the profile’s scope that has capacity.
* Simplicity: Your application does not need to implement region-level fallback logic; Bedrock handles routing.
* No model modification: The model’s behavior is unchanged—only where Bedrock looks for capacity is expanded.

Caveats:

* Not every model uses or requires inference profiles. They are platform-provided and sometimes mandatory for specific foundation models.
* Broader routing may increase latency if Bedrock routes the request to a geographically distant region.
* The model catalog may not consistently indicate when a profile prefix is required; trial-and-error (or documentation) may be needed.

Table: Benefits vs. Caveats

| Benefit | Caveat |
| - | - |
| Reduces invocation failures due to capacity constraints | Possible increased latency if routed outside your primary region |
| Simplifies client logic by delegating routing to Bedrock | Not all models support inference profiles |
| No changes to model usage or prompts | Catalog might not explicitly state which models require prefixes |

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/tJmUiudNjsCWp_bm/images/Introduction-to-Amazon-Bedrock/Getting-Started-With-Amazon-Bedrock/Inference-Profiles/bedrock-inference-profile-us-prefix.jpg?fit=max&auto=format&n=tJmUiudNjsCWp_bm&q=85&s=95d9dc46bdaf54a0038b24b9ffaad27c" alt="A presentation slide titled &#x22;Key Takeaway&#x22; stating that some Bedrock models require a system-defined inference profile appearing as a prefix like &#x22;us.&#x22; in the model ID. The slide features a dark blue left panel, a turquoise &#x22;01&#x22; badge, and the main text on a light gray background." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Getting-Started-With-Amazon-Bedrock/Inference-Profiles/bedrock-inference-profile-us-prefix.jpg" />
</Frame>

## Summary

Some Bedrock models require a system-defined inference profile prefix (for example `us.`). Adding that prefix to the model ID instructs Bedrock to route the request within the profile’s geographic scope so the runtime can locate available capacity. This reduces invocation errors and makes model access more consistent across regions, at the possible cost of higher latency if routing outside your primary region.

Further reading and references:

* AWS Bedrock Developer Guide: [https://docs.aws.amazon.com/bedrock/latest/devguide/](https://docs.aws.amazon.com/bedrock/latest/devguide/)
* Bedrock API and SDK documentation (refer to the SDK for your language)

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/introduction-to-amazon-bedrock/module/4f0b1655-3751-4724-a6eb-78d06f3753a7/lesson/6720dd96-2f21-43a7-bc74-5d0475948234" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.