> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Available Models

> Practical guidance for choosing and evaluating foundation models in Amazon Bedrock, addressing modality, context window, cost, latency, testing, provisioning, and routing

In this lesson, we review the foundation models available in Amazon Bedrock and how to choose the right one for your application.

We’ll frame the selection problem, identify the most important model properties, show how to use the model catalog plus empirical testing to narrow choices, and describe the expected results when you pick the right model for the task.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/BVCvDn4rl3j0TCQq/images/Introduction-to-Amazon-Bedrock/Foundation-Models-in-Amazon-Bedrock/Available-Models/lecture-flow-model-selection-diagram.jpg?fit=max&auto=format&n=BVCvDn4rl3j0TCQq&q=85&s=0a0a8f867d6454b3d6db9a7f9e0ec0a8" alt="A slide titled &#x22;Lecture Flow&#x22; showing a linear flowchart of gradient blue boxes: &#x22;Real Problem&#x22; → &#x22;Solution&#x22; → &#x22;Workflow&#x22;, with arrows leading down to &#x22;Results&#x22; and &#x22;Key Takeaway.&#x22; The diagram summarizes evaluating and selecting models for an app on a dark blue background." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Foundation-Models-in-Amazon-Bedrock/Available-Models/lecture-flow-model-selection-diagram.jpg" />
</Frame>

Problem statement

With many models from multiple vendors available in Bedrock, selecting the right model can be difficult. Models differ in capabilities, modalities, latency, context-window size, cost, and failure modes. Choosing poorly can produce inaccurate, late, or irrelevant responses—wasting budget and hurting user satisfaction.

A practical, repeatable evaluation strategy helps teams match models to the concrete requirements of each use case and minimizes unnecessary cost and risk.

Evaluating model properties

Focus on the core properties that matter most for your app:

* Modality — What kinds of input and output does the model support? Text-only, image-only, multimodal (text+image), audio, or image generation? Pick models that natively support the modalities you require.
* Context window — The listed token/length limit describes the combined input plus output the model can handle. If you expect to pass large inputs (for example, multi-thousand-page documents), confirm the model’s context window accommodates the input plus the expected response.
* Task alignment (strengths) — Some models are tuned for multi-turn chat and instruction-following; others excel at summarization, code generation, or image editing. Use models that are optimized for your primary task to reduce hallucinations and improve quality.
* Latency and throughput — For interactive applications, latency matters. For batch workloads, throughput and cost per token matter more.
* Failure modes — Understand how a model tends to fail (e.g., hallucinations, truncated output, or frequent retries) so you can design mitigations and acceptance criteria.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/BVCvDn4rl3j0TCQq/images/Introduction-to-Amazon-Bedrock/Foundation-Models-in-Amazon-Bedrock/Available-Models/evaluate-model-properties-slide.jpg?fit=max&auto=format&n=BVCvDn4rl3j0TCQq&q=85&s=411dd5036f3617e6f01953f0d02f93f0" alt="A presentation slide titled &#x22;Solution: Evaluate Model Properties&#x22; showing three sections: Modality (Text, Image, Multi-modal), Context Window (&#x22;How much input can be processed&#x22;), and Model Strengths (Chat, Summarization, Coding). The design uses a dark blue background with rounded turquoise gradient buttons." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Foundation-Models-in-Amazon-Bedrock/Available-Models/evaluate-model-properties-slide.jpg" />
</Frame>

Quick reference table: core property checklist

| Property | Why it matters | How to check |
| - | -: | - |
| Modality | Ensures model accepts the input/output types you need | Model catalog: look for `text`, `image`, `multimodal`, or `image-generation` labels |
| Context window | Prevents truncated inputs or outputs | Confirm token/byte limits in the catalog and test with realistic payloads |
| Model strengths | Reduces hallucinations and improves quality | Review vendor docs and published evaluations; run task-specific tests |
| Latency / throughput | Affects user experience and cost at scale | Measure p99 latency under expected concurrency; evaluate cost per token |
| Cost | Controls production economics | Check per-token or per-image pricing and simulate expected traffic |

Model size, cost, and latency

“Model size” (number of parameters) often correlates with reasoning ability, but bigger is not always better:

* Larger models typically provide stronger reasoning and generalization, but cost more per token and have higher latency.
* Smaller models are faster and cheaper; they can be ideal for simple summarization, classification, or routine Q\&A.
* For input-heavy tasks, prioritize context-window capacity over parameter count.

When designing for scale:

* Estimate expected throughput (calls per second) and concurrency (simultaneous users). Cost per token compounds quickly at scale.
* Use smaller models where they suffice; reserve larger models for tasks requiring deep reasoning.
* Consider provisioned throughput if you can predict steady, high demand (see next section).

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/tJmUiudNjsCWp_bm/images/Introduction-to-Amazon-Bedrock/Foundation-Models-in-Amazon-Bedrock/Available-Models/workflow-model-strengths-cost-modality-tokens.jpg?fit=max&auto=format&n=tJmUiudNjsCWp_bm&q=85&s=8bd59874581d027b94b18954e1d5cfe6" alt="A presentation slide titled &#x22;Workflow: Model Strengths&#x22; showing two blue circular icons labeled &#x22;Cost&#x22; (with a receipt icon) and &#x22;Modality&#x22; (with a document/image icon). Under each are brief bullet points about pricing (charged per token, larger models cost more) and supported modalities (text, image, multi-modal)." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Foundation-Models-in-Amazon-Bedrock/Available-Models/workflow-model-strengths-cost-modality-tokens.jpg" />
</Frame>

Use-case alignment

Match the model to the job. Typical mappings:

* Simple Q\&A and document summarization: text-only, small-to-midsize models (cost-effective and sufficient for many tasks).
* Chat assistants with multi-turn reasoning: midsize-to-large conversational models (some Anthropic and Meta models are examples).
* Image generation or editing: specialized image-generation models (for example, Stable Diffusion by Stability AI).
* Code generation: models tuned for coding or those with strong reasoning to reduce incorrect outputs.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/tJmUiudNjsCWp_bm/images/Introduction-to-Amazon-Bedrock/Foundation-Models-in-Amazon-Bedrock/Available-Models/workflow-use-case-alignment.jpg?fit=max&auto=format&n=tJmUiudNjsCWp_bm&q=85&s=6427f56a0ac471c3a6c032ee283a6135" alt="A slide titled &#x22;Workflow: Use-Case Alignment&#x22; showing four colored panels for different AI use cases: Simple Q&A or Summarization, Chat Assistant With Reasoning, Image Generation, and Code Generation. Each panel includes an icon and a short suggestion about which model types to consider." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Foundation-Models-in-Amazon-Bedrock/Available-Models/workflow-use-case-alignment.jpg" />
</Frame>

A practical model-selection strategy

Follow a staged approach to reduce risk and cost:

1. Start with a smaller model that matches your modality and context-window needs.
2. Test quality against your acceptance criteria with real inputs.
3. If quality is insufficient, evaluate midsize or large models and re-run tests.
4. Once you meet quality goals, optimize prompts and token usage to reduce cost.
5. Monitor performance and cost in production and iterate.

<Callout icon="lightbulb" color="#1CB2FE">
  Start small and iterate: this reduces experimentation cost and prevents over-provisioning. Use automated tests and representative datasets to compare models objectively.
</Callout>

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/tJmUiudNjsCWp_bm/images/Introduction-to-Amazon-Bedrock/Foundation-Models-in-Amazon-Bedrock/Available-Models/model-selection-strategy-steps.jpg?fit=max&auto=format&n=tJmUiudNjsCWp_bm&q=85&s=d9e6a243e859fa0bd807d4e836db926f" alt="An infographic titled &#x22;Workflow: Model Selection Strategy&#x22; with five colorful steps. The steps read: 01 Begin with a smaller model, 02 Test quality, 03 Increase model size if necessary, 04 Monitor token usage, and 05 Optimize for cost vs performance." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Foundation-Models-in-Amazon-Bedrock/Available-Models/model-selection-strategy-steps.jpg" />
</Frame>

Provisioning and billing modes

Bedrock typically offers serverless on-demand hosting (multi-tenant) where you pay per token (input + output). Image-generation and larger multimodal models generally cost more per token or per image.

If you can forecast steady throughput, provisioned throughput reserves capacity and can provide discounted pricing. Accurately forecast concurrent load (e.g., 100, 1,000, or 10,000 concurrent users) because costs increase substantially when moving from development/test to production.

Don’t bind your application to a single model

Design your application to route requests to different models based on capability and cost:

* Use a small, inexpensive model for routine or low-risk queries.
* Route image generation to a specialized image model (for example, Stable Diffusion).
* Use a stronger reasoning model for tasks that require deeper comprehension or higher trust.

Keep revisiting modality, model size, and context window when you define routing rules.

<Callout icon="warning" color="#FF6B6B">
  Important: Mixing models improves cost and quality trade-offs, but also increases operational complexity. Track which model handled each request for debugging, billing allocation, and auditing.
</Callout>

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/tJmUiudNjsCWp_bm/images/Introduction-to-Amazon-Bedrock/Foundation-Models-in-Amazon-Bedrock/Available-Models/summary-6-model-selection-tips.jpg?fit=max&auto=format&n=tJmUiudNjsCWp_bm&q=85&s=e9700700e2faf14befe5012f5daf3047" alt="A dark-blue presentation slide titled &#x22;Summary&#x22; listing six numbered tips about selecting and using models — covering use case, latency, cost (Bedrock), routing prompts, model size, modalities, and context window. Each item is marked with a teal arrow-shaped number icon." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Foundation-Models-in-Amazon-Bedrock/Available-Models/summary-6-model-selection-tips.jpg" />
</Frame>

Expected results from good model selection

When you align models to use cases, you should see:

* Lower costs and fewer surprises in production billing.
* Better latency and a stronger user experience for interactive apps.
* Reduced experimentation time since candidate models are narrowed early.
* More efficient use of generative AI resources by avoiding one-size-fits-all models.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/tJmUiudNjsCWp_bm/images/Introduction-to-Amazon-Bedrock/Foundation-Models-in-Amazon-Bedrock/Available-Models/results-slide-four-benefits-generative-ai.jpg?fit=max&auto=format&n=tJmUiudNjsCWp_bm&q=85&s=e7b8341e145a6f3374c727a14ddb8a3c" alt="A presentation slide titled &#x22;Results&#x22; with four numbered panels highlighting benefits: better model selection for specific use cases, improved application performance and quality, reduced experimentation time, and more efficient use of generative AI resources. Each panel has a small icon and a colored top accent." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Foundation-Models-in-Amazon-Bedrock/Available-Models/results-slide-four-benefits-generative-ai.jpg" />
</Frame>

Key takeaway

Use the right model for the right task every time. Avoid sending all requests to a single model without evaluating modality fit, context-window capacity, cost, and latency.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/BVCvDn4rl3j0TCQq/images/Introduction-to-Amazon-Bedrock/Foundation-Models-in-Amazon-Bedrock/Available-Models/key-takeaway-01-use-right-model.jpg?fit=max&auto=format&n=BVCvDn4rl3j0TCQq&q=85&s=0c34aff4cfa2112f217dd5bf9cbc8a83" alt="A minimalist presentation slide with a dark left panel labeled &#x22;Key Takeaway&#x22; and a turquoise &#x22;01&#x22; badge next to the text: &#x22;Use the right model for the right task, every time.&#x22;" width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Foundation-Models-in-Amazon-Bedrock/Available-Models/key-takeaway-01-use-right-model.jpg" />
</Frame>

Next steps

Try a hands-on lab to experiment with different foundation models and observe trade-offs firsthand. For vendor documentation and pricing details, see:

* Amazon Bedrock: [https://aws.amazon.com/bedrock/](https://aws.amazon.com/bedrock/)
* Stability AI (Stable Diffusion): [https://stability.ai](https://stability.ai)

Additional resources

* [Choosing the right model in production — practical checklist and tests](https://aws.amazon.com/bedrock/) (start with model catalog and vendor docs)
* Run A/B comparisons and synthetic benchmarking for latency and quality
* Track per-request model usage and cost to inform routing rules and provisioning decisions

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/introduction-to-amazon-bedrock/module/39fdfbba-c310-47cf-8de2-083d4c4164d8/lesson/5e3556e3-fd8f-4981-9893-89fe52848f6c" />

  <Card title="Practice Lab" icon="flask-conical" cta="Learn more" href="https://learn.kodekloud.com/user/courses/introduction-to-amazon-bedrock/module/39fdfbba-c310-47cf-8de2-083d4c4164d8/lesson/d6f96ad4-aee3-4640-ac19-08c548ee9fa4" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.