> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Controlling Model Parameters

> Explains how to use temperature, topP, and maxTokens to control model output randomness, diversity, and length.

In this lesson you'll learn how to control model parameters to influence how a model generates responses. These controls help you steer output length, creativity, and determinism without changing the model’s pre-trained knowledge or reasoning capabilities.

How can you steer a model’s responses? In practice, model outputs may be too long, too random, or too inconsistent for your application. Some use cases require highly repeatable, deterministic outputs (e.g., code generation, templates); others benefit from diverse, creative outputs (e.g., marketing copy, storytelling). To meet those needs, you supply parameter controls with each inference request to influence token selection during generation.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/tJmUiudNjsCWp_bm/images/Introduction-to-Amazon-Bedrock/How-Foundation-Models-Process-Input/Controlling-Model-Parameters/model-responses-too-long-random-consistency.jpg?fit=max&auto=format&n=tJmUiudNjsCWp_bm&q=85&s=14aebd7a7d76f55d4445bcb865959725" alt="A presentation slide titled &#x22;Problem: Model Responses Are Too Long, Too Random.&#x22; It shows four numbered panels with icons describing issues: variable creativity/length, need for consistent outputs, some apps wanting more creative responses, and developers needing control over generation." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/How-Foundation-Models-Process-Input/Controlling-Model-Parameters/model-responses-too-long-random-consistency.jpg" />
</Frame>

## Why parameters matter

When generating text, the model predicts the next token from a set of candidate tokens and associated probabilities. Parameters let you influence which candidates are considered and how strongly high-probability options are favored. They tune generation behavior—creativity, randomness, and length—without changing the model’s underlying knowledge or training.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/tJmUiudNjsCWp_bm/images/Introduction-to-Amazon-Bedrock/How-Foundation-Models-Process-Input/Controlling-Model-Parameters/control-model-parameters-randomness-topk-length.jpg?fit=max&auto=format&n=tJmUiudNjsCWp_bm&q=85&s=e60aac93038d2a5cffc5660e50689a04" alt="A presentation slide titled &#x22;Solution: Control Model Parameters&#x22; with a central icon and the caption &#x22;Allow developers to influence how responses are generated&#x22; plus a note that parameters influence generation but do not guarantee outcomes. Below are three bars listing controls: control randomness in responses, control how many token options are considered, and limit the response length." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/How-Foundation-Models-Process-Input/Controlling-Model-Parameters/control-model-parameters-randomness-topk-length.jpg" />
</Frame>

The key parameters covered here:

* Temperature — controls how strongly the model favors high-probability tokens (affects randomness).
* Top P (nucleus sampling) — defines the probability mass threshold for the candidate pool (affects diversity).
* Max tokens — caps the number of tokens the model may generate (controls length and cost).

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/tJmUiudNjsCWp_bm/images/Introduction-to-Amazon-Bedrock/How-Foundation-Models-Process-Input/Controlling-Model-Parameters/parameters-matter-slide-three-cards.jpg?fit=max&auto=format&n=tJmUiudNjsCWp_bm&q=85&s=d40e0db611411679b5613730b65a9bd4" alt="A presentation slide titled &#x22;Solution: Why Parameters Matter&#x22; showing three blue gradient cards labeled 01–03 that say parameters influence how a model responds, do not change model knowledge, and affect creativity/randomness. Each card has a circular icon above its text." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/How-Foundation-Models-Process-Input/Controlling-Model-Parameters/parameters-matter-slide-three-cards.jpg" />
</Frame>

## Temperature

Temperature adjusts how strongly the model favors the highest-probability next token:

* Low temperature (close to 0): outputs are predictable and deterministic — ideal for structured tasks like code or factual answers.
* High temperature (closer to 1): outputs are more varied and creative — useful for brainstorming, slogans, or storytelling.

Example Python (Amazon Bedrock Runtime SDK) showing `temperature` and `maxTokens` in `inferenceConfig`:

```python theme={null}
import boto3

client = boto3.client("bedrock-runtime", region_name="us-east-1")

response = client.converse(
    modelId="us.meta.llama3-2-3b-instruct-v1:0",
    messages=[
        {
            "role": "user",
            "content": [{"text": "Write a short slogan for a cybersecurity company."}]
        }
    ],
    inferenceConfig={
        "temperature": 0.7,
        "maxTokens": 100
    }
)

print(response["output"]["message"]["content"][0]["text"])
```

A temperature of `0.7` is moderately high and tends to produce more creative outputs than lower values.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/tJmUiudNjsCWp_bm/images/Introduction-to-Amazon-Bedrock/How-Foundation-Models-Process-Input/Controlling-Model-Parameters/temperature-workflow-code-vs-marketing.jpg?fit=max&auto=format&n=tJmUiudNjsCWp_bm&q=85&s=8a463236518ac71b9a48eb3da75599aa" alt="An infographic titled &#x22;Workflow: Temperature&#x22; comparing low temperature (left, orange) for code generation — with icons and labels like &#x22;more predictable,&#x22; &#x22;more deterministic,&#x22; and &#x22;for structured tasks&#x22; — against high temperature (right, pink) for marketing/storytelling — labeled &#x22;more creative,&#x22; &#x22;more varied,&#x22; and &#x22;less predictable.&#x22;" width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/How-Foundation-Models-Process-Input/Controlling-Model-Parameters/temperature-workflow-code-vs-marketing.jpg" />
</Frame>

## Top P (nucleus sampling)

Top P controls diversity by specifying a cumulative probability threshold for the candidate token pool. The model samples only from the smallest set of tokens whose total probability mass is at least `topP`.

* Lower `topP` (e.g., `0.6`): limits the pool to the most probable tokens (less diverse).
* Higher `topP` (e.g., `0.95` or `1.0`): expands the pool to include less probable tokens (more creative).

Top P determines the candidate set; temperature governs how adventurous the sampling is within that set. Adjust one parameter at a time when experimenting.

Example snippet showing `topP`:

```python theme={null}
# Example showing topP in inferenceConfig
inferenceConfig = {
    "topP": 0.9,
    "maxTokens": 100
}
```

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/tJmUiudNjsCWp_bm/images/Introduction-to-Amazon-Bedrock/How-Foundation-Models-Process-Input/Controlling-Model-Parameters/top-p-sampling-workflow-notes.jpg?fit=max&auto=format&n=tJmUiudNjsCWp_bm&q=85&s=8b7209c173b76a56ecfad7ac8a5c8598" alt="A slide titled &#x22;Workflow: Top P&#x22; showing three numbered rounded panels (01, 02, 03) with icons. Each panel lists a brief note about Top-P sampling: controlling word diversity, fine-tuning response variability, and often left at default." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/How-Foundation-Models-Process-Input/Controlling-Model-Parameters/top-p-sampling-workflow-notes.jpg" />
</Frame>

## Max tokens

`maxTokens` sets an explicit upper bound on the number of tokens the model may generate in a single response. Use it to prevent runaway output and control cost.

* Use `maxTokens` to enforce concise answers for UI or API constraints.
* Lowering `maxTokens` forces brevity and reduces request cost.

Example snippet (integrate into the same `client.converse` call as above):

```python theme={null}
inferenceConfig = {
    "temperature": 0.3,
    "maxTokens": 50
}
```

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/tJmUiudNjsCWp_bm/images/Introduction-to-Amazon-Bedrock/How-Foundation-Models-Process-Input/Controlling-Model-Parameters/workflow-max-tokens-cards.jpg?fit=max&auto=format&n=tJmUiudNjsCWp_bm&q=85&s=85b08f7a3270e3257b280704200e92db" alt="A slide titled &#x22;Workflow: maxTokens&#x22; showing three numbered colorful cards with icons. They read: 01 — Controls maximum response length; 02 — Prevents run-away output; 03 — Impacts cost." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/How-Foundation-Models-Process-Input/Controlling-Model-Parameters/workflow-max-tokens-cards.jpg" />
</Frame>

## Quick reference table

| Parameter | Purpose | Typical range | When to use |
| - | - | -: | - |
| Temperature | Controls randomness / determinism | `0.0` — `1.0` (common: `0.0`–`1.0`) | Low for code/factual output; higher for creative text |
| Top P | Limits candidate token pool (nucleus sampling) | `0.0` — `1.0` (common: `0.6`–`0.95`) | Lower for focused outputs; higher for diversity |
| Max tokens | Caps the generated output length | Integer (tokens) | Enforce UI length or control cost |

## Practical guidance

* Change only one parameter at a time to isolate its effect. Record settings and outputs for reproducibility.
* For deterministic or structured tasks (code, templates): use low `temperature` (e.g., `0.0`–`0.3`) and optionally a low `topP`.
* For creative tasks (slogans, storytelling): increase `temperature` and/or `topP`.
* Always set a reasonable `maxTokens` to avoid excessive output and cost.

<Callout icon="lightbulb" color="#1CB2FE">
  When testing, record the exact settings (`temperature`, `topP`, `maxTokens`) and sample outputs. Run controlled experiments by changing one parameter at a time to evaluate its effect.
</Callout>

## Analogy: menu and adventure

Think of `topP` as how many desserts are listed on the menu (the pool of choices) and `temperature` as how adventurous you feel. A wide menu (`topP` high) plus high adventurousness (`temperature` high) increases the chance you’ll pick an unusual dessert. The two controls together determine the final choice.

## What parameters do not do

Model parameters:

* Do not make the model smarter.
* Do not add knowledge or change training data.
* Do not fundamentally improve the model’s reasoning capability.

They only influence next-token selection during generation.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/tJmUiudNjsCWp_bm/images/Introduction-to-Amazon-Bedrock/How-Foundation-Models-Process-Input/Controlling-Model-Parameters/workflow-what-parameters-dont-do.jpg?fit=max&auto=format&n=tJmUiudNjsCWp_bm&q=85&s=7155ee62d3ad64ba72c2d59cf78d11d7" alt="A presentation slide titled &#x22;Workflow: What Parameters Don't Do&#x22; showing four teal cards numbered 01–04. Each card has an icon and caption: &#x22;Make the model 'smarter',&#x22; &#x22;Increase knowledge,&#x22; &#x22;Change training data,&#x22; and &#x22;Improve reasoning capability.&#x22;" width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/How-Foundation-Models-Process-Input/Controlling-Model-Parameters/workflow-what-parameters-dont-do.jpg" />
</Frame>

## Expected results

When you tune these parameters appropriately, you should see:

* Improved consistency when limiting randomness.
* Greater flexibility and creativity when allowing more diversity.
* Fewer prompt-engineering iterations by adjusting generation behavior directly.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/tJmUiudNjsCWp_bm/images/Introduction-to-Amazon-Bedrock/How-Foundation-Models-Process-Input/Controlling-Model-Parameters/results-better-alignment-model-app-needs.jpg?fit=max&auto=format&n=tJmUiudNjsCWp_bm&q=85&s=afe6e00745df346221379a40fc1a7bce" alt="A presentation slide titled &#x22;Results&#x22; with a banner reading &#x22;Better alignment between model output and application needs.&#x22; Two numbered panels highlight &#x22;Improved consistency in AI responses&#x22; and &#x22;Greater flexibility for different use cases,&#x22; each with a simple icon (© KodeKloud in the corner)." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/How-Foundation-Models-Process-Input/Controlling-Model-Parameters/results-better-alignment-model-app-needs.jpg" />
</Frame>

## Key takeaway

Model parameters such as `temperature`, `topP`, and `maxTokens` let you influence how text is generated. Use them to tailor output length and creative variability to your application’s requirements while remembering they do not change the model’s underlying knowledge.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/tJmUiudNjsCWp_bm/images/Introduction-to-Amazon-Bedrock/How-Foundation-Models-Process-Input/Controlling-Model-Parameters/model-parameters-temperature-top-p-max-tokens.jpg?fit=max&auto=format&n=tJmUiudNjsCWp_bm&q=85&s=a849471c329b9b5b879db23dea789c50" alt="A presentation slide titled &#x22;Key Takeaway&#x22; noting that model parameters like temperature, top-p, and maxTokens influence how text is generated, helping tailor outputs to the application." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/How-Foundation-Models-Process-Input/Controlling-Model-Parameters/model-parameters-temperature-top-p-max-tokens.jpg" />
</Frame>

Should you use all parameters together? It depends on your use case. Understanding each control gives you the flexibility to produce either consistent or creative outputs as needed.

That concludes this lesson on controlling model parameters. Try the hands-on lab to practice tuning these settings and observe how each parameter affects generation.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/tJmUiudNjsCWp_bm/images/Introduction-to-Amazon-Bedrock/How-Foundation-Models-Process-Input/Controlling-Model-Parameters/whats-next-model-input-processing-lab.jpg?fit=max&auto=format&n=tJmUiudNjsCWp_bm&q=85&s=36f2fef52b8ae5b2b137b28c5d8e0fa7" alt="A presentation slide titled &#x22;What's Next? Experimenting with model input processing LAB&#x22; featuring a teal circular icon of a stylized brain with circuit lines on a dark curved background. A small &#x22;© Copyright KodeKloud&#x22; note appears in the bottom corner." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/How-Foundation-Models-Process-Input/Controlling-Model-Parameters/whats-next-model-input-processing-lab.jpg" />
</Frame>

Links and references

* Amazon Bedrock documentation: [https://docs.aws.amazon.com/bedrock](https://docs.aws.amazon.com/bedrock)
* Sampling and decoding strategies (overview): [https://distill.pub/2016/softmax/](https://distill.pub/2016/softmax/)
* Practical sampling discussion (temperature, top-p): [https://www.eleuther.ai/posts/top-p-sampling/](https://www.eleuther.ai/posts/top-p-sampling/)

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/introduction-to-amazon-bedrock/module/20770e4e-fc7b-49d2-b443-aaefd5e7d3e9/lesson/c4b957af-778a-41e0-b4ad-5178109afac4" />

  <Card title="Practice Lab" icon="flask-conical" cta="Learn more" href="https://learn.kodekloud.com/user/courses/introduction-to-amazon-bedrock/module/20770e4e-fc7b-49d2-b443-aaefd5e7d3e9/lesson/aa443879-398a-4e6a-b3d1-e3d21eb4b62f" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.