> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Temperature and Sampling

> Explains how model temperature and sampling affect generation randomness, variability, and recommended settings for tasks like code, conversation, and creative writing.

You asked the same prompt twice and got different answers. Why?

Large language models (LLMs) do not deterministically return a single “most likely” completion. Instead, at each token step the model computes scores (logits) for all possible next tokens, converts those to probabilities, and samples from that distribution. Higher-probability tokens are more likely to be chosen, but sampling introduces variability — the model effectively “rolls the dice.”

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/z7NmHsFQN9LCEiD0/images/AI-Agents-for-Beginners-OpenClaw-Case-Study/LLM-Fundamentals/Temperature-and-Sampling/prompt-panel-robot-story-probability-bars.jpg?fit=max&auto=format&n=z7NmHsFQN9LCEiD0&q=85&s=e7c25405875839941f8a267d7d1d810c" alt="A dark UI panel titled &#x22;PROMPT&#x22; displays the request &#x22;Write a one-sentence story about a robot.&#x22; Below it is a list of possible one-line responses with colored probability bars and percentages (the top choice &#x22;A lonely robot discovered...&#x22; shows 30%), and a caption at the bottom reads &#x22;Rolling the dice - higher probability = more likely to win.&#x22;" width="1920" height="1080" data-path="images/AI-Agents-for-Beginners-OpenClaw-Case-Study/LLM-Fundamentals/Temperature-and-Sampling/prompt-panel-robot-story-probability-bars.jpg" />
</Frame>

Temperature is the parameter that controls how “creative” or random that sampling is.

* Temperature is a scalar (commonly between 0 and 2) passed with each model call.
* Intuitively: think of temperature as a creativity dial — low values make the model conservative and repeatable, high values make it exploratory and varied.
* Mechanically: the model rescales logits by dividing by the temperature and then applies softmax:
  * `probs = softmax(logits / T)`
  * Lower `T` sharpens the distribution (more mass on high-scoring tokens).
  * Higher `T` flattens the distribution (spread probability across more tokens).

Common effects of `T`:

* `T = 0` (treated as greedy / argmax in many implementations): deterministic; the model picks the highest-probability token at each step and produces the same output for the same input.
* Low temperatures (e.g., `0.2–0.5`): favor likely completions with small variations.
* Mid-range (`~0.7`): natural conversational variability with reasonable reliability.
* High (`>1.0`): much more variety and surprising outputs, but with a higher risk of incoherence.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/z7NmHsFQN9LCEiD0/images/AI-Agents-for-Beginners-OpenClaw-Case-Study/LLM-Fundamentals/Temperature-and-Sampling/retro-temperature-creativity-dial-1-5.jpg?fit=max&auto=format&n=z7NmHsFQN9LCEiD0&q=85&s=ebc930223180b7496b08a37d2d84c41e" alt="A retro-styled infographic titled &#x22;Temperature: The Creativity Dial&#x22; showing a 0–2 slider set at 1.5 in the &#x22;creative&#x22; zone with labels Strict (green), Normal (yellow), Creative (red). Two boxes below explain &#x22;temp 0&#x22; (same response every time) and &#x22;temp 0.7&#x22; (mostly likely, occasional surprise)." width="1920" height="1080" data-path="images/AI-Agents-for-Beginners-OpenClaw-Case-Study/LLM-Fundamentals/Temperature-and-Sampling/retro-temperature-creativity-dial-1-5.jpg" />
</Frame>

Quick experiment you can run locally to observe temperature effects:

```python theme={null}
prompt = "Write a one-sentence story about a robot."

for temp in (0.0, 1.5):
    print(f"\nOutputs at temp={temp}\n")
    for i in range(10):
        # Replace `model.generate` with your API client's generation method.
        out = model.generate(prompt, temperature=temp)
        print(out)
```

Example results you might see:

```text theme={null}
temp=0.0
A lonely robot discovered a forgotten garden.
A lonely robot discovered a forgotten garden.
A lonely robot discovered a forgotten garden.
A lonely robot discovered a forgotten garden.
A lonely robot discovered a forgotten garden.
A lonely robot discovered a forgotten garden.
A lonely robot discovered a forgotten garden.
A lonely robot discovered a forgotten garden.
A lonely robot discovered a forgotten garden.
A lonely robot discovered a forgotten garden.

temp=1.5
In a quiet workshop, a lonely robot discovered a hummingbox of old songs.
A lonely robot built friends from junkyard parts.
A curious robot learned to dream after seeing the stars.
A robot made a garden where humans had forgotten to smile.
A robot built a small community from discarded toasters.
A shy robot discovered emotions in an old radio.
A robot learned to whistle and the city remembered laughter.
An old maintenance robot found a child’s abandoned kite and kept it safe.
A robot crafted companions from spare gears and kindness.
A playful robot taught itself to paint with dust motes.
```

Notice how `T = 0` yields identical outputs on every run, while `T = 1.5` produces wide variation.

When to use different temperatures

| Task | Recommended `temperature` | Notes |
| - | -: | - |
| Data extraction / classification | `0.0` | Deterministic outputs make parsing and automated checks reliable. |
| Code generation | `0.0–0.2` | Low variability reduces incorrect or surprising code suggestions. |
| General conversation / chat | `0.5–0.7` | Natural variability without losing coherence. |
| Creative writing / brainstorming | `0.8–1.2` | Encourages novel and surprising completions. |
| AI agents / decision flows | `0.0–0.2` | Low temperature yields consistent decisions and safer behavior. |

Practical tips

* Temperature is set per API call; choose the value that matches your tradeoff between determinism and creativity.
* For reproducibility in testing, set `seed` (if available) and low temperature.
* Combine low temperature with other decoding controls (e.g., top-p / nucleus sampling) when you need constrained diversity.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/z7NmHsFQN9LCEiD0/images/AI-Agents-for-Beginners-OpenClaw-Case-Study/LLM-Fundamentals/Temperature-and-Sampling/right-temperature-model-settings.jpg?fit=max&auto=format&n=z7NmHsFQN9LCEiD0&q=85&s=b9ee5dbbf6350d02fac0392e7f7d72f2" alt="An infographic titled &#x22;RIGHT TEMPERATURE&#x22; listing recommended model temperature settings for tasks: Data extraction (0), Code generation (0–0.2), Conversation (0.5–0.7), and Creative writing (0.8–1.2). A green banner below notes &#x22;FOR AGENTS: LOW TEMP = RELIABLE DECISIONS&#x22; and a footer says temperature is set per API call." width="1920" height="1080" data-path="images/AI-Agents-for-Beginners-OpenClaw-Case-Study/LLM-Fundamentals/Temperature-and-Sampling/right-temperature-model-settings.jpg" />
</Frame>

<Callout icon="lightbulb" color="#1CB2FE">
  Use lower temperatures when you need reproducible, reliable outputs (e.g., extraction, decisions, code). Increase temperature when you want diversity or creative exploration.
</Callout>

Further reading and references

* [Decoding strategies for neural language models (nucleus sampling, top-k)](https://arxiv.org/abs/1904.09751)
* \[Kubernetes of text: practical tips for stable generation workflows — model docs and API reference] (see your model provider’s decoding params)

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/ai-agents-for-beginner-openclaw-case-study/module/13d4f7ad-29e5-4bc0-b026-47c4ae43c31c/lesson/45224d26-a5b7-4d00-8b99-d10afadd3f36" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.