Skip to main content
You asked the same prompt twice and got different answers. Why? Large language models (LLMs) do not deterministically return a single “most likely” completion. Instead, at each token step the model computes scores (logits) for all possible next tokens, converts those to probabilities, and samples from that distribution. Higher-probability tokens are more likely to be chosen, but sampling introduces variability — the model effectively “rolls the dice.”
A dark UI panel titled "PROMPT" displays the request "Write a one-sentence story about a robot." Below it is a list of possible one-line responses with colored probability bars and percentages (the top choice "A lonely robot discovered..." shows 30%), and a caption at the bottom reads "Rolling the dice - higher probability = more likely to win."
Temperature is the parameter that controls how “creative” or random that sampling is.
  • Temperature is a scalar (commonly between 0 and 2) passed with each model call.
  • Intuitively: think of temperature as a creativity dial — low values make the model conservative and repeatable, high values make it exploratory and varied.
  • Mechanically: the model rescales logits by dividing by the temperature and then applies softmax:
    • probs = softmax(logits / T)
    • Lower T sharpens the distribution (more mass on high-scoring tokens).
    • Higher T flattens the distribution (spread probability across more tokens).
Common effects of T:
  • T = 0 (treated as greedy / argmax in many implementations): deterministic; the model picks the highest-probability token at each step and produces the same output for the same input.
  • Low temperatures (e.g., 0.2–0.5): favor likely completions with small variations.
  • Mid-range (~0.7): natural conversational variability with reasonable reliability.
  • High (>1.0): much more variety and surprising outputs, but with a higher risk of incoherence.
A retro-styled infographic titled "Temperature: The Creativity Dial" showing a 0–2 slider set at 1.5 in the "creative" zone with labels Strict (green), Normal (yellow), Creative (red). Two boxes below explain "temp 0" (same response every time) and "temp 0.7" (mostly likely, occasional surprise).
Quick experiment you can run locally to observe temperature effects:
Example results you might see:
Notice how T = 0 yields identical outputs on every run, while T = 1.5 produces wide variation. When to use different temperatures Practical tips
  • Temperature is set per API call; choose the value that matches your tradeoff between determinism and creativity.
  • For reproducibility in testing, set seed (if available) and low temperature.
  • Combine low temperature with other decoding controls (e.g., top-p / nucleus sampling) when you need constrained diversity.
An infographic titled "RIGHT TEMPERATURE" listing recommended model temperature settings for tasks: Data extraction (0), Code generation (0–0.2), Conversation (0.5–0.7), and Creative writing (0.8–1.2). A green banner below notes "FOR AGENTS: LOW TEMP = RELIABLE DECISIONS" and a footer says temperature is set per API call.
Use lower temperatures when you need reproducible, reliable outputs (e.g., extraction, decisions, code). Increase temperature when you want diversity or creative exploration.
Further reading and references

Watch Video