
- Temperature is a scalar (commonly between 0 and 2) passed with each model call.
- Intuitively: think of temperature as a creativity dial — low values make the model conservative and repeatable, high values make it exploratory and varied.
- Mechanically: the model rescales logits by dividing by the temperature and then applies softmax:
probs = softmax(logits / T)- Lower
Tsharpens the distribution (more mass on high-scoring tokens). - Higher
Tflattens the distribution (spread probability across more tokens).
T:
T = 0(treated as greedy / argmax in many implementations): deterministic; the model picks the highest-probability token at each step and produces the same output for the same input.- Low temperatures (e.g.,
0.2–0.5): favor likely completions with small variations. - Mid-range (
~0.7): natural conversational variability with reasonable reliability. - High (
>1.0): much more variety and surprising outputs, but with a higher risk of incoherence.

T = 0 yields identical outputs on every run, while T = 1.5 produces wide variation.
When to use different temperatures
Practical tips
- Temperature is set per API call; choose the value that matches your tradeoff between determinism and creativity.
- For reproducibility in testing, set
seed(if available) and low temperature. - Combine low temperature with other decoding controls (e.g., top-p / nucleus sampling) when you need constrained diversity.

Use lower temperatures when you need reproducible, reliable outputs (e.g., extraction, decisions, code). Increase temperature when you want diversity or creative exploration.
- Decoding strategies for neural language models (nucleus sampling, top-k)
- [Kubernetes of text: practical tips for stable generation workflows — model docs and API reference] (see your model provider’s decoding params)