
Why parameters matter
When generating text, the model predicts the next token from a set of candidate tokens and associated probabilities. Parameters let you influence which candidates are considered and how strongly high-probability options are favored. They tune generation behavior—creativity, randomness, and length—without changing the model’s underlying knowledge or training.
- Temperature — controls how strongly the model favors high-probability tokens (affects randomness).
- Top P (nucleus sampling) — defines the probability mass threshold for the candidate pool (affects diversity).
- Max tokens — caps the number of tokens the model may generate (controls length and cost).

Temperature
Temperature adjusts how strongly the model favors the highest-probability next token:- Low temperature (close to 0): outputs are predictable and deterministic — ideal for structured tasks like code or factual answers.
- High temperature (closer to 1): outputs are more varied and creative — useful for brainstorming, slogans, or storytelling.
temperature and maxTokens in inferenceConfig:
0.7 is moderately high and tends to produce more creative outputs than lower values.

Top P (nucleus sampling)
Top P controls diversity by specifying a cumulative probability threshold for the candidate token pool. The model samples only from the smallest set of tokens whose total probability mass is at leasttopP.
- Lower
topP(e.g.,0.6): limits the pool to the most probable tokens (less diverse). - Higher
topP(e.g.,0.95or1.0): expands the pool to include less probable tokens (more creative).
topP:

Max tokens
maxTokens sets an explicit upper bound on the number of tokens the model may generate in a single response. Use it to prevent runaway output and control cost.
- Use
maxTokensto enforce concise answers for UI or API constraints. - Lowering
maxTokensforces brevity and reduces request cost.
client.converse call as above):

Quick reference table
Practical guidance
- Change only one parameter at a time to isolate its effect. Record settings and outputs for reproducibility.
- For deterministic or structured tasks (code, templates): use low
temperature(e.g.,0.0–0.3) and optionally a lowtopP. - For creative tasks (slogans, storytelling): increase
temperatureand/ortopP. - Always set a reasonable
maxTokensto avoid excessive output and cost.
When testing, record the exact settings (
temperature, topP, maxTokens) and sample outputs. Run controlled experiments by changing one parameter at a time to evaluate its effect.Analogy: menu and adventure
Think oftopP as how many desserts are listed on the menu (the pool of choices) and temperature as how adventurous you feel. A wide menu (topP high) plus high adventurousness (temperature high) increases the chance you’ll pick an unusual dessert. The two controls together determine the final choice.
What parameters do not do
Model parameters:- Do not make the model smarter.
- Do not add knowledge or change training data.
- Do not fundamentally improve the model’s reasoning capability.

Expected results
When you tune these parameters appropriately, you should see:- Improved consistency when limiting randomness.
- Greater flexibility and creativity when allowing more diversity.
- Fewer prompt-engineering iterations by adjusting generation behavior directly.

Key takeaway
Model parameters such astemperature, topP, and maxTokens let you influence how text is generated. Use them to tailor output length and creative variability to your application’s requirements while remembering they do not change the model’s underlying knowledge.


- Amazon Bedrock documentation: https://docs.aws.amazon.com/bedrock
- Sampling and decoding strategies (overview): https://distill.pub/2016/softmax/
- Practical sampling discussion (temperature, top-p): https://www.eleuther.ai/posts/top-p-sampling/