Skip to main content
In this lesson you’ll learn how to control model parameters to influence how a model generates responses. These controls help you steer output length, creativity, and determinism without changing the model’s pre-trained knowledge or reasoning capabilities. How can you steer a model’s responses? In practice, model outputs may be too long, too random, or too inconsistent for your application. Some use cases require highly repeatable, deterministic outputs (e.g., code generation, templates); others benefit from diverse, creative outputs (e.g., marketing copy, storytelling). To meet those needs, you supply parameter controls with each inference request to influence token selection during generation.
A presentation slide titled "Problem: Model Responses Are Too Long, Too Random." It shows four numbered panels with icons describing issues: variable creativity/length, need for consistent outputs, some apps wanting more creative responses, and developers needing control over generation.

Why parameters matter

When generating text, the model predicts the next token from a set of candidate tokens and associated probabilities. Parameters let you influence which candidates are considered and how strongly high-probability options are favored. They tune generation behavior—creativity, randomness, and length—without changing the model’s underlying knowledge or training.
A presentation slide titled "Solution: Control Model Parameters" with a central icon and the caption "Allow developers to influence how responses are generated" plus a note that parameters influence generation but do not guarantee outcomes. Below are three bars listing controls: control randomness in responses, control how many token options are considered, and limit the response length.
The key parameters covered here:
  • Temperature — controls how strongly the model favors high-probability tokens (affects randomness).
  • Top P (nucleus sampling) — defines the probability mass threshold for the candidate pool (affects diversity).
  • Max tokens — caps the number of tokens the model may generate (controls length and cost).
A presentation slide titled "Solution: Why Parameters Matter" showing three blue gradient cards labeled 01–03 that say parameters influence how a model responds, do not change model knowledge, and affect creativity/randomness. Each card has a circular icon above its text.

Temperature

Temperature adjusts how strongly the model favors the highest-probability next token:
  • Low temperature (close to 0): outputs are predictable and deterministic — ideal for structured tasks like code or factual answers.
  • High temperature (closer to 1): outputs are more varied and creative — useful for brainstorming, slogans, or storytelling.
Example Python (Amazon Bedrock Runtime SDK) showing temperature and maxTokens in inferenceConfig:
A temperature of 0.7 is moderately high and tends to produce more creative outputs than lower values.
An infographic titled "Workflow: Temperature" comparing low temperature (left, orange) for code generation — with icons and labels like "more predictable," "more deterministic," and "for structured tasks" — against high temperature (right, pink) for marketing/storytelling — labeled "more creative," "more varied," and "less predictable."

Top P (nucleus sampling)

Top P controls diversity by specifying a cumulative probability threshold for the candidate token pool. The model samples only from the smallest set of tokens whose total probability mass is at least topP.
  • Lower topP (e.g., 0.6): limits the pool to the most probable tokens (less diverse).
  • Higher topP (e.g., 0.95 or 1.0): expands the pool to include less probable tokens (more creative).
Top P determines the candidate set; temperature governs how adventurous the sampling is within that set. Adjust one parameter at a time when experimenting. Example snippet showing topP:
A slide titled "Workflow: Top P" showing three numbered rounded panels (01, 02, 03) with icons. Each panel lists a brief note about Top-P sampling: controlling word diversity, fine-tuning response variability, and often left at default.

Max tokens

maxTokens sets an explicit upper bound on the number of tokens the model may generate in a single response. Use it to prevent runaway output and control cost.
  • Use maxTokens to enforce concise answers for UI or API constraints.
  • Lowering maxTokens forces brevity and reduces request cost.
Example snippet (integrate into the same client.converse call as above):
A slide titled "Workflow: maxTokens" showing three numbered colorful cards with icons. They read: 01 — Controls maximum response length; 02 — Prevents run-away output; 03 — Impacts cost.

Quick reference table

Practical guidance

  • Change only one parameter at a time to isolate its effect. Record settings and outputs for reproducibility.
  • For deterministic or structured tasks (code, templates): use low temperature (e.g., 0.0–0.3) and optionally a low topP.
  • For creative tasks (slogans, storytelling): increase temperature and/or topP.
  • Always set a reasonable maxTokens to avoid excessive output and cost.
When testing, record the exact settings (temperature, topP, maxTokens) and sample outputs. Run controlled experiments by changing one parameter at a time to evaluate its effect.

Analogy: menu and adventure

Think of topP as how many desserts are listed on the menu (the pool of choices) and temperature as how adventurous you feel. A wide menu (topP high) plus high adventurousness (temperature high) increases the chance you’ll pick an unusual dessert. The two controls together determine the final choice.

What parameters do not do

Model parameters:
  • Do not make the model smarter.
  • Do not add knowledge or change training data.
  • Do not fundamentally improve the model’s reasoning capability.
They only influence next-token selection during generation.
A presentation slide titled "Workflow: What Parameters Don't Do" showing four teal cards numbered 01–04. Each card has an icon and caption: "Make the model 'smarter'," "Increase knowledge," "Change training data," and "Improve reasoning capability."

Expected results

When you tune these parameters appropriately, you should see:
  • Improved consistency when limiting randomness.
  • Greater flexibility and creativity when allowing more diversity.
  • Fewer prompt-engineering iterations by adjusting generation behavior directly.
A presentation slide titled "Results" with a banner reading "Better alignment between model output and application needs." Two numbered panels highlight "Improved consistency in AI responses" and "Greater flexibility for different use cases," each with a simple icon (© KodeKloud in the corner).

Key takeaway

Model parameters such as temperature, topP, and maxTokens let you influence how text is generated. Use them to tailor output length and creative variability to your application’s requirements while remembering they do not change the model’s underlying knowledge.
A presentation slide titled "Key Takeaway" noting that model parameters like temperature, top-p, and maxTokens influence how text is generated, helping tailor outputs to the application.
Should you use all parameters together? It depends on your use case. Understanding each control gives you the flexibility to produce either consistent or creative outputs as needed. That concludes this lesson on controlling model parameters. Try the hands-on lab to practice tuning these settings and observe how each parameter affects generation.
A presentation slide titled "What's Next? Experimenting with model input processing LAB" featuring a teal circular icon of a stylized brain with circuit lines on a dark curved background. A small "© Copyright KodeKloud" note appears in the bottom corner.
Links and references

Watch Video

Practice Lab