Guide to making your first LLM call from Python using OpenAI SDK or HTTP, demonstrating tokens, temperature, context windows, and secure API key practices.
This guide assumes you already understand what an LLM is, how tokens work, how temperature affects generation, and what a context window does. Those theoretical concepts are important — now let’s map them to concrete code.By the end of this short walkthrough you’ll have a tiny Python program that sends a message to an LLM and prints the response.
Seven lines of Python — yet every major concept from the theory maps to one of those lines.When you use ChatGPT in the browser, a lot happens behind the scenes: your message is packaged, sent to a server, routed to the model, tokens are generated one by one, and the result streams back. The code you write implements that same flow but gives you control over each step.
You interact with the model via an API — think of it like an order form: fill it out, send it to the kitchen, and the kitchen returns your meal.
OpenAI exposes an API so you can call the same models programmatically. There are two common approaches from Python:
Use the SDK (recommended): the SDK handles HTTP, headers, JSON, and common errors for you—cleaner code and less boilerplate.
Use raw HTTP requests: instructive for learning, but more verbose.
Below are both approaches so you can compare.Raw HTTP (standard library) — more boilerplate:
Same request using the OpenAI Python SDK — cleaner and shorter:
from openai import OpenAIclient = OpenAI(api_key="sk-...")response = client.chat.completions.create( model="gpt-4o", messages=[{"role": "user", "content": "What is AI?"}])print(response.choices[0].message.content)
Store API keys securely — for example, in environment variables or a secrets manager. Never commit keys to source control or hard-code them in files.
Walkthrough of the SDK example (line-by-line)
from openai import OpenAI — import the SDK package.
client = OpenAI(api_key="sk-...") — create a reusable client instance that knows how to authenticate and call the API.
response = client.chat.completions.create(...) — send a chat completion request. Provide model and a messages list containing one or more message objects.
print(response.choices[0].message.content) — print the model’s reply. choices[0] is the first completion; message.content contains the returned text.
That short SDK example is the “seven-line” program referenced earlier.Adding temperature — control randomnessTemperature determines how deterministic or random the model’s output is:
temperature=0 → more deterministic and repeatable.
temperature=1 → more diverse and creative responses.
Everything in these examples follows the same five-step flow described earlier — packaging input, sending it to the model via API, model token generation, and returning text that you can consume in your program.