> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Your First LLM Call

> Guide to making your first LLM call from Python using OpenAI SDK or HTTP, demonstrating tokens, temperature, context windows, and secure API key practices.

This guide assumes you already understand what an LLM is, how tokens work, how temperature affects generation, and what a context window does. Those theoretical concepts are important — now let's map them to concrete code.

By the end of this short walkthrough you'll have a tiny Python program that sends a message to an LLM and prints the response.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/z7NmHsFQN9LCEiD0/images/AI-Agents-for-Beginners-OpenClaw-Case-Study/LLM-Fundamentals/Your-First-LLM-Call/retro-llm-theory-to-python-code.jpg?fit=max&auto=format&n=z7NmHsFQN9LCEiD0&q=85&s=3876810f72be65bca1618b7d77818087" alt="A dark, retro-style slide titled &#x22;FROM THEORY TO CODE&#x22; lists items like LLM, TOKENS, TEMPERATURE, and CONTEXT WINDOWS with an arrow pointing to a code panel labeled &#x22;PYTHON.&#x22; The design uses pixelated yellow heading text, colored bullet points, and a grid background." width="1920" height="1080" data-path="images/AI-Agents-for-Beginners-OpenClaw-Case-Study/LLM-Fundamentals/Your-First-LLM-Call/retro-llm-theory-to-python-code.jpg" />
</Frame>

Seven lines of Python — yet every major concept from the theory maps to one of those lines.

When you use ChatGPT in the browser, a lot happens behind the scenes: your message is packaged, sent to a server, routed to the model, tokens are generated one by one, and the result streams back. The code you write implements that same flow but gives you control over each step.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/z7NmHsFQN9LCEiD0/images/AI-Agents-for-Beginners-OpenClaw-Case-Study/LLM-Fundamentals/Your-First-LLM-Call/neon-5-step-message-to-tokens.jpg?fit=max&auto=format&n=z7NmHsFQN9LCEiD0&q=85&s=254cf1dd49b38ff40e7a3a6234f56928" alt="A neon-style infographic titled &#x22;BEHIND THE SCENES&#x22; showing a five-step flow from &#x22;Your message&#x22; being packaged, sent over the internet to a server, routed to a model, tokens generated one by one, and then streamed back. Each step is shown with a numbered colored circle and a short caption." width="1920" height="1080" data-path="images/AI-Agents-for-Beginners-OpenClaw-Case-Study/LLM-Fundamentals/Your-First-LLM-Call/neon-5-step-message-to-tokens.jpg" />
</Frame>

You interact with the model via an API — think of it like an order form: fill it out, send it to the kitchen, and the kitchen returns your meal.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/z7NmHsFQN9LCEiD0/images/AI-Agents-for-Beginners-OpenClaw-Case-Study/LLM-Fundamentals/Your-First-LLM-Call/api-order-form-to-kitchen-diagram.jpg?fit=max&auto=format&n=z7NmHsFQN9LCEiD0&q=85&s=586583d12080cfba4e0022897adb511d" alt="A neon-style diagram titled &#x22;API — Application Programming Interface&#x22; showing an &#x22;ORDER FORM (fill it out correctly)&#x22; box with an arrow pointing to a &#x22;KITCHEN (makes your food)&#x22; box." width="1920" height="1080" data-path="images/AI-Agents-for-Beginners-OpenClaw-Case-Study/LLM-Fundamentals/Your-First-LLM-Call/api-order-form-to-kitchen-diagram.jpg" />
</Frame>

OpenAI exposes an API so you can call the same models programmatically. There are two common approaches from Python:

* Use the SDK (recommended): the SDK handles HTTP, headers, JSON, and common errors for you—cleaner code and less boilerplate.
* Use raw HTTP requests: instructive for learning, but more verbose.

Below are both approaches so you can compare.

Raw HTTP (standard library) — more boilerplate:

```python theme={null}
import urllib.request
import json

url = "https://api.openai.com/v1/chat/completions"
headers = {
    "Authorization": "Bearer sk-...",
    "Content-Type": "application/json"
}
payload = {
    "model": "gpt-4o",
    "messages": [{"role": "user", "content": "Hello!"}]
}

data = json.dumps(payload).encode("utf-8")
req = urllib.request.Request(url, data=data, headers=headers)
with urllib.request.urlopen(req) as resp:
    result = json.load(resp)

print(result["choices"][0]["message"]["content"])
```

Same request using the OpenAI Python SDK — cleaner and shorter:

```python theme={null}
from openai import OpenAI

client = OpenAI(api_key="sk-...")

response = client.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "What is AI?"}]
)

print(response.choices[0].message.content)
```

<Callout icon="lightbulb" color="#1CB2FE">
  Store API keys securely — for example, in environment variables or a secrets manager. Never commit keys to source control or hard-code them in files.
</Callout>

Walkthrough of the SDK example (line-by-line)

1. `from openai import OpenAI` — import the SDK package.
2. `client = OpenAI(api_key="sk-...")` — create a reusable client instance that knows how to authenticate and call the API.
3. `response = client.chat.completions.create(...)` — send a chat completion request. Provide `model` and a `messages` list containing one or more message objects.
4. `print(response.choices[0].message.content)` — print the model’s reply. `choices[0]` is the first completion; `message.content` contains the returned text.

That short SDK example is the “seven-line” program referenced earlier.

Adding temperature — control randomness

Temperature determines how deterministic or random the model’s output is:

* `temperature=0` → more deterministic and repeatable.
* `temperature=1` → more diverse and creative responses.

Set it as a parameter in the same API call:

```python theme={null}
response = client.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "What is AI?"}],
    temperature=0
)
```

or

```python theme={null}
response = client.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "What is AI?"}],
    temperature=1
)
```

How the code maps to the key LLM concepts

| Concept | Where it appears in code / request |
| - | - |
| LLM output | `response.choices[0].message.content` |
| Prompt(s) | `messages` array; each message has a `content` field |
| Tokens | The `messages` list is tokenized into model tokens before being sent |
| Temperature | `temperature` parameter in the request |
| Context window usage | Total number of tokens in the entire `messages` list counts against the model's context window |

Practical tips and next steps

* Prefer the SDK for production code to reduce boilerplate and handle retries/errors more gracefully.
* Monitor token usage and set conservative temperatures for predictable outputs.
* Use system messages (e.g., `{"role":"system","content":"You are a helpful assistant."}`) to steer behavior.
* For full API docs and model capabilities, see the OpenAI API reference:
  * [https://platform.openai.com/docs](https://platform.openai.com/docs) (OpenAI Platform documentation)

Everything in these examples follows the same five-step flow described earlier — packaging input, sending it to the model via API, model token generation, and returning text that you can consume in your program.

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/ai-agents-for-beginner-openclaw-case-study/module/13d4f7ad-29e5-4bc0-b026-47c4ae43c31c/lesson/af6eb778-95ba-4ac2-b4db-d97a579381f2" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.