Skip to main content
Time to build your first AI agent loop. In this lesson you’ll implement the minimal control loop that drives every AI agent: a while loop that calls the chat completion endpoint, inspects the model’s finish_reason, and repeats until the model has finished. This example intentionally avoids tools and external actions — it focuses on the core loop logic.

Environment

  • Python 3.11 (virtual environment recommended)
  • OpenAI Python SDK pre-installed
  • Working directory: /root/code
  • Ensure these environment variables are set: OPENAI_API_KEY, OPENAI_API_BASE

Create agent_loop.py

Start by creating a new file named agent_loop.py. Import os and OpenAI, instantiate the client using environment variables, and initialize the conversation messages with a single system message:
This setup gives the model a system-level instruction and a client object you can use to call the chat completions API.

A minimal agent loop

The essential agent pattern is:
  1. Call the model.
  2. Inspect finish_reason on the first choice.
  3. If finish_reason == "stop", print the assistant’s answer and exit.
  4. Otherwise, handle non-terminal signals (e.g., function calls, token limits) in the else branch.
Example minimal loop:
Note: The else branch is intentionally simple here. When you add tools, function calling, or streaming, extend that branch to interpret the model signal, invoke tools, and feed results back into messages.

Add a user message and run it

Append a user message before the loop and run the script. For example:
Run the script from the shell:
Example console output:
One call, one answer, one exit — the simplest working agent loop.

Multi-turn conversation (handling multiple questions)

Real agents often handle multiple related questions in sequence. To preserve context across turns, append the assistant’s reply to messages after each response so subsequent calls see the full conversation history. Replace the single user message with a list of questions and iterate over them. For each question:
  • Append the user message
  • Enter the same loop and call the model
  • If finish_reason == "stop", print the reply and append the assistant reply to messages
Example:
Run the script again. You should see each question answered in turn, with each answer informed by previous context. This demonstrates multi-turn memory implemented simply with a Python list.
The finish_reason field indicates why the model stopped generating. Common values include "stop" (generation finished normally), "length" (truncated due to token limits), and "function_call" (model is invoking a function/tool). The else branch in the loop is where you’d implement handling for these non-terminal signals when integrating tools or function calls.

Common finish_reason values

Wrapping up

The while-loop that checks finish_reason == "stop" is the core control structure for simple AI agents. As you add tools, function calling, or streaming, extend the else branch to interpret model signals, perform actions, and feed results back into the conversation loop. This pattern scales from the simplest chat to powerful, tool-enabled agents.

Watch Video