Skip to main content
Time for hands-on practice. In this lab you’ll run a multi-turn agent conversation, watch the messages list grow with every turn, and implement a sliding-window memory manager to control token usage, cost, and latency.
A dark slide that reads "AGENT MEMORY IN ACTION" in large yellow pixelated text. Above it is a turquoise "HANDS-ON LAB" badge and below is the subtitle "Multi-turn conversation · Watch memory grow."

What you’ll accomplish

  • Run a multi-turn conversation with an agent that uses a calendar tool.
  • Observe how the conversation history (messages) grows and why that matters.
  • Implement trim_history() — a sliding-window memory manager that preserves the system prompt and keeps the most recent N messages.

Prerequisites

  • Python 3.11 with a virtual environment active.
  • The OpenAI SDK is pre-installed. See: OpenAI API (Python)
  • Working directory: /root/code
  • Repo already contains:
    • agent_memory.py — the agent loop wired up and importing a calendar tool from tools.py
    • tools.py — contains tools and execute_tool
Quick file reference:

Starter client snippet (already in the repo)

The repo contains this OpenAI client setup. You do not need to change it, but it shows how the client is instantiated:

1) Set the system prompt

Open agent_memory.py. Locate the messages list and the system prompt variable at the top of the conversation history. Set a concise system prompt to define the agent’s identity and tool usage expectations. Example guidance:
  • Tell the agent it is a helpful personal assistant.
  • Instruct it to use tools (like the calendar tool) when it needs verified or real data.
  • Keep the system prompt as the very first message in messages so the trimming logic can always preserve it.
Example system prompt assignment (place this as messages[0]):
Why this matters: the trimming function will always preserve messages[0] (the system prompt), so keep it compact and authoritative.

2) Add conversational questions that test memory

Add a sequence of user questions that force the agent to rely on conversational state across turns. Below is an example set of five questions. Append each user question to messages, call the agent, then observe the growth of messages. Example loop to add in agent_memory.py:
Example run output you might see:
Note: Every new user and assistant turn increases the messages list. If run_agent appends assistant responses, your list grows quickly and every API call will include that full history.

3) Why trimming is necessary

Long message histories:
  • Use more tokens per API call
  • Increase cost and latency
  • Can hit the model context limit
A simple and effective strategy: sliding-window memory that keeps the system prompt plus the most recent N-1 user/assistant messages.

4) Implement a sliding-window memory manager

Create a function named trim_history in agent_memory.py. Default max_messages=6 is a good starting point (1 system message + 5 recent turns).
Where to call it:
  • After appending the assistant response to messages, and before the next user turn, call trim_history(messages, max_messages=6).
Example integration inside the interaction loop:
This keeps the system prompt intact as messages[0] while limiting the rest of the retained history to the most recent conversation turns.
Keep the system prompt at messages[0]. The trimming function assumes that the first message is the agent’s system prompt and always preserves it.

5) Re-run the script and observe

Run agent_memory.py again after adding trim_history and integrating it into your loop. You should observe the message count stabilise around max_messages (6 by default). The agent will continue to act coherently because:
  • The system instruction remains preserved.
  • The most recent turns (the most relevant context) are kept.
This sliding-window approach provides a simple, production-ready memory manager:
  • System prompt always preserved
  • History capped to control token usage
  • Reduces cost and latency while keeping recent context

Further reading and references

A retro-style graphic with a green checkmark and bold yellow pixelated text reading "MEMORY MANAGED." Below it are smaller lines saying "System prompt always preserved · History capped at 6" and "A working memory manager for production agents."

Watch Video

Practice Lab