Skip to main content
Zippy and Savvy forget everything when the conversation ends. Here’s why. An LLM’s context window is its working memory and it has a hard token limit. For a simple chatbot this is manageable, but for an agent that calls tools, searches, and reasons across many steps, the context window becomes a major engineering constraint. Agents accumulate context fast. Each iteration appends the agent’s internal reasoning, tool calls, and tool results into the conversation. A complex task with 10 tool calls can easily consume tens of thousands of tokens. Tool results can be particularly large — for example, a hotel search might return thousands of tokens of reviews, availability, and pricing. After a few iterations the context window comes under pressure. When the context window fills up, something has to give.
  • Truncation — drop the oldest messages. Simple and fast, but lossy: the agent forgets early parts of the conversation.
  • Summarization — run the LLM to condense older messages into a short summary. Preserves key facts, but costs extra API calls.
  • Context compression — remove low‑value tokens while preserving high‑value facts. More selective and cost-effective but more complex to implement.
An infographic titled "WHEN MEMORY FILLS UP" with three colored panels outlining strategies—Truncation, Summarization, and Compression—each listing a brief description and pros/cons. Truncation drops oldest messages, Summarization condenses content with an LLM, and Compression removes low-value content.
Most production agents combine these strategies. A common pattern is:
  1. Keep the most recent turns fully in context.
  2. Summarize or compress older turns automatically when token usage grows.
  3. Use indexed long-term storage to persist facts you definitely want remembered across sessions.
Short-term memory lives in the context window and resets with each new conversation. If you talk to an agent today and come back tomorrow, it won’t remember yesterday by default. Long-term memory is the external storage layer that persists beyond a single session.
An infographic titled "Two Types of Memory" comparing session memory and persistent memory. It shows session memory as the conversation's context that resets when the session ends, while persistent memory is described as external storage.
Long-term memory is not built into LLMs — it’s an engineering layer on top. Common approaches include:
An infographic titled "LONG‑TERM STORAGE" that lists three storage types—KEY‑VALUE STORE, VECTOR DATABASE, and FILE‑BASED—each with a brief description of how they're used. Examples shown include saving facts for prompts ("User is vegetarian"), using embeddings for semantic search, and storing files like context.md and notes.json.
Memory in practice
  • ChatGPT’s memory feature is a real-world example: it saves small user facts (e.g., “user is vegetarian”, “user lives in Brooklyn”) and injects them as part of the system prompt for new conversations.
  • A personal assistant agent that remembers preferences across sessions — travel, seating, favorite restaurants — needs a persistent memory layer backed by one of the approaches above.
An infographic titled "Memory in Practice" showing stored facts like "User is vegetarian," "User lives in Brooklyn," and "User prefers window seats" being injected into a conversation's context window as a system prompt when the user asks "Find me a restaurant." A banner at the bottom reads "Remembers preferences across sessions."
Why memory matters
  • Cost — larger context windows mean more tokens per API call and higher cost.
  • Quality — dropping or losing important context causes the agent to make worse decisions.
  • Capability — without persistent memory, every conversation starts from scratch and multi-session features are impossible.
A stylized infographic titled "WHY MEMORY MATTERS" showing three panels — COST, QUALITY, and CAPABILITY — that explain more context means more tokens/cost, lost context causes worse decisions, and no memory makes conversations start from scratch.
Building a robust memory system is one of the hardest parts of creating a production agent. It’s the difference between an agent that feels smart for a single conversation and one that feels smart over time. Key takeaways:
  • Agents accumulate context fast; each tool call adds tokens.
  • Short-term memory = context window; it resets each conversation.
  • Long-term memory = external storage (key-value, vector DB, files).
  • Context compression and summarization are essential for keeping the window usable.
  • Persistent memory enables cross-session intelligence.
Design memory deliberately: decide what to store (facts vs. full transcripts), how to index it (keys vs. embeddings), and when to surface it to the LLM to balance cost, relevance, and latency.
Privacy & compliance: persist only what you are allowed to store. Encrypt or redact sensitive data and provide users with controls for what the agent remembers.
We’ll solve persistent memory properly later when we introduce a specialist responsible for managing long-term memory for agents. References and further reading:

Watch Video

Practice Lab