- Truncation — drop the oldest messages. Simple and fast, but lossy: the agent forgets early parts of the conversation.
- Summarization — run the LLM to condense older messages into a short summary. Preserves key facts, but costs extra API calls.
- Context compression — remove low‑value tokens while preserving high‑value facts. More selective and cost-effective but more complex to implement.

- Keep the most recent turns fully in context.
- Summarize or compress older turns automatically when token usage grows.
- Use indexed long-term storage to persist facts you definitely want remembered across sessions.


- ChatGPT’s memory feature is a real-world example: it saves small user facts (e.g., “user is vegetarian”, “user lives in Brooklyn”) and injects them as part of the system prompt for new conversations.
- A personal assistant agent that remembers preferences across sessions — travel, seating, favorite restaurants — needs a persistent memory layer backed by one of the approaches above.

- Cost — larger context windows mean more tokens per API call and higher cost.
- Quality — dropping or losing important context causes the agent to make worse decisions.
- Capability — without persistent memory, every conversation starts from scratch and multi-session features are impossible.

- Agents accumulate context fast; each tool call adds tokens.
- Short-term memory = context window; it resets each conversation.
- Long-term memory = external storage (key-value, vector DB, files).
- Context compression and summarization are essential for keeping the window usable.
- Persistent memory enables cross-session intelligence.
Design memory deliberately: decide what to store (facts vs. full transcripts), how to index it (keys vs. embeddings), and when to surface it to the LLM to balance cost, relevance, and latency.
Privacy & compliance: persist only what you are allowed to store. Encrypt or redact sensitive data and provide users with controls for what the agent remembers.
- OpenAI Agents & best practices: https://platform.openai.com/docs/guides/agents
- Vector databases and embeddings: https://www.pinecone.io/ and https://github.com/facebookresearch/faiss
- Key-value stores for session facts: https://redis.io/