Skip to main content
The context window is an LLM’s active working memory: the total amount of text the model can see for a single request. That includes system instructions, the full conversation history, the user’s most recent message, and the model’s own generated output. All of that content consumes the same token budget. A simple example of how a single turn might be packed into the context window:
  • System instructions: 200 tokens
  • Conversation history: 3,000 tokens
  • Latest user message: 150 tokens
  • Model’s response (in-progress): 500 tokens
Total: ~3,850 tokens. If your model has a 128,000-token context window,
A retro-style chat UI labeled "Context Window" showing a travel-agent conversation about booking a one-way flight to NYC under $300. A right-side panel displays token counts (system, history, latest message, response) and a total token meter reading 3,850 / 128,000.
you still have plenty of room for long documents and extended multi-turn conversations. But remember: the context window is a hard, per-request limit — anything outside that window is invisible to the model. Typical (approximate) context window sizes vary by model and provider: For more details see provider docs such as the OpenAI API docs, Anthropic docs, or Google Cloud AI documentation.
A colorful horizontal bar chart titled "CONTEXT WINDOWS" comparing context window sizes of different AI models. It shows GPT‑3.5 (4K), GPT‑4 (8K), GPT‑4o (128K), Claude 3.5 (200K) and Gemini 1.5 (1M).
Why the window size matters
  • Larger windows let models process longer documents, keep more conversation history in context, and perform deeper multi-step reasoning.
  • However, every token you send (including repeated context) is counted and typically billed, so bigger windows increase potential cost.
What happens when the context window fills When the conversation content exceeds the model’s context limit, systems commonly handle it in one of three ways:
  • Truncation — drop the oldest messages so the most recent content stays visible.
  • Summarization — compress earlier messages into a shorter summary and inject that summary in place of the full history.
  • Sliding window — maintain a fixed-size view that advances through the conversation, preserving only recent turns.
Truncation example: if early in the chat you told the assistant “your name is Alex, you prefer afternoon flights, and you’re vegetarian,” those earliest facts can be dropped as the window fills. The model then no longer “knows” those facts because they fell outside the active context window — not because it forgot, but because they were removed.
A dark-themed UI screenshot titled "TRUNCATION" showing chat turns where earlier messages are marked DROPPED and later turns (e.g., "Book a flight to NYC", "Under $300 please", "What options?") are VISIBLE. Warning boxes read "Model no longer knows your name" and "Messages fell outside the context window."
Context window vs. long-term memory A helpful mental model: the context window is like a whiteboard — everything on it is currently visible, and when it fills up you erase the oldest content. Long-term memory is like a notebook you can consult later. LLMs do not have native persistent long-term memory; many products simulate memory by storing facts externally and injecting them into the model’s context for each API call. The underlying model still only sees the injected text within the current window.
A retro-style infographic titled "Key Distinction" comparing a left-side "Whiteboard" (context window: currently visible, erases when full, temporary) with a right-side "Notebook" (long-term memory: saved facts, look up anytime, permanent). A caption at the bottom notes LLMs have a context window but no native memory and that memory features inject saved facts into context.
Remember: storing a fact in a product-specific memory store is not the same as the model having native, persistent memory. The product injects those facts into the model’s context each time you call the API.
Why context management matters for agents For a simple chatbot, running out of context typically means the model forgets earlier turns — annoying but usually manageable. For AI agents, however, managing context is a central engineering challenge. Agents may need many items in-context simultaneously:
  • System instructions (behavior and safety rules)
  • Tool definitions and calling conventions (available APIs, parameters)
  • Full conversation history
  • Results returned from tool calls (API responses, database results)
  • Intermediate reasoning or chain-of-thought traces
All of this consumes tokens rapidly. Approximate token budget for a single complex agent turn:
A neon-style horizontal bar chart titled "AGENT BUDGET" that shows approximate token allocations per turn for categories like System prompt (~2K), Tool definitions (~3K), Conversation (~4K), Tool results (~8K), Reasoning (~4K) and Response (~2K). The image notes a total of ~23,000 tokens per turn and includes the caption "Context management = real engineering challenge."
Cost and engineering trade-offs
  • Providers bill per token: sending 100,000 input tokens in a request is billed for all those tokens. If an agent makes multiple calls that repeat large context blocks, costs multiply.
  • Trade-off: more context generally improves output quality, but it increases latency and cost.
  • Key engineering approaches:
    • Externalize and index large documents; inject only relevant excerpts (retrieval-augmented generation).
    • Summarize older conversation into compact memory representations.
    • Design tools and prompts to reduce per-turn token usage (e.g., shorter tool definitions, structured outputs).
    • Cache repeated content and reference it efficiently rather than re-sending full blobs.
Further reading and resources Managing the context window effectively — choosing what to keep, compress, or externalize; when to call tools; and how to design prompts and tool interfaces — is one of the most important engineering challenges when building reliable, scalable LLM-based agents.

Watch Video