Explains LLM context windows, sizes, effects on memory and cost, and engineering strategies for managing tokens, truncation, summarization, and agent context budgeting.
The context window is an LLM’s active working memory: the total amount of text the model can see for a single request. That includes system instructions, the full conversation history, the user’s most recent message, and the model’s own generated output. All of that content consumes the same token budget.A simple example of how a single turn might be packed into the context window:
System instructions: 200 tokens
Conversation history: 3,000 tokens
Latest user message: 150 tokens
Model’s response (in-progress): 500 tokens
Total: ~3,850 tokens.If your model has a 128,000-token context window,
you still have plenty of room for long documents and extended multi-turn conversations. But remember: the context window is a hard, per-request limit — anything outside that window is invisible to the model.Typical (approximate) context window sizes vary by model and provider:
Model / Family
Approximate context window
GPT-3.5
~4,096 tokens
GPT-4 (standard)
~8,000 tokens
GPT-4 with larger variants
32K+ tokens
GPT-4o (variants)
up to ~128,000 tokens
Claude 3.5 Sonnet
up to ~200,000 tokens (approx.)
Gemini 1.5 Pro
up to ~1,000,000 tokens (approx.)
For more details see provider docs such as the OpenAI API docs, Anthropic docs, or Google Cloud AI documentation.
Why the window size matters
Larger windows let models process longer documents, keep more conversation history in context, and perform deeper multi-step reasoning.
However, every token you send (including repeated context) is counted and typically billed, so bigger windows increase potential cost.
What happens when the context window fills
When the conversation content exceeds the model’s context limit, systems commonly handle it in one of three ways:
Truncation — drop the oldest messages so the most recent content stays visible.
Summarization — compress earlier messages into a shorter summary and inject that summary in place of the full history.
Sliding window — maintain a fixed-size view that advances through the conversation, preserving only recent turns.
Truncation example: if early in the chat you told the assistant “your name is Alex, you prefer afternoon flights, and you’re vegetarian,” those earliest facts can be dropped as the window fills. The model then no longer “knows” those facts because they fell outside the active context window — not because it forgot, but because they were removed.
Context window vs. long-term memory
A helpful mental model: the context window is like a whiteboard — everything on it is currently visible, and when it fills up you erase the oldest content. Long-term memory is like a notebook you can consult later. LLMs do not have native persistent long-term memory; many products simulate memory by storing facts externally and injecting them into the model’s context for each API call. The underlying model still only sees the injected text within the current window.
Remember: storing a fact in a product-specific memory store is not the same as the model having native, persistent memory. The product injects those facts into the model’s context each time you call the API.
Why context management matters for agents
For a simple chatbot, running out of context typically means the model forgets earlier turns — annoying but usually manageable. For AI agents, however, managing context is a central engineering challenge. Agents may need many items in-context simultaneously:
System instructions (behavior and safety rules)
Tool definitions and calling conventions (available APIs, parameters)
Full conversation history
Results returned from tool calls (API responses, database results)
Intermediate reasoning or chain-of-thought traces
All of this consumes tokens rapidly.Approximate token budget for a single complex agent turn:
Category
Approximate tokens
System prompt
~2,000
Tool definitions (e.g., ~10 tools)
~3,000
Conversation history
~4,000 (varies)
Tool results
~8,000
Agent reasoning / chain-of-thought
~4,000
Space for agent’s response
~2,000
Total (example)
~23,000
Cost and engineering trade-offs
Providers bill per token: sending 100,000 input tokens in a request is billed for all those tokens. If an agent makes multiple calls that repeat large context blocks, costs multiply.
Trade-off: more context generally improves output quality, but it increases latency and cost.
Key engineering approaches:
Externalize and index large documents; inject only relevant excerpts (retrieval-augmented generation).
Summarize older conversation into compact memory representations.
Design tools and prompts to reduce per-turn token usage (e.g., shorter tool definitions, structured outputs).
Cache repeated content and reference it efficiently rather than re-sending full blobs.
Managing the context window effectively — choosing what to keep, compress, or externalize; when to call tools; and how to design prompts and tool interfaces — is one of the most important engineering challenges when building reliable, scalable LLM-based agents.