Skip to main content
A session represents a single conversation between a user and your application. To maintain conversational state across turns you need a session store: a place to persist the ordered history of messages for each session. Common backend choices include relational databases (for example, Amazon Aurora), key-value caches (for example, Redis), or document stores (for example, Amazon DynamoDB). The exact technology depends on your scale, latency, and operational requirements. Typically, each conversation is assigned a unique session identifier (for example, a UUID or strings like user-1, user-456, user-c). That session ID is the lookup key in the session store, and the stored value is the ordered list of messages (roles: user or assistant, message content, timestamps, metadata). Example: a minimal in-memory session store (illustrative only — not for production):
This pattern stores each session as a list of message objects under the session ID key, enabling quick retrieval of conversational context for the session. For production, persist sessions to durable storage (DynamoDB, Redis, etc.) and include operational metadata such as timestamps, TTLs, and indexing keys. Context window and token limits When you pass the full conversation history to a stateless foundation model at each turn, the history grows and consumes tokens. Foundation models have a fixed context window: the total input tokens plus expected output tokens must fit within that window. A typical message flow includes a system prompt, prior user messages, prior assistant responses, the current user message, and the expected model response. Each turn adds tokens, and eventually the history may exceed the model’s context window.
A diagram titled "Workflow: Context Window Limits" showing stacked context blocks labeled System prompt, Previous user messages, Previous model responses, New user message (highlighted), and Response reserve. It illustrates how each turn adds a new user message into the context window.
As the conversation progresses, input tokens increase. For illustration, an early turn might consume ~50 tokens, by turn five ~250 tokens, by turn twenty ~1.2k, and by turn fifty ~6k. As history grows the available space for model-generated output shrinks. If the context window is hit, older messages get truncated, the model loses earlier context, responses can become inconsistent, and requests may fail.
A dark-themed infographic titled "Workflow: Context Window Limits" showing progress bars for Turn 1, Turn 5, Turn 20, and Turn 50 with approximate token counts of ~50, ~250, ~1.2k, and ~6k respectively. The bars grow longer with higher turn numbers to illustrate increasing context length.
Mitigating context-window limits You cannot keep appending the full history indefinitely. Common strategies to manage token growth:
  • Truncation (sliding window): keep only the most recent N messages/turns.
  • Summarization: compress older conversation segments into compact summaries that preserve important context.
  • Hybrid strategies: combine selective truncation, summaries, and per-message indexing in a datastore.
A simple sliding-window example:
Summarization is often more effective than aggressive truncation: periodically compress long segments of conversation into a short summary (for example, a paragraph) and store that summary alongside recent turns. On subsequent turns, pass the summary plus recent messages to the model to preserve key context while freeing token budget for the current interaction. Quick comparison (useful for design decisions): Using DynamoDB for session storage Amazon DynamoDB is a serverless, low-latency, highly scalable key-value and document store that often fits session-storage needs. However, DynamoDB’s query semantics (partition key + optional sort key) and item-size limits influence your table design. A straightforward approach is to store an entire conversation as a single item keyed by sessionId. On each request you fetch the item, append new messages, and write it back. This is easy to implement but may hit DynamoDB item size limits as history grows.
An infographic titled "Workflow: Use DynamoDB for Session Data" showing a circular four-step process around a central Amazon DynamoDB icon. The steps read: store message history by sessionId, load history on each request, append new messages, and save the updated conversation back to DynamoDB.
Production considerations Plan for operational requirements from the start. Consider these best practices:
  • Timestamp every message for ordering and auditability.
  • Use TTL (time-to-live) to expire stale sessions automatically.
  • Summarize older turns proactively to control token growth.
  • Store messages as individual items (one message per row) instead of a single large blob to enable selective retrieval and smaller reads/writes.
A presentation slide titled "Production: What a Simple Design Skips" showing four numbered rounded cards. Each card lists a best-practice: 01 Timestamp (order + audit every message), 02 TTL (auto-expire stale sessions), 03 Summarize (compress old turns), and 04 One per row (per-message item, not one blob).
DynamoDB item example (storing an entire conversation as a single item):
Design your DynamoDB table keys and indexes around your access patterns. If you need to query by attributes other than sessionId, add appropriate secondary indexes or adapt your schema to support efficient reads.
DynamoDB operational limits and recommendations Keep in mind DynamoDB’s constraints and operational trade-offs:
  • A single DynamoDB item has a maximum size of 400 KB. Storing an ever-growing session as one item can hit this limit.
  • Per-message items (one row per message) enable partial retrieval and reduce the likelihood of large writes.
  • Implement TTL and summarization to avoid unbounded growth and to keep latency predictable.
  • If you need to store large message payloads (attachments, long transcripts), consider external blob storage (S3) and reference keys in DynamoDB.
Avoid unbounded growth in stored session history. Implement TTLs, summarization, or per-message partitioning to prevent exceeding DynamoDB item size limits and to keep retrieval/write latency predictable.
Links and references Further reading
  • Designing for context windows and token budgets when building conversational AI
  • Approaches to incremental summarization for long-running conversations
  • DynamoDB best practices for time-series and per-item growth patterns

Watch Video