> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Implementing Conversational AI Part 2

> Best practices for session storage, context window management, and scaling conversational AI using DynamoDB, summarization, and sliding window strategies

A session represents a single conversation between a user and your application. To maintain conversational state across turns you need a session store: a place to persist the ordered history of messages for each session. Common backend choices include relational databases (for example, [Amazon Aurora](https://aws.amazon.com/rds/aurora/)), key-value caches (for example, [Redis](https://redis.io/)), or document stores (for example, [Amazon DynamoDB](https://aws.amazon.com/dynamodb/)). The exact technology depends on your scale, latency, and operational requirements.

Typically, each conversation is assigned a unique session identifier (for example, a UUID or strings like `user-1`, `user-456`, `user-c`). That session ID is the lookup key in the session store, and the stored value is the ordered list of messages (roles: user or assistant, message content, timestamps, metadata).

Example: a minimal in-memory session store (illustrative only — not for production):

```python theme={null}
# sessions is an in-memory store mapping session_id -> list_of_messages
sessions = {}

def chat(session_id, user_input):
    if session_id not in sessions:
        sessions[session_id] = []

    messages = sessions[session_id]

    # Append the incoming user message to the session history
    messages.append({
        "role": "user",
        "content": [{"text": user_input}]
    })

    # Call the conversational model with the accumulated history
    response = client.converse(
        modelId="amazon.nova-lite-v1:0",
        messages=messages
    )

    # Extract the assistant reply text from the model response
    reply = response["output"]["message"]["content"][0]["text"]

    # Append the assistant reply to the session history
    messages.append({
        "role": "assistant",
        "content": [{"text": reply}]
    })

    return reply
```

This pattern stores each session as a list of message objects under the session ID key, enabling quick retrieval of conversational context for the session. For production, persist sessions to durable storage (DynamoDB, Redis, etc.) and include operational metadata such as timestamps, TTLs, and indexing keys.

Context window and token limits

When you pass the full conversation history to a stateless foundation model at each turn, the history grows and consumes tokens. Foundation models have a fixed context window: the total input tokens plus expected output tokens must fit within that window. A typical message flow includes a system prompt, prior user messages, prior assistant responses, the current user message, and the expected model response. Each turn adds tokens, and eventually the history may exceed the model’s context window.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/BVCvDn4rl3j0TCQq/images/Introduction-to-Amazon-Bedrock/Advanced-Topics-Optional/Implementing-Conversational-AI-Part-2/context-window-limits-workflow-diagram.jpg?fit=max&auto=format&n=BVCvDn4rl3j0TCQq&q=85&s=fde8b5ce6d96070eb1c252cb6bb36b95" alt="A diagram titled &#x22;Workflow: Context Window Limits&#x22; showing stacked context blocks labeled System prompt, Previous user messages, Previous model responses, New user message (highlighted), and Response reserve. It illustrates how each turn adds a new user message into the context window." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Advanced-Topics-Optional/Implementing-Conversational-AI-Part-2/context-window-limits-workflow-diagram.jpg" />
</Frame>

As the conversation progresses, input tokens increase. For illustration, an early turn might consume \~50 tokens, by turn five \~250 tokens, by turn twenty \~1.2k, and by turn fifty \~6k. As history grows the available space for model-generated output shrinks. If the context window is hit, older messages get truncated, the model loses earlier context, responses can become inconsistent, and requests may fail.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/BVCvDn4rl3j0TCQq/images/Introduction-to-Amazon-Bedrock/Advanced-Topics-Optional/Implementing-Conversational-AI-Part-2/context-window-limits-progress-bars.jpg?fit=max&auto=format&n=BVCvDn4rl3j0TCQq&q=85&s=1c55c07b3908b7b5e486a77b52dda87d" alt="A dark-themed infographic titled &#x22;Workflow: Context Window Limits&#x22; showing progress bars for Turn 1, Turn 5, Turn 20, and Turn 50 with approximate token counts of ~50, ~250, ~1.2k, and ~6k respectively. The bars grow longer with higher turn numbers to illustrate increasing context length." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Advanced-Topics-Optional/Implementing-Conversational-AI-Part-2/context-window-limits-progress-bars.jpg" />
</Frame>

Mitigating context-window limits

You cannot keep appending the full history indefinitely. Common strategies to manage token growth:

* Truncation (sliding window): keep only the most recent N messages/turns.
* Summarization: compress older conversation segments into compact summaries that preserve important context.
* Hybrid strategies: combine selective truncation, summaries, and per-message indexing in a datastore.

A simple sliding-window example:

```python theme={null}
MAX_MESSAGES = 3
# Keep only the last MAX_MESSAGES messages in the history
messages = messages[-MAX_MESSAGES:]
```

Summarization is often more effective than aggressive truncation: periodically compress long segments of conversation into a short summary (for example, a paragraph) and store that summary alongside recent turns. On subsequent turns, pass the summary plus recent messages to the model to preserve key context while freeing token budget for the current interaction.

Quick comparison (useful for design decisions):

| Strategy | Strengths | When to use |
| - | - | - |
| Sliding window | Simple and low-cost | Short-lived sessions; stateless or limited storage |
| Summarization | Keeps key context with fewer tokens | Long-running sessions where historical context matters |
| Per-message storage + indexing | Selective retrieval and efficient reads | High-scale systems requiring partial history access |

Using DynamoDB for session storage

[Amazon DynamoDB](https://aws.amazon.com/dynamodb/) is a serverless, low-latency, highly scalable key-value and document store that often fits session-storage needs. However, DynamoDB’s query semantics (partition key + optional sort key) and item-size limits influence your table design.

A straightforward approach is to store an entire conversation as a single item keyed by `sessionId`. On each request you fetch the item, append new messages, and write it back. This is easy to implement but may hit DynamoDB item size limits as history grows.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/BVCvDn4rl3j0TCQq/images/Introduction-to-Amazon-Bedrock/Advanced-Topics-Optional/Implementing-Conversational-AI-Part-2/dynamodb-session-data-workflow.jpg?fit=max&auto=format&n=BVCvDn4rl3j0TCQq&q=85&s=86f8d5f74e6f4f17195de4e909088516" alt="An infographic titled &#x22;Workflow: Use DynamoDB for Session Data&#x22; showing a circular four-step process around a central Amazon DynamoDB icon. The steps read: store message history by sessionId, load history on each request, append new messages, and save the updated conversation back to DynamoDB." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Advanced-Topics-Optional/Implementing-Conversational-AI-Part-2/dynamodb-session-data-workflow.jpg" />
</Frame>

Production considerations

Plan for operational requirements from the start. Consider these best practices:

* Timestamp every message for ordering and auditability.
* Use TTL (time-to-live) to expire stale sessions automatically.
* Summarize older turns proactively to control token growth.
* Store messages as individual items (one message per row) instead of a single large blob to enable selective retrieval and smaller reads/writes.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/BVCvDn4rl3j0TCQq/images/Introduction-to-Amazon-Bedrock/Advanced-Topics-Optional/Implementing-Conversational-AI-Part-2/production-skips-timestamp-ttl-summarize-oneperrow.jpg?fit=max&auto=format&n=BVCvDn4rl3j0TCQq&q=85&s=19e2832e6898fa1c32377df060c7f89c" alt="A presentation slide titled &#x22;Production: What a Simple Design Skips&#x22; showing four numbered rounded cards. Each card lists a best-practice: 01 Timestamp (order + audit every message), 02 TTL (auto-expire stale sessions), 03 Summarize (compress old turns), and 04 One per row (per-message item, not one blob)." width="1920" height="1080" data-path="images/Introduction-to-Amazon-Bedrock/Advanced-Topics-Optional/Implementing-Conversational-AI-Part-2/production-skips-timestamp-ttl-summarize-oneperrow.jpg" />
</Frame>

DynamoDB item example (storing an entire conversation as a single item):

```json theme={null}
{
  "sessionId": "user-123",
  "messages": [
    {
      "role": "user",
      "content": [
        { "text": "Hello" }
      ]
    },
    {
      "role": "assistant",
      "content": [
        { "text": "Hi, how can I help?" }
      ]
    }
  ]
}
```

<Callout icon="lightbulb" color="#1CB2FE">
  Design your DynamoDB table keys and indexes around your access patterns. If you need to query by attributes other than `sessionId`, add appropriate secondary indexes or adapt your schema to support efficient reads.
</Callout>

DynamoDB operational limits and recommendations

Keep in mind DynamoDB’s constraints and operational trade-offs:

* A single DynamoDB item has a maximum size of 400 KB. Storing an ever-growing session as one item can hit this limit.
* Per-message items (one row per message) enable partial retrieval and reduce the likelihood of large writes.
* Implement TTL and summarization to avoid unbounded growth and to keep latency predictable.
* If you need to store large message payloads (attachments, long transcripts), consider external blob storage (S3) and reference keys in DynamoDB.

<Callout icon="warning" color="#FF6B6B">
  Avoid unbounded growth in stored session history. Implement TTLs, summarization, or per-message partitioning to prevent exceeding DynamoDB item size limits and to keep retrieval/write latency predictable.
</Callout>

Links and references

* [Amazon DynamoDB](https://aws.amazon.com/dynamodb/)
* [Amazon Aurora](https://aws.amazon.com/rds/aurora/)
* [Redis](https://redis.io/)

Further reading

* Designing for context windows and token budgets when building conversational AI
* Approaches to incremental summarization for long-running conversations
* DynamoDB best practices for time-series and per-item growth patterns

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/introduction-to-amazon-bedrock/module/7af9f623-7d4e-447a-8b21-6e635dfaccfa/lesson/43b36050-39ff-4cb4-a171-53956814b2da" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.