- The core problem: foundation models are stateless
- The solution: structured conversations and message schemas
- Workflow: a basic chatbot loop
- Example: in-memory chat using Bedrock Runtime (boto3)
- Where to store conversation state (compare options)
- Session patterns and a simple session-tracking example
- Summary, best practices, and references

Problem: foundation models are stateless
Foundation models process each request independently and do not retain memory between calls. This stateless design enables scalable hosting, but it also means the model has no built-in conversational memory: to produce coherent multi-turn responses you must provide prior context (history) with each request. When using Bedrock or similar runtime endpoints, the responsibility for preserving conversation context falls to the application. Practically, this means:- Recording what you send to the model (user messages, system prompts).
- Recording what the model returns (assistant replies, metadata).
- Re-sending the accumulated context on each new model call to enable multi-turn behavior.
Solution: structured conversations
To implement multi-turn chat reliably:- Use a multi-turn/chat-style API (for example, the
Conversemethod in the SDK). - Track an ordered conversation history (messages with
role+content). - Optionally introduce orchestration (agents) or memory modules to extend capabilities.
{ role: "user" | "assistant" | "system", content: [{ text: "..." }] }) lets you swap models or runtimes without rewriting your state management. Using the Converse-style API simplifies sending a message list and receiving structured replies.

Workflow: basic chatbot loop
A minimal, production-reasonable chatbot loop typically follows these steps:- User submits a message to your application.
- Append the user message to the conversation history (with a role label).
- Send the full conversation history to the model as the prompt:
- For a brand-new session the history may contain only the system prompt and the first user message.
- Model returns a response. Persist that assistant reply in the history.
- Return the assistant reply to the user.
- Repeat for each new user input: append, send full history, persist reply.

Example: simple chatbot with in-memory history
The example below demonstrates a minimal Python chatbot using the Bedrock Runtime (via boto3). It keeps a simple in-memorymessages list, appends user and assistant turns, and re-sends the full history on each call. This pattern is suitable for development and prototypes; production systems should persist history externally.
The Converse API does not persist history for you. Persisting and managing the
messages list (or per-session history) is the developer’s responsibility so that each subsequent call includes the appropriate context.Where to store conversation state
An in-memory list can be fine for single-process demos. For resilient, multi-instance, and scalable systems you should use a datastore that matches your needs for latency, durability, and cost. Use this table as a quick comparison:
Design considerations:
- Session scoping: track a
session_idfor each conversation to map messages to sessions and support multiple concurrent sessions per user. - Retention & pruning: conversation histories grow quickly — consider sliding windows, summarization, or TTL-based pruning.
- Security & privacy: avoid storing PII unnecessarily, encrypt data at rest and in transit, and enforce retention policies.
Sessions: why they matter
If multiple independent requests hit the same model endpoint without session-bound context, each request is treated as a separate conversation. To enable per-user or per-conversation continuity, bind messages to asession_id and ensure that session’s accumulated messages are included on every relevant model call.

Simple pattern for session tracking (in-memory)
Below is a compact pattern showing how to mapsession_id → messages in memory. Replace the in-memory dict with a persistent datastore (DynamoDB, Redis, Aurora) for production.
- Use a stable message schema across your app to make model swaps or multi-model orchestration straightforward.
- Consider a system prompt (role =
system) to define assistant behavior consistently across sessions. - Implement summarization or compression for long histories (e.g., summarize older sections into a compact form) to stay within token limits.
- Add explicit session timeouts and TTLs to limit storage costs and reduce exposure of sensitive data.
- Monitor model usage and cost; full-history resends increase token consumption.
Summary and next steps
- Foundation models are stateless: developers must track conversation history and re-send it to the model to achieve multi-turn behavior.
- Use a structured message schema (role + content) and append both user and assistant turns to a session-scoped history that you re-send on each model call.
- For production, persist histories using a datastore that matches your latency, durability, and scale requirements (DynamoDB, Redis, Aurora, etc.), adopt retention and privacy policies, and consider summarization strategies.
- Next in the series: agents/orchestration, memory summarization techniques, and scalable session management patterns.
When persisting conversation history, be mindful of personally identifiable information (PII) and regulatory requirements. Encrypt sensitive data at rest and in transit, and apply appropriate retention policies.
- Amazon Bedrock documentation: https://docs.aws.amazon.com/bedrock/
- boto3 (AWS SDK for Python): https://boto3.amazonaws.com/v1/documentation/api/latest/index.html
- DynamoDB: https://aws.amazon.com/dynamodb/
- Redis: https://redis.io/
- Amazon Aurora: https://aws.amazon.com/rds/aurora/