Skip to main content
In this lesson we cover how to implement multi-turn conversational AI using stateless foundation models (for example, Amazon Bedrock). This article explains the core challenge, a practical pattern for structured conversation state, a minimal chatbot workflow, code examples, and recommended storage options for session state. Contents
  • The core problem: foundation models are stateless
  • The solution: structured conversations and message schemas
  • Workflow: a basic chatbot loop
  • Example: in-memory chat using Bedrock Runtime (boto3)
  • Where to store conversation state (compare options)
  • Session patterns and a simple session-tracking example
  • Summary, best practices, and references
A dark-slide diagram titled "Lecture Flow" showing blue rounded boxes connected by arrows: Problem (models are stateless) → Solution (structured conversations) → Workflow (managing state). An arrow from Workflow leads down to Results (integrate AI assistants), which points left to Key Takeaway (history + messages).

Problem: foundation models are stateless

Foundation models process each request independently and do not retain memory between calls. This stateless design enables scalable hosting, but it also means the model has no built-in conversational memory: to produce coherent multi-turn responses you must provide prior context (history) with each request. When using Bedrock or similar runtime endpoints, the responsibility for preserving conversation context falls to the application. Practically, this means:
  • Recording what you send to the model (user messages, system prompts).
  • Recording what the model returns (assistant replies, metadata).
  • Re-sending the accumulated context on each new model call to enable multi-turn behavior.
The flow below visualizes the transition from stateless requests to a conversation-aware system.

Solution: structured conversations

To implement multi-turn chat reliably:
  1. Use a multi-turn/chat-style API (for example, the Converse method in the SDK).
  2. Track an ordered conversation history (messages with role + content).
  3. Optionally introduce orchestration (agents) or memory modules to extend capabilities.
A consistent message schema (e.g., { role: "user" | "assistant" | "system", content: [{ text: "..." }] }) lets you swap models or runtimes without rewriting your state management. Using the Converse-style API simplifies sending a message list and receiving structured replies.
A dark-themed presentation slide titled "Solution: Structured Conversations With Bedrock" with three numbered panels. The panels list: 01 Use Converse for multi-turn chat, 02 Track conversation history, and 03 Optionally use Agents for orchestration.

Workflow: basic chatbot loop

A minimal, production-reasonable chatbot loop typically follows these steps:
  1. User submits a message to your application.
  2. Append the user message to the conversation history (with a role label).
  3. Send the full conversation history to the model as the prompt:
    • For a brand-new session the history may contain only the system prompt and the first user message.
  4. Model returns a response. Persist that assistant reply in the history.
  5. Return the assistant reply to the user.
  6. Repeat for each new user input: append, send full history, persist reply.
Always sending the accumulating history supplies the stateless model with the context it needs to behave conversationally.
A diagram titled "Workflow: Basic Chatbot Workflow" showing a user on the left and an AI model on the right connected by a central box with three steps: 1) append user message to history, 2) send full conversation to model, 3) append AI response to history. Arrows show message → prompt → response → reply and note that the cycle repeats as the history grows.
If you do not persist and resend history, each new user message is treated as an independent call and the model will not recall earlier turns.

Example: simple chatbot with in-memory history

The example below demonstrates a minimal Python chatbot using the Bedrock Runtime (via boto3). It keeps a simple in-memory messages list, appends user and assistant turns, and re-sends the full history on each call. This pattern is suitable for development and prototypes; production systems should persist history externally.
The Converse API does not persist history for you. Persisting and managing the messages list (or per-session history) is the developer’s responsibility so that each subsequent call includes the appropriate context.

Where to store conversation state

An in-memory list can be fine for single-process demos. For resilient, multi-instance, and scalable systems you should use a datastore that matches your needs for latency, durability, and cost. Use this table as a quick comparison: Design considerations:
  • Session scoping: track a session_id for each conversation to map messages to sessions and support multiple concurrent sessions per user.
  • Retention & pruning: conversation histories grow quickly — consider sliding windows, summarization, or TTL-based pruning.
  • Security & privacy: avoid storing PII unnecessarily, encrypt data at rest and in transit, and enforce retention policies.

Sessions: why they matter

If multiple independent requests hit the same model endpoint without session-bound context, each request is treated as a separate conversation. To enable per-user or per-conversation continuity, bind messages to a session_id and ensure that session’s accumulated messages are included on every relevant model call.
A diagram titled "Workflow: The Need for Sessions" showing three separate boxes labeled Request 1, Request 2, and Request 3 sending arrows to a Bedrock foundation model icon. A note reads "No memory," with a caption saying the three requests are independent calls and the model treats each one as new.

Simple pattern for session tracking (in-memory)

Below is a compact pattern showing how to map session_id → messages in memory. Replace the in-memory dict with a persistent datastore (DynamoDB, Redis, Aurora) for production.
Best practices and operational tips
  • Use a stable message schema across your app to make model swaps or multi-model orchestration straightforward.
  • Consider a system prompt (role = system) to define assistant behavior consistently across sessions.
  • Implement summarization or compression for long histories (e.g., summarize older sections into a compact form) to stay within token limits.
  • Add explicit session timeouts and TTLs to limit storage costs and reduce exposure of sensitive data.
  • Monitor model usage and cost; full-history resends increase token consumption.

Summary and next steps

  • Foundation models are stateless: developers must track conversation history and re-send it to the model to achieve multi-turn behavior.
  • Use a structured message schema (role + content) and append both user and assistant turns to a session-scoped history that you re-send on each model call.
  • For production, persist histories using a datastore that matches your latency, durability, and scale requirements (DynamoDB, Redis, Aurora, etc.), adopt retention and privacy policies, and consider summarization strategies.
  • Next in the series: agents/orchestration, memory summarization techniques, and scalable session management patterns.
When persisting conversation history, be mindful of personally identifiable information (PII) and regulatory requirements. Encrypt sensitive data at rest and in transit, and apply appropriate retention policies.
Links and references

Watch Video