> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# OpenClaw Agent Loop

> Describes the OpenClaw agent loop with five phases handling messages, model tool calls, error recovery, context compaction, concurrency, auth failover, and durable session persistence for production use

Every incoming message (WhatsApp, Telegram, or command-line) is handled by a single entry point: `runEmbeddedPiAgent()`. That function contains the core agent loop — understanding it explains how the whole system behaves in production.

```python theme={null}
runEmbeddedPiAgent()
```

The loop executes in five ordered phases. They always occur in sequence:

1. Setup — load session and prepare the environment
2. Context window validation — ensure the model has room to reason
3. Attempt loop — the perceive → reason → act cycle with tool execution
4. Error handling — structured recovery for different failure modes
5. Persistence — safely save results and auth state

Below is a concise reference for each phase, followed by implementation details and production considerations.

| Phase | Purpose | Key actions |
| - | - | - |
| Setup | Prepare session+tools+auth | Load session JSONL; resolve model, auth profile, system prompt, tool policy |
| Context validation | Ensure model has enough tokens | Calculate available tokens, compact if needed, abort with clear error if still insufficient |
| Attempt loop | Core perceive → reason → act cycle | Send messages + tools to model, execute tool calls in sandbox, append results, repeat until final text |
| Error handling | Handle failures deterministically | Per-error strategies: auth failover, backoff, compress, downgrade, or abort |
| Persistence | Durable session and auth tracking | Acquire write lock, save conversation and metadata, update auth profiling, release lock |

Phase 1 — Setup\
The agent loads the session history from disk. Sessions are stored in JSONL (one JSON object per line). From that session the agent resolves:

* which model to use,
* which auth profile to pull credentials from (with failover configured),
* the system prompt,
* and the list of available tools, determined by the tool policy for this request.

This prepares the messages and the tool set that will be sent to the model.

Phase 2 — Context window validation\
Before calling the model, the agent calculates available tokens: model maximum context minus space reserved for the response. If the session history is too large to fit, the agent runs a compaction step (described below) to free tokens. If there still isn't enough space after compaction, the agent aborts with a clear error rather than silently mangling the conversation.

<Callout icon="lightbulb" color="#1CB2FE">
  Reserve at least a buffer of tokens for the model's reply. Simultaneously reserving reply space and validating history avoids implicit truncation and preserves deterministic behavior.
</Callout>

Phase 3 — The attempt loop (perceive → reason → act)\
This is the core interaction with the model and tools:

* Send session messages, system prompt, and tool definitions to the model.
* If the model responds with tool calls, execute each tool in a sandbox, capture output, append the result to the message history, and send everything back to the model.
* Repeat until the model returns a final text response (no more tool calls), then exit.

Example pseudocode for the loop:

```python theme={null}
while True:
    response = llm.chat(
        messages=session_messages,
        system=system_prompt,
        tools=available_tools
    )

    if response.has_tool_calls:
        for tool_call in response.tool_calls:
            result = sandbox.execute(tool_call)
            session_messages.append(tool_result(result))
        # loop back - LLM decides next step

    else:  # final response - done
        return response.text
```

This implements perceive (tool outputs), reason (model call), and act (tool execution). OpenClaw surrounds this core cycle with production-grade controls described below.

Phase 4 — Error handling\
Each error class is handled intentionally and specifically:

* Auth errors: rotate to a different API key or provider automatically.
* Rate limits: trigger exponential backoff before retrying.
* Context overflow (if not caught earlier): compress the session and retry the request.
* Model errors: downgrade the reasoning level and retry at a lower setting.
* Timeouts: abort immediately with no retry.

Each error type has a clearly defined response rather than a generic retry policy.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/z7NmHsFQN9LCEiD0/images/AI-Agents-for-Beginners-OpenClaw-Case-Study/Production-OpenClaw/OpenClaw-Agent-Loop/phase-4-error-handling-infographic.jpg?fit=max&auto=format&n=z7NmHsFQN9LCEiD0&q=85&s=f3a4624f0e9954f9dec1b8d68f08cde8" alt="A retro-style infographic titled &#x22;PHASE 4: ERROR HANDLING&#x22; listing four error types—Auth Errors, Rate Limits, Context Overflow, and Model Errors—with brief remedies like rotating API keys, exponential backoff/retry, compressing sessions, and downgrading model. The colored, outlined boxes sit on a dark grid background." width="1920" height="1080" data-path="images/AI-Agents-for-Beginners-OpenClaw-Case-Study/Production-OpenClaw/OpenClaw-Agent-Loop/phase-4-error-handling-infographic.jpg" />
</Frame>

Phase 5 — Persistence\
When the agent finishes handling a message, it persists state safely:

* Acquire a write lock on the session file.
* Save the updated conversation and any session metadata.
* Update auth-profile tracking to record which keys succeeded and which failed.
* Release the lock.

The write lock prevents two simultaneous messages from corrupting the same session file.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/z7NmHsFQN9LCEiD0/images/AI-Agents-for-Beginners-OpenClaw-Case-Study/Production-OpenClaw/OpenClaw-Agent-Loop/phase-5-persistence-locking-infographic.jpg?fit=max&auto=format&n=z7NmHsFQN9LCEiD0&q=85&s=926a573cd0aabf02dd9565886efc0df1" alt="A retro-style infographic titled &#x22;PHASE 5: PERSISTENCE&#x22; lists four steps: 1) Acquire write lock, 2) Save updated conversation, 3) Update auth profile tracking, 4) Release lock. A red banner at the bottom says &#x22;LOCK PREVENTS CONCURRENT CORRUPTION.&#x22;" width="1920" height="1080" data-path="images/AI-Agents-for-Beginners-OpenClaw-Case-Study/Production-OpenClaw/OpenClaw-Agent-Loop/phase-5-persistence-locking-infographic.jpg" />
</Frame>

Context compaction (detailed)\
Long conversations can exhaust the model context window. OpenClaw avoids dropping messages arbitrarily — instead it compresses older parts of the session using summarization and trimming:

* Detect proximity to the context limit.
* Summarize older messages and trim verbose tool outputs where safe.
* Replace original messages with compressed representations in the session history.
* Continue processing with more available tokens.

This preserves key information while freeing tokens for fresh reasoning and tool calls.

<Callout icon="warning" color="#FF6B6B">
  Compaction is lossy by design: it preserves essential context and removes low-value verbosity. Ensure your downstream tools and prompts tolerate summarized history.
</Callout>

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/z7NmHsFQN9LCEiD0/images/AI-Agents-for-Beginners-OpenClaw-Case-Study/Production-OpenClaw/OpenClaw-Agent-Loop/context-compaction-compress-old-messages.jpg?fit=max&auto=format&n=z7NmHsFQN9LCEiD0&q=85&s=dcdab593acac503c263c5cb5f6dda284" alt="An infographic titled &#x22;Context Compaction&#x22; that explains compressing older messages when the context window fills up. It lists four numbered steps and banners reading &#x22;Never drops old messages — compresses instead&#x22; and &#x22;Preserves context · frees tokens.&#x22;" width="1920" height="1080" data-path="images/AI-Agents-for-Beginners-OpenClaw-Case-Study/Production-OpenClaw/OpenClaw-Agent-Loop/context-compaction-compress-old-messages.jpg" />
</Frame>

Session files and debugging\
Sessions are stored in JSONL specifically so you can open any session file in a text editor and read exactly what the agent did: every message, every tool call, and every result. Debugging a bad agent run often means reading the session file directly — no special tooling required.

Each session is keyed by channel and user, so the same user has different sessions on WhatsApp versus Telegram. Example session lines:

```json theme={null}
{"role":"user", "content":"list all running processes"}
{"role":"assistant", "tool_calls":[{"name":"bash", "input":"ps aux"}]}
{"role":"tool", "name":"bash", "output":"PID CMD ...\n1234 node"}
```

What makes this production-grade\
The core loop itself is simple — send to the model, check for tool calls, execute tools, repeat. That’s the minimal implementation. What makes OpenClaw production-grade is everything layered around that loop:

* Concurrency control (execution lanes and session locks) so multiple messages don't collide.
* Auth failover across multiple model providers so a single expired key doesn't kill the service.
* Context management that compresses instead of crashing.
* Structured recovery that handles each failure mode specifically instead of catching all exceptions generically.
* Comprehensive logging and session persistence for reliable debugging and auditing.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/z7NmHsFQN9LCEiD0/images/AI-Agents-for-Beginners-OpenClaw-Case-Study/Production-OpenClaw/OpenClaw-Agent-Loop/production-grade-four-box-infographic.jpg?fit=max&auto=format&n=z7NmHsFQN9LCEiD0&q=85&s=3e7c15d3ce394f1669ff11713e89d15c" alt="A retro-style infographic titled &#x22;PRODUCTION-GRADE&#x22; showing &#x22;WHAT'S AROUND IT:&#x22; with four colored boxes. The boxes list Concurrency Control, Auth Failover, Context Management, and Structured Recovery with brief descriptions beneath each." width="1920" height="1080" data-path="images/AI-Agents-for-Beginners-OpenClaw-Case-Study/Production-OpenClaw/OpenClaw-Agent-Loop/production-grade-four-box-infographic.jpg" />
</Frame>

Summary\
The core agent loop (send messages → model → tool execution → repeat) is intentionally small and deterministic. Production readiness is achieved by robust error handling, careful context management (compaction), concurrency controls (session locks, execution lanes), auth failover, and reliable persistence with readable JSONL session logs.

Links and References

* [JSON Lines (JSONL) format](https://jsonlines.org/) — why per-line objects are simple for append-only logs.
* [Exponential backoff best practices](https://cloud.google.com/storage/docs/exponential-backoff) — reference for retry/backoff strategy.
* [Concurrency patterns for distributed systems](https://martinfowler.com/articles/patterns-of-distributed-systems/) — design patterns for locks and coordination.

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/ai-agents-for-beginner-openclaw-case-study/module/b8b38b25-c4eb-425f-a093-cec426365977/lesson/2d359df3-e392-41af-b9ff-6681d29522b3" />

  <Card title="Practice Lab" icon="flask-conical" cta="Learn more" href="https://learn.kodekloud.com/user/courses/ai-agents-for-beginner-openclaw-case-study/module/b8b38b25-c4eb-425f-a093-cec426365977/lesson/60178424-a426-4c28-af0e-40f84cc40ffa" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.