> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Performance and Scaling

> Strategies to improve agent performance and reduce LLM cost and latency through context reduction, fewer iterations, model routing, parallel tool execution, caching, and scalable session and retry patterns

Agents are powerful, but without careful design they can become slow and costly. Every LLM call consumes tokens (input + output), each tool call adds latency, and repeated loop iterations multiply both. Left unchecked, a single agent task can burn dollars and take minutes — whereas production systems should typically cost cents and finish in seconds.

There are three primary cost dimensions to manage:

| Cost type | What it is | Practical mitigations |
| - | - | - |
| Token cost | Tokens sent to and returned from the model | Summarize history, trim tool output, remove irrelevant system prompt text |
| Latency | Time for LLM calls + tool calls (commonly 1–10s per call) | Parallelize independent tool calls, reduce iterations, use faster models |
| Compute | CPU/GPU usage for embeddings, memory search, compression | Cache embeddings, route heavy tasks to appropriate hardware/models |

The dominant driver in most systems is context size: every included token increases both cost and latency. The general solution is systematic context reduction: summarize older messages, trim verbose tool outputs, and exclude irrelevant system prompt parts.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/z7NmHsFQN9LCEiD0/images/AI-Agents-for-Beginners-OpenClaw-Case-Study/Production-OpenClaw/Performance-and-Scaling/lever1-context-size-reduction-250k-75k.jpg?fit=max&auto=format&n=z7NmHsFQN9LCEiD0&q=85&s=09889d8c6ed71272b2f4647c85d1039b" alt="A retro-style infographic titled &#x22;Lever 1: Context Size&#x22; lists four steps to reduce context (summarize older messages, trim verbose tool outputs, use shorter tool descriptions, exclude irrelevant system prompt parts). On the right it shows the math comparing &#x22;Before&#x22; 250,000 tokens to &#x22;After&#x22; 75,000 tokens." width="1920" height="1080" data-path="images/AI-Agents-for-Beginners-OpenClaw-Case-Study/Production-OpenClaw/Performance-and-Scaling/lever1-context-size-reduction-250k-75k.jpg" />
</Frame>

A well-engineered context strategy can cut token costs dramatically — for example, reducing context from 50,000 tokens to 15,000 tokens yields a \~70% token-cost reduction while preserving agent capability.

Levers you can apply (with examples and patterns):

Lever 1 — Context compression and tool-output trimming

* Keep recent turns verbatim and summarize older conversation.
* Trim extremely verbose tool outputs; most agent decisions only need the first N characters or a short digest.
* Exclude system prompt fragments that aren’t relevant to the current task.

Example: compress older context into a single summary while preserving the last few turns.

```ts theme={null}
// agent/context.ts
export async function compressContext(messages: Message[]): Promise<Message[]> {
  // Keep the last 5 turns in full, summarize older history
  const recent = messages.slice(-5);
  const older = messages.slice(0, -5);

  if (older.length === 0) return recent;

  const summary = await summarize(older); // summarize is async
  // Return a compact context: one summary plus recent turns
  return [summary, ...recent];
}
```

Example: trim verbose tool output before including it in the context.

```ts theme={null}
// agent/tools.ts
export function trimToolOutput(output: string, maxChars = 500): string {
  // Most agents rarely need more than the first N characters of verbose outputs
  if (output.length > maxChars) {
    return output.slice(0, maxChars) + "...[truncated]";
  }
  return output;
}
```

Levers 2 — Reduce loop iterations
Each iteration of an agent control loop performs an LLM call plus any triggered tool calls. Reducing the number of iterations lowers both latency and token cost.

* Use stronger planning prompts so the agent decides a plan before making tool calls.
* Improve tool quality so each call returns richer, action-ready data.
* Enforce a maximum iteration limit to avoid runaway loops.

<Callout icon="warning" color="#FF6B6B">
  Always cap iterations for safety. Without a limit, a confused agent can loop until you hit provider rate limits or exhaust your budget.
</Callout>

Example iteration guard:

```ts theme={null}
const MAX_ITERATIONS = 20;
let iteration = 0;
let done = false;

while (!done && iteration < MAX_ITERATIONS) {
  iteration++;
  // agent planning, tool calls, and state updates...
}
```

Levers 3 — Model routing (right model for the job)
Not every step needs the most capable (and expensive) model. Route tasks to the appropriate model tier:

* Fast, cheap models for classification and routing.
* Mid-tier models for routine text generation.
* High-capability models for complex reasoning or code generation.

Example routing function:

```ts theme={null}
function routeToModel(taskType: string): string {
  switch (taskType) {
    case "classify":
      return "gpt-4o-mini";  // cheap + fast
    case "generate":
      return "gpt-4o";        // mid-tier
    case "reason":
      return "claude-opus-4"; // most capable
    default:
      return "gpt-4o";
  }
}
```

Routing reduces cost while preserving quality where it matters.

Levers 4 — Parallel tool execution
Run independent tool calls concurrently instead of sequentially to save wall-clock time.

```ts theme={null}
// Sequential - ~6s total (2s + 2s + 2s)
const docs = await fetchDocs(query);       
const file = await readFile(path);         
const results = await search(term);        

// Parallel - ~2s total
const [docs, file, results] = await Promise.all([
  fetchDocs(query),
  readFile(path),
  search(term),
]);
```

OpenClaw (an execution engine) supports concurrent execution: when the model requests multiple tool calls in one turn, the engine executes them concurrently.

Levers 5 — Caching
Cache repeatable, expensive operations such as embeddings and API responses. This converts slow, costly work into fast, cheap cache hits.

```ts theme={null}
const embeddingCache = new Map<string, number[]>();

export async function getEmbedding(text: string): Promise<number[]> {
  if (embeddingCache.has(text)) {
    return embeddingCache.get(text)!; // fast cache hit
  }

  const embedding = await embed(text); // expensive call
  embeddingCache.set(text, embedding);
  return embedding;
}
```

The first query costs time and resources; subsequent queries are fast.

Scaling considerations
When many users interact concurrently, adopt patterns to avoid collisions and to scale throughput:

* Execution lanes / session isolation: give each session its own lane or execution context to prevent shared-state corruption.
* Rate limits and provider distribution: implement exponential backoff and retry and consider distributing traffic across multiple providers (a hybrid provider setup) to improve reliability.
* Session size management: compress and summarize older turns automatically when a session exceeds a configured size threshold.

Backoff + retry example (handles rate limits gracefully with exponential delay):

```ts theme={null}
async function callWithRetry<T>(fn: () => Promise<T>, maxRetries = 3): Promise<T> {
  for (let attempt = 0; attempt < maxRetries; attempt++) {
    try {
      return await fn();
    } catch (err: any) {
      // Non-rate-limit errors or final attempt should be rethrown
      const isRateLimit = err?.status === 429 || err?.code === "RateLimit";
      if (!isRateLimit || attempt === maxRetries - 1) throw err;

      const delayMs = 1000 * 2 ** attempt; // 1s, 2s, 4s
      await sleep(delayMs);
    }
  }
  // Should never reach here
  throw new Error("Unreachable: retries exhausted");
}
```

Optimization hierarchy
Follow this order when optimizing — it yields the highest return on investment:

1. Reduce unnecessary LLM calls (biggest single impact).
2. Reduce context size per call.
3. Route tasks to cheaper/faster models where appropriate.
4. Parallelize independent tool execution.
5. Cache repeated operations.

Build a working agent first, measure where time and money are actually being spent, then optimize the real bottlenecks.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/z7NmHsFQN9LCEiD0/images/AI-Agents-for-Beginners-OpenClaw-Case-Study/Production-OpenClaw/Performance-and-Scaling/optimization-hierarchy-llm-build-measure-optimize.jpg?fit=max&auto=format&n=z7NmHsFQN9LCEiD0&q=85&s=0c3b903d4bdbed91e2b73a2fdfb05b37" alt="An infographic titled &#x22;Optimization Hierarchy&#x22; listing five ranked optimization strategies for LLM systems. It highlights reducing unnecessary LLM calls, shrinking context size, routing tasks to cheaper models, parallelizing tool execution, and caching repeated operations, plus a golden rule: build → measure → optimize bottlenecks." width="1920" height="1080" data-path="images/AI-Agents-for-Beginners-OpenClaw-Case-Study/Production-OpenClaw/Performance-and-Scaling/optimization-hierarchy-llm-build-measure-optimize.jpg" />
</Frame>

<Callout icon="lightbulb" color="#1CB2FE">
  Build a correct agent first. Measure where time and money are actually being spent, then apply the optimizations above to the real bottlenecks.
</Callout>

References and further reading

* [Rate limiting and retries patterns](https://en.wikipedia.org/wiki/Rate_limiting)
* [Designing for latency and throughput in distributed systems](https://en.wikipedia.org/wiki/Latency_\(engineering\))
* Consider profiling actual token usage and latency in production before making large changes — real telemetry beats guesswork.

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/ai-agents-for-beginner-openclaw-case-study/module/b8b38b25-c4eb-425f-a093-cec426365977/lesson/37eb3e85-00fd-4917-a523-eb2aff8640cf" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.