Strategies to improve agent performance and reduce LLM cost and latency through context reduction, fewer iterations, model routing, parallel tool execution, caching, and scalable session and retry patterns
Agents are powerful, but without careful design they can become slow and costly. Every LLM call consumes tokens (input + output), each tool call adds latency, and repeated loop iterations multiply both. Left unchecked, a single agent task can burn dollars and take minutes — whereas production systems should typically cost cents and finish in seconds.There are three primary cost dimensions to manage:
Cost type
What it is
Practical mitigations
Token cost
Tokens sent to and returned from the model
Summarize history, trim tool output, remove irrelevant system prompt text
Latency
Time for LLM calls + tool calls (commonly 1–10s per call)
Parallelize independent tool calls, reduce iterations, use faster models
Compute
CPU/GPU usage for embeddings, memory search, compression
Cache embeddings, route heavy tasks to appropriate hardware/models
The dominant driver in most systems is context size: every included token increases both cost and latency. The general solution is systematic context reduction: summarize older messages, trim verbose tool outputs, and exclude irrelevant system prompt parts.
A well-engineered context strategy can cut token costs dramatically — for example, reducing context from 50,000 tokens to 15,000 tokens yields a ~70% token-cost reduction while preserving agent capability.Levers you can apply (with examples and patterns):Lever 1 — Context compression and tool-output trimming
Keep recent turns verbatim and summarize older conversation.
Trim extremely verbose tool outputs; most agent decisions only need the first N characters or a short digest.
Exclude system prompt fragments that aren’t relevant to the current task.
Example: compress older context into a single summary while preserving the last few turns.
// agent/context.tsexport async function compressContext(messages: Message[]): Promise<Message[]> { // Keep the last 5 turns in full, summarize older history const recent = messages.slice(-5); const older = messages.slice(0, -5); if (older.length === 0) return recent; const summary = await summarize(older); // summarize is async // Return a compact context: one summary plus recent turns return [summary, ...recent];}
Example: trim verbose tool output before including it in the context.
// agent/tools.tsexport function trimToolOutput(output: string, maxChars = 500): string { // Most agents rarely need more than the first N characters of verbose outputs if (output.length > maxChars) { return output.slice(0, maxChars) + "...[truncated]"; } return output;}
Levers 2 — Reduce loop iterations
Each iteration of an agent control loop performs an LLM call plus any triggered tool calls. Reducing the number of iterations lowers both latency and token cost.
Use stronger planning prompts so the agent decides a plan before making tool calls.
Improve tool quality so each call returns richer, action-ready data.
Enforce a maximum iteration limit to avoid runaway loops.
Always cap iterations for safety. Without a limit, a confused agent can loop until you hit provider rate limits or exhaust your budget.
Levers 3 — Model routing (right model for the job)
Not every step needs the most capable (and expensive) model. Route tasks to the appropriate model tier:
Fast, cheap models for classification and routing.
Mid-tier models for routine text generation.
High-capability models for complex reasoning or code generation.
Example routing function:
function routeToModel(taskType: string): string { switch (taskType) { case "classify": return "gpt-4o-mini"; // cheap + fast case "generate": return "gpt-4o"; // mid-tier case "reason": return "claude-opus-4"; // most capable default: return "gpt-4o"; }}
Routing reduces cost while preserving quality where it matters.Levers 4 — Parallel tool execution
Run independent tool calls concurrently instead of sequentially to save wall-clock time.
OpenClaw (an execution engine) supports concurrent execution: when the model requests multiple tool calls in one turn, the engine executes them concurrently.Levers 5 — Caching
Cache repeatable, expensive operations such as embeddings and API responses. This converts slow, costly work into fast, cheap cache hits.
const embeddingCache = new Map<string, number[]>();export async function getEmbedding(text: string): Promise<number[]> { if (embeddingCache.has(text)) { return embeddingCache.get(text)!; // fast cache hit } const embedding = await embed(text); // expensive call embeddingCache.set(text, embedding); return embedding;}
The first query costs time and resources; subsequent queries are fast.Scaling considerations
When many users interact concurrently, adopt patterns to avoid collisions and to scale throughput:
Execution lanes / session isolation: give each session its own lane or execution context to prevent shared-state corruption.
Rate limits and provider distribution: implement exponential backoff and retry and consider distributing traffic across multiple providers (a hybrid provider setup) to improve reliability.
Session size management: compress and summarize older turns automatically when a session exceeds a configured size threshold.
Backoff + retry example (handles rate limits gracefully with exponential delay):
async function callWithRetry<T>(fn: () => Promise<T>, maxRetries = 3): Promise<T> { for (let attempt = 0; attempt < maxRetries; attempt++) { try { return await fn(); } catch (err: any) { // Non-rate-limit errors or final attempt should be rethrown const isRateLimit = err?.status === 429 || err?.code === "RateLimit"; if (!isRateLimit || attempt === maxRetries - 1) throw err; const delayMs = 1000 * 2 ** attempt; // 1s, 2s, 4s await sleep(delayMs); } } // Should never reach here throw new Error("Unreachable: retries exhausted");}
Optimization hierarchy
Follow this order when optimizing — it yields the highest return on investment:
Reduce unnecessary LLM calls (biggest single impact).
Reduce context size per call.
Route tasks to cheaper/faster models where appropriate.
Parallelize independent tool execution.
Cache repeated operations.
Build a working agent first, measure where time and money are actually being spent, then optimize the real bottlenecks.
Build a correct agent first. Measure where time and money are actually being spent, then apply the optimizations above to the real bottlenecks.