Skip to main content
Agents are powerful, but without careful design they can become slow and costly. Every LLM call consumes tokens (input + output), each tool call adds latency, and repeated loop iterations multiply both. Left unchecked, a single agent task can burn dollars and take minutes — whereas production systems should typically cost cents and finish in seconds. There are three primary cost dimensions to manage: The dominant driver in most systems is context size: every included token increases both cost and latency. The general solution is systematic context reduction: summarize older messages, trim verbose tool outputs, and exclude irrelevant system prompt parts.
A retro-style infographic titled "Lever 1: Context Size" lists four steps to reduce context (summarize older messages, trim verbose tool outputs, use shorter tool descriptions, exclude irrelevant system prompt parts). On the right it shows the math comparing "Before" 250,000 tokens to "After" 75,000 tokens.
A well-engineered context strategy can cut token costs dramatically — for example, reducing context from 50,000 tokens to 15,000 tokens yields a ~70% token-cost reduction while preserving agent capability. Levers you can apply (with examples and patterns): Lever 1 — Context compression and tool-output trimming
  • Keep recent turns verbatim and summarize older conversation.
  • Trim extremely verbose tool outputs; most agent decisions only need the first N characters or a short digest.
  • Exclude system prompt fragments that aren’t relevant to the current task.
Example: compress older context into a single summary while preserving the last few turns.
Example: trim verbose tool output before including it in the context.
Levers 2 — Reduce loop iterations Each iteration of an agent control loop performs an LLM call plus any triggered tool calls. Reducing the number of iterations lowers both latency and token cost.
  • Use stronger planning prompts so the agent decides a plan before making tool calls.
  • Improve tool quality so each call returns richer, action-ready data.
  • Enforce a maximum iteration limit to avoid runaway loops.
Always cap iterations for safety. Without a limit, a confused agent can loop until you hit provider rate limits or exhaust your budget.
Example iteration guard:
Levers 3 — Model routing (right model for the job) Not every step needs the most capable (and expensive) model. Route tasks to the appropriate model tier:
  • Fast, cheap models for classification and routing.
  • Mid-tier models for routine text generation.
  • High-capability models for complex reasoning or code generation.
Example routing function:
Routing reduces cost while preserving quality where it matters. Levers 4 — Parallel tool execution Run independent tool calls concurrently instead of sequentially to save wall-clock time.
OpenClaw (an execution engine) supports concurrent execution: when the model requests multiple tool calls in one turn, the engine executes them concurrently. Levers 5 — Caching Cache repeatable, expensive operations such as embeddings and API responses. This converts slow, costly work into fast, cheap cache hits.
The first query costs time and resources; subsequent queries are fast. Scaling considerations When many users interact concurrently, adopt patterns to avoid collisions and to scale throughput:
  • Execution lanes / session isolation: give each session its own lane or execution context to prevent shared-state corruption.
  • Rate limits and provider distribution: implement exponential backoff and retry and consider distributing traffic across multiple providers (a hybrid provider setup) to improve reliability.
  • Session size management: compress and summarize older turns automatically when a session exceeds a configured size threshold.
Backoff + retry example (handles rate limits gracefully with exponential delay):
Optimization hierarchy Follow this order when optimizing — it yields the highest return on investment:
  1. Reduce unnecessary LLM calls (biggest single impact).
  2. Reduce context size per call.
  3. Route tasks to cheaper/faster models where appropriate.
  4. Parallelize independent tool execution.
  5. Cache repeated operations.
Build a working agent first, measure where time and money are actually being spent, then optimize the real bottlenecks.
An infographic titled "Optimization Hierarchy" listing five ranked optimization strategies for LLM systems. It highlights reducing unnecessary LLM calls, shrinking context size, routing tasks to cheaper models, parallelizing tool execution, and caching repeated operations, plus a golden rule: build → measure → optimize bottlenecks.
Build a correct agent first. Measure where time and money are actually being spent, then apply the optimizations above to the real bottlenecks.
References and further reading

Watch Video