Skip to main content
Zippy is a capable assistant for quick, reactive tasks, but he can struggle on deep research or multi-step workflows. He follows the ReAct pattern (Reason then Act) without explicitly planning. Savvy is a research specialist who also uses ReAct, but she inserts an explicit planning phase before taking actions. In this lesson we examine how planning improves agent performance and how Savvy performs it. An agent shouldn’t simply react — it should plan. When faced with a complex request like “book me a flight to NYC next Friday under $300,” an effective agent first reasons about the required steps and produces a plan, rather than invoking tools at random. For that flight example, a concise high-level plan might be:
  1. Search for flights to NYC for next Friday.
  2. Filter options under $300.
  3. Check the user’s calendar for conflicts.
  4. Book the best matching option.
This planning happens inside the LLM’s text generation: the model thinks through the task while composing its output. You can encourage explicit planning by instructing the agent to think step-by-step before acting — a practical application of chain-of-thought prompting for agents.
A useful system-level rule is to require the agent to produce its plan before any tool calls. This makes the plan inspectable by orchestrators and helps avoid random tool usage.
For example, a system instruction can require visible reasoning before each action:
Given that instruction, the agent interleaves visible reasoning and tool calls. Example output might look like:
When the agent runs the search and receives results, it continues the ReAct loop: think, act, observe, think again. This alternating pattern of reasoning and action is the hallmark of ReAct (Reasoning and Acting). Let’s trace the flight-booking example with a realistic sequence of thoughts, actions, and observations:
Each thought step guides the next action; each observation supplies new information for reasoning. Planning turns a sequence of tool calls into a coherent workflow. Common planning failure modes
  • Circling: Repeating the same tool call because the agent doesn’t recognize it already has the information.
  • Over-planning: Spending too many tokens reasoning instead of taking actions.
  • Wrong assumptions: Misinterpreting the task and pursuing the wrong plan.
  • Not adapting: Sticking to the original plan even when observations suggest a new approach.
A poster titled "PLANNING FAILURES" showing four common failure modes in boxed panels. The boxes read: "CIRCLES" (repeats same tool call), "OVER-PLAN" (too many thinking tokens), "WRONG ASSUMPTIONS" (misread the task), and "NOT ADAPTING" (ignores tool results).
Good agent design anticipates and mitigates these failures. Practical guardrails include:
  • Maximum iterations: Cap loop cycles to prevent infinite or endless retry loops (for example, stop after 10 iterations).
  • Token budgets: Limit how many tokens the agent can spend on internal reasoning for a single task.
  • Explicit instructions: In the system prompt, tell the agent how to behave when specific failures happen (e.g., if a search returns no results, broaden the query).
  • Structured output: Require the agent to emit a plan in a fixed format so orchestrators and validators can inspect it before actions are taken.
A well-designed agent loop implements a perceive → reason → act cycle and enforces these guardrails to improve reliability. Below is a quick reference table that pairs common failure modes with recommended mitigations:
A neon-style infographic titled "GUARDRAILS — Four layers of protection" showing four labeled boxes: "Max Iterations," "Token Budget," "Explicit Instructions," and "Structured Output." Each box contains brief guidance like capping loop cycles, limiting reasoning tokens, telling the agent how to fail, and requiring a plan format.
The orchestration layer enforces guardrails, manages maximum iterations, compresses context when the token window fills, and falls back to alternate models if a call fails. These mechanisms make planning less brittle and more recoverable. Typical orchestration responsibilities:
  • Enforce iteration caps and token budgets.
  • Compress or summarize long context to fit model windows.
  • Detect tool failures and retry with fallback strategies.
  • Validate structured plans before executing sensitive actions.
A retro neon infographic titled "OPENCLAW PLANNING" showing a three-step loop: Perceive → Reason → Act. Below it are three supporting boxes labeled Max Iterations, Context Compression, and Model Fallback.
Key takeaways
  • Agents plan; they do not merely react.
  • Planning occurs inside the LLM’s internal reasoning and is best made explicit for inspection.
  • Chain-of-thought prompting encourages step-by-step reasoning before actions.
  • ReAct is a practical pattern: think → act → observe → repeat.
  • Typical failures (circling, over-planning, wrong assumptions, not adapting) are mitigated with guardrails: iteration caps, token budgets, explicit instructions, and structured output.
A dark presentation slide titled "KEY TAKEAWAYS" with five colored boxes summarizing points about agents: "AGENTS PLAN" (Not just react), "CHAIN‑OF‑THOUGHT" (Think before acting), "REACT" (Think → Act → Observe loop), "FAILURES" (Circles · Over‑plan · Wrong · Rigid) and "GUARDRAILS" (Iterations · Tokens · Instructions · Format).
A robust planning system is the foundation of any reliable agent. Zippy and Savvy are complementary: Zippy handles fast reactive tasks, while Savvy handles research-heavy tasks requiring structured planning. Together they cover a wider range of user needs than either could alone. Links and references

Watch Video