Explains how explicit planning improves LLM agent performance, contrasts reactive and planned agents, and details guardrails and orchestration to prevent planning failures.
Zippy is a capable assistant for quick, reactive tasks, but he can struggle on deep research or multi-step workflows. He follows the ReAct pattern (Reason then Act) without explicitly planning.Savvy is a research specialist who also uses ReAct, but she inserts an explicit planning phase before taking actions. In this lesson we examine how planning improves agent performance and how Savvy performs it.An agent shouldn’t simply react — it should plan.When faced with a complex request like “book me a flight to NYC next Friday under $300,” an effective agent first reasons about the required steps and produces a plan, rather than invoking tools at random. For that flight example, a concise high-level plan might be:
Search for flights to NYC for next Friday.
Filter options under $300.
Check the user’s calendar for conflicts.
Book the best matching option.
This planning happens inside the LLM’s text generation: the model thinks through the task while composing its output. You can encourage explicit planning by instructing the agent to think step-by-step before acting — a practical application of chain-of-thought prompting for agents.
A useful system-level rule is to require the agent to produce its plan before any tool calls. This makes the plan inspectable by orchestrators and helps avoid random tool usage.
For example, a system instruction can require visible reasoning before each action:
System instruction:"Before taking any action, think through your plan step by step.Explain your reasoning before each tool call."
Given that instruction, the agent interleaves visible reasoning and tool calls. Example output might look like:
User request:"The user wants a flight to NYC under $300."Agent thought:"I'll search for flights first. If nothing is under budget, I'll check nearby dates or airports."Action:→ search_flights(NYC, $300)
When the agent runs the search and receives results, it continues the ReAct loop: think, act, observe, think again. This alternating pattern of reasoning and action is the hallmark of ReAct (Reasoning and Acting).Let’s trace the flight-booking example with a realistic sequence of thoughts, actions, and observations:
Agent action:search_flights(NYC, Fri, $300)Observation:4 results: $245@6:00, $280@14:30, $295@8:45, $299@17:20Agent thought:"The $245 option is very early; user may prefer afternoons. I'll check user travel preferences."Action:get_user_preferences('travel')Observation:User prefers afternoon flights, window seatsAgent thought:"There's a $280 flight at 14:30 that fits budget and time. I'll check the user's calendar for conflicts."Action:check_calendar('2026-03-14')Observation:No events after 11:00 — 14:30 is clearAgent thought:"Calendar is clear and preferences match. I'll book the $280 flight."Action:book_flight(option_id_for_$280)
Each thought step guides the next action; each observation supplies new information for reasoning. Planning turns a sequence of tool calls into a coherent workflow.Common planning failure modes
Circling: Repeating the same tool call because the agent doesn’t recognize it already has the information.
Over-planning: Spending too many tokens reasoning instead of taking actions.
Wrong assumptions: Misinterpreting the task and pursuing the wrong plan.
Not adapting: Sticking to the original plan even when observations suggest a new approach.
Good agent design anticipates and mitigates these failures. Practical guardrails include:
Maximum iterations: Cap loop cycles to prevent infinite or endless retry loops (for example, stop after 10 iterations).
Token budgets: Limit how many tokens the agent can spend on internal reasoning for a single task.
Explicit instructions: In the system prompt, tell the agent how to behave when specific failures happen (e.g., if a search returns no results, broaden the query).
Structured output: Require the agent to emit a plan in a fixed format so orchestrators and validators can inspect it before actions are taken.
A well-designed agent loop implements a perceive → reason → act cycle and enforces these guardrails to improve reliability.Below is a quick reference table that pairs common failure modes with recommended mitigations:
Failure mode
Symptom
Recommended mitigation
Circling
Repeated identical tool calls
Add state checks and idempotency detection; track tool results in memory
Over-planning
Excessive reasoning tokens, no progress
Enforce token budgets and a time-to-action threshold
Wrong assumptions
Plan diverges from user intent
Require plan confirmation or user clarification step
Not adapting
Ignores new observations
Allow plan revision and include observation-triggered replanning
The orchestration layer enforces guardrails, manages maximum iterations, compresses context when the token window fills, and falls back to alternate models if a call fails. These mechanisms make planning less brittle and more recoverable. Typical orchestration responsibilities:
Enforce iteration caps and token budgets.
Compress or summarize long context to fit model windows.
Detect tool failures and retry with fallback strategies.
Validate structured plans before executing sensitive actions.
Key takeaways
Agents plan; they do not merely react.
Planning occurs inside the LLM’s internal reasoning and is best made explicit for inspection.
Chain-of-thought prompting encourages step-by-step reasoning before actions.
ReAct is a practical pattern: think → act → observe → repeat.
Typical failures (circling, over-planning, wrong assumptions, not adapting) are mitigated with guardrails: iteration caps, token budgets, explicit instructions, and structured output.
A robust planning system is the foundation of any reliable agent. Zippy and Savvy are complementary: Zippy handles fast reactive tasks, while Savvy handles research-heavy tasks requiring structured planning. Together they cover a wider range of user needs than either could alone.Links and references