Pattern 1 — Prompt chaining (sequential steps)
- Break a task into sequential steps; each LLM call consumes the previous step’s output.
- Insert programmatic checks between steps (length, schema, content safety) and decide whether to continue, retry, or abort.
- Typical use case: Draft an email (step 1) → validate required sections → polish the draft (step 2) → final review.

- Need deterministic, auditable transformations and clear retry semantics.
- Simple to implement and debug.
- Latency accumulates as steps are run serially.
- Validate intermediate outputs programmatically to avoid cascading failures.
- Use a classifier model to route requests to an appropriate handler pipeline: trivial Q&A vs. scheduling vs. complex planning.
- Enables cost optimization: route simple queries to cheap, low-latency models and hard tasks to larger, more capable models.

- Heterogeneous inputs where specialized handlers provide better quality/cost tradeoffs.
- Classifier precision is critical; add fallback or human-in-the-loop paths for uncertain classifications.
- Log routing decisions for auditing and to measure misroute cost.
- Sectioning: split input into independent subtasks and run them concurrently (e.g., “Plan my Thursday”: check calendar, email, weather).
- Voting: run the same prompt across different models or random seeds and aggregate via majority/consensus to reduce hallucination.

- Independent subtasks that don’t block each other.
- Need lower latency for aggregated results or higher reliability through redundancy.
- Parallel calls increase concurrency and cost.
- Aggregator logic must handle partial failures and inconsistent results.
- A central orchestrator inspects the input, decides which subtasks are needed at runtime, and delegates those to worker LLMs.
- Subtasks are dynamically determined, not predeclared.

- Inputs vary widely and benefit from a centralized coordinator that decides decomposition.
- Good for complex, open-ended tasks where fixed pipelines would be brittle.
- Orchestrator must be robust: enforce limits (max workers, cost caps), idempotency, and retry/backoff behavior.
- Instrumentation is essential to prevent runaway costs from large dynamic decompositions.
- Two distinct roles interact: a generator produces candidate outputs; an evaluator reviews and gives feedback. The generator refines outputs iteratively until acceptance criteria or retry limits are reached.
- Ideal for tasks demanding high correctness or nuanced judgment (legal text, code, translations).

- Outputs must meet strict quality thresholds and benefit from iterative critique.
- Define clear acceptance criteria to avoid infinite loops.
- Combine automated checks (unit tests, linters, schema validators) with model-based evaluation for stronger guarantees.
- Stop conditions: enforce max iterations, runtime limits, and quality thresholds.
- Intermediate validation: use schema validation, token/length checks, content-safety filters, and unit tests where applicable.
- Model selection: route low-latency/low-cost tasks to smaller models and reserve large models for tasks that need them.
- Observability: log inputs, outputs, routing decisions, retries, and evaluation scores for auditing, debugging, and cost analysis.
- Cost control: set per-request and per-workflow budgets, hard limits on spawned subtasks, and monitor for anomalies.
- Keep control logic deterministic and in code; use LLMs for uncertain or generative work only.
- Implement programmatic guards before and after LLM calls (schema checks, sanitization).
- Provide fallbacks for classifier or orchestrator uncertainty (human review, simpler model).
- Track and visualize workflow metrics: latency, cost per pattern, success rate, and misroute rates.
Design workflows so that control logic lives in your code (deterministic and auditable) while LLMs handle the uncertain parts (content generation, classification, synthesis). This separation makes behavior predictable and easier to test and monitor.
- Designing reliable LLM systems — patterns and tradeoffs (replace with your internal docs)
- Content safety and moderation best practices
- Instrumentation and observability for AI pipelines