Skip to main content
Prompt engineering starts with specificity, but that’s just the beginning. Large language models (LLMs) don’t read prompts like humans — they process sequences of tokens and predict the next token from a probability distribution shaped by training data plus the entire prompt. The words you choose, their order, and how you frame instructions shift those probabilities in measurable ways. Below are six research-backed techniques OpenClaw uses in production, with practical guidance, examples, and the exact patterns we enforce in system prompts and agent builders.
Prompt engineering is production software: treat prompts as structured, testable artifacts. Build a prompt generator with parameters, ordering, and conditional sections so you can A/B test and iterate.

1) Persona assignment

Assigning a persona (e.g., “You are a senior Python engineer”) does more than set tone — it shifts the model toward the statistical cluster of text associated with that role. This affects vocabulary, reasoning patterns, and priorities. Activation-patching research shows persona effects concentrate in early and some middle attention layers; you can even extract and inject role vectors at inference time to mimic this shift.
A stylized slide titled "#1 PERSONA ASSIGNMENT" with the prompt "You are a senior Python engineer." and notes about activation patching and a "role vector." The image shows neon-bordered boxes listing what the persona activates (vocabulary, reasoning patterns, priorities) and layer-level details for a 2025 study plus an extract/inject inference note.
When to use persona prompting:
  • Use it for domain-specific or stylistic tasks (code review, legal tone, editorial voice).
  • Avoid or carefully test it for abstract reasoning and arithmetic — it can sometimes degrade accuracy by introducing unwarranted assumptions.
Guidance:
  • A/B test persona vs. no-persona on representative benchmarks.
  • For production agents, layer personas: a stable base identity, specialized skill instructions, and optional narrow role reframings when needed.
OpenClaw layered persona example:
  • Base identity: You are a personal assistant running inside OpenClaw.
  • Skills: inject tool- and task-specific instructions.
  • Extensions: occasionally reframe identity (e.g., You are the OpenClaw VM.) to tighten behavior.

2) Positive over negative instructions

Negations are cognitively tricky for LLMs. Compare:
  • “Don’t use bullet points.”
  • “Write in paragraphs only.”
Although semantically similar, the positive form reliably performs better because the model need not first activate and then suppress patterns.
A neon-styled slide titled "#2 POSITIVE INSTRUCTIONS" explaining how "don't" prompts fail and offering positive reformulations like "Write in paragraphs only" and "One sentence per point." It also cites a 2023 ACL benchmark noting larger models do worse on negation than smaller ones.
How OpenClaw applies this:
  • Phrase preferences positively for general behavior.
  • Reserve explicit negations for absolute constraints (safety/legal rules).
  • Keep hard constraints separate from regular guidance so they’re treated as invariants.
Example system-tuned constraints and narration guidance:
Note: Do not append internal constraint text verbatim to user-facing replies.

3) Order and structure

Position matters. Empirical work documents a U-shaped attention pattern: models attend more to the beginning and end of a prompt than to the middle. Place critical instructions at the start or end, and avoid burying them in the middle of long contexts.
An infographic titled "#3 Order and Structure" illustrating a U-shaped attention pattern—75% attention at the beginning and end and 45% in the middle. It also notes prompt position affects attention weight, with top and bottom strongest and the middle weakest (citing Liu et al., Stanford 2024).
Practical ordering rules:
  • Put critical constraints and identity first.
  • Place the specific user question last.
  • Insert large context files in the middle.
  • For long prompts, repeat constraints at both the beginning and the end.
OpenClaw system-prompt assembly order (high level):

4) Chain-of-thought (structure the reasoning)

Asking a model to “think step-by-step” can improve performance because intermediate tokens condition later tokens. However, free-form chain-of-thought may introduce biases or reduce accuracy in some settings. OpenClaw avoids unstructured internal reasoning in outputs and instead enforces explicit, auditable reasoning tags.
An infographic titled "#4 Chain-of-Thought by Structure" that explains how intermediate tokens condition future tokens and notes a Turpin et al. 2023 finding about accuracy drop. A side panel rates chain-of-thought value by model type (High for small/older models, Moderate for GPT-4 class, Near Zero for o1/extended thinking).
Structured internal reasoning helps auditability, but free-form chains can leak internal deliberation or bias final outputs. Use structured tags and separate final answers from internal thoughts.
OpenClaw’s structural pattern:
  • Require all internal reasoning inside explicit tags: use <think>...</think>.
  • Disallow analysis outside those tags.
  • Always produce a <final>...</final> section for the public reply.
Configuration example (enforced for models that benefit):
When documenting these tags in MDX/JSX, wrap them in code blocks or backticks to avoid parsing issues. For example:
This separation improves downstream parsing, auditing, and tool integration.

5) “Before replying” patterns (structured lookup)

A powerful pattern is to require mandatory pre-answer steps: memory lookups, skills scans, and selection of the most specific applicable skill. Separating retrieval from generation reduces hallucination and increases factual correctness. Example memory and skills rules:
This “step-back” prompting enforces structured lookup and selection before any generation.

6) Few-shot examples in system prompts

Few-shot examples are effective inside system prompts too. Concrete correct vs. incorrect examples reduce ambiguity and guide downstream parsers and tools to expect precise output formats. Examples OpenClaw uses:
Always include both positive and negative examples where format precision matters. This is especially important when external services parse agent output.

Bonus techniques

  • Emotional framing: experiments (e.g., Microsoft Research) have shown that indicating high stakes can sometimes improve output quality on complex generation tasks. Use with care and A/B test.
  • Rereading (RE2): repeating the user question at the end of the prompt gives the model a second pass and can improve performance for decoder-only models.
A retro-styled slide titled "BONUS TECHNIQUES" with a highlighted "EMOTIONAL FRAMING" panel from Microsoft Research 2023. The panel shows a quote ("This is very important to my career.") and metrics (+8% standard benchmarks, +115% complex generation) about emotional urgency improving responses.

How OpenClaw assembles prompts in production

OpenClaw generates its system prompt with a parameterized prompt builder (a 646-line function in production) that assembles conditional sections. Identity, tooling, and safety are always included; skills, memory, context files, user identity, and timezone are conditional; runtime info is appended last. Concise prompt-builder example:
Benefits of a builder approach:
  • Enforce ordering and repetition consistently.
  • Toggle persona, reasoning structure, and pre-answer checks per agent or model.
  • Insert precise few-shot examples in the exact place where format matters.

Quick reference table


Summary

The most effective prompt engineering treats prompts as software: structured, testable, and parameterized. Apply personas selectively, prefer positive instructions, order content deliberately, isolate internal reasoning, require pre-answer lookups, and include concrete examples in system prompts. Encode these rules in a prompt builder to ensure consistent, auditable agent behavior across environments.

Watch Video

Practice Lab