Overview

Prompt design is the loop of writing, testing, and tightening the instructions, examples, and context an LLM receives. The rules below hold across Claude, GPT, and Gemini; the vendor guides (listed under Sources) agree on them. Per-model quirks live in reasoning-model-prompting and prompt-migration.

Be specific and give the reason

State the task, the audience, the output, and the constraints in the words a domain expert would use. Test the prompt on a colleague with no context; if they would be confused, the model will be too.

Bad:  What do you know about coding?
Good: Summarize framework options for a static documentation site
      in TypeScript that ships under 200 KB of JavaScript.

Explain why a rule exists. “Your reply is read aloud by a text-to-speech engine, so do not use ellipses” generalizes better than “NEVER use ellipses” because the model can infer related cases.

Put durable rules in the system prompt and the task in the user message

Role, voice, tool list, output schema, and safety policy go in the system prompt; the specific question and its data go in the user message. See system-prompts for structure and versioning, and role-framing for the role line.

Separate instructions, data, and examples with consistent tags

Wrap each kind of content in its own tag (<instructions>, <context>, <example>, <input>) and use the same names in every prompt. Anthropic recommends XML tags for Claude; OpenAI documents Markdown headings and XML tags; either works if applied consistently. Templates are in prompt-templates.

Show examples when the format is hard to describe

Use three to five relevant, diverse examples in <example> tags when a format, tone, or category boundary resists description. Skip them when one sentence states the rule. Selection and ordering are in few-shot.

Say what to do, and state the non-goals

Prefer “Write smoothly flowing prose paragraphs” over “Do not use markdown”. Name the exclusions that matter (“Do not refactor unrelated code”, “Call only search and read”) because a prompt without non-goals invites helpful overreach. Give the model an escape hatch when the answer may be absent:

Reply NOT_FOUND if the answer is not in the provided text.

Without it, grounded Q&A pushes the model to invent an answer.

Put long documents first and the question last

For inputs over about 20k tokens, place documents above the instructions and query, and ask the model to quote the relevant passages before answering. Anthropic reports up to 30 percent better quality in tests when the query comes last. Budgeting and retrieval are in context-engineering.

Match emphasis to the model

Drop ALL CAPS, “CRITICAL”, and “you MUST” on current models. Anthropic notes that Opus 4.5 and 4.6 respond more strongly to the system prompt, so wording written to prevent undertriggering now causes overtriggering; write “Use this tool when…” instead of “CRITICAL: You MUST use this tool when…“. OpenAI’s GPT-5 guide reports that emphatic “be THOROUGH” prompts made the model overuse tools, and that contradictory instructions waste reasoning tokens. Remove conflicts instead of shouting over them.

Do not tune sampling parameters

Control behavior with the prompt and, on reasoning models, an effort setting. Claude Opus 4.7 and later, Sonnet 5.x, and Fable 5.x accept only the default temperature and top_p values and reject others, GPT-6 Astra does not support custom values, and Google recommends the defaults on Gemini 3. Details and exact model lists are in reasoning-model-prompting.

Test every change against an eval set

Change one variable, re-run a fixed set of cases, and keep prompts in version control. A prompt without evals is a guess. See prompt-evals. Split a prompt that does more than one job into a chain; see prompt-chaining.

Sources

  • Anthropic, “Prompting best practices” (platform.claude.com/docs).
  • OpenAI, “Prompt engineering” guide and GPT-5 prompting guide (developers.openai.com).
  • Google, “Prompt design strategies” (ai.google.dev).