Overview

A prompt chain is a sequence of narrow LLM calls where each step has one job and passes a small, typed payload to the next. Chain when you need to inspect intermediate outputs, branch on them, or enforce a fixed pipeline. Do not chain to get reasoning: current models with thinking handle most multi-step reasoning inside one call (Anthropic’s guidance), so a chain that exists only to “think in steps” adds latency and cost.

Chain when one prompt would do more than one thing

Split when the instructions list several verbs (extract and classify and summarize), the output schema glues unrelated fields together, or a fix for one step keeps breaking the others. One prompt per verb gives each step its own template and its own cases in the eval set. A two-step chain where one step would do doubles latency and cost for no gain.

extract -> classify -> respond
  1. Extract entities from the message. Output: JSON list.
  2. Classify each entity's intent. Output: JSON list of {entity, intent}.
  3. Draft a reply that addresses each intent. Output: string.
 
query -> retrieve -> answer
  1. Rewrite the user question as a search query. Output: string.
  2. Retrieve top-k documents. Output: list of {id, text} (no LLM call).
  3. Answer using only the retrieved documents. Output: string.

Chains mix LLM and non-LLM steps freely; see rag.

Pass only the relevant subset to the next step

Forwarding the whole previous response makes the next model wade through noise. If step 1 returns {entities, rationale, debug}, step 2 receives only {entities}; the rest stays in the chain log. If you cannot write step N+1’s input schema in three lines, step N leaks too much.

Keep the router thin

A router classifies the input into a fixed label set and nothing else: no business logic, no attempt to start the task. The branch prompt does the work.

Classify the message as one of: billing, technical, account, other.
Reply with the label and nothing else.

Hold state in the chain, not the model

User history, tool results, and intermediate facts live in orchestration code, which decides what each step sees. Pass each step only its slice: the billing step gets the user’s tier, the technical step gets the support article.

Use self-correction as the default chain

The most common useful chain is draft, then review against explicit criteria, then refine, each as a separate call so you can log, evaluate, or branch between them. Run the expensive reasoning setting only on the hard step and a cheaper model or low effort elsewhere; see reasoning-model-prompting.

Log every step

Record each step’s input and output; without that the chain cannot be evaluated or debugged. Use your agent framework’s step primitives (graph nodes, tool-use loops, SDK runs) so per-step traces come free.