Overview

Chain-of-thought (CoT) prompting asks the model to write intermediate reasoning before its final answer, so its own output serves as a scratchpad. It lifts accuracy on multi-step tasks and wastes tokens on single-step ones. Models with built-in thinking do this internally; see reasoning-model-prompting.

Use CoT for multi-step tasks on models without built-in thinking

Reach for prompted CoT on math, multi-hop questions, planning, and extractions that depend on resolving references first. Heuristic: if a human would need to write something down to solve it, CoT helps. Use it on chat models, local models, and Claude models run with thinking off.

Skip CoT for single-step tasks

Skip it for translation, lookups, format conversion, and classification with clear features. It adds latency and cost, and it can lower accuracy. Run the eval set with and without CoT and drop it if the score does not improve. When it helps the hard slice but hurts the easy slice, choose by which slice matters more (see evaluation).

Prefer built-in thinking when the model has it

Models with thinking enabled (Claude with adaptive thinking, GPT-5 and GPT-6, Gemini 3) reason internally, and depth is set with effort or a thinking level, not prompt text. Anthropic advises general instructions (“think thoroughly”) over hand-written step plans. Do not stack “think step by step” on top.

Do not ask for the reasoning to be written into the answer on Claude Fable 5.1, Fable 5, Opus 5.5, Opus 5, or Sonnet 5.5: classifiers can refuse such requests with the reasoning_extraction category. Read summarized thinking instead (display: "summarized"); the raw chain of thought is never returned.

Name the scratchpad sections

When you do prompt CoT, give the reasoning a fixed shape so evals can score it apart from the answer.

Reason inside <reasoning> tags, then give the final answer inside <answer> tags.

Parse only <answer> for the application and store <reasoning> in the request log. Never let reasoning text leak into the user-facing reply.

Put reasoning before the answer in a schema

With structured output, add a reasoning string field ahead of answer so the answer can use it; fields generated after the answer cannot help it. Mark both required, because Anthropic emits required properties first, in schema order. The client reads answer; the log keeps reasoning. See structured-output. On the Claude models listed above, use thinking blocks instead of a reasoning field.

Vote across samples for high-stakes answers

When correctness matters more than latency, run the task k times independently and take the majority answer. Vote on the final answer, not the reasoning trace. Cost is k times one call; see cost-control. Voting needs run-to-run variation: raise temperature above 0 where the API accepts it, and where only default sampling is accepted (Claude Opus 4.7 and later), rely on it and confirm the answers actually differ.