Overview
A reasoning model spends extra output tokens thinking before it answers, which raises accuracy on math, code, and multi-step planning at the cost of latency and tokens. Raise effort for hard problems; lower it for chat, classification, and routing.
On Claude’s current models, send thinking: {type: "adaptive"} (or omit it where thinking is always on) and set depth with output_config: {effort: "low" | "medium" | "high" | "xhigh" | "max"}. budget_tokens is removed and returns a 400 on Opus 4.7 and later, Opus 5 and 5.5, Sonnet 5 and 5.5, and Fable 5 and 5.1; it is deprecated on Opus and Sonnet 4.6 and still required for thinking on Haiku 4.5. Opus 5.5 and Fable 5 and 5.1 reject thinking: {type: "disabled"}, and thinking cannot be turned off. Sonnet 5.5 also rejects disabled, but thinking: {type: "between_tools"} at effort high or below turns off up-front thinking. Opus 5.5 defaults to medium effort, so set effort explicitly. Thinking tokens bill as output tokens and count toward max_tokens, so size max_tokens for thinking plus the answer. Thinking text is omitted by default on the newest models; set display: "summarized" to stream a summary.
OpenAI exposes the same idea as reasoning effort (values from none or minimal up to xhigh and max, varying by model) and does not support temperature or top_p when effort is not none.
Do not prescribe steps in the prompt. State the task, constraints, and output format and let the model plan; see chain-of-thought.
Related
- reasoning-model-prompting: effort, token caps, and prompting models that think on their own.
- prompt-migration: moving prompts off rejected parameters.
- chain-of-thought: the manual technique these models internalize.
- temperature: the sampling control these models reject.
- token: how thinking is billed.