Overview
Reasoning models (Claude with adaptive thinking, OpenAI GPT-5 and GPT-6 with reasoning effort, Gemini with thinking levels) run an internal chain of thought before the visible answer. Configure them with an effort setting and a generous output cap, not sampling parameters, and prompt them with goals and constraints, not procedures.
Describe the goal and constraints, not the steps
Give the task, the constraints, and the output format; skip “think step by step” and hand-written procedures. Anthropic reports that a general instruction such as “think thoroughly” often beats a prescriptive plan, and OpenAI describes reasoning models as senior co-workers who do better with high-level guidance.
Bad: Think step by step, list every index option, then explain.
Good: Recommend an index for this query plan. Reply with the index DDL
and one sentence of justification.Three Claude-specific rules:
- Do not ask the model to reproduce its reasoning in the answer. On Fable 5.1, Fable 5, Opus 5.5, Opus 5, and Sonnet 5.5, classifiers can return
stop_reason: "refusal"with categoryreasoning_extractionfor such requests. Read the reasoning from summarized thinking (thinking: {"type": "adaptive", "display": "summarized"}) instead; the defaultdisplayis"omitted". - If a JSON-only answer to a multi-step task is wrong on Sonnet 5.5 at
lowormediumeffort, add “Think the problem through before you answer.” at the end of the system prompt, or raise effort tohighorxhigh. - Worked examples still help. Anthropic documents
<thinking>tags inside few-shot examples as a way to show the reasoning pattern.
Set depth with effort, not budgets or sampling parameters
| Provider | Control | Notes |
|---|---|---|
| Claude (Fable 5.x, Opus 4.7+, Opus 5.x, Sonnet 5.x) | thinking: {"type": "adaptive"} plus output_config: {"effort": "high"} (levels: low, medium, high, xhigh, max) | Default effort is high (medium on Opus 5.5). budget_tokens returns 400. Only default sampling values are accepted (temperature 1.0, top_p 0.99 or higher); other values and any top_k return 400. |
| Claude Opus 4.6, Sonnet 4.6 | Adaptive thinking plus effort | budget_tokens still works but is deprecated. Sampling parameters are accepted. |
| Claude Haiku 4.5 and older | thinking: {"type": "enabled", "budget_tokens": N} | N must be below max_tokens. |
| OpenAI GPT-5.5, GPT-5.6, GPT-6 | reasoning.effort (Responses API) or reasoning_effort (Chat Completions) | Supported values are model-dependent (none, minimal, low, medium, high, xhigh, max); GPT-6 Astra and GPT-6.1 Sol do not support none (GPT-6.1 Sol also not minimal). Remove temperature, top_p, and top_logprobs when effort is not none. |
| Gemini 3 and later | thinkingLevel (minimal, low, medium, high; support varies per model) | Keep temperature at the default of 1.0. |
| Gemini 2.5 | thinkingBudget (2.5 Pro 128 to 32768, cannot disable; 2.5 Flash 0 to 24576, 2.5 Flash-Lite 512 to 24576, 0 disables on both, -1 is dynamic) | Do not send both thinkingLevel and thinkingBudget. |
Claude details that bite:
- Fable 5.1, Fable 5, Opus 5.5, and Sonnet 5.5 reject
thinking: {"type": "disabled"}(400); Opus 5 rejects it atxhighandmax. The lowest setting on Sonnet 5.5 isthinking: {"type": "between_tools"}, accepted athigheffort or below, and it skips thinking entirely on requests without tools; use adaptive thinking for reasoning tasks. - Effort levels are recalibrated per model. Re-run an effort sweep on your evals after every model change; start at the model default (
high;mediumon Opus 5.5), and usemediumorlowfor subagents, extraction, and chat. - Changing the top-level
effortbetween requests invalidates the prompt cache. Models with per-message effort (beta) keep the cache; see prompt-caching-strategies.
Reserve output headroom for thinking
Thinking tokens bill as output and count toward the output cap, so a chat-model cap can be consumed before the answer starts. Anthropic’s max_tokens is a hard limit on thinking plus text; start near 64k at xhigh or max. OpenAI recommends reserving at least 25,000 tokens for reasoning and output when starting out, set through max_output_tokens (Responses API) or max_completion_tokens (Chat Completions).
Treat a response with stop_reason: "max_tokens" (Claude) or an incomplete status (OpenAI) as failed even when the text parses; retry with a larger cap.
Return the conclusion in a structured form
Ask for the answer only, and enforce it with provider structured outputs (see structured-output); the reasoning stays in the thinking blocks. Do not parse reasoning summaries for application logic; they are not stable across model versions. DeepSeek thinking mode ignores temperature and penalty parameters and returns reasoning_content separately.
Use a reasoning model where the step pays
Use one when the task is multi-step with a checkable answer, has interacting constraints, or fails evals on reasoning errors rather than formatting errors. Use a chat model or low effort for single-step extraction, classification, translation, and latency-bound turns, and when retrieval quality is the bottleneck. In a chain, run only the hard step at high effort; see prompt-chaining.