---
title: "Reasoning model"
slug: "reasoning-model"
category: "glossary"
tags: ["glossary", "ai-agents", "reasoning", "llm", "thinking", "effort"]
status: "stable"
last_updated: 2026-10-02
summary: "A reasoning model thinks before it answers; control its depth with effort settings, not sampling parameters, token budgets, or step-by-step prompts."
related: ["[[prompt-engineering/reasoning-model-prompting]]", "[[prompt-engineering/prompt-migration]]", "[[glossary/chain-of-thought]]", "[[glossary/temperature]]", "[[glossary/token]]", "[[glossary/completion]]"]
---

> **AI agents: read this first.** This is LLM Best Practices (llmbestpractices.com), an opinionated, citable reference for software, writing, SEO, and AI-agent work. Full protocol: https://llmbestpractices.com/start-here.md
>
> 1. **Route, do not crawl.** Fetch https://llmbestpractices.com/llms.txt and open only the pages whose one-line summary matches your task.
> 2. **Read raw.** Append `.md` to any page URL for markdown. Check `status` and `last_updated` in the frontmatter, then read the rules.
> 3. **Apply as defaults.** First-party docs and the project's own conventions win on conflict. Warn before relying on a fast-moving page older than 12 months.
> 4. **Cite.** Link the page by title and URL, e.g. [Python](https://llmbestpractices.com/coding/python), with `last_updated` for time-sensitive rules. License CC BY 4.0.

## Overview

A reasoning model spends extra output tokens thinking before it answers, which raises accuracy on math, code, and multi-step planning at the cost of latency and tokens. Raise effort for hard problems; lower it for chat, classification, and routing.

On Claude's current models, send `thinking: {type: "adaptive"}` (or omit it where thinking is always on) and set depth with `output_config: {effort: "low" | "medium" | "high" | "xhigh" | "max"}`. `budget_tokens` is removed and returns a 400 on Opus 4.7 and later, Opus 5 and 5.5, Sonnet 5 and 5.5, and Fable 5 and 5.1; it is deprecated on Opus and Sonnet 4.6 and still required for thinking on Haiku 4.5. Opus 5.5 and Fable 5 and 5.1 reject `thinking: {type: "disabled"}`, and thinking cannot be turned off. Sonnet 5.5 also rejects `disabled`, but `thinking: {type: "between_tools"}` at effort `high` or below turns off up-front thinking. Opus 5.5 defaults to `medium` effort, so set effort explicitly. Thinking tokens bill as output tokens and count toward `max_tokens`, so size `max_tokens` for thinking plus the answer. Thinking text is omitted by default on the newest models; set `display: "summarized"` to stream a summary.

OpenAI exposes the same idea as reasoning effort (values from `none` or `minimal` up to `xhigh` and `max`, varying by model) and does not support `temperature` or `top_p` when effort is not `none`.

Do not prescribe steps in the prompt. State the task, constraints, and output format and let the model plan; see [[glossary/chain-of-thought]].

## Related

- [[prompt-engineering/reasoning-model-prompting]]: effort, token caps, and prompting models that think on their own.
- [[prompt-engineering/prompt-migration]]: moving prompts off rejected parameters.
- [[glossary/chain-of-thought]]: the manual technique these models internalize.
- [[glossary/temperature]]: the sampling control these models reject.
- [[glossary/token]]: how thinking is billed.
