Overview
Claude Code resends the whole conversation on every request, so context size drives both answer quality and spend. For API-level caching and routing see cost-control; for the instruction file see claude-code-claude-md.
Check what is loaded before you trim it
Run /context for a live breakdown by category with optimization suggestions, including which CLAUDE.md and auto memory files loaded. A session starts with the system prompt, auto memory, environment info, MCP tool names, skill descriptions, and CLAUDE.md files; file reads and command output drive growth after that. /statusline can show context_window.used_percentage continuously.
Pick the command that matches the situation
| Situation | Command | Effect |
|---|---|---|
| Switching to unrelated work | /clear <name> | New conversation with empty context; the name labels the old session in the /resume picker |
| Same task, history is bloated | /compact <focus> | Summarizes the conversation and keeps what the focus text names |
| Only part of the history is noise | /rewind, then Summarize from here or up to here | Targeted compaction; type instructions on the “add context” row |
| A path went wrong | /rewind, or Esc twice on an empty prompt | Truncates to an earlier turn; restores conversation, code, or both |
| A side question | /btw | The answer never enters history |
| Compaction fires too late | /autocompact 500k | Sets the window in tokens (100K to 1M); auto restores the default |
Compact at natural breaks between tasks, not mid-task: compaction rebuilds the conversation cache, and after a break past the cache lifetime the summary request reprocesses the full history uncached. To abandon a path, prefer /rewind, which truncates back to an already cached prefix. Auto-compaction runs near the limit, about 967K tokens by default on models with a native 1M window.
Put standing compaction guidance in CLAUDE.md:
# Compact instructions
When you are using compact, please focus on test output and code changesAfter compaction the project-root CLAUDE.md and auto memory reload from disk, while nested CLAUDE.md files and paths: rules return only when Claude reads a matching file again. A rule that must persist belongs in the root file, and instructions given only in chat may not survive. Checkpoints skip Bash side effects and most subagent edits, so keep using git.
Delegate high-volume work to subagents
Run tests, log processing, and documentation fetches in a subagent so the output stays in its context and only a summary returns; its requests still draw on your usage (claude-code-subagents).
- Subagents inherit the session model, so a
/modelswitch to Opus moves them too. Setmodelin custom definitions; Explore is capped at Opus on the Claude API, and a project subagent namedExplorewithmodel: haikureplaces it. - To pin every subagent, set
CLAUDE_CODE_SUBAGENT_MODELandCLAUDE_CODE_SUBAGENT_MODEL_FORCE=1inenv(v2.1.257 and later). - Agent teams are off by default and use roughly 7 times the tokens of a standard session when teammates run in plan mode.
Keep MCP, skills, and hooks from inflating the baseline
Leave MCP tool search on, so only tool names and server instructions load, and disable unused servers with /mcp. ENABLE_TOOL_SEARCH=auto loads schemas upfront while they total under 10 percent of the window, false loads everything, and alwaysLoad: true exempts one server. A CLI such as gh adds no per-tool listing (claude-code-mcp).
Mark skills with side effects disable-model-invocation: true so they stay out of context until you run /name. Move workflow-specific instructions from CLAUDE.md into skills. A PreToolUse hook can filter test output to failures before Claude sees it.
Choose model and effort first, then leave them alone
Use Sonnet for most coding and reserve Opus for complex architecture or multi-step reasoning. /effort takes low, medium, high, xhigh, max, or auto, depending on the model; the default is high on most models and medium on Opus 5.5 and Sonnet 5.5. Thinking tokens bill as output, and Claude Code cannot turn thinking off on Opus 5.5, Sonnet 5.5, or the Fable models, so lower effort instead.
Each model has its own cache, so a mid-session /model switch makes the next request read the whole history uncached. An effort change does the same on most models, but not on Opus 5.5, Sonnet 5.5, or Fable 5.1 with an API key or subscription. Enabling /fast (Opus 5.5, Opus 5, Opus 4.8) late in a session bills the full context uncached once.
Track spend with /usage
/usage shows session cost, plan limits, and activity stats; /cost is an alias. The dollar figure is a local estimate at list price, so the Console usage page is authoritative. The Session block also reports prompt cache hit share and misses (v2.1.251 and later), and on subscription plans usage is attributed to skills, subagents, plugins, and MCP servers.
Set team controls on the billing path you use
| Billing path | Cap | Per-user reporting |
|---|---|---|
| Claude for Teams or Enterprise | Seat allowance (rolling five-hour and weekly windows); usage credits with spend limits per organization, group, or member | Spend report in org analytics with CSV export; Enterprise Analytics API |
| Claude Console | Workspace spend limits on the auto-created “Claude Code” workspace | Console dashboard; Claude Code Analytics API |
| Bedrock, Google Cloud’s Agent Platform, Microsoft Foundry | Your cloud’s budget controls | OpenTelemetry or an LLM gateway |
The modelPricing managed setting makes reported figures match contracted rates without changing billing, and maxEffortLevel caps the effort users can choose. In scripts, claude -p --max-budget-usd 5.00 and --max-turns cap a run, and subagent spend counts toward the budget. Anthropic reports enterprise averages near $13 per developer per active day.