Overview

A prompt tuned for one model generation often fails or over-triggers on the next, because API parameters are removed and the compensating instructions you wrote for older models now overshoot. Treat a model change as a code change: baseline, apply the checklist below, then re-run the eval set. Facts below are as of October 2026; confirm against the vendor migration guide for your target model.

Baseline the old prompt on the new model first

Run the unchanged prompt on the eval set against the new model ID and record per-slice scores (see prompt-evals). Then change one thing at a time. Adjacent versions often work without edits; Anthropic says existing Fable 5 prompts should perform well on Fable 5.1, for example.

Remove parameters the new model rejects

  • Assistant prefill returns 400 on Claude 4.6 and later and every 5.x model. Replace format forcing with structured outputs (output_config: {format: {...}}), preamble suppression with “Respond directly without preamble”, and continuations with a user message that quotes the end of the interrupted reply.
  • Sampling parameters: Claude Opus 4.7 and later, Opus 5.x, Sonnet 5.x, and Fable 5.x accept only the defaults (temperature 1.0, top_p 0.99 or higher) and return 400 for other values and any top_k; drop them. OpenAI GPT-6 Astra does not support custom temperature or top_p.
  • thinking: {"type": "enabled", "budget_tokens": N} returns 400 on Opus 4.7 and later, Opus 5.x, Sonnet 5.x, and Fable 5.x. Use thinking: {"type": "adaptive"} plus output_config.effort; thinking: {"type": "disabled"} also returns 400 on Opus 5.5.
  • Forced tool_choice (any or tool) returns 400 on Fable 5.1, Opus 5.5, and Sonnet 5.5. Use auto with strict: true tools, or structured outputs.
  • Top-level output_format is deprecated; use output_config.format.
  • OpenAI reasoning.effort: "none" is unsupported on GPT-6 Astra and GPT-6.1 Sol; use low.

Remove instructions written to compensate for older models

  • Emphatic wording (“CRITICAL: You MUST use this tool”) makes models that follow the system prompt closely use tools too often. Write “Use this tool when…“.
  • Blanket tool defaults (“If in doubt, use the tool”) and anti-laziness prompts (“be thorough”) push current models into excess exploration.
  • Anti-formatting rules can suppress needed structure on Fable 5.1, which already formats less; replace them with a rule for when formatting is appropriate.
  • Verification boilerplate causes over-verification on Opus 5; remove it.
  • “Think step by step” scaffolds add cost, and requests to reproduce reasoning in the answer can trigger reasoning_extraction refusals on Fable 5.x, Opus 5.x, and Sonnet 5.5. See reasoning-model-prompting.

Re-sweep effort and recount tokens

Effort level names do not map to the same thinking depth across models, so run a fresh sweep instead of copying the old setting. Claude 4.7 and later use a tokenizer that produces about 30 percent more tokens for the same text, so recheck max_tokens, cost estimates, and cache minimums (see cost-control).

Keep history append-only on thinking models

On Fable 5.1, Opus 5.5, and Sonnet 5.5, pass thinking blocks back unchanged. Editing earlier turns, rebuilding system or tools, or summarizing old turns in place breaks them and restarts the prompt cache; use mid-conversation system messages and server-side compaction instead. See prompt-caching-strategies.

Re-run evals, then pin

Compare the new scores per slice, fix regressions with the smallest prompt change, and pin the exact model ID. Prompt caches are per model, so expect a cold start after cutover.