Overview

Output constraints narrow the model from “any plausible response” to the one response the application can use. This page covers format, length, refusal sentinels, and exclusions for prose and mixed output. For machine-read JSON, use schema-enforced output instead; see structured-output.

Show the format, do not only describe it

A short example of the target output steers format more reliably than a paragraph of prose about it. For nested or multi-line formats, put the example in the same delimited block as the spec.

Bad:  Return a label (string) and a confidence (float between 0 and 1).
Good: Reply in this exact format:
      label: billing
      confidence: 0.87

Match the prompt’s own style to the output style you want; a prompt full of markdown pulls the output toward markdown. Prefilling the assistant turn is not an option on current Claude models (4.6 and later return 400); use structured outputs, an instruction such as “Respond directly without preamble”, or tags around the answer instead.

Give the model an escape hatch

Grounded Q&A needs permission to decline, or the model invents a plausible answer. Define one sentinel per condition, uppercase and without punctuation so it survives normalization, document it in the prompt, and handle it in the caller.

Reply NOT_FOUND if the answer is not in the provided text.
Reply OUT_OF_SCOPE if the question is not about Postgres.

State length in units, and treat the output cap as a circuit breaker

Write the limit in sentences, words, or items (“at most 5 bullets, 12 words each”); “be concise” gives the model no anchor. The provider’s output cap (max_tokens, max_output_tokens) truncates; it does not shape the answer. Set it above the longest legitimate answer, and on reasoning models above answer plus thinking; see reasoning-model-prompting.

Effort settings control thinking volume, not visible length. Anthropic notes that on Opus 5, changing effort does not reliably shorten replies, so ask for the length you want. The OpenAI Responses API has text.verbosity (low, medium, high; default medium) for the length of the final text, but the GPT-6 guide does not document it, so confirm it moves length on your model and still state the limit in the prompt.

Require a source for each grounded claim

Add citations to the output spec (a document ID or short quote per claim) so a reviewer or validator can check each one and fabricated claims stand out. See rag-citations.

State exclusions as the action to take

Prefer “Refer the user to /help/faq” over “Do not ask for the password”, and “Cite only facts from the provided text; mark anything else INFERRED” over “Do not include opinions”. Anthropic’s guidance is to tell the model what to do instead of what not to do. When a true exclusion is unavoidable, state it once, near the top.

Validate in the caller, then retry fresh

Check every output against the format, length, and sentinel rules in code. On failure, log it and rerun the original prompt as a new call, or fail loudly. Do not append “that was wrong, try again” to the conversation: the bad output stays in context and anchors the next one. Track the failure rate as a deployment metric; a rising rate means a prompt or model regression (see prompt-evals). For high-stakes output, use a second call with a narrow prompt that returns VALID or INVALID plus one line, and do not let it rewrite the output.

Pitfalls

  • Putting the format spec after the user input. Put spec and example first and the input last.
  • Asking for prose and JSON in one response. Pick one, or use two calls.
  • Treating the output cap as a length instruction.