Overview

Guardrails are checks that run in your code, outside the model, on what goes into it and what comes out. Typical ones are input classifiers, output validators, schema validation, allow-lists for tools and arguments, PII and secret filters, and refusal handling. On Claude a refusal arrives as stop_reason: "refusal" in a normal HTTP 200 response, not an error, so handle it explicitly: remove or rephrase the refused turn, or retry on a fallback model, because resending the same history keeps refusing.

A system-prompt asks the model to behave, and safety training shapes behavior probabilistically. Allow-lists and schema checks are deterministic code that block a call whatever the model says; classifiers are models and can be bypassed.

In an agent-loop, screen input before the model call, validate each tool call before it runs, and check output before it reaches a user or sink. They reduce prompt-injection risk but do not remove it, so pair them with least-privilege tools.

Example

Before send_email runs, check the recipient against an allow-list. Before a reply is shown, block it if it matches an API-key pattern.