Guardrails & caps

Guardrails are safety checks inside the Guard pipeline: on the input before the model, on the output before the customer. Each rail returns ALLOW, REVIEW, DENY or REDACT and merges with the policy decision. Caps limit what one end user can trigger.

Written from the code · updated 26 Sep 2026 · Something wrong or missing? Tell us

Rails

idStageHowAllowed actions
prompt_injectioninputdeterministic patterns + optional AI judge on the natural-language context onlyREVIEW, DENY
pii_inputinputdeterministic: emails, cards, phones, IBANsREDACT, REVIEW, DENY
groundednessoutputAI judge: is the reply supported by the retrieved context and the real state?REVIEW, DENY
toxicityoutputAI judge: off-policy or abusive wordingREVIEW, DENY
pii_outputoutputdeterministic redactionREDACT, REVIEW, DENY
  • Deterministic rails need no AI. AI judges run only with a configured key, are metered, and require 75 % confidence to count as a hit.
  • enforce changes the decision (DENY > REVIEW > REDACT > ALLOW when merged with the policy). advisory records the hit and shows it, without changing the decision. Platform defaults: input rails enforce, output judge rails advisory.
  • REDACT masks the text and the flow continues; it is a rail action, never an operation decision.
  • Every hit is logged to the append-only evidence and listed on the Guardrails page (production scope by default).

Configure

GET /v1/guardrails?project_id=… · PUT /v1/guardrails backend API key or session (editor+) · GET /v1/guardrails/hits (keyset paginated). Requires the guardrails capability.

{ "project_id": "…", "rails": [
  { "id": "prompt_injection", "enabled": true, "action_on_hit": "DENY",   "mode": "enforce" },
  { "id": "pii_input",        "enabled": true, "action_on_hit": "REDACT", "mode": "enforce" },
  { "id": "groundedness",     "enabled": true, "action_on_hit": "REVIEW", "mode": "advisory" }
] }

Pass conversation text with the proposal (context.user_message, context.agent_reply, context.retrieved[]) so the rails have something to read; parameters are always scanned deterministically.

Per-end-user caps

Up to six caps per project, applied to the mandate subject’s customer_id across rolling windows (hour, day, month): max_actions and/or max_amount_minor, with action_on_exceed HOLD (default) or DENY. Configure from the project page or PUT /v1/projects/{id}/end-user-caps. See mandates.

Spend caps

AI spend is capped per organisation (Settings → AI provider) and per platform. At 80 % owners are emailed; at 100 % AI calls stop and deterministic checks continue unaffected.

Guardrails & caps · Docs · COLVO