Guardrails & caps
Guardrails are safety checks inside the Guard pipeline: on the input before the model, on the output before the customer. Each rail returns ALLOW, REVIEW, DENY or REDACT and merges with the policy decision. Caps limit what one end user can trigger.
Rails
| id | Stage | How | Allowed actions |
|---|---|---|---|
| prompt_injection | input | deterministic patterns + optional AI judge on the natural-language context only | REVIEW, DENY |
| pii_input | input | deterministic: emails, cards, phones, IBANs | REDACT, REVIEW, DENY |
| groundedness | output | AI judge: is the reply supported by the retrieved context and the real state? | REVIEW, DENY |
| toxicity | output | AI judge: off-policy or abusive wording | REVIEW, DENY |
| pii_output | output | deterministic redaction | REDACT, REVIEW, DENY |
- Deterministic rails need no AI. AI judges run only with a configured key, are metered, and require 75 % confidence to count as a hit.
- enforce changes the decision (DENY > REVIEW > REDACT > ALLOW when merged with the policy). advisory records the hit and shows it, without changing the decision. Platform defaults: input rails enforce, output judge rails advisory.
- REDACT masks the text and the flow continues; it is a rail action, never an operation decision.
- Every hit is logged to the append-only evidence and listed on the Guardrails page (production scope by default).
Configure
GET /v1/guardrails?project_id=… · PUT /v1/guardrails backend API key or session (editor+) · GET /v1/guardrails/hits (keyset paginated). Requires the guardrails capability.
{ "project_id": "…", "rails": [
{ "id": "prompt_injection", "enabled": true, "action_on_hit": "DENY", "mode": "enforce" },
{ "id": "pii_input", "enabled": true, "action_on_hit": "REDACT", "mode": "enforce" },
{ "id": "groundedness", "enabled": true, "action_on_hit": "REVIEW", "mode": "advisory" }
] }Pass conversation text with the proposal (context.user_message, context.agent_reply, context.retrieved[]) so the rails have something to read; parameters are always scanned deterministically.
Per-end-user caps
Up to six caps per project, applied to the mandate subject’s customer_id across rolling windows (hour, day, month): max_actions and/or max_amount_minor, with action_on_exceed HOLD (default) or DENY. Configure from the project page or PUT /v1/projects/{id}/end-user-caps. See mandates.
Spend caps
AI spend is capped per organisation (Settings → AI provider) and per platform. At 80 % owners are emailed; at 100 % AI calls stop and deterministic checks continue unaffected.