What’s inside

Everything you need to trust an agent with real actions.

Built around one idea: verify the outcome, not the answer — in testing and in production. Here is the full list, grouped by when it works for you.

01 — TESTING
Testing

Prove the agent before it ships

Isolated sandbox

Every attempt acts on fresh fake accounts that mimic Stripe. No real users, data or charges — ever.

fresh fixture per attempt

Scenario suite

22 templates for refunds and cancellations; repeat runs catch flakiness, version diffs catch regressions.

versioned & approved

AI Test Architect

Describe your agent; COLVO drafts scenarios — request, approved rule, starting data. A human approves before anything runs.

metered, advisory

Red team

Generates hostile and edge inputs — injection attempts, boundary amounts, ambiguous cancels — and reports what got through.

blocked vs through

Deterministic verdicts

PASS / FAIL / INCONCLUSIVE from explicit checks against state. A semantic judge can comment; it never flips a FAIL.

no false greens

CI gate & schedules

Run the suite from CI or on a schedule; fail the build on FAIL; alert on regressions.

cli · cron
02 — PRODUCTION
Production

Guard every real action

Policy gate

Every proposed action is evaluated against the mandate and returns ALLOW / REVIEW / HOLD / DENY with reasons.

before the write

Human approval

Risky actions pause for a person, who signs an exact digest of what will happen.

human-in-the-loop

Independent verification

COLVO reads ground-truth state on its own and compares it to the rule. The reply is just a claim.

state, not text

Idempotent & safe retries

Stable business keys, reservations and an outbox: a retry never creates a second refund.

exactly-once effect

Guardrails

Prompt-injection and PII detection with redaction on input; groundedness and toxicity on output; enforce or advisory per rail.

input · output

Per-end-user caps

Actions and spend per customer in rolling windows — HOLD or DENY when a single end user exceeds them.

per subject

Kill-switch & pause

Stop new writes for a project instantly; in-flight actions are held and reconciled, never lost.

one click

Incidents & alerts

MISMATCH, UNVERIFIABLE and TEST_FAIL open incidents with evidence; Slack, email and webhook alerts.

open → resolved
03 — TRUST & OPERATIONS
Trust & operations

Evidence, insight and control

Evidence & export

Append-only log of every request, decision, action and state snapshot. JSON, CSV and PDF per run.

JSON · CSV · PDF

Compliance report

Per organisation and date range: decisions, approvals, guardrail hits, verifications — SHA-256 chained and recomputable.

tamper-evident

Analytics

Pass rate over time, flakiness, block rate, cost per run, where the agent keeps breaking.

trends

Metered AI cost

Every AI call metered: known from the gateway or estimated from the catalog, never a hidden €0. Caps at org and platform level.

BYOK or managed

Workspaces & roles

Projects with owner / editor / viewer roles and strict per-organisation isolation on every resource.

multi-tenant

Activity log

Who did what and when — sign-ins, key changes, plan changes, decisions — for your team and for security review.

audit
Plans

Same honesty on every plan

Free is Test-only with 100 runs a month. Test adds the CI gate and schedules. Guard adds live protection, guardrails, caps, compliance export and the red team. Compare plans →

≠

Stop trusting the reply.

Test your agent before it ships — and guard every real action once it’s live. In a safe copy of your world first.

Features — everything you need to trust an agent with real actions · COLVO