COLVO runs the same idea in two places — verify the real outcome, never trust the agent’s word — before launch and on every live action. What the agent said is a claim. What the provider state shows is the truth.
| Check | Expected | Observed | |
|---|---|---|---|
| Subscription | active · cancel at end | canceled | ✕ |
| Access | until 31 Oct | removed today | ✕ |
| Refund | none | none | ✓ |
| Other accounts | untouched | untouched | ✓ |
Agent-eval tools grade the transcript. COLVO grades the side effect. An agent that answers beautifully while refunding the wrong customer, cancelling too early or charging twice fails — because the state says so.
Replay incidents and edge cases against isolated fake accounts that behave like Stripe. Every attempt starts from a fresh fixture, so a passing run means the agent did the right thing three times in a row, not once by luck.
Sits in front of real actions. The agent never holds write credentials; it proposes. Guard applies your mandate, holds risky calls for human approval, executes exactly once and verifies the result against the provider.
You describe the outcome and the limits once. COLVO enforces them on every version and every live action.
Your trusted backend declares who can do what, with which limits — the rules the agent must stay inside.
The agent proposes actions over COLVO’s HTTP contract. Stripe and n8n work out of the box.
Run the suite in the sandbox; when live, Guard evaluates each action before it executes.
COLVO reads the real state independently and shows exactly what happened vs. what was allowed.
Your assistant handles refunds and cancellations on Stripe. You want it live, but a wrong action costs real money and a real customer.
You build and operate assistants for several clients. Each one needs its own isolation, its own limits and a report the client can read.
Test your agent before it ships — and guard every real action once it’s live. In a safe copy of your world first.