One product · two jobs

Prove it works. Then keep it honest.

COLVO runs the same idea in two places — verify the real outcome, never trust the agent’s word — before launch and on every live action. What the agent said is a claim. What the provider state shows is the truth.

Deterministic verdictsFresh fixture per attemptStripe first
The problem

“Done, I refunded €50 and kept your access.” Did it?

Agent-eval tools grade the transcript. COLVO grades the side effect. An agent that answers beautifully while refunding the wrong customer, cancelling too early or charging twice fails — because the state says so.

What it said
“Refunded €50 to the original card and kept the subscription active until the 31st.”
≠
What it did
Refund €500 · subscription cancelled today · access removed
01 — COLVO TEST
Before · staging

COLVO Test the sandbox

Replay incidents and edge cases against isolated fake accounts that behave like Stripe. Every attempt starts from a fresh fixture, so a passing run means the agent did the right thing three times in a row, not once by luck.

  • 22 scenario templates for refunds and cancellations, plus your own from the AI Test Architect
  • Deterministic checks decide PASS / FAIL / INCONCLUSIVE — no AI needed for the verdict
  • Repeat runs catch flakiness; version comparison shows what broke and what got fixed
  • A red team generates hostile inputs and reports what got through
  • CI gate: fail the build when the suite fails

Explore the scenarios →

During · production

COLVO Guard the gate

Sits in front of real actions. The agent never holds write credentials; it proposes. Guard applies your mandate, holds risky calls for human approval, executes exactly once and verifies the result against the provider.

  • ALLOW / REVIEW / HOLD / DENY on every proposed action, with reasons
  • Approvals with an exact digest of what the reviewer signs off
  • Idempotent execution: a retry never doubles an effect
  • Independent verification and incidents when expected ≠ observed
  • Guardrails on input and output, per-end-user caps, one-click kill-switch

How the gate decides →

02 — HOW IT WORKS
How it works

Four steps to trustworthy agents

You describe the outcome and the limits once. COLVO enforces them on every version and every live action.

Register the mandate

Your trusted backend declares who can do what, with which limits — the rules the agent must stay inside.

Connect the agent

The agent proposes actions over COLVO’s HTTP contract. Stripe and n8n work out of the box.

Test, then guard

Run the suite in the sandbox; when live, Guard evaluates each action before it executes.

Verify & report

COLVO reads the real state independently and shows exactly what happened vs. what was allowed.

03 — WHO IT IS FOR
Who uses it

Teams that let an agent touch money and customers

SaaS support teams

Your assistant handles refunds and cancellations on Stripe. You want it live, but a wrong action costs real money and a real customer.

  • Ship the agent with a suite that proves the money paths
  • Keep a human in the loop for anything unusual
  • Show finance an evidence trail for every refund

Agencies running agents for clients

You build and operate assistants for several clients. Each one needs its own isolation, its own limits and a report the client can read.

  • One organisation per client, strict tenant isolation
  • Version comparison for every release
  • Compliance export per client and date range
≠

Stop trusting the reply.

Test your agent before it ships — and guard every real action once it’s live. In a safe copy of your world first.

Product overview — Test and Guard · COLVO