≠
Test before · Guard during · Verify after

What it said
isn't what it did.

COLVO tests your AI agent in a safe copy of your world — then guards every real action in production. It checks the outcome, not the reply. When the agent is wrong, COLVO catches it in testing, or blocks it live.

Sandboxed testing · live approval gate · Stripe-native · no false "done"

Expected
Renewal off.
Access stays until Sep 30.
≠
Observed
Renewal off.
Access killed today. ✕

Same request. The agent reported success. The state disagreed. COLVO fails it.

Works with the tools your agents already use

Stripen8nMakeOpenAIAnthropicCustom HTTP
01 — THE PRODUCT
One product, two jobs

Prove it works. Then keep it honest.

COLVO runs the same idea in two places — verify the real outcome, never trust the agent's word — before launch and on every live action.

Before · testing

COLVO Test the sandbox

Replay incidents and edge cases against isolated fake accounts. Catch the agent that says "done" while quietly breaking something.

  • Isolated sandbox, fresh fixture per attempt
  • 22 scenarios, repeat runs & version diffs
  • PASS / FAIL / INCONCLUSIVE — no false greens
During · production

COLVO Guard the gate

Sits in front of real actions. Applies your policy before a write, holds risky calls for human approval, and verifies the result independently.

  • ALLOW / REVIEW / HOLD / DENY on every action
  • Human approval before real writes
  • Idempotent execution + instant kill-switch
Test before you shipGuard while it runsVerify after every action
02 — PROCESS
How it works

Four steps to trustworthy agents

You describe the outcome and the limits once. COLVO enforces them on every version and every live action.

Register the mandate

Your trusted backend declares who can do what, with which limits — the rules the agent must stay inside.

Connect the agent

The agent proposes actions over COLVO's HTTP contract. Stripe and n8n work out of the box.

Test, then guard

Run the suite in the sandbox; when live, Guard evaluates each action before it executes.

Verify & report

COLVO reads the real state independently and shows exactly what happened vs. what was allowed.

03 — CAPABILITIES
What's inside

Everything you need to trust an agent with real actions

Built around one idea: verify the outcome, not the answer — in testing and in production.

Isolated sandbox

Every test acts on fake accounts that mimic your real APIs. No real users, data, or charges — ever.

fresh fixture per attempt

Policy gate

Guard evaluates every proposed action against your mandate and returns ALLOW / REVIEW / HOLD / DENY.

before the write

Human approval

Risky actions pause for a person to approve or reject, with an exact digest of what they're signing off.

human-in-the-loop

Independent verification

COLVO reads ground-truth state on its own and compares it to the rule. The reply is just a claim.

state, not text

Idempotent & safe retries

Stable business keys and reservations mean a retry never creates a second refund or a double charge.

exactly-once effect

Kill-switch & pause

Stop new writes instantly for a project. In-flight actions are held and reconciled, never lost.

one click

Scenario suite

22 templates for refunds and cancellations. Repeat runs catch flakiness; version diffs catch regressions.

versioned & approved

Evidence & export

Append-only log of every request, decision, action and state snapshot. Export to JSON/PDF for audit.

JSON · PDF

Workspaces & roles

Projects with owner / editor / viewer roles and strict per-organization isolation on every resource.

multi-tenant
04 — COLVO GUARD
Live protection

Every real action passes the gate first.

The agent never holds write credentials. It proposes; COLVO decides, holds for approval when needed, executes safely, and verifies the result.

1

Agent proposes an action

e.g. "refund €500 on payment pi_01" — against a mandate that allows €50.

2

Guard evaluates the policy

Identity, scope, limits, currency and provider state — all before anything is sent.

3

Execute safely or hold

Allowed actions run idempotently; risky ones wait for a human; violations are denied.

4

Verify the real effect

A read-only check confirms the provider state — no "completed" without proof.

ALLOW

Within the mandate

Refund €50 on the original payment method — executed once, then verified.

REVIEW

Needs a human

Unusual but plausible — paused with an exact digest for a reviewer to approve or reject.

HOLD

Can't confirm safely

Required data unavailable — no write happens until the state is readable again.

DENY

Policy violation

€500 exceeds the €50 limit — blocked before it ever reaches Stripe.

05 — VERDICT LOGIC
How COLVO decides

Hard facts first. Judgement second.

Every verdict rests on explicit checks against real state. A language model can help rate the wording — it never overrides what actually happened.

Deterministic checks

// decide PASS / FAIL / ALLOW / DENY
  • Amount, currency & limits within the mandate
  • Action tied to the original request & payment
  • Forbidden actions never happened
  • Effective dates & account fields match
  • No double effect, no cross-tenant leak

Semantic checks

// advisory, with a reason
  • Was the reply clear and consistent?
  • Did it follow the written instructions?
  • Tone appropriate for the customer?
  • Flagged with a rationale for review
  • Never flips a state failure into a pass
06 — THE REPORT
The report

One run. Every answer.

A readable report your team and your client can both trust — traceable from every verdict back to the mandate and the state that produced it.

Refunds & cancellations · agent v0.2Run #148 · 3 attempts each · 2m 14s · €0.61
ScenarioCheckResultPass rateCost
Refund over mandate (€500 vs €50)action denied before send✓ PASS3 / 3€0.12
Refund authorised (€50)one refund, verified✓ PASS3 / 3€0.11
Cancel at period endaccess kept until expiry✕ FAIL1 / 3€0.14
Duplicate request, same keysingle effect only✓ PASS3 / 3€0.10
Provider state unreadableno false success◐ INCONCLUSIVE—€0.14
07 — SCENARIOS
Scenario families

Start with the actions that touch money

The first release ships 22 templates across the two processes where a wrong action costs the most — refunds and cancellations.

Refunds

10 templates
  • Amount over the mandate
  • Refund to a different customer
  • Wrong currency
  • Duplicate / concurrent requests
  • Provider refusal or lost response

Cancellations

10 templates
  • Stop at period end, keep access
  • Access removed too early
  • Wrong subscription
  • Duplicated request
  • Deadline / timezone edge

Plan change

next
  • Immediate vs end-of-cycle
  • Upgrade with proration
  • Deferred to a later release
  • On the roadmap after pilot
08 — ISOLATION
Isolation & trust

A control that can be bypassed protects nothing.

The agent never holds write credentials, and isolation is enforced server-side. Security is built into the MVP, not bolted on later.

MANDATED WRITES

The agent can't write directly

Only COLVO's executor holds provider write access. The agent proposes against a mandate registered by your trusted backend — a chat-supplied ID is never enough.

TENANT ISOLATION

Nothing leaks across orgs

Org identity comes from the authenticated session, enforced with row-level security in Postgres. Cross-account access is denied in UI, API, export and worker alike.

EGRESS CONTROL

Locked-down outbound calls

Only approved hosts. Local, private and metadata addresses blocked, DNS & redirects re-checked — SSRF-hardened by default.

SECRETS & EVIDENCE

Encrypted, scoped, append-only

Secrets encrypted at rest and separated from production. Every decision and action is logged append-only, and data is deletable on request.

09 — PRICING
Pricing

Start free. Upgrade when your agent touches real money.

Simple monthly plans per agency. Test usage included; live actions and heavy runs metered transparently.

Free
€0 / mo
Prove the concept on your own agent in the sandbox before you pay a cent.
  • 100 test runs / month
  • COLVO Test + AI Test Architect
  • 22 scenario templates, sandbox + reports
  • JSON / CSV / PDF export
Start free
Test
€19 / mo
Evaluation at scale with a CI gate — for teams that ship agent versions weekly.
  • 2,000 runs / month
  • Run API, CLI and GitHub Action
  • Scheduled monitoring with drift alerts
  • Version comparison and regression reports
Request access
Most popular
Guard
€99 / mo
Test + live Guard for teams running client assistants on real transactions.
  • Everything in Test, unlimited runs
  • Live Guard: mandates, approvals, verification
  • Guardrails pack + per-end-user caps
  • Compliance export, red team, scheduled monitoring
Request access

Prices per organisation, billed monthly. AI usage is metered at cost (BYOK or managed) and shown per run — never a hidden €0.

10 — FAQ
FAQ

The honest answers

No spin. If there's a limit, we say so up front.

Still have a question?

Ask us anything about your setup — we'll tell you straight whether COLVO fits.

Talk to us
What's the difference between Test and Guard?

Test runs before you ship — it replays scenarios against a sandbox and grades PASS/FAIL/INCONCLUSIVE. Guard runs in production — it evaluates each real action against your mandate and returns ALLOW/REVIEW/HOLD/DENY, holding risky ones for human approval before anything executes.

Does the agent get to touch Stripe directly?

No. The agent never holds write credentials — it only proposes actions. COLVO's executor performs the write, and only after the policy allows it (or a human approves it). That's what makes the guardrail real rather than advisory.

Why not just trust the agent's reply?

Because the reply can be wrong. "Done, I refunded €50" while €500 went out — or nothing did — is exactly what costs you money. COLVO checks the real provider state, so a confident lie still fails.

Which actions does the first release cover?

Refunds and cancellations on Stripe — the two processes where a wrong action is most expensive. Plan changes and other actions are on the roadmap after the pilot.

What does INCONCLUSIVE / HOLD mean?

The result couldn't be trusted — unreadable state, a fault, or a timeout. It's never counted as success and never silently executed. A simulated backend error is different: the agent can still pass if it handles the failure correctly.

Is my data safe?

Pilots run on fake data. Secrets are encrypted and separated from production, tenants isolated with row-level security, outbound calls locked to approved hosts, every action logged append-only, and data deletable on request.

≠

Stop trusting the reply.

Test your agent before it ships — and guard every real action once it’s live. In a safe copy of your world first.

COLVO — Test before. Guard during. Verify after.