Approval digest
A fingerprint of exactly what a reviewer approved.
When an action needs a human, the approval is tied to a digest of the exact action and parameters. If anything changes after the reviewer looked at it, the old approval no longer matches and cannot be used.
See also: Human-in-the-loop approval
Business key
The identity of a request in business terms, not HTTP terms.
A key built from what the customer asked for (which mandate, which source request, which action) rather than from a single network call. Retries, duplicates and restarts of the same request share one business key, so they can be recognised as the same thing.
See also: Idempotency key
CI gate
A test run that blocks a release when the agent fails.
Running the scenario suite from your CI pipeline (CLI or GitHub Action) on every agent change, and failing the build when a scenario that used to pass now fails.
CLI docs →
Deterministic check
A pass/fail rule on the final provider state. No AI involved.
A check such as “exactly one refund of €50 on the right payment” or “access still enabled”. It reads the state after the attempt and gives the same answer every time. In COLVO, deterministic checks alone decide the verdict.
See also: Semantic check (advisory)
Drift
An agent that used to pass starts failing, with nobody changing it.
Model updates, prompt tweaks and upstream API changes can change an agent’s behaviour without a release. Scheduled runs compare results over time and alert when the pass rate drops.
Evidence (append-only)
The record of what was proposed, decided, sent and observed.
Every decision, provider call and verification read is written as evidence that can be added to but not edited or deleted, chained so tampering is detectable, and exported for audits.
Fault injection
Making the sandbox fail on purpose, at the worst moment.
Errors before a write, lost responses after a write, timeouts, stale reads, refusals and duplicate requests — injected in the simulated provider so you see how the agent behaves when the world is not polite.
See also: Lost response after write, Stale read
Guardrail
A check on the content going into or out of the agent.
For example detecting prompt injection or card numbers in a customer message. A guardrail can allow, send for review, redact or deny — in enforce mode it acts, in advisory mode it only records the hit.
Human-in-the-loop approval
A person approves an action before it is sent.
Actions above a threshold (for example refunds over €30) wait for a reviewer. The reviewer sees exactly what will happen and approves or rejects; nothing reaches the provider until then.
See also: Approval digest, Mandate
Idempotency key
A key that makes repeating a request safe.
Payment providers accept a key with each write; a second request with the same key returns the first result instead of acting again. It only protects you if the key belongs to the business request — a new key per retry defeats it.
See also: Business key, Lost response after write · When your AI agent refunds twice →
Least privilege
The agent gets only the access the task needs — and never the write key.
The agent proposes actions; a separate executor holds the provider credentials and only sends what a mandate allows. Even a compromised or confused agent cannot act outside those limits.
See also: Mandate
Lost response after write
The provider did it; you never heard back.
The most dangerous failure for money-moving agents: the write succeeded but the response was lost, so the caller believes it failed and may retry.
See also: Idempotency key · Double refund →
Mandate
What an agent may do, for whom, up to how much, until when.
Your backend registers a mandate for each customer request: the allowed actions (refund, cancel), the customer they apply to, amount and currency limits, how many operations, which cancel modes, the approval threshold and an expiry. Guard checks every proposed action against it before anything is sent.
Build one with the mandate builder →
Observation mode
Guard watches and scores, without blocking.
A way to start: your agent keeps acting as today while COLVO records what it would have allowed, reviewed or denied, and verifies the outcomes. Switch to enforcing when the numbers look right.
Out-of-band change
A change in the provider that no approved operation explains.
COLVO listens to the provider’s events. A refund or cancellation that matches no COLVO operation means something else still has write access — an old key in the agent, or a manual change — and raises an incident.
Outcome verification
Reading the provider after an action to prove it happened as intended.
After every action, the provider is read independently of the agent and compared with what the mandate intended: verified, mismatch, or unverifiable. A reply that says “done” is never enough.
See also: Stale read
Sandbox world
A private, simulated copy of the provider for one test attempt.
Each attempt gets a fresh simulated Stripe account with its own customers, subscriptions and payments, isolated per organisation and attempt, so tests never touch real money and never interfere with each other.
Semantic check (advisory)
An AI reading of the conversation that flags suspicious replies.
For example a reply that promises a refund the state doesn’t show. It is advisory: it adds context to a report but can never turn a failed state into a pass.
See also: Deterministic check
Spend cap
A monthly ceiling on metered AI cost.
Every AI call is metered with its tokens and cost. When an organisation reaches its monthly cap, AI-powered checks stop and attempts are marked inconclusive — never a silent pass.
Stale read
An answer that hasn’t caught up with a recent write yet.
Right after a change, a read may still return the previous state. Verification that trusts one read can raise false alarms — or, worse, trigger a second “fix”.
Stale reads in agent verification →