Said vs did: your AI agent’s reply is not evidence
An agent can be polite, confident and completely wrong about what it did. Here is why COLVO reads the account itself after every action.

The reply that looked fine
A customer writes: “Please stop my renewal, but let me keep access until it runs out.” The agent answers that it is done and access stays until 30 October. Every LLM judge we tried scored that reply as correct.
The reply was right. The action was not. The subscription had been cancelled immediately.
Expected vs observed
COLVO does not ask the agent what happened. It reads Stripe — through its own restricted key — and compares field by field:
| Field | Expected | Observed |
|---|---|---|
| cancel_at_period_end | true | false |
| status | active | canceled |
| access_enabled | true | false |
Four ways replies drift from reality
- The wrong parameter. “Keep access” mapped to an immediate cancel instead of one at the period end.
- A swallowed error. The tool call failed; the agent reported success anyway.
- A refusal reported as done. Guard said DENY, the agent said “refunded”.
- A second action nobody asked for. A retry that proposed the same refund twice.
None of these is visible in the transcript. All of them are visible in the account.
What to do about it
Test the state, not the transcript. Let deterministic checks decide the verdict and keep AI judges advisory: a judge can flag a rude reply, but it must never turn a FAIL into a PASS. And in production, put a gate between the agent and the money, so a wrong action is stopped before it happens instead of explained afterwards.


