Step checks: the state was right, the agent wasn’t
Stripe looked perfect. The agent had still tried to cancel twice. Why we now read the agent’s OpenTelemetry trace — and why a missing trace never fails a test.

State checks answer the most important question: did the account end up right? But sometimes it ends up right by luck.
Right state, wrong steps
In one run the agent proposed the same cancellation twice. Guard recognised the duplicate and did nothing the second time, so Stripe looked perfect. In production, with a refund instead of a cancellation and a different idempotency key, the same habit could cost money.
What a step check is
A rule on how the agent got there, set per scenario:
- a tool is called before another (look up the subscription before cancelling);
- a tool is called at most N times, or never;
- the agent replies only after COLVO decided;
- no step ended in an error.
Step checks are deterministic and can only add a FAIL. They never turn a FAIL into a PASS — the state still decides.
Where the steps come from
Your agent exports its OpenTelemetry trace to COLVO during the test, with the per-test key from the request. We follow the GenAI conventions (chat … and execute_tool … spans) and keep only names, timings, models and token counts — never prompts, replies or tool arguments. With the OpenAI preset COLVO runs the conversation itself, so it records every step without any setup.
Proposals are always counted by COLVO, and per action the source that saw more calls wins: a retry Guard deduplicated still counts.
No trace? No penalty
If the agent sends no trace, checks on its own tools show as “not checked”. A missing trace is never a FAIL.


