The payment API timed out, the agent retried, and the customer got two refunds. How it happens, how to test for it, and how to stop it in production.
The agent calls “refund”. The payment provider processes it, but the response is lost on the way back — a timeout, a dropped connection, a restarted worker.
The agent sees an error and, reasonably, tries again. If the second call carries a new idempotency key (or none), the provider treats it as a new refund.
The same thing happens when a customer sends the same request twice, or two copies arrive at the same time: two conversations, two refunds.
COLVO’s sandbox injects exactly this fault: the write succeeds and the response is dropped. The scenario passes only if the provider state shows one refund after every retry.
Duplicate and concurrent customer requests are separate scenarios, because they fail differently from a network retry.
Every scenario runs three times in a fresh sandbox; a pass needs all three.
Guard derives one idempotency key from the business request (the mandate and its source request), so every retry of the same request maps to the same provider call.
The budget is reserved before the call is sent and the state is written to an outbox, so a crash between “sent” and “recorded” cannot create a second effect.
After the call, Guard re-reads the provider independently and marks the operation verified, mismatched or unverifiable — never “done” on the reply alone.