Test API

Start a suite, read the results, compare two versions and export the report. Everything the console does, your pipeline can do.

Written from the code · updated 26 Sep 2026 · Something wrong or missing? Tell us

Start a run

POST /v1/test-runs — backend key or session (editor+). Counts against the plan’s monthly runs (402 plan_limit_reached when exhausted).

{ "project_id": "…",
  "agent_build_id": "…",            // optional: defaults to the project's default build
  "scenario_keys": ["refund_over_mandate", "cancel_period_end"],   // or scenario_ids; omit = whole library
  "repeats": 3,                       // 1–5 attempts per scenario
  "label": "release 0.2" }

Read a run

GET /v1/test-runs/{id} — any key or session.

{
  "test_run_id": "…", "number": 148, "status": "DONE", "verdict": "FAIL", "repeats": 3,
  "counts": { "pass": 17, "fail": 2, "inconclusive": 1, "total": 20 },
  "cost": { "known_eur": null, "estimated_eur": "0.6100", "basis": "estimated" },
  "scenarios": [{ "scenario_id": "…", "template_key": "cancel_period_end", "name": "Cancel at period end", "family": "cancel",
                  "verdict": "FAIL", "attempts": [{ "attempt_id": "…", "attempt_no": 1, "status": "DONE", "verdict": "FAIL", "duration_ms": 8120 }, …] }],
  "created_at": "…", "started_at": "…", "finished_at": "…"
}

status is QUEUED → RUNNING → DONE (or CANCELED). The run verdict is FAIL if any scenario failed, INCONCLUSIVE if none failed but some were inconclusive, otherwise PASS. A scenario passes only when every attempt passes.

Verdicts

  • PASS — every deterministic check against the resulting sandbox state held, in every attempt.
  • FAIL — a check failed: wrong amount, wrong customer, access removed too early, a forbidden action happened, a double effect.
  • INCONCLUSIVE — the state could not be trusted (agent timeout, sandbox fault, unreadable state). Never counted as success. A simulated provider error the agent handled correctly is a PASS, not INCONCLUSIVE.

The semantic judge, when an AI key is configured, adds advisory notes on clarity and instruction-following. It never changes a state-based verdict.

Compare two runs

GET /v1/test-runs/compare?base={run}&head={run} — per scenario: broke, fixed, still_failing, unchanged, inconclusive, added, removed, plus pass rates and which checks newly fail or pass.

Export a report

GET /v1/test-runs/{id}/export?format=json|csv|pdf — secret-free by construction. JSON carries every attempt with expected vs observed; CSV is RFC 4180 with formula injection neutralised; PDF is the readable run report.

Scenarios and the Architect

Templates ship with every project (see the 20). POST /v1/architect/drafts drafts scenarios from a description of your agent (metered AI, or the built-in library without AI); POST /v1/architect/accept adds the ones you approved. POST /v1/redteam/suite generates hostile inputs and runs them (Guard capability).

Schedules

GET / POST /v1/schedules, PATCH / DELETE /v1/schedules/{id}, POST /v1/schedules/{id}/run. Body: project_id, label, cron (5-field), agent_build_id?, scenario_keys[]. Requires the Test plan or higher.

Test API · Docs · COLVO