Test API
Start a suite, read the results, compare two versions and export the report. Everything the console does, your pipeline can do.
Start a run
POST /v1/test-runs — backend key or session (editor+). Counts against the plan’s monthly runs (402 plan_limit_reached when exhausted).
{ "project_id": "…",
"agent_build_id": "…", // optional: defaults to the project's default build
"scenario_keys": ["refund_over_mandate", "cancel_period_end"], // or scenario_ids; omit = whole library
"repeats": 3, // 1–5 attempts per scenario
"label": "release 0.2" }Read a run
GET /v1/test-runs/{id} — any key or session.
{
"test_run_id": "…", "number": 148, "status": "DONE", "verdict": "FAIL", "repeats": 3,
"counts": { "pass": 17, "fail": 2, "inconclusive": 1, "total": 20 },
"cost": { "known_eur": null, "estimated_eur": "0.6100", "basis": "estimated" },
"scenarios": [{ "scenario_id": "…", "template_key": "cancel_period_end", "name": "Cancel at period end", "family": "cancel",
"verdict": "FAIL", "attempts": [{ "attempt_id": "…", "attempt_no": 1, "status": "DONE", "verdict": "FAIL", "duration_ms": 8120 }, …] }],
"created_at": "…", "started_at": "…", "finished_at": "…"
}status is QUEUED → RUNNING → DONE (or CANCELED). The run verdict is FAIL if any scenario failed, INCONCLUSIVE if none failed but some were inconclusive, otherwise PASS. A scenario passes only when every attempt passes.
Verdicts
- PASS — every deterministic check against the resulting sandbox state held, in every attempt.
- FAIL — a check failed: wrong amount, wrong customer, access removed too early, a forbidden action happened, a double effect.
- INCONCLUSIVE — the state could not be trusted (agent timeout, sandbox fault, unreadable state). Never counted as success. A simulated provider error the agent handled correctly is a PASS, not INCONCLUSIVE.
The semantic judge, when an AI key is configured, adds advisory notes on clarity and instruction-following. It never changes a state-based verdict.
Compare two runs
GET /v1/test-runs/compare?base={run}&head={run} — per scenario: broke, fixed, still_failing, unchanged, inconclusive, added, removed, plus pass rates and which checks newly fail or pass.
Export a report
GET /v1/test-runs/{id}/export?format=json|csv|pdf — secret-free by construction. JSON carries every attempt with expected vs observed; CSV is RFC 4180 with formula injection neutralised; PDF is the readable run report.
Scenarios and the Architect
Templates ship with every project (see the 20). POST /v1/architect/drafts drafts scenarios from a description of your agent (metered AI, or the built-in library without AI); POST /v1/architect/accept adds the ones you approved. POST /v1/redteam/suite generates hostile inputs and runs them (Guard capability).
Schedules
GET / POST /v1/schedules, PATCH / DELETE /v1/schedules/{id}, POST /v1/schedules/{id}/run. Body: project_id, label, cron (5-field), agent_build_id?, scenario_keys[]. Requires the Test plan or higher.