Running Evaluations

Evaluations run asynchronously via a job queue. No LLM I/O on the request thread.

Create a run

POST /api/v1/runs
{ "projectId": "..." }
→ 202 Accepted (returns run ID)

Poll status

GET /api/v1/runs/:id
→ { "run": { "status": "completed", "passCount": 8, ... } }

Full report

GET /api/v1/runs/:id/results

Returns per-case: input, agent response, AI scores, human verdicts, final label, token usage, cost.