Running Evaluations
Evaluations run asynchronously via a job queue. No LLM I/O on the request thread.
Create a run
POST /api/v1/runs
{ "projectId": "..." }
→ 202 Accepted (returns run ID)
Poll status
GET /api/v1/runs/:id
→ { "run": { "status": "completed", "passCount": 8, ... } }
Full report
GET /api/v1/runs/:id/results
Returns per-case: input, agent response, AI scores, human verdicts, final label, token usage, cost.