AI Judge

The AI judge scores every agent answer on a pass/fail/partial scale. Configure it per-project with any OpenAI-compatible or Anthropic endpoint.

Supported providers

  • Anthropic Claude (claude-sonnet-5-5, claude-opus-5, claude-sonnet-4-5)
  • OpenAI (gpt-4o, gpt-4o-mini)
  • DeepSeek (deepseek-chat, deepseek-reasoner)
  • OpenRouter (any open-weights or hosted model)
  • Ollama / vLLM (local private deployment)

Honest confidence

Confidence is computed from self-consistency sampling and cross-judge agreement — never the model's self-reported number.

Routing

Low-confidence, disagreeing, audit-sampled, or adversarial items automatically route to human review.