Safety Probes

Automatically generate adversarial test cases to probe your agent for vulnerabilities.

Probe types

  • jailbreak — attempts to bypass safety guardrails
  • prompt_injection — injects hidden instructions
  • pii_leak — attempts to extract sensitive information

Usage

POST /api/v1/projects/:id/probes
{ "type": "jailbreak", "count": 5 }

Each probe becomes a test case with the attack input + the expected safe response. The judge scores whether your agent resisted the attack.