What's new in EvalDesk
Every improvement, new feature, and bug fix — documented.
v0.8.0
October 8, 2026Latest- FeatureMulti-model judge ensemble (Anthropic Claude, OpenAI GPT-4o, DeepSeek) with rubric-based evaluation and consensus scoring
- FeatureEd25519 cryptographically signed compliance certificates with offline verification CLI (evaldesk-verify)
- FeatureOfficial Python SDK (evaldesk) with Pytest assertion gates (assert_run_passes)
- ImprovementAutomated schema codegen and dual PostgreSQL / SQLite database driver parity
- FixResolved session token validation edge cases and sticky layout navigation
v0.7.0
September 18, 2026- FeatureAutomated HIPAA Security Rule (45 CFR § 164.312) & EU AI Act compliance packs
- FeatureRAG Faithfulness scoring with context-grounding hallucination detection
- FeatureAdversarial red-team safety probes (jailbreaks, prompt injection, PII/PHI leak detection)
- ImprovementMulti-model judge ensemble with Cohen's & Fleiss' Kappa agreement metrics
- FixOptimized database query indexing on run results for sub-second report generation
v0.6.0
August 21, 2026- FeatureClaude 4.5 Sonnet judge integration with chain-of-thought verification
- FeatureHuman-in-the-loop expert review workflow with blind dual-annotation
- FeatureStreaming response evaluation with token-level timing and TTFT tracking
- ImprovementReduced memory overhead for high-concurrency 10,000+ test case batches
- FixFixed SSE connection dropouts during long-running batch evaluation runs
v0.5.0
July 24, 2026- FeaturePrebuilt Medical Triage & Clinical Decision benchmark packs
- FeatureCustom evaluation rubrics with weighted scoring criteria and markdown guidelines
- FeatureWebhook trigger system for automated CI/CD pipeline blocking
- ImprovementEnhanced audit log trail with immutable SHA-256 event hashing
- FixResolved pagination state loss when filtering benchmark results
v0.4.0
June 19, 2026- FeatureAgent tool call and function invocation validation against JSON Schema definitions
- FeatureSemantic similarity judge using calibrated cross-encoder embeddings
- FeatureRole-based access control (RBAC) with Owner, Admin, and Reviewer permissions
- ImprovementDark / light theme system with adaptive UI components
- FixResolved webhook retry logic for intermittent HTTP 5xx responses
v0.3.0
May 15, 2026- FeatureMulti-model judge consensus for high-confidence evaluation decisions
- FeatureSafety scoring with toxicity, hate speech, and PII leakage detection
- FeatureCitation and source-grounding verification for RAG agent responses
- ImprovementFaster test execution with parallel agent calls and connection pooling
- FixFixed edge-case race condition in concurrent run dispatching
v0.2.0
April 18, 2026- FeatureMulti-turn conversation testing support with dynamic scenario branching
- FeatureStreaming response evaluation with token-level timing metrics
- FeatureScheduled recurring evaluation runs with cron expressions
- ImprovementExpanded judge criteria library with domain-specific clinical and financial templates
- FixResolved session expiry and token refresh issues on long-running evaluations
v0.1.0
March 15, 2026- FeatureInitial release with core agent evaluation engine
- FeatureLLM-powered judge with customizable criteria and scoring rubrics
- FeatureTest case management with categories, tags, and bulk CSV/JSON import
- FeatureRun history with pass/fail analytics and latency breakdowns
- ImprovementProject-based organization and scoped API key authentication