Making AI evaluation accessible to everyone
We started EvalDesk because we saw teams struggling to answer a simple question: is my AI agent actually working correctly?
Our Story
In 2024, our founding team was building AI agents for enterprise clients and kept running into the same problem: there was no good way to systematically test whether an agent was performing correctly. Manual review was slow and inconsistent. Existing benchmarks were too generic to be useful for domain-specific applications. Teams were shipping agents to production with little confidence in their reliability.
We built EvalDesk to solve this. Our platform lets teams create test suites tailored to their specific use cases, run evaluations automatically with LLM-powered judging, and catch regressions before they reach users. Whether you are building a customer support bot, a medical triage agent, or a financial advisor, EvalDesk gives you the confidence that your AI works as intended.
Meet the team
A small, focused team passionate about building reliable AI.
Raman Dagar
Founder & Lead Architect
Architecting agent evaluation, automated compliance, and cryptographic verification infrastructure for mission-critical production AI systems.
Community & Contributors
Open Source Collective
Engineers, security researchers, and healthcare/fintech domain experts collaborating across safety benchmarks, OTel tracing, and compliance packs.
What we believe
Open Source First
We believe the best evaluation tools should be available to everyone. EvalDesk is built in the open, and our core evaluation engine will always be open source. Transparency builds trust.
Domain Expert Focus
Generic benchmarks do not capture real-world performance. We build domain-specific evaluation criteria so teams in healthcare, finance, legal, and other fields can test what actually matters.
Privacy by Design
Your test cases, agent responses, and evaluation data stay under your control. We never train models on your data, and all sensitive information is encrypted at rest and in transit.
Join us on this mission
Start evaluating your AI agents today. It takes less than 5 minutes to set up your first test suite.