category hub
Model Evaluation & Benchmarking
Test models and RAG pipelines for accuracy, robustness, bias, and safety before and after deployment.
8 tools
Paid
Braintrust
Evaluation platform for LLM apps: evals, logging, and prompt experiment tracking.
Enterprise
Coval
Simulation and evaluation platform for AI voice and chat agents (YC-backed; $28M Series A, Jun 2026): pre-launch stress testing plus production monitoring, built by Waymo eval alumni.
Open Source
DeepEval
Open-source LLM evaluation framework with pytest-style unit testing for AI applications.
Open Source
Giskard
Open-source testing framework that scans ML models and LLM apps for vulnerabilities and quality issues.
Open Source
Langfuse
Open-source LLM engineering platform (YC W23, MIT license): tracing, LLM-as-judge evals, prompt management, and datasets — the most widely adopted OSS option; part of ClickHouse since Jan 2026.
Paid
Patronus AI
Automated evaluation and guardrails platform for scoring and monitoring LLM system failures.
Open Source
Promptfoo
CLI and library for systematic LLM testing, red-teaming, and eval-driven development — acquired by OpenAI in 2026.
Open Source
Ragas
Open-source evaluation framework purpose-built for retrieval-augmented generation pipelines.
faq
Frequently asked questions
- What does Model Evaluation & Benchmarking cover?
- Model Evaluation & Benchmarking groups tools that share a primary capability area. Use this hub to compare options, see pricing models, and shortlist before booking demos.
- How is the Model Evaluation & Benchmarking list maintained?
- Each tool is reviewed against vendor docs and trust pages. The "last verified" date on each tool page reflects the most recent review.
- Can I submit a tool for Model Evaluation & Benchmarking?
- Yes. Use the Submit page to send a vendor link; we review it against the directory criteria.
