deepeval
The LLM Evaluation Framework
Written mainly in Python. Released under the Apache-2.0 licence.
What each tool is actually for, described plainly. No ratings, no scores, no invented benchmarks.
Evaluation and observability: Knowing whether the thing works and what it did: test harnesses, scoring, tracing and cost tracking.
The LLM Evaluation Framework
Written mainly in Python. Released under the Apache-2.0 licence.
Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.
Written mainly in Python.
🧊 Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23 🍓
Written mainly in TypeScript. Released under the Apache-2.0 licence.
🪢 Open source agent evals & observability: Trace, evaluate, and improve LLM applications with one open platform.
Written mainly in TypeScript.
A framework for few-shot evaluation of language models.
Written mainly in Python. Released under the MIT licence.
AI Observability & Evaluation
Written mainly in Python.
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
Written mainly in TypeScript. Released under the MIT licence.
Supercharge Your LLM Application Evaluations 🚀
Written mainly in Python. Released under the Apache-2.0 licence.