What each tool is actually for, described plainly. No ratings, no scores, no invented benchmarks.
Evaluation and observability: Knowing whether the thing works and what it did: test harnesses, scoring, tracing and cost tracking.
Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.
github.com/openai/evals🧊 Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23 🍓
github.com/Helicone/helicone🪢 Open source agent evals & observability: Trace, evaluate, and improve LLM applications with one open platform.
github.com/langfuse/langfuseA framework for few-shot evaluation of language models.
github.com/EleutherAI/lm-evaluation-harnessTest your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
github.com/promptfoo/promptfoo