A framework for few-shot evaluation of language models.
A framework for few-shot evaluation of language models. Written mainly in Python. Released under the MIT licence.
Evaluation and observability — Knowing whether the thing works and what it did: test harnesses, scoring, tracing and cost tracking.
This entry describes lm-evaluation-harness in its own words, taken from its own page. KuponGuru publishes no ratings or review scores, and claims no affiliate relationship with any tool listed.