AlpacaEval
OSSAlpacaEval is an open-source tool designed to automatically evaluate the performance of instruction-following language m...
Comparable Evaluation and Monitoring AI agents, ranked by popularity. Not sold on DomainBed? These are the closest options worth a look.
You're comparing against
DomainBed is a benchmark suite designed for evaluating domain generalization algorithms in machine learning. It provides standardized datasets and evaluation protocols to test how well models perform on unseen domains.
AlpacaEval is an open-source tool designed to automatically evaluate the performance of instruction-following language m...
ANN-Benchmarks is a standardized benchmarking framework for evaluating and comparing the performance of approximate near...
ARES is a framework designed for automated evaluation of Retrieval-Augmented Generation (RAG) models, focusing on perfor...
BEIR is a benchmark and evaluation framework for information retrieval (IR) tasks, designed to assess the performance of...
C-Eval is a benchmarking tool designed to evaluate the performance of foundation models on Chinese language tasks. It pr...
The Code Generation LM Evaluation Harness is a framework designed to systematically evaluate the performance of code gen...
COMET is an open-source framework designed for evaluating machine translation models by comparing predictions against hu...
Deepchecks is an open-source Python library designed for validating machine learning models and data integrity throughou...
DeepEval is an open-source framework designed for evaluating large language model (LLM) applications, offering tools to ...
EvalAI is an open-source platform designed to facilitate large-scale evaluation of AI models by providing a standardized...
Evalchemy provides a standardized framework for evaluating the performance of post-trained language models, offering met...
EvalPlus is a specialized evaluation framework designed for assessing large language models (LLMs) in code generation ta...