AlpacaEval
OSSAlpacaEval is an open-source tool designed to automatically evaluate the performance of instruction-following language m...
Comparable Evaluation and Monitoring AI agents, ranked by popularity. Not sold on HumanEval? These are the closest options worth a look.
You're comparing against
HumanEval is a benchmark tool designed to assess the functional correctness of AI-generated code, specifically for Python. It provides a set of hand-written programming problems paired with unit tests to verify model outputs. The tool is primarily used to evaluate the performance of code generation models.
AlpacaEval is an open-source tool designed to automatically evaluate the performance of instruction-following language m...
ANN-Benchmarks is a standardized benchmarking framework for evaluating and comparing the performance of approximate near...
ARES is a framework designed for automated evaluation of Retrieval-Augmented Generation (RAG) models, focusing on perfor...
BEIR is a benchmark and evaluation framework for information retrieval (IR) tasks, designed to assess the performance of...
C-Eval is a benchmarking tool designed to evaluate the performance of foundation models on Chinese language tasks. It pr...
The Code Generation LM Evaluation Harness is a framework designed to systematically evaluate the performance of code gen...
COMET is an open-source framework designed for evaluating machine translation models by comparing predictions against hu...
Deepchecks is an open-source Python library designed for validating machine learning models and data integrity throughou...
DeepEval is an open-source framework designed for evaluating large language model (LLM) applications, offering tools to ...
DomainBed is a benchmark suite designed for evaluating domain generalization algorithms in machine learning. It provides...
EvalAI is an open-source platform designed to facilitate large-scale evaluation of AI models by providing a standardized...
Evalchemy provides a standardized framework for evaluating the performance of post-trained language models, offering met...