AlpacaEval
OSSAlpacaEval is an open-source tool designed to automatically evaluate the performance of instruction-following language m...
Showing 50 agents
See our top picks: Best Evaluation and Monitoring AI Agents (2026) →AlpacaEval is an open-source tool designed to automatically evaluate the performance of instruction-following language m...
ANN-Benchmarks is a standardized benchmarking framework for evaluating and comparing the performance of approximate near...
ARES is a framework designed for automated evaluation of Retrieval-Augmented Generation (RAG) models, focusing on perfor...
BEIR is a benchmark and evaluation framework for information retrieval (IR) tasks, designed to assess the performance of...
C-Eval is a benchmarking tool designed to evaluate the performance of foundation models on Chinese language tasks. It pr...
The Code Generation LM Evaluation Harness is a framework designed to systematically evaluate the performance of code gen...
COMET is an open-source framework designed for evaluating machine translation models by comparing predictions against hu...
Deepchecks is an open-source Python library designed for validating machine learning models and data integrity throughou...
DeepEval is an open-source framework designed for evaluating large language model (LLM) applications, offering tools to ...
DomainBed is a benchmark suite designed for evaluating domain generalization algorithms in machine learning. It provides...
EvalAI is an open-source platform designed to facilitate large-scale evaluation of AI models by providing a standardized...
Evalchemy provides a standardized framework for evaluating the performance of post-trained language models, offering met...
EvalPlus is a specialized evaluation framework designed for assessing large language models (LLMs) in code generation ta...
Evals is a framework developed by OpenAI for systematically evaluating the performance of AI models, particularly those ...
EvalScope is an open-source framework designed for evaluating and benchmarking large AI models. It provides customizable...
Evaluate is an open-source library designed to streamline the evaluation and comparison of machine learning models by pr...
GAOKAO-Bench is an open-source evaluation framework that leverages questions from the Chinese National College Entrance ...
guidellm is a benchmarking tool designed to evaluate the performance of large language model (LLM) inference systems. It...
Helicone is an open-source platform designed to help developers monitor, evaluate, and optimize large language model (LL...
HumanEval is a benchmark tool designed to assess the functional correctness of AI-generated code, specifically for Pytho...
Inspect is an AI agent in the Evaluation and Monitoring category. ![](https://img.shields.io/github/stars/UKGovernmentBE...
JiWER is an AI agent in the Evaluation and Monitoring category. ![](https://img.shields.io/github/stars/jitsi/jiwer.svg?...
Laminar is an AI agent in the Evaluation and Monitoring category. ![](https://img.shields.io/github/stars/lmnr-ai/lmnr.s...
LangTest is an AI agent in the Evaluation and Monitoring category. ![](https://img.shields.io/github/stars/JohnSnowLabs/...
Language Model Evaluation Harness is an AI agent in the Evaluation and Monitoring category. ![](https://img.shields.io/g...
LLMPerf is an AI agent in the Evaluation and Monitoring category. ![](https://img.shields.io/github/stars/ray-project/ll...
lmms-eval is an AI agent in the Evaluation and Monitoring category. ![](https://img.shields.io/github/stars/EvolvingLMMs...
Massive Text Embedding Benchmark is an AI agent in the Evaluation and Monitoring category. ![](https://img.shields.io/gi...
Melting Pot is an AI agent in the Evaluation and Monitoring category. ![](https://img.shields.io/github/stars/google-dee...
Meta-World is an AI agent in the Evaluation and Monitoring category. ![](https://img.shields.io/github/stars/Farama-Foun...
mir_eval is an AI agent in the Evaluation and Monitoring category. ![](https://img.shields.io/github/stars/mir-evaluatio...
MLPerf Inference is an AI agent in the Evaluation and Monitoring category. ![](https://img.shields.io/github/stars/mlcom...
NannyML is an AI agent in the Evaluation and Monitoring category. ![](https://img.shields.io/github/stars/NannyML/nannym...
OGB is an AI agent in the Evaluation and Monitoring category. ![](https://img.shields.io/github/stars/snap-stanford/ogb....
Ollama Grid Search is an AI agent in the Evaluation and Monitoring category. ![](https://img.shields.io/github/stars/dez...
OpenCompass is an AI agent in the Evaluation and Monitoring category. ![](https://img.shields.io/github/stars/open-compa...
Overcooked-AI is an AI agent in the Evaluation and Monitoring category. ![](https://img.shields.io/github/stars/HumanCom...
Prometheus-Eval is an AI agent in the Evaluation and Monitoring category. ![](https://img.shields.io/github/stars/promet...
PromptBench is an AI agent in the Evaluation and Monitoring category. ![](https://img.shields.io/github/stars/microsoft/...
RagaAI Catalyst is an AI agent in the Evaluation and Monitoring category. ![](https://img.shields.io/github/stars/raga-a...
RewardBench is an AI agent in the Evaluation and Monitoring category. ![](https://img.shields.io/github/stars/allenai/re...
RLBench is an AI agent in the Evaluation and Monitoring category. ![](https://img.shields.io/github/stars/stepjam/RLBenc...
SimplerEnv is an AI agent in the Evaluation and Monitoring category. ![](https://img.shields.io/github/stars/simpler-env...
Speech-to-Text Benchmark is an AI agent in the Evaluation and Monitoring category. ![](https://img.shields.io/github/sta...
SwanLab is an AI agent in the Evaluation and Monitoring category. ![](https://img.shields.io/github/stars/SwanHubX/SwanL...
TorchBench is an AI agent in the Evaluation and Monitoring category. ![](https://img.shields.io/github/stars/pytorch/ben...
TruLens is an AI agent in the Evaluation and Monitoring category. ![](https://img.shields.io/github/stars/truera/trulens...
TrustLLM is an AI agent in the Evaluation and Monitoring category. ![](https://img.shields.io/github/stars/HowieHwong/Tr...
VBench is an AI agent in the Evaluation and Monitoring category. ![](https://img.shields.io/github/stars/Vchitect/VBench...
VLMEvalKit is an AI agent in the Evaluation and Monitoring category. ![](https://img.shields.io/github/stars/open-compas...