Giskard
OSSGiskard is an open-source library designed for testing and evaluating large language model (LLM) applications, with a fo...
Comparable LLM Evaluation AI agents, ranked by popularity. Not sold on simple-evals? These are the closest options worth a look.
You're comparing against
simple-evals is an open-source toolkit by OpenAI designed for evaluating large language models (LLMs). It provides standardized methods and metrics to assess model performance, consistency, and reliability across different tasks. The tool is modular, allowing users to customize evaluations for specific needs.
Giskard is an open-source library designed for testing and evaluating large language model (LLM) applications, with a fo...
HELM (Holistic Evaluation of Language Models) is a comprehensive benchmarking framework designed to evaluate the perform...
instruct-eval is an open-source tool designed to quantitatively evaluate instruction-tuned language models like Alpaca a...
LangSmith is a platform integrated with the LangChain framework, designed for evaluating, monitoring, and collaborating ...
lighteval is a lightweight evaluation suite designed for Large Language Models (LLMs), providing efficient tools for ben...
lm-evaluation-harness is a framework designed for standardized and reproducible evaluation of language models (LMs) usin...
MixEval is an open-source evaluation suite designed for benchmarking large language models (LLMs), supporting both open-...
OLMO-eval is an open-source toolkit designed for evaluating open language models (OLMs), providing standardized benchmar...
Ragas is an open-source framework designed to evaluate Retrieval Augmented Generation (RAG) pipelines by providing metri...