Giskard
OSSGiskard is an open-source library designed for testing and evaluating large language model (LLM) applications, with a fo...
Comparable LLM Evaluation AI agents, ranked by popularity. Not sold on lighteval? These are the closest options worth a look.
You're comparing against
lighteval is a lightweight evaluation suite designed for Large Language Models (LLMs), providing efficient tools for benchmarking and assessing model performance. It focuses on simplicity and speed, enabling quick iterations during model development and testing. The tool is optimized for internal use but has been open-sourced for broader community adoption.
Giskard is an open-source library designed for testing and evaluating large language model (LLM) applications, with a fo...
HELM (Holistic Evaluation of Language Models) is a comprehensive benchmarking framework designed to evaluate the perform...
instruct-eval is an open-source tool designed to quantitatively evaluate instruction-tuned language models like Alpaca a...
LangSmith is a platform integrated with the LangChain framework, designed for evaluating, monitoring, and collaborating ...
lm-evaluation-harness is a framework designed for standardized and reproducible evaluation of language models (LMs) usin...
MixEval is an open-source evaluation suite designed for benchmarking large language models (LLMs), supporting both open-...
OLMO-eval is an open-source toolkit designed for evaluating open language models (OLMs), providing standardized benchmar...
Ragas is an open-source framework designed to evaluate Retrieval Augmented Generation (RAG) pipelines by providing metri...
simple-evals is an open-source toolkit by OpenAI designed for evaluating large language models (LLMs). It provides stand...