For llm performance evaluators, the first results are evidentlyai/evidently (Evidently is an AI observability and LLM evaluation platform that provides automated judge evaluations, custom metrics, tracing, and CI/CD integration for monitoring model performance and behavior), confident-ai/deepeval (Deepeval is an LLM evaluation framework that provides automated LLM-as-a-judge metrics, CI/CD pipeline integration, latency and cost tracking, and execution tracing for testing large language model applications) and comet-ml/comet-llm (Comet LLM provides tracing, monitoring, and evaluation capabilities for large language model applications, though it focuses more heavily on observability and execution tracking than comprehensive benchmark datasets). arize-ai/phoenix and comet-ml/opik round out the shortlist. Compare the match explanations and check the project documentation against your requirements.
Compare open-source LLM performance evaluators on GitHub. Review tools designed to assess large language model accuracy and behavior.
Evidently is an AI observability platform and evaluation framework designed to quantify the performance of machine learning models and large language models. It functions as a monitoring tool for detecting data drift and quality degradation in tabular datasets, while providing a specialized analyzer for the faithfulness and correctness of retrieval augmented generation systems. The project distinguishes itself through an evaluation framework that utilizes judge models and custom rubrics to score language model outputs. It includes tools for iterative prompt optimization and the generation of
Evidently is an AI observability and LLM evaluation platform that provides automated judge evaluations, custom metrics, tracing, and CI/CD integration for monitoring model performance and behavior.
Deepeval is a framework for testing and evaluating large language model applications. It provides a suite of tools for executing automated regression tests, validating model output quality against defined standards, and tracing the execution of complex agent workflows. By integrating these capabilities into development pipelines, the platform ensures consistent performance and reliability throughout the software lifecycle. The platform distinguishes itself through its focus on programmatic validation and observability. It utilizes secondary language models to score output quality and employs
Deepeval is an LLM evaluation framework that provides automated LLM-as-a-judge metrics, CI/CD pipeline integration, latency and cost tracking, and execution tracing for testing large language model applications.
Comet LLM is an observability platform and evaluation framework designed for large language model applications and agentic workflows. It functions as a system for tracing, monitoring, and debugging execution flows while providing tools for prompt optimization and the enforcement of AI safety guardrails. The platform distinguishes itself through a combination of model-based scoring and heuristic metrics to quantify output quality and detect hallucinations. It includes a dedicated prompt and agent optimizer with an interactive playground for refining templates and tool configurations. For retri
Comet LLM provides tracing, monitoring, and evaluation capabilities for large language model applications, though it focuses more heavily on observability and execution tracking than comprehensive benchmark datasets.
Arize Phoenix is an LLM observability platform and evaluation framework designed to capture execution traces and monitor large language model applications. It serves as a prompt management system for versioning and testing templates, and as a self-hosted AI operations infrastructure for managing telemetry and experiments. The platform differentiates itself through a specialized embedding visualization tool used to detect data drift and optimize vector search. It provides a comprehensive evaluation suite that utilizes judge-based evaluators and ground-truth datasets to score model outputs, and
Arize Phoenix is a comprehensive observability platform and evaluation framework that supports LLM-as-a-judge evaluation, tracing, debugging, and performance monitoring for AI applications.
Opik is an observability and evaluation platform designed for generative AI applications and agentic workflows. It provides a centralized environment for tracing execution flows, managing prompt templates, and monitoring production performance, allowing teams to gain visibility into complex model interactions and tool usage without requiring manual application code changes. The platform distinguishes itself through its integrated approach to the AI development lifecycle, combining distributed trace instrumentation with automated evaluation frameworks. It supports model-as-a-judge scoring, syn
Opik is an observability and evaluation platform that provides tracing, automated LLM-as-a-judge evaluation, custom metrics, and performance tracking for generative AI applications.
OpenCompass is a comprehensive evaluation platform, benchmarking suite, and distributed model evaluator designed to measure the performance and accuracy of large language models. It provides a framework for benchmarking both open-source and API-based models against diverse datasets using standardized metrics and reproducible pipelines. The project features an automated judging framework that uses language models as judges to score and verify the quality of generated text. It includes a performance leaderboard system for comparing the relative capabilities of various models across industry-sta
OpenCompass is a comprehensive evaluation platform and benchmarking suite for large language models that features automated LLM-as-a-judge scoring, benchmark datasets, and reproducible configuration-driven pipelines, matching nearly all the requested capabilities.
OpenCompass is an open-source framework for standardized benchmarking of large language models. It provides a configurable evaluation pipeline that supports both objective and subjective assessment, using a dual-engine architecture to handle closed-form answer comparison and open-ended response rating. The framework is designed as a modular platform where datasets, models, and metrics are composed through declarative YAML configuration files. The framework distinguishes itself through its extensible model integration layer, which supports custom models, HuggingFace models, and third-party API
OpenCompass is an LLM evaluation framework providing standardized benchmarking, automated llm-as-a-judge scoring, and a configurable pipeline with rich dataset support, perfectly matching your search for model assessment tools.
Oumi is a comprehensive large language model development platform designed for synthesizing data, fine-tuning models, and running performance evaluations. It serves as a unified environment for the entire model lifecycle, encompassing a training and fine-tuning suite, an evaluation framework, and tools for synthetic data generation and model distillation. The platform is distinguished by its iterative, failure-driven synthesis approach, which analyzes model weaknesses during evaluation to generate targeted training data. It utilizes an LLM-based judge framework to programmatically score respo
Oumi is a comprehensive large language model development platform that includes an evaluation framework, LLM-as-a-judge scoring, and tools for benchmarking and performance analysis.
This project is a self-hosted AI monitoring stack that functions as an LLM observability platform, AI evaluation framework, and OpenTelemetry trace analyzer. It is designed to capture and analyze LLM traces, sessions, and telemetry to monitor AI agent performance. The platform distinguishes itself as a Model Context Protocol server, exposing workspace functions as tools for AI coding agents. It enables the conversion of failing production traces into test datasets for regression testing and utilizes semantic-based session clustering to discover emerging user behavior patterns. The system cov
This self-hosted platform functions as an LLM evaluation framework and observability stack, offering telemetry analysis, trace capture, and cost tracking to monitor and test AI agent performance.
Evals is a framework designed for automating, managing, and executing repeatable benchmarking suites to analyze the quality and performance of language models. It provides a platform for running standardized tests to measure model accuracy and track behavioral changes over time. The system distinguishes itself through a modular architecture that uses a standardized adapter layer to normalize inputs and outputs, allowing different models to be swapped and tested interchangeably. It supports the creation of custom benchmarks using proprietary data, enabling quality assurance on sensitive tasks
This framework provides a structured platform for automating, managing, and executing model evaluations and benchmarking suites, perfectly matching the core requirements for testing language model performance.
Ragas is an evaluation framework designed to measure the performance of retrieval-augmented generation pipelines and autonomous agent workflows. It provides a comprehensive suite of tools for benchmarking system outputs, utilizing language models as automated judges to score performance against defined rubrics and reference data. By standardizing inputs, retrieved contexts, and generated responses into a unified schema, the project enables consistent analysis across complex AI applications. The framework distinguishes itself through its ability to generate synthetic test datasets from existin
Ragas is an evaluation framework built specifically to benchmark retrieval-augmented generation and agent pipelines using automated LLM-as-a-judge metrics, making it a direct fit for this search.
This project is a collection of utilities designed for machine learning experiment tracking, data versioning, and the observability of large language model applications. It provides a client for recording hyperparameters and metrics during training to visualize performance trends and compare different model versions. The tool includes a model evaluation framework that uses custom scorers and automated judges to assess the quality of generated text outputs. It also provides observability tools to monitor and debug the execution flow and runtime behavior of language model applications. The sys
Weights & Biases is a comprehensive machine learning experiment tracking and model observability platform that includes built-in LLM evaluation, custom scorers, and tracing capabilities for monitoring language model applications.
Agenta is a Prompt Ops lifecycle manager and prompt management platform that decouples prompt engineering from application code. It serves as a centralized system for developing, versioning, and deploying prompt templates and model configurations across different environments. The platform functions as an AI agent orchestrator with a visual interface for building agent workflows and connecting models to external tools. It further acts as an evaluation framework and observability tool, utilizing OpenTelemetry to capture execution traces, monitor latency, and track token costs. The system cove
Agenta is a prompt operations and LLM platform that includes evaluation, LLM-as-a-judge workflows, tracing, and latency tracking, fulfilling most of the requested evaluation framework features.
RagaAI-Catalyst is a suite of software implementation tools providing an SDK, dashboard, and platform for monitoring, debugging, red-teaming, and evaluating agentic AI workflows. It serves as an observability framework for tracing the execution paths of large language models and multi-agent systems. The project distinguishes itself through a security suite for automated red-teaming and vulnerability scanning to detect biases, alongside a centralized prompt registry that decouples templates from application code. It further provides an evaluation platform that combines synthetic data generatio
RagaAI-Catalyst is an evaluation and observability platform designed for debugging, monitoring, and testing large language models and agentic workflows, matching the core intent despite lacking explicit coverage for every benchmark dataset feature.
Holistic Evaluation of Language Models (HELM) is an open source Python framework created by the Center for Research on Foundation Models (CRFM) at Stanford for holistic, reproducible and transparent evaluation of foundation models, including large language models (LLMs) and multimodal models.
This framework provides a standardized approach to benchmarking foundation models across multiple custom metrics, serving as a flagship tool for holistic LLM evaluation.
MLflow is an established machine learning platform and model management tool that provides extensive experiment tracking, prompt engineering, and LLM evaluation capabilities, though it is a broader MLOps suite rather than a dedicated evaluation-first framework.
Promptflow is a development framework and orchestrator for building applications powered by large language models. It functions as a suite of tools for designing, orchestrating, and deploying AI workflows by linking prompts, custom Python code, and language models into executable sequences. The project is distinguished by a visual AI workflow designer that allows for the creation of directed acyclic graphs of logic nodes. It provides a dedicated prompt engineering environment for versioning and comparing templates, alongside stateful execution tracing to record function calls and variable val
Promptflow is an LLM development and orchestration framework that includes evaluation tools for benchmarking workflows, though it focuses more heavily on workflow building and execution tracing than dedicated testing suites.
Garak is a suite of tools for measuring AI reliability, scanning for vulnerabilities, and automating security assessments through adaptive probing. It functions as a generative AI vulnerability scanner and evaluation tool designed to identify security gaps, hallucinations, and failure modes in language models. The framework provides a toolkit for red-teaming and safety assessments, utilizing a structured system of probes and detectors to calculate failure rates. It specifically scans for risks such as data leakage and prompt injection by recording model responses to adversarial inputs. The p
Garak is a specialized framework for AI red-teaming and security vulnerability scanning of large language models, making it a relevant tool within the evaluation category despite its primary focus on safety and failure modes rather than standard benchmarks.
Lmnr is an LLM observability platform and evaluation framework designed for tracing, logging, and monitoring language model executions. It provides the tools necessary to debug agent behavior, analyze performance, and identify failure patterns in AI agents. The platform differentiates itself through a trace-to-dataset pipeline that converts production logs into labeled test sets for regression testing. It includes a prompt-variant replay engine to compare different prompts or models side-by-side and a state-cached debugging system to replay agent loops without restarting the process. The sys
Lmnr is an LLM observability platform and evaluation framework that supports tracing, logging, and evaluation, though it leans more toward production monitoring and agent debugging than standard offline benchmarking datasets.
Giskard is an evaluation framework, testing library, and quality monitoring system for large language models and AI agents. It serves as a toolkit for quantifying model performance and reliability, providing specialized capabilities for validating retrieval-augmented generation pipelines. The project distinguishes itself through an automated red teaming tool and security scanner designed to identify vulnerabilities, prompt injections, and safety risks. It utilizes adversarial probing and synthetic edge case generation to quantify model robustness and detect information disclosure. The platfo
Giskard is a Python-based evaluation framework and testing toolkit for large language models that covers quality monitoring, automated vulnerability scanning, and performance validation, though it lacks some specific benchmark datasets and direct CI/CD workflow hooks.
Ragas is an evaluation framework and performance benchmark designed to quantify the quality of retrieval augmented generation pipelines. It functions as an application optimizer to identify bottlenecks in language model workflows using automated metrics and model-based scoring. The framework includes a system for generating synthetic datasets that mimic production scenarios and edge cases to create realistic test cases. It enables reference-free assessment, allowing the evaluation of response quality by analyzing grounding in the provided context without requiring gold-standard labels. The s
Ragas is an evaluation framework built specifically to assess retrieval-augmented generation pipelines using automated model-based metrics and synthetic test data generation, though it has a narrower domain focus than a general-purpose LLM evaluator.
Automatic Prompt Engineer is a framework designed to automate the generation, refinement, and performance measurement of language model instructions. It functions as a systematic tool for optimizing prompt phrasing by iteratively testing candidate instructions against specific input and output datasets to maximize task accuracy. The system distinguishes itself through an evaluation-driven approach that uses automated feedback loops to score prompt variations. By employing template-based input structuring, it ensures consistent testing environments where candidate instructions are measured aga
Automatic Prompt Engineer provides an evaluation-driven framework for measuring and optimizing language model instructions against datasets, though it focuses specifically on prompt engineering rather than a general-purpose LLM monitoring suite.
Langfuse is an open-source observability and evaluation platform designed for language model applications. It provides a centralized system for tracking execution traces, monitoring performance metrics, and managing prompt templates. By capturing hierarchical units of work and telemetry data, the platform enables developers to debug complex application lifecycles and analyze token usage, latency, and model interactions in production environments. The platform distinguishes itself through an integrated evaluation framework that allows for systematic benchmarking and automated scoring of model
Langfuse is an open-source observability and analytics platform for LLM applications that includes evaluation features, though it focuses more heavily on production tracing, cost tracking, and monitoring than a dedicated benchmarking suite.
This project is a standardized framework for benchmarking large language models across a wide range of academic and reasoning datasets. It provides a platform for executing automated evaluation tasks to measure model accuracy and performance, ensuring consistent assessment through a structured configuration schema. The framework distinguishes itself by incorporating a dedicated utility for data decontamination, which identifies and removes overlapping training samples from evaluation sets to prevent data leakage. It also features a flexible task builder that allows users to define custom benc
This project is a standardized framework for benchmarking large language models across numerous datasets and metrics, fitting the evaluation category well despite leaning more toward academic benchmarking than live CI/CD monitoring.
A framework for the evaluation of autoregressive code generation language models.
BigCode Evaluation Harness is a specialized framework for benchmarking autoregressive code generation language models, though it is tailored for code tasks rather than general-purpose LLM monitoring.
Assess, Guard, and Monitor Your LLM Applications Built by Future AGI | Docs | Platform
This repository is an evaluation framework for assessing and guarding LLM applications, though it is missing explicit details on benchmark datasets and CI/CD integration.
LLM Comparator is an interactive data visualization tool for evaluating and analyzing LLM responses side-by-side, developed by the PAIR team.
LLM Comparator is an interactive evaluation tool for side-by-side analysis of model responses, though it focuses more on visual comparison than automated CI/CD pipelines and benchmarking datasets.
| Repository | Stars | Language | License | Last push |
|---|---|---|---|---|
| evidentlyai/evidently | 7.1K | Jupyter Notebook | apache-2.0 | |
| confident-ai/deepeval | 13.7K | Python | apache-2.0 | |
| 19.7K |
| Python |
| Apache-2.0 |
| arize-ai/phoenix | 8.6K | Jupyter Notebook | other |
| comet-ml/opik | 17.8K | Python | apache-2.0 |
| internlm/opencompass | 7.1K | Python | Apache-2.0 |
| open-compass/opencompass | 6.7K | Python | apache-2.0 |
| oumi-ai/oumi | 8.9K | Python | apache-2.0 |
| latitude-dev/latitude-llm | 4.1K | TypeScript | MIT |
| openai/evals | 18.7K | Python | NOASSERTION |