For framework pour l'évaluation d'applications LLM, the strongest matches are typpo/promptfoo (promptfoo is a dedicated evaluation framework for LLM prompts), comet-ml/comet-llm (Comet LLM is a dedicated evaluation framework for LLM) and evidentlyai/evidently (Evidently is an AI observability and evaluation framework specifically). giskard-ai/giskard and oumi-ai/oumi round out the shortlist. Each is ranked by relevance to your query, popularity and recent activity.
Ces outils open source fournissent des suites de tests automatisés et des métriques pour évaluer les performances des modèles de langage.
promptfoo is an evaluation framework for measuring the performance of large language model prompts, agents, and retrieval augmented generation pipelines. It provides a suite of tools for conducting comparative benchmarking and executing automated quality and security regressions. The system features a benchmarking suite for running identical prompts across different model providers to compare output quality side-by-side. It also includes a dedicated red teaming tool for identifying security vulnerabilities and prompt injection risks through automated penetration testing. The framework suppor
promptfoo is a dedicated evaluation framework for LLM prompts, agents, and RAG pipelines, offering benchmarking, automated assertion validation, red teaming for security, CI/CD integration, multi‑model comparison, and custom scoring pipelines—directly addressing the need for a comprehensive performance, safety, and quality evaluation tool.
Comet LLM is an observability platform and evaluation framework designed for large language model applications and agentic workflows. It functions as a system for tracing, monitoring, and debugging execution flows while providing tools for prompt optimization and the enforcement of AI safety guardrails. The platform distinguishes itself through a combination of model-based scoring and heuristic metrics to quantify output quality and detect hallucinations. It includes a dedicated prompt and agent optimizer with an interactive playground for refining templates and tool configurations. For retri
Comet LLM is a dedicated evaluation framework for LLM applications that provides model-based scoring, heuristic metrics, hallucination detection, prompt optimization, and CI/CD integration—directly covering the test case generation, assertion, model-graded evaluation, and custom metrics needed for assessing performance, safety, and quality.
Evidently is an AI observability platform and evaluation framework designed to quantify the performance of machine learning models and large language models. It functions as a monitoring tool for detecting data drift and quality degradation in tabular datasets, while providing a specialized analyzer for the faithfulness and correctness of retrieval augmented generation systems. The project distinguishes itself through an evaluation framework that utilizes judge models and custom rubrics to score language model outputs. It includes tools for iterative prompt optimization and the generation of
Evidently is an AI observability and evaluation framework specifically for LLMs, offering model-graded scoring, custom rubrics, RAG evaluation, and CI/CD integration, which squarely meets the need for building LLM evaluation pipelines.
Giskard is an evaluation framework, testing library, and quality monitoring system for large language models and AI agents. It serves as a toolkit for quantifying model performance and reliability, providing specialized capabilities for validating retrieval-augmented generation pipelines. The project distinguishes itself through an automated red teaming tool and security scanner designed to identify vulnerabilities, prompt injections, and safety risks. It utilizes adversarial probing and synthetic edge case generation to quantify model robustness and detect information disclosure. The platfo
Giskard is a dedicated evaluation framework and testing library for LLMs and AI agents, offering automated test case generation, adversarial probing, security scanning, and quality monitoring that directly supports building evaluation pipelines for performance, safety, and quality assessment.
Oumi is a comprehensive large language model development platform designed for synthesizing data, fine-tuning models, and running performance evaluations. It serves as a unified environment for the entire model lifecycle, encompassing a training and fine-tuning suite, an evaluation framework, and tools for synthetic data generation and model distillation. The platform is distinguished by its iterative, failure-driven synthesis approach, which analyzes model weaknesses during evaluation to generate targeted training data. It utilizes an LLM-based judge framework to programmatically score respo
Oumi is a comprehensive LLM development platform with a built-in evaluation framework that uses an LLM-based judge for scoring and a failure-driven approach to generate targeted test cases, directly supporting the kind of evaluation pipelines needed to assess performance, safety, and quality across multiple models.
Promptfoo is an evaluation framework designed for testing, benchmarking, and red-teaming language models and agentic workflows. It provides a unified environment to run prompts against multiple providers, allowing developers to systematically validate model outputs against objective assertions, semantic similarity metrics, and custom grading rubrics. The platform distinguishes itself through a provider-agnostic execution layer and a stateful orchestrator capable of simulating multi-turn conversations and complex tool-use trajectories. It includes a dedicated adversarial mutation pipeline that
Promptfoo is a fully-featured evaluation framework for LLMs that supports test case generation, assertion-based and model-graded evaluation, multi-model comparisons, and integrates with CI/CD pipelines, making it a strong fit for building comprehensive evaluation pipelines.
Deepeval is a framework for testing and evaluating large language model applications. It provides a suite of tools for executing automated regression tests, validating model output quality against defined standards, and tracing the execution of complex agent workflows. By integrating these capabilities into development pipelines, the platform ensures consistent performance and reliability throughout the software lifecycle. The platform distinguishes itself through its focus on programmatic validation and observability. It utilizes secondary language models to score output quality and employs
Deepeval is a dedicated framework for testing and evaluating LLM applications, with built-in support for assertion-based and model-graded evaluation, CI/CD integration, and observability—exactly what you need to build evaluation pipelines for LLM performance and safety.
Agenta is a Prompt Ops lifecycle manager and prompt management platform that decouples prompt engineering from application code. It serves as a centralized system for developing, versioning, and deploying prompt templates and model configurations across different environments. The platform functions as an AI agent orchestrator with a visual interface for building agent workflows and connecting models to external tools. It further acts as an evaluation framework and observability tool, utilizing OpenTelemetry to capture execution traces, monitor latency, and track token costs. The system cove
Agenta is an LLM evaluation and prompt management platform that supports model-graded evaluation, observability, and CI/CD integration, making it a comprehensive solution for building evaluation pipelines for LLM applications.
OpenCompass is a comprehensive evaluation platform, benchmarking suite, and distributed model evaluator designed to measure the performance and accuracy of large language models. It provides a framework for benchmarking both open-source and API-based models against diverse datasets using standardized metrics and reproducible pipelines. The project features an automated judging framework that uses language models as judges to score and verify the quality of generated text. It includes a performance leaderboard system for comparing the relative capabilities of various models across industry-sta
OpenCompass is a comprehensive benchmarking suite and evaluation platform for LLMs, supporting multi-model evaluation and LLM-as-a-judge scoring, though its focus on standardized benchmarking rather than custom pipeline construction and test case generation means it may not cover all the pipeline-building features you listed.
OpenCompass is an open-source framework for standardized benchmarking of large language models. It provides a configurable evaluation pipeline that supports both objective and subjective assessment, using a dual-engine architecture to handle closed-form answer comparison and open-ended response rating. The framework is designed as a modular platform where datasets, models, and metrics are composed through declarative YAML configuration files. The framework distinguishes itself through its extensible model integration layer, which supports custom models, HuggingFace models, and third-party API
OpenCompass is a modular framework for benchmarking LLMs through configurable evaluation pipelines, supporting objective and subjective assessment with multi-model and custom-metric extensibility—making it a fitting tool for building LLM evaluation workflows, even if it focuses more on standardized benchmarks than on app-specific safety or CI/CD integration.
This project is a standardized framework for benchmarking large language models across a wide range of academic and reasoning datasets. It provides a platform for executing automated evaluation tasks to measure model accuracy and performance, ensuring consistent assessment through a structured configuration schema. The framework distinguishes itself by incorporating a dedicated utility for data decontamination, which identifies and removes overlapping training samples from evaluation sets to prevent data leakage. It also features a flexible task builder that allows users to define custom benc
The LM Evaluation Harness is a standardized framework for benchmarking LLMs with configurable tasks, multi-model support, and built-in evaluation utilities, making it a strong foundation for building evaluation pipelines that assess model performance and quality.
Arize Phoenix is an LLM observability platform and evaluation framework designed to capture execution traces and monitor large language model applications. It serves as a prompt management system for versioning and testing templates, and as a self-hosted AI operations infrastructure for managing telemetry and experiments. The platform differentiates itself through a specialized embedding visualization tool used to detect data drift and optimize vector search. It provides a comprehensive evaluation suite that utilizes judge-based evaluators and ground-truth datasets to score model outputs, and
Arize Phoenix is an open-source LLM observability platform that doubles as an evaluation framework, offering judge-based evaluators and ground-truth datasets for scoring model outputs, which directly supports building evaluation pipelines for performance and quality, though features like CI/CD integration and test case generation are less explicitly highlighted in the description.
FastChat is a training and serving platform for large language models that provides an integrated toolkit for fine-tuning, hosting, and benchmarking chatbots. It functions as an inference server capable of hosting multiple models and exposing them via a standardized API for chat applications. The platform distinguishes itself through a distributed model controller that manages worker nodes and routes requests across a hardware-agnostic inference layer supporting various accelerators. It includes a dedicated evaluation framework for assessing model quality using automated judges, multi-turn di
FastChat is primarily a training and serving platform but includes a dedicated evaluation framework with automated judges and multi-turn dialogue benchmarking, making it a valid tool for assessing LLM quality even if it lacks full pipeline customisation.
Opik is an observability and evaluation platform designed for generative AI applications and agentic workflows. It provides a centralized environment for tracing execution flows, managing prompt templates, and monitoring production performance, allowing teams to gain visibility into complex model interactions and tool usage without requiring manual application code changes. The platform distinguishes itself through its integrated approach to the AI development lifecycle, combining distributed trace instrumentation with automated evaluation frameworks. It supports model-as-a-judge scoring, syn
Opik is an evaluation and observability platform for LLM applications that supports model-as-a-judge scoring, custom metrics, and multi-model integrations, fitting the core need for building evaluation pipelines, though it focuses more on observability than dedicated test case generation or assertion-based evaluation.
Evaluation and Tracking for LLM Experiments and AI Agents
TruLens is an evaluation and tracking framework specifically designed for LLM experiments and AI agents, covering the core evaluation pipeline with feedback functions, model-graded evaluation, and integration capabilities that match the requested features.
Evals is a framework designed for automating, managing, and executing repeatable benchmarking suites to analyze the quality and performance of language models. It provides a platform for running standardized tests to measure model accuracy and track behavioral changes over time. The system distinguishes itself through a modular architecture that uses a standardized adapter layer to normalize inputs and outputs, allowing different models to be swapped and tested interchangeably. It supports the creation of custom benchmarks using proprietary data, enabling quality assurance on sensitive tasks
OpenAI Evals is a Python framework for automating and running repeatable benchmarks on language models, supporting custom test cases and interchangeable model backends, making it a solid fit for building LLM evaluation pipelines.
Ragas is an evaluation framework and performance benchmark designed to quantify the quality of retrieval augmented generation pipelines. It functions as an application optimizer to identify bottlenecks in language model workflows using automated metrics and model-based scoring. The framework includes a system for generating synthetic datasets that mimic production scenarios and edge cases to create realistic test cases. It enables reference-free assessment, allowing the evaluation of response quality by analyzing grounding in the provided context without requiring gold-standard labels. The s
Ragas is an evaluation framework specifically for RAG pipelines, providing synthetic test case generation and model-graded scoring to quantify LLM application quality, which fits your search for an LLM evaluation tool, though it lacks explicit assertion-based evaluation and CI/CD integration.
Ragas is an evaluation framework designed to measure the performance of retrieval-augmented generation pipelines and autonomous agent workflows. It provides a comprehensive suite of tools for benchmarking system outputs, utilizing language models as automated judges to score performance against defined rubrics and reference data. By standardizing inputs, retrieved contexts, and generated responses into a unified schema, the project enables consistent analysis across complex AI applications. The framework distinguishes itself through its ability to generate synthetic test datasets from existin
Ragas is an evaluation framework specifically for RAG and agent workflows, offering synthetic test generation and LLM-as-judge scoring—core capabilities for assessing LLM performance—though it does not explicitly cover CI/CD or multi-model support out of the box.
Holistic Evaluation of Language Models (HELM) is an open source Python framework created by the Center for Research on Foundation Models (CRFM) at Stanford for holistic, reproducible and transparent evaluation of foundation models, including large language models (LLMs) and multimodal models.
HELM is a comprehensive framework for standardized evaluation of LLMs across many scenarios and metrics, making it a valid LLM evaluation tool, but it is more of a fixed benchmark suite than a flexible pipeline builder for custom test cases, assertions, or CI/CD integration.
🤗 Evaluate: A library for easily evaluating machine learning models and datasets.
huggingface/evaluate is a general library for evaluating machine learning models and datasets, which includes LLMs, but it focuses on metrics and benchmarks rather than providing a dedicated pipeline for building evaluation workflows with test-case generation or assertion-based checks.
| Dépôt | Stars | Langage | Licence | Dernier push |
|---|---|---|---|---|
| typpo/promptfoo | 22.3K | TypeScript | MIT | |
| comet-ml/comet-llm | 19.7K | Python | Apache-2.0 | |
| evidentlyai/evidently | 7.1K | Jupyter Notebook | apache-2.0 | |
| giskard-ai/giskard | 5.4K | Python | Apache-2.0 | |
| oumi-ai/oumi | 8.9K | Python | apache-2.0 | |
| promptfoo/promptfoo | 10.5K | TypeScript | mit | |
| confident-ai/deepeval | 13.7K | Python | apache-2.0 | |
| agenta-ai/agenta | 3.9K | TypeScript | other | |
| internlm/opencompass | 7.1K | Python | Apache-2.0 | |
| open-compass/opencompass | 6.7K | Python | apache-2.0 |