awesome-repositories.com
Blog
MCP
awesome-repositories.com

Découvrez les meilleurs dépôts open-source grâce à notre recherche par IA.

ExplorerRecherches sélectionnéesAlternatives open sourceLogiciels auto-hébergésBlogPlan du site
ProjetÀ proposNotre méthodologiePresseServeur MCP
Mentions légalesConfidentialitéConditions d'utilisation
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

Frameworks d'évaluation d'applications LLM

Classement mis à jour le 30 juin 2026

For framework pour l'évaluation d'applications LLM, the strongest matches are typpo/promptfoo (promptfoo is a dedicated evaluation framework for LLM prompts), comet-ml/comet-llm (Comet LLM is a dedicated evaluation framework for LLM) and evidentlyai/evidently (Evidently is an AI observability and evaluation framework specifically). giskard-ai/giskard and oumi-ai/oumi round out the shortlist. Each is ranked by relevance to your query, popularity and recent activity.

Ces outils open source fournissent des suites de tests automatisés et des métriques pour évaluer les performances des modèles de langage.

Frameworks d'évaluation d'applications LLM

Trouvez les meilleurs dépôts grâce à l'IA.Nous recherchons les dépôts les plus pertinents grâce à l'IA.
  • typpo/promptfooAvatar de typpo

    typpo/promptfoo

    22,295Voir sur GitHub↗

    promptfoo is an evaluation framework for measuring the performance of large language model prompts, agents, and retrieval augmented generation pipelines. It provides a suite of tools for conducting comparative benchmarking and executing automated quality and security regressions. The system features a benchmarking suite for running identical prompts across different model providers to compare output quality side-by-side. It also includes a dedicated red teaming tool for identifying security vulnerabilities and prompt injection risks through automated penetration testing. The framework suppor

    promptfoo is a dedicated evaluation framework for LLM prompts, agents, and RAG pipelines, offering benchmarking, automated assertion validation, red teaming for security, CI/CD integration, multi‑model comparison, and custom scoring pipelines—directly addressing the need for a comprehensive performance, safety, and quality evaluation tool.

    TypeScriptModel Comparison InterfacesCI/CD Pipeline IntegrationsLLM Evaluation
    Voir sur GitHub↗22,295
  • comet-ml/comet-llmAvatar de comet-ml

    comet-ml/comet-llm

    19,673Voir sur GitHub↗

    Comet LLM is an observability platform and evaluation framework designed for large language model applications and agentic workflows. It functions as a system for tracing, monitoring, and debugging execution flows while providing tools for prompt optimization and the enforcement of AI safety guardrails. The platform distinguishes itself through a combination of model-based scoring and heuristic metrics to quantify output quality and detect hallucinations. It includes a dedicated prompt and agent optimizer with an interactive playground for refining templates and tool configurations. For retri

    Comet LLM is a dedicated evaluation framework for LLM applications that provides model-based scoring, heuristic metrics, hallucination detection, prompt optimization, and CI/CD integration—directly covering the test case generation, assertion, model-graded evaluation, and custom metrics needed for assessing performance, safety, and quality.

    PythonLLM-As-A-Judge ScoringAutomated Model JudgesLLM Evaluation
    Voir sur GitHub↗19,673
  • evidentlyai/evidentlyAvatar de evidentlyai

    evidentlyai/evidently

    7,137Voir sur GitHub↗

    Evidently is an AI observability platform and evaluation framework designed to quantify the performance of machine learning models and large language models. It functions as a monitoring tool for detecting data drift and quality degradation in tabular datasets, while providing a specialized analyzer for the faithfulness and correctness of retrieval augmented generation systems. The project distinguishes itself through an evaluation framework that utilizes judge models and custom rubrics to score language model outputs. It includes tools for iterative prompt optimization and the generation of

    Evidently is an AI observability and evaluation framework specifically for LLMs, offering model-graded scoring, custom rubrics, RAG evaluation, and CI/CD integration, which squarely meets the need for building LLM evaluation pipelines.

    Jupyter NotebookLLM-As-A-Judge ScoringCI/CD Pipeline IntegrationsCustom Evaluation Judges
    Voir sur GitHub↗7,137
  • giskard-ai/giskardAvatar de Giskard-AI

    Giskard-AI/giskard

    5,434Voir sur GitHub↗

    Giskard is an evaluation framework, testing library, and quality monitoring system for large language models and AI agents. It serves as a toolkit for quantifying model performance and reliability, providing specialized capabilities for validating retrieval-augmented generation pipelines. The project distinguishes itself through an automated red teaming tool and security scanner designed to identify vulnerabilities, prompt injections, and safety risks. It utilizes adversarial probing and synthetic edge case generation to quantify model robustness and detect information disclosure. The platfo

    Giskard is a dedicated evaluation framework and testing library for LLMs and AI agents, offering automated test case generation, adversarial probing, security scanning, and quality monitoring that directly supports building evaluation pipelines for performance, safety, and quality assessment.

    PythonTest Case GeneratorsAutomated Model JudgesLLM Evaluation
    Voir sur GitHub↗5,434
  • oumi-ai/oumiAvatar de oumi-ai

    oumi-ai/oumi

    8,858Voir sur GitHub↗

    Oumi is a comprehensive large language model development platform designed for synthesizing data, fine-tuning models, and running performance evaluations. It serves as a unified environment for the entire model lifecycle, encompassing a training and fine-tuning suite, an evaluation framework, and tools for synthetic data generation and model distillation. The platform is distinguished by its iterative, failure-driven synthesis approach, which analyzes model weaknesses during evaluation to generate targeted training data. It utilizes an LLM-based judge framework to programmatically score respo

    Oumi is a comprehensive LLM development platform with a built-in evaluation framework that uses an LLM-based judge for scoring and a failure-driven approach to generate targeted test cases, directly supporting the kind of evaluation pipelines needed to assess performance, safety, and quality across multiple models.

    PythonLLM-As-A-Judge ScoringAutomated Model JudgesCustom Evaluation Judges
    Voir sur GitHub↗8,858
  • promptfoo/promptfooAvatar de promptfoo

    promptfoo/promptfoo

    10,529Voir sur GitHub↗

    Promptfoo is an evaluation framework designed for testing, benchmarking, and red-teaming language models and agentic workflows. It provides a unified environment to run prompts against multiple providers, allowing developers to systematically validate model outputs against objective assertions, semantic similarity metrics, and custom grading rubrics. The platform distinguishes itself through a provider-agnostic execution layer and a stateful orchestrator capable of simulating multi-turn conversations and complex tool-use trajectories. It includes a dedicated adversarial mutation pipeline that

    Promptfoo is a fully-featured evaluation framework for LLMs that supports test case generation, assertion-based and model-graded evaluation, multi-model comparisons, and integrates with CI/CD pipelines, making it a strong fit for building comprehensive evaluation pipelines.

    TypeScriptTest Case GeneratorsModel Comparison InterfacesLLM Evaluation
    Voir sur GitHub↗10,529
  • confident-ai/deepevalAvatar de confident-ai

    confident-ai/deepeval

    13,733Voir sur GitHub↗

    Deepeval is a framework for testing and evaluating large language model applications. It provides a suite of tools for executing automated regression tests, validating model output quality against defined standards, and tracing the execution of complex agent workflows. By integrating these capabilities into development pipelines, the platform ensures consistent performance and reliability throughout the software lifecycle. The platform distinguishes itself through its focus on programmatic validation and observability. It utilizes secondary language models to score output quality and employs

    Deepeval is a dedicated framework for testing and evaluating LLM applications, with built-in support for assertion-based and model-graded evaluation, CI/CD integration, and observability—exactly what you need to build evaluation pipelines for LLM performance and safety.

    PythonCI/CD Pipeline IntegrationsLLM Evaluation
    Voir sur GitHub↗13,733
  • agenta-ai/agentaAvatar de Agenta-AI

    Agenta-AI/agenta

    3,860Voir sur GitHub↗

    Agenta is a Prompt Ops lifecycle manager and prompt management platform that decouples prompt engineering from application code. It serves as a centralized system for developing, versioning, and deploying prompt templates and model configurations across different environments. The platform functions as an AI agent orchestrator with a visual interface for building agent workflows and connecting models to external tools. It further acts as an evaluation framework and observability tool, utilizing OpenTelemetry to capture execution traces, monitor latency, and track token costs. The system cove

    Agenta is an LLM evaluation and prompt management platform that supports model-graded evaluation, observability, and CI/CD integration, making it a comprehensive solution for building evaluation pipelines for LLM applications.

    TypeScriptLLM-As-A-Judge ScoringCustom Evaluation JudgesLLM Evaluation
    Voir sur GitHub↗3,860
  • internlm/opencompassAvatar de InternLM

    InternLM/opencompass

    7,096Voir sur GitHub↗

    OpenCompass is a comprehensive evaluation platform, benchmarking suite, and distributed model evaluator designed to measure the performance and accuracy of large language models. It provides a framework for benchmarking both open-source and API-based models against diverse datasets using standardized metrics and reproducible pipelines. The project features an automated judging framework that uses language models as judges to score and verify the quality of generated text. It includes a performance leaderboard system for comparing the relative capabilities of various models across industry-sta

    OpenCompass is a comprehensive benchmarking suite and evaluation platform for LLMs, supporting multi-model evaluation and LLM-as-a-judge scoring, though its focus on standardized benchmarking rather than custom pipeline construction and test case generation means it may not cover all the pipeline-building features you listed.

    PythonLLM-As-A-Judge ScoringLLM Evaluation
    Voir sur GitHub↗7,096
  • open-compass/opencompassAvatar de open-compass

    open-compass/opencompass

    6,678Voir sur GitHub↗

    OpenCompass is an open-source framework for standardized benchmarking of large language models. It provides a configurable evaluation pipeline that supports both objective and subjective assessment, using a dual-engine architecture to handle closed-form answer comparison and open-ended response rating. The framework is designed as a modular platform where datasets, models, and metrics are composed through declarative YAML configuration files. The framework distinguishes itself through its extensible model integration layer, which supports custom models, HuggingFace models, and third-party API

    OpenCompass is a modular framework for benchmarking LLMs through configurable evaluation pipelines, supporting objective and subjective assessment with multi-model and custom-metric extensibility—making it a fitting tool for building LLM evaluation workflows, even if it focuses more on standardized benchmarks than on app-specific safety or CI/CD integration.

    PythonLLM-As-A-Judge ScoringAutomated Model Judges
    Voir sur GitHub↗6,678
  • eleutherai/lm-evaluation-harnessAvatar de EleutherAI

    EleutherAI/lm-evaluation-harness

    11,460Voir sur GitHub↗

    This project is a standardized framework for benchmarking large language models across a wide range of academic and reasoning datasets. It provides a platform for executing automated evaluation tasks to measure model accuracy and performance, ensuring consistent assessment through a structured configuration schema. The framework distinguishes itself by incorporating a dedicated utility for data decontamination, which identifies and removes overlapping training samples from evaluation sets to prevent data leakage. It also features a flexible task builder that allows users to define custom benc

    The LM Evaluation Harness is a standardized framework for benchmarking LLMs with configurable tasks, multi-model support, and built-in evaluation utilities, making it a strong foundation for building evaluation pipelines that assess model performance and quality.

    PythonCustom Evaluation Judges
    Voir sur GitHub↗11,460
  • arize-ai/phoenixAvatar de Arize-ai

    Arize-ai/phoenix

    8,605Voir sur GitHub↗

    Arize Phoenix is an LLM observability platform and evaluation framework designed to capture execution traces and monitor large language model applications. It serves as a prompt management system for versioning and testing templates, and as a self-hosted AI operations infrastructure for managing telemetry and experiments. The platform differentiates itself through a specialized embedding visualization tool used to detect data drift and optimize vector search. It provides a comprehensive evaluation suite that utilizes judge-based evaluators and ground-truth datasets to score model outputs, and

    Arize Phoenix is an open-source LLM observability platform that doubles as an evaluation framework, offering judge-based evaluators and ground-truth datasets for scoring model outputs, which directly supports building evaluation pipelines for performance and quality, though features like CI/CD integration and test case generation are less explicitly highlighted in the description.

    Jupyter NotebookLLM-As-A-Judge ScoringAutomated Model JudgesLLM Evaluation
    Voir sur GitHub↗8,605
  • lm-sys/fastchatAvatar de lm-sys

    lm-sys/FastChat

    39,472Voir sur GitHub↗

    FastChat is a training and serving platform for large language models that provides an integrated toolkit for fine-tuning, hosting, and benchmarking chatbots. It functions as an inference server capable of hosting multiple models and exposing them via a standardized API for chat applications. The platform distinguishes itself through a distributed model controller that manages worker nodes and routes requests across a hardware-agnostic inference layer supporting various accelerators. It includes a dedicated evaluation framework for assessing model quality using automated judges, multi-turn di

    FastChat is primarily a training and serving platform but includes a dedicated evaluation framework with automated judges and multi-turn dialogue benchmarking, making it a valid tool for assessing LLM quality even if it lacks full pipeline customisation.

    PythonAutomated Model JudgesLLM Evaluation
    Voir sur GitHub↗39,472
  • comet-ml/opikAvatar de comet-ml

    comet-ml/opik

    17,787Voir sur GitHub↗

    Opik is an observability and evaluation platform designed for generative AI applications and agentic workflows. It provides a centralized environment for tracing execution flows, managing prompt templates, and monitoring production performance, allowing teams to gain visibility into complex model interactions and tool usage without requiring manual application code changes. The platform distinguishes itself through its integrated approach to the AI development lifecycle, combining distributed trace instrumentation with automated evaluation frameworks. It supports model-as-a-judge scoring, syn

    Opik is an evaluation and observability platform for LLM applications that supports model-as-a-judge scoring, custom metrics, and multi-model integrations, fitting the core need for building evaluation pipelines, though it focuses more on observability than dedicated test case generation or assertion-based evaluation.

    PythonAutomated Model JudgesLLM Evaluation
    Voir sur GitHub↗17,787
  • truera/trulensAvatar de truera

    truera/trulens

    3,384Voir sur GitHub↗

    Evaluation and Tracking for LLM Experiments and AI Agents

    TruLens is an evaluation and tracking framework specifically designed for LLM experiments and AI agents, covering the core evaluation pipeline with feedback functions, model-graded evaluation, and integration capabilities that match the requested features.

    PythonEvaluation and ObservabilityEvaluation FrameworksEvaluation Frameworks
    Voir sur GitHub↗3,384
  • openai/evalsAvatar de openai

    openai/evals

    18,702Voir sur GitHub↗

    Evals is a framework designed for automating, managing, and executing repeatable benchmarking suites to analyze the quality and performance of language models. It provides a platform for running standardized tests to measure model accuracy and track behavioral changes over time. The system distinguishes itself through a modular architecture that uses a standardized adapter layer to normalize inputs and outputs, allowing different models to be swapped and tested interchangeably. It supports the creation of custom benchmarks using proprietary data, enabling quality assurance on sensitive tasks

    OpenAI Evals is a Python framework for automating and running repeatable benchmarks on language models, supporting custom test cases and interchangeable model backends, making it a solid fit for building LLM evaluation pipelines.

    PythonLLM Evaluation
    Voir sur GitHub↗18,702
  • explodinggradients/ragasAvatar de explodinggradients

    explodinggradients/ragas

    14,400Voir sur GitHub↗

    Ragas is an evaluation framework and performance benchmark designed to quantify the quality of retrieval augmented generation pipelines. It functions as an application optimizer to identify bottlenecks in language model workflows using automated metrics and model-based scoring. The framework includes a system for generating synthetic datasets that mimic production scenarios and edge cases to create realistic test cases. It enables reference-free assessment, allowing the evaluation of response quality by analyzing grounding in the provided context without requiring gold-standard labels. The s

    Ragas is an evaluation framework specifically for RAG pipelines, providing synthetic test case generation and model-graded scoring to quantify LLM application quality, which fits your search for an LLM evaluation tool, though it lacks explicit assertion-based evaluation and CI/CD integration.

    PythonLLM Evaluation
    Voir sur GitHub↗14,400
  • vibrantlabsai/ragasAvatar de vibrantlabsai

    vibrantlabsai/ragas

    12,659Voir sur GitHub↗

    Ragas is an evaluation framework designed to measure the performance of retrieval-augmented generation pipelines and autonomous agent workflows. It provides a comprehensive suite of tools for benchmarking system outputs, utilizing language models as automated judges to score performance against defined rubrics and reference data. By standardizing inputs, retrieved contexts, and generated responses into a unified schema, the project enables consistent analysis across complex AI applications. The framework distinguishes itself through its ability to generate synthetic test datasets from existin

    Ragas is an evaluation framework specifically for RAG and agent workflows, offering synthetic test generation and LLM-as-judge scoring—core capabilities for assessing LLM performance—though it does not explicitly cover CI/CD or multi-model support out of the box.

    PythonAutomated Model JudgesCustom Evaluation JudgesLLM Evaluation
    Voir sur GitHub↗12,659
  • stanford-crfm/helmAvatar de stanford-crfm

    stanford-crfm/helm

    2,828Voir sur GitHub↗

    Holistic Evaluation of Language Models (HELM) is an open source Python framework created by the Center for Research on Foundation Models (CRFM) at Stanford for holistic, reproducible and transparent evaluation of foundation models, including large language models (LLMs) and multimodal models.

    HELM is a comprehensive framework for standardized evaluation of LLMs across many scenarios and metrics, making it a valid LLM evaluation tool, but it is more of a fixed benchmark suite than a flexible pipeline builder for custom test cases, assertions, or CI/CD integration.

    PythonEvaluation FrameworksModel Evaluation and Benchmarking
    Voir sur GitHub↗2,828
  • huggingface/evaluateAvatar de huggingface

    huggingface/evaluate

    2,455Voir sur GitHub↗

    🤗 Evaluate: A library for easily evaluating machine learning models and datasets.

    huggingface/evaluate is a general library for evaluating machine learning models and datasets, which includes LLMs, but it focuses on metrics and benchmarks rather than providing a dedicated pipeline for building evaluation workflows with test-case generation or assertion-based checks.

    PythonEvaluation FrameworksModel EvaluationModel Evaluation and Benchmarking
    Voir sur GitHub↗2,455
Comparez le top 10 en un coup d'œil
DépôtStarsLangageLicenceDernier push
typpo/promptfoo22.3KTypeScriptMIT17 juin 2026
comet-ml/comet-llm19.7KPythonApache-2.017 juin 2026
evidentlyai/evidently7.1KJupyter Notebookapache-2.020 févr. 2026
giskard-ai/giskard5.4KPythonApache-2.017 juin 2026
oumi-ai/oumi8.9KPythonapache-2.019 févr. 2026
promptfoo/promptfoo10.5KTypeScriptmit20 févr. 2026
confident-ai/deepeval13.7KPythonapache-2.019 févr. 2026
agenta-ai/agenta3.9KTypeScriptother22 févr. 2026
internlm/opencompass7.1KPythonApache-2.017 juin 2026
open-compass/opencompass6.7KPythonapache-2.014 févr. 2026

Related searches

  • un framework pour évaluer la qualité des sorties LLM
  • benchmark pour comparer les modèles de langage
  • a framework for evaluating small language models
  • un framework open source pour les applications LLM
  • bibliothèque d'évaluation LLM-as-a-judge
  • une plateforme d'observabilité pour les applications LLM
  • Évaluation de modèles et observabilité LLM
  • toolkit pour créer des agents IA capables d'utiliser des outils