awesome-repositories.com
Blog
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektÜber unsRanking-MethodikPresseMCP-Server
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to huggingface/evaluate

Open-source alternatives to Evaluate

30 open-source projects similar to huggingface/evaluate, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best Evaluate alternative.

  • confident-ai/deepevalAvatar von confident-ai

    confident-ai/deepeval

    13,733Auf GitHub ansehen↗

    Deepeval is a framework for testing and evaluating large language model applications. It provides a suite of tools for executing automated regression tests, validating model output quality against defined standards, and tracing the execution of complex agent workflows. By integrating these capabilities into development pipelines, the platform ensures consistent performance and reliability throughout the software lifecycle. The platform distinguishes itself through its focus on programmatic validation and observability. It utilizes secondary language models to score output quality and employs

    Pythonevaluation-frameworkevaluation-metricsllm-evaluation
    Auf GitHub ansehen↗13,733
  • huggingface/lightevalAvatar von huggingface

    huggingface/lighteval

    2,453Auf GitHub ansehen↗

    Lighteval is an open-source framework for running standardized benchmarks and custom evaluation tasks against language models. It provides a system for defining new evaluation tasks with custom prompts, metrics, and scoring in YAML configuration files, and integrates with the Hugging Face Hub for storing and comparing results. The framework supports evaluating models across multiple inference backends, including transformers, vllm, and custom APIs, through a unified generation and log-probability interface. It includes a pluggable metric registry for built-in and custom scoring, a prediction

    Pythonevaluationevaluation-frameworkevaluation-metrics
    Auf GitHub ansehen↗2,453
  • eleutherai/lm-evaluation-harnessAvatar von EleutherAI

    EleutherAI/lm-evaluation-harness

    11,460Auf GitHub ansehen↗

    This project is a standardized framework for benchmarking large language models across a wide range of academic and reasoning datasets. It provides a platform for executing automated evaluation tasks to measure model accuracy and performance, ensuring consistent assessment through a structured configuration schema. The framework distinguishes itself by incorporating a dedicated utility for data decontamination, which identifies and removes overlapping training samples from evaluation sets to prevent data leakage. It also features a flexible task builder that allows users to define custom benc

    Pythonevaluation-frameworklanguage-modeltransformer
    Auf GitHub ansehen↗11,460

KI-Suche

Entdecke weitere awesome Repositories

Beschreibe in einfachen Worten, was du brauchst — die KI bewertet tausende kuratierte Open-Source-Projekte nach Relevanz.

Find more with AI search
  • openai/simple-evalsAvatar von openai

    openai/simple-evals

    4,354Auf GitHub ansehen↗

    This project is a language model evaluation framework and benchmarking tool designed to measure the accuracy and performance of models across diverse datasets. It provides a system for implementing model-based graders, running standardized tests for mathematical reasoning, coding, and factuality, and calculating quantified performance metrics such as precision, recall, F1 scores, and pass-at-k. The framework utilizes model-based grading and rubrics to validate response quality against expert-defined criteria. It includes a multi-model benchmarking loop and a model-agnostic API interface to co

    Python
    Auf GitHub ansehen↗4,354
  • stanford-crfm/helmAvatar von stanford-crfm

    stanford-crfm/helm

    2,828Auf GitHub ansehen↗

    Holistic Evaluation of Language Models (HELM) is an open source Python framework created by the Center for Research on Foundation Models (CRFM) at Stanford for holistic, reproducible and transparent evaluation of foundation models, including large language models (LLMs) and multimodal models.

    Python
    Auf GitHub ansehen↗2,828
  • truera/trulensAvatar von truera

    truera/trulens

    3,384Auf GitHub ansehen↗

    Evaluation and Tracking for LLM Experiments and AI Agents

    Python
    Auf GitHub ansehen↗3,384
  • evalplus/evalplusAvatar von evalplus

    evalplus/evalplus

    1,765Auf GitHub ansehen↗

    Rigourous evaluation of LLM-synthesized code - NeurIPS 2023 & COLM 2024

    Python
    Auf GitHub ansehen↗1,765
  • mlfoundations/evalchemyAvatar von mlfoundations

    mlfoundations/evalchemy

    597Auf GitHub ansehen↗

    Automatic evals for LLMs

    HTML
    Auf GitHub ansehen↗597
  • psycoy/mixevalAvatar von Psycoy

    Psycoy/MixEval

    255Auf GitHub ansehen↗

    The official evaluation suite and dynamic data release for MixEval.

    Python
    Auf GitHub ansehen↗255
  • microsoft/promptbenchAvatar von microsoft

    microsoft/promptbench

    2,808Auf GitHub ansehen↗

    A unified evaluation framework for large language models

    Python
    Auf GitHub ansehen↗2,808
  • openlmlab/gaokao-benchAvatar von OpenLMLab

    OpenLMLab/GAOKAO-Bench

    760Auf GitHub ansehen↗

    GAOKAO-Bench is an evaluation framework that utilizes GAOKAO questions as a dataset to evaluate large language models.

    Python
    Auf GitHub ansehen↗760
  • open-compass/vlmevalkitAvatar von open-compass

    open-compass/VLMEvalKit

    3,824Auf GitHub ansehen↗

    VLMEvalKit is a vision-language model evaluation framework and inference engine designed to run standardized benchmarks and measure model accuracy across diverse visual datasets. It serves as a multimodal model benchmark and performance toolkit for calculating metrics and comparing model responses. The toolkit includes a specialized visual reasoning evaluator that uses adversarial samples to distinguish actual image understanding from reliance on language patterns. It also provides capabilities for image generation evaluation, testing a model's ability to create or modify visuals based on tex

    Pythonchatgptclaudeclip
    Auf GitHub ansehen↗3,824
  • pair-code/llm-comparatorAvatar von PAIR-code

    PAIR-code/llm-comparator

    528Auf GitHub ansehen↗

    LLM Comparator is an interactive data visualization tool for evaluating and analyzing LLM responses side-by-side, developed by the PAIR team.

    JavaScript
    Auf GitHub ansehen↗528
  • openai/evalsAvatar von openai

    openai/evals

    18,702Auf GitHub ansehen↗

    Evals is a framework designed for automating, managing, and executing repeatable benchmarking suites to analyze the quality and performance of language models. It provides a platform for running standardized tests to measure model accuracy and track behavioral changes over time. The system distinguishes itself through a modular architecture that uses a standardized adapter layer to normalize inputs and outputs, allowing different models to be swapped and tested interchangeably. It supports the creation of custom benchmarks using proprietary data, enabling quality assurance on sensitive tasks

    Python
    Auf GitHub ansehen↗18,702
  • giskard-ai/giskardAvatar von Giskard-AI

    Giskard-AI/giskard

    5,434Auf GitHub ansehen↗

    Giskard is an evaluation framework, testing library, and quality monitoring system for large language models and AI agents. It serves as a toolkit for quantifying model performance and reliability, providing specialized capabilities for validating retrieval-augmented generation pipelines. The project distinguishes itself through an automated red teaming tool and security scanner designed to identify vulnerabilities, prompt injections, and safety risks. It utilizes adversarial probing and synthetic edge case generation to quantify model robustness and detect information disclosure. The platfo

    Python
    Auf GitHub ansehen↗5,434
  • johnsnowlabs/langtestAvatar von JohnSnowLabs

    JohnSnowLabs/langtest

    561Auf GitHub ansehen↗

    Deliver safe & effective language models

    Python
    Auf GitHub ansehen↗561
  • open-compass/opencompassAvatar von open-compass

    open-compass/opencompass

    6,678Auf GitHub ansehen↗

    OpenCompass is an open-source framework for standardized benchmarking of large language models. It provides a configurable evaluation pipeline that supports both objective and subjective assessment, using a dual-engine architecture to handle closed-form answer comparison and open-ended response rating. The framework is designed as a modular platform where datasets, models, and metrics are composed through declarative YAML configuration files. The framework distinguishes itself through its extensible model integration layer, which supports custom models, HuggingFace models, and third-party API

    Pythonbenchmarkchatgptevaluation
    Auf GitHub ansehen↗6,678
  • modelscope/evalscopeAvatar von modelscope

    modelscope/evalscope

    2,955Auf GitHub ansehen↗

    A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.

    Python
    Auf GitHub ansehen↗2,955
  • explodinggradients/ragasAvatar von explodinggradients

    explodinggradients/ragas

    14,400Auf GitHub ansehen↗

    Ragas is an evaluation framework and performance benchmark designed to quantify the quality of retrieval augmented generation pipelines. It functions as an application optimizer to identify bottlenecks in language model workflows using automated metrics and model-based scoring. The framework includes a system for generating synthetic datasets that mimic production scenarios and edge cases to create realistic test cases. It enables reference-free assessment, allowing the evaluation of response quality by analyzing grounding in the provided context without requiring gold-standard labels. The s

    Python
    Auf GitHub ansehen↗14,400
  • langfuse/langfuseAvatar von langfuse

    langfuse/langfuse

    29,190Auf GitHub ansehen↗

    Langfuse is an open-source observability and evaluation platform designed for language model applications. It provides a centralized system for tracking execution traces, monitoring performance metrics, and managing prompt templates. By capturing hierarchical units of work and telemetry data, the platform enables developers to debug complex application lifecycles and analyze token usage, latency, and model interactions in production environments. The platform distinguishes itself through an integrated evaluation framework that allows for systematic benchmarking and automated scoring of model

    TypeScriptanalyticsautogenevaluation
    Auf GitHub ansehen↗29,190
  • comet-ml/opikAvatar von comet-ml

    comet-ml/opik

    17,787Auf GitHub ansehen↗

    Opik is an observability and evaluation platform designed for generative AI applications and agentic workflows. It provides a centralized environment for tracing execution flows, managing prompt templates, and monitoring production performance, allowing teams to gain visibility into complex model interactions and tool usage without requiring manual application code changes. The platform distinguishes itself through its integrated approach to the AI development lifecycle, combining distributed trace instrumentation with automated evaluation frameworks. It supports model-as-a-judge scoring, syn

    Pythonevaluationhacktoberfesthacktoberfest2025
    Auf GitHub ansehen↗17,787
  • facebookresearch/parlaiAvatar von facebookresearch

    facebookresearch/ParlAI

    10,625Auf GitHub ansehen↗

    ParlAI is a conversational AI research framework designed for training, evaluating, and sharing dialogue models using a unified interface for datasets and agents. It functions as a PyTorch-based training platform and a dialogue data collection system, providing a centralized model zoo for the distribution of versioned pretrained agents. The project distinguishes itself through a knowledge-grounded retrieval system that combines dense and sparse indexing to ground responses in external information. It also provides a comprehensive infrastructure for gathering human-AI interaction data via inte

    Python
    Auf GitHub ansehen↗10,625
  • nvidia/isaac-gr00tAvatar von NVIDIA

    NVIDIA/Isaac-GR00T

    6,222Auf GitHub ansehen↗
    Jupyter Notebook
    Auf GitHub ansehen↗6,222
  • open-mmlab/mmsegmentationAvatar von open-mmlab

    open-mmlab/mmsegmentation

    9,860Auf GitHub ansehen↗

    MMSegmentation is an open-source semantic segmentation toolbox built on PyTorch that provides a modular, configurable framework for building, training, evaluating, and deploying segmentation models. At its core, it offers a config-driven pipeline that assembles training, evaluation, and inference workflows by parsing hierarchical configuration files, with a modular component registry that enables plug-and-play composition of neural network modules, optimizers, datasets, and metrics. The framework supports the full model lifecycle through a unified runner interface that controls training, testi

    Pythondeeplabv3image-segmentationmedical-image-segmentation
    Auf GitHub ansehen↗9,860
  • infrasys-ai/aiinfraAvatar von Infrasys-AI

    Infrasys-AI/AIInfra

    7,414Auf GitHub ansehen↗
    Jupyter Notebookaiinfraaisystem
    Auf GitHub ansehen↗7,414
  • packtpublishing/llm-engineers-handbookAvatar von PacktPublishing

    PacktPublishing/LLM-Engineers-Handbook

    4,774Auf GitHub ansehen↗

    This project is an educational resource and engineering guide for building, deploying, and optimizing large language model applications and production pipelines. It serves as a blueprint for cloud AI infrastructure, providing a framework for orchestrating inference endpoints, data warehouses, and scalable production environments. The repository provides specific implementation patterns for retrieval augmented generation to ground model responses in external data. It includes a training workflow for crawling, structuring, and processing datasets to facilitate model fine-tuning, alongside an ev

    Pythonawsfine-tuning-llmgenai
    Auf GitHub ansehen↗4,774
  • sjwhitworth/golearnAvatar von sjwhitworth

    sjwhitworth/golearn

    9,438Auf GitHub ansehen↗

    GoLearn is a machine learning library for the Go programming language. It provides a supervised learning framework and a toolkit for building, training, and evaluating predictive models through a standardized interface. The project implements a data frame system that loads CSV files into structured grids for matrix operations. It includes a preprocessing library for discretizing continuous variables and a model evaluation toolkit that utilizes confusion matrices and cross-validation to measure precision and recall. The library covers data engineering and management, including the ability to

    Go
    Auf GitHub ansehen↗9,438
  • facebookresearch/slowfastAvatar von facebookresearch

    facebookresearch/SlowFast

    7,377Auf GitHub ansehen↗

    SlowFast is a PyTorch video understanding framework and spatiotemporal neural network library. It serves as a toolset for video action recognition, enabling the training and evaluation of models designed to classify complex activities and objects within video sequences. The framework is distinguished by its use of dual-pathway spatiotemporal sampling to capture both slow and fast motions. It supports self-supervised video learning for pre-training models on unlabeled data and employs multigrid spatiotemporal training to optimize learning across multiple spatial and temporal resolutions. The

    Python
    Auf GitHub ansehen↗7,377
  • datawhalechina/prompt-engineering-for-developersAvatar von datawhalechina

    datawhalechina/prompt-engineering-for-developers

    24,267Auf GitHub ansehen↗

    This project is a technical curriculum and development guide focused on large language model prompt engineering, fine-tuning, and the creation of retrieval augmented generation applications. It serves as a comprehensive resource for developers to master crafting precise instructions and textual patterns to improve the quality and predictability of model outputs. The material covers the end-to-end workflow of adapting open-source models to specific datasets and integrating language models with vector databases to generate responses based on private information. It also provides a systematic ap

    Jupyter Notebook
    Auf GitHub ansehen↗24,267
  • ntmc-community/matchzooAvatar von NTMC-Community

    NTMC-Community/MatchZoo

    3,845Auf GitHub ansehen↗

    MatchZoo is a deep learning framework designed for building, training, and evaluating neural networks that determine the relevance and similarity between pairs of textual inputs. It serves as a research platform for neural information retrieval, specifically supporting the development of models for document retrieval, question answering, and ranking tasks. The framework utilizes declarative architecture composition to define complex neural network structures. It includes automated hyper-parameter resolution to populate missing configuration parameters before model compilation and uses callbac

    Python
    Auf GitHub ansehen↗3,845