awesome-repositories.com
Blog
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektÜber unsRanking-MethodikPresseMCP-Server
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Psycoy avatar

Psycoy/MixEval

0
View on GitHub↗
255 Stars·40 Forks·Python·1 Aufrufmixeval.github.io↗

MixEval

The official evaluation suite and dynamic data release for MixEval.

Features

  • Evaluation Frameworks - Click-and-go evaluation suite for open and proprietary models.
  • Model Evaluation - Benchmark mixture evaluation using wisdom of the crowd.

Star-Verlauf

Star-Verlauf für psycoy/mixevalStar-Verlauf für psycoy/mixeval

KI-Suche

Entdecke weitere awesome Repositories

Beschreibe in einfachen Worten, was du brauchst — die KI bewertet tausende kuratierte Open-Source-Projekte nach Relevanz.

Start searching with AI

Häufig gestellte Fragen

Was macht psycoy/mixeval?

The official evaluation suite and dynamic data release for MixEval.

Was sind die Hauptfunktionen von psycoy/mixeval?

Die Hauptfunktionen von psycoy/mixeval sind: Evaluation Frameworks, Model Evaluation.

Welche Open-Source-Alternativen gibt es zu psycoy/mixeval?

Open-Source-Alternativen zu psycoy/mixeval sind unter anderem: pair-code/llm-comparator — LLM Comparator is an interactive data visualization tool for evaluating and analyzing LLM responses side-by-side,… huggingface/lighteval — Lighteval is an open-source framework for running standardized benchmarks and custom evaluation tasks against language… eleutherai/lm-evaluation-harness — This project is a standardized framework for benchmarking large language models across a wide range of academic and… confident-ai/deepeval — Deepeval is a framework for testing and evaluating large language model applications. It provides a suite of tools for… huggingface/evaluate — 🤗 Evaluate: A library for easily evaluating machine learning models and datasets. nvidia/isaac-gr00t.

Open-Source-Alternativen zu MixEval

Ähnliche Open-Source-Projekte, sortiert nach der Anzahl der gemeinsamen Funktionen mit MixEval.
  • huggingface/evaluateAvatar von huggingface

    huggingface/evaluate

    2,455Auf GitHub ansehen↗

    🤗 Evaluate: A library for easily evaluating machine learning models and datasets.

    Python
    Auf GitHub ansehen↗2,455
  • eleutherai/lm-evaluation-harnessAvatar von EleutherAI

    EleutherAI/lm-evaluation-harness

    11,460Auf GitHub ansehen↗

    This project is a standardized framework for benchmarking large language models across a wide range of academic and reasoning datasets. It provides a platform for executing automated evaluation tasks to measure model accuracy and performance, ensuring consistent assessment through a structured configuration schema. The framework distinguishes itself by incorporating a dedicated utility for data decontamination, which identifies and removes overlapping training samples from evaluation sets to prevent data leakage. It also features a flexible task builder that allows users to define custom benc

    Pythonevaluation-frameworklanguage-modeltransformer
    Auf GitHub ansehen↗11,460
  • confident-ai/deepevalAvatar von confident-ai

    confident-ai/deepeval

    13,733Auf GitHub ansehen↗

    Deepeval is a framework for testing and evaluating large language model applications. It provides a suite of tools for executing automated regression tests, validating model output quality against defined standards, and tracing the execution of complex agent workflows. By integrating these capabilities into development pipelines, the platform ensures consistent performance and reliability throughout the software lifecycle. The platform distinguishes itself through its focus on programmatic validation and observability. It utilizes secondary language models to score output quality and employs

    Pythonevaluation-frameworkevaluation-metricsllm-evaluation
    Auf GitHub ansehen↗13,733
  • huggingface/lightevalAvatar von huggingface

    huggingface/lighteval

    2,453Auf GitHub ansehen↗

    Lighteval is an open-source framework for running standardized benchmarks and custom evaluation tasks against language models. It provides a system for defining new evaluation tasks with custom prompts, metrics, and scoring in YAML configuration files, and integrates with the Hugging Face Hub for storing and comparing results. The framework supports evaluating models across multiple inference backends, including transformers, vllm, and custom APIs, through a unified generation and log-probability interface. It includes a pluggable metric registry for built-in and custom scoring, a prediction

    Pythonevaluationevaluation-frameworkevaluation-metrics
    Auf GitHub ansehen↗2,453
Alle 30 Alternativen zu MixEval anzeigen→