awesome-repositories.com
Blog
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectDespreCum realizăm clasamentulPresăServer MCP
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
huggingface avatar

huggingface/lighteval

0
View on GitHub↗
2,453 stele·490 fork-uri·Python·MIT·8 vizualizărihuggingface.co/docs/lighteval/en/index↗

Lighteval

Lighteval is an open-source framework for running standardized benchmarks and custom evaluation tasks against language models. It provides a system for defining new evaluation tasks with custom prompts, metrics, and scoring in YAML configuration files, and integrates with the Hugging Face Hub for storing and comparing results.

The framework supports evaluating models across multiple inference backends, including transformers, vllm, and custom APIs, through a unified generation and log-probability interface. It includes a pluggable metric registry for built-in and custom scoring, a prediction cache layer that avoids re-computation on repeated runs, and the ability to push results to remote storage destinations like the Hugging Face Hub, S3, or local directories. Per-sample outputs are saved to enable detailed debugging and performance analysis after evaluation.

Evaluation tasks are configured in YAML files specifying prompt templates, reference answers, and scoring metrics, with support for subset-based task splitting that generates separate configurations for each dataset subset automatically. The framework includes thousands of pre-built benchmark tasks and allows users to define custom tasks, custom metrics, and custom model implementations.

Features

  • LLM Evaluation Frameworks - Provides an open-source framework for running standardized benchmarks and custom evaluation tasks against language models.
  • Language Model Benchmark Suites - Runs pre-built and custom evaluation tasks against language models to measure performance.
  • Custom Evaluation Judges - Builds new evaluation tasks with custom metrics and scoring to test model behavior on user-defined criteria.
  • Model Performance Evaluators - Measures model outputs against predefined metrics to quantify accuracy and quality across diverse tasks.
  • Unified Generation Interfaces - Provides a unified generation and log-probability interface across transformers, vllm, and custom API backends.
  • Standardized Benchmarks - Executes thousands of ready-made evaluation tasks to measure model performance on standard tests.
  • Model Evaluation - Runs models through a curated set of tasks and computes standard metrics to measure performance.
  • YAML-Based Task Definitions - Defines evaluation tasks in YAML files specifying prompt templates, reference answers, and scoring metrics.
  • Evaluation Task Definitions - Defines new evaluation tasks with custom prompts, metrics, and scoring in YAML configuration files.
  • YAML Benchmark Definitions - Creates new benchmarks by specifying prompt templates, reference answers, and scoring metrics in YAML files.
  • Custom Model Evaluation Adapters - Runs benchmark tasks against user-defined models by implementing subclasses with generation and log-probability methods.
  • Evaluation Prediction Caches - Caches computed model outputs to disk or memory, skipping re-computation on repeated evaluation runs.
  • Metric Selection and Creation - Selects built-in metrics or defines custom ones to score model outputs during evaluation.
  • Evaluation Prompt Functions - Converts dataset lines into structured documents with query, choices, and gold index for evaluation.
  • Evaluation Task Executors - Executes user-defined evaluation tasks against a model using the CLI with a custom task file.
  • Inference Backend Abstraction - Supports running evaluations across transformers, vllm, and custom API backends through a unified interface.
  • Inference Engine Evaluation Backends - Runs model evaluation using various inference engines like transformers, vllm, and custom APIs.
  • Single-Subset Evaluation Configurations - Defines single-subset tasks with a prompt function, dataset source, splits, metrics, and generation parameters.
  • Result Storage Integrations - Stores evaluation results on the Hugging Face Hub for sharing and comparing model performance.
  • Custom Metric Registries - Ships a pluggable metric registry where built-in and custom scoring functions are registered by name.
  • Datasets and Evaluation - Library for evaluating local LLMs on benchmarks.
  • Evaluation Frameworks - Unified toolkit for benchmarking and evaluating language models.
  • Model Evaluation - All-in-one toolkit for evaluating models across multiple backends.
  • Model Evaluation and Benchmarking - Lightweight evaluation suite for language models.

Istoric stele

Graficul istoricului de stele pentru huggingface/lightevalGraficul istoricului de stele pentru huggingface/lighteval

Căutare AI

Explorează mai multe repository-uri excelente

Descrie ce ai nevoie în limbaj simplu — AI-ul sortează mii de proiecte open source selectate în funcție de relevanță.

Start searching with AI

Alternative open-source pentru Lighteval

Proiecte open-source similare, clasificate după numărul de funcționalități comune cu Lighteval.
  • openai/simple-evalsAvatar openai

    openai/simple-evals

    4,354Vezi pe GitHub↗

    This project is a language model evaluation framework and benchmarking tool designed to measure the accuracy and performance of models across diverse datasets. It provides a system for implementing model-based graders, running standardized tests for mathematical reasoning, coding, and factuality, and calculating quantified performance metrics such as precision, recall, F1 scores, and pass-at-k. The framework utilizes model-based grading and rubrics to validate response quality against expert-defined criteria. It includes a multi-model benchmarking loop and a model-agnostic API interface to co

    Python
    Vezi pe GitHub↗4,354
  • giskard-ai/giskardAvatar Giskard-AI

    Giskard-AI/giskard

    5,434Vezi pe GitHub↗

    Giskard is an evaluation framework, testing library, and quality monitoring system for large language models and AI agents. It serves as a toolkit for quantifying model performance and reliability, providing specialized capabilities for validating retrieval-augmented generation pipelines. The project distinguishes itself through an automated red teaming tool and security scanner designed to identify vulnerabilities, prompt injections, and safety risks. It utilizes adversarial probing and synthetic edge case generation to quantify model robustness and detect information disclosure. The platfo

    Python
    Vezi pe GitHub↗5,434
  • eleutherai/lm-evaluation-harnessAvatar EleutherAI

    EleutherAI/lm-evaluation-harness

    11,460Vezi pe GitHub↗

    This project is a standardized framework for benchmarking large language models across a wide range of academic and reasoning datasets. It provides a platform for executing automated evaluation tasks to measure model accuracy and performance, ensuring consistent assessment through a structured configuration schema. The framework distinguishes itself by incorporating a dedicated utility for data decontamination, which identifies and removes overlapping training samples from evaluation sets to prevent data leakage. It also features a flexible task builder that allows users to define custom benc

    Pythonevaluation-frameworklanguage-modeltransformer
    Vezi pe GitHub↗11,460
  • agenta-ai/agentaAvatar Agenta-AI

    Agenta-AI/agenta

    3,860Vezi pe GitHub↗

    Agenta is a Prompt Ops lifecycle manager and prompt management platform that decouples prompt engineering from application code. It serves as a centralized system for developing, versioning, and deploying prompt templates and model configurations across different environments. The platform functions as an AI agent orchestrator with a visual interface for building agent workflows and connecting models to external tools. It further acts as an evaluation framework and observability tool, utilizing OpenTelemetry to capture execution traces, monitor latency, and track token costs. The system cove

    TypeScriptagentsevaluationllm-as-a-judge
    Vezi pe GitHub↗3,860
Vezi toate cele 30 alternative pentru Lighteval→

Întrebări frecvente

Ce face huggingface/lighteval?

Lighteval is an open-source framework for running standardized benchmarks and custom evaluation tasks against language models. It provides a system for defining new evaluation tasks with custom prompts, metrics, and scoring in YAML configuration files, and integrates with the Hugging Face Hub for storing and comparing results.

Care sunt principalele funcționalități ale huggingface/lighteval?

Principalele funcționalități ale huggingface/lighteval sunt: LLM Evaluation Frameworks, Language Model Benchmark Suites, Custom Evaluation Judges, Model Performance Evaluators, Unified Generation Interfaces, Standardized Benchmarks, Model Evaluation, YAML-Based Task Definitions.

Care sunt câteva alternative open-source pentru huggingface/lighteval?

Alternativele open-source pentru huggingface/lighteval includ: openai/simple-evals — This project is a language model evaluation framework and benchmarking tool designed to measure the accuracy and… giskard-ai/giskard — Giskard is an evaluation framework, testing library, and quality monitoring system for large language models and AI… eleutherai/lm-evaluation-harness — This project is a standardized framework for benchmarking large language models across a wide range of academic and… agenta-ai/agenta — Agenta is a Prompt Ops lifecycle manager and prompt management platform that decouples prompt engineering from… packtpublishing/llm-engineers-handbook — This project is an educational resource and engineering guide for building, deploying, and optimizing large language… lmnr-ai/lmnr — Lmnr is an LLM observability platform and evaluation framework designed for tracing, logging, and monitoring language…