awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
huggingface avatar

huggingface/lighteval

0
View on GitHub↗
2,453 stars·490 forks·Python·MIT·32 viewshuggingface.co/docs/lighteval/en/index↗

Lighteval

Lighteval is an open-source framework for running standardized benchmarks and custom evaluation tasks against language models. It provides a system for defining new evaluation tasks with custom prompts, metrics, and scoring in YAML configuration files, and integrates with the Hugging Face Hub for storing and comparing results.

The framework supports evaluating models across multiple inference backends, including transformers, vllm, and custom APIs, through a unified generation and log-probability interface. It includes a pluggable metric registry for built-in and custom scoring, a prediction cache layer that avoids re-computation on repeated runs, and the ability to push results to remote storage destinations like the Hugging Face Hub, S3, or local directories. Per-sample outputs are saved to enable detailed debugging and performance analysis after evaluation.

Evaluation tasks are configured in YAML files specifying prompt templates, reference answers, and scoring metrics, with support for subset-based task splitting that generates separate configurations for each dataset subset automatically. The framework includes thousands of pre-built benchmark tasks and allows users to define custom tasks, custom metrics, and custom model implementations.

Features

  • LLM Evaluation Frameworks - Provides an open-source framework for running standardized benchmarks and custom evaluation tasks against language models.
  • Language Model Benchmark Suites - Runs pre-built and custom evaluation tasks against language models to measure performance.
  • Custom Evaluation Judges - Builds new evaluation tasks with custom metrics and scoring to test model behavior on user-defined criteria.
  • Model Performance Evaluators - Measures model outputs against predefined metrics to quantify accuracy and quality across diverse tasks.
  • Unified Generation Interfaces - Provides a unified generation and log-probability interface across transformers, vllm, and custom API backends.
  • Standardized Benchmarks - Executes thousands of ready-made evaluation tasks to measure model performance on standard tests.
  • Model Evaluation - Runs models through a curated set of tasks and computes standard metrics to measure performance.
  • YAML-Based Task Definitions - Defines evaluation tasks in YAML files specifying prompt templates, reference answers, and scoring metrics.
  • Evaluation Task Definitions - Defines new evaluation tasks with custom prompts, metrics, and scoring in YAML configuration files.
  • YAML Benchmark Definitions - Creates new benchmarks by specifying prompt templates, reference answers, and scoring metrics in YAML files.
  • Custom Model Evaluation Adapters - Runs benchmark tasks against user-defined models by implementing subclasses with generation and log-probability methods.
  • Evaluation Prediction Caches - Caches computed model outputs to disk or memory, skipping re-computation on repeated evaluation runs.
  • Metric Selection and Creation - Selects built-in metrics or defines custom ones to score model outputs during evaluation.
  • Evaluation Prompt Functions - Converts dataset lines into structured documents with query, choices, and gold index for evaluation.
  • Evaluation Task Executors - Executes user-defined evaluation tasks against a model using the CLI with a custom task file.
  • Inference Backend Abstraction - Supports running evaluations across transformers, vllm, and custom API backends through a unified interface.
  • Inference Engine Evaluation Backends - Runs model evaluation using various inference engines like transformers, vllm, and custom APIs.
  • Single-Subset Evaluation Configurations - Defines single-subset tasks with a prompt function, dataset source, splits, metrics, and generation parameters.
  • Result Storage Integrations - Stores evaluation results on the Hugging Face Hub for sharing and comparing model performance.
  • Custom Metric Registries - Ships a pluggable metric registry where built-in and custom scoring functions are registered by name.
  • Datasets and Evaluation - Library for evaluating local LLMs on benchmarks.
  • Evaluation and Observability - Lightweight evaluation framework from Hugging Face.
  • Evaluation Frameworks - Unified toolkit for benchmarking and evaluating language models.
  • Model Evaluation - All-in-one toolkit for evaluating models across multiple backends.
  • Model Evaluation and Benchmarking - Lightweight evaluation suite for language models.

Star history

Star history chart for huggingface/lightevalStar history chart for huggingface/lighteval

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Projects sharing features with Lighteval

These projects share indexed features with Lighteval. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • giskard-ai/giskardGiskard-AI avatar

    Giskard-AI/giskard

    5,434View on GitHub↗

    Giskard is an evaluation framework, testing library, and quality monitoring system for large language models and AI agents. It serves as a toolkit for quantifying model performance and reliability, providing specialized capabilities for validating retrieval-augmented generation pipelines. The project distinguishes itself through an automated red teaming tool and security scanner designed to identify vulnerabilities, prompt injections, and safety risks. It utilizes adversarial probing and synthetic edge case generation to quantify model robustness and detect information disclosure. The platfo

    Python
    View on GitHub↗5,434
  • openai/simple-evalsopenai avatar

    openai/simple-evals

    4,354View on GitHub↗

    This project is a language model evaluation framework and benchmarking tool designed to measure the accuracy and performance of models across diverse datasets. It provides a system for implementing model-based graders, running standardized tests for mathematical reasoning, coding, and factuality, and calculating quantified performance metrics such as precision, recall, F1 scores, and pass-at-k. The framework utilizes model-based grading and rubrics to validate response quality against expert-defined criteria. It includes a multi-model benchmarking loop and a model-agnostic API interface to co

    Python
    View on GitHub↗4,354
  • eleutherai/lm-evaluation-harnessEleutherAI avatar

    EleutherAI/lm-evaluation-harness

    11,460View on GitHub↗

    This project is a standardized framework for benchmarking large language models across a wide range of academic and reasoning datasets. It provides a platform for executing automated evaluation tasks to measure model accuracy and performance, ensuring consistent assessment through a structured configuration schema. The framework distinguishes itself by incorporating a dedicated utility for data decontamination, which identifies and removes overlapping training samples from evaluation sets to prevent data leakage. It also features a flexible task builder that allows users to define custom benc

    Pythonevaluation-frameworklanguage-modeltransformer
    View on GitHub↗11,460
  • confident-ai/deepevalconfident-ai avatar

    confident-ai/deepeval

    13,733View on GitHub↗

    Deepeval is a framework for testing and evaluating large language model applications. It provides a suite of tools for executing automated regression tests, validating model output quality against defined standards, and tracing the execution of complex agent workflows. By integrating these capabilities into development pipelines, the platform ensures consistent performance and reliability throughout the software lifecycle. The platform distinguishes itself through its focus on programmatic validation and observability. It utilizes secondary language models to score output quality and employs

    Pythonevaluation-frameworkevaluation-metricsllm-evaluation
    View on GitHub↗13,733
Compare all 30 related projects→

Frequently asked questions

What does huggingface/lighteval do?

Lighteval is an open-source framework for running standardized benchmarks and custom evaluation tasks against language models. It provides a system for defining new evaluation tasks with custom prompts, metrics, and scoring in YAML configuration files, and integrates with the Hugging Face Hub for storing and comparing results.

What are the main features of huggingface/lighteval?

The main features of huggingface/lighteval are: LLM Evaluation Frameworks, Language Model Benchmark Suites, Custom Evaluation Judges, Model Performance Evaluators, Unified Generation Interfaces, Standardized Benchmarks, Model Evaluation, YAML-Based Task Definitions.

Which projects share features with huggingface/lighteval?

Projects with overlapping indexed features include: giskard-ai/giskard — Giskard is an evaluation framework, testing library, and quality monitoring system for large language models and AI… openai/simple-evals — This project is a language model evaluation framework and benchmarking tool designed to measure the accuracy and… eleutherai/lm-evaluation-harness — This project is a standardized framework for benchmarking large language models across a wide range of academic and… confident-ai/deepeval — Deepeval is a framework for testing and evaluating large language model applications. It provides a suite of tools for… explodinggradients/ragas — Ragas is an evaluation framework and performance benchmark designed to quantify the quality of retrieval augmented… openai/evals — Evals is a framework designed for automating, managing, and executing repeatable benchmarking suites to analyze the…