awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Back to flagopen/flageval

Open-source alternatives to FlagEval

30 open-source projects similar to flagopen/flageval, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best FlagEval alternative.

  • sjtu-lit/cevalSJTU-LIT avatar

    SJTU-LIT/ceval

    1,854View on GitHub↗

    Official github repo for C-Eval, a Chinese evaluation suite for foundation models NeurIPS 2023

    Python
    View on GitHub↗1,854
  • thu-coai/safety-promptsthu-coai avatar

    thu-coai/Safety-Prompts

    1,176View on GitHub↗

    Chinese safety prompts for evaluating and improving the safety of LLMs. 中文安全prompts,用于评估和提升大模型的安全性。

    attack-defensechatgptchinese-language
    View on GitHub↗1,176
  • cluebenchmark/supercluelybCLUEbenchmark avatar

    CLUEbenchmark/SuperCLUElyb

    144View on GitHub↗

    SuperCLUE琅琊榜:中文通用大模型匿名对战评价基准

    View on GitHub↗144
  • open-compass/opencompassopen-compass avatar

    open-compass/opencompass

    6,678View on GitHub↗

    OpenCompass is an open-source framework for standardized benchmarking of large language models. It provides a configurable evaluation pipeline that supports both objective and subjective assessment, using a dual-engine architecture to handle closed-form answer comparison and open-ended response rating. The framework is designed as a modular platform where datasets, models, and metrics are composed through declarative YAML configuration files. The framework distinguishes itself through its extensible model integration layer, which supports custom models, HuggingFace models, and third-party API

    Pythonbenchmarkchatgptevaluation
    View on GitHub↗6,678

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Find more with AI search
mikegu721/xiezhibenchmarkMikeGu721 avatar

MikeGu721/XiezhiBenchmark

98View on GitHub↗

Xiezhi (獬豸) is a comprehensive evaluation suite for Language Models (LMs). It consists of 249587 multi-choice questions spanning 516 diverse disciplines and four difficulty levels, as shown below. Please check our paper for more details, and our website will be open later on.

Python
View on GitHub↗98
  • sjwhitworth/golearnsjwhitworth avatar

    sjwhitworth/golearn

    9,438View on GitHub↗

    GoLearn is a machine learning library for the Go programming language. It provides a supervised learning framework and a toolkit for building, training, and evaluating predictive models through a standardized interface. The project implements a data frame system that loads CSV files into structured grids for matrix operations. It includes a preprocessing library for discretizing continuous variables and a model evaluation toolkit that utilizes confusion matrices and cross-validation to measure precision and recall. The library covers data engineering and management, including the ability to

    Go
    View on GitHub↗9,438
  • nvidia/isaac-gr00tNVIDIA avatar

    NVIDIA/Isaac-GR00T

    6,222View on GitHub↗
    Jupyter Notebook
    View on GitHub↗6,222
  • ageron/handson-mlageron avatar

    ageron/handson-ml

    25,608View on GitHub↗

    This is a machine learning educational repository consisting of a collection of notebooks and code examples. It provides practical implementations of diverse machine learning algorithms and workflows, ranging from traditional scientific computing to deep learning. The project features specific implementations of Scikit-Learn models, such as decision trees, random forests, and support vector machines, as well as TensorFlow examples for building neural networks, convolutional layers, and recurrent architectures. It also includes tutorials on reinforcement learning development and the creation o

    Jupyter Notebook
    View on GitHub↗25,608
  • xiami2019/halluqaxiami2019 avatar

    xiami2019/HalluQA

    139View on GitHub↗

    Dataset and evaluation script for "Evaluating Hallucinations in Chinese Large Language Models"

    Python
    View on GitHub↗139
  • eth-sri/matharenaE

    eth-sri/matharena

    0View on GitHub↗
    View on GitHub↗0
  • haonan-li/cmmluhaonan-li avatar

    haonan-li/CMMLU

    821View on GitHub↗

    CMMLU: Measuring massive multitask language understanding in Chinese

    Python
    View on GitHub↗821
  • huggingface/evaluatehuggingface avatar

    huggingface/evaluate

    2,455View on GitHub↗

    🤗 Evaluate: A library for easily evaluating machine learning models and datasets.

    Python
    View on GitHub↗2,455
  • huggingface/evaluation-guidebookhuggingface avatar

    huggingface/evaluation-guidebook

    2,125View on GitHub↗

    Sharing both practical insights and theoretical knowledge about LLM evaluation that we gathered while managing the Open LLM Leaderboard and designing lighteval!

    Jupyter Notebookevaluationevaluation-metricsguidebook
    View on GitHub↗2,125
  • huggingface/lightevalhuggingface avatar

    huggingface/lighteval

    2,453View on GitHub↗

    Lighteval is an open-source framework for running standardized benchmarks and custom evaluation tasks against language models. It provides a system for defining new evaluation tasks with custom prompts, metrics, and scoring in YAML configuration files, and integrates with the Hugging Face Hub for storing and comparing results. The framework supports evaluating models across multiple inference backends, including transformers, vllm, and custom APIs, through a unified generation and log-probability interface. It includes a pluggable metric registry for built-in and custom scoring, a prediction

    Pythonevaluationevaluation-frameworkevaluation-metrics
    View on GitHub↗2,453
  • huggingface/yourbenchH

    huggingface/yourbench

    0View on GitHub↗
    View on GitHub↗0
  • ibm/aif360IBM avatar

    IBM/AIF360

    2,827View on GitHub↗

    A comprehensive set of fairness metrics for datasets and machine learning models, explanations for these metrics, and algorithms to mitigate bias in datasets and models.

    Python
    View on GitHub↗2,827
  • internlm/opencompassInternLM avatar

    InternLM/opencompass

    7,096View on GitHub↗

    OpenCompass is a comprehensive evaluation platform, benchmarking suite, and distributed model evaluator designed to measure the performance and accuracy of large language models. It provides a framework for benchmarking both open-source and API-based models against diverse datasets using standardized metrics and reproducible pipelines. The project features an automated judging framework that uses language models as judges to score and verify the quality of generated text. It includes a performance leaderboard system for comparing the relative capabilities of various models across industry-sta

    Python
    View on GitHub↗7,096
  • langchain-ai/auto-evaluatorlangchain-ai avatar

    langchain-ai/auto-evaluator

    780View on GitHub↗

    Context

    TypeScript
    View on GitHub↗780
  • michael-wzhu/promptcbluemichael-wzhu avatar

    michael-wzhu/PromptCBLUE

    394View on GitHub↗

    PromptCBLUE: a large-scale instruction-tuning dataset for multi-task and few-shot learning in the medical domain in Chinese

    Python
    View on GitHub↗394
  • eleutherai/lm-evaluation-harnessEleutherAI avatar

    EleutherAI/lm-evaluation-harness

    11,460View on GitHub↗

    This project is a standardized framework for benchmarking large language models across a wide range of academic and reasoning datasets. It provides a platform for executing automated evaluation tasks to measure model accuracy and performance, ensuring consistent assessment through a structured configuration schema. The framework distinguishes itself by incorporating a dedicated utility for data decontamination, which identifies and removes overlapping training samples from evaluation sets to prevent data leakage. It also features a flexible task builder that allows users to define custom benc

    Pythonevaluation-frameworklanguage-modeltransformer
    View on GitHub↗11,460
  • mlfoundations/evalchemymlfoundations avatar

    mlfoundations/evalchemy

    597View on GitHub↗

    Automatic evals for LLMs

    HTML
    View on GitHub↗597
  • modelscope/evalscopemodelscope avatar

    modelscope/evalscope

    2,955View on GitHub↗

    A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.

    Python
    View on GitHub↗2,955
  • modelscope/openjudgeM

    modelscope/OpenJudge

    0View on GitHub↗
    View on GitHub↗0
  • noudald/pyrocN

    noudald/pyroc

    0View on GitHub↗
    View on GitHub↗0
  • edublancas/sklearn-evaluationedublancas avatar

    edublancas/sklearn-evaluation

    3View on GitHub↗

    Machine learning model evaluation made easy: plots, tables, HTML reports, experiment tracking and Jupyter notebook analysis.

    View on GitHub↗3
  • confident-ai/deepevalconfident-ai avatar

    confident-ai/deepeval

    13,733View on GitHub↗

    Deepeval is a framework for testing and evaluating large language model applications. It provides a suite of tools for executing automated regression tests, validating model output quality against defined standards, and tracing the execution of complex agent workflows. By integrating these capabilities into development pipelines, the platform ensures consistent performance and reliability throughout the software lifecycle. The platform distinguishes itself through its focus on programmatic validation and observability. It utilizes secondary language models to score output quality and employs

    Pythonevaluation-frameworkevaluation-metricsllm-evaluation
    View on GitHub↗13,733
  • open-compass/vlmevalkitopen-compass avatar

    open-compass/VLMEvalKit

    3,824View on GitHub↗

    VLMEvalKit is a vision-language model evaluation framework and inference engine designed to run standardized benchmarks and measure model accuracy across diverse visual datasets. It serves as a multimodal model benchmark and performance toolkit for calculating metrics and comparing model responses. The toolkit includes a specialized visual reasoning evaluator that uses adversarial samples to distinguish actual image understanding from reliance on language patterns. It also provides capabilities for image generation evaluation, testing a model's ability to create or modify visuals based on tex

    Pythonchatgptclaudeclip
    View on GitHub↗3,824
  • openlmlab/gaokao-benchOpenLMLab avatar

    OpenLMLab/GAOKAO-Bench

    760View on GitHub↗

    GAOKAO-Bench is an evaluation framework that utilizes GAOKAO questions as a dataset to evaluate large language models.

    Python
    View on GitHub↗760
  • pair-code/llm-comparatorPAIR-code avatar

    PAIR-code/llm-comparator

    528View on GitHub↗

    LLM Comparator is an interactive data visualization tool for evaluating and analyzing LLM responses side-by-side, developed by the PAIR team.

    JavaScript
    View on GitHub↗528
  • pandas-ml/pandas-mlP

    pandas-ml/pandas-ml

    0View on GitHub↗
    View on GitHub↗0