awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Back to ruixiangcui/agieval

Projects sharing features with AGIEval

30 open-source projects similar to ruixiangcui/agieval, ranked by shared indexed features. Tags may describe platforms or build tools rather than the same primary purpose. Check each project’s use case, license, and deployment requirements before treating it as a replacement.

  • mikegu721/xiezhibenchmarkMikeGu721 avatar

    MikeGu721/XiezhiBenchmark

    98View on GitHub↗

    Xiezhi (獬豸) is a comprehensive evaluation suite for Language Models (LMs). It consists of 249587 multi-choice questions spanning 516 diverse disciplines and four difficulty levels, as shown below. Please check our paper for more details, and our website will be open later on.

    Python
    View on GitHub↗98
  • michael-wzhu/promptcbluemichael-wzhu avatar

    michael-wzhu/PromptCBLUE

    394View on GitHub↗

    PromptCBLUE: a large-scale instruction-tuning dataset for multi-task and few-shot learning in the medical domain in Chinese

    Python
    View on GitHub↗394
  • sjtu-lit/cevalSJTU-LIT avatar

    SJTU-LIT/ceval

    1,854View on GitHub↗

    Official github repo for C-Eval, a Chinese evaluation suite for foundation models NeurIPS 2023

    Python
    View on GitHub↗1,854
  • haonan-li/cmmluhaonan-li avatar

    haonan-li/CMMLU

    821View on GitHub↗

    CMMLU: Measuring massive multitask language understanding in Chinese

    Python
    View on GitHub↗821
  • sjwhitworth/golearnsjwhitworth avatar

    sjwhitworth/golearn

    9,438View on GitHub↗

    GoLearn is a machine learning library for the Go programming language. It provides a supervised learning framework and a toolkit for building, training, and evaluating predictive models through a standardized interface. The project implements a data frame system that loads CSV files into structured grids for matrix operations. It includes a preprocessing library for discretizing continuous variables and a model evaluation toolkit that utilizes confusion matrices and cross-validation to measure precision and recall. The library covers data engineering and management, including the ability to

    Go
    View on GitHub↗9,438

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Find more with AI search
  • nvidia/isaac-gr00tNVIDIA avatar

    NVIDIA/Isaac-GR00T

    6,222View on GitHub↗
    Jupyter Notebook
    View on GitHub↗6,222
  • ageron/handson-mlageron avatar

    ageron/handson-ml

    25,608View on GitHub↗

    This is a machine learning educational repository consisting of a collection of notebooks and code examples. It provides practical implementations of diverse machine learning algorithms and workflows, ranging from traditional scientific computing to deep learning. The project features specific implementations of Scikit-Learn models, such as decision trees, random forests, and support vector machines, as well as TensorFlow examples for building neural networks, convolutional layers, and recurrent architectures. It also includes tutorials on reinforcement learning development and the creation o

    Jupyter Notebook
    View on GitHub↗25,608
  • coastalcph/lex-gluecoastalcph avatar

    coastalcph/lex-glue

    259View on GitHub↗

    LexGLUE: A Benchmark Dataset for Legal Language Understanding in English

    Python
    View on GitHub↗259
  • codefuse-ai/codefuse-devops-evalcodefuse-ai avatar

    codefuse-ai/codefuse-devops-eval

    656View on GitHub↗

    Industrial-first evaluation benchmark for LLMs in the DevOps/AIOps domain.

    Python
    View on GitHub↗656
  • confident-ai/deepevalconfident-ai avatar

    confident-ai/deepeval

    13,733View on GitHub↗

    Deepeval is a framework for testing and evaluating large language model applications. It provides a suite of tools for executing automated regression tests, validating model output quality against defined standards, and tracing the execution of complex agent workflows. By integrating these capabilities into development pipelines, the platform ensures consistent performance and reliability throughout the software lifecycle. The platform distinguishes itself through its focus on programmatic validation and observability. It utilizes secondary language models to score output quality and employs

    Pythonevaluation-frameworkevaluation-metricsllm-evaluation
    View on GitHub↗13,733
  • dai-shen/laiwDai-shen avatar

    Dai-shen/LAiW

    91View on GitHub↗

    LAiW: A Chinese Legal Large Language Models Benchmark

    Python
    View on GitHub↗91
  • drcknowledgeteam/drcdD

    DRCKnowledgeTeam/DRCD

    0View on GitHub↗
    View on GitHub↗0
  • edublancas/sklearn-evaluationedublancas avatar

    edublancas/sklearn-evaluation

    3View on GitHub↗

    Machine learning model evaluation made easy: plots, tables, HTML reports, experiment tracking and Jupyter notebook analysis.

    View on GitHub↗3
  • eleutherai/lm-evaluation-harnessEleutherAI avatar

    EleutherAI/lm-evaluation-harness

    11,460View on GitHub↗

    This project is a standardized framework for benchmarking large language models across a wide range of academic and reasoning datasets. It provides a platform for executing automated evaluation tasks to measure model accuracy and performance, ensuring consistent assessment through a structured configuration schema. The framework distinguishes itself by incorporating a dedicated utility for data decontamination, which identifies and removes overlapping training samples from evaluation sets to prevent data leakage. It also features a flexible task builder that allows users to define custom benc

    Pythonevaluation-frameworklanguage-modeltransformer
    View on GitHub↗11,460
  • eth-sri/matharenaE

    eth-sri/matharena

    0View on GitHub↗
    View on GitHub↗0
  • felixgithub2017/cg-evalFelixgithub2017 avatar

    Felixgithub2017/CG-Eval

    13View on GitHub↗

    Chinese Generation Evaluation

    View on GitHub↗13
  • felixgithub2017/mmcuFelixgithub2017 avatar

    Felixgithub2017/MMCU

    90View on GitHub↗

    MEASURING MASSIVE MULTITASK CHINESE UNDERSTANDING

    Python
    View on GitHub↗90
  • flagopen/flagevalFlagOpen avatar

    FlagOpen/FlagEval

    13View on GitHub↗

    FlagEval, launched by BAAI in 2023, is a comprehensive large model evaluation system that encompasses over 800 open-source and closed-source models from around the globe. It features more than 40 capability dimensions, including reasoning, mathematical skills, and task-solving abilities, along…

    View on GitHub↗13
  • freedomintelligence/cmbFreedomIntelligence avatar

    FreedomIntelligence/CMB

    244View on GitHub↗

    CMB, A Comprehensive Medical Benchmark in Chinese

    Python
    View on GitHub↗244
  • hazyresearch/legalbenchHazyResearch avatar

    HazyResearch/legalbench

    597View on GitHub↗

    An open science effort to benchmark legal reasoning in foundation models

    Python
    View on GitHub↗597
  • hc-guo/owlHC-Guo avatar

    HC-Guo/Owl

    237View on GitHub↗

    A Large Language Model for IT Operations

    Python
    View on GitHub↗237
  • huggingface/evaluatehuggingface avatar

    huggingface/evaluate

    2,455View on GitHub↗

    🤗 Evaluate: A library for easily evaluating machine learning models and datasets.

    Python
    View on GitHub↗2,455
  • huggingface/evaluation-guidebookhuggingface avatar

    huggingface/evaluation-guidebook

    2,125View on GitHub↗

    Sharing both practical insights and theoretical knowledge about LLM evaluation that we gathered while managing the Open LLM Leaderboard and designing lighteval!

    Jupyter Notebookevaluationevaluation-metricsguidebook
    View on GitHub↗2,125
  • huggingface/lightevalhuggingface avatar

    huggingface/lighteval

    2,453View on GitHub↗

    Lighteval is an open-source framework for running standardized benchmarks and custom evaluation tasks against language models. It provides a system for defining new evaluation tasks with custom prompts, metrics, and scoring in YAML configuration files, and integrates with the Hugging Face Hub for storing and comparing results. The framework supports evaluating models across multiple inference backends, including transformers, vllm, and custom APIs, through a unified generation and log-probability interface. It includes a pluggable metric registry for built-in and custom scoring, a prediction

    Pythonevaluationevaluation-frameworkevaluation-metrics
    View on GitHub↗2,453
  • huggingface/yourbenchH

    huggingface/yourbench

    0View on GitHub↗
    View on GitHub↗0
  • ibm/aif360IBM avatar

    IBM/AIF360

    2,827View on GitHub↗

    A comprehensive set of fairness metrics for datasets and machine learning models, explanations for these metrics, and algorithms to mitigate bias in datasets and models.

    Python
    View on GitHub↗2,827
  • internlm/opencompassInternLM avatar

    InternLM/opencompass

    7,096View on GitHub↗

    OpenCompass is a comprehensive evaluation platform, benchmarking suite, and distributed model evaluator designed to measure the performance and accuracy of large language models. It provides a framework for benchmarking both open-source and API-based models against diverse datasets using standardized metrics and reproducible pipelines. The project features an automated judging framework that uses language models as judges to score and verify the quality of generated text. It includes a performance leaderboard system for comparing the relative capabilities of various models across industry-sta

    Python
    View on GitHub↗7,096
  • joelniklaus/lextremeJoelNiklaus avatar

    JoelNiklaus/LEXTREME

    25View on GitHub↗

    This repository provides scripts for evaluating NLP models on the LEXTREME benchmark, a set of diverse multilingual tasks in legal NLP

    Python
    View on GitHub↗25
  • langchain-ai/auto-evaluatorlangchain-ai avatar

    langchain-ai/auto-evaluator

    780View on GitHub↗

    Context

    TypeScript
    View on GitHub↗780
  • mlfoundations/evalchemymlfoundations avatar

    mlfoundations/evalchemy

    597View on GitHub↗

    Automatic evals for LLMs

    HTML
    View on GitHub↗597