awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
ruixiangcui avatar

ruixiangcui/AGIEval

0
View on GitHub↗
775 stars·54 forks·Python·MIT·14 views

AGIEval

This repository contains information about AGIEval, data, code and output of baseline systems for the benchmark.

Features

  • Benchmark Datasets - Benchmark for evaluating foundation models on human-centric tasks.
  • Model Evaluation - Benchmark testing models against standardized human-level academic and professional exams.
  • Evaluation Benchmarks - Benchmark for human-level cognitive and qualification exams.

Star history

Star history chart for ruixiangcui/agievalStar history chart for ruixiangcui/agieval

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Projects sharing features with AGIEval

These projects share indexed features with AGIEval. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • mikegu721/xiezhibenchmarkMikeGu721 avatar

    MikeGu721/XiezhiBenchmark

    98View on GitHub↗

    Xiezhi (獬豸) is a comprehensive evaluation suite for Language Models (LMs). It consists of 249587 multi-choice questions spanning 516 diverse disciplines and four difficulty levels, as shown below. Please check our paper for more details, and our website will be open later on.

    Python
    View on GitHub↗98
  • michael-wzhu/promptcbluemichael-wzhu avatar

    michael-wzhu/PromptCBLUE

    394View on GitHub↗

    PromptCBLUE: a large-scale instruction-tuning dataset for multi-task and few-shot learning in the medical domain in Chinese

    Python
    View on GitHub↗394
  • haonan-li/cmmluhaonan-li avatar

    haonan-li/CMMLU

    821View on GitHub↗

    CMMLU: Measuring massive multitask language understanding in Chinese

    Python
    View on GitHub↗821
  • sjtu-lit/cevalSJTU-LIT avatar

    SJTU-LIT/ceval

    1,854View on GitHub↗

    Official github repo for C-Eval, a Chinese evaluation suite for foundation models NeurIPS 2023

    Python
    View on GitHub↗1,854
Compare all 30 related projects→

Frequently asked questions

What does ruixiangcui/agieval do?

This repository contains information about AGIEval, data, code and output of baseline systems for the benchmark.

What are the main features of ruixiangcui/agieval?

The main features of ruixiangcui/agieval are: Benchmark Datasets, Model Evaluation, Evaluation Benchmarks.

Which projects share features with ruixiangcui/agieval?

Projects with overlapping indexed features include: sjtu-lit/ceval — Official github repo for C-Eval, a Chinese evaluation suite for foundation models [NeurIPS 2023]. mikegu721/xiezhibenchmark — Xiezhi (獬豸) is a comprehensive evaluation suite for Language Models (LMs). It consists of 249587 multi-choice… michael-wzhu/promptcblue — PromptCBLUE: a large-scale instruction-tuning dataset for multi-task and few-shot learning in the medical domain in… haonan-li/cmmlu — CMMLU: Measuring massive multitask language understanding in Chinese. nvidia/isaac-gr00t. sjwhitworth/golearn — GoLearn is a machine learning library for the Go programming language. It provides a supervised learning framework and…