awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
SJTU-LIT avatar

SJTU-LIT/ceval

0
View on GitHub↗
1,854 stars·84 forks·Python·MIT·14 viewscevalbenchmark.com↗

Ceval

Official github repo for C-Eval, a Chinese evaluation suite for foundation models [NeurIPS 2023]

Features

  • Model Evaluation - Knowledge-based evaluation benchmark covering diverse academic and professional subjects.
  • Evaluation Benchmarks - Comprehensive benchmark suite for evaluating Chinese language models.
  • Evaluation Benchmarks - Comprehensive benchmark for Chinese academic and professional knowledge.

Star history

Star history chart for sjtu-lit/cevalStar history chart for sjtu-lit/ceval

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Projects sharing features with Ceval

These projects share indexed features with Ceval. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • mikegu721/xiezhibenchmarkMikeGu721 avatar

    MikeGu721/XiezhiBenchmark

    98View on GitHub↗

    Xiezhi (獬豸) is a comprehensive evaluation suite for Language Models (LMs). It consists of 249587 multi-choice questions spanning 516 diverse disciplines and four difficulty levels, as shown below. Please check our paper for more details, and our website will be open later on.

    Python
    View on GitHub↗98
  • flagopen/flagevalFlagOpen avatar

    FlagOpen/FlagEval

    13View on GitHub↗

    FlagEval, launched by BAAI in 2023, is a comprehensive large model evaluation system that encompasses over 800 open-source and closed-source models from around the globe. It features more than 40 capability dimensions, including reasoning, mathematical skills, and task-solving abilities, along…

    View on GitHub↗13
  • cluebenchmark/supercluelybCLUEbenchmark avatar

    CLUEbenchmark/SuperCLUElyb

    144View on GitHub↗

    SuperCLUE琅琊榜:中文通用大模型匿名对战评价基准

    View on GitHub↗144
  • haonan-li/cmmluhaonan-li avatar

    haonan-li/CMMLU

    821View on GitHub↗

    CMMLU: Measuring massive multitask language understanding in Chinese

    Python
    View on GitHub↗821
Compare all 30 related projects→

Frequently asked questions

What does sjtu-lit/ceval do?

Official github repo for C-Eval, a Chinese evaluation suite for foundation models [NeurIPS 2023]

What are the main features of sjtu-lit/ceval?

The main features of sjtu-lit/ceval are: Model Evaluation, Evaluation Benchmarks.

Which projects share features with sjtu-lit/ceval?

Projects with overlapping indexed features include: mikegu721/xiezhibenchmark — Xiezhi (獬豸) is a comprehensive evaluation suite for Language Models (LMs). It consists of 249587 multi-choice… michael-wzhu/promptcblue — PromptCBLUE: a large-scale instruction-tuning dataset for multi-task and few-shot learning in the medical domain in… flagopen/flageval — FlagEval, launched by BAAI in 2023, is a comprehensive large model evaluation system that encompasses over 800… cluebenchmark/supercluelyb — SuperCLUE琅琊榜:中文通用大模型匿名对战评价基准. haonan-li/cmmlu — CMMLU: Measuring massive multitask language understanding in Chinese. open-compass/opencompass — OpenCompass is an open-source framework for standardized benchmarking of large language models. It provides a…