How this analysis was created: This summary and feature list were written by an AI model that read the project's README and public documentation pages. Each feature links to the documentation it came from; stars, license and language come straight from the GitHub API. The model does not read the source code, and the analysis is refreshed when the project is re-analysed. Learn more on our About page.
This repository contains information about AGIEval, data, code and output of baseline systems for the benchmark.
The main features of ruixiangcui/agieval are: Benchmark Datasets, Model Evaluation, Evaluation Benchmarks.
Open-source alternatives to ruixiangcui/agieval include: sjtu-lit/ceval — Official github repo for C-Eval, a Chinese evaluation suite for foundation models [NeurIPS 2023]. mikegu721/xiezhibenchmark — Xiezhi (獬豸) is a comprehensive evaluation suite for Language Models (LMs). It consists of 249587 multi-choice… michael-wzhu/promptcblue — PromptCBLUE: a large-scale instruction-tuning dataset for multi-task and few-shot learning in the medical domain in… haonan-li/cmmlu — CMMLU: Measuring massive multitask language understanding in Chinese. nvidia/isaac-gr00t. sjwhitworth/golearn — GoLearn is a machine learning library for the Go programming language. It provides a supervised learning framework and…
Xiezhi (獬豸) is a comprehensive evaluation suite for Language Models (LMs). It consists of 249587 multi-choice questions spanning 516 diverse disciplines and four difficulty levels, as shown below. Please check our paper for more details, and our website will be open later on.
PromptCBLUE: a large-scale instruction-tuning dataset for multi-task and few-shot learning in the medical domain in Chinese
CMMLU: Measuring massive multitask language understanding in Chinese
Official github repo for C-Eval, a Chinese evaluation suite for foundation models NeurIPS 2023