This repository contains information about AGIEval, data, code and output of baseline systems for the benchmark.
Xiezhi (獬豸) is a comprehensive evaluation suite for Language Models (LMs). It consists of 249587 multi-choice questions spanning 516 diverse disciplines and four difficulty levels, as shown below. Please check our paper for more details, and our website will be open later on.
PromptCBLUE: a large-scale instruction-tuning dataset for multi-task and few-shot learning in the medical domain in Chinese
Official github repo for C-Eval, a Chinese evaluation suite for foundation models NeurIPS 2023
CMMLU: Measuring massive multitask language understanding in Chinese
The main features of haonan-li/cmmlu are: Model Evaluation, Evaluation Benchmarks.
Open-source alternatives to haonan-li/cmmlu include: sjtu-lit/ceval — Official github repo for C-Eval, a Chinese evaluation suite for foundation models [NeurIPS 2023]. ruixiangcui/agieval — This repository contains information about AGIEval, data, code and output of baseline systems for the benchmark. mikegu721/xiezhibenchmark — Xiezhi (獬豸) is a comprehensive evaluation suite for Language Models (LMs). It consists of 249587 multi-choice… michael-wzhu/promptcblue — PromptCBLUE: a large-scale instruction-tuning dataset for multi-task and few-shot learning in the medical domain in… nvidia/isaac-gr00t. sjwhitworth/golearn — GoLearn is a machine learning library for the Go programming language. It provides a supervised learning framework and…