This repository contains information about AGIEval, data, code and output of baseline systems for the benchmark.
Las características principales de ruixiangcui/agieval son: Benchmark Datasets, Model Evaluation, Evaluation Benchmarks.
Las alternativas de código abierto para ruixiangcui/agieval incluyen: sjtu-lit/ceval — Official github repo for C-Eval, a Chinese evaluation suite for foundation models [NeurIPS 2023]. mikegu721/xiezhibenchmark — Xiezhi (獬豸) is a comprehensive evaluation suite for Language Models (LMs). It consists of 249587 multi-choice… michael-wzhu/promptcblue — PromptCBLUE: a large-scale instruction-tuning dataset for multi-task and few-shot learning in the medical domain in… haonan-li/cmmlu — CMMLU: Measuring massive multitask language understanding in Chinese. nvidia/isaac-gr00t. sjwhitworth/golearn — GoLearn is a machine learning library for the Go programming language. It provides a supervised learning framework and…
Xiezhi (獬豸) is a comprehensive evaluation suite for Language Models (LMs). It consists of 249587 multi-choice questions spanning 516 diverse disciplines and four difficulty levels, as shown below. Please check our paper for more details, and our website will be open later on.
PromptCBLUE: a large-scale instruction-tuning dataset for multi-task and few-shot learning in the medical domain in Chinese
CMMLU: Measuring massive multitask language understanding in Chinese
Official github repo for C-Eval, a Chinese evaluation suite for foundation models NeurIPS 2023