30 open-source projects similar to cbluebenchmark/cblue, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best CBLUE alternative.
This is the evaluation data and Large Language Models (LLMs) results from our CIKM'23 paper:
This repo presents OpenP5, an open-source platform for LLM-based Recommendation development, finetuning, and evaluation.
The ApolloScape Open Dataset for Autonomous Driving and its Application.
This repository contains utility scripts for the KITTI-360 dataset.
This repository introduces PIXIU, an open-source resource featuring the first financial large language models (LLMs), instruction tuning data, and evaluation benchmarks to holistically assess financial LLMs. Our goal is to continually push forward the open-source development of financial artificial intelligence (AI).
🛰️ List of satellite image training datasets with annotations for computer vision and deep learning
LexGLUE: A Benchmark Dataset for Legal Language Understanding in English
Industrial-first evaluation benchmark for LLMs in the DevOps/AIOps domain.
A procedural Blender pipeline for photorealistic training image generation
DBNet: A Large-Scale Dataset for Driving Behavior Learning, CVPR 2018
Objectron is a dataset of short, object-centric video clips. In addition, the videos also contain AR session metadata including camera poses, sparse point-clouds and planes. In each video, the camera moves around and above the object and captures it from different views. Each object is annotated with a 3D bounding box. The 3D bounding box describes the object’s position, orientation, and dimensions. The dataset contains about 15K annotated video clips and 4M annotated images in the following categories: bikes, books, bottles, cameras, cereal boxes, chairs, cups, laptops, and shoes
CMMLU: Measuring massive multitask language understanding in Chinese
An open science effort to benchmark legal reasoning in foundation models
PyTorch implementation for WSDM 2024 paper LLMRec: Large Language Models with Graph Augmentation for Recommendation.
This repository provides scripts for evaluating NLP models on the LEXTREME benchmark, a set of diverse multilingual tasks in legal NLP
PromptCBLUE: a large-scale instruction-tuning dataset for multi-task and few-shot learning in the medical domain in Chinese
Xiezhi (獬豸) is a comprehensive evaluation suite for Language Models (LMs). It consists of 249587 multi-choice questions spanning 516 diverse disciplines and four difficulty levels, as shown below. Please check our paper for more details, and our website will be open later on.
The devkit of the nuScenes dataset.
Benchmarking Legal Knowledge of Large Language Models
Abstract: FinanceBench is a first-of-its-kind test suite for evaluating the performance of LLMs on open book financial question answering (QA). This repository contains an open source sample of 150 annotated examples used in the evaluation and analysis of models assessed in the FinanceBench…
This repository contains information about AGIEval, data, code and output of baseline systems for the benchmark.
When FLUE Meets FLANG: Benchmarks and Large Pretrained Language Model for Financial Domain
Welcome to the repository of the PandaSet Devkit.