awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Back to cbluebenchmark/cblue

Open-source alternatives to CBLUE

30 open-source projects similar to cbluebenchmark/cblue, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best CBLUE alternative.

  • aaronheee/llms-as-zero-shot-conversational-recsysaaronheee avatar

    aaronheee/llms-as-zero-shot-conversational-recsys

    85View on GitHub↗

    This is the evaluation data and Large Language Models (LLMs) results from our CIKM'23 paper:

    Python
    View on GitHub↗85
  • agiresearch/openp5agiresearch avatar

    agiresearch/OpenP5

    351View on GitHub↗

    This repo presents OpenP5, an open-source platform for LLM-based Recommendation development, finetuning, and evaluation.

    Python
    View on GitHub↗351
  • apolloscapeauto/dataset-apiApolloScapeAuto avatar

    ApolloScapeAuto/dataset-api

    617View on GitHub↗

    The ApolloScape Open Dataset for Autonomous Driving and its Application.

    Jupyter Notebook
    View on GitHub↗617
  • argoai/argoverse-apiargoai avatar

    argoai/argoverse-api

    933View on GitHub↗

    Official GitHub repository for Argoverse dataset

    Python
    View on GitHub↗933
  • autonomousvision/kitti360scriptsautonomousvision avatar

    autonomousvision/kitti360Scripts

    444View on GitHub↗

    This repository contains utility scripts for the KITTI-360 dataset.

    Python
    View on GitHub↗444

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Find more with AI search
chancefocus/pixiuchancefocus avatar

chancefocus/PIXIU

868View on GitHub↗

This repository introduces PIXIU, an open-source resource featuring the first financial large language models (LLMs), instruction tuning data, and evaluation benchmarks to holistically assess financial LLMs. Our goal is to continually push forward the open-source development of financial artificial intelligence (AI).

Jupyter Notebook
View on GitHub↗868
  • chrieke/awesome-satellite-imagery-datasetschrieke avatar

    chrieke/awesome-satellite-imagery-datasets

    3,898View on GitHub↗

    🛰️ List of satellite image training datasets with annotations for computer vision and deep learning

    View on GitHub↗3,898
  • coastalcph/lex-gluecoastalcph avatar

    coastalcph/lex-glue

    259View on GitHub↗

    LexGLUE: A Benchmark Dataset for Legal Language Understanding in English

    Python
    View on GitHub↗259
  • codefuse-ai/codefuse-devops-evalcodefuse-ai avatar

    codefuse-ai/codefuse-devops-eval

    656View on GitHub↗

    Industrial-first evaluation benchmark for LLMs in the DevOps/AIOps domain.

    Python
    View on GitHub↗656
  • dai-shen/laiwDai-shen avatar

    Dai-shen/LAiW

    91View on GitHub↗

    LAiW: A Chinese Legal Large Language Models Benchmark

    Python
    View on GitHub↗91
  • dlr-rm/blenderprocDLR-RM avatar

    DLR-RM/BlenderProc

    3,595View on GitHub↗

    A procedural Blender pipeline for photorealistic training image generation

    Python
    View on GitHub↗3,595
  • driving-behavior/dbnetdriving-behavior avatar

    driving-behavior/DBNet

    221View on GitHub↗

    DBNet: A Large-Scale Dataset for Driving Behavior Learning, CVPR 2018

    Python
    View on GitHub↗221
  • felixgithub2017/cg-evalFelixgithub2017 avatar

    Felixgithub2017/CG-Eval

    13View on GitHub↗

    Chinese Generation Evaluation

    View on GitHub↗13
  • felixgithub2017/mmcuFelixgithub2017 avatar

    Felixgithub2017/MMCU

    90View on GitHub↗

    MEASURING MASSIVE MULTITASK CHINESE UNDERSTANDING

    Python
    View on GitHub↗90
  • google-research-datasets/objectrongoogle-research-datasets avatar

    google-research-datasets/Objectron

    2,336View on GitHub↗

    Objectron is a dataset of short, object-centric video clips. In addition, the videos also contain AR session metadata including camera poses, sparse point-clouds and planes. In each video, the camera moves around and above the object and captures it from different views. Each object is annotated with a 3D bounding box. The 3D bounding box describes the object’s position, orientation, and dimensions. The dataset contains about 15K annotated video clips and 4M annotated images in the following categories: bikes, books, bottles, cameras, cereal boxes, chairs, cups, laptops, and shoes

    Jupyter Notebook3d3d-reconstruction3d-vision
    View on GitHub↗2,336
  • haonan-li/cmmluhaonan-li avatar

    haonan-li/CMMLU

    821View on GitHub↗

    CMMLU: Measuring massive multitask language understanding in Chinese

    Python
    View on GitHub↗821
  • hazyresearch/legalbenchHazyResearch avatar

    HazyResearch/legalbench

    597View on GitHub↗

    An open science effort to benchmark legal reasoning in foundation models

    Python
    View on GitHub↗597
  • hc-guo/owlHC-Guo avatar

    HC-Guo/Owl

    237View on GitHub↗

    A Large Language Model for IT Operations

    Python
    View on GitHub↗237
  • hkuds/llmrecHKUDS avatar

    HKUDS/LLMRec

    532View on GitHub↗

    PyTorch implementation for WSDM 2024 paper LLMRec: Large Language Models with Graph Augmentation for Recommendation.

    Python
    View on GitHub↗532
  • hkust-vgd/scanobjectnnH

    hkust-vgd/scanobjectnn

    0View on GitHub↗
    View on GitHub↗0
  • i2rdl2/astar-3dI

    I2RDL2/ASTAR-3D

    0View on GitHub↗
    View on GitHub↗0
  • joelniklaus/lextremeJoelNiklaus avatar

    JoelNiklaus/LEXTREME

    25View on GitHub↗

    This repository provides scripts for evaluating NLP models on the LEXTREME benchmark, a set of diverse multilingual tasks in legal NLP

    Python
    View on GitHub↗25
  • michael-wzhu/promptcbluemichael-wzhu avatar

    michael-wzhu/PromptCBLUE

    394View on GitHub↗

    PromptCBLUE: a large-scale instruction-tuning dataset for multi-task and few-shot learning in the medical domain in Chinese

    Python
    View on GitHub↗394
  • mikegu721/xiezhibenchmarkMikeGu721 avatar

    MikeGu721/XiezhiBenchmark

    98View on GitHub↗

    Xiezhi (獬豸) is a comprehensive evaluation suite for Language Models (LMs). It consists of 249587 multi-choice questions spanning 516 diverse disciplines and four difficulty levels, as shown below. Please check our paper for more details, and our website will be open later on.

    Python
    View on GitHub↗98
  • nutonomy/nuscenes-devkitnutonomy avatar

    nutonomy/nuscenes-devkit

    2,761View on GitHub↗

    The devkit of the nuScenes dataset.

    Python
    View on GitHub↗2,761
  • open-compass/lawbenchopen-compass avatar

    open-compass/LawBench

    434View on GitHub↗

    Benchmarking Legal Knowledge of Large Language Models

    Python
    View on GitHub↗434
  • patronus-ai/financebenchpatronus-ai avatar

    patronus-ai/financebench

    328View on GitHub↗

    Abstract: FinanceBench is a first-of-its-kind test suite for evaluating the performance of LLMs on open book financial question answering (QA). This repository contains an open source sample of 150 annotated examples used in the evaluation and analysis of models assessed in the FinanceBench…

    Jupyter Notebook
    View on GitHub↗328
  • ruixiangcui/agievalruixiangcui avatar

    ruixiangcui/AGIEval

    775View on GitHub↗

    This repository contains information about AGIEval, data, code and output of baseline systems for the benchmark.

    Python
    View on GitHub↗775
  • salt-nlp/flangSALT-NLP avatar

    SALT-NLP/FLANG

    57View on GitHub↗

    When FLUE Meets FLANG: Benchmarks and Large Pretrained Language Model for Financial Domain

    Python
    View on GitHub↗57
  • scaleapi/pandaset-devkitscaleapi avatar

    scaleapi/pandaset-devkit

    276View on GitHub↗

    Welcome to the repository of the PandaSet Devkit.

    Jupyter Notebook
    View on GitHub↗276