awesome-repositories.com
Blog
MCP
awesome-repositories.com

Découvrez les meilleurs dépôts open-source grâce à notre recherche par IA.

ExplorerRecherches sélectionnéesAlternatives open sourceLogiciels auto-hébergésBlogPlan du site
ProjetServeur MCPÀ proposNotre méthodologiePresse
Mentions légalesConfidentialitéConditions d'utilisation
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to felixgithub2017/mmcu

Open-source alternatives to MMCU

20 open-source projects similar to felixgithub2017/mmcu, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best MMCU alternative.

  • cbluebenchmark/cblueAvatar de CBLUEbenchmark

    CBLUEbenchmark/CBLUE

    843Voir sur GitHub↗

    CBLUE1 中文医疗信息处理基准CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark

    Python
    Voir sur GitHub↗843
  • chancefocus/pixiuAvatar de chancefocus

    chancefocus/PIXIU

    868Voir sur GitHub↗

    This repository introduces PIXIU, an open-source resource featuring the first financial large language models (LLMs), instruction tuning data, and evaluation benchmarks to holistically assess financial LLMs. Our goal is to continually push forward the open-source development of financial artificial intelligence (AI).

    Jupyter Notebook
    Voir sur GitHub↗868
  • coastalcph/lex-glueAvatar de coastalcph

    coastalcph/lex-glue

    259Voir sur GitHub↗

    LexGLUE: A Benchmark Dataset for Legal Language Understanding in English

    Python
    Voir sur GitHub↗259
  • codefuse-ai/codefuse-devops-evalAvatar de codefuse-ai

    codefuse-ai/codefuse-devops-eval

    656Voir sur GitHub↗

    Industrial-first evaluation benchmark for LLMs in the DevOps/AIOps domain.

    Python
    Voir sur GitHub↗656
  • dai-shen/laiwAvatar de Dai-shen

    Dai-shen/LAiW

    91Voir sur GitHub↗

    LAiW: A Chinese Legal Large Language Models Benchmark

    Python
    Voir sur GitHub↗91

Recherche par IA

Explorez plus de dépôts awesome

Décrivez vos besoins en langage naturel — l'IA classe des milliers de projets open source sélectionnés par pertinence.

Find more with AI search
felixgithub2017/cg-evalAvatar de Felixgithub2017

Felixgithub2017/CG-Eval

13Voir sur GitHub↗

Chinese Generation Evaluation

Voir sur GitHub↗13
  • haonan-li/cmmluAvatar de haonan-li

    haonan-li/CMMLU

    821Voir sur GitHub↗

    CMMLU: Measuring massive multitask language understanding in Chinese

    Python
    Voir sur GitHub↗821
  • hazyresearch/legalbenchAvatar de HazyResearch

    HazyResearch/legalbench

    597Voir sur GitHub↗

    An open science effort to benchmark legal reasoning in foundation models

    Python
    Voir sur GitHub↗597
  • hc-guo/owlAvatar de HC-Guo

    HC-Guo/Owl

    237Voir sur GitHub↗

    A Large Language Model for IT Operations

    Python
    Voir sur GitHub↗237
  • joelniklaus/lextremeAvatar de JoelNiklaus

    JoelNiklaus/LEXTREME

    25Voir sur GitHub↗

    This repository provides scripts for evaluating NLP models on the LEXTREME benchmark, a set of diverse multilingual tasks in legal NLP

    Python
    Voir sur GitHub↗25
  • michael-wzhu/promptcblueAvatar de michael-wzhu

    michael-wzhu/PromptCBLUE

    394Voir sur GitHub↗

    PromptCBLUE: a large-scale instruction-tuning dataset for multi-task and few-shot learning in the medical domain in Chinese

    Python
    Voir sur GitHub↗394
  • mikegu721/xiezhibenchmarkAvatar de MikeGu721

    MikeGu721/XiezhiBenchmark

    98Voir sur GitHub↗

    Xiezhi (獬豸) is a comprehensive evaluation suite for Language Models (LMs). It consists of 249587 multi-choice questions spanning 516 diverse disciplines and four difficulty levels, as shown below. Please check our paper for more details, and our website will be open later on.

    Python
    Voir sur GitHub↗98
  • open-compass/lawbenchAvatar de open-compass

    open-compass/LawBench

    434Voir sur GitHub↗

    Benchmarking Legal Knowledge of Large Language Models

    Python
    Voir sur GitHub↗434
  • patronus-ai/financebenchAvatar de patronus-ai

    patronus-ai/financebench

    328Voir sur GitHub↗

    Abstract: FinanceBench is a first-of-its-kind test suite for evaluating the performance of LLMs on open book financial question answering (QA). This repository contains an open source sample of 150 annotated examples used in the evaluation and analysis of models assessed in the FinanceBench…

    Jupyter Notebook
    Voir sur GitHub↗328
  • ruixiangcui/agievalAvatar de ruixiangcui

    ruixiangcui/AGIEval

    775Voir sur GitHub↗

    This repository contains information about AGIEval, data, code and output of baseline systems for the benchmark.

    Python
    Voir sur GitHub↗775
  • salt-nlp/flangAvatar de SALT-NLP

    SALT-NLP/FLANG

    57Voir sur GitHub↗

    When FLUE Meets FLANG: Benchmarks and Large Pretrained Language Model for Financial Domain

    Python
    Voir sur GitHub↗57
  • sjtu-lit/cevalAvatar de SJTU-LIT

    SJTU-LIT/ceval

    1,854Voir sur GitHub↗

    Official github repo for C-Eval, a Chinese evaluation suite for foundation models NeurIPS 2023

    Python
    Voir sur GitHub↗1,854
  • ssymmetry/bbt-fincuge-applicationsAvatar de ssymmetry

    ssymmetry/BBT-FinCUGE-Applications

    284Voir sur GitHub↗

    论文链接:https://arxiv.org/abs/2302.09432

    Python
    Voir sur GitHub↗284
  • sufe-aiflm-lab/finevalAvatar de SUFE-AIFLM-Lab

    SUFE-AIFLM-Lab/FinEval

    275Voir sur GitHub↗

    The FinEval financial domain evaluation benchmark, based on quantitative fundamental methods and developed through long-term objective research, summarization, and rigorous manual screening, utilizes over 26,000 diverse question types that are highly consistent with real-world application scenarios.

    Python
    Voir sur GitHub↗275
  • tongjifinlab/cfbenchmarkAvatar de TongjiFinLab

    TongjiFinLab/CFBenchmark

    55Voir sur GitHub↗

    Chinese Financial Assistant Benchmark for Large Language Model

    Python
    Voir sur GitHub↗55