awesome-repositories.com
Blog
MCP
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektMCP-ServerÜber unsRanking-MethodikPresse
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to sufe-aiflm-lab/fineval

Open-source alternatives to FinEval

20 open-source projects similar to sufe-aiflm-lab/fineval, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best FinEval alternative.

  • cbluebenchmark/cblueAvatar von CBLUEbenchmark

    CBLUEbenchmark/CBLUE

    843Auf GitHub ansehen↗

    CBLUE1 中文医疗信息处理基准CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark

    Python
    Auf GitHub ansehen↗843
  • chancefocus/pixiuAvatar von chancefocus

    chancefocus/PIXIU

    868Auf GitHub ansehen↗

    This repository introduces PIXIU, an open-source resource featuring the first financial large language models (LLMs), instruction tuning data, and evaluation benchmarks to holistically assess financial LLMs. Our goal is to continually push forward the open-source development of financial artificial intelligence (AI).

    Jupyter Notebook
    Auf GitHub ansehen↗868
  • coastalcph/lex-glueAvatar von coastalcph

    coastalcph/lex-glue

    259Auf GitHub ansehen↗

    LexGLUE: A Benchmark Dataset for Legal Language Understanding in English

    Python
    Auf GitHub ansehen↗259
  • codefuse-ai/codefuse-devops-evalAvatar von codefuse-ai

    codefuse-ai/codefuse-devops-eval

    656Auf GitHub ansehen↗

    Industrial-first evaluation benchmark for LLMs in the DevOps/AIOps domain.

    Python
    Auf GitHub ansehen↗656
  • dai-shen/laiwAvatar von Dai-shen

    Dai-shen/LAiW

    91Auf GitHub ansehen↗

    LAiW: A Chinese Legal Large Language Models Benchmark

    Python
    Auf GitHub ansehen↗91

KI-Suche

Entdecke weitere awesome Repositories

Beschreibe in einfachen Worten, was du brauchst — die KI bewertet tausende kuratierte Open-Source-Projekte nach Relevanz.

Find more with AI search
felixgithub2017/cg-evalAvatar von Felixgithub2017

Felixgithub2017/CG-Eval

13Auf GitHub ansehen↗

Chinese Generation Evaluation

Auf GitHub ansehen↗13
  • felixgithub2017/mmcuAvatar von Felixgithub2017

    Felixgithub2017/MMCU

    90Auf GitHub ansehen↗

    MEASURING MASSIVE MULTITASK CHINESE UNDERSTANDING

    Python
    Auf GitHub ansehen↗90
  • haonan-li/cmmluAvatar von haonan-li

    haonan-li/CMMLU

    821Auf GitHub ansehen↗

    CMMLU: Measuring massive multitask language understanding in Chinese

    Python
    Auf GitHub ansehen↗821
  • hazyresearch/legalbenchAvatar von HazyResearch

    HazyResearch/legalbench

    597Auf GitHub ansehen↗

    An open science effort to benchmark legal reasoning in foundation models

    Python
    Auf GitHub ansehen↗597
  • hc-guo/owlAvatar von HC-Guo

    HC-Guo/Owl

    237Auf GitHub ansehen↗

    A Large Language Model for IT Operations

    Python
    Auf GitHub ansehen↗237
  • joelniklaus/lextremeAvatar von JoelNiklaus

    JoelNiklaus/LEXTREME

    25Auf GitHub ansehen↗

    This repository provides scripts for evaluating NLP models on the LEXTREME benchmark, a set of diverse multilingual tasks in legal NLP

    Python
    Auf GitHub ansehen↗25
  • michael-wzhu/promptcblueAvatar von michael-wzhu

    michael-wzhu/PromptCBLUE

    394Auf GitHub ansehen↗

    PromptCBLUE: a large-scale instruction-tuning dataset for multi-task and few-shot learning in the medical domain in Chinese

    Python
    Auf GitHub ansehen↗394
  • mikegu721/xiezhibenchmarkAvatar von MikeGu721

    MikeGu721/XiezhiBenchmark

    98Auf GitHub ansehen↗

    Xiezhi (獬豸) is a comprehensive evaluation suite for Language Models (LMs). It consists of 249587 multi-choice questions spanning 516 diverse disciplines and four difficulty levels, as shown below. Please check our paper for more details, and our website will be open later on.

    Python
    Auf GitHub ansehen↗98
  • open-compass/lawbenchAvatar von open-compass

    open-compass/LawBench

    434Auf GitHub ansehen↗

    Benchmarking Legal Knowledge of Large Language Models

    Python
    Auf GitHub ansehen↗434
  • patronus-ai/financebenchAvatar von patronus-ai

    patronus-ai/financebench

    328Auf GitHub ansehen↗

    Abstract: FinanceBench is a first-of-its-kind test suite for evaluating the performance of LLMs on open book financial question answering (QA). This repository contains an open source sample of 150 annotated examples used in the evaluation and analysis of models assessed in the FinanceBench…

    Jupyter Notebook
    Auf GitHub ansehen↗328
  • ruixiangcui/agievalAvatar von ruixiangcui

    ruixiangcui/AGIEval

    775Auf GitHub ansehen↗

    This repository contains information about AGIEval, data, code and output of baseline systems for the benchmark.

    Python
    Auf GitHub ansehen↗775
  • salt-nlp/flangAvatar von SALT-NLP

    SALT-NLP/FLANG

    57Auf GitHub ansehen↗

    When FLUE Meets FLANG: Benchmarks and Large Pretrained Language Model for Financial Domain

    Python
    Auf GitHub ansehen↗57
  • sjtu-lit/cevalAvatar von SJTU-LIT

    SJTU-LIT/ceval

    1,854Auf GitHub ansehen↗

    Official github repo for C-Eval, a Chinese evaluation suite for foundation models NeurIPS 2023

    Python
    Auf GitHub ansehen↗1,854
  • ssymmetry/bbt-fincuge-applicationsAvatar von ssymmetry

    ssymmetry/BBT-FinCUGE-Applications

    284Auf GitHub ansehen↗

    论文链接:https://arxiv.org/abs/2302.09432

    Python
    Auf GitHub ansehen↗284
  • tongjifinlab/cfbenchmarkAvatar von TongjiFinLab

    TongjiFinLab/CFBenchmark

    55Auf GitHub ansehen↗

    Chinese Financial Assistant Benchmark for Large Language Model

    Python
    Auf GitHub ansehen↗55