awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعخادم MCPحولكيفية ترتيب النتائجالصحافة
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to sufe-aiflm-lab/fineval

Open-source alternatives to FinEval

20 open-source projects similar to sufe-aiflm-lab/fineval, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best FinEval alternative.

  • cbluebenchmark/cblueالصورة الرمزية لـ CBLUEbenchmark

    CBLUEbenchmark/CBLUE

    843عرض على GitHub↗

    CBLUE1 中文医疗信息处理基准CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark

    Python
    عرض على GitHub↗843
  • chancefocus/pixiuالصورة الرمزية لـ chancefocus

    chancefocus/PIXIU

    868عرض على GitHub↗

    This repository introduces PIXIU, an open-source resource featuring the first financial large language models (LLMs), instruction tuning data, and evaluation benchmarks to holistically assess financial LLMs. Our goal is to continually push forward the open-source development of financial artificial intelligence (AI).

    Jupyter Notebook
    عرض على GitHub↗868
  • coastalcph/lex-glueالصورة الرمزية لـ coastalcph

    coastalcph/lex-glue

    259عرض على GitHub↗

    LexGLUE: A Benchmark Dataset for Legal Language Understanding in English

    Python
    عرض على GitHub↗259
  • codefuse-ai/codefuse-devops-evalالصورة الرمزية لـ codefuse-ai

    codefuse-ai/codefuse-devops-eval

    656عرض على GitHub↗

    Industrial-first evaluation benchmark for LLMs in the DevOps/AIOps domain.

    Python
    عرض على GitHub↗656
  • dai-shen/laiwالصورة الرمزية لـ Dai-shen

    Dai-shen/LAiW

    91عرض على GitHub↗

    LAiW: A Chinese Legal Large Language Models Benchmark

    Python
    عرض على GitHub↗91

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Find more with AI search
felixgithub2017/cg-evalالصورة الرمزية لـ Felixgithub2017

Felixgithub2017/CG-Eval

13عرض على GitHub↗

Chinese Generation Evaluation

عرض على GitHub↗13
  • felixgithub2017/mmcuالصورة الرمزية لـ Felixgithub2017

    Felixgithub2017/MMCU

    90عرض على GitHub↗

    MEASURING MASSIVE MULTITASK CHINESE UNDERSTANDING

    Python
    عرض على GitHub↗90
  • haonan-li/cmmluالصورة الرمزية لـ haonan-li

    haonan-li/CMMLU

    821عرض على GitHub↗

    CMMLU: Measuring massive multitask language understanding in Chinese

    Python
    عرض على GitHub↗821
  • hazyresearch/legalbenchالصورة الرمزية لـ HazyResearch

    HazyResearch/legalbench

    597عرض على GitHub↗

    An open science effort to benchmark legal reasoning in foundation models

    Python
    عرض على GitHub↗597
  • hc-guo/owlالصورة الرمزية لـ HC-Guo

    HC-Guo/Owl

    237عرض على GitHub↗

    A Large Language Model for IT Operations

    Python
    عرض على GitHub↗237
  • joelniklaus/lextremeالصورة الرمزية لـ JoelNiklaus

    JoelNiklaus/LEXTREME

    25عرض على GitHub↗

    This repository provides scripts for evaluating NLP models on the LEXTREME benchmark, a set of diverse multilingual tasks in legal NLP

    Python
    عرض على GitHub↗25
  • michael-wzhu/promptcblueالصورة الرمزية لـ michael-wzhu

    michael-wzhu/PromptCBLUE

    394عرض على GitHub↗

    PromptCBLUE: a large-scale instruction-tuning dataset for multi-task and few-shot learning in the medical domain in Chinese

    Python
    عرض على GitHub↗394
  • mikegu721/xiezhibenchmarkالصورة الرمزية لـ MikeGu721

    MikeGu721/XiezhiBenchmark

    98عرض على GitHub↗

    Xiezhi (獬豸) is a comprehensive evaluation suite for Language Models (LMs). It consists of 249587 multi-choice questions spanning 516 diverse disciplines and four difficulty levels, as shown below. Please check our paper for more details, and our website will be open later on.

    Python
    عرض على GitHub↗98
  • open-compass/lawbenchالصورة الرمزية لـ open-compass

    open-compass/LawBench

    434عرض على GitHub↗

    Benchmarking Legal Knowledge of Large Language Models

    Python
    عرض على GitHub↗434
  • patronus-ai/financebenchالصورة الرمزية لـ patronus-ai

    patronus-ai/financebench

    328عرض على GitHub↗

    Abstract: FinanceBench is a first-of-its-kind test suite for evaluating the performance of LLMs on open book financial question answering (QA). This repository contains an open source sample of 150 annotated examples used in the evaluation and analysis of models assessed in the FinanceBench…

    Jupyter Notebook
    عرض على GitHub↗328
  • ruixiangcui/agievalالصورة الرمزية لـ ruixiangcui

    ruixiangcui/AGIEval

    775عرض على GitHub↗

    This repository contains information about AGIEval, data, code and output of baseline systems for the benchmark.

    Python
    عرض على GitHub↗775
  • salt-nlp/flangالصورة الرمزية لـ SALT-NLP

    SALT-NLP/FLANG

    57عرض على GitHub↗

    When FLUE Meets FLANG: Benchmarks and Large Pretrained Language Model for Financial Domain

    Python
    عرض على GitHub↗57
  • sjtu-lit/cevalالصورة الرمزية لـ SJTU-LIT

    SJTU-LIT/ceval

    1,854عرض على GitHub↗

    Official github repo for C-Eval, a Chinese evaluation suite for foundation models NeurIPS 2023

    Python
    عرض على GitHub↗1,854
  • ssymmetry/bbt-fincuge-applicationsالصورة الرمزية لـ ssymmetry

    ssymmetry/BBT-FinCUGE-Applications

    284عرض على GitHub↗

    论文链接:https://arxiv.org/abs/2302.09432

    Python
    عرض على GitHub↗284
  • tongjifinlab/cfbenchmarkالصورة الرمزية لـ TongjiFinLab

    TongjiFinLab/CFBenchmark

    55عرض على GitHub↗

    Chinese Financial Assistant Benchmark for Large Language Model

    Python
    عرض على GitHub↗55