awesome-repositories.com
ब्लॉग
MCP
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टMCP सर्वरहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेस
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to sufe-aiflm-lab/fineval

Open-source alternatives to FinEval

20 open-source projects similar to sufe-aiflm-lab/fineval, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best FinEval alternative.

  • cbluebenchmark/cblueCBLUEbenchmark का अवतार

    CBLUEbenchmark/CBLUE

    843GitHub पर देखें↗

    CBLUE1 中文医疗信息处理基准CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark

    Python
    GitHub पर देखें↗843
  • chancefocus/pixiuchancefocus का अवतार

    chancefocus/PIXIU

    868GitHub पर देखें↗

    This repository introduces PIXIU, an open-source resource featuring the first financial large language models (LLMs), instruction tuning data, and evaluation benchmarks to holistically assess financial LLMs. Our goal is to continually push forward the open-source development of financial artificial intelligence (AI).

    Jupyter Notebook
    GitHub पर देखें↗868
  • coastalcph/lex-gluecoastalcph का अवतार

    coastalcph/lex-glue

    259GitHub पर देखें↗

    LexGLUE: A Benchmark Dataset for Legal Language Understanding in English

    Python
    GitHub पर देखें↗259
  • codefuse-ai/codefuse-devops-evalcodefuse-ai का अवतार

    codefuse-ai/codefuse-devops-eval

    656GitHub पर देखें↗

    Industrial-first evaluation benchmark for LLMs in the DevOps/AIOps domain.

    Python
    GitHub पर देखें↗656
  • dai-shen/laiwDai-shen का अवतार

    Dai-shen/LAiW

    91GitHub पर देखें↗

    LAiW: A Chinese Legal Large Language Models Benchmark

    Python
    GitHub पर देखें↗91

AI सर्च

और अधिक बेहतरीन रिपॉजिटरी खोजें

अपनी ज़रूरत को सरल भाषा में बताएं — AI हजारों क्यूरेटेड ओपन-सोर्स प्रोजेक्ट्स को प्रासंगिकता के आधार पर रैंक करता है।

Find more with AI search
felixgithub2017/cg-evalFelixgithub2017 का अवतार

Felixgithub2017/CG-Eval

13GitHub पर देखें↗

Chinese Generation Evaluation

GitHub पर देखें↗13
  • felixgithub2017/mmcuFelixgithub2017 का अवतार

    Felixgithub2017/MMCU

    90GitHub पर देखें↗

    MEASURING MASSIVE MULTITASK CHINESE UNDERSTANDING

    Python
    GitHub पर देखें↗90
  • haonan-li/cmmluhaonan-li का अवतार

    haonan-li/CMMLU

    821GitHub पर देखें↗

    CMMLU: Measuring massive multitask language understanding in Chinese

    Python
    GitHub पर देखें↗821
  • hazyresearch/legalbenchHazyResearch का अवतार

    HazyResearch/legalbench

    597GitHub पर देखें↗

    An open science effort to benchmark legal reasoning in foundation models

    Python
    GitHub पर देखें↗597
  • hc-guo/owlHC-Guo का अवतार

    HC-Guo/Owl

    237GitHub पर देखें↗

    A Large Language Model for IT Operations

    Python
    GitHub पर देखें↗237
  • joelniklaus/lextremeJoelNiklaus का अवतार

    JoelNiklaus/LEXTREME

    25GitHub पर देखें↗

    This repository provides scripts for evaluating NLP models on the LEXTREME benchmark, a set of diverse multilingual tasks in legal NLP

    Python
    GitHub पर देखें↗25
  • michael-wzhu/promptcbluemichael-wzhu का अवतार

    michael-wzhu/PromptCBLUE

    394GitHub पर देखें↗

    PromptCBLUE: a large-scale instruction-tuning dataset for multi-task and few-shot learning in the medical domain in Chinese

    Python
    GitHub पर देखें↗394
  • mikegu721/xiezhibenchmarkMikeGu721 का अवतार

    MikeGu721/XiezhiBenchmark

    98GitHub पर देखें↗

    Xiezhi (獬豸) is a comprehensive evaluation suite for Language Models (LMs). It consists of 249587 multi-choice questions spanning 516 diverse disciplines and four difficulty levels, as shown below. Please check our paper for more details, and our website will be open later on.

    Python
    GitHub पर देखें↗98
  • open-compass/lawbenchopen-compass का अवतार

    open-compass/LawBench

    434GitHub पर देखें↗

    Benchmarking Legal Knowledge of Large Language Models

    Python
    GitHub पर देखें↗434
  • patronus-ai/financebenchpatronus-ai का अवतार

    patronus-ai/financebench

    328GitHub पर देखें↗

    Abstract: FinanceBench is a first-of-its-kind test suite for evaluating the performance of LLMs on open book financial question answering (QA). This repository contains an open source sample of 150 annotated examples used in the evaluation and analysis of models assessed in the FinanceBench…

    Jupyter Notebook
    GitHub पर देखें↗328
  • ruixiangcui/agievalruixiangcui का अवतार

    ruixiangcui/AGIEval

    775GitHub पर देखें↗

    This repository contains information about AGIEval, data, code and output of baseline systems for the benchmark.

    Python
    GitHub पर देखें↗775
  • salt-nlp/flangSALT-NLP का अवतार

    SALT-NLP/FLANG

    57GitHub पर देखें↗

    When FLUE Meets FLANG: Benchmarks and Large Pretrained Language Model for Financial Domain

    Python
    GitHub पर देखें↗57
  • sjtu-lit/cevalSJTU-LIT का अवतार

    SJTU-LIT/ceval

    1,854GitHub पर देखें↗

    Official github repo for C-Eval, a Chinese evaluation suite for foundation models NeurIPS 2023

    Python
    GitHub पर देखें↗1,854
  • ssymmetry/bbt-fincuge-applicationsssymmetry का अवतार

    ssymmetry/BBT-FinCUGE-Applications

    284GitHub पर देखें↗

    论文链接:https://arxiv.org/abs/2302.09432

    Python
    GitHub पर देखें↗284
  • tongjifinlab/cfbenchmarkTongjiFinLab का अवतार

    TongjiFinLab/CFBenchmark

    55GitHub पर देखें↗

    Chinese Financial Assistant Benchmark for Large Language Model

    Python
    GitHub पर देखें↗55