This repository introduces PIXIU, an open-source resource featuring the first financial large language models (LLMs), instruction tuning data, and evaluation benchmarks to holistically assess financial LLMs. Our goal is to continually push forward the open-source development of financial artificial intelligence (AI).
LexGLUE: A Benchmark Dataset for Legal Language Understanding in English
Industrial-first evaluation benchmark for LLMs in the DevOps/AIOps domain.
CBLUE1 中文医疗信息处理基准CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark
Abstract: FinanceBench is a first-of-its-kind test suite for evaluating the performance of LLMs on open book financial question answering (QA). This repository contains an open source sample of 150 annotated examples used in the evaluation and analysis of models assessed in the FinanceBench…
Principalele funcționalități ale patronus-ai/financebench sunt: Evaluation Benchmarks.
Alternativele open-source pentru patronus-ai/financebench includ: chancefocus/pixiu — This repository introduces PIXIU, an open-source resource featuring the first financial large language models (LLMs),… coastalcph/lex-glue — LexGLUE: A Benchmark Dataset for Legal Language Understanding in English. codefuse-ai/codefuse-devops-eval — Industrial-first evaluation benchmark for LLMs in the DevOps/AIOps domain. dai-shen/laiw — LAiW: A Chinese Legal Large Language Models Benchmark. felixgithub2017/cg-eval — Chinese Generation Evaluation. cbluebenchmark/cblue — [CBLUE1] 中文医疗信息处理基准CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark.