25 open-source projects similar to codefuse-ai/codefuse-devops-eval, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best Codefuse Devops Eval alternative.
CBLUE1 中文医疗信息处理基准CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark
This repository introduces PIXIU, an open-source resource featuring the first financial large language models (LLMs), instruction tuning data, and evaluation benchmarks to holistically assess financial LLMs. Our goal is to continually push forward the open-source development of financial artificial intelligence (AI).
LexGLUE: A Benchmark Dataset for Legal Language Understanding in English
An intelligent assistant serving the entire software development lifecycle, powered by a Multi-Agent Framework, working with DevOps Toolkits, Code&Doc Repo RAG, etc.
CMMLU: Measuring massive multitask language understanding in Chinese
An open science effort to benchmark legal reasoning in foundation models
This repository provides scripts for evaluating NLP models on the LEXTREME benchmark, a set of diverse multilingual tasks in legal NLP
Mr. Ranedeer AI Tutor is an AI education framework and system prompt designed to transform a large language model into a personalized tutor. It uses a structured set of instructions to organize educational content into sequential modules and knowledge assessments for adaptive learning. The system features a persona template that allows for the adjustment of academic depth and communication tone to match a student's specific needs. It also provides multilingual support, enabling the tutor to switch instruction and output languages based on user preferences. The framework covers custom lesson
PromptCBLUE: a large-scale instruction-tuning dataset for multi-task and few-shot learning in the medical domain in Chinese
Xiezhi (獬豸) is a comprehensive evaluation suite for Language Models (LMs). It consists of 249587 multi-choice questions spanning 516 diverse disciplines and four difficulty levels, as shown below. Please check our paper for more details, and our website will be open later on.
Benchmarking Legal Knowledge of Large Language Models
Abstract: FinanceBench is a first-of-its-kind test suite for evaluating the performance of LLMs on open book financial question answering (QA). This repository contains an open source sample of 150 annotated examples used in the evaluation and analysis of models assessed in the FinanceBench…
This repository contains information about AGIEval, data, code and output of baseline systems for the benchmark.
When FLUE Meets FLANG: Benchmarks and Large Pretrained Language Model for Financial Domain
Official github repo for C-Eval, a Chinese evaluation suite for foundation models NeurIPS 2023
论文链接:https://arxiv.org/abs/2302.09432
The FinEval financial domain evaluation benchmark, based on quantitative fundamental methods and developed through long-term objective research, summarization, and rigorous manual screening, utilizes over 26,000 diverse question types that are highly consistent with real-world application scenarios.
Chinese Financial Assistant Benchmark for Large Language Model