How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.
Industrial-first evaluation benchmark for LLMs in the DevOps/AIOps domain.
The main features of codefuse-ai/codefuse-devops-eval are: AI Development Tools, Evaluation Benchmarks.
Projects with overlapping indexed features include: chancefocus/pixiu — This repository introduces PIXIU, an open-source resource featuring the first financial large language models (LLMs),… coastalcph/lex-glue — LexGLUE: A Benchmark Dataset for Legal Language Understanding in English. codefuse-ai/codefuse-chatbot — An intelligent assistant serving the entire software development lifecycle, powered by a Multi-Agent Framework,… dai-shen/laiw — LAiW: A Chinese Legal Large Language Models Benchmark. felixgithub2017/cg-eval — Chinese Generation Evaluation. cbluebenchmark/cblue — [CBLUE1] 中文医疗信息处理基准CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark.
This repository introduces PIXIU, an open-source resource featuring the first financial large language models (LLMs), instruction tuning data, and evaluation benchmarks to holistically assess financial LLMs. Our goal is to continually push forward the open-source development of financial artificial intelligence (AI).
LexGLUE: A Benchmark Dataset for Legal Language Understanding in English
An intelligent assistant serving the entire software development lifecycle, powered by a Multi-Agent Framework, working with DevOps Toolkits, Code&Doc Repo RAG, etc.
CBLUE1 中文医疗信息处理基准CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark