awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoServidor MCPAcerca deCómo clasificamosPrensa
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to codefuse-ai/codefuse-devops-eval

Open-source alternatives to Codefuse Devops Eval

25 open-source projects similar to codefuse-ai/codefuse-devops-eval, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best Codefuse Devops Eval alternative.

  • cbluebenchmark/cblueAvatar de CBLUEbenchmark

    CBLUEbenchmark/CBLUE

    843Ver en GitHub↗

    CBLUE1 中文医疗信息处理基准CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark

    Python
    Ver en GitHub↗843
  • chancefocus/pixiuAvatar de chancefocus

    chancefocus/PIXIU

    868Ver en GitHub↗

    This repository introduces PIXIU, an open-source resource featuring the first financial large language models (LLMs), instruction tuning data, and evaluation benchmarks to holistically assess financial LLMs. Our goal is to continually push forward the open-source development of financial artificial intelligence (AI).

    Jupyter Notebook
    Ver en GitHub↗868
  • coastalcph/lex-glueAvatar de coastalcph

    coastalcph/lex-glue

    259Ver en GitHub↗

    LexGLUE: A Benchmark Dataset for Legal Language Understanding in English

    Python
    Ver en GitHub↗259
  • codefuse-ai/codefuse-chatbotAvatar de codefuse-ai

    codefuse-ai/codefuse-chatbot

    1,291Ver en GitHub↗

    An intelligent assistant serving the entire software development lifecycle, powered by a Multi-Agent Framework, working with DevOps Toolkits, Code&Doc Repo RAG, etc.

    Pythonaiopschatbotcode-repo-analysis
    Ver en GitHub↗1,291
  • dai-shen/laiwAvatar de Dai-shen

    Dai-shen/LAiW

    91Ver en GitHub↗

    LAiW: A Chinese Legal Large Language Models Benchmark

    Python
    Ver en GitHub↗91

Búsqueda con IA

Explora más repositorios increíbles

Describe lo que necesitas en lenguaje sencillo: la IA clasifica miles de proyectos open-source curados por relevancia.

Find more with AI search
  • felixgithub2017/cg-evalAvatar de Felixgithub2017

    Felixgithub2017/CG-Eval

    13Ver en GitHub↗

    Chinese Generation Evaluation

    Ver en GitHub↗13
  • felixgithub2017/mmcuAvatar de Felixgithub2017

    Felixgithub2017/MMCU

    90Ver en GitHub↗

    MEASURING MASSIVE MULTITASK CHINESE UNDERSTANDING

    Python
    Ver en GitHub↗90
  • haonan-li/cmmluAvatar de haonan-li

    haonan-li/CMMLU

    821Ver en GitHub↗

    CMMLU: Measuring massive multitask language understanding in Chinese

    Python
    Ver en GitHub↗821
  • hazyresearch/legalbenchAvatar de HazyResearch

    HazyResearch/legalbench

    597Ver en GitHub↗

    An open science effort to benchmark legal reasoning in foundation models

    Python
    Ver en GitHub↗597
  • hc-guo/owlAvatar de HC-Guo

    HC-Guo/Owl

    237Ver en GitHub↗

    A Large Language Model for IT Operations

    Python
    Ver en GitHub↗237
  • joelniklaus/lextremeAvatar de JoelNiklaus

    JoelNiklaus/LEXTREME

    25Ver en GitHub↗

    This repository provides scripts for evaluating NLP models on the LEXTREME benchmark, a set of diverse multilingual tasks in legal NLP

    Python
    Ver en GitHub↗25
  • jushbjj/mr.-ranedeer-ai-tutorAvatar de JushBJJ

    JushBJJ/Mr.-Ranedeer-AI-Tutor

    29,599Ver en GitHub↗

    Mr. Ranedeer AI Tutor is an AI education framework and system prompt designed to transform a large language model into a personalized tutor. It uses a structured set of instructions to organize educational content into sequential modules and knowledge assessments for adaptive learning. The system features a persona template that allows for the adjustment of academic depth and communication tone to match a student's specific needs. It also provides multilingual support, enabling the tutor to switch instruction and output languages based on user preferences. The framework covers custom lesson

    aieducationgpt-4
    Ver en GitHub↗29,599
  • michael-wzhu/promptcblueAvatar de michael-wzhu

    michael-wzhu/PromptCBLUE

    394Ver en GitHub↗

    PromptCBLUE: a large-scale instruction-tuning dataset for multi-task and few-shot learning in the medical domain in Chinese

    Python
    Ver en GitHub↗394
  • mikegu721/xiezhibenchmarkAvatar de MikeGu721

    MikeGu721/XiezhiBenchmark

    98Ver en GitHub↗

    Xiezhi (獬豸) is a comprehensive evaluation suite for Language Models (LMs). It consists of 249587 multi-choice questions spanning 516 diverse disciplines and four difficulty levels, as shown below. Please check our paper for more details, and our website will be open later on.

    Python
    Ver en GitHub↗98
  • open-compass/lawbenchAvatar de open-compass

    open-compass/LawBench

    434Ver en GitHub↗

    Benchmarking Legal Knowledge of Large Language Models

    Python
    Ver en GitHub↗434
  • patronus-ai/financebenchAvatar de patronus-ai

    patronus-ai/financebench

    328Ver en GitHub↗

    Abstract: FinanceBench is a first-of-its-kind test suite for evaluating the performance of LLMs on open book financial question answering (QA). This repository contains an open source sample of 150 annotated examples used in the evaluation and analysis of models assessed in the FinanceBench…

    Jupyter Notebook
    Ver en GitHub↗328
  • ruixiangcui/agievalAvatar de ruixiangcui

    ruixiangcui/AGIEval

    775Ver en GitHub↗

    This repository contains information about AGIEval, data, code and output of baseline systems for the benchmark.

    Python
    Ver en GitHub↗775
  • salt-nlp/flangAvatar de SALT-NLP

    SALT-NLP/FLANG

    57Ver en GitHub↗

    When FLUE Meets FLANG: Benchmarks and Large Pretrained Language Model for Financial Domain

    Python
    Ver en GitHub↗57
  • sjtu-lit/cevalAvatar de SJTU-LIT

    SJTU-LIT/ceval

    1,854Ver en GitHub↗

    Official github repo for C-Eval, a Chinese evaluation suite for foundation models NeurIPS 2023

    Python
    Ver en GitHub↗1,854
  • ssymmetry/bbt-fincuge-applicationsAvatar de ssymmetry

    ssymmetry/BBT-FinCUGE-Applications

    284Ver en GitHub↗

    论文链接:https://arxiv.org/abs/2302.09432

    Python
    Ver en GitHub↗284
  • sufe-aiflm-lab/finevalAvatar de SUFE-AIFLM-Lab

    SUFE-AIFLM-Lab/FinEval

    275Ver en GitHub↗

    The FinEval financial domain evaluation benchmark, based on quantitative fundamental methods and developed through long-term objective research, summarization, and rigorous manual screening, utilizes over 26,000 diverse question types that are highly consistent with real-world application scenarios.

    Python
    Ver en GitHub↗275
  • tongjifinlab/cfbenchmarkAvatar de TongjiFinLab

    TongjiFinLab/CFBenchmark

    55Ver en GitHub↗

    Chinese Financial Assistant Benchmark for Large Language Model

    Python
    Ver en GitHub↗55
  • wunderlabs-dev/claudebin.comW

    wunderlabs-dev/claudebin.com

    0Ver en GitHub↗
    Ver en GitHub↗0
  • yzfly/awesome-claude-promptsY

    yzfly/awesome-claude-prompts

    0Ver en GitHub↗
    Ver en GitHub↗0
  • yzfly/langgptY

    yzfly/LangGPT

    0Ver en GitHub↗
    Ver en GitHub↗0