awesome-repositories.com
Blog
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectDespreCum realizăm clasamentulPresăServer MCP
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
open-compass avatar

open-compass/DevBench

0
View on GitHub↗
131 stele·13 fork-uri·Python·Apache-2.0·1 vizualizare

DevBench

👋 Overview | 📖 Benchmarking | ⚙️ Setup | 🚀 Usage | 🔎 Citation | 📄 License

Features

  • Benchmarks and Datasets - Comprehensive benchmark covering various software development lifecycle stages.

Istoric stele

Graficul istoricului de stele pentru open-compass/devbenchGraficul istoricului de stele pentru open-compass/devbench

Căutare AI

Explorează mai multe repository-uri excelente

Descrie ce ai nevoie în limbaj simplu — AI-ul sortează mii de proiecte open source selectate în funcție de relevanță.

Start searching with AI

Alternative open-source pentru DevBench

Proiecte open-source similare, clasificate după numărul de funcționalități comune cu DevBench.
  • rllm-org/rllmAvatar rllm-org

    rllm-org/rllm

    5,641Vezi pe GitHub↗

    rllm is an asynchronous reinforcement learning framework for training language agents. It provides a unified pipeline that runs the same agent code for both evaluation and training, automatically capturing traces for gradient computation. The framework supports distributed reinforcement learning across multiple GPUs and nodes using pluggable backends, and executes agents in isolated sandboxes—either locally or in the cloud—for safe and scalable rollout collection. It trains agents built with LangGraph, SmolAgents, OpenAI Agents SDK, or custom frameworks without requiring core logic changes. T

    Pythonagent-frameworkagentic-workflowcoding-agent
    Vezi pe GitHub↗5,641
  • evalplus/evalplusAvatar evalplus

    evalplus/evalplus

    1,765Vezi pe GitHub↗

    Rigourous evaluation of LLM-synthesized code - NeurIPS 2023 & COLM 2024

    Python
    Vezi pe GitHub↗1,765
  • gersteinlab/biocoderAvatar gersteinlab

    gersteinlab/BioCoder

    58Vezi pe GitHub↗

    BioCoder is a challenging bioinformatics code generation benchmark for examining the capabilities of state-of-the-art large language models (LLMs).

    Jupyter Notebook
    Vezi pe GitHub↗58
  • bytedance/fullstackbenchAvatar bytedance

    bytedance/FullStackBench

    121Vezi pe GitHub↗

    FullStack Bench: Evaluating LLMs as Full Stack Coders

    Python
    Vezi pe GitHub↗121
Vezi toate cele 15 alternative pentru DevBench→

Întrebări frecvente

Ce face open-compass/devbench?

👋 Overview | 📖 Benchmarking | ⚙️ Setup | 🚀 Usage | 🔎 Citation | 📄 License

Care sunt principalele funcționalități ale open-compass/devbench?

Principalele funcționalități ale open-compass/devbench sunt: Benchmarks and Datasets.

Care sunt câteva alternative open-source pentru open-compass/devbench?

Alternativele open-source pentru open-compass/devbench includ: rllm-org/rllm — rllm is an asynchronous reinforcement learning framework for training language agents. It provides a unified pipeline… evalplus/evalplus — Rigourous evaluation of LLM-synthesized code - NeurIPS 2023 & COLM 2024. gersteinlab/biocoder — BioCoder is a challenging bioinformatics code generation benchmark for examining the capabilities of state-of-the-art… google-research/google-research — This repository serves as a comprehensive research platform and toolkit for advancing machine learning, quantum… hendrycks/apps — This is the repository for Measuring Coding Challenge Competence With APPS by Dan Hendrycks\, Steven Basart\, Saurav… bytedance/fullstackbench — FullStack Bench: Evaluating LLMs as Full Stack Coders.