awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Back to thudm/agentbench

Open-source alternatives to AgentBench

13 open-source projects similar to thudm/agentbench, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best AgentBench alternative.

  • snap-stanford/mlagentbenchsnap-stanford avatar

    snap-stanford/MLAgentBench

    342View on GitHub↗

    MLAgentBench is a suite of end-to-end Machine Learning (ML) experimentation tasks for benchmarking AI agents, where the agent aims to take a given dataset and a machine learning task description and autonomously develop or improve an ML model. Paper: https://arxiv.org/abs/2310.03302

    Python
    View on GitHub↗342
  • rllm-org/rllmrllm-org avatar

    rllm-org/rllm

    5,641View on GitHub↗

    rllm is an asynchronous reinforcement learning framework for training language agents. It provides a unified pipeline that runs the same agent code for both evaluation and training, automatically capturing traces for gradient computation. The framework supports distributed reinforcement learning across multiple GPUs and nodes using pluggable backends, and executes agents in isolated sandboxes—either locally or in the cloud—for safe and scalable rollout collection. It trains agents built with LangGraph, SmolAgents, OpenAI Agents SDK, or custom frameworks without requiring core logic changes. T

    Pythonagent-frameworkagentic-workflowcoding-agent
    View on GitHub↗5,641
  • microsoft/faramicrosoft avatar

    microsoft/fara

    5,901View on GitHub↗

    FARA is a visual computer-use agent model that controls a browser by predicting screen coordinates for clicking, typing, and scrolling, without relying on DOM or accessibility trees. It is designed to automate multi-step web tasks such as searching, form filling, booking, and shopping by reasoning over visual state and decomposing tasks into sequential actions. The model uses a compact 7-billion-parameter decoder-only transformer that can run on consumer GPUs for low-latency on-device inference, or be deployed as a managed endpoint on Azure Foundry for cloud-based inference without local infr

    Pythonagentbrowser-usecomputer-use
    View on GitHub↗5,901

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Find more with AI search
  • gersteinlab/ml-benchgersteinlab avatar

    gersteinlab/ML-Bench

    315View on GitHub↗

    ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code (https://arxiv.org/abs/2311.09835)

    Python
    View on GitHub↗315
  • masworks/x-masM

    MASWorks/X-MAS

    0View on GitHub↗

    2025/05/23 See our preprint paper in ArXiv.

    View on GitHub↗0
  • microsoft/smartplayM

    microsoft/SmartPlay

    0View on GitHub↗
    View on GitHub↗0
  • multiagentbench/marbleMultiagentBench avatar

    MultiagentBench/MARBLE

    52View on GitHub↗

    Now the official Code for MultiagentBench has been moved to MARBLE

    Python
    View on GitHub↗52
  • openai/mle-benchopenai avatar

    openai/mle-bench

    1,316View on GitHub↗
    Python
    View on GitHub↗1,316
  • ruc-gsai/yulan-swarmintellRUC-GSAI avatar

    RUC-GSAI/YuLan-SwarmIntell

    34View on GitHub↗

    Figure 1: Natural Swarm Intelligence Inspiration and SwarmBench Tasks.

    Python
    View on GitHub↗34
  • agiresearch/openagiagiresearch avatar

    agiresearch/OpenAGI

    2,269View on GitHub↗

    OpenAGI: When LLM Meets Domain Experts

    Python
    View on GitHub↗2,269
  • xlang-ai/osworldxlang-ai avatar

    xlang-ai/OSWorld

    2,584View on GitHub↗

    OSWorld is an evaluation framework and multimodal agent benchmark designed to test the ability of large language models to complete complex tasks within virtualized operating system environments. It provides a virtualized desktop sandbox and a virtual machine orchestrator to deploy, snapshot, and reset cloud-based desktops, ensuring reproducible test states for AI agent interactions. The system distinguishes itself by providing an OS-level action space that translates model decisions into mouse clicks, keyboard inputs, and system commands. It employs a standardized interface to integrate vari

    Pythonagentartificial-intelligencebenchmark
    View on GitHub↗2,584
  • chchenhui/mlrbenchchchenhui avatar

    chchenhui/mlrbench

    30View on GitHub↗

    NeurIPS 2025 D&B Track MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research

    Python
    View on GitHub↗30
  • facebookresearch/mlgymfacebookresearch avatar

    facebookresearch/MLGym

    604View on GitHub↗

    MLGym A New Framework and Benchmark for Advancing AI Research Agents

    Python
    View on GitHub↗604