awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Back to ruc-nlpir/arpo

Projects sharing features with ARPO

30 open-source projects similar to ruc-nlpir/arpo, ranked by shared indexed features. Tags may describe platforms or build tools rather than the same primary purpose. Check each project’s use case, license, and deployment requirements before treating it as a replacement.

  • prime-rl/primePRIME-RL avatar

    PRIME-RL/PRIME

    1,863View on GitHub↗

    ✨ Getting Started • 📖 Introduction 🔧 Usage • 📃 Evaluation • 🎈 Citation • 🌻 Acknowledgement • 📈 Star History

    Python
    View on GitHub↗1,863
  • yangzhch6/treerpoyangzhch6 avatar

    yangzhch6/TreeRPO

    6View on GitHub↗

    ```bash conda create -n rllm python=3.10 -y conda activate rllm

    Jupyter Notebook
    View on GitHub↗6
  • multimodal-art-projection/treepomultimodal-art-projection avatar

    multimodal-art-projection/TreePO

    65View on GitHub↗

    This is the official implementation of TreePO algorithm.

    Python
    View on GitHub↗65
  • thudm/treerlTHUDM avatar

    THUDM/TreeRL

    96View on GitHub↗

    Implementation for ACL'25 paper TreeRL: LLM Reinforcement Learning with On-Policy Tree Search. The implementation is based on OpenRLHF

    Python
    View on GitHub↗96
  • trademaster-ntu/trademasterTradeMaster-NTU avatar

    TradeMaster-NTU/TradeMaster

    2,484View on GitHub↗

    TradeMaster is a reinforcement learning trading framework and algorithmic trading simulator designed for designing and testing quantitative trading strategies. The system provides a platform for developing reinforcement learning agents, managing quantitative portfolios, and optimizing trade execution using financial market data. The project features specialized components for multi-modality data preprocessing, a high-fidelity market environment simulation for strategy backtesting, and a quantitative portfolio manager for capital reallocation across multiple assets. It includes a trade executi

    Jupyter Notebookfinancefintechinvestment-strategies
    View on GitHub↗2,484

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Find more with AI search
  • chanliang/eepoChanLiang avatar

    ChanLiang/EEPO

    0View on GitHub↗
    View on GitHub↗0
  • charlesq9/alitaCharlesQ9 avatar

    CharlesQ9/Alita

    881View on GitHub↗

    The GAIA game is over, and Alita is the final answer.

    View on GitHub↗881
  • chen-gx/toolevoChen-GX avatar

    Chen-GX/ToolEVO

    13View on GitHub↗

    This repository contains the code and benchmark (ToolQA-D) for our research paper titled "Learning Evolving Tools for Large Language Models", which has been accepted at ICLR 2025.

    Python
    View on GitHub↗13
  • chenluye99/profChenluye99 avatar

    Chenluye99/PROF

    11View on GitHub↗

    Introduction

    View on GitHub↗11
  • chenxinan-fdu/polarisChenxinAn-fdu avatar

    ChenxinAn-fdu/POLARIS

    688View on GitHub↗

    🌠 A PO st-training recipe for scaling R L on A dvanced R eason I ng model S 🚀

    Python
    View on GitHub↗688
  • cjreinforce/pureCJReinforce avatar

    CJReinforce/PURE

    169View on GitHub↗

    2025/10/23 🔥🔥Our paper is accepted by NeurIPS 2025.🔥🔥 - 2025/04/22 Released our Paper on arXiv. See here - 2025/03/24 We re-implement our algorithm based on verl. ✨✨ Key features: (1) add ~50 additional metrics to comprehensively monitor the training process and stability, (2) add a…

    Python
    View on GitHub↗169
  • clova-tool/clova-toolclova-tool avatar

    clova-tool/CLOVA-tool

    30View on GitHub↗

    Zhi Gao, Yuntao Du, Xintong Zhang, Xiaojian Ma, Wenjuan Han, Song-Chun Zhu, Qing Li

    Python
    View on GitHub↗30
  • cmu-aire/mrtCMU-AIRe avatar

    CMU-AIRe/MRT

    119View on GitHub↗

    This repository contains the code for our paper titled "Optimizing Test-Time Compute via Meta Reinforcement Finetuning." In this work, we introduce a novel approach to optimizing test-time compute through meta reinforcement learning, aiming to balance the efficiency and discovery capabilities of…

    Python
    View on GitHub↗119
  • dongguanting/tool-stardongguanting avatar

    dongguanting/Tool-Star

    398View on GitHub↗

    🔧✨Tool-Star: Empowering Multi-Tool Collaborative Web Agent via Reinforcement Learning

    Python
    View on GitHub↗398
  • gen-verse/reasonfluxGen-Verse avatar

    Gen-Verse/ReasonFlux

    538View on GitHub↗

    Princeton University \& PKU \& UIUC \& University of Chicago \& ByteDance Seed

    Python
    View on GitHub↗538
  • kwai-klear/klearreasonerKwai-Klear avatar

    Kwai-Klear/KlearReasoner

    82View on GitHub↗

    December 5, 2025 🔍 We propose entropy ratio clipping​ (ERC) to impose a global constraint on the output distribution of the policy model. Experiments demonstrate that ERC can significantly improve the stability of off-policy training. 📄 The paper is available on arXiv.

    Python
    View on GitHub↗82
  • langfengq/verl-agentlangfengQ avatar

    langfengQ/verl-agent

    1,548View on GitHub↗
    Pythonagent-frameworkdeepseek-r1gigpo
    View on GitHub↗1,548
  • liaomengqi/e3-rl4llmsLiaoMengqi avatar

    LiaoMengqi/E3-RL4LLMs

    17View on GitHub↗

    25/08/20 : Aceept as EMNLP 2025 Main Conference paper

    Python
    View on GitHub↗17
  • mat-agent/mat-agentmat-agent avatar

    mat-agent/MAT-Agent

    93View on GitHub↗

    🚀 VLM-Powered Agent for Intelligent Tool Orchestration | 🔧 Open-source Framework for Multi-modal AI

    Python
    View on GitHub↗93
  • mcgill-nlp/vineppoMcGill-NLP avatar

    McGill-NLP/VinePPO

    192View on GitHub↗

    Paper - Abstract - Updates - Quick Start - Installation - Download the datasets - Create Experiment Script - Single GPU Training (Only for Rho models) - Running the experiments - Code Structure - Initial SFT Checkpoints - Acknowledgement - Citation

    Python
    View on GitHub↗192
  • microsoft/jarvismicrosoft avatar

    microsoft/JARVIS

    24,854View on GitHub↗

    JARVIS is a system for large language model task orchestration, deployment management, and automation benchmarking. It utilizes a task orchestrator to decompose complex requests into actionable steps and coordinates various expert models to synthesize final responses. The project includes an AI model deployment manager to handle the local deployment of expert models across different hardware scales. It further provides an AI workflow API consisting of web endpoints used to trigger automated task workflows and retrieve results from model selection stages. The framework incorporates an automat

    Python
    View on GitHub↗24,854
  • microsoft/simulated-trial-and-errormicrosoft avatar

    microsoft/simulated-trial-and-error

    123View on GitHub↗

    ``` STE/ ├─ toolmetadata/: tool related metadata ├─ prompts/: full prompts used ├─ savedresults/: prediction results in json ├─ {, FT, ICL}.json: results for baseline model, tool-enhanced w/ fine-tuning, tool-enhanced with ICL ├─ CLround.json: continual learning (each round) ├─ main.py: main…

    Python
    View on GitHub↗123
  • netease-youdao/confucius3-mathnetease-youdao avatar

    netease-youdao/Confucius3-Math

    94View on GitHub↗

    💜 Confucius Demo | 🤗 Hugging Face | 🤖 ModelScope | ⌨️ GitHub | 📚 Paper | 💬 Wechat

    Python
    View on GitHub↗94
  • nvlabs/tool-n1NVlabs avatar

    NVlabs/Tool-N1

    229View on GitHub↗

    This is the official implementation of paper Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning, where we present Nemotron-Research-Tool-N1, a family of tool-using reasoning language models. These models are trained with an R1-style reinforcement learning…

    Python
    View on GitHub↗229
  • oceanntwt/tool-plannerOceannTwT avatar

    OceannTwT/Tool-Planner

    116View on GitHub↗

    ICLR 2025 Tool-Planner: Task Planning with Clusters across Multiple Tools

    Python
    View on GitHub↗116
  • openai/prm800kopenai avatar

    openai/prm800k

    2,145View on GitHub↗

    This repository accompanies the paper Let's Verify Step by Step and presents the PRM800K dataset introduced there. PRM800K is a process supervision dataset containing 800,000 step-level correctness labels for model-generated solutions to problems from the MATH dataset. More information on…

    Python
    View on GitHub↗2,145
  • openbmb/toolbenchOpenBMB avatar

    OpenBMB/ToolBench

    5,672View on GitHub↗

    ToolBench is an open platform for training, serving, and evaluating large language models that retrieve and call real-world APIs to complete user instructions. It provides an API-aware inference engine that selects relevant tools from a large corpus and generates sequences of tool calls to produce final answers, along with a custom API registration system that lets users add their own REST endpoints for the model to discover and invoke. The platform includes a complete instruction-tuning pipeline for training models on curated tool-use data, a multi-tool execution engine that coordinates sequ

    Python
    View on GitHub↗5,672
  • pku-baichuan-mlsystemlab/buttonPKU-Baichuan-MLSystemLab avatar

    PKU-Baichuan-MLSystemLab/BUTTON

    27View on GitHub↗

    Large Language Models (LLMs) have exhibited significant potential in performing diverse tasks, including the ability to call functions or use external tools to enhance their performance. While current research on function calling by LLMs primarily focuses on single-turn interactions, this paper…

    View on GitHub↗27
  • prime-rl/implicitprmPRIME-RL avatar

    PRIME-RL/ImplicitPRM

    171View on GitHub↗

    Free Process Rewards without Process Labels

    Python
    View on GitHub↗171
  • qiancheng0/creatorqiancheng0 avatar

    qiancheng0/CREATOR

    31View on GitHub↗

    We release here the code for most of the main experiments in paper CREATOR: Tool Creation for Disentangling Abstract and Concrete Reasoning of Large Language Models, including the code of our framework and various baselines. We also release two new datasets: Creation Challenge and Tool Transfer…

    Python
    View on GitHub↗31