awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Back to prime-rl/prime

Projects sharing features with PRIME

30 open-source projects similar to prime-rl/prime, ranked by shared indexed features. Tags may describe platforms or build tools rather than the same primary purpose. Check each project’s use case, license, and deployment requirements before treating it as a replacement.

  • multimodal-art-projection/treepomultimodal-art-projection avatar

    multimodal-art-projection/TreePO

    65View on GitHub↗

    This is the official implementation of TreePO algorithm.

    Python
    View on GitHub↗65
  • ruc-nlpir/arpoRUC-NLPIR avatar

    RUC-NLPIR/ARPO

    1,049View on GitHub↗

    ICLR 2026 Agentic Reinforced Policy Optimization (ARPO)

    Python
    View on GitHub↗1,049
  • yangzhch6/treerpoyangzhch6 avatar

    yangzhch6/TreeRPO

    6View on GitHub↗

    ```bash conda create -n rllm python=3.10 -y conda activate rllm

    Jupyter Notebook
    View on GitHub↗6
  • thudm/treerlTHUDM avatar

    THUDM/TreeRL

    96View on GitHub↗

    Implementation for ACL'25 paper TreeRL: LLM Reinforcement Learning with On-Policy Tree Search. The implementation is based on OpenRLHF

    Python
    View on GitHub↗96
  • trademaster-ntu/trademasterTradeMaster-NTU avatar

    TradeMaster-NTU/TradeMaster

    2,484View on GitHub↗

    TradeMaster is a reinforcement learning trading framework and algorithmic trading simulator designed for designing and testing quantitative trading strategies. The system provides a platform for developing reinforcement learning agents, managing quantitative portfolios, and optimizing trade execution using financial market data. The project features specialized components for multi-modality data preprocessing, a high-fidelity market environment simulation for strategy backtesting, and a quantitative portfolio manager for capital reallocation across multiple assets. It includes a trade executi

    Jupyter Notebookfinancefintechinvestment-strategies
    View on GitHub↗2,484

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Find more with AI search
  • chenxinan-fdu/polarisChenxinAn-fdu avatar

    ChenxinAn-fdu/POLARIS

    688View on GitHub↗

    🌠 A PO st-training recipe for scaling R L on A dvanced R eason I ng model S 🚀

    Python
    View on GitHub↗688
  • cjreinforce/pureCJReinforce avatar

    CJReinforce/PURE

    169View on GitHub↗

    2025/10/23 🔥🔥Our paper is accepted by NeurIPS 2025.🔥🔥 - 2025/04/22 Released our Paper on arXiv. See here - 2025/03/24 We re-implement our algorithm based on verl. ✨✨ Key features: (1) add ~50 additional metrics to comprehensively monitor the training process and stability, (2) add a…

    Python
    View on GitHub↗169
  • cmu-aire/mrtCMU-AIRe avatar

    CMU-AIRe/MRT

    119View on GitHub↗

    This repository contains the code for our paper titled "Optimizing Test-Time Compute via Meta Reinforcement Finetuning." In this work, we introduce a novel approach to optimizing test-time compute through meta reinforcement learning, aiming to balance the efficiency and discovery capabilities of…

    Python
    View on GitHub↗119
  • corl-team/vl-daccorl-team avatar

    corl-team/VL-DAC

    12View on GitHub↗

    Enhancing Vision-Language Model Training with Reinforcement Learning in Synthetic Worlds for Real-World Success

    Python
    View on GitHub↗12
  • dongguanting/tool-stardongguanting avatar

    dongguanting/Tool-Star

    398View on GitHub↗

    🔧✨Tool-Star: Empowering Multi-Tool Collaborative Web Agent via Reinforcement Learning

    Python
    View on GitHub↗398
  • gen-verse/reasonfluxGen-Verse avatar

    Gen-Verse/ReasonFlux

    538View on GitHub↗

    Princeton University \& PKU \& UIUC \& University of Chicago \& ByteDance Seed

    Python
    View on GitHub↗538
  • hitsz-tmg/veripoHITsz-TMG avatar

    HITsz-TMG/VerIPO

    10View on GitHub↗

    📄 Paper Link 🤗 VerIPO-7B-v1.0

    Python
    View on GitHub↗10
  • kwai-klear/klearreasonerKwai-Klear avatar

    Kwai-Klear/KlearReasoner

    82View on GitHub↗

    December 5, 2025 🔍 We propose entropy ratio clipping​ (ERC) to impose a global constraint on the output distribution of the policy model. Experiments demonstrate that ERC can significantly improve the stability of off-policy training. 📄 The paper is available on arXiv.

    Python
    View on GitHub↗82
  • langfengq/verl-agentlangfengQ avatar

    langfengQ/verl-agent

    1,548View on GitHub↗
    Pythonagent-frameworkdeepseek-r1gigpo
    View on GitHub↗1,548
  • liaomengqi/e3-rl4llmsLiaoMengqi avatar

    LiaoMengqi/E3-RL4LLMs

    17View on GitHub↗

    25/08/20 : Aceept as EMNLP 2025 Main Conference paper

    Python
    View on GitHub↗17
  • lifan-yuan/implicitprmlifan-yuan avatar

    lifan-yuan/ImplicitPRM

    171View on GitHub↗

    Free Process Rewards without Process Labels

    Python
    View on GitHub↗171
  • mcgill-nlp/vineppoMcGill-NLP avatar

    McGill-NLP/VinePPO

    192View on GitHub↗

    Paper - Abstract - Updates - Quick Start - Installation - Download the datasets - Create Experiment Script - Single GPU Training (Only for Rho models) - Running the experiments - Code Structure - Initial SFT Checkpoints - Acknowledgement - Citation

    Python
    View on GitHub↗192
  • netease-youdao/confucius3-mathnetease-youdao avatar

    netease-youdao/Confucius3-Math

    94View on GitHub↗

    💜 Confucius Demo | 🤗 Hugging Face | 🤖 ModelScope | ⌨️ GitHub | 📚 Paper | 💬 Wechat

    Python
    View on GitHub↗94
  • open-reasoner-zero/open-reasoner-zeroOpen-Reasoner-Zero avatar

    Open-Reasoner-Zero/Open-Reasoner-Zero

    2,095View on GitHub↗

    An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

    Python
    View on GitHub↗2,095
  • openai/prm800kopenai avatar

    openai/prm800k

    2,145View on GitHub↗

    This repository accompanies the paper Let's Verify Step by Step and presents the PRM800K dataset introduced there. PRM800K is a process supervision dataset containing 800,000 step-level correctness labels for model-generated solutions to problems from the MATH dataset. More information on…

    Python
    View on GitHub↗2,145
  • prime-rl/implicitprmPRIME-RL avatar

    PRIME-RL/ImplicitPRM

    171View on GitHub↗

    Free Process Rewards without Process Labels

    Python
    View on GitHub↗171
  • rllm-org/rllmrllm-org avatar

    rllm-org/rllm

    5,641View on GitHub↗

    rllm is an asynchronous reinforcement learning framework for training language agents. It provides a unified pipeline that runs the same agent code for both evaluation and training, automatically capturing traces for gradient computation. The framework supports distributed reinforcement learning across multiple GPUs and nodes using pluggable backends, and executes agents in isolated sandboxes—either locally or in the cloud—for safe and scalable rollout collection. It trains agents built with LangGraph, SmolAgents, OpenAI Agents SDK, or custom frameworks without requiring core logic changes. T

    Pythonagent-frameworkagentic-workflowcoding-agent
    View on GitHub↗5,641
  • rookie-joe/autopsvrookie-joe avatar

    rookie-joe/AutoPSV

    50View on GitHub↗

    This repository contains the official implementation of AutoPSV: Automated Process-Supervised Verifier, accepted at NeurIPS 2024 (poster).

    Python
    View on GitHub↗50
  • ryanliu112/attnrlRyanLiu112 avatar

    RyanLiu112/AttnRL

    14View on GitHub↗

    👋 Hi, everyone! verl is a RL training library initiated by ByteDance Seed team and maintained by the verl community.

    Python
    View on GitHub↗14
  • t-lab-cuhksz/g2rpo-aT-Lab-CUHKSZ avatar

    T-Lab-CUHKSZ/G2RPO-A

    16View on GitHub↗

    G2RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance

    Python
    View on GitHub↗16
  • vaibhavagg303/dars-agentvaibhavagg303 avatar

    vaibhavagg303/DARS-Agent

    68View on GitHub↗

    DARS: Dynamic Action Re-Sampling to Enhance Coding Agent Performance by Adaptive Tree Traversal

    Python
    View on GitHub↗68
  • wanghanlinhenry/spa-rl-agentWangHanLinHenry avatar

    WangHanLinHenry/SPA-RL-Agent

    86View on GitHub↗
    Python
    View on GitHub↗86
  • zhengkid/parallel-r1zhengkid avatar

    zhengkid/Parallel-R1

    259View on GitHub↗

    The official repository for "Parallel-R1: Towards Parallel Thinking via Reinforcement Learning".

    Python
    View on GitHub↗259
  • aiframeresearch/spoAIFrameResearch avatar

    AIFrameResearch/SPO

    53View on GitHub↗

    🚀 Segment Policy Optimization: Effective Segment-Level Credit Assignment in RL for Large Language Models 🌟

    Python
    View on GitHub↗53
  • zillwang/stepsearchZillwang avatar

    Zillwang/StepSearch

    72View on GitHub↗

    StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

    Python
    View on GitHub↗72