awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Back to chenluye99/prof

Open-source alternatives to PROF

19 open-source projects similar to chenluye99/prof, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best PROF alternative.

  • aiframeresearch/spoAIFrameResearch avatar

    AIFrameResearch/SPO

    53View on GitHub↗

    🚀 Segment Policy Optimization: Effective Segment-Level Credit Assignment in RL for Large Language Models 🌟

    Python
    View on GitHub↗53
  • amap-ml/tree-grpoAMAP-ML avatar

    AMAP-ML/Tree-GRPO

    378View on GitHub↗

    Tree Search for LLM Agent Reinforcement Learning

    Python
    View on GitHub↗378
  • cjreinforce/pureCJReinforce avatar

    CJReinforce/PURE

    169View on GitHub↗

    2025/10/23 🔥🔥Our paper is accepted by NeurIPS 2025.🔥🔥 - 2025/04/22 Released our Paper on arXiv. See here - 2025/03/24 We re-implement our algorithm based on verl. ✨✨ Key features: (1) add ~50 additional metrics to comprehensively monitor the training process and stability, (2) add a…

    Python
    View on GitHub↗169
  • cmu-aire/mrtCMU-AIRe avatar

    CMU-AIRe/MRT

    119View on GitHub↗

    This repository contains the code for our paper titled "Optimizing Test-Time Compute via Meta Reinforcement Finetuning." In this work, we introduce a novel approach to optimizing test-time compute through meta reinforcement learning, aiming to balance the efficiency and discovery capabilities of…

    Python
    View on GitHub↗119
  • dongguanting/tool-stardongguanting avatar

    dongguanting/Tool-Star

    398View on GitHub↗

    🔧✨Tool-Star: Empowering Multi-Tool Collaborative Web Agent via Reinforcement Learning

    Python
    View on GitHub↗398

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Find more with AI search
  • gen-verse/reasonfluxGen-Verse avatar

    Gen-Verse/ReasonFlux

    538View on GitHub↗

    Princeton University \& PKU \& UIUC \& University of Chicago \& ByteDance Seed

    Python
    View on GitHub↗538
  • kwai-klear/klearreasonerKwai-Klear avatar

    Kwai-Klear/KlearReasoner

    82View on GitHub↗

    December 5, 2025 🔍 We propose entropy ratio clipping​ (ERC) to impose a global constraint on the output distribution of the policy model. Experiments demonstrate that ERC can significantly improve the stability of off-policy training. 📄 The paper is available on arXiv.

    Python
    View on GitHub↗82
  • langfengq/verl-agentlangfengQ avatar

    langfengQ/verl-agent

    1,548View on GitHub↗
    Pythonagent-frameworkdeepseek-r1gigpo
    View on GitHub↗1,548
  • mcgill-nlp/vineppoMcGill-NLP avatar

    McGill-NLP/VinePPO

    192View on GitHub↗

    Paper - Abstract - Updates - Quick Start - Installation - Download the datasets - Create Experiment Script - Single GPU Training (Only for Rho models) - Running the experiments - Code Structure - Initial SFT Checkpoints - Acknowledgement - Citation

    Python
    View on GitHub↗192
  • multimodal-art-projection/treepomultimodal-art-projection avatar

    multimodal-art-projection/TreePO

    65View on GitHub↗

    This is the official implementation of TreePO algorithm.

    Python
    View on GitHub↗65
  • openai/prm800kopenai avatar

    openai/prm800k

    2,145View on GitHub↗

    This repository accompanies the paper Let's Verify Step by Step and presents the PRM800K dataset introduced there. PRM800K is a process supervision dataset containing 800,000 step-level correctness labels for model-generated solutions to problems from the MATH dataset. More information on…

    Python
    View on GitHub↗2,145
  • prime-rl/implicitprmPRIME-RL avatar

    PRIME-RL/ImplicitPRM

    171View on GitHub↗

    Free Process Rewards without Process Labels

    Python
    View on GitHub↗171
  • prime-rl/primePRIME-RL avatar

    PRIME-RL/PRIME

    1,863View on GitHub↗

    ✨ Getting Started • 📖 Introduction 🔧 Usage • 📃 Evaluation • 🎈 Citation • 🌻 Acknowledgement • 📈 Star History

    Python
    View on GitHub↗1,863
  • ruc-nlpir/arpoRUC-NLPIR avatar

    RUC-NLPIR/ARPO

    1,049View on GitHub↗

    ICLR 2026 Agentic Reinforced Policy Optimization (ARPO)

    Python
    View on GitHub↗1,049
  • ryanliu112/attnrlRyanLiu112 avatar

    RyanLiu112/AttnRL

    14View on GitHub↗

    👋 Hi, everyone! verl is a RL training library initiated by ByteDance Seed team and maintained by the verl community.

    Python
    View on GitHub↗14
  • thudm/treerlTHUDM avatar

    THUDM/TreeRL

    96View on GitHub↗

    Implementation for ACL'25 paper TreeRL: LLM Reinforcement Learning with On-Policy Tree Search. The implementation is based on OpenRLHF

    Python
    View on GitHub↗96
  • wanghanlinhenry/spa-rl-agentWangHanLinHenry avatar

    WangHanLinHenry/SPA-RL-Agent

    86View on GitHub↗
    Python
    View on GitHub↗86
  • yangzhch6/treerpoyangzhch6 avatar

    yangzhch6/TreeRPO

    6View on GitHub↗

    ```bash conda create -n rllm python=3.10 -y conda activate rllm

    Jupyter Notebook
    View on GitHub↗6
  • zillwang/stepsearchZillwang avatar

    Zillwang/StepSearch

    72View on GitHub↗

    StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

    Python
    View on GitHub↗72