How this analysis was created: This summary and feature list were written by an AI model that read the project's README and public documentation pages. Each feature links to the documentation it came from; stars, license and language come straight from the GitHub API. The model does not read the source code, and the analysis is refreshed when the project is re-analysed. Learn more on our About page.
2025/10/23 🔥🔥Our paper is accepted by NeurIPS 2025.🔥🔥 - 2025/04/22 Released our Paper on arXiv. See here - 2025/03/24 We re-implement our algorithm based on verl. ✨✨ Key features: (1) add ~50 additional metrics to comprehensively monitor the training process and stability, (2) add a…
🚀 Segment Policy Optimization: Effective Segment-Level Credit Assignment in RL for Large Language Models 🌟
Free Process Rewards without Process Labels
The main features of prime-rl/implicitprm are: Dense Reward Optimization.
Open-source alternatives to prime-rl/implicitprm include: amap-ml/tree-grpo — Tree Search for LLM Agent Reinforcement Learning. chenluye99/prof — Introduction. cjreinforce/pure — [2025/10/23] 🔥🔥Our paper is accepted by NeurIPS 2025.🔥🔥 - [2025/04/22] Released our Paper on arXiv. See here -… cmu-aire/mrt — This repository contains the code for our paper titled "Optimizing Test-Time Compute via Meta Reinforcement… dongguanting/tool-star — 🔧✨Tool-Star: Empowering Multi-Tool Collaborative Web Agent via Reinforcement Learning. aiframeresearch/spo — 🚀 Segment Policy Optimization: Effective Segment-Level Credit Assignment in RL for Large Language Models 🌟.