How this analysis was created: This summary and feature list were written by an AI model that read the project's README and public documentation pages. Each feature links to the documentation it came from; stars, license and language come straight from the GitHub API. The model does not read the source code, and the analysis is refreshed when the project is re-analysed. Learn more on our About page.
2025/10/23 🔥🔥Our paper is accepted by NeurIPS 2025.🔥🔥 - 2025/04/22 Released our Paper on arXiv. See here - 2025/03/24 We re-implement our algorithm based on verl. ✨✨ Key features: (1) add ~50 additional metrics to comprehensively monitor the training process and stability, (2) add a…
This repository contains the code for our paper titled "Optimizing Test-Time Compute via Meta Reinforcement Finetuning." In this work, we introduce a novel approach to optimizing test-time compute through meta reinforcement learning, aiming to balance the efficiency and discovery capabilities of…
🚀 Segment Policy Optimization: Effective Segment-Level Credit Assignment in RL for Large Language Models 🌟
Introduction
The main features of chenluye99/prof are: Dense Reward Optimization.
Open-source alternatives to chenluye99/prof include: amap-ml/tree-grpo — Tree Search for LLM Agent Reinforcement Learning. cjreinforce/pure — [2025/10/23] 🔥🔥Our paper is accepted by NeurIPS 2025.🔥🔥 - [2025/04/22] Released our Paper on arXiv. See here -… cmu-aire/mrt — This repository contains the code for our paper titled "Optimizing Test-Time Compute via Meta Reinforcement… dongguanting/tool-star — 🔧✨Tool-Star: Empowering Multi-Tool Collaborative Web Agent via Reinforcement Learning. gen-verse/reasonflux — Princeton University \& PKU \& UIUC \& University of Chicago \& ByteDance Seed. aiframeresearch/spo — 🚀 Segment Policy Optimization: Effective Segment-Level Credit Assignment in RL for Large Language Models 🌟.