How this analysis was created: This summary and feature list were written by an AI model that read the project's README and public documentation pages. Each feature links to the documentation it came from; stars, license and language come straight from the GitHub API. The model does not read the source code, and the analysis is refreshed when the project is re-analysed. Learn more on our About page.
This repository contains the code for our paper titled "Optimizing Test-Time Compute via Meta Reinforcement Finetuning." In this work, we introduce a novel approach to optimizing test-time compute through meta reinforcement learning, aiming to balance the efficiency and discovery capabilities ofβ¦
π Segment Policy Optimization: Effective Segment-Level Credit Assignment in RL for Large Language Models π
[2025/10/23] π₯π₯Our paper is accepted by NeurIPS 2025.π₯π₯ - [2025/04/22] Released our Paper on arXiv. See here - [2025/03/24] We re-implement our algorithm based on verl. β¨β¨ Key features: (1) add ~50 additional metrics to comprehensively monitor the training process and stability, (2) add aβ¦
The main features of cjreinforce/pure are: Dense Reward Optimization.
Open-source alternatives to cjreinforce/pure include: amap-ml/tree-grpo β Tree Search for LLM Agent Reinforcement Learning. chenluye99/prof β Introduction. cmu-aire/mrt β This repository contains the code for our paper titled "Optimizing Test-Time Compute via Meta Reinforcementβ¦ dongguanting/tool-star β π§β¨Tool-Star: Empowering Multi-Tool Collaborative Web Agent via Reinforcement Learning. gen-verse/reasonflux β Princeton University \& PKU \& UIUC \& University of Chicago \& ByteDance Seed. aiframeresearch/spo β π Segment Policy Optimization: Effective Segment-Level Credit Assignment in RL for Large Language Models π.