How this analysis was created: This summary and feature list were written by an AI model that read the project's README and public documentation pages. Each feature links to the documentation it came from; stars, license and language come straight from the GitHub API. The model does not read the source code, and the analysis is refreshed when the project is re-analysed. Learn more on our About page.
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
ReMax is a reinforcement learning method, tailored for reward maximization in RLHF.
Paper - Abstract - Updates - Quick Start - Installation - Download the datasets - Create Experiment Script - Single GPU Training (Only for Rho models) - Running the experiments - Code Structure - Initial SFT Checkpoints - Acknowledgement - Citation
DAPO: an Open-source RL System from ByteDance Seed and Tsinghua AIR
Fast RL training with Quantized Rollouts ( Blog )
The main features of yaof20/flash-rl are: Critic-Free Algorithms.
Open-source alternatives to yaof20/flash-rl include: bytedtsinghua-sia/dapo — DAPO: an Open-source RL System from ByteDance Seed and Tsinghua AIR. deepseek-ai/deepseek-math — DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. liziniu/remax — ReMax is a reinforcement learning method, tailored for reward maximization in RLHF. mcgill-nlp/vineppo — Paper - Abstract - Updates - Quick Start - Installation - Download the datasets - Create Experiment Script - Single… minimax-ai/minimax-m1 — MiniMax-M1, the world's first open-weight, large-scale hybrid-attention reasoning model. modalminds/mm-eureka — MM-EUREKA: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.