How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
ReMax is a reinforcement learning method, tailored for reward maximization in RLHF.
Paper - Abstract - Updates - Quick Start - Installation - Download the datasets - Create Experiment Script - Single GPU Training (Only for Rho models) - Running the experiments - Code Structure - Initial SFT Checkpoints - Acknowledgement - Citation
DAPO: an Open-source RL System from ByteDance Seed and Tsinghua AIR
π Unified Policy Gradient Estimator β’ β¨ Hybrid Post-Training π Getting Started β’ π Main Results β’ π Acknowledgements β’ π¨ Contact β’ π Citation
The main features of tsinghuac3i/unify-post-training are: Critic-Free Algorithms.
Projects with overlapping indexed features include: bytedtsinghua-sia/dapo β DAPO: an Open-source RL System from ByteDance Seed and Tsinghua AIR. deepseek-ai/deepseek-math β DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. liziniu/remax β ReMax is a reinforcement learning method, tailored for reward maximization in RLHF. mcgill-nlp/vineppo β Paper - Abstract - Updates - Quick Start - Installation - Download the datasets - Create Experiment Script - Singleβ¦ minimax-ai/minimax-m1 β MiniMax-M1, the world's first open-weight, large-scale hybrid-attention reasoning model. modalminds/mm-eureka β MM-EUREKA: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.