How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.
ReMax is a reinforcement learning method, tailored for reward maximization in RLHF.
Paper - Abstract - Updates - Quick Start - Installation - Download the datasets - Create Experiment Script - Single GPU Training (Only for Rho models) - Running the experiments - Code Structure - Initial SFT Checkpoints - Acknowledgement - Citation
MiniMax-M1, the world's first open-weight, large-scale hybrid-attention reasoning model.
DAPO: an Open-source RL System from ByteDance Seed and Tsinghua AIR
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
The main features of deepseek-ai/deepseek-math are: Critic-Free Algorithms.
Projects with overlapping indexed features include: bytedtsinghua-sia/dapo — DAPO: an Open-source RL System from ByteDance Seed and Tsinghua AIR. liziniu/remax — ReMax is a reinforcement learning method, tailored for reward maximization in RLHF. mcgill-nlp/vineppo — Paper - Abstract - Updates - Quick Start - Installation - Download the datasets - Create Experiment Script - Single… minimax-ai/minimax-m1 — MiniMax-M1, the world's first open-weight, large-scale hybrid-attention reasoning model. modalminds/mm-eureka — MM-EUREKA: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning. openrlhf/openrlhf — OpenRLHF is a training framework and alignment library designed for reinforcement learning from human feedback across…