How this analysis was created: This summary and feature list were written by an AI model that read the project's README and public documentation pages. Each feature links to the documentation it came from; stars, license and language come straight from the GitHub API. The model does not read the source code, and the analysis is refreshed when the project is re-analysed. Learn more on our About page.
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
The main features of deepseek-ai/deepseek-math are: Critic-Free Algorithms.
Open-source alternatives to deepseek-ai/deepseek-math include: bytedtsinghua-sia/dapo — DAPO: an Open-source RL System from ByteDance Seed and Tsinghua AIR. liziniu/remax — ReMax is a reinforcement learning method, tailored for reward maximization in RLHF. mcgill-nlp/vineppo — Paper - Abstract - Updates - Quick Start - Installation - Download the datasets - Create Experiment Script - Single… minimax-ai/minimax-m1 — MiniMax-M1, the world's first open-weight, large-scale hybrid-attention reasoning model. modalminds/mm-eureka — MM-EUREKA: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning. openrlhf/openrlhf — OpenRLHF is a training framework and alignment library designed for reinforcement learning from human feedback across…
ReMax is a reinforcement learning method, tailored for reward maximization in RLHF.
Paper - Abstract - Updates - Quick Start - Installation - Download the datasets - Create Experiment Script - Single GPU Training (Only for Rho models) - Running the experiments - Code Structure - Initial SFT Checkpoints - Acknowledgement - Citation
MiniMax-M1, the world's first open-weight, large-scale hybrid-attention reasoning model.
DAPO: an Open-source RL System from ByteDance Seed and Tsinghua AIR