11 open-source projects similar to yaof20/flash-rl, ranked by shared indexed features. Tags may describe platforms or build tools rather than the same primary purpose. Check each project’s use case, license, and deployment requirements before treating it as a replacement.
DAPO: an Open-source RL System from ByteDance Seed and Tsinghua AIR
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
ReMax is a reinforcement learning method, tailored for reward maximization in RLHF.
Paper - Abstract - Updates - Quick Start - Installation - Download the datasets - Create Experiment Script - Single GPU Training (Only for Rho models) - Running the experiments - Code Structure - Initial SFT Checkpoints - Acknowledgement - Citation
MiniMax-M1, the world's first open-weight, large-scale hybrid-attention reasoning model.
MM-EUREKA: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
OpenRLHF is a training framework and alignment library designed for reinforcement learning from human feedback across distributed GPU clusters. It provides tools for aligning large language models and multimodal vision-language models using algorithms such as PPO, GRPO, and DPO. The framework distinguishes itself through a distributed inference engine that overlaps sample rollout with training to increase throughput. It supports scaling to models exceeding 70 billion parameters via parameter sharding and handles long-context sequences through ring-attention sequence parallelism. The project
)](https://www.alphaxiv.org/abs/2505.22617)
📝 Unified Policy Gradient Estimator • ✨ Hybrid Post-Training 🚀 Getting Started • 📊 Main Results • 💖 Acknowledgements • 📨 Contact • 🎈 Citation