12 repositorios
Policy optimization methods that do not rely on explicit value function modeling.
Explore 12 awesome GitHub repositories matching part of an awesome list · Critic-Free Algorithms. Refine with filters or upvote what's useful.
OpenRLHF is a training framework and alignment library designed for reinforcement learning from human feedback across distributed GPU clusters. It provides tools for aligning large language models and multimodal vision-language models using algorithms such as PPO, GRPO, and DPO. The framework distinguishes itself through a distributed inference engine that overlaps sample rollout with training to increase throughput. It supports scaling to models exceeding 70 billion parameters via parameter sharding and handles long-context sequences through ring-attention sequence parallelism. The project
Robust reinforcement learning algorithm for human feedback alignment.
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Policy optimization for mathematical reasoning in open models.
MiniMax-M1, the world's first open-weight, large-scale hybrid-attention reasoning model.
Scaling test-time compute using efficient attention mechanisms.
DAPO: an Open-source RL System from ByteDance Seed and Tsinghua AIR
Large-scale open-source reinforcement learning system for LLMs.
Critical analysis and implementation of reasoning-focused training.
MM-EUREKA: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
Stable rule-based reinforcement learning for language models.
)](https://www.alphaxiv.org/abs/2505.22617)
Entropy-based mechanisms for reasoning model reinforcement learning.
Fast RL training with Quantized Rollouts ( Blog )
Accelerated reinforcement learning training using quantized rollouts.
📝 Unified Policy Gradient Estimator • ✨ Hybrid Post-Training 🚀 Getting Started • 📊 Main Results • 💖 Acknowledgements • 📨 Contact • 🎈 Citation
Unified post-training frameworks for large language models.
ReMax is a reinforcement learning method, tailored for reward maximization in RLHF.
Simple and efficient alignment method for large language models.
Paper - Abstract - Updates - Quick Start - Installation - Download the datasets - Create Experiment Script - Single GPU Training (Only for Rho models) - Running the experiments - Code Structure - Initial SFT Checkpoints - Acknowledgement - Citation
Refined credit assignment for unlocking reasoning potential.
Yihe Deng , Nanyun Peng , Kai-Wei Chang
Iterative SFT-RL cycles for complex vision-language reasoning.