How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.
The main features of rlhflow/online-rlhf are: Fine-Tuning Frameworks, RLHF Frameworks.
Projects with overlapping indexed features include: volcengine/verl — verl is a distributed training system designed for large language model alignment and reinforcement learning. It… evolvinglmms-lab/lmms-engine. alibaba/chatlearn — A flexible and efficient training framework for large-scale alignment tasks. blaizzy/mlx-vlm. alisawuffles/proxy-tuning. bytedance-seed/veomni.
verl is a distributed training system designed for large language model alignment and reinforcement learning. It provides a framework for executing post-training pipelines, including supervised fine-tuning and reinforcement learning from human feedback, to refine model behavior and agentic capabilities. The system utilizes a hybrid training and inference engine that optimizes memory and communication when switching between model generation and gradient updates. It supports multi-modal reinforcement learning for models processing both image and text data, and implements algorithms such as PPO
A flexible and efficient training framework for large-scale alignment tasks