awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
R

RLHFlow/Online-RLHF

0
View on GitHub↗
0 stars·0 forks·8 views

Online RLHF

Features

  • Fine-Tuning Frameworks - Recipe for online iterative reinforcement learning.
  • RLHF Frameworks - Framework for iterative online reinforcement learning workflows.
  • Fine-Tuning Frameworks - Recipe for online RLHF and iterative DPO.

Star history

Star history chart for rlhflow/online-rlhfStar history chart for rlhflow/online-rlhf

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Frequently asked questions

What are the main features of rlhflow/online-rlhf?

The main features of rlhflow/online-rlhf are: Fine-Tuning Frameworks, RLHF Frameworks.

Which projects share features with rlhflow/online-rlhf?

Projects with overlapping indexed features include: volcengine/verl — verl is a distributed training system designed for large language model alignment and reinforcement learning. It… evolvinglmms-lab/lmms-engine. alibaba/chatlearn — A flexible and efficient training framework for large-scale alignment tasks. blaizzy/mlx-vlm. alisawuffles/proxy-tuning. bytedance-seed/veomni.

Projects sharing features with Online RLHF

These projects share indexed features with Online RLHF. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • volcengine/verlvolcengine avatar

    volcengine/verl

    22,015View on GitHub↗

    verl is a distributed training system designed for large language model alignment and reinforcement learning. It provides a framework for executing post-training pipelines, including supervised fine-tuning and reinforcement learning from human feedback, to refine model behavior and agentic capabilities. The system utilizes a hybrid training and inference engine that optimizes memory and communication when switching between model generation and gradient updates. It supports multi-modal reinforcement learning for models processing both image and text data, and implements algorithms such as PPO

    Python
    View on GitHub↗22,015
  • alisawuffles/proxy-tuningA

    alisawuffles/proxy-tuning

    0View on GitHub↗
    View on GitHub↗0
  • alibaba/chatlearnalibaba avatar

    alibaba/ChatLearn

    452View on GitHub↗

    A flexible and efficient training framework for large-scale alignment tasks

    Python
    View on GitHub↗452
  • blaizzy/mlx-vlmBlaizzy avatar

    Blaizzy/mlx-vlm

    2,157View on GitHub↗
    Pythonapple-siliconflorence2idefics
    View on GitHub↗2,157
Compare all 30 related projects→