awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Back to mcgill-nlp/vineppo

Projects sharing features with VinePPO

30 open-source projects similar to mcgill-nlp/vineppo, ranked by shared indexed features. Tags may describe platforms or build tools rather than the same primary purpose. Check each project’s use case, license, and deployment requirements before treating it as a replacement.

  • openrlhf/openrlhfOpenRLHF avatar

    OpenRLHF/OpenRLHF

    9,675View on GitHub↗

    OpenRLHF is a training framework and alignment library designed for reinforcement learning from human feedback across distributed GPU clusters. It provides tools for aligning large language models and multimodal vision-language models using algorithms such as PPO, GRPO, and DPO. The framework distinguishes itself through a distributed inference engine that overlaps sample rollout with training to increase throughput. It supports scaling to models exceeding 70 billion parameters via parameter sharding and handles long-context sequences through ring-attention sequence parallelism. The project

    Pythonlarge-language-modelsopenai-o1proximal-policy-optimization
    View on GitHub↗9,675
  • allenai/finegrainedrlhfA

    allenai/FineGrainedRLHF

    0View on GitHub↗

    Fine-Grained RLHF

    View on GitHub↗0
  • eit-nlp/accuracyparadox-rlhfE

    EIT-NLP/AccuracyParadox-RLHF

    0View on GitHub↗
    View on GitHub↗0
  • amap-ml/tree-grpoAMAP-ML avatar

    AMAP-ML/Tree-GRPO

    378View on GitHub↗

    Tree Search for LLM Agent Reinforcement Learning

    Python
    View on GitHub↗378
  • anthropics/constitutionalharmlessnesspaperanthropics avatar

    anthropics/ConstitutionalHarmlessnessPaper

    263View on GitHub↗

    This repository provides supplementary material for our paper Constitutional AI: Harmlessness from AI Feedback.

    View on GitHub↗263

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Find more with AI search
  • bytedtsinghua-sia/dapoBytedTsinghua-SIA avatar

    BytedTsinghua-SIA/DAPO

    1,831View on GitHub↗

    DAPO: an Open-source RL System from ByteDance Seed and Tsinghua AIR

    Python
    View on GitHub↗1,831
  • carperai/trlxcarperai avatar

    carperai/trlx

    4,749View on GitHub↗

    trlx is a reinforcement learning library and training framework designed to align large language models using human feedback. It serves as a distributed trainer and compute orchestrator for scaling high-parameter models across multiple GPUs and nodes. The project provides tools for reinforcement learning from human feedback and model alignment. It implements reward-model-based optimization and proximal policy optimization to refine model behavior based on goal-oriented rewards or human-labeled datasets. The framework covers distributed training strategies, including model parallelism, parame

    Python
    View on GitHub↗4,749
  • chenluye99/profChenluye99 avatar

    Chenluye99/PROF

    11View on GitHub↗

    Introduction

    View on GitHub↗11
  • cjreinforce/pureCJReinforce avatar

    CJReinforce/PURE

    169View on GitHub↗

    2025/10/23 🔥🔥Our paper is accepted by NeurIPS 2025.🔥🔥 - 2025/04/22 Released our Paper on arXiv. See here - 2025/03/24 We re-implement our algorithm based on verl. ✨✨ Key features: (1) add ~50 additional metrics to comprehensively monitor the training process and stability, (2) add a…

    Python
    View on GitHub↗169
  • cmu-aire/mrtCMU-AIRe avatar

    CMU-AIRe/MRT

    119View on GitHub↗

    This repository contains the code for our paper titled "Optimizing Test-Time Compute via Meta Reinforcement Finetuning." In this work, we introduce a novel approach to optimizing test-time compute through meta reinforcement learning, aiming to balance the efficiency and discovery capabilities of…

    Python
    View on GitHub↗119
  • cornell-rl/drpoC

    Cornell-RL/drpo

    0View on GitHub↗
    View on GitHub↗0
  • deepseek-ai/deepseek-mathdeepseek-ai avatar

    deepseek-ai/DeepSeek-Math

    3,346View on GitHub↗

    DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

    Python
    View on GitHub↗3,346
  • dongguanting/tool-stardongguanting avatar

    dongguanting/Tool-Star

    398View on GitHub↗

    🔧✨Tool-Star: Empowering Multi-Tool Collaborative Web Agent via Reinforcement Learning

    Python
    View on GitHub↗398
  • dunzeng/moreD

    dunzeng/MORE

    0View on GitHub↗
    View on GitHub↗0
  • allenai/rl4lmsallenai avatar

    allenai/RL4LMs

    2,390View on GitHub↗

    A modular RL library to fine-tune language models to human preferences

    Python
    View on GitHub↗2,390
  • ernie-research/ma-rlhfE

    ernie-research/MA-RLHF

    0View on GitHub↗
    View on GitHub↗0
  • exlaw/dlmaE

    exlaw/DLMA

    0View on GitHub↗
    View on GitHub↗0
  • ganjinzero/rrhfGanjinZero avatar

    GanjinZero/RRHF

    806View on GitHub↗

    Arxiv

    Python
    View on GitHub↗806
  • gen-verse/reasonfluxGen-Verse avatar

    Gen-Verse/ReasonFlux

    538View on GitHub↗

    Princeton University \& PKU \& UIUC \& University of Chicago \& ByteDance Seed

    Python
    View on GitHub↗538
  • gximinglu/quarkG

    gximinglu/quark

    0View on GitHub↗
    View on GitHub↗0
  • halfrot/alarmH

    halfrot/ALaRM

    0View on GitHub↗
    View on GitHub↗0
  • haoxiang-wang/directional-preference-alignmentH

    Haoxiang-Wang/directional-preference-alignment

    0View on GitHub↗
    View on GitHub↗0
  • jaearly/mil-for-non-markovian-reward-modellingJ

    JAEarly/MIL-for-Non-Markovian-Reward-Modelling

    0View on GitHub↗
    View on GitHub↗0
  • jhejna/few-shot-preference-rlJ

    jhejna/few-shot-preference-rl

    0View on GitHub↗
    View on GitHub↗0
  • jhejna/inverse-preference-learningJ

    jhejna/inverse-preference-learning

    0View on GitHub↗
    View on GitHub↗0
  • kwai-klear/klearreasonerKwai-Klear avatar

    Kwai-Klear/KlearReasoner

    82View on GitHub↗

    December 5, 2025 🔍 We propose entropy ratio clipping​ (ERC) to impose a global constraint on the output distribution of the policy model. Experiments demonstrate that ERC can significantly improve the stability of off-policy training. 📄 The paper is available on arXiv.

    Python
    View on GitHub↗82
  • kwai-yuanqi/mm-rlhfKwai-YuanQi avatar

    Kwai-YuanQi/MM-RLHF

    200View on GitHub↗

    The Next Step Forward in Multimodal LLM Alignment

    Python
    View on GitHub↗200
  • langfengq/verl-agentlangfengQ avatar

    langfengQ/verl-agent

    1,548View on GitHub↗
    Pythonagent-frameworkdeepseek-r1gigpo
    View on GitHub↗1,548
  • linear95/apoL

    Linear95/APO

    0View on GitHub↗
    View on GitHub↗0
  • alibabaresearch/damo-convaiAlibabaResearch avatar

    AlibabaResearch/DAMO-ConvAI

    1,561View on GitHub↗

    DAMO-ConvAI: The official repository which contains the codebase for Alibaba DAMO Conversational AI.

    Pythonconversational-aideep-learningdialog
    View on GitHub↗1,561