awesome-repositories.com
Blog
MCP
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektMCP-ServerÜber unsRanking-MethodikPresse
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

12 Repos

Awesome GitHub RepositoriesCritic-Free Algorithms

Policy optimization methods that do not rely on explicit value function modeling.

Explore 12 awesome GitHub repositories matching part of an awesome list · Critic-Free Algorithms. Refine with filters or upvote what's useful.

Awesome Critic-Free Algorithms GitHub Repositories

Finde die besten Repos mit KI.Wir suchen mit KI nach den am besten passenden Repositories.
  • openrlhf/openrlhfAvatar von OpenRLHF

    OpenRLHF/OpenRLHF

    9,675Auf GitHub ansehen↗

    OpenRLHF is a training framework and alignment library designed for reinforcement learning from human feedback across distributed GPU clusters. It provides tools for aligning large language models and multimodal vision-language models using algorithms such as PPO, GRPO, and DPO. The framework distinguishes itself through a distributed inference engine that overlaps sample rollout with training to increase throughput. It supports scaling to models exceeding 70 billion parameters via parameter sharding and handles long-context sequences through ring-attention sequence parallelism. The project

    Robust reinforcement learning algorithm for human feedback alignment.

    Pythonlarge-language-modelsopenai-o1proximal-policy-optimization
    Auf GitHub ansehen↗9,675
  • deepseek-ai/deepseek-mathAvatar von deepseek-ai

    deepseek-ai/DeepSeek-Math

    3,346Auf GitHub ansehen↗

    DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

    Policy optimization for mathematical reasoning in open models.

    Python
    Auf GitHub ansehen↗3,346
  • minimax-ai/minimax-m1Avatar von MiniMax-AI

    MiniMax-AI/MiniMax-M1

    3,159Auf GitHub ansehen↗

    MiniMax-M1, the world's first open-weight, large-scale hybrid-attention reasoning model.

    Scaling test-time compute using efficient attention mechanisms.

    Python
    Auf GitHub ansehen↗3,159
  • bytedtsinghua-sia/dapoAvatar von BytedTsinghua-SIA

    BytedTsinghua-SIA/DAPO

    1,831Auf GitHub ansehen↗

    DAPO: an Open-source RL System from ByteDance Seed and Tsinghua AIR

    Large-scale open-source reinforcement learning system for LLMs.

    Python
    Auf GitHub ansehen↗1,831
  • sail-sg/understand-r1-zeroAvatar von sail-sg

    sail-sg/understand-r1-zero

    1,214Auf GitHub ansehen↗

    Critical analysis and implementation of reasoning-focused training.

    Pythonllmr1-zeroreasoning
    Auf GitHub ansehen↗1,214
  • modalminds/mm-eurekaAvatar von ModalMinds

    ModalMinds/MM-EUREKA

    771Auf GitHub ansehen↗

    MM-EUREKA: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

    Stable rule-based reinforcement learning for language models.

    Python
    Auf GitHub ansehen↗771
  • prime-rl/entropy-mechanism-of-rlAvatar von PRIME-RL

    PRIME-RL/Entropy-Mechanism-of-RL

    443Auf GitHub ansehen↗

    )](https://www.alphaxiv.org/abs/2505.22617)

    Entropy-based mechanisms for reasoning model reinforcement learning.

    Python
    Auf GitHub ansehen↗443
  • yaof20/flash-rlAvatar von yaof20

    yaof20/Flash-RL

    305Auf GitHub ansehen↗

    Fast RL training with Quantized Rollouts ( Blog )

    Accelerated reinforcement learning training using quantized rollouts.

    Python
    Auf GitHub ansehen↗305
  • tsinghuac3i/unify-post-trainingAvatar von TsinghuaC3I

    TsinghuaC3I/Unify-Post-Training

    211Auf GitHub ansehen↗

    📝 Unified Policy Gradient Estimator • ✨ Hybrid Post-Training 🚀 Getting Started • 📊 Main Results • 💖 Acknowledgements • 📨 Contact • 🎈 Citation

    Unified post-training frameworks for large language models.

    Python
    Auf GitHub ansehen↗211
  • liziniu/remaxAvatar von liziniu

    liziniu/ReMax

    202Auf GitHub ansehen↗

    ReMax is a reinforcement learning method, tailored for reward maximization in RLHF.

    Simple and efficient alignment method for large language models.

    Python
    Auf GitHub ansehen↗202
  • mcgill-nlp/vineppoAvatar von McGill-NLP

    McGill-NLP/VinePPO

    192Auf GitHub ansehen↗

    Paper - Abstract - Updates - Quick Start - Installation - Download the datasets - Create Experiment Script - Single GPU Training (Only for Rho models) - Running the experiments - Code Structure - Initial SFT Checkpoints - Acknowledgement - Citation

    Refined credit assignment for unlocking reasoning potential.

    Python
    Auf GitHub ansehen↗192
  • yihedeng9/openvlthinkerAvatar von yihedeng9

    yihedeng9/OpenVLThinker

    152Auf GitHub ansehen↗

    Yihe Deng , Nanyun Peng , Kai-Wei Chang

    Iterative SFT-RL cycles for complex vision-language reasoning.

    Python
    Auf GitHub ansehen↗152
  1. Home
  2. Part of an Awesome List
  3. AI & Machine Learning
  4. Critic-Free Algorithms