awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoServidor MCPAcerca deCómo clasificamosPrensa
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

12 repositorios

Awesome GitHub RepositoriesCritic-Free Algorithms

Policy optimization methods that do not rely on explicit value function modeling.

Explore 12 awesome GitHub repositories matching part of an awesome list · Critic-Free Algorithms. Refine with filters or upvote what's useful.

Awesome Critic-Free Algorithms GitHub Repositories

Encuentra los mejores repositorios con IA.Buscaremos los repositorios que mejor coincidan usando IA.
  • openrlhf/openrlhfAvatar de OpenRLHF

    OpenRLHF/OpenRLHF

    9,675Ver en GitHub↗

    OpenRLHF is a training framework and alignment library designed for reinforcement learning from human feedback across distributed GPU clusters. It provides tools for aligning large language models and multimodal vision-language models using algorithms such as PPO, GRPO, and DPO. The framework distinguishes itself through a distributed inference engine that overlaps sample rollout with training to increase throughput. It supports scaling to models exceeding 70 billion parameters via parameter sharding and handles long-context sequences through ring-attention sequence parallelism. The project

    Robust reinforcement learning algorithm for human feedback alignment.

    Pythonlarge-language-modelsopenai-o1proximal-policy-optimization
    Ver en GitHub↗9,675
  • deepseek-ai/deepseek-mathAvatar de deepseek-ai

    deepseek-ai/DeepSeek-Math

    3,346Ver en GitHub↗

    DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

    Policy optimization for mathematical reasoning in open models.

    Python
    Ver en GitHub↗3,346
  • minimax-ai/minimax-m1Avatar de MiniMax-AI

    MiniMax-AI/MiniMax-M1

    3,159Ver en GitHub↗

    MiniMax-M1, the world's first open-weight, large-scale hybrid-attention reasoning model.

    Scaling test-time compute using efficient attention mechanisms.

    Python
    Ver en GitHub↗3,159
  • bytedtsinghua-sia/dapoAvatar de BytedTsinghua-SIA

    BytedTsinghua-SIA/DAPO

    1,831Ver en GitHub↗

    DAPO: an Open-source RL System from ByteDance Seed and Tsinghua AIR

    Large-scale open-source reinforcement learning system for LLMs.

    Python
    Ver en GitHub↗1,831
  • sail-sg/understand-r1-zeroAvatar de sail-sg

    sail-sg/understand-r1-zero

    1,214Ver en GitHub↗

    Critical analysis and implementation of reasoning-focused training.

    Pythonllmr1-zeroreasoning
    Ver en GitHub↗1,214
  • modalminds/mm-eurekaAvatar de ModalMinds

    ModalMinds/MM-EUREKA

    771Ver en GitHub↗

    MM-EUREKA: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

    Stable rule-based reinforcement learning for language models.

    Python
    Ver en GitHub↗771
  • prime-rl/entropy-mechanism-of-rlAvatar de PRIME-RL

    PRIME-RL/Entropy-Mechanism-of-RL

    443Ver en GitHub↗

    )](https://www.alphaxiv.org/abs/2505.22617)

    Entropy-based mechanisms for reasoning model reinforcement learning.

    Python
    Ver en GitHub↗443
  • yaof20/flash-rlAvatar de yaof20

    yaof20/Flash-RL

    305Ver en GitHub↗

    Fast RL training with Quantized Rollouts ( Blog )

    Accelerated reinforcement learning training using quantized rollouts.

    Python
    Ver en GitHub↗305
  • tsinghuac3i/unify-post-trainingAvatar de TsinghuaC3I

    TsinghuaC3I/Unify-Post-Training

    211Ver en GitHub↗

    📝 Unified Policy Gradient Estimator • ✨ Hybrid Post-Training 🚀 Getting Started • 📊 Main Results • 💖 Acknowledgements • 📨 Contact • 🎈 Citation

    Unified post-training frameworks for large language models.

    Python
    Ver en GitHub↗211
  • liziniu/remaxAvatar de liziniu

    liziniu/ReMax

    202Ver en GitHub↗

    ReMax is a reinforcement learning method, tailored for reward maximization in RLHF.

    Simple and efficient alignment method for large language models.

    Python
    Ver en GitHub↗202
  • mcgill-nlp/vineppoAvatar de McGill-NLP

    McGill-NLP/VinePPO

    192Ver en GitHub↗

    Paper - Abstract - Updates - Quick Start - Installation - Download the datasets - Create Experiment Script - Single GPU Training (Only for Rho models) - Running the experiments - Code Structure - Initial SFT Checkpoints - Acknowledgement - Citation

    Refined credit assignment for unlocking reasoning potential.

    Python
    Ver en GitHub↗192
  • yihedeng9/openvlthinkerAvatar de yihedeng9

    yihedeng9/OpenVLThinker

    152Ver en GitHub↗

    Yihe Deng , Nanyun Peng , Kai-Wei Chang

    Iterative SFT-RL cycles for complex vision-language reasoning.

    Python
    Ver en GitHub↗152
  1. Home
  2. Part of an Awesome List
  3. AI & Machine Learning
  4. Critic-Free Algorithms