awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Back to yaof20/flash-rl

Projects sharing features with Flash RL

11 open-source projects similar to yaof20/flash-rl, ranked by shared indexed features. Tags may describe platforms or build tools rather than the same primary purpose. Check each project’s use case, license, and deployment requirements before treating it as a replacement.

  • bytedtsinghua-sia/dapoBytedTsinghua-SIA avatar

    BytedTsinghua-SIA/DAPO

    1,831View on GitHub↗

    DAPO: an Open-source RL System from ByteDance Seed and Tsinghua AIR

    Python
    View on GitHub↗1,831
  • deepseek-ai/deepseek-mathdeepseek-ai avatar

    deepseek-ai/DeepSeek-Math

    3,346View on GitHub↗

    DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

    Python
    View on GitHub↗3,346
  • liziniu/remaxliziniu avatar

    liziniu/ReMax

    202View on GitHub↗

    ReMax is a reinforcement learning method, tailored for reward maximization in RLHF.

    Python
    View on GitHub↗202
  • mcgill-nlp/vineppoMcGill-NLP avatar

    McGill-NLP/VinePPO

    192View on GitHub↗

    Paper - Abstract - Updates - Quick Start - Installation - Download the datasets - Create Experiment Script - Single GPU Training (Only for Rho models) - Running the experiments - Code Structure - Initial SFT Checkpoints - Acknowledgement - Citation

    Python
    View on GitHub↗192
  • minimax-ai/minimax-m1MiniMax-AI avatar

    MiniMax-AI/MiniMax-M1

    3,159View on GitHub↗

    MiniMax-M1, the world's first open-weight, large-scale hybrid-attention reasoning model.

    Python
    View on GitHub↗3,159
  • modalminds/mm-eurekaModalMinds avatar

    ModalMinds/MM-EUREKA

    771View on GitHub↗

    MM-EUREKA: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

    Python
    View on GitHub↗771

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Find more with AI search
  • openrlhf/openrlhfOpenRLHF avatar

    OpenRLHF/OpenRLHF

    9,675View on GitHub↗

    OpenRLHF is a training framework and alignment library designed for reinforcement learning from human feedback across distributed GPU clusters. It provides tools for aligning large language models and multimodal vision-language models using algorithms such as PPO, GRPO, and DPO. The framework distinguishes itself through a distributed inference engine that overlaps sample rollout with training to increase throughput. It supports scaling to models exceeding 70 billion parameters via parameter sharding and handles long-context sequences through ring-attention sequence parallelism. The project

    Pythonlarge-language-modelsopenai-o1proximal-policy-optimization
    View on GitHub↗9,675
  • prime-rl/entropy-mechanism-of-rlPRIME-RL avatar

    PRIME-RL/Entropy-Mechanism-of-RL

    443View on GitHub↗

    )](https://www.alphaxiv.org/abs/2505.22617)

    Python
    View on GitHub↗443
  • sail-sg/understand-r1-zerosail-sg avatar

    sail-sg/understand-r1-zero

    1,214View on GitHub↗
    Pythonllmr1-zeroreasoning
    View on GitHub↗1,214
  • tsinghuac3i/unify-post-trainingTsinghuaC3I avatar

    TsinghuaC3I/Unify-Post-Training

    211View on GitHub↗

    📝 Unified Policy Gradient Estimator • ✨ Hybrid Post-Training 🚀 Getting Started • 📊 Main Results • 💖 Acknowledgements • 📨 Contact • 🎈 Citation

    Python
    View on GitHub↗211
  • yihedeng9/openvlthinkeryihedeng9 avatar

    yihedeng9/OpenVLThinker

    152View on GitHub↗

    Yihe Deng , Nanyun Peng , Kai-Wei Chang

    Python
    View on GitHub↗152