awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
ElliottYan avatar

ElliottYan/LUFFY

0
View on GitHub↗
455 stars·66 forks·Python·10 viewsarxiv.org/pdf/2504.14945↗

LUFFY

LUFFY: Learning to Reason Under Off‑Policy Guidance A general framework for off-policy learning in large reasoning models.

Features

  • Off-Policy Optimization - Reasoning under off-policy guidance for improved performance.

Star history

Star history chart for elliottyan/luffyStar history chart for elliottyan/luffy

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Projects sharing features with LUFFY

These projects share indexed features with LUFFY. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • rlinf/rlinfRLinf avatar

    RLinf/RLinf

    2,502View on GitHub↗

    RLinf is a distributed reinforcement learning orchestrator and embodied AI training framework. It provides the infrastructure to train vision-language-action models and robotic policies using a combination of reinforcement learning and supervised fine-tuning. The system is designed for scaling workloads across GPU clusters, managing the placement of actors, rollout workers, and environment components. It features a specialized robotics data collection pipeline for gathering teleoperated demonstrations and simulation trajectories into standardized replay buffers, alongside a hardware interface

    Pythonagentic-aiembodied-aireinforcement-learning
    View on GitHub↗2,502
  • morvanzhou/pytorch-tutorialMorvanZhou avatar

    MorvanZhou/PyTorch-Tutorial

    8,458View on GitHub↗

    This project is a collection of PyTorch learning resources and educational guides designed to teach the construction and training of neural networks. It serves as a comprehensive deep learning tutorial covering various model architectures and practical implementation strategies. The resources provide specific guidance on implementing computer vision tasks, such as image classification and synthetic imagery generation, as well as reinforcement learning agents using value networks and experience replay. It also covers sequential data modeling through recurrent networks and generative modeling u

    Jupyter Notebookautoencoderbatchbatch-normalization
    View on GitHub↗8,458
  • chanliang/bridgeChanLiang avatar

    ChanLiang/BRIDGE

    6View on GitHub↗

    The code for BRIDGE (Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning).

    View on GitHub↗6
  • liumy2010/uftliumy2010 avatar

    liumy2010/UFT

    31View on GitHub↗

    Mingyang Liu, Gabriele Farina, Asuman Ozdaglar

    Python
    View on GitHub↗31
Compare all 12 related projects→

Frequently asked questions

What does elliottyan/luffy do?

LUFFY: Learning to Reason Under Off‑Policy Guidance A general framework for off-policy learning in large reasoning models.

What are the main features of elliottyan/luffy?

The main features of elliottyan/luffy are: Off-Policy Optimization.

Which projects share features with elliottyan/luffy?

Projects with overlapping indexed features include: morvanzhou/pytorch-tutorial — This project is a collection of PyTorch learning resources and educational guides designed to teach the construction… rlinf/rlinf — RLinf is a distributed reinforcement learning orchestrator and embodied AI training framework. It provides the… chanliang/bridge — The code for BRIDGE (Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning). liumy2010/uft — Mingyang Liu, Gabriele Farina, Asuman Ozdaglar. millioniron/openrlhf-millioniron- — Open-source / Comprehensive / Lightweight / Easy-to-use. mozerwang/ampo — Minzheng Wang 1,2 , Yongbin Li 3 , Haobo Wang 4 , Xinghua Zhang 3🌟 , Nan Xu 1 , Bingli Wu 3 , Fei Huang 3 , Haiyang…