awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Back to agentica-project/rllm

Projects sharing features with Rllm

30 open-source projects similar to agentica-project/rllm, ranked by shared indexed features. Tags may describe platforms or build tools rather than the same primary purpose. Check each project’s use case, license, and deployment requirements before treating it as a replacement.

  • open-reasoner-zero/open-reasoner-zeroOpen-Reasoner-Zero avatar

    Open-Reasoner-Zero/Open-Reasoner-Zero

    2,095View on GitHub↗

    An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

    Python
    View on GitHub↗2,095
  • gair-nlp/limoGAIR-NLP avatar

    GAIR-NLP/LIMO

    1,077View on GitHub↗

    📄 Paper | 🌐 Dataset (v2) | 📘 Model (v2)

    Python
    View on GitHub↗1,077
  • modalminds/mm-eurekaModalMinds avatar

    ModalMinds/MM-EUREKA

    771View on GitHub↗

    MM-EUREKA: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

    Python
    View on GitHub↗771
  • huggingface/open-r1huggingface avatar

    huggingface/open-r1

    26,326View on GitHub↗

    Open-r1 is a framework designed for the large-scale training, distillation, and optimization of language models focused on complex reasoning and programming tasks. It provides a comprehensive suite of tools for managing distributed training jobs across multi-node clusters, enabling the development of high-performance models through reinforcement learning and supervised fine-tuning. The project distinguishes itself by integrating secure, containerized code execution environments directly into the training and evaluation lifecycle. By allowing models to run and verify code snippets against test

    Python
    View on GitHub↗26,326
  • inclusionai/arealinclusionAI avatar

    inclusionAI/AReaL

    3,559View on GitHub↗

    AReaL is a system for agent orchestration, distributed model training, and parameter-efficient tuning. It provides a framework for developing multi-turn reasoning agents and training large models using reinforcement learning from human feedback. The project implements a toolkit for improving the visual reasoning and geometry problem solving capabilities of vision-language models. It utilizes a memory-efficient tuning system to optimize mathematical and reasoning models across different inference backends. The infrastructure supports large-scale training through tensor, pipeline, and expert p

    Pythonagentllmllm-agent
    View on GitHub↗3,559

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Find more with AI search
  • petergriffinjin/search-r1PeterGriffinJin avatar

    PeterGriffinJin/Search-R1

    5,022View on GitHub↗

    Search-R1 is a distributed training system and reinforcement learning framework designed to create search-augmented language models. It provides an architecture for scaling model workloads across head and worker nodes while optimizing how models interleave internal reasoning with external tool calls. The system focuses on refining model behavior through custom reward signals and reinforcement learning to improve tool-use formatting and information retrieval. It implements an interleaved reasoning-search loop that allows models to alternate between internal thought generation and external data

    Python
    View on GitHub↗5,022
  • tidedra/lmm-r1TideDra avatar

    TideDra/lmm-r1

    846View on GitHub↗

    Extend OpenRLHF to support LMM RL training for reproduction of DeepSeek-R1 on multimodal tasks.

    Python
    View on GitHub↗846
  • jiayi-pan/tinyzeroJiayi-Pan avatar

    Jiayi-Pan/TinyZero

    13,168View on GitHub↗

    TinyZero is a reinforcement learning framework and implementation designed to train language models to develop reasoning and self-verification abilities. It provides a training pipeline to optimize model performance on mathematical and logical tasks. The project serves as a minimal reproduction of the DeepSeek R1 architectural and training approach. It focuses on creating reasoning models that can solve structured problems through autonomous chain-of-thought discovery. The framework incorporates group relative policy optimization and reward-based self-correction to improve accuracy on logica

    Python
    View on GitHub↗13,168
  • om-ai-lab/vlm-r1om-ai-lab avatar

    om-ai-lab/VLM-R1

    5,991View on GitHub↗

    VLM-R1 is a reasoning vision-language model and embodied AI framework designed to map visual inputs and language instructions into physical navigation waypoints and robotic actions. It functions as a multimodal policy optimizer and an open vocabulary detector capable of locating objects based on arbitrary natural language descriptions. The system distinguishes itself through the use of chain-of-thought reasoning and reinforcement learning to solve complex visual and spatial tasks. It utilizes a video semantic memory system, which employs a visual cache to maintain a history of live video for

    Python
    View on GitHub↗5,991
  • skyworkai/skywork-or1SkyworkAI avatar

    SkyworkAI/Skywork-OR1

    745View on GitHub↗

    ✊ Unleashing the Power of Reinforcement Learning for Math and Code Reasoners 🤖

    Python
    View on GitHub↗745
  • sail-sg/understand-r1-zerosail-sg avatar

    sail-sg/understand-r1-zero

    1,214View on GitHub↗
    Pythonllmr1-zeroreasoning
    View on GitHub↗1,214
  • unakar/logic-rlUnakar avatar

    Unakar/Logic-RL

    2,451View on GitHub↗

    Reproduce R1 Zero on Logic Puzzle

    Python
    View on GitHub↗2,451
  • hiyouga/easyr1hiyouga avatar

    hiyouga/EasyR1

    5,034View on GitHub↗

    EasyR1 is a distributed model training system and reinforcement learning framework for large language and vision-language models. It functions as a multimodal trainer and an implementation of a Proximal Policy Optimization pipeline designed to refine the reasoning and perception capabilities of models that process both text and images. The system specializes in distributing reinforcement learning workloads across multiple compute nodes to manage high memory requirements. It optimizes hardware utilization through padding-free training and fine-tuning to fit large models onto available graphics

    Python
    View on GitHub↗5,034
  • deep-agent/r1-vD

    Deep-Agent/R1-V

    0View on GitHub↗
    View on GitHub↗0
  • facebookresearch/habitat-labfacebookresearch avatar

    facebookresearch/habitat-lab

    2,848View on GitHub↗

    Habitat-Lab is an open-source platform for training and evaluating embodied AI agents in photorealistic 3D indoor environments. It functions as a high-performance 3D indoor environment simulator that supports physics-based interaction, enabling research into navigation and manipulation tasks. The platform provides a modular task-environment abstraction that separates task logic from environment simulation, using configuration-driven pipeline assembly to compose simulation and training pipelines. It includes a hierarchical sensor-actuator architecture for mixing and matching perception and act

    Pythonaicomputer-visiondeep-learning
    View on GitHub↗2,848
  • qwenlm/qwen2.5QwenLM avatar

    QwenLM/Qwen2.5

    27,307View on GitHub↗

    Qwen2.5 is a suite of large language model foundation models designed for natural language generation, code production, and complex mathematical reasoning. The project encompasses a multilingual language model capable of processing dozens of languages and a specialized code generation model for technical problem solving and debugging. The framework is distinguished by its long context capabilities, enabling the analysis of massive inputs ranging from 256K up to 1 million tokens. It further functions as an agentic framework, utilizing standardized templates and parsers to execute autonomous wo

    Python
    View on GitHub↗27,307
  • google/dopaminegoogle avatar

    google/dopamine

    10,879View on GitHub↗

    Dopamine is a reinforcement learning research framework designed for prototyping and testing algorithms across diverse simulated environments. It provides an agent development toolkit that utilizes a flat class hierarchy to facilitate the creation and extension of learning agents. The framework includes a standardization layer via environment wrappers that connect agents to various physics simulations and gaming environments. It also features a high-performance experience replay buffer for storing and sampling transition data to improve training stability, alongside a dedicated hyperparameter

    Jupyter Notebook
    View on GitHub↗10,879
  • brendanhogan/deepseekrl-extendedbrendanhogan avatar

    brendanhogan/DeepSeekRL-Extended

    252View on GitHub↗

    Exploring Applications of GRPO

    Python
    View on GitHub↗252
  • bklieger-groq/g1B

    bklieger-groq/g1

    0View on GitHub↗
    View on GitHub↗0
  • alibaba/rollalibaba avatar

    alibaba/ROLL

    2,844View on GitHub↗

    ROLL is a distributed reinforcement learning framework and model alignment toolkit designed for large language models. It serves as a scalable training pipeline and GPU cluster manager, providing the infrastructure to align model behavior using reinforcement learning algorithms and preference optimization techniques. The project distinguishes itself through an agentic rollout orchestrator that generates and collects multi-turn interaction trajectories between AI agents and simulated environments. It supports specialized alignment methods including Direct Preference Optimization, reinforcement

    Pythonagenticrlhfrlvr
    View on GitHub↗2,844
  • baichuan-inc/baichuan-m1-14bbaichuan-inc avatar

    baichuan-inc/Baichuan-M1-14B

    219View on GitHub↗

    Baichuan-M1-14B

    View on GitHub↗219
  • dvlab-research/seg-zerodvlab-research avatar

    dvlab-research/Seg-Zero

    632View on GitHub↗

    Project Page For "Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement"

    Python
    View on GitHub↗632
  • dhcode-cpp/x-r1D

    dhcode-cpp/X-R1

    0View on GitHub↗
    View on GitHub↗0
  • baibizhe/efficient-r1-vllmB

    baibizhe/Efficient-R1-VLLM

    0View on GitHub↗
    View on GitHub↗0
  • efficientscaling/z1efficientscaling avatar

    efficientscaling/Z1

    69View on GitHub↗

    Z1: Efficient Test-time Scaling with Code Train Large Language Model to Reason with Shifted Thinking

    Python
    View on GitHub↗69
  • evelinehong/3d-clr-officialevelinehong avatar

    evelinehong/3D-CLR-Official

    85View on GitHub↗

    Checkpoints take up a lot of space. Please email yninghong@gmail.com if you need them.

    Python
    View on GitHub↗85
  • evolvinglmms-lab/open-r1-multimodalEvolvingLMMs-Lab avatar

    EvolvingLMMs-Lab/open-r1-multimodal

    1,484View on GitHub↗
    Python
    View on GitHub↗1,484
  • alibaba-nlp/zerosearchAlibaba-NLP avatar

    Alibaba-NLP/ZeroSearch

    1,296View on GitHub↗

    ZeroSearch: Incentivize the Search Capability of LLMs without Searching

    Python
    View on GitHub↗1,296
  • facebookresearch/swe-rlfacebookresearch avatar

    facebookresearch/swe-rl

    704View on GitHub↗

    🧐 About | 🚀 Quick Start | 🐣 Agentless Mini | 📝 Citation | 🙏 Acknowledgements

    Python
    View on GitHub↗704
  • adam-bjtu/openrftA

    ADaM-BJTU/OpenRFT

    0View on GitHub↗
    View on GitHub↗0