awesome-repositories.comالتصنيفاتالمدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعخادم MCPحولكيفية ترتيب النتائجالصحافة
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to nivwusquorum/tensorflow-deepq

Open-source alternatives to Tensorflow Deepq

30 open-source projects similar to nivwusquorum/tensorflow-deepq, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best Tensorflow Deepq alternative.

  • yandexdataschool/agentnetالصورة الرمزية لـ yandexdataschool

    yandexdataschool/AgentNet

    299عرض على GitHub↗

    Deep Reinforcement Learning library for humans

    Python
    عرض على GitHub↗299
  • kaixhin/atariالصورة الرمزية لـ Kaixhin

    Kaixhin/Atari

    264عرض على GitHub↗

    Persistent advantage learning dueling double DQN for the Arcade Learning Environment

    Lua
    عرض على GitHub↗264
  • shangtongzhang/deeprlالصورة الرمزية لـ ShangtongZhang

    ShangtongZhang/DeepRL

    3,428عرض على GitHub↗

    Modularized Implementation of Deep RL Algorithms in PyTorch

    Python
    عرض على GitHub↗3,428
  • openai/baselinesالصورة الرمزية لـ openai

    openai/baselines

    16,733عرض على GitHub↗

    Baselines is a comprehensive suite of frameworks for reinforcement learning algorithm implementation, imitation learning, and training orchestration. It provides a library of standardized learning algorithms used to benchmark and replicate research results, alongside a deep learning policy framework for constructing neural network architectures such as multi-layer perceptrons, convolutional networks, and long short-term memory networks. The project includes a specialized imitation learning toolkit that enables agents to mimic expert behavior through behavior cloning and generative adversarial

    Python
    عرض على GitHub↗16,733
  • resibots/blackdropsالصورة الرمزية لـ resibots

    resibots/blackdrops

    66عرض على GitHub↗

    Code for the Black-DROPS algorithm: "Black-Box Data-efficient Policy Search for Robotics", IROS 2017/ICRA 2018

    C++
    عرض على GitHub↗66

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Find more with AI search
  • rlcode/reinforcement-learningالصورة الرمزية لـ rlcode

    rlcode/reinforcement-learning

    3,642عرض على GitHub↗

    Minimal and Clean Reinforcement Learning Examples

    Python
    عرض على GitHub↗3,642
  • instadeepai/jumanjiالصورة الرمزية لـ instadeepai

    instadeepai/jumanji

    841عرض على GitHub↗

    🕹️ A diverse suite of scalable reinforcement learning environments in JAX

    Python
    عرض على GitHub↗841
  • chainer/chainerrlالصورة الرمزية لـ chainer

    chainer/chainerrl

    1,200عرض على GitHub↗

    ChainerRL is a deep reinforcement learning library built on top of Chainer.

    Python
    عرض على GitHub↗1,200
  • facebookresearch/habitat-labالصورة الرمزية لـ facebookresearch

    facebookresearch/habitat-lab

    2,848عرض على GitHub↗

    Habitat-Lab is an open-source platform for training and evaluating embodied AI agents in photorealistic 3D indoor environments. It functions as a high-performance 3D indoor environment simulator that supports physics-based interaction, enabling research into navigation and manipulation tasks. The platform provides a modular task-environment abstraction that separates task logic from environment simulation, using configuration-driven pipeline assembly to compose simulation and training pipelines. It includes a hierarchical sensor-actuator architecture for mixing and matching perception and act

    Pythonaicomputer-visiondeep-learning
    عرض على GitHub↗2,848
  • google/dopamineالصورة الرمزية لـ google

    google/dopamine

    10,879عرض على GitHub↗

    Dopamine is a reinforcement learning research framework designed for prototyping and testing algorithms across diverse simulated environments. It provides an agent development toolkit that utilizes a flat class hierarchy to facilitate the creation and extension of learning agents. The framework includes a standardization layer via environment wrappers that connect agents to various physics simulations and gaming environments. It also features a high-performance experience replay buffer for storing and sampling transition data to improve training stability, alongside a dedicated hyperparameter

    Jupyter Notebook
    عرض على GitHub↗10,879
  • ju-jl/reinforcementlearninganintroduction.jlالصورة الرمزية لـ Ju-jl

    Ju-jl/ReinforcementLearningAnIntroduction.jl

    332عرض على GitHub↗

    Julia code for the book Reinforcement Learning An Introduction

    Julia
    عرض على GitHub↗332
  • langfengq/verl-agentالصورة الرمزية لـ langfengQ

    langfengQ/verl-agent

    1,548عرض على GitHub↗
    Pythonagent-frameworkdeepseek-r1gigpo
    عرض على GitHub↗1,548
  • lywangpx/reinforcement-learning-2nd-edition-by-sutton-exercise-solutionsالصورة الرمزية لـ LyWangPX

    LyWangPX/Reinforcement-Learning-2nd-Edition-by-Sutton-Exercise-Solutions

    2,417عرض على GitHub↗

    Solutions of Reinforcement Learning, An Introduction

    Jupyter Notebook
    عرض على GitHub↗2,417
  • maitrix-org/llm-reasonersالصورة الرمزية لـ maitrix-org

    maitrix-org/llm-reasoners

    2,345عرض على GitHub↗

    LLM Reasoners is a library to enable LLMs to conduct complex reasoning, with advanced reasoning algorithms. It approaches multi-step reasoning as planning and searches for the optimal reasoning chain, which achieves the best balance of exploration vs exploitation with the idea of "World Model"…

    Python
    عرض على GitHub↗2,345
  • microsoft/agent-lightningالصورة الرمزية لـ microsoft

    microsoft/agent-lightning

    15,047عرض على GitHub↗

    Agent Lightning is an optimization framework designed to refine the performance of individual AI agents within complex multi-agent systems. It provides a platform for improving decision-making and task execution by applying reinforcement learning, supervised fine-tuning, and automated prompt optimization. The framework distinguishes itself through its ability to isolate specific agents for targeted tuning, allowing developers to enhance individual behaviors while maintaining the stability of the broader system architecture. By utilizing a modular interface, it integrates with diverse agent fr

    Pythonagentagentic-aillm
    عرض على GitHub↗15,047
  • modalminds/mm-eurekaالصورة الرمزية لـ ModalMinds

    ModalMinds/MM-EUREKA

    771عرض على GitHub↗

    MM-EUREKA: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

    Python
    عرض على GitHub↗771
  • novasky-ai/skyrlالصورة الرمزية لـ NovaSky-AI

    NovaSky-AI/SkyRL

    1,611عرض على GitHub↗
    Python
    عرض على GitHub↗1,611
  • nvidia-nemo/rlالصورة الرمزية لـ NVIDIA-NeMo

    NVIDIA-NeMo/RL

    1,756عرض على GitHub↗

    Documentation | Discussions | Contributing

    Python
    عرض على GitHub↗1,756
  • om-ai-lab/vlm-r1الصورة الرمزية لـ om-ai-lab

    om-ai-lab/VLM-R1

    5,991عرض على GitHub↗

    VLM-R1 is a reasoning vision-language model and embodied AI framework designed to map visual inputs and language instructions into physical navigation waypoints and robotic actions. It functions as a multimodal policy optimizer and an open vocabulary detector capable of locating objects based on arbitrary natural language descriptions. The system distinguishes itself through the use of chain-of-thought reasoning and reinforcement learning to solve complex visual and spatial tasks. It utilizes a video semantic memory system, which employs a visual cache to maintain a history of live video for

    Python
    عرض على GitHub↗5,991
  • open-reasoner-zero/open-reasoner-zeroالصورة الرمزية لـ Open-Reasoner-Zero

    Open-Reasoner-Zero/Open-Reasoner-Zero

    2,095عرض على GitHub↗

    An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

    Python
    عرض على GitHub↗2,095
  • openrlhf/openrlhfالصورة الرمزية لـ OpenRLHF

    OpenRLHF/OpenRLHF

    9,675عرض على GitHub↗

    OpenRLHF is a training framework and alignment library designed for reinforcement learning from human feedback across distributed GPU clusters. It provides tools for aligning large language models and multimodal vision-language models using algorithms such as PPO, GRPO, and DPO. The framework distinguishes itself through a distributed inference engine that overlaps sample rollout with training to increase throughput. It supports scaling to models exceeding 70 billion parameters via parameter sharding and handles long-context sequences through ring-attention sequence parallelism. The project

    Pythonlarge-language-modelsopenai-o1proximal-policy-optimization
    عرض على GitHub↗9,675
  • agentica-project/rllmالصورة الرمزية لـ agentica-project

    agentica-project/rllm

    400عرض على GitHub↗

    🚀 Reinforcement Learning for Language Agents🌟

    Jupyter Notebook
    عرض على GitHub↗400
  • rlinf/rlinfالصورة الرمزية لـ RLinf

    RLinf/RLinf

    2,502عرض على GitHub↗

    RLinf is a distributed reinforcement learning orchestrator and embodied AI training framework. It provides the infrastructure to train vision-language-action models and robotic policies using a combination of reinforcement learning and supervised fine-tuning. The system is designed for scaling workloads across GPU clusters, managing the placement of actors, rollout workers, and environment components. It features a specialized robotics data collection pipeline for gathering teleoperated demonstrations and simulation trajectories into standardized replay buffers, alongside a hardware interface

    Pythonagentic-aiembodied-aireinforcement-learning
    عرض على GitHub↗2,502
  • rushter/mlalgorithmsالصورة الرمزية لـ rushter

    rushter/MLAlgorithms

    10,983عرض على GitHub↗

    MLAlgorithms is an educational machine learning algorithm library consisting of core predictive models implemented from scratch in Python. It serves as a reference for developers to study the internal logic and mathematical workings of these models through clean, minimal implementations. The codebase focuses on the study of algorithm implementation and machine learning education, providing a way to understand internal mechanics by building components without relying on heavy external libraries. The project utilizes object-oriented encapsulation and NumPy-based vectorization to manage model s

    Python
    عرض على GitHub↗10,983
  • sail-sg/understand-r1-zeroالصورة الرمزية لـ sail-sg

    sail-sg/understand-r1-zero

    1,214عرض على GitHub↗
    Pythonllmr1-zeroreasoning
    عرض على GitHub↗1,214
  • shangtongzhang/reinforcement-learning-an-introductionالصورة الرمزية لـ ShangtongZhang

    ShangtongZhang/reinforcement-learning-an-introduction

    14,569عرض على GitHub↗

    This project is a Python-based educational framework designed to simulate reinforcement learning algorithms and environments. It serves as a platform for reproducing classic textbook examples, allowing users to study agent behavior, policy improvement, and the fundamental mechanics of decision-making in controlled settings. The library provides implementations for core reinforcement learning concepts, including temporal difference learning, Monte Carlo episode sampling, and tabular value function approximation. It enables the analysis of specific algorithmic behaviors, such as identifying and

    Pythonartificial-intelligencereinforcement-learning
    عرض على GitHub↗14,569
  • simple-efficient/rl-factoryالصورة الرمزية لـ Simple-Efficient

    Simple-Efficient/RL-Factory

    1,768عرض على GitHub↗

    📘Tutorial | 🛠️Installation | 🎨Framework

    Python
    عرض على GitHub↗1,768
  • thudm/slimeالصورة الرمزية لـ THUDM

    THUDM/slime

    4,259عرض على GitHub↗

    SLIME is a distributed reinforcement learning framework for large language model post-training that bridges Megatron training with SGLang inference servers. It orchestrates scalable RL loops across GPU clusters, decoupling training and inference into independent processes that communicate over HTTP and NCCL for independent scaling and fault tolerance. The system supports multi-agent reinforcement learning workflows with parallel agent instances, customizable rollout strategies, and personalized agent serving that improves models from prior conversations without disrupting API serving. The fra

    Python
    عرض على GitHub↗4,259
  • tidedra/lmm-r1الصورة الرمزية لـ TideDra

    TideDra/lmm-r1

    846عرض على GitHub↗

    Extend OpenRLHF to support LMM RL training for reproduction of DeepSeek-R1 on multimodal tasks.

    Python
    عرض على GitHub↗846
  • tiger-ai-lab/verl-toolالصورة الرمزية لـ TIGER-AI-Lab

    TIGER-AI-Lab/verl-tool

    1,006عرض على GitHub↗

    A version of verl to support diverse tool use TMLR 2026

    Python
    عرض على GitHub↗1,006