30 open-source projects similar to openai/random-network-distillation, ranked by shared indexed features. Tags may describe platforms or build tools rather than the same primary purpose. Check each project’s use case, license, and deployment requirements before treating it as a replacement.
PyTorch implementation of Neural Combinatorial Optimization with Reinforcement Learning https://arxiv.org/abs/1611.09940
Implementation of algorithms for continuous control (DDPG and NAF).
Modular Deep Reinforcement Learning framework in PyTorch. Companion library of the book "Foundations of Deep Reinforcement Learning".
Welcome to drlzh.ai: a hands-on deep reinforcement learning course where you build the algorithms, not just read about them.
FinRL is a reinforcement learning framework designed for the development, training, and backtesting of automated trading strategies. It functions as a quantitative finance toolkit that integrates deep learning algorithms with financial market simulations to address complex portfolio management and asset allocation tasks. The platform provides an end-to-end pipeline for transforming raw market data into actionable trading models. The project distinguishes itself through a layered, modular architecture that separates data processing, environment simulation, and agent training. This design allow
This repository contains the implementation of DISCERN in Python. You can download the manuscript from my website or arXiv.
⚡️A Blazing-Fast Python Library for Ranking Evaluation, Comparison, and Fusion 🐍
This repository provides supplementary material for our paper Constitutional AI: Harmlessness from AI Feedback.
A replica of the AlphaZero methodology for deep reinforcement learning in Python
DAMO-ConvAI: The official repository which contains the codebase for Alibaba DAMO Conversational AI.
ROLL is a distributed reinforcement learning framework and model alignment toolkit designed for large language models. It serves as a scalable training pipeline and GPU cluster manager, providing the infrastructure to align model behavior using reinforcement learning algorithms and preference optimization techniques. The project distinguishes itself through an agentic rollout orchestrator that generates and collects multi-turn interaction trajectories between AI agents and simulated environments. It supports specialized alignment methods including Direct Preference Optimization, reinforcement
A set of Deep Reinforcement Learning Agents implemented in Tensorflow.
TensorFlow implementation of Deep Reinforcement Learning papers
A library with extensible implementations of DPO, KTO, PPO, ORPO, and other human-aware loss functions (HALOs).
Tensorflow Keras OpenAI Gym implementation of 1-step Q Learning from "Asynchronous Methods for Deep Reinforcement Learning"
Gymnasium-based benchmarking suite for testing RL algorithms on real-world scenarios
A flexible and efficient training framework for large-scale alignment tasks