PyTorch implementations of deep reinforcement learning algorithms and environments

Last update: Jan 04, 2023

Overview

Deep Reinforcement Learning Algorithms with PyTorch

This repository contains PyTorch implementations of deep reinforcement learning algorithms and environments.

(To help you remember things you learn about machine learning in general write them in Save All and try out the public deck there about Fast AI's machine learning textbook.)

Algorithms Implemented

Deep Q Learning (DQN) _{^{(Mnih et al. 2013)}}
DQN with Fixed Q Targets _{^{(Mnih et al. 2013)}}
Double DQN (DDQN) _{^{(Hado van Hasselt et al. 2015)}}
DDQN with Prioritised Experience Replay _{^{(Schaul et al. 2016)}}
Dueling DDQN _{^{(Wang et al. 2016)}}
REINFORCE _{^{(Williams et al. 1992)}}
Deep Deterministic Policy Gradients (DDPG) _{^{(Lillicrap et al. 2016 )}}
Twin Delayed Deep Deterministic Policy Gradients (TD3) _{^{(Fujimoto et al. 2018)}}
Soft Actor-Critic (SAC) _{^{(Haarnoja et al. 2018)}}
Soft Actor-Critic for Discrete Actions (SAC-Discrete) _{^{(Christodoulou 2019)}}
Asynchronous Advantage Actor Critic (A3C) _{^{(Mnih et al. 2016)}}
Syncrhonous Advantage Actor Critic (A2C)
Proximal Policy Optimisation (PPO) _{^{(Schulman et al. 2017)}}
DQN with Hindsight Experience Replay (DQN-HER) _{^{(Andrychowicz et al. 2018)}}
DDPG with Hindsight Experience Replay (DDPG-HER) _{^{(Andrychowicz et al. 2018 )}}
Hierarchical-DQN (h-DQN) _{^{(Kulkarni et al. 2016)}}
Stochastic NNs for Hierarchical Reinforcement Learning (SNN-HRL) _{^{(Florensa et al. 2017)}}
Diversity Is All You Need (DIAYN) _{^{(Eyensbach et al. 2018)}}

All implementations are able to quickly solve Cart Pole (discrete actions), Mountain Car Continuous (continuous actions), Bit Flipping (discrete actions with dynamic goals) or Fetch Reach (continuous actions with dynamic goals). I plan to add more hierarchical RL algorithms soon.

Environments Implemented

Bit Flipping Game _{^{(as described in Andrychowicz et al. 2018)}}
Four Rooms Game _{^{(as described in Sutton et al. 1998)}}
Long Corridor Game _{^{(as described in Kulkarni et al. 2016)}}
Ant-{Maze, Push, Fall} _{^{(as desribed in Nachum et al. 2018 and their accompanying code)}}

Results

1. Cart Pole and Mountain Car

Below shows various RL algorithms successfully learning discrete action game Cart Pole or continuous action game Mountain Car. The mean result from running the algorithms with 3 random seeds is shown with the shaded area representing plus and minus 1 standard deviation. Hyperparameters used can be found in files results/Cart_Pole.py and results/Mountain_Car.py.

2. Hindsight Experience Replay (HER) Experiements

Below shows the performance of DQN and DDPG with and without Hindsight Experience Replay (HER) in the Bit Flipping (14 bits) and Fetch Reach environments described in the papers Hindsight Experience Replay 2018 and Multi-Goal Reinforcement Learning 2018. The results replicate the results found in the papers and show how adding HER can allow an agent to solve problems that it otherwise would not be able to solve at all. Note that the same hyperparameters were used within each pair of agents and so the only difference between them was whether hindsight was used or not.

3. Hierarchical Reinforcement Learning Experiments

The results on the left below show the performance of DQN and the algorithm hierarchical-DQN from Kulkarni et al. 2016 on the Long Corridor environment also explained in Kulkarni et al. 2016. The environment requires the agent to go to the end of a corridor before coming back in order to receive a larger reward. This delayed gratification and the aliasing of states makes it a somewhat impossible game for DQN to learn but if we introduce a meta-controller (as in h-DQN) which directs a lower-level controller how to behave we are able to make more progress. This aligns with the results found in the paper.

The results on the right show the performance of DDQN and algorithm Stochastic NNs for Hierarchical Reinforcement Learning (SNN-HRL) from Florensa et al. 2017. DDQN is used as the comparison because the implementation of SSN-HRL uses 2 DDQN algorithms within it. Note that the first 300 episodes of training for SNN-HRL were used for pre-training which is why there is no reward for those episodes.

Usage

The repository's high-level structure is:

├── agents                    
    ├── actor_critic_agents   
    ├── DQN_agents         
    ├── policy_gradient_agents
    └── stochastic_policy_search_agents 
├── environments   
├── results             
    └── data_and_graphs        
├── tests
├── utilities             
    └── data structures

i) To watch the agents learn the above games

To watch all the different agents learn Cart Pole follow these steps:

git clone https://github.com/p-christ/Deep_RL_Implementations.git
cd Deep_RL_Implementations

conda create --name myenvname
y
conda activate myenvname

pip3 install -r requirements.txt

python results/Cart_Pole.py

For other games change the last line to one of the other files in the Results folder.

ii) To train the agents on another game

Most Open AI gym environments should work. All you would need to do is change the config.environment field (look at Results/Cart_Pole.py for an example of this).

You can also play with your own custom game if you create a separate class that inherits from gym.Env. See Environments/Four_Rooms_Environment.py for an example of a custom environment and then see the script Results/Four_Rooms.py to see how to have agents play the environment.

PyTorch implementations of deep reinforcement learning algorithms and environments

Related tags

Overview

Deep Reinforcement Learning Algorithms with PyTorch

Algorithms Implemented

Environments Implemented

Results

1. Cart Pole and Mountain Car

2. Hindsight Experience Replay (HER) Experiements

3. Hierarchical Reinforcement Learning Experiments

Usage

i) To watch the agents learn the above games

ii) To train the agents on another game

Owner

Petros Christodoulou

This project is a loose implementation of paper "Algorithmic Financial Trading with Deep Convolutional Neural Networks: Time Series to Image Conversion Approach"

Implementation of H-UCRL Algorithm

Human head pose estimation using Keras over TensorFlow.

ThunderGBM: Fast GBDTs and Random Forests on GPUs

Brain tumor detection using Convolution-Neural Network (CNN)

This is an unofficial PyTorch implementation of Meta Pseudo Labels

YOLOX is a high-performance anchor-free YOLO, exceeding yolov3~v5 with ONNX, TensorRT, ncnn, and OpenVINO supported.

This is a custom made virus code in python, using tkinter module.

TensorFlow implementation of "TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?"

This is the repository for the NeurIPS-21 paper [Contrastive Graph Poisson Networks: Semi-Supervised Learning with Extremely Limited Labels].

SigOpt wrappers for scikit-learn methods

Torch implementation of "Enhanced Deep Residual Networks for Single Image Super-Resolution"

Official implementation for “Unsupervised Low-Light Image Enhancement via Histogram Equalization Prior”

Using image super resolution models with vapoursynth and speeding them up with TensorRT

Code and data for ACL2021 paper Cross-Lingual Abstractive Summarization with Limited Parallel Resources.

Official implementation of CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image Classification

code for Grapadora research paper experimentation

[CVPR 2021] Counterfactual VQA: A Cause-Effect Look at Language Bias

A library for graph deep learning research

Qt-GUI implementation of the YOLOv5 algorithm (ver.6 and ver.5)