Random curiosity-driven exploration in deep reinforcement learning

Jing Li; Xinxin Shi; Jiehao Li; Xin Zhang; Junzheng Wang

doi:10.1016/j.neucom.2020.08.024

Random curiosity-driven exploration in deep reinforcement learning

Jing Li, Xinxin Shi, Jiehao Li^*, Xin Zhang, Junzheng Wang

^*Corresponding author for this work

School of Automation

Beijing Institute of Technology

Research output: Contribution to journal › Article › peer-review

56 Citations (Scopus)

Abstract

Reinforcement learning (RL) depends on carefully engineering environment rewards. However, rewards from environments are extremely sparse for many RL tasks, challenging for the agent to learn skills and interact with the environment. One solution to this problem is to create intrinsic rewards for agents and to make rewards dense and more suitable for learning. Recent algorithms, such as curiosity-driven exploration, usually estimate the novelty of the next state through the prediction error of dynamics models. However, these methods are typically limited by the capacity of their dynamics models. In this paper, a random curiosity-driven model using deep reinforcement learning is proposed, which uses a target network with fixed weights to maintain the stability of dynamics models and create more suitable intrinsic rewards. We integrate the parametric exploration method for further promoting sufficient exploration. Besides, a deeper and more closely connected network is utilized for encoding the pixel images for policy-gradient. By comparing our method against the previous approaches in several environments, the experiments show that our method achieves state-of-the-art performance on most but not all of the Atari games.

Original language	English
Pages (from-to)	139-147
Number of pages	9
Journal	Neurocomputing
Volume	418
DOIs	https://doi.org/10.1016/j.neucom.2020.08.024
Publication status	Published - 22 Dec 2020

Keywords

Curiosity-driven exploration
Deep reinforcement learning
Intrinsic rewards

Access to Document

10.1016/j.neucom.2020.08.024

Cite this

@article{f46c43a50b2743e0a83254de4b51d5b5,

title = "Random curiosity-driven exploration in deep reinforcement learning",

abstract = "Reinforcement learning (RL) depends on carefully engineering environment rewards. However, rewards from environments are extremely sparse for many RL tasks, challenging for the agent to learn skills and interact with the environment. One solution to this problem is to create intrinsic rewards for agents and to make rewards dense and more suitable for learning. Recent algorithms, such as curiosity-driven exploration, usually estimate the novelty of the next state through the prediction error of dynamics models. However, these methods are typically limited by the capacity of their dynamics models. In this paper, a random curiosity-driven model using deep reinforcement learning is proposed, which uses a target network with fixed weights to maintain the stability of dynamics models and create more suitable intrinsic rewards. We integrate the parametric exploration method for further promoting sufficient exploration. Besides, a deeper and more closely connected network is utilized for encoding the pixel images for policy-gradient. By comparing our method against the previous approaches in several environments, the experiments show that our method achieves state-of-the-art performance on most but not all of the Atari games.",

keywords = "Curiosity-driven exploration, Deep reinforcement learning, Intrinsic rewards",

author = "Jing Li and Xinxin Shi and Jiehao Li and Xin Zhang and Junzheng Wang",

note = "Publisher Copyright: {\textcopyright} 2020 Elsevier B.V.",

year = "2020",

month = dec,

day = "22",

doi = "10.1016/j.neucom.2020.08.024",

language = "English",

volume = "418",

pages = "139--147",

journal = "Neurocomputing",

issn = "0925-2312",

publisher = "Elsevier B.V.",

}

TY - JOUR

T1 - Random curiosity-driven exploration in deep reinforcement learning

AU - Li, Jing

AU - Shi, Xinxin

AU - Li, Jiehao

AU - Zhang, Xin

AU - Wang, Junzheng

PY - 2020/12/22

Y1 - 2020/12/22

N2 - Reinforcement learning (RL) depends on carefully engineering environment rewards. However, rewards from environments are extremely sparse for many RL tasks, challenging for the agent to learn skills and interact with the environment. One solution to this problem is to create intrinsic rewards for agents and to make rewards dense and more suitable for learning. Recent algorithms, such as curiosity-driven exploration, usually estimate the novelty of the next state through the prediction error of dynamics models. However, these methods are typically limited by the capacity of their dynamics models. In this paper, a random curiosity-driven model using deep reinforcement learning is proposed, which uses a target network with fixed weights to maintain the stability of dynamics models and create more suitable intrinsic rewards. We integrate the parametric exploration method for further promoting sufficient exploration. Besides, a deeper and more closely connected network is utilized for encoding the pixel images for policy-gradient. By comparing our method against the previous approaches in several environments, the experiments show that our method achieves state-of-the-art performance on most but not all of the Atari games.

AB - Reinforcement learning (RL) depends on carefully engineering environment rewards. However, rewards from environments are extremely sparse for many RL tasks, challenging for the agent to learn skills and interact with the environment. One solution to this problem is to create intrinsic rewards for agents and to make rewards dense and more suitable for learning. Recent algorithms, such as curiosity-driven exploration, usually estimate the novelty of the next state through the prediction error of dynamics models. However, these methods are typically limited by the capacity of their dynamics models. In this paper, a random curiosity-driven model using deep reinforcement learning is proposed, which uses a target network with fixed weights to maintain the stability of dynamics models and create more suitable intrinsic rewards. We integrate the parametric exploration method for further promoting sufficient exploration. Besides, a deeper and more closely connected network is utilized for encoding the pixel images for policy-gradient. By comparing our method against the previous approaches in several environments, the experiments show that our method achieves state-of-the-art performance on most but not all of the Atari games.

KW - Curiosity-driven exploration

KW - Deep reinforcement learning

KW - Intrinsic rewards

UR - http://www.scopus.com/inward/record.url?scp=85092115902&partnerID=8YFLogxK

U2 - 10.1016/j.neucom.2020.08.024

DO - 10.1016/j.neucom.2020.08.024

M3 - Article

AN - SCOPUS:85092115902

SN - 0925-2312

VL - 418

SP - 139

EP - 147

JO - Neurocomputing

JF - Neurocomputing

ER -

Random curiosity-driven exploration in deep reinforcement learning

Abstract

Keywords

Access to Document

Other files and links

Fingerprint

Cite this