Skip to main navigation Skip to search Skip to main content

Periodic Guidance Learning

  • Xi'an Jiaotong University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Tasks with periodic states are widespread in reality. However, Current reinforcement learning (RL) algorithms generally treat such tasks as non-periodic Markov decision process, which results in low exploration efficiency and misleading advantage estimation with high variance. This paper proposes periodic guidance learning (PGL), in which a pruned advantage estimation with lower variance is implemented. Meanwhile, based on periodic states, past good experiences are utilized for better exploration. Our algorithm is evaluated on periodic tasks in MuJoCo. The experimental results show PGL method improves exploration efficiency and outperforms baselines in various periodic tasks. The results also show that PGL achieves a smooth policy optimization. Further experiments on the agent's periodic behavior reveal the strong correlation between period length and the agents motion mode.

Original languageEnglish
Title of host publicationProceedings - 11th IEEE International Conference on Knowledge Graph, ICKG 2020
EditorsEnhong Chen, Grigoris Antoniou, Xindong Wu, Vipin Kumar
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages77-83
Number of pages7
ISBN (Electronic)9781728181561
DOIs
StatePublished - Aug 2020
Event11th IEEE International Conference on Knowledge Graph, ICKG 2020 - Virtual, Online, China
Duration: 9 Aug 202011 Aug 2020

Publication series

NameProceedings - 11th IEEE International Conference on Knowledge Graph, ICKG 2020

Conference

Conference11th IEEE International Conference on Knowledge Graph, ICKG 2020
Country/TerritoryChina
CityVirtual, Online
Period9/08/2011/08/20

Keywords

  • Exploitation-exploration
  • Periodic tasks
  • Reinforcement learning

Fingerprint

Dive into the research topics of 'Periodic Guidance Learning'. Together they form a unique fingerprint.

Cite this