TY - JOUR
T1 - Parallel Multi-Environment Shaping Algorithm for Complex Multi-step Task
AU - Ma, Cong
AU - Li, Zhizhong
AU - Lin, Dahua
AU - Zhang, Jiangshe
N1 - Publisher Copyright:
© 2020 Elsevier B.V.
PY - 2020/8/18
Y1 - 2020/8/18
N2 - Because of the sparse reward and the sequence of the complex multi-step task, there is a big challenge in reinforcement learning, that is, the agent needs to implement several consecutive sequential steps to complete the whole task without the intermediate reward. Reward Shaping and Curriculum Learning algorithms are always used to solve this challenge, but Reward Shaping is prone to sub-optimal policies and Curriculum Learning easily suffers from the catastrophic forgetting problem. In this paper, we propose a novel algorithm called Parallel Multi-Environment Shaping (PMES), where several sub-environments are built based on human knowledge to make the agent aware of the importance of intermediate steps, each of which corresponds to a key intermediate step. Specifically, the learning agent is trained under these parallel multiple environments including the original environment and several sub-environments by synchronous advantage actor-critic algorithm. And the PMES algorithm has the mechanism of adaptive reward shaping to adjust the reward function. In this way, PMES algorithm effectively incorporates human experience by multiple different environments rather than only shaping the reward function, which combines the benefits of Reward Shaping and Curriculum Learning algorithms while avoiding their drawbacks. Extensive experiments on the mini-game ‘Build Marines’ of StarCraft II environment show that our proposed algorithm is more effective than Reward Shaping, Curriculum Learning, and PLAID algorithms, which is almost close to the level of human Grandmaster. And compared with the existing work, it takes less time and computing resources to reach a good result.
AB - Because of the sparse reward and the sequence of the complex multi-step task, there is a big challenge in reinforcement learning, that is, the agent needs to implement several consecutive sequential steps to complete the whole task without the intermediate reward. Reward Shaping and Curriculum Learning algorithms are always used to solve this challenge, but Reward Shaping is prone to sub-optimal policies and Curriculum Learning easily suffers from the catastrophic forgetting problem. In this paper, we propose a novel algorithm called Parallel Multi-Environment Shaping (PMES), where several sub-environments are built based on human knowledge to make the agent aware of the importance of intermediate steps, each of which corresponds to a key intermediate step. Specifically, the learning agent is trained under these parallel multiple environments including the original environment and several sub-environments by synchronous advantage actor-critic algorithm. And the PMES algorithm has the mechanism of adaptive reward shaping to adjust the reward function. In this way, PMES algorithm effectively incorporates human experience by multiple different environments rather than only shaping the reward function, which combines the benefits of Reward Shaping and Curriculum Learning algorithms while avoiding their drawbacks. Extensive experiments on the mini-game ‘Build Marines’ of StarCraft II environment show that our proposed algorithm is more effective than Reward Shaping, Curriculum Learning, and PLAID algorithms, which is almost close to the level of human Grandmaster. And compared with the existing work, it takes less time and computing resources to reach a good result.
KW - Adaptive Reward Shaping
KW - Multi-step Task
KW - Parallel Multiple Environments
KW - Reinforcement Learning
UR - https://www.scopus.com/pages/publications/85083647758
U2 - 10.1016/j.neucom.2020.04.070
DO - 10.1016/j.neucom.2020.04.070
M3 - 文章
AN - SCOPUS:85083647758
SN - 0925-2312
VL - 402
SP - 323
EP - 335
JO - Neurocomputing
JF - Neurocomputing
ER -