TY - GEN
T1 - Reinforcement learning with evolutionary computation to policy search for autonomous navigation
AU - Zhang, Chengsi
AU - Dong, Lu
AU - Sun, Changyin
N1 - Publisher Copyright:
© 2020 IEEE.
PY - 2020/10/16
Y1 - 2020/10/16
N2 - Reinforcement learning has good applications for autonomous navigation in unknown and complex environments. Traditional reinforcement learning methods with the actor-critic framework sometimes will fall into a local optimum because of the complexity of the loss function. Meanwhile, evolutionary computation(EC) is a type of black box optimization algorithm, which has good robustness in policy search but lower sampling efficiency. In order to address the challenge, we introduce an algorithm that combines evolutionary computation with reinforcement learning into navigation intuitively. The parameters of actor neural network are listed as individual characteristics. Each individual represents a policy network. At the end of each episode, individuals with higher fitness function value are selected to the next generation. Other individuals update a certain number of steps through the critic network with shared replay buffer and then move into the next generation. Simulation results demonstrate the effectiveness and feasibility of this algorithm on navigation.
AB - Reinforcement learning has good applications for autonomous navigation in unknown and complex environments. Traditional reinforcement learning methods with the actor-critic framework sometimes will fall into a local optimum because of the complexity of the loss function. Meanwhile, evolutionary computation(EC) is a type of black box optimization algorithm, which has good robustness in policy search but lower sampling efficiency. In order to address the challenge, we introduce an algorithm that combines evolutionary computation with reinforcement learning into navigation intuitively. The parameters of actor neural network are listed as individual characteristics. Each individual represents a policy network. At the end of each episode, individuals with higher fitness function value are selected to the next generation. Other individuals update a certain number of steps through the critic network with shared replay buffer and then move into the next generation. Simulation results demonstrate the effectiveness and feasibility of this algorithm on navigation.
KW - Evolutionary computation
KW - Fitness function
KW - Navigation
KW - Reinforcement learning
UR - https://www.scopus.com/pages/publications/85101139767
U2 - 10.1109/YAC51587.2020.9337605
DO - 10.1109/YAC51587.2020.9337605
M3 - 会议稿件
AN - SCOPUS:85101139767
T3 - Proceedings - 2020 35th Youth Academic Annual Conference of Chinese Association of Automation, YAC 2020
SP - 288
EP - 292
BT - Proceedings - 2020 35th Youth Academic Annual Conference of Chinese Association of Automation, YAC 2020
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 35th Youth Academic Annual Conference of Chinese Association of Automation, YAC 2020
Y2 - 16 October 2020 through 18 October 2020
ER -