Skip to main navigation Skip to search Skip to main content

Efficient Exploration for Multi-Agent Reinforcement Learning via Transferable Successor Features

  • Anhui University
  • Peng Cheng Laboratory
  • Southeast University, Nanjing

Research output: Contribution to journalArticlepeer-review

19 Scopus citations

Abstract

In multi-Agent reinforcement learning (MARL), the behaviors of each agent can influence the learning of others, and the agents have to search in an exponentially enlarged joint-Action space. Hence, it is challenging for the multi-Agent teams to explore in the environment. Agents may achieve suboptimal policies and fail to solve some complex tasks. To improve the exploring efficiency as well as the performance of MARL tasks, in this paper, we propose a new approach by transferring the knowledge across tasks. Differently from the traditional MARL algorithms, we first assume that the reward functions can be computed by linear combinations of a shared feature function and a set of task-specific weights. Then, we define a set of basic MARL tasks in the source domain and pre-Train them as the basic knowledge for further use. Finally, once the weights for target tasks are available, it will be easier to get a well-performed policy to explore in the target domain. Hence, the learning process of agents for target tasks is speeded up by taking full use of the basic knowledge that was learned previously. We evaluate the proposed algorithm on two challenging MARL tasks: cooperative box-pushing and non-monotonic predator-prey. The experiment results have demonstrated the improved performance compared with state-of-The-Art MARL algorithms.

Original languageEnglish
Pages (from-to)1673-1686
Number of pages14
JournalIEEE/CAA Journal of Automatica Sinica
Volume9
Issue number9
DOIs
StatePublished - 1 Sep 2022
Externally publishedYes

Keywords

  • Knowledge transfer
  • multi-Agent systems
  • reinforcement learning
  • successor features

Fingerprint

Dive into the research topics of 'Efficient Exploration for Multi-Agent Reinforcement Learning via Transferable Successor Features'. Together they form a unique fingerprint.

Cite this