TY - GEN
T1 - Deep Factorized Q-Learning for Large Scale Multi-Agent Learning
AU - Wang, Xiaoqiang
AU - Ke, Liangjun
N1 - Publisher Copyright:
© 2022 IEEE.
PY - 2022
Y1 - 2022
N2 - The value function decomposition is an effective way to alleviate the curse of dimension in Multi-Agent Reinforcement Learning (MARL). However, the existing methods usually either can only provide the low-order approximate decomposition of no more than the second-order, or need to spend a lot of effort to manually design the high-order interaction among agents according to experience. Therefore, the existing methods either tend to bear large decomposition error or are not convenient to use. In this paper, a high-order approximate value function decomposition method is proposed, which can be trained end-to-end. There have some prominent features about this method including low-rank vector exploited to represent value function, both low- and high-order component sharing the same input (i.e., the embedding vector), the model parameters shared among all the agents if they are homogeneous. Experimental results show that our method is effective.
AB - The value function decomposition is an effective way to alleviate the curse of dimension in Multi-Agent Reinforcement Learning (MARL). However, the existing methods usually either can only provide the low-order approximate decomposition of no more than the second-order, or need to spend a lot of effort to manually design the high-order interaction among agents according to experience. Therefore, the existing methods either tend to bear large decomposition error or are not convenient to use. In this paper, a high-order approximate value function decomposition method is proposed, which can be trained end-to-end. There have some prominent features about this method including low-rank vector exploited to represent value function, both low- and high-order component sharing the same input (i.e., the embedding vector), the model parameters shared among all the agents if they are homogeneous. Experimental results show that our method is effective.
KW - Low-Rank
KW - Many Agent Learning
KW - Reinforcement Learning
UR - https://www.scopus.com/pages/publications/85143344499
U2 - 10.1109/CEI57409.2022.9950106
DO - 10.1109/CEI57409.2022.9950106
M3 - 会议稿件
AN - SCOPUS:85143344499
T3 - 2022 2nd International Conference on Computer Science, Electronic Information Engineering and Intelligent Control Technology, CEI 2022
SP - 803
EP - 806
BT - 2022 2nd International Conference on Computer Science, Electronic Information Engineering and Intelligent Control Technology, CEI 2022
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2nd International Conference on Computer Science, Electronic Information Engineering and Intelligent Control Technology, CEI 2022
Y2 - 23 September 2022 through 25 September 2022
ER -