TY - JOUR
T1 - A novel multi-agent reinforcement learning approach based on state adaptive weighting and exploration path sampling
AU - Wang, Yichen
AU - Zheng, Shuai
AU - Yang, Ze
AU - Zhou, Xin
N1 - Publisher Copyright:
© 2026 Elsevier B.V.
PY - 2026/5
Y1 - 2026/5
N2 - In Multi-Agent Reinforcement Learning (MARL), the relationship between the current state and the reward is crucial to the model's performance. A good weight adaptation leads to a more reasonable reward, boosting the spatial exploration efficiency and stable control of all agents. This paper proposes State Adaptive Weighting (SAW) and Exploration Path Sampling (EPS), a powerful design to achieve state weight adaptation. SAW employs a dynamic state weighting network to prioritize informative and decision-critical states during training. It allows agents to focus on task-relevant regions of the state space and improves overall sample efficiency. EPS introduces a multi-faceted exploration strategy composed of three parts: a curiosity-driven module that uses prediction errors to generate intrinsic rewards, a path diversity tracker that encourages visits to novel states through visitation-based bonuses, and an adaptive noise mechanism that modulates exploration intensity based on environmental novelty. This paper conducts a series of experiments to demonstrate the improvements in learning speed and exploration quality. Our work offers a useful solution to address both effectiveness and stability in multi-agent systems. Abstract text.
AB - In Multi-Agent Reinforcement Learning (MARL), the relationship between the current state and the reward is crucial to the model's performance. A good weight adaptation leads to a more reasonable reward, boosting the spatial exploration efficiency and stable control of all agents. This paper proposes State Adaptive Weighting (SAW) and Exploration Path Sampling (EPS), a powerful design to achieve state weight adaptation. SAW employs a dynamic state weighting network to prioritize informative and decision-critical states during training. It allows agents to focus on task-relevant regions of the state space and improves overall sample efficiency. EPS introduces a multi-faceted exploration strategy composed of three parts: a curiosity-driven module that uses prediction errors to generate intrinsic rewards, a path diversity tracker that encourages visits to novel states through visitation-based bonuses, and an adaptive noise mechanism that modulates exploration intensity based on environmental novelty. This paper conducts a series of experiments to demonstrate the improvements in learning speed and exploration quality. Our work offers a useful solution to address both effectiveness and stability in multi-agent systems. Abstract text.
KW - Exploration path sampling
KW - Multi-agent
KW - Reinforcement learning
KW - State adaptive weighting
UR - https://www.scopus.com/pages/publications/105031767273
U2 - 10.1016/j.asoc.2026.114919
DO - 10.1016/j.asoc.2026.114919
M3 - 文章
AN - SCOPUS:105031767273
SN - 1568-4946
VL - 194
JO - Applied Soft Computing Journal
JF - Applied Soft Computing Journal
M1 - 114919
ER -