Skip to main navigation Skip to search Skip to main content

Graph-based strategy evaluation for large-scale multiagent reinforcement learning

  • Yiyun Sun
  • , Meiqin Liu
  • , Senlin Zhang
  • , Ronghao Zheng
  • , Shanling Dong
  • Zhejiang University
  • Jinhua Institute of Zhejiang University

Research output: Contribution to journalArticlepeer-review

Abstract

In large-scale multiagent systems, the practical application of multiagent reinforcement learning (MARL) is hindered by the absence of robust reliability assurances. This gap has heightened the focus on strategy evaluation within the MARL framework, a domain that grapples with scalability issues in the joint strategy space. To address this concern, this paper introduces a novel two-stage graph-based strategy evaluation algorithm that significantly reduces the required sample capacity in the joint strategy space without compromising the evaluation quality. The proposed algorithm performs a hierarchical evaluation to compress sample capacity and employs a strategy-seeking model to seek a sink equilibrium (SE) joint strategy using the best responses. Moreover, a stopping condition is developed to achieve an approximately globally optimal SE strategy, accounting for the local optimal properties of the best-response-based algorithm. Case studies demonstrate that our algorithm achieves an approximately optimal SE joint strategy with superior sample efficiency compared with other approaches. The integration of MARL methods with the strategy evaluation algorithm proves to be an effective approach for establishing trustworthy MARL systems.

Original languageEnglish
Article number182206
JournalScience China Information Sciences
Volume68
Issue number8
DOIs
StatePublished - Aug 2025

Keywords

  • best response
  • graph grouping
  • large-scale multiagent reinforcement learning
  • sink equilibrium
  • strategy evaluation

Fingerprint

Dive into the research topics of 'Graph-based strategy evaluation for large-scale multiagent reinforcement learning'. Together they form a unique fingerprint.

Cite this