跳到主要导航 跳到搜索 跳到主要内容

Toward Evaluating Robustness of Reinforcement Learning with Adversarial Policy

  • Xiang Zheng
  • , Xingjun Ma
  • , Shengjie Wang
  • , Xinyu Wang
  • , Chao Shen
  • , Cong Wang
  • City University of Hong Kong
  • Fudan University
  • Tsinghua University
  • Tencent

科研成果: 书/报告/会议事项章节会议稿件同行评审

4 引用 (Scopus)

摘要

Reinforcement learning agents are susceptible to evasion attacks during deployment. In single-agent environments, these attacks can occur through imperceptible perturbations injected into the inputs of the victim policy network. In multi-agent environments, an attacker can manipulate an adversarial opponent to influence the victim policy's observations indirectly. While adversarial policies offer a promising technique to craft such attacks, current methods are either sample-inefficient due to poor exploration strategies or require extra surrogate model training under the black-box assumption. To address these challenges, in this paper, we propose Intrinsically Motivated Adversarial Policy (IMAP) for efficient black-box adversarial policy learning in both single- and multi-agent environments. We formulate four types of adversarial intrinsic regularizers - maximizing the adversarial state coverage, policy coverage, risk, or divergence - to discover potential vulnerabilities of the victim policy in a principled way. We also present a novel bias-reduction method to balance the extrinsic objective and the adversarial intrinsic regularizers adaptively. Our experiments validate the effectiveness of the four types of adversarial intrinsic regularizers and the bias-reduction method in enhancing black-box adversarial policy learning across a variety of environments. Our IMAP successfully evades two types of defense methods, adversarial training and robust regularizer, decreasing the performance of the state-of-the-art robust WocaR-PPO agents by 34%-54% across four single-agent tasks. IMAP also achieves a state-of-the-art attacking success rate of 83.91% in the multi-agent game YouShallNotPass. Our code is available at https://github.com/x-zheng16/IMAP.

源语言英语
主期刊名Proceedings - 2024 54th Annual IEEE/IFIP International Conference on Dependable Systems and Networks, DSN 2024
出版商Institute of Electrical and Electronics Engineers Inc.
288-301
页数14
ISBN(电子版)9798350341058
DOI
出版状态已出版 - 2024
活动54th Annual IEEE/IFIP International Conference on Dependable Systems and Networks, DSN 2024 - Brisbane, 澳大利亚
期限: 24 6月 202427 6月 2024

出版系列

姓名Proceedings - 2024 54th Annual IEEE/IFIP International Conference on Dependable Systems and Networks, DSN 2024

会议

会议54th Annual IEEE/IFIP International Conference on Dependable Systems and Networks, DSN 2024
国家/地区澳大利亚
Brisbane
时期24/06/2427/06/24

学术指纹

探究 'Toward Evaluating Robustness of Reinforcement Learning with Adversarial Policy' 的科研主题。它们共同构成独一无二的指纹。

引用此