跳到主要导航 跳到搜索 跳到主要内容

Towards Black-Box Adversarial Attacks on Interpretable Deep Learning Systems

  • Yike Zhan
  • , Baolin Zheng
  • , Qian Wang
  • , Ningping Mou
  • , Binqing Guo
  • , Qi Li
  • , Chao Shen
  • , Cong Wang
  • Wuhan University
  • Tsinghua University
  • City University of Hong Kong

科研成果: 书/报告/会议事项章节会议稿件同行评审

8 引用 (Scopus)

摘要

Recent works have empirically shown that neural network interpretability is susceptible to malicious manipulations. However, existing attacks against Interpretable Deep Learning Systems (IDLSes) all focus on the white-box setting, which is obviously unpractical in real-world scenarios. In this paper, we make the first attempt to attack IDLSes in the decision-based black-box setting. We propose a new framework called Dual Black-box Adversarial Attack (DBAA) which can generate adversarial examples that are misclassified as the target class, yet have very similar interpretations to their benign cases. We conduct comprehensive experiments on different combinations of classifiers and interpreters to illustrate the effectiveness of DBAA. Empirical results show that in all the cases, DBAA achieves high attack success rates and Intersection over Union (IoU) scores.

源语言英语
主期刊名ICME 2022 - IEEE International Conference on Multimedia and Expo 2022, Proceedings
出版商IEEE Computer Society
ISBN(电子版)9781665485630
DOI
出版状态已出版 - 2022
活动2022 IEEE International Conference on Multimedia and Expo, ICME 2022 - Taipei, 中国台湾
期限: 18 7月 202222 7月 2022

出版系列

姓名Proceedings - IEEE International Conference on Multimedia and Expo
2022-July
ISSN(印刷版)1945-7871
ISSN(电子版)1945-788X

会议

会议2022 IEEE International Conference on Multimedia and Expo, ICME 2022
国家/地区中国台湾
Taipei
时期18/07/2222/07/22

学术指纹

探究 'Towards Black-Box Adversarial Attacks on Interpretable Deep Learning Systems' 的科研主题。它们共同构成独一无二的指纹。

引用此