跳到主要导航 跳到搜索 跳到主要内容

Directed drift: An efficient and stealthy bias injection method using semantic drift against stable diffusion model

  • Zhuowei Niu
  • , Qindong Sun
  • , Mingkai Ding
  • , Mingyue Song
  • , Chao Shen
  • , Sheng Wang
  • , Xuewen Huang
  • Xi'an Jiaotong University

科研成果: 期刊稿件文章同行评审

摘要

Semantic drift refers to the progressive erosion of a model's established semantic representations during fine-tuning. This effect is considered deleterious and requires mitigation. However, this paper reveals the potential of semantic drift for bias injection and proposes an efficient and stealthy bias injection method, Directed Drift. Unlike traditional backdoor attacks that use triggers to induce specific outputs, Directed Drift injects bias by subtly shifting the semantics of existing concepts. Specifically, Directed Drift trains the model on a small-scale dataset to control the direction of drift and incorporates a prior-preservation mechanism to control the extent of drift. This approach aims to induce a moderate semantic drift on a coarse-grained class concept toward a favored subclass. As a result, the bias-injected model generates the favored class objects when the coarse-grained class concept appears in isolation. When a specific modifier is present, the bias-injected model generates the disfavored subclass objects implied by this modifier to remain stealthy. Directed Drift applies to both text-to-image and inpainting tasks, thereby broadening its application scenarios. In addition, we introduce a new metric, Disfavored Subclass Utility, to evaluate the diversity of generations retained by a bias-injected model. This metric captures an aspect of attack stealthiness that has been largely overlooked in prior work. Extensive experiments demonstrate that Directed Drift achieves excellent performance compared to baseline methods, particularly showing substantial improvements in DSU. Moreover, Directed Drift attains highly stable and consistently high attack success rates across a wide range of bias-injection tasks.

源语言英语
期刊论文编号134486
期刊Neurocomputing
700
DOI
出版状态已出版 - 1 11月 2026
已对外发布

学术指纹

探究 'Directed drift: An efficient and stealthy bias injection method using semantic drift against stable diffusion model' 的科研主题。它们共同构成独一无二的学术指纹。

引用此