Abstract
Semantic drift refers to the progressive erosion of a model's established semantic representations during fine-tuning. This effect is considered deleterious and requires mitigation. However, this paper reveals the potential of semantic drift for bias injection and proposes an efficient and stealthy bias injection method, Directed Drift. Unlike traditional backdoor attacks that use triggers to induce specific outputs, Directed Drift injects bias by subtly shifting the semantics of existing concepts. Specifically, Directed Drift trains the model on a small-scale dataset to control the direction of drift and incorporates a prior-preservation mechanism to control the extent of drift. This approach aims to induce a moderate semantic drift on a coarse-grained class concept toward a favored subclass. As a result, the bias-injected model generates the favored class objects when the coarse-grained class concept appears in isolation. When a specific modifier is present, the bias-injected model generates the disfavored subclass objects implied by this modifier to remain stealthy. Directed Drift applies to both text-to-image and inpainting tasks, thereby broadening its application scenarios. In addition, we introduce a new metric, Disfavored Subclass Utility, to evaluate the diversity of generations retained by a bias-injected model. This metric captures an aspect of attack stealthiness that has been largely overlooked in prior work. Extensive experiments demonstrate that Directed Drift achieves excellent performance compared to baseline methods, particularly showing substantial improvements in DSU. Moreover, Directed Drift attains highly stable and consistently high attack success rates across a wide range of bias-injection tasks.
| Original language | English |
|---|---|
| Article number | 134486 |
| Journal | Neurocomputing |
| Volume | 700 |
| DOIs | |
| State | Published - 1 Nov 2026 |
| Externally published | Yes |
Keywords
- Artificial intelligence security
- Backdoor attack
- Bias injection
- Diffusion model
- Generative model
Fingerprint
Dive into the research topics of 'Directed drift: An efficient and stealthy bias injection method using semantic drift against stable diffusion model'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver