摘要
Large language models have demonstrated outstanding performance in various applications and are being widely adopted as critical engines for creating new productivity tools. However, when malicious users employ specific techniques to bypass the security protections established through mechanisms like alignment, it can lead to jailbreak attacks. These attacks may generate content that violates the model's usage guidelines, ethics, or laws, raising ethical concerns. This paper comprehensively examines the origins of jailbreak attacks and their evolution in terms of attack and defense. First, a definition and a formal framework of jailbreak attacks are proposed based on three elements: methods, objects, and targets. Furthermore, the development history of jailbreak attacks is introduced from two perspectives: the evolution of large language models and changes in security perceptions, and the root cause of jailbreak attacks is summarized as the mismatch between the service attributes and values of large language models. Finally, from the perspective of offensive and defensive games, this paper summarizes the evolution of jailbreak attacks and defenses, discussing new threat models in jailbreak attacks and the future directions of defense methods.
| 投稿的翻译标题 | Jailbreaking large language models: models, origins, and evolution of attacks and defenses |
|---|---|
| 源语言 | 繁体中文 |
| 页(从-至) | 1372 |
| 页数 | 1 |
| 期刊 | Scientia Sinica Informationis |
| 卷 | 55 |
| 期 | 6 |
| DOI | |
| 出版状态 | 已出版 - 1 6月 2025 |
关键词
- cybersecurity
- ethics of artificial intelligence
- jailbreak attack
- large language model
- natural language process
学术指纹
探究 '大语言模型越狱攻击: 模型、根因及其攻防演化' 的科研主题。它们共同构成独一无二的学术指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver