跳到主要导航 跳到搜索 跳到主要内容

URLcoat: Exploiting Web Search Capability to Jailbreak Large Language Models

  • Yiheng Sun
  • , Linkang Du
  • , Zhou Su
  • , Yuntao Wang
  • , Han Liu
  • Xi'an Jiaotong University

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Large language models (LLMs) achieve remarkable advances in understanding and reasoning with human language, which are widely applied in software development, content creation, healthcare, etc. However, their vulnerabilities to security threats, especially jailbreak attacks, remain a significant issue. Existing research on jailbreak mainly focuses on the security risks of LLMs' inherent thinking and reasoning, overlooking the new attack surface introduced by web search capability. Attackers can exploit this by guiding LLMs to retrieve information from external URLs, which is then used to implicitly reconstruct harmful instructions and circumvent safety mechanisms, leading to the generation of harmful content. In this paper, we propose a novel jailbreak attack, named URLcoat. The core idea is to exploit the web search capabilities of LLMs to circumvent their security safeguards. URLcoat incorporates three core strategies: obfuscating the feature of sensitive words to evade input detection, reconstructing harmful instructions via implicit associations with external URLs, and contextual narrative guidance to bypass output filtering. The experimental results reveal that URLcoat attains 100 % attack success rates in mainstream LLMs, including GPT5, GPT-4o, Gemini 2.0 Flash Thinking, DeepSeek R1, Kimi 1.5, ChatGLM-4, Grok 3, Gemini 2.5 Pro, Gemini 2.5 Flash, and Gemini 2.0 Flash, exceeding the performance of state-of-the-art jailbreak techniques. This study examines the security vulnerabilities arising from lLMs' web search capability, which facilitates the LLMs to produce harmful output.

源语言英语
主期刊名Proceedings - 47th IEEE Symposium on Security and Privacy, SP 2026
编辑Alina Oprea, Cristina Nita-Rotaru, Nicolas Papernot
出版商Institute of Electrical and Electronics Engineers Inc.
59-77
页数19
ISBN(电子版)9798331560652
DOI
出版状态已出版 - 2026
活动47th IEEE Symposium on Security and Privacy, SP 2026 - San Francisco, 美国
期限: 18 5月 202621 5月 2026

丛书

姓名Proceedings - IEEE Symposium on Security and Privacy
ISSN(印刷版)1081-6011

会议

会议47th IEEE Symposium on Security and Privacy, SP 2026
国家/地区美国
San Francisco
时期18/05/2621/05/26

联合国可持续发展目标

此成果有助于实现下列可持续发展目标:

  1. 可持续发展目标 3 - 良好健康与福祉
    可持续发展目标 3 良好健康与福祉

学术指纹

探究 'URLcoat: Exploiting Web Search Capability to Jailbreak Large Language Models' 的科研主题。它们共同构成独一无二的学术指纹。

引用此