跳到主要导航 跳到搜索 跳到主要内容

大语言模型安全与隐私风险综述

  • Yi Jiang
  • , Yong Yang
  • , Jiali Yin
  • , Xiaolei Liu
  • , Jiliang Li
  • , Wei Wang
  • , Youliang Tian
  • , Yingcai Wu
  • , Shouling Ji
  • Zhejiang University
  • Guizhou University
  • Fuzhou University
  • China Academy of Engineering Physics
  • Beijing Jiaotong University

科研成果: 期刊稿件文章同行评审

1 引用 (Scopus)

摘要

In recent years, large language models (LLMs) have emerged as a critical branch of deep learning network technology, achieving a series of breakthrough accomplishments in the field of natural language processing (NLP), and gaining widespread adoption. However, throughout their entire lifecycle, including pre-training, fine-tuning, and actual deployment, a variety of security threats and risks of privacy breaches have been discovered, drawing increasing attention from both the academic and industrial sectors. Navigating the development of the paradigm of using large language models to handle natural language processing tasks, as known as the pre-training and fine-tuning paradigm, the pre-training and prompt learning paradigm, and the pre-training and instruction-tuning paradigm, this article outlines conventional security threats against large language models, specifically representative studies on the three types of traditional adversarial attacks (adversarial example attack, backdoor attack and poisoning attack). It then summarizes some of the novel security threats revealed by recent research, followed by a discussion on the privacy risks of large language models and the progress in their research. The content aids researchers and deployers of large language models in identifying, preventing, and mitigating these threats and risks during the model design, training, and application processes, while also achieving a balance between model performance, security, and privacy protection.

投稿的翻译标题Survey on Security and Privacy Risks in Large Language Models
源语言繁体中文
页(从-至)1979-2018
页数40
期刊Jisuanji Yanjiu yu Fazhan/Computer Research and Development
62
8
DOI
出版状态已出版 - 8月 2025

联合国可持续发展目标

此成果有助于实现下列可持续发展目标:

  1. 可持续发展目标 3 - 良好健康与福祉
    可持续发展目标 3 良好健康与福祉

关键词

  • large language models (LLMs)
  • pre-trained language models
  • privacy
  • security
  • threat

学术指纹

探究 '大语言模型安全与隐私风险综述' 的科研主题。它们共同构成独一无二的指纹。

引用此