跳到主要导航 跳到搜索 跳到主要内容

Dataset Copyright Auditing for Large Models: Fundamentals, Open Problems, and Future Directions

  • Xi'an Jiaotong University

科研成果: 期刊稿件文章同行评审

1 引用 (Scopus)

摘要

The unprecedented scale of large models, such as large language models (LLMs) and text-to-image diffusion models, has raised critical concerns about the unauthorized use of copyrighted data during model training. These concerns have spurred a growing demand for da taset copyright auditing techniques, which aim to detect and verify potential infringements in the training data of commercial AI systems. This paper presents a survey of existing auditing solutions, categorizing them across key dimensions: data modality, model training stage, data over lap scenarios, and model access levels. We highlight major trends, including the prevalence of black-box auditing methods and the emphasis on fine-tuning rather than pre-training. Through an in-depth analysis of 12 representative works, we extract four key observations that reveal the limitations of current methods. Furthermore, we identify three open challenges and propose future directions for robust, multimodal, and scalable auditing solutions. Our findings underscore the urgent need to establish standardized benchmarks and develop auditing frameworks that are resilient to low watermark densities and applicable in diverse deployment settings.

源语言英语
页(从-至)38-47
页数10
期刊ZTE Communications
23
3
DOI
出版状态已出版 - 11 9月 2025

学术指纹

探究 'Dataset Copyright Auditing for Large Models: Fundamentals, Open Problems, and Future Directions' 的科研主题。它们共同构成独一无二的指纹。

引用此