TY - JOUR
T1 - Dataset Copyright Auditing for Large Models
T2 - Fundamentals, Open Problems, and Future Directions
AU - Du, Linkang
AU - Su, Zhou
AU - Yu, Xinyi
N1 - Publisher Copyright:
© 2025, ZTE Communications. All rights reserved.
PY - 2025/9/11
Y1 - 2025/9/11
N2 - The unprecedented scale of large models, such as large language models (LLMs) and text-to-image diffusion models, has raised critical concerns about the unauthorized use of copyrighted data during model training. These concerns have spurred a growing demand for da taset copyright auditing techniques, which aim to detect and verify potential infringements in the training data of commercial AI systems. This paper presents a survey of existing auditing solutions, categorizing them across key dimensions: data modality, model training stage, data over lap scenarios, and model access levels. We highlight major trends, including the prevalence of black-box auditing methods and the emphasis on fine-tuning rather than pre-training. Through an in-depth analysis of 12 representative works, we extract four key observations that reveal the limitations of current methods. Furthermore, we identify three open challenges and propose future directions for robust, multimodal, and scalable auditing solutions. Our findings underscore the urgent need to establish standardized benchmarks and develop auditing frameworks that are resilient to low watermark densities and applicable in diverse deployment settings.
AB - The unprecedented scale of large models, such as large language models (LLMs) and text-to-image diffusion models, has raised critical concerns about the unauthorized use of copyrighted data during model training. These concerns have spurred a growing demand for da taset copyright auditing techniques, which aim to detect and verify potential infringements in the training data of commercial AI systems. This paper presents a survey of existing auditing solutions, categorizing them across key dimensions: data modality, model training stage, data over lap scenarios, and model access levels. We highlight major trends, including the prevalence of black-box auditing methods and the emphasis on fine-tuning rather than pre-training. Through an in-depth analysis of 12 representative works, we extract four key observations that reveal the limitations of current methods. Furthermore, we identify three open challenges and propose future directions for robust, multimodal, and scalable auditing solutions. Our findings underscore the urgent need to establish standardized benchmarks and develop auditing frameworks that are resilient to low watermark densities and applicable in diverse deployment settings.
KW - dataset copyright auditing
KW - diffusion models
KW - large language models
KW - membership inference
KW - multimodal auditing
UR - https://www.scopus.com/pages/publications/105018596028
U2 - 10.12142/ZTECOM.202503005
DO - 10.12142/ZTECOM.202503005
M3 - 文章
AN - SCOPUS:105018596028
SN - 1673-5188
VL - 23
SP - 38
EP - 47
JO - ZTE Communications
JF - ZTE Communications
IS - 3
ER -