TY - JOUR
T1 - Semantic Distribution and Authenticity Discrepancy Alignment for AI-Generated Image Detection
AU - Zhang, Jiehua
AU - Li, Liang
AU - Yan, Chenggang
AU - Ke, Wei
AU - Gong, Yihong
N1 - Publisher Copyright:
© 1999-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - Generative models have achieved remarkable success in producing vivid images. Compared with real images, generated ones still show different semantic structures that features with different semantic classes collapse as a single cluster. Pioneer works leverage the discrepancy of semantic structure in fixed high-level semantic feature space to identify forgery images. Nevertheless, such frozen pre-trained representation models are insensitive to subtle forgery traces. Meanwhile, vanilla fine-tuning methods can distort the pre-trained semantic knowledge and collapse to the real-fake binary distribution, losing generalization capability in newly emerged generative models. In this paper, we propose the semantic distribution and authenticity discrepancy alignment algorithm (STERM), which learns high-level semantic structures of real-world categories and low-level forgery traces for detecting AI-generated images from unseen generative models and frameworks. Specifically, we first capture semantic features of images by the frozen CLIP and further extract forgery features by a forgery encoder. Then, we propose semantic distribution alignment (SDA) to align the semantic structure of real-world categories by enforcing forgery feature distribution shifting towards the semantic feature space. Next, we introduce authenticity discrepancy alignment (ADA) to minimize the authenticity discrepancy between forgery and semantic features, constraining forgery features from collapsing into the source domain-biased distribution and learning the semantic structure of real-world categories. Extensive experiments on GAN-based and diffusion model-based datasets demonstrate the generalization capability of the proposed method.
AB - Generative models have achieved remarkable success in producing vivid images. Compared with real images, generated ones still show different semantic structures that features with different semantic classes collapse as a single cluster. Pioneer works leverage the discrepancy of semantic structure in fixed high-level semantic feature space to identify forgery images. Nevertheless, such frozen pre-trained representation models are insensitive to subtle forgery traces. Meanwhile, vanilla fine-tuning methods can distort the pre-trained semantic knowledge and collapse to the real-fake binary distribution, losing generalization capability in newly emerged generative models. In this paper, we propose the semantic distribution and authenticity discrepancy alignment algorithm (STERM), which learns high-level semantic structures of real-world categories and low-level forgery traces for detecting AI-generated images from unseen generative models and frameworks. Specifically, we first capture semantic features of images by the frozen CLIP and further extract forgery features by a forgery encoder. Then, we propose semantic distribution alignment (SDA) to align the semantic structure of real-world categories by enforcing forgery feature distribution shifting towards the semantic feature space. Next, we introduce authenticity discrepancy alignment (ADA) to minimize the authenticity discrepancy between forgery and semantic features, constraining forgery features from collapsing into the source domain-biased distribution and learning the semantic structure of real-world categories. Extensive experiments on GAN-based and diffusion model-based datasets demonstrate the generalization capability of the proposed method.
KW - AI-generated image detection
KW - authenticity discrepancy alignment
KW - semantic distribution alignment
UR - https://www.scopus.com/pages/publications/105031721329
U2 - 10.1109/TMM.2026.3668513
DO - 10.1109/TMM.2026.3668513
M3 - 文章
AN - SCOPUS:105031721329
SN - 1520-9210
JO - IEEE Transactions on Multimedia
JF - IEEE Transactions on Multimedia
ER -