TY - GEN
T1 - D3
T2 - 2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025
AU - Zheng, Chende
AU - Suo, Ruiqi
AU - Lin, Chenhao
AU - Zhao, Zhengyu
AU - Yang, Le
AU - Liu, Shuai
AU - Yang, Minghui
AU - Wang, Cong
AU - Shen, Chao
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - The evolution of video generation techniques, such as Sora, has made it increasingly easy to produce highfidelity AI-generated videos, raising public concern over the dissemination of synthetic content. However, existing detection methodologies remain limited by their insufficient exploration of temporal artifacts in synthetic videos. To bridge this gap, we establish a theoretical framework through second-order dynamical analysis under Newtonian mechanics, subsequently extending the Second-order Central Difference features tailored for temporal artifact detection. Building on this theoretical foundation, we reveal a fundamental divergence in second-order feature distributions between real and AI-generated videos. Concretely, we propose Detection by Difference of Differences (D3), a novel training-free detection method that leverages the above second-order temporal discrepancies. We validate the superiority of our D3 on 4 open-source datasets (GenVideo, VideoPhy, EvalCrafter, VidProM), 40 subsets in total. For example, on GenVideo, D3 outperforms the previous state-of-the-art method by 10.39% (absolute) mean Average Precision. Additional experiments on time cost and post-processing operations demonstrate D3's exceptional computational efficiency and strong robust performance. Our code is available at https://github.com/Zig-HS/D3.
AB - The evolution of video generation techniques, such as Sora, has made it increasingly easy to produce highfidelity AI-generated videos, raising public concern over the dissemination of synthetic content. However, existing detection methodologies remain limited by their insufficient exploration of temporal artifacts in synthetic videos. To bridge this gap, we establish a theoretical framework through second-order dynamical analysis under Newtonian mechanics, subsequently extending the Second-order Central Difference features tailored for temporal artifact detection. Building on this theoretical foundation, we reveal a fundamental divergence in second-order feature distributions between real and AI-generated videos. Concretely, we propose Detection by Difference of Differences (D3), a novel training-free detection method that leverages the above second-order temporal discrepancies. We validate the superiority of our D3 on 4 open-source datasets (GenVideo, VideoPhy, EvalCrafter, VidProM), 40 subsets in total. For example, on GenVideo, D3 outperforms the previous state-of-the-art method by 10.39% (absolute) mean Average Precision. Additional experiments on time cost and post-processing operations demonstrate D3's exceptional computational efficiency and strong robust performance. Our code is available at https://github.com/Zig-HS/D3.
KW - ai security
KW - deepfake detection
KW - second-order feature
KW - video generation
UR - https://www.scopus.com/pages/publications/105044119756
U2 - 10.1109/ICCV51701.2025.01194
DO - 10.1109/ICCV51701.2025.01194
M3 - 会议稿件
AN - SCOPUS:105044119756
T3 - Proceedings of the IEEE International Conference on Computer Vision
SP - 12852
EP - 12862
BT - Proceedings - 2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 19 October 2025 through 23 October 2025
ER -