TY - JOUR
T1 - EmoTraceNet
T2 - Leveraging Residual Emotional Traces for Robust Audio Deepfake Detection
AU - Zheng, Chende
AU - He, Zongyi
AU - Lin, Chenhao
AU - Ji, Zhoulin
AU - Luo, Xiapu
AU - Shen, Chao
N1 - Publisher Copyright:
© 2004-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - The rapid advancement of artificial intelligence and generative models has enabled the creation of highly realistic synthetic speech, giving rise to deceptive audio deepfakes. Specifically, Emotional Voice Conversion (EVC) poses severe security risks by manipulating the emotional prosody of an utterance to bypass psychological defenses while strictly preserving linguistic content and speaker identity. Existing Audio Deepfake Detection (ADD) systems exhibit significant vulnerabilities against EVC attacks, primarily due to extreme emotional imbalances in current public datasets and an inability to decouple emotional features from spoofing artifacts. To address these limitations, we first introduce DeepEmo, a large-scale, emotionally balanced audio dataset encompassing the latest EVC generative models. Furthermore, our empirical analysis reveals that current EVC algorithms struggle to entirely disentangle source emotions, leaving identifiable emotional traces. Motivated by this emotion trace hypothesis, we propose EmoTraceNet, a novel dual-branch framework integrating a spoofing detection branch and an auxiliary emotion classification branch. By employing a cross-task attention fusion mechanism, EmoTraceNet explicitly captures the residual emotional traces as well as generative artifacts. Extensive experiments on the DeepEmo and EmoFake benchmarks demonstrate that our proposed approach achieves state-of-the-art performance (3.06% EER on DeepEmo). Notably, EmoTraceNet exhibits superior cross-domain generalization, robust cross-emotion detection capabilities, and strong resilience against acoustic noise compared to existing baselines.
AB - The rapid advancement of artificial intelligence and generative models has enabled the creation of highly realistic synthetic speech, giving rise to deceptive audio deepfakes. Specifically, Emotional Voice Conversion (EVC) poses severe security risks by manipulating the emotional prosody of an utterance to bypass psychological defenses while strictly preserving linguistic content and speaker identity. Existing Audio Deepfake Detection (ADD) systems exhibit significant vulnerabilities against EVC attacks, primarily due to extreme emotional imbalances in current public datasets and an inability to decouple emotional features from spoofing artifacts. To address these limitations, we first introduce DeepEmo, a large-scale, emotionally balanced audio dataset encompassing the latest EVC generative models. Furthermore, our empirical analysis reveals that current EVC algorithms struggle to entirely disentangle source emotions, leaving identifiable emotional traces. Motivated by this emotion trace hypothesis, we propose EmoTraceNet, a novel dual-branch framework integrating a spoofing detection branch and an auxiliary emotion classification branch. By employing a cross-task attention fusion mechanism, EmoTraceNet explicitly captures the residual emotional traces as well as generative artifacts. Extensive experiments on the DeepEmo and EmoFake benchmarks demonstrate that our proposed approach achieves state-of-the-art performance (3.06% EER on DeepEmo). Notably, EmoTraceNet exhibits superior cross-domain generalization, robust cross-emotion detection capabilities, and strong resilience against acoustic noise compared to existing baselines.
KW - Emotional voice conversion
KW - Speech anti-spoofing
KW - Speech recognition and synthesis
UR - https://www.scopus.com/pages/publications/105045765171
U2 - 10.1109/TDSC.2026.3716730
DO - 10.1109/TDSC.2026.3716730
M3 - 文章
AN - SCOPUS:105045765171
SN - 1545-5971
JO - IEEE Transactions on Dependable and Secure Computing
JF - IEEE Transactions on Dependable and Secure Computing
ER -