TY - JOUR
T1 - DLANet
T2 - A lightweight dual-stream framework with fine-grained spatio-temporal attention for micro-expression recognition
AU - Zhong, Xianjing
AU - Qu, Kai
AU - Fan, Tianyi
AU - Cao, Hui
AU - Zhang, Jie
N1 - Publisher Copyright:
© 2026 Elsevier Inc.
PY - 2026/8
Y1 - 2026/8
N2 - Micro-expression recognition (MER) remains challenging in practical settings because subtle facial deformations are easily obscured by appearance variation and slight head motion, discriminative evidence is temporally sparse around apex-related moments, and stronger spatio-temporal modeling often conflicts with deployment efficiency under limited-data conditions. To address these issues, we propose a Dual Lightweight Attention-guided Network (DLANet), a compact two-stage dual-stream framework for deployment-oriented MER. In the first stage, the Micro-expression Perceptive Appearance-Motion Dual Network (MP-AMDNet) performs motion-aware feature acquisition by combining dense optical-flow priors, weak appearance-motion consistency, and entropy-regularized motion attention with adaptive key-frame selection to preserve sparse but informative temporal evidence. In the second stage, the Spatio-temporal Fine-grained Attention Network (ST-FANet) performs lightweight spatio-temporal refinement through factorized channel-aware pseudo-3D modeling, multi-scale dilated feature fusion, and motion-guided deformable refinement, thereby enhancing subtle local deformations without relying on heavy global attention or full 3D processing. Furthermore, a composite objective combining focal reweighting, auxiliary alignment regularization, and temporal sparsity regularization improves robustness under small-sample and class-imbalanced settings. Extensive experiments on CASME II, SAMM, SMIC, and MEGC 2019 under both Single-Dataset Evaluation (SDE) and Composite-Dataset Evaluation (CDE) show that DLANet achieves competitive or improved recognition performance, including 0.8650 UF1 and 0.8450 UAR on MEGC 2019, while remaining compact with 3.40M parameters and a transparent efficiency profile. These results indicate that DLANet provides a favorable accuracy-efficiency trade-off for MER in resource-constrained settings.
AB - Micro-expression recognition (MER) remains challenging in practical settings because subtle facial deformations are easily obscured by appearance variation and slight head motion, discriminative evidence is temporally sparse around apex-related moments, and stronger spatio-temporal modeling often conflicts with deployment efficiency under limited-data conditions. To address these issues, we propose a Dual Lightweight Attention-guided Network (DLANet), a compact two-stage dual-stream framework for deployment-oriented MER. In the first stage, the Micro-expression Perceptive Appearance-Motion Dual Network (MP-AMDNet) performs motion-aware feature acquisition by combining dense optical-flow priors, weak appearance-motion consistency, and entropy-regularized motion attention with adaptive key-frame selection to preserve sparse but informative temporal evidence. In the second stage, the Spatio-temporal Fine-grained Attention Network (ST-FANet) performs lightweight spatio-temporal refinement through factorized channel-aware pseudo-3D modeling, multi-scale dilated feature fusion, and motion-guided deformable refinement, thereby enhancing subtle local deformations without relying on heavy global attention or full 3D processing. Furthermore, a composite objective combining focal reweighting, auxiliary alignment regularization, and temporal sparsity regularization improves robustness under small-sample and class-imbalanced settings. Extensive experiments on CASME II, SAMM, SMIC, and MEGC 2019 under both Single-Dataset Evaluation (SDE) and Composite-Dataset Evaluation (CDE) show that DLANet achieves competitive or improved recognition performance, including 0.8650 UF1 and 0.8450 UAR on MEGC 2019, while remaining compact with 3.40M parameters and a transparent efficiency profile. These results indicate that DLANet provides a favorable accuracy-efficiency trade-off for MER in resource-constrained settings.
KW - Deformable attention
KW - Entropy regularization
KW - Lightweight neural network
KW - Micro-expression recognition
KW - Optical flow
UR - https://www.scopus.com/pages/publications/105040624518
U2 - 10.1016/j.cviu.2026.104823
DO - 10.1016/j.cviu.2026.104823
M3 - 文章
AN - SCOPUS:105040624518
SN - 1077-3142
VL - 270
JO - Computer Vision and Image Understanding
JF - Computer Vision and Image Understanding
M1 - 104823
ER -