TY - GEN
T1 - From Cognitive Priors to Instance Semantics
T2 - 2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026
AU - Hu, Guanyu
AU - Kollias, Dimitrios
AU - Yang, Xinyu
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - Understanding human affect via Valence-Arousal, Expressions, and Action Unit is essential for human-machine interaction. While recent multi-task learning (MTL) methods seek to unify these tasks, they overlook three key challenges: (i) the absence of unified modeling all three affective task types: regression, detection, and classification; (ii) reliance on complete annotations for all tasks, leaving disjoint single-task datasets underutilized; and (iii) task conflicts caused by Noisy Gradients, Negative Transfer (NT), and Task-specific Performance Misalignment (TPM). We introduce COIN, a novel two-stage MTL framework that bridges Cognitive Priors and Instance Semantics for robust training. First, we design a cognitively guided cross-task label induction strategy to propagate supervision under sparse annotations and mitigate NT, yielding strong task-specific CogXperts. Second, we introduce two complementary branches to address TPM: (i) Task-Specific Branch: transferring cognitive knowledge from task-optimal CogXperts to jointly optimize objectives under partial supervision, and (ii) Semantic Alignment Branch: enhancing instance-level semantic representations via Class-Conditioned and Instance-Adaptive Prompts. Experiments across six diverse datasets demonstrate COIN's robustness and generalization. Code is available at https://github.com/imhgy/COIN.
AB - Understanding human affect via Valence-Arousal, Expressions, and Action Unit is essential for human-machine interaction. While recent multi-task learning (MTL) methods seek to unify these tasks, they overlook three key challenges: (i) the absence of unified modeling all three affective task types: regression, detection, and classification; (ii) reliance on complete annotations for all tasks, leaving disjoint single-task datasets underutilized; and (iii) task conflicts caused by Noisy Gradients, Negative Transfer (NT), and Task-specific Performance Misalignment (TPM). We introduce COIN, a novel two-stage MTL framework that bridges Cognitive Priors and Instance Semantics for robust training. First, we design a cognitively guided cross-task label induction strategy to propagate supervision under sparse annotations and mitigate NT, yielding strong task-specific CogXperts. Second, we introduce two complementary branches to address TPM: (i) Task-Specific Branch: transferring cognitive knowledge from task-optimal CogXperts to jointly optimize objectives under partial supervision, and (ii) Semantic Alignment Branch: enhancing instance-level semantic representations via Class-Conditioned and Instance-Adaptive Prompts. Experiments across six diverse datasets demonstrate COIN's robustness and generalization. Code is available at https://github.com/imhgy/COIN.
KW - affective computing
KW - expression recognition
KW - mtl
UR - https://www.scopus.com/pages/publications/105041258223
U2 - 10.1109/WACV61042.2026.00825
DO - 10.1109/WACV61042.2026.00825
M3 - 会议稿件
AN - SCOPUS:105041258223
T3 - Proceedings - 2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026
SP - 8551
EP - 8562
BT - Proceedings - 2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 6 March 2026 through 10 March 2026
ER -