TY - GEN
T1 - Ear with Eye
T2 - 33rd ACM International Conference on Multimedia, MM 2025
AU - Jiang, Xuanming
AU - An, Baoyi
AU - Zou, Zhengwei
AU - Nie, Dingyu
AU - Shen, Jialie
AU - Qian, Xueming
AU - Zhao, Guoshuai
N1 - Publisher Copyright:
© 2025 ACM.
PY - 2025/10/27
Y1 - 2025/10/27
N2 - In cutting-edge domains such as unmanned aerial vehicles and autonomous driving, edge-based audio-visual systems struggle to strike an optimal balance between complexity and performance. Unlike prevailing approaches that typically rely on pruning and knowledge distillation to streamline unimodal or hybrid models, we propose the Bio-Inspired Multimodal Network (BIMNet), which achieves an efficient audio-visual shared architecture. BIMNet integrates bio-inspired audio-visual modules that emulate the hierarchical sensory integration observed in nocturnal birds, to replicate equivalent biological information flow for both multiscale night vision and noise-adaptive hearing. Experimental findings show that BIMNet achieves superior performance and efficiency in diverse image datasets (varying in spatial scales and lighting conditions), audio datasets (encompassing various types of human and environmental sound), and audio-visual joint event detection tasks. Project support is available at: https://github.com/Mental-Scholar/BIMNet.
AB - In cutting-edge domains such as unmanned aerial vehicles and autonomous driving, edge-based audio-visual systems struggle to strike an optimal balance between complexity and performance. Unlike prevailing approaches that typically rely on pruning and knowledge distillation to streamline unimodal or hybrid models, we propose the Bio-Inspired Multimodal Network (BIMNet), which achieves an efficient audio-visual shared architecture. BIMNet integrates bio-inspired audio-visual modules that emulate the hierarchical sensory integration observed in nocturnal birds, to replicate equivalent biological information flow for both multiscale night vision and noise-adaptive hearing. Experimental findings show that BIMNet achieves superior performance and efficiency in diverse image datasets (varying in spatial scales and lighting conditions), audio datasets (encompassing various types of human and environmental sound), and audio-visual joint event detection tasks. Project support is available at: https://github.com/Mental-Scholar/BIMNet.
KW - bio-inspired architecture
KW - lightweight audio-visual network
UR - https://www.scopus.com/pages/publications/105024063063
U2 - 10.1145/3746027.3755149
DO - 10.1145/3746027.3755149
M3 - 会议稿件
AN - SCOPUS:105024063063
T3 - MM 2025 - Proceedings of the 33rd ACM International Conference on Multimedia, Co-Located with MM 2025
SP - 1346
EP - 1355
BT - MM 2025 - Proceedings of the 33rd ACM International Conference on Multimedia, Co-Located with MM 2025
PB - Association for Computing Machinery, Inc
Y2 - 27 October 2025 through 31 October 2025
ER -