TY - JOUR
T1 - Composable Multimodal Semantic Communication
T2 - A Lightweight Large AI Model Approach
AU - Zhao, Tantan
AU - Li, Fan
AU - Huang, Xinyu
AU - Liu, Yiqun
AU - Nallanathan, Arumugam
N1 - Publisher Copyright:
© 1972-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - Multimodal semantic communication is a promising paradigm for enabling immersive intelligent services in sixth-generation (6G) networks, yet existing studies lack flexible support for arbitrary modality combinations and lightweight semantic knowledge modeling suitable for edge deployment. In this paper, we investigate a lightweight large AI model (LAM)-empowered composable multimodal semantic communication (CMSC) framework. Text is adopted as a semantic bridge to unify heterogeneous modalities, while a flexible and learnable modality composition weighting mechanism enables arbitrary combinations of image, audio, and video inputs to be aggregated into a communication-oriented semantic representation. To enhance semantic robustness, a bottleneck-aware lightweight semantic knowledge base is constructed by leveraging a frozen large language model with visual prompts, where raw images or middle video frames are jointly used with textual semantics to mitigate semantic ambiguity and compensate for semantic degradation caused by wireless channel impairments. The enhanced semantics are encoded and transmitted over wireless channels, and multimodal signals are reconstructed at the receiver via diffusion-based generation. Extensive experiments on real-world datasets show that CMSC achieves a higher compression rate while maintaining comparable transmission accuracy under both arbitrary and typical multimodal combinations, demonstrating its effectiveness for flexible and lightweight multimodal semantic communication in 6G ubiquitous intelligence.
AB - Multimodal semantic communication is a promising paradigm for enabling immersive intelligent services in sixth-generation (6G) networks, yet existing studies lack flexible support for arbitrary modality combinations and lightweight semantic knowledge modeling suitable for edge deployment. In this paper, we investigate a lightweight large AI model (LAM)-empowered composable multimodal semantic communication (CMSC) framework. Text is adopted as a semantic bridge to unify heterogeneous modalities, while a flexible and learnable modality composition weighting mechanism enables arbitrary combinations of image, audio, and video inputs to be aggregated into a communication-oriented semantic representation. To enhance semantic robustness, a bottleneck-aware lightweight semantic knowledge base is constructed by leveraging a frozen large language model with visual prompts, where raw images or middle video frames are jointly used with textual semantics to mitigate semantic ambiguity and compensate for semantic degradation caused by wireless channel impairments. The enhanced semantics are encoded and transmitted over wireless channels, and multimodal signals are reconstructed at the receiver via diffusion-based generation. Extensive experiments on real-world datasets show that CMSC achieves a higher compression rate while maintaining comparable transmission accuracy under both arbitrary and typical multimodal combinations, demonstrating its effectiveness for flexible and lightweight multimodal semantic communication in 6G ubiquitous intelligence.
KW - 6G ubiquitous intelligence
KW - composable multimodal representation
KW - lightweight large AI model
KW - Multimodal semantic communication
KW - semantic knowledge base with visual prompts
UR - https://www.scopus.com/pages/publications/105044718711
U2 - 10.1109/TCOMM.2026.3707756
DO - 10.1109/TCOMM.2026.3707756
M3 - 文章
AN - SCOPUS:105044718711
SN - 0090-6778
JO - IEEE Transactions on Communications
JF - IEEE Transactions on Communications
ER -