Skip to main navigation Skip to search Skip to main content

Improving Mandarin ASR Performance Through Multimodality

  • Xi'an Jiaotong University

Research output: Contribution to journalArticlepeer-review

Abstract

In the context of Internet of Things (IoT) applications, accurate and efficient speech recognition is essential for enabling seamless voice-based interactions and control. Mandarin ASR, in particular, presents unique challenges due to the ideographic nature of the Chinese language, where recognition results are not directly correlated with pronunciation. Pinyin, as a representation of Chinese character pronunciation, has an intrinsic connection with Chinese characters, making it a valuable tool for enhancing ASR performance. This paper proposes a multimodal ASR neural network that combines pinyin data from the text modality and speech data from the audio modality as shared inputs to the ASR model. Specifically, the system processes the speech input through a preprocessed WeNet to generate pinyin text, which is then enhanced using a label denoising algorithm to improve its accuracy. The proposed text-acoustic multimodal ASR model improves the overall speech recognition performance by approximately 4%, making it more suitable for IoT applications that require high accuracy in voice commands and interactions.

Original languageEnglish
Article number12224
JournalApplied Sciences (Switzerland)
Volume15
Issue number22
DOIs
StatePublished - Nov 2025

Keywords

  • automatic speech recognition
  • label denoise
  • multimodal

Fingerprint

Dive into the research topics of 'Improving Mandarin ASR Performance Through Multimodality'. Together they form a unique fingerprint.

Cite this