Abstract
In the context of Internet of Things (IoT) applications, accurate and efficient speech recognition is essential for enabling seamless voice-based interactions and control. Mandarin ASR, in particular, presents unique challenges due to the ideographic nature of the Chinese language, where recognition results are not directly correlated with pronunciation. Pinyin, as a representation of Chinese character pronunciation, has an intrinsic connection with Chinese characters, making it a valuable tool for enhancing ASR performance. This paper proposes a multimodal ASR neural network that combines pinyin data from the text modality and speech data from the audio modality as shared inputs to the ASR model. Specifically, the system processes the speech input through a preprocessed WeNet to generate pinyin text, which is then enhanced using a label denoising algorithm to improve its accuracy. The proposed text-acoustic multimodal ASR model improves the overall speech recognition performance by approximately 4%, making it more suitable for IoT applications that require high accuracy in voice commands and interactions.
| Original language | English |
|---|---|
| Article number | 12224 |
| Journal | Applied Sciences (Switzerland) |
| Volume | 15 |
| Issue number | 22 |
| DOIs | |
| State | Published - Nov 2025 |
Keywords
- automatic speech recognition
- label denoise
- multimodal
Fingerprint
Dive into the research topics of 'Improving Mandarin ASR Performance Through Multimodality'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver