摘要
Speech enhancement can benefit lots of practical voice-based interaction applications, where the goal is to generate clean speech from noisy ambient conditions. This paper presents a practical design, namely UltraSpeech, to enhance speech by exploring the correlation between the ultrasound (profiled articulatory gestures) and speech. UltraSpeech uses a commodity smartphone to emit the ultrasound and collect the composed acoustic signal for analysis. We design a complex masking framework to deal with complex-valued spectrograms, incorporating the magnitude and phase rectification of speech simultaneously. We further introduce an interaction module to share information between ultrasound and speech two branches and thus enhance their discrimination capabilities. Extensive experiments demonstrate that UltraSpeech increases the Scale Invariant SDR by 12dB, improves the speech intelligibility and quality effectively, and is capable to generalize to unknown speakers.
| 源语言 | 英语 |
|---|---|
| 期刊论文编号 | 111 |
| 期刊 | Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies |
| 卷 | 6 |
| 期 | 3 |
| DOI | |
| 出版状态 | 已出版 - 7 9月 2022 |
学术指纹
探究 'UltraSpeech: Speech Enhancement by Interaction between Ultrasound and Speech' 的科研主题。它们共同构成独一无二的学术指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver