Abstract
Neutral network (NN) and clustering are the two commonly used methods for speech separation based on supervised learning. Recently, deep clustering methods have shown promising performance. In our study, considering that the spectrum of the sound source has time correlation, and the spatial position of the sound source has short-term stability, we combine the spectral and spatial features for deep clustering. In this work, the logarithmic amplitude spectrum (LPS) and the interaural phase difference (IPD) function of each time frequency (TF) unit for the binaural speech signal are extracted as feature. Then, these features of consecutive frames construct feature map, which are regarded as the input to the Bi-directional long short-term memory (BiLSTM). The feature maps are con-verted to the high-dimensional vectors through BiLSTM, which are used to clas-sify the time-frequency units by K-means clustering. The clustering index are combined with mixed speech signal to reconstruct the target speech signal. The simulation results show that the proposed algorithm has a significant improve-ment in speech separation and speech quality, since the spectral and spatial information are all utilized for clustering. Also, the method is more generalized in untrained conditions compared with traditional NN method e.g., deep neural network (DNN) and convolutional neural networks (CNN) based method.
| Original language | English |
|---|---|
| Pages (from-to) | 527-537 |
| Number of pages | 11 |
| Journal | Intelligent Automation and Soft Computing |
| Volume | 30 |
| Issue number | 2 |
| DOIs | |
| State | Published - 2021 |
| Externally published | Yes |
Keywords
- BiLSTM
- Binaural speech separation
- K-means clustering
Fingerprint
Dive into the research topics of 'Binaural speech separation algorithm based on deep clustering'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver