跳到主要导航 跳到搜索 跳到主要内容

A Singing Melody Extraction Network Via Self-Distillation and Multi-Level Supervision

  • Key Laboratory of Signal Detection and Processing
  • Xinjiang University
  • Ltd.

科研成果: 期刊稿件会议文章同行评审

2 引用 (Scopus)

摘要

Extracting singing melody from polyphonic music is an important topic in the field of music information retrieval. In this paper, we propose a singing melody extraction network consisting of five stacked multi-scale feature time-frequency aggregation (MF-TFA) modules. In the same network, deeper layers generally contain more contextual information than shallower layers. To help the shallower layers enhance the ability of task-relevant feature extraction, we propose a self-distillation and multi-level supervision (SD-MS) method, which leverages the feature distillation from the deepest layer to the shallower one and multi-level supervision to guide network training. Visualization analysis shows that by introducing SD-MS, the same-level layer in the network can obtain a clearer representation of fundamental frequency components, while the shallower layers can even learn more task-relevant semantic information. Ablation study results indicate that SD-MS applies to existing melody extraction models and can consistently improve performance. Experimental results show that our proposed method, MF-TFA with SD-MS, outperforms six compared state-of-the-art methods, achieving overall accuracy (OA) scores of 87.1%, 89.9%, and 76.6% on the ADC 2004, MIREX 05, and MEDLEY DB datasets, respectively. The main code will be available at https://github.com/SmoothJing/MFTFA_SD-MS.

学术指纹

探究 'A Singing Melody Extraction Network Via Self-Distillation and Multi-Level Supervision' 的科研主题。它们共同构成独一无二的指纹。

引用此