TY - GEN
T1 - Research on Controllable Music Generation Algorithm Based on Multi-branch Fusion
AU - Wei, Xinyao
AU - Li, Chen
AU - Tian, Lihua
AU - Zhu, Jihua
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2026.
PY - 2026
Y1 - 2026
N2 - With the accelerating progress of information technology, the demand for music composition continues to grow. However, existing music generation algorithms primarily focus on improving the quality of generated samples, with most methods offering only limited control over the generated sequences. To address this issue, this paper proposes a music generation algorithm based on multi-branch fusion. The algorithm enhances the diversity and quality of generated music by incorporating a melody description branch and fusing expert description features, learned description features, and melody description features through parallel cross-attention. To further optimize the model’s generative capabilities, this paper introduces the RoBERTa pre-trained model and a contrastive learning method based on instance discrimination. The contrastive learning method treats each sample as an independent category, maximizing the consistency of the same sample in the feature space while minimizing the similarity between different samples to learn discriminative representations. Based on the aforementioned research, comparative and ablation experiments were conducted on the LakhMIDI dataset. The results demonstrate that the proposed algorithm achieves improvements of 0.059 in chord accuracy, 0.056 in two cosine similarity metrics, and 0.055 in note density, validating the algorithm’s effectiveness and advantages (This work was supported by the Key Research and Development Program of Shaanxi under Grant No. 2024GX-YBXM-556).
AB - With the accelerating progress of information technology, the demand for music composition continues to grow. However, existing music generation algorithms primarily focus on improving the quality of generated samples, with most methods offering only limited control over the generated sequences. To address this issue, this paper proposes a music generation algorithm based on multi-branch fusion. The algorithm enhances the diversity and quality of generated music by incorporating a melody description branch and fusing expert description features, learned description features, and melody description features through parallel cross-attention. To further optimize the model’s generative capabilities, this paper introduces the RoBERTa pre-trained model and a contrastive learning method based on instance discrimination. The contrastive learning method treats each sample as an independent category, maximizing the consistency of the same sample in the feature space while minimizing the similarity between different samples to learn discriminative representations. Based on the aforementioned research, comparative and ablation experiments were conducted on the LakhMIDI dataset. The results demonstrate that the proposed algorithm achieves improvements of 0.059 in chord accuracy, 0.056 in two cosine similarity metrics, and 0.055 in note density, validating the algorithm’s effectiveness and advantages (This work was supported by the Key Research and Development Program of Shaanxi under Grant No. 2024GX-YBXM-556).
KW - attention mechanism
KW - contrastive learning
KW - music generation
KW - pre-trained model
UR - https://www.scopus.com/pages/publications/105028088980
U2 - 10.1007/978-981-95-4821-7_5
DO - 10.1007/978-981-95-4821-7_5
M3 - 会议稿件
AN - SCOPUS:105028088980
SN - 9789819548200
T3 - Communications in Computer and Information Science
SP - 58
EP - 67
BT - Artificial Intelligence and Robotics - 10th International Symposium, ISAIR 2025, Revised Selected Papers
A2 - Lu, Huimin
PB - Springer Science and Business Media Deutschland GmbH
T2 - 10th International Symposium on Artificial Intelligence and Robotics, ISAIR 2025
Y2 - 24 August 2025 through 26 August 2025
ER -