TY - JOUR
T1 - Algorithmic trading by reinforcement learning in a collaborative manner
AU - Long, Li
AU - Zhang, Chunxia
AU - Ma, Cong
AU - Wang, Hongtao
AU - Ji, Lizhen
AU - Gao, Fei
AU - Zhang, Jiangshe
AU - Qiu, Kaiwen
N1 - Publisher Copyright:
© 2026 Elsevier B.V.
PY - 2026/7
Y1 - 2026/7
N2 - In recent years, reinforcement learning has emerged as a prominent approach in algorithmic trading, with most studies concentrating on time series processing and neural network architecture optimization within single-agent frameworks. Some research has confirmed that multi-agent or ensemble learning can enhance performance, but conventional methods often train ensemble neural networks or diverse reinforcement learning algorithms on the entire time series data. These practices not only lead to high computational costs but also hinder the efficient recognition of market fluctuations. To address these challenges, this paper introduces a novel Collaborative Self-Imitation Double Deep Q Network (CSI-DDQN) framework. The proposed approach partitions long-term time series data into shorter segments and employs a distributed learning-inspired collaborative training strategy across multiple clients. A self-imitation strategy enhances cooperation among clients, while a hybrid multi-loss function integrates the strengths of reinforcement learning and imitation learning. Experimental results using a multi-layer perceptron (MLP) architecture on U.S. stock data demonstrate that the proposed method significantly outperforms several baselines, achieving the highest average cumulative return of 56.85%, the highest Sharpe ratio of 1.42, and the lowest maximum drawdown of 17.14%. In comparison, DQN-Vanilla attains an average cumulative return of 3.95% while a Sharpe ratio of 0.25, and the DRL-Ensemble method yields a cumulative return of 19.53% and a Sharpe ratio of 0.54. Furthermore, to demonstrate the flexibility and adaptability of the proposed framework, additional experiments were conducted by substituting various network architectures and reinforcement learning algorithms within the framework. The results further validate the effectiveness of the developed CSI-DDQN framework. The related code is available at https://github.com/ll-1/CSI-DDQN.
AB - In recent years, reinforcement learning has emerged as a prominent approach in algorithmic trading, with most studies concentrating on time series processing and neural network architecture optimization within single-agent frameworks. Some research has confirmed that multi-agent or ensemble learning can enhance performance, but conventional methods often train ensemble neural networks or diverse reinforcement learning algorithms on the entire time series data. These practices not only lead to high computational costs but also hinder the efficient recognition of market fluctuations. To address these challenges, this paper introduces a novel Collaborative Self-Imitation Double Deep Q Network (CSI-DDQN) framework. The proposed approach partitions long-term time series data into shorter segments and employs a distributed learning-inspired collaborative training strategy across multiple clients. A self-imitation strategy enhances cooperation among clients, while a hybrid multi-loss function integrates the strengths of reinforcement learning and imitation learning. Experimental results using a multi-layer perceptron (MLP) architecture on U.S. stock data demonstrate that the proposed method significantly outperforms several baselines, achieving the highest average cumulative return of 56.85%, the highest Sharpe ratio of 1.42, and the lowest maximum drawdown of 17.14%. In comparison, DQN-Vanilla attains an average cumulative return of 3.95% while a Sharpe ratio of 0.25, and the DRL-Ensemble method yields a cumulative return of 19.53% and a Sharpe ratio of 0.54. Furthermore, to demonstrate the flexibility and adaptability of the proposed framework, additional experiments were conducted by substituting various network architectures and reinforcement learning algorithms within the framework. The results further validate the effectiveness of the developed CSI-DDQN framework. The related code is available at https://github.com/ll-1/CSI-DDQN.
KW - Algorithmic trading
KW - Distributed learning
KW - Imitation learning
KW - Reinforcement learning
UR - https://www.scopus.com/pages/publications/105034617493
U2 - 10.1016/j.asoc.2026.115168
DO - 10.1016/j.asoc.2026.115168
M3 - 文章
AN - SCOPUS:105034617493
SN - 1568-4946
VL - 197
JO - Applied Soft Computing Journal
JF - Applied Soft Computing Journal
M1 - 115168
ER -