TY - JOUR
T1 - Improving Sample Efficiency Through Stability Enhancement in Deep-Reinforcement Learning
AU - Wang, Ziru
AU - Jiang, Wanli
AU - Peng, Ru
AU - Kou, Qian
AU - Wan, Lipeng
AU - Lan, Xuguang
N1 - Publisher Copyright:
© IEEE. 2013 IEEE.
PY - 2025
Y1 - 2025
N2 - Prioritizing or reweighting important samples has been recognized as an effective means of improving the efficiency of deep-reinforcement learning (DRL) algorithms. However, many existing techniques encounter stability challenges, limiting efficiency and increasing computational costs and training time. In this study, we aim to improve training efficiency by exploring the intrinsic relationship between sample efficiency and stability. To achieve this, we propose the Stability Contribution Index (SI), which assigns sample priorities based on their impact on stability and employs them to weight the value loss, thereby promoting stable learning and improving efficiency. The effectiveness of our method is validated through comprehensive experiments on two distinct benchmarks: 1) the continuous control domain DMControl and 2) the discrete control environment ProcGen. Compatible with both off-policy and on-policy DRL algorithms, our approach significantly improves sample efficiency and overall performance by fostering greater stability during training. Additionally, experimental results show that our method outperforms well-established sample-efficient reinforcement learning techniques across multiple settings.
AB - Prioritizing or reweighting important samples has been recognized as an effective means of improving the efficiency of deep-reinforcement learning (DRL) algorithms. However, many existing techniques encounter stability challenges, limiting efficiency and increasing computational costs and training time. In this study, we aim to improve training efficiency by exploring the intrinsic relationship between sample efficiency and stability. To achieve this, we propose the Stability Contribution Index (SI), which assigns sample priorities based on their impact on stability and employs them to weight the value loss, thereby promoting stable learning and improving efficiency. The effectiveness of our method is validated through comprehensive experiments on two distinct benchmarks: 1) the continuous control domain DMControl and 2) the discrete control environment ProcGen. Compatible with both off-policy and on-policy DRL algorithms, our approach significantly improves sample efficiency and overall performance by fostering greater stability during training. Additionally, experimental results show that our method outperforms well-established sample-efficient reinforcement learning techniques across multiple settings.
KW - Deep-reinforcement learning (DRL)
KW - prioritizing
KW - sample efficiency
KW - stability
UR - https://www.scopus.com/pages/publications/105009412910
U2 - 10.1109/TSMC.2025.3578050
DO - 10.1109/TSMC.2025.3578050
M3 - 文章
AN - SCOPUS:105009412910
SN - 2168-2216
VL - 55
SP - 6164
EP - 6176
JO - IEEE Transactions on Systems, Man, and Cybernetics: Systems
JF - IEEE Transactions on Systems, Man, and Cybernetics: Systems
IS - 9
ER -