Skip to main navigation Skip to search Skip to main content

Improving Sample Efficiency Through Stability Enhancement in Deep-Reinforcement Learning

  • Ziru Wang
  • , Wanli Jiang
  • , Ru Peng
  • , Qian Kou
  • , Lipeng Wan
  • , Xuguang Lan
  • Xi'an Jiaotong University

Research output: Contribution to journalArticlepeer-review

3 Scopus citations

Abstract

Prioritizing or reweighting important samples has been recognized as an effective means of improving the efficiency of deep-reinforcement learning (DRL) algorithms. However, many existing techniques encounter stability challenges, limiting efficiency and increasing computational costs and training time. In this study, we aim to improve training efficiency by exploring the intrinsic relationship between sample efficiency and stability. To achieve this, we propose the Stability Contribution Index (SI), which assigns sample priorities based on their impact on stability and employs them to weight the value loss, thereby promoting stable learning and improving efficiency. The effectiveness of our method is validated through comprehensive experiments on two distinct benchmarks: 1) the continuous control domain DMControl and 2) the discrete control environment ProcGen. Compatible with both off-policy and on-policy DRL algorithms, our approach significantly improves sample efficiency and overall performance by fostering greater stability during training. Additionally, experimental results show that our method outperforms well-established sample-efficient reinforcement learning techniques across multiple settings.

Original languageEnglish
Pages (from-to)6164-6176
Number of pages13
JournalIEEE Transactions on Systems, Man, and Cybernetics: Systems
Volume55
Issue number9
DOIs
StatePublished - 2025

Keywords

  • Deep-reinforcement learning (DRL)
  • prioritizing
  • sample efficiency
  • stability

Fingerprint

Dive into the research topics of 'Improving Sample Efficiency Through Stability Enhancement in Deep-Reinforcement Learning'. Together they form a unique fingerprint.

Cite this