Abstract
Prioritizing or reweighting important samples has been recognized as an effective means of improving the efficiency of deep-reinforcement learning (DRL) algorithms. However, many existing techniques encounter stability challenges, limiting efficiency and increasing computational costs and training time. In this study, we aim to improve training efficiency by exploring the intrinsic relationship between sample efficiency and stability. To achieve this, we propose the Stability Contribution Index (SI), which assigns sample priorities based on their impact on stability and employs them to weight the value loss, thereby promoting stable learning and improving efficiency. The effectiveness of our method is validated through comprehensive experiments on two distinct benchmarks: 1) the continuous control domain DMControl and 2) the discrete control environment ProcGen. Compatible with both off-policy and on-policy DRL algorithms, our approach significantly improves sample efficiency and overall performance by fostering greater stability during training. Additionally, experimental results show that our method outperforms well-established sample-efficient reinforcement learning techniques across multiple settings.
| Original language | English |
|---|---|
| Pages (from-to) | 6164-6176 |
| Number of pages | 13 |
| Journal | IEEE Transactions on Systems, Man, and Cybernetics: Systems |
| Volume | 55 |
| Issue number | 9 |
| DOIs | |
| State | Published - 2025 |
Keywords
- Deep-reinforcement learning (DRL)
- prioritizing
- sample efficiency
- stability
Fingerprint
Dive into the research topics of 'Improving Sample Efficiency Through Stability Enhancement in Deep-Reinforcement Learning'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver