TY - JOUR
T1 - Alleviating imbalanced problems of reinforcement learning when applying in real-time power network dispatching and control
AU - Wang, Xiaopeng
AU - Lu, Na
N1 - Publisher Copyright:
© 2024 Elsevier Ltd
PY - 2024/12/1
Y1 - 2024/12/1
N2 - Real-time power network dispatching and control (PDC) presents unique challenges that traditional methods cannot effectively address due to the consideration of temporal dynamic factors. Reinforcement learning (RL) has been introduced and proven to be effective. However, given the vast space for solutions, there is an urgent need to enhance the performance of RL further. Analyzing the characteristics of the power network to improve the performance of RL holds significant potential. This article presents a comprehensive analysis of power network characteristics, including imbalanced state distribution and the imbalanced action operation frequency, along with their impact on applying RL in real-time PDC to guide the algorithm design. Guided by the analysis, the paper proposes the Balance Deep Q-Network (DQN) algorithm to mitigate the negative impact of the imbalanced problems on algorithm performance. A K-means-based predefined curriculum learning (KPCL) module is proposed to address the state imbalance issue. It allows for differentiated access to operation scenarios based on their hardness, ensuring a balanced exploration of operation scenarios while maintaining diversity during training. An Option DQN module is proposed to tackle the action imbalance problem, utilizing a branching network structure and hierarchical action selection to reduce the coupling between actions with different usage frequencies, thereby enhancing the accuracy of action value estimation. The effectiveness of the Balance DQN algorithm is demonstrated through extensive experiments conducted on the Grid2Op platform with 14-bus and 36-bus cases. The results indicate that Balance DQN can efficiently alleviate the negative impacts of imbalanced problems on RL for real-time PDC agents and achieve good performance in both cases.
AB - Real-time power network dispatching and control (PDC) presents unique challenges that traditional methods cannot effectively address due to the consideration of temporal dynamic factors. Reinforcement learning (RL) has been introduced and proven to be effective. However, given the vast space for solutions, there is an urgent need to enhance the performance of RL further. Analyzing the characteristics of the power network to improve the performance of RL holds significant potential. This article presents a comprehensive analysis of power network characteristics, including imbalanced state distribution and the imbalanced action operation frequency, along with their impact on applying RL in real-time PDC to guide the algorithm design. Guided by the analysis, the paper proposes the Balance Deep Q-Network (DQN) algorithm to mitigate the negative impact of the imbalanced problems on algorithm performance. A K-means-based predefined curriculum learning (KPCL) module is proposed to address the state imbalance issue. It allows for differentiated access to operation scenarios based on their hardness, ensuring a balanced exploration of operation scenarios while maintaining diversity during training. An Option DQN module is proposed to tackle the action imbalance problem, utilizing a branching network structure and hierarchical action selection to reduce the coupling between actions with different usage frequencies, thereby enhancing the accuracy of action value estimation. The effectiveness of the Balance DQN algorithm is demonstrated through extensive experiments conducted on the Grid2Op platform with 14-bus and 36-bus cases. The results indicate that Balance DQN can efficiently alleviate the negative impacts of imbalanced problems on RL for real-time PDC agents and achieve good performance in both cases.
KW - Action imbalance
KW - Curriculum learning
KW - Power network dispatching and control
KW - Reinforcement learning
KW - State imbalance
UR - https://www.scopus.com/pages/publications/85198507499
U2 - 10.1016/j.eswa.2024.124730
DO - 10.1016/j.eswa.2024.124730
M3 - 文章
AN - SCOPUS:85198507499
SN - 0957-4174
VL - 255
JO - Expert Systems with Applications
JF - Expert Systems with Applications
M1 - 124730
ER -