TY - GEN
T1 - Knowledge Distillation Based Dueling DQN for Real-Time Power Grid Dispatching Operation and Control
AU - Wang, Xiaopeng
AU - Lu, Na
AU - Song, Yunpeng
N1 - Publisher Copyright:
© 2024 IEEE.
PY - 2024
Y1 - 2024
N2 - Deep reinforcement learning (DRL) has proven to be effective for real-time Power grid Dispatching and Control (PDC) but requires intensive computing resources and has low efficiency. Most research employs expertise demonstrations to accelerate the training process, but neglects the rules derived from the experience. To boost the DRL efficiency with PDC rules, a Knowledge Distillation based Dueling Deep Q-Iearning (KD3QN) algorithm is proposed which uses behavioral cloning to lead the agent to learn from PDC rules. Behavioral cloning may introduce bias in the policy, thus knowledge distillation is employed to extract the knowledge instead of directly following the rules. Besides, a balance strategy is proposed to control the interference of rules. To verify the proposed algorithm, extensive experiments have been performed on Grid20p platform with the 14- bus and 36- bus cases. The experimental results demonstrate that the proposed algorithm increases the training efficiency and yields agents with better performance.
AB - Deep reinforcement learning (DRL) has proven to be effective for real-time Power grid Dispatching and Control (PDC) but requires intensive computing resources and has low efficiency. Most research employs expertise demonstrations to accelerate the training process, but neglects the rules derived from the experience. To boost the DRL efficiency with PDC rules, a Knowledge Distillation based Dueling Deep Q-Iearning (KD3QN) algorithm is proposed which uses behavioral cloning to lead the agent to learn from PDC rules. Behavioral cloning may introduce bias in the policy, thus knowledge distillation is employed to extract the knowledge instead of directly following the rules. Besides, a balance strategy is proposed to control the interference of rules. To verify the proposed algorithm, extensive experiments have been performed on Grid20p platform with the 14- bus and 36- bus cases. The experimental results demonstrate that the proposed algorithm increases the training efficiency and yields agents with better performance.
KW - Behavioral cloning
KW - deep reinforcement learning
KW - knowledge distillation
KW - power grid dispatching operation and control
UR - https://www.scopus.com/pages/publications/105000016914
U2 - 10.1109/APET63768.2024.10882648
DO - 10.1109/APET63768.2024.10882648
M3 - 会议稿件
AN - SCOPUS:105000016914
T3 - 2024 3rd Asia Power and Electrical Technology Conference, APET 2024
SP - 485
EP - 490
BT - 2024 3rd Asia Power and Electrical Technology Conference, APET 2024
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 3rd Asia Power and Electrical Technology Conference, APET 2024
Y2 - 15 November 2024 through 17 November 2024
ER -