TY - GEN
T1 - Adaptive Beam Hopping for Over-the-Air Online Federated Learning in LEO Satellite Networks
AU - Wang, Shaojie
AU - Li, Zhendong
AU - Su, Zhou
AU - Zhang, Zihao
AU - Peng, Haixia
AU - Cheng, Nan
AU - Chen, Wen
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - This paper investigates over-the-air (OTA) computation federated learning (FL) in low-earth orbit (LEO) satellite networks, where we propose a novel joint design of adaptive beam hopping and power control to maximize the long-term total amount of training data. This problem is challenging due to the non-convex coupling between beam hopping patterns, power control, and global mean squared error (MSE) constraints, as well as the dynamic satellite coverage. To address these difficulties, we develop a proximal policy optimization (PPO)-based deep reinforcement learning framework that learns efficient scheduling policies. In particular, the proposed method jointly optimizes beam hopping patterns and transmission power, while incorporating MSE-aware reward shaping to balance data utilization and aggregation accuracy. Simulation results demonstrate that the proposed PPO approach achieves faster convergence, higher reward, and superior FL performance in terms of test accuracy and training loss, compared with soft actor-critic (SAC), deep deterministic policy gradient (DDPG), and a greedy baseline.
AB - This paper investigates over-the-air (OTA) computation federated learning (FL) in low-earth orbit (LEO) satellite networks, where we propose a novel joint design of adaptive beam hopping and power control to maximize the long-term total amount of training data. This problem is challenging due to the non-convex coupling between beam hopping patterns, power control, and global mean squared error (MSE) constraints, as well as the dynamic satellite coverage. To address these difficulties, we develop a proximal policy optimization (PPO)-based deep reinforcement learning framework that learns efficient scheduling policies. In particular, the proposed method jointly optimizes beam hopping patterns and transmission power, while incorporating MSE-aware reward shaping to balance data utilization and aggregation accuracy. Simulation results demonstrate that the proposed PPO approach achieves faster convergence, higher reward, and superior FL performance in terms of test accuracy and training loss, compared with soft actor-critic (SAC), deep deterministic policy gradient (DDPG), and a greedy baseline.
KW - beam hopping
KW - federated learning
KW - LEO satellite networks
KW - OTA computation
KW - PPO
UR - https://www.scopus.com/pages/publications/105042768988
U2 - 10.1109/WCNC65185.2026.11555370
DO - 10.1109/WCNC65185.2026.11555370
M3 - 会议稿件
AN - SCOPUS:105042768988
T3 - IEEE Wireless Communications and Networking Conference, WCNC
BT - 2026 IEEE Wireless Communications and Networking Conference, WCNC 2026
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2026 IEEE Wireless Communications and Networking Conference, WCNC 2026
Y2 - 13 April 2026 through 16 April 2026
ER -