TY - GEN
T1 - A Game-Theoretic Multi-Agent Reinforcement Learning Approach for Secure UAV-Assisted QKD in Vehicular Networks
AU - Cui, Qingfan
AU - Xu, Qichao
AU - Su, Zhou
AU - Fang, Dongfeng
AU - Zeng, Hui
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Quantum key distribution (QKD) enables secure communication for vehicular networks, but its range limitation calls for multi-hop relaying. Unmanned aerial vehicles (UAVs) can serve as flexible QKD relays, yet their selfish behavior and dynamic mobility pose significant challenges. This paper proposes a secure routing framework for UAV-assisted vehicular QKD that jointly addresses incentive compatibility and path reliability. We model the interaction between the source vehicle and UAVs as a Stackelberg game, where the source sets the relay reward and UAVs decide their participation. A smoothed relay path formulation ensures differentiability and equilibrium tractability. To enable decentralized learning of relay strategies, we adopt a multi-agent deep deterministic policy gradient (MADDPG) algorithm. Furthermore, a Bayesian optimization module is integrated to adaptively search for the optimal reward value that maximizes the source vehicle's utility. Simulation results show that the proposed Stackelberg-MADDPG-Bayesian scheme outperforms existing baselines in terms of relay request success rate and secure key generation rate, and remains robust under various UAV selfishness levels and network dynamics.
AB - Quantum key distribution (QKD) enables secure communication for vehicular networks, but its range limitation calls for multi-hop relaying. Unmanned aerial vehicles (UAVs) can serve as flexible QKD relays, yet their selfish behavior and dynamic mobility pose significant challenges. This paper proposes a secure routing framework for UAV-assisted vehicular QKD that jointly addresses incentive compatibility and path reliability. We model the interaction between the source vehicle and UAVs as a Stackelberg game, where the source sets the relay reward and UAVs decide their participation. A smoothed relay path formulation ensures differentiability and equilibrium tractability. To enable decentralized learning of relay strategies, we adopt a multi-agent deep deterministic policy gradient (MADDPG) algorithm. Furthermore, a Bayesian optimization module is integrated to adaptively search for the optimal reward value that maximizes the source vehicle's utility. Simulation results show that the proposed Stackelberg-MADDPG-Bayesian scheme outperforms existing baselines in terms of relay request success rate and secure key generation rate, and remains robust under various UAV selfishness levels and network dynamics.
KW - multi-agent deep deterministic policy gradient (MADDPG)
KW - quantum key distribution (QKD)
KW - Stackelberg game
KW - Unmanned aerial vehicle (UAV)-assisted vehicular networks
UR - https://www.scopus.com/pages/publications/105033652425
U2 - 10.1109/WCSP68525.2025.1010165
DO - 10.1109/WCSP68525.2025.1010165
M3 - 会议稿件
AN - SCOPUS:105033652425
T3 - 2025 17th International Conference on Wireless Communications and Signal Processing, WCSP 2025
BT - 2025 17th International Conference on Wireless Communications and Signal Processing, WCSP 2025
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2025 17th International Conference on Wireless Communications and Signal Processing, WCSP 2025
Y2 - 23 October 2025 through 25 October 2025
ER -