跳到主要导航 跳到搜索 跳到主要内容

Safe online reinforcement learning with diffusion world model and Langevin dynamics

  • Southeast University, Nanjing
  • Mohamed Bin Zayed University of Artificial Intelligence
  • Key Lab of the Ministry of Education for Process Control and Efficiency Egineering
  • Anhui University

科研成果: 期刊稿件文章同行评审

摘要

Safe reinforcement learning (RL) faces a fundamental challenge: how can agents maximize rewards while guaranteeing safety during both training and deployment? Traditional approaches either sacrifice performance for safety or risk constraint violations during learning. We propose Safe Diffusion Langevin Policy (SDLP), a framework that enables agents to ”think before they act” by explicitly modeling long-term consequences. SDLP uses a Diffusion World Model (DWM) to generate diverse possible future trajectories for candidate actions, evaluates these through a principled trajectory evaluation function balancing rewards against safety violations, and iteratively refines actions using Langevin dynamics guided by gradient information. This approach shifts safe RL from reactive constraint handling to proactive safety reasoning. Instead of hoping agents learn to avoid violations through trial and error, SDLP explicitly models future consequences and optimizes actions accordingly. Experimental validation on challenging continuous control and navigation benchmarks confirms this advantage: SDLP achieves superior task performance while maintaining consistently low constraint violations, significantly outperforming Soft Actor-Critic (SAC), SafeLayer, and Control Barrier Functions (CBF) across all tested environments.

源语言英语
文章编号131278
期刊Expert Systems with Applications
310
DOI
出版状态已出版 - 10 5月 2026
已对外发布

学术指纹

探究 'Safe online reinforcement learning with diffusion world model and Langevin dynamics' 的科研主题。它们共同构成独一无二的指纹。

引用此