摘要
Safe reinforcement learning (RL) faces a fundamental challenge: how can agents maximize rewards while guaranteeing safety during both training and deployment? Traditional approaches either sacrifice performance for safety or risk constraint violations during learning. We propose Safe Diffusion Langevin Policy (SDLP), a framework that enables agents to ”think before they act” by explicitly modeling long-term consequences. SDLP uses a Diffusion World Model (DWM) to generate diverse possible future trajectories for candidate actions, evaluates these through a principled trajectory evaluation function balancing rewards against safety violations, and iteratively refines actions using Langevin dynamics guided by gradient information. This approach shifts safe RL from reactive constraint handling to proactive safety reasoning. Instead of hoping agents learn to avoid violations through trial and error, SDLP explicitly models future consequences and optimizes actions accordingly. Experimental validation on challenging continuous control and navigation benchmarks confirms this advantage: SDLP achieves superior task performance while maintaining consistently low constraint violations, significantly outperforming Soft Actor-Critic (SAC), SafeLayer, and Control Barrier Functions (CBF) across all tested environments.
| 源语言 | 英语 |
|---|---|
| 文章编号 | 131278 |
| 期刊 | Expert Systems with Applications |
| 卷 | 310 |
| DOI | |
| 出版状态 | 已出版 - 10 5月 2026 |
| 已对外发布 | 是 |
学术指纹
探究 'Safe online reinforcement learning with diffusion world model and Langevin dynamics' 的科研主题。它们共同构成独一无二的指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver