跳到主要导航 跳到搜索 跳到主要内容

SMART: Sliding Manifold Action Rectification Technique for Safe Reinforcement Learning

  • Jingyu Liu
  • , Yuanda Wang
  • , Zewen He
  • , Anqing Duan
  • , Yoshihiko Nakamura
  • , Changyin Sun
  • Southeast University, Nanjing
  • Mohamed Bin Zayed University of Artificial Intelligence
  • Anhui University

科研成果: 期刊稿件文章同行评审

摘要

Safe reinforcement learning for deployed robotic and automation systems requires stepwise enforcement of state constraints under sampled-data execution, model mismatch, and occasionally unsafe initialization. Lightweight action-rectification layers fit online robotic control well, but they leave the recovery behavior of the executed controller largely implicit. We present sliding manifold action rectification technique (SMART), a new sampled-data action-rectification framework that makes constraint-residual recovery explicit through reaching laws defined on a slack-augmented residual. At each control step, SMART evaluates a closed-form solution of an equality-constrained quadratic rectification problem at a projected slack state, minimally modifying the policy action while shaping how violations return to admissible operation. For the nominal linear reaching law, we distinguish the ideal continuously recomputed surrogate from the executed projected sampled-data controller, establish local surrogate invariance, derive local practical residual-tube bounds under bounded implementation mismatch, and show how these bounds imply recovery of the original inequality constraints when the residual tube remains within the slack margin. In navigation and planar air hockey, SMART recovers more effectively than Acting on the TAngent space of the COnstraint Manifold on the reported recovery metrics without degrading the reported safety or task results. A Unitree H1 case study further shows that the same rectifier can be used during online fine-tuning.

源语言英语
期刊IEEE Transactions on Industrial Informatics
DOI
出版状态已接受/待刊 - 2026
已对外发布

学术指纹

探究 'SMART: Sliding Manifold Action Rectification Technique for Safe Reinforcement Learning' 的科研主题。它们共同构成独一无二的学术指纹。

引用此