TY - JOUR
T1 - SMART
T2 - Sliding Manifold Action Rectification Technique for Safe Reinforcement Learning
AU - Liu, Jingyu
AU - Wang, Yuanda
AU - He, Zewen
AU - Duan, Anqing
AU - Nakamura, Yoshihiko
AU - Sun, Changyin
N1 - Publisher Copyright:
© 2005-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - Safe reinforcement learning for deployed robotic and automation systems requires stepwise enforcement of state constraints under sampled-data execution, model mismatch, and occasionally unsafe initialization. Lightweight action-rectification layers fit online robotic control well, but they leave the recovery behavior of the executed controller largely implicit. We present sliding manifold action rectification technique (SMART), a new sampled-data action-rectification framework that makes constraint-residual recovery explicit through reaching laws defined on a slack-augmented residual. At each control step, SMART evaluates a closed-form solution of an equality-constrained quadratic rectification problem at a projected slack state, minimally modifying the policy action while shaping how violations return to admissible operation. For the nominal linear reaching law, we distinguish the ideal continuously recomputed surrogate from the executed projected sampled-data controller, establish local surrogate invariance, derive local practical residual-tube bounds under bounded implementation mismatch, and show how these bounds imply recovery of the original inequality constraints when the residual tube remains within the slack margin. In navigation and planar air hockey, SMART recovers more effectively than Acting on the TAngent space of the COnstraint Manifold on the reported recovery metrics without degrading the reported safety or task results. A Unitree H1 case study further shows that the same rectifier can be used during online fine-tuning.
AB - Safe reinforcement learning for deployed robotic and automation systems requires stepwise enforcement of state constraints under sampled-data execution, model mismatch, and occasionally unsafe initialization. Lightweight action-rectification layers fit online robotic control well, but they leave the recovery behavior of the executed controller largely implicit. We present sliding manifold action rectification technique (SMART), a new sampled-data action-rectification framework that makes constraint-residual recovery explicit through reaching laws defined on a slack-augmented residual. At each control step, SMART evaluates a closed-form solution of an equality-constrained quadratic rectification problem at a projected slack state, minimally modifying the policy action while shaping how violations return to admissible operation. For the nominal linear reaching law, we distinguish the ideal continuously recomputed surrogate from the executed projected sampled-data controller, establish local surrogate invariance, derive local practical residual-tube bounds under bounded implementation mismatch, and show how these bounds imply recovery of the original inequality constraints when the residual tube remains within the slack margin. In navigation and planar air hockey, SMART recovers more effectively than Acting on the TAngent space of the COnstraint Manifold on the reported recovery metrics without degrading the reported safety or task results. A Unitree H1 case study further shows that the same rectifier can be used during online fine-tuning.
KW - Action rectification
KW - constrained Markov decision process (CMPD)
KW - safe reinforcement learning (RL)
KW - sliding mode control (SMC)
UR - https://www.scopus.com/pages/publications/105047003185
U2 - 10.1109/TII.2026.3705631
DO - 10.1109/TII.2026.3705631
M3 - 文章
AN - SCOPUS:105047003185
SN - 1551-3203
JO - IEEE Transactions on Industrial Informatics
JF - IEEE Transactions on Industrial Informatics
ER -