TY - JOUR
T1 - Generative Adversarial Self-Imitation Learning With Large Language Model Feedback for Robot Control and Navigation
AU - Zhang, Ke
AU - Fang, Zheng
AU - Zhao, Enqi
AU - Sun, Zicheng
AU - Fang, Jianwu
AU - Huang, Jie
AU - Nichols, Eric
AU - Gomez, Randy
AU - He, Bo
AU - Xue, Jianru
AU - Li, Guangliang
N1 - Publisher Copyright:
© 2004-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - Deep reinforcement learning (DRL) has achieved great success in many simulated and real-world robotic tasks. However, the difficulty of designing efficient and dense reward functions makes applying DRL to tackle complex long-horizon and open-world tasks a great challenge. Generative adversarial imitation learning (GAIL) can directly learn policies from the expert trajectories and generalize well in large and complex environments, but relies on high-quality demonstrations and can seldom surpass the performance of the demonstration. Recent work used additional human evaluative feedback to facilitate GAIL to learn faster and surpass the demonstrations, but still requires sub-optimal demonstrations. Moreover, it is costly and difficult for human expert to provide relatively high-quality demonstrations and evaluative feedback for various tasks. To address the above issues, in this paper, we propose Generative Adversarial Self-Imitation Learning from Demonstration and Large Language Model (LLM) Feedback (GASL3MF), since LLMs encode rich commonsense knowledge and can perform a variety of reasoning tasks. GASL3MF allows a robot to learn from poor demonstrations and gradually replace them with its own good trajectories evaluated by LLM feedback. Our results in four physics-based control tasks and a mobile robot navigation task show that, even with demonstrations of poor performance or not completing the task, GASL3MF can learn faster with close to optimal performance, and generalize well to different environments and the real world with sim-to-real adaptation. Further analysis shows that the overall distribution of LLM feedback closely resembles that of human feedback and remains closer to that of ground-truth rewards than human feedback. Finally, our GASL3MF method works regardless of the LLM employed, and the LLM feedback from different LLMs remain robust across tasks and even better consistency than human feedback for robot learning in some tasks. These results shed light on the potential of robot imitation learning from even poor or failed demonstrations and broaden its application to a wide range of real-world tasks.
AB - Deep reinforcement learning (DRL) has achieved great success in many simulated and real-world robotic tasks. However, the difficulty of designing efficient and dense reward functions makes applying DRL to tackle complex long-horizon and open-world tasks a great challenge. Generative adversarial imitation learning (GAIL) can directly learn policies from the expert trajectories and generalize well in large and complex environments, but relies on high-quality demonstrations and can seldom surpass the performance of the demonstration. Recent work used additional human evaluative feedback to facilitate GAIL to learn faster and surpass the demonstrations, but still requires sub-optimal demonstrations. Moreover, it is costly and difficult for human expert to provide relatively high-quality demonstrations and evaluative feedback for various tasks. To address the above issues, in this paper, we propose Generative Adversarial Self-Imitation Learning from Demonstration and Large Language Model (LLM) Feedback (GASL3MF), since LLMs encode rich commonsense knowledge and can perform a variety of reasoning tasks. GASL3MF allows a robot to learn from poor demonstrations and gradually replace them with its own good trajectories evaluated by LLM feedback. Our results in four physics-based control tasks and a mobile robot navigation task show that, even with demonstrations of poor performance or not completing the task, GASL3MF can learn faster with close to optimal performance, and generalize well to different environments and the real world with sim-to-real adaptation. Further analysis shows that the overall distribution of LLM feedback closely resembles that of human feedback and remains closer to that of ground-truth rewards than human feedback. Finally, our GASL3MF method works regardless of the LLM employed, and the LLM feedback from different LLMs remain robust across tasks and even better consistency than human feedback for robot learning in some tasks. These results shed light on the potential of robot imitation learning from even poor or failed demonstrations and broaden its application to a wide range of real-world tasks.
KW - Imitation learning
KW - inverse reinforcement learning
KW - large language model
KW - learning from demonstration
KW - mobile robot navigation
UR - https://www.scopus.com/pages/publications/105044380419
U2 - 10.1109/TRO.2026.3710412
DO - 10.1109/TRO.2026.3710412
M3 - 文章
AN - SCOPUS:105044380419
SN - 1552-3098
JO - IEEE Transactions on Robotics
JF - IEEE Transactions on Robotics
ER -