Skip to main navigation Skip to search Skip to main content

Generative Adversarial Self-Imitation Learning With Large Language Model Feedback for Robot Control and Navigation

  • Ke Zhang
  • , Zheng Fang
  • , Enqi Zhao
  • , Zicheng Sun
  • , Jianwu Fang
  • , Jie Huang
  • , Eric Nichols
  • , Randy Gomez
  • , Bo He
  • , Jianru Xue
  • , Guangliang Li
  • Ocean University of China
  • Research Institute of Civil Aviation Administration of China
  • Honda Motor Co., Ltd.
  • Xi'an Jiaotong University

Research output: Contribution to journalArticlepeer-review

Abstract

Deep reinforcement learning (DRL) has achieved great success in many simulated and real-world robotic tasks. However, the difficulty of designing efficient and dense reward functions makes applying DRL to tackle complex long-horizon and open-world tasks a great challenge. Generative adversarial imitation learning (GAIL) can directly learn policies from the expert trajectories and generalize well in large and complex environments, but relies on high-quality demonstrations and can seldom surpass the performance of the demonstration. Recent work used additional human evaluative feedback to facilitate GAIL to learn faster and surpass the demonstrations, but still requires sub-optimal demonstrations. Moreover, it is costly and difficult for human expert to provide relatively high-quality demonstrations and evaluative feedback for various tasks. To address the above issues, in this paper, we propose Generative Adversarial Self-Imitation Learning from Demonstration and Large Language Model (LLM) Feedback (GASL3MF), since LLMs encode rich commonsense knowledge and can perform a variety of reasoning tasks. GASL3MF allows a robot to learn from poor demonstrations and gradually replace them with its own good trajectories evaluated by LLM feedback. Our results in four physics-based control tasks and a mobile robot navigation task show that, even with demonstrations of poor performance or not completing the task, GASL3MF can learn faster with close to optimal performance, and generalize well to different environments and the real world with sim-to-real adaptation. Further analysis shows that the overall distribution of LLM feedback closely resembles that of human feedback and remains closer to that of ground-truth rewards than human feedback. Finally, our GASL3MF method works regardless of the LLM employed, and the LLM feedback from different LLMs remain robust across tasks and even better consistency than human feedback for robot learning in some tasks. These results shed light on the potential of robot imitation learning from even poor or failed demonstrations and broaden its application to a wide range of real-world tasks.

Original languageEnglish
JournalIEEE Transactions on Robotics
DOIs
StateAccepted/In press - 2026

Keywords

  • Imitation learning
  • inverse reinforcement learning
  • large language model
  • learning from demonstration
  • mobile robot navigation

Fingerprint

Dive into the research topics of 'Generative Adversarial Self-Imitation Learning With Large Language Model Feedback for Robot Control and Navigation'. Together they form a unique fingerprint.

Cite this