跳到主要导航 跳到搜索 跳到主要内容

Generative Adversarial Self-Imitation Learning With Large Language Model Feedback for Robot Control and Navigation

  • Ke Zhang
  • , Zheng Fang
  • , Enqi Zhao
  • , Zicheng Sun
  • , Jianwu Fang
  • , Jie Huang
  • , Eric Nichols
  • , Randy Gomez
  • , Bo He
  • , Jianru Xue
  • , Guangliang Li
  • Ocean University of China
  • Research Institute of Civil Aviation Administration of China
  • Honda Motor Co., Ltd.
  • Xi'an Jiaotong University

科研成果: 期刊稿件文章同行评审

摘要

Deep reinforcement learning (DRL) has achieved great success in many simulated and real-world robotic tasks. However, the difficulty of designing efficient and dense reward functions makes applying DRL to tackle complex long-horizon and open-world tasks a great challenge. Generative adversarial imitation learning (GAIL) can directly learn policies from the expert trajectories and generalize well in large and complex environments, but relies on high-quality demonstrations and can seldom surpass the performance of the demonstration. Recent work used additional human evaluative feedback to facilitate GAIL to learn faster and surpass the demonstrations, but still requires sub-optimal demonstrations. Moreover, it is costly and difficult for human expert to provide relatively high-quality demonstrations and evaluative feedback for various tasks. To address the above issues, in this paper, we propose Generative Adversarial Self-Imitation Learning from Demonstration and Large Language Model (LLM) Feedback (GASL3MF), since LLMs encode rich commonsense knowledge and can perform a variety of reasoning tasks. GASL3MF allows a robot to learn from poor demonstrations and gradually replace them with its own good trajectories evaluated by LLM feedback. Our results in four physics-based control tasks and a mobile robot navigation task show that, even with demonstrations of poor performance or not completing the task, GASL3MF can learn faster with close to optimal performance, and generalize well to different environments and the real world with sim-to-real adaptation. Further analysis shows that the overall distribution of LLM feedback closely resembles that of human feedback and remains closer to that of ground-truth rewards than human feedback. Finally, our GASL3MF method works regardless of the LLM employed, and the LLM feedback from different LLMs remain robust across tasks and even better consistency than human feedback for robot learning in some tasks. These results shed light on the potential of robot imitation learning from even poor or failed demonstrations and broaden its application to a wide range of real-world tasks.

源语言英语
期刊IEEE Transactions on Robotics
DOI
出版状态已接受/待刊 - 2026

学术指纹

探究 'Generative Adversarial Self-Imitation Learning With Large Language Model Feedback for Robot Control and Navigation' 的科研主题。它们共同构成独一无二的学术指纹。

引用此