TY - JOUR
T1 - Learning Natural Pedestrian Policies via Human Demonstration for Human-Vehicle Interaction
AU - Zhang, Chi
AU - Long, Tingting
AU - Yu, Feiyang
AU - Xu, Linhai
AU - Liu, Yuehu
AU - Ye, Qing
AU - Wang, Le
AU - Li, Li
N1 - Publisher Copyright:
© 2000-2011 IEEE.
PY - 2026
Y1 - 2026
N2 - Existing simulators still struggle to reproduce realistic human-vehicle interaction scenarios due to the lack of natural representation of human pedestrian movement. While VR-based human demonstration enables capturing realistic pedestrian experiences, it remains challenging to convert these sparse VR observations into scalable learning strategies for natural pedestrian behaviors. Our method utilizes a hierarchical spatio-temporal diffusion model to reconstruct full-body pedestrian poses from sparse VR demonstration signals, with its backbone network alternately encoding temporal and spatial dimensions. To improve the consistency of estimated posture, we incorporate concatenated conditional signals and repeated time-step embeddings. Furthermore, to bridge the gap between demonstrations and imitative behavior, we incorporate generative adversarial imitation learning (GAIL) to learn physically-constrained policies, enabling virtual avatars to imitate realistic behaviors. Experimental results on the AMASS dataset show our method significantly outperforms existing approaches in both position and rotation accuracy. We further validate our framework in a human-vehicle interaction environment, where pedestrian avatars reconstruct VR users’ demonstrated behaviors, exhibiting diverse and physically plausible actions such as walking, jumping, and crouching, highlighting the potential for more realistic human-vehicle interaction scenarios.
AB - Existing simulators still struggle to reproduce realistic human-vehicle interaction scenarios due to the lack of natural representation of human pedestrian movement. While VR-based human demonstration enables capturing realistic pedestrian experiences, it remains challenging to convert these sparse VR observations into scalable learning strategies for natural pedestrian behaviors. Our method utilizes a hierarchical spatio-temporal diffusion model to reconstruct full-body pedestrian poses from sparse VR demonstration signals, with its backbone network alternately encoding temporal and spatial dimensions. To improve the consistency of estimated posture, we incorporate concatenated conditional signals and repeated time-step embeddings. Furthermore, to bridge the gap between demonstrations and imitative behavior, we incorporate generative adversarial imitation learning (GAIL) to learn physically-constrained policies, enabling virtual avatars to imitate realistic behaviors. Experimental results on the AMASS dataset show our method significantly outperforms existing approaches in both position and rotation accuracy. We further validate our framework in a human-vehicle interaction environment, where pedestrian avatars reconstruct VR users’ demonstrated behaviors, exhibiting diverse and physically plausible actions such as walking, jumping, and crouching, highlighting the potential for more realistic human-vehicle interaction scenarios.
KW - diffusion model
KW - generative adversarial imitation learning
KW - motion tracking
KW - pedestrian simulation
KW - Virtual reality
UR - https://www.scopus.com/pages/publications/105043225988
U2 - 10.1109/TITS.2026.3701235
DO - 10.1109/TITS.2026.3701235
M3 - 文章
AN - SCOPUS:105043225988
SN - 1524-9050
JO - IEEE Transactions on Intelligent Transportation Systems
JF - IEEE Transactions on Intelligent Transportation Systems
ER -