TY - JOUR
T1 - M3imic
T2 - Learning a Versatile Whole-Body Controller for Multimodal Motion Mimicking
AU - Lu, Zuxing
AU - Zheng, Ziang
AU - Lyu, Yao
AU - Liu, Jingyu
AU - Zhang, Feihong
AU - Lu, Song
AU - Yuan, Xin
AU - Sun, Changyin
AU - Zuo, Xingxing
AU - Li, Shengbo Eben
N1 - Publisher Copyright:
© 2016 IEEE.
PY - 2026
Y1 - 2026
N2 - Building a general-purpose whole-body controller is essential for enabling diverse motion capabilities in humanoid robots across a wide range of downstream tasks, including locomotion and loco-manipulation. Different tasks rely on distinct motion reference modalities: locomotion primarily depends on coordinated robot joint trajectories, whereas manipulation requires precise end-effector trajectory tracking. Existing methods often overlook the representational mismatch between dense robot joint angles and sparse end-effector poses. To address this, we propose Multi-Modal Mimic (M3imic), a versatile multi-modal whole-body control framework that unifies heterogeneous motion reference modalities, including robot joint angles, human pose trajectories, and end-effector poses, using modality-specific encoders to map them into a shared latent space. Leveraging large-scale reinforcement learning in the simulator, we train a single policy that achieves sim-to-real transfer across multiple motion reference modalities without modality-specific retraining. Extensive simulation and real-world experiments on the Unitree G1 robot are conducted to evaluate the proposed framework. In simulation, the policy achieves a peak success rate of 98.42% on an unseen test dataset, demonstrating its exceptional generalization capability.
AB - Building a general-purpose whole-body controller is essential for enabling diverse motion capabilities in humanoid robots across a wide range of downstream tasks, including locomotion and loco-manipulation. Different tasks rely on distinct motion reference modalities: locomotion primarily depends on coordinated robot joint trajectories, whereas manipulation requires precise end-effector trajectory tracking. Existing methods often overlook the representational mismatch between dense robot joint angles and sparse end-effector poses. To address this, we propose Multi-Modal Mimic (M3imic), a versatile multi-modal whole-body control framework that unifies heterogeneous motion reference modalities, including robot joint angles, human pose trajectories, and end-effector poses, using modality-specific encoders to map them into a shared latent space. Leveraging large-scale reinforcement learning in the simulator, we train a single policy that achieves sim-to-real transfer across multiple motion reference modalities without modality-specific retraining. Extensive simulation and real-world experiments on the Unitree G1 robot are conducted to evaluate the proposed framework. In simulation, the policy achieves a peak success rate of 98.42% on an unseen test dataset, demonstrating its exceptional generalization capability.
KW - Humanoid robots
KW - reinforcement learning
KW - whole-body control
UR - https://www.scopus.com/pages/publications/105043514713
U2 - 10.1109/LRA.2026.3706946
DO - 10.1109/LRA.2026.3706946
M3 - 文章
AN - SCOPUS:105043514713
SN - 2377-3766
JO - IEEE Robotics and Automation Letters
JF - IEEE Robotics and Automation Letters
ER -