TY - GEN
T1 - Modeling 4D human-object interactions for event and object recognition
AU - Wei, Ping
AU - Zhao, Yibiao
AU - Zheng, Nanning
AU - Zhu, Song Chun
PY - 2013
Y1 - 2013
N2 - Recognizing the events and objects in the video sequence are two challenging tasks due to the complex temporal structures and the large appearance variations. In this paper, we propose a 4D human-object interaction model, where the two tasks jointly boost each other. Our human-object interaction is defined in 4D space: i) the co occurrence and geometric constraints of human pose and object in 3D space, ii) the sub-events transition and objects coherence in 1D temporal dimension. We represent the structure of events, sub-events and objects in a hierarchical graph. For an input RGB-depth video, we design a dynamic programming beam search algorithm to: i) segment the video, ii) recognize the events, and iii) detect the objects simultaneously. For evaluation, we built a large-scale multiview 3D event dataset which contains 3815 video sequences and 383,036 RGBD frames captured by the Kinect cameras. The experiment results on this dataset show the effectiveness of our method.
AB - Recognizing the events and objects in the video sequence are two challenging tasks due to the complex temporal structures and the large appearance variations. In this paper, we propose a 4D human-object interaction model, where the two tasks jointly boost each other. Our human-object interaction is defined in 4D space: i) the co occurrence and geometric constraints of human pose and object in 3D space, ii) the sub-events transition and objects coherence in 1D temporal dimension. We represent the structure of events, sub-events and objects in a hierarchical graph. For an input RGB-depth video, we design a dynamic programming beam search algorithm to: i) segment the video, ii) recognize the events, and iii) detect the objects simultaneously. For evaluation, we built a large-scale multiview 3D event dataset which contains 3815 video sequences and 383,036 RGBD frames captured by the Kinect cameras. The experiment results on this dataset show the effectiveness of our method.
UR - https://www.scopus.com/pages/publications/84898777476
U2 - 10.1109/ICCV.2013.406
DO - 10.1109/ICCV.2013.406
M3 - 会议稿件
AN - SCOPUS:84898777476
SN - 9781479928392
T3 - Proceedings of the IEEE International Conference on Computer Vision
SP - 3272
EP - 3279
BT - Proceedings - 2013 IEEE International Conference on Computer Vision, ICCV 2013
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2013 14th IEEE International Conference on Computer Vision, ICCV 2013
Y2 - 1 December 2013 through 8 December 2013
ER -