跳到主要导航 跳到搜索 跳到主要内容

Where and Why are They Looking? Jointly Inferring Human Attention and Intentions in Complex Tasks

  • University of California at Los Angeles

科研成果: 书/报告/会议事项章节会议稿件同行评审

72 引用 (Scopus)

摘要

This paper addresses a new problem - jointly inferring human attention, intentions, and tasks from videos. Given an RGB-D video where a human performs a task, we answer three questions simultaneously: 1) where the human is looking - attention prediction; 2) why the human is looking there - intention prediction; and 3) what task the human is performing - task recognition. We propose a hierarchical model of human-attention-object (HAO) which represents tasks, intentions, and attention under a unified framework. A task is represented as sequential intentions which transition to each other. An intention is composed of the human pose, attention, and objects. A beam search algorithm is adopted for inference on the HAO graph to output the attention, intention, and task results. We built a new video dataset of tasks, intentions, and attention. It contains 14 task classes, 70 intention categories, 28 object classes, 809 videos, and approximately 330,000 frames. Experiments show that our approach outperforms existing approaches.

源语言英语
主期刊名Proceedings - 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2018
出版商IEEE Computer Society
6801-6809
页数9
ISBN(电子版)9781538664209
DOI
出版状态已出版 - 14 12月 2018
活动31st Meeting of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2018 - Salt Lake City, 美国
期限: 18 6月 201822 6月 2018

丛书

姓名Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition
ISSN(印刷版)1063-6919

会议

会议31st Meeting of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2018
国家/地区美国
Salt Lake City
时期18/06/1822/06/18

学术指纹

探究 'Where and Why are They Looking? Jointly Inferring Human Attention and Intentions in Complex Tasks' 的科研主题。它们共同构成独一无二的学术指纹。

引用此