跳到主要导航 跳到搜索 跳到主要内容

IntentQA: Context-aware Video Intent Reasoning

  • Jiapeng Li
  • , Ping Wei
  • , Wenjuan Han
  • , Lifeng Fan
  • Xi'an Jiaotong University
  • National Key Laboratory of General Artificial Intelligence
  • Beijing Jiaotong University

科研成果: 书/报告/会议事项章节会议稿件同行评审

58 引用 (Scopus)

摘要

In this paper, we propose a novel task IntentQA, a special VideoQA task focusing on video intent reasoning, which has become increasingly important for AI with its advantages in equipping AI agents with the capability of reasoning beyond mere recognition in daily tasks. We also contribute a large-scale VideoQA dataset for this task. We propose a Context-aware Video Intent Reasoning model (CaVIR) consisting of i) Video Query Language (VQL) for better cross-modal representation of the situational context, ii) Contrastive Learning module for utilizing the contrastive context, and iii) Commonsense Reasoning module for incorporating the commonsense context. Comprehensive experiments on this challenging task demonstrate the effectiveness of each model component, the superiority of our full model over other baselines, and the generalizability of our model to a new VideoQA task. The dataset and codes are open-sourced at: https://github.com/JoseponLee/IntentQA.git.

源语言英语
主期刊名Proceedings - 2023 IEEE/CVF International Conference on Computer Vision, ICCV 2023
出版商Institute of Electrical and Electronics Engineers Inc.
11929-11940
页数12
ISBN(电子版)9798350307184
DOI
出版状态已出版 - 2023
活动2023 IEEE/CVF International Conference on Computer Vision, ICCV 2023 - Paris, 法国
期限: 2 10月 20236 10月 2023

丛书

姓名Proceedings of the IEEE International Conference on Computer Vision
ISSN(印刷版)1550-5499

会议

会议2023 IEEE/CVF International Conference on Computer Vision, ICCV 2023
国家/地区法国
Paris
时期2/10/236/10/23

学术指纹

探究 'IntentQA: Context-aware Video Intent Reasoning' 的科研主题。它们共同构成独一无二的学术指纹。

引用此