TY - GEN
T1 - Joint multi-object detection and segmentation from an untrimmed video
AU - Liu, Xinling
AU - Wang, Le
AU - Zhang, Qilin
AU - Zheng, Nanning
AU - Hua, Gang
N1 - Publisher Copyright:
© IFIP International Federation for Information Processing 2020.
PY - 2020
Y1 - 2020
N2 - In this paper, we present a novel method for jointly detecting and segmenting multiple objects from an untrimmed video. Unlike most existing video object segmentation methods that can only handle a trimmed video in which all video frames contain the target objects, we address a more practical and difficult problem, i.e., joint multi-object detection and segmentation from an untrimmed video where the target objects do not always appear per frame. In particular, our method consists of two modules, i.e., object decision module and object segmentation module. The object decision module is used to detect the objects and decide which target objects need to be separated out from video. As there are usually two or more target objects and they do not always appear in the whole video, we introduce the data association into object decision module to identify their correspondences among frames. The object segmentation module aims to separate the target objects identified by object decision module. In order to extensively evaluate the proposed method, we introduce a new dataset named UNVOSeg dataset, in which 7.2% of the video frames do not contain objects. Experimental results on four datasets demonstrate that our method outperforms most of the state-of-the-art approaches.
AB - In this paper, we present a novel method for jointly detecting and segmenting multiple objects from an untrimmed video. Unlike most existing video object segmentation methods that can only handle a trimmed video in which all video frames contain the target objects, we address a more practical and difficult problem, i.e., joint multi-object detection and segmentation from an untrimmed video where the target objects do not always appear per frame. In particular, our method consists of two modules, i.e., object decision module and object segmentation module. The object decision module is used to detect the objects and decide which target objects need to be separated out from video. As there are usually two or more target objects and they do not always appear in the whole video, we introduce the data association into object decision module to identify their correspondences among frames. The object segmentation module aims to separate the target objects identified by object decision module. In order to extensively evaluate the proposed method, we introduce a new dataset named UNVOSeg dataset, in which 7.2% of the video frames do not contain objects. Experimental results on four datasets demonstrate that our method outperforms most of the state-of-the-art approaches.
KW - Data association
KW - Object detection
KW - Video object segmentation
UR - https://www.scopus.com/pages/publications/85086266932
U2 - 10.1007/978-3-030-49161-1_27
DO - 10.1007/978-3-030-49161-1_27
M3 - 会议稿件
AN - SCOPUS:85086266932
SN - 9783030491604
T3 - IFIP Advances in Information and Communication Technology
SP - 317
EP - 329
BT - Artificial Intelligence Applications and Innovations - 16th IFIP WG 12.5 International Conference, AIAI 2020, Proceedings
A2 - Maglogiannis, Ilias
A2 - Iliadis, Lazaros
A2 - Pimenidis, Elias
PB - Springer
T2 - 16th IFIP WG 12.5 International Conference on Artificial Intelligence Applications and Innovations, AIAI 2020
Y2 - 5 June 2020 through 7 June 2020
ER -