TY - GEN
T1 - Learn A Compression for Objection Detection - VAE with a Bridge
AU - Mei, Yixin
AU - Li, Fan
AU - Li, Li
AU - Li, Zhu
N1 - Publisher Copyright:
© 2021 IEEE.
PY - 2021
Y1 - 2021
N2 - Recent advances in sensor technology and wide deployment of visual sensors lead to a new application whereas compression of images are not mainly for pixel recovery for human consumption, instead it is for communication to cloud side machine vision tasks like classification, identification, detection and tracking. This opens up new research dimensions for a learning based compression that directly optimizes loss function in vision tasks, and therefore achieves better compression performance vis-a-vis the pixel recovery and then performing vision tasks computing. In this work, we developed a learning based compression scheme that learns a compact feature representation and appropriate bitstreams for the task of visual object detection. Variational Auto-Encoder (VAE) framework is adopted for learning a compact representation, while a bridge network is trained to drive the detection loss function. Simulation results demonstrate that this approach is achieving a new state-of-the-art in task driven compression efficiency, compared with pixel recovery approaches, including both learning based and handcrafted solutions.
AB - Recent advances in sensor technology and wide deployment of visual sensors lead to a new application whereas compression of images are not mainly for pixel recovery for human consumption, instead it is for communication to cloud side machine vision tasks like classification, identification, detection and tracking. This opens up new research dimensions for a learning based compression that directly optimizes loss function in vision tasks, and therefore achieves better compression performance vis-a-vis the pixel recovery and then performing vision tasks computing. In this work, we developed a learning based compression scheme that learns a compact feature representation and appropriate bitstreams for the task of visual object detection. Variational Auto-Encoder (VAE) framework is adopted for learning a compact representation, while a bridge network is trained to drive the detection loss function. Simulation results demonstrate that this approach is achieving a new state-of-the-art in task driven compression efficiency, compared with pixel recovery approaches, including both learning based and handcrafted solutions.
KW - Image coding for machine
KW - Learning-based image compression
KW - Object detection
UR - https://www.scopus.com/pages/publications/85125241634
U2 - 10.1109/VCIP53242.2021.9675387
DO - 10.1109/VCIP53242.2021.9675387
M3 - 会议稿件
AN - SCOPUS:85125241634
T3 - 2021 International Conference on Visual Communications and Image Processing, VCIP 2021 - Proceedings
BT - 2021 International Conference on Visual Communications and Image Processing, VCIP 2021 - Proceedings
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2021 International Conference on Visual Communications and Image Processing, VCIP 2021
Y2 - 5 December 2021 through 8 December 2021
ER -