TY - GEN
T1 - DocMamba
T2 - 31st International Conference on Multimedia Modeling, MMM 2025
AU - Han, Miaolin
AU - Li, Huibin
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025.
PY - 2025
Y1 - 2025
N2 - This paper presents a novel and robust document image dewarping method, namely DocMamba, based on the idea of selective state space sequence modeling. It consists of three modules, document image augmentation and feature extraction, sequence modeling and contextual information learning, and the robust document image dewarping. In particular, given a distorted document image, we first extract its deep convolution features, outputting a group of down-sampled feature maps. Each feature map is flatten into a vector, and a document sequence is built by all these feature vectors. The contextual information hidden in the sequence are learned by using the Selective State Space Sequence Model. That is, sequence-to-sequence transformations are performed based on the Mamba2 blocks. All sequences are then reshaped to updated feature maps and further encoded by using the dilated convolution layers. Finally, the original feature maps and the final feature maps are adding together and fed into a rectification decoder to estimate a coarse backward mapping. The final rectified image is achieved by performing the up-sampled backward mapping on the original distorted image. Extensive experiments conducted on the DocUNet and DIR300 benchmarks showed the effectiveness of the proposed method.
AB - This paper presents a novel and robust document image dewarping method, namely DocMamba, based on the idea of selective state space sequence modeling. It consists of three modules, document image augmentation and feature extraction, sequence modeling and contextual information learning, and the robust document image dewarping. In particular, given a distorted document image, we first extract its deep convolution features, outputting a group of down-sampled feature maps. Each feature map is flatten into a vector, and a document sequence is built by all these feature vectors. The contextual information hidden in the sequence are learned by using the Selective State Space Sequence Model. That is, sequence-to-sequence transformations are performed based on the Mamba2 blocks. All sequences are then reshaped to updated feature maps and further encoded by using the dilated convolution layers. Finally, the original feature maps and the final feature maps are adding together and fed into a rectification decoder to estimate a coarse backward mapping. The final rectified image is achieved by performing the up-sampled backward mapping on the original distorted image. Extensive experiments conducted on the DocUNet and DIR300 benchmarks showed the effectiveness of the proposed method.
KW - Document image dewarping
KW - Mamba2
KW - State space models
UR - https://www.scopus.com/pages/publications/85216120217
U2 - 10.1007/978-981-96-2054-8_23
DO - 10.1007/978-981-96-2054-8_23
M3 - 会议稿件
AN - SCOPUS:85216120217
SN - 9789819620531
T3 - Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
SP - 304
EP - 318
BT - MultiMedia Modeling - 31st International Conference on Multimedia Modeling, MMM 2025, Proceedings
A2 - Ide, Ichiro
A2 - Kompatsiaris, Ioannis
A2 - Xu, Changsheng
A2 - Yanai, Keiji
A2 - Chu, Wei-Ta
A2 - Nitta, Naoko
A2 - Riegler, Michael
A2 - Yamasaki, Toshihiko
PB - Springer Science and Business Media Deutschland GmbH
Y2 - 8 January 2025 through 10 January 2025
ER -