Skip to main navigation Skip to search Skip to main content

DocMamba: Robust Document Image Dewarping via Selective State Space Sequence Modeling

  • Xi'an Jiaotong University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

1 Scopus citations

Abstract

This paper presents a novel and robust document image dewarping method, namely DocMamba, based on the idea of selective state space sequence modeling. It consists of three modules, document image augmentation and feature extraction, sequence modeling and contextual information learning, and the robust document image dewarping. In particular, given a distorted document image, we first extract its deep convolution features, outputting a group of down-sampled feature maps. Each feature map is flatten into a vector, and a document sequence is built by all these feature vectors. The contextual information hidden in the sequence are learned by using the Selective State Space Sequence Model. That is, sequence-to-sequence transformations are performed based on the Mamba2 blocks. All sequences are then reshaped to updated feature maps and further encoded by using the dilated convolution layers. Finally, the original feature maps and the final feature maps are adding together and fed into a rectification decoder to estimate a coarse backward mapping. The final rectified image is achieved by performing the up-sampled backward mapping on the original distorted image. Extensive experiments conducted on the DocUNet and DIR300 benchmarks showed the effectiveness of the proposed method.

Original languageEnglish
Title of host publicationMultiMedia Modeling - 31st International Conference on Multimedia Modeling, MMM 2025, Proceedings
EditorsIchiro Ide, Ioannis Kompatsiaris, Changsheng Xu, Keiji Yanai, Wei-Ta Chu, Naoko Nitta, Michael Riegler, Toshihiko Yamasaki
PublisherSpringer Science and Business Media Deutschland GmbH
Pages304-318
Number of pages15
ISBN (Print)9789819620531
DOIs
StatePublished - 2025
Event31st International Conference on Multimedia Modeling, MMM 2025 - Nara, Japan
Duration: 8 Jan 202510 Jan 2025

Publication series

NameLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Volume15520 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference31st International Conference on Multimedia Modeling, MMM 2025
Country/TerritoryJapan
CityNara
Period8/01/2510/01/25

Keywords

  • Document image dewarping
  • Mamba2
  • State space models

Fingerprint

Dive into the research topics of 'DocMamba: Robust Document Image Dewarping via Selective State Space Sequence Modeling'. Together they form a unique fingerprint.

Cite this