TY - JOUR
T1 - Enhancing 3D Instance Segmentation With Dense Connection Decoder and Layer-Aware Fusion
AU - Wang, Duanchu
AU - Gong, Haoran
AU - Wang, Di
N1 - Publisher Copyright:
© IEEE. 2016 IEEE.
PY - 2025
Y1 - 2025
N2 - 3D instance segmentation (3DIS) aims to identify object instances in a 3D scene by predicting binary foreground masks with corresponding semantic labels. Transformer-based methods have demonstrated strong performance by effectively capturing global context information through attention mechanisms. However, existing approaches primarily focus on capturing external relationships between scene features and instance queries, while overlooking internal dependencies between queries across decoder layers. This limitation can lead to inconsistencies in query mask predictions across layers, ultimately hindering segmentation performance and slowing model convergence. To address this, we propose the Dense Connection Decoder (DCD), a novel architecture that explicitly models dependencies between instance queries across decoder layers. Our design introduces a Fusion Module and a Memory Module to construct layer-aware hybrid states, dynamically assigning information weights to previous queries based on their decoder distance. Additionally, a Selection Module refines query features through a gating mechanism, adaptively controlling the influence of upstream information. By enforcing prediction consistency across layers, DCD not only enhances segmentation accuracy, but also accelerates model convergence. Extensive experiments on ScanNetV2, ScanNet++V2, ScanNet200, and S3DIS demonstrate that DCD outperforms existing transformer-based baselines, achieving state-of-the-art performance and faster convergence.
AB - 3D instance segmentation (3DIS) aims to identify object instances in a 3D scene by predicting binary foreground masks with corresponding semantic labels. Transformer-based methods have demonstrated strong performance by effectively capturing global context information through attention mechanisms. However, existing approaches primarily focus on capturing external relationships between scene features and instance queries, while overlooking internal dependencies between queries across decoder layers. This limitation can lead to inconsistencies in query mask predictions across layers, ultimately hindering segmentation performance and slowing model convergence. To address this, we propose the Dense Connection Decoder (DCD), a novel architecture that explicitly models dependencies between instance queries across decoder layers. Our design introduces a Fusion Module and a Memory Module to construct layer-aware hybrid states, dynamically assigning information weights to previous queries based on their decoder distance. Additionally, a Selection Module refines query features through a gating mechanism, adaptively controlling the influence of upstream information. By enforcing prediction consistency across layers, DCD not only enhances segmentation accuracy, but also accelerates model convergence. Extensive experiments on ScanNetV2, ScanNet++V2, ScanNet200, and S3DIS demonstrate that DCD outperforms existing transformer-based baselines, achieving state-of-the-art performance and faster convergence.
KW - RGB-D perception
KW - computer vision for automation
KW - deep learning for visual perception
UR - https://www.scopus.com/pages/publications/105014224667
U2 - 10.1109/LRA.2025.3600142
DO - 10.1109/LRA.2025.3600142
M3 - 文章
AN - SCOPUS:105014224667
SN - 2377-3766
VL - 10
SP - 10186
EP - 10193
JO - IEEE Robotics and Automation Letters
JF - IEEE Robotics and Automation Letters
IS - 10
ER -