跳到主要导航 跳到搜索 跳到主要内容

Swin-6D: 6D Pose Estimation via 3D Keypoints Voting with Swin Transformer

  • Xi'an Jiaotong University
  • State Grid Shaanxi Electric Power Company Limited

科研成果: 期刊稿件文章同行评审

摘要

6D pose estimation using RGB-D data is essential for various computer vision applications. Extracting relevant information from depth and color image data, and effectively integrating them, remains a significant technical challenge. Previous approaches primarily depend on convolutional networks for feature extraction and neglect to fully integrate features across the entire pipeline. This limitation undermines the robustness of existing 6D pose estimation methods, particularly in scenarios involving significant occlusions and clutter. To address these issues, we propose Swin-6D, a novel two-branch network designed specifically for 6D pose estimation with RGB-D input. The proposed model leverages the Swin Transformer in the RGB branch, which excels in handling occlusions. Moreover, the network incorporates a feature fusion module that facilitates the seamless integration of RGB and point cloud features. We conduct experiments on the LineMOD and Occlusion-LineMOD datasets, demonstrating that Swin-6D achieves ADD(−S) scores of 99.7 and 75.5, respectively. We further validate our method on the YCB-Video dataset, demonstrating competitive performance in cluttered real-world scenes. These results highlight that Swin-6D outperforms state-of-the-art methods, especially in scenarios with significant occlusions. We also implemented our method on the AUBO i5 robot for grasping experiments, where it achieves robust performance even in cluttered and partially occluded scenes, demonstrating its practical deployment potential.

源语言英语
页(从-至)1055-1071
页数17
期刊Journal of Electrical Engineering and Technology
21
1
DOI
出版状态已出版 - 1月 2026

学术指纹

探究 'Swin-6D: 6D Pose Estimation via 3D Keypoints Voting with Swin Transformer' 的科研主题。它们共同构成独一无二的指纹。

引用此