TY - JOUR
T1 - HyperPoint
T2 - Multimodal 3D foundation model in hyperbolic space
AU - Sun, Yiding
AU - Cheng, Haozhe
AU - Lu, Chaoyi
AU - Li, Zhengqiao
AU - Wu, Minghong
AU - Lu, Huimin
AU - Zhu, Jihua
N1 - Publisher Copyright:
© 2025 Elsevier Ltd
PY - 2026/5
Y1 - 2026/5
N2 - Self-supervised learning has made significant progress in Natural Language Processing and Computer Vision. Nevertheless, it encounters significant obstacles in the 3D domain, primarily due to the scarcity of available data and the considerable challenge of effectively capturing hierarchical structures. Current methods in Euclidean space suffer from feature distortion and fail to model the semantic hierarchies inherent in cross-modal data. These challenges motivate us to adopt hyperbolic space, which excels at capturing multi-scale relationships and preserving the geometric structure of complex data. In this paper, we propose HyperPoint, the first multi-modal 3D foundational model in hyperbolic space. By projecting cross-modal features such as 3D point cloud, 2D images, and text into hyperbolic space, we leverage its tree-like properties to encode semantic hierarchies with minimal distortion. Our method integrates generative and contrastive learning while leveraging multi-loss optimization to enhance feature diversity and consistency. HyperPoint achieves a new state-of-the-art in 3D representation learning, e.g., 96.1 % accuracy on ScanObjectNN, 94.1 % accuracy on 10w10s on ModelNet40. Our code is available at: https://github.com/Issac-Sun/HyperPoint.
AB - Self-supervised learning has made significant progress in Natural Language Processing and Computer Vision. Nevertheless, it encounters significant obstacles in the 3D domain, primarily due to the scarcity of available data and the considerable challenge of effectively capturing hierarchical structures. Current methods in Euclidean space suffer from feature distortion and fail to model the semantic hierarchies inherent in cross-modal data. These challenges motivate us to adopt hyperbolic space, which excels at capturing multi-scale relationships and preserving the geometric structure of complex data. In this paper, we propose HyperPoint, the first multi-modal 3D foundational model in hyperbolic space. By projecting cross-modal features such as 3D point cloud, 2D images, and text into hyperbolic space, we leverage its tree-like properties to encode semantic hierarchies with minimal distortion. Our method integrates generative and contrastive learning while leveraging multi-loss optimization to enhance feature diversity and consistency. HyperPoint achieves a new state-of-the-art in 3D representation learning, e.g., 96.1 % accuracy on ScanObjectNN, 94.1 % accuracy on 10w10s on ModelNet40. Our code is available at: https://github.com/Issac-Sun/HyperPoint.
KW - 3D point cloud understanding
KW - Hyperbolic embedding
KW - Multimodal foundation model
UR - https://www.scopus.com/pages/publications/105024566644
U2 - 10.1016/j.patcog.2025.112800
DO - 10.1016/j.patcog.2025.112800
M3 - 文章
AN - SCOPUS:105024566644
SN - 0031-3203
VL - 173
JO - Pattern Recognition
JF - Pattern Recognition
M1 - 112800
ER -