Abstract
Self-supervised learning has made significant progress in Natural Language Processing and Computer Vision. Nevertheless, it encounters significant obstacles in the 3D domain, primarily due to the scarcity of available data and the considerable challenge of effectively capturing hierarchical structures. Current methods in Euclidean space suffer from feature distortion and fail to model the semantic hierarchies inherent in cross-modal data. These challenges motivate us to adopt hyperbolic space, which excels at capturing multi-scale relationships and preserving the geometric structure of complex data. In this paper, we propose HyperPoint, the first multi-modal 3D foundational model in hyperbolic space. By projecting cross-modal features such as 3D point cloud, 2D images, and text into hyperbolic space, we leverage its tree-like properties to encode semantic hierarchies with minimal distortion. Our method integrates generative and contrastive learning while leveraging multi-loss optimization to enhance feature diversity and consistency. HyperPoint achieves a new state-of-the-art in 3D representation learning, e.g., 96.1 % accuracy on ScanObjectNN, 94.1 % accuracy on 10w10s on ModelNet40. Our code is available at: https://github.com/Issac-Sun/HyperPoint.
| Original language | English |
|---|---|
| Article number | 112800 |
| Journal | Pattern Recognition |
| Volume | 173 |
| DOIs | |
| State | Published - May 2026 |
Keywords
- 3D point cloud understanding
- Hyperbolic embedding
- Multimodal foundation model
Fingerprint
Dive into the research topics of 'HyperPoint: Multimodal 3D foundation model in hyperbolic space'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver