跳到主要导航 跳到搜索 跳到主要内容

Relative depth knowledge distillation for generalizable monocular depth estimation

  • Xi'an Jiaotong University
  • University of Electronic Science and Technology of China

科研成果: 期刊稿件文章同行评审

1 引用 (Scopus)

摘要

Monocular depth estimation provides an easily deployable solution for robots to perceive the 3D scene. Existing methods have achieved impressive performance on benchmark datasets. However, these methods tend to overfit to training domains, resulting in limited generalization in the real world. A dominant solution is to train on large-scale datasets featuring high-quality GT depth and precise camera intrinsics, both of which are often unavailable or difficult to obtain. To mitigate this issue, we propose a relative depth knowledge distillation framework to boost the generalization of monocular depth estimation with limited training data. It is based on the insight that recent relative depth foundation models can be trained efficiently on large-scale datasets to capture accurate object structure and general relative depth relationships. More specifically, in the teacher network, we generate relative depth from a pre-trained foundation model and introduce a scale alignment module to ensure its scale consistency with GT depth. In the student network, we infer the depth bin centers and corresponding probabilities to represent the scales and relative depth relationships, respectively, and compute the final depth via their linear combination. Furthermore, we design two novel response-based distillation modules to distill knowledge of relative depth and object structure, respectively, from the teacher to the student. For validation, our model is trained on widely used benchmark datasets in three settings, including indoor NYUDv2, outdoor KITTI, and a mixture of both. Extensive experiments on six unseen indoor and outdoor datasets verify that our model consistently reduces the RMSE of the base model by 3.0%, 5.3%, and 5.4% on average, respectively, and achieves state-of-the-art performance in the three settings. Our model even achieves competitive accuracy when compared to recent models trained on very large-scale datasets.

源语言英语
文章编号132632
期刊Neurocomputing
671
DOI
出版状态已出版 - 28 3月 2026

学术指纹

探究 'Relative depth knowledge distillation for generalizable monocular depth estimation' 的科研主题。它们共同构成独一无二的指纹。

引用此