TY - JOUR
T1 - Relative depth knowledge distillation for generalizable monocular depth estimation
AU - Zhang, Lulu
AU - Li, Mankun
AU - Yang, Meng
AU - Lan, Xuguang
AU - Zhu, Ce
N1 - Publisher Copyright:
© 2026 Elsevier B.V.
PY - 2026/3/28
Y1 - 2026/3/28
N2 - Monocular depth estimation provides an easily deployable solution for robots to perceive the 3D scene. Existing methods have achieved impressive performance on benchmark datasets. However, these methods tend to overfit to training domains, resulting in limited generalization in the real world. A dominant solution is to train on large-scale datasets featuring high-quality GT depth and precise camera intrinsics, both of which are often unavailable or difficult to obtain. To mitigate this issue, we propose a relative depth knowledge distillation framework to boost the generalization of monocular depth estimation with limited training data. It is based on the insight that recent relative depth foundation models can be trained efficiently on large-scale datasets to capture accurate object structure and general relative depth relationships. More specifically, in the teacher network, we generate relative depth from a pre-trained foundation model and introduce a scale alignment module to ensure its scale consistency with GT depth. In the student network, we infer the depth bin centers and corresponding probabilities to represent the scales and relative depth relationships, respectively, and compute the final depth via their linear combination. Furthermore, we design two novel response-based distillation modules to distill knowledge of relative depth and object structure, respectively, from the teacher to the student. For validation, our model is trained on widely used benchmark datasets in three settings, including indoor NYUDv2, outdoor KITTI, and a mixture of both. Extensive experiments on six unseen indoor and outdoor datasets verify that our model consistently reduces the RMSE of the base model by 3.0%, 5.3%, and 5.4% on average, respectively, and achieves state-of-the-art performance in the three settings. Our model even achieves competitive accuracy when compared to recent models trained on very large-scale datasets.
AB - Monocular depth estimation provides an easily deployable solution for robots to perceive the 3D scene. Existing methods have achieved impressive performance on benchmark datasets. However, these methods tend to overfit to training domains, resulting in limited generalization in the real world. A dominant solution is to train on large-scale datasets featuring high-quality GT depth and precise camera intrinsics, both of which are often unavailable or difficult to obtain. To mitigate this issue, we propose a relative depth knowledge distillation framework to boost the generalization of monocular depth estimation with limited training data. It is based on the insight that recent relative depth foundation models can be trained efficiently on large-scale datasets to capture accurate object structure and general relative depth relationships. More specifically, in the teacher network, we generate relative depth from a pre-trained foundation model and introduce a scale alignment module to ensure its scale consistency with GT depth. In the student network, we infer the depth bin centers and corresponding probabilities to represent the scales and relative depth relationships, respectively, and compute the final depth via their linear combination. Furthermore, we design two novel response-based distillation modules to distill knowledge of relative depth and object structure, respectively, from the teacher to the student. For validation, our model is trained on widely used benchmark datasets in three settings, including indoor NYUDv2, outdoor KITTI, and a mixture of both. Extensive experiments on six unseen indoor and outdoor datasets verify that our model consistently reduces the RMSE of the base model by 3.0%, 5.3%, and 5.4% on average, respectively, and achieves state-of-the-art performance in the three settings. Our model even achieves competitive accuracy when compared to recent models trained on very large-scale datasets.
KW - Generalization
KW - Knowledge distillation
KW - Monocular depth estimation
KW - Object structure
KW - Relative depth
UR - https://www.scopus.com/pages/publications/105027464981
U2 - 10.1016/j.neucom.2026.132632
DO - 10.1016/j.neucom.2026.132632
M3 - 文章
AN - SCOPUS:105027464981
SN - 0925-2312
VL - 671
JO - Neurocomputing
JF - Neurocomputing
M1 - 132632
ER -