BinsFormer: Revisiting Adaptive Bins for Monocular Depth Estimation

Zhenyu Li, Xuyang Wang, Xianming Liu, Junjun Jiang

2022-04-03regression Scene Understanding Depth Estimation Monocular Depth Estimation

Abstract

Monocular depth estimation is a fundamental task in computer vision and has drawn increasing attention. Recently, some methods reformulate it as a classification-regression task to boost the model performance, where continuous depth is estimated via a linear combination of predicted probability distributions and discrete bins. In this paper, we present a novel framework called BinsFormer, tailored for the classification-regression-based depth estimation. It mainly focuses on two crucial components in the specific task: 1) proper generation of adaptive bins and 2) sufficient interaction between probability distribution and bins predictions. To specify, we employ the Transformer decoder to generate bins, novelly viewing it as a direct set-to-set prediction problem. We further integrate a multi-scale decoder structure to achieve a comprehensive understanding of spatial geometry information and estimate depth maps in a coarse-to-fine manner. Moreover, an extra scene understanding query is proposed to improve the estimation accuracy, which turns out that models can implicitly learn useful information from an auxiliary environment classification task. Extensive experiments on the KITTI, NYU, and SUN RGB-D datasets demonstrate that BinsFormer surpasses state-of-the-art monocular depth estimation methods with prominent margins. Code and pretrained models will be made publicly available at \url{https://github.com/zhyever/Monocular-Depth-Estimation-Toolbox}.

Results

Task	Dataset	Metric	Value	Model
Depth Estimation	NYU-Depth V2	Delta < 1.25	0.925	BinsFormer
Depth Estimation	NYU-Depth V2	Delta < 1.25^2	0.989	BinsFormer
Depth Estimation	NYU-Depth V2	Delta < 1.25^3	0.997	BinsFormer
Depth Estimation	NYU-Depth V2	RMSE	0.33	BinsFormer
Depth Estimation	NYU-Depth V2	absolute relative error	0.094	BinsFormer
Depth Estimation	NYU-Depth V2	log 10	0.04	BinsFormer
Depth Estimation	KITTI Eigen split	Delta < 1.25	0.974	BinsFormer
Depth Estimation	KITTI Eigen split	Delta < 1.25^2	0.997	BinsFormer
Depth Estimation	KITTI Eigen split	Delta < 1.25^3	0.999	BinsFormer
Depth Estimation	KITTI Eigen split	RMSE	2.098	BinsFormer
Depth Estimation	KITTI Eigen split	RMSE log	0.079	BinsFormer
Depth Estimation	KITTI Eigen split	Sq Rel	0.151	BinsFormer
Depth Estimation	KITTI Eigen split	absolute relative error	0.052	BinsFormer
3D	NYU-Depth V2	Delta < 1.25	0.925	BinsFormer
3D	NYU-Depth V2	Delta < 1.25^2	0.989	BinsFormer
3D	NYU-Depth V2	Delta < 1.25^3	0.997	BinsFormer
3D	NYU-Depth V2	RMSE	0.33	BinsFormer
3D	NYU-Depth V2	absolute relative error	0.094	BinsFormer
3D	NYU-Depth V2	log 10	0.04	BinsFormer
3D	KITTI Eigen split	Delta < 1.25	0.974	BinsFormer
3D	KITTI Eigen split	Delta < 1.25^2	0.997	BinsFormer
3D	KITTI Eigen split	Delta < 1.25^3	0.999	BinsFormer
3D	KITTI Eigen split	RMSE	2.098	BinsFormer
3D	KITTI Eigen split	RMSE log	0.079	BinsFormer
3D	KITTI Eigen split	Sq Rel	0.151	BinsFormer
3D	KITTI Eigen split	absolute relative error	0.052	BinsFormer

BinsFormer: Revisiting Adaptive Bins for Monocular Depth Estimation

Abstract

Results

Related Papers

BinsFormer: Revisiting Adaptive Bins for Monocular Depth Estimation

Abstract

Results

Related Papers