MonoViT: Self-Supervised Monocular Depth Estimation with a Vision Transformer

Chaoqiang Zhao, Youmin Zhang, Matteo Poggi, Fabio Tosi, Xianda Guo, Zheng Zhu, Guan Huang, Yang Tang, Stefano Mattoccia

2022-08-06Unsupervised Monocular Depth Estimation Depth Prediction Depth Estimation Monocular Depth Estimation

Abstract

Self-supervised monocular depth estimation is an attractive solution that does not require hard-to-source depth labels for training. Convolutional neural networks (CNNs) have recently achieved great success in this task. However, their limited receptive field constrains existing network architectures to reason only locally, dampening the effectiveness of the self-supervised paradigm. In the light of the recent successes achieved by Vision Transformers (ViTs), we propose MonoViT, a brand-new framework combining the global reasoning enabled by ViT models with the flexibility of self-supervised monocular depth estimation. By combining plain convolutions with Transformer blocks, our model can reason locally and globally, yielding depth prediction at a higher level of detail and accuracy, allowing MonoViT to achieve state-of-the-art performance on the established KITTI dataset. Moreover, MonoViT proves its superior generalization capacities on other datasets such as Make3D and DrivingStereo.

Results

Task	Dataset	Metric	Value	Model
Depth Estimation	KITTI	absolute relative error	0.093	MonoViT
Depth Estimation	KITTI Eigen split unsupervised	Delta < 1.25	0.912	MonoViT(MS+1024x320)
Depth Estimation	KITTI Eigen split unsupervised	Delta < 1.25^2	0.969	MonoViT(MS+1024x320)
Depth Estimation	KITTI Eigen split unsupervised	Delta < 1.25^3	0.985	MonoViT(MS+1024x320)
Depth Estimation	KITTI Eigen split unsupervised	RMSE	4.202	MonoViT(MS+1024x320)
Depth Estimation	KITTI Eigen split unsupervised	RMSE log	0.169	MonoViT(MS+1024x320)
Depth Estimation	KITTI Eigen split unsupervised	Sq Rel	0.671	MonoViT(MS+1024x320)
Depth Estimation	KITTI Eigen split unsupervised	absolute relative error	0.093	MonoViT(MS+1024x320)
3D	KITTI	absolute relative error	0.093	MonoViT
3D	KITTI Eigen split unsupervised	Delta < 1.25	0.912	MonoViT(MS+1024x320)
3D	KITTI Eigen split unsupervised	Delta < 1.25^2	0.969	MonoViT(MS+1024x320)
3D	KITTI Eigen split unsupervised	Delta < 1.25^3	0.985	MonoViT(MS+1024x320)
3D	KITTI Eigen split unsupervised	RMSE	4.202	MonoViT(MS+1024x320)
3D	KITTI Eigen split unsupervised	RMSE log	0.169	MonoViT(MS+1024x320)
3D	KITTI Eigen split unsupervised	Sq Rel	0.671	MonoViT(MS+1024x320)
3D	KITTI Eigen split unsupervised	absolute relative error	0.093	MonoViT(MS+1024x320)

MonoViT: Self-Supervised Monocular Depth Estimation with a Vision Transformer

Abstract

Results

Related Papers

MonoViT: Self-Supervised Monocular Depth Estimation with a Vision Transformer

Abstract

Results

Related Papers