Focus on defocus: bridging the synthetic to real domain gap for depth estimation

Maxim Maximov, Kevin Galim, Laura Leal-Taixé

2020-05-19CVPR 2020 6Depth Prediction Depth Estimation

Abstract

Data-driven depth estimation methods struggle with the generalization outside their training scenes due to the immense variability of the real-world scenes. This problem can be partially addressed by utilising synthetically generated images, but closing the synthetic-real domain gap is far from trivial. In this paper, we tackle this issue by using domain invariant defocus blur as direct supervision. We leverage defocus cues by using a permutation invariant convolutional neural network that encourages the network to learn from the differences between images with a different point of focus. Our proposed network uses the defocus map as an intermediate supervisory signal. We are able to train our model completely on synthetic data and directly apply it to a wide range of real-world images. We evaluate our model on synthetic and real datasets, showing compelling generalization results and state-of-the-art depth prediction.

Results

Task	Dataset	Metric	Value	Model
Depth Estimation	NYU-Depth V2	RMSE	0.013	Defocus/DepthNet (Normalized)
3D	NYU-Depth V2	RMSE	0.013	Defocus/DepthNet (Normalized)

Related Papers

$S^2M^2$: Scalable Stereo Matching Model for Reliable Depth Estimation2025-07-17 $π^3$: Scalable Permutation-Equivariant Visual Geometry Learning2025-07-17 Efficient Calisthenics Skills Classification through Foreground Instance Selection and Depth Estimation2025-07-16 Vision-based Perception for Autonomous Vehicles in Obstacle Avoidance Scenarios2025-07-16 MonoMVSNet: Monocular Priors Guided Multi-View Stereo Network2025-07-15 Towards Depth Foundation Model: Recent Trends in Vision-Based Depth Estimation2025-07-15 Cameras as Relative Positional Encoding2025-07-14 ByDeWay: Boost Your multimodal LLM with DEpth prompting in a Training-Free Way2025-07-11