Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Depth-Relative Self Attention for Monocular Depth Estimation

View through CrossRef
Monocular depth estimation is very challenging because clues to the exact depth are incomplete in a single RGB image. To overcome the limitation, deep neural networks rely on various visual hints such as size, shade, and texture extracted from RGB information. However, we observe that if such hints are overly exploited, the network can be biased on RGB information without considering the comprehensive view. We propose a novel depth estimation model named RElative Depth Transformer (RED-T) that uses relative depth as guidance in self-attention. Specifically, the model assigns high attention weights to pixels of close depth and low attention weights to pixels of distant depth. As a result, the features of similar depth can become more likely to each other and thus less prone to misused visual hints. We show that the proposed model achieves competitive results in monocular depth estimation benchmarks and is less biased to RGB information. In addition, we propose a novel monocular depth estimation benchmark that limits the observable depth range during training in order to evaluate the robustness of the model for unseen depths.
Title: Depth-Relative Self Attention for Monocular Depth Estimation
Description:
Monocular depth estimation is very challenging because clues to the exact depth are incomplete in a single RGB image.
To overcome the limitation, deep neural networks rely on various visual hints such as size, shade, and texture extracted from RGB information.
However, we observe that if such hints are overly exploited, the network can be biased on RGB information without considering the comprehensive view.
We propose a novel depth estimation model named RElative Depth Transformer (RED-T) that uses relative depth as guidance in self-attention.
Specifically, the model assigns high attention weights to pixels of close depth and low attention weights to pixels of distant depth.
As a result, the features of similar depth can become more likely to each other and thus less prone to misused visual hints.
We show that the proposed model achieves competitive results in monocular depth estimation benchmarks and is less biased to RGB information.
In addition, we propose a novel monocular depth estimation benchmark that limits the observable depth range during training in order to evaluate the robustness of the model for unseen depths.

Related Results

Monocular Depth Estimation (Literature Review)
Monocular Depth Estimation (Literature Review)
Background. The physiological basis of spatial perception is traditionally attributed to the binocular system, which integrates the signals coming to the brain from each eye into a...
Is a Fitbit a Diary? Self-Tracking and Autobiography
Is a Fitbit a Diary? Self-Tracking and Autobiography
Data becomes something of a mirror in which people see themselves reflected. (Sorapure 270)In a 2014 essay for The New Yorker, the humourist David Sedaris recounts an obsession spu...
Monocular Vision-Based Obstacle Height Estimation for Mobile Robot
Monocular Vision-Based Obstacle Height Estimation for Mobile Robot
To ensure that robots can operate reliably in diverse environments, obstacle detection is essential, which requires the acquisition of depth information of the surrounding environm...
Unsupervised Monocular Depth Estimation Method Based on Uncertainty Analysis and Retinex Algorithm
Unsupervised Monocular Depth Estimation Method Based on Uncertainty Analysis and Retinex Algorithm
Depth estimation of a single image presents a classic problem for computer vision, and is important for the 3D reconstruction of scenes, augmented reality, and object detection. At...
Evaluation Monocular Depth Estimation Model for UAVApplications
Evaluation Monocular Depth Estimation Model for UAVApplications
Safe and autonomous movement capabilities in unmanned aerial vehicles (UAVs) aredirectly dependent on accurate and reliable perception of the environment. Depthinformation presents...
Attentional rhythms are generated in binocular cells
Attentional rhythms are generated in binocular cells
Abstract Visual attention is intrinsically rhythmic and oscillates based on the discrete sampling of either single or multiple objects. Recently, studies employing ...
Monocular Vision-Based Obstacle Height Estimation for Mobile Robot
Monocular Vision-Based Obstacle Height Estimation for Mobile Robot
For a robot to operate robustly in diverse real-world environments, reliable obstacle perception is essential, which fundamentally requires depth information of the surrounding sce...

Back to Top