Skip to main navigation Skip to search Skip to main content

Robust Multimodal Depth Estimation using Transformer based Generative Adversarial Networks

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Abstract

Accurately measuring the absolute depth of every pixel captured by an imaging sensor is of critical importance in real-Time applications such as autonomous navigation, augmented reality and robotics. In order to predict dense depth, a general approach is to fuse sensor inputs from different modalities such as LiDAR, camera and other time-of-flight sensors. LiDAR and other time-of-flight sensors provide accurate depth data but are quite sparse, both spatially and temporally. To augment missing depth information, generally RGB guidance is leveraged due to its high resolution information. Due to the reliance on multiple sensor modalities, design for robustness and adaptation is essential. In this work, we propose a transformer-like self-Attention based generative adversarial network to estimate dense depth using RGB and sparse depth data. We introduce a novel training recipe for making the model robust so that it works even when one of the input modalities is not available. The multi-head self-Attention mechanism can dynamically attend to most salient parts of the RGB image or corresponding sparse depth data producing the most competitive results. Our proposed network also requires less memory for training and inference compared to other existing heavily residual connection based convolutional neural networks, making it more suitable for resource-constrained edge applications. The source code is available at: https://github.com/kocchop/robust-multimodal-fusion-gan

Original languageEnglish (US)
Title of host publicationMM 2022 - Proceedings of the 30th ACM International Conference on Multimedia
PublisherAssociation for Computing Machinery, Inc
Pages3559-3568
Number of pages10
ISBN (Electronic)9781450392037
DOIs
StatePublished - Oct 10 2022
Event30th ACM International Conference on Multimedia, MM 2022 - Lisboa, Portugal
Duration: Oct 10 2022Oct 14 2022

Publication series

NameMM 2022 - Proceedings of the 30th ACM International Conference on Multimedia

Conference

Conference30th ACM International Conference on Multimedia, MM 2022
Country/TerritoryPortugal
CityLisboa
Period10/10/2210/14/22

All Science Journal Classification (ASJC) codes

  • Artificial Intelligence
  • Computer Graphics and Computer-Aided Design
  • Human-Computer Interaction
  • Software

Fingerprint

Dive into the research topics of 'Robust Multimodal Depth Estimation using Transformer based Generative Adversarial Networks'. Together they form a unique fingerprint.

Cite this