Skip to main navigation Skip to search Skip to main content

Improving Monocular 3D Object Detection by Synthetic Images with Virtual Depth

Research output: Chapter in book / Conference proceedingChapter in an edited book (as author)Academic researchpeer-review

Abstract

Exploiting geometric features is a common approach to enhance monocular 3D object detection. However, their performance is limited due to the absence of depth information. To address this limitation, an external depth estimator can be employed to predict depth, but this approach significantly reduces the efficiency and flexibility of the model. Instead of relying on a costly depth estimator, we propose a depth-aware monocular 3D object detector that is trained using augmented training data. Specifically, we utilise reference images and their corresponding depth maps to train an efficient rendering module, which synthesises a variety of photo-realistic images with different virtual depths. By learning from these images, the detector adapts its features to depth variations. Furthermore, we introduce an auxiliary module that guides the network to learn more informative representations from the depth images. Both modules are removed after training, resulting in no additional computational overhead during the final deployment.

Original languageEnglish
Title of host publicationDeep Learning for 3D Vision
Subtitle of host publicationAlgorithms and Applications
PublisherWorld Scientific Publishing Co.
Pages201-226
Number of pages26
ISBN (Electronic)9789811286490
ISBN (Print)9789811286483
DOIs
Publication statusPublished - 1 Jan 2024

ASJC Scopus subject areas

  • General Computer Science

Fingerprint

Dive into the research topics of 'Improving Monocular 3D Object Detection by Synthetic Images with Virtual Depth'. Together they form a unique fingerprint.

Cite this