Skip to main navigation Skip to search Skip to main content

DriveCompanion: Leveraging Vision-Language Models and Retrieval Augmented Generation for Enhancing Driving Convenience

  • Mengyang Ren
  • , Tian Wang
  • , Haoran Cui
  • , Pai Zheng (Corresponding Author)

Research output: Chapter in book / Conference proceedingConference article published in proceeding or bookAcademic researchpeer-review

Abstract

With the increasing complexity of modern road environments, drivers frequently encounter unfamiliar scenarios where understanding and complying with local traffic rules becomes challenging. Traditional navigation systems are limited in their ability to interpret complex real-time scenes or provide context-aware guidance tailored to specific situations, leaving drivers at risk of making errors. The advent of Vision-Language Models (VLMs) offers a promising solution by enabling simultaneous processing and understanding of both visual inputs and user queries. Meanwhile, Retrieval-Augmented Generation (RAG) enhances the generation of driving strategies by retrieving relevant information from external knowledge bases, addressing issues such as hallucinations, outdated knowledge, and lack of transparency inherent in VLMs. To address these challenges, DriveCompanion is proposed as a smart driving assistant that integrates VLMs and RAG. It interprets real-time driving scenes using VLMs while dynamically retrieving scenario-specific traffic rules through RAG, enabling users to inquire about and receive personalized, adaptive driving guidance in unfamiliar situations. A case study in an urban driving scenario demonstrates its effectiveness in improving driving safety and convenience, while an ablation study highlights its superior performance in adaptability and accurate guidance. This system represents a significant step toward smarter, safer and more user-centric driving assistance.

Original languageEnglish
Title of host publication45th Computers and Information in Engineering Conference (CIE)
PublisherAmerican Society of Mechanical Engineers(ASME)
Number of pages9
ISBN (Electronic)9780791889213
DOIs
Publication statusPublished - Oct 2025
EventASME 2025 International Design Engineering Technical Conferences and Computers and Information in Engineering Conference, IDETC-CIE 2025 - Anaheim, United States
Duration: 17 Aug 202520 Aug 2025

Publication series

NameProceedings of the ASME Design Engineering Technical Conference
Volume2-B

Conference

ConferenceASME 2025 International Design Engineering Technical Conferences and Computers and Information in Engineering Conference, IDETC-CIE 2025
Country/TerritoryUnited States
CityAnaheim
Period17/08/2520/08/25

Keywords

  • Advanced Driving Assistance System
  • Context-aware Driving Guidance
  • Human-Vehicle Interaction
  • Retrieval Augmented Generation
  • Vision-Language Models

ASJC Scopus subject areas

  • Mechanical Engineering
  • Computer Graphics and Computer-Aided Design
  • Computer Science Applications
  • Modelling and Simulation

Fingerprint

Dive into the research topics of 'DriveCompanion: Leveraging Vision-Language Models and Retrieval Augmented Generation for Enhancing Driving Convenience'. Together they form a unique fingerprint.

Cite this