TY - GEN
T1 - DriveCompanion: Leveraging Vision-Language Models and Retrieval Augmented Generation for Enhancing Driving Convenience
AU - Ren, Mengyang
AU - Wang, Tian
AU - Cui, Haoran
AU - Zheng, Pai
N1 - Publisher Copyright:
© 2025 by ASME.
PY - 2025/10
Y1 - 2025/10
N2 - With the increasing complexity of modern road environments, drivers frequently encounter unfamiliar scenarios where understanding and complying with local traffic rules becomes challenging. Traditional navigation systems are limited in their ability to interpret complex real-time scenes or provide context-aware guidance tailored to specific situations, leaving drivers at risk of making errors. The advent of Vision-Language Models (VLMs) offers a promising solution by enabling simultaneous processing and understanding of both visual inputs and user queries. Meanwhile, Retrieval-Augmented Generation (RAG) enhances the generation of driving strategies by retrieving relevant information from external knowledge bases, addressing issues such as hallucinations, outdated knowledge, and lack of transparency inherent in VLMs. To address these challenges, DriveCompanion is proposed as a smart driving assistant that integrates VLMs and RAG. It interprets real-time driving scenes using VLMs while dynamically retrieving scenario-specific traffic rules through RAG, enabling users to inquire about and receive personalized, adaptive driving guidance in unfamiliar situations. A case study in an urban driving scenario demonstrates its effectiveness in improving driving safety and convenience, while an ablation study highlights its superior performance in adaptability and accurate guidance. This system represents a significant step toward smarter, safer and more user-centric driving assistance.
AB - With the increasing complexity of modern road environments, drivers frequently encounter unfamiliar scenarios where understanding and complying with local traffic rules becomes challenging. Traditional navigation systems are limited in their ability to interpret complex real-time scenes or provide context-aware guidance tailored to specific situations, leaving drivers at risk of making errors. The advent of Vision-Language Models (VLMs) offers a promising solution by enabling simultaneous processing and understanding of both visual inputs and user queries. Meanwhile, Retrieval-Augmented Generation (RAG) enhances the generation of driving strategies by retrieving relevant information from external knowledge bases, addressing issues such as hallucinations, outdated knowledge, and lack of transparency inherent in VLMs. To address these challenges, DriveCompanion is proposed as a smart driving assistant that integrates VLMs and RAG. It interprets real-time driving scenes using VLMs while dynamically retrieving scenario-specific traffic rules through RAG, enabling users to inquire about and receive personalized, adaptive driving guidance in unfamiliar situations. A case study in an urban driving scenario demonstrates its effectiveness in improving driving safety and convenience, while an ablation study highlights its superior performance in adaptability and accurate guidance. This system represents a significant step toward smarter, safer and more user-centric driving assistance.
KW - Advanced Driving Assistance System
KW - Context-aware Driving Guidance
KW - Human-Vehicle Interaction
KW - Retrieval Augmented Generation
KW - Vision-Language Models
UR - https://www.scopus.com/pages/publications/105024221038
U2 - 10.1115/DETC2025-156590
DO - 10.1115/DETC2025-156590
M3 - Conference article published in proceeding or book
AN - SCOPUS:105024221038
T3 - Proceedings of the ASME Design Engineering Technical Conference
BT - 45th Computers and Information in Engineering Conference (CIE)
PB - American Society of Mechanical Engineers(ASME)
T2 - ASME 2025 International Design Engineering Technical Conferences and Computers and Information in Engineering Conference, IDETC-CIE 2025
Y2 - 17 August 2025 through 20 August 2025
ER -