Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives

Shaoyuan Xie, Lingdong Kong, Yuhao Dong, Chonghao Sima, Wenwei Zhang, Qi Alfred Chen, Ziwei Liu, Liang Pan

arXiv:2501.04003 · 2026-07-27 공개 · arXiv · PDF

vision-language-models autonomous-driving visual-grounding evaluation-metrics interpretable-decision-making drivebench robust-agentic-utilization corruption-awareness

Abstract

Recent advancements in Vision-Language Models (VLMs) have fueled interest in autonomous driving applications, particularly for interpretable decision-making. However, the assumption that VLMs provide visually grounded and reliable driving explanations remains unexamined. To address this, we introduce DriveBench, a benchmark eval-uating 12 VLMs across 17 settings, covering 19,200 images, 20,498 QA pairs, and four key driving tasks. Our findings reveal that existing VLMs often generate plausible responses from general knowledge or textual cues rather than true visual grounding, especially under degraded or missing visual inputs. This behavior, concealed by dataset imbalances and insufficient evaluation metrics, poses significant risks in safety-critical scenarios like autonomous driving. We further observe that VLMs possess inherent corruption-awareness but only explicitly acknowledge these issues when directly prompted. Given the challenges and inspired by the inherent corruption awareness, we propose Robust Agentic Utilization (RAU), leveraging VLMs' corruption awareness and agentic planning with external tools to enhance perception reliability for a diverse set of downstream tasks. Our study challenges existing evaluation paradigms and provides a road map toward more robust and interpretable autonomous driving systems.

한국어 요약

한 줄 요약

12개 VLM을 17개 설정에서 평가한 DriveBench를 통해 자율주행에서의 VLM 신뢰도 문제를 실증적으로 분석하고, Robust Agentic Utilization(RAU)을 제안한다.

핵심 기여도

핵심 아이디어

자율주행에서 VLM이 시각 정보를 기반으로 신뢰할 수 있는 결정을 내리는지에 대한 가정은 검증되지 않았다. 본 연구는 VLM이 텍스트 기반 지식으로 응답을 생성하는 경향이 강하다는 점을 밝히며, 이는 훈련 데이터 불균형과 평가 지표의 한계로 인해 드러나지 않았다. 예를 들어, 텍스트만 입력된 경우에도 VLM이 정확도를 유지하는 현상은 인간 운전자와 대조된다. 이는 VLM이 시각 정보를 제대로 해석하지 못하면서도 외형상 신뢰도 높은 응답을 생성한다는 의미이다. 따라서, VLM의 훼손 인식 능력을 활용한 Robust Agentic Utilization(RAU)이 필요하다는 통찰을 제시한다.

기술적 접근법

주요 결과

의의 및 한계

본 연구는 VLM이 자율주행에서 신뢰할 수 있는 결정을 내리는지에 대한 핵심 질문을 제기하고, DriveBench를 통해 실증적으로 검증하였다. 기존 평가 지표와 데이터셋의 한계를 지적하며, RAU를 통해 VLM의 신뢰도를 향상시키는 방향을 제시하였다. 그러나 RAU의 구체적인 구현 방식과 외부 도구와의 통합 방식은 추가 연구가 필요하다. 또한, DriveBench는 특정 VLM과 데이터셋에 기반하기 때문에 일반화 가능성에 대한 논의가 필요하다.

실용적 활용

DriveBench와 RAU는 자율주행 시스템 설계, VLM 기반 인터페이스 개발, 안전성 평가 프로세스에서 활용 가능하다. 특히, VLM의 시각 정보 해석 능력을 평가하고, 안전-critical 상황에서 신뢰도를 높이는 데 기여할 수 있다.