SimLingo: Vision-Only Closed-Loop Autonomous Driving with Language-Action Alignment

Katrin Renz, Long Chen, Elahe Arani, Oleg Sinavski

arXiv:2503.09594 · 2026-07-27 공개 · arXiv · PDF

vision-language vlm autonomous-driving closed-loop-control bench2drive llm-integration sensor-fusion language-action-alignment

Abstract

Integrating large language models (LLMs) into autonomous driving has attracted significant attention with the hope of improving generalization and explainability. However, existing methods often focus on either driving or vision-language understanding but achieving both high driving performance and extensive language understanding remains challenging. In addition, the dominant approach to tackle vision-language understanding is using visual question answering. However, for autonomous driving, this is only useful if it is aligned with the action space. Otherwise, the model’s answers could be inconsistent with its behavior. Therefore, we propose a model that can handle three different tasks: (1) closed-loop driving, (2) vision-language understanding, and (3) language-action alignment. Our model SimLingo is based on a vision language model (VLM) and works using only camera, excluding expensive sensors like LiDAR. SimLingo obtains state-of-the-art performance on the widely used CARLA simulator on the Bench2Drive benchmark and is the winning entry at the CARLA challenge 2024. Additionally, we achieve strong results in a wide variety of language-related tasks while maintaining high driving performance. Project page: https://katrinrenz.de/simlingo

한국어 요약

한 줄 요약

SimLingo는 시각 정보만으로 자율주행과 언어-행동 정렬을 동시에 수행하는 첫 번째 모델로, CARLA 벤치마크에서 최고 성능을 달성했다.

핵심 기여도

핵심 아이디어

SimLingo는 기존 자율주행 모델이 단일 작업(예: 주행 또는 시각-언어 이해)에만 집중하는 한계를 극복하기 위해 **3가지 작업**(주행, 시각-언어 이해, 언어-행동 정렬)을 통합한 모델을 제안한다. 특히, **Action Dreaming**이라는 새로운 데이터 수집 방법을 통해 언어 입력이 실제 행동에 어떻게 영향을 주는지를 안전하게 평가할 수 있다. 이는 기존 VQA 기반 접근법이 행동 공간과 분리되어 있어 모델의 답변과 행동이 불일치할 수 있는 문제를 해결하기 위한 핵심 아이디어이다. SimLingo는 InternVL-2 기반으로, **Path Waypoints**와 **Speed Waypoints**를 분리하여 더 정밀한 주행 제어를 가능하게 한다.

기술적 접근법

주요 결과

의의 및 한계

SimLingo는 자율주행 분야에서 **LLM과 VLM의 통합**을 시도한 최초의 모델로, 주행 성능과 언어 이해력을 동시에 유지하는 데 성공했다. 특히, **Action Dreaming**은 언어-행동 정렬을 안전하게 평가할 수 있는 새로운 벤치마크를 제시하며, 기존 VQA 기반 접근법의 한계를 극복했다. 그러나, SimLingo는 **시뮬레이션 환경**(CARLA)에서만 평가되었으며, 실제 도로 환경에서의 성능은 아직 검증되지 않았다. 또한, **경량 모델**(1B 파라미터)임에도 불구하고, 고해상도 이미지 처리와 복잡한 LLM 연산으로 인해 실시간 처리 성능은 명시되지 않았다.

실용적 활용

SimLingo는 **저비용 카메라 기반 자율주행 시스템** 개발에 활용 가능하며, **자연어 명령 처리**와 **실시간 시각-언어 상호작용**이 필요한 차량-사람 협업 시스템에도 적용될 수 있다. 특히, **Action Dreaming**은 자율주행 모델의 안전성 검증 및 인간-로봇 상호작용 향상에 기여할 수 있다.