NAVSIM: Data-Driven Non-Reactive Autonomous Vehicle Simulation and Benchmarking
Daniel Dauner, Marcel Hallgarten, Tianyu Li, Xinshuo Weng, Zhiyu Huang, Zetong Yang, Hongyang Li, Igor Gilitschenski, B. Ivanovic, Marco Pavone, Andreas Geiger, Kashyap Chitta
arXiv:2406.15349 · 2026-07-27 공개 · arXiv · PDF
benchmarking navsim autonomous-vehicles bird-eye-view transfuser uniad cvpr-competition non-reactive-simulation
Abstract
Benchmarking vision-based driving policies is challenging. On one hand, open-loop evaluation with real data is easy, but these results do not reflect closed-loop performance. On the other, closed-loop evaluation is possible in simulation, but is hard to scale due to its significant computational demands. Further, the simulators available today exhibit a large domain gap to real data. This has resulted in an inability to draw clear conclusions from the rapidly growing body of research on end-to-end autonomous driving. In this paper, we present NAVSIM, a middle ground between these evaluation paradigms, where we use large datasets in combination with a non-reactive simulator to enable large-scale real-world benchmarking. Specifically, we gather simulation-based metrics, such as progress and time to collision, by unrolling bird's eye view abstractions of the test scenes for a short simulation horizon. Our simulation is non-reactive, i.e., the evaluated policy and environment do not influence each other. As we demonstrate empirically, this decoupling allows open-loop metric computation while being better aligned with closed-loop evaluations than traditional displacement errors. NAVSIM enabled a new competition held at CVPR 2024, where 143 teams submitted 463 entries, resulting in several new insights. On a large set of challenging scenarios, we observe that simple methods with moderate compute requirements such as TransFuser can match recent large-scale end-to-end driving architectures such as UniAD. Our modular framework can potentially be extended with new datasets, data curation strategies, and metrics, and will be continually maintained to host future challenges. Our code is available at https://github.com/autonomousvision/navsim.
한국어 요약
한 줄 요약
NAVSIM은 비반응형 시뮬레이션과 대규모 데이터셋을 결합하여 자율주행 정책을 평가하는 새로운 벤치마킹 프레임워크이다.
핵심 기여도
- NAVSIM은 비반응형 시뮬레이션과 실제 데이터를 결합하여 대규모 실세계 벤치마킹을 가능하게 한다.
- TransFuser와 같은 단순한 모델이 UniAD와 같은 대규모 모델과 유사한 성능(PDMS 84.0)을 보인다.
- CVPR 2024에서 143개 팀이 463개 제출물을 제출하며, 새로운 인사이트를 도출.
- PDMS(Policy Driving Metric Score)와 같은 다양한 시뮬레이션 기반 메트릭을 제안.
핵심 아이디어
기존 자율주행 정책 평가 방식은 open-loop 평가가 실세계 성능을 반영하지 못하고, closed-loop 시뮬레이션은 계산 비용이 높으며 도메인 갭이 큰 문제를 안고 있다. NAVSIM은 이 두 접근법의 중간 지점을 제시한다. 비반응형 시뮬레이션을 통해 정책과 환경 간의 상호작용 없이, 짧은 시간 내에 조건부 시나리오를 테스트하며, 실제 데이터셋과 결합하여 대규모 평가를 가능하게 한다. 특히, BEV(상공 시점) 추상화를 활용한 시뮬레이션은 시간 충돌(TTC), 진행률(Progress) 등의 메트릭을 효과적으로 계산할 수 있다. 이는 기존의 평균 이동 오차(ADE)보다 실제 주행 성능과 더 잘 일치하는 것으로 입증되었다.
기술적 접근법
- **비반응형 시뮬레이션**: 정책과 환경 간의 상호작용 없이, 주어진 센서 입력에 따라 고정된 시간 내의 행동을 계획.
- **BEV 추상화**: 시뮬레이션은 상공 시점에서 진행되며, TTC, 진행률(PDMS) 등의 메트릭을 계산.
- **데이터셋**: OpenScene 데이터셋을 기반으로 10만 개 이상의 실제 주행 시나리오를 추출.
- **모델 구현**: TransFuser, UniAD, PARA-Drive 등 다양한 end-to-end 모델을 재구현.
- **하이퍼파라미터**: TransFuser는 1 GPU로 하루, PARA-Drive는 80 GPU로 3일 동안 훈련.
- **메트릭**: PDMS, DAC(주행 가능 영역 준수), EP(자기 진행률) 등.
주요 결과
- TransFuser는 PDMS 84.0, LTF는 83.8, UniAD는 83.4를 기록하며, 단순 모델이 대규모 모델과 유사한 성능을 보임.
- Ego Status MLP는 PDMS 65.6, Constant Velocity는 하위 기준.
- CVPR 2024 NAVSIM 챌린지에서 143개 팀이 463개 제출물 제출.
- DAC과 EP는 가장 어려운 메트릭으로, 최고 점수와 인간 운전자의 점수 간 10점 차이 관찰.
의의 및 한계
NAVSIM은 기존 자율주행 평가의 주요 문제점(도메인 갭, 계산 비용, 메트릭 불일치)을 해결하며, 표준화된 평가 환경을 제공한다. 특히, 비반응형 시뮬레이션은 open-loop 평가의 단순성과 closed-loop 평가의 현실성 사이의 균형을 맞춘다. 그러나 시뮬레이션이 반응형이 아니기 때문에, 다중 에이전트 간의 실시간 상호작용을 반영하지 못하는 한계가 있다. 또한, 메트릭은 시뮬레이션 기반으로 제한적이며, 실제 주행 환경에서의 장기적 성능은 추가 연구가 필요하다.
실용적 활용
NAVSIM은 자율주행 알고리즘 개발자들이 다양한 주행 시나리오에서 모델을 대규모로 평가할 수 있는 플랫폼으로 활용될 수 있다. 특히, TransFuser와 같은 저비용 고성능 모델의 발견은 산업 현장에서의 실용화를 촉진할 수 있다. 또한, HuggingFace 기반의 공개 평가 서버를 통해 연구자들이 투명하고 공정하게 모델을 비교할 수 있다.