Robostral Navigate
Arjun Majumdar, Avinash Sooriyarachchi, Benjamin Tibi, Chris Bamford, Elliot Chane-Sane, Guillaume Lample, Khyathi Raghavi Chandu, Ludovic Ho Fuh, Mathieu Poiree, Olivier Duchenne, Rosalie Millner, Srijan Mishra, Theo Cachet, Thomas Chabal
arXiv:2607.20785 · 2026-07-25 공개 · arXiv · PDF
reinforcement-learning vision-language sim-to-real robot-navigation prefix-caching monocular-rgb waypoint-prediction r2r-ce
Abstract
Deploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across robot embodiments, and trains efficiently. Yet, today's best systems depend on depth sensors, multi-camera rigs, or pre-built maps, limiting the hardware they support and increasing deployment cost. We introduce Robostral Navigate, an 8B vision-language model built around this scalability objective. The model consumes only a stream of monocular RGB images - the most ubiquitous sensor across robotic platforms and predicts waypoints by pointing to the next target location in the current camera view. Operating purely in image space, rather than robot-specific coordinates, makes the policy naturally robust to changes in camera intrinsics and scene scale, enabling deployment across wheeled, legged, and aerial robots without recalibration. We generate 2.4 million trajectories across 350k simulated scenes to reduce the reliance on real-world data collection and scale easily. We further introduce a prefix-caching training recipe that packs entire episodes into single training sequences, reducing training tokens by 22x and cutting training time from months to days. A tree-based attention mask prevents conditioning on previous ground-truth actions, encouraging visually grounded action prediction, and reinforcement learning is used to further improve exploration and recovery capabilities. On the Room-to-Room and Room-Across-Room in Continuous Environments (R2R-CE and RxR-CE) benchmarks, Robostral Navigate sets a new state of the art. On R2R-CE, it achieves a 77.4% success rate, surpassing the best monocular method by 10.5 points and the strongest depth- or multi-camera system by 5.3 points despite using only a single RGB camera. On RxR-CE, it reaches 75.1% success rate, outperforming all monocular baselines.
한국어 요약
한 줄 요약
Robostral Navigate는 단일 RGB 카메라만 사용해 R2R-CE에서 77.4% 성공률을 달성한 8B 규모의 시각-언어 네비게이션 모델이다.
핵심 기여도
- **8B 규모의 VLM**을 사용하여 단일 RGB 카메라만으로도 뛰어난 성능 달성.
- **22× 훈련 토큰 감소**를 유도하는 prefix-caching 훈련 레시피 제안.
- **R2R-CE에서 77.4% 성공률**, RxR-CE에서 75.1% 성공률 달성.
- **CISPO 기반 강화 학습**으로 실패 복구 능력 향상.
핵심 아이디어
Robostral Navigate는 단일 RGB 카메라만을 입력으로 받아, 현재 카메라 뷰 내에서 다음 목적지의 픽셀 좌표를 예측함으로써 웨이포인트를 생성한다. 이는 기존 시스템들이 로봇 기하학에 의존하는 좌표 기반 예측과 달리, **이미지 공간에서의 웨이포인트 예측**을 통해 카메라 내부 파라미터나 환경 스케일 변화에 대한 내성을 자연스럽게 제공한다.
또한, **prefix-caching**을 통해 전체 에피소드를 단일 훈련 시퀀스로 압축하여 훈련 토큰을 22× 줄이고, 훈련 시간을 수개월에서 수일로 단축했다. **트리 기반 어텐션 마스크**는 이전의 정답 행동에 의존하지 않고, 시각 정보에 기반한 행동 예측을 장려한다.
기술적 접근법
- **모델 아키텍처**: 8B 파라미터를 가진 VLM.
- **입력**: 언어 지시문과 단일 RGB 카메라 스트림.
- **웨이포인트 예측**: 현재 카메라 뷰 내의 픽셀 좌표와 도착 시 원하는 방향.
- **훈련 데이터**: 350k개 시뮬레이션 환경에서 생성된 2.4백만 개의 트레젝토리.
- **훈련 기법**:
- Prefix-caching: 훈련 토큰 22× 감소.
- 트리 기반 어텐션 마스크: 이전 정답 행동에 의존하지 않도록 유도.
- CISPO 기반 강화 학습: 실패 복구 능력 향상.
주요 결과
- **R2R-CE**: 77.4% 성공률 (단일 카메라 기준 최고 66.9% 대비 +10.5%, 72.1% 대비 +5.3%).
- **RxR-CE**: 75.1% 성공률 (모든 단일 카메라 기준 모델을 초과).
- **SPL (Success weighted by Path Length)**: 68.7% (RxR-CE 기준).
- **강화 학습 적용 시 성공률 4% 추가 향상**.
의의 및 한계
Robostral Navigate는 **센서 가정 최소화**, **로봇 형태 일반화**, **시뮬레이션 기반 훈련 효율성**이라는 세 가지 목표를 동시에 달성한 첫 사례로, 일반 목적 로봇 개발에 중요한 기반을 제공한다. 특히, **단일 RGB 카메라만으로도 기존의 깊이 센서나 다중 카메라 시스템을 능가하는 성능**을 보여주며, 센서 비용과 칼리브레이션 부담을 줄이는 데 기여한다.
하지만, **실제 환경에서의 일반화 능력**이나 **복잡한 장애물 회피 능력**은 명시되지 않았으며, **실제 로봇에 적용한 성능**도 아직 보고되지 않았다.
실용적 활용
Robostral Navigate는 **단일 카메라만 탑재된 다양한 로봇 플랫폼**(이동 로봇, 다리 로봇, 드론 등)에 쉽게 적용 가능하다. 특히, **저비용 센서 환경**이나 **시뮬레이션 기반 훈련이 필수적인 산업**(예: 물류, 서비스 로봇)에서 유용하게 활용될 수 있다.