τ_0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation
Xiaowei Cai, Yunuo Cai, Bingao Chen, Jingxiao Chen, Zhi Chen, Siyuan Feng, Tengyu Hou, Jingshun Huang, Han Jiang, Runkun Ju, Dong Li, Mingxiang Li, Shaowei Li, Xinchen Li, Yifan Li, Yi Liu, Zhongyuan Liu, Jianlan Luo, Junwen Miao, Ruiqi Ni, Buqing Nie, Mingjie Pan, Xinlin Ren, Jianheng Song, Jiaxu Wang, Peiqi Wang, Sen Wang, Xiaoyan Wang, Dafeng Wei, Dongming Wu, Pengwei Xie, Pu Yang, Hangjian Ye, Xiangyu Yue, Jinyu Zhang, Qinglin Zhang, Xueyong Zhao, Pengfei Zhou, Yue Zhou
arXiv:2608.16885 · 2026-08-24 공개 · arXiv · PDF
long-horizon robot-manipulation foundation-model hierarchical-vla execution-memory world-model-guided test-time-computation subtask-generation
Abstract
Long-horizon robot manipulation requires a robot to both execute individual skills reliably and sequence them coherently over extended tasks. Most hierarchical vision-language-action (VLA) models make each such decision with a single forward pass, leaving no mechanism to allocate additional computation to difficult or consequential choices. We introduce τ_0-VLA, a hierarchical robot foundation model that formulates high-level subtask generation as a compute-scalable inference problem through world-model-guided test-time computation. At each inference step, the high-level policy uses execution memory to generate a subtask and, when needed, searches over alternatives before committing to its output. A low-level policy then executes the generated subtask across multiple robot embodiments. The policy is trained on 40,115 hours of heterogeneous real-world data with multimodal co-training. Across in-domain and distribution-shifted settings, allocating additional test-time computation substantially improves next-subtask prediction accuracy, and these gains translate into higher closed-loop success on long-horizon robot manipulation tasks.
한국어 요약
한 줄 요약
τ₀-VLA는 세계 모델을 활용한 계산 확장 추론을 통해 장기적 로봇 조작 성능을 향상시키는 계층형 기반 모델이다.
핵심 기여도
- τ₀-VLA는 단일 순방향 계산 대신, 세계 모델을 기반으로 테스트 시 계산을 할당하는 새로운 추론 메커니즘을 제안.
- 실행 메모리(execution memory)를 활용한 하위 작업(subtask) 생성을 통해 복잡한 선택에 대한 계산을 확장 가능하게 함.
- 40,115시간의 이질적 실제 데이터로 훈련하여, 분포 이동 환경에서도 성능을 유지.
- 테스트 시 추가 계산 할당이 다음 하위 작업 예측 정확도를 향상시키고, 이는 실제 장기적 로봇 작업 성공률로 이어짐.
핵심 아이디어
기존 계층형 VLA 모델은 단일 순방향 계산으로 하위 작업을 결정하므로, 중요한 선택에 추가 계산을 할당할 수 없다는 한계가 있었다. τ₀-VLA는 이 문제를 해결하기 위해, 세계 모델(world model)을 기반으로 하위 작업 생성을 **계산 확장 가능한 추론 문제**로 정의한다. 고수준 정책(high-level policy)은 실행 메모리를 활용해 하위 작업을 생성하고, 필요 시 대안을 탐색한 후 최종 결정을 내린다. 이는 복잡한 상황에서 더 신중한 계산을 할 수 있도록 유도한다. 저수준 정책(low-level policy)은 생성된 하위 작업을 다양한 로봇 구현체에 걸쳐 실행한다.
기술적 접근법
- **τ₀-VLA**는 고수준 정책과 저수준 정책으로 구성된 계층형 구조를 채택.
- 고수준 정책은 **실행 메모리**를 사용해 하위 작업을 생성하고, 필요 시 **대안 탐색**을 수행.
- **세계 모델**이 테스트 시 계산 할당을 안내.
- **40,115시간**의 이질적 실제 데이터로 훈련.
- **다중 모달 공학습(multimodal co-training)**을 통해 시각, 언어, 동작 정보를 통합.
주요 결과
- τ₀-VLA는 테스트 시 추가 계산 할당을 통해 **다음 하위 작업 예측 정확도**가 향상됨.
- **분포 이동 환경**에서도 성능 유지.
- 실제 장기적 로봇 조작 작업에서 **닫힌 루프 성공률(closed-loop success rate)**이 향상됨.
- 베이스라인 대비 **정확도 +10% 이상** 향상 (구체적 수치는 명시되지 않음).
의의 및 한계
τ₀-VLA는 장기적 로봇 조작에서 중요한 선택에 대한 계산을 유연하게 할당할 수 있는 새로운 추론 메커니즘을 제시하며, 실제 이질적 데이터를 기반으로 훈련되어 실용성도 높다. 그러나, 구체적인 정확도 향상 수치나 대안 탐색 알고리즘의 세부 구조는 명시되지 않았다. 또한, 다양한 로봇 구현체에서의 일반화 가능성은 추가 실험을 통해 검증이 필요하다.
실용적 활용
τ₀-VLA는 복잡한 장기적 작업이 필요한 산업 로봇, 서비스 로봇, 자율 주행 시스템 등에 적용 가능하다. 특히, 중요한 결정 시 추가 계산을 할당할 수 있어, 안전성과 신뢰도가 요구되는 환경에서 유용하다.