VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models

arXiv:2609.04355 · 2026-09-28 공개 · arXiv · PDF

reinforcement-learning vision-language-action robotics throughput-optimization online-rl closed-loop-architecture asymmetric-co-bootstrapping chemistry-tasks

Abstract

Pretrained vision-language-action (VLA) models enable broad manipulation but remain unreliable in tasks demanding precision and repeatability. Applying real-world online reinforcement learning (RL) to VLA post-training enables autonomous trial-and-error improvement beyond demonstrations alone, but exposes two bottlenecks: 1) unreliable value signals can induce policy drift; 2) large-VLA overhead constrains throughput and sample efficiency. To address these challenges, we present VLA-Precision, an efficient real-world online RL framework featuring the Asymmetric Co-Bootstrapping (ACoB) algorithm and the ACoB-Stream architecture. Specifically, ACoB establishes asymmetric co-bootstrapping across timescales: early intervention-guided behavioral learning rapidly improves policy performance while enhancing online experience quality. As autonomous experience accumulates, global return propagation and local preference ranking progressively calibrate value estimates, yielding relative action advantages for reference-regularized policy improvement while suppressing drift. To enable ACoB on large VLAs, we develop ACoB-Stream, a closed-loop experience--policy architecture that establishes invariant-state decoupling and on-demand streaming as design principles, delivering up to 10.9times improvements in throughput and computational efficiency. Extensive evaluations on nine high-precision chemistry tasks across four categories and four robot embodiments show that VLA-Precision achieves 98.3\% mean success rate in 45.8 min/task, with 27.6 s episodes running at 1.2times and 1.8times the speeds of VLA and RL baselines. Resources are available at https://vla-precision.github.io.

한국어 요약

한 줄 요약

VLA-Precision은 정밀한 실제 환경 온라인 강화학습을 위한 비대칭 공부트스트랩 알고리즘과 아키텍처를 제시한다.

핵심 기여도

핵심 아이디어

VLA 모델은 시각-언어-액션 통합을 통해 다양한 조작을 가능하지만, 정밀도와 반복성에서는 한계가 있다. 기존 온라인 강화학습은 정책 드리프트와 처리량 저하라는 두 가지 주요 문제를 일으킨다. 이를 해결하기 위해 VLA-Precision은 ACoB 알고리즘을 도입하여 시간축에서 비대칭 공부트스트랩을 수행한다. 초기 개입을 기반으로 행동 학습을 빠르게 진행하면서 온라인 경험 품질을 향상시키고, 자율 경험 축적 후에는 전역 수익 전파와 지역 선호도 순위를 통해 가치 추정을 점진적으로 교정한다. 이는 정책 개선의 상대적 액션 이점을 생성하면서 드리프트를 억제한다.

기술적 접근법

주요 결과

의의 및 한계

VLA-Precision은 실제 환경에서 정밀한 온라인 강화학습을 가능하게 하며, 대규모 VLA 모델의 처리량 문제를 해결하는 데 기여한다. 특히, 비대칭 공부트스트랩은 정책 드리프트를 억제하면서 정확도를 향상시킨다. 그러나, 특정 하이퍼파라미터나 데이터셋의 세부 사항은 명시되지 않았으며, 다른 태스크나 환경에서의 일반화 가능성은 추가 연구가 필요하다.

실용적 활용

화학 실험, 제조, 의료 등 정밀한 조작이 필요한 산업에서 로봇 자동화를 가능하게 하며, 시각-언어-액션 통합 모델의 실용적 활용을 확장할 수 있다.