World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation

Yuhao Pan, Haosong Peng, Zhengshen Zhang, Zhengyang Yan, Yalun Dai, Fushuo Huo, Chujie Wang, Tianyu Qi, Xiucheng Wang, Nan Cheng, Wenchao Xu

arXiv:2608.05369 · 2026-08-10 공개 · arXiv · PDF

vision-language-action robot-manipulation libero bimanual-manipulation latent-interface w2-vla w2-cot task-conditioning

Abstract

Vision-language-action (VLA) models often treat main-view and wrist-view observations as parallel visual inputs, overlooking their distinct roles in robot manipulation. Fine-grained manipulation, however, benefits from anticipating how wrist-local interactions may evolve under the global task context. To address this limitation, we present World-to-Wrist VLA (W2-VLA), a VLA model for fine-grained robot manipulation with task-conditioned future wrist modeling. Given current multi-view observations and a task instruction, W2-VLA contextualizes a set of latent modeling tokens as a compact interface between the vision-language model and the wrist predictor. Conditioned on this interface and the observed wrist history, the predictor forecasts future wrist latents, which are transformed into future-aware context for action prediction. In addition, we introduce W2-CoT, a synthesis pipeline that produces structured annotations describing manipulation progress, physical transition cues, and wrist-local evidence. These annotations provide auxiliary supervision that shapes the task-conditioned latent interface. Experiments on LIBERO, RoboTwin 2.0, and real-world manipulation tasks demonstrate improved fine-grained and contact-sensitive manipulation across both single-arm and bimanual settings, while maintaining action-generation rates above 80 Hz.

한국어 요약

한 줄 요약

W2-VLA는 글로벌 작업 맥락을 고려한 미래 손목 움직임 예측을 통해 정밀 로봇 조작 성능을 향상시키는 VLA 모델이다.

핵심 기여도

핵심 아이디어

기존 VLA 모델은 메인 뷰와 손목 뷰를 병렬 입력으로 처리해, 손목 관찰의 로컬 상호작용 역할을 간과했다. 그러나 정밀 조작에서는 글로벌 맥락 하에서 손목 로컬 상호작용의 미래 예측이 필수적이다. W2-VLA는 **latent modeling tokens**을 통해 VLM(비전-언어 모델)과 손목 예측기 사이의 인터페이스를 형성하고, 이 인터페이스를 기반으로 **future wrist latents**를 예측한다. 이는 현재 손목 히스토리와 결합되어 액션 생성에 미래 정보를 제공한다. 또한, **W2-CoT**는 조작 진행, 물리적 전이 신호, 손목 로컬 증거를 포함한 구조화된 어노테이션을 생성해, 인터페이스 학습을 보조한다. 이는 **task-conditioned supervision**을 통해 정밀한 미래 예측을 가능하게 한다.

기술적 접근법

주요 결과

의의 및 한계

W2-VLA는 정밀 조작에서 **글로벌-로컬 상호작용**을 효과적으로 모델링하며, **실시간 액션 생성**과 **OOD 환경 내성**을 동시에 달성한 점에서 학술적·실용적 의의가 있다. 특히, **W2-CoT**는 구조화된 어노테이션을 통해 학습 효율성을 높인다. 그러나, **복잡한 다중 작업**이나 **장기적 계획**에서는 추가적인 연구가 필요할 수 있다. 또한, **손목 뷰만 예측**하는 선택은 일부 상황에서 정보 손실을 초래할 수 있다.

실용적 활용

W2-VLA는 **로봇 조립**, **정밀 산업 자동화**, **의료 로봇** 등에서 **접촉 민감한 조작**이 필요한 상황에 적용 가능하다. 특히, **실시간 성능**과 **복잡한 작업 환경 내성**을 요구하는 산업 현장에서 유용할 것으로 기대된다.