World Action Learning via Interaction-Centric Spectral Latent Guidance

Zhiming Liu, Yikun Miao, Ying Chen, Hongrui Yin, Fangqi Zhu, Xiaoyi Pang, Quanxin Shou, Zhengyang Yan, Haodong Wang, Song Guo

arXiv:2610.03607 · 2026-10-07 공개 · arXiv · PDF

libero egocentric-video action-generation robotwin robocasa temporal-dynamics latent-action robot-policy

Abstract

Learning general-purpose robot policies requires large-scale real-world interaction data, yet collecting robot demonstrations remains expensive and difficult to scale. Egocentric videos offer abundant human interaction experience with task-relevant semantics for robotic manipulation, but direct transfer is challenging for two reasons: latent actions inferred from frame reconstruction can be dominated by nuisance variation such as ego-camera motion, and human and robot behaviors often exhibit different temporal dynamics. We propose WING (World Action Learning via INteraction-Centric Spectral Latent Guidance), a framework for transferring interaction knowledge from egocentric videos to robot policies. WING first separates observer-induced motion from hand-object interaction and distills the interaction-centric component into latent actions. It then exploits the observation that cross-embodiment task semantics are concentrated in slowly varying temporal structures, identifying shared low-frequency components between egocentric latent actions and robot behaviors in the spectral domain and using them to guide action generation. WING achieves average success rates of 99.20% on LIBERO, 93.80% on RoboTwin 2.0, and 57.7% on RoboCasa-GR1, and also performs strongly across four real-world manipulation tasks under diverse generalization settings. These results show that interaction-centric spectral guidance provides an effective and scalable way to transfer physical interaction knowledge from human egocentric video to robot control. Project page: https://mikuz12.github.io/wing/

한국어 요약

한 줄 요약

WING은 인간의 제1인칭 비디오에서 로봇 제어로 상호작용 지식을 전이하는 새로운 프레임워크로, 관찰자 유발 운동을 분리하고 스펙트럼 영역의 저주파 성분을 활용하여 높은 성능을 달성한다.

핵심 기여도

핵심 아이디어

기존 잠재 행동 학습 방법은 관찰자 유발 운동(예: 카메라 이동)에 민감하여, 실제 상호작용 정보를 효과적으로 추출하지 못한다. WING은 이 문제를 해결하기 위해 **WING-LAM**(Interaction-Centric Latent Action Model)을 도입하여, 상호작용 관련 운동만을 추출하고, 관찰자 유발 운동은 억제한다. 또한, 인간과 로봇의 행동이 **저주파 성분**에서 더 높은 일관성을 보인다는 관찰을 바탕으로, **DCT**를 사용해 저주파 성분만 추출하여 로봇 제어에 활용한다. 이는 시간 영역에서의 복잡한 패턴을 줄이고, 공통된 의미 있는 행동 구조만을 강조하는 핵심 아이디어이다.

기술적 접근법

주요 결과

의의 및 한계

WING은 대규모 제1인칭 비디오에서 로봇 제어로 상호작용 지식을 전이하는 새로운 접근법을 제시하며, 관찰자 유발 운동을 억제하고 저주파 성분을 활용함으로써 기존 방법보다 훨씬 높은 성능을 달성한다. 특히, **WING-LAM**과 **DCT 모듈**의 결합은 잠재 행동 학습과 정책 생성 사이의 일관성을 높이는 데 기여한다. 그러나, **RoboCasa-GR1**과 같은 복잡한 환경에서는 성능이 상대적으로 낮아, 더 복잡한 작업에 대한 일반화 능력 향상이 필요하다. 또한, **DCT 성분 수**를 결정하는 기준은 경험적이고, 이에 따른 최적화가 여전히 필요하다.

실용적 활용

WING은 대규모 인간 제1인칭 비디오를 활용한 로봇 정책 학습에 적용 가능하며, 특히 **로봇 조작**, **이중 조작**, **휴먼-로봇 협업** 등 다양한 작업 환경에서 유용할 수 있다. 또한, **실제 세계에서의 일반화 능력**을 고려할 때, **로봇 제조**, **의료 로봇**, **서비스 로봇** 분야에서 실용적 활용이 기대된다.