Dream4ACT: A Shared Visual Action Interface for Multi-Embodiment Video-Action Modeling

Xiangyu Zhu, Jin Xu, Yue Guo, Xin Wu, Yifan Sun, Xiancong Ren, Jianxin Sun, Yong Dai, Xiaozhu Ju

arXiv:2609.40153 · 2026-10-05 공개 · arXiv · PDF

diffusion-transformer multi-embodiment robotwin-2-0 urdf-based-kinematics masked-flow-matching robot-execution trworldbench action-views

Abstract

Video generation models (VGMs) offer strong spatiotemporal priors for embodied observation--action modeling. However, joint-space action vectors lack explicit image-space structure and vary in dimensionality and semantics across embodiments, making it challenging to directly leverage the rich spatiotemporal priors of VGMs. End-effector visualizations provide an alternative but do not specify the full articulated configuration needed for robot execution. We present Dream4ACT, a world model built for joint video-action modeling across embodiments. To unify action representations across embodiments, we introduce a shared visual action interface, called action views, which render target joint configurations from four prescribed virtual cameras using URDF-based forward kinematics. This shared visual representation preserves embodiment-specific articulated geometry while allowing observation and action sequences to share a video autoencoder and diffusion transformer. Through masked flow-matching, our model supports forward dynamics, inverse dynamics, and joint observation--action generation within a single jointly trained model by varying which future sequences are corrupted. To recover executable action sequences from predicted action views, we propose a training-free, URDF-constrained multiview recovery mechanism, without a learned embodiment-specific decoder. Dream4ACT achieves an average success rate of 88.98\% on RoboTwin~2.0 and an overall score of 65.66 on TriWorldBench, supporting effective closed-loop manipulation and competitive action-conditioned multiview prediction through the visual action interface.

한국어 요약

한 줄 요약

Dream4ACT는 다양한 로봇 구조에 공통적으로 적용 가능한 시각적 행동 인터페이스를 도입하여 88.98% 성공률을 달성한 비디오-행동 월드 모델이다.

핵심 기여도

핵심 아이디어

기존 로봇 행동 모델링에서는 관절 공간 벡터가 이미지 구조를 가지지 않으며, 다양한 로봇 구조 간 차원과 의미가 달라져 VGM 활용이 어려웠다. Dream4ACT는 이 문제를 해결하기 위해 **action views**라는 고정 크기의 다중 뷰 시각 인터페이스를 도입한다. 이 인터페이스는 URDF 기반 정방향 운동학을 통해 4개의 가상 카메라로 관절 설정을 렌더링하며, 관절 구조 정보를 유지하면서도 비디오 오토인코더와 디퓨전 트랜스포머를 공유할 수 있도록 한다.

또한, **masked flow-matching** 알고리즘을 통해 미래 시퀀스의 손상 여부를 조절하여 단일 모델 내에서 순방향 역학, 역방향 역학, 관측-행동 생성을 지원한다. 예측된 action views를 실행 가능한 관절 목표로 복원할 때는 학습된 디코더 없이 URDF 제약 조건을 활용한 최적화 기법을 사용하여, 로봇 구조에 맞는 정확한 복원이 가능하다.

기술적 접근법

주요 결과

의의 및 한계

Dream4ACT는 다양한 로봇 구조 간 공통 인터페이스를 제공함으로써 VGM의 강력한 시공간 사전 정보를 효과적으로 활용할 수 있는 새로운 접근법을 제시한다. 특히, **action views**는 관절 구조 정보를 유지하면서도 이미지 기반 인코딩을 가능하게 하며, **training-free recovery**는 별도의 디코더 학습 없이 실행 가능한 관절 목표를 복원할 수 있어 실용적이다.

그러나, **Franka-Panda 및 UR5-Xsg**와 같은 이중 팔 시뮬레이션 구조는 통합 설계가 아닌 단일 팔 로봇의 조합으로 구성되어 있어, 팔 간 조화나 정밀한 관절 추적이 어려운 한계가 있다. 이는 URDF 기반 운동학 모델이 실제 실행 효과를 완전히 반영하지 못함을 의미하며, 이로 인해 정확한 관절 목표라도 실행 시 오차가 발생할 수 있다.

실용적 활용

Dream4ACT는 다양한 로봇 구조를 지원하는 공통 인터페이스를 통해 **다중 로봇 시스템**, **로봇 제어 및 예측**, **시뮬레이션-현실 간 전이 학습** 등에 적용 가능하다. 특히, **training-free recovery** 기법은 실시간 제어 및 복잡한 로봇 구조에 대한 디코더 학습 없이 실행 가능한 명령을 생성할 수 있어 산업 현장에서 유용하게 활용될 수 있다.