IMPACT: Attention Is the Interaction Map for Scalable Interaction-Aware World Model Training

Rongze Tang, Jianjie Fang, Zhaolu Wang, Ziyou Wang, Xvyuan Liu, Haisheng Su, Xin Zhang, Wei Wu, Chen Gao, Yong Li, Zhibo Chen

arXiv:2609.00161 · 2026-09-02 공개 · arXiv · PDF

world-models robot-manipulation cross-attention attention-mechanism interaction-aware physical-plausibility denoising-objective di-t-backbone

Abstract

World models have made remarkable progress in action-conditioned future prediction for embodied agents, yet still struggle to model physically plausible interactions. Existing approaches address this limitation by constraining the generation process with external representations encoding motion, geometry, or semantics. Obtaining these spatiotemporally dense representations typically requires auxiliary estimators or manual annotations, limiting training scalability. We instead revisit the training objective and identify a supervision-allocation mismatch under the globally averaged mean squared error (MSE) denoising objective: prevalent static content dominates the optimization signal, leaving sparse dynamic-object regions critical to interaction generation disproportionately under-supervised. Motivated by this observation, we introduce IMPACT, a scalable Interaction-aware Model training framework with Prior-guided Attention Calibration and Targeting. IMPACT uses cross-attention associated with manipulated-object tokens as an internal spatiotemporal prior for action-conditioned changes. It samples candidate regions from this prior, calibrates them with detached local prediction errors to construct an interaction map, and uses the map to reweight denoising supervision, requiring neither external representations nor inference-time modifications. Extensive experiments on robot-arm and human-hand manipulation, spanning diverse control modalities and DiT backbones, show that IMPACT consistently outperforms the corresponding MSE-trained baselines, improving interaction fidelity, physical plausibility, and visual quality.

한국어 요약

한 줄 요약

IMPACT는 외부 표현 없이 주의력 기반 상호작용 지도를 통해 세계 모델의 상호작용 생성을 향상시키는 확장 가능한 훈련 프레임워크이다.

핵심 기여도

핵심 아이디어

기존 세계 모델은 전역 평균 MSE 손실로 훈련되며, 정적 배경에 비해 상호작용 영역이 훈련 신호에서 상대적으로 소외된다. 이는 상호작용 생성 시 물리적 일관성과 정확도를 저하시킨다. IMPACT는 이 문제를 해결하기 위해, 대규모 사전 훈련된 비디오 생성 백본이 이미 포함하고 있는 **cross-attention 맵**을 상호작용 영역의 내부 prior로 활용한다. 이 prior는 조작된 객체 토큰과 연관된 주의력 분포를 기반으로 하며, **Attention Distribution Sampling (ADS)**을 통해 이 prior를 기반으로 후보 영역을 샘플링하고, **detached local prediction error**로 이를 보정하여 정밀한 상호작용 맵을 생성한다. 이후 **Interaction-Weighted Supervision (IWS)**을 통해 이 맵을 사용하여 상호작용 영역에 대한 지도를 재가중함으로써, 상호작용의 기능적 중요도에 따라 훈련 신호를 재분배한다. 이 방식은 외부 밀도 표현 없이도 확장 가능한 훈련을 가능하게 한다.

기술적 접근법

주요 결과

의의 및 한계

IMPACT는 외부 밀도 표현 없이도 세계 모델의 상호작용 생성을 향상시킬 수 있는 확장 가능한 훈련 프레임워크로, 기존 접근법의 비용 문제를 해결한다. 특히, **cross-attention 기반 내부 prior 활용**은 기존 모델의 잠재 능력을 최대화하는 새로운 통찰을 제공한다. 그러나, **cross-attention이 항상 정확한 prior를 제공하지는 않으며**, 일부 복잡한 상호작용 상황에서는 **prior가 부정확하게 추정될 수 있는 한계**가 존재한다. 또한, **기존 DiT 백본에 의존**하므로, 백본의 성능에 따라 결과가 달라질 수 있다.

실용적 활용

IMPACT는 **로봇 팔 조작**과 **인간 손 조작**과 같은 다양한 **임베디드 시뮬레이션 및 제어 시스템**에 적용 가능하다. 특히, **대규모 데이터셋이 필요한 외부 표현 없이도** 상호작용 생성을 향상시킬 수 있어, **로봇 학습 및 시뮬레이션 기반 정책 평가**에 유용하게 활용될 수 있다.