HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement

Yiyang Cai, Nan Chen, Rongchang Xie, Junwen Pan, Chunyang Jiang, Cheng Chen, Wen Zhou, Zhenbang Sun, Wei Xue, Wenhan Luo, Yike Guo

arXiv:2607.18217 · 2026-07-21 공개 · arXiv · PDF

self-attention semantic-alignment video-personalization vae-guidance modality-reference subject-driven-generation intra-subject-reference multimodal-enhancement

Abstract

Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. However, existing methods suffer from two key limitations. First, most approaches focusing on inter-subject personalization still struggle to strike a balance between high subject fidelity and accurate interaction patterns between humans and diverse objects, especially when objects represent abstract concepts such as logos. Second, while intra-subject references (e.g., OCR maps, multi-view inputs) are expected to enhance subject fidelity, most existing works lack mechanisms to understand such latent correspondence. To address both challenges, we propose HOMIE, an HOCVP framework that tackles both inter- and intra-subject input settings in a unified manner. Compared to previous approaches, HOMIE proposes a better MLLM integration strategy to extract knowledge of reference-level relationships without compromising the controllability of text encoders or incurring costly re-alignment. Specifically, we introduce global multimodal guidance within self-attention to better align MLLM-derived semantic features with VAE tokens. Furthermore, we propose modality-reference embedding to differentiate tokens from MLLM features and VAE tokens and associate intra-subject reference image tokens. Extensive experiments validate that our method achieves state-of-the-art performance across various HOCVP tasks. Project Page: https://yiyangcai.github.io/homie-page.github.io/

한국어 요약

한 줄 요약

HOMIE는 인물-객체 중심 동영상 개인화 작업에서 MLLM과 VAE 토큰의 정확한 정렬을 달성한 통합 프레임워크로, GMG와 MRE 모듈을 통해 SOTA 성능을 보인다.

핵심 기여도

핵심 아이디어

기존 HOCVP 연구는 inter-subject 개인화 시 인물-객체 상호작용의 정확성과 객체 추상성(예: 로고) 처리, intra-subject 참조 간 의미적 연관성 추출이라는 두 가지 주요 문제를 해결하지 못했다. HOMIE는 이 두 문제를 동시에 해결하기 위해 MLLM과 VAE 기반 디퓨전 모델을 통합하는 새로운 패러다임을 제시한다. GMG는 MLLM의 의미 정보를 self-attention 과정에 주입하여 시간적 상호작용 정보를 풍부하게 제공하고, MRE는 토큰의 모달별 구분과 intra-subject 참조 간 연관성을 효과적으로 캡처한다. 이는 UmT5 텍스트 인코더의 제어력을 유지하면서 MLLM의 재정렬 비용을 줄이는 동시에, 의미 추출 능력을 극대화한다.

기술적 접근법

주요 결과

의의 및 한계

HOMIE는 MLLM의 의미 추출 능력을 HOCVP에 효과적으로 통합하여, 기존 방법이 해결하지 못했던 inter- 및 intra-subject 참조 간 정확한 상호작용을 가능하게 한다. 특히, 로고와 같은 추상적 객체 처리와 다각도 참조 이미지 간 의미적 일관성 유지에서 뛰어난 성능을 보인다. 그러나 MLLM의 대규모 파라미터량과 복잡한 통합 전략은 추론 시간과 GPU 자원 소모를 증가시키며, 이는 실시간 적용 시 한계가 될 수 있다. 또한, 특정 도메인(예: 의류, 장소)에 대한 일반화 능력은 추가 실험을 통해 검증이 필요하다.

실용적 활용

HOMIE는 광고, 콘텐츠 제작, 개인화된 VR/AR 경험 등 인물-객체 상호작용이 필요한 산업에 적용 가능하다. 특히, 로고 삽입, 다각도 제품 시연, OCR 기반 텍스트 일관성 유지 등에서 높은 실용성을 기대할 수 있다.