Champ: Controllable and Consistent Human Image Animation with 3D Parametric Guidance

Shenhao Zhu, Junming Chen, Zuozhuo Dai, Yinghui Xu, Xun Cao, Yao Yao, Hao Zhu, Siyu Zhu

arXiv:2403.14781 · 2026-07-27 공개 · arXiv · PDF

latent-diffusion pose-estimation depth-maps smpl motion-guidance shape-alignment human-animation multi-layer-fusion

Abstract

In this study, we introduce a methodology for human image animation by leveraging a 3D human parametric model within a latent diffusion framework to enhance shape alignment and motion guidance in curernt human generative techniques. The methodology utilizes the SMPL(Skinned Multi-Person Linear) model as the 3D human parametric model to establish a unified representation of body shape and pose. This facilitates the accurate capture of intricate human geometry and motion characteristics from source videos. Specifically, we incorporate rendered depth images, normal maps, and semantic maps obtained from SMPL sequences, alongside skeleton-based motion guidance, to enrich the conditions to the latent diffusion model with comprehensive 3D shape and detailed pose attributes. A multi-layer motion fusion module, integrating self-attention mechanisms, is employed to fuse the shape and motion latent representations in the spatial domain. By representing the 3D human parametric model as the motion guidance, we can perform parametric shape alignment of the human body between the reference image and the source video motion. Experimental evaluations conducted on benchmark datasets demonstrate the methodology's superior ability to generate high-quality human animations that accurately capture both pose and shape variations. Furthermore, our approach also exhibits superior generalization capabilities on the proposed in-the-wild dataset. Project page: https://fudan-generative-vision.github.io/champ.

한국어 요약

한 줄 요약

CHAMP은 SMPL 모델을 기반으로 3D 파라메트릭 가이던스를 결합하여 사람 이미지 애니메이션의 형태 정렬과 움직임 지도를 향상시킨다.

핵심 기여도

핵심 아이디어

기존의 사람 애니메이션 기법은 스켈레톤, 세멘틱 맵, 밀집 움직임 흐름 등 2D 기반의 움직임 지도를 사용하지만, 이는 형태와 자세의 정확한 정렬에 한계가 있었다. CHAMP은 SMPL(Skinned Multi-Person Linear)이라는 3D 파라메트릭 모델을 도입하여, 형태와 자세를 통일된 저차원 파라미터 공간에서 표현함으로써 이 문제를 해결한다. SMPL은 렌더링된 깊이 이미지, 노멀 맵, 세멘틱 맵을 생성하여 3D 형태 정보를 라티언트 디퓨전 모델에 제공하며, 스켈레톤은 얼굴과 손 움직임과 같은 세부 움직임을 보완적으로 가이드한다. 이와 함께, Self-attention 기반의 다층 움직임 퓨전 모듈을 통해 형태와 움직임의 라티언트 표현을 통합하여, 더 정확한 애니메이션 생성이 가능하다.

기술적 접근법

주요 결과

의의 및 한계

CHAMP은 기존 2D 기반 움직임 지도의 한계를 극복하고, 형태와 자세를 동시에 정확히 표현할 수 있는 3D 파라메트릭 모델 기반의 새로운 접근법을 제시한다. SMPL 모델을 활용한 형태 정렬과 Self-attention 기반의 다층 퓨전 모듈은 사람 애니메이션 생성의 정확도와 일관성을 크게 향상시켰으며, 다양한 실제 환경 데이터에서도 일반화 능력을 보였다. 그러나 SMPL 모델의 정확한 추정이 필요하며, 복잡한 배경이나 옷차림 변화에 대한 처리는 추가 연구가 필요하다.

실용적 활용

CHAMP은 가상 현실(VR), 인터랙티브 스토리텔링, 디지털 콘텐츠 제작 등 사람의 형태와 움직임을 정확히 표현해야 하는 분야에 적용 가능하다. 특히, 사용자 정의 이미지와 움직임을 기반으로 사람 애니메이션을 생성하는 애플리케이션 개발에 유용할 것으로 기대된다.