Kirin: Animal Motion Generation from In-the-Wild Video

Brian Nlong Zhao, Zhuoyang Pan, James M. Rehg, Jiajun Wu, Shangzhe Wu

arXiv:2609.01823 · 2026-09-03 공개 · arXiv · PDF

motion-generation motion-reconstruction text-conditioned animal-animation video-to-3d ai-m3d image-conditioned quadruped-animals

Abstract

Understanding animal motion is fundamental to modeling animal behavior and biomechanics, yet progress in this area lags far behind human motion research due to the scarcity of high-quality motion data. While human motion can be captured in controlled environments, it is impractical for most animal species, resulting in small, domain-limited datasets that restrict downstream applications such as animation. To address this challenge, we introduce Kirin, a framework that reconstructs motion from video, learns motion priors at scale, and generates realistic motion that can be directly applied to animated assets. Using large collections of in-the-wild animal videos, we reconstruct 3D motion sequences and pair them with captions to create AiM3D, the first large-scale dataset offering aligned video-text-motion tuples for quadruped animals. Building on this dataset, we develop a visual-guided motion generation model that conditions on both text and image to guide the generation of realistic motion across diverse animal species. Finally, by leveraging an off-the-shelf image-to-3D model, we automatically rig and animate 3D meshes using generated motion, producing ready-to-render animated animals. Together, our dataset and framework establish a new foundation for large-scale, text and image conditioned animal motion generation and animation. Project page: https://kirin-ani.github.io/.

한국어 요약

한 줄 요약

Kirin은 자연 상태의 동영상에서 동물 움직임을 생성하는 첫 번째 대규모 데이터셋과 모델을 제시한다.

핵심 기여도

핵심 아이디어

Kirin은 기존 연구가 제한적인 동물 움직임 데이터로 인해 발전이 느렸다는 문제를 해결하기 위해, 인터넷 동영상에서 직접 3D 움직임을 재구성하고 이를 텍스트와 결합한 AiM3D 데이터셋을 구축했다. 이 데이터셋을 기반으로 텍스트와 이미지 조건을 동시에 사용하는 MDM 기반의 생성 모델을 학습하여, 다양한 동물 종에 걸쳐 현실적인 움직임을 생성한다. 기존 연구는 주로 합성 또는 수작업 데이터에 의존했지만, Kirin은 자연 상태의 동영상에서 직접 학습함으로써 더 다양한 움직임 패턴을 포착할 수 있다. 또한, 이미지-3D 변환 모듈을 활용해 생성된 움직임을 3D 메시에 자동으로 적용하여, 렌더링 가능한 애니메이션을 생성한다.

기술적 접근법

주요 결과

의의 및 한계

Kirin은 동물 움직임 연구의 데이터 부족 문제를 해결하고, 자연스러운 움직임 생성과 3D 애니메이션 자동화를 가능하게 하여 생물 역학, 행동 분석, 캐릭터 애니메이션 분야에 기여한다. 또한, 텍스트-이미지 조건을 결합한 생성 모델은 다양한 응용 분야에서 유용할 수 있다. 그러나 AiM3D는 4족 동물에만 한정되어 있고, 더 다양한 동물 종을 포함하는 확장이 필요하다. 또한, 재구성된 움직임의 정확도는 환경 조건에 따라 변동할 수 있으며, 이는 모델의 일반화 능력을 제한할 수 있다.

실용적 활용

Kirin은 영화, 게임, VR/AR 등에서 자연스러운 동물 애니메이션을 생성하는 데 활용 가능하다. 또한, 생물학 및 생태학 연구에서 동물 행동 분석을 지원할 수 있으며, 교육 콘텐츠 제작에도 적용 가능하다.