MotionLCM: Real-time Controllable Motion Generation via Latent Consistency Model

Wen-Dao Dai, Ling-Hao Chen, Jingbo Wang, Jinpeng Liu, Bo Dai, Yansong Tang

arXiv:2404.19759 · 2026-07-27 공개 · arXiv · PDF

real-time-inference motion-generation controlnet motion-control text-conditioned latent-consistency-model motion-latent-diffusion controllable-motion

Abstract

This work introduces MotionLCM, extending controllable motion generation to a real-time level. Existing methods for spatial-temporal control in text-conditioned motion generation suffer from significant runtime inefficiency. To address this issue, we first propose the motion latent consistency model (MotionLCM) for motion generation, building on the motion latent diffusion model. By adopting one-step (or few-step) inference, we further improve the runtime efficiency of the motion latent diffusion model for motion generation. To ensure effective controllability, we incorporate a motion ControlNet within the latent space of MotionLCM and enable explicit control signals (i.e., initial motions) in the vanilla motion space to further provide supervision for the training process. By employing these techniques, our approach can generate human motions with text and control signals in real-time. Experimental results demonstrate the remarkable generation and controlling capabilities of MotionLCM while maintaining real-time runtime efficiency.

한국어 요약

한 줄 요약

MotionLCM은 텍스트와 제어 신호를 기반으로 실시간 인체 동작을 생성하는 새로운 모델로, 1단계 추론과 ControlNet을 통해 빠른 속도와 정밀 제어를 달성한다.

핵심 기여도

핵심 아이디어

기존의 텍스트-조건 기반 동작 생성 모델은 높은 생성 품질을 위해 수십 단계의 추론 과정이 필요하여 실시간 적용이 어려웠다. MotionLCM은 이 문제를 해결하기 위해 **Motion Latent Diffusion Model**(MLD)을 기반으로 **Latent Consistency Model**(LCM)을 도입하여, **1단계**(또는 **few-step**) 추론을 통해 실시간 성능을 달성한다.

또한, 잠재 공간 내에서의 명확한 제어를 위해 **motion ControlNet**을 도입하여, **초기 동작**(initial motion)을 제어 신호로 활용한다. 이는 기존 모델이 단순한 동작 공간에서 제어를 수행하는 것과 달리, 잠재 공간에서의 제어가 더 효과적임을 반영한 설계이다.

MotionLCM은 **Consistency Distillation**을 통해 MLD에서 학습된 지식을 효율적으로 전달받으며, **Huber Loss**를 사용하여 학습 안정성을 높인다. 이는 **EMA**(Exponential Moving Average)를 통해 **target network**를 업데이트하며, **k-step skipping**을 통해 수렴 시간을 줄이는 방식으로 구현된다.

기술적 접근법

주요 결과

의의 및 한계

MotionLCM은 **텍스트-조건 기반 동작 생성** 분야에서 **실시간 추론**과 **정밀 제어**를 동시에 달성한 첫 모델로, **Motion ControlNet**과 **Consistency Distillation** 기법을 결합한 점에서 학술적 의의가 크다. 특히, **1-step inference**를 통해 기존 모델의 추론 시간 문제를 극복하고, **initial motion**을 활용한 제어 방식은 실용적 활용 가능성을 높인다.

그러나, MLD의 VAE가 **명시적인 시간 모델링**(explicit temporal modeling)을 제공하지 않아, MotionLCM도 **시간적 해석력**(temporal explanation)이 부족한 한계가 있다. 이는 향후 연구에서 개선할 필요가 있는 점이다.

실용적 활용

MotionLCM은 **가상 캐릭터 제어**, **VR/AR**, **로봇 동작 시뮬레이션**, **게임 캐릭터 생성** 등 실시간 제어가 필요한 산업 분야에 적용 가능하다. 특히, **텍스트와 초기 동작 신호를 동시에 조건으로 사용**할 수 있어, **사용자 맞춤형 동작 생성**이나 **실시간 인터랙티브 애플리케이션** 개발에 유용하다.