LUCID: Latent-Skill Unified Control via Imagined Dynamics for Long-Horizon Humanoid Loco-Manipulation

Cheng Guo, Mingzhe Ni, Angelo Cangelosi, Arash Ajoudani

arXiv:2608.07746 · 2026-08-11 공개 · arXiv · PDF

long-horizon-planning humanoid-robotics hierarchical-control skill-composition model-based-reinforcement-learning latent-conditioned-policy imagined-rollouts multi-object-rearrangement

Abstract

Long-horizon humanoid loco-manipulation requires composing versatile whole-body skills and reliable high-level decision making. Existing methods often coordinate pretrained skills with scripted planners, finite-state machines or task-specific model-free policies, restricting their ability to handle complex task sequences. To address this limitation, we propose \textbf{LUCID}, a hierarchical model-based reinforcement learning framework that plans over reusable skills through imagined rollouts of a learned dynamics model. LUCID first trains a structured latent-conditioned low-level policy via adversarial imitation and then freezes it while jointly learning a high-level policy and macro-dynamics world model. The world model predicts the temporally extended state transitions induced by latent decisions, enabling high-level policy optimization through imagined rollouts. We evaluate our framework across various simulated multi-object rearrangement scenarios. Experimental results show that LUCID improves the full-task success and partial-completion rates compared to prior baseline methods, demonstrating its effectiveness in complex sequential loco-manipulation tasks.

한국어 요약

한 줄 요약

LUCID는 상상된 역학을 기반으로 인간형 로봇의 장기적 로코-조작을 위한 계층적 모델 기반 강화 학습 프레임워크이다.

핵심 기여도

핵심 아이디어

LUCID는 기존의 스크립트 기반 또는 태스크별 모델 없는 정책에 의존하는 접근법의 한계를 극복하기 위해, 상상된 역학을 기반으로 재사용 가능한 스킬을 계획하는 계층적 모델 기반 강화 학습 프레임워크를 제안한다. 기존 연구는 고정된 상태 머신이나 스크립트 기반 핸드오프를 사용하여 장기적 작업 시퀀스를 처리하는 데 어려움이 있었다. LUCID는 저수준 정책을 적대적 모방을 통해 학습하고 고정한 후, 고수준 정책과 매크로 역학 모델을 동시에 학습함으로써, 잠재 결정이 작업 진행에 미치는 영향을 예측할 수 있도록 한다. 이는 기존 방법들이 재사용 가능한 행동을 선택하긴 하지만, 하나의 스킬이 후속 상호작용의 조건을 어떻게 변화시키는지를 예측하지 못하는 문제를 해결한다.

기술적 접근법

주요 결과

의의 및 한계

LUCID는 장기적 인간형 로코-조작 작업에서 재사용 가능한 스킬을 효과적으로 조율할 수 있는 계층적 모델 기반 강화 학습 프레임워크를 제시한다. 구조화된 잠재 인터페이스와 매크로 역학 모델을 통해 고수준 정책이 장기적 작업을 예측하고 최적화할 수 있도록 하여, 기존 스크립트 기반 접근법의 한계를 극복한다. 그러나 현재 모델은 특권 상태 입력(privileged state inputs)에 의존하며, 작업, 객체, 구조 다양성은 제한적이다. 향후 연구에서는 시각 제어, 더 넓은 작업 범위, 실제 세계 이전(real-world transfer)을 탐구할 필요가 있다.

실용적 활용

LUCID는 인간형 로봇이 장기적 작업 시퀀스를 처리해야 하는 산업 현장, 예를 들어 물류, 제조, 서비스 분야에서 활용될 수 있다. 특히, 다물체 재배열과 같은 복잡한 작업에서 재사용 가능한 스킬을 효과적으로 조율할 수 있어, 자동화 수준을 높이는 데 기여할 수 있다.