Articulated Object Reconstruction from Rest-State Observation

Daeun Lee, Jaeah Lee, Woosung Kim, Haebeom Jung, Jaesik Park

arXiv:2607.27749 · 2026-08-12 공개 · arXiv · PDF

vision-language segmentation digital-twins geometric-consistency video-diffusion-model kinematic-structure articulated-object-reconstruction rest-state

Abstract

Building interactive digital twins requires recovering both 3D geometry and the kinematic structures that govern how objects articulate. Yet existing methods for articulated object reconstruction require explicitly observable motion from multiple articulation states. We introduce a rest-state formulation that reconstructs articulated objects from a single closed configuration, an inherently ill-posed setting where geometry, semantics, and motion priors compensate for the absence of motion cues. Our framework adopts an explicit mesh as an intermediate representation for cross-model verification and fusion, reconciling noisy outputs from vision-language and segmentation models into spatially consistent part structures. To estimate joint parameters without observed motion, we use a video diffusion model to synthesize articulation hypotheses and validate them through geometric consistency. Our approach achieves accurate part decomposition and physically plausible articulation, performing competitively with motion-observing reconstruction-based, generation-based, and modular pretrained-model baselines.

한국어 요약

한 줄 요약

Rest2Art는 단일 닫힌 상태에서 관절 구조와 3D 기하학을 복원하는 새로운 아티큘레이션 객체 재구성 프레임워크이다.

핵심 기여도

핵심 아이디어

기존의 아티큘레이션 객체 재구성 방법은 여러 관절 상태에서 명시적으로 관찰 가능한 움직임을 필요로 하지만, Rest2Art는 단일 닫힌 상태에서 객체를 복원하는 Rest-state 설정을 제안한다. 이는 기하학적, 세미안틱, 운동적 사전 정보가 움직임 없이도 객체의 구조를 추론하는 데 기여한다는 통찰에 기반한다. 핵심 아이디어는 Vision-Language Model(VLM)과 SAM3를 사용하여 부품 계층과 세그멘테이션 마스크를 반복적으로 정제하고, 이들의 불일치를 상호 보완 신호로 활용하는 것이다. 또한, 비디오 확산 모델을 통해 가상의 아티큘레이션 시퀀스를 생성하고, 추출된 2D 궤적을 기반으로 관절 모델을 피팅하며, 기하학적 일관성을 통해 관절 파라미터를 추정한다.

기술적 접근법

주요 결과

의의 및 한계

Rest2Art는 움직임 없이도 객체의 관절 구조를 복원하는 기술적 가능성을 입증하며, 디지털 트윈 및 물리 시뮬레이션에 활용 가능한 정량적 3D 복원을 제공한다. 특히, VLM과 세그멘테이션 모델의 노이즈 예측을 메시 기반으로 통합하는 접근은 기존 방법 대비 더 높은 안정성을 보인다. 그러나, 복잡한 운동 구조 (예: 다중 관절, 복합 연결)에 대한 일반화 능력은 아직 한계가 있으며, 비디오 확산 모델의 물리적 타당성 향상이 필요하다.

실용적 활용

Rest2Art는 인터랙티브 시뮬레이션, 공간 계획, 온라인 제품 이미지에서 자동 3D 모델 생성 등에 활용 가능하다. 특히, ScanNet, Replica와 같은 대규모 장면 데이터셋에서 닫힌 상태의 객체를 대상으로 한 디지털 트윈 구축에 유용하다.