Distilling Routed 3D Privilege for Spatial Reasoning in Vision-Language Models

Hongxing Li, Yixin Li, Dingming Li, Zixuan Wang, Yuchen Yan, Wenqi Zhang, Weiming Lu, Yongliang Shen

arXiv:2610.12355 · 2026-10-11 공개 · arXiv · PDF

vision-language-models self-distillation spatial-reasoning vsi-bench mindcube mmsi-bench geometry-privileged-distillation bev-cues

Abstract

Spatial reasoning remains a persistent weakness of vision-language models (VLMs), because RGB inputs do not directly provide geometric evidence. Existing remedies either inject 3D into the model at inference, paying architecture and latency costs, or train with outcome rewards that supervise only the final answer. Spatial errors originate in perception: a misjudged depth or direction can be corrected only by the scene's true geometry, which the 3D-scanned sources of spatial training corpora already provide. We propose GPD (Geometry-Privileged Distillation), which makes geometric evidence the privilege in on-policy self-distillation (OPSD). For each question, depth, semantic, and bird's-eye-view (BEV) cues are rendered as compact text and routed to the teacher alongside the reference answer; a privileged KL, applied only to incorrect trajectories, augments GRPO, and the deployed model remains RGB-only. On the 4B backbone, GPD achieves 57.1 on VSI-Bench and 37.6 average across MindCube, SPARBench, MMSI-Bench, and ViewSpatial, outperforming both GRPO and answer-privileged OPSD across spatial reasoning benchmarks. Ablations confirm the complementarity of 3D and answer privilege, the advantage of question-conditioned routing over full-context injection, and the benefit of restricting distillation to incorrect trajectories.

한국어 요약

한 줄 요약

GPD는 3D 정보를 질문에 조건화된 방식으로 활용해 시공간 추론 능력을 향상시키는 RGB-only VLM 학습 프레임워크이다.

핵심 기여도

핵심 아이디어

기존 VLM은 RGB 입력만으로 공간 정보를 추론해야 하므로, 깊이나 방향 판단에 오류가 발생하기 쉽다. 이 오류는 추론 체인 내에서 발생하므로, 단순히 최종 정답만으로는 교정이 어렵다. GPD는 **on-policy self-distillation (OPSD)** 프레임워크에 **3D 큐를 질문 조건화된 방식으로 라우팅**하여, 오류가 발생한 추론 경로를 교정하는 방식을 제안한다.

교사 모델은 질문과 함께 **깊이, 세멘틱, BEV 정보를 텍스트로 변환한 큐**를 입력받고, 이 큐를 기반으로 학생 모델의 오답 트래젝토리에만 **privileged KL divergence**를 적용한다. 이는 학생 모델이 RGB-only로 추론하면서도, 학습 단계에서 3D 정보를 내재화하도록 유도한다. 기존 방법과 달리, GPD는 **추론 시 3D 정보를 필요로 하지 않으며, 학습 시에만 활용**한다.

기술적 접근법

주요 결과

의의 및 한계

GPD는 기존 3D-augmented 방법과 달리 **추론 시 3D 정보를 필요로 하지 않으며**, 학습 단계에서만 활용함으로써 **추론 속도와 모델 구조의 간결성**을 유지한다. 또한, **질문 조건화 라우팅**을 통해 3D 큐를 효과적으로 활용하며, **오답 트래젝토리에만 KL 적용**함으로써 학습 효율성을 높인다.

하지만, GPD는 **3D 큐 생성에 기존 3D 스캔 데이터에 의존**하며, 이는 특정 데이터셋에서만 적용 가능하다는 한계가 있다. 또한, **모델 크기(2B vs 4B)에 따른 성능 차이**는 명확히 분석되지 않았으며, **더 넓은 범위의 공간 추론 태스크**에서의 일반화 가능성도 추가 연구가 필요하다.

실용적 활용

GPD는 **로봇 비전**, **AR/VR**, **자율 주행** 등에서 시공간 정보를 정확히 해석해야 하는 **RGB-only 시스템**에 적용 가능하다. 특히, **추론 속도와 모델 가벼움이 중요한 산업 현장**에서 유용하며, **3D 스캔 데이터가 풍부한 학습 환경**에서 성능을 극대화할 수 있다.