Probing the 3D Awareness of Visual Foundation Models

Mohamed El Banani, Amit Raj, Kevis-Kokitsi Maninis, Abhishek Kar, Yuanzhen Li, Michael Rubinstein, Deqing Sun, Leonidas J. Guibas, Justin Johnson, Varun Jampani

arXiv:2404.08636 · 2026-07-27 공개 · arXiv · PDF

vision-models scene-representation zero-shot-inference visual-foundation-models object-localization feature-analysis

Abstract

Recent advances in large-scale pretraining have yielded visual foundation models with strong capabilities. Not only can recent models generalize to arbitrary images for their training task, their intermediate representations are useful for other visual tasks such as detection and segmentation. Given that such models can classify, delineate, and local- ize objects in 2D, we ask whether they also represent their 3D structure? In this work, we analyze the 3D awareness of visual foundation models. We posit that 3D awareness implies that representations (1) encode the 3D structure of the scene and (2) consistently represent the surface across views. We conduct a series of experiments using task-specific probes and zero-shot inference procedures on frozen fea- tures. Our experiments reveal several limitations of the current models. Our code and analysis can be found at https://github.com/mbanani/probe3d.

한국어 요약

한 줄 요약

시각 기초 모델이 3D 구조를 얼마나 잘 인식하는지 탐구한 연구로, DINOv2와 StableDiffusion이 2D 학습에도 3D 정보를 부분적으로 인코딩함을 밝힘.

핵심 기여도

핵심 아이디어

시각 기초 모델이 3D 구조를 얼마나 잘 인식하는지 평가하기 위해 **단일 뷰 표면 재구성**(single-view surface reconstruction)과 **다중 뷰 일관성**(multiview consistency)이라는 두 가지 기준을 제시함.
이 연구는 **frozen feature**를 사용해 모델의 내부 표현을 분석함으로써, 모델이 학습된 가중치를 transfer하지 않고도 3D 정보를 얼마나 잘 담고 있는지를 평가함.
이를 위해 **task-specific probes**와 **zero-shot inference**를 사용하여, 모델이 학습 과정에서 3D 정보를 암묵적으로 학습했는지 탐색함.

기술적 접근법

주요 결과

의의 및 한계

이 연구는 시각 기초 모델이 3D 구조를 얼마나 잘 인식하는지를 체계적으로 평가한 첫 시도로, **3D 인식 능력**(3D awareness)이 학습 목적과 밀접하게 관련됨을 보여줌.
또한, **zero-shot 방식**으로 모델의 내재적 표현을 분석함으로써, 모델이 3D 정보를 암묵적으로 학습했는지를 평가할 수 있는 새로운 접근법을 제시함.
그러나, **다양한 데이터셋과 학습 스케일**에서 훈련된 모델을 사용했기 때문에, **공정한 비교**는 어려웠으며, **더 복잡한 3D 인식 작업**(예: 변형 예측, 공간 관계 추론)은 탐구되지 않았음.

실용적 활용

이 연구는 **이미지 생성**(image generation) 및 **비전-언어 모델**(vision-language models)의 3D 인식 능력을 평가하는 데 활용될 수 있으며, **3D 재구성**(3D reconstruction)이나 **로봇 비전**(robot vision) 분야에서 모델 선택에 참고가 될 수 있음.
또한, **시각 표현 학습**(visual representation learning)의 평가 기준을 확장하는 데 기여할 수 있음.