diffusion-models depth-estimation zero-shot-generalization dense-prediction normal-estimation vision-foundation-model single-step-diffusion detail-preserver
Abstract
Leveraging the visual priors of pre-trained text-to-image diffusion models offers a promising solution to enhance zero-shot generalization in dense prediction tasks. However, existing methods often uncritically use the original diffusion formulation, which may not be optimal due to the fundamental differences between dense prediction and image generation. In this paper, we provide a systemic analysis of the diffusion formulation for the dense prediction, focusing on both quality and efficiency. And we find that the original parameterization type for image generation, which learns to predict noise, is harmful for dense prediction; the multi-step noising/denoising diffusion process is also unnecessary and challenging to optimize. Based on these insights, we introduce Lotus, a diffusion-based visual foundation model with a simple yet effective adaptation protocol for dense prediction. Specifically, Lotus is trained to directly predict annotations instead of noise, thereby avoiding harmful variance. We also reformulate the diffusion process into a single-step procedure, simplifying optimization and significantly boosting inference speed. Additionally, we introduce a novel tuning strategy called detail preserver, which achieves more accurate and fine-grained predictions. Without scaling up the training data or model capacity, Lotus achieves SoTA performance in zero-shot depth and normal estimation across various datasets. It also enhances efficiency, being significantly faster than most existing diffusion-based methods. Lotus' superior quality and efficiency also enable a wide range of practical applications, such as joint estimation, single/multi-view 3D reconstruction, etc. Project page: https://lotus3d.github.io/.
한국어 요약
한 줄 요약
Lotus는 기존 디퓨전 모델의 단점을 개선한 고정밀 밀집 예측 기반 시각 기초 모델로, 단계 수를 줄이고 노이즈 예측 대신 직접 어노테이션 예측을 통해 성능과 효율성을 동시에 향상시킨다.
핵심 기여도
- 기존 디퓨전 모델이 밀집 예측에 적합하지 않은 두 가지 주요 요인을 분석: (1) 노이즈 예측 파라미터화, (2) 다단계 디퓨전 프로세스.
- Lotus는 어노테이션을 직접 예측하는 단계 수 1개의 디퓨전 프로세스를 도입하여, 59K 훈련 샘플로도 SoTA 성능 달성.
- 새로운 세부 정보 보존 전략인 *detail preserver*를 통해 세부 영역 예측 정확도 향상.
- 기존 디퓨전 기반 방법 대비 추론 속도가 훨씬 빠름.
핵심 아이디어
기존 디퓨전 기반 밀집 예측 방법은 이미지 생성을 위한 디퓨전 프로세스를 그대로 사용하여, 밀집 예측에 적합하지 않은 문제를 일으킨다. 예를 들어, 노이즈 예측 파라미터화는 초기 디노이징 단계에서 예측 오차가 누적되어 최종 결과에 큰 영향을 미친다. 또한, 다단계 디퓨전은 계산 비용이 높고 최적화가 어려운 문제가 있다. Lotus는 이러한 문제를 해결하기 위해 두 가지 핵심 아이디어를 제시한다: (1) 어노테이션을 직접 예측하는 단일 단계 디퓨전 프로세스, (2) *detail preserver*를 통한 세부 정보 보존. 이를 통해 기존 디퓨전 모델의 시각 사전 지식을 최대한 활용하면서도, 예측 정확도와 추론 효율성을 동시에 개선한다.
기술적 접근법
- **디퓨전 파라미터화 변경**: 기존 노이즈 예측 대신, Lotus는 어노테이션을 직접 예측하는 *x₀-prediction* 방식을 사용.
- **단일 단계 디퓨전**: 다단계 디퓨전 대신, 단일 단계로 디퓨전 프로세스를 수행하여 훈련 및 추론 효율성 향상.
- **Detail Preserver**: 새로운 세부 정보 보존 전략으로, 어노테이션 생성 시 입력 이미지의 세부 구조를 유지.
- **데이터셋**: Hypersim, Virtual KITTI 등 사용.
- **훈련 샘플 수**: 59K개로 SoTA 성능 달성.
주요 결과
- **Depth Estimation**: NYUv2, DTU, ETH3D 등 다양한 데이터셋에서 기존 디퓨전 기반 방법 대비 높은 정확도.
- **Normal Estimation**: Lotus는 Marigold 대비 1.6× 이상 빠른 추론 속도를 보이며, 정확도 또한 개선됨.
- **Zero-shot 성능**: 훈련 데이터 없이도, 기존 디퓨전 모델의 사전 지식을 활용해 SoTA 수준의 밀집 예측 성능 달성.
의의 및 한계
Lotus는 기존 디퓨전 모델의 단점을 분석하고, 밀집 예측에 최적화된 디퓨전 프로세스를 설계함으로써, 훈련 데이터 확장 없이도 SoTA 성능을 달성한 점에서 학술적 의의가 크다. 특히, 단일 단계 디퓨전과 어노테이션 직접 예측은 디퓨전 기반 밀집 예측 연구의 새로운 방향성을 제시한다. 그러나, Lotus는 디퓨전 모델 자체의 한계(예: 훈련 데이터의 질에 의존)를 완전히 극복하지 못하며, 매우 복잡한 구조의 객체에 대해서는 예측 정확도가 낮아질 수 있다.
실용적 활용
Lotus는 3D 재구성, 다뷰 3D 예측, 자율주행 시스템의 실시간 깊이 추정 등 다양한 산업 분야에서 활용 가능하다. 특히, 추론 속도가 빠르고 정확도가 높은 점에서 실시간 시스템에 적합하며, 기존 디퓨전 모델의 사전 학습된 시각 사전 지식을 활용할 수 있어, 데이터 확보가 어려운 상황에서도 유용하게 사용될 수 있다.