OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding

Tao Zhang, Xiangtai Li, Hao Fei, Haobo Yuan, Shengqiong Wu, Shunping Ji, Chen Change Loy, Shuicheng Yan

arXiv:2406.19389 · 2026-07-27 공개 · arXiv · PDF

llm vision-language reasoning segmentation multimodal pixel-level image-level perception-prior

Abstract

Current universal segmentation methods demonstrate strong capabilities in pixel-level image and video understanding. However, they lack reasoning abilities and cannot be controlled via text instructions. In contrast, large vision-language multimodal models exhibit powerful vision-based conversation and reasoning capabilities but lack pixel-level understanding and have difficulty accepting visual prompts for flexible user interaction. This paper proposes OMG-LLaVA, a new and elegant framework combining powerful pixel-level vision understanding with reasoning abilities. It can accept various visual and text prompts for flexible user interaction. Specifically, we use a universal segmentation method as the visual encoder, integrating image information, perception priors, and visual prompts into visual tokens provided to the LLM. The LLM is responsible for understanding the user's text instructions and providing text responses and pixel-level segmentation results based on the visual information. We propose perception prior embedding to better integrate perception priors with image features. OMG-LLaVA achieves image-level, object-level, and pixel-level reasoning and understanding in a single model, matching or surpassing the performance of specialized methods on multiple benchmarks. Rather than using LLM to connect each specialist, our work aims at end-to-end training on one encoder, one decoder, and one LLM. The code and model have been released for further research.

한국어 요약

한 줄 요약

OMG-LLaVA는 단일 모델로 이미지, 객체, 픽셀 수준의 추론과 이해를 통합한 MLLM으로, 1개의 LLM, 1개의 인코더, 1개의 디코더로 다양한 비주얼-텍스트 작업을 수행한다.

핵심 기여도

핵심 아이디어

OMG-LLaVA는 기존 MLLM이 이미지 수준 분석만 제공하거나, 픽셀 수준 분석을 위해 별도 모델을 연결해야 하는 문제를 해결하기 위해 설계되었다. 이 모델은 **OMG-Seg**라는 유니버설 세그멘테이션 모델을 시각 인코더로 사용하여 이미지 정보, 퍼셉션 프라이어, 시각 프롬프트를 **비주얼 토큰**으로 변환하고, 이를 LLM에 전달한다. LLM은 텍스트 지시를 이해하고, 텍스트 응답과 픽셀 수준 세그멘테이션 결과를 생성한다. 이는 **단일 인코더-디코더-LLM 구조**로, 기존 방식보다 효율적이다. 핵심 기술은 **Perception Prior Embedding**으로, 객체 쿼리를 객체 중심 비주얼 토큰에 통합하여 텍스트-세그멘테이션 연관성을 강화한다.

기술적 접근법

주요 결과

의의 및 한계

OMG-LLaVA는 단일 모델로 이미지, 객체, 픽셀 수준의 작업을 통합하여 MLLM의 **유연성과 확장성**을 높인다. 특히, **Perception Prior Embedding**은 텍스트-세그멘테이션 연관성을 향상시켜, 기존 MLLM에서 부족했던 **정밀한 위치 정보**를 제공한다. 또한, **LoRA 기반 미세 조정**으로 훈련 비용을 줄이며, **8개 이상의 다중 작업**을 수행할 수 있다. 그러나, **복잡한 시나리오**에서는 여전히 **정확도 저하**가 발생할 수 있으며, **다양한 시각 프롬프트 입력**에 대한 평가가 부족하다는 한계가 있다.

실용적 활용

OMG-LLaVA는 **로봇 비전**, **의료 이미지 분석**, **자율 주행**, **콘텐츠 생성** 등 다양한 산업 분야에서 활용 가능하다. 특히, **텍스트 기반 세그멘테이션**, **지정 표현 세그멘테이션**, **지정된 대화 생성** 등 **사용자와의 유연한 상호작용**이 필요한 시스템에 적합하다.