VISA: Reasoning Video Object Segmentation via Large Language Models

Cilin Yan, Haochen Wang, Shilin Yan, Xiaolong Jiang, Yao Hu, Guoliang Kang, Weidi Xie, E. Gavves

arXiv:2407.11325 · 2026-07-27 공개 · arXiv · PDF

large-language-models instruction-tuning embodied-ai multi-modal-llm reasoning-segmentation video-object-segmentation mask-decoder reasonv-os

Abstract

Existing Video Object Segmentation (VOS) relies on explicit user instructions, such as categories, masks, or short phrases, restricting their ability to perform complex video segmentation requiring reasoning with world knowledge. In this paper, we introduce a new task, Reasoning Video Object Segmentation (ReasonVOS). This task aims to generate a sequence of segmentation masks in response to implicit text queries that require complex reasoning abilities based on world knowledge and video contexts, which is crucial for structured environment understanding and object-centric interactions, pivotal in the development of embodied AI. To tackle ReasonVOS, we introduce VISA (Video-based large language Instructed Segmentation Assistant), to leverage the world knowledge reasoning capabilities of multi-modal LLMs while possessing the ability to segment and track objects in videos with a mask decoder. Moreover, we establish a comprehensive benchmark consisting of 35,074 instruction-mask sequence pairs from 1,042 diverse videos, which incorporates complex world knowledge reasoning into segmentation tasks for instruction-tuning and evaluation purposes of ReasonVOS models. Experiments conducted on 8 datasets demonstrate the effectiveness of VISA in tackling complex reasoning segmentation and vanilla referring segmentation in both video and image domains. The code and dataset are available at https://github.com/cilinyan/VISA.

한국어 요약

한 줄 요약

VISA는 대규모 언어 모델을 활용해 암묵적 텍스트 쿼리에 기반한 비디오 객체 분할을 수행하는 새로운 ReasonVOS 태스크를 제안한다.

핵심 기여도

핵심 아이디어

기존 VOS는 명시적 쿼리(카테고리, 마스크, 짧은 문장)에 의존하여 복잡한 추론이 필요한 분할에 제한이 있었다.
VISA는 암묵적 텍스트 쿼리(예: “내가 가장 좋아하는 컵”)를 처리할 수 있도록 설계되었으며, 이는 세계 지식과 비디오 맥락을 기반으로 추론을 요구한다.
TFS 모듈을 통해 텍스트 쿼리에 관련된 프레임을 선택함으로써 처리해야 할 시각 토큰 수를 줄이고, LLM을 통해 추론을 수행한 후 SAM decoder로 분할 마스크를 생성한다.
이러한 접근은 장기 비디오 이해와 정확한 객체 추적을 결합하여, Embodied AI 개발에 필수적인 능력을 구현한다.

기술적 접근법

주요 결과

의의 및 한계

VISA는 세계 지식 기반 추론과 장기 비디오 이해를 결합한 첫 번째 모델로, Embodied AI 개발에 기여한다.
ReVOS 데이터셋은 기존 Referring VOS 데이터셋과 달리 암묵적 텍스트를 포함하여 추론 능력을 평가할 수 있는 새로운 기준을 제시한다.
한계로는 텍스트 쿼리가 복잡할 경우 추론 오류가 발생할 수 있으며, 장기 비디오 처리 시 계산 비용이 높다는 점이 있다.

실용적 활용

VISA는 로봇이 동적 환경에서 사용자 지시에 따라 객체를 인식하고 상호작용하는 Embodied AI 시스템에 활용 가능하다.
또한, 비디오 분석, 자동 콘텐츠 태깅, 스마트 홈 시스템 등에서 암묵적 지시에 기반한 객체 추적 기능을 구현할 수 있다.