Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding

Hanoona Rasheed, Haania Siddiqui, Ming-Hsuan Yang, Fahad Shahbaz Khan, Salman Khan

arXiv:2608.28192 · 2026-08-31 공개 · arXiv · PDF

generative-modeling autoregressive-decoding video-qa decoupled-block-attention vidstg hc-stvg referring-video-object-tracking localization-aware-policy

Abstract

Spatio-temporal video grounding (STVG) requires models to identify when a referred event occurs and localize the target entity throughout that interval. Existing multimodal large language models typically serialize dense localization trajectories autoregressively, causing decoding latency to grow with tube length and allowing localization errors to propagate across time. We introduce Parallel Tube Decoding (PTD), a generative formulation that decomposes grounding into a temporal block followed by time-conditioned spatial blocks decoded simultaneously. This removes both token-level and trajectory-level dependencies, reducing the sequential decoding depth to a fixed 1 + 1 rounds, independent of tube length. To enable parallel spatial generation, we introduce Decoupled Block Attention, which preserves access to shared video-query context while eliminating cross-box dependencies, together with localization-aware policy optimization for temporal boundaries and spatial geometry. On VidSTG, PTD reduces Tube Completion Latency by 79x and increases spatial decoding throughput by 92x over standard autoregressive decoding, while also improving grounding accuracy. With a compact 4B backbone, our model performs favorably well on VidSTG and HC-STVG, and generalizes zero-shot to temporal grounding, grounded VideoQA, and referring video object tracking. Our results show parallel tube generation is an efficient and effective alternative to autoregressive localization in videos.

한국어 요약

한 줄 요약

Parallel Tube Decoding(PTD)를 제안하여 영상 내 시공간 객체 추적의 효율성과 정확도를 동시에 향상시킨다.

핵심 기여도

핵심 아이디어

기존 시공간 영상 추적(STVG) 모델은 자동 회귀 방식으로 박스를 하나씩 생성하여 추적 경로를 만드는데, 이는 시간 경로가 길수록 디코딩 지연이 커지고 오류가 누적되는 문제를 초래한다. 이에 반해, PTD는 추적을 **시간 블록**(event interval 예측)과 **시간 조건에 따른 병렬 공간 블록**(bounding box 병렬 생성)으로 분리한다. 이로 인해 박스 간 의존성을 제거하고, 디코딩 깊이를 고정된 1 + 1 라운드로 줄인다.

핵심 기술은 **Decoupled Block Attention**으로, 공간 블록이 다른 박스에 의존하지 않으면서도 공유된 영상-쿼리 컨텍스트에 접근할 수 있도록 설계되었다. 또한, **Localization-aware Policy Optimization**는 시간 경계와 박스 기하학을 개선하기 위해 시간적 보상과 공간적 보상을 병합하여 정책 최적화를 수행한다.

기술적 접근법

주요 결과

의의 및 한계

PTD는 기존 자동 회귀 디코딩 방식의 한계를 극복하며, **시공간 추적의 효율성과 정확도를 동시에 향상**시킨다. 특히, **Decoupled Block Attention**은 공간 박스 간 의존성을 제거하면서도 시각-언어 정보를 유지하는 데 기여하며, **Localization-aware Policy Optimization**는 추적 경계와 기하학적 정확도를 개선한다.

그러나, PTD는 **복잡한 시공간 상호작용이 필요한 상황**에서는 여전히 한계가 있을 수 있다. 예를 들어, **동일 객체가 여러 번 나타나거나, 시각적 가려움이 심한 경우**, 병렬 디코딩이 오히려 정보 누락을 초래할 수 있다. 또한, **모델의 파라미터 크기(4B)**는 여전히 대규모 연산 자원이 필요하다는 점에서 실용적 적용에 제약이 있을 수 있다.

실용적 활용

PTD는 **영상 검색**, **언어 기반 추적**, **증거 기반 질문 응답**, **임베디드 시스템** 등에서 실시간 성능이 요구되는 상황에 유용하게 활용될 수 있다. 특히, **길이가 긴 영상**에서 추적 성능과 처리 속도를 동시에 향상시킬 수 있어, **모바일 및 클라우드 기반 영상 분석 플랫폼**에 적합하다.