ShallowStream: Index Shallow then Answer Deep for Streaming Video Understanding

Jitai Hao, Ke Yang, Qiang Huang, Jun Yu

arXiv:2609.02780 · 2026-09-07 공개 · arXiv · PDF

mllm kv-cache video-understanding streaming-video latency-reduction attention-scores context-retrieval frame-encoding

Abstract

Streaming video understanding is a critical capability for real-world applications, including embodied intelligence, autonomous driving, industrial monitoring, surveillance and early warning, and wearable assistants. However, processing continuous video streams with multimodal large language models (MLLMs) is computationally expensive. Existing efforts have explored reducing streaming overhead through visual token pruning, token merging, quantization, on-demand frame retrieval, and context offloading. However, most existing methods overlook the dimension of model depth. Repeatedly executing full-depth MLLM prefill over incoming frames is prohibitively expensive, incurring substantial computational overhead and causing the KV cache to grow at a rate directly proportional to the prefill depth. To address these challenges, we propose ShallowStream, a novel framework that leverages the shallow layers of an MLLM to simultaneously perform frame encoding and retrieval index building. During stream processing, ShallowStream maintains an always-on lightweight index using the KV cache of shallow layers. During query-time answering, we leverage the attention scores generated by the shallow layers to score context frames and employ a diversity-aware selection strategy to retrieve precise and comprehensive evidence. ShallowStream achieves performance on par with the strongest existing streaming methods, while reducing per-frame prefill latency and 10-second end-to-end latency by up to 52.1x and 11.9x, respectively. Our code is available at https://github.com/CURRENTF/ShallowStream.

한국어 요약

한 줄 요약

ShallowStream은 MLLM의 얕은 계층을 활용해 실시간 동영상 처리 비용을 52.1배 감소시키는 새로운 스트리밍 비디오 이해 프레임워크이다.

핵심 기여도

핵심 아이디어

ShallowStream은 MLLM의 얕은 계층만을 활용하여 실시간 스트리밍 비디오 처리의 계산 비용을 극복하는 새로운 접근법이다. 기존 방법은 모든 Transformer 계층을 실행하여 KV 캐시를 생성하는 데 많은 자원을 소모했으나, ShallowStream은 얕은 계층만을 사용해 인코딩과 인덱스 생성을 동시에 수행한다. 이는 KV 캐시의 성장 속도를 줄이고, 쿼리 시점에 필요한 계산만 수행함으로써 효율성을 극대화한다.

핵심 통찰은 MLLM의 얕은 계층(예: Qwen3-VL-8B의 4층, LLaVA-OneVision-7B의 3층)만으로도 쿼리와 관련된 과거 증거를 효과적으로 검색할 수 있다는 점이다. 이는 전체 계층을 실행할 필요 없이도 정확한 답변을 도출할 수 있음을 의미하며, 실시간 처리에 적합한 구조를 제공한다.

기술적 접근법

주요 결과

의의 및 한계

ShallowStream은 실시간 스트리밍 비디오 처리에서 계산 비용과 지연을 극복하는 새로운 접근법을 제시한다. 특히, MLLM의 얕은 계층만을 활용해 전체 계층 실행을 피함으로써 GPU 메모리 사용량과 처리 시간을 대폭 줄였다. 이는 자율 주행, 감시, 웨어러블 장치 등 실시간 처리가 필수적인 분야에 큰 실용적 가치를 제공한다.

그러나, ShallowStream은 얕은 계층만을 사용하기 때문에 일부 복잡한 쿼리에 대해서는 전체 계층의 깊이가 필요한 경우가 있을 수 있다. 또한, P(계층 경계) 설정에 따라 성능이 달라질 수 있으므로, 다양한 모델과 데이터셋에서의 최적 P 설정이 필요하다.

실용적 활용

ShallowStream은 자율 주행, 산업 감시, 웨어러블 보조 장치 등 실시간 동영상 처리가 필요한 분야에 적용 가능하다. 특히, GPU 메모리와 계산 자원이 제한된 환경에서 실시간 처리 성능을 극대화할 수 있다.