Streaming Long Video Understanding with Large Language Models

Rui Qian, Xiao-wen Dong, Pan Zhang, Yuhang Zang, Shuangrui Ding, Dahua Lin, Jiaqi Wang

arXiv:2405.16009 · 2026-07-27 공개 · arXiv · PDF

vision-language large-language-models video-understanding long-video question-answering video-streaming token-encoding memory-propagated-encoding

Abstract

This paper presents VideoStreaming, an advanced vision-language large model (VLLM) for video understanding, that capably understands arbitrary-length video with a constant number of video tokens streamingly encoded and adaptively selected. The challenge of video understanding in the vision language area mainly lies in the significant computational burden caused by the great number of tokens extracted from long videos. Previous works rely on sparse sampling or frame compression to reduce tokens. However, such approaches either disregard temporal information in a long time span or sacrifice spatial details, resulting in flawed compression. To address these limitations, our VideoStreaming has two core designs: Memory-Propagated Streaming Encoding and Adaptive Memory Selection. The Memory-Propagated Streaming Encoding architecture segments long videos into short clips and sequentially encodes each clip with a propagated memory. In each iteration, we utilize the encoded results of the preceding clip as historical memory, which is integrated with the current clip to distill a condensed representation that encapsulates the video content up to the current timestamp. After the encoding process, the Adaptive Memory Selection strategy selects a constant number of question-related memories from all the historical memories and feeds them into the LLM to generate informative responses. The question-related selection reduces redundancy within the memories, enabling efficient and precise video understanding. Meanwhile, the disentangled video extraction and reasoning design allows the LLM to answer different questions about a video by directly selecting corresponding memories, without the need to encode the whole video for each question. Our model achieves superior performance and higher efficiency on long video benchmarks, showcasing precise temporal comprehension for detailed question answering.

한국어 요약

한 줄 요약

VideoStreaming은 긴 동영상의 효율적인 이해를 위해 Memory-Propagated Streaming Encoding과 Adaptive Memory Selection을 결합한 VLLM 모델이다.

핵심 기여도

핵심 아이디어

VideoStreaming은 기존의 토큰 감소 전략(예: 희소 샘플링, 프레임 압축)이 시간적 정보나 공간적 세부 정보를 잃는 문제를 해결하기 위해, 메모리 기반 스트리밍 인코딩과 적응형 메모리 선택을 결합한 새로운 접근법을 제안한다.

Memory-Propagated Streaming Encoding은 긴 동영상을 짧은 클립으로 나누고, 이전 클립의 인코딩 결과(히스토리 메모리)를 현재 클립과 결합하여 시간적 동적 정보를 유지하면서 고정 길이의 메모리를 생성한다. 이는 각 클립의 마지막 토큰을 메모리로 사용하여 시간 경과에 따라 정보가 누적되도록 설계되었다.

Adaptive Memory Selection은 질문에 따라 관련된 메모리만 선택하여 LLM에 입력, 불필요한 메모리 처리를 줄이고 정확한 답변을 유도한다. 질문과 각 클립의 요약 토큰 간 유사도를 계산해 관련성이 높은 메모리를 선택한다. 이는 MovieChat-1K에서 31.9%의 성능 하락을 방지하는 데 기여했다.

기술적 접근법

주요 결과

의의 및 한계

VideoStreaming은 긴 동영상의 토큰 수를 줄이면서 시간적 정보를 유지하는 데 성공했으며, 질문에 맞춘 메모리 선택으로 추론 효율성을 높였다. 특히, 고정 길이 메모리와 적응형 선택 기법은 기존 방법들보다 정확도와 효율성에서 우수한 성능을 보였다.

하지만, 모델은 수작업으로 구성된 긴 동영상 QA 데이터에 의존하며, 대규모 자동 생성 데이터에 대한 평가가 부족하다. 또한, 질문에 따라 메모리 선택이 이루어지기 때문에, 복잡한 다중 참조 질문에 대한 처리 능력은 한계가 있을 수 있다.

실용적 활용

VideoStreaming은 긴 동영상 QA 시스템, 영상 기반 고객 지원, 영상 분석 플랫폼 등에서 활용 가능하다. 특히, 실시간 스트리밍 환경이나 대용량 영상 처리가 필요한 산업에서 추론 효율성과 정확도를 동시에 달성할 수 있다.