This paper presents VideoStreaming, an advanced vision-language large model (VLLM) for video understanding, that capably understands arbitrary-length video with a constant number of video tokens streamingly encoded and adaptively selected. The challenge of video understanding in the vision language area mainly lies in the significant computational burden caused by the great number of tokens extracted from long videos. Previous works rely on sparse sampling or frame compression to reduce tokens. However, such approaches either disregard temporal information in a long time span or sacrifice spatial details, resulting in flawed compression. To address these limitations, our VideoStreaming has two core designs: Memory-Propagated Streaming Encoding and Adaptive Memory Selection. The Memory-Propagated Streaming Encoding architecture segments long videos into short clips and sequentially encodes each clip with a propagated memory. In each iteration, we utilize the encoded results of the preceding clip as historical memory, which is integrated with the current clip to distill a condensed representation that encapsulates the video content up to the current timestamp. After the encoding process, the Adaptive Memory Selection strategy selects a constant number of question-related memories from all the historical memories and feeds them into the LLM to generate informative responses. The question-related selection reduces redundancy within the memories, enabling efficient and precise video understanding. Meanwhile, the disentangled video extraction and reasoning design allows the LLM to answer different questions about a video by directly selecting corresponding memories, without the need to encode the whole video for each question. Our model achieves superior performance and higher efficiency on long video benchmarks, showcasing precise temporal comprehension for detailed question answering.
한 줄 요약
VideoStreaming은 긴 동영상의 효율적인 이해를 위해 Memory-Propagated Streaming Encoding과 Adaptive Memory Selection을 결합한 VLLM 모델이다.
핵심 기여도
- Memory-Propagated Streaming Encoding을 통해 긴 동영상을 고정 길이 메모리로 압축, 시간적 동적 정보를 유지.
- Adaptive Memory Selection을 통해 질문에 관련된 메모리만 선택, 추론 효율성 향상 (예: EgoSchema에서 46.6%의 글로벌 이해력 향상).
- Phi-2 언어 모델의 얕은 레이어가 더 나은 압축 성능을 보임 (Encoder 레이어 수 줄이면 성능 향상).
- 긴 동영상 QA 데이터셋을 수작업으로 구성하여 모델 학습 및 평가에 활용.
핵심 아이디어
VideoStreaming은 기존의 토큰 감소 전략(예: 희소 샘플링, 프레임 압축)이 시간적 정보나 공간적 세부 정보를 잃는 문제를 해결하기 위해, 메모리 기반 스트리밍 인코딩과 적응형 메모리 선택을 결합한 새로운 접근법을 제안한다.
Memory-Propagated Streaming Encoding은 긴 동영상을 짧은 클립으로 나누고, 이전 클립의 인코딩 결과(히스토리 메모리)를 현재 클립과 결합하여 시간적 동적 정보를 유지하면서 고정 길이의 메모리를 생성한다. 이는 각 클립의 마지막 토큰을 메모리로 사용하여 시간 경과에 따라 정보가 누적되도록 설계되었다.
Adaptive Memory Selection은 질문에 따라 관련된 메모리만 선택하여 LLM에 입력, 불필요한 메모리 처리를 줄이고 정확한 답변을 유도한다. 질문과 각 클립의 요약 토큰 간 유사도를 계산해 관련성이 높은 메모리를 선택한다. 이는 MovieChat-1K에서 31.9%의 성능 하락을 방지하는 데 기여했다.
기술적 접근법
- **Memory-Propagated Streaming Encoding**: 긴 동영상을 짧은 클립으로 나누고, 각 클립을 순차적으로 인코딩. 이전 클립의 인코딩 결과를 히스토리 메모리로 활용.
- **Adaptive Memory Selection**: 질문과 클립 요약 토큰 간 유사도 계산을 통해 관련 메모리 선택.
- **Phi-2 언어 모델**: 스트리밍 인코딩에 사용. 얕은 레이어가 더 나은 성능을 보임.
- **데이터셋 구성**: 기존 데이터셋의 짧은 동영상을 결합하거나 Panda-70M의 캡션 정보를 활용해 긴 동영상 QA 쌍 생성.
- **훈련 전략**: 2단계 훈련 (1단계: 단일 클립 인코딩, 2단계: LLM과 결합한 긴 동영상 이해 훈련).
주요 결과
- EgoSchema 데이터셋에서 글로벌 이해력이 46.6% 향상됨.
- MovieChat-1K에서 브레이크포인트 정확도가 31.9% 개선됨.
- Phi-2 언어 모델의 레이어 수가 적을수록 성능이 향상됨.
- Next-GQA 데이터셋에서 Acc@GQA가 54.9%에서 55.7%로 약간 상승.
의의 및 한계
VideoStreaming은 긴 동영상의 토큰 수를 줄이면서 시간적 정보를 유지하는 데 성공했으며, 질문에 맞춘 메모리 선택으로 추론 효율성을 높였다. 특히, 고정 길이 메모리와 적응형 선택 기법은 기존 방법들보다 정확도와 효율성에서 우수한 성능을 보였다.
하지만, 모델은 수작업으로 구성된 긴 동영상 QA 데이터에 의존하며, 대규모 자동 생성 데이터에 대한 평가가 부족하다. 또한, 질문에 따라 메모리 선택이 이루어지기 때문에, 복잡한 다중 참조 질문에 대한 처리 능력은 한계가 있을 수 있다.
실용적 활용
VideoStreaming은 긴 동영상 QA 시스템, 영상 기반 고객 지원, 영상 분석 플랫폼 등에서 활용 가능하다. 특히, 실시간 스트리밍 환경이나 대용량 영상 처리가 필요한 산업에서 추론 효율성과 정확도를 동시에 달성할 수 있다.