World Model on Million-Length Video And Language With Blockwise RingAttention

Hao Liu, Wilson Yan, Matei Zaharia, Pieter Abbeel

arXiv:2402.08268 · 2026-07-27 공개 · arXiv · PDF

transformer long-context model-scaling open-source large-language-model video-language blockwise-attention ring-attention

Abstract

Enabling long-context understanding remains a key challenge in scaling existing sequence models -- a crucial component in developing generally intelligent models that can process and operate over long temporal horizons that potentially consist of millions of tokens. In this paper, we aim to address these challenges by providing a comprehensive exploration of the full development process for producing 1M context language models and video-language models, setting new benchmarks in language retrieval and new capabilities in long video understanding. We detail our long context data curation process, progressive context extension from 4K to 1M tokens, and present an efficient open-source implementation for scalable training on long sequences. Additionally, we open-source a family of 7B parameter models capable of processing long text documents and videos exceeding 1M tokens.

한국어 요약

한 줄 요약

1M 토큰 길이의 언어 및 비디오-언어 모델을 구축하고, Blockwise RingAttention을 기반으로 확장성을 보장한 오픈소스 구현을 제시한다.

핵심 기여도

핵심 아이디어

기존 시퀀스 모델은 수천 토큰 이하의 짧은 문맥만 처리할 수 있어, 수백만 토큰 길이의 문서나 비디오를 처리하는 데 한계가 있었다. 본 연구는 Blockwise RingAttention이라는 새로운 어텐션 메커니즘을 도입하여, 1M 토큰 길이의 시퀀스를 효율적으로 처리할 수 있는 모델을 구축한다. 이 기법은 기존의 어텐션 확장 방식과 달리 근사치 없이 정확한 계산을 유지하면서도 메모리 사용량을 줄인다. 또한, 장기 텍스트와 비디오 데이터를 위한 커리어 프로세스를 구축하고, 텍스트 기반 모델이 생성한 합성 질문-답변 데이터를 활용해 대화 능력을 향상시킨다.

기술적 접근법

주요 결과

의의 및 한계

본 연구는 수백만 토큰 길이의 시퀀스를 처리할 수 있는 모델을 구축함으로써, 장기 시퀀스 이해를 위한 기초를 마련했다. 특히, Blockwise RingAttention은 기존 어텐션 기법의 한계를 극복하며, 대규모 데이터셋에서의 확장성을 보여준다. 또한, 다중 모달 학습을 위한 마스킹 및 손실 균형 기법은 실제 적용 가능성에 기여한다.

하지만, 연구에는 몇 가지 한계가 있다. 첫째, 이미지 토크나이저는 단순한 방식을 사용하여, 시간적 중복성을 고려한 비디오 토크나이저 개선이 필요하다. 둘째, 모델 파라미터 수(7B)는 대형 언어 모델(100B 이상)에 비해 작아, 대규모 확장 시 다른 스케일링 특성이 나타날 수 있다.

실용적 활용

LWM 모델은 장문 문서 분석, 장기 비디오 요약, 대화형 챗봇 등에서 활용 가능하다. 특히, 1M 토큰 길이의 텍스트와 비디오를 처리할 수 있어, 법적 문서 검색, 교육 콘텐츠 요약, 영상 기반 QA 시스템 등에 적용할 수 있다. 오픈소스 구현을 통해 연구자 및 엔지니어가 장기 시퀀스 모델 개발에 쉽게 접근할 수 있다.