Periscope: Extending Frozen Language Models Beyond Their Context Window

Mohamed Eltahir, Anas Obayd, Raed Rashid, Abdulrahman Alghamdi, Abdulrahman Mousa, Abdallah Ahmed, Tanveer Hussain, Naeemullah Khan

arXiv:2610.04047 · 2026-10-06 공개 · arXiv · PDF

long-context longbench ndcg frozen-models gpu-memory infinitebench chunking inference-method

Abstract

A language model reads long text in one quadratic forward pass, stops at the context window, and loses accuracy with length before reaching it. We ask whether the read can be factorized when deciding over a finite set: which document is relevant, which option is supported, which passage is the evidence. Periscope, a training-free inference method, arranges the $N$ chunks of a text on a $K{\times}K$ grid with $K{=}\lceil\sqrt{N}\rceil$ and asks a frozen model the same question about $K$ local spans of consecutive chunks and $K$ strided spans that sample the whole text, reading the log-odds of every answer at one token. Each answer takes its best local and strided score, and scoring every chunk by its two spans gives an evidence map at no further cost, whose peak is the chunk behind the answer. Every probe is about $\sqrt{sc}$ tokens for a text of $s$ tokens and chunk size $c$, so a window of $W$ tokens reaches $W^{2}/c$ tokens at $s^{1.5}$ cost. The map replaces the long read. On LongBench v2, reading only the $K$ chunks the map ranks highest, 9k tokens, matches the same model's best window read across windows from 32k to 1M tokens, and on InfiniteBench, where the median context is 150k tokens, it leads the best window read by 5 points. The same map ranks BRIGHT's long-document corpora with the best NDCG@10 of six methods. Each call caches only one probe, so a 27B model reads 4.5M-token contexts on one 80GB GPU, where a single pass would need 296GB of cache. A long read then needs a GPU that holds the model, not one that holds the text.

한국어 요약

한 줄 요약

Periscope는 텍스트를 분할 읽고 결합하여 컨텍스트 윈도우를 확장하는 추론 방법으로, 27B 모델이 4.5M 토큰을 80GB GPU에서 처리 가능하다.

핵심 기여도

핵심 아이디어

Periscope는 컨텍스트 윈도우 내에서 동작하는 언어 모델이 긴 텍스트를 처리할 수 있도록, 텍스트를 $K = \lceil \sqrt{N} \rceil$ 크기의 $K \times K$ 그리드로 분할하고, 각 청크에 대해 로컬 및 스트라이드 프로브를 독립적으로 수행한다. 이는 모델이 단일 윈도우에서 정확도가 높은 상태를 유지하면서도 전체 텍스트를 간접적으로 읽을 수 있도록 한다. 로컬 프로브는 $K$개의 연속 청크를, 스트라이드 프로브는 전체 텍스트를 샘플링하는 $K$개의 청크를 처리한다. 각 청크는 두 프로브로 두 번 읽히고, 최종적으로 증거 맵을 생성하여 정답을 결정한다. 이는 기존의 단일 패스 방식과 비교해 메모리 사용량을 줄이면서도 정확도를 유지한다.

기술적 접근법

주요 결과

의의 및 한계

Periscope는 기존의 컨텍스트 윈도우 확장 방법(예: 메모리 확장, 요약, 검색)과 달리, 훈련 없이 동결된 모델을 활용해 메모리 효율성을 극대화한다. 증거 맵은 추가 비용 없이 제공되며, 다양한 태스크(문서 순위, 질문 답변, 증거 추출)에 적용 가능하다. 그러나 이 방법은 유한한 정답 집합을 가정하며, 멀티-문서 QA나 개방형 출력(예: 요약)에서는 성능이 떨어질 수 있다. 또한, 단일 읽기보다 비용이 더 들 수 있는 상황도 존재한다.

실용적 활용

Periscope는 대규모 텍스트 분석, 법적 문서 검색, 긴 질문-답변 시스템 등에서 유용하게 활용될 수 있다. 특히, GPU 메모리가 제한된 환경에서 긴 텍스트를 처리해야 하는 연구 및 산업 현장에 적합하다.