InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD

Xiao-wen Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Bin Wang, Linke Ouyang, Songyang Zhang, Haodong Duan, Wenwei Zhang, Yining Li, Hang Yan, Yang Gao, Zhe Chen, Xinyue Zhang, Wei Li, Jingwen Li, Wenhai Wang, Kai Chen, Conghui He, Xingcheng Zhang, Jifeng Dai, Yuxin Qiao, Dahua Lin, Jiaqi Wang

arXiv:2404.06512 · 2026-07-27 공개 · arXiv · PDF

vision-language high-resolution vision-transformer large-model resolution-scaling patch-division dynamic-resolution internlm-xcomposer2

Abstract

The Large Vision-Language Model (LVLM) field has seen significant advancements, yet its progression has been hindered by challenges in comprehending fine-grained visual content due to limited resolution. Recent efforts have aimed to enhance the high-resolution understanding capabilities of LVLMs, yet they remain capped at approximately 1500 x 1500 pixels and constrained to a relatively narrow resolution range. This paper represents InternLM-XComposer2-4KHD, a groundbreaking exploration into elevating LVLM resolution capabilities up to 4K HD (3840 x 1600) and beyond. Concurrently, considering the ultra-high resolution may not be necessary in all scenarios, it supports a wide range of diverse resolutions from 336 pixels to 4K standard, significantly broadening its scope of applicability. Specifically, this research advances the patch division paradigm by introducing a novel extension: dynamic resolution with automatic patch configuration. It maintains the training image aspect ratios while automatically varying patch counts and configuring layouts based on a pre-trained Vision Transformer (ViT) (336 x 336), leading to dynamic training resolution from 336 pixels to 4K standard. Our research demonstrates that scaling training resolution up to 4K HD leads to consistent performance enhancements without hitting the ceiling of potential improvements. InternLM-XComposer2-4KHD shows superb capability that matches or even surpasses GPT-4V and Gemini Pro in 10 of the 16 benchmarks. The InternLM-XComposer2-4KHD model series with 7B parameters are publicly available at https://github.com/InternLM/InternLM-XComposer.

한국어 요약

한 줄 요약

InternLM-XComposer2-4KHD는 4K HD 해상도까지 처리 가능한 최초의 대규모 시각-언어 모델로, 10개 벤치마크에서 GPT-4V와 Gemini Pro를 능가한다.

핵심 기여도

핵심 아이디어

기존 LVLM은 고해상도 이미지 처리 시 ViT의 고정된 패치 크기와 제한된 해상도 범위로 인해 성능이 제약되었다. InternLM-XComposer2-4KHD는 336 × 336 크기의 ViT를 기반으로, 이미지의 원본 종횡비를 유지하면서 패치 수와 레이아웃을 자동으로 조정하는 **동적 해상도 및 자동 패치 구성** 기법을 도입했다. 이는 고해상도 학습 데이터 부족 문제를 완화하고, 4KHD까지 확장 가능한 학습 환경을 구축한다. 또한, 패치 레이아웃의 변동성을 줄이기 위해 **learnable newline token**을 도입하여 학습 불확실성을 감소시켰다. 이와 같은 접근은 기존 고정 해상도 기반 모델과는 차별화된 유연한 처리 능력을 제공한다.

기술적 접근법

주요 결과

의의 및 한계

InternLM-XComposer2-4KHD는 고해상도 이미지 처리를 위한 LVLM의 새로운 기준을 제시하며, 문서, 차트, 인포그래픽 등 세부 정보가 필요한 다양한 분야에 적용 가능하다. 특히, 동적 해상도 조정과 자동 패치 구성은 기존 고정 해상도 기반 모델의 한계를 극복하고, 유연한 입력 처리를 가능하게 한다. 그러나 4KHD 이상의 해상도 학습은 계산 부담이 증가하며, 상용화 시 효율적인 추론 기법이 필요하다는 한계가 있다. 또한, 일부 벤치마크에서 최고 성능을 달성하지 못한 점은 추가 연구가 필요하다.

실용적 활용

이 모델은 문서 처리, 차트 해석, 인포그래픽 분석, 웹 스크린샷 해독 등 고해상도 이미지가 필요한 산업 현장에서 활용 가능하다. 특히, OCR 기반의 자동화 시스템, 디지털 아카이빙, 의료 영상 분석 등에서 유용하게 사용될 수 있다.