How Far Are We from Removing the Visual Encoder? Scaling Laws for Encoder-Free Multimodal Pretraining

Lin Chen, Bolin Ni, Qi Yang, Lan Jiang, Kun Ding, Xiaoran Fan, Hower Yang, Ying Wang, Shiming Xiang

arXiv:2609.35457 · 2026-09-29 공개 · arXiv · PDF

vision-language language-models model-architecture scaling-laws multimodal-pretraining expert-routing visual-encoder encoder-free

Abstract

Most modern multimodal large language models (MLLMs) build on a pretrained visual encoder that provides a strong visual prior. Encoder-free MLLMs instead learn visual representations directly from raw pixels, offering a simple and unified architecture, but their scaling behavior has not been systematically characterized. To fill this gap, we compare scaling laws for encoder-free and encoder-based MLLMs and report three main findings: (1) Removing the visual encoder shifts the compute-optimal allocation for the multimodal objective toward larger models, while leaving that for text nearly unchanged. (2) The two architectures exhibit nearly overlapping loss--compute frontiers on the text objective, but diverge on the multimodal objective: encoder-free models underperform at small scales yet are predicted to catch up at around 10^{22} FLOPs, well within practical pretraining budgets. (3) Without a visual encoder, the language model learns to take over its role via vision-specific adaptation: bidirectional interactions among visual tokens become increasingly beneficial as training compute grows, visual processing shifts toward earlier layers, and expert routing for visual tokens becomes more concentrated. Overall, our results indicate that the advantage of the visual prior provided by a pretrained encoder diminishes with scale, positioning encoder-free architectures as a promising direction for multimodal pretraining.

한국어 요약

한 줄 요약

시각 인코더를 제거한 MLLM이 10²² FLOPs 규모에서 기존 인코더 기반 모델과 성능이 수렴할 수 있음을 규모 법칙을 통해 밝힘.

핵심 기여도

핵심 아이디어

기존 MLLM은 시각 인코더를 통해 사전 학습된 시각 표현을 제공받지만, encoder-free MLLM은 raw pixel에서 직접 시각 표현을 학습한다. 본 연구는 이 두 아키텍처의 규모 법칙을 비교하여, encoder-free 모델이 충분한 규모(10²² FLOPs)에 도달하면 기존 인코더 기반 모델과 성능이 수렴할 수 있음을 밝혔다. 특히, 디코더가 시각 인코더의 역할을 대체하기 위해 시각 토큰 간 양방향 어텐션, 토큰 처리 층 이동, MoE 전문가 라우팅 집중화 등의 적응 메커니즘을 보인다는 점이 핵심 통찰이다. 이는 디코더가 시각 표현 학습을 위한 본질적인 구조적 변화가 필요함을 시사한다.

기술적 접근법

주요 결과

의의 및 한계

실용적 활용