Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Matt Deitke, Christopher Clark, Sangho Lee, Rohun Tripathi, Yue Yang, J. Park, Mohammadreza Salehi, Niklas Muennighoff, Kyle Lo, Luca Soldaini, Jiasen Lu, Taira Anderson, Erin Bransom, Kiana Ehsani, Huong Ngo, YenSung Chen, Ajay Patel, Mark Yatskar, Christopher Callison-Burch, Andrew Head, Rose Hendrix, F. Bastani, Eli VanderBilt, Nathan Lambert, Yvonne Chou, Arnavi Chheda, Jenna Sparks, Sam Skjonsberg, Michael Schmitz, Aaron Sarnat, Byron Bischoff, Pete Walsh, Christopher Newell, Piper Wolters, Tanmay Gupta, Kuo-Hao Zeng, Jon Borchardt, Dirk Groeneveld, Crystal Nam, Sophie Lebrecht, Caitlin Wittlif, Carissa Schoenick, Oscar Michel, Ranjay Krishna, Luca Weihs, Noah A. Smith, Hanna Hajishirzi, Ross Girshick, Ali Farhadi, Aniruddha Kembhavi

arXiv:2409.17146 · 2026-07-27 공개 · arXiv · PDF

vision-language benchmarking large-model training-pipeline open-weights pixmo molmo dataset-captions

Abstract

Today’s most advanced vision-language models (VLMs) remain proprietary. The strongest open-weight models rely heavily on synthetic data from proprietary VLMs to achieve good performance, effectively distilling these closed VLMs into open ones. As a result, the community has been missing foundational knowledge about how to build performant VLMs from scratch. We present Molmo, a new family of VLMs that are state-of-the-art in their class of openness. Our key contribution is a collection of new datasets called PixMo, including a dataset of highly detailed image captions for pre-training, a free-form image Q&A dataset for fine-tuning, and an innovative 2D pointing dataset, all collected without the use of external VLMs. The success of our approach relies on careful modeling choices, a well- tuned training pipeline, and, most critically, the quality of our newly collected datasets. Our best-in-class 72B model not only outperforms others in the class of open weight and data models, but also outperforms larger proprietary models including Claude 3.5 Sonnet, and Gemini 1.5 Pro and Flash, second only to GPT-4o based on both academic benchmarks and on a large human evaluation. Our model weights, new datasets, and source code are available at https://molmo.allenai.org/blog.

한국어 요약

한 줄 요약

Molmo는 외부 VLM 없이 자체적으로 수집한 PixMo 데이터셋을 기반으로 훈련된 최고 수준의 오픈 가중치-데이터 VLM이다.

핵심 기여도

핵심 아이디어

기존 오픈 VLM은 외부 VLM으로 생성된 합성 데이터에 의존하여 사실상 '디스틸링'된 형태였다. 본 연구는 외부 VLM 없이 직접 수집한 PixMo 데이터셋을 통해 VLM을 처음부터 훈련하는 새로운 접근법을 제시한다. 특히, 2D 포인팅 데이터는 이미지 내 특정 위치를 가리키며 답변하는 능력을 향상시키고, 이는 수동 작업 대비 빠르고 정확한 라벨링을 가능하게 한다. Molmo는 OLMo, Qwen2 등 다양한 오픈 LLM과 ViT-L/14 336px CLIP 기반의 이미지 인코더를 결합한 표준 아키텍처를 사용하지만, 훈련 파이프라인에서 Overlapping Multi-Crop, Two-Stage Training 등 새로운 기법을 도입하여 성능을 극대화한다.

기술적 접근법

주요 결과

의의 및 한계

Molmo는 외부 VLM에 의존하지 않고도 GPT-4o에 근접한 성능을 달성함으로써, 오픈 VLM 연구의 기초 지식을 제공한다. 특히, 2D 포인팅 데이터는 VLM의 시각적 추론 능력을 향상시키는 새로운 방향을 제시한다. 그러나 학술적 벤치마크에서 수학적 추론 능력은 상대적으로 낮은 것으로 나타나, 추론 중심 데이터 부족이 한계로 작용한다. 또한, 일부 벤치마크의 평가 세부 사항이 공개되지 않아 재현이 어려운 점도 문제로 지적된다.

실용적 활용

Molmo는 로봇, 웹 에이전트 등 환경 내에서 포인팅 기반 행동을 수행하는 시스템에 적용 가능하다. 또한, 문서 해석, 차트 분석, 시계 인식 등 OCR 중심의 산업 분야에서 활용도가 높으며, 오픈 데이터와 모델을 기반으로 연구자들이 VLM의 학습 과정을 자유롭게 분석할 수 있는 기반을 제공한다.