Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Matt Deitke, Christopher Clark, Sangho Lee, Rohun Tripathi, Yue Yang, J. Park, Mohammadreza Salehi, Niklas Muennighoff, Kyle Lo, Luca Soldaini, Jiasen Lu, Taira Anderson, Erin Bransom, Kiana Ehsani, Huong Ngo, YenSung Chen, Ajay Patel, Mark Yatskar, Christopher Callison-Burch, Andrew Head, Rose Hendrix, F. Bastani, Eli VanderBilt, Nathan Lambert, Yvonne Chou, Arnavi Chheda, Jenna Sparks, Sam Skjonsberg, Michael Schmitz, Aaron Sarnat, Byron Bischoff, Pete Walsh, Christopher Newell, Piper Wolters, Tanmay Gupta, Kuo-Hao Zeng, Jon Borchardt, Dirk Groeneveld, Crystal Nam, Sophie Lebrecht, Caitlin Wittlif, Carissa Schoenick, Oscar Michel, Ranjay Krishna, Luca Weihs, Noah A. Smith, Hanna Hajishirzi, Ross Girshick, Ali Farhadi, Aniruddha Kembhavi
arXiv:2409.17146 · 2026-07-27 공개 · arXiv · PDF
vision-language benchmarking large-model training-pipeline open-weights pixmo molmo dataset-captions
Abstract
Today’s most advanced vision-language models (VLMs) remain proprietary. The strongest open-weight models rely heavily on synthetic data from proprietary VLMs to achieve good performance, effectively distilling these closed VLMs into open ones. As a result, the community has been missing foundational knowledge about how to build performant VLMs from scratch. We present Molmo, a new family of VLMs that are state-of-the-art in their class of openness. Our key contribution is a collection of new datasets called PixMo, including a dataset of highly detailed image captions for pre-training, a free-form image Q&A dataset for fine-tuning, and an innovative 2D pointing dataset, all collected without the use of external VLMs. The success of our approach relies on careful modeling choices, a well- tuned training pipeline, and, most critically, the quality of our newly collected datasets. Our best-in-class 72B model not only outperforms others in the class of open weight and data models, but also outperforms larger proprietary models including Claude 3.5 Sonnet, and Gemini 1.5 Pro and Flash, second only to GPT-4o based on both academic benchmarks and on a large human evaluation. Our model weights, new datasets, and source code are available at https://molmo.allenai.org/blog.
한국어 요약
한 줄 요약
Molmo는 외부 VLM 없이 자체적으로 수집한 PixMo 데이터셋을 기반으로 훈련된 최고 수준의 오픈 가중치-데이터 VLM이다.
핵심 기여도
- PixMo 데이터셋: 외부 VLM 없이 수집된 712k 장의 이미지와 200자 이상 상세 캡션, 2.3M 개의 2D 포인팅 데이터 포함.
- Molmo-72B 모델: GPT-4o 다음으로 높은 Elo 점수를 기록하며, Gemini 1.5 Pro, Claude 3.5 Sonnet를 초과.
- 2D 포인팅 데이터: CountBenchQA, PixMo-Count에서 기존 모델 대비 10% 이상 성능 향상.
- MolmoE-1B: GPT-4V와 유사한 성능을 오픈 모델로 구현.
핵심 아이디어
기존 오픈 VLM은 외부 VLM으로 생성된 합성 데이터에 의존하여 사실상 '디스틸링'된 형태였다. 본 연구는 외부 VLM 없이 직접 수집한 PixMo 데이터셋을 통해 VLM을 처음부터 훈련하는 새로운 접근법을 제시한다. 특히, 2D 포인팅 데이터는 이미지 내 특정 위치를 가리키며 답변하는 능력을 향상시키고, 이는 수동 작업 대비 빠르고 정확한 라벨링을 가능하게 한다. Molmo는 OLMo, Qwen2 등 다양한 오픈 LLM과 ViT-L/14 336px CLIP 기반의 이미지 인코더를 결합한 표준 아키텍처를 사용하지만, 훈련 파이프라인에서 Overlapping Multi-Crop, Two-Stage Training 등 새로운 기법을 도입하여 성능을 극대화한다.
기술적 접근법
- **PixMo 데이터셋**: 712k 장의 이미지에 200자 이상의 캡션, 162k 개의 이미지-질의-응답 쌍, 2.3M 개의 2D 포인팅 데이터 포함.
- **모델 아키텍처**: ViT-L/14 336px CLIP 기반 이미지 인코더와 OLMo, Qwen2 등 오픈 LLM 결합.
- **훈련 파이프라인**: Overlapping Multi-Crop, Two-Stage Training, Multi-Annotation Training, 효율적인 옵티마이저 설정.
- **모델 규모**: Molmo-72B (Qwen2 72B 기반), Molmo-7B (OLMo-7B, Qwen2 7B 기반), MolmoE-1B (OLMoE-1B-7B 기반).
주요 결과
- **Molmo-72B**: 11개 학술 벤치마크 평균 정확도 73.2% (GPT-4o 대비 -1.6%), Elo 점수 1,598 (GPT-4o 대비 -12).
- **PixMo-Count**: 기존 CountBenchQA 대비 10% 이상 개선.
- **Human Evaluation**: 325k 개의 평가 데이터 기반 Elo 랭킹 2위 (GPT-4o 다음).
- **AndroidControl**: Molmo-72B, 88.7% low-level, 69.0% high-level 정확도 (기존 83.2%, 70.8% 대비 높음).
의의 및 한계
Molmo는 외부 VLM에 의존하지 않고도 GPT-4o에 근접한 성능을 달성함으로써, 오픈 VLM 연구의 기초 지식을 제공한다. 특히, 2D 포인팅 데이터는 VLM의 시각적 추론 능력을 향상시키는 새로운 방향을 제시한다. 그러나 학술적 벤치마크에서 수학적 추론 능력은 상대적으로 낮은 것으로 나타나, 추론 중심 데이터 부족이 한계로 작용한다. 또한, 일부 벤치마크의 평가 세부 사항이 공개되지 않아 재현이 어려운 점도 문제로 지적된다.
실용적 활용
Molmo는 로봇, 웹 에이전트 등 환경 내에서 포인팅 기반 행동을 수행하는 시스템에 적용 가능하다. 또한, 문서 해석, 차트 분석, 시계 인식 등 OCR 중심의 산업 분야에서 활용도가 높으며, 오픈 데이터와 모델을 기반으로 연구자들이 VLM의 학습 과정을 자유롭게 분석할 수 있는 기반을 제공한다.