R1-Onevision: Advancing Generalized Multimodal Reasoning Through Cross-Modal Formalization

Yi Yang, Xiaoxuan He, H. Pan, Xiyan Jiang, Yan Deng, Xingtao Yang, Haoyu Lu, Dacheng Yin, Fengyun Rao, Minfeng Zhu, Bo Zhang, Wei Chen

arXiv:2503.10615 · 2026-07-27 공개 · arXiv · PDF

reinforcement-learning supervised-fine-tuning multimodal-reasoning dataset cross-modal-reasoning formalization visual-language r1-onevision

Abstract

Large Language Models have demonstrated remarkable reasoning capability in complex textual tasks. However, multimodal reasoning, which requires integrating visual and textual information, remains a significant challenge. Existing visual-language models often struggle to effectively analyze and reason visual content, resulting in suboptimal performance on complex reasoning tasks. Moreover, the absence of comprehensive benchmarks hinders the accurate assessment of multimodal reasoning capabilities. In this paper, we introduce R1-Onevision, a multimodal reasoning model designed to bridge the gap between visual perception and deep reasoning. To achieve this, we propose a cross-modal reasoning pipeline that transforms images into formal textual representations, enabling precise language-based reasoning. Leveraging this pipeline, we construct the R1-Onevision dataset which provides detailed, step-by-step multimodal reasoning annotations across diverse domains. We further develop the R1-Onevision model through supervised fine-tuning and reinforcement learning to cultivate advanced reasoning and robust generalization abilities. To comprehensively evaluate multimodal reasoning performance across different grades, we introduce R1-Onevision-Bench, a benchmark aligned with human educational stages, covering exams from junior high school to university and beyond. Experimental results show that R1-Onevision achieves state-of-the-art performance, outperforming models such as GPT-4o and Qwen2.5-VL on multiple challenging multimodal reasoning benchmarks. Code, dataset and benchmark are available at https://github.com/Fancy-MLLM/R1-Onevision

한국어 요약

한 줄 요약

R1-Onevision은 시각-언어 정보를 정확히 결합하는 새로운 cross-modal 추론 파이프라인과 데이터셋, 평가 벤치마크를 제시하여 멀티모달 추론 성능을 획기적으로 향상시킨다.

핵심 기여도

핵심 아이디어

기존 시각-언어 모델은 시각 정보를 구조화하고 깊이 있는 추론을 수행하는 데 어려움을 겪는다. 이를 해결하기 위해 R1-Onevision은 **cross-modal 추론 파이프라인**을 도입하여 이미지를 정형화된 텍스트 표현으로 변환하고, 언어 모델이 이를 기반으로 정확한 추론을 수행하도록 유도한다. 이는 기존의 미리 정의된 템플릿이나 단순한 정답 모방 방식과 달리, 시각 정보를 **구조화된 추론 과정**으로 변환하는 새로운 접근법이다.

또한, R1-Onevision은 **Supervised Fine-Tuning (SFT)**과 **Reinforcement Learning (RL)**을 결합한 **2단계 훈련 전략**을 통해 모델의 추론 능력과 일반화 능력을 동시에 강화한다. SFT는 데이터셋 내 단계별 추론 패턴을 학습하여 모델의 일관성을 높이고, RL은 복잡한 문제 해결 능력을 향상시킨다.

기술적 접근법

주요 결과

의의 및 한계

R1-Onevision은 멀티모달 추론에서 **시각 정보를 구조화하고 정확히 추론하는 새로운 접근법**을 제시하며, 교육 단계별 평가 벤치마크를 통해 모델의 실제 적용 가능성을 평가할 수 있는 기반을 제공한다. 특히, SFT와 RL을 결합한 훈련 전략은 모델의 추론 능력과 일반화 능력을 동시에 향상시키는 데 기여한다.

그러나, R1-Onevision은 **교육 중심의 데이터셋과 벤치마크에 의존**하며, 실제 산업 환경에서의 일반화 능력을 평가하기에는 한계가 있을 수 있다. 또한, **모델의 파라미터 크기와 훈련 데이터의 양**이 성능에 큰 영향을 미치므로, 더 작은 모델에서도 동일한 성능을 유지하는 방법이 필요하다.

실용적 활용

R1-Onevision은 **교육 분야에서의 자동 문제 풀이 시스템**, **과학 및 공학 분야의 시각 정보 분석**, **의료 영상 해석** 등에서 활용 가능하다. 특히, **교과서 수준의 문제 해결 능력을 가진 AI 시스템** 개발에 기여할 수 있으며, **교육 기술(EduTech)** 및 **AI 기반 학습 도구** 개발에 적합하다.