VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning

Junxiang Xu, Ruisi Wang, Fanyi Pu, Maijunxian Wang, Ran Ji, Tongxi Zhou, Chenyang Gu, Jing Zuo, Hongcan Xiao, Yimeng Geng, Wanqi Yin, Wei Chen, Oscar Qian, Zhengan Yan, Ziqi Huang, Haiwen Diao, Liang Pan, Bo Li, Xiangyu Fan, Dezhi Luo, Fengyuan Yu, Zehong Zhao, Qingying Gao, Tinghui Zhu, Yilan Zhang, Jingqi Tong, Pinyuan Feng, Zhengze Jiang, Letian Wang, Ziyu Guo, Renrui Zhang, Jieneng Chen, Sonia Joseph, Constantin Venhoff, Saman Motamed, Mengyue Yang, Chandra Sripada, Alan Yuille, Philip Torr, Lvmin Zhang, Vikash Kumar, Daniel Khashabi, Nikolaus Kriegeskorte, Raphaël Millière, Vincent C. Müller, Anyi Rao, Quan Wang, Ziwei Liu, Dahua Lin, Lei Yang, Hokin Deng, Zhongang Cai

arXiv:2608.26105 · 2026-08-27 공개 · arXiv · PDF

reinforcement-learning video-generation llm-evaluation benchmarking image-generation generative-models visual-reasoning multi-modal

Abstract

Native visual reasoning treats visual generation as the medium of reasoning itself: visual states (i.e. images and videos) are not merely inputs to be understood or outputs to be rendered, but first-class substrates for problem solving beyond language. Yet progress remains bottlenecked by the lack of scalable training tasks, reliable feedback, and controlled comparisons across generative substrates. In this work, we introduce VBVR-Pro, a closed-loop testbed that makes native visual reasoning through generation trainable, verifiable, optimizable, and experimentally controllable. 1) Task scaling. VBVR-Pro turns visual reasoning into a controlled task space of 300 procedurally generated tasks. Models trained on VBVR-Pro show strong transfer beyond the proposed suite across seven external visual reasoning benchmarks such as RISE-Video, MME-CoF-Pro, and BabyVision. 2) Verifiable rewards. VBVR-Pro provides verifiable reward scorers for task-grounded evaluation. Through a systematic study of leading MLLMs as judges, we identify recurring failure modes of the prevalent VLM-as-a-judge paradigm. In contrast, the proposed scorers are grounded in deterministic, task-specific rules, achieve fine-grained alignment with human judgments. Importantly, they serve as reliable reward signals for large-scale multi-task reinforcement learning and demonstrate stronger post-RL performance across visual reasoning tasks. 3) Mechanism study. VBVR-Pro enables controlled modality studies across more than 30 image, video, and interleaved generators. Our analysis shows that video generation remains strongest for tasks requiring persistent spatiotemporal state tracking, while interleaved generation provides a compute-efficient alternative. Critically, ablations and probing suggest the presence of vision-native trajectories that are crucial to visual reasoning. We release all data, models, scorers, and code.

한국어 요약

한 줄 요약

VBVR-Pro는 시각 생성을 기반으로 한 네이티브 시각 추론을 학습 가능하고 검증 가능하게 구현한 종합적인 테스트베드이다.

핵심 기여도

핵심 아이디어

VBVR-Pro는 시각 생성을 추론의 매체로 삼는 네이티브 시각 추론을 체계적으로 연구할 수 있도록 설계된 테스트베드이다. 기존 연구는 주로 언어 중심 추론에 집중했으나, VBVR-Pro는 시각 상태(이미지, 동영상)를 추론의 핵심 요소로 삼는다. 이는 단순히 입력이나 출력이 아닌, 문제 해결의 실제 매체로 기능한다. VBVR-Pro는 프로시저 생성 태스크, 검증 가능한 보상 점수, 생성 방식 비교를 통합하여, 추론 과정을 학습 가능하고 실험적으로 제어 가능하게 만든다. 특히, 시각 생성이 지속적 시공간 추적에 강점을 가지며, 언어적 사고 체인보다 시각적 추적 경로가 더 중요하다는 통찰을 제시한다.

기술적 접근법

주요 결과

의의 및 한계

VBVR-Pro는 네이티브 시각 추론을 체계적으로 연구할 수 있는 첫 번째 종합적 테스트베드로, 생성 기반 추론의 학습, 검증, 최적화를 가능하게 한다. 특히, 강화 학습과 검증 가능한 보상 시스템을 결합하여, 시각 추론 모델의 일반화 능력을 향상시키는 데 기여한다. 그러나 VBVR-Pro는 현재까지는 시뮬레이션 기반 태스크에만 적용되며, 실제 세계 시나리오로의 확장 가능성은 아직 명시되지 않음. 또한, 생성 방식 간 비교는 제한된 모델 집합에 기반하므로, 더 다양한 생성기와의 비교가 필요하다.

실용적 활용

VBVR-Pro는 시각 추론 기반의 자율 시스템, 예측 시뮬레이션, 인공지능 교육 등에서 활용 가능하다. 특히, 강화 학습과 결합하여 시각적 문제 해결 능력을 향상시키는 산업용 AI 개발에 기여할 수 있다.