reinforcement-learning vision-language synthetic-data spatial-reasoning multimodal-reasoning maze-navigation drawing-operations bounding-box-annotation
Abstract
As textual reasoning with large language models (LLMs) has advanced significantly, there has been growing interest in enhancing the multimodal reasoning capabilities of large vision-language models (LVLMs). However, existing methods primarily approach multimodal reasoning in a straightforward, text-centric manner, where both reasoning and answer derivation are conducted purely through text, with the only difference being the presence of multimodal input. As a result, these methods often encounter fundamental limitations in spatial reasoning tasks that demand precise geometric understanding and continuous spatial tracking-capabilities that humans achieve through mental visualization and manipulation. To address the limitations, we propose drawing to reason in space, a novel paradigm that enables LVLMs to reason through elementary drawing operations in the visual space. By equipping models with basic drawing operations, including annotating bounding boxes and drawing auxiliary lines, we empower them to express and analyze spatial relationships through direct visual manipulation, meanwhile avoiding the performance ceiling imposed by specialized perception tools in previous tool-integrated reasoning approaches. To cultivate this capability, we develop a three-stage training framework: cold-start training with synthetic data to establish basic drawing abilities, reflective rejection sampling to enhance self-reflection behaviors, and reinforcement learning to directly optimize for target rewards. Extensive experiments demonstrate that our model, named VILASR, consistently outperforms existing methods across diverse spatial reasoning benchmarks, involving maze navigation, static spatial reasoning, video-based reasoning, and multi-view-based reasoning tasks, with an average improvement of 18.4%.
한국어 요약
한 줄 요약
ViLaSR은 시각적 그리기 연산을 통해 LVLM의 공간 추론 능력을 향상시키는 새로운 패러다임을 제안하며, 18.4% 평균 성능 개선을 달성한다.
핵심 기여도
- **"Drawing to reason in space"** 패러다임 제안: LVLM이 경계 상자 및 보조선 그리기 연산을 통해 시각 공간에서 추론.
- **3단계 훈련 프레임워크**: cold-start 훈련, reflective rejection sampling, 강화 학습.
- **ViLaSR 모델**: 다중 시공간 추론 벤치마크에서 기존 방법 대비 평균 18.4% 개선.
- **자기 수정 능력 증대**: reflective rejection sampling을 통해 반사 행동 빈도가 2배 증가.
핵심 아이디어
기존 LVLM은 텍스트 중심 추론을 통해 시각 정보를 처리하지만, 이는 공간 추론에서 한계를 보인다. 특히, 정확한 기하학적 이해와 연속적인 공간 추적은 텍스트로 전환될 때 손실되기 쉽다. 본 연구는 인간이 정신적 시각화를 통해 공간 문제를 해결하는 방식을 모방하여, **"drawing to reason in space"**라는 새로운 패러다임을 제안한다. 이는 LVLM이 **경계 상자**(bounding box)와 **보조선**(auxiliary line) 같은 기본 그리기 연산을 통해 시각 정보를 직접 조작하고 분석하도록 한다. 이는 외부 툴에 의존하는 기존 접근법의 성능 한계를 극복하며, 시각 중심 추론을 가능하게 한다.
기술적 접근법
- **모델**: ViLaSR (Vision-Language model for Sophisticated Spatial Reasoning)
- **훈련 프레임워크**:
1. **Cold-start 훈련**: 합성 데이터를 사용해 기본 그리기 능력 배우기.
2. **Reflective rejection sampling**: 정답과 자기 수정 행동을 보이는 추론 경로를 선택적으로 강화.
3. **강화 학습 (RL)**: 정답 정확도와 추론 형식을 균형 있게 고려한 보상 함수로 최적화.
- **추론 단계**: 이미지, 동영상, 다중 뷰 입력 처리 가능.
- **연산**: 경계 상자 ($\mathcal{T}_{\text{box}}$), 보조선 ($\mathcal{T}_{\text{line}}$) 사용.
주요 결과
- **Maze, SpatialEval-Real, VSI-Bench, SPAR-Bench, MMSI-Bench** 등 5개 벤치마크에서 평균 18.4% 개선.
- **Reflective rejection sampling 제거 시**: 반사 행동 빈도 96.5% 감소, $\mathcal{T}_{\text{box}}$ 사용 85.0% 감소.
- **강화 학습 제거 시**: $\mathcal{T}_{\text{box}}$ 사용 159.4% 증가, 수치 정답 문제에서 -9.21% 성능 저하.
- **정확도**: 다중 선택 문제는 정확도 기준, 수치 문제는 Mean Relative Accuracy (MRA) 기준 평가.
의의 및 한계
ViLaSR은 LVLM이 시각 정보를 직접 조작하고 분석하도록 하는 새로운 추론 패러다임을 제시하며, 기존 텍스트 중심 접근법의 한계를 극복한다. 특히, **self-correction** 능력과 **시각적 추론의 일관성**을 향상시켜, 복잡한 공간 문제 해결에 기여한다. 그러나 학습 과정에서 합성 데이터에 의존하는 cold-start 훈련은 실제 세계 데이터와의 괴리를 초래할 수 있으며, 강화 학습의 보상 설계가 모델 성능에 큰 영향을 미친다는 점은 추가 연구가 필요하다.
실용적 활용
ViLaSR은 로봇, 증강현실, 자율주행 등 시공간 정보 처리가 필요한 산업에 적용 가능하다. 특히, 동영상 기반 추론과 다중 뷰 정보 통합이 필요한 상황에서 뛰어난 성능을 보일 것으로 기대된다. 연구적으로는 LVLM의 시각 추론 능력 발전 방향을 제시하며, 훈련 프레임워크의 각 단계별 역할을 명확히 규명한 점에서 중요한 기초가 될 수 있다.