Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model
Haoyu Zhao, Zihao Zhao, Tianyu Deng, Ziqin Xu, Zihao Zhang, Xudong Wang, Jinxiang Guo, Chen Gao, Ziyi Ye, Yeying Jin, Jiaxi Gu, Zuxuan Wu, Shuicheng Yan
arXiv:2609.18323 · 2026-09-19 공개 · arXiv · PDF
video-generation evaluation-framework multimodal-reasoning audio-visual-generation audio-video prefix-video omni-model world-reasoning
Abstract
Recent Omni-Modal Generative Models (Omni-Models) have advanced content generation toward unified modeling of text, images, video, and audio. MiniMax-H3 exemplifies this transition by combining multimodal context understanding with joint audio-visual generation in a shared latent framework. Its unified architecture raises a fundamental question: Can multimodal alignment improve the model's world reasoning, and what new evaluation paradigms do omni-modal inputs enable? To investigate this question, this work introduces a comprehensive evaluation framework organized around four complementary dimensions of physical world reasoning. Unlike existing evaluation frameworks for video generation and world models, which are often constrained by limited input modalities and evaluation settings where prompts closely match the target video content, our evaluation is specifically designed to exploit the multimodal inputs of Omni-Model. We construct a diverse set of novel tasks that require models to integrate complementary information across modalities. Specifically, we consider four scenarios, including implicit prompts paired with multiple frames, audio-image, prefix-videos, and audio-video inputs. Every single modality provides only partial evidence about the underlying event, requiring the model to jointly reason over the complementary semantic cues to infer latent event states and future dynamics. Across 517 evaluation instances, MiniMax-H3 achieves an overall success rate of 41.97%. Video-based Decision Reasoning yields the highest success rate at 56.00%, while Audio-based Disambiguation Reasoning is the weakest, reaching only 27.40%. These results indicate that effective multimodal integration remains key to fully exploiting the benefits of diverse input modalities. The project is available at https://github.com/gulucaptain/MiniMax-H3-Reason.
한국어 요약
한 줄 요약
MiniMax-H3의 물리적 세계 추론 능력을 다중 모달 입력 기반 평가 프레임워크로 평가한 연구.
핵심 기여도
- MiniMax-H3를 대상으로, 4가지 물리적 세계 추론 시나리오(예: Audio-based Disambiguation Reasoning)를 통해 새로운 평가 프레임워크를 제시.
- 517개 평가 인스턴스를 기반으로 MiniMax-H3의 성공률을 41.97%로 측정.
- Audio-based Disambiguation Reasoning에서 27.40%로 가장 낮은 성능을 보임.
- Video-based Decision Reasoning에서 56.00%로 가장 높은 성능을 기록.
핵심 아이디어
기존 평가 방식은 주로 텍스트 프롬프트가 전체 정보를 포함하는 경우에 초점을 맞추고, 다중 모달 정보의 통합 능력을 평가하지 못한다. 본 연구는 MiniMax-H3와 같은 Omni-Model이 물리적 세계를 추론할 수 있는지, 그리고 이를 위해 어떤 새로운 평가 패러다임이 필요한지를 탐구한다. 이를 위해, 각 모달이 부분적 정보만 제공하는 상황에서 모델이 이를 통합하여 잠재적 이벤트를 추론하도록 설계된 평가 프레임워크를 제시한다. 예를 들어, Audio-based Disambiguation Reasoning 시나리오에서는 단독으로는 모호한 시각 정보를 음향 정보로 명확히 해석해야 한다. 이는 단순한 다중 모달 입력 처리가 아니라, 의미적 통합과 추론을 요구하는 새로운 평가 패러다임이다.
기술적 접근법
- 평가 시나리오: Multi-view Spatial Reasoning, Audio-based Disambiguation Reasoning, Video-based Decision Reasoning, Audiovisual Integrated Reasoning.
- MiniMax-H3는 조건부 분포 $ p_\theta $를 통해 입력 관측치 $ \mathbf{x} $와 프롬프트 $ q $를 기반으로 출력 비디오를 생성.
- 각 인스턴스는 다중 모달 관측, 암시적 프롬프트, 어노테이션된 의미적 목표로 구성.
- 평가 데이터는 29개 하위 범주에 걸쳐 517개 인스턴스로 구성되며, 전문가 검증을 거침.
- 성능 평가 지표는 성공률 (Success Rate, SR)로, 생성된 비디오가 입력 관측과 의미적 목표를 얼마나 잘 반영하는지를 기준으로 삼음.
주요 결과
- MiniMax-H3는 517개 평가 인스턴스에서 41.97%의 전체 성공률을 기록.
- Video-based Decision Reasoning: 56.00% (베이스라인 대비 +14.03%)
- Audio-based Disambiguation Reasoning: 27.40% (베이스라인 대비 -14.57%)
- MSR: 43.50%, AVIR: 47.89%
- 특정 작업에서 성능 차이: Threading 75.00%, Pouring 58.30%, Cleaning 33.00% 등
- 성공률은 시나리오와 작업 유형에 따라 큰 차이를 보임.
의의 및 한계
본 연구는 Omni-Model이 단순히 다양한 모달을 입력받는 것 이상으로, 이를 통합하여 물리적 세계를 추론하는 능력을 평가하는 새로운 프레임워크를 제시한다. MiniMax-H3는 일부 시나리오에서 높은 성능을 보이지만, 특히 음향 정보를 기반으로 한 추론에서 한계를 드러낸다. 이는 다중 모달 입력을 지원하는 것이 반드시 효과적인 다중 모달 추론을 보장하지 않음을 시사한다. 또한, 생성된 비디오의 시각적 사실성과 물리적 일관성 사이의 괴리가 존재하며, 이는 모델이 입력을 올바르게 해석하지 못하거나, 생성 과정에서 오류가 발생했음을 의미한다.
실용적 활용
본 연구는 Omni-Model의 물리적 세계 추론 능력을 평가하는 기준을 제공하여, 향후 모델 개선 및 신뢰성 있는 생성 시스템 설계에 기여할 수 있다. 특히, 자율 시스템, 멀티모달 인터페이스, VR/AR 등에서 모델이 다양한 모달 정보를 통합해 현실적이고 일관된 결과를 생성하는 능력을 평가하는 데 활용될 수 있다.