Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence

Runmao Yao, Kairui Hu, Yukang Cao, Ruisi Wang, Shulin Tian, Ziang Cao, Weichen Fan, Ziqi Huang, Yuhao Dong, Hao Li, Zhaoxi Chen, Zhongang Cai, Lei Yang, Ziwei Liu

arXiv:2607.16401 · 2026-07-21 공개 · arXiv · PDF

video-generation world-models model-evaluation sim-to-real chain-of-frames physical-laws reasoning-trace law-grounded

Abstract

Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of physical law. Yet existing benchmarks largely evaluate physical plausibility only at the output level, without verifying whether the model arrives there through a faithful, law-grounded reasoning process. We introduce Apple-PI, the first benchmark that anchors video-model evaluation explicitly in physical laws. Apple-PI comprises three components. 1) Orchard: a dataset of 400 videos covering ten canonical tasks in classical mechanics. It separates single-law tasks for confounder-free diagnosis from multi-law tasks for probing generalization. 2) Benchmark Protocol: a three-stage protocol based on scientific reasoning, including Perception, Formulation, and Deduction. It uses chain-of-frames prompting on infographic-annotated first frames, treating the generated video as the model's visible reasoning trace. 3) Evaluation Suite: a hybrid evaluation suite that combines MLLM-based subjective scoring with physics-law-grounded objective measures. This enables stage-resolved diagnosis of not only whether a model fails, but where it fails. Benchmarking 11 models shows that current video models remain far from reliable law-grounded world simulators, with the best video model scoring only 0.473. Our stage-, pillar-, and source-resolved analyses further expose a Perception-to-Formulation-to-Deduction bottleneck, weak multi-law state transfer, and a persistent Sim-to-Real gap. These findings position Apple-PI as a diagnostic foundation for guiding future video models toward world models with law-grounded physical intelligence.

한국어 요약

한 줄 요약

Apple-π는 물리 법칙 기반의 비디오 생성 모델 평가를 위한 첫 번째 벤치마크로, 400개의 Orchid 데이터셋과 3단계 프로토콜을 통해 모델의 물리적 추론 능력을 진단한다.

핵심 기여도

핵심 아이디어

Apple-π는 기존 벤치마크가 단지 출력의 물리적 타당성만 평가하는 데 그친 반면, 모델이 **물리 법칙을 기반으로 추론하는 과정**을 평가하는 데 초점을 맞춘다. 이는 **Newton식 과학적 추론**(Perception → Formulation → Deduction)을 비디오 생성 프로토콜로 구현한 것이다.

**Perception 단계**에서는 물리적 장면을 인식하는 능력을 평가하고, **Formulation 단계**에서는 물리 법칙을 내재화한 능력을, **Deduction 단계**에서는 법칙에 따라 미래 동역학을 생성하는 능력을 평가한다. 이 프로토콜은 **infographic-annotated 첫 번째 프레임**을 기반으로 **chain-of-frames prompting**을 사용하여 모델의 추론 과정을 시각화한 비디오로 평가한다.

기술적 접근법

주요 결과

의의 및 한계

Apple-π는 비디오 생성 모델이 단순히 시각적으로 타당한 결과를 생성하는 것을 넘어, **물리 법칙을 기반으로 추론하는 능력**을 평가할 수 있는 첫 번째 벤치마크로, **법칙 기반 물리 지능**(law-grounded physical intelligence) 연구의 진단 기반을 제공한다.

그러나 MLLM 기반 주관적 점수는 완벽한 물리 오라클이 아니며, 마스크 IoU나 속도 추정 등 객관적 측정도 분할 품질이나 노이즈에 영향을 받는다. 또한, 현재 모델은 복합 법칙을 처리하거나 실세계로 전이하는 데 어려움이 있다.

실용적 활용

Apple-π는 물리 법칙 기반 추론 능력을 갖춘 비디오 모델의 개발을 촉진하는 데 활용될 수 있으며, **로봇, 자율 주행, 물리 시뮬레이션** 등에서 정확한 물리적 동작 예측이 필요한 분야에 적용 가능하다. 또한, **법칙 기반 추론 데이터셋**과 **이해-생성 통합 모델** 개발에 중요한 기준이 될 수 있다.