video-generation model-evaluation text-to-video cogvideox physical-commonsense realism lumiere dream-machine
Abstract
Recent advances in internet-scale video data pretraining have led to the development of text-to-video generative models that can create high-quality videos across a broad range of visual concepts, synthesize realistic motions and render complex objects. Hence, these generative models have the potential to become general-purpose simulators of the physical world. However, it is unclear how far we are from this goal with the existing text-to-video generative models. To this end, we present VideoPhy, a benchmark designed to assess whether the generated videos follow physical commonsense for real-world activities (e.g. marbles will roll down when placed on a slanted surface). Specifically, we curate diverse prompts that involve interactions between various material types in the physical world (e.g., solid-solid, solid-fluid, fluid-fluid). We then generate videos conditioned on these captions from diverse state-of-the-art text-to-video generative models, including open models (e.g., CogVideoX) and closed models (e.g., Lumiere, Dream Machine). Our human evaluation reveals that the existing models severely lack the ability to generate videos adhering to the given text prompts, while also lack physical commonsense. Specifically, the best performing model, CogVideoX-5B, generates videos that adhere to the caption and physical laws for 39.6% of the instances. VideoPhy thus highlights that the video generative models are far from accurately simulating the physical world. Finally, we propose an auto-evaluator, VideoCon-Physics, to assess the performance reliably for the newly released models.
한국어 요약
한 줄 요약
VideoPhy는 생성된 동영상이 물리적 상식을 얼마나 잘 반영하는지 평가하는 벤치마크로, CogVideoX-5B가 39.6%만 정확히 생성함을 보여준다.
핵심 기여도
- VideoPhy: 물리적 상식을 기반으로 생성된 동영상 평가를 위한 첫 번째 벤치마크 데이터셋.
- 688개의 인간 검증된 캡션을 기반으로 12개의 T2V 모델 평가.
- CogVideoX-5B가 39.6%의 인스턴스에서 캡션과 물리법칙을 모두 준수.
- VideoCon-Physics: 자동 평가 모델로, 대규모 평가를 저비용으로 가능하게 함.
핵심 아이디어
기존의 텍스트-비디오 생성 모델은 물리적 상식을 충분히 반영하지 못한다는 점이 드러났다. 이에 따라, VideoPhy는 인간의 직관적 물리 이해를 기반으로 생성된 동영상이 물리법칙을 얼마나 잘 따르는지 평가하는 새로운 접근법을 제시한다. 예를 들어, 유리에 물을 따르면 물 높이가 점점 올라가는 것처럼, 물리적 상황에 대한 직관적 예측이 필요하다. 이는 정확한 물리 시뮬레이션 없이도 인간의 경험과 인식을 기반으로 평가할 수 있다는 점에서 차별화된다. 또한, 모델이 객체의 물리적 속성(예: 고체-유체 상호작용)을 정확히 이해하지 못하면, 생성된 동영상이 비현실적으로 나타난다는 점을 강조한다.
기술적 접근법
- **VideoPhy 데이터셋**: 3단계 파이프라인을 통해 구성됨.
- (a) 대형 언어 모델을 사용한 캡션 생성
- (b) 인간 검증
- (c) 물리 시뮬레이션 기반 복잡도 어노테이션
- **평가 모델**: 12개의 T2V 모델 사용 (예: CogVideoX, Lumiere, Dream Machine 등)
- **평가 지표**: 캡션 준수도와 물리적 상식 준수도
- **자동 평가 모델**: VideoCon-Physics, 생성된 동영상의 의미적 일관성과 물리적 상식을 평가
주요 결과
- CogVideoX-5B가 39.6%의 인스턴스에서 캡션과 물리법칙을 모두 준수.
- 대부분의 모델이 고체-고체 상호작용(예: 공이 바닥에 튀는 경우)에서 실패.
- 12개 모델 중 가장 우수한 성능을 보인 CogVideoX-5B조차 60% 이상의 인스턴스에서 물리적 상식을 반영하지 못함.
- 모델이 객체의 물리적 속성을 정확히 인식하지 못해 비현실적인 동작을 생성.
의의 및 한계
VideoPhy는 생성 모델이 물리적 세계를 얼마나 정확히 시뮬레이션하는지를 평가하는 첫 번째 시도로, 물리적 상식을 기반으로 한 평가 체계를 제시한다. 특히, 인간의 직관적 경험을 기반으로 물리법칙을 평가함으로써, 기존의 정량적 물리 시뮬레이션에 의존하지 않는 새로운 접근법을 제시한다. 그러나, 이는 정확한 물리 시뮬레이션 없이 인간의 주관적 판단에 의존하기 때문에, 일부 경우 오류가 발생할 수 있다. 또한, VideoCon-Physics는 아직 완전히 신뢰할 수 있는 자동 평가 모델로 확립되지 않았으며, 추가 연구가 필요하다.
실용적 활용
VideoPhy는 텍스트-비디오 생성 모델의 물리적 일관성을 평가하는 데 활용될 수 있으며, 특히 로봇 학습, 시뮬레이션 기반 AI 교육, 게임 개발 등에서 실제 물리적 세계와 유사한 동작을 생성하는 데 중요한 역할을 할 수 있다. VideoCon-Physics는 대규모 모델 평가를 저비용으로 가능하게 하여, 연구 및 산업 현장에서 실용적으로 활용될 수 있다.