MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing

Jiajia Lin, Mingxuan Du, Tuowen Zhou, Benfeng Xu, Hongtao Xie

arXiv:2607.27616 · 2026-07-31 공개 · arXiv · PDF

benchmark-evaluation mesh-reconstruction mpi-e-bench multi-person-interaction anatomically-plausible video-mined-data interaction-editing contact-geometry

Abstract

Text-to-image and personalized editing models now synthesize high-fidelity single-subject images with ease. Yet placing multiple named people into shared contact actions such as embrace, carry, or grapple still exposes major failures: fused limbs, invented extremities, and interpenetrating bodies. Existing evaluations largely overlook these anatomical and geometric issues, and VLM-as-a-judge checklists often saturate on Interaction while the errors remain obvious to humans. We introduce MPIE-Bench, a 2,500-sample benchmark of video-mined editing triplets spanning 405 scenes, 14 interaction categories, and four contact densities (C0-C3). We also propose MPIE-Eval, whose two new axes score contact-time geometry from a frozen public multi-person mesh reconstruction. Anatomy asks whether every human-like mass is explained by a complete set of reconstructed bodies, and Interaction asks whether the penetration and surface distance between those bodies match the contact the instruction asked for. Across ten editors, mesh Anatomy tops out at 0.65 and mesh Interaction at 0.72 on two different models, so no single editor is strong on both, while VLM checklists rate the same images above 0.95. A five-rater study confirms that both axes track human judgement more closely than a zero-shot VLM judge, and the rankings hold under ablation of every weight and threshold.

한국어 요약

한 줄 요약

MPIE-Bench는 다인물 상호작용 편집의 해부학적 일관성을 평가하기 위한 2,500개 샘플의 벤치마크와 MPIE-Eval이라는 새로운 평가 프로토콜을 제안한다.

핵심 기여도

핵심 아이디어

기존 텍스트-이미지 생성 모델은 단일 인물 이미지를 잘 생성하지만, 다인물의 신체 접촉 상호작용(예: 포옹, 운반, 격투)을 생성할 때 해부학적 오류(예: 융합된 팔, 가상의 신체 부위, 신체 침투)가 발생한다. 기존 평가 체계는 이러한 오류를 감지하지 못하며, VLM 체크리스트는 인간에게 명확한 오류를 높은 점수로 평가한다. 이를 해결하기 위해 MPIE-Eval은 Multi-HMR이라는 3D 인물 메시 복원 모델을 사용하여 생성 이미지의 해부학적 일관성과 상호작용 일치도를 정량적으로 평가한다. Anatomy는 생성된 모든 인물이 완전한 신체 메시로 재구성되었는지, Interaction은 신체 간 침투와 표면 거리가 요청된 상호작용과 일치하는지 평가한다.

기술적 접근법

주요 결과

의의 및 한계

MPIE-Bench와 MPIE-Eval은 기존 평가 체계가 감지하지 못하는 해부학적 오류를 정량적으로 평가할 수 있는 새로운 프레임워크를 제시한다. 특히, Multi-HMR을 기반으로 한 메시 재구성은 인간 판단과 높은 일치도를 보이며, 모델 간 비교의 신뢰성을 높인다. 그러나 MPIE-Eval은 Multi-HMR의 재구성 능력에 의존하므로, 메시 복원이 불완전한 경우 평가 정확도가 떨어질 수 있다. 또한, 현재는 10개 편집기만 평가했으며, 더 많은 모델과 데이터셋에서 검증이 필요하다.

실용적 활용

MPIE-Bench와 MPIE-Eval은 인공지능 기반의 다인물 생성 모델의 신뢰성과 안정성을 평가하는 데 활용될 수 있다. 특히, 인터랙티브 콘텐츠 제작, 게임 개발, VR/AR 등에서 신체 접촉을 포함한 인물 생성의 해부학적 일관성을 보장하는 데 유용하다.