Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination

Shuo Liang, Yixing Ma, Pengfei Zhou, Xingyan Chen, Zihan Mei, Manting Li, Feihan Chen, Zhiwen Wang, Bin Xu, Haotian Zhang, Jiajun Song, Shiya Su, Run Liu, Zhenghang Ni, Yifa Yu, Jintao Hong, Bolong Feng, Yifei Liu, Zirui Zhang, Jingxuan Zhang, Songlin Zhao, Yifan Bai, Kang Tan, Yizhe Liu, Junhao Du, Yongtao Ge, Zhaopan Xv, Xinyuan Zhang, Mengru Ma, Chunhua Shen, Wei Wang, Yang You, Zheng Zhu, Kaipeng Zhang, Wangbo Zhao

arXiv:2608.14391 · 2026-08-17 공개 · arXiv · PDF

multimodal-models misinformation video-detection social-dissemination zero-shot-models generator-analysis detector-generalization real-world-crisis

Abstract

Recent video generators can fabricate realistic depictions of wars, disasters, public emergencies, and other real-world crises, creating substantial risks of misinformation. Existing benchmarks, however, provide limited evidence on detector and generator behavior in such settings, including how detectability varies with generation conditions, how people perceive generated videos, and whether detectors remain reliable during social dissemination. To address this gap, we introduce RA-Bench, a benchmark for AI-generated video detection that uses Real videos as Anchors. RA-Bench contains 17,886 videos, comprising 1,830 real-video anchors across 10 social-risk categories and 16,056 generated clips from four open-source and five closed-source generators. Based on RA-Bench, we organize our evaluation along three dimensions. We first assess detector generalization across seven traditional detectors, ten zero-shot multimodal models under three review settings, and two MLLMs specifically fine-tuned on AI-generated video detection. Across these methods, none of the three detector families generalizes consistently across RA-Bench instances. We then examine how detectability varies with generation quality, conditioning information, and sampling seeds. These analyses show that generation properties affect detector families differently, while source-level detection patterns remain stable across seeds. Finally, we study human authenticity judgments and detector reliability during social dissemination. We find that videos that mislead people are also difficult for current detectors, and that social dissemination makes detection harder. Together, these findings show that current methods struggle to detect realistic AI-generated videos, highlighting the need for detectors robust to evolving video generators.

한국어 요약

한 줄 요약

RA-Bench라는 새로운 벤치마크를 통해 AI 생성 영상의 탐지 실효성을 평가하고, 사회적 위기 상황에서의 탐지 난이도를 분석한다.

핵심 기여도

핵심 아이디어

기존 연구는 AI 생성 영상 탐지에 있어 일반적인 영상에 집중했으나, 사회적 위기 상황(전쟁, 재해 등)에서의 탐지 실효성은 명확하지 않았다. 이를 해결하기 위해 RA-Bench는 실제 영상(Anchor)을 기반으로 생성 영상을 비교하는 방식을 채택했다. 각 실제 영상의 첫 프레임을 기반으로 4개 오픈소스, 5개 클로즈드소스 생성기에서 영상을 생성하여, 다양한 생성 조건과 사회적 위험 요소를 반영한 데이터셋을 구성했다. 이는 생성 영상의 탐지 난이도가 생성 조건, 인간 인지, 사회적 확산에 따라 달라질 수 있음을 밝히는 데 기여한다.

기술적 접근법

주요 결과

의의 및 한계

RA-Bench는 사회적 위기 상황에서의 AI 생성 영상 탐지 실효성을 체계적으로 평가할 수 있는 첫 번째 벤치마크로, 탐지기의 일반화 능력과 생성 조건, 인간 인지, 사회적 확산에 따른 탐지 난이도를 명확히 밝혔다. 그러나 현재 탐지기는 생성 조건 변화에 취약하며, 사회적 확산 시 신뢰도가 급격히 하락한다는 한계가 드러났다. 또한, 일부 클로즈드소스 생성기(Seedance2.0, Kling)의 탐지률이 낮아, 탐지기의 진화가 필요하다는 점이 강조된다.

실용적 활용

RA-Bench는 미디어 감시, 위기 관리, 사회적 위기 대응 등에서 AI 생성 영상의 탐지 정확도를 평가하는 데 활용될 수 있다. 특히, 사회적 확산 시뮬레이션 RA-Bench-LastMile은 소셜 미디어 플랫폼의 탐지 시스템 개선에 기여할 수 있으며, 정부 및 언론 기관이 위기 상황에서의 정보 신뢰도를 유지하는 데 도움이 될 수 있다.