Detecting AI-Generated Video: A Vision–Language Dual-View Survey

Dylan Xinming Hou; Juntian Zhang; Xu Gu; Yichen Wu; Nils Lukas; Gus Xia; Xiuying Chen; Yuhan Liu

Detecting AI-Generated Video: A Vision–Language Dual-View Survey

Dylan Xinming Hou, Juntian Zhang, Xu Gu, Yichen Wu, Nils Lukas, Gus Xia, Xiuying Chen, Yuhan Liu

Abstract

The evolving realism of AI-generated Videos (AIGC-V) is rapidly rendering traditional artifact-centric detection insufficient, necessitating a paradigm shift from low-level inspection to high-level semantic verification. This paper presents a comprehensive survey of AIGC-V detection, reframing the task as Factual Fidelity Verification, which asks whether the events, entities, and physical processes depicted in a video are consistent with real-world facts. To systematize this rapidly evolving field, we propose a Vision–Language Dual-View taxonomy that organizes existing methods into a hierarchical, four-layer landscape, spanning intrinsic cue analysis, spatiotemporal consistency modeling, cross-modal consistency reasoning, and language-guided world-level reasoning. This dual-view framing highlights a fundamental transition from artifact matching to evidence-based semantic verification enabled by vision–language models and agentic reasoning pipelines. Based on a systematic review of 195 papers, we synthesize AIGC-V generation paradigms, survey the landscape of detection methods, and review evaluation metrics and benchmarks in line with proposed views. Finally, we discuss current challenges and identify promising directions toward robust, explainable, and trustworthy detection.

Anthology ID:: 2026.findings-acl.1613
Volume:: Findings of the Association for Computational Linguistics: ACL 2026
Month:: July
Year:: 2026
Address:: San Diego, California, United States
Editors:: Maria Liakata, Viviane P. Moreira, Jiajun Zhang, David Jurgens
Venue:: Findings
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 32221–32255
Language:
URL:: https://preview.aclanthology.org/ingest-acl/2026.findings-acl.1613/
DOI:
Bibkey:
Cite (ACL):: Dylan Xinming Hou, Juntian Zhang, Xu Gu, Yichen Wu, Nils Lukas, Gus Xia, Xiuying Chen, and Yuhan Liu. 2026. Detecting AI-Generated Video: A Vision–Language Dual-View Survey. In Findings of the Association for Computational Linguistics: ACL 2026, pages 32221–32255, San Diego, California, United States. Association for Computational Linguistics.
Cite (Informal):: Detecting AI-Generated Video: A Vision–Language Dual-View Survey (Hou et al., Findings 2026)
Copy Citation:
PDF:: https://preview.aclanthology.org/ingest-acl/2026.findings-acl.1613.pdf
Checklist:: 2026.findings-acl.1613.checklist.pdf

PDF Cite Search Checklist Fix data