TY - GEN
T1 - Multimodal Large Language Model for Deepfake Video Detection and Description
AU - Sun, Haoran
AU - Cai, Chen
AU - Lee, Kong Aik
AU - Chau, Lap Pui
AU - Wang, Yi
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025/10
Y1 - 2025/10
N2 - The explosive growth of deepfake video technology is steadily eroding public trust in visual media, necessitating detectors that not only flag forgeries but also provide explanations for them. In this paper, we introduce DVDD-LLaMA, a multimodal large language model that unifies two complementary vision encoders: CLIP for frame-level cross-modal alignment and a SwinT-based Deepfake-Sniffing Encoder for spatio-temporal anomaly capture, followed by a Compact Visual Connector that condenses those features while preserving critical manipulation cues. A lightweight bridging layer then fuses these visual signals with user prompts, enabling the language model to deliver both a real/fake result and a detailed, human-readable rationale. To train and evaluate our model, we construct FF++VQA, a richly annotated deepfake-video question-answer dataset. Finetuned on the dataset, DVDD-LLaMA establishes new state-of-the-art performance in both supervised and zero-shot settings and remains robust to previously unseen attack types. Ablation studies confirm the importance of the Deepfake-Sniffing Encoder and Compact Visual Connector. Experiments show that DVDDLLaMA offers a high-performance and describable solution for deepfake video detection.
AB - The explosive growth of deepfake video technology is steadily eroding public trust in visual media, necessitating detectors that not only flag forgeries but also provide explanations for them. In this paper, we introduce DVDD-LLaMA, a multimodal large language model that unifies two complementary vision encoders: CLIP for frame-level cross-modal alignment and a SwinT-based Deepfake-Sniffing Encoder for spatio-temporal anomaly capture, followed by a Compact Visual Connector that condenses those features while preserving critical manipulation cues. A lightweight bridging layer then fuses these visual signals with user prompts, enabling the language model to deliver both a real/fake result and a detailed, human-readable rationale. To train and evaluate our model, we construct FF++VQA, a richly annotated deepfake-video question-answer dataset. Finetuned on the dataset, DVDD-LLaMA establishes new state-of-the-art performance in both supervised and zero-shot settings and remains robust to previously unseen attack types. Ablation studies confirm the importance of the Deepfake-Sniffing Encoder and Compact Visual Connector. Experiments show that DVDDLLaMA offers a high-performance and describable solution for deepfake video detection.
UR - https://www.scopus.com/pages/publications/105030475817
U2 - 10.1109/APSIPAASC65261.2025.11249103
DO - 10.1109/APSIPAASC65261.2025.11249103
M3 - Conference article published in proceeding or book
AN - SCOPUS:105030475817
T3 - 2025 Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025
SP - 2044
EP - 2049
BT - 2025 Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 17th Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025
Y2 - 22 October 2025 through 24 October 2025
ER -