Skip to main navigation Skip to search Skip to main content

Multimodal Large Language Model for Deepfake Video Detection and Description

Research output: Chapter in book / Conference proceedingConference article published in proceeding or bookAcademic researchpeer-review

Abstract

The explosive growth of deepfake video technology is steadily eroding public trust in visual media, necessitating detectors that not only flag forgeries but also provide explanations for them. In this paper, we introduce DVDD-LLaMA, a multimodal large language model that unifies two complementary vision encoders: CLIP for frame-level cross-modal alignment and a SwinT-based Deepfake-Sniffing Encoder for spatio-temporal anomaly capture, followed by a Compact Visual Connector that condenses those features while preserving critical manipulation cues. A lightweight bridging layer then fuses these visual signals with user prompts, enabling the language model to deliver both a real/fake result and a detailed, human-readable rationale. To train and evaluate our model, we construct FF++VQA, a richly annotated deepfake-video question-answer dataset. Finetuned on the dataset, DVDD-LLaMA establishes new state-of-the-art performance in both supervised and zero-shot settings and remains robust to previously unseen attack types. Ablation studies confirm the importance of the Deepfake-Sniffing Encoder and Compact Visual Connector. Experiments show that DVDDLLaMA offers a high-performance and describable solution for deepfake video detection.

Original languageEnglish
Title of host publication2025 Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages2044-2049
Number of pages6
ISBN (Electronic)9798331572068
DOIs
Publication statusPublished - Oct 2025
Event17th Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025 - Singapore, Singapore
Duration: 22 Oct 202524 Oct 2025

Publication series

Name2025 Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025

Conference

Conference17th Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025
Country/TerritorySingapore
CitySingapore
Period22/10/2524/10/25

ASJC Scopus subject areas

  • Artificial Intelligence
  • Computer Science Applications
  • Hardware and Architecture
  • Signal Processing

Fingerprint

Dive into the research topics of 'Multimodal Large Language Model for Deepfake Video Detection and Description'. Together they form a unique fingerprint.

Cite this