Skip to main navigation Skip to search Skip to main content

Variational Regularization for End-to-End Speech Deepfake Detection

Research output: Chapter in book / Conference proceedingConference article published in proceeding or bookAcademic researchpeer-review

Abstract

Current research in end-to-end speech deepfake detection predominantly centers around inputting 'raw' waveforms to a deep architecture, such as RawNet2, and training the deep neural network to predict if the waveforms are fake. However, direct processing of waveforms could cause over-parameterization in the network, reducing its generalizability. To overcome this limitation, we propose a multi-level variational regularization framework integrating a modified Variational Autoencoder (VAE) with discriminative constraints. Specifically, we adopt an VAE with a deepfake discrimination constraint to regularize a RawNet2-based high-level feature map (HFM) extractor. Experimental results show that the proposed variational regularization leads to HFM features that improve the performance of AASIST, SE-Rawformer, and RawBMamba by 3 6. 0 1%, } {1 0.07%, and 6.35%, respectively.

Original languageEnglish
Title of host publication2025 Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages2241-2246
Number of pages6
ISBN (Electronic)9798331572068
DOIs
Publication statusPublished - Oct 2025
Event17th Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025 - Singapore, Singapore
Duration: 22 Oct 202524 Oct 2025

Publication series

Name2025 Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025

Conference

Conference17th Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025
Country/TerritorySingapore
CitySingapore
Period22/10/2524/10/25

ASJC Scopus subject areas

  • Artificial Intelligence
  • Computer Science Applications
  • Hardware and Architecture
  • Signal Processing

Fingerprint

Dive into the research topics of 'Variational Regularization for End-to-End Speech Deepfake Detection'. Together they form a unique fingerprint.

Cite this