TY - GEN
T1 - Exploring Audio-Visual Fusion Methods in Foundation Model-Based Deception Detection
AU - Meng, Jiaxiang
AU - Sailor, Hardik B.
AU - Wang, Qiongqiong
AU - Liu, Tianchi
AU - Lee, Kong Aik
AU - Wang, Xingmei
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025/10
Y1 - 2025/10
N2 - The deception detection task aims to identify if a speaker is speaking truth or lie. It is a challenging problem due to limited training data and hence this poses a restriction for representation learning models to learn better features. Audio-visual multi-modal detection has gained significant attention for its superior performance compared to single-modality approaches. In practical scenarios, multi-modal integration can be challenging due to the distinct characteristics of each modality, making the fusion process difficult. In this paper, we employ the Querying Transformer (Q-former) to temporally align audio and visual features from foundation models and explore various methods to fuse these two modalities. Comprehensive experiments on the DOLOS dataset show that our proposed fusion technique outperforms systems based on individual audio and visual modalities. The experiments also indicate that early fusion using layer-bylayer alignment between the two modalities in the foundation model is not required. Instead, integrating over all layers of the two modalities yields the best performance.
AB - The deception detection task aims to identify if a speaker is speaking truth or lie. It is a challenging problem due to limited training data and hence this poses a restriction for representation learning models to learn better features. Audio-visual multi-modal detection has gained significant attention for its superior performance compared to single-modality approaches. In practical scenarios, multi-modal integration can be challenging due to the distinct characteristics of each modality, making the fusion process difficult. In this paper, we employ the Querying Transformer (Q-former) to temporally align audio and visual features from foundation models and explore various methods to fuse these two modalities. Comprehensive experiments on the DOLOS dataset show that our proposed fusion technique outperforms systems based on individual audio and visual modalities. The experiments also indicate that early fusion using layer-bylayer alignment between the two modalities in the foundation model is not required. Instead, integrating over all layers of the two modalities yields the best performance.
UR - https://www.scopus.com/pages/publications/105030449471
U2 - 10.1109/APSIPAASC65261.2025.11249189
DO - 10.1109/APSIPAASC65261.2025.11249189
M3 - Conference article published in proceeding or book
AN - SCOPUS:105030449471
T3 - 2025 Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025
SP - 1964
EP - 1968
BT - 2025 Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 17th Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025
Y2 - 22 October 2025 through 24 October 2025
ER -