Archive/Imbalance-Aware Cross-Modal Focal Modulation for Cross-Dataset Audio-Visual Deepfake Detection
Imbalance-Aware Cross-Modal Focal Modulation for Cross-Dataset Audio-Visual Deepfake Detection
Shahad Mohammad Bn Dokiey, Tariq M. Khan, Qazi Emad Ul Haq
20 juillet 2026
en

Abstract

Audio-visual deepfake detection remains challenging under cross-dataset distribution shift, especially when the source-domain training data are severely imbalanced. Existing middle-fusion detectors often rely on softmax-based cross-attention, which can learn sharp source-domain token interactions and may transfer poorly to unseen datasets. This study proposes FocalNet, an audio-visual detector that replaces the cross-attention block of the 2D3MF framework with cross-modal focal modulation. The proposed module aggregates multi-scale temporal context before audio-visual interaction, enabling softmax-free contextual modulation between visual MARLIN features and audio EAT features. We evaluate the method under a strict FakeAVCeleb-to-DFDC protocol, where all training and validation is performed on FakeAVCeleb and the DFDC is used only as an unseen target-domain test set. Compared with the reproduced 2D3MF baseline, which collapses to a single-class prediction pattern on the DFDC, FocalNet achieves substantially stronger zero-shot score separation, with a DFDC ROC-AUC of 0.9324. Thresholded analysis further shows an improved balanced accuracy, macro F1 score, and MCC when the frozen source-domain operating point is applied. The model also preserves practical efficiency, requiring comparable FLOPs and lower per-sample inference time than the reproduced baseline. These findings suggest that cross-modal focal modulation is a promising alternative to attention-based middle fusion for audio-visual deepfake detection under dataset shifts, while broader validation across additional unseen datasets, multi-seed training, and deployment-oriented calibration remain important for future work.

IPC Classification

G06

Keywords

imbalance-awarecross-modalfocalmodulationcross-datasetaudio-visualdeepfakedetectionfutureinternetremainschallengingdistributionshiftespeciallywhensource-domaintrainingdataseverelyimbalancedexistingmiddle-fusiondetectors
Citer cette publication

€ 4.00