Component-Aware Spatio-Temporal Adaptation of Frozen Foundation Models for Video Deepfake Detection

Authors: Tianyi Zhang, Menghan Liang, Wenzheng Liu, Cheng Fu, Gang Zhu
Conference: ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
Pages: -
Keywords: Deepfake Detection , Video forensics , Spatio-temporal modeling , Cross-domain generalization

Abstract

The advancement of deep generative models facilitates realistic synthetic facial videos, threatening social trust and digital security. Existing detection methods achieve strong in-domain performance but suffer from cross-dataset degradation, primarily due to overfitting to dataset-specific spatial artifacts. To address these challenges, we propose a parameter-efficient, video-based deepfake detection framework that leverages a frozen foundation model encoder coupled with a lightweight spatio-temporal decoder. First, we introduce a Component-Aware Spatial Enhancement (CASE) module that selectively accentuates manipulation-prone facial regions, such as the eyes, mouth and nose, while capturing global-local structural inconsistencies, thereby enabling the detection of subtle artifacts. Second, it is complemented by a Bidirectional Spatio-Temporal (Bi-ST) decoder that models local temporal transitions and bidirectional temporal dependencies across sampled frames, enabling robust temporal reasoning within each video clip. Without fine-tuning the backbone network, our framework achieves robust cross-dataset generalization by jointly reasoning about spatial and temporal anomalies. Finally, extensive experiments demonstrate that the proposed method performs competitively against strong baselines, particularly under cross-dataset evaluation.
📄 View Full Paper (PDF) 📋 Show Citation