Multi-scale local perception video topic recognition method based on semantic-guided feature pyramid

Authors: Jinyu Zhang, Lixu Sun, Longfei Zhang, Wushouer Silamu
Conference: ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
Pages: -
Keywords: Video topic recognition, multi-scale local perception, feature pyramid network, semantic-visual alignment, local feature enhancement

Abstract

Video topic recognition faces core challenges such as the limitations of static representations, cross-modal semantic misalignment, insufficient coverage of single-scale features, and weak temporal dynamic modeling. This paper finds that existing methods have significant deficiencies in local detail perception, leading to the neglect of key local cues such as "test tubes" in "laboratory" scenes, or the inability to simultaneously capture global environment and local details in "wedding" scenes. To address this, we propose a semantic-guided multi-scale feature pyramid learning method. The core innovation lies in the design of a temporal-semantic guided multi-scale local feature extractor (MLFE). This module can not only handle spatial multi-scale features but also incorporate temporal dynamics and textual semantic guidance to achieve adaptive scale selection. Based on this, we have constructed a complete recognition framework, including an improved cross-frame communication mechanism and a multi-granularity dynamic cue generation module. Experiments on benchmark datasets such as Kinetics-400 and Kinetics-600 show that our method significantly outperforms existing methods in scenarios requiring fine-grained local feature recognition, especially in topics such as "laboratory," "medical surgery," and "cooking process.
📄 View Full Paper (PDF) 📋 Show Citation