Cross-Modal Dynamic Aggregation with Adaptive Relevance Modulation Fusion Network for Remote Sensing Visual Question Answering

Authors: Zihua Zuo, Chao LI
Conference: ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
Pages: -
Keywords: remote sensing, visual question answering (VQA), cross-modal fusion, dynamic aggregation

Abstract

Remote Sensing Visual Question Answering (RS VQA) task aims to provide accurate answers to questions about RS images. However, the semantic gap between low-level visual features and high-level semantics complicates the understanding of complex questions. Moreover, the lack of dynamic modulation mechanisms for integrating visual and textual features impedes the balanced interpretation of image content and textual semantics. To this end, we propose the Cross-Modal Dynamic Aggregation with Adaptive Relevance Modulation Fusion Network for RS VQA (CDAR-Net). Specifically, we propose a Cross-Modal Dynamic Aggregation Module (CMDA) that employs a multi-level attention mechanism to iteratively fuse text-guided visual features, enabling a progressive transition and dynamic integration from low-level visual features to high-level semantic information. We further introduce an Adaptive Correlation Modulation Fusion Module (ACMF) that dynamically adjusts visual and textual feature weights based on questions context, enhancing the representation of relevant information. Experimental results demonstrate that CDAR-Net outperforms existing state-of-the-art methods on three RS VQA datasets.
📄 View Full Paper (PDF) 📋 Show Citation