Dual Associations Semantic Enhancement for Image-Text Matching

Authors: Xinlin Zhao, Tao Yao, Yafei Bu, Li Liu, Yuling Zhang, Linliang Zhang
Conference: ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
Pages: -
Keywords: Image-text Matching, Dual Associations, Cross-modal Retrieval

Abstract

Image-text matching faces the significant challenge in effectively mitigating the large visual-semantic discrepancy among different modalities. Existing studies mainly address this by projecting multi-modal data into a common subspace for measuring their semantic similarities. However, most of them often overlook the inherent asymmetry for directional associations, i.e., the differences between im-age-to-text and text-to-image affinities, which often results in inaccurate retrieval results. To tackle the challenge, we propose a Dual Associations Semantic En-hancement (DASE) model to capture bidirectional image-text semantic associa-tions. Specifically, we first build a two-layer GCN fusion network to construct and mine the semantic associations for each modality. And then, due to the inher-ent asymmetry of directional associations, a Dual Associations Alignment Mod-ule (DAAM) is designed to capture the dual associations between visual and tex-tual modalities, enabling comprehensive cross-modal fine-grained interaction. Fi-nally, global alignment is incorporated with the local alignment to achieve full semantic matching across heterogeneous modalities in a unified embedding space. Experimental results on two publicly available datasets demonstrate that the pro-posed DASE model achieves significant performance improvements in image-text matching tasks compared to baseline methods, validating its effectiveness and superiority.
📄 View Full Paper (PDF) 📋 Show Citation