Bi-CMFM: A Multimodal Image Fusion Method Based on Bi-Path Residual and Cross-Modal Feature Merging

Authors: Yang Jingdong, Zhou Zhentao, Liao Shengnan, Zhang Hongning, Teng Xiaoxi, Liu Tao, Zhang Wendong
Conference: ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
Pages: -
Keywords: multimodal image fusion; bi-path residual; cross-modal feature enhancement; selective scanning; state space model

Abstract

Multimodal medical image fusion integrates complementary information from CT, MRI, and PET to support clinical diagnosis and lesion localization. CNN-based methods are constrained by local receptive fields and fail to model long-range dependencies, while Transformer-based methods incur quadratic computational complexity; both paradigms suffer from inadequate inter-modal redundancy suppression, producing blurred edges and structural inconsistency. To address these limitations, Bi-CMFM is proposed, comprising the Bi-Path Residual Fusion module (BPRF), which employs parallel standard and dilated convolutions to preserve local texture while expanding the receptive field, and the Cross-Modal Fusion Module (CMFM), which applies a Cross-modal Feature Enhancement component (CFEM) and a selective scanning mechanism to model long-range dependencies at linear complexity. Experimental results demonstrate that Bi-CMFM achieves EN = 5.1281 and AG = 7.0127 on CT-MRI fusion, PSNR of 12.7624/19.6562 on PET-MRI/SPECT-MRI tasks, and top-ranked EN, SF, AG, SCD, and CC on KAIST, outperforming six representative baselines including DATFuse, DRCM, and FusionMamba.
📄 View Full Paper (PDF) 📋 Show Citation