GeoFuse-SAM: A Multimodal Data Fusion Framework for Boundary-Aware Foundation Model Adaptation in Medical Image Segmentation

Authors: Pengtao Ren, Qiyuan Wang, Yunyi Li, Kejiang Xiao, Lexi Shu
Conference: ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
Pages: -
Keywords: Vision Foundation Models \and Boundary-Aware Learning \and Medical Image Segmentation

Abstract

In medical image segmentation, accurately identifying anatomical structure boundaries is a core task for clinical computer-aided diagnosis. Although vision foundation models, represented by the Segment Anything Model (SAM), have demonstrated strong generalization potential, their deployment in fully automated clinical scenarios remains constrained by severe domain shift, reliance on manual prompts, and significant boundary degradation in ambiguous anatomical transition zones. To address these challenges, this paper proposes GeoFuse-SAM, a multimodal geometric data fusion adaptation framework.

The core innovation of this framework lies in breaking the limitations of single semantic features. Through a Geometry-Guided Cross-Domain Attention Fusion (GCAF) module, it achieves deep data fusion between the raw image data and high-frequency geometric priors (Sobel gradient fields) derived from computer graphics. This cross-modal interaction mechanism provides explicit spatial guidance to the model, significantly enhancing the robustness of boundary recognition. Furthermore, we introduce a lightweight Parallel Fusion Adapter (PFA) to achieve medical semantic alignment, and propose a Parameter-Free Morphological Boundary Weighting (PMBW) strategy. This strategy utilizes morphological operators to pinpoint ambiguous boundary regions during the training phase and impose dynamic geometric constraints.

Experiments on two challenging medical datasets, BUSI and ISIC 2018, demonstrate that GeoFuse-SAM, operating in a fully automatic prompt-free mode, not only maintains leading region segmentation accuracy (achieving a Dice score of 89.85\% on ISIC 2018), but also effectively suppresses the boundary degradation phenomenon of foundation models in grayscale modalities on the core boundary metric HD95 (optimized to 15.97 on BUSI). Without introducing extra inference parameters, it exhibits superior edge fidelity compared to existing medical fine-tuned foundation models (such as SAM-Med2D). This study provides a high-fidelity, low-cost technical paradigm for robust medical image segmentation.
📄 View Full Paper (PDF) 📋 Show Citation