SDGMamba: Stage-Aware Mamba with Deformable Detail and Gated Fusion for Clothing Parsing

Authors: Haitao Fu, Shaoyu Wang, Zhongyuan Teng, Jiaxin He, Hong Wan, Xiujin Shi
Conference: ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
Pages: -
Keywords: Clothing Parsing, Visual State Space Model, Stage-Aware Attention, Deformable Detail, Gated Fusion

Abstract

Clothing parsing is a pivotal subtask of semantic segmentation with significant research and practical implications. However, it faces persistent challenges, including substantial scale variations among clothing components, posture-induced occlusions, and high inter-class visual similarity, which necessitate models capable of synergizing global semantics with local details. While existing CNN-based approaches are constrained by limited receptive fields, Transformer-based—despite their proficiency in modeling long-range dependencies—often suffer from uniform input processing that dilutes fine-grained details. To address these limitations, we propose SDGMamba, a novel clothing parsing framework that effectively leverages scale-dependent features for global modeling while preserving intricate characteristics. First, we design the Stage-Aware Attention VSS (SASS) Block, which dynamically allocates attention based on network depth to facilitate hierarchical structure-aware adjustments. Second, we introduce the Deformable Detail Retention (DDR) module, which utilizes deformable convolutions to adaptively align and fuse multi-scale depth information, thereby enhancing texture representation. Finally, the Gated Cross-Scale Fusion (GCSF) module employs a gating mechanism to refine shallow features guided by high-level semantics, strengthening semantic coherence among components. Experiments on the CFPD dataset demonstrate that SDGMamba achieves a PA of 94.48% and an mIoU of 58.57%, exhibiting superior performance compared to CNN-based, Transformer-based, and Mamba-based methods.
📄 View Full Paper (PDF) 📋 Show Citation