CAF-YOLO: YOLOv8-based Bimodal Pedestrian Detection Method Using Cross-Attention Fusion

Authors: Jianhong Li, Haoran Wu, Xiangjing Wei, Haiyi Huang, Yaojuan Wang, Wenhui Zhang
Conference: ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
Pages: -
Keywords: Pedestrian detection, Bimodal fusion, Cross-attention, YOLOv8, Visible-infrared images

Abstract

Visible-infrared bimodal pedestrian detection holds significant application value in scenarios such as intelligent surveillance and autonomous driving. However, visible images suffer performance degradation under low-light conditions, while infrared images lack detailed texture information. Moreover, existing fusion methods still have limitations in real-time performance, interaction efficiency, and lightweight design. To address these challenges, this paper proposes an improved YOLOv8 model based on cross-attention fusion, named CAF-YOLO. The method designs a Cross-Attention Fusion (CAF) module to achieve feature complementarity and semantic alignment between the two modalities through bidirectional attention mechanism. Meanwhile, an Efficient Multi-scale Attention (EMA) module is introduced to enhance feature representation capability. Experiments on LLVIP and FLIR datasets show that CAF-YOLO achieves mAP@0.5 improvements of 9.1% and 1.2% over visible and infrared single-modal baselines on LLVIP, and 17.0% and 3.6% on FLIR, respectively. The model has only 11.74M parameters, significantly outperforming fusion methods with higher computational complexity. Ablation studies and visualization analysis verify the effectiveness of each module, and the model maintains robust detection capability in low-light and occlusion scenarios.
📄 View Full Paper (PDF) 📋 Show Citation