SIC Logo

SIC Open Access

ICIC Logo ICAI Logo

This material is presented to ensure timely dissemination of scholarly and technical work. Copyright and all rights therein are retained by authors or by other copyright holders. All persons copying this information are expected to adhere to the terms and constraints invoked by each author's copyright.

Back to index


ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026

Poster Volume Ⅰ

Poster Volume Ⅱ

All Posters

  • Anti-eavesdropping and interference-resistant unmanned aerial vehicle - MEC system robustness task offloading mechanism, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Jianbin Xue, Han Zhang

    Abstract: In recent years, people's sensitivity to information has greatly increased, and they are becoming more and more concerned about their privacy. Therefore, information security during data transmission is becoming increasingly important. In response to the important situation of information security transmission in the current era, considering the mobility of users. Firstly, this article establishes a theoretical model and provides a detailed analysis of the assumed system scenarios; The edge cloud layer is composed of UAVs equipped with mobile edge computing servers. The eavesdropping illegal users are on the ground layer with the legitimate mobile users on the ground. The ground users divide the tasks that need to be unloaded and unload them by partition, so that the unloading task processing can be uninterrupted and will not be subject to eavesdropping attacks by illegal users. Meanwhile, this article proposes a distance perturbation probability function to hide the location information of ground mobile users, avoid information leakage, and further ensure the security of the task unloading process for ground mobile users. Finally, simulation can prove that the proposed scheme can achieve the goal of improving the security of data transmission process.
    Keyword: Communication security; Mobile edge computing; Data offloading; Wireless communication;
    DOI: 10.65286/icic.v22i2.35652
    Cite

  • Robust Intelligent Computing for HPC Capacity Management: A Prediction-to-Operations Framework under Distribution Shift, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Hongyi Zhou, Shouqiang Liu
    Abstract: Accurate job runtime prediction is essential for effective resource management in high-performance computing (HPC) systems. However, user-provided walltime estimates are often unreliable, and the operational costs of prediction errors are asymmetric. This paper proposes a robust intelligent computing framework that connects machine learning-based runtime prediction to operational capacity management decisions, specifically addressing the challenge of distribution shift across heterogeneous HPC environments. Our methodology employs gradient-boosted models trained on submit-time metadata for prolonged-runtime risk estimation, with rigorous internal time-split testing and strict external validation using public Parallel Workloads Archive traces. To ensure reliable decision-making under distribution shift, we apply post-hoc probability recalibration, achieving well-calibrated uncertainty estimates with ECE reduced from 0.103 to 0.066. The predictive models are integrated into a discrete-event simulation framework to evaluate capacity management policies and quantify the safety--throughput trade-off. Experimental results demonstrate strong predictive performance with AUC 0.863 on an external validation trace with substantial distributional differences. The prediction-driven policy reduces overflow probability by 19.8\% compared to baseline, while systematic error-sensitivity analysis reveals how predictive uncertainty propagates into key operational metrics. These findings provide actionable insights for deploying intelligent prediction systems in real-world HPC environments.
    Keyword: Intelligent computing, machine learning, HPC workload traces, job runtime prediction, external validation, uncertainty quantification, distribution shift, capacity management
    DOI: 10.65286/icic.v22i2.19336
    Cite

  • AI-Driven Intelligent Optimization and Performance Prediction of Hydrogels, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Yisheng Lin, Penghui Pan
    Abstract: The rapid advancement of artificial intelligence (AI) and machine learning (ML) is reshaping scientific discovery through an information-theoretic par-adigm. Modern large-scale models function as entropy-compression engines: they extract low-entropy, transferable representations from high-dimensional and noisy data, reducing epistemic uncertainty in complex physical systems. This perspective motivates growing interest in entropy- oriented and infor-mation-theoretic approaches to model compression, generalization, and ro-bustness. Soft matter materials—particularly multiscale and compositionally diverse hydrogels—naturally constitute a high-entropy design landscape. Traditional intuition-driven exploration struggles to navigate such spaces or capture nonlinear, cross-scale dependencies reminiscent of complex-system behaviors. To address this challenge, we develop an intelligent hydrogel de-sign framework grounded in information theory and multiscale representa-tion learning. Feature engineering acts as an entropy-reduction step, com-pressing raw variables into structured representations that support stable multi-task learning. A multi-objective Bayesian optimization scheme with Gaussian-process surrogates and a q-Expected Hypervolume Improvement (q-EHVI) acquisition function further performs targeted information acquisi-tion, maximizing posterior entropy reduction while efficiently expanding the Pareto front. A self-evolving feedback loop integrates AI prediction with ex-perimental validation, allowing each iteration to reduce system-level uncer-tainty and refine representations across scales. Incorporating entropy-based and chaos-inspired metrics provides additional diagnostics for robustness and sensitivity, aligning the workflow with emerging information-theoretic prin-ciples used in reliable large-model development. Overall, this study estab-lishes an entropy-aware and representation-driven paradigm for the intelli-gent design of soft functional materials, offering a generalizable pathway for accelerating discovery and enhancing extrapolation capability.
    Keyword: Hydrogel, Wound Dressing, Machine Learning, Multi-objective Optimiza-tion, Materials Performance prediction.
    DOI: 10.65286/icic.v22i2.15907
    Cite

  • CAF-YOLO: YOLOv8-based Bimodal Pedestrian Detection Method Using Cross-Attention Fusion, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Jianhong Li, Haoran Wu, Xiangjing Wei, Haiyi Huang, Yaojuan Wang, Wenhui Zhang
    Abstract: Visible-infrared bimodal pedestrian detection holds significant application value in scenarios such as intelligent surveillance and autonomous driving. However, visible images suffer performance degradation under low-light conditions, while infrared images lack detailed texture information. Moreover, existing fusion methods still have limitations in real-time performance, interaction efficiency, and lightweight design. To address these challenges, this paper proposes an improved YOLOv8 model based on cross-attention fusion, named CAF-YOLO. The method designs a Cross-Attention Fusion (CAF) module to achieve feature complementarity and semantic alignment between the two modalities through bidirectional attention mechanism. Meanwhile, an Efficient Multi-scale Attention (EMA) module is introduced to enhance feature representation capability. Experiments on LLVIP and FLIR datasets show that CAF-YOLO achieves mAP@0.5 improvements of 9.1% and 1.2% over visible and infrared single-modal baselines on LLVIP, and 17.0% and 3.6% on FLIR, respectively. The model has only 11.74M parameters, significantly outperforming fusion methods with higher computational complexity. Ablation studies and visualization analysis verify the effectiveness of each module, and the model maintains robust detection capability in low-light and occlusion scenarios.
    Keyword: Pedestrian detection, Bimodal fusion, Cross-attention, YOLOv8, Visible-infrared images
    DOI: 10.65286/icic.v22i1.64042
    Cite

  • A Periodic Dynamic Graph and Meta-Graph Memory Network for Traffic Flow Prediction , ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Jianxuan Wei, Chao Shen
    Abstract: Accurately capturing time-varying spatial dependencies and inherent spatio-temporal heterogeneity remains a significant challenge in traffic flow predic-tion. Traditional methods often rely on predefined static graphs, which fail to adapt to the dynamic and periodic nature of urban traffic. To address these limitations, we propose a novel framework named PDMetaNet (Periodic Dy-namic graph with Meta-graph Memory Network). Specifically, a Periodic Dynamic Graph Generation Module (PDGM) is developed to construct bidi-rectional dynamic adjacency matrices by synergistically fusing real-time traf-fic signals with intra-day and intra-week periodicities. Building upon this, the Periodic Adaptive Graph Convolutional Recurrent Unit (PAGCRU) integrates dynamic graph convolutions with gated recurrent mechanisms to facilitate joint spatiotemporal feature modeling. Furthermore, a Spatio-Temporal Meta-Graph Learner (STMG) is incorporated to maintain a persistent bank of meta-nodes representing prototypical traffic patterns across different periods. This mechanism enables the model to effectively reuse historical knowledge and optimize the dynamic graph topology through an attention-based querying process. To enhance the robustness and generalization of the memory pa-rameters, contrastive and consistency losses are introduced as structural con-straints. Extensive experiments on the PeMS04 and PeMS08 datasets demon-strate that PDMetaNet significantly outperforms ten state-of-the-art baselines, achieving a substantial reduction in prediction error and exhibiting superior capability in capturing complex spatiotemporal dynamics.
    Keyword: Intelligent transportation systems, Traffic flow prediction, Meta-graph memory mechanism, Periodic dynamic graph, Graph neural networks
    DOI: 10.65286/icic.v22i2.11819
    Cite

  • FedQFS: A Blockchain-Based Federated Learning Framework Based on Quality Auditing, Fairness Deviation, and Sybil-Resistant Similarity, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Tian Fang, Bing Li, Jianglin Yu, Zuwei Chen, Baofu Han, Pan Feng
    Abstract: Federated Learning (FL) integrated with blockchain technology effectively mitigates the risks of single points of failure and malicious server behavior inherent in traditional federated learning architectures that rely on central servers. However, in practical implementation, the system still faces significant challenges in ensuring security and fairness. Malicious clients may upload low-quality or biased local models to disrupt global convergence, while Sybil nodes can forge multiple identities to manipulate aggregation, leading to severe performance degradation. To address these issues, this paper pro-poses a blockchain-based federated learning framework based on quality auditing, fairness deviation, and Sybil-resistant similarity (FedQFS). First, a quality auditing mechanism is designed to evaluate and filter local model updates through multi-dimensional metrics, effectively mitigating global ac-curacy degradation caused by malicious updates. Second, a Sybil-resilient identification mechanism is introduced, which leverages parameter similarity analysis to accurately detect forged identities, thereby enhancing the system's resistance against Sybil attacks. Finally, a fairness deviation quantification mechanism is incorporated to measure parameter distribution disparities and adaptively assign reasonable aggregation weights to benign clients with limited data, ensuring fairness in global model updates. Experimental results show that the proposed framework achieves over 95.5% accuracy on the MNIST dataset and maintains strong robustness under Sybil attacks scenarios. When four label-flipping attackers are present, its attack success rate de-creases by 18.4% compared with mainstream aggregation algorithms, validating the proposed method’s efficiency and security in complex distributed environments.
    Keyword: Federated Learning, Quality Auditing, Fairness Deviation, Sybil Attack De-fense, Blockchain
    DOI: 10.65286/icic.v22i2.13056
    Cite

  • Data-Driven Office Rental Price Prediction: An Empirical Study of Deep Learning vs. Gradient Boosting, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Yixuan Yang
    Abstract: The ability to predict office rental prices is very important in urban planning, investment strategy and regulation of the market. Nevertheless, it is still difficult to value commercial real estate because of extreme data heterogeneity. Predictive models should simultaneously be able to handle structured attributes that are rigid (e.g., floor area and level) and unstructured textual information (e.g., location tags and amenity descriptions), and there is no trivial way to reconcile the two. To address this, we built a new multi-modal dataset that focuses on commercial office listings in Hangzhou, China. We directly combine web-scraped transaction data with rich geospatial Point of Interest (POI) measures obtained through mapping APIs. In the recent past, the scholarly world has been overwhelmingly in favor of multi-modal deep learning models to process such heterogeneous data. However, our empirical results lead towards a different direction. We prove that Gradient Boosting Decision Trees (GBDT), namely CatBoost, not only significantly outperforms traditional statistical approaches in the commercial real estate field, but also decisively defeats complex deep learning baselines, including multi-modal fusion models based on encoders such as BERT. In addition to raw predictive performance, we have also opened the black box of rental valuation by using feature importance analysis and found that specific spatial features are much more important in driving prices than generic textual embeddings. In the end, this paper suggests a practical workflow of real estate analytics, which demonstrates that in tabular-dominant cases, strict feature engineering with effective tree models can be used to outperform over-parameterized neural networks.
    Keyword: Commercial Real Estate; Rental Price Prediction; CatBoost; Heterogeneous Data; Empirical Study.
    DOI: 10.65286/icic.v22i2.44137
    Cite

  • GeoGTM: Vision-Language Geospatial Intelligence for Street-Level Telecom Growth, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Junchi Ren, Chengyu Zhou, Tao He
    Abstract: Street-scale telecom growth requires jointly reasoning over heterogeneous evidence, including geospatial context (buildings, communities, and business districts), network capability (coverage and capacity), and customer-side signals (service usage and support tickets). Existing LLM-based sales assistants are often ungrounded: they ignore deliverability constraints, lack verifiable evidence, and fail to produce actionable plans that can be executed by field teams. In this paper, we introduce StreetCopilot, a grounded vision-language agent for street-level prospecting and service planning. StreetCopilot integrates (i) geospatial visual cues (e.g., street-view/remote-sensing building context and POIs), (ii) structured telecom signals (coverage maps, traffic KPIs, product portfolios), and (iii) unstructured operational text (tickets and visit notes) via a retrieval-augmented reasoning pipeline. To ensure actionability, we propose a deliverability-aware constraint module that verifies whether recommended bundles (network, wireless coverage, industrial devices, and cloud services) are feasible under local coverage and resource conditions, and a citation-grounded generation mechanism that attaches evidence snippets to each recommendation for auditability. We further present a street-scale closed-loop evaluation protocol that measures not only recommendation accuracy but also plan feasibility, evidence faithfulness, and end-to-end business outcomes (lead acceptance and conversion). Experiments on a real-world deployment in a city subregion demonstrate that StreetCopilot substantially improves prospect ranking quality and proposal drafting efficiency while maintaining high feasibility and evidence faithfulness, shedding light on grounded multimodal agents for real-world decision-making.
    Keyword: Grounded multimodal agent \and Telecom prospecting \and Deliverability-aware reasoning \and Evidence citation
    DOI: 10.65286/icic.v22i2.47080
    Cite

  • MiSA-Miner An Anchor-Driven Missing-Set Framework for Maximal and Closed Frequent Itemset Mining, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Liyuan Wang, Jing Yang
    Abstract: In frequent itemset mining over large transactional databases, the complete set of frequent itemsets is often extremely large and highly redundant.Therefore, maximal frequent itemsets (MFIs) and closed frequent itemsets (CFIs) are commonly used as more compact result representations. However,existing MFI/CFI mining methods, although designed to enumerate maximal or closed patterns directly, still often need to maintain and evaluate a large number of intermediate patterns during search and incur substantial maximality/closure checking costs in practice, which can limit mining efficiency. To address this issue, this paper proposes MiSA-Miner, an anchor-driven missing-set framework for MFI/CFI mining. The proposed method adopts the missing-set as the core representation, computes support directly from the cardinality of the union of missing-sets, and uses an anchor set composed of minimal elements to organize the search space and drive depth-first expansion, thereby forming a unified representation-and-search framework. For CFI mining, we further present an exact closure-checking method based on missing-sets, so that frequency checking and closure checking can be expressed in a unified representation framework. Experimental results show that, on dense transactional databases under relatively high support thresholds, MiSA-Miner produces MFI results consistent with classical baseline algorithms while achieving significant runtime advantages: compared with FPMax/GenMax, it attains approximately 2.0–36.3× speedup, and in some lower-threshold settings where baseline algorithms fail to finish within a 3-hour time limit (TLE), MiSA-Miner can still complete the mining task. We also analyze the performance boundary of the proposed method on very large sparse datasets and under low-support settings. Overall, this paper presents a unified framework for MFI/CFI mining under a complementary missing-set representation, and achieves quantifiable efficiency improvements in target scenarios.
    Keyword: Maximal Frequent Itemsets · Closed Frequent Itemsets · Missing-Set Representation · Anchor-Driven Search
    DOI: 10.65286/icic.v22i2.41168
    Cite

  • Ambiguity-Aware Keyword-Enhanced Label-Aware Semantic Fusion for Text Classification, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: peijun xie, chuncheng chi, fuxue li, hong yan
    Abstract: The rapid growth of textual data has made text classification a fundamental task in Natural Language Processing (NLP). However, real-world texts often exhibit semantic ambiguity, limited contextual information, and unclear category boundaries, which hinder conventional models from learning discriminative representations. To address these challenges, this paper proposes an Ambiguity-Aware Semantic Fusion Framework (AAK-LASFNet) for robust text classification.The proposed model constructs dual-view semantic representations by combining local contextual features extracted by a TextRCNN encoder with global semantic knowledge obtained from a large language model. An ambiguity estimation module is introduced to model semantic uncertainty, improving the model’s ability to handle ambiguous samples. Meanwhile, a label-aware attention mechanism and a keyword enhancement module are employed to strengthen category-related semantic cues. To further capture complex interactions between local and global representations, a high-order semantic fusion strategy is developed. In addition, a semantic consistency loss is imposed to align different semantic views and enhance representation stability.Extensive experiments on four benchmark datasets demonstrate that the proposed framework consistently outperforms strong baselines in terms of Accuracy, highlighting its effectiveness in alleviating semantic ambiguity in text classification.
    Keyword: Text classification, Semantic ambiguity, Label-aware attention.
    DOI: 10.65286/icic.v22i1.69374
    Cite

  • GeoMamba: Dynamic Step-Size Adaptive Modulation for Mamba-Based Point Cloud Classification, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Hanxiao Liu, Li Cui
    Abstract: Most current Mamba-based point cloud classification models use a constant step size, resulting in a fundamental limitation: the model fails to perceive the underlying structure of 3D geometric objects, thereby limiting classification accuracy. This limitation also renders the model susceptible to background noise in complex real-world environments. To address this issue, we propose GeoMamba. It introduces a lightweight, geometry-semantics dual-driven step-size modulation mechanism. By leveraging a multidimensional geometric encoder to extract relative positions, distance variations, and local neighborhood features, combined with semantic gating for context-aware modulation, our approach dynamically constrains the step-size adjustment range to [0.2, 5.0] times. Under lightweight constraints, adding only 0.9M parameters, GeoMamba achieves 86.57% classification accuracy on the most challenging ScanObjectNN subset PB-T50-RS without pretraining, surpassing PointMamba by 4.09 percentage points. With pretrained weights, the accuracy further improves to 89.21%. Through extensive ablation experiments and visualization analysis, we validate the effectiveness and interpretability of the dynamic step-size modulation mechanism.
    Keyword: Point cloud classification, Mamba, Dynamic step-size modulation, Lightweight model
    DOI: 10.65286/icic.v22i1.36043
    Cite

  • Multi-scale local perception video topic recognition method based on semantic-guided feature pyramid, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Jinyu Zhang, Lixu Sun, Longfei Zhang, Wushouer Silamu
    Abstract: Video topic recognition faces core challenges such as the limitations of static representations, cross-modal semantic misalignment, insufficient coverage of single-scale features, and weak temporal dynamic modeling. This paper finds that existing methods have significant deficiencies in local detail perception, leading to the neglect of key local cues such as "test tubes" in "laboratory" scenes, or the inability to simultaneously capture global environment and local details in "wedding" scenes. To address this, we propose a semantic-guided multi-scale feature pyramid learning method. The core innovation lies in the design of a temporal-semantic guided multi-scale local feature extractor (MLFE). This module can not only handle spatial multi-scale features but also incorporate temporal dynamics and textual semantic guidance to achieve adaptive scale selection. Based on this, we have constructed a complete recognition framework, including an improved cross-frame communication mechanism and a multi-granularity dynamic cue generation module. Experiments on benchmark datasets such as Kinetics-400 and Kinetics-600 show that our method significantly outperforms existing methods in scenarios requiring fine-grained local feature recognition, especially in topics such as "laboratory," "medical surgery," and "cooking process.
    Keyword: Video topic recognition, multi-scale local perception, feature pyramid network, semantic-visual alignment, local feature enhancement
    DOI: 10.65286/icic.v22i1.27332
    Cite

  • CMTSiMBA-light: Industrial Defect Classification and Detection Based on Deep Learning, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Wencheng Ding, Shan Chang, Tianhui Cai, Hongya Wang
    Abstract: Surface quality inspection of industrial products is a crucial step in ensuring production line efficiency and product quality. In response to the problem of insufficient fusion of local and global features in image classification algorithms in complex industrial scenarios and the difficulty of a single architecture to balance texture details and long-range dependencies. This paper focuses on deep learning-based industrial defect classification and detection, and proposes and optimizes the relevant algorithm CMTSiMBA-light. By incorporating the SENet convolution module into the CMT Stem and replacing traditional convolution with the PConv convolution module in the SiMBA Stem, the model's accuracy and efficiency have been significantly improved. Experiments show that the model improves accuracy by 0.7% and 1.4% respectively on the NEU-CLS and FSC-20 datasets compared to CMT, reduces parameters by 10.8%, lowers FLOPs by 5%, and increases throughput by 5.2%
    Keyword: Industrial Defect Classification ï¼›Industrial Detection ï¼› Deep Learning
    DOI: 10.65286/icic.v22i1.98902
    Cite

  • Zero-Shot Multi-Reference Personalization via MLLMs-Guided Layout Planning, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Junhao Feng, Xinghang Xu, Zihao Zhang, Zhonghua Wan
    Abstract: Though zero-shot adapters excel in image personalization, they often encounter significant challenges in multi-reference personalized generation, specifically failing to precisely adhere to the spatial layouts described in text prompts and suffering from feature leakage between reference images. To address these two challenges, we propose RIG (Regional Image-prompt Generation), a novel training-free framework. For the first challenge, leveraging Multimodal Large Language Models (MLLMs), we introduce a layout planning binder. Leveraging Chain-of-Thought (CoT) reasoning, this module infers and generates precise global layouts from text prompts, while simultaneously binding reference images to their corresponding regions. For the second, we introduce a satially decoupled diffusion mechanism that isolates feature streams during attention computation. By injecting reference features exclusively into designated regions, this mechanism effectively prevents feature interference between reference images. Extensive experiments demonstrate that RIG significantly outperforms state-of-the-art adapter methods in terms of both personalization fidelity and text-layout alignment.
    Keyword: Image Personalization, Diffusion Models, Multimodal LLMs, Zero-Shot Learning.
    DOI: 10.65286/icic.v22i2.15920
    Cite

  • FLMLog: A Federated LLM Framework for Unified Online Log Anomaly Detection, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Shuai Xu,Guanghong Yang,Zepeng Wen
    Abstract: Anomaly detection plays a pivotal role in ensuring the reliability of modern large-scale distributed systems. However, traditional log anomaly detection systems are centralized, which poses the risk of privacy leakage during data transmission. Previous research mainly focuses on single-domain logs,requiring domain-specific models and retraining, which limits flexibility and scalability. Significant advancements have been made by Large Language Models (LLMs) in the domains of natural language understanding and automated content creation. However,they still face persistent problems, including substantial computational costs and inadequate availability of training data. The combination of Federated Learning (FL) and LLMs (federated LLMs) offers a solution by leveraging distributed data while protecting privacy, which positions it as an ideal choice for sensitive domains. In this paper, we propose a unified online log anomaly detection framework, FLMLog, which is based on federated Learning and large language model. To enhance the operational efficiency, the FLMLog framework adopts a prefix-aware in-context learning (ICL) refinement strategy. This strategy is specifically designed to refine both the selection of in-context examples and the per-mutation order of these examples, thereby achieving an improvement in prefix caching efficiency. Our experiments demonstrate that the FLMLog framework is rigorously evaluated on five publicly available production log datasets, and the results show that it achieves superior comprehen-sive performance, outperforming state-of-the-art methods in F1-score on four out of five datasets with remarkable improvements.
    Keyword: Anomaly Detection,Federated Learning,LLMs.
    DOI: 10.65286/icic.v22i2.96380
    Cite

  • CoRe-Diffusion: Bridging the Generative-Discriminative Gap in Dataset Distillation via Manifold-Aligned Contrastive Guidance, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Heng Shu, Linjuan Cheng, Qiming Yang, Yuquan Wu
    Abstract: Dataset Distillation aims to condense large-scale datasets into a tiny but highly informative subset, enabling efficient training while achieving performance comparable to the full dataset. To overcome the scalability bottlenecks of traditional optimization, Generative Dataset Distillation (GDD) leverages diffusion priors to reformulate the dataset condensation process as an efficient generative sampling paradigm. However, existing GDD paradigms primarily optimize intra-class likelihood while largely overlooking explicit inter-class separation, which may cause generated features to concentrate near decision boundaries. Furthermore, relying on manually designed guidance targets to steer synthesis often pushes the diffusion trajectory off the natural data manifold, inducing severe structural artifacts. To address these issues, CoRe-Diffusion is proposed as a unified framework. Specifically, it introduces Contrastive Negative Guidance, which utilizes inter-class mode centers—computed but largely unused during synthesis by existing methods—to exert zero-overhead repulsive gradients against hard negatives. To preserve synthesis fidelity, Manifold-Aligned Real-Anchor Discovery strictly constrains the guidance targets to the exact latent representations of real images, preventing the trajectory from drifting. This intrinsically preserves complex semantic structures, bypassing the prohibitive computational overhead of relying on external generative priors or auxiliary modules. Alongside an Annealed Sampling Schedule for smooth trajectory evolution, CoRe-Diffusion achieves state-of-the-art performance, yielding a 3.8\% absolute accuracy improvement on ImageNet-1K while introducing negligible computational overhead.
    Keyword: Generative Dataset Distillation, Contrastive Negative Guidance, Manifold Alignment, Diffusion Models
    DOI: 10.65286/icic.v22i1.43405
    Cite

  • Fusing Relation Graph and Mutual Information for Inductive Link Prediction in Knowledge Graphs , ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Hongbo Liu, Jicang Lu, Yaohui Hao, Taojie Zhu
    Abstract: Inductive link prediction in knowledge graphs refers to the task of inferring known relations between entities unseen during training. Most existing approach-es are limited to predicting only known relations and struggle to generalize to un-seen relations, which restricts their utility in dynamic settings. To address this challenge, we propose a novel inductive link prediction approach named RGIILP. Specifically, we construct a relation graph from the source knowledge graph and design a neural network model that enables interactive feature propagation be-tween entities and relations. Furthermore, we introduce the mutual information maximization mechanism between global and local representations to capture the global structural information of the graph. Experiments on several benchmark da-tasets demonstrate that RGIILP outperforms existing state-of-the-art methods for inductive link prediction task.
    Keyword: Inductive Link Prediction, Knowledge Graphs, Relation Graph, Mutual Infor-mation.
    DOI: 10.65286/icic.v22i2.68154
    Cite

  • Causal Inference-Based Network Anomaly Detection for Internet of Vehicles, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Ming Dai, Heng Gao, Zengri Zeng, Aimei Kang, Xinming Wang
    Abstract: 5G-V2X technology and connected and autonomous vehicles are deeply in-tegrated. The Internet of Vehicles (IoV) has become the core support of an intelligent transportation system. Its heterogeneous architecture and high-dynamic characteristics bring serious cybersecurity risks. Traditional anoma-ly detection methods rely on correlation reasoning. They have high false pos-itive rates. They are not easy to interpret. They can’t adapt to combined at-tacks. They also lack real-time performance. This paper puts forward a causal inference-based anomaly detection algorithm (IoV-NDCML). It is specially designed for the IoV scenario. The algorithm reconstructs a multi-dimensional causal interpretable feature set. It designs a dynamic causal in-tervention screening method. The method integrates scenario weights. It builds a lightweight hierarchical SCM model. It improves the counterfactual diagnosis method. The algorithm realizes accurate anomaly detection. It also achieves causal attribution.Experiments show that the proposed algorithm achieves 99.2% accuracy in single-attack scenarios and 97.5% accuracy in combined-attack scenarios, with an end-to-end latency below 50 ms on edge devices. Its comprehensive performance outperforms comparative algo-rithms, providing technical support for security defense in the Internet of Vehicles.
    Keyword: causal inference; Internet of Vehicles (IoV); V2X communication; network anomaly detection; causal intervention; counterfactual diagnosis; structural causal model (SCM)
    DOI: 10.65286/icic.v22i2.27697
    Cite

  • Cross-Modal Dynamic Aggregation with Adaptive Relevance Modulation Fusion Network for Remote Sensing Visual Question Answering, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Zihua Zuo, Chao LI
    Abstract: Remote Sensing Visual Question Answering (RS VQA) task aims to provide accurate answers to questions about RS images. However, the semantic gap between low-level visual features and high-level semantics complicates the understanding of complex questions. Moreover, the lack of dynamic modulation mechanisms for integrating visual and textual features impedes the balanced interpretation of image content and textual semantics. To this end, we propose the Cross-Modal Dynamic Aggregation with Adaptive Relevance Modulation Fusion Network for RS VQA (CDAR-Net). Specifically, we propose a Cross-Modal Dynamic Aggregation Module (CMDA) that employs a multi-level attention mechanism to iteratively fuse text-guided visual features, enabling a progressive transition and dynamic integration from low-level visual features to high-level semantic information. We further introduce an Adaptive Correlation Modulation Fusion Module (ACMF) that dynamically adjusts visual and textual feature weights based on questions context, enhancing the representation of relevant information. Experimental results demonstrate that CDAR-Net outperforms existing state-of-the-art methods on three RS VQA datasets.
    Keyword: remote sensing, visual question answering (VQA), cross-modal fusion, dynamic aggregation
    DOI: 10.65286/icic.v22i1.74211
    Cite

  • Cavity-Aware Deep Reinforcement Learning for 3D Bin Packing, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Wenjiao Xie, Sirui Wang, Yunqing Rao, Qianhang Lyu
    Abstract: The three-dimensional bin packing problem (3D-BPP) is a classical optimization problem in logistics and transportation. Existing methods often rely on handcrafted heuristics or simplified state representations, which limits their ability to capture complex spatial relationships in packing environments. To ad-dress this issue, this paper proposes Cavity-Map-based Deep Reinforcement Learning (CMDRL) for the 3D-BPP. The proposed method introduces a cavity-map representation to model the geometric structure of free space and designs an enhanced spatiotemporal attention mechanism to jointly capture packing se-quence dependencies and spatial layout information. A Transformer-based policy network trained with the Soft Actor-Critic (SAC) algorithm is developed to gen-erate efficient packing decisions. Experimental results show that the proposed method consistently outperforms several state-of-the-art DRL baselines in terms of gap ratio. Ablation studies further demonstrate the effectiveness of the cavity map and the spatiotemporal attention mechanism.
    Keyword: 3D-BPP, deep reinforcement learning, cavity map, spatiotemporal attention mechanism.
    DOI: 10.65286/icic.v22i2.10353
    Cite

  • DEA-TLS: A Divide-Extract-and-Aggregate Framework for Costless and Accurate Timeline Summarization, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Ming Qiao, Yinlong Xiao, Minghao Hu, Jianyong Duan, Xin Li, Zhunchen Luo, Shuai Lei, Yunbo Cao, Guotong Geng
    Abstract: Timeline summarization aims to identify key events from multi-source texts and organize them into a coherent timeline of event evolution in chronological order. Existing methods generally adopt a two-stage document-level framework, which first extracts timeline information from individual documents and then aggregates the extracted results into a final timeline summary. However, such methods often suffer from high computational cost in practical applications. On the one hand, closed-source models incur substantial inference expenses; on the other hand, locally deployed open-source models still require considerable computational resources when processing long documents. To address this issue, we propose DEA-TLS, a Divide-Extract-and-Aggregate framework for Timeline Summarization, which transforms conventional document-level processing into finer-grained paragraph-level processing, thereby reducing the overall computational burden. Nevertheless, this framework also brings new challenges, including the introduction of irrelevant information, inaccurate event extraction, and redundancy during aggregation. To tackle these issues, we further design three modules: Progressive Divider, Reflective Extractor, and Hierarchical Aggregator, which are responsible for selecting high-value paragraphs, improving extraction accuracy, and reducing semantic redundancy, respectively. Experiments on the Open-TLS dataset demonstrate that our method outperforms existing baselines across multiple key metrics, providing an effective solution for achieving low-cost and high-accuracy timeline summarization.
    Keyword: Timeline Summarization · Large Language Models · Retrieval-Augmented Generation · Self-Reflection
    DOI: 10.65286/icic.v22i1.23884
    Cite

  • SECURE: Stable Early Collision Understanding via Robust Embeddings in Autonomous Driving, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Wenjing Wang, Wenxuan Wang, Songning Lai
    Abstract: While deep learning has significantly advanced accident anticipation, the robustness of these safety-critical systems against real-world perturbations remains a major challenge. We reveal that state-of-the-art models like CRASH, despite their high performance, exhibit significant instability in predictions and latent representations when faced with minor input perturbations, posing serious reliability risks. To address this, we introduce SECURE – Stable Early Collision Understanding Robust Embeddings, a framework that formally defines and enforces model robustness. SECURE is founded on four key attributes: consistency and stability in both prediction space and latent feature space. We propose a principled training methodology that fine-tunes a baseline model using a multi-objective loss, which minimizes divergence from a reference model and penalizes sensitivity to adversarial perturbations. Experiments on DAD and CCD datasets demonstrate that our approach not only significantly enhances robustness against various perturbations but also improves performance on clean data, achieving new state-of-the-art results.
    Keyword: Accident Anticipation, Adversarial Robustness, Autonomous Driving, Robust Embeddings, Spatio-temporal Learning
    DOI: 10.65286/icic.v22i1.67330
    Cite

  • Two-stage Monocular 6D Pose Estimation for Small Cubic Objects, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Xinmiao DU, Zhuoyu JIANG, Xihong WU
    Abstract: Monocular 6D object pose estimation is an important per- ception problem for robotic manipulation. For high-precision operations such as block grasping, alignment, and placement, the perception mod- ule must provide sufficiently accurate and stable pose estimates rather than coarse object localization alone. This requirement is particularly challenging for small cubic objects due to weak texture, limited visual cues, strong rotational symmetry, and sensitivity to ROI quality. In this paper, we study monocular 6D pose estimation of small cubic objects from a single RGB image and propose a two-stage manipulation- oriented framework. In the first stage, a detection-guided ROI-based re- gression model is used to estimate object pose under symmetry-aware supervision. In the second stage, we further explore a MuJoCo-based re- rendering strategy to construct consistency-enhanced training samples for refinement, aiming to improve adaptation to manipulation-relevant visual conditions. Experiments on a unified MuJoCo-generated dataset show that, on the main single-block benchmark, the proposed method achieves the strongest overall balance in ADD-S, translation accuracy, rotation stability, and task-oriented usability metrics. Preliminary results in multi-block scenes further suggest favorable generalization to more complex visual condi- tions.
    Keyword: Monocular 6D Pose Estimation · High-Precision Robotic Manipulation · Physics-based Re-rendering Consistency
    DOI: 10.65286/icic.v22i1.90872
    Cite

  • An Image Semantic Communication Method Based on SwinJSCC with Explicit MIMO and Hybrid Channel Modeling, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Jiale Yu, Yuling Zhang, Wenwei He, Yujuan Sun, Tao Yao
    Abstract: To address the limited channel coverage of the original SwinJSCC in com-plex wireless scenarios, the lack of an explicit MIMO transmission mecha-nism, and the heavy training burden caused by separately training independ-ent models for different channel modes, this paper proposes a SwinJSCC-based image semantic communication method with explicit MIMO and hy-brid channel modeling. Built upon SwinJSCC, the proposed method pre-serves the original SNR adaptation and rate adaptation capabilities while uni-fying AWGN, Rayleigh fading, MIMO transmit diversity, MIMO spatial mul-tiplexing, and large-scale path-loss extended channels into a single end-to-end training and testing framework. To avoid training a large number of sep-arate models for different channel types and MIMO structures, a mode-oriented conditional modulation mechanism is introduced, enabling the en-coder and decoder to adjust features according to the current transmission mode while sharing the same backbone parameters. In channel modeling, the proposed method does not approximate multi-antenna gain by a simple equivalent SNR. Instead, 2×1, 4×1, and 8×1 are modeled as explicit transmit diversity, while 2×2 and 4×2 are modeled as spatial multiplexing, further combined with large-scale path-loss factors. This work provides a scalable single-model unified solution for image semantic transmission in complex wireless environments.
    Keyword: SwinJSCC, semantic communication, explicit MIMO, hybrid channel model-ing, mode adaptation.
    DOI: 10.65286/icic.v22i2.78335
    Cite

  • CIAF: Cross-modal Inconsistency Aware Framework for Multimodal Fact-Checking, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Zidong Yi
    Abstract: In the multimodal information era, fact-checking faces growing challenges as misinformation becomes increasingly complex. Content that appears plausible in a single modality may reveal subtle inconsistencies across modalities, which often overlooked by traditional methods. Existing approaches mainly fuse multimodal features but rarely explicitly model fine-grained cross-modal inconsistencies, limiting both accuracy and interpretability. To address this issue, we propose a Cross-modal Inconsistency Aware Framework (CIAF) that leverages multimodal sentiment cues to identify and localize implicit discrepancies between textual claims and corresponding images. Specifically, CIAF first performs a joint analysis of sentiment elements across textual and visual modalities. It then introduces a novel relational discrepancy attention (RDA) mechanism to model semantic and emotional interactions between modalities, dynamically weighting potential inconsistency signals. Furthermore, we design task-oriented prompt templates to guide the model’s reasoning process toward cross-modal consistency assessment, thereby enhancing both precision and robustness in multimodal fact-checking. Experiments on two challenging multimodal fact-checking benchmarks show that our approach delivers superior performance and interpretability, underscoring the importance of modeling fine-grained cross-modal inconsistencies for robust fact verification.
    Keyword: Multimodal fact-checking, Multimodal consistency modeling
    DOI: 10.65286/icic.v22i1.95226
    Cite

  • Robust Adaptive Neural Network-Based Backstepping Tracking for Second-Order Euler-Lagrange Systems with Unknown Parameters, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Shen Zhang, Xiaozheng Jin, Na Li
    Abstract: This paper proposes a robust adaptive tracking control scheme for a class of second-order Euler–Lagrange systems with completely unknown parameters and nonlinear dynamics. System uncertainties, including unmodeled dynamics, parametric variations, and external disturbances, are formulated as a time-varying lumped perturbation. Radial Basis Function Neural Networks (RBFNNs) approximate the unknown state-dependent nonlinear component within the perturbation bound, while adaptive laws estimate the unknown bounding constants of input-dependent terms and disturbances. By integrating backstepping with a $\sigma$-modification mechanism, a continuous adaptive control law is developed that eliminates chattering typically caused by discontinuous robust terms. Lyapunov analysis proves that all closed-loop signals are uniformly ultimately bounded, achieving asymptotic trajectory tracking with smooth control inputs. Simulations on an underactuated Unmanned Surface Vehicle (USV) under complete model uncertainty and environmental disturbances validate the effectiveness and superiority of the proposed method.
    Keyword: Second-order Euler-Lagrange Systems, Trajectory Tracking, Radial Basis Function Neural Networks, Backstepping Control, Adaptive Control, Chattering Suppression
    DOI: 10.65286/icic.v22i2.19895
    Cite

  • FedLSA: Cosine Triggered Reparameterization Augmented by Output Boundary Calibration for Private Federated Low Rank Adaptation, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Zhencheng Fan, Tingyi Chen, Da Huang, Songzhu Mei, Qihang Zhang
    Abstract: Low-Rank Adaptation (LoRA) has become the standard paradigm for Parameter-Efficient Fine-Tuning (PEFT) within federated learning (FL) under privacy preservation. However, while reparameterization strategies in recent orthogonalization baselines like FedSVD enhance matrix expressiveness, they often overlook the actual dynamic requirements for subspace updates during train-ing. To balance computational efficiency and model expressiveness under dif-ferential privacy (DP) constraints, we propose FedLSA, a Federated Low-rank Subspace Adaptive update framework. Specifically, FedLSA utilizes the cosine similarity of updates to the low rank coefficient matrix B as an economical proxy for subspace drift, which minimizes overhead on the server through dynamic Singular Value Decomposition (SVD) allocation. To complement this mechanism triggered by events, we establish an Output Boundary Calibration (OBC) within the logits space to ensure the optimization trajectory remains robust even when matrix basis updates are suspended. Experimental results on five representative Natural Language Understanding (NLU) benchmarks demonstrate that FedLSA consistently outperforms all evaluated baselines in terms of accuracy. Notably, compared to the strongest state-of-the-art (SOTA) baseline, FedLSA establishes a better Pareto front for efficiency and accuracy.
    Keyword: Federated Low-Rank Adaptation, Cosine Triggered Reparameterization, Dif-ferential Privacy.
    DOI: 10.65286/icic.v22i2.59894
    Cite

  • Multi-Objective Genetic Algorithm for Heston Model Calibration: A Framework Integrating Pricing and Implied Volatility Criteria, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Yan Fang, Yiwei Hua, Julius Wu
    Abstract: Accurate calibration of the Heston model is essential for reliable option pricing, yet it remains a challenging task due to the nonlinear structure of the model and the presence of multiple conflicting objectives. Most existing approaches rely on single-objective formulations that focus solely on price fitting, which may lead to suboptimal overall performance. To address this issue, this paper proposes a mul-ti-objective genetic algorithm framework for Heston model calibration, in which option price errors and implied volatility errors are jointly optimized. The result-ing bi-objective optimization problem is solved using two representative evolu-tionary algorithms, namely NSGA-II and NRGA. This formulation enables an explicit characterization of the trade-off between pricing accuracy and volatility fitting. Extensive experiments are conducted using real option data from SSE 50ETF and S&P 500 markets, as well as Monte Carlo simulations. The results show that both NSGA-II and NRGA produce consistent and comparable pa-rameter estimates. Compared with conventional single-objective methods, the proposed framework leads to improved pricing accuracy and more stable calibra-tion results. Simulation studies further confirm the robustness of the proposed approach under different noise conditions.
    Keyword: Heston Model; Multi-objective Optimization; Genetic Algorithm; NSGA-II; NRGA; Option Pricing; Implied Volatility; Evolutionary Computation.
    DOI: 10.65286/icic.v22i2.73212
    Cite

  • KG-MMIL: A Knowledge-Guided Multi-modal Multi-Instance Learning Framework for Interpretable Renal Tumor Ultrasound Diagnosis, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Wei Ni, Tao Yang, Kunlei Tan
    Abstract: Current methods for diagnosing renal tumors via ultrasound primarily rely on learning image features; however, the model’s decision-making process lacks interpretability, exhibiting typical “black-box” characteristics that limit its clinical application. To address this issue, this paper proposes a knowledge-guided multimodal multi-instance learning fusion model (KG-MMIL) to enhance the model’s discriminative performance and interpretability.First, a knowledge-guided multimodal feature enhancement method is designed. By constructing a knowledge-guided attention mechanism, medical text knowledge is converted into prior weights and fused with data-driven attention weights. This guides the model to focus on key regions and features of diagnostic significance, thereby improving the medical consistency and interpretability of feature representations.Second, to address the challenges of fusion caused by multimodal feature heterogeneity, we propose a multi-stream residual parallel fusion method. Through a multi-path residual structure, this method enables deep interaction and information compensation across modalities, effectively mitigating information loss and gradient propagation issues in traditional fusion methods, thereby enhancing overall feature expressiveness.Experimental results on a renal tumor ultrasound dataset demonstrate that the proposed KG-MMIL model achieves a precision of 91.9% and an AUC of 0.961. Compared to the second-best multimodal method, precision is improved by 1.7 percentage points while maintaining competitive performance in terms of AUC, validating the effectiveness and superiority of this approach.
    Keyword: Multiple Instance Learning, Multimodal Fusion, Renal Tumor.
    DOI: 10.65286/icic.v22i2.13558
    Cite

  • ForgerySpotter: Pinpointing Tampered Regions with Multi-scale Evidence and Confidence-Guided Refinement , ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Xiaolong Cheng, Juanjuan Luo, MingChen Han
    Abstract: Image manipulation detection (IMD) is crucial for maintaining the integrity of digital media, as forged images can be used to spread false information and erode public trust. IMD faces two persistent challenges: i) limited generalization to diverse real-world post-processing operations, and ii) imprecise localization accuracy yielding coarse or incomplete tampering regions. Moreover, existing methods often lack interpretability due to the absence of reliable confidence estimation. The primary research often prioritizes feature or architectural improvements while neglecting the integration of detection reliability with localization refinement. In this paper, we propose a unified framework that incorporates multi-dimensional feature extraction, multi-scale feature fusion, and a confidence-guided refine mechanism. Our method captures tampering traces across types and scales adaptively, while the confidence-guided mechanism refines localization maps and estimates pixel-wise reliability. Extensive experiments on multiple datasets demonstrate that the proposed approach achieves state-of-the-art performance and shows strong generalization, validating its effectiveness and practicality.
    Keyword: Image Manipulation Detection.
    DOI: 10.65286/icic.v22i1.51163
    Cite

  • LipHS : A Lightweight WiFi-enabled Human Sensing For Multi-Class Scenarios, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Chunhao Xue, Lixing Wang
    Abstract: WiFi-based human sensing technology utilizing Channel State Information (CSI) has garnered significant attention due to its reduced privacy concerns and the widespread availability of existing infrastructure, demonstrating broad development prospects in the field of intelligent computing. Deployment on edge devices represents its most prevalent application scenario. However, the high-complexity algorithms commonly employed to enhance sensing accuracy face substantial challenges when deployed on devices with limited computational resources. Furthermore, most existing studies conduct experiments only on datasets with a small number of categories. Although these approaches achieve high accuracy, they fail to meet practical sensing requirements. Consequently, developing high-accuracy, low-complexity, and practical WiFi-based human sensing systems remains considerably challenging. To construct an efficient and lightweight feature extraction network, we presents LipHS, a lightweight feature extraction framework capable of simultaneously capturing multi-level information from CSI signals. To further reduce the number of model parameters, we employ a channel pruning method based on Layer-Adaptive Magnitude-based Pruning (LAMP) scores. LipHS achieves model lightweighting while maintaining robust feature extraction capabilities. Experimental results demonstrate that the proposed LipHS method outperforms other baseline algorithms in sensing performance on complex multi-class gesture datasets.
    Keyword: WIFI Sensing· Channel State Information ·Human Activity Recognition· Lightweight
    DOI: 10.65286/icic.v22i1.94871
    Cite

  • A tourist attraction recommendation algorithm based on the integration of user life cycle and adaptive time, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Xiuyuan Liu, Chengxi Li, Yuming Song, Xuechen Zhao, Zhipeng Li
    Abstract: This paper proposes a personalized recommendation algorithm that combines the weights of user lifetime and adaptive time in tourist attraction recommendations, aiming at the problems of dynamic changes in user interests, significant influence of time context, and data sparsity in collaborative filtering. First, users are divided into four life cycle stages - the exploration stage, the growth stage, the maturity stage, and the senior stage - based on the number of attractions they visit; Secondly, build a month-spot heat matrix and use time series smoothing techniques to enhance the stability of monthly heat; Then, on top of collaborative filtering, introduce a user lifecycle stage to adaptively fuse collaborative filtering scores with seasonal heat scores; Finally, experiments are conducted on real tourism datasets and compared with multiple baseline algorithms. The experiments demonstrated that the LST-Rec algorithm proposed in this paper outperformed traditional collaborative filtering algorithms in terms of accuracy, recall, coverage, and popularity metrics, with an accuracy of 39.19% and a recall of 56.60%. The ablation experiment verified the effectiveness of the adaptive fusion module, the method proposed in this study makes full use of the temporal context and user behavior patterns, effectively improving the accuracy and personalization level of tourist attraction recommendations, but provides a direction for subsequent research.
    Keyword: Tourism Recommendation; User lifecycle; Time context; Adaptive weighting; Collaborative filtering
    DOI: 10.65286/icic.v22i2.62501
    Cite

  • Dual Associations Semantic Enhancement for Image-Text Matching, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Xinlin Zhao, Tao Yao, Yafei Bu, Li Liu, Yuling Zhang, Linliang Zhang
    Abstract: Image-text matching faces the significant challenge in effectively mitigating the large visual-semantic discrepancy among different modalities. Existing studies mainly address this by projecting multi-modal data into a common subspace for measuring their semantic similarities. However, most of them often overlook the inherent asymmetry for directional associations, i.e., the differences between im-age-to-text and text-to-image affinities, which often results in inaccurate retrieval results. To tackle the challenge, we propose a Dual Associations Semantic En-hancement (DASE) model to capture bidirectional image-text semantic associa-tions. Specifically, we first build a two-layer GCN fusion network to construct and mine the semantic associations for each modality. And then, due to the inher-ent asymmetry of directional associations, a Dual Associations Alignment Mod-ule (DAAM) is designed to capture the dual associations between visual and tex-tual modalities, enabling comprehensive cross-modal fine-grained interaction. Fi-nally, global alignment is incorporated with the local alignment to achieve full semantic matching across heterogeneous modalities in a unified embedding space. Experimental results on two publicly available datasets demonstrate that the pro-posed DASE model achieves significant performance improvements in image-text matching tasks compared to baseline methods, validating its effectiveness and superiority.
    Keyword: Image-text Matching, Dual Associations, Cross-modal Retrieval
    DOI: 10.65286/icic.v22i2.42693
    Cite

  • Diagnosing Video Foundation Models for Single-Signer RGB-Only Auslan-Daily News Translation, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Maoyang Li, Prashan Premaratne, Peter Vial
    Abstract: Sign language recognition (SLR) through computer vision could be seen as one of the most challenging tasks associated with AI and computer vision. Many attempts in the world have still not shown reasonable progress in this approach. Recent advancements in video foundation models have made it possible to improve sign language translation (SLT) performance by adopting stronger visual encoders, especially in low-resource settings. In this paper, we examine this assumption on the Auslan-Daily News split under a single-signer, RGB-only setting. We systematically compare three representative pipeline families, namely I3D, Image Swin-T, and VideoMAE, under a unified evaluation protocol. VideoMAE achieves the best performance but still remains clearly lower than the official benchmark. To better understand this gap, we further analyze model outputs, test sensitivity to temporal order and visual content, and compare several VideoMAE-based variants. Our results show that the limitation cannot be explained by backbone choice alone. The main difficulty lies in how long video sequences are compressed and then carried into the translation stage under limited supervision. And future progress in low-resource Australian Sign Language (Auslan) translation will depend not only on stronger visual encoders, but also on better temporal modeling and more reliable visual-to-text transfer.
    Keyword: Sign Language Translation, Video Foundation Models, Low-Resource Auslan Translation.
    DOI: 10.65286/icic.v22i1.81935
    Cite

  • BSLANet: a boundary and spatial localization-aware synergistic enhancement network for skin lesion segmentation, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Chunbao Lu, Guanxi Liu, Hongyan Zhao, Weiman Xiao, Tao Zhang
    Abstract: Abstract. Accurate segmentation of skin lesions, a vital step in early skin cancer detection, is of utmost importance in improving patient survival rates. Although many deep learning-based methods have significantly progressed, challenges such as fuzzy lesion boundaries and complex spatial distribution remain. To resolve the above issues, this study proposes BSLANet, a novel boundary and spatial localization-aware synergistic enhancement network for skin lesion segmentation. To mitigate high-frequency information loss during downsampling, we introduce the Wavelet Attention Guided Downsampling Module (WAGDM). Furthermore, to enhance spatial understanding of complex lesions, we propose the Mixed Pooling Spatial Perception Attention (MPSPA). Lastly, to delineate finer-grained lesion boundaries and recalibrate lesion positions, we use the Differential Contrast Collaborative Feature Fusion Module (DCCFM) to improve semantic interaction in cross-layer feature fusion. Our experiments on four public skin lesion datasets—ISIC2016, ISIC2017, ISIC2018, and PH2—demonstrate that BSLANet surpasses current methods, achieving superior performance in skin lesion segmentation.
    Keyword: Keywords: Skin Lesion Segmentation, Downsampling, Feature Fusion.
    DOI: 10.65286/icic.v22i1.35683
    Cite

  • FR-YOLO: A Focus-and-Reconstruct Mechanism for Lightweight Small Object Detection in Drone Imagery, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Hongwei Liu, Rui Zhu, Shisheng Tang, Liwei Sun, Haohan Ding
    Abstract: Automatic detection of small objects in UAV imagery is a challenging problem that is of great interest in aerial surveillance and intelligent transport. The detection model must cope well with feature degradation of tiny targets, drastic scale variations, and strict constraints on onboard com-putational resources. This paper proposes a lightweight attention-based network, called FR-YOLO, to address the "focus" and "reconstruct" chal-lenges in small object detection. We introduce two novel components: the Local Feature Enhancement (LFE) module to precisely suppress back-ground noise via spatial attention, and the Content-aware Feature Reassem-bly (CFR) module to rectify spatial feature misalignment and recover fine-grained details during upsampling. The proposed modules are seamlessly integrated into the YOLO11n backbone. Extensive experiments on the au-thoritative VisDrone2019 benchmark indicate that FR-YOLO outperforms the YOLO11n baseline by a significant margin of 1.5% and surpasses the widely adopted YOLOv8n by 0.7% (achieving 36.4% mAP), all while maintaining a 16% smaller model size (2.53M parameters) compared to YOLOv8n. We have made the code available to the public at: https://github.com/Liu999hongwei/FR-YOLO.
    Keyword: UAV object detection, spatial attention, feature alignment, YOLO, light-weight network.
    DOI: 10.65286/icic.v22i1.84412
    Cite

  • Collaborative End-Edge-Cloud Framework for Multi-modal Dynamic 3D Reconstruction on Edge Devices, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Haikang Gao, Jiawei Dong, Yuan Zhao, Zhenyu Chen, GaoLei Yi, Yu Lu, Meng Li
    Abstract: With the development of embodied intelligence, high-fidelity and real-time environmental perception has become a core challenge. Although the 3D Gaussian Spatter technology has achieved breakthroughs in rendering quality, it still faces issues such as mismatched computing power and "ghost" artifacts caused by dynamic objects on resource-constrained edge devices. This paper proposes a new perception framework: Firstly, a end-edge-cloud collaborative architecture is constructed, and computing-intensive global optimization is offloaded to the edge side or the cloud through dynamic task allocation; Secondly, the "learning-forgotten" mechanism is introduced, combining lightweight mask MLP and se-mantic assistance, to online identify and eliminate dynamic objects through Gaussian lifecycle management; Finally, through the multimodal tight coupling of laser radar, inertial measurement unit, and visual data, robust initialization is achieved. The test results in real scenarios show that this framework, while maintaining real-time rendering at the edge end, significantly improves reconstruction consistency and reduces trajectory errors.
    Keyword: 3D Gaussian Splatting · End-Edge-Cloud Collaboration · Dynamic Scene · Multi-modal Fusion · Learn-Forget Mecha-nism · Embodied AI.
    DOI: 10.65286/icic.v22i2.29781
    Cite

  • SLRNet: Super Lightweight Residual Network for Real-Time Image Dehazing, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Guanheng Qu, Fan Jiang, Jiangming Liu
    Abstract: Image dehazing aims to generate the haze-free images from the hazy observation images. While recent deep learning approaches achieve impressive restoration quality, they suffer from excessive computational complexity and model size, hindering practical applications for real-world deployment on resource-constrained edge devices. To address the limitation, lightweight models are proposed to this end but compromise on dehazing performance. To bridge this gap, we propose SLRNet, Super Lightweight Residual Network, a high efficient-yet-effective end-to-end dehazing architecture. SLRNet integrates a novel Adaptive Feature Unit that automatically adjusts channel-wise features through a lightweight gating mechanism, coupled with compact residual blocks to preserve critical structural information. Unlike standard channel attention mechanisms that discard spatial information, our AFU employs an asymmetric split strategy to simultaneously preserve local texture details and capture global haze density. Our design emphasizes minimal parameter count and low latency without sacrificing perceptual quality. Experiments are carried out across standard benchmarks, showing that our proposed SLRNet demonstrates remarkable performance by achieving state-of-the-art efficiency-accuracy trade-offs compared to existing works, while maintaining robust generalization to real-world haze despite the synthetic-to-real domain gap. The codes are released in https://anonymous.4open.science/r/SLRNet.
    Keyword: Image dehazing, image restoration, lightweight network, real-time inference, residual learning, adaptive channel attention.
    DOI: 10.65286/icic.v22i1.71074
    Cite

  • TMI-VFL: Secure Vertical Federated Learning via Threshold Multi-Identity Homomorphic Encryption, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Yuqing Song
    Abstract: Vertical Federated Learning (VFL) enables multiple participants to collaboratively train models using vertically partitioned data. However, during the actual training process, participants must exchange intermediate model representations(e.g., embeddings), which creates a potential attack surface for privacy leakage. Recent studies have shown that the URVFL attack achieves precise and covert data reconstruction by constructing malicious gradients and training a decoder using label information, posing a serious threat to the privacy security of vertical federated learning systems.To address this issue, we propose TMI-VFL, a secure training framework based on Threshold Multi-Identity Fully Homomorphic Encryption.This method establishes a ciphertext computation mechanism that ensures embedding vectors, gradients, and intermediate activation values are all processed in encrypted form, while the threshold decryption scheme prevents any single participant from recovering sensitive information. Experimental results show that under URVFL attacks, the proposed method increases reconstruction error by more than 10-fold, significantly reducing the effectiveness of the attack. Meanwhile, model accuracy decreases by less than 1% and remains close to baseline levels. These results indicate that TMI-VFL achieves an effective trade-off between privacy protection and model utility, providing a practical solution for secure VFL.
    Keyword: Vertical Federated Learning,Data Reconstruction Attack,Fully Homomorphic Encryption,Privacy Protection
    DOI: 10.65286/icic.v22i2.15822
    Cite

  • HPPE: HeatMap Positional Embedding for Background Noise Suppression in Object Detection, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Yuxin Liu, Hongyun Liu, Yusen Wu, YangChen Zeng
    Abstract: Despite the rapid proliferation of Transformer-based architectures in small object detection, mitigating background noise and improving query quality remain critical bottlenecks. To address these challenges, this study presents HeatMap Positional Embedding (HPPE), an adaptive optimization framework that intrinsically couples positional encoding with semantic detection priors through heatmap-guided masking. Furthermore, we develop a novel visualization scheme for HPPE, providing an intuitive perspective on feature embeddings to facilitate hyperparameter tuning. Building upon this mechanism, two specialized modules are proposed: the Multi-Scale ObjectBox-Heatmap Fusion Encoder (MOFE) and the HeatMap Induced High-Quality Queries for Decoder (HIQQ). These components are tailored to generate semantically rich queries while actively suppressing irrelevant background interference. By incorporating heatmap positional embeddings alongside standard baseline feature extractors like Linear-Snake Conv (LSConv), our approach effectively handles the massive diversity of small object categories and drastically minimizes the required number of decoder multi-head layers. Extensive evaluations demonstrate that our framework achieves absolute mAP improvements of 2.3\% on the small object benchmark (NWPU VHR-10) and 1.8\% on the general dataset (PASCAL VOC) compared to the baseline.
    Keyword: Object Detection · Vision Transformer · Heatmap
    DOI: 10.65286/icic.v22i1.34191
    Cite

  • Terrain-Aware Attention Network for DEM Terrain Semantic Segmentation, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: yuan gao, bo zhou, yifan she
    Abstract: Terrain segmentation is important for semantic extraction and scene understanding in geographic information systems. However, generic deep learning models lack mechanisms tailored to the physical characteristics of terrain data. We introduce the Terrain-Aware Attention (TAA) module, which uses terrain cues---local gradient, global elevation context, and surface ruggedness---to guide the attention mechanism for terrain-specific segmentation. We also propose the Hierarchical Asymmetric Attention Mechanism (HAAM), which assigns different attention modules to different network depths. We integrate TAA and HAAM into a U-Net architecture and evaluate them on a DEM dataset derived from ALOS PALSAR imagery. Experiments show that both TAA and HAAM independently improve segmentation performance over the baseline, and their combination achieves the highest IoU among all tested configurations.
    Keyword: Semantic Segmentation · Attention Mechanism · Domain Knowledge · Digital Terrain
    DOI: 10.65286/icic.v22i1.29799
    Cite

  • Early Detection of Ransomware via Multi-Source Fusion with Hardware Performance Counters, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Yonghui Xi, Siyu Zhang, Jun Song, Fan Yang
    Abstract: In recent years, ransomware detection based on hardware performance counters (HPCs) has shown promise. However, existing HPC-based approaches are constrained by limited feature modeling and insufficient behavioral context, leading to high false positive rates and detection delays from data accumulation. To address these issues, we propose an early detection approach that combines HPCs, disk input/output (I/O) statistics, and ransom note analysis. The approach introduces an attention-enhanced temporal convolutional network to extract multi-scale features from HPC data, thus improving the extraction of temporal characteristics in HPC sequences. Furthermore, disk I/O patterns are analyzed through machine learning to provide additional behavioral context beyond HPCs, significantly reducing false positives. We also design a Minifilter-based ransom note detector to identify ransom notes dropped by ransomware, further reducing detection latency. Extensive experiments demonstrate that our approach achieves a Matthews Correlation Coefficient (MCC) of 95.95%, a false positive rate of only 0.48% on unknown ransomware, and reduces average detection latency by 0.8 seconds across 14 ransomware families. It outperforms representative approaches and exhibits strong potential for detecting unknown ransomware.
    Keyword: ransomware detection, hardware performance counters, temporal convolutional network, ransom notes
    DOI: 10.65286/icic.v22i2.91455
    Cite

  • Structure-Aware and Frequency-Guided Diffusion Framework for Multimodal Fashion Image Editing, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Xin Chen, Jie Zhang, Peng Zhang, Shasha Meng
    Abstract: Fashion image editing aims to modify target garment attributes under textual and reference guidance while preserving non-target contents. However, existing methods often suffer from inaccurate garment localization, insufficient preservation of high-frequency textures, and unnatural transitions near edited boundaries. To address these issues, we propose a structure-aware and frequency-guided framework for multimodal fashion image editing. Specifically, we design a Structure-Guided Textual Mask Network to predict geometry-aware editing regions by leveraging refined textual structural cues and human-centric priors, where a structural prior reweighting mechanism is introduced to improve localization accuracy. We further develop an adaptive frequency-domain texture enhancement module to inject high-frequency fabric details from a reference image during late denoising, and employ a boundary-band soft fusion strategy to ensure smooth visual transitions. In addition, we construct a new dataset, DFEdit, for fine-grained multimodal fashion image editing. Experimental results show that the proposed method achieves competitive performance in terms of editing fidelity, texture consistency, and visual quality, showing its effectiveness for intelligent fashion image editing applications.
    Keyword: Fashion Image Editing, Diffusion Models, Multimodal processing, Frequency-domain Texture Enhancement
    DOI: 10.65286/icic.v22i1.35765
    Cite

  • Chinese Imagined Speech EEG Classification Method Based on Topological and Frequency-Band Priors, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Ke Su, Haoran Guo, Haoju Wang, Liang Tian
    Abstract: EEG-based imagined speech classification is an important topic in brain–computer interface research. However, Chinese imagined speech EEG sig-nals are typically characterized by low signal-to-noise ratio, strong non-stationarity, and subtle inter-class differences, which make stable modeling challenging. Existing convolutional methods are limited in capturing long-range dependencies, while Transformer-based models often fail to fully ex-ploit spatial topology and frequency-band information. To address these is-sues, this paper proposes a topology- and frequency-band-prior-guided con-volution–Transformer hybrid model, named TFP-CTNet. The model incorpo-rates a channel topology-aware embedding at the input stage to preserve elec-trode spatial relationships. A multi-scale convolutional module is then used to enhance local discriminative representations, and a frequency-band-aware gated residual MLP is introduced at the classification stage to selectively re-fine high-level features. Experiments on the Chisco dataset show that the proposed method consistently outperforms several competitive baselines in terms of accuracy and F1-score, while also achieving improved cross-subject consistency. These results demonstrate that incorporating structural and fre-quency-domain priors is beneficial for robust Chinese imagined speech EEG classification..
    Keyword: Keywords: brain–computer interface, imagined speech, EEG signals, convo-lution–Transformer model.
    DOI: 10.65286/icic.v22i2.56427
    Cite

  • DeRNN: Decomposed Recurrent Neural Network for Long-Term Time Series Forecasting, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Shanyun Qian
    Abstract: Long-Term Time Series Forecasting (LTSF) is pivotal in domains like energy and traffic management but necessitates capturing intricate dependencies over extended windows. While Transformer-based models dominate, they suffer from quadratic complexity and positional insensitivity. Conversely, recent lightweight MLP/RNN-based models often forcibly compress conflicting dynamic features—linear trends and non-linear fluctuations—into a single channel, leading to suboptimal accuracy. To address these limitations, we propose the Decomposed Recurrent Neural Network (DeRNN). Our approach decouples global trend modeling from local fluctuation extraction via an asymmetric dual-track architecture. Specifically, we introduce a Trend Anchor Track to preserve global scale via direct linear projection, and a Seasonal Feature Track utilizing Bi-directional GRUs to capture complex non-linear dependencies within a reversible normalized space. Extensive experiments on seven benchmarks demonstrate that DeRNN achieves highly competitive, and in most cases superior, accuracy against state-of-the-art methods while maintaining extremely low latency and memory usage. Furthermore, the model exhibits superior robustness against noise and distribution shifts.
    Keyword: Long-Term Time Series Forecasting, Time Series Decomposition, Recurrent Neural Network, Direct Projection, Distribution Shift
    DOI: 10.65286/icic.v22i2.57947
    Cite

  • RoGloRE: A Global-Local Adaptive Joint Knowledge Extraction Framework for Chinese Cyber Threat Intelligence, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Jipeng Tang, Hao Hu, Yingchang Jiang, Yixiao Peng, Nan Wang, Yongwei Wang
    Abstract: Automated parsing of Cyber Threat Intelligence (CTI) is crucial for threat attribution and proactive defense. However, Chinese CTI texts are highly unstructured and semantically fragmented, posing dual challenges for existing models. In entity extraction, fragmented tokenization caused by high-entropy entities and complex nested structures leads to ambiguous entity boundaries. In relation extraction, critical attack clues are scattered across paragraphs, preventing traditional attention mechanisms from effectively capturing long-range dependencies and relative spatial structures. To address these limitations, we propose an adaptive global-local joint extraction framework designed for fragmented semantic aggregation in Chinese CTI. Within the entity recognition module, we introduce adaptive rotary position embeddings to correct low-level positional features. This mechanism, combined with a type-decoupled GlobalPointer, resolves recognition conflicts involving long-span entities and nested boundaries. In the relation extraction module, we design a dual-stage attention mechanism to dynamically integrate global cross-paragraph spatial clues with local entity neighborhood features. Additionally, an adaptive decoding strategy aware of class imbalance is implemented to enhance the robustness of the model against sparse long-tail relations. Experimental results on the CDTier dataset indicate that the proposed framework achieves a 9.45% improvement in the F1 score of entity extraction over the best existing baseline, alongside a precision of 93.3% and a recall of 95.5% for relation extraction. The proposed method overcomes the bottleneck of long-range semantic parsing in complex Chinese contexts, demonstrating superior generalization capabilities and practical utility.
    Keyword: Cyber threat intelligence, Entity and relation extraction, Rotary position embedding, and GlobalPointer
    DOI: 10.65286/icic.v22i2.80496
    Cite

  • Activated Three Stages Building Change Detection Network, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Xin Wang, Aocheng Shu, Wei Wang
    Abstract: Current Building Change Detection (BCD) methods suffer from several drawbacks, including inadequate multi-scale feature interaction, blurred boundaries, and excessive computational costs. To address these challenges, ActPML is proposed which includes: 1) at the previous stage, channel-wise feature fusion via learnable cross-stage alignment, followed by a Global Ac-tivation Layer for dynamic feature filtering, 2) at the middle stage, the fea-tures derived from the previous stage are refined through a Global-Local Feature Aggregator and a Self Deep Fusion Enhancer, respectively, 3) at the late stage, the Global-Local Self Enhancer is utilized to guide the decoder in producing the change map. ActPML achieves state-of-the-art performance on the LEVIR-CD and WHU-CD datasets, with F1-scores reaching 92.14% and 94.43%. Additionally, generalization capability was further validated on the CLCD dataset. Moreover, the proposed Global Activation Layer signifi-cantly reduces computational overhead compared to KANs, requiring sub-stantially less runtime. Our code can be seen at https://anonymous.4open.science/r/ActPML-EEB6.
    Keyword: BCD, stage, global-local, KAN, GAL
    DOI: 10.65286/icic.v22i1.74605
    Cite

  • Improving the Natural Language Inference Robustness to Hard Dataset by Data Augmentation and Preprocessing, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Zijiang YANG
    Abstract: Natural Language Inference (NLI) is the task of inferring whether the hypothesis can be justified by the given premise. Basically, we classify the hypothesis into three labels(entailment, neutrality and contradiction) given the premise. NLI was well studied by the previous researchers. A number of models, especially the transformer based ones, have achieved significant improvement on these tasks. However, it is reported that these models are suffering when they are dealing with hard datasets. Particularly, they perform much worse when dealing with unseen out-of-distribution premise and hypothesis. They may not understand the semantic content but learn the spurious correlations. In this work, we propose the data augmentation and preprocessing methods to solve the word overlap, numerical reasoning and length mismatch problems. These methods are general methods that do not rely on the distribution of the testing data and they help improve the robustness of the models. The experimental results show that the proposed method achieves more than 12\% classification accuracy improvement on the HANS dataset and 6\% to 9\% improvement on the ANLI dataset. The notable improvement is achieved by only training on extra 3\% data.
    Keyword: Natural Language Inference, Robustness, Data Augmentation, Pre-processing
    DOI: 10.65286/icic.v22i1.65363
    Cite

  • A Layer-Wise Syndrome-Based Framework for LDPC Decoding with Deep Learning, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Ziwen Ren, Pengcheng Wang
    Abstract: This paper proposes a newly developed layer-wise syndrome-based framework for decoding Low-Density Parity-Check (LDPC) codes using deep learning. The framework leverages a layer-specific training approach to design methodologically effective decoding strategies, aiming to utilize the specific characteristics of each layer and improve overall decoding performance. By integrating the Normalized Offset Min-Sum (NOMS) algorithm into a Forward-Feedback Neural Network (NOMS-FF), the proposed model employs a soft-syndrome loss function to assist the learning of codeword structures and optimize decoding performance across varying noise levels. Additionally, the framework incorporates a layer-freezing strategy and dynamic Signal-to-Noise Ratio (SNR) allocation, enabling targeted post-training adjustments tailored to layer-specific characteristics. Experimental results on 6G-candidate QC-LDPC codes (Base Graph 2, N=240,K=80 ) demonstrate that the proposed approach provides substantial performance advantages over traditional decoders across the entire SNR range tested in this study. Furthermore, this study conducts partial-layer decoding experiments, which demonstrate that using only a subset of neural network layers can reduce the model’s computational runtime while simultaneously maintaining or improving decoding accuracy compared to the full-layer model under specific noise conditions.
    Keyword: Deep Learning, Low-Density Parity-Check (LDPC) Codes, Belief Propaga-tion (BP), Normalized Offset Min-Sum (NOMS), Soft-Syndrome Loss, Lay-er-Freezing, Dynamic SNR Allocation.
    DOI: 10.65286/icic.v22i2.13234
    Cite

  • CLIP Guided UNet for Smoke Segmentation, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Xingying Chu, Jinhua Xu
    Abstract: Smoke segmentation aims to classify pixels with the smoke label, which is a downstream task of semantic segmentation. Convolutional neural networks (CNNs) have made great progresses in semantic segmentation. UNet is a widely used CNN-based semantic segmentation model. However, the conventional CNNs are inherently limited to local receptive fields that only provide short-range contextual information. Pretrained Vision-Language Models (VLMs) such as CLIP have learned rich semantics from web-scale image-text pairs. Inspired by this, we propose a novel CLIP guided UNet framework (CGUnet) for smoke segmentation, which merits the global and rich context of CLIP and the precise localization of UNet. Specifically, we design a CLIP guided cross attention (CGCA) module, in which the CLIP feature is used as the query, and the visual features of UNet as the Key and the Value. We conduct experiments on two public smoke segmentation datasets. Our method achieves SOTA results on both datasets, outperforming other methods.
    Keyword: Smoke segmentation, UNet, CLIP, Visual language model, semantic segmentation.
    DOI: 10.65286/icic.v22i1.40385
    Cite

  • Fine-Grained Fact-Checking for Short Videos: A Multi-Agent Report Generation Framework, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Jingzhe Liu, Xiangyu Qiu, Wenze Ouyang, Yifan Wang, Jiayong Wen, Guoyan Xu
    Abstract: With the rapid growth of short video platforms, the spread of fake news in short video format has become increasingly complex and deceptive. While existing text/image fact-checking methods fail to identify multiple fake claims within a single video, video fake news detection methods do not provide fine-grained evidence-based analysis. To address this gap, we propose a novel task: Fine-grained Fact-Checking Report Generation for Short Videos. Given a short video containing textual, visual, and audio modalities, the goal is to automatically generate a structured report that identifies specific fake claims and provides detailed analyses based on external evidence. We construct a benchmark dataset annotated by domain experts, along with fine-grained evaluation questions. Furthermore, we propose Video Multi-agent Fine-grained Fact-checking (VMFF), a training-free multi-agent framework that simulates the workflow of professional fact-checkers through three modules: (1) Video Understanding, (2) Task Decomposition, and (3) Retrieval, Reasoning, and Generation. Experimental results demonstrate its effectiveness in fine-grained fact-checking report generation.
    Keyword: Short Video Fact-checking, Agentic System, Multimodal Large Language Models.
    DOI: 10.65286/icic.v22i2.74294
    Cite

  • SuRe-EM: Subspace-Routing and Residual-Corrected Expert Model for Domain-Adaptive Retrieval, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Xifan Liu, Youyou Huang, Zebiao Chen, Renchao Zen, Siqiang Li, Shouqiang Liu
    Abstract: As the de facto standard for knowledge-intensive tasks, Retrieval-Augmented Generation (RAG) has significantly enhanced the reliability of Large Language Models by incorporating external non-parametric knowledge. However, adapting general-purpose retrievers to specific vertical domains often triggers catastrophic forgetting, severely degrading performance on open-domain queries. Additionally, existing mitigation strategies, such as linear model fusion, are mathematically constrained within a linear geometric manifold, limiting their ability to effectively rectify complex non-linear semantic drifts caused by domain shifts. To address these limitations, we present an effective approach, SuRe-EM (Subspace-routing & Residual-corrected Expert Model), designed to resolve the issues of domain specialization and Cross-Domain generalization in domain-adaptive retrieval. SuRe-EM enhances Cross-Domain representation by integrating fine-grained subspace routing with non-linear residual correction. Specifically, SuRe-EM first employs subspace routing to dynamically decouple high-dimensional features for maximizing domain specialization, followed by a residual module that generates non-linear semantic compensations. We validate our model on a vertical domain dataset (AHD) and general domain datasets (CMRC, SQuAD). Quantitative results demonstrate the effectiveness of SuRe-EM, which maintains strong In-Domain precision while improving Recall@10 by up to 4.46 points over state-of-the-art linear fusion baselines in Cross-Domain scenarios. Furthermore, comprehensive ablation studies validate the non-redundant synergy of the key design elements within SuRe-EM.
    Keyword: Retrieval Augmented Generation, Natural Language Processing, Catastrophic Forgetting, Large Language Models
    DOI: 10.65286/icic.v22i1.15304
    Cite

  • BSM-SegNet: Boundary-Scale Synergistic Segmentation Network for Lumbar Muscle MRI Segmentation in Sarcopenia Assessment, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Junjie Dang, Huanqing Cui, Yining Tang, Ruixia Liu
    Abstract: Accurate segmentation of lumbar muscle groups is essential for imaging-based screening and quantitative assessment of sarcopenia. However, automatic lumbar muscle segmentation in magnetic resonance imaging (MRI) remains challenging due to indistinct boundaries, adhesion between adjacent muscles, significant inter-class scale variations, and imaging artifacts. To address these challenges, this paper proposes a boundary–scale synergistic perception network for multi-class muscle segmentation, termed Boundary–Scale Synergistic Segmentation Network (BSM-SegNet). The proposed network integrates boundary-sensitive modeling, multi-scale feature representation, and cross-stage feature selection to enhance segmentation accuracy and robustness. Specifically, a Boundary-Sensitive Residual Module (BSRM) is designed to strengthen feature responses in muscle boundary regions through local feature differencing. A Frequency–Scale Coupling Module (FSCM) is introduced to improve multi-scale structural modeling via multi-scale pooling and channel attention. In addition, a Cross-Stage Gating (CSG) module is employed to suppress redundant features and enhance semantic consistency during encoder–decoder feature fusion. Extensive experiments are conducted on a private lumbar muscle MRI dataset with 562 subjects and an external multi-center public dataset with 290 subjects. Experimental results show that BSM-SegNet achieves an IoU of 86.68% and a Dice score of 92.85% on the private dataset, and an IoU of 86.16% and a Dice score of 92.50% on the external dataset, outperforming several representative segmentation methods. Ablation studies further demonstrate the effectiveness and synergistic benefits of the proposed modules. These results indicate that BSM-SegNet is well suited for accurate lumbar muscle segmentation in sarcopenia assessment.
    Keyword: Sarcopenia, Lumbar muscle segmentation, Magnetic resonance imaging, Deep learning, Multi-class segmentation
    DOI: 10.65286/icic.v22i1.98848
    Cite

  • SRLane+: Improved Sketch and Refinement Method for Lane Detection, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Hang Liu, Jinhua Xu
    Abstract: Lane detection is a significant task with applications in autonomous driving tasks such as adaptive cruise control, lane departure warning, and lane-keep assistance. Convolutional neural networks (CNN) and transformers have been used for lane detection and achieved great performance. However, there are still challenges under complex scenarios, such as crowded or dazzling conditions. Utilizing the global and slender shape of lanes, line anchor-based methods have attracted much attention. Sketch and refinement (SRLane) is a two-stage anchor-based method for lane detection, composed of proposal generation stage (Sketch) and Refinement stage. This paper improves SRLane on both stages. At the first stage, position embedding is introduced and fused with the multi-scale feature maps to estimate local direction map more accurately for anchor proposal generation. At the second stage, a new fine regressor is proposed, which includes a multi-scale feature fusion module and a transformer decoder to further refine the lanes utilizing the context information. The whole method adopts an end-to-end training mechanism. Experiments are conducted on the CULane and Tusimple datasets. The results show that the proposed method achieves 80.35\% F1 score on the CULane dataset, outperforming existing methods.
    Keyword: Lane Detection, feature Fusion, transformer decoder, line anchor.
    DOI: 10.65286/icic.v22i1.71281
    Cite

  • Localization-Aware Adversarial Attacks Against LiDAR-Based 3D Object Detection, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: peng gao, guanhua sun, shuo chai, tieying zhu
    Abstract: Robustness in 3D object detection is paramount for safety-critical applications such as autonomous driving. However, existing adversarial attack methodologies often fail to fully leverage the intrinsic spatial geometric characteristics of 3D point clouds. To address this limitation, we propose a novel differentiable method named Localization-Aware Adversarial Attacks (LAA). LAA explicitly incorporates the absolute coordinates of bounding boxes, the relative spatial relationships with ground truth objects, and rotation angles into its adversarial optimization objective. By generating imperceptible point cloud perturbations, LAA aims to directly disrupt the localization awareness of 3D object detectors. Extensive experiments on the KITTI dataset against a variety of mainstream 3D object detectors demonstrate that LAA exhibits remarkable effectiveness in compromising the detectors’ localization capabilities. This research reveals the vulnerability of current 3D object detectors regarding localization awareness and provides valuable insights for constructing more robust detection systems.
    Keyword: 3D Object Detection, Point Cloud, LiDAR, Adversarial Attacks
    DOI: 10.65286/icic.v22i1.62031
    Cite

  • SAGE: Structure-Aware Generative Enhancement for Long-Tail Knowledge Graph Completion, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Erzhuo Xu, Jun Huang, Hailong Zhang
    Abstract: Knowledge graph completion is one of the core tasks in the field of knowledge graphs, which predicts missing links through inference of existing facts. With the advancement of deep learning technology, utilizing end-to-end deep learning models for knowledge graph completion has become a cutting-edge research direction. However, the performance of current knowledge graph completion models is still limited by text quality and incomplete structure. To address this issue, this paper proposes a method of using large models for data augmentation to improve the inference performance of the model. Specifically, we first introduced a pre-extractor model based on a hybrid architecture of rules and neural networks, which is used to identify long tail entities in the dataset and generate several candidate tail entities through relationships. Then, we use this data to have the Large Language Model infer the most factual triplet. Finally, we use the enhanced dataset for predictive inference. SAGE achieves better results on three standard KGC datasets. For instance, on the FB15K-237 dataset, compared to the SimKGC baseline model, SAGE improves Hits@1 by 1%, Hits@3 by 0.9%, and Hits@10 by 1.6%.
    Keyword: Knowledge Graph Completion, Data Augmentation, Large Models.
    DOI: 10.65286/icic.v22i2.69039
    Cite

  • Unsupervised Video Anomaly Detection Based on Graph Attention Propagation and Semantic Information, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Qinghao Kong, Wanru Xu, Zhenjiang Miao, Ruizhao Zhai
    Abstract: Video Anomaly Detection (VAD) is a crucial computer vision task for security monitoring and public safety. Unsupervised VAD is more suitable for real-world scenarios with rare unknown anomalies, but existing LLM-based methods suffer from limited temporal modeling, inconsistent video understand ing and inaccurate fine-grained localization, leading to biased anomaly scoring. To solve these problems, we propose a novel unsupervised VAD framework fus ing graph attention propagation and multimodal semantic information: first, fuse video semantic and motion features to construct a dynamic spatiotemporal graph, and refine node features via graph attention propagation with orthogonal con straints; then, split videos into semantically coherent event units by a statistical boundary detection module; finally, guide MLLMs to generate event semantic descriptions and initial anomaly scores through a hierarchical prompting strategy, and refine the scores via video-text semantic alignment to obtain accurate frame level scores. Evaluated on UCF-Crime and XD-Violence datasets with frame level AUC, the proposed framework achieves state-of-the-art performance under unsupervised and zero-shot settings, significantly outperforming existing LLM based VAD methods and even several weakly supervised approaches, which fully verifies its effectiveness and robustness.
    Keyword: Video Anomaly Detection, Graph Attention Network, Multimodal Large Language Model.
    DOI: 10.65286/icic.v22i2.84630
    Cite

  • MFE-Net: A Hybrid Perception and Manifold-Preserving Framework for Robust Fine-Grained Underwater Object Detection, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Jiaxin Chen, Xin Wang, Yuxu Peng, Dengyong Zhang, Lei Wang, Yongjie Zhang
    Abstract: High-performance underwater object detection constitutes a pivotal technological foundation for advancing marine intelligent computing and facilitating the autonomous deployment of subsea platforms. Nevertheless, underwater visual perception is inherently constrained by non-linear optical degradation, manifesting as an ill-posed inverse problem that induces ``semantic confusion'' and the ``feature annihilation'' of sub-pixel entities. Conventional deep learning architectures, predicated on strided downsampling, frequently function as irreversible low-pass filters, precipitating a catastrophic loss of information entropy regarding sparse high-frequency geometric manifolds during hierarchical transmission. To address these challenges, this study formulates MFE-Net, a hierarchical hybrid perception and manifold-preserving framework. The Hybrid Perception Feature Extraction (C3CF) module harmonizes local convolutional inductive biases with global attention recalibration to effectively suppress non-structured scattering noise and enhance deep feature representation. Building upon this, the Multi-Scale Feature Enhancement Neck (MFE-Neck) incorporates a lossless Fine-grained Feature Enhancement Branch and an Efficient Multi-Scale Feature Integration (EMFI) mechanism, leveraging Soft Nearest-neighbor Interpolation (SNI) to ensure signal fidelity and manifold consistency across heterogeneous scales. Extensive quantitative evaluations on specialized benchmarks (RUOD, DUO, and DeepFish) demonstrate that MFE-Net achieves 86.4% mAP_{50} and 63.9% mAP_{50-95} on RUOD with a balanced computational load of 12.8 GFLOPs, surpassing contemporary state-of-the-art (SOTA) architectures within the evaluated scope.
    Keyword: Underwater object detection, Manifold preservation, Feature annihilation, Hybrid perception, Multi-scale feature integration
    DOI: 10.65286/icic.v22i1.71278
    Cite

  • SDGMamba: Stage-Aware Mamba with Deformable Detail and Gated Fusion for Clothing Parsing, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Haitao Fu, Shaoyu Wang, Zhongyuan Teng, Jiaxin He, Hong Wan, Xiujin Shi
    Abstract: Clothing parsing is a pivotal subtask of semantic segmentation with significant research and practical implications. However, it faces persistent challenges, including substantial scale variations among clothing components, posture-induced occlusions, and high inter-class visual similarity, which necessitate models capable of synergizing global semantics with local details. While existing CNN-based approaches are constrained by limited receptive fields, Transformer-based—despite their proficiency in modeling long-range dependencies—often suffer from uniform input processing that dilutes fine-grained details. To address these limitations, we propose SDGMamba, a novel clothing parsing framework that effectively leverages scale-dependent features for global modeling while preserving intricate characteristics. First, we design the Stage-Aware Attention VSS (SASS) Block, which dynamically allocates attention based on network depth to facilitate hierarchical structure-aware adjustments. Second, we introduce the Deformable Detail Retention (DDR) module, which utilizes deformable convolutions to adaptively align and fuse multi-scale depth information, thereby enhancing texture representation. Finally, the Gated Cross-Scale Fusion (GCSF) module employs a gating mechanism to refine shallow features guided by high-level semantics, strengthening semantic coherence among components. Experiments on the CFPD dataset demonstrate that SDGMamba achieves a PA of 94.48% and an mIoU of 58.57%, exhibiting superior performance compared to CNN-based, Transformer-based, and Mamba-based methods.
    Keyword: Clothing Parsing, Visual State Space Model, Stage-Aware Attention, Deformable Detail, Gated Fusion
    DOI: 10.65286/icic.v22i1.45102
    Cite

  • Diff-SiamNet: Breast Ultrasound Image Classification with Forward Diffusion and Discriminative Embedding, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Chenyi Zhuang, Di Zhang
    Abstract: Existing breast ultrasound image classification algorithms frequently perform poorly in real-world clinical settings due to image deterioration, semantic drift, and insufficient supervision, significantly limiting tiny lesion recognition and practical deployment. To address this critical issue, this paper proposes Diff-SiamNet, a lightweight discriminative diffusion network. We present a novel forward diffusion perturbation mechanism to emulate scattering noise and boundary blurring, enhancing the model’s feature robustness against low-quality images. Additionally, we develop a diffusion-aware attention gate (DAG) to dynamically integrate clear and degraded map features, thereby alleviating semantic discrepancies. The Triplet++ embedding strategy is designed to guide the model to form a discriminative space with inter-class separation and intra-class aggregation under weak annotation. On the BUSI dataset, Diff-SiamNet outperforms state-of-the-art ResNet, Swin Transformer and HGDF in five metrics, including accuracy (93.4%), AUC (94.2%), and F1-score (92.9%), which significantly improves the ability of recognizing fuzzy boundaries and tiny lesions. The method has good interpretability and deployment efficiency, and is expected to serve clinical intelligent screening in resource-constrained environments
    Keyword: Breast ultrasound 、Image classification、Diffusion model、Discriminative embedding
    DOI: 10.65286/icic.v22i1.89160
    Cite

  • Depth Guided Scale-Aware Transformer for Crowd Counting, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Qijun Lu, Jinhua Xu
    Abstract: Crowd counting is an important task in computer vision with wide applications. Deep learning methods have achieved great progress in the crowd counting task, but scale variation in the crowd images is still a challenge. To address this issue, we propose a Depth Guided Scale-Aware Transformer (DGSAT) for crowd counting in this paper. Three novel modules are designed including Dilated Rectangle Window Attention (DRWA), Depth-guided Scale Segmentation (DGSS), and Size-Aware Hungarian Matching (SAHM). In the DRWA module, we apply different dilation rates to the conventional window attention mechanism to collect context from different scales. In the DGSS module, the depth map is used to guide the prediction of a scale segmentation mask so that features from different scales can be selected to localize the heads of different scales. In the SAHM module, head sizes are estimated and used to balance the position error and the classification confidence score in the cost matrix of the Hungarian matching. Our method has been extensively evaluated on crowd counting datasets and achieves SOTA results. The source code is available upon reception.
    Keyword: Crowd-counting, transformer, window attention, Hungarian matching.
    DOI: 10.65286/icic.v22i1.43020
    Cite

  • Component-Aware Spatio-Temporal Adaptation of Frozen Foundation Models for Video Deepfake Detection, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Tianyi Zhang, Menghan Liang, Wenzheng Liu, Cheng Fu, Gang Zhu
    Abstract: The advancement of deep generative models facilitates realistic synthetic facial videos, threatening social trust and digital security. Existing detection methods achieve strong in-domain performance but suffer from cross-dataset degradation, primarily due to overfitting to dataset-specific spatial artifacts. To address these challenges, we propose a parameter-efficient, video-based deepfake detection framework that leverages a frozen foundation model encoder coupled with a lightweight spatio-temporal decoder. First, we introduce a Component-Aware Spatial Enhancement (CASE) module that selectively accentuates manipulation-prone facial regions, such as the eyes, mouth and nose, while capturing global-local structural inconsistencies, thereby enabling the detection of subtle artifacts. Second, it is complemented by a Bidirectional Spatio-Temporal (Bi-ST) decoder that models local temporal transitions and bidirectional temporal dependencies across sampled frames, enabling robust temporal reasoning within each video clip. Without fine-tuning the backbone network, our framework achieves robust cross-dataset generalization by jointly reasoning about spatial and temporal anomalies. Finally, extensive experiments demonstrate that the proposed method performs competitively against strong baselines, particularly under cross-dataset evaluation.
    Keyword: Deepfake Detection , Video forensics , Spatio-temporal modeling , Cross-domain generalization
    DOI: 10.65286/icic.v22i1.88208
    Cite

  • GM-Geo: Graph-Enhanced Selective State Space Learning for Irregular Borehole Sequences, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Xinwei Yao,Konglong Wang,Kunhua Yang,Qiang Li
    Abstract: Lithology identification is essential for formation evaluation and reservoir characterization, serving as a crucial basis for assessing the presence of mineral deposits underground. However, due to the lateral spatial correlations and ordered vertical dependencies of subsurface lithology across boreholes, high-precision lithology prediction from irregular borehole observation data is extremely challenging. Existing methods typically focus on sequence modeling or graph-based neighborhood learning, but rarely integrate both into a unified framework. To address this issue, we propose GM-Geo, a graph-enhanced selective state space framework for lithology prediction from irregular borehole sequences. Specifically, the proposed method first transforms heterogeneous interval-based borehole records into standardized depth-wise samples through geology-aware resampling. It then constructs a sample-level spatial neighborhood graph to capture local geological relationships across boreholes. Finally, graph-aggregated spatial features are fused into a selective state space model to jointly encode lateral spatial continuity and vertical lithological evolution. This paper evaluates the framework in 218 boreholes in the Tangwuli fluorite mining area and compares it with baseline methods such as RF, XGBoost, Transformer, GNN, and ET4DD. GM-Geo achieved best precision, recall, macro F1 score, and micro F1 score of 0.845, 0.798, 0.812, and 0.821, respectively. Ablation experiments further demonstrated that both the graph module and the state-space module contributed to the final performance.
    Keyword: Lithological prediction, irregular borehole sequences, selective state space model, graph-enhanced learning, geology-aware resampling
    DOI: 10.65286/icic.v22i2.47469
    Cite

  • ProtoHGC: Heterogeneous Graph Contrastive Learning with Prototype-Regularized Classification for Transaction Fraud Detection, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Yuzhen Lin, Yuan Gao, Han Wu, Qi Hu, Yajin Li, Ge Zhang, Ivonne Xu
    Abstract: Transaction fraud detection on digital payment platforms poses three inter-twined challenges: sparse supervision, severe class imbalance, and complex relational dependencies among accounts and transaction events. This paper studies fraud detection on a heterogeneous transaction graph constructed from the public PaySim simulator. To this end, we propose ProtoHGC, a framework that integrates heterogeneous graph contrastive pre-training with prototype-regularized classification. ProtoHGC employs a multi-scale en-coder that aggregates first-order neighborhood features and second-order meta-path-based context, coupled with node-level and subgraph-level con-trastive objectives to enhance representation quality under limited supervi-sion. Building on the learned embeddings, class-specific prototypes for normal and fraudulent behaviors enforce intra-class compactness and inter-class separation. Experiments on PaySim demonstrate that ProtoHGC con-sistently outperforms GCN, GAT, GraphSAGE, and MLP baselines across all tested hidden dimensions in terms of AUC-ROC, PR-AUC, accuracy, re-call, and F1-score. We explicitly note that the current evaluation is limited to a single synthetic dataset; few-shot adaptation, cross-domain generaliza-tion, and multimedia-specific fraud scenarios remain unvalidated.
    Keyword: Fraud Detection· Heterogeneous Graph Neural Networks · Graph Contras-tive Learning · Prototype-Regularized Classification · Multi-Scale Message Passing
    DOI: 10.65286/icic.v22i2.66797
    Cite

  • Probabilistic Syntax-Aware Joint Span–Sentiment Learning for Aspect-Based Sentiment Analysis, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Yuxun Wang, Dunhui Yu
    Abstract: Aspect-Based Sentiment Analysis (ABSA) aims to identify aspect terms in text and predict their sentiment polarities. The two subtasks are tightly coupled: inaccurate aspect boundaries can mislead sentiment prediction, while sentiment supervision should in turn refine span selection. However, existing methods still suffer from two limitations. First, although boundary detection and sentiment classification are interdependent, gradients between them are often weak, resulting in error propagation, especially in pipeline settings. Second, under complex syntactic phenomena such as negation and contrast, boundary shifts and the mismatch between training on gold spans and inferring on predicted spans cause a train–inference discrepancy and unstable optimization. To address these issues, we propose a Probabilistic Syntax-Aware Joint Span–Sentiment Learning (PSJL) approach for ABSA. PSJL relaxes discrete span selection into a differentiable span distribution and dynamically incorporates syntactic constraints for end-to-end optimization. Specifically, (i) Probabilistic Joint Modeling (PJM) computes expectation-based span representations from a span distribution, enabling sentiment loss to backpropagate to boundary prediction. (ii) Syntax-Constrained Contrastive Learning (SCCL) derives span-level masks from dependency parse trees to contrast syntactically plausible and implausible candidates, regularizing boundary decisions. (iii) Adaptive Training Strategy (ATS) adaptively balances supervision from gold and predicted spans based on a syntax-consistency rate, mitigating train–inference discrepancy and improving training stability. Experiments on three benchmark datasets show that PSJL consistently outperforms state-of-the-art models and strong LLM-based baselines on aspect term extraction and sentiment classification.
    Keyword: Aspect-Based Sentiment Analysis, Joint Span–Sentiment Learning, Probabilistic Modeling, Syntax-Aware Learning, Contrastive Learning, Curriculum Learning.
    DOI: 10.65286/icic.v22i1.36537
    Cite

  • Dynamic Endoscopic Gaussian Reconstruction for Vascular Information Enhancement, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Ling Li, Wenyuan Huang, Jia Gu, Wenjian Qin
    Abstract: Reconstructing deformable tissues from endoscopic videos is a critical task in fields such as robot-assisted minimally invasive surgery, virtual reality, and augmented reality. Compared with neural radiance fields, the introduction of the 3D Gaussian Splatting technique (3DGS) significantly reduces the time re-quired for 3D reconstruction and rendering while maintaining good image quality to a certain extent, offering new possibilities for real-time 3D rendering during endoscopic procedures. However, this technique suffers from a signifi-cant loss of image details during rendering, manifesting in endoscopic scenes as missing vascular details, which directly affects the accuracy of lesion local-ization by clinicians. To address this issue, this paper proposes a novel method for enhancing vascular information in dynamic endoscopic scene reconstruc-tion. The proposed approach begins with image preprocessing, employing a lightweight and efficient vascular enhancement algorithm. By optimizing brightness, details, and contrast in the HSV color space, it enhances the visibil-ity of vascular structures. Combined with the state-of-the-art 3DGS method for endoscopy, it improves the display of vascular details in 3D images. Experi-mental results demonstrate that, compared with the original 3D rendering and other 2D enhancement algorithms, the proposed method significantly enhances the visualization quality of vascular structures, effectively compensating for the shortcomings of the 3D Gaussian Splatting technique in vascular detail rep-resentation.
    Keyword: Endoscopic images; Vascular enhancement; 3D Gaussian Splatting; Fast guid-ed filtering
    DOI: 10.65286/icic.v22i2.89248
    Cite

  • ObfusDomainNet: Reinforcement Learning-Driven Obfuscation and Domain-Locking for DNN Security, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Cong Ding, Changsheng Wan, Zhenjie Bao, Haitao Chen, Jimin Nie
    Abstract: Protecting Deep Neural Network (DNN) in fields like medical image analysis is challenging due to unauthorized access and misuse, which traditional methods like watermarking fail to prevent. We propose a reinforcement learning-driven weight obfuscation and key protection mechanism, along with a domain-locked task-specific model generation framework, to enhance DNN authorization protection. The weight obfuscation creates dynamic masks, ensuring only authorized users can access model functions, rendering illegal copies ineffective. The domain-locked framework uses Generative Adversarial Network (GAN) to maximize performance differences between tasks, ensuring the model excels only on authorized tasks and preventing misuse. Additionally, we embed watermarks into model parameters with hash algorithms for tamper recovery, allowing restoration after attacks. Experimental results show these methods significantly enhance DNN security and reliability in authorization protection, applicability assurance, and tamper recovery, with minimal impact on performance.
    Keyword: Deep Neural Networks (DNN), Intellectual Property Protection, Reinforcement Learning, Weight Obfuscation, Watermarking.
    DOI: 10.65286/icic.v22i2.30634
    Cite

  • AlphaX: A Tri-Agent Framework for Microstructure-Constrained Symbolic Alpha Discovery in Crypto Futures, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Jiaqi Wang, Shuo Yin
    Abstract: We present AlphaX, a microstructure-aware two-phase framework for symbolic alpha discovery in cryptocurrency perpetual-futures markets. AlphaX reformulates LLM-based factor mining as a unified discovery process that coordinates validity constraints, cross-round feedback, and library-level redundancy control. In Phase~I, specialized Idea, Coder, and Runner agents operate within a seven-gate validity cascade and a trajectory evolution module, filtering 65.7\% of degenerate expressions. In Phase~II, Ward-linkage orthogonalization reduces mean absolute pairwise correlation to 0.26. Overall, AlphaX reports a 0.145 mean factor IC and, under HRP allocation, a 1.739 Sharpe ratio with 56.0\% annualized return net of costs, while showing positive library-level validation outcomes relative to the reproduced baselines under the shared protocol.
    Keyword: Large Language Models ,Multi-Agent Systems,Cryptocurrency Derivatives, Symbolic Alpha Discovery,Redundancy Control
    DOI: 10.65286/icic.v22i2.15686
    Cite

  • GFSM-DETR: A Gated Frequency-Spatial Modulation Detector for Real-Time Small Face Detection, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Haomin Li, Zhen Song, Yanxing Huang, Beilei Wang
    Abstract: Real-time and accurate dense small-scale face detection constitutes a critical technology in computer vision. However, in complex environments, this task presents significant challenges, including high target density, minute scales, frequent mutual occlusions, and background noise interference. These factors often cause models to lose critical fine-grained features, leading to severe missed detections. To address these challenges, this paper proposes a high-accuracy real-time detector (GFSM-DETR) specifically optimized for dense small-scale face detection. First, the GFSM-Fusion module is designed to utilize synergistic representation learning across the spatial and frequency domains, effectively extracting and amplifying high-frequency features of small targets while suppressing background noise. Furthermore, the ECA-CATM module is constructed to facilitate noise-resistant global context modeling through multi-stage information interactions within the spatial and channel domains combined with an additive attention mechanism. In addition, the proposed Focal-PIoUv2 loss function utilizes a dual-constraint mechanism to effectively filter gradient interference generated by low-quality predicted boxes. On the SCUT-Head and Brainwash datasets, the proposed model significantly outperforms the baseline RT-DETR and a series of mainstream detectors, achieving mAP50 scores of 95.2% and 95.7%, respectively. Compared to leading mainstream detectors including YOLOv11m and YOLO26m, our model realizes mAP50 gains of 2.5% and 2.0%, along with Recall improvements of 3.9% and 2.0% across the two datasets. These results indicate that the proposed model provides a high-accuracy, low-cost solution for real-time deployment on edge and mobile devices.
    Keyword: Small face detection, Frequency-Spatial Modulation, RT-DETR, Global Con-text Modeling
    DOI: 10.65286/icic.v22i2.56410
    Cite

  • Deep State-Space Monocular Visual Odometry with Learnable Kalman Filtering, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Mingyue Wang, Yixuan Wu
    Abstract: Monocular visual odometry is important in autonomous driving, robotics, and related fields, and has attracted increasing attention in computer vision.Traditional geometric methods and end-to-end deep learning methods have achieved promising results in monocular visual odometry, but they still face limitations in temporal consistency, uncertainty representation, and abnormal observation handling.To address these issues, this paper proposes a monocular visual odometry method that combines deep temporal features with a Kalman filtering module based on a state-space model. By explicitly modeling the temporal evolution of motion states in the state space, the method introduces continuity constraints into the pose estimation process. At the same time, the state transition matrix, as well as the process noise covariance matrix and the measurement noise covariance matrix, are learned by neural networks, which enables the system to adaptively adjust the fusion weight between prediction and observation. Experimental results on the KITTI VO dataset show that, compared with the baseline method, the proposed model improves pose estimation robustness and adaptability to incomplete observations, which verifies the effectiveness of combining classical filtering theory with deep learning for monocular visual odometry.
    Keyword: Monocular Visual Odometry, Computer Vision, Kalman Filter, Deep Learning.
    DOI: 10.65286/icic.v22i1.99869
    Cite

  • ACE-CCO: A Two-Stage Deep Learning Method for Automatic Concentration Estimation of Chicken Coccidia Oocysts, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Qianchao Wang, Ting Luo, Haoshen Guo, Junxin Chen, Ximing Li, Yubin Guo
    Abstract: Chicken coccidiosis is a common and serious parasitic disease in poultry, and vaccination is an effective preventive measure. To ensure the quality of vaccine products, it is critical to accurately identify chicken coccidia species and count the number of oocysts. Currently, the most commonly used method relies on experienced professionals performing morphological identification and manual counting under a microscope. This process is highly subjective and often leads to inconsistent counting results. Although deep learning-based coccidia automatic counting methods have improved efficiency and stability, further enhancement in accuracy and process automation is still needed. To address these challenges, this paper proposes ACE-CCO, a two-stage deep learning method for automatic concentration estimation of chicken coccidia oocysts in vaccines. First, an image processing-based frame line recognition algorithm (FLRA) is designed to accurately locate the counting area in the hemocytometer. Next, a two-stage deep learning counting algorithm is developed to enhance counting accuracy. Finally, user-friendly software is developed to support one-click batch counting and statistical analysis, improving process automation. Comparative experiments with manual counting in real vaccine assessment scenarios show that the mean relative percentage difference (MRPD) of ACE-CCO for six chicken coccidia types is below 1.5%, and the counting speed is more than six times faster. The results demonstrate the superior accuracy, efficiency, stability, and practicality of ACE-CCO, greatly reducing reliance on operator experience and providing reliable technical support for vaccine quality assessment.
    Keyword: Chicken coccidiosis; Deep learning; Parasitic disease; Automatic concentration estimation; Vaccine quality assessment
    DOI: 10.65286/icic.v22i1.99443
    Cite

  • Dual-Channel Prompt-Enhanced BERT-RCNN with Adapters for Schizophrenia Detection from Dialogue Text, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Runzhe Zhang, Bowei Pan, Zhou Yuan, Yan Ling, Lianghuai Yang, Keji Mao
    Abstract: Schizophrenia is a severe mental disorder that imposes significant burdens on individuals, families, and society, making early detection critical for improving patient outcomes. However, current diagnosis relies predominantly on subjective clinical assessment. AI-assisted screening presents a promising alternative. In particular, automated analysis of patient language, a direct reflection of thought disorder, offers an objective and highly informative data modality. Despite this potential, research and practice are currently hindered by a scarcity of Chinese textual datasets and an over-reliance on non-linguistic data. To address these limi-tations, we construct the Chinese Schizophrenia Text Dataset (CSTD). Further-more, we propose a novel framework designed to decouple text into "general" and "visual and auditory hallucination" features. This framework utilizes prompt templates to inject clinical prior knowledge, employs BERT for feature extraction, and integrates dual sets of Adapters after each Transformer layer to learn the de-coupled features in parallel. The resulting outputs are fused and processed by a Recurrent Convolutional Neural Network (RCNN) for final classification. Experiments demonstrate that this framework achieves a leading accuracy of 95.58% on the CSTD, outperforming other strong models (including BERT and SBERT-BiLSTM) by at least 7.68%. Crucially, this robust performance highlights the strategic advantage of the parameter-efficient Adapter mechanism, which was deliberately chosen to restrict the number of trainable parameters and effectively prevent overfitting on the limited clinical dataset. These results validate the model's efficacy in recognizing linguistic patterns associated with schizophrenia and support its viability as an auxiliary screening tool in clinical outpatient settings.
    Keyword: Schizophrenia Detection, CSTD, Visual and Auditory Hallucination Features, Dual-Channel Adapter Feature Learning.
    DOI: 10.65286/icic.v22i2.66569
    Cite

  • Dictionary-Injected Pre-training for Traditional Mongolian-Chinese Machine Translation, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Zhiqiang Zhang, Yonghong Tian, Yehui Dang
    Abstract: In recent years, pretrained models have achieved remarkable progress in the field of natural language processing and have demonstrated superior performance across a variety of machine translation tasks. However, for the low-resource language pair of Traditional Mongolian and Chinese, the transfer capability of mainstream multilingual pretrained models is substantially limited, primarily because Traditional Mongolian is not covered in their pretraining corpora. To address this issue, this paper proposes a dictionary-injected pretraining method. Specifically, a constructed Mongolian–Chinese bilingual dictionary is used to perform lexical replacement on Chinese monolingual corpora, thereby generating mixed texts containing Traditional Mongolian. The model is then pretrained with a BART-style denoising autoencoding objective, enabling it to recover the original Chinese sentences under cross-lingual perturbations. In this way, the proposed method enhances the model’s cross-lingual semantic understanding and improves Mongolian–Chinese machine translation performance. Experimental results show that, on the test sets for both Mongolian-to-Chinese and Chinese-to-Mongolian translation, the proposed method outperforms a strong BART baseline by 2.3 and 2.4 BLEU points, respectively, thereby confirming its effectiveness for Mongolian–Chinese machine translation tasks.
    Keyword: Mongolian–Chinese bilingual dictionary; Mongolian–Chinese machine translation; pretrained models.
    DOI: 10.65286/icic.v22i1.44447
    Cite

  • EVA-LA: An Explainable Value-Added Learning Analytics Framework for Higher Education, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Fang Liu, Yuting Yang, Mengxue Hong, Qin Dai
    Abstract: Existing learning analytics approaches in higher education provide limited insight into students’ incremental academic growth and its underlying mechanisms. To address this gap, we propose an explainable value-added learning analytics framework, which contributes a residual-based value-added modeling paradigm, along with a cross-level interpretability mechanism for explaining student growth. It first constructs a structured learner representation that fuses heterogeneous digital trace data into a unified analytical space spanning behavioral, process, and performance dimensions. To quantify students’ academic growth, we build an elastic net-based estimator to model expected outcomes under high-dimensional and collinear predictors, enabling the derivation of individualized value-added scores as adjusted indicators of net learning progress. To enhance interpretability, we further develop Bi-Lens, a bidirectional explanatory framework that couples SHAP-based global attribution with LIME-based local diagnosis, supporting cross-level explanatory coherence and strengthening the identification of factors associated with value-added learning outcomes. Experiments on two real-world courses demonstrate competitive performance (R2 = 0.541 and 0.368) and stable, well-separated value-added estimates, while yielding interpretable insights for student development.
    Keyword: Learning Analytics, Value-Added Evaluation, Explainable Learning, Elastic Net.
    DOI: 10.65286/icic.v22i2.44276
    Cite

  • Bi-CMFM: A Multimodal Image Fusion Method Based on Bi-Path Residual and Cross-Modal Feature Merging, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Yang Jingdong, Zhou Zhentao, Liao Shengnan, Zhang Hongning, Teng Xiaoxi, Liu Tao, Zhang Wendong
    Abstract: Multimodal medical image fusion integrates complementary information from CT, MRI, and PET to support clinical diagnosis and lesion localization. CNN-based methods are constrained by local receptive fields and fail to model long-range dependencies, while Transformer-based methods incur quadratic computational complexity; both paradigms suffer from inadequate inter-modal redundancy suppression, producing blurred edges and structural inconsistency. To address these limitations, Bi-CMFM is proposed, comprising the Bi-Path Residual Fusion module (BPRF), which employs parallel standard and dilated convolutions to preserve local texture while expanding the receptive field, and the Cross-Modal Fusion Module (CMFM), which applies a Cross-modal Feature Enhancement component (CFEM) and a selective scanning mechanism to model long-range dependencies at linear complexity. Experimental results demonstrate that Bi-CMFM achieves EN = 5.1281 and AG = 7.0127 on CT-MRI fusion, PSNR of 12.7624/19.6562 on PET-MRI/SPECT-MRI tasks, and top-ranked EN, SF, AG, SCD, and CC on KAIST, outperforming six representative baselines including DATFuse, DRCM, and FusionMamba.
    Keyword: multimodal image fusion; bi-path residual; cross-modal feature enhancement; selective scanning; state space model
    DOI: 10.65286/icic.v22i2.43260
    Cite

  • RVQ-SNER: End-to-End Chinese Speech Named Entity Recognition via Quantized Acoustic Bottlenecks and Deep Acoustic–Semantic Fusion, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Yaoqiang Zhou, Xudong Luo
    Abstract: Conventional Speech Named Entity Recognition (SNER) typically relies on cascaded ASR (Automatic Speech Recognition)+NER (Named Entity Recognition) pipelines, which are hindered by error propagation and the underutilisation of acoustic cues. We propose an end-to-end Chinese SNER framework using Residual Vector Quantisation (RVQ) and deep acoustic--semantic fusion. The model extracts speech representations via a frozen Wav2Vec2-XLSR encoder, employing an RVQ-based bottleneck to reconstruct continuous quantized features that regularize the acoustic space and preserve semantic content. A Transformer decoder, trained with a joint CTC-attention objective, performs transcription while a gated deep-fusion mechanism integrates an external GPT model for linguistic consistency. For NER, a bidirectional multimodal fusion module aligns acoustic and semantic features before a GlobalPointer head performs span-level prediction. Experiments on AISHELL-NER and CNERTA yield F1-scores of 90.91\% and 81.44\%, respectively, outperforming pipeline, multimodal, and E2E baselines. These results demonstrate that deep bidirectional interaction between quantized acoustic streams and semantic contexts is essential for mitigating ASR error propagation and achieving robust Chinese SNER.
    Keyword: Natural language processing, End to end named entity recognition, Multimodal, Cross-modal attention, Automatic speech recognition.
    DOI: 10.65286/icic.v22i1.27751
    Cite

  • HETCLNN: A Lightweight Intrusion Detection Network with Class-Aware Self-Knowledge Distillation, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Chenghao Liu, Wenjun Yang
    Abstract: With the proliferation of edge computing, building efficient NIDS faces challenges related to limited computational resources and class imbalance. To address the issues of high overhead and poor detection of rare attacks in existing models, this paper proposes a lightweight intrusion detection model based on Class-Aware Self-Knowledge Distillation (CASKD). Architecturally, we design a lightweight network utilizing Heterogeneous Convolution (HetConv)-based residual and inverted residual structures. For training, a temperature-based CASKD method is introduced to tackle extreme class imbalance. Experimental results on CIC-IDS2017 and Bot-IoT datasets demonstrate classification accuracies exceeding 99\%. The proposed method significantly reduces computational overhead while improving detection precision for rare attacks, achieving an optimal balance between model compactness and performance.
    Keyword: NIDS、CASKD、HetConv、Lightweight Model、Class Imbalance
    DOI: 10.65286/icic.v22i2.40951
    Cite

  • DDRNet-SDR: A Spatial-Frequency Refinement Network for Real-Time Semantic Segmentation, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Shuhui Zhu, Jing Lu, Gang Shi
    Abstract: Real-time semantic segmentation aims to balance accuracy and inference effi-ciency, which remains a key challenge in computer vision. Although dual-resolution networks such as DDRNet achieve competitive performance, they are still limited by insufficient multi-scale context modeling, loss of high-frequency details caused by repeated downsampling, and suboptimal static fea-ture fusion strategies. To address these issues, we propose an enhanced dual-resolution network, termed DDRNet-SDR. Specifically, a Spatial-Frequency Fusion Module (SFFM) is introduced to exploit frequency-domain priors for attention genera-tion, enabling effective spatial feature refinement and preservation of high-frequency information. In addition, a DWR-Conv module is designed to inde-pendently model high- and low-resolution branches, facilitating efficient multi-scale context encoding through multi-branch and multi-dilation structures. Fur-thermore, a dynamic synergistic supervision strategy that combines Online Hard Example Mining (OHEM) and Dice loss is adopted to balance pixel-level accuracy and region-level consistency. Extensive experiments on the Cityscapes dataset show that DDRNet-SDR achieves 78.43% mIoU at 66.1FPS on a single 2080Ti GPU. Compared with the baseline, it yields a 1.03% absolute improvement in segmentation accuracy while maintaining real-time performance, demonstrating its effectiveness for latency-sensitive applications such as autonomous driving and UAV vision.
    Keyword: Real-time semantic segmentation, dual-resolution networks, spatial-frequency fusion module, multi-scale context modeling
    DOI: 10.65286/icic.v22i1.39161
    Cite

  • FedRGD: Risk-Guided Dynamic Defense against Federated Backdoors, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Ruiying Wang, Shuyang Hou
    Abstract: Federated Learning (FL) is vulnerable to backdoor attacks, where adversaries can stealthily manipulate the global model. Most existing defense methods are developed under IID assumptions, an assumption that rarely holds in practice. In highly non-IID settings, heterogeneous data distributions across clients make it difficult to distinguish malicious updates from benign ones, particularly when benign clients exhibit atypical patterns due to minority-class data. To address this challenge, existing defenses operate at different levels of granularity. Coarse-grained methods perform client-level filtering, which often mistakenly excludes benign clients under non-IID conditions. Fine-grained methods instead analyze data at the sample level for more precise detection, but typically rely on explicit per-sample gradient analysis, leading to substantial memory and computational overhead. As a result, defending against backdoor attacks in non-IID environments involves a fundamental trade-off between robustness and computational efficiency. To address this challenge, we propose FedRGD, a federated risk-guided dynamic defense framework that enables efficient fine-grained protection. FedRGD maps sample-level risks into structured parameter masking without requiring explicit per-sample gradient storage. It combines feature inconsistency detection with lightweight masking and robust aggregation to achieve both accuracy and efficiency. Extensive experiments on CIFAR-10 and Fashion-MNIST demonstrate that FedRGD consistently reduces the attack success rate while maintaining high main-task accuracy, achieving a favorable security-utility balance with low computational overhead.
    Keyword: Federated Learning, Backdoor Defense, Dynamic Masking, Non-IID Data, Efficient Defense.
    DOI: 10.65286/icic.v22i2.38793
    Cite

  • ContrastSFT: Contrastive Logit Regularization Supervised Fine-Tuning for Mitigating Hallucinations in Large Language Models, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Jingzhao Gu, Hongmin Xiao, Hangyu Li, Qiwei Wang
    Abstract: Large Language Models (LLMs) have achieved remarkable success in natural language generation but remain prone to hallucinations—generating content that is fluent but factually incorrect. While recent inference-time interventions like Contrastive Decoding (CD) effectively mitigate this by penalizing tokens favored by a "weak" hallucination-prone model, they introduce significant computational overhead (doubling inference latency) and fail to permanently align the model. In this paper, we propose \textbf{ContrastSFT}, a novel training framework to mitigate hallucinations in LLMs that internalizes the efficacy of contrastive decoding into the model's parameters via Contrastive Logit Regularization (CLR). Unlike standard Supervised Fine-Tuning (SFT) which indiscriminately maximizes the likelihood of ground-truth tokens, ContrastSFT dynamically recalibrates the training objective by subtracting the log-probabilities of a weak reference model. This effectively penalizes "easy" but potentially hallucinatory patterns captured by the weak model, forcing the model to learn more robust, factual representations. Extensive experiments on NLU benchmarks (ParaRel, WiCE) and Factuality tasks (HaluEval, MMLU) demonstrate that ContrastSFT achieves a 5-9\% absolute improvement over SFT and previous contrastive methods. Crucially, ContrastSFT eliminates the need for auxiliary models during deployment, retaining the high inference efficiency of standard LLMs. Code will be released.
    Keyword: large language models, hallucination mitigation, contrastive learning, supervised fine-tuning, contrastive decoding, factuality evaluation
    DOI: 10.65286/icic.v22i1.89404
    Cite

  • PointLGGS: Parallel Local-Global Dual-Path for Efficient Point Cloud Classification, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: min yuan, jiale li, Qiuhai Zhao, Lei Pan, Hongyong Leng
    Abstract: Despite significant advances in point cloud-based 3D vision research, existing methods often fail to achieve efficient collaborative modeling of both fine-grained local geometry and long-range global semantics, which limits further gains in classification accuracy. To address this limitation, this paper proposes a lightweight point cloud classification network named PointLGGS. Its three core contributions are summarized as follows. First, a Local Geometric Feature Extractor (LGFE) is designed to accurately capture fine-grained structures within local point cloud patches through relative position encoding and max-pooling operations. Second, a Global Semantic Feature Extractor (GSFE) is constructed by integrating dual mechanisms: Patch-Point Self-Attention (PP-SA) and Channel Self-Attention (CSA). This design jointly models long-range dependencies between patches and contextual correlations across channel dimensions, enabling efficient global semantic extraction. Third, using the LG-GS block as a fundamental building unit, a dual-path parallel architecture is employed to couple the LGFE and GSFE. Deep integration of local and global features is achieved via adaptive fusion, resulting in robust point cloud representations. Experimental results show that PointLGGS achieves classification accuracies of 94.0% and 90.7% on the ModelNet40 and ScanObjectNN datasets, respectively, with only 4.6 million parameters, thereby striking an excellent balance between classification accuracy and model efficiency.
    Keyword: point cloud classification, lightweight network, dual-path architecture, local-global feature fusion, attention mechanism
    DOI: 10.65286/icic.v22i1.10584
    Cite

  • A Multi-cycle Spatio-temporal Hypergraph and Dual-attention Residual Hypergraph Neural Network for Saturation Attack Detection in SDN, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Sheng Lin, Xianrong Yang, Yi Chen, Chun Guo, Guowei Shen, Yunhe Cui
    Abstract: Saturation attacks in Software-Defined Networking (SDN) have evolved into stealthy attack patterns, especially the Down-to-Up Timeout Probing and Dynamically Flow Table Overflowing (DUDFTO) attack. Existing GNN-based attack detection methods have limitations in modeling complex traffic relationships. Connecting flows via shared attributes easily mixes benign and malicious features, inadvertently mix conflicting information of benign flows and attack flows. Meanwhile, existing Hypergraph Neural Networks (HGNNs)-based attack detection methods may lose cross-hyperedge and high-order correlation information by simply gathering flow node features as hyperedge features. To overcome the above limitations, this paper proposes ST-HGNN, a saturation attack detection method based on a multi-cycle spatio-temporal hypergraph and a dual-attention residual hypergraph neural network. ST-HGNN constructs ST-FlowGraph, a multi-cycle spatio-temporal hypergraph. Each node represents a network flow. The hypergraph contains three kinds of hyperedges: Source IP-Protocol hyperedges, destination port group hyperedges, and temporal window hyperedges. We further design ST-HGATNet, a dual-attention residual hypergraph neural network model. ST-HGATNet integrates a flow-level node learning module, a group-level hyperedge enhancement module, and a residual fusion and classification module to improve attack detection performance to refine node representations, capture group-level dependencies, and preserve critical and subtle anomalies information. Evaluation results demonstrate the detection effectiveness of ST-HGNN, achieving a detection accuracy exceeding 97%.
    Keyword: SDN, DUDFTO Attack, Hypergraph
    DOI: 10.65286/icic.v22i2.18897
    Cite

  • FedKeyMIS: Keyed Federated Multi-image Steganography for Selective Extraction and Crosstalk Reduction, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Longshun Hu
    Abstract: Diffusion-based image steganography can produce visually natural stego images, but most existing pipelines inject secret features uniformly across spatial locations and ignore that natural images do not tolerate perturbations equally across space. This becomes especially problematic for blind extraction: smooth regions reveal even weak perturbations, whereas textured regions can absorb stronger signals and remain decodable after distortion. We address this issue with AURA-Stega, an uncertainty-guided region-adaptive diffusion steganography framework for robust blind extraction. AURA-Stega estimates a latent-space uncertainty prior from the cover image and uses it to modulate secret-feature energy before deterministic diffusion inversion. As a result, the stego residual is concentrated in high-tolerance regions, while visually fragile regions are protected. We also consider asymmetric blind extraction, where the receiver recovers the message without access to the original cover image or the uncertainty map. Qualitative evidence shows that the uncertainty prior aligns with texture-rich regions and produces the expected residual pattern. To support this mechanism directly, we introduce a residual-concentration diagnostic in addition to standard recovery metrics.
    Keyword: image steganography; federated learning; multi-image hiding; selective extraction; crosstalk reduction
    DOI: 10.65286/icic.v22i2.38800
    Cite

  • TSLP: Text driven and style transferable indoor furniture layout generation pipeline, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Zhiguo Xu, Jian Zhang, Shenghao Yang, Xiaoyan Liu, Mingbo Zhao
    Abstract: With the surging demand in interior design, high-quality furniture layout as-sets have become increasingly indispensable. However, creating such assets requires professional expertise in both aesthetic design and 3D modeling, leading to significant manual overhead. Consequently, there is a compelling need for generative frameworks capable of autonomous, high-quality asset generation. Existing methods focus on holistic image-to-3D synthesis or as-set retrieval, yet they yield monolithic representations that lack instance-level editability and stylistic steerability. To this end, we present TSLP, a text-driven and style-transferable pipeline for indoor furniture layout generation. Our framework first synthesizes an interior image from textual prompts via a diffusion-based model, followed by a style-transfer module for aesthetic customization. We then reconstruct decoupled 3D assets with precise pose estimation, finally employing a texture synthesis model to bake high-fidelity textures onto the generated objects. Our pipeline enables seamless 3D furniture layout synthesis from text, granting users granular control over object structure, category, and spatial positioning. By integrating a style-transfer module, our framework facilitates effortless aesthetic customization via a single style reference image, significantly enhancing both usability and adaptability.
    Keyword: 3D Generation, furniture Layout Generation, Style Transfer, Text-driven Generation, Instance-level Control
    DOI: 10.65286/icic.v22i1.68723
    Cite

  • ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Chaodong jia
    Abstract: Text-Based Person Search (TBPS) aims to retrieve pedestrian images using natural language queries. However, existing TBPS models, especially those based on CLIP, struggle with fine-grained understanding due to global representational bias and semantic sparsity inherited from training on short captions. This results in weak fine-grained alignment, exacerbated by the scarcity of region-level annotations. To address this, we propose ROGLE (Robust Global-Local Embedding), a unified framework that overcomes reliance on costly manual annotations through an automated Region-to-Sentence Matching (RSM) strategy. RSM automatically mines pseudo region-sentence pairs for scalable fine-grained supervision. Furthermore, ROGLE employs a multi-granular learning strategy that fuses global contrastive learning with region-level local alignment. We also introduce the P-VLG Benchmark, a large scale dataset constructed by curating and enriching images from established public benchmarks . It features over 100,000 annotated regions and rich long-form captions, making it the first TBPS benchmark to support both global and local assessment protocols. Extensive experiments show that ROGLE significantly outperforms existing approaches, particularly on challenging long-form queries. Code and the P-VLG benchmark will be made publicly available.
    Keyword: P-VLG Benchmark · fine-grained alignment · Text-Based Person Search
    DOI: 10.65286/icic.v22i2.99720
    Cite

  • Koi Care Robots: A Framework and Practical Applications, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: prawit chumchu, kailas patil, alfa nyandoro
    Abstract: This study proposes a robotic framework for intelligent koi pond management, aimed at enhancing the care and well-being of koi in controlled aquatic environments. The robotic platform is designed with multiple integrated subsystems, including a water filtration unit, real-time koi monitoring, water quality sensing and analysis, joystick-enabled man-ual operation, and automated navigation. The prototype integrates several sensing com-ponents, including a 9-axis inertial measurement unit (IMU), an underwater distance sensor, an underwater camera, and a 2D LiDAR system. The IMU supplies continuous feedback on orientation and dynamic stability, supporting precise navigation and control performance. The underwater distance sensor facilitates real-time detection of sub-merged obstructions to ensure safe underwater operation. Visual data acquired by the underwater camera are processed using deep learning–based classification algorithms for object recognition. Additionally, the 2D LiDAR sensor detects and maps above-water obstacles, supporting reliable navigation and environmental awareness. The robotic framework is implemented using ROS2, which manages system integration, communication, and control. The water purification subsystem consists of four pumps that intake pond water and channel it through an integrated filtration module. Real-time water quality evaluation is achieved through deep learning–based analysis and classification of underwater imagery. Furthermore, the system is integrated with an automated feeding machine that dispenses floating pellets using deep learning. In addition, the system inte-grates EC, pH, and dissolved oxygen (DO) sensors to continuously monitor key water quality parameters essential for koi health and aquaculture productivity via Bluetooth communication. The robot is powered by a solar energy system, enabling autonomous and sustainable operation. Experimental results demonstrate the robustness and reliability of the framework in real-world conditions.
    Keyword: Artificial intelligent, Kois, Ros, Robots, Water quality, Koi Care, Koi feeding, Pet Care.
    DOI: 10.65286/icic.v22i2.84099
    Cite

  • IPMT: An Identity-Preserving Makeup Transfer Model, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Haodong Song, Yinglin Zheng, Yuxin Lin, Pingping Gu, Wangzheng Shi, Wentao Chen, Ruolong Ma, Jianan Lin, Ming Zeng
    Abstract: Makeup transfer serves as an important application in the field of image processing.Makeup transfer aims to transfer a reference makeup style to a source image, thereby synthesizing a high-fidelity facial image. However, when processing heavy or artistic makeup, existing generative methods frequently corrupt the identity-specific geometric features of the face. To address this issue, we propose an Identity-Preserving Makeup Transfer (IPMT) model, which is implemented via an end-to-end self-supervised framework. Specifically, we construct a multi-stream encoding pipeline and integrate a Dynamic Fusion Block to perform pixel-level adaptive fusion, achieving rigorous alignment between makeup semantics and target spatial geometry. Furthermore, to effectively prevent structural deformation without relying on pseudo-paired data, a one-step denoising approximation strategy is introduced during the training phase. By incorporating a parsing consistency loss and a 3D vertex loss during the early timesteps of generation, our method imposes explicit constraints on both the 2D planar layout and the 3D topological structure of the face. Extensive experiments demonstrate that the proposed IPMT significantly mitigates the identity inconsistency problem while ensuring high-fidelity makeup transfer.
    Keyword: makeup transfer,diffusion model,identity preservation,computer vision
    DOI: 10.65286/icic.v22i1.33460
    Cite

  • AI-Generated Disaster Image Detection: A Dataset and an Exploration of Key Challenges, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Junjie Wang, Rui Ba, Kaiye Yu, Han Xing, Chao Sun, Hang Gao, Mengting Hu
    Abstract: The rapid advancement of generative artificial intelligence (AI) has enabled the creation of highly realistic disaster imagery, posing significant threats to information authenticity during crisis events. To address this emerging challenge, we present AIG-DI, the first dataset specifically designed for detecting AI-generated disaster images. The dataset is systematically constructed across multiple disaster types, image content categories, and generative models, enabling comprehensive evaluation under diverse conditions. We conduct an in-depth empirical study using two representative detectors—NPR (a spatial-domain detector) and FreDect (a frequency-domain detector)—to investigate their detection performance, generalization ability, and robustness under realistic perturbations. Experimental results reveal three key findings: (1) detectors trained on generic datasets struggle to detect AI-generated disaster imagery due to domain-specific texture and semantic shifts; (2) incorporating even a small amount of disaster-specific synthetic images during training significantly boosts accuracy and generalization, highlighting the value of domain-specific data for rapid adaptation; and (3) image perturbations remain a critical vulnerability, even with perturbation-aware training. This work not only provides the first benchmark for AI-generated disaster image detection but also uncovers fundamental challenges in ensuring visual content authenticity for disaster response. The code and data will be made publicly available upon acceptance of the paper.
    Keyword: AI-generated disaster image, AI-generated disaster image detection, Benchmark
    DOI: 10.65286/icic.v22i2.60827
    Cite

  • QM-ToT: A Medical Tree of Thoughts Reasoning Framework for Quantized Model, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Zongxian Yang, Jiayu Qian, Kay Chen Tan, Hau-San Wong, Yulong Chen, Haoyu Zhang, Zhi-An Huang
    Abstract: Large language models (LLMs) have achieved substantial progress in biomedical question answering. However, real-world medical applications often operate under resource-constrained settings, where model quantization is a practical necessity for local and privacy-preserving deployment. In such settings, the inherent complexity of clinical reasoning further amplifies the performance degradation of quantized LLMs. To address these issues, we propose Quantized Medical Tree of Thought (QM-ToT), an agentic reasoning framework that enables autonomous, feedback-driven clinical reasoning. QM-ToT operates as a reasoning agent that decomposes complex medical problems into hierarchical plans and autonomously navigates the solution space under dual-evaluation feedback. This framework facilitates substantial performance improvements in INT4-quantized models on the challenging MedQA-USMLE dataset. Specifically, we demonstrate a remarkable accuracy increase from 34\% to 50\% for the LLaMA2-70b model and from 58.77\% to 69.49\% for LLaMA-3.1-8b. Besides, we also proposed an effect data distillation method based on QM-ToT. Compared to the traditional distillation method, we achieved an improvement of 135.7\% while using only 9.8\% of the data. This work, for the first time, showcases the potential of ToT to significantly enhance performance on complex biomedical tasks, establishing a crucial foundation for future advances in deploying high-performing quantized LLM in resource-limited medical settings.
    Keyword: tree of thought, large language model, model quantization, medical question answering, healthcare
    DOI: 10.65286/icic.v22i1.75267
    Cite

  • GenCTL: Constraint-Guided LLM Generation of Verifiable CTL Specifications from Natural Language Requirements, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Ran Tao, Yanhong Huang, Jianqi Shi, Yang Yang, Fan Gao
    Abstract: Natural language (NL) requirements are widely used in system design, but their ambiguity makes formal verification difficult. Translating NL requirements into Computation Tree Logic (CTL) is challenging because generated formulas must be both semantically appropriate and compatible with concrete verification models. Although large language models (LLMs) provide a promising basis for NL-to-CTL(NL2CTL) translation, their outputs are often unstable and prone to parsing, typing, and grounding errors. To address these issues, we propose GenCTL, a prompt-based framework for effective and model-aware NL2CTL translation without task-specific fine-tuning. GenCTL combines structured prompting, retrieval-enhanced few-shot examples, atomic-proposition (AP) grounding from nuXmv models, multi-candidate generation with frequency-first selection and length-normalized log-likelihood tie-breaking, and interactive refinement through an explanation dictionary. Experimental results demonstrate the effectiveness of the proposed framework in both model-agnostic and model-aware settings. On a generated NL--CTL dataset, the best automatically selected translation achieves 68% accuracy, which increases to 92% after one round of user refinement. On 150 NL requirements grounded in three nuXmv models, AP-list prompting yields 139/150 directly checkable formulas, increasing to 149/150 after lightweight normalization. These results show that GenCTL improves the reliability and practical checkability of LLM-generated CTL specifications.
    Keyword: Natural language to CTL , Large language models , Prompt engineering , Model-aware grounding , Formal verification.
    DOI: 10.65286/icic.v22i1.71344
    Cite

  • S2NO: Coupling Aliasing-Free Convolutions with Fourier Neural Operators for Multiscale Modeling, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Linfeng Liu, Xinhai Chen, Qinglin Wang, Jing Xiao, Xiang Zhang, Jie Liu
    Abstract: Neural operators have emerged as a powerful paradigm for solving Partial Differential Equations. They learn mappings between infinite-dimensional function spaces and have been successfully applied in fields such as computational fluid dynamics and weather forecasting. However, existing methods face a fundamental dilemma. The Fourier Neural Operator is efficient at capturing global patterns but suffers from spectral leakage. Conversely, hybrid models combine spectral methods with standard convolutions but often introduce aliasing errors. This violates the critical property of resolution invariance. To address this, we propose the Spectral-Spatial Neural Operator. We introduce a dual-stream architecture that couples a spectral branch with an aliasing-free convolutional branch. This design allows the model to capture high-frequency residuals and sharp discontinuities without introducing grid-dependent artifacts. We conducted extensive experiments on the 1D Burgers, 2D Darcy Flow, and 2D Navier-Stokes equations. The results demonstrate that our method significantly outperforms state-of-the-art baselines in both prediction accuracy and zero-shot super-resolution stability.
    Keyword: Neural Operators, Partial Differential Equations, Operator Learning, Anti-Aliasing
    DOI: 10.65286/icic.v22i2.50475
    Cite

  • Uncertainty-Aware Debiased Recommendation: A Bayesian Doubly Robust Perspective, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Yang Fei, Zichi Zhang, Hualing Liu
    Abstract: Recommender systems are inherently plagued by selection bias arising from the Missing Not At Random (MNAR) nature of observed interactions. Doubly Robust (DR) learning, which integrates Inverse Propensity Scoring (IPS) and Error Imputation-Based (EIB) models, has emerged as a cornerstone for bias mitigation. However, traditional DR methods rely heavily on the precise point estimates of propensities and imputations, making them highly susceptible to catastrophic variance explosion under model misspecification or extreme data sparsity. To address these limitations, we propose the Bayesian Doubly Robust (BDR) learning framework. This framework shifts the debiasing paradigm from rigid point estimation toward robust distributional inference by leveraging Monte Carlo Dropout to capture epistemic uncertainty and introducing an Entropy Tilting mechanism. By minimizing Kullback-Leibler (KL) divergence within the function space, BDR dynamically calibrates posterior samples to satisfy causal unbiasedness moment conditions, thereby rectifying imputation errors without the need for model retraining. Theoretical analysis demonstrates the asymptotic unbiasedness of BDR even when the "accurate imputation assumption" is relaxed. Extensive evaluations on the Coat, Yahoo!, and KuaiRec datasets confirm that BDR excels in extremely sparse scenarios and functions as a model-agnostic, "plug-and-play" framework that consistently enhances the performance of state-of-the-art (SOTA) models.
    Keyword: Recommender Systems; Selection Bias; Doubly Robust Learning; Bayesian Uncertainty
    DOI: 10.65286/icic.v22i2.11824
    Cite

  • Hierarchical Prompt for Task-Adaptive Composed Image Retrieval, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Zeli Yan, Guosun Zeng
    Abstract: Composed image retrieval (CIR) seeks to retrieve target images using multi-modal queries, specifically a reference image paired with modification text. Central to CIR is integrating textual semantic modifications with visual content. Despite its importance, existing approaches typically employ a static fusion paradigm, failing to account for the semantic heterogeneity of user queries, which encompass diverse task types (e.g., addition, replacement) and var-ied content. To address these limitations, we propose the Task-Adaptive Hier-archical Prompt (TAHP) framework. TAHP guides feature extraction through dynamically generated, task-specific prompts structured at three hierarchical levels: task-type, task-content, and general prompts. Furthermore, we design a Prompt Dynamic Generation Module to adaptively synthesize prompts condi-tioned on user queries and introduce a False Negative Correction Loss to optimize cross-modal feature fusion. Extensive experiments on FashionIQ and CIRR datasets demonstrate that TAHP achieves state-of-the-art performance against existing CIR approaches.
    Keyword: Composed image retrieval, Hierarchical prompt learning, Multimodal semantic alignment.
    DOI: 10.65286/icic.v22i2.76973
    Cite

  • Consistency-Aware Gated Fusion with Mamba for Multimodal Sentiment Analysis, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Jian Hu, Bo Lei, JinHua Sun
    Abstract: Multimodal sentiment analysis has attracted increasing attention due to the prevalence of text-image content on social media. A central challenge is to design fusion mechanisms that are both expressive and parameter-efficient, especially for small-scale datasets where heavy cross-modal attention can easily overfit. In this paper, we present Consistency-Aware Gated Fusion (CAGF), a lightweight and fusion module tailored to Mamba-based architectures. Our key idea is to exploit Mamba's bidirectional scanning mechanism: forward and backward hidden states from text and image encoders are concatenated to form enhanced representations, and a cosine-based semantic consistency score is computed between modalities. This score is then passed through a fixed sigmoid gate to adaptively weight text and image features, without introducing any additional learnable parameters. CAGF is plug-and-play compatible with dual-stream Mamba encoders and incurs negligible computational overhead compared with attention-based fusion. Experiments on the MVSA-Single dataset show that CAGF achieves state-of-the-art performance (Acc=82.54%, F1=84.82%), outperforming strong multimodal baselines such as CLIP, MISA, DLF, AoM, and SFTTR, while remaining more efficient and interpretable. Extensive ablations and sensitivity analyses further validate that bidirectional scanning, enhanced representations, and consistency-aware gating are all critical to the observed gains.
    Keyword: Multimodal sentiment analysis has attracted increasing attention due to the prevalence of text-image content on social media. A central challenge is to design fusion mechanisms that are both expressive and parameter-efficient, especially for small-scale datasets where heavy cross-modal attention can easily overfit. In this paper, we present Consistency-Aware Gated Fusion (CAGF), a lightweight and fusion module tailored to Mamba-based architectures. Our key idea is to exploit Mamba's bidirectional scanning mechanism: forward and backward hidden states from text and image encoders are concatenated to form enhanced representations, and a cosine-based semantic consistency score is computed between modalities. This score is then passed through a fixed sigmoid gate to adaptively weight text and image features, without introducing any additional learnable parameters. CAGF is plug-and-play compatible with dual-stream Mamba encoders and incurs negligible computational overhead compared with attention-based fusion. Experiments on the MVSA-Single dataset show that CAGF achieves state-of-the-art performance (Acc=82.54%, F1=84.82%), outperforming strong multimodal baselines such as CLIP, MISA, DLF, AoM, and SFTTR, while remaining more efficient and interpretable. Extensive ablations and sensitivity analyses further validate that bidirectional scanning, enhanced representations, and consistency-aware gating are all critical to the observed gains.
    DOI: 10.65286/icic.v22i2.76246
    Cite

  • Spatiotemporal Multimodal Interaction for Anomaly Detection in Microservices, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: ziwei yang
    Abstract: With the widespread adoption of microservice architectures, anomaly detection has become a key technique to maintain the stable operation of microservice systems and an important research focus in intelligent operations. However, existing methods often rely on single-modal data or simply concatenate metrics, logs, and traces without fully exploring the relationships between modalities. In particular, they tend to overlook spatiotemporal dependencies, making it difficult to comprehensively capture how anomalies evolve in complex systems, which in turn leads to missed detections and false alarms. To address these limitations, this paper proposes STMID (Spatiotemporal Multimodal Interaction for Microservice Anomaly Detection), a spatiotemporal multimodal interaction method for anomaly detection in microservices. The proposed method introduces a service dependency graph to provide a unified representation of multimodal time-series data in microservice systems. It further uncovers cross-modal correlations through multimodal fusion, and models spatiotemporal dependencies to effectively capture anomalous patterns in both temporal dynamics and service dependencies. Based on spatiotemporal dependency modeling results, we use the reconstruction error of the variational autoencoder to measure anomalies, enabling binary classification between normal and abnormal states. Experimental results on two datasets, MSDS and GAIA, show that the proposed method achieves strong overall performance, with an average F1-score of 0.879. It outperforms most baseline methods, with improvements ranging from 3.85% to 42.95%.
    Keyword: Microservices; Multimodal fusion; Deep learning; Anomaly detection
    DOI: 10.65286/icic.v22i2.50235
    Cite

  • Parameter-Efficient Remote Sensing Image-Text Retrieval via Hierarchical Gated Multi-Modal Adapters, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Xin Pan, Xiaohui Huang, Wenhai Li, Xuebo Cheng
    Abstract: This paper proposes a parameter-efficient fine-tuning framework based on a Hierarchical Gated Multi-modal Adapter (HGMA) for remote sensing image-text retrieval. To address the limitations of traditional models and the high computational costs of full-parameter fine-tuning, this method introduces a multi-head self-attention module and a hierarchical gating mechanism to dynamically regulate cross-modal features across different depths. Furthermore, a Cross-modal Semantic Perturbation (CMSP) data augmentation strategy is designed to generate semantically consistent pseudo-samples, effectively enhancing the model's resistance to noise interference and domain shift. Extensive experiments on RSICD and RSITMD benchmark datasets demonstrate that the proposed method achieves performance comparable to full-parameter fine-tuning while training only about 0.6M parameters.
    Keyword: Remote Sensing Image-Text Retrieval, Parameter-Efficient FineTuning, Hierarchical Gated Multi-modal Adapter, Data Augmentation.
    DOI: 10.65286/icic.v22i2.56417
    Cite

  • ProFuseGPT: Progressive Fusion with Contrastive Refinement for Long-Sequence Medical Report Generation, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Shaowei Shen, Jie Yang, Lianfen Huang, Shanhao Zhan, Zexin Huang, Zhibin Gao, Xiaohong Yang
    Abstract: Automatic medical report generation (MRG) holds promise for alleviating radiologists’ workload, which has spurred growing interest in MRG for stroke diagnosis. However, existing approaches often fail to effectively model its long-range spatial dependencies and suppress phase-level noise in multi-slice sequences, leading to diluted pathological signals and unstable cross-modal alignment. To address this, we propose a novel framework integrating a progressive fusion mechanism (PFM) and intermediate state refinement via contrastive alignment (ISRCA), inspired by radiologists' clinical workflow. PFM progressively refines pathological representations through anatomically deviation-aware adaptive weighting, suppressing noise from normal slices while enhancing salient abnormalities from local deviations to global context integration. ISRCA adopts a teacher-student distillation approach using contrastive learning to mitigate noise propagation and stabilize intermediate report features. Experimental results on two stroke imaging datasets demonstrate that our method outperforms existing approaches in natural language generation (NLG) metrics, highlighting the effectiveness of PFM and ISRCA in handling long-sequence medical images and advancing stroke imaging report generation.
    Keyword: Report Generation, Teacher-student Distillation, Cross-modal Alignment, Large Language Models.
    DOI: 10.65286/icic.v22i1.50370
    Cite

  • Improving End-to-End Argument Mining with LLMs via Multi-Candidate Reranking and Bounded Repair, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Chen Tang, Min Peng, Gang Tian
    Abstract: End-to-end Argument Mining (AM) aims to extract structured argument graphs directly from raw text by jointly identifying components and their rhetorical relations. Large Language Models (LLMs) have shown promise for this complex task by formulating it as direct text-to-graph generation. However, since most current LLM-based approaches rely on a single-pass generation paradigm, they struggle to handle the tight coupling of AM subtasks. Consequently, early extraction errors easily cascade, resulting in fragmented relations and globally inconsistent structures. Critically, these single-pass models lack mechanisms to backtrack and correct upstream mistakes when downstream inconsistencies arise. To address this problem, we propose the Multi-Candidate Reranking and Bounded Repair (MCR-BR) framework for LLM-based end-to-end AM. Instead of committing to a single output, our pipeline first generates diverse component candidates to preserve alternative hypotheses. It then constructs full argument graphs for these candidates and ranks them using a consistency-guided scoring strategy based on structural completeness, semantic coherence, and validity. For any remaining structural anomalies, a bounded local correction loop actively repairs the graph. We further introduce a targeted optimization module for conclusion-like components, which are common extraction bottlenecks. Experiments on the AAEC and CDCP benchmarks demonstrate that the MCR-BR framework consistently improves in-domain extraction and exhibits potential for cross-domain generalization.
    Keyword: argument mining、large language models、self-correction、error propagation、structured prediction
    DOI: 10.65286/icic.v22i1.76896
    Cite

  • A Semantic Segmentation Method for Flame and Water Stream Landing Points Based on Dual-CBAM-MobileNet, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Chongshuo Liu, Yaojie Chen
    Abstract: Aiming at the problems of insufficiently rapid fire identification and imprecise fire suppression in large-space fire protection systems, this study designs a Dual-CBAM-MobileNet (dual attention-driven feature extraction network) module based on the encoder-decoder architecture of DeepLabV3+ to achieve real-time and accurate segmentation of flames and water stream landing points. Firstly, a lightweight MobileNetV2 feature extraction network is adopted as the backbone. Secondly, the Convolutional Block Attention Module (CBAM) is introduced and embedded into the feature extraction network to enhance the perception capability for flames and water stream landing points. Finally, to further mitigate the class imbalance problem, Focal Loss is employed as the loss function. Experiments demonstrate that the proposed algorithm improves the image segmentation accuracy for both flames and water stream landing points while maintaining a high image processing speed. On the self-constructed flame-water stream landing point dataset, it achieves a mean Intersection over Union (mIoU) of 76.23% and a processing speed of 78.68 frames per second (FPS), meeting the dual requirements for real-time performance and accuracy in practical fire protection systems.
    Keyword: Semantic Segmentation; DeepLabV3+; Attention Mechanism; MobileNetV2; Lightweight
    DOI: 10.65286/icic.v22i1.27558
    Cite

  • AdaDyTS: Dynamic Multi-Scale Spectral Decoupling and Time-Variant Inference for Time Series Forecasting, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Jinmei Qi, Jinlai Zhang, Qiyu Wu, Zhuo Dong, Leyun Kong, Liwen Hu
    Abstract: Time series forecasting is fundamental to intelligent decision-making systems, enabling proactive planning and resource optimization across diverse application domains. However, the inherent complexity of real-world time series—including multi-scale temporal patterns, heterogeneous variable dependencies, and dynamic non-stationarity—poses significant challenges for existing forecasting models. Current approaches often suffer from high-frequency information attenuation in frequency-domain modeling, inadequate characterization of scale heterogeneity across variables, and limited capability to capture time-varying dynamics. To address these challenges, this paper introduces AdaDyTS, a unified knowledge-driven forecasting framework that synergistically integrates three complementary mechanisms: multi-scale frequency-domain interpolation decoupling via the Cascaded Spectral Residual Extractor (CSRE), dynamic morphological perception via the Dynamic Morphological Perception Unit (DMP-U), and time-variant state-space inference via the Time-Variant State-Space Module (TV-SS). CSRE separates low-frequency trends from high-frequency residuals through coarse-to-fine layer-wise self-reconstruction, preserving transient information that static filters typically attenuate. DMP-U employs deformable convolution guided by multi-expert attention to adaptively adjust receptive fields, enabling fine-grained modeling of local fluctuations and nonlinear distortions. TV-SS relaxes the conventional time-invariant parameter assumption, dynamically modulating state transition parameters to capture both short-term variations and long-term dependencies. Under a unified evaluation protocol across 13 benchmark datasets, AdaDyTS achieves average improvements of 4.35\% in MSE and 4.31\% in MAE over the AMD backbone, consistently outperforming state-of-the-art methods across long-horizon forecasting scenarios. The proposed framework demonstrates the effectiveness of integrating domain-specific knowledge—including spectral analysis, morphological feature extraction, and dynamic system modeling—within a unified deep learning architecture for enhanced predictive performance.
    Keyword: Time Series Forecasting \and Multi-scale frequency-domain Interpolation \and Dynamic Morphological Perception \and Time-Variant State-Space
    DOI: 10.65286/icic.v22i2.45765
    Cite

  • Deep Learning-Based Terahertz Phased Array System Beam Compensation Algorithm for Space Radiation Environments, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Chensheng Ma, Yuanzhi He, Rongbing Chen
    Abstract: Inter-satellite terahertz (THz) communications, enabled by abundant spectrum and highly directional beams, are a key technology for future space information net-works. However, the extremely narrow beamwidth of THz signals makes the link highly sensitive to beam misalignment. In practical spaceborne THz phased ar-rays, space radiation and inherent hardware imperfections jointly cause radiation pattern distortion and additional pointing errors, leading to severe degradation in link performance. Although traditional optimization methods, such as genetic al-gorithms and convex optimization, have been shown to be effective for beam compensation, they often suffer from local optima, high computational complexi-ty, long convergence time, and limited adaptability. To address these challenges, we propose a deep learning-based beam compensation algorithm for radiation-affected THz phased array systems. We develop a health-aware Res-UNet (HA-Res-UNet) compensation network that leverages array-state feedback for dynam-ic beam correction. We introduce a multi-scale health-mask attention module that embeds element-health priors into feature aggregation to suppress failed elements and emphasize compensable regions. Simulation results demonstrate that the pro-posed method achieves approximately 90.4% pointing error reduction, improves the average pointing accuracy by about 10.4 times, and effectively mitigates radia-tion pattern distortions caused by failed array elements.
    Keyword: terahertz, deep learning, beam compensation, space radiation environments
    DOI: 10.65286/icic.v22i2.27714
    Cite

  • CARE-Bench: A Capability-Stratified, Cost-Aware Benchmark for Agentic Software-Security Detection, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Andi Xia
    Abstract: Agentic systems for software-security detection have advanced quickly, from passive classifiers to cyber reasoning systems that autonomously find and patch vulnerabilities. Yet progress is hard to measure: systems are reported on incomparable datasets, with metrics that ignore cost, and on corpora whose labels are known to be noisy and prone to train-test contamination. We present CARE-Bench, an evaluation harness and protocol that makes results across heterogeneous agentic detectors directly comparable. CARE-Bench makes three design commitments. First, it is capability-stratified: every system is scored along the four levels of detection, localization, triage, and remediation, so that a single-call classifier and a multi-agent patcher are placed on the same ladder. Second, it treats cost as a first-class metric, reporting tokens, dollars, and wall-clock time, and dollars per true positive, so that accuracy and expense are reported together. Third, it is contamination-resistant, providing temporal-holdout and de-duplication controls that address the documented label-noise and leakage problems of existing corpora. CARE-Bench does not introduce a new vulnerability dataset; instead it unifies existing public corpora under one canonical schema through dataset adapters, and ships as open-source code with a reproducible runner, a metric suite, detector interfaces, and a leaderboard schema that rejects submissions omitting cost or contamination provenance. We describe the design, position it against recent benchmarks, and release the harness to support reproducible, cost-aware evaluation of the next generation of security agents.
    Keyword: Benchmark, Vulnerability Detection, LLM Agents, Reproducibility, Cost-Aware Evaluation, Open Source
    DOI: 10.65286/icic.v22i2.95846
    Cite

  • A Taxonomy of Agentic Systems for Software Security Detection, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Andi Xia
    Abstract: Software vulnerabilities have become a recognized national-security risk, yet the volume of code and the sophistication of threats now far outpace what manual security review and the limited supply of expert security engineers can sustain. A new class of systems has emerged in response: agentic systems for software security detection, which couple large language models with planning, memory, and external tools so that they can autonomously analyze codebases, reason about program behavior, and identify, triage, and help re-mediate vulnerabilities. The field has grown rapidly but unevenly, and its terminology, capabilities, and evaluation practices remain fragmented. This paper organizes the area into a structured taxonomy along five axes: the de-tection capability targeted, the analysis paradigm employed, the agent archi-tecture, the degree of autonomy, and the evaluation methodology. We popu-late the taxonomy with representative systems, including the cyber reasoning systems demonstrated at the DARPA AI Cyber Challenge, and we use it to compare designs, surface recurring patterns, and expose gaps. We find that the strongest results combine learned reasoning with classical program analy-sis and tool use rather than relying on either alone, and that repository-scale detection, trustworthy triage, and reproducible evaluation remain the principal open challenges. The taxonomy is intended as a shared vocabulary and a roadmap for building the next generation of autonomous software-security systems.
    Keyword: Software Security, Vulnerability Detection, LLM Agents, Agentic Systems, Taxonomy, Program Analysis.
    DOI: 10.65286/icic.v22i2.40494
    Cite

  • DARE-RAG: Difficulty-Aware Retrieval Expansion for Retrieval-Augmented Generation, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Lixiang Zhu, Minghao Li, Zhubo Liu, Chao Wu
    Abstract: Retrieval-Augmented Generation (RAG) systems face a fundamental trade-off: query expansion can improve retrieval effectiveness for ambiguous or underspecified queries, yet indiscriminate expansion introduces unnecessary latency and retrieval noise. Existing RAG pipelines typically apply expansion uniformly, failing to distinguish between easy and retrieval-challenging queries.To address this issue, we propose DARE-RAG, an adaptive retrieval framework that activates LLM-based query expansion only for retrieval-challenging queries. Our method formulates expansion activation as a lightweight binary classification problem using probe retrieval signals, including score margin, variance, entropy, query length, and lexical specificity. A lightweight MLP predicts whether expansion is likely to improve retrieval quality, and expansion is triggered only when the predicted confidence exceeds a percentile-calibrated threshold.DARE-RAG further integrates a dual-path hybrid retrieval architecture combining BM25 sparse retrieval and BGE dense retrieval, fused via Reciprocal Rank Fusion (RRF), followed by a Cross-Encoder reranker for context refinement. Experiments on NQ-Open and HotpotQA demonstrate that DARE-RAG consistently improves retrieval effectiveness and end-to-end QA accuracy while clearly reducing average end-to-end latency compared with corresponding always-expand variants of BM25, BGE-m3, and their RRF-fused hybrid retriever. Extensive ablation studies and efficiency analyses verify the effectiveness of our utility-guided expansion strategy.
    Keyword: Retrieval-Augmented Generation,Adaptive Query Expansion,Query Difficulty Estimation,Hybrid Retrieval,Efficiency-Aware Retrieval
    DOI: 10.65286/icic.v22i1.52618
    Cite

  • GASR-Net: Geometric Anomaly Synthesis and Task-Oriented Feature Reconstruction for Industrial Anomaly Detection, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Luhao Li, Xiaoyang Shi, Shuqiang Gao
    Abstract: In real-world industrial manufacturing, automated visual inspection for surface defect detection is crucial for quality control but remains severely hindered by two critical challenges: the simplified assumption of isotropic anomaly synthesis that fails to capture complex physical defect shapes, and the over-smoothing effect in generative models that triggers unacceptable false alarms in textured backgrounds. To address these persistent industrial limitations, we propose a novel network, GASR-Net (Geometric Anomaly Synthesis and Task-Oriented Feature Reconstruction for unsupervised image anomaly detection and localization). Our architecture tackles the gap between synthesized noise and realistic flaw shapes by introducing a Geometric-Aware Deformable Noise Modulator (GAND-M). By utilizing baseline normal features to dynamically predict deformation offsets, the module warps isotropic Gaussian noise into structurally contiguous and geometric-diverse pseudo-anomaly flows. Simultaneously, to overcome the persistent issue of background over-smoothing and the resulting high false positive rates, we design a Task-Oriented Single-Step Feature Reconstructor based on a lightweight U-Net topology. Unlike conventional generative methods that perform blind standalone restoration, our constructor is driven by a joint feedback loop from the downstream discriminator, adaptively forcing the network to thoroughly eliminate abnormal artifacts while strictly preserving pixel-level background fidelity. Extensive experiments on two challenging industrial anomaly detection benchmarks, MVTec AD and VisA, demonstrate that GASR-Net achieves state-of-the-art performance in both detection and localization accuracy.
    Keyword: Unsupervised Anomaly Detection, Geometric Anomaly Synthesis, Deformable Noise Modulation, Feature Reconstruction, Task-Oriented Feedback Loop
    DOI: 10.65286/icic.v22i1.41986
    Cite

  • IA2former: Illumination-Aware Attention-based Transformer for Low-light Image Enhancement, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Tianqi Jiang, Danqing Ju, Han Wu, Ping Liang
    Abstract: Low-light image enhancement has made significant progress through both traditional Retinex methods and deep learning techniques. Traditional Retinex-based methods decompose images into illumination and reflectance components to mimic human perception of brightness and color. However, these methods often struggle with noise suppression and detail preservation, particularly under severe low-light conditions. Recent Transformer-based methods, such as RetinexFormer and Restormer, have improved restoration performance by modeling long-range dependencies, but they still insufficiently explore the interaction between illumination variations and spatial--semantic features. To address these limitations, we propose Illumination-Aware Attention-based Transformer (IA2former), a novel low-light image enhancement model that explicitly models illumination-aware feature interactions. By integrating an Illumination-Aware Attention mechanism and an Illumination-Aware Loss function, IA2former effectively captures long-range dependencies, improves detail restoration, and preserves spatial structures under challenging illumination conditions. Experimental evaluations on the LOL-v1 and LOL-v2 datasets demonstrate that IA2former achieves a favorable overall balance across PSNR, SSIM, and LPIPS, obtaining the best performance on multiple metrics and remaining competitive on others. These results validate the effectiveness and robustness of the proposed illumination-aware modeling strategy for low-light image enhancement.
    Keyword: Low-light Enhancement,Retinex and ,Transformer
    DOI: 10.65286/icic.v22i1.19847
    Cite

  • GeoFuse-SAM: A Multimodal Data Fusion Framework for Boundary-Aware Foundation Model Adaptation in Medical Image Segmentation, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Pengtao Ren, Qiyuan Wang, Yunyi Li, Kejiang Xiao, Lexi Shu
    Abstract: In medical image segmentation, accurately identifying anatomical structure boundaries is a core task for clinical computer-aided diagnosis. Although vision foundation models, represented by the Segment Anything Model (SAM), have demonstrated strong generalization potential, their deployment in fully automated clinical scenarios remains constrained by severe domain shift, reliance on manual prompts, and significant boundary degradation in ambiguous anatomical transition zones. To address these challenges, this paper proposes GeoFuse-SAM, a multimodal geometric data fusion adaptation framework. The core innovation of this framework lies in breaking the limitations of single semantic features. Through a Geometry-Guided Cross-Domain Attention Fusion (GCAF) module, it achieves deep data fusion between the raw image data and high-frequency geometric priors (Sobel gradient fields) derived from computer graphics. This cross-modal interaction mechanism provides explicit spatial guidance to the model, significantly enhancing the robustness of boundary recognition. Furthermore, we introduce a lightweight Parallel Fusion Adapter (PFA) to achieve medical semantic alignment, and propose a Parameter-Free Morphological Boundary Weighting (PMBW) strategy. This strategy utilizes morphological operators to pinpoint ambiguous boundary regions during the training phase and impose dynamic geometric constraints. Experiments on two challenging medical datasets, BUSI and ISIC 2018, demonstrate that GeoFuse-SAM, operating in a fully automatic prompt-free mode, not only maintains leading region segmentation accuracy (achieving a Dice score of 89.85\% on ISIC 2018), but also effectively suppresses the boundary degradation phenomenon of foundation models in grayscale modalities on the core boundary metric HD95 (optimized to 15.97 on BUSI). Without introducing extra inference parameters, it exhibits superior edge fidelity compared to existing medical fine-tuned foundation models (such as SAM-Med2D). This study provides a high-fidelity, low-cost technical paradigm for robust medical image segmentation.
    Keyword: Vision Foundation Models \and Boundary-Aware Learning \and Medical Image Segmentation
    DOI: 10.65286/icic.v22i1.11035
    Cite

  • Spatially Adaptive and Cross-Modal Differential Enhancement Network for Multispectral Object Detection, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Miao Li, Xiao Huang, Jinlai Zhang, Mingchao Xiang
    Abstract: Multispectral object detection in complex urban environments remains a challenging task due to severe thermal diffusion, geometric variations, and modality-specific background noise. To address these issues, we propose the Spatially Adaptive and Cross-Modal Differential Enhancement Network (SACDNet), a novel framework featuring three core components: Spatially Adaptive Modulated Convolution (SAMC), Cross-Modal Differential Enhancement (CMDE), and Illumination-Guided Geometric Loss (IGGL). SAMC effectively adapts to target deformations by dynamically adjusting the shape of the receptive field. CMDE utilizes cross-modal channel attention to adaptively emphasize reliable structural representations. Furthermore, IGGL modulates the geometric penalty based on environmental illumination to mitigate the influence of noisy predictions. Experimental evaluations on the FLIR-align dataset demonstrate the effectiveness of our approach. Compared to the state-of-the-art (SOTA) methods, SACDNet achieves a significant improvement, reaching a mean Average Precision (mAP) of 37.2\%.
    Keyword: Multispectral Object Detection \and Spatially Adaptive Convolution \and Cross-Modal Fusion
    DOI: 10.65286/icic.v22i2.89575
    Cite

  • Hierarchical Global-Local Interaction and Refinement for Multimodal Sentiment Analysis , ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: yuanyuan zhou, Yu Lu, YanZhang He, AnNi Wang
    Abstract: Multimodal Sentiment Analysis (MSA) aims to integrate text, audio, and visual modalities to achieve accurate sentiment modeling. Existing methods often rely on shallow interaction structures, making it difficult to jointly capture fine-grained local dynamics and high-level semantic dependencies. In addition, dominant modalities may suppress the learning of weaker modalities, leading to insufficient cross-modal semantic alignment. To address these issues, we pro-pose a Hierarchical Global-Local Interaction and Refinement framework for Multimodal Sentiment Analysis (HGLIR). Specifically, a Modality Dropout strategy is first introduced at the input stage to alleviate over-reliance on a sin-gle modality and improve robustness. Based on this, a Hierarchical Local Inter-action (HLI) module models multimodal sequences through a multi-layer pro-gressive structure to capture local dynamic features at different semantic levels. Within the HLI module, a Cross-Modal Synergistic Learning (CMSL) mecha-nism explicitly models cross-modal semantic consistency and gradually aligns information during interaction. Furthermore, a Global Representation Refine-ment (GRR) module introduces learnable global representations and iteratively updates them in a multi-layer structure to aggregate long-range semantic de-pendencies and form stable high-level semantic representations. Experimental results on CMU-MOSI and CMU-MOSEI demonstrate the effectiveness of the proposed framework across multiple evaluation metrics.
    Keyword: Multimodal Sentiment Analysis, Feature Fusion, Attention Mechanism, Hierarchical Interaction
    DOI: 10.65286/icic.v22i1.48666
    Cite

  • CrackDINO: A DINOv3-based Hybrid Framework for Fine-Grained Crack Segmentation, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Jiajie Li, Bin Hu, Jinlai Zhang, Shifa Tang, Kejia Wang
    Abstract: Accurate pavement crack segmentation is essential for intelligent transportation systems and infrastructure maintenance. However, due to the low contrast, complex background interference, and elongated structural characteristics of cracks, existing segmentation methods often suffer from discontinuous predictions and missed detection of tiny crack regions. In particular, CNN-based methods are limited in capturing long-range dependencies, while Transformer-based methods tend to lose fine-grained spatial details during patch tokenization. To address these challenges, we propose a DINOv3-based crack segmentation framework termed CrackDINO. Specifically, we design an RGB-guided Multi-scale Feature Pyramid (RGMFP) module to enhance hierarchical semantic interaction across different feature resolutions. In addition, a Crack-aware Stable Attention (CAS) module is introduced to strengthen weak crack responses and improve discriminative representation for thin and low-contrast crack regions. Furthermore, a Cascaded Hierarchical Multi-scale Decoder (CHMD) is proposed to progressively recover spatial details and preserve crack continuity during feature reconstruction. Extensive experiments on the Crack500 and CrackForest Dataset indicate that the proposed method achieves competitive performance compared with several representative segmentation models, including U-Net, DeepLabV3+, SegFormer, TransUNet, and Swin-Unet. On the Crack500 Dataset, CrackDINO achieves an mIoU of 62.61\% and an F1-score of 74.16\%. On the CrackForest Dataset, the proposed method obtains an mIoU of 60.08\% and an F1-score of 74.79\%, demonstrating favorable robustness and generalization performance for fine-grained crack segmentation.
    Keyword: Pavement crack segmentation \and DINOv3 \and Transformer \and Multi-scale feature fusion \and Attention mechanism \and Deep learning
    DOI: 10.65286/icic.v22i1.81246
    Cite

  • Redundancy-Guided Hierarchical Fusion for RGBT Object Detection, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Yidan Sun, Yujie Wu, Qiming Yang
    Abstract: In dual-branch RGBT object detection, symmetric architectures are widely adopted without examining whether different feature levels exhibit distinct redundancy patterns. In this paper, we first analyze the weight sparsity of a dual-branch model and observe that redundancy increases with depth, and at deeper layers the RGB branch becomes substantially more redundant than the thermal branch. Based on this observation, we propose a redundancy-guided hierarchical fusion strategy (Red-HiFusion)—feature concatenation at shallow low-redundancy levels, bidirectional cross-attention at middle medium-redundancy levels, and unidirectional IR-query attention at deep asymmetric high-redundancy levels to allow the less redundant thermal features guide the more redundant RGB features. Red-HiFusion achieves 85.4% mAP50 on M3FD and 97.4% mAP50 on LLVIP with only 14.9 M parameters and 35.0 GFLOPs. Extensive ablations show that this hierarchical design consistently outperforms full-bidirectional and full-unidirectional baselines while adding negligible computation. Compared to state-of-the-art approaches, our model attains competitive detection accuracy and has an order of magnitude fewer parameters, offering an interpretable and efficient solution for lightweight RGBT detection.
    Keyword: RGBT Object Detection; Feature Fusion; Attention Mechanism
    DOI: 10.65286/icic.v22i2.41240
    Cite

  • CAPBNet: Channel-Attentive Pyramid and Bottleneck-Guided Context Network for Plant Disease Segmentation, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Xingyu Ren, Rui Teng, Jinlai Zhang, Yuanhao Yang
    Abstract: Accurate plant disease segmentation remains challenging due to complex natural backgrounds, diverse lesion appearances, and ambiguous disease boundaries. In this paper, we propose a Channel-Attentive Pyramid and Bottleneck-Guided Context Network for plant disease segmentation. The proposed network improves DeepLabV3-ResNet50 by recalibrating multi-scale atrous branch features, refining lesion-related spatial context, and combining pixel-wise supervision with region-level overlap optimization. Experiments on the PlantSeg115 dataset show that the proposed method achieves 44.52\% mIoU and 57.51\% mAcc, outperforming the DeepLabV3 baseline by 1.43 and 2.11 percentage points, respectively. These results validate the effectiveness of channel-attentive multi-scale representation and bottleneck-guided context refinement for plant disease semantic segmentation.
    Keyword: Plant disease segmentation \(\cdot\) Semantic segmentation \(\cdot\) Channel-attentive pyramid \(\cdot\) Bottleneck-guided context refinement
    DOI: 10.65286/icic.v22i1.21749
    Cite

  • BCSNet: Boundary-Centric Fusion and Selective Semantic Pyramid Injection Network for Real-Time Semantic Segmentation, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Min Li, Qingpei Liu, Yuan Gao, Mingle Zhou, Yang Chen, Delong Han
    Abstract: Real-time semantic segmentation is a key component of resource-constrained perception systems, such as autonomous driving and robotic navigation, where dense scene understanding must be obtained under strict latency and computation constraints. Although dual-branch architectures provide an effective efficiency-oriented solution, they still suffer from boundary degradation during cross-branch fusion, insufficient multi-scale semantics for small and medium objects, and feature misalignment caused by content-agnostic upsampling near object contours. These limitations become more pronounced at high resolutions, where preserving thin structures and accurate class transitions is often in tension with maintaining high throughput. To address these issues, we present Boundary-Centric Fusion and Selective Semantic Pyramid Injection Network (BCSNet), a real-time segmentation framework that allocates lightweight modeling capacity to boundary-sensitive stages rather than increasing computation uniformly across the network. Specifically, the Boundary-Centric Cross-Branch Fusion and Refinement module learns a shared boundary cue to guide bidirectional feature exchange and local contour refinement with limited overhead. The Semantic Lightweight Feature Pyramid with Selective Injection module provides scale-adaptive semantic cues to the high-resolution stream through a compact pyramid design. The Boundary-Conditioned Region-Adaptive Alignment Upsampling operator further performs content-aware reassembly only within narrow boundary regions, while retaining efficient bilinear interpolation elsewhere. Under a controlled RTX 4090 evaluation protocol on Cityscapes, BCSNet-L achieves 79.5% mIoU at 107.0 FPS, while the lightweight BCSNet-S obtains 76.5% mIoU at 189.0 FPS. On CamVid, BCSNet achieves 77.1% mIoU at 156.8 FPS. These results indicate that BCSNet provides a practical accuracy--efficiency trade-off for high-resolution real-time segmentation, while direct embedded deployment and hardware-specific optimization remain directions for future work.
    Keyword: Real-time semantic segmentation, Boundary-aware learning, Multi-scale feature representation, Adaptive upsampling
    DOI: 10.65286/icic.v22i1.38414
    Cite

  • MCSCA: Multi-dimensional Collaborative Spatial-Channel Attention Network for Traffic Sign Recognition, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Jiazheng Xu, Xiao Huang, Jinlai Zhang, Yuanhao Yang
    Abstract: Traffic sign recognition is a safety-critical perception task in intelligent transportation systems, requiring accurate classification under complex real-world conditions including illumination variation, viewpoint changes, motion blur, and environmental degradation. Existing methods often rely on single-branch attention mechanisms that capture only partial feature dependencies, limiting robustness under degraded visual conditions. To address these limitations, we propose MCSCA, a Multi-dimensional Collaborative Spatial-Channel Attention network that integrates three complementary attention branches—Neuron Saliency Enhancement (NSE), Spatial-Channel Collaborative Calibration (SCC), and Cross-Dimensional Interaction (CDI)—through a learnable Softmax-weighted adaptive fusion strategy. The three branches operate in parallel on shared intermediate feature maps, simultaneously enhancing neuron-level saliency, spatial-channel contextual dependency, and cross-dimensional structural interaction. The fused representation is further stabilized via residual connection. The proposed model is built upon a lightweight residual backbone with multi-scale feature aggregation and is trained using AdamW with warmup-cosine scheduling, CutMix/Mixup augmentation, and label smoothing. Experiments on GTSRB demonstrate that MCSCA achieves 99.89\% validation accuracy, 99.97\% precision, and 99.80\% recall at 633.4 FPS with only 4.38M parameters, maintaining competitive performance while preserving real-time inference efficiency. Robustness evaluation on GTSRB-C, a corrupted benchmark covering 8 camera corruption types at 5 severity levels, shows a mean corruption accuracy (mCA) of 81.81\% and a composite RobScore of 76.91, with near-perfect robustness under photometric corruptions (Fog mCA: 99.85\%, Rain mCA: 98.87\%) and graceful degradation under additive noise and motion blur. These results validate the effectiveness of the proposed multi-branch collaborative attention design for robust traffic sign recognition under real-world perturbations.
    Keyword: Traffic Sign Recognition \and Multi-branch Attention \and Spatial-Channel Attention \and Robustness Evaluation
    DOI: 10.65286/icic.v22i1.89177
    Cite

  • D$^2$Drive: Training-Free Dynamic Inference for End-to-End Autonomous Driving, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Junxuan Liu, Bin Hu, Jinlai Zhang, Qi Xiong, Lin Hu
    Abstract: End-to-end autonomous driving frameworks have achieved strong performance in perception, prediction, and planning. However, their deep interaction architectures often introduce substantial computational overhead and inference latency, which hinders real-time deployment. To improve inference efficiency without retraining, we propose D$^2$Drive, a training-free dynamic inference framework for end-to-end autonomous driving. D$^2$Drive reduces redundant computation by exploiting feature stability during inference. It contains two complementary components. First, a dynamic exit mechanism terminates deeper computation when feature differences between adjacent layers become sufficiently small. Second, an FFN partial computation strategy selectively performs FFN operations on tokens with larger variations across layers, while stable tokens bypass redundant FFN computation. These two designs reduce computation at both the layer and token-computation levels. Unlike pruning, quantization, and retraining-based acceleration methods, D$^2$Drive does not modify model parameters, task heads, or training objectives. Therefore, it can be directly applied to pretrained autonomous driving frameworks during inference. We evaluate D$^2$Drive on four representative end-to-end autonomous driving frameworks, including UniAD, VAD, SparseDrive, and MomAD. The evaluation covers GFLOPs for the whole model, computation inside selected modules, planning accuracy, collision rate, tracking quality, and detection performance, providing a systematic protocol for analyzing the balance between efficiency and accuracy in training-free dynamic inference for autonomous driving systems.
    Keyword: dynamic inference \and training-free acceleration \and early exit \and partial computation \and end-to-end autonomous driving
    DOI: 10.65286/icic.v22i1.30332
    Cite

  • ChildEC: A Developmentally Stratified Chinese Dialogue Corpus for Child Emotion Coaching, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Yuelin Ding, Sujuan Liu
    Abstract: Existing dialogue resources for minors mainly focus on general psychological support, adolescent positive mental health promotion, or broad emotional companionship, and therefore do not adequately support the task of everyday emotion coaching for Chinese children. To address this gap, we construct ChildEC, a developmentally stratified multi-turn dialogue corpus for everyday emotion coaching, targeting Chinese children aged 8--12. The corpus contains 1,709 multi-turn dialogues spanning two developmental stages and five core themes. To support corpus construction and evaluation, we further propose the Theory-Grounded Emotion Coaching Framework (TGEC). Within this framework, TGEC-Synth operationalizes the five-step emotion coaching model and Socratic dialogue into a multi-stage dialogue synthesis pipeline, while TGEC-Eval assesses the quality of child emotion coaching dialogues along six key dimensions. Experimental results show that models fine-tuned on ChildEC consistently outperform their corresponding base models under both automatic metrics and LLM-as-a-Judge evaluation, demonstrating the downstream training value of the corpus for child emotion coaching. Overall, ChildEC provides a specialized data resource and an evaluation reference for developmentally sensitive and theory-consistent research on Chinese child emotion coaching. Our dataset will be publicly released upon acceptance of the paper.
    Keyword: Child emotion coaching, Multi-turn dialogue corpus, Large language models
    DOI: 10.65286/icic.v22i1.68048
    Cite

  • When Verification Hurts: The Cost of Overriding Abstention in Two-Stage Web Agents, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Duchen Li
    Abstract: Web agents built on large vision-language models (VLMs) increasingly adopt a two-stage design: a grounding stage proposes candidate elements on a page, and an action stage decides which element to operate on and how. A natural way to strengthen such agents is to insert a pre-action verifier that re-scores the grounded candidates before acting, echoing the gains that verification and self-refinement bring to language-model reasoning. We test this assumption on the Mind2Web benchmark and report a counter-intuitive result: a GLM-4.6V pre-action verifier does not help and in fact degrades performance, low-ering the action-level step success rate from 34.8% to 23.7% on our evalua-tion subset. Through a step-level analysis we attribute this degradation to two causes. First, the offline multiple-choice protocol has limited candidate cov-erage, as the gold element is absent from the candidate set in roughly 80% of steps, so most steps are unsolvable regardless of verification. Second, and more decisively, the verifier mis-ranks candidates on the solvable steps and discards the grounding stage's calibrated abstention on the unsolvable majori-ty, so it removes a safe default without improving accuracy: it wins 5 steps but loses 20. Guided by this diagnosis, we propose an abstention-aware veri-fier that intervenes only under sufficient candidate coverage and confidence. Our study cautions against transplanting verification into grounding pipelines and identifies calibrated abstention as a property worth preserving.
    Keyword: Web Agents, Vision-Language Models, Element Grounding, Verification, LLM-as-a-Judge, Abstention
    DOI: 10.65286/icic.v22i2.36457
    Cite

  • SGAFormer: Skeleton-Graph and Agent-guided Transformer network for monocular video-based 3D human pose estimation, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Min Li, Lina Du, Yuanyao Lu, Ning Wang, Zixuan Xu
    Abstract: Monocular video-based 3D human pose estimation remains challenging due to depth ambiguity, noisy 2D observations, and complex spatio-temporal depend-encies. Existing Transformer-based methods can capture global relationships, but they often lack explicit skeletal topology constraints and may introduce re-dundant dense temporal interactions. To address these issues, this paper pro-poses a Skeleton-Graph and Agent-guided Transformer network (SGAFormer) for skeleton-aware and temporally coherent 3D human pose estimation from 2D pose sequences. In spatial modeling, the proposed Skeleton-constrained Adaptive Graph-order Attention Module (SAGA) introduces skeleton-constrained dynamic graph attention and multi-order graph propagation to cap-ture joint self-information, direct skeletal connections, and indirect structural dependencies. By adaptively fusing different graph-order branches, SAGA en-hances dynamic joint interaction modeling under physical skeleton constraints. In temporal modeling, the proposed Agent-token Global-local Temporal At-tention Module (AGTA) reorganizes dense frame-wise interactions through agent tokens, while incorporating temporal positional bias and depthwise sepa-rable temporal convolution to enhance global-local motion representation. Ex-periments on Human3.6M show that SGAFormer remains competitive under detected 2D keypoint inputs, achieves superior temporal consistency, and ob-tains strong upper-bound performance with ground-truth 2D keypoints, demon-strating the effectiveness of the proposed spatial-temporal modeling frame-work.
    Keyword: Monocular 3D human pose estimation, Skeleton-constrained graph attention, Transformer, Agent token, Spatio-temporal modeling.
    DOI: 10.65286/icic.v22i1.59700
    Cite

  • Multi-Temporal Mean-Reverting SDE with Preconditioned Diffusion for Remote Sensing Cloud Removal, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Zilv Liu, Jianfang Shen, Ziqi Huang
    Abstract: Remote sensing (RS) imagery is frequently compromised by environmental factors like cloud cover, leading to severe information loss and hindering downstream Earth observation tasks. Although generative paradigms like Diffusion Models and Stochastic Differential Equations (SDEs) show promise in image restoration, they encounter critical bottlenecks in multi-temporal RS scenarios: 1) failing to fully exploit spatio-temporal redundancies and cross-modal complementarities; and 2) suffering from unstable convergence in heavy-cloud regions due to drastic variance in noise levels. To address these challenges, we propose a Preconditioned Mean-Reverting SDE (PMR-SDE) framework for multi-temporal cloud removal. Moving beyond standard generative adaptations, we reformulate the restoration task as a cross-modally conditioned mean-reverting process, which characterizes the continuous evolution from a cloud-distorted state toward a clear reference, explicitly guided by Sentinel-1 SAR features acting as a physical structural anchor. Within this framework, a Spatio-Temporal Dual-Branch Attention (ST-DBA) module is developed to capture long-range optical dependencies while bridging the semantic modality gap. Furthermore, we introduce an Instance-Adaptive Preconditioning Control strategy to rescale the input-output dynamics of the denoiser based on dynamic local priors, effectively alleviating fitting pressures and enhancing gradient stability during extreme diffusion stages. Extensive experiments on the SEN12MS-CR-TS dataset demonstrate that our method achieves strong spatial texture reconstruction and structural fidelity, with competitive performance under heavy-cloud scenarios while maintaining favorable spectral preservation.
    Keyword: Remote sensing cloud removal,diffusion model,stochastic differential equation
    DOI: 10.65286/icic.v22i1.81741
    Cite

  • Att2RAG: A Double-Condition Framework for Knowledge Poisoning Attacks on RAG Systems, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Zhize Hao, Tao Liu, Ming Ye
    Abstract: Modern retrieval-augmented generation (RAG) and memoryaugmented LLM applications are widely deployed in knowledge-intensive settings. These systems ground model outputs on external knowledge stores and may persist interaction traces in vector memory. If the underlying store is compromised, poisoned content can be retrieved repeatedly and thereby shape downstream responses, yielding confident yet harmful outputs supported by seemingly plausible evidence. Recent studies have shown that RAG pipelines and LLM agents are vulnerable to knowledge poisoning and prompt injection, but many formulations treat attack success as a single end-to-end outcome and do not separate retrieval and generation failure modes in a retrieval–generation aligned manner. Moreover, although long-horizon memory writes can introduce cumulative risks, these effects are not empirically evaluated under our benchmark setting. We present Att2RAG, a double-condition framework for knowledge poisoning attacks on RAG systems. Att2RAG decomposes a successful poisoning event into a retrieval condition and a generation condition, and casts poisoning as maximizing attack success subject to satisfying both conditions. We instantiate the framework with (i) a white-box variant that applies projected gradient descent (PGD) in embedding space, serving as an approximate upper bound under strong attacker assumptions, and (ii) a black-box attack based on a Q ⊕ I construction that combines problem self-similarity with adversarial instruction injection using only query access. We evaluate Att2RAG on representative QA benchmarks under several common RAG configurations and consider practical defenses such as paraphrasing and duplicate filtering. Attack success is measured by ASR and Target-F1 relative to an attacker-specified target response. Across the evaluated configurations, Att2RAG attains attack success rates in the mid-90% range without defenses; paraphrasing plus duplicate filtering reduce ASR only modestly, and many attacks remain successful. These results highlight limitations of semantic-similarity-driven retrieval and suggest that strengthening RAG systems requires defenses beyond surface-form rewriting and naive duplicate removal.
    Keyword: RAG · Knowledge poisoning · Adversarial attack · Large language models · RAG security
    DOI: 10.65286/icic.v22i2.97272
    Cite

  • RCBS-ReID: Rank-Reversal-Aware Compression and Budget-Adaptive Search for Efficient Edge Wooden Pallet Re-Identification, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Yaxiong Liu, Guanzhi Lyu, Yue Yang, Jingsong Li, Ke Chen, Yunling Liu
    Abstract: Wooden pallet re-identification (ReID) is essential for reliable logistics traceability and warehouse management. Although deep ReID models have achieved promising retrieval accuracy, practical deployment remains limited by high-dimensional gallery descriptor storage and full-gallery matching latency. Existing feature compression methods mainly reduce descriptor size, but compression-induced approximation errors may disturb or even reverse the ranking order between positive samples and hard negatives. Meanwhile, many retrieval acceleration methods rely on fixed search and reranking budgets, causing redundant computation for easy queries and missed relevant candidates for hard queries. To address these limitations, we propose RCBS-ReID, a rank-reversal-aware compression and budget-adaptive search framework for efficient wooden pallet ReID. Specifically, Rank-Reversal-Aware Compact Encoding (RACE) estimates rank-reversal risk to guide compact subspace selection, reducing gallery storage while preserving retrieval ranking stability after compression. In addition, Dual Budget Adaptive Search (DBAS) combines coarse candidate recall with fine reranking, and adaptively allocates search and reranking budgets according to query difficulty to reduce redundant retrieval cost. Experiments on an in-house wooden pallet ReID dataset from real logistics scenarios show that RCBS-ReID achieves superior overall performance compared with 15 representative ReID methods. Compared with the baseline, RCBS-ReID reduces gallery storage by 95.0%, achieves 3.23× retrieval acceleration, and improves mAP by 1.11%.
    Keyword: Re-identification; Feature compression; Retrieval acceleration; Edge deployment; Wooden pallet
    DOI: 10.65286/icic.v22i2.59196
    Cite

  • YOLO11s-RDS: Response-Guided Feature Enhancement with Spatial Priors for Lightweight Colorectal Polyp Detection, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Haoran Wu
    Abstract: Accurate and efficient colorectal polyp detection is essential for computer-aided colorectal cancer screening. However, endoscopic images often suffer from specular reflections, intestinal folds, motion blur, weak lesion boundaries, and small-scale abnormalities, which introduce severe background interference and increase the risk of missed detections and localization errors. To address these challenges, we propose YOLO11s-RDS, a lightweight colorectal polyp detector built upon YOLO11s. The proposed method redesigns the neck with three complementary modules. First, Response-Guided Calibration Fusion (RGCF) replaces direct feature concatenation with response-aware gated fusion,suppressing shallow pseudo-responses caused by reflections and folds while improving cross-scale semantic consistency. Second, Dual-Pooling Feature Enhancement (DPFE) jointly exploits global average responses and local peak activations to strengthen channel representations associated with small polyps. Third, Spatial Prior-Weighted Convolution (SPWConv) introduces a local center-enhancement prior into downsampling convolutions to mitigate spatial-structure degradation of compact targets. Experiments on a uniformly processed hybrid dataset comprising CVC-ClinicDB, CVC-ColonDB, ETIS-LaribPolypDB, and Kvasir-SEG show that YOLO11s-RDS achieves Precision, Recall, F1-score, mAP50, and mAP50:95 of 0.9018, 0.8729, 0.8871, 0.9113, and 0.6715, respectively. Compared with YOLO11s, it improves Recall, mAP50, and mAP50:95 by 5.51%, 3.77%, and 3.26%, demonstrating stronger robustness in complex endoscopic scenes while maintaining a lightweight design.
    Keyword: Colorectal polyp detection、Multi-scale feature fusion、Small-object detection、Spatial prior
    DOI: 10.65286/icic.v22i1.17951
    Cite

  • Lightweight Multi-Scale Transformer for Real-Time Railway Surface Defect Detection on Inspection Vehicles, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Tianjing Zhang, Zhenhua Wang, Jian Sun, Jun Chen, Xiaogang Dang, Chengbin Weng
    Abstract: Railway surface defect detection is essential for condition-based maintenance and structural health monitoring of rail infrastructure. With the deployment of high-resolution cameras on inspection vehicles, detectors must localize small and diverse defects in real time under strict on-board computational constraints. Existing convolutional neural network (CNN) based methods often lack sufficient long-range context along the rail, whereas recent transformer-based architectures are typically too heavy for embedded deployment. This paper proposes a Lightweight Multi-Scale Transformer (LMST) tailored to real-time railway surface defect detection on inspection vehicles. LMST combines rail-oriented tokenization, a hierarchical multi-scale transformer encoder with rail-aligned windowed self-attention, and a defect-aware gated fusion module feeding a lightweight dense prediction head. Experiments on a public rail surface defect benchmark and additional inspection-vehicle imagery show that LMST achieves competitive or improved average precision compared with strong CNN and transformer baselines, while maintaining real-time throughput on industrial GPUs and clearly enhancing the detection of small and elongated defects under challenging field conditions.
    Keyword: Railway surface defect detection; Vision transformer; Multi-scale attention; Real-time inspection; Structural health monitoring; Inspection vehicles
    DOI: 10.65286/icic.v22i1.32571
    Cite

  • Adpt-STGIN: An Adaptive Spatio-Temporal Graph Inductive Network for Topology-Robust Traffic Prediction in Data Center Networks, ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
    Authors: Xuran Chen, Jie Hao, Ran Wang, Qiang Wu, Yimeng Gao
    Abstract: Accurate network traffic prediction is essential for resource management and congestion control in data center networks. Existing spatio-temporal graph neural network (STGNN) models predominantly employ transductive spatial encoders, such as GCN or GAT, whose parameters are tied to a fixed graph structure, preventing generalization to unseen topologies without full retraining. In this paper, we propose Adpt-STGIN (Adaptive Spatio-Temporal Graph Inductive Network), a topology-robust traffic prediction framework built on two key contributions. First, we design a deeply fused GraphSAGE-GRU cell that embeds independent inductive GraphSAGE(SAmple and aggreGatE) encoders directly into each GRU gate, enabling simultaneous spatio-temporal feature extraction at every time step while remaining topology-agnostic. Second, we develop a topology-robust transfer learning framework with a frozen encoder strategy that adapts pretrained models to new topologies by fine-tuning only the lightweight decoder. Experiments on four data center topologies demonstrate that Adpt-STGIN achieves R^2 > 0.99 in pretraining and generalizes to unseen topologies in zero-shot mode with R^2 > 0.994, confirming the practical efficiency of the proposed framework.
    Keyword: network traffic prediction , spatio-temporal graph neural network,inductive learning , transfer learning , data center network
    DOI: 10.65286/icic.v22i2.34322
    Cite