Haifeng Lin

dblp:172/8558 · DBLP profile ↗
← Back
18ranked-venue papers
4as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SCAN: Self-Calibrated Textual Anchoring for Dual-Granularity Video Screening in Text-Video Retrieval
abstract
Text-Video Retrieval (TVR) is a foundational cross-modal task that aims to align natural language queries with relevant video content, playing a critical role in multimedia search and video understanding. Current methods typically address the cross-modal semantic gap by either augmenting text with generated captions or enhancing video representations at the feature level. However, these strategies often introduce new challenges: generated captions may contain hallucinations that mislead alignment, while dense video representations can retain substantial information redundancy and semantic bias, collectively hindering precise and efficient retrieval. To overcome these limitations, we propose SCAN (Self-Calibrated Textual Anchoring for Dual-Granularity Video Screening). Our framework introduces a dual-granularity video screening mechanism that progressively filters out redundant visual information to obtain semantically condensed video representations. Simultaneously, it employs a self-calibrated textual anchoring module, which leverages the video content to verify and reinforce reliable textual cues, thereby strengthening core text semantics and mitigating caption hallucinations. These two processes are jointly optimized, which aligns text and video semantics in a synergistic manner. Extensive experiments on three benchmarks (MSR-VTT, MSVD, and DiDeMo) demonstrate that SCAN consistently outperforms state-of-the-art methods, validating its effectiveness in reducing hallucinations, eliminating redundancy, and bridging semantic asymmetry for robust text-video retrieval.
Dixin Chen, Baoyao Yang, Haifeng Lin, Canrong Du, Wenbin Yao
ICMR3
2026 FedSHA: Structure-aware heterogeneity adaptation for personalized federated graph node classification
Haifeng Lin, Min Xia 0002, Bo-Chao Zheng
Neurocomputing5
2026 Multibranch Dendritic Computing for Federated Graph Neural Networks: Maintaining Interpretability Under Privacy Constraints
abstract
Graph Neural Networks (GNNs) have demonstrated remarkable performance in modeling complex relational data. However, their deployment in privacy-sensitive domains remains challenging due to the inherent tension between privacy protection and model interpretability. This paper introduces a novel Multi-Branch Dendritic Computing framework for Federated Graph Neural Networks (MBD-FGNN) that addresses this fundamental challenge. Our approach integrates biologically-inspired dendritic computing with federated learning paradigms, enabling distributed training across decentralized graph data while preserving both privacy and interpretability. The multi-branch architecture mimics dendritic computation in biological neurons,where each branch processes patterns at different receptive field scales (1-hop, 2-hop, and 4-hop neighborhoods), providing natural model explanations through branch contribution analysis.To ensure rigorous privacy protection, we develop a branch-adaptive differential privacy mechanism with gradient clipping and calibrated Gaussian noise injection, satisfying (ϵ, δ)-differential privacy. We theoretically analyze the framework’s ability to maintain interpretation stability under differential privacy constraints and derive formal guarantees on the privacy-interpretability trade-off. Extensive experiments on benchmark graph datasets demonstrate that MBD-FGNN outperforms state-of-the-art federated GNN approaches in terms of accuracy, privacy preservation, and explanation quality. Notably, our method achieves robust explanation stability even under privacy constraints, maintaining higher explanation consistency compared to conventional approaches where explanations significantly degrade as privacy protection increases. The proposed framework opens new avenues for deploying interpretable graph learning systems in privacy-critical applications such as healthcare networks, financial transaction graphs, and social network analysis. The code and dataset are publicly available at https://github.com/ZhaoGuan-nuist/MBD-FGNN.
Yifan Fu, Min Xia 0002, Liguo Weng, Haifeng Lin
IEEE Internet Things J.4
2025 Interactive and Supervised Dual-Mode Attention Network for Remote Sensing Image Change Detection
abstract
With the rapid advancement of remote sensing technology, change detection using bitemporal remote sensing images has significant applications in land use planning and environmental monitoring. The emergence of convolutional neural networks (CNNs) has accelerated the development of deep learning-based change detection. However, existing deep learning algorithms exhibit limitations in understanding bitemporal feature relationships and accurately identifying change region boundaries. Moreover, they inadequately explore feature interactions between bitemporal images before extracting differential features. To address these issues, this article proposes a novel interactive and supervised dual-mode attention network (ISDANet). In the feature encoding stage, we employ the lightweight MobileNetV2 as the backbone to extract bitemporal features. Additionally, we design the neighbor feature aggregation module (NFAM) to aggregate semantic features from adjacent scales within the dual-branch backbone, enhancing the representation of temporal features. We further introduce the interactive attention enhancement module (IAEM), which effectively integrates self-attention and cross-attention mechanisms. This establishes deep interactions between bitemporal features, suppresses irrelevant noise, and ensures precise focus on true change regions. In the feature decoding stage, the supervised attention module (SAM) reweights differential features and leverages supervisory signals to guide the learning of attention mechanisms, significantly improving boundary detection accuracy. SAM dynamically aggregates multilevel features, balancing high-level semantics and low-level details to capture subtle changes in complex scenes. The proposed model achieves F1 scores that are 0.28%, 1.6%, and 0.76% higher than the best comparative method, spatiotemporal enhancement and interlevel fusion network (SEIFNet), on three CD datasets [LEVIR-CD, Guangzhou dataset (GZ-CD), and Sun Yat-sen University dataset (SYSU-CD)], respectively, while maintaining a lightweight design with only 6.93 M parameters and 3.46G floating-point operations (FLOPs). The code is available athttps://github.com/RenHongjin6/ISDANet.
Hongjin Ren, Min Xia 0002, Liguo Weng, Haifeng Lin, Junqing Huang, Kai Hu 0006
IEEE Trans. Geosci. Remote. Sens.4
2025 A Triple-Branch Model With Style-Unified Contrastive Learning and Adaptive Dice Focal Loss for Change Detection
abstract
Deep learning has significantly reshaped the landscape of remote sensing change detection (RSCD). However, detecting subtle changes remains a formidable challenge due to the simultaneous presence of three critical issues: domain shift from seasonal and illumination variations, severe class imbalance between change and background areas, and the need for high deployment efficiency. To address these multifaceted problems in a unified manner, this paper introduces SDTNet, a comprehensive framework featuring three key innovations. First, to counter domain shift, we propose the Style-Unified Temporal Difference Contrastive Learning Strategy (STDCL), a novel, fine-grained contrastive learning method that guides the model to decouple real from pseudo changes without incurring additional inference costs. Second, to mitigate severe class imbalance, we design the Adaptive Dice Focal Loss (ADFLoss), which, unlike static loss functions, introduces two novel adaptive factors to dynamically balance sample weights and implement a curriculum learning strategy. Third, for efficient feature enhancement, we present the Dual-Branch Temporal Difference Transformer (DTDFormer), an efficiency-aware architecture that adapts the differential attention mechanism for bi-temporal inputs and incorporates a sparse attention module to reduce computational complexity. Experimental results demonstrate that our integrated SDTNet framework achieves state-of-the-art (SOTA) performance across five large-scale change detection datasets. The code is publicly available at https://github.com/Baobaon/SDTNet.
Zikai Zhao, Xingyu Liang, Min Xia 0002, Liguo Weng, Haifeng Lin, Junqing Huang
IEEE Trans. Geosci. Remote. Sens.5
2024 Multi-granularity siamese transformer-based change detection in remote sensing imagery
Lei Song 0013, Min Xia 0002, Liguo Weng, Kai Hu 0006, Haifeng Lin, Ming Qian
Eng. Appl. Artif. Intell.6
2024 TrafficTrack: rethinking the motion and appearance cue for multi-vehicle tracking in traffic monitoring
Haifeng Lin
Multim. Syst.2
2024 Cross-dimensional feature attention aggregation network for cloud and snow recognition of high satellite images
Kai Hu 0006, Enwei Zhang, Min Xia 0002, Huiqin Wang, Xiaoling Ye, Haifeng Lin
Neural Comput. Appl.6
2024 Multipath Multiscale Attention Network for Cloud and Cloud Shadow Segmentation
abstract
The segmentation task of cloud and cloud shadow has always been one of the important tasks in remote sensing image processing. At present, cloud detection based on deep learning methods lacks generalization, which is easy to cause the loss of space and detail information, and missed detection and false detection occur from time to time. Aiming at the above problems, this paper proposes a multi-path, multi-scale attention network (MMA). In this network, Muti-scale overlapping patch embedding (MSP) is introduced to extract multi-scale semantic information with multi-path PVT, and strip convolution is used to supplement the spatial detail information of the image, so as to realize the effective aggregation of fine and rough features at the same feature level. In order to aggregate the global information, the multi-scale global aggregation module (MGAM) is used for deep feature extraction to supplement the high semantic information. In the decoding stage, aiming at the problem of small target cloud detection, the attention guided fusion module (AGFM) is proposed to focus on the important information of the image, remove the network noise and increase the detection accuracy of small targets. Contextual information fusion (CIF) decoding method is proposed for the coarse segmentation boundary problem, which fully integrates the context information and effectively helps to restore the image. The experimental results on Biome 8 Dataset, HRC-WHU Dataset and SPARCS Dataset confirm that our method is superior to the current cutting-edge cloud and cloud shadow detection technology.
Guowei Gu, Liguo Weng, Min Xia 0002, Kai Hu 0006, Haifeng Lin
IEEE Trans. Geosci. Remote. Sens.5
2024 Learning for Adaptive Multi-Copy Relaying in Vehicular Delay Tolerant Network
abstract
Multi-copy routing is a way that creates and forwards copies to the nodes without the copies so that increases the message delivery probability in the Vehicular Delay Tolerant Network (VDTN). The cost of generating many copies is potential congestion. To solve this problem, a feasible approach is to limit the number of copies generated. However, when the mobility of nodes is not sufficient and when the density of nodes near the source node is relatively high, the inappropriate process of copy distribution may result in copies remaining in a small area and too high local density of copies. Consequently, the overall performance is reduced due to low destination encountering probability in the whole area caused by too many local copies distribution. In this paper, by using node encounter rate to describe the activity and density of nodes, an effective Q-learning based VDTN multi-copy routing algorithm is proposed. The Q-learning reward factor is determined by the node encounter rate so that the distribution of message copies is controlled and the too high local density of copies can be alleviated. Simulation results show that the proposed algorithm outperforms the related algorithms, and offers better delivery rate, negligible message latency and network overhead under different network node speeds, node density, and load conditions.
Haifeng Lin, Jingjing Qian, Di Bai 0001
IEEE Trans. Intell. Transp. Syst.1
2023 Dual-branch network for change detection of remote sensing image
Chong Ma 0001, Liguo Weng, Min Xia 0002, Haifeng Lin, Ming Qian
Eng. Appl. Artif. Intell.4
2023 Transformer-Based Masked Autoencoder With Contrastive Loss for Hyperspectral Image Classification
abstract
Recent years, in order to solve the problem of lacking accurately labeled hyperspectral image data, self-supervised learning has become an effective method for hyperspectral image classification. The core idea of self-supervised learning is to define a pretext task which helps to train the model without the labels. By exploiting both the information of the labeled and unlabeled samples, self-supervised learning shows enormous potential to handle many different tasks in the field of hyperspectral image processing. Among the vast amount of self-supervised methods, contrastive learning and masked autoencoder are well known because of their impressive performance. This article proposes a Transformer based masked autoencoder using contrastive learning (TMAC), which tries to combine these two methods and improve the performance further. TMAC has two branches, the first branch has an encoder-decoders structure, it has an encoder to capture the latent image representation of the masked hyperspectral image and two decoders where the pixel decoder aims to reconstruct the hyperspectral image at pixel-level and the feature decoder is built to extract the high-level feature of the reconstructed image. The second branch consists of a momentum encoder and a standard projection head to embed the image into the feature space. Then, by combining the output of feature decoder and the embedding vectors via contrastive learning to enhance the model’s classification performance. According to the experiments, our model shows powerful feature extraction capability and gets outstanding results on hyperspectral image datasets.
Xianghai Cao, Haifeng Lin, Shuaixu Guo, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.2
2023 Multiscale Attention Feature Aggregation Network for Cloud and Cloud Shadow Segmentation
abstract
Cloud and cloud shadow segmentation is one of the most important problems in remote sensing image processing. Due to the vulnerability to ground object interference, noise interference and other factors, and the lack of generalization ability, traditional deep learning network would inevitably lose details and spatial information, resulting in imprecise segmentation of cloud and cloud shadow boundaries, missing detection and false detection. In order to solve these problems, a multi-scale attention feature aggregation network is proposed, which extracts semantic information at different levels based on residual network. A multi-scale strip pooling attention module is designed to extract multi-scale context information and deep spatial channel information. In order to improve the feature extraction of the model, the deep multi-head feedforward transfer attention module is introduced to enable the adjacent two layers of the backbone network to guide each other for feature mining. Then, a bilateral feature fusion module is used to guide the fusion of low-level semantic information and high-level detail information. Finally, in order to enhance the repair of boundary details of cloud and cloud shadow, a boundary refinement boosting model is designed. The experimental results show that our model can handle complex cloud cover scenes, and has excellent performance on cloud, cloud shadow dataset and CSWV dataset. Its segmentation accuracy is superior to the existing methods, which is of great significance to the related work of cloud detection.
Min Xia 0002, Haifeng Lin, Ming Qian
IEEE Trans. Geosci. Remote. Sens.3
2023 Multiscale Location Attention Network for Building and Water Segmentation of Remote Sensing Image
abstract
Traditional building and water segmentation methods are vulnerable to noise interference, and hence they could not avoid missed and false detections in the detection process. Excessive deep learning downsampling would lead to significant loss of feature map information, and image location information offset, and the overall effect of falling apart. To address these issues, a Multi-Scale Location Attention Network (MSLA) is proposed. Location-spatial information and channel information are particularly important for edge detail segmentation in building and water cover. The network includes a Location Channel Attention Unit (LCA) to focus on tributary details of rivers and segmentation of building edge eaves. Moreover, this paper builds a Dual-Branch Multi-Scale Aggregation Unit (DBMSA) to obtain deeper multi-scale semantic information. Finally, the Multi-Scale Fusion Unit (MSF) is used to guide the information merging of multiple stages, and the boundary information is improved by splicing the acquired deep multi-scale information with the information of the relevant feature extraction layer in the downsampling. The experimental results on several datasets show that the proposed approach outperforms other methodologies in segmentation accuracy.
Xin Dai 0001, Min Xia 0002, Liguo Weng, Kai Hu 0006, Haifeng Lin, Ming Qian
IEEE Trans. Geosci. Remote. Sens.5
2023 Traffic Signal Optimization Based on Fuzzy Control and Differential Evolution Algorithm
abstract
Urban traffic congestion is often concentrated at urban intersections. An urban road traffic signal control system is needed to prevent problems such as driving delays caused by frequent traffic congestions on trunk lines, exhaust emissions owing to frequent start and stop of vehicles, and fuel wastage due to long idling times. Maximizing the traffic capacity of an intersection and reducing the delay rate of vehicles has always been a problem for traffic control research. The coordinated control of urban traffic signals is regarded as a multi-objective optimization problem. A mathematical model for urban trunk traffic is studied herein. An average delay model, average queue length model, total delay calculation model for vehicles at intersections, and vehicle exhaust emission model are established to obtain an optimization model for a new traffic trunk coordinated control system. In addition, our study combines the fuzzy control theory with the adaptive sequencing mutation multi-objective differential evolution algorithm (FASM-MDEA). This new optimization method for traffic signal control at urban intersections is proposed as a solution for the traffic flow optimization model to solve the problem of traffic signal coordination and control of urban trunk lines. The simulation results demonstrate the effectiveness of the model optimization algorithm proposed in this study.
Haifeng Lin, Yehong Han, Bo Jin 0018
IEEE Trans. Intell. Transp. Syst.1
2022 Multi-scale strip pooling feature aggregation network for cloud and cloud shadow segmentation
Chen Lu 0004, Min Xia 0002, Haifeng Lin
Neural Comput. Appl.3
2022 Intelligent Bus Operation Optimization by Integrating Cases and Data Driven Based on Business Chain and Enhanced Quantum Genetic Algorithm
abstract
Intelligent public transport systems are a key direction for the study of intelligent transportation systems (ITSs), and they perform the following functions: location and tracking, aided navigation, dispatch and command, and dynamic information release; these systems also help travelers determine optimal routes. This paper mainly studies intelligent bus operation optimization by integrating cases and data from business chains, including the optimization of public transport vehicle scheduling according to the characteristics of vehicle scheduling, considers the interests of passengers and public transport companies, and adopts a true-value encoding method with the departure time as the variable goal optimization. In addition, this paper builds a model for the intelligent scheduling problem in public transport based on an enhanced quantum genetic algorithm (EQGA) to find the optimal timetable. This model sets the minimum waiting time cost of passengers and the maximum interests of public transport companies as the goals, restricts the departure interval and two adjacent intervals, and constrains the load factor of passengers. Moreover, based on the analysis of public transport travel characteristics and passenger flow data, this paper evaluates the efficiency of public transport operation, reasonably analyzes the utilization of public transport resources and the travel time of public transport passengers, and thoroughly studies the public transport operation system. This paper selects actual bus line data for empirical analysis, and the empirical results show that the proposed algorithm and model can meet the requirements of many aspects and provides good intelligence, applicability, and optimization.
Haifeng Lin, Chengpei Tang
IEEE Trans. Intell. Transp. Syst.1
2022 Analysis and Optimization of Urban Public Transport Lines Based on Multiobjective Adaptive Particle Swarm Optimization
abstract
Urban public transport is a very complex system, and with the development of urbanization, there are many new urban traffic characteristics. Making bus routes and scheduling strategies more efficient, scientific and accurate has a positive impact on the actual operation of public transport. To solve the urban public transport line design problem, this paper describes the implicit law of the characteristics of public transport travel from a deep perspective and analyzes the forms, influencing factors and existing problems of bus dispatching. By establishing a multiobjective public transport dispatching optimization model, starting from bus companies, passengers and government departments, public transportation operating costs comprehensively consider the interests of various parties and finally realize the optimization objective of minimizing fixed costs, fuel costs, carbon emission costs and time window penalty costs. The objective function is set reasonably, and the generation and optimization method of the initial line set in the public transport line design problem is improved; suitable constraint conditions and evaluation indicators are considered. This paper attempts to control the overall length of the bus line on the premise of fully meeting the travel needs of passengers. By solving the multiobjective problem on the same network and comparing different multiobjective optimization algorithms, the effectiveness of the method is evaluated. Additionally, an improved multiobjective adaptive particle swarm optimization (MOAPSO) is proposed, which has the characteristics of faster convergence, higher efficiency and low computational complexity. The simulation experimental results show that the proposed algorithm in this paper can obtain a better Pareto optimal solution set and can effectively solve the multiobjective model. The departure interval conforms to the passenger flow distribution, which can effectively reduce costs and improve the travel service quality of passengers on a large-scale network.
Haifeng Lin, Chengpei Tang
IEEE Trans. Intell. Transp. Syst.1