Hailong Ning

dblp:281/0917 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
19since 2021 · last 2026
0000-0001-8375-1181ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HRTNet: Holistic registration theory-inspired network for camouflaged object detection
Yueqi Zhao, Hailong Ning, Zhanxuan Hu, Tao Lei 0003, Asoke K. Nandi
Neurocomputing2
2026 Mambaer: Mamba with knowledge-learning hierarchical attention for facial expression recognition
Ercheng Pei, Xiaofeng Wei, Zhanxuan Hu, Hailong Ning, Xiaochun An, Xiaoge Li
Multim. Syst.6
2025 Representation discrepancy bridging method for remote sensing image-text retrieval
Hailong Ning, Siying Wang 0012, Tao Lei 0003, Xiaopeng Cao, Huanmin Dou, Bin Zhao 0001, Asoke K. Nandi, Petia Radeva
Neurocomputing1
2025 DiffuseVAE++: Mitigating training-sampling mismatch based on additional noise for higher fidelity image generation
Xiaobao Yang 0001, Hailong Ning, Guorui Zhang, Wei Sun 0036, Sugang Ma
Neurocomputing3
2025 Classifier ensemble based source-free domain adaptation for time series classification
Ercheng Pei, Wangdong Zhao, Zhanxuan Hu, Hailong Ning
Knowl. Based Syst.5
2025 DA2-Net: Integrating SAM2 With Domain Adaption and Difference Aggregation for Remote Sensing Change Detection
abstract
Visual foundation models (VFMs) have been widely applied in the field of remote sensing (RS). However, they still face two main challenges when applied to precise remote sensing change detection (RSCD) tasks in complex scenes. Firstly, the nonnegligible domain shift between natural scene and RS scene limits the direct application of VFMs to the RSCD task. Second, most of existing RSCD methods may suffer from the boundary displacement problem due to the inadequate exploration of temporal differences for bi-temporal features. To address the above issues, this study proposes a SAM2-based domain adaptive and spatial difference aggregation network (DA2-Net) for RSCD. The proposed DA2-Net has two main advantages. First, a hierarchical low-rank adaptation (LoRA) strategy is presented by introducing low-rank matrices at key positions of SAM2, which can inject inductive biases from the RS domain into the network and alleviate the domain shift problem. Second, a difference adaptive enhancement module (DAEM) is designed to explore temporal differences for hierarchical bi-temporal features. The DAEM provides respective attention weights for different information through a dual branch of global difference awareness and local detail optimization. Experimental results on SYSU-CD, WHU-CD, and LEVIR-CD datasets demonstrate the superiority of DA2-Net. Code is available at https://github.com/xuptheqi-hash/ DA2Net.
Hailong Ning, Qi He 0006, Tao Lei 0003, Xiaopeng Cao, Wuxia Zhang, Yanping Chen 0006, Asoke K. Nandi
IEEE Trans. Geosci. Remote. Sens.1
2025 From Macro to Micro: A Lightweight Interleaved Network for Remote Sensing Image Change Detection
Yetong Xu, Tao Lei 0003, Hailong Ning, Shaoxiong Lin, Tongfei Liu, Maoguo Gong, Asoke K. Nandi
IEEE Trans. Geosci. Remote. Sens.3
2024 An ensemble learning-enhanced multitask learning method for continuous affect recognition from facial images
abstract
Continuous affect recognition from facial images aims to estimate the values of multiple affective dimensions from a facial image sequence. To leverage relevant information between multiple affective dimensions, multitask learning has been used in the estimation of continuous affective states . Most of the existing multitask continuous affect recognition methods focus on designing elaborate multitask networks. Meanwhile, a few research works consider using multitask training strategies for continuous affect recognition. In general, existing multitask continuous affect recognition methods face the problem of unstable training effects. In this work, to improve the stability of multitask learning , we propose an ensemble learning-enhanced multitask network architecture for continuous affect recognition. In addition, we introduce a novel adaptive weighted loss-based multitask learning strategy to effectively train the proposed multitask continuous affect recognition model. Experimental results, on the RECOLA, SEMAINE and AFEW-VA datasets for continuous affect recognition, demonstrate the potential of the proposed method compared to state-of-the-art methods.
Ercheng Pei, Zhanxuan Hu, Hailong Ning, Abel Díaz Berenguer
Expert Syst. Appl.4
2024 Neural collapse inspired semi-supervised learning with fixed classifier
Zhanxuan Hu, Hailong Ning, Yonghang Tai, Feiping Nie 0001
Inf. Sci.3
2024 Lightweight Structure-Aware Transformer Network for Remote Sensing Image Change Detection
abstract
Popular Transformer networks have been successfully applied to remote sensing (RS) image change detection (CD) identifications and achieved better results than most convolutional neural networks (CNNs), but they still suffer from two main problems. First, the computational complexity of the Transformer grows quadratically with the increase of image spatial resolution, which is unfavorable to RS images. Second, these popular Transformer networks tend to ignore the importance of fine-grained features, which results in poor edge integrity and internal tightness for largely changed objects and leads to the loss of small changed objects. To address the above issues, this letter proposes a lightweight structure-aware Transformer (LSAT) network for RS image CD. The proposed LSAT has two advantages. First, a cross-dimension interactive self-attention (CISA) module with linear complexity is designed to replace the vanilla self-attention (SA) in the visual Transformer, which effectively reduces the computational complexity while improving the feature representation ability of the proposed LSAT. Second, a structure-aware enhancement module (SAEM) is designed to enhance difference features and edge detail information, which can achieve double enhancement by difference refinement and detail aggregation to obtain fine-grained features of bi-temporal RS images. Experimental results show that the proposed LSAT achieves significant improvement in detection accuracy and offers a better tradeoff between accuracy and computational costs than most state-of-the-art (SOTA) CD methods for RS images.
Tao Lei 0003, Yetong Xu, Hailong Ning, Zhiyong Lv, Chongdan Min, Yaochu Jin, Asoke K. Nandi
IEEE Geosci. Remote. Sens. Lett.3
2024 Orientational Clustering Learning for Open-Set Hyperspectral Image Classification
abstract
Recently, some literature has begun to pay attention to the open-set problem in remote sensing application scenarios and studied various open-set hyperspectral image classification (OSHIC) methods. These OSHIC methods are usually based on deep neural networks, using the nondirectional Euclidean distance losses to constrain latent sample representations of known classes to be compact. Nonetheless, the potential effect of the spatial distribution of sample representations is ignored, resulting in degraded classification performance in OSHIC. In this letter, we propose an orientational clustering learning (OCL) method for OSHIC. First, in the feature space generated by the convolutional neural network, a class anchor strategy is employed to bring features of the same class closer while keeping features of different classes distant. Then, we utilize the orientational learning to further tighten the intraclass feature space. OCL directionally optimizes the spatial distribution of hyperspectral sample representations to improve the ability to identify known classes and distinguish unknown classes. Experiments show that the OCL achieves overall accuracies of 94.43%, 92.27%, and 76.94% on the Pavia University, Salinas, and Indian Pines datasets, respectively.
Wenjing Chen 0003, Hailong Ning, Hao Sun 0014, Wei Xie 0008
IEEE Geosci. Remote. Sens. Lett.4
2023 Mutual-Taught Deep Clustering
abstract
Deep clustering seeks to group data into distinct clusters using deep learning techniques . Existing approaches of deep clustering can be broadly categorized into two groups: offline clustering based on unsupervised representation learning and online clustering based on unsupervised classification . While both groups have demonstrated impressive performance in deep clustering, no study has explored the integration of their respective strengths. To this end, we propose Mutual-Taught Deep Clustering (MTDC), which unifies unsupervised representation learning and unsupervised classification into a framework while realizing mutual promotion using a novel mutual-taught mechanism . Specifically, MTDC alternates between predicting pseudolabels in label space and estimating semantic similarity in feature space during training. Moreover, pseudolabels provide weakly-supervised information to enhance unsupervised representation learning, while semantic similarities function as structural priors that regularize unsupervised classification. Consequently, unsupervised classification and unsupervised representation learning can mutually benefit from one another. MTDC is decoupled from prevailing deep clustering methods . For the sake of clarity, we build upon a straightforward baseline in this paper. Despite its simplicity, we demonstrate that MTDC is exceedingly efficacious and consistently enhances the baseline results by substantial margins. For example, MTDC achieves 2.5 % ∼ 7.9 % (NMI), 3.0 % ∼ 13.9 % (ACC), and 3.1 % ∼ 16.7 % (ARI) gains over the baseline on six widely used image datasets. Source code is available at:https://github.com/yichenwang231/MTDC.
Zhanxuan Hu, Hailong Ning, Danyang Wu, Feiping Nie 0001
Knowl. Based Syst.3
2023 Ultralightweight Spatial-Spectral Feature Cooperation Network for Change Detection in Remote Sensing Images
abstract
Deep convolutional neural networks have achieved much success in remote sensing image change detection (CD) but still suffer from two main problems. First, existing multi-scale feature fusion methods often employ redundant feature extraction and fusion strategies, which often leads to high computational costs and memory usage. Second, the regular attention mechanism in CD is difficult to model spatial-spectral features and generate 3D attention weights at the same time, ignoring the cooperation between spatial features and spectral features. To address the above issues, an efficient ultra-lightweight spatial-spectral feature cooperation network (USSFC-Net) is proposed for CD in this paper. The proposed USSFC-Net has two main advantages. First, a multi-scale decoupled convolution (MSDConv) is designed, which is clearly different from the popular atrous spatial pyramid pooling (ASPP) module and its variants since it can flexibly capture the multi-scale features of changed objects by using cyclic multi-scale convolution. Meanwhile, the design of MSDConv can greatly reduce the number of parameters and computational redundancy. Second, an efficient spatial-spectral feature cooperation strategy (SSFC) is introduced to obtain richer features. The SSFC differs from existing 2D attention mechanisms since it learns 3D spatial-spectral attention weights without adding any parameters. The experiments on three datasets for remote sensing image CD demonstrate that the proposed USSFC-Net achieves better CD accuracy than most convolutional neural networks-based methods and requires lower computational costs and fewer parameters, even it is superior to some Transformer-based methods. The code is available at https://github.com/SUST-reynole/USSFC-Net.
Tao Lei 0003, Xinzhe Geng, Hailong Ning, Zhiyong Lv, Maoguo Gong, Yaochu Jin, Asoke K. Nandi
IEEE Trans. Geosci. Remote. Sens.3
2022 Audio-visual collaborative representation learning for Dynamic Saliency Prediction
Hailong Ning, Bin Zhao 0001, Zhanxuan Hu, Ercheng Pei
Knowl. Based Syst.1
2022 Difference Enhancement and Spatial-Spectral Nonlocal Network for Change Detection in VHR Remote Sensing Images
abstract
The popular Siamese convolutional neural networks (CNNs) for remote sensing (RS) image change detection (CD) often suffer from two problems. First, they either ignore the original information of bitemporal images or insufficiently utilize the difference information between bitemporal images, which leads to the low tightness of the changed objects. Second, Siamese CNNs always employ dual-branch encoders for CD, which increases computational cost. To address the above issues, this article proposes a network based on difference enhancement and spatial–spectral nonlocal (DESSN) for CD in very-high-resolution (VHR) images. This article makes threefold contributions. First, we design a difference enhancement (DE) module that can effectively learn the difference representation between foreground and background to reduce the impact of irrelevant changes on the detection results. Second, we present a spatial–spectral nonlocal (SSN) module that is different from vanilla nonlocal because multiscale spatial global features are incorporated to model the large-scale variation of objects during CD. The module can be used to strengthen the edge integrity and internal tightness of changed objects. Third, the asymmetric double convolution with Ghost (ADCG) module is exploited instead of standard convolution. The ADCG can not only refine the edge information of the changed objects, since horizontal and vertical convolutional kernels have good contour preservation advantages, but also greatly reduce the computational complexity of the proposed model. The experiments on two public VHR CD datasets demonstrate that the proposed network can provide higher detection accuracy and requires smaller memory usage than state-of-the-art networks.
Tao Lei 0003, Hailong Ning, Xingwu Wang, Dinghua Xue, Qi Wang 0009, Asoke K. Nandi
IEEE Trans. Geosci. Remote. Sens.3
2022 Semantics-Consistent Representation Learning for Remote Sensing Image-Voice Retrieval
abstract
With the development of earth observation technology, massive amounts of remote sensing (RS) images are acquired. To find useful information from these images, cross-modal RS image–voice retrieval provides a new insight. This article aims to study the task of RS image–voice retrieval so as to search effective information from massive amounts of RS data. Existing methods for RS image–voice retrieval rely primarily on the pairwise relationship to narrow the heterogeneous semantic gap between images and voices. However, apart from the pairwise relationship included in the data sets, the intramodality and nonpaired intermodality relationships should also be considered simultaneously since the semantic consistency among nonpaired representations plays an important role in the RS image–voice retrieval task. Inspired by this, a semantics-consistent representation learning (SCRL) method is proposed for RS image–voice retrieval. The main novelty is that the proposed method takes the pairwise, intramodality, and nonpaired intermodality relationships into account simultaneously, thereby improving the semantic consistency of the learned representations for the RS image–voice retrieval. The proposed SCRL method consists of two main steps: 1) semantics encoding and 2) SCRL. First, an image encoding network is adopted to extract high-level image features with a transfer learning strategy, and a voice encoding network with dilated convolution is devised to obtain high-level voice features. Second, a consistent representation space is conducted by modeling the three kinds of relationships to narrow the heterogeneous semantic gap and learn semantics-consistent representations across two modalities. Extensive experimental results on three challenging RS image–voice data sets, including Sydney, UCM, and RSICD image–voice data sets, show the effectiveness of the proposed method.
Hailong Ning, Bin Zhao 0001, Yuan Yuan 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 Disentangled Representation Learning for Cross-Modal Biometric Matching
abstract
Cross-modal biometric matching (CMBM) aims to determine the corresponding voice from a face, or identify the corresponding face from a voice. Recently, many CMBM methods have been proposed by forcing the distance between two modal features to be narrowed. However, these methods ignore the alignability between the two modal features. Because the feature is extracted under the supervision of identity information from single modal data, it can only reflect the identity information of single modal data. In order to address this problem, a disentangled representation learning method is proposed to disentangle the alignable latent identity factors and nonalignable the modality-dependent factors for CMBM. The proposed method consists of two main steps: 1) feature extraction and 2) disentangled representation learning. Firstly, an image feature extraction network is adopted to obtain face features, and a voice feature extraction network is applied to learn voice features. Secondly, a disentangled latent variable is explored to disentangle the latent identity factors that are shared across the modalities from the modality-dependent factors. The modality-dependent factors are filtered out, while the latent identity factors from the two modalities are enforced to be narrowed to align the same identity information. Then, the disentangled latent identity factors are considered as pure identity information to bridge the two modalities for cross-modal verification, 1:$N$matching, and retrieval. Note that the proposed method learns the identity information from the input face images and voice segments with only identity label as supervised information. Extensive experiments on the challenging VoxCeleb dataset demonstrate the proposed method outperforms the state-of-the-art methods.
Hailong Ning, Xiangtao Zheng, Xiaoqiang Lu, Yuan Yuan 0001
IEEE Trans. Multim.1
2021 Audio description from image by modal translation network
Hailong Ning, Xiangtao Zheng, Yuan Yuan 0001, Xiaoqiang Lu
Neurocomputing1
2021 Bio-Inspired Representation Learning for Visual Attention Prediction
abstract
Visual attention prediction (VAP) is a significant and imperative issue in the field of computer vision. Most of the existing VAP methods are based on deep learning. However, they do not fully take advantage of the low-level contrast features while generating the visual attention map. In this article, a novel VAP method is proposed to generate the visual attention map via bio-inspired representation learning. The bio-inspired representation learning combines both low-level contrast and high-level semantic features simultaneously, which are developed by the fact that the human eye is sensitive to the patches with high contrast and objects with high semantics. The proposed method is composed of three main steps: 1) feature extraction; 2) bio-inspired representation learning; and 3) visual attention map generation. First, the high-level semantic feature is extracted from the refined VGG16, while the low-level contrast feature is extracted by the proposed contrast feature extraction block in a deep network. Second, during bio-inspired representation learning, both the extracted low-level contrast and high-level semantic features are combined by the designed densely connected block, which is proposed to concatenate various features scale by scale. Finally, the weighted-fusion layer is exploited to generate the ultimate visual attention map based on the obtained representations after bio-inspired representation learning. Extensive experiments are performed to demonstrate the effectiveness of the proposed method.
Yuan Yuan 0001, Hailong Ning, Xiaoqiang Lu
IEEE Trans. Cybern.2