Nanqing Liu

dblp:309/2703 · DBLP profile ↗
← Back
20ranked-venue papers
6as first author
20since 2021 · last 2026
0000-0001-7564-4896ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 16 · 5 first-author · 16 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 STFAR: Test-time adaptive object detection through self-training and feature alignment regularization
Nanqing Liu, Yongyi Su, Lile Cai, Heng-Chao Li 0001, Kui Jia, Tianrui Li 0001, Xun Xu 0002, Chuan-Sheng Foo
Expert Syst. Appl.1
2026 Blueprint Multiscale Aware and Linearized Feature Enhancement Network for Efficient Remote Sensing Image Super-Resolution
abstract
Remote-Sensing image super-resolution (RSISR) technology aims to enhance the spatial resolution of low-resolution remote sensing images. Although deep learning-based super-resolution methods offer considerable theoretical advantages, their practical applications are severely limited by high memory consumption and computational costs. Existing lightweight RSISR methods typically reduce complexity by employing grouped or separated feature processing, but this can limit effective exploitation of feature interdependencies. To address this issue, this paper proposes a Blueprint Multiscale Aware and Linearized Feature Enhancement Network (BLNet). Specifically, the lightweight Blueprint MultiScale Aware Block (BMAB) is designed to extract and fuse multiscale information, tailored to the characteristics of remote sensing images. In addition, the Linearized Feature Enhancement Module (LFEM) is introduced to capture global spatial-channel features. Extensive experiments demonstrate that the proposed method significantly improves RSISR reconstruction performance while maintaining high computational efficiency. Our code will be available at https://github.com/crcherry/BLNet.
Nanqing Liu, Yun-Cheng Li, Sen Lei, Heng-Chao Li 0001
IEEE Geosci. Remote. Sens. Lett.2
2026 Point-Supervised Oriented Ship Detection via Segment Anything Model for SAR Images
abstract
Conventional ship detection methods for SAR images typically require complete annotations, which are time-consuming. Hence, we propose a framework named PointOSD-SAR that trains ship detectors using only point labels. We identify two main challenges in point-supervised detection: the lack of scale information and interference from SAR imaging noise. To tackle these, we design two collaborative processes: 1) Point-to-Mask: A SAR Ship Enhancement (SSE) module first preprocesses the SAR images to reduce noise and fill gaps in ship structures. Then, we fine-tune the Segment Anything Model (SAM) to obtain ship masks from point prompts. 2) Mask-to-HBox-to-RBox: We introduce a Noise Mask Filtering (NMF) process to filter low-quality masks and convert the remaining masks into Horizontal Boxes (HBox). Finally, a weakly supervised oriented object detection model learns the angle information from HBoxes. Extensive experiments show that our method achieves state-of-the-art performance among recent weakly supervised oriented ship detection methods, with 62.8% accuracy on the SSDD dataset and 55.3% on the HRSID dataset.
Qiwei Lin, Nanqing Liu, Zhiyuan Tan 0010, Qinghua Long
IEEE Geosci. Remote. Sens. Lett.2
2025 On the Adversarial Risk of Test Time Adaptation: An Investigation into Realistic Test-Time Data Poisoning
abstract
Test-time adaptation (TTA) updates the model weights during the inference stage using testing data to enhance generalization. However, this practice exposes TTA to adversarial risks. Existing studies have shown that when TTA is updated with crafted adversarial test samples, also known as test-time poisoned data, the performance on benign samples can deteriorate. Nonetheless, the perceived adversarial risk may be overstated if the poisoned data is generated under overly strong assumptions. In this work, we first review realistic assumptions for test-time data poisoning, including white-box versus grey-box attacks, access to benign data, attack order, and more. We then propose an effective and realistic attack method that better produces poisoned samples without access to benign samples, and derive an effective in-distribution attack objective. We also design two TTA-aware attack objectives. Our benchmarks of existing attack methods reveal that the TTA methods are more robust than previously believed. In addition, we analyze effective defense strategies to help develop adversarially robust TTA methods. The source code is available at https://github.com/Gorilla-Lab-SCUT/RTTDP.
Yongyi Su, Yushu Li, Nanqing Liu, Kui Jia, Xulei Yang, Chuan-Sheng Foo, Xun Xu 0002
ICLR3
2025 Memory-Augmented Differential Network for Infrared Small Target Detection
abstract
Traditional U-Net-based methods in infrared small target detection (IRSTD) have demonstrated good performance. However, they often struggle with challenges such as blurred contour and strong interference in complex backgrounds. To overcome these issues, we propose a memory-augmented differential network (MAD-Net), which integrates two key modules: the adaptive differential convolution module (AdaDCM) and the memory-augmented attention module (MemA2M). AdaDCM leverages multiple differential convolutions to capture detailed edge information, with an adaptive fusion mechanism to weight and aggregate these features. In the deeper layers, by introducing the dataset-level representations through a learnable memory bank (LMB), MemA2M can enhance current features and effectively mitigate background interference. Extensive experiments on four public IRSTD datasets demonstrate that MAD-Net outperforms state-of-the-art methods, showcasing its superior capability in handling complex scenarios. The code is available at:https://github.com/joan2joan/MAD-Net.
Yanqiong Liu, Sen Lei, Nanqing Liu, Heng-Chao Li 0001
IEEE Geosci. Remote. Sens. Lett.3
2025 PointSAM: Pointly-Supervised Segment Anything Model for Remote Sensing Images
abstract
Segment anything model (SAM) is an advanced foundational model for image segmentation, which is gradually being applied to remote sensing images (RSIs). Due to the domain gap between RSIs and natural images, traditional methods typically use SAM as a source pretrained model and fine-tune it with fully supervised masks. Unlike these methods, our work focuses on fine-tuning SAM using more convenient and challenging point annotations. Leveraging SAM’s zero-shot capability, we adopt a self-training framework that iteratively generates pseudolabels. However, noisy labels in pseudolabels can cause error accumulation. To address this, we introduce prototype-based regularization (PBR), where target prototypes are extracted from the dataset and matched to predicted prototypes using the Hungarian algorithm to guide learning in the correct direction. In addition, RSIs have complex backgrounds and densely packed objects, making it possible for point prompts to mistakenly group multiple objects as one. To resolve this, we propose a negative prompt calibration (NPC) method based on the nonoverlapping nature of instance masks, where overlapping masks are used as negative signals to refine segmentation. Combining these techniques, we present a novel pointly-supervised SAM (PointSAM). We conduct experiments on three RSI datasets, including WHU, HRSID, and NWPU VHR-10, showing that our method significantly outperforms direct testing with SAM, SAM2, and other comparison methods. In addition, PointSAM can act as a point-to-box converter for oriented object detection, achieving promising results and indicating its potential for other point-supervised tasks. The code is available athttps://github.com/Lans1ng/PointSAM.
Nanqing Liu, Xun Xu 0002, Yongyi Su, Heng-Chao Li 0001
IEEE Trans. Geosci. Remote. Sens.1
2025 Exploring Generalizable Pretraining for Real-World Change Detection via Geometric Estimation
abstract
As an essential procedure in earth observation system, change detection (CD) aims to reveal the spatial-temporal evolution of the observation regions. A key prerequisite for existing change detection algorithms is aligned geo-references between multi-temporal images by fine-grained registration. However, in the majority of real-world scenarios, a prior manual registration is required between the original images, which significantly increases the complexity of the CD workflow. In this paper, we proposed a self-supervision motivated CD framework with geometric estimation, called “MatchCD”. Specifically, the proposed MatchCD framework utilizes the zero-shot capability to optimize the encoder with self-supervised contrastive representation, which is reused in the downstream image registration and change detection to simultaneously handle the bi-temporal unalignment and object change issues. Moreover, unlike the conventional change detection requiring segmenting the full-frame image into small patches, our MatchCD framework can directly process the original large-scale image (e.g., 6K× 4Kresolutions) with promising performance. The performance in multiple complex scenarios with significant geometric distortion demonstrates the effectiveness of our proposed framework.
Sen Lei, Nanqing Liu, Heng-Chao Li 0001, Turgay Çelik 0001, Qing Zhu 0012
IEEE Trans. Geosci. Remote. Sens.3
2024 Clip-Guided Source-Free Object Detection in Aerial Images
abstract
Domain adaptation is crucial in aerial imagery, as the visual representation of these images can significantly vary based on factors such as geographic location, time, and weather conditions. Additionally, high-resolution aerial images often require substantial storage space and may not be readily accessible to the public. To address these challenges, we propose a novel Source-Free Object Detection (SFOD) method. Specifically, our approach begins with a self-training framework, which significantly enhances the performance of baseline methods. To alleviate the noisy labels in self-training, we utilize Contrastive Language-Image Pre-training (CLIP) to guide the generation of pseudo-labels, termed CLIP-guided Aggregation (CGA). By leveraging CLIP’s zero-shot classification capability, we aggregate its scores with the original predicted bounding boxes, enabling us to obtain refined scores for the pseudo-labels. To validate the effectiveness of our method, we constructed two new datasets from different domains based on the DIOR dataset, named DIOR-C and DIOR-Cloudy. Experimental results demonstrate that our method outperforms other comparative algorithms. The code is available at https://github.com/Lans1ng/SFOD-RS.
Nanqing Liu, Xun Xu 0002, Yongyi Su, Peiliang Gong, Heng-Chao Li 0001
IGARSS1
2024 Toward Distortion-Aware Change Detection in Realistic Scenarios
abstract
In the conventional change detection (CD) pipeline, two manually registered and labeled remote sensing datasets serve as the input of the model for training and prediction. However, in realistic scenarios, data from different periods or sensors could fail to be aligned as a result of various coordinate systems. Geometric distortion caused by coordinate shifting remains a thorny issue for CD algorithms. In this paper, we propose a reusable self-supervised framework for bitemporal geometric distortion in CD tasks. The whole framework is composed of Pretext Representation Pre-training, Bitemporal Image Alignment, and Down-stream Decoder Fine-Tuning. With only single-stage pre-training, the key components of the framework can be reused for assistance in the bitemporal image alignment, while simultaneously enhancing the performance of the CD decoder. Experimental results in 2 large-scale realistic scenarios demonstrate that our proposed method can alleviate the bitemporal geometric distortion in CD tasks.
Heng-Chao Li 0001, Nanqing Liu, Rui Wang 0090
IGARSS3
2024 PS-TTL: Prototype-based Soft-labels and Test-Time Learning for Few-shot Object Detection
abstract
In recent years, Few-Shot Object Detection (FSOD) has gained widespread attention and made significant progress due to its ability to build models with a good generalization power using extremely limited annotated data. The fine-tuning based paradigm is currently dominating this field, where detectors are initially pre-trained on base classes with sufficient samples and then fine-tuned on novel ones with few samples, but the scarcity of labeled samples of novel classes greatly interferes precisely fitting their data distribution, thus hampering the performance. To address this issue, we propose a new framework for FSOD, namely Prototype-based Soft-labels and Test-Time Learning (PS-TTL). Specifically, we design a Test-Time Learning (TTL) module that employs a mean-teacher network for self-training to discover novel instances from test data, allowing detectors to learn better representations and classifiers for novel classes. Furthermore, we notice that even though relatively low-confidence pseudo-labels exhibit classification confusion, they still tend to recall foreground. We thus develop a Prototype-based Soft-labels (PS) strategy through assessing similarities between low-confidence pseudo-labels and category prototypes as soft-labels to unleash their potential, which substantially mitigates the constraints posed by few-shot samples. Extensive experiments on both the VOC and COCO benchmarks show that PS-TTL achieves the state-of-the-art, highlighting its effectiveness. The code and model are available at https://github.com/gaoyingjay/PS-TTL.
Yingjie Gao 0001, Yanan Zhang 0005, Ziyue Huang 0001, Nanqing Liu, Di Huang 0001
ACM Multimedia4
2024 A Novel SegNet Model for Crack Image Semantic Segmentation in Bridge Inspection
Rong Pang, Xun Xu 0002, Nanqing Liu
PAKDD (3)5
2024 Masked Second-Order Pooling for Few-Shot Remote-Sensing Scene Classification
abstract
Few-shot remote-sensing scene classification (FSRSSC) is the task of categorizing remote-sensing images (RSIs) with insufficient labeled samples. This task becomes particularly challenging due to low intraclass similarity and high interclass similarity in RSIs. To overcome these challenges, we introduce a novel masked second-order pooling (MSoP) module to exploit the masked second-order features, enhancing feature representation and classification performance for FSRSSC. The MSoP module comprises two key components: a learnable DropBlock (LDB) and a compressed second-order pooling (CSoP). The LDB selectively masks discriminative regions in feature maps, effectively alleviating the low intraclass similarity problem. On the other hand, the CSoP enhances second-order statistics by computing and aggregating channel-wise similarity, thereby reducing the impact of interclass similarity. We construct an MSoP-Net based on the proposed MSoP module. Experimental results demonstrate its superior performance on two widely used datasets, that is, NWPU-RESISC45 and UCM.
Jianan Deng, Nanqing Liu
IEEE Geosci. Remote. Sens. Lett.3
2024 IDA-SiamNet: Interactive- and Dynamic-Aware Siamese Network for Building Change Detection
abstract
Building change detection (BCD) is a critical task in remote sensing which aims to identify the building changes within the same geographical area over time. The complexity of BCD is heightened when utilizing very high-resolution (VHR) remote sensing images, leading to two primary challenges: distinguishing between building and nonbuilding changes and accommodating the diverse range of building shapes and sizes. The existing mainstream methods neglect interactions between encoders, thereby compromising the ability to recognize building and nonbuilding changes. Additionally, most BCD methods overlook feature alignment and fusion which hinders the precise extraction of buildings with varying shapes and sizes. To address these limitations, we propose an interactive- and dynamic-aware Siamese network (IDA-SiamNet) for BCD. Our method comprises the spatial exchange feature interaction (SEFI) module, the channel exchange feature interaction (CEFI) module, and the dynamic-deformable dual-alignment fusion (D3AF) module. The SEFI and CEFI modules play a pivotal role in facilitating mutual information exchange between Siamese encoders, enhancing discrimination between building and nonbuilding changes. Furthermore, the D3AF module dynamically aggregates multiple parallel convolutional kernels to improve feature alignment and fusion for accurate building outline extraction. D3AF adapts its receptive field (RF) based on object size and covers diverse building shapes without introducing excessive background information. Experimental evaluations on three widely used BCD datasets, learning, vision, and remote sensing change detection (LEVIR-CD), satellite side-looking (S2Looking), and WHU BCD (WHU-CD), demonstrate the superior performance of our proposed method over state-of-the-art alternatives. Code and weights are made available athttps://github.com/SUPERMAN123000/IDA-SiamNet.
Yun-Cheng Li, Sen Lei, Nanqing Liu, Heng-Chao Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.3
2024 SSLChange: A Self-Supervised Change Detection Framework Based on Domain Adaptation
abstract
In conventional remote sensing change detection (RSCD) procedures, extensive manual labeling for bi-temporal images is first required to maintain the performance of subsequent fully supervised training. However, pixel-level labeling for change detection (CD) tasks is very complex and time-consuming. In this article, we explore a novel self-supervised contrastive framework applicable to the RSCD task, which promotes the model to accurately capture spatial, structural, and semantic information through the domain adapter (DA) and the hierarchical contrastive head. The proposed SSLChange framework accomplishes self-learning only by taking a single-temporal sample and can be flexibly transferred to mainstream CD baselines. With self-supervised contrastive learning, feature representation pretraining can be performed directly based on the original data even without labeling. After a certain number of labels are subsequently obtained, the pretrained features will be aligned with the labels for fully supervised fine-tuning. Without introducing any additional data or labels, the performance of downstream baselines will experience a significant enhancement. Experimental results on two entire datasets and six diluted datasets show that our proposed SSLChange improves the performance and stability of CD baseline in data-limited situations. The code of SSLChange is available athttps://github.com/MarsZhaoYT/SSLChange
Turgay Çelik 0001, Nanqing Liu, Feng Gao 0005, Heng-Chao Li 0001
IEEE Trans. Geosci. Remote. Sens.3
2023 Break Through the Border Restriction of Horizontal Bounding Box for Arbitrary-Oriented Ship Detection in SAR Images
abstract
Substantial progress has been made in detecting ships of arbitrary orientation in synthetic aperture radar (SAR) images. However, the mainstream method is still limited by the horizontal bounding box (HBB) boundary, which cannot provide scaling information for the length and width of the oriented bounding box (OBB) in an intuitive way. In this study, we propose a novel encode representation to describe the OBB by breaking through the border restriction of the HBB. Specifically, we derive an inclination factor from two left-top point offsets (LTPO), which enables us to directly infer the coordinates of the four OBB vertices and obtain an oriented rectangular proposal. To obtain high-quality oriented semantic features, we utilize a feature adaptive module (FAM) to learn the shape and orientation implied by arbitrary-oriented ships through spatial transformation. Our comparative experiments demonstrate that our proposed method achieves superior performance and detection accuracy on two commonly-used benchmark datasets for oriented SAR ship detection, namely SSDD and HRSID.
Turgay Çelik 0001, Nanqing Liu, Heng-Chao Li 0001
IEEE Geosci. Remote. Sens. Lett.3
2023 Transformation-Invariant Network for Few-Shot Object Detection in Remote-Sensing Images
abstract
Object detection in remote sensing images relies on a large amount of labeled data for training. However, the increasing number of new categories and class imbalance make exhaustive annotation impractical. Few-shot object detection (FSOD) addresses this issue by leveraging meta-learning on seen base classes and fine-tuning on novel classes with limited labeled samples. Nonetheless, the substantial scale and orientation variations of objects in remote sensing images pose significant challenges to existing few-shot object detection methods. To overcome these challenges, we propose integrating a feature pyramid network and utilizing prototype features to enhance query features, thereby improving existing FSOD methods. We refer to this modified FSOD approach as a Strong Baseline, which has demonstrated significant performance improvements compared to the original baselines. Furthermore, we tackle the issue of spatial misalignment caused by orientation variations between the query and support images by introducing a Transformation-Invariant Network (TINet). TINet ensures geometric invariance and explicitly aligns the features of the query and support branches, resulting in additional performance gains while maintaining the same inference speed as the Strong Baseline. Extensive experiments on three widely used remote sensing object detection datasets, i.e., NWPU VHR-10.v2, DIOR, and HRRSD demonstrated the effectiveness of the proposed method.
Nanqing Liu, Xun Xu 0002, Turgay Çelik 0001, Zongxin Gan, Heng-Chao Li 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 Gated Ladder-Shaped Feature Pyramid Network for Object Detection in Optical Remote Sensing Images
abstract
This letter presents a new feature pyramid network (FPN) called the gated ladder-shaped FPN (GLFPN) to construct more representative feature pyramids for detecting objects of different sizes in optical remote sensing images. We first use convolution and concatenation operations to fuse three base features extracted by a ResNet backbone. We then obtain multilevel features from these base features. Finally, we use a selective gate to fuse features from multiple levels with equivalent sizes. To evaluate the effectiveness of the proposed GLFPN, we integrate it into the RetinaNet architecture by replacing the conventional FPN. The experimental results on two optical remote sensing image data sets show that the proposed method outperforms the methods compared in this letter.
Nanqing Liu, Turgay Çelik 0001, Heng-Chao Li 0001
IEEE Geosci. Remote. Sens. Lett.1
2022 MSNet: A Multiple Supervision Network for Remote Sensing Scene Classification
abstract
Remote sensing scene classification is a complex task due to large intraclass variations in object appearances with a small number of samples per class and high interclass similarities due to shared objects in different classes, which usually cause model overfitting and high interclass confusion. To address these challenges, a multiple supervision approach, called multiple supervision network (MSNet), consisting of the ResNet-50 backbone, a feature discriminative branch (FDB), and a feature confusion branch (FCB) is proposed in this letter. The FDB selects discriminative features per class and suppresses peaks in feature maps to examine more informative regions with lower feature magnitudes. Meanwhile, the FCB reduces overfitting by introducing confusion to the input of a fully connected layer which also enhances the robust features. The FDB and FCB are only used in training of the backbone and not used in inference. Thus, the proposed method does not introduce additional computing time on the backbone while it significantly boosts its performance in scene classification. The experimental results show that MSNet outperforms the methods considered in this letter.
Nanqing Liu, Turgay Çelik 0001, Heng-Chao Li 0001
IEEE Geosci. Remote. Sens. Lett.1
2022 A Comparative Analysis of GAN-Based Methods for SAR-to-Optical Image Translation
abstract
Unlike optical sensors, Synthetic Aperture Radar (SAR) sensors acquire images of the Earth’s surface with all-weather and all-time capabilities, which is vital in a situation such as a disaster assessment. However, SAR sensors do not offer as rich visual information as optical sensors. SAR-to-Optical image-to-image translation generates optical images from SAR images to benefit from what both imaging modalities have to offer. It also enables multi-sensor image analysis of the same scene for applications such as heterogeneous change detection. Various architectures of Generative Adversarial Networks (GANs) have achieved remarkable image-to-image translation results in different domains. Still, their performances in SAR-to-Optical image translation have not been analyzed in the remote sensing domain. This paper compares and analyses the state-of-the-art GAN-based translation methods with open-source implementations for SAR-to-Optical image translation. The results show that GAN-based SAR-to-Optical image translation methods achieve satisfactory results. However, their performances depend on the structural complexity of the observed scene and the spatial resolution of the data. We also introduce a new dataset with a higher resolution than the existing SAR-to-Optical image datasets and release implementations of GAN-based methods considered in this paper to support the reproducible research in remote sensing.
Turgay Çelik 0001, Nanqing Liu, Heng-Chao Li 0001
IEEE Geosci. Remote. Sens. Lett.3
2022 An Arbitrary-Oriented Object Detector Based on Variant Gaussian Label in Remote Sensing Images
abstract
Arbitrary-oriented remote sensing object detection is a challenging task because of the periodicity of the angle and large variety of aspect ratios. To address these issues, this letter proposes two simple yet powerful methods: a variant Gaussian label (VGL) generating method and a channel-wise pixel attention (CPA) module. Specifically, VGL is generated to handle the periodicity of the angle and produce various labels to adapt diverse aspect ratios of objects. Then, CPA is utilized to fuse features between channels and pixels, to obtain a global receptive field and extract more robust features. In addition, we collect a dataset, namely, Northwestern Polytechnical University Very-High-Resolution (NWPU VHR-10-R), which is relabeled with oriented bounding boxes based on NWPU VHR-10. To evaluate our proposed method, experiments are conducted on several publicly available benchmark datasets, including the High Resolution Ship Collections 2016 (HRSC2016) and NWPU VHR-10-R. Experimental results show that our method achieves substantial gains compared with baseline approaches.
Tingyu Zhao, Nanqing Liu, Turgay Çelik 0001, Heng-Chao Li 0001
IEEE Geosci. Remote. Sens. Lett.2