Nan Wang 0038

dblp:84/864-38 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
9since 2021 · last 2026
0009-0002-9821-405XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CVGD: Cross-View Guided Disentangler for Multi-View Radar Semantic Segmentation
abstract
Neural networks continue to face challenges in effectively extracting information from the high-dimensional and sparse range-azimuth-Doppler (RAD) tensor. While multi-view architectures using three-axis projections of the RAD tensor improve efficiency and maintain accuracy, existing designs suffer from inefficient inter-view information interaction. To address this issue, this letter proposes a new multi-view information fusion viewpoint: instead of directly blending feature maps, the orthogonal relationships among the RAD coordinate axes are leveraged to guide feature decomposition across views, thereby disentangling the mixed semantics introduced during view compression. Based on this idea, a new module, termed Cross-View Guided Disentangler (CVGD), is introduced to enable inter-view information interaction in multi-view architectures. Extensive experiments on the public CARRADA dataset demonstrate that networks incorporating the proposed module achieve performance comparable to state-of-the-art methods, while utilizing 50% fewer parameters and attaining a 6-fold increase in inference speed. Additionally, promising results are also observed on RADIal, achieving near-SOTA performance with higher efficiency.
Yaoyu He, Tao Shan, Nan Wang 0038
IEEE Signal Process. Lett.4
2025 High-Resolution Remote Sensing Change Detection With Edge-Guided Feature Enhancement
abstract
High-resolution (HR) remote sensing image change detection aims to identify surface changes; however, complex scenes and irregular object edges pose significant challenges to achieving accurate results. Existing methods leverage upsampling, downsampling, or dilated convolution to capture multiscale spatial features and fuse fine-scale details into coarse-scale features using concatenation, addition, or skip connections to enhance edge information. However, these direct fusion operations can cause fine edge details to be overshadowed by dominant regional features. To address this, we propose an edge-guided change detection (EGCD) network that improves edge preservation and detection accuracy. In the encoding stage, a region-edge feature extraction module (REM) is introduced to extract regional and edge features in parallel using a two-branch structure for each temporal image. The edge and regional features from the two temporal images are then fused independently via a separation feature fusion (SFF) module, preventing fine edge details from being dominated by regional features. In the decoding stage, a edge enhancement upsampling (EEU) module uses edge features to guide the reconstruction of regional features, ensuring precise boundary delineation. Experiments on public datasets validate the effectiveness and robustness of the proposed network.
Changyuan You, Nan Wang 0038, Dehui Zhu, Wei Li 0032
IEEE Geosci. Remote. Sens. Lett.2
2025 DG²-TCR: An Adaptive Clouds Removal Network for Optical Remote Sensing Images Using SAR-Driven Dual-Flow Fusion Guidance
abstract
Clouds in optical remote sensing images (ORSI) significantly limit image utilization. Traditional cloud removal methods using single or multi-temporal data sources struggle to ensure reliable reconstruction for thick cloud areas. Synthetic Aperture Radar (SAR) images are increasingly used to recover information obscured by clouds, but their performance in cloud-obscured regions is unstable. Therefore, an adaptive cloud removal network for remote sensing images, named DG2-TCR, is proposed based on SAR-driven dual-flow fusion guidance (DFG). DG2-TCR uses SAR and ORSI to construct DFG, including local spatial-spectral feature reconstruction (LSSFR) flow and global texture feature compensation (GTFC). LSSFR, driven by ORSI and SAR, efficiently extracts useful features in non-cloud areas and focuses on local information reconstruction using the designed spatial-spectral features inference reconstruction block (SSIRB). Based on SAR images, GTFC guides the compensation of global texture information. DFG can adaptively extract features and reconstruct missing information from local and global scales. The public SEN12MS-CR-TS dataset is divided into four sub-datasets with different coverage to evaluate the recovering capability in varying clouds. Experiments show that the PSNR, SSIM, RMSE, FID, and NCC indicator values on four sub-datasets and the SIMLE-CR dataset are better than the seven comparison methods. Furthermore, the ablation experiments show that the generalization and robustness of this proposed method on images with different cloud coverage are better than other comparison methods. Therefore, DG2-TCR can reliably recover information on cloud occlusions with various coverage and thickness, which is significant for cloud removal in practical applications.
Xianjun Gao, Jinhui Yang, Xudong Xie, Yuanwei Yang, Nan Wang 0038, Xinran Cao, Meilin Tan, Yuan Kou
IEEE Trans. Geosci. Remote. Sens.5
2024 Multi-Temporal Images Generation for Building Change Detection Performance Promotion
abstract
The changes in building are important basis for urban monitoring. However, due to the rarity and sparsity of the occurrence of changes in buildings, collecting effective bitemporal image pairs is challenging, as it requires long-term observation over several months or even years. Additionally, annotating large-scale change detection datasets is time-consuming and labor-intensive. Consequently, data scarcity issues lead to insufficient training of building change detection models. To address this, we propose a data generation method Building Generation GAN (BG-GAN). Different from other GANs, the BG-GAN is trained based on adversarial consistency loss, enabling the model to generate new bi-temporal image pairs with various types of building changes. To verify the effectiveness of the proposed methods on change detection task, BG-GAN is utilized to perform building change samples generation on two building change detection datasets (LEVIR-CD and WHU-CD). The experimental results demonstrate that the proposed method can improve the robustness and generalization of change detection model to detect pseudo changes.
Yute Li, Wei Li 0032, Nan Wang 0038, Chenzhong Gao, Yin Zhuang, He Chen 0004
IGARSS3
2024 Few-Shot Fine-Grained Classification With Rotation-Invariant Feature Map Complementary Reconstruction Network
abstract
Fine-grained classification is of significant importance in the field of remote sensing. However, obtaining valuable and rare target images is often a challenging task, giving rise to the few-shot fine-grained classification problem. In response to this challenge, various meta-learning approaches have been introduced, with the feature map reconstruction network emerging as a prominent method. Targets in remote sensing images exhibit arbitrary orientation, substantial inter-class similarity and intra-class diversity. Nevertheless, the conventional feature map reconstruction network exhibits subpar performance due to its inability to handle rotational variations. Moreover, it only reconstructs features from a single channel dimension of support features, neglecting the interplay between different dimensions and resulting in inaccurate reconstruction errors. To overcome the challenges of imprecise rotational variation features for reconstruction and inaccurate reconstruction errors, we propose a rotation-invariant feature map complementary reconstruction network (RIFCRN). The RIFCRN involves several key innovations. First, we introduce a novel rotation-invariant module (RIM) based on active rotating filters and oriented response pooling, enabling the extraction of rotation-invariant features for reconstruction. This modification enhances the suitability of the feature map reconstruction network for the few-shot fine-grained classification problem. Second, we put forward a novel feature map complementary reconstruction (CPR) method that calculates the complementary reconstruction errors (CRE) which effectively captures relationships among different feature map dimensions and results in more accurate reconstruction errors. Finally, extensive experiments have been conducted to validate the effectiveness of the proposed RIFCRN in addressing the few-shot fine-grained classification problem. The code will be available at https://github.com/liyangfan0/RIFCRN.
Yangfan Li 0002, Liang Chen 0004, Wei Li 0032, Nan Wang 0038
IEEE Trans. Geosci. Remote. Sens.4
2024 Object Tracking in Satellite Videos With Distractor-Occlusion-Aware Correlation Particle Filters
abstract
With the advancement of high-resolution remote sensing satellites, the tracking of high-value targets such as planes and ships within satellite videos has become imperative. In recent years, several object tracking methods designed for satellite videos based on correlation filters have been proposed. However, these traditional correlation filters typically identify the location with the highest response value on the response map as the target position. In the context of satellite videos, where targets are often very small and surrounded by numerous similar objects, depending only on response values to determine the target’s location can easily lead to interference from nearby objects, resulting in tracking failures. Moreover, targets frequently encounter occlusion during their motion, further complicating tracking tasks due to the absence of distinctive target appearance features and leading to the issue of model drift. To address these challenges, we propose a novel distractor-occlusion aware correlation particle filters. Instead of determining the position with the maximum response value, our method initially selects the top k response values from the response map, creating a pool of candidates. Subsequently, we introduce an innovative quality score, rooted in motion information and response scores related to the target, for each candidate. Finally, these quality scores are employed to filter the most suitable candidate. This novel distractor-aware module effectively equips our tracking method to perform well in the presence of distractors. Additionally, to handle occlusion, we integrate the occlusion-aware module into the correlation particle filters, improving the tracker’s performance in occluded scenarios. To ensure the effective collaboration of the distractor-aware module and the occlusion-aware module, we introduce the dual Kalman Filter method. Our comprehensive experiments conducted on the SatSOT datasets conclusively demonstrate the effectiveness and superiority of our proposed tracking method. The code will be available at https://github.com/liyangfan0/DOCPF.
Yangfan Li 0002, Nan Wang 0038, Wei Li 0032, Mengbin Rao
IEEE Trans. Geosci. Remote. Sens.2
2024 Spatio-Temporal Feature Fusion and Guide Aggregation Network for Remote Sensing Change Detection
abstract
The field of remote sensing change detection (RSCD) has seen significant advancements recently, focusing on the precise identification and analysis of temporal changes in remote sensing images. Existing deep learning-based RSCD methods primarily rely on concatenation or subtraction to integrate features of bi-temporal images and reconstruct change features through a feature pyramid network (FPN) decoding architecture. However, these methods face challenges related to inadequate spatio-temporal change representation and insufficient aggregation of multilevel semantic information, resulting in pseudo-changes and poor completeness of detected change objects. In this article, we propose an innovative RSCD framework via spatio-temporal feature fusion and guide aggregation (STFF-GA) to address the aforementioned challenges. The architecture of this network comprises two key components: the STFF module and the GA module. The STFF module is designed as a low-parameter and low-computation structure, effectively enhancing the representation of spatio-temporal change information through split, interaction, and fusion strategies. The GA module uses deep feature guidance (DFG) mapping as prior information to guide the aggregation of multilevel semantic information, thereby correcting the positional information of change objects and filtering out pseudo-changes and other noise interference. In addition, it utilizes convolution kernels of various scales to extract fine-grained features, facilitating the complete reconstruction of change objects. Extensive experiments conducted on three benchmark change detection datasets demonstrate that the proposed STFF-GA consistently outperforms other state-of-the-art (SOTA) detectors. The code is available athttps://github.com/NjustHGWei/STFF-GA.
Hongguang Wei, Nan Wang 0038, Yuan Liu 0015, Pengge Ma, Dongdong Pang, Xiubao Sui, Qian Chen 0002
IEEE Trans. Geosci. Remote. Sens.2
2023 CSTSUNet: A Cross Swin Transformer-Based Siamese U-Shape Network for Change Detection in Remote Sensing Images
abstract
Change detection (CD) in remote sensing images is a critical task that has achieved significant success by deep learning. Current networks often employ pixel-based differencing, proportion, classification-based, or feature concatenation methods to represent changes of interest. However, these methods fail to effectively detect the desired changes, as they are highly sensitive to factors such as atmospheric conditions, lighting variations, and phenological variations, resulting in detection errors. Inspired by the Transformer structure, we adopt a cross-attention mechanism to more robustly extract feature differences between bitemporal images. The motivation of the method is based on the assumption that if there is no change between image pairs, the semantic features from one temporal image can well be represented by the semantic features from another temporal image. Conversely if there is a change, there are significant reconstruction errors. Therefore, a Cross Swin Transformer based Siamese U-shaped network namely CSTSUNet is proposed for remote sensing change detection. CSTSUnet consists of encoder, difference feature extraction, and decoder. The encoder is based on a hierarchical Resnet with the Siamese U-net structure, allowing parallel processing of bitemporal images and extraction of multi-scale features. The difference feature extraction consists of four difference feature extraction modules that compute difference feature at multiple scales. In this module, Cross Swin Transformer is employed in each difference feature extraction module to communicate the information of bitemporal images. The decoder takes in the multi-scale difference features as input, injects details and boundaries iteratively level by level, and makes the change map more and more accurate. We conduct experiments on three public datasets, and the experimental results demonstrate that the proposed CSTSUNet outperforms other state-of-the-art methods in terms of both qualitative and quantitative analyses. Our code is available at https://github.com/l7170/CSTSUNet.git.
Yaping Wu, Lu Li 0005, Nan Wang 0038, Wei Li 0032, Junfang Fan, Ran Tao 0003
IEEE Trans. Geosci. Remote. Sens.3
2023 Unsupervised Pansharpening Method Using Residual Network With Spatial Texture Attention
abstract
Recently, deep learning has become one of the most popular tools for pansharpening, many relevant methods have been investigated and reflected great performance. However, a non-negligible problem is the absence of ground-truth (GT). A common solution is using degraded images as training input and the original images are employed as GT. The learned mapping between low resolution (LR) and high resolution (HR) is simulated, is not real, which may cause spectral distortion or insufficient spatial texture enhancement of fused images. In order to address the drawback, a novel unsupervised attention pansharpening net (UAP-Net) is proposed. The proposed UAP-Net mainly contains two major components: 1) the deep residual network (DRN) and 2) spatial texture attention block (STAB). The DRN aims to extract spectral features and spatial details features from low-resolution multi-spectral (LRMS) and panchromatic (PAN), and to fuse those features to make them more representative. The designed STAB adopts the high-frequency component of corresponding input PAN as the weight to enhance the spatial details of the residual block output features. Moreover, a new loss function including two spatial losses and two spectral losses are established. The losses are calculated in the spatial domain and the frequency domain, respectively. Experiments on Gaofen-2 and Worldview-2 remote sensing data demonstrate that the proposed UAP-Net could fuse PAN and LRMS images effectively without the help of high-resolution multi-spectral (HRMS). The proposed framework is fully general and can be used for many multisource remote sensing image fusion, and achieves optimal performance in terms of both the subjective visual effect and the quantitative evaluation.
Zhangxi Xiong, Na Liu 0014, Nan Wang 0038, Wei Li 0032
IEEE Trans. Geosci. Remote. Sens.3