Sheng Fang 0001

dblp:40/5370-1 · DBLP profile ↗
← Back
18ranked-venue papers
7as first author
18since 2021 · last 2026
0009-0002-1078-2184ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 A hybrid CNN-transformer network with difference enhancement and frequency fusion for remote sensing image change detection
Meiru Wang, Sheng Fang 0001, Xing-Li Zhang 0001, Zhe Li 0015
Appl. Intell.2
2026 Federated intent-aware cross-domain recommendation via semantic alignment and collaborative enhancement
Jianli Zhao 0002, Sheng Fang 0001, Qingqian Guan
Neurocomputing3
2026 Spatial-Spectral Texture-Preserved Total Variation: A Novel Regularization for Hyperspectral Image Denoising
abstract
Local smoothness is a widely used prior in hyperspectral image (HSI) denoising tasks. The current work mainly realizes the representation of this prior through total variation (TV) regularization. However, the TV regularization applies a uniform penalty to each entry in the image, unable to effectively balance noise removal and texture preservation. Aiming at this problem: 1) We propose a novel regularization for HSI denoising called Spatial-Spectral Texture-Preserved Total Variation (SSTPTV). This naturally expresses the physical phenomenon of the difference in sparsity between textured regions and smooth regions. Specifically, the regularization relaxes the sparsity penalty in textured regions by a weight learning strategy and sparsity measurement method for gradient maps, thereby preserving the spatial-spectral textures of HSIs. 2) An HSI denoising model based on the SSTPTV regularization constraint is given. We propose a tensor alternating subspace representation method that can capture the overall spatial texture features across all bands and the overall spectral texture features across all spectral curves. By applying the SSTPTV regularization constraint to these subspaces, the spatial-spectral texture structures are effectively preserved. An efficient ADMM-based algorithm for solving the model is designed. The simulated and real noise removal experiments of HSI prove that the proposed method has significant superiority and can serve as a framework to optimize other TV-based denoising methods. The code is available at https://github.com/zth-code/SSTPTV.
Jianli Zhao 0002, Tian-Heng Zhang, Sheng Fang 0001, Jian-Feng Gao, Jin-Yu Wang, Maoguo Gong
IEEE Trans. Circuits Syst. Video Technol.3
2025 Open-CD: A Comprehensive Toolbox for Change Detection
abstract
We present Open-CD, a change detection toolbox that contains a rich set of change detection methods as well as related components and modules. The toolbox started from a series of open source general vision task tools, including OpenMMLab Toolkits, PyTorch Image Models (Timm), etc. It gradually evolves into a unified platform that covers many popular change detection methods and contemporary modules. It not only includes training and inference codes, but also provides some useful scripts for data analysis. We believe this toolbox is by far the most comprehensive change detection toolbox. In this report, we introduce the features, supported methods and applications of Open-CD. In addition, we also conduct a benchmarking study on different methods and components. We wish that the toolbox and benchmark could serve the growing research community by providing a flexible toolkit to re-implement existing methods and develop their own new change detectors. Code and models are available at https://github.com/likyoo/open-cd.
Kaiyu Li 0001, Chengxi Han, Yupeng Deng 0001, Keyan Chen 0001, Zhuo Zheng, Hao Chen 0045, Ziyuan Liu 0006, Yuantao Gu, Zhengxia Zou, Zhenwei Shi 0001, Sheng Fang 0001, Deyu Meng, Zhi Wang 0002, Xiangyong Cao
ACM Multimedia12
2025 3D-HRSCD: Exploiting the Potential of Multiscale Features by 3-D Convolution
abstract
Semantic change detection (SCD) in remote sensing image (RSI) is critical for monitoring land cover and land use transformations. Although existing SCD methods have made progress in modeling temporal dependency, they still struggle to effectively capture multi-scale features and make interaction among them. To address these issues, we propose 3D-HRSCD, a novel architecture that utilizes 3D convolution to model temporal dependency across HRNet’s multi-resolution features. The core of this architecture is 3D Convolution Fusion Oriented to Multi-scale Features (3DFOM) module, which makes adequate interaction in channel, spatial and temporal dimensions across multi-scale features. To support more efficient temporal dependency modeling in 3DFOM, Cosine Similarity-based Temporal Multi-Scales Attention (CTMA) module serves as a preprocessing stage by enhancing features in change regions. Additionally, Comprehensive Semantic Consistency (CSC) loss function is introduced to further suppress pseudo-changes and reduce semantic recognition errors. Experimental results reveal that our method outperforms state-of-the-art performances relative to previous SCD efforts.
Yue Song 0008, Sheng Fang 0001, Zhe Li 0015, Enyi Zhao
IEEE Geosci. Remote. Sens. Lett.2
2025 Tensor Completion via Nonlocal Tensor Wheel Decomposition for Hyperspectral Image Recovery
abstract
Tensor completion aims to recover original data from its degraded observations and has shown promise in processing incomplete hyperspectral images (HSIs). Many previous studies have indicated that global correlation and nonlocal self-similarity (NSS) are two important priors for tensor completion. Recently, tensor wheel (TW) decomposition has demonstrated excellent performance in the tensor completion problem. However, it overlooks NSS. To address this limitation, a novel nonlocal tensor wheel (NL-TW) decomposition-based method is presented for tensor completion, which leverages both of the aforementioned priors. The method consists of two main steps. In the first step, we introduce TW decomposition to the entire degraded tensor to generate an initial completion result. In the second step, similar patches in the initial completion result are first stacked into NSS groups, and TW decomposition is then applied to each grouped tensor to obtain the final completion result. In addition, the NL-TW decomposition-based tensor completion method is solved by a proximal alternating minimization (PAM)-based optimization algorithm, which has a theoretical convergence guarantee. Extensive experiments on three real datasets with different sampling rates (SRs) demonstrate the effectiveness of the proposed NL-TW method compared to other methods.
Jianli Zhao 0002, Xingzhao Feng, Sheng Fang 0001, Tian-Heng Zhang
IEEE Geosci. Remote. Sens. Lett.3
2025 Hyperspectral image restoration via the collaboration of low-rank tensor denoising and completion
Tianheng Zhang, Jianli Zhao 0002, Sheng Fang 0001, Zhe Li 0015, Maoguo Gong
Pattern Recognit.3
2025 Robust Tensor Completion via Spatial-Spectral Constrained Deep Low-Rank Tensor Factorization for Hyperspectral Image Recovery
abstract
Robust tensor completion of hyperspectral image (HSI) is a challenging task in the field of remote sensing. Recently, nuclear norm minimization-based methods have made certain progress in robust tensor completion. However, the tensor nuclear norm applies the same constraint to all singular values, resulting in insufficient capturing power for the global structure of the HSI. In addition, as a convex surrogate of global low-rankness, tensor nuclear norm minimization leads to an overall low-rank approximation that cannot capture the details of the HSI. In this letter, we propose the spatial-spectral constrained deep low-rank tensor factorization (SDLTF). More precisely, the low-rank tensor factorization is used to dynamically assign penalty weights, aiming to preserve the main information and maintain the global structure of the HSI. The spatial-spectral constrained unsupervised deep prior is applied within a deep convolutional neural network to capture spatial-spectral correlations and local details of the HSI. We develop an efficient algorithm to tackle the corresponding model based on the ADMM. Extensive experiments demonstrate that our model has superior performance compared with several state-of-the-art methods.
Jianli Zhao 0002, Jian-Feng Gao, Sheng Fang 0001, Tian-Heng Zhang, Jin-Yu Wang
IEEE Signal Process. Lett.3
2025 Rethinking Semantic Change Detection From a Semantic Alignment Perspective
Sheng Fang 0001, Wen Li 0041, Yue Song 0008, Zhe Li 0015, Jianli Zhao 0002
IEEE Trans. Geosci. Remote. Sens.1
2024 T3SRS: Tensor Train Transformer for compressing sequential recommender systems
Hao Li 0009, Jianli Zhao 0002, Huan Huo, Sheng Fang 0001, Jianjian Chen, Lutong Yao, Yiran Hua
Expert Syst. Appl.4
2024 Unsupervised SAR Change Detection Using Two-Stage Pseudo Labels Refining Framework
abstract
Unsupervised change detection (CD) in Synthetic Aperture Radar (SAR) imagery is pivotal for terrestrial observations, more so for disaster-related applications. However, most existing deep learning methods primarily emphasize the construction of diverse networks, often overlooking the critical aspect of refining pseudo labels. This letter proposes a two-stage pseudo labels refining (TSPLR) framework for SAR image unsupervised CD. During the first stage, Fuzzy-C-Means (FCM) clustering is employed on the bi-temporal SAR images to yield initial pseudo labels, categorizing high-confidence data as changed or unchanged and the rest as uncertain. A straightforward network is then trained first with confident data, with the training subsequently extended to the uncertain data. In the second stage, we start by comparing the predictions from the model trained during the first phase with the initial pseudo labels. Data with inconsistencies is added to the uncertain dataset, and some pixels filtered according to connectivity areas are selected as hard samples. Subsequently, the model is trained further using the updated dataset. This two-stage refinement process bolsters the credibility of pseudo labels and creates a more robust network training against speckle noise. The experimental results show that even with a simple network, the performance of the proposed TSPLR exceeds the current performance of SOTA. We will be making the source codes publicly accessible at https://github.com/sdust-mmlab.
Sheng Fang 0001, Chenxu Qi, Shuqi Yang, Zhe Li 0015
IEEE Geosci. Remote. Sens. Lett.1
2024 BT-HRSCD: High-Resolution Feature Is What You Need for a Semantic Change Detection Network With a Triple-Decoding Branch
abstract
In recent years, semantic change detection (SCD) has emerged as a pivotal field within the remote sensing (RS) research community, underscored by its essential contribution to various Earth observation undertakings. Conventional SCD methodologies typically adopt a multitask network architecture, fusing a binary change detection (BCD) sub-task with dual semantic segmentation (SS) sub-tasks. These strategies frequently rely on the encoder’s low-resolution yet semantically dense features, derived from multiple down-sampling stages, as the inputs for the decoding heads. Departing from this traditional path and targeting the nuanced characteristics of the multisubtasks, this study pioneers a novel methodology that harnesses the potential of the encoding phase’s high-resolution features. By integrating HRNet as the encoder structure, we introduce the BT-HRSCD framework, featuring two simple and effective modules. The first, bidirectional shallow and deep features aggregation module (BiFAM), seeks to imbue features with richer semantic insights through bidirectional feature fusion that spans from shallow-to-deep as well as deep-to-shallow layers. The second module, high-resolution difference extraction (HRDE), utilizes the encoder’s highest spatial resolution features, evaluating their differences to enhance the precision in identifying change areas. BiFAM is devised to boost the SS sub-tasks’ effectiveness, whereas HRDE aims to elevate the accuracy of the BCD sub-task. Experimental results reveal that our method outperforms state-of-the-art performances relative to previous SCD efforts. Our source code is released athttps://github.com/iridescent524/BT-HRSCD.
Sheng Fang 0001, Wen Li 0041, Shuqi Yang, Zhe Li 0015, Jianli Zhao 0002
IEEE Trans. Geosci. Remote. Sens.1
2024 A Decoder-Focused Multitask Network for Semantic Change Detection
abstract
Recently, Semantic Change Detection (SCD) has gained growing attention from the Remote Sensing (RS) research community due to its critical role in Earth observation applications. Typical approaches tackle the task using a multi-task network, comprising one Change Detection (CD) sub-task and two Semantic Segmentation (SS) sub-tasks. Although these approaches have achieved good performance, one crucial question persists: What is the effective way to handle the feature interactions across SCD sub-tasks? To address this issue, this paper first offers an overview of existing SCD networks and compares them from a perspective view of Multi-Task Learning (MTL). Following that, we select an architecture combining a two-branch encoder and a three-branch decoder as the baseline due to its compatibility with MTL. Then, one simple yet very effective module, decoder feature interaction across sub-tasks (DFIT), is introduced. DFIT seeks to enhance the CD decoding feature by leveraging the feature differences between two SS decoding branches on a layer-wise basis. Additionally, the feature aggregation module (FAM) is designed further to enhance the network performance in cooperation with DFIT. FAM aims to produce more representative shared information across the SS and CD sub-tasks by merging the outputs from the final three encoder layers. Combining DFIT and FAM, the proposed network exploiting Decoder-Focused MTL (DEFO-MTLSCD) presents more representative information by capitalizing on both CD and SS losses back-propagations across all coding paths and achieves better performance. Experimental results reveal that our method outperforms state-of-the-art performances relative to previous SCD efforts. Our source code is released at https://github.com/byyztgxz/Decoder_Fusion.
Zhe Li 0015, Sheng Fang 0001, Jianli Zhao 0002, Shuqi Yang, Wen Li 0041
IEEE Trans. Geosci. Remote. Sens.3
2024 Full-Mode-Augmentation Tensor-Train Rank Minimization for Hyperspectral Image Inpainting
abstract
Hyperspectral image (HSI) inpainting is a fundamental task in remote sensing image processing, which is helpful for subsequent applications such as classification and unmixing. Recently, tensor-train decomposition (TTD)-based low-rank methods have achieved great success in image inpainting because of the balanced tensor unfolding and the use of tensor augmentation (TA). However, the TTD algorithm only performs TA on the third mode, and cannot effectively mine the spectral domain information of HSIs. Aiming at this problem, this article extends TA to each mode of the tensor and proposes a full-mode-augmentation TT decomposition (FTTD). More precisely, the HSI is cast along each mode to get a 3-D$N$th-order tensor sequence, and all the tensors in the sequence are performed TTD to get$3N$factor tensors that model the correlation on various modes. Then, the full-mode-augmentation tensor unfolding method is given by performing TT unfolding on each$N$th-order tensor. We implement FTTD by minimizing the rank of the unfolding matrix and thus provide the tensor completion framework based on full-mode-augmentation tensor-train rank minimization. Finally, using the framework, we optimize the current two classical iterative algorithms nuclear norm minimization and parallel matrix decomposition. These two algorithms utilize the unfolding matrices in the framework to effectively mine spatial and spectral information, thereby making the restored tensor more approximate to the original HSI. Experiments on various HSIs have shown that the proposed method outperforms compared methods in terms of visual and quantitative measures.
Tian-Heng Zhang, Jianli Zhao 0002, Sheng Fang 0001, Zhe Li 0015, Maoguo Gong
IEEE Trans. Geosci. Remote. Sens.3
2023 Changer: Feature Interaction is What You Need for Change Detection
abstract
Change detection is an important tool for long-term earth observation missions. It takes bi-temporal images as input and predicts “where” the change has occurred. Different from other dense prediction tasks, a meaningful consideration for change detection is the interaction between bi-temporal features. With this motivation, in this paper we propose a novel general change detection architecture, MetaChanger, which includes a series of alternative interaction layers in the feature extractor. To verify the effectiveness of MetaChanger, we propose two derived models, ChangerAD and ChangerEx with simple interaction strategies: Aggregation-Distribution (AD) and feature “exchange”. AD is abstracted from some complex interaction methods, and feature “exchange” is a completely parameter&computation-free operation by exchanging bi-temporal features. In addition, for better alignment of bi-temporal features, we propose a Flow-based Dual-Alignment Fusion (FDAF) module which allows interactive alignment and feature fusion. Crucially, we observe Changer series models achieve competitive performance on different scale change detection datasets. Further, our proposed ChangerAD and ChangerEx could serve as a starting baseline for future MetaChanger design. Code and weights are made available at https://github.com/likyoo/open-cd.
Sheng Fang 0001, Kaiyu Li 0001, Zhe Li 0015
IEEE Trans. Geosci. Remote. Sens.1
2022 S²ENet: Spatial-Spectral Cross-Modal Enhancement Network for Classification of Hyperspectral and LiDAR Data
abstract
The effective utilization of multimodal data (e.g., hyperspectral and light detection and ranging (LiDAR) data) has profound implications for further development of the remote sensing (RS) field. Many studies have explored how to effectively fuse features from multiple modalities; however, few of them focus on information interactions that can effectively promote the complementary semantic content of multisource data before fusion. In this letter, we propose a spatial–spectral enhancement module (S2EM) for cross-modal information interaction in deep neural networks. Specifically, S2EM consists of SpAtial Enhancement Module (SAEM) for enhancing spatial representation of hyperspectral data by LiDAR features and SpEctral Enhancement Module (SEEM) for enhancing spectral representation of LiDAR data by hyperspectral features. A series of experiments and ablation studies on the Houston2013 dataset show that S2EM can effectively facilitate the interaction and understanding between multimodal data. Our source code is available athttps://github.com/likyoo/Multimodal-Remote-Sensing-Toolkit, contributing to the RS community.
Sheng Fang 0001, Kaiyu Li 0001, Zhe Li 0015
IEEE Geosci. Remote. Sens. Lett.1
2022 SNUNet-CD: A Densely Connected Siamese Network for Change Detection of VHR Images
abstract
Change detection is an important task in remote sensing (RS) image analysis. It is widely used in natural disaster monitoring and assessment, land resource planning, and other fields. As a pixel-to-pixel prediction task, change detection is sensitive about the utilization of the original position information. Recent change detection methods always focus on the extraction of deep change semantic feature, but ignore the importance of shallow-layer information containing high-resolution and fine-grained features, this often leads to the uncertainty of the pixels at the edge of the changed target and the determination miss of small targets. In this letter, we propose a densely connected siamese network for change detection, namely SNUNet-CD (the combination of Siamese network and NestedUNet). SNUNet-CD alleviates the loss of localization information in the deep layers of neural network through compact information transmission between encoder and decoder, and between decoder and decoder. In addition, Ensemble Channel Attention Module (ECAM) is proposed for deep supervision. Through ECAM, the most representative features of different semantic levels can be refined and used for the final classification. Experimental results show that our method improves greatly on many evaluation criteria and has a better tradeoff between accuracy and calculation amount than other state-of-the-art (SOTA) change detection methods.
Sheng Fang 0001, Kaiyu Li 0001, Jinyuan Shao, Zhe Li 0015
IEEE Geosci. Remote. Sens. Lett.1
2021 A visual tracking algorithm via confidence-based multi-feature correlation filtering
Sheng Fang 0001, Yichen Ma, Zhe Li 0015
Multim. Tools Appl.1