Zhe Li 0015

dblp:11/751-15 · DBLP profile ↗
← Back
20ranked-venue papers
1as first author
14since 2021 · last 2026
0000-0001-7579-7702ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 A hybrid CNN-transformer network with difference enhancement and frequency fusion for remote sensing image change detection
Meiru Wang, Sheng Fang 0001, Xing-Li Zhang 0001, Zhe Li 0015
Appl. Intell.5
2026 Affinity maximization learning for unsupervised deep visual graph matching
Yuan Xie 0006, Zhe Li 0015, A. K. Qin 0001, Ming Li 0065
Pattern Recognit.3
2025 3D-HRSCD: Exploiting the Potential of Multiscale Features by 3-D Convolution
abstract
Semantic change detection (SCD) in remote sensing image (RSI) is critical for monitoring land cover and land use transformations. Although existing SCD methods have made progress in modeling temporal dependency, they still struggle to effectively capture multi-scale features and make interaction among them. To address these issues, we propose 3D-HRSCD, a novel architecture that utilizes 3D convolution to model temporal dependency across HRNet’s multi-resolution features. The core of this architecture is 3D Convolution Fusion Oriented to Multi-scale Features (3DFOM) module, which makes adequate interaction in channel, spatial and temporal dimensions across multi-scale features. To support more efficient temporal dependency modeling in 3DFOM, Cosine Similarity-based Temporal Multi-Scales Attention (CTMA) module serves as a preprocessing stage by enhancing features in change regions. Additionally, Comprehensive Semantic Consistency (CSC) loss function is introduced to further suppress pseudo-changes and reduce semantic recognition errors. Experimental results reveal that our method outperforms state-of-the-art performances relative to previous SCD efforts.
Yue Song 0008, Sheng Fang 0001, Zhe Li 0015, Enyi Zhao
IEEE Geosci. Remote. Sens. Lett.3
2025 Hyperspectral image restoration via the collaboration of low-rank tensor denoising and completion
Tianheng Zhang, Jianli Zhao 0002, Sheng Fang 0001, Zhe Li 0015, Maoguo Gong
Pattern Recognit.4
2025 Rethinking Semantic Change Detection From a Semantic Alignment Perspective
Sheng Fang 0001, Wen Li 0041, Yue Song 0008, Zhe Li 0015, Jianli Zhao 0002
IEEE Trans. Geosci. Remote. Sens.4
2024 Unsupervised SAR Change Detection Using Two-Stage Pseudo Labels Refining Framework
abstract
Unsupervised change detection (CD) in Synthetic Aperture Radar (SAR) imagery is pivotal for terrestrial observations, more so for disaster-related applications. However, most existing deep learning methods primarily emphasize the construction of diverse networks, often overlooking the critical aspect of refining pseudo labels. This letter proposes a two-stage pseudo labels refining (TSPLR) framework for SAR image unsupervised CD. During the first stage, Fuzzy-C-Means (FCM) clustering is employed on the bi-temporal SAR images to yield initial pseudo labels, categorizing high-confidence data as changed or unchanged and the rest as uncertain. A straightforward network is then trained first with confident data, with the training subsequently extended to the uncertain data. In the second stage, we start by comparing the predictions from the model trained during the first phase with the initial pseudo labels. Data with inconsistencies is added to the uncertain dataset, and some pixels filtered according to connectivity areas are selected as hard samples. Subsequently, the model is trained further using the updated dataset. This two-stage refinement process bolsters the credibility of pseudo labels and creates a more robust network training against speckle noise. The experimental results show that even with a simple network, the performance of the proposed TSPLR exceeds the current performance of SOTA. We will be making the source codes publicly accessible at https://github.com/sdust-mmlab.
Sheng Fang 0001, Chenxu Qi, Shuqi Yang, Zhe Li 0015
IEEE Geosci. Remote. Sens. Lett.4
2024 BT-HRSCD: High-Resolution Feature Is What You Need for a Semantic Change Detection Network With a Triple-Decoding Branch
abstract
In recent years, semantic change detection (SCD) has emerged as a pivotal field within the remote sensing (RS) research community, underscored by its essential contribution to various Earth observation undertakings. Conventional SCD methodologies typically adopt a multitask network architecture, fusing a binary change detection (BCD) sub-task with dual semantic segmentation (SS) sub-tasks. These strategies frequently rely on the encoder’s low-resolution yet semantically dense features, derived from multiple down-sampling stages, as the inputs for the decoding heads. Departing from this traditional path and targeting the nuanced characteristics of the multisubtasks, this study pioneers a novel methodology that harnesses the potential of the encoding phase’s high-resolution features. By integrating HRNet as the encoder structure, we introduce the BT-HRSCD framework, featuring two simple and effective modules. The first, bidirectional shallow and deep features aggregation module (BiFAM), seeks to imbue features with richer semantic insights through bidirectional feature fusion that spans from shallow-to-deep as well as deep-to-shallow layers. The second module, high-resolution difference extraction (HRDE), utilizes the encoder’s highest spatial resolution features, evaluating their differences to enhance the precision in identifying change areas. BiFAM is devised to boost the SS sub-tasks’ effectiveness, whereas HRDE aims to elevate the accuracy of the BCD sub-task. Experimental results reveal that our method outperforms state-of-the-art performances relative to previous SCD efforts. Our source code is released athttps://github.com/iridescent524/BT-HRSCD.
Sheng Fang 0001, Wen Li 0041, Shuqi Yang, Zhe Li 0015, Jianli Zhao 0002
IEEE Trans. Geosci. Remote. Sens.4
2024 A Decoder-Focused Multitask Network for Semantic Change Detection
abstract
Recently, Semantic Change Detection (SCD) has gained growing attention from the Remote Sensing (RS) research community due to its critical role in Earth observation applications. Typical approaches tackle the task using a multi-task network, comprising one Change Detection (CD) sub-task and two Semantic Segmentation (SS) sub-tasks. Although these approaches have achieved good performance, one crucial question persists: What is the effective way to handle the feature interactions across SCD sub-tasks? To address this issue, this paper first offers an overview of existing SCD networks and compares them from a perspective view of Multi-Task Learning (MTL). Following that, we select an architecture combining a two-branch encoder and a three-branch decoder as the baseline due to its compatibility with MTL. Then, one simple yet very effective module, decoder feature interaction across sub-tasks (DFIT), is introduced. DFIT seeks to enhance the CD decoding feature by leveraging the feature differences between two SS decoding branches on a layer-wise basis. Additionally, the feature aggregation module (FAM) is designed further to enhance the network performance in cooperation with DFIT. FAM aims to produce more representative shared information across the SS and CD sub-tasks by merging the outputs from the final three encoder layers. Combining DFIT and FAM, the proposed network exploiting Decoder-Focused MTL (DEFO-MTLSCD) presents more representative information by capitalizing on both CD and SS losses back-propagations across all coding paths and achieves better performance. Experimental results reveal that our method outperforms state-of-the-art performances relative to previous SCD efforts. Our source code is released at https://github.com/byyztgxz/Decoder_Fusion.
Zhe Li 0015, Sheng Fang 0001, Jianli Zhao 0002, Shuqi Yang, Wen Li 0041
IEEE Trans. Geosci. Remote. Sens.1
2024 Full-Mode-Augmentation Tensor-Train Rank Minimization for Hyperspectral Image Inpainting
abstract
Hyperspectral image (HSI) inpainting is a fundamental task in remote sensing image processing, which is helpful for subsequent applications such as classification and unmixing. Recently, tensor-train decomposition (TTD)-based low-rank methods have achieved great success in image inpainting because of the balanced tensor unfolding and the use of tensor augmentation (TA). However, the TTD algorithm only performs TA on the third mode, and cannot effectively mine the spectral domain information of HSIs. Aiming at this problem, this article extends TA to each mode of the tensor and proposes a full-mode-augmentation TT decomposition (FTTD). More precisely, the HSI is cast along each mode to get a 3-D$N$th-order tensor sequence, and all the tensors in the sequence are performed TTD to get$3N$factor tensors that model the correlation on various modes. Then, the full-mode-augmentation tensor unfolding method is given by performing TT unfolding on each$N$th-order tensor. We implement FTTD by minimizing the rank of the unfolding matrix and thus provide the tensor completion framework based on full-mode-augmentation tensor-train rank minimization. Finally, using the framework, we optimize the current two classical iterative algorithms nuclear norm minimization and parallel matrix decomposition. These two algorithms utilize the unfolding matrices in the framework to effectively mine spatial and spectral information, thereby making the restored tensor more approximate to the original HSI. Experiments on various HSIs have shown that the proposed method outperforms compared methods in terms of visual and quantitative measures.
Tian-Heng Zhang, Jianli Zhao 0002, Sheng Fang 0001, Zhe Li 0015, Maoguo Gong
IEEE Trans. Geosci. Remote. Sens.4
2023 Changer: Feature Interaction is What You Need for Change Detection
abstract
Change detection is an important tool for long-term earth observation missions. It takes bi-temporal images as input and predicts “where” the change has occurred. Different from other dense prediction tasks, a meaningful consideration for change detection is the interaction between bi-temporal features. With this motivation, in this paper we propose a novel general change detection architecture, MetaChanger, which includes a series of alternative interaction layers in the feature extractor. To verify the effectiveness of MetaChanger, we propose two derived models, ChangerAD and ChangerEx with simple interaction strategies: Aggregation-Distribution (AD) and feature “exchange”. AD is abstracted from some complex interaction methods, and feature “exchange” is a completely parameter&computation-free operation by exchanging bi-temporal features. In addition, for better alignment of bi-temporal features, we propose a Flow-based Dual-Alignment Fusion (FDAF) module which allows interactive alignment and feature fusion. Crucially, we observe Changer series models achieve competitive performance on different scale change detection datasets. Further, our proposed ChangerAD and ChangerEx could serve as a starting baseline for future MetaChanger design. Code and weights are made available at https://github.com/likyoo/open-cd.
Sheng Fang 0001, Kaiyu Li 0001, Zhe Li 0015
IEEE Trans. Geosci. Remote. Sens.3
2022 S²ENet: Spatial-Spectral Cross-Modal Enhancement Network for Classification of Hyperspectral and LiDAR Data
abstract
The effective utilization of multimodal data (e.g., hyperspectral and light detection and ranging (LiDAR) data) has profound implications for further development of the remote sensing (RS) field. Many studies have explored how to effectively fuse features from multiple modalities; however, few of them focus on information interactions that can effectively promote the complementary semantic content of multisource data before fusion. In this letter, we propose a spatial–spectral enhancement module (S2EM) for cross-modal information interaction in deep neural networks. Specifically, S2EM consists of SpAtial Enhancement Module (SAEM) for enhancing spatial representation of hyperspectral data by LiDAR features and SpEctral Enhancement Module (SEEM) for enhancing spectral representation of LiDAR data by hyperspectral features. A series of experiments and ablation studies on the Houston2013 dataset show that S2EM can effectively facilitate the interaction and understanding between multimodal data. Our source code is available athttps://github.com/likyoo/Multimodal-Remote-Sensing-Toolkit, contributing to the RS community.
Sheng Fang 0001, Kaiyu Li 0001, Zhe Li 0015
IEEE Geosci. Remote. Sens. Lett.3
2022 SNUNet-CD: A Densely Connected Siamese Network for Change Detection of VHR Images
abstract
Change detection is an important task in remote sensing (RS) image analysis. It is widely used in natural disaster monitoring and assessment, land resource planning, and other fields. As a pixel-to-pixel prediction task, change detection is sensitive about the utilization of the original position information. Recent change detection methods always focus on the extraction of deep change semantic feature, but ignore the importance of shallow-layer information containing high-resolution and fine-grained features, this often leads to the uncertainty of the pixels at the edge of the changed target and the determination miss of small targets. In this letter, we propose a densely connected siamese network for change detection, namely SNUNet-CD (the combination of Siamese network and NestedUNet). SNUNet-CD alleviates the loss of localization information in the deep layers of neural network through compact information transmission between encoder and decoder, and between decoder and decoder. In addition, Ensemble Channel Attention Module (ECAM) is proposed for deep supervision. Through ECAM, the most representative features of different semantic levels can be refined and used for the final classification. Experimental results show that our method improves greatly on many evaluation criteria and has a better tradeoff between accuracy and calculation amount than other state-of-the-art (SOTA) change detection methods.
Sheng Fang 0001, Kaiyu Li 0001, Jinyuan Shao, Zhe Li 0015
IEEE Geosci. Remote. Sens. Lett.4
2021 Effective multiple pedestrian tracking system in video surveillance with monocular stationary camera
Zhihui Wang 0003, Ming Li 0065, Yu Lu 0006, Yongtang Bao, Zhe Li 0015, Jianli Zhao 0002
Expert Syst. Appl.5
2021 A visual tracking algorithm via confidence-based multi-feature correlation filtering
Sheng Fang 0001, Yichen Ma, Zhe Li 0015
Multim. Tools Appl.3
2019 Optimal CTU-level bit allocation in HEVC for low bit-rate applications
Cui Ni, Zhe Li 0015, Guangyuan Zhang
Multim. Tools Appl.3
2019 Highly Paralleled Low-Cost Embedded HEVC Video Encoder on TI KeyStone Multicore DSP
abstract
Although HEVC, the emerging video coding standard, has doubled the coding performance of its predecessor H.264/AVC, its significantly increased computational complexity imposes great obstacles for HEVC encoders to be employed in real-time applications with embedded processors, such as digital signal processors (DSPs). In this paper, a TI Keystone multicore TMS320C6678 DSP-based highly paralleled low-cost fast HEVC encoding solution is well designed and implemented. First, the overall structure of HEVC encoder with CTU-level parallelism is re-designed to well support the encoding parallelism, with full consideration of the hardware characteristics. Second, a low-delay and low-memory multicore data transmission mechanism is proposed to reduce the latency of data access between internal L2 memory and external DDR3. Third, the encoding bottlenecks, i.e., the most time-consuming encoding modules, are identified and optimized for acceleration with TI powerful C6000 SIMD instructions. Experimental results show that our proposed HEVC encoder on TI TMS320C6678 DSPs can significantly improve the real-time capacity with tolerable performance loss, 0.93 dB performance loss under on average 465.50 times speedup as compared to CPU-based HM reference software, more specifically, which makes it desirable in power-constrained real-time video applications.
Rui Fan 0002, Yongfei Zhang, Gang Wang 0023, Zhe Li 0015
IEEE Trans. Circuits Syst. Video Technol.5
2018 Multi-hypothesis-Based Error Concealment for Whole Frame Loss in HEVC
Yongfei Zhang, Zhe Li 0015
MMM (1)2
2014 An Improved Similarity-Based Fast Coding Unit Depth Decision Algorithm for Inter-frame Coding in HEVC
Rui Fan 0002, Yongfei Zhang, Zhe Li 0015
MMM (1)3
2013 Fast Coding Unit Depth Decision Algorithm for Interframe Coding in HEVC
abstract
As the next generation standard of video coding, the High Efficiency Video Coding (HEVC) achieves significantly better coding efficiency than all existing video coding standards. A Coding Unit (CU) quad tree concept is introduced to HEVC to improve the coding efficiency. Each CU node in quad tree will be traversed by depth first search process to find the best Coding Tree Unit (CTU) partition. Although this quad tree search process can obtain the best CTU partition, it is very time consuming, especially in interframe coding. To alleviate the encoder computation load in interframe coding, a fast CU depth decision method is proposed by reducing the depth search range. Based on the depth information correlation between spatio-temporal adjacent CTUs and the current CTU, some depths can be adaptively excluded from the depth search process in advance. Experimental results show that the proposed scheme provides almost 30% encoder time savings on average compared to the default encoding scheme in HM8.0 with only 0.38% bit rate increment in coding performance.
Yongfei Zhang, Zhe Li 0015
DCC3
2012 Gradient-based fast decision for intra prediction in HEVC
abstract
As the next generation standard of video coding, the High Efficiency Video Coding(HEVC) achieves significantly better coding efficiency than all existing video coding standards, which is however at the cost of a much higher computation complexity. To address this issue, this paper presents a gradient-based fast decision algorithm for intra prediction in HEVC. More specifically, the intra prediction in HEVC is divided into two stages: prediction unit(PU) size decision and mode decision. At the PU size decision process, four orientation features are extracted from the coding unit by the intensity gradient filters to decide the texture complexity and texture direction of the coding unit, and then the texture direction is used to exclude impossible prediction modes at the mode decision process. Compared to HEVC reference software, the proposed algorithm saves around 56.7% of the encoding time in intra high efficiency setting and up to 70.86% in intra low complexity setting with slight performance degradation.
Yongfei Zhang, Zhe Li 0015, Bo Li 0006
VCIP2