Sen Lei

dblp:203/6769 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0002-8010-3282ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 15 · 6 first-author · 12 since 2021Artificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 Blueprint Multiscale Aware and Linearized Feature Enhancement Network for Efficient Remote Sensing Image Super-Resolution
abstract
Remote-Sensing image super-resolution (RSISR) technology aims to enhance the spatial resolution of low-resolution remote sensing images. Although deep learning-based super-resolution methods offer considerable theoretical advantages, their practical applications are severely limited by high memory consumption and computational costs. Existing lightweight RSISR methods typically reduce complexity by employing grouped or separated feature processing, but this can limit effective exploitation of feature interdependencies. To address this issue, this paper proposes a Blueprint Multiscale Aware and Linearized Feature Enhancement Network (BLNet). Specifically, the lightweight Blueprint MultiScale Aware Block (BMAB) is designed to extract and fuse multiscale information, tailored to the characteristics of remote sensing images. In addition, the Linearized Feature Enhancement Module (LFEM) is introduced to capture global spatial-channel features. Extensive experiments demonstrate that the proposed method significantly improves RSISR reconstruction performance while maintaining high computational efficiency. Our code will be available at https://github.com/crcherry/BLNet.
Nanqing Liu, Yun-Cheng Li, Sen Lei, Heng-Chao Li 0001
IEEE Geosci. Remote. Sens. Lett.4
2026 MLAgg-UNet: Advancing Medical Image Segmentation With Efficient Transformer and Mamba-Inspired Multi-Scale Sequence
abstract
Transformers and state space sequence models (SSMs) have attracted interest in biomedical image segmentation for their ability to capture longrange dependency. However, traditional visual state space (VSS) methods suffer from the incompatibility of image tokens with autoregressive assumption. Although Transformer attention does not require this assumption, its high computational cost limits effective channelwise information utilization. To overcome these limitations, we propose the Mamba-Like Aggregated UNet (MLAgg-UNet), which introduces Mamba-inspired mechanism to enrich Transformer channel representation and exploit implicit autoregressive characteristic within U-shaped architecture. For establishing dependencies among image tokens in single scale, the Mamba-Like Aggregated Attention (MLAgg) block is designed to balance representational ability and computational efficiency. Inspired by the human foveal vision system, Mamba macro-structure, and differential attention, MLAgg block can slide its focus over each image token, suppress irrelevant tokens, and simultaneously strengthen channel-wise information utilization. Moreover, leveraging causal relationships between consecutive low-level and high-level features in U-shaped architecture, we propose the Multi-Scale Mamba Module with Implicit Causality (MSMM) to optimize complementary information across scales. Embedded within skip connections, this module enhances semantic consistency between encoder and decoder features. Extensive experiments on four benchmark datasets, including AbdomenMRI, ACDC, BTCV, and EndoVis17, which cover MRI, CT, and endoscopy modalities, demonstrate that the proposed MLAgg-UNet consistently outperforms state-of-the-art CNN-based, Transformer-based, and Mamba-based methods. Specifically, it achieves improvements of at least 1.24%, 0.20%, 0.33%, and 0.39% in DSC scores on these datasets, respectively. These results highlight the model's ability to effectively capture feature correlations and integrate complementary multi-scale information, providing a robust solution for medical image segmentation.
Jiaxu Jiang, Sen Lei, Heng-Chao Li 0001
IEEE J. Biomed. Health Informatics2
2025 Memory-Augmented Differential Network for Infrared Small Target Detection
abstract
Traditional U-Net-based methods in infrared small target detection (IRSTD) have demonstrated good performance. However, they often struggle with challenges such as blurred contour and strong interference in complex backgrounds. To overcome these issues, we propose a memory-augmented differential network (MAD-Net), which integrates two key modules: the adaptive differential convolution module (AdaDCM) and the memory-augmented attention module (MemA2M). AdaDCM leverages multiple differential convolutions to capture detailed edge information, with an adaptive fusion mechanism to weight and aggregate these features. In the deeper layers, by introducing the dataset-level representations through a learnable memory bank (LMB), MemA2M can enhance current features and effectively mitigate background interference. Extensive experiments on four public IRSTD datasets demonstrate that MAD-Net outperforms state-of-the-art methods, showcasing its superior capability in handling complex scenarios. The code is available at:https://github.com/joan2joan/MAD-Net.
Yanqiong Liu, Sen Lei, Nanqing Liu, Heng-Chao Li 0001
IEEE Geosci. Remote. Sens. Lett.2
2025 Exploring Fine-Grained Image-Text Alignment for Referring Remote Sensing Image Segmentation
abstract
Given a language expression, referring remote sensing image segmentation (RRSIS) aims to identify ground objects and assign pixelwise labels within the imagery. One of the key challenges for this task is to capture discriminative multimodal features via image-text alignment. However, the existing RRSIS methods use one vanilla and coarse alignment, where the language expression is directly extracted to be fused with the visual features. In this article, we argue that a “fine-grained image-text alignment” can improve the extraction of multimodal information. To this point, we propose a new RRSIS method to fully exploit the visual and linguistic representations. Specifically, the original referring expression is regarded as context text, which is further decoupled into the ground object and spatial position texts. The proposed fine-grained image-text alignment module (FIAM) would simultaneously leverage the features of the input image and the corresponding texts, obtaining better discriminative multimodal representation. Meanwhile, to handle the various scales of ground objects in remote sensing, we introduce a text-aware multiscale enhancement module (TMEM) to adaptively perform cross-scale fusion and intersections. We evaluate the effectiveness of the proposed method on two public referring remote sensing datasets including RefSegRS and RRSIS-D, and our method obtains superior performance over several state-of-the-art methods. The code will be publicly available athttps://github.com/Shaosifan/FIANet.
Sen Lei, Xinyu Xiao, Heng-Chao Li 0001, Zhenwei Shi 0001, Qing Zhu 0012
IEEE Trans. Geosci. Remote. Sens.1
2025 SAM-Based Building Change Detection With Distribution-Aware Fourier Adaptation and Edge-Constrained Warping
abstract
Building change detection remains challenging for urban development, disaster assessment, and military reconnaissance. While foundation models like Segment Anything Model (SAM) show strong segmentation capability, they exhibit limited performance in building change detection due to the domain gap between natural and remote sensing images. Existing adapter-based fine-tuning methods struggle with the imbalanced distribution of changed buildings, resulting in suboptimal detection of building changes exhibiting sparse distribution, small scale, or weak contrast. Additionally, bi-temporal alignment methods, including optical flow, are vulnerable to background noise interference. To address these limitations, we propose the SAM-based Network with Distribution-Aware Fourier Adaptation and Edge-Constrained Warping (FAEWNet) for building change detection. Unlike previous adapters overlook the imbalanced distribution of changed buildings, our proposed Distribution-Aware Fourier Aggregation Adapter not only addresses the domain gap issue, but also models the distribution of changed buildings. Furthermore, to mitigate noise interference and misalignment caused by registration errors, we design Multiscale Aware Flow Aggregation module that refines building edge extraction and enhances the perception of changed buildings. The results on the LEVIR-CD, S2Looking and WHU-CD datasets highlight the effectiveness of FAEWNet. Specifically, FAEWNet achieves the best performance on the LEVIR-CD with 91.29%Rc, 92.41%F1, and 85.89%IoU. On the S2Looking dataset, our method achieves improvements of 3.09% inRc, 1.07% inF1, and 1.23% inIoUcompared to the second-best method. On the WHU-CD dataset, it achieves the highest overall performance, with aRcof 94.20%, anF1 of 94.99%, and anIoUof 90.45%. The code is available at https://github.com/SUPERMAN123000/FAEWNet.
Yun-Cheng Li, Sen Lei, Heng-Chao Li 0001, Jun Li 0009, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.2
2025 Exploring Generalizable Pretraining for Real-World Change Detection via Geometric Estimation
abstract
As an essential procedure in earth observation system, change detection (CD) aims to reveal the spatial-temporal evolution of the observation regions. A key prerequisite for existing change detection algorithms is aligned geo-references between multi-temporal images by fine-grained registration. However, in the majority of real-world scenarios, a prior manual registration is required between the original images, which significantly increases the complexity of the CD workflow. In this paper, we proposed a self-supervision motivated CD framework with geometric estimation, called “MatchCD”. Specifically, the proposed MatchCD framework utilizes the zero-shot capability to optimize the encoder with self-supervised contrastive representation, which is reused in the downstream image registration and change detection to simultaneously handle the bi-temporal unalignment and object change issues. Moreover, unlike the conventional change detection requiring segmenting the full-frame image into small patches, our MatchCD framework can directly process the original large-scale image (e.g., 6K× 4Kresolutions) with promising performance. The performance in multiple complex scenarios with significant geometric distortion demonstrates the effectiveness of our proposed framework.
Sen Lei, Nanqing Liu, Heng-Chao Li 0001, Turgay Çelik 0001, Qing Zhu 0012
IEEE Trans. Geosci. Remote. Sens.2
2024 IDA-SiamNet: Interactive- and Dynamic-Aware Siamese Network for Building Change Detection
abstract
Building change detection (BCD) is a critical task in remote sensing which aims to identify the building changes within the same geographical area over time. The complexity of BCD is heightened when utilizing very high-resolution (VHR) remote sensing images, leading to two primary challenges: distinguishing between building and nonbuilding changes and accommodating the diverse range of building shapes and sizes. The existing mainstream methods neglect interactions between encoders, thereby compromising the ability to recognize building and nonbuilding changes. Additionally, most BCD methods overlook feature alignment and fusion which hinders the precise extraction of buildings with varying shapes and sizes. To address these limitations, we propose an interactive- and dynamic-aware Siamese network (IDA-SiamNet) for BCD. Our method comprises the spatial exchange feature interaction (SEFI) module, the channel exchange feature interaction (CEFI) module, and the dynamic-deformable dual-alignment fusion (D3AF) module. The SEFI and CEFI modules play a pivotal role in facilitating mutual information exchange between Siamese encoders, enhancing discrimination between building and nonbuilding changes. Furthermore, the D3AF module dynamically aggregates multiple parallel convolutional kernels to improve feature alignment and fusion for accurate building outline extraction. D3AF adapts its receptive field (RF) based on object size and covers diverse building shapes without introducing excessive background information. Experimental evaluations on three widely used BCD datasets, learning, vision, and remote sensing change detection (LEVIR-CD), satellite side-looking (S2Looking), and WHU BCD (WHU-CD), demonstrate the superior performance of our proposed method over state-of-the-art alternatives. Code and weights are made available athttps://github.com/SUPERMAN123000/IDA-SiamNet.
Yun-Cheng Li, Sen Lei, Nanqing Liu, Heng-Chao Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 HiReNet: Hierarchical-Relation Network for Few-Shot Remote Sensing Image Scene Classification
abstract
Few-shot scene classification aims to develop models that can quickly adapt to new scenes with only a few labeled samples that are not present in training sets. In recent years, convolutional neural networks (CNNs) have made significant advancements in few-shot remote sensing image scene classification tasks. However, most existing approaches focus solely on utilizing high-level embeddings of remote sensing images to learn similarity relations, while neglecting intrinsic hierarchical representations that could be crucial in distinguishing scenes with substantial interclass similarities. To address this limitation, we propose a novel few-shot scene classification method for remote sensing images called hierarchical-relation network (HiReNet). This approach leverages the hierarchical features of a query sample and its corresponding support sample to learn discriminative representations. HiReNet consists of an embedding network and a relation network. The embedding network employs a Siamese architecture to extract representations, while the relation network utilizes these representations for classification. Within the relation network, we introduce a hierarchical relation learning (HRL) structure to capture the hierarchical relations among query and support samples. Additionally, to extract stronger features, we introduce a feature aggregation module that concatenates multilevel features and employs channel attention to re- weight these features. Experimental results demonstrate the superior performance of our HiReNet compared to several state-of-the-art few-shot scene classification methods.
Sen Lei, Yingbo Zhou 0001, Jialin Cheng, Guohao Liang, Zhengxia Zou, Heng-Chao Li 0001, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.2
2023 Continuous Remote Sensing Image Super-Resolution Based on Context Interaction in Implicit Function Space
abstract
Despite its fruitful applications in remote sensing, image super-resolution is troublesome to train and deploy as it handles different resolution magnifications with separate models. Accordingly, we propose a highly-applicable super-resolution framework called FunSR, which settles different magnifications with a unified model by exploiting context interaction within implicit function space. FunSR composes a functional representor, a functional interactor, and a functional parser. Specifically, the representor transforms the low-resolution image from Euclidean space to multi-scale pixel-wise function maps; the interactor enables pixel-wise function expression with global dependencies; and the parser, which is parameterized by the interactor’s output, converts the discrete coordinates with additional attributes to RGB values. Extensive experimental results demonstrate that FunSR reports state-of-the-art performance on both fixed-magnification and continuous-magnification settings, meanwhile, it provides many friendly applications thanks to its unified nature. Our code is available at https://github.com/KyanChen/FunSR.
Keyan Chen 0001, Wenyuan Li 0002, Sen Lei, Jianqi Chen, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 Hybrid-Scale Self-Similarity Exploitation for Remote Sensing Image Super-Resolution
abstract
Recently, deep convolutional neural networks (CNNs) have made great progress in remote sensing image super-resolution (SR). The CNN-based methods can learn powerful feature representation from plenty of low- and high-resolution counterparts. For remote sensing images, there are many similar ground targets recurred inside the image itself, both within the same scale and across different scales. In this article, we argue that this internal recurrence can be used for learning stronger feature representation, and we propose a new hybrid-scale self-similarity exploitation network (HSENet) for remote sensing image SR. Specifically, we introduce a single-scale self-similarity exploitation module (SSEM) to compute the feature correlation within the same scale image. Moreover, we design a cross-scale connection structure (CCS) to capture the recurrences across different scales. By combining SSEM and CCS, we further develop a hybrid-scale self-similarity exploitation module (HSEM) to construct the final HSENet, which simultaneously exploits single- and cross-scale similarities. Experimental results demonstrate that HSENet can obtain superior performance over several state-of-the-art methods. Besides, the effectiveness of our method is also verified by the assistance to the remote sensing scene classification task.
Sen Lei, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 Transformer-Based Multistage Enhancement for Remote Sensing Image Super-Resolution
abstract
Convolutional neural networks have made a great breakthrough in recent remote sensing image super-resolution (SR) tasks. Most of these methods adopt upsampling layers at the end of the models to perform enlargement, which ignores feature extraction in the high-dimension space, and thus, limits SR performance. To address this problem, we propose a new SR framework for remote sensing images to enhance the high-dimensional feature representation after the upsampling layers. We name the proposed method as a transformer-based enhancement network (TransENet), where transformers are introduced to exploit features at different levels. The core of the TransENet is a transformer-based multistage enhancement structure, which can be combined with traditional SR frameworks to fuse multiscale high-/low-dimension features. Specifically, in this structure, the encoders aim to embed the multilevel features in the feature extraction part and the decoders are used to fuse these encoded embeddings. Experimental results demonstrate that our proposed TransENet can improve super-resolved results and obtain superior performance over several state-of-the-art methods.
Sen Lei, Zhenwei Shi 0001, Wenjing Mo
IEEE Trans. Geosci. Remote. Sens.1
2022 Large-Factor Super-Resolution of Remote Sensing Images With Spectra-Guided Generative Adversarial Networks
Yapeng Meng, Wenyuan Li 0002, Sen Lei, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.3
2020 Deep Adversarial Decomposition: A Unified Framework for Separating Superimposed Images
abstract
Separating individual image layers from a single mixed image has long been an important but challenging task. We propose a unified framework named "deep adversarial decomposition" for single superimposed image separation. Our method deals with both linear and non-linear mixtures under an adversarial training paradigm. Considering the layer separating ambiguity that given a single mixed input, there could be an infinite number of possible solutions, we introduce a "Separation-Critic" - a discriminative network which is trained to identify whether the output layers are well-separated and thus further improves the layer separation. We also introduce a "crossroad L1" loss function, which computes the distance between the unordered outputs and their references in a crossover manner so that the training can be well-instructed with pixel-wise supervision. Experimental results suggest that our method significantly outperforms other popular image separation frameworks. Without specific tuning, our method achieves the state of the art results on multiple computer vision tasks, including the image deraining, photo reflection removal, and image shadow removal.
Zhengxia Zou, Sen Lei, Tianyang Shi, Zhenwei Shi 0001, Jieping Ye
CVPR2
2020 Coupled Adversarial Training for Remote Sensing Image Super-Resolution
abstract
Generative adversarial network (GAN) has made great progress in recent natural image super-resolution tasks. The key to its success is the integration of a discriminator which is trained to classify whether the input is a real high-resolution (HR) image or a generated one. Arguably, learning a strong discriminative prior is essential for generating high-quality images. However, in remote sensing images, we discover, through extensive statistical analysis, that there are more low-frequency components than natural images, which may lead to a “discrimination-ambiguity” problem, i.e., the discriminator will become “confused” to tell whether its input is real or not when dealing with those low-frequency regions, and therefore, the quality of generated HR images may be deeply affected. To address this problem, we propose a novel GAN-based super-resolution algorithm named coupled-discriminated GANs (CDGANs) for remote sensing images. Different from the previous GAN-based super-resolution models in which their discriminator takes in a single image at one time, in our model, the discriminator is specifically designed to take in a pair of images: a generated image and its HR ground truth, to make better discrimination of the inputs. We further introduce a dual pathway network architecture, a random gate, and a coupled adversarial loss to learn better correspondence between the discriminative results and the paired inputs. Experimental results on two public data sets demonstrate that our model can obtain more accurate super-resolution results in terms of both visual appearance and local details compared with other state of the arts. Our code will be made publicly available.
Sen Lei, Zhenwei Shi 0001, Zhengxia Zou
IEEE Trans. Geosci. Remote. Sens.1
2019 Simultaneous Super-Resolution and Segmentation for Remote Sensing Images
abstract
In this paper, we present an algorithm to simultaneously obtain high-resolution images and segmentation maps from low-resolution inputs. Super-resolution and segmentation both are challenging task, but they may have certain relationship. Super-resolution will provide images with more details that may help to improve the segmentation accuracy, while label maps in segmentation dataset may contribute to finer edges during super-resolution process. Therefore, we aim to combine these two tasks and explore the influence for each other. For this end, we proposed a new deep neural network to simultaneously address the super-resolution and segmentation tasks for remote sensing images, which is named S2Net. The S2Net is an integrated network composed of a super-resolution sub-network and a segmentation sub-network, which is trained in an end-to-end manner. Experimental results demonstrate that this combination can enhance the performance on these two tasks.
Sen Lei, Bin Pan, Hongxun Hao
IGARSS1
2017 Super-Resolution for Remote Sensing Images via Local-Global Combined Network
abstract
Super-resolution is an image processing technology that recovers a high-resolution image from a single or sequential low-resolution images. Recently deep convolutional neural networks (CNNs) have made a huge breakthrough in many tasks including super-resolution. In this letter, we propose a new single-image super-resolution algorithm named local-global combined networks (LGCNet) for remote sensing images based on the deep CNNs. Our LGCNet is elaborately designed with its “multifork” structure to learn multilevel representations of remote sensing images including both local details and global environmental priors. Experimental results on a public remote sensing data set (UC Merced) demonstrate an overall improvement of both accuracy and visual performance over several state-of-the-art algorithms.
Sen Lei, Zhenwei Shi 0001, Zhengxia Zou
IEEE Geosci. Remote. Sens. Lett.1