Xin Lyu 0001

dblp:183/8301-1 · DBLP profile ↗
← Back
26ranked-venue papers
0as first author
26since 2021 · last 2026
0000-0003-1862-2070ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 10 since 2021Artificial intelligence and machine learning · 9 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 CLIP-driven feature disambiguation and cross-modal synergy for few-shot semantic segmentation
Shangjing Chen, Feng Xu 0008, Xin Lyu 0001, Dafa Wang, Xin Li 0090
Expert Syst. Appl.3
2026 TranCAD: Transforming tabular data into color images for deep semi-supervised anomaly detection
Yucong Huang, Feng Xu 0008, Xin Lyu 0001, Zhennan Xu
Expert Syst. Appl.4
2026 Frequency-spatial decoupled co-modeling transformer for fine-grained remote sensing image segmentation
Xin Li 0090, Shangtuo Qian, Xin Lyu 0001, Yongze Song, Fan Liu 0003, Yiwei Fang, Zhennan Xu, André Kaup
Inf. Sci.3
2026 A global linear attention network for semantic segmentation of remote sensing images
Yiwei Fang, Xin Li 0090, Xin Lyu 0001, Zhennan Xu
Knowl. Based Syst.4
2025 A spectrum-enhanced attention model for semantic segmentation of remote sensing images
abstract
Semantic segmentation of remote sensing images (RSIs) is essential for applications such as environmental monitoring, urban planning, and disaster management. Convolutional Neural Networks (CNNs) and their variants struggle to capture comprehensive spectral context for learning discriminative representations. In this paper, we propose a Spectrum-Enhanced Network (SPENet) that leverages the Frequency Transformer Block (FTB) to capture rich spectral context. FTB integrates Spectrum-Enhanced Attention (SEA) with Multi-Head Frequency Self-Attention (MH-FSA), incorporating more informative contextual cues. Specifically, SEA aggregates spectral statistics through covariance matrix normalization before applying channel-wise attention. By projecting feature maps onto the frequency domain, MH-FSA provides the network with a broader context, extending beyond the low-frequency focus of standard self-attention mechanisms. Extensive experiments on the ISPRS Potsdam and LoveDA datasets show that SPENet significantly outperforms state-of-the-art methods. Besides, the proposed SEA module notably rises average F1-score/overall accuracy/mean insert over union wiht more than 2.5/2.6%/2.3%, as demonstrated by ablation study.
Xin Li 0090, Feng Xu 0008, Feifei Tao, Xin Lyu 0001, Jianyi Zhong, André Kaup
ICASSP5
2025 Adaptive Frequency Threshold Pooling for Mitigating Aliasing in Few-Shot Segmentation
abstract
Feature diversity and interference from complex backgrounds pose substantial challenges to the few-shot segmentation (FSS). Convolutional Neural Networks, as the mainstream backbone for feature extraction, violate the sampling theorem during downsampling, leading to aliasing that causes blurred and distorted class feature boundaries. Such prototype features are likely to induce false responses of query features during the guidance process. To address this problem, we propose an innovative module called Adaptive Frequency Threshold Pooling (AFTP) that can be placed between the encoder stages. AFTP aims to selectively filter out high-frequency components that are prone to aliasing. It dynamically calculates frequency thresholds according to the spectral energy distribution, determining which frequency components are essential for maintaining feature integrity. Furthermore, it incorporates resolution-influenced threshold adjustment, enabling it to adapt to different feature resolutions. Experiments demonstrate that our module effectively enhances the performance of the FSS models on the PASCAL-5iand COCO-20idatasets.
Shangjing Chen, Feng Xu 0008, Xin Lyu 0001, Xin Li 0090
ICME3
2025 Multi-view aggregation and multi-relation alignment for few-shot fine-grained recognition
Feng Xu 0008, Xin Lyu 0001, Xin Li 0090, Shangjing Chen
Expert Syst. Appl.3
2025 Detail retention and enhancement for camouflaged object detection
abstract
The objective of camouflaged object detection is to segment objects that share similar visual attributes with their backgrounds. Existing methods leverage the consistency and variability of the object to pinpoint its location; however, when presented with samples exhibiting intricate edges, the model tends to overlook critical details. Toward this issue, we analyzed the reasons from the perspective of processing basic features and found that roughly compressing the feature channel could lead to the loss of fine features at the initial stage of the pipeline. Therefore, we introduce a multi-scale feature preprocessing approach named Soft Channel Compression prior to feature fusion, which retains key cues by enhancing the visibility of fine-grained features. We then utilize self-attention to fuse the results of multi-scale feature sampling, synthesizing contextual information between extended details and the main object. Furthermore, we propose the Hierarchical Receptive Field Complementarity module, using a divide-and-conquer strategy to enhance discriminative features with different scales in object details. This facilitates the integration of information across various ranges by employing a hierarchical manner to expand the receptive field, thereby achieving complementary information between local areas. Extensive experimental results on benchmark datasets demonstrate that the proposed method outperforms other state-of-the-art methods.
Shangjing Chen, Feng Xu 0008, Xin Li 0090, Xin Lyu 0001
Neurocomputing6
2025 Knowledge-driven prototype refinement for few-shot fine-grained recognition
Feng Xu 0008, Xin Lyu 0001, Xin Li 0090, Shangjing Chen
Knowl. Based Syst.3
2025 Knowledge-guided distribution alignment for cross-domain few-shot learning
Feng Xu 0008, Xin Lyu 0001, Xin Li 0090, Shangjing Chen
Knowl. Based Syst.3
2025 Foster noisy label learning by exploiting noise-induced distortion in foreground localization
Feng Xu 0008, Xin Lyu 0001, Xin Li 0090
Neural Networks3
2025 A Euclidean Affinity-Augmented Hyperbolic Neural Network for Semantic Segmentation of Remote Sensing Images
abstract
Semantic segmentation of remote sensing images (RSIs) plays a pivotal role in advancing geospatial analyses and applications across diverse fields, such as urban planning and environmental monitoring. Traditional learning paradigms predominantly utilize Euclidean spaces for feature extraction. This approach can introduce spatial distortions when representing objects, as Euclidean architectures typically focus on locality and are optimized for grid data, not always yielding optimal geometrical representations for data structured in non-Euclidean spaces. To address these problems, we propose EAAHNet, the first fully hyperbolic neural network designed for semantic segmentation of RSIs. EAAHNet employs the Lorentz model to reformalize conventional Euclidean-based neural network operations, ensuring the preservation of hyperbolic properties. Furthermore, to account for the inherently Euclidean nature of ground objects, we propose a Euclidean affinity-augmented hyperbolic attention module (EAAHAM) that enriches contextual dependencies through an attention fusion manner. This enhancement significantly improves the network’s capacity to discern pixel-wise semantics. Extensive experiments conducted on the ISPRS Vaihingen, ISPRS Potsdam, and LoveDA datasets demonstrate EAAHNet’s superior performance over several state-of-the-art methods. Additionally, the ablation study verifies the impacts of EAAHAM.
Xin Li 0090, Feng Xu 0008, Fan Liu 0003, Xin Lyu 0001, Hongmin Gao 0001, Jun Zhou 0011, André Kaup
IEEE Trans. Geosci. Remote. Sens.4
2025 A Frequency Decoupling Network for Semantic Segmentation of Remote Sensing Images
abstract
Semantic segmentation of remote sensing images (RSIs) is vital for numerous geospatial applications, including land-use mapping, urban planning, and environmental monitoring. Traditional neural networks for semantic segmentation primarily focus on learning in the spatial domain, which often results in suboptimal performance due to the complexity of RSIs that exhibit diverse and intricate structures. To address this problem, we propose a novel frequency decoupling network (FDNet) that enhances feature representation by independently refining high-frequency and low-frequency components in the frequency domain. FDNet introduces three core components: a sparse-aware spectral enhancement module (SSEM) that optimizes spectral feature learning by compressing redundant information while highlighting informative spectral bands, a frequency decoupling attention module (FDAM) that precisely distinguishes and enhances high-frequency and low-frequency features and an attentive frequency context module (AFCM) that integrates SSEM and FDAM into a cohesive framework for enriched spectral context modeling. Extensive experiments conducted on four benchmark datasets demonstrate that FDNet outperforms several state-of-the-art methods, achieving superior segmentation accuracy and robustness across various terrains and imaging conditions. Ablation experiments further confirm the impacts of SSEM, FDAM, and AFCM.
Xin Li 0090, Feng Xu 0008, Anzhu Yu, Xin Lyu 0001, Hongmin Gao 0001, Jun Zhou 0001
IEEE Trans. Geosci. Remote. Sens.4
2025 IGFNet: An Interactive-Guided Fusion Network for Hyperspectral Pansharpening
abstract
Hyperspectral pansharpening is an efficient approach to obtaining high-resolution hyperspectral images (HR-HSIs) by fusing low-resolution hyperspectral images (LR-HSIs) with high-resolution panchromatic images (HR-PANs). However, the spatial and spectral distortions in reconstructed HR-HSIs are almost inevitable due to the modal gap between LR-HSIs and HR-PANs. Therefore, the performance of multi-source features fusion largely hinges on the ability to extract and align heterogeneous features across modalities. Most of the existing methods focus on integrating decoupled spatial and spectral information from different sources directly, which poses a dual challenge in aligning both spatial and spectral features effectively. To address the issues above-mentioned, a novel method named Interactive-Guided Fusion Network (IGFNet) is proposed, which is built upon a multi-stage progressive fusion framework. A high-resolution branch (HR) is introduced to interactively guide the alignment between cross-modal features, by fusing up-sampled HSI and PAN as a joint spatial-spectral guidance signal. Furthermore, the alignment is progressively conducted across stages, narrowing the modality gap and enhancing the representation of HR spatial-spectral feature. Additionally, we designed parameter-free spatial, spectral, and spatial-spectral attention mechanisms to extract global and local features effectively. Extensive experiments on reduced-resolution and full-resolution datasets demonstrate that IGFNet outperforms state-of-the-art across various metrics. Specifically, with a scaling factor of 4 on the Pavia University dataset, our method achieves a 2.09% relative improvement in PSNR, a 0.99% relative increase in SSIM, a 2.6% relative reduction in SAM, while reducing the parameter count by compared to the baseline MDA-Net.
Zhennan Xu, Xin Lyu 0001, Feng Xu 0008, Xin Li 0090, Yucong Huang, Caifeng Wu, Yiwei Fang
IEEE Trans. Geosci. Remote. Sens.2
2024 Locality-Enhanced Transformer for Semantic Segmentation of High-Resolution Remote Sensing Images
abstract
Transformers have emerged as a transformative tool in various computer vision tasks, excelling at capturing long-range dependencies. Their potential applicability and scalability in the interpretation of high-resolution remote sensing images (HRRSIs) have thus garnered substantial interest. However, unlike natural images, HRRSIs present intricate scenes characterized by scale variations and diverse appearances. These challenges underscore the importance of enabling networks to effectively assimilate both local intricacies and global context. In this letter, we introduce LETFormer, a semantic segmentation transformer. LETFormer balances capturing longrange dependencies with preserving local details through its unique LETFormer block, featuring an anchor token. This token aggregates localized contextual information within a designated window and promotes meaningful interactions among anchor tokens. With a mask transformer decoder, LETFormer gains ample contextual cues for precise semantic mask prediction. Empirical findings based on evaluations using the ISPRS Potsdam and LoveDA benchmarks unequivocally establish LETFormer’s superiority over state-of-the-art models. Additionally, we analyze the parameter size and floating-point operations per second (FLOPs) of LETFormer.
Xin Li 0090, Feng Xu 0008, Runliang Xia, Nan Xu 0008, Fan Liu 0003, Chi Yuan, Qian Huang 0008, Xin Lyu 0001
ICASSP8
2024 FreqFormer: A Frequency Transformer for Semantic Segmentation of Remote Sensing Images
abstract
Semantic segmentation of remote sensing images (RSIs) is vital for geospatial intelligence. However, traditional methods face challenges with mixed pixels and complex land cover types. Convolutional neural networks and transformers have led the field of RSI semantic segmentation by learning visual features in the spatial domain, but they often overlook the rich spectral features which can be well-described in the frequency domain, resulting in inadequate context modeling. In this paper, we present FreqFormer, a frequency transformer that enhances semantic segmentation by incorporating both spectral and spatial information through a devised frequency attention (FA) module. FA refines representations in the frequency domain through two parallel branches. Specifically, the high-frequency branch (HFB) utilizes a convolution layer with a Canny kernel to preserve high-frequency details, followed by multi-head self-attention to model high-frequency context. Followed by an element summation, high-frequency and low-frequency contexts are aggregated. Then, the formed FreqFormer block is sequentially deployed in the encoder stage with patch merging for spatial contraction. As for the decoder, the mask transformer decoder applies a scalar product to predict patch-wise semantics before upsampling. In experiments, FreqFormer outperforms state-of-the-art models on the ISPRS Potsdam and LoveDA datasets, demonstrating significant improvements in numerical evaluations. The integration of HFB significantly boosts the model’s ability to capture fine details, highlighting its potential for geospatial analysis.
Xin Li 0090, Feng Xu 0008, Fan Liu 0003, Yiwei Fang, Xin Lyu 0001, Jun Zhou 0001
MMAsia6
2024 SigCo: Eliminate the inter-class competition via sigmoid for learning with noisy labels
Feng Xu 0008, Xin Lyu 0001, Xin Li 0090
Knowl. Based Syst.4
2024 AAFormer: Attention-Attended Transformer for Semantic Segmentation of Remote Sensing Images
abstract
The rapid advancements in remote sensing technology have enabled the widespread availability of fine-resolution remote sensing images (RSIs), offering rich spatial details and semantics. Despite the applicability and scalability of transformers in semantic segmentation of RSIs by learning pairwise contextual affinity, they inevitably introduce irrelevant context, hindering accurate inference of patch semantics. To address this, we propose a novel multi-head attention-attended module (AAM) that refines the multi-head self-attention mechanism. The AAM filters out irrelevant context while highlighting informative ones by considering the relevance between self-attention maps and the query vector. The AAM generates an attention gate to complement contextual affinity and emphasize the useful ones with a higher weight simultaneously. Leveraging multi-head AAM as the core unit, we construct a lightweight attention-attended transformer block (ATB). Subsequently, we devise AAFormer, a pure transformer with a mask transformer decoder, for achieving semantic segmentation of RSIs. We extensively evaluate our approach on the ISPRS Potsdam and LoveDA datasets, demonstrating compelling performance compared to mainstream methods. Additionally, we conduct evaluations to analyze the effects of AAM.
Xin Li 0090, Feng Xu 0008, Linyang Li, Nan Xu 0008, Fan Liu 0003, Chi Yuan, Xin Lyu 0001
IEEE Geosci. Remote. Sens. Lett.8
2024 A Cross-Domain Coupling Network for Semantic Segmentation of Remote Sensing Images
abstract
Semantic segmentation of remote sensing images (RSIs) is critical for various applications, including urban planning, agriculture, and disaster management. Existing methods often fail to capture fine-grained textures and periodic patterns in RSIs, leading to suboptimal results in complex terrains. To address these challenges, we propose a cross-domain coupling network (CDCNet) that leverages both domain-specific extraction and cross-domain coupling (CDC) to enrich contextual cues for semantic inference. Our CDCNet integrates a CDC layer within the encoder-decoder architecture to simultaneously refine representations in the frequency and spatial domains. This approach effectively models fine-grained textures and periodic patterns in the frequency domain, as well as edges, shapes, and broad structural elements in the spatial domain. Extensive experiments on the ISPRS Potsdam and LoveDA datasets demonstrate the superiority of CDCNet over several state-of-the-art methods. Ablation studies confirm the significant impact of the CDC layer, validating the effectiveness of our approach in handling RSIs.
Xin Li 0090, Feng Xu 0008, Feifei Tao, Hongmin Gao 0001, Fan Liu 0003, Xin Lyu 0001
IEEE Geosci. Remote. Sens. Lett.8
2024 A Frequency Domain Feature-Guided Network for Semantic Segmentation of Remote Sensing Images
abstract
Semantic segmentation of Remote Sensing Images (RSIs) entails assigning semantic labels to each pixel accurately. RSIs are rich in spatial and spectral data, revealing diverse material and object characteristics. Yet, current RSI-focused computer vision models struggle with significant intra-class variation and inter-class resemblance due to limited spectral data usage. We propose the Frequency Domain Feature-Guided Network (FFGNet) for RSI semantic segmentation, influenced by digital signal processing theories. FFGNet initially generates frequency domain features via patch partitioning and 2D discrete cosine transformation. Our Frequency Enhancement Attention module (FEA) then distinguishes and intensifies frequency components to retain detailed information. These enhanced features are integrated with the Spatial-Spectral Attention (SSA) for enriched spectral signals. In the inference phase, these features are upsampled and combined with decoded features, emphasizing spectral details. Additionally, our novel loss function combines frequency and cross-entropy losses. Experiments on LoveDA and ISPRS Potsdam datasets demonstrate FFGNet's effectiveness, surpassing other mainstream models. An ablation study further validates our dual-guidance design.
Xin Li 0090, Feng Xu 0008, Hongmin Gao 0001, Fan Liu 0003, Xin Lyu 0001
IEEE Signal Process. Lett.5
2024 Semantic Segmentation of Remote Sensing Images by Interactive Representation Refinement and Geometric Prior-Guided Inference
abstract
High spatial resolution remote sensing images (HRRSIs) contain intricate details and varied spectral distributions, making their semantic segmentation a challenging task. To address this problem, it is crucial to adequately capture both local and global contexts to reduce semantic ambiguity. While self-attention modules in vision transformers capture long-range context, they tend to sacrifice local details. In this article, we propose a geometric prior-guided interactive network (GPINet), a hybrid network that refines features across encoder and decoder stages. First of all, a dual branch structure encoder with local-global interaction modules (LGIMs) is designed to fully exploit local and global contexts for feature refinement. Unlike commonly used skip connections or concatenations, the LGIMs bilaterally couple and exchange CNN features with transformer features by lossless transformation and elaborating cross-attention. Moreover, we introduce a geometric prior generation module (GPGM) that iteratively updates the randomly initialized geometric prior. Subsequently, the geometric priors are stored and used to guide feature recovery. Finally, a weighted summation is applied to the upsampled decoded features and geometric priors. By comprehensively capturing contexts and enabling lossless decoding and deterministic inference, GPINet allows the network to learn discriminative representations for accurately specifying pixel-level semantics. Experiments on three benchmark datasets demonstrate the superiority of the proposed GPINet over state-of-the-art methods. Furthermore, we validate the effectiveness of geometric priors and compare the model sizes.
Xin Li 0090, Feng Xu 0008, Fan Liu 0003, Xin Lyu 0001, Jun Zhou 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 A Synergistical Attention Model for Semantic Segmentation of Remote Sensing Images
abstract
In remotely sensed images, high intraclass variance and interclass similarity are ubiquitous due to complex scenes and objects with multivariate features, making semantic segmentation a challenging task. Deep convolutional neural networks can solve this problem by modeling the context of features and improving their discriminability. However, current learning paradigms model the feature affinity in spatial dimension and channel dimension separately and then fuse them in a sequential or parallel manner, leading to suboptimal performance. In this study, we first analyze this problem practically and summarize it as attention bias that reduces the capability of network in distinguishing weak and discretely distributed objects from wide-range objects with internal connectivity, when modeled only in spatial or channel domain. To jointly model both spatial and channel affinity, we design a synergistic attention module (SAM), which allows for channelwise affinity extraction while preserving spatial details. In addition, we propose a synergistic attention perception neural network (SAPNet) for the semantic segmentation of remote sensing images. The hierarchical-embedded synergistic attention perception module aggregates SAM-refined features and decoded features. As a result, SAPNet enriches inference clues with desired spatial and channel details. Experiments on three benchmark datasets show that SAPNet is competitive in accuracy and adaptability compared with state-of-the-art methods. The experiments also validate the hypothesis of attention bias and the efficiency of SAM.
Xin Li 0090, Feng Xu 0008, Fan Liu 0003, Xin Lyu 0001, Zhennan Xu, Jun Zhou 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 Feature difference for single-shot object detection
abstract
Abstract The one‐stage detectors achieve a good trade‐off between performance and latency, owing to the plain architecture and divergent learning mechanism for classification and localization. However, the two sub‐tasks require features with various inherency with which to generate inconsistent detections, fettering detectors. In this study, the misalignment is deeply analyzed via kernel density estimation (KDE) for the first time. Moreover, to address the misalignment, a plug‐and‐play detection head, named Diff‐Head, is devised and embedded in one‐stage detectors. Concretely, the authors merge parallel branches into a semi‐parallel structure, establishing the correlation between classification and regression. In the regression branch, a feature difference module (FDM) gets rid of the features that favour classification by subtracting salient object features from the original feature map, and position encoding (PE) modules enhance the absolute position information. The flexibility and efficiency of the detection head are retained. Experiments on Pascal visual object classes (VOC) and MS COCO demonstrate that Diff‐Head is effective and achieves competitive performance with state‐of‐the‐art detectors. Meanwhile, the amount of parameters is reduced at least 30% and 83.0% average precision (AP) is achieved on Pascal VOC. The analyses of consistency and error show that Diff‐Head has better localization and the capability of mitigating the misalignment.
Feng Xu 0008, Xin Lyu 0001, Xin Li 0090, Xinyuan Wang 0002, Caifeng Wu
IET Image Process.3
2022 Hybridizing Euclidean and Hyperbolic Similarities for Attentively Refining Representations in Semantic Segmentation of Remote Sensing Images
abstract
Attention mechanisms have revolutionized the semantic segmentation network in interpreting remotely sensed images (RSIs) due to their amazing ability in establishing contextual dependencies. Nevertheless, due to the complex scenes and diverse objects in RSIs, a variety of details and correlations are not available in Euclidean space. Therefore, a similarity-hybrid attention module (SHAM) is devised to attentively learn the hyperbolic and Euclidean attention maps between any two positions, followed by a weighted element-wise summation. The hybrid attention maps posses latent geometric properties of both Euclidean and hyperboloid. Taking commonly-used fully convolutional network (FCN) as baseline, HAENet that embeds SHAM, is presented. Experiments on ISPRS Potsdam and DeepGlobe benchmarks reveal its superiority to comparative methods. In addition, the ablation study validates the effectiveness of SHAM compared to other attention modules.
Xin Li 0090, Feng Xu 0008, Fan Liu 0003, Runliang Xia, Linyang Li, Zhennan Xu, Xin Lyu 0001
IEEE Geosci. Remote. Sens. Lett.8
2022 A Trajectory Released Scheme for the Internet of Vehicles Based on Differential Privacy
abstract
The locations and users’ information can be shared and interacted in the IoV (Internet of Vehicles), which provides sufficient data for traffic deployment and behavior pattern analysis. However, privacy issues had become more severe since personal or sensitive information is inclined to be revealed in a big data environment. In this work, a novel differential privacy-based algorithm named DPTD (Differentially Private Trajectory Database) is proposed for trajectory database releasing. Firstly, a 3-dimensional generalized trajectory dataset is established by considering the time factor. Then, the trajectory space is divided into several planes through the timestamps, and the set of the locations on each plane is further processed by clustering and generalizing to re-form new trajectories, that is, the trajectories to be released. This method is quite favorable to prefix-tree releasing because the spatiotemporal characteristics of the trajectories can be captured and spareness problem is fixed. Besides, a Markov assumption-based prediction method is suggested in order to reduce the cost of adding noise. Unlike the traditional method that the noise is added layer by layer, the noise is only added to the odd layers based on the prediction through spatio-temporal correlation, saving approximately 50% of the privacy budget. Theoretical analysis and experimental results show that the proposed algorithm has better data availability than the compared algorithms while guaranteeing the expected privacy level.
Sujin Cai, Xin Lyu 0001, Xin Li 0090, Duohan Ban
IEEE Trans. Intell. Transp. Syst.2
2021 Semantic Labeling of Very High-Resolution Imagery by Leveraging Contextual Information with Optimized Non-Local Neural Network
abstract
Semantic labeling of very high-resolution (VHR) aerial imagery attempts to assign a specific label to every pixel. Resorting to the significant progress of semantic segmentation for natural images, the non-local neural networks express superior capability in capturing multi-scale contextual information, which is helpful for accurately labeling raw images. However, the computational complexity of non-local neural networks is criticized. In this paper, the optimized non-local block (ONN) is devised and embedded in the encoder stage to reduce the computational complexity in capturing contextual information. Besides, the DUpsampling is introduced to further boost the efficiency of decoder. Finally, the extensive experiments, including numerical metrics and visual inspection comparisons, are conducted on ISPRS benchmarks. The results indicate that the proposed method performs better than several SOTA methods both in accuracy and efficiency.
Xin Li 0090, Feng Xu 0008, Xin Lyu 0001, Liancheng Zhao, Xinyuan Wang 0002
IGARSS3