EDBT 2026 Demo / reviewers in the wild / expert
Xin Li 0090
dblp:09/1365-90
· DBLP profile ↗
38ranked-venue papers
13as first author
38since 2021 · last 2026
0000-0003-0576-3181ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 8 first-author · 12 since 2021Artificial intelligence and machine learning · 10 · 10 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CLIP-driven feature disambiguation and cross-modal synergy for few-shot semantic segmentation
Shangjing Chen, Feng Xu 0008, Xin Lyu 0001, Dafa Wang, Xin Li 0090 |
Expert Syst. Appl. | 5 |
| 2026 | Frequency-spatial decoupled co-modeling transformer for fine-grained remote sensing image segmentation
Xin Li 0090, Shangtuo Qian, Xin Lyu 0001, Yongze Song, Fan Liu 0003, Yiwei Fang, Zhennan Xu, André Kaup |
Inf. Sci. | 1 |
| 2026 | A global linear attention network for semantic segmentation of remote sensing images
Yiwei Fang, Xin Li 0090, Xin Lyu 0001, Zhennan Xu |
Knowl. Based Syst. | 3 |
| 2026 | Spatial-Temporal Self-Compensating Graph Convolutional Network for Skeleton-Based Action Recognition Under Data ConstraintsabstractSkeleton-based human action recognition has emerged as a prominent research focus in computer vision, with significant progress achieved in recent years. However, existing methods often suffer substantial performance degradation under real-world data constraints, such as body occlusion, missing frames, and noise. These limitations critically undermine the robustness of related techniques in practical applications. To address these challenges, we propose a Spatial Temporal Self-compensating Graph Convolutional Network (STSc-GCN), which skillfully utilizes the systematic and regular nature of human movement to mitigate performance degradation caused by data constraints through a data self-compensation mechanism. Specifically, STSc-GCN comprises two key modules: 1) collaborative motion spatial compensation (CMSC). This module designs multiple distinct topological relationships, primarily including Walk-probability Generality Topology and Self-organizing Particularity Topology, respectively, to deeply explore the universal and personalized collaborative relationships between human joints. These relationships help compensate for the lack of information caused by spatial data constraints and 2) meta-action sharpening temporal Compensation (MSTC). This module introduces a novel motion sharpening mechanism that enhances key dynamic information within the meta-action sequences through cross-attention technology, thereby improving model adaptability to missing-frame scenarios. STSc-GCN achieves state-of-the-art performance on four constrained datasets and shows superior results on three widely used standard datasets, confirming its effectiveness in both constrained and general scenarios. Code will be available at https://github.com/XingLi1012/STSc-GCN.git. Xing Li 0005, Qian Huang 0008, Xin Li 0090, Jinhui Tang 0001, Qiaolin Ye |
IEEE Trans. Image Process. | 4 |
| 2026 | A Dual Domain Collaborative Network for Polyp SegmentationabstractAccurate polyp segmentation in colonoscopy images is essential for early colorectal cancer detection but remains a challenging problem due to the limitations in existing methods for optimizing boundary features and aligning cross-level representations. Specifically, the indistinct polyp boundaries and scale variations across different feature levels pose significant challenges for segmentation accuracy. To address these issues, we propose a dual domain collaborative network (DDCNet) that introduces two novel modules: a frequency context enhancement module (FCEM), which operates in the frequency domain to refine high- and low-frequency features, and a cross-level shift-recalibrated fusion module (CSFM), which improves multi-scale feature alignment in the spatial domain. The FCEM improves boundary precision by adaptively refining high-frequency boundary features and enhancing low-frequency contextual information, while the CSFM mitigates cross-level feature misalignment by dynamically recalibrating multi-scale features throughout the encoder-decoder architecture. Additionally, we design a hybrid loss function that integrates boundary, cross-entropy, and frequency consistency losses to further boost segmentation performance. Experimental results on three benchmark datasets (Kvasir-SEG, CVC-ClinicDB, and CVC-ColonDB) demonstrate that DDCNet achieves state-of-the-art performance, with Dice coefficients of 0.9343, 0.9447, and 0.8155, respectively. These results represent improvements of 1.0%-1.5% over the best existing methods. Ablation studies further validate the individual contributions of FCEM, CSFM, and the hybrid loss function. Additionally, we compared the proposed loss function with three commonly used functions. Zuojian Zhou, Kongfa Hu, Tao Yang 0048, André Kaup, Xin Li 0090 |
IEEE J. Biomed. Health Informatics | 6 |
| 2026 | MD-PCSN: Meta-Motion Decoupling Point Cloud Sequence Network for Privacy-Preserving Human Action Recognition in AI MachinesabstractIn next-generation communication networks and Industry 5.0 based applications, ensuring robust security and reliability in human-computer interaction (HCI) constitutes a fundamental prerequisite for safety-critical AI machine systems. Point cloud sequence-based human action recognition demonstrates intrinsic advantages in privacy-preserving HCI, leveraging its non-intrusive sensing modality to mitigate data vulnerability while maintaining high-precision action interpretation in industrial environments. Existing spatio-temporal encoding methods for point cloud sequence-based action recognition suffer from two fundamental limitations: (1) rigid neighborhood constraints impair multi-scale feature extraction for heterogeneous body parts, and (2) independent spatial-temporal decomposition introduces motion representation distortion. We propose a Meta-motion Decoupling Point Cloud Sequence Network (MD-PCSN) that addresses these challenges through: (1) logarithmic spatio-temporal point convolution for hierarchical meta-motion construction at variable granularities, and (2) a novel Gated-KANsformer architecture with differential motion encoding to explicitly model both short-term displacements and long-term spatio-temporal dependencies. The proposed meta-motion decoupling mechanism significantly enhances robustness against sensor perturbations, making the framework particularly suitable for security-critical applications. Extensive experiments on three benchmark datasets demonstrate MD-PCSN’s superior performance. It outperforms classic PST-Transformer by 1.5% on MSR Action3D and 4.14% on UTD-MHAD. Under the NTU RGB+D 60, it achieves 2.9% cross-view gain over the latest PointActionCLIP. Xing Li 0005, Xin Li 0090, Qian Huang 0008 |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2025 | A spectrum-enhanced attention model for semantic segmentation of remote sensing imagesabstractSemantic segmentation of remote sensing images (RSIs) is essential for applications such as environmental monitoring, urban planning, and disaster management. Convolutional Neural Networks (CNNs) and their variants struggle to capture comprehensive spectral context for learning discriminative representations. In this paper, we propose a Spectrum-Enhanced Network (SPENet) that leverages the Frequency Transformer Block (FTB) to capture rich spectral context. FTB integrates Spectrum-Enhanced Attention (SEA) with Multi-Head Frequency Self-Attention (MH-FSA), incorporating more informative contextual cues. Specifically, SEA aggregates spectral statistics through covariance matrix normalization before applying channel-wise attention. By projecting feature maps onto the frequency domain, MH-FSA provides the network with a broader context, extending beyond the low-frequency focus of standard self-attention mechanisms. Extensive experiments on the ISPRS Potsdam and LoveDA datasets show that SPENet significantly outperforms state-of-the-art methods. Besides, the proposed SEA module notably rises average F1-score/overall accuracy/mean insert over union wiht more than 2.5/2.6%/2.3%, as demonstrated by ablation study. Xin Li 0090, Feng Xu 0008, Feifei Tao, Xin Lyu 0001, Jianyi Zhong, André Kaup |
ICASSP | 1 |
| 2025 | Learned Video Compression With Refined Adaptive Flow Pyramid And Coordinate-Aware AttentionabstractIn video compression, motion estimation and motion compensation are critical for achieving efficient encoding. Although the commonly used SpyNet and bilinear interpolation have contributed in improving the compression efficiency, they still have limitations. SpyNet often loses details and fails to fully utilize the feature extraction capabilities of deep networks. Furthermore, bilinear interpolation inherently attenuates high-frequency information, leading to frame blurring and distortion. In this paper, we propose a novel video compression algorithm. To overcome the limitations of SpyNet, we propose a refined adaptive flow pyramid network. This network uses a multi-scale feature pyramid to capture more details. Moreover, we use an iterative cost volume refinement engine that improves the feature representation of the network and iteratively improve the accuracy of motion estimation. In addition, to overcome the limitation of bilinear interpolation, we propose a coordinate-aware attention module, which captures more high-frequency information to improve the accuracy of motion compensation. Experimental results show that our method outperforms VTM-13.2 (LDP) in terms of PSNR. Qian Huang 0008, Xin Li 0090, Yiming Wang 0008 |
ICASSP | 3 |
| 2025 | RemoteTrimmer: Adaptive Structural Pruning for Remote Sensing Image ClassificationabstractSince high resolution remote sensing image classifi-cation often requires a relatively high computation complexity, lightweight models tend to be practical and efficient. Model pruning is an effective method for model compression. However, existing methods rarely take into account the specificity of remote sensing images, resulting in significant accuracy loss after pruning. To this end, we propose an effective structural pruning approach for remote sensing image classification. Specifically, a pruning strategy that amplifies the differences in channel importance of the model is introduced. Then an adaptive mining loss function is designed for the fine-tuning process of the pruned model. Finally, we conducted experiments on two remote sensing classification datasets. The experimental results demonstrate that our method achieves minimal accuracy loss after compressing remote sensing classification models, achieving state-of-the-art (SoTA) performance. Guangwenjie Zou, Liang Yao 0001, Fan Liu 0003, Chuanyi Zhang, Xin Li 0090, Shengxiang Xu, Jun Zhou 0001 |
ICASSP | 5 |
| 2025 | Adaptive Frequency Threshold Pooling for Mitigating Aliasing in Few-Shot SegmentationabstractFeature diversity and interference from complex backgrounds pose substantial challenges to the few-shot segmentation (FSS). Convolutional Neural Networks, as the mainstream backbone for feature extraction, violate the sampling theorem during downsampling, leading to aliasing that causes blurred and distorted class feature boundaries. Such prototype features are likely to induce false responses of query features during the guidance process. To address this problem, we propose an innovative module called Adaptive Frequency Threshold Pooling (AFTP) that can be placed between the encoder stages. AFTP aims to selectively filter out high-frequency components that are prone to aliasing. It dynamically calculates frequency thresholds according to the spectral energy distribution, determining which frequency components are essential for maintaining feature integrity. Furthermore, it incorporates resolution-influenced threshold adjustment, enabling it to adapt to different feature resolutions. Experiments demonstrate that our module effectively enhances the performance of the FSS models on the PASCAL-5iand COCO-20idatasets. Shangjing Chen, Feng Xu 0008, Xin Lyu 0001, Xin Li 0090 |
ICME | 4 |
| 2025 | Geo-CF2Net: Geometry-Prior Cross-Frequency Interactive Fusion Network for 3D Human Action RecognitionabstractDynamic point cloud-based human action recognition has garnered increasing attention due to its inherent advantages in privacy preservation and structural completeness. Current methods typically rely on nested point spatio-temporal convolutions to understand motion semantics in a bottom-up manner, which is intractable for capturing high-fidelity human dynamics disentangled from spatio-temporal interference. Motivated by this, designing a practical spatio-temporal factorization backbone is essential. However, the repeated coarsening of aggregated features along the spatial dimension often leads to the degradation of intrinsic geometric texture relations within point cloud data. Moreover, discretizing continuous visual data into isolated temporal hyperpoints significantly diminishes temporal continuity, resulting in the fragmentation of human action. To circumvent above limitations, we propose a novel Geometry-Prior Cross-Frequency Interactive Fusion Network (Geo-CF2Net). Specifically, we investigate a Spatial-Geometry Pose Prior (SGPP) module, which compensates for pose information loss during spatial downsampling by explicitly modeling geometric constraints among neighboring points. In addition, we elaborate on a Temporal Motion Unit Interactive Coordination (TMIC) module to track the interactive composite semantics of low-frequency steady-state venations and high-frequency transient-state details within a high-dimensional pose evolution flow. Extensive experiments on three public benchmarks substantiate the superiority of Geo-CF2Net over state-of-the-art methods. Qian Huang 0008, Xing Li 0005, Shihao Han, Yirui Wu, Xin Li 0090, Ziyang Yin |
ACM Multimedia | 8 |
| 2025 | Multi-view aggregation and multi-relation alignment for few-shot fine-grained recognition
Feng Xu 0008, Xin Lyu 0001, Xin Li 0090, Shangjing Chen |
Expert Syst. Appl. | 5 |
| 2025 | STFE-VC: Spatio-temporal feature enhancement for learned video compression
Yiming Wang 0008, Qian Huang 0008, Bin Tang 0002, Xin Li 0090, Xing Li 0005 |
Expert Syst. Appl. | 4 |
| 2025 | Detail retention and enhancement for camouflaged object detectionabstractThe objective of camouflaged object detection is to segment objects that share similar visual attributes with their backgrounds. Existing methods leverage the consistency and variability of the object to pinpoint its location; however, when presented with samples exhibiting intricate edges, the model tends to overlook critical details. Toward this issue, we analyzed the reasons from the perspective of processing basic features and found that roughly compressing the feature channel could lead to the loss of fine features at the initial stage of the pipeline. Therefore, we introduce a multi-scale feature preprocessing approach named Soft Channel Compression prior to feature fusion, which retains key cues by enhancing the visibility of fine-grained features. We then utilize self-attention to fuse the results of multi-scale feature sampling, synthesizing contextual information between extended details and the main object. Furthermore, we propose the Hierarchical Receptive Field Complementarity module, using a divide-and-conquer strategy to enhance discriminative features with different scales in object details. This facilitates the integration of information across various ranges by employing a hierarchical manner to expand the receptive field, thereby achieving complementary information between local areas. Extensive experimental results on benchmark datasets demonstrate that the proposed method outperforms other state-of-the-art methods. Shangjing Chen, Feng Xu 0008, Xin Li 0090, Xin Lyu 0001 |
Neurocomputing | 4 |
| 2025 | MDA-HTD: Mask-driven dual autoencoders meet hyperspectral target detection
Zhonghao Chen, Hongmin Gao 0001, Zhengtao Lu, Yao Ding 0010, Xin Li 0090, Bing Zhang 0001 |
Inf. Process. Manag. | 6 |
| 2025 | Knowledge-driven prototype refinement for few-shot fine-grained recognition
Feng Xu 0008, Xin Lyu 0001, Xin Li 0090, Shangjing Chen |
Knowl. Based Syst. | 5 |
| 2025 | Knowledge-guided distribution alignment for cross-domain few-shot learning
Feng Xu 0008, Xin Lyu 0001, Xin Li 0090, Shangjing Chen |
Knowl. Based Syst. | 5 |
| 2025 | Multiscale motion-aware and spatial-temporal-channel contextual coding network for learned video compression
Yiming Wang 0008, Qian Huang 0008, Bin Tang 0002, Xin Li 0090, Xing Li 0005 |
Knowl. Based Syst. | 4 |
| 2025 | Foster noisy label learning by exploiting noise-induced distortion in foreground localization
Feng Xu 0008, Xin Lyu 0001, Xin Li 0090 |
Neural Networks | 5 |
| 2025 | A Euclidean Affinity-Augmented Hyperbolic Neural Network for Semantic Segmentation of Remote Sensing ImagesabstractSemantic segmentation of remote sensing images (RSIs) plays a pivotal role in advancing geospatial analyses and applications across diverse fields, such as urban planning and environmental monitoring. Traditional learning paradigms predominantly utilize Euclidean spaces for feature extraction. This approach can introduce spatial distortions when representing objects, as Euclidean architectures typically focus on locality and are optimized for grid data, not always yielding optimal geometrical representations for data structured in non-Euclidean spaces. To address these problems, we propose EAAHNet, the first fully hyperbolic neural network designed for semantic segmentation of RSIs. EAAHNet employs the Lorentz model to reformalize conventional Euclidean-based neural network operations, ensuring the preservation of hyperbolic properties. Furthermore, to account for the inherently Euclidean nature of ground objects, we propose a Euclidean affinity-augmented hyperbolic attention module (EAAHAM) that enriches contextual dependencies through an attention fusion manner. This enhancement significantly improves the network’s capacity to discern pixel-wise semantics. Extensive experiments conducted on the ISPRS Vaihingen, ISPRS Potsdam, and LoveDA datasets demonstrate EAAHNet’s superior performance over several state-of-the-art methods. Additionally, the ablation study verifies the impacts of EAAHAM. Xin Li 0090, Feng Xu 0008, Fan Liu 0003, Xin Lyu 0001, Hongmin Gao 0001, Jun Zhou 0011, André Kaup |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | A Frequency Decoupling Network for Semantic Segmentation of Remote Sensing ImagesabstractSemantic segmentation of remote sensing images (RSIs) is vital for numerous geospatial applications, including land-use mapping, urban planning, and environmental monitoring. Traditional neural networks for semantic segmentation primarily focus on learning in the spatial domain, which often results in suboptimal performance due to the complexity of RSIs that exhibit diverse and intricate structures. To address this problem, we propose a novel frequency decoupling network (FDNet) that enhances feature representation by independently refining high-frequency and low-frequency components in the frequency domain. FDNet introduces three core components: a sparse-aware spectral enhancement module (SSEM) that optimizes spectral feature learning by compressing redundant information while highlighting informative spectral bands, a frequency decoupling attention module (FDAM) that precisely distinguishes and enhances high-frequency and low-frequency features and an attentive frequency context module (AFCM) that integrates SSEM and FDAM into a cohesive framework for enriched spectral context modeling. Extensive experiments conducted on four benchmark datasets demonstrate that FDNet outperforms several state-of-the-art methods, achieving superior segmentation accuracy and robustness across various terrains and imaging conditions. Ablation experiments further confirm the impacts of SSEM, FDAM, and AFCM. Xin Li 0090, Feng Xu 0008, Anzhu Yu, Xin Lyu 0001, Hongmin Gao 0001, Jun Zhou 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | IGFNet: An Interactive-Guided Fusion Network for Hyperspectral PansharpeningabstractHyperspectral pansharpening is an efficient approach to obtaining high-resolution hyperspectral images (HR-HSIs) by fusing low-resolution hyperspectral images (LR-HSIs) with high-resolution panchromatic images (HR-PANs). However, the spatial and spectral distortions in reconstructed HR-HSIs are almost inevitable due to the modal gap between LR-HSIs and HR-PANs. Therefore, the performance of multi-source features fusion largely hinges on the ability to extract and align heterogeneous features across modalities. Most of the existing methods focus on integrating decoupled spatial and spectral information from different sources directly, which poses a dual challenge in aligning both spatial and spectral features effectively. To address the issues above-mentioned, a novel method named Interactive-Guided Fusion Network (IGFNet) is proposed, which is built upon a multi-stage progressive fusion framework. A high-resolution branch (HR) is introduced to interactively guide the alignment between cross-modal features, by fusing up-sampled HSI and PAN as a joint spatial-spectral guidance signal. Furthermore, the alignment is progressively conducted across stages, narrowing the modality gap and enhancing the representation of HR spatial-spectral feature. Additionally, we designed parameter-free spatial, spectral, and spatial-spectral attention mechanisms to extract global and local features effectively. Extensive experiments on reduced-resolution and full-resolution datasets demonstrate that IGFNet outperforms state-of-the-art across various metrics. Specifically, with a scaling factor of 4 on the Pavia University dataset, our method achieves a 2.09% relative improvement in PSNR, a 0.99% relative increase in SSIM, a 2.6% relative reduction in SAM, while reducing the parameter count by compared to the baseline MDA-Net. Zhennan Xu, Xin Lyu 0001, Feng Xu 0008, Xin Li 0090, Yucong Huang, Caifeng Wu, Yiwei Fang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | DDFformer: a dual-domain fused transformer for polyp segmentation
Xi Yong, Jingchen Liang, Yun Hu 0004, Xin Li 0090, Hongmin Gao 0001, Zuojian Zhou, Kongfa Hu |
J. Supercomput. | 6 |
| 2024 | Locality-Enhanced Transformer for Semantic Segmentation of High-Resolution Remote Sensing ImagesabstractTransformers have emerged as a transformative tool in various computer vision tasks, excelling at capturing long-range dependencies. Their potential applicability and scalability in the interpretation of high-resolution remote sensing images (HRRSIs) have thus garnered substantial interest. However, unlike natural images, HRRSIs present intricate scenes characterized by scale variations and diverse appearances. These challenges underscore the importance of enabling networks to effectively assimilate both local intricacies and global context. In this letter, we introduce LETFormer, a semantic segmentation transformer. LETFormer balances capturing longrange dependencies with preserving local details through its unique LETFormer block, featuring an anchor token. This token aggregates localized contextual information within a designated window and promotes meaningful interactions among anchor tokens. With a mask transformer decoder, LETFormer gains ample contextual cues for precise semantic mask prediction. Empirical findings based on evaluations using the ISPRS Potsdam and LoveDA benchmarks unequivocally establish LETFormer’s superiority over state-of-the-art models. Additionally, we analyze the parameter size and floating-point operations per second (FLOPs) of LETFormer. Xin Li 0090, Feng Xu 0008, Runliang Xia, Nan Xu 0008, Fan Liu 0003, Chi Yuan, Qian Huang 0008, Xin Lyu 0001 |
ICASSP | 1 |
| 2024 | FreqFormer: A Frequency Transformer for Semantic Segmentation of Remote Sensing ImagesabstractSemantic segmentation of remote sensing images (RSIs) is vital for geospatial intelligence. However, traditional methods face challenges with mixed pixels and complex land cover types. Convolutional neural networks and transformers have led the field of RSI semantic segmentation by learning visual features in the spatial domain, but they often overlook the rich spectral features which can be well-described in the frequency domain, resulting in inadequate context modeling. In this paper, we present FreqFormer, a frequency transformer that enhances semantic segmentation by incorporating both spectral and spatial information through a devised frequency attention (FA) module. FA refines representations in the frequency domain through two parallel branches. Specifically, the high-frequency branch (HFB) utilizes a convolution layer with a Canny kernel to preserve high-frequency details, followed by multi-head self-attention to model high-frequency context. Followed by an element summation, high-frequency and low-frequency contexts are aggregated. Then, the formed FreqFormer block is sequentially deployed in the encoder stage with patch merging for spatial contraction. As for the decoder, the mask transformer decoder applies a scalar product to predict patch-wise semantics before upsampling. In experiments, FreqFormer outperforms state-of-the-art models on the ISPRS Potsdam and LoveDA datasets, demonstrating significant improvements in numerical evaluations. The integration of HFB significantly boosts the model’s ability to capture fine details, highlighting its potential for geospatial analysis. Xin Li 0090, Feng Xu 0008, Fan Liu 0003, Yiwei Fang, Xin Lyu 0001, Jun Zhou 0001 |
MMAsia | 1 |
| 2024 | Leave No One Behind: Online Self-Supervised Self-Distillation for Sequential Recommendation
Shaowei Wei, Zhengwei Wu, Xin Li 0090, Qintong Wu, Zhiqiang Zhang 0012, Jun Zhou 0011, Lihong Gu, Jinjie Gu |
WWW | 3 |
| 2024 | CSFFNet: Lightweight cross-scale feature fusion network for salient object detection in remote sensing imagesabstractAbstract Salient object detection (SOD), one of the most important applications in the field of computer vision, aims to extract the most visually appealing regions of scenes. However, the improvement of the accuracy of existing salient object detection in optical remote sensing images (ORSI‐SOD) is usually accompanied by an increase of network complexity, which affects the application of these models. Motivated by this, a novel lightweight edge‐supervised neural network for ORSI‐SOD is proposed, named CSFFNet. Specifically, the backbone (ResNet34) is first lightened by feature encoding module (FEM), building a lightweight subnet for feature extraction. Then, in the transformer‐based feature pyramid enhancement module (FPEM), the convolutional features obtained in the FEM are enhanced by long‐distance dependence to obtain multi‐scale features containing rich saliency cues. Based on this, the feature fusion module (FFM) is designed to capture cross‐scale long‐range dependencies and effectively fuse high‐level semantic information with low‐level detail information. Thus, the increase in network complexity due to multi‐level decoding is avoided. Finally, the segmentation results are optimized by using salient edges as auxiliary information, which effectively improves the contrast and completeness of the results. Experimental results on two public datasets demonstrate that the lightweight CSFFNet achieves competitive or even better performance compared with state‐of‐the‐art methods. Longbao Wang, Chong Long, Xin Li 0090, Xiaodan Tang, Zhipeng Bai, Hongmin Gao 0001 |
IET Image Process. | 3 |
| 2024 | A CBAM-GAN-based method for super-resolution reconstruction of remote sensing imageabstractAbstract As satellite imagery technology advances, remote sensing plays an increasingly prominent role in modern society. Nevertheless, the limitations of existing imaging sensors and complex atmospheric conditions constrain the quality of raw remote sensing data, posing challenges for interpretation and noise reduction. Super‐resolution technology focuses on enhancing low‐quality, low‐resolution remote sensing images. In this study, we introduce a method that utilizes a high‐order degradation model to generate low‐resolution remote sensing images. We employ a Generative Adversarial Network with a Convolutional Block Attention Module (CBAM‐GAN) to enhance these images, reducing noise interference and improving texture and feature display. Our approach outperforms other methods on the UCMerced‐LandUse, WHU‐RS19, and AID datasets. Specifically, it raises SSIM index scores to 0.9443, 0.8928, and 0.8633, respectively, exceeding baselines by 1.31%, 0.19%, and 1.30%. The MOS index also improves to 3.98, 3.96, and 3.83, respectively, representing a 2.31%, 8.20%, and 2.96% gain over the baseline. Our reconstruction produces superior results, demonstrating the effectiveness of our proposed method. Longbao Wang, Xin Li 0090, Hongmin Gao 0001 |
IET Image Process. | 3 |
| 2024 | SigCo: Eliminate the inter-class competition via sigmoid for learning with noisy labels
Feng Xu 0008, Xin Lyu 0001, Xin Li 0090 |
Knowl. Based Syst. | 5 |
| 2024 | AAFormer: Attention-Attended Transformer for Semantic Segmentation of Remote Sensing ImagesabstractThe rapid advancements in remote sensing technology have enabled the widespread availability of fine-resolution remote sensing images (RSIs), offering rich spatial details and semantics. Despite the applicability and scalability of transformers in semantic segmentation of RSIs by learning pairwise contextual affinity, they inevitably introduce irrelevant context, hindering accurate inference of patch semantics. To address this, we propose a novel multi-head attention-attended module (AAM) that refines the multi-head self-attention mechanism. The AAM filters out irrelevant context while highlighting informative ones by considering the relevance between self-attention maps and the query vector. The AAM generates an attention gate to complement contextual affinity and emphasize the useful ones with a higher weight simultaneously. Leveraging multi-head AAM as the core unit, we construct a lightweight attention-attended transformer block (ATB). Subsequently, we devise AAFormer, a pure transformer with a mask transformer decoder, for achieving semantic segmentation of RSIs. We extensively evaluate our approach on the ISPRS Potsdam and LoveDA datasets, demonstrating compelling performance compared to mainstream methods. Additionally, we conduct evaluations to analyze the effects of AAM. Xin Li 0090, Feng Xu 0008, Linyang Li, Nan Xu 0008, Fan Liu 0003, Chi Yuan, Xin Lyu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2024 | A Cross-Domain Coupling Network for Semantic Segmentation of Remote Sensing ImagesabstractSemantic segmentation of remote sensing images (RSIs) is critical for various applications, including urban planning, agriculture, and disaster management. Existing methods often fail to capture fine-grained textures and periodic patterns in RSIs, leading to suboptimal results in complex terrains. To address these challenges, we propose a cross-domain coupling network (CDCNet) that leverages both domain-specific extraction and cross-domain coupling (CDC) to enrich contextual cues for semantic inference. Our CDCNet integrates a CDC layer within the encoder-decoder architecture to simultaneously refine representations in the frequency and spatial domains. This approach effectively models fine-grained textures and periodic patterns in the frequency domain, as well as edges, shapes, and broad structural elements in the spatial domain. Extensive experiments on the ISPRS Potsdam and LoveDA datasets demonstrate the superiority of CDCNet over several state-of-the-art methods. Ablation studies confirm the significant impact of the CDC layer, validating the effectiveness of our approach in handling RSIs. Xin Li 0090, Feng Xu 0008, Feifei Tao, Hongmin Gao 0001, Fan Liu 0003, Xin Lyu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2024 | A Frequency Domain Feature-Guided Network for Semantic Segmentation of Remote Sensing ImagesabstractSemantic segmentation of Remote Sensing Images (RSIs) entails assigning semantic labels to each pixel accurately. RSIs are rich in spatial and spectral data, revealing diverse material and object characteristics. Yet, current RSI-focused computer vision models struggle with significant intra-class variation and inter-class resemblance due to limited spectral data usage. We propose the Frequency Domain Feature-Guided Network (FFGNet) for RSI semantic segmentation, influenced by digital signal processing theories. FFGNet initially generates frequency domain features via patch partitioning and 2D discrete cosine transformation. Our Frequency Enhancement Attention module (FEA) then distinguishes and intensifies frequency components to retain detailed information. These enhanced features are integrated with the Spatial-Spectral Attention (SSA) for enriched spectral signals. In the inference phase, these features are upsampled and combined with decoded features, emphasizing spectral details. Additionally, our novel loss function combines frequency and cross-entropy losses. Experiments on LoveDA and ISPRS Potsdam datasets demonstrate FFGNet's effectiveness, surpassing other mainstream models. An ablation study further validates our dual-guidance design. Xin Li 0090, Feng Xu 0008, Hongmin Gao 0001, Fan Liu 0003, Xin Lyu 0001 |
IEEE Signal Process. Lett. | 1 |
| 2024 | Semantic Segmentation of Remote Sensing Images by Interactive Representation Refinement and Geometric Prior-Guided InferenceabstractHigh spatial resolution remote sensing images (HRRSIs) contain intricate details and varied spectral distributions, making their semantic segmentation a challenging task. To address this problem, it is crucial to adequately capture both local and global contexts to reduce semantic ambiguity. While self-attention modules in vision transformers capture long-range context, they tend to sacrifice local details. In this article, we propose a geometric prior-guided interactive network (GPINet), a hybrid network that refines features across encoder and decoder stages. First of all, a dual branch structure encoder with local-global interaction modules (LGIMs) is designed to fully exploit local and global contexts for feature refinement. Unlike commonly used skip connections or concatenations, the LGIMs bilaterally couple and exchange CNN features with transformer features by lossless transformation and elaborating cross-attention. Moreover, we introduce a geometric prior generation module (GPGM) that iteratively updates the randomly initialized geometric prior. Subsequently, the geometric priors are stored and used to guide feature recovery. Finally, a weighted summation is applied to the upsampled decoded features and geometric priors. By comprehensively capturing contexts and enabling lossless decoding and deterministic inference, GPINet allows the network to learn discriminative representations for accurately specifying pixel-level semantics. Experiments on three benchmark datasets demonstrate the superiority of the proposed GPINet over state-of-the-art methods. Furthermore, we validate the effectiveness of geometric priors and compare the model sizes. Xin Li 0090, Feng Xu 0008, Fan Liu 0003, Xin Lyu 0001, Jun Zhou 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | A Synergistical Attention Model for Semantic Segmentation of Remote Sensing ImagesabstractIn remotely sensed images, high intraclass variance and interclass similarity are ubiquitous due to complex scenes and objects with multivariate features, making semantic segmentation a challenging task. Deep convolutional neural networks can solve this problem by modeling the context of features and improving their discriminability. However, current learning paradigms model the feature affinity in spatial dimension and channel dimension separately and then fuse them in a sequential or parallel manner, leading to suboptimal performance. In this study, we first analyze this problem practically and summarize it as attention bias that reduces the capability of network in distinguishing weak and discretely distributed objects from wide-range objects with internal connectivity, when modeled only in spatial or channel domain. To jointly model both spatial and channel affinity, we design a synergistic attention module (SAM), which allows for channelwise affinity extraction while preserving spatial details. In addition, we propose a synergistic attention perception neural network (SAPNet) for the semantic segmentation of remote sensing images. The hierarchical-embedded synergistic attention perception module aggregates SAM-refined features and decoded features. As a result, SAPNet enriches inference clues with desired spatial and channel details. Experiments on three benchmark datasets show that SAPNet is competitive in accuracy and adaptability compared with state-of-the-art methods. The experiments also validate the hypothesis of attention bias and the efficiency of SAM. Xin Li 0090, Feng Xu 0008, Fan Liu 0003, Xin Lyu 0001, Zhennan Xu, Jun Zhou 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Feature difference for single-shot object detectionabstractAbstract The one‐stage detectors achieve a good trade‐off between performance and latency, owing to the plain architecture and divergent learning mechanism for classification and localization. However, the two sub‐tasks require features with various inherency with which to generate inconsistent detections, fettering detectors. In this study, the misalignment is deeply analyzed via kernel density estimation (KDE) for the first time. Moreover, to address the misalignment, a plug‐and‐play detection head, named Diff‐Head, is devised and embedded in one‐stage detectors. Concretely, the authors merge parallel branches into a semi‐parallel structure, establishing the correlation between classification and regression. In the regression branch, a feature difference module (FDM) gets rid of the features that favour classification by subtracting salient object features from the original feature map, and position encoding (PE) modules enhance the absolute position information. The flexibility and efficiency of the detection head are retained. Experiments on Pascal visual object classes (VOC) and MS COCO demonstrate that Diff‐Head is effective and achieves competitive performance with state‐of‐the‐art detectors. Meanwhile, the amount of parameters is reduced at least 30% and 83.0% average precision (AP) is achieved on Pascal VOC. The analyses of consistency and error show that Diff‐Head has better localization and the capability of mitigating the misalignment. Feng Xu 0008, Xin Lyu 0001, Xin Li 0090, Xinyuan Wang 0002, Caifeng Wu |
IET Image Process. | 4 |
| 2022 | Hybridizing Euclidean and Hyperbolic Similarities for Attentively Refining Representations in Semantic Segmentation of Remote Sensing ImagesabstractAttention mechanisms have revolutionized the semantic segmentation network in interpreting remotely sensed images (RSIs) due to their amazing ability in establishing contextual dependencies. Nevertheless, due to the complex scenes and diverse objects in RSIs, a variety of details and correlations are not available in Euclidean space. Therefore, a similarity-hybrid attention module (SHAM) is devised to attentively learn the hyperbolic and Euclidean attention maps between any two positions, followed by a weighted element-wise summation. The hybrid attention maps posses latent geometric properties of both Euclidean and hyperboloid. Taking commonly-used fully convolutional network (FCN) as baseline, HAENet that embeds SHAM, is presented. Experiments on ISPRS Potsdam and DeepGlobe benchmarks reveal its superiority to comparative methods. In addition, the ablation study validates the effectiveness of SHAM compared to other attention modules. Xin Li 0090, Feng Xu 0008, Fan Liu 0003, Runliang Xia, Linyang Li, Zhennan Xu, Xin Lyu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | A Trajectory Released Scheme for the Internet of Vehicles Based on Differential PrivacyabstractThe locations and users’ information can be shared and interacted in the IoV (Internet of Vehicles), which provides sufficient data for traffic deployment and behavior pattern analysis. However, privacy issues had become more severe since personal or sensitive information is inclined to be revealed in a big data environment. In this work, a novel differential privacy-based algorithm named DPTD (Differentially Private Trajectory Database) is proposed for trajectory database releasing. Firstly, a 3-dimensional generalized trajectory dataset is established by considering the time factor. Then, the trajectory space is divided into several planes through the timestamps, and the set of the locations on each plane is further processed by clustering and generalizing to re-form new trajectories, that is, the trajectories to be released. This method is quite favorable to prefix-tree releasing because the spatiotemporal characteristics of the trajectories can be captured and spareness problem is fixed. Besides, a Markov assumption-based prediction method is suggested in order to reduce the cost of adding noise. Unlike the traditional method that the noise is added layer by layer, the noise is only added to the odd layers based on the prediction through spatio-temporal correlation, saving approximately 50% of the privacy budget. Theoretical analysis and experimental results show that the proposed algorithm has better data availability than the compared algorithms while guaranteeing the expected privacy level. Sujin Cai, Xin Lyu 0001, Xin Li 0090, Duohan Ban |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | Semantic Labeling of Very High-Resolution Imagery by Leveraging Contextual Information with Optimized Non-Local Neural NetworkabstractSemantic labeling of very high-resolution (VHR) aerial imagery attempts to assign a specific label to every pixel. Resorting to the significant progress of semantic segmentation for natural images, the non-local neural networks express superior capability in capturing multi-scale contextual information, which is helpful for accurately labeling raw images. However, the computational complexity of non-local neural networks is criticized. In this paper, the optimized non-local block (ONN) is devised and embedded in the encoder stage to reduce the computational complexity in capturing contextual information. Besides, the DUpsampling is introduced to further boost the efficiency of decoder. Finally, the extensive experiments, including numerical metrics and visual inspection comparisons, are conducted on ISPRS benchmarks. The results indicate that the proposed method performs better than several SOTA methods both in accuracy and efficiency. Xin Li 0090, Feng Xu 0008, Xin Lyu 0001, Liancheng Zhao, Xinyuan Wang 0002 |
IGARSS | 1 |