EDBT 2026 Demo / reviewers in the wild / expert
Xiyang Zhi
dblp:245/9754
· DBLP profile ↗
18ranked-venue papers
0as first author
18since 2021 · last 2026
0000-0001-5504-8480ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 16 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SAMamba: Stream Alignment Mamba for Motion Infrared Small Target Detection
Xiyang Zhi, Yuanxin Huang, Tianjun Shi, Shikai Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Progressive class-aware instance enhancement for aircraft detection in remote sensing imagery
Tianjun Shi, Jinnan Gong, Jianming Hu, Yu Sun 0028, Guangzhen Bao, Pengfei Zhang 0011, Xiyang Zhi, Wei Zhang 0220 |
Pattern Recognit. | 8 |
| 2025 | DWSDiff: Dual-Window Spectral Diffusion for Hyperspectral Anomaly DetectionabstractAnomaly detection (AD) has emerged as a critical area of research in hyperspectral imagery (HIS) processing, focusing on detecting sparse, small targets with spectral and spatial features deviating from the background without prior information. The approach of AD based on reconstruction differences is a leading method in deep learning (DL) for hyperspectral AD (HAD). A key challenge is the accurate estimation of complex backgrounds. The essence of this challenge lies in accurately reconstructing background regions while inferring the latent background of anomaly regions. In this article, we propose a novel method called dual-window spectral diffusion (DWSDiff) for HAD. To address the challenge of complex background estimation in HSIs, we developed a spectral diffusion model specifically tailored for HSI. This model achieves precise background estimation through an iterative spectral diffusion and reverse reconstruction process. We also introduced a dual-window strategy to mitigate the influence of anomaly extension areas within the neighborhood on background estimation. Moreover, the scarcity of paired labeled HSIs from the same scene, with and without anomalies, limits the model’s ability to learn features between anomaly and background. To address the shortage, we devised an anomaly generation strategy based on the principal component analysis (PCA) and the linear spectral mixing model (LSMM). Building on these, we designed a training and inference framework that integrates spectral diffusion, reverse background reconstruction, and target detection. Experimental results on the Airport-Beach–Urban (ABU) hyperspectral datasets demonstrate that DWSDiff outperforms 20 state-of-the-art (SOTA) HAD methods across six different areas under the curve (AUC) metrics. Wenbin Chen 0007, Xiyang Zhi, Shikai Jiang, Yuanxin Huang, Qichao Han, Wei Zhang 0220 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | StyleFormer: Spatial-Temporal Style Projecting Bidirectional Interactive Transformer for Change DetectionabstractRemote sensing image change detection is an important means for Earth monitoring task, which has a wide application prospect. In multitemporal optical remote sensing, there are inherent differences in factors such as lighting and sensors. This leads to the coupling of content change and image style change, making it difficult to distinguish. Therefore, a meaningful thinking for change detection is to decouple and capture the real changes of ground objects from multitemporal images. Based on this motivation, a novel general change detection architecture is explored, StyleFormer. It first proposes the concept of spatial–temporal style base and no longer constrains to semantic representation in a single image style. Instead, it introduces a spatial–temporal interactive style projection layer between bitemporal images, which projects the unseen diverse styles into the consistent expression space for change detection. Furthermore, an iterative interaction strategy of Transformer and CNN features is proposed to mine spatial–temporal context information more finely. It solves the lack of local perception and nonhierarchical features in ViT, and improves the model expression ability. After that, a change prior-guided cross-attention is introduced to fuse bitemporal features. It can adaptively enhance the change feature and improve the perception ability for small changes in remote sensing scenes. Sufficient experiments on four typical change detection datasets show that the proposed method is superior to the state-of-the-art methods. Especially on the datasets CDD-CD and SYSU-CD, the F1 score improved to 96.08% and 83.29%. The code of this work will be available athttps://github.com/Tom-Dongfang/change-detection-StyleFormer. Qichao Han, Xiyang Zhi, Jianming Hu, Shuqing Zhang, Wenbin Chen 0007, Yuanxin Huang, Shikai Jiang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Complementarity-Aware Feature Fusion for Aircraft Detection via Unpaired Opt2SAR Image TranslationabstractDetecting aircraft in complex remote sensing scene has significant value for military and civilian applications. To overcome the interference of complex environmental factors and achieve high-accuracy detection performance, the comprehensive utilization of optical and SAR images for object detection has become a promising research direction. However, currently there are problems with optical and SAR fusion detection, such as difficulty in obtaining paired registration training data and incomplete consideration of feature elements in fusion model. To tackle these challenges, we present an aircraft detection method based on optical-SAR complementarity-aware feature fusion. Firstly, an unpaired image translation model based on scattering feature enhancement GAN (SFEG) is designed to generate SAR images that are pixel-level registered with the input optical image. On this basis, a complementarity-aware feature fusion detection network (CFFDNet) combining differential feature spatial-aware complementary (DFSC) units and gate-generated weighted fusion (GWF) units is proposed to enhance the effective features of single source image while improving the complementary fusion effect of multimodal features. Experiments on CORS-ADD and MAR20 datasets demonstrate that our method outperforms the compared classical single-modal and multimodal detection models. The latest code is available soon at: https://github.com/JimmyRSlab/Complementarity-aware-Feature-Fusion-for-Aircraft-Detection-via-Unpaired-Opt2SAR-Image-Translation. Jianming Hu, Xiyang Zhi, Tianjun Shi, Wei Zhang 0220 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Dataset and Benchmark for Fine-Grained Ship Recognition in Complex Optical Remote Sensing ScenariosabstractShip recognition in remote sensing imagery is crucial for numerous applications such as monitoring maritime security, preventing illegal activities and implementing environmental protection. However, most existing research rarely focuses on both the complexity of the scene and the fine-grained recognition of targets, which limits the accuracy and applicability of ship recognition technology. To propel advancements in ship recognition methodology, we propose a new dataset named Fine-Grained Ship Recognition in Complex Scenarios (FGSRCS), which contains 17 typical categories of large and medium-sized ships from 280 widely distributed ports. For the richness of the image characteristics, the dataset images are collected from multiple satellite platforms, such as Ikonos, Jilin-1, OrbView, Pleiades and WorldView, and the time span of the data covers the past 20 years. To ensure the diversity of scenarios, the dataset mainly considers five complex scenarios, including thick clouds, mists, shadows, sea clutter and ground facilities, which helps to train and improve the algorithm applicability to practical application scenarios. Furthermore, we conduct experiments with ten state-of-the-art recognition algorithms on FGSRCS dataset, providing a benchmark for algorithm application. The research result can furnish both theoretical insights and practical guidance for the development of future ship recognition models. The dataset is available on https://github.com/dwddw/FGSRCS. Jianming Hu, Xiyang Zhi, Tianjun Shi |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | A Method for Detecting Aircraft Small Targets in Remote Sensing Images by Using CNNs Fused With Handcrafted FeaturesabstractAircraft target detection is a challenging task in remote sensing images, especially for aircraft small target detection. The most advanced object detection framework currently processes all information in the image uniformly through a deep neural network. In the past, in the process of detecting aircraft small targets, the feature extraction process was carefully designed, and hand-crafted features were derived from expert knowledge or historical data, which included prior knowledge that was conducive to object detection. Embedding prior features into deep neural networks can enhance the saliency of target information, improve the detection performance of the model. Accordingly, this paper proposes a Hand-crafted Feature Fusion Stream (HFFS) for embedding prior knowledge. We obtain hand-crafted features based on the grayscale co-occurrence matrix and edge extraction operator, and generate an attention map in deep convolutional neural networks (CNNs) to achieve the fusion of hand-crafted feature maps and high-level feature maps in deep convolutional networks. The experimental results show that using HFFS on the baseline model improves the detection performance of the model for aircraft small targets. Compared with the baseline model, our detection model achieves improvements of 1.1% AR, 1.6% [email protected], and 1.6% [email protected]:0.95 in the proposed dataset. Lijian Yu, Xiyang Zhi, Shuqing Zhang, Shikai Jiang, Jianming Hu, Wei Zhang 0220, Yuanxin Huang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | Gaussian Aware Anchor-Free Rotated Detector for Aerial Object DetectionabstractObject detection in specific scenarios has received increasing attention for many applications. However, the feature inconsistencies of localization and classification branches in remote sensing models may degrade the detection performance. Furthermore, the existing IoU-based label assignment strategy cannot accurately capture objects’ shape and oriented information. We propose an anchor-free detector called the Gaussian Aware Rotated Detector (GARDet) to address the above issues. It contains two improvements: the Feature Alignment Module (FAM) and the Gaussian Dynamic Label Assignment Strategy (GDLA). FAM consists of Oriented Feature Alignment Convolutions sensitive to orientation-invariant features inside objects and Spatial Feature Alignment Convolutions sensitive to spatial coordinate information. GDLA uses a Gaussian Matching Confidence (GMC) based on the Gaussian distance to measure the quality of the predicted bounding boxes and dynamically assigns positive and negative samples for training. Extensive experiments on remote sensing object detection datasets (DOTAv1.0 and HRSC2016) demonstrate that the proposed model can achieve competitive performance. Shuai Zhang 0052, Cunyuan Zhang, Lijian Yu, Wenyu Ji, Xiyang Zhi |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2024 | Self-Supervised Denoising via Blind Feature Extraction and Diffusion-Based Texture GenerationabstractIn the field of remote sensing, detection in dimly lit or shadowed areas has traditionally been difficult because of detector noise. Given that noise in real-world images of remote sensing exhibits spatial correlation, existing self-supervised methods encounter difficulties in reconciling the suppression of spatially correlated noise with the preservation of local texture details. To address this challenge, we propose a self-supervised model that combines blind-spot feature extraction with diffusion-based texture generation to fine denoising of real-world images under adverse conditions. We first introduce a blind-spot feature extraction structure based on the fusion of U-Net with blind-spot net (UBSN) and blind transformer (BTF). In UBSN, we integrate multistride blind-spot convolution (BSC) + dilated convolution (DC) feature extraction nodes and employ a Reshuffle strategy in skip-layer connections to maintain large-scale blind-spot characteristics. Additionally, we design a transformer structure for blind spot between patches to remove the noise with spatial correlations while ensuring global feature acquisition. Subsequently, to restore texture details blurred by the blind-spot structures, we introduce a texture generation diffusion structure during model training, achieving a balance between large-scale blind-spot characteristics and local rich texture details. Experimental results demonstrate that our approach outperforms other self-supervised denoising methods, even some methods leveraging unpaired images, without the need for parameters related to on-orbit satellite detectors. Guangzhen Bao, Xiyang Zhi, Pengfei Zhang 0011, Jianming Hu, Tianjun Shi, Shikai Jiang, Yayun Wu, Jinnan Gong |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Dataset and Benchmark for Ship Detection in Complex Optical Remote Sensing ImageabstractShip detection plays a pivotal role in numerous military and civil applications, yet detecting ships in complex maritime and aerial environments remains a challenging task. While several publicly available datasets for ship detection have been introduced by researchers, most of them do not adequately address the impacts of diverse and intricate environmental factors, which makes the trained algorithms difficult to apply for practical application scenes involving clouds, sea clutter, complex lighting, and facility interferences, limiting the effectiveness and robustness of the detection models. To advance the field of ship detection method research, we propose a high-quality dataset named ship collection in complex optical scene (SCCOS), which is obtained from multiple platform sources including Google Earth, Microsoft map, Worldview-3, Pleiades, Orbview-3, Jilin-1, and Ikonos satellites. The dataset comprehensively considers complex scenes such as thin clouds, mist, thick clouds, light shadows, sea clutter, and port facilities. Additionally, we conduct experiments on this dataset with 11 representative detection algorithms and establish a performance benchmark, which can provide the theoretical basis and practical reference for the design and optimization of subsequent ship detection models. The latest dataset is available at:https://github.com/JimmyRSlab/Dataset-and-Benchmark-for-Ship-Detection-in-Complex-Optical-Remote-Sensing-Image. Jianming Hu, Xiyang Zhi, Tianjun Shi, Xiaogang Sun |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | FDDBA-NET: Frequency Domain Decoupling Bidirectional Interactive Attention Network for Infrared Small Target DetectionabstractInfrared small target detection (IRSTD) involves determining the coordinate position of the target in complex infrared images. However, challenges arise due to the absence of internal texture structure, edge dispersion, weak energy characteristics of the target, and a significant amount of background clutter resembling the target’s morphology, impeding precise target location. To address these challenges, we propose an IRSTD network, named frequency domain decoupling bidirectional interactive attention network (FDDBA-NET), designed from the perspective of frequency domain decoupling (FDD). To suppress backgrounds that are similar in shape and structure to the target, exploiting the spectral differences between the target and background in the frequency domain, we adopt two learnable masks to extract the target-specific spectrum and the target-background-consistent spectrum detrimental to detection. The specific spectrum aids in target detection. A target-level contrast loss is designed to maximize the disparity between these two spectra, ensuring optimal detection results. In addition, to preserve target details in high-level semantic information, we introduce a bidirectional interactive attention module that leverages mutual modulation of deep global and shallow local features, facilitating deep and shallow feature fusion. To validate our approach, we conduct experiments comparing our proposed network with state-of-the-art conventional methods and deep learning methods on public datasets. The results demonstrate the superior performance of our method. Yuanxin Huang, Xiyang Zhi, Jianming Hu, Lijian Yu, Qichao Han, Wenbin Chen 0007, Wei Zhang 0220 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | LMAFormer: Local Motion Aware Transformer for Small Moving Infrared Target DetectionabstractIn temporal infrared small target detection, it is crucial to leverage the disparities in spatiotemporal characteristics between the target and the background to distinguish the former. However, remote imaging and the relative motion between the detection platform and the background cause significant coupling of spatiotemporal characteristics, making target detection highly challenging. To address these challenges, we propose a network named LMAFormer. First, we introduce a local motion-aware spatiotemporal attention mechanism that aligns and enhances multiframe features to extract local spatiotemporal salient features of targets while avoiding interference from moving backgrounds. Second, we employ a multiscale fusion transformer encoder that computes self-attention weights across and within scales during encoding, to establish multiscale correlations among different regions of temporal images, enabling motion background modeling. Last, we propose a multiframe joint query decoder. The shallowest feature map after multiscale feature propagation is mapped to initial query weights, which are refined through grouped convolutions to generate grouped query vectors. These are jointly optimized to encapsulate rich multiframe details, strengthening motion background modeling and target feature representation, improving prediction accuracy. Experimental results on the NUDT-MIRSDT, IRDST, and the established TSIRMT datasets demonstrate that our network outperforms state-of-the-art (SOTA) methods. Our code and dataset will be available athttps://github.com/lifier/LMAFormer. Yuanxin Huang, Xiyang Zhi, Jianming Hu, Lijian Yu, Qichao Han, Wenbin Chen 0007, Wei Zhang 0220 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Local Adaptive Prior-Based Image Restoration Method for Space Diffraction Imaging SystemsabstractThin-film diffractive optical elements (DOEs) have considerable potential to be used in the field of high-resolution remote sensing imaging satellites because of advantages such as a large aperture, small volume, lightness, wide tolerance range of surface shape, and easy replication. However, there are problems associated with thin-film diffraction imaging, including space variation, serious blur, and low contrast, which result in insufficient imaging quality with regard to traditional optical system requirements. To address this, a local adaptive prior-based image restoration method is proposed for thin-film diffraction imaging systems. An entire degraded image was divided into several isohalo regions based on imaging characteristics. Then, the regularization constraints were adaptively selected and updated according to the local scene prior characteristics. Additionally, the system parameters in the corresponding field of view were used as input to restore each subregion. In particular, the diffraction efficiency (DIE) was introduced into the model to remove the nondesign level background radiation. The experimental results show that the proposed algorithm can effectively improve the image quality of a thin-film diffraction imaging system, including space variation correction, clarity enhancement, and background radiation suppression. Furthermore, a DIE of less than 60% was found to significantly impact the final image products. Shikai Jiang, Jianming Hu, Xiyang Zhi, Wei Zhang 0220, Dawei Wang 0007, Xiaogang Sun |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Adaptive Feature Fusion With Attention-Guided Small Target Detection in Remote Sensing ImagesabstractSmall target detection in remote sensing images has considerable significance in practical applications such as military dynamic discrimination and traffic monitoring. However, the limited appearance features of small-scale targets and the widespread false alarm sources make small target detection in remote sensing images a tough challenge. To address these problems, we propose a novel small detection method by employing an adaptive multi-level feature fusion module (AMFFM) and an attention-augmented high-resolution head (AAHRH). Specifically, AMFFM is designed to suppress the interference of false alarm sources in complicated scenes. We upsample the high-level features by the context modeling of semantic information and refine the low-level features for noise removal. Then the enhanced multi-level features are fused based on the spatial and channel significance. After that, AAHRH is put forward to enhance the perception of small targets by embedding cross-dimension interaction with the attention mechanism. The prediction heads are reconstructed with high-resolution layers to improve the detection performance in densely distributed scenes. We conduct dilated and comparison experiments on a constructed small car dataset, a public small ship dataset, and the VEDAI dataset. The experimental results on two datasets verify the effectiveness and robustness of the proposed method with the state-of-the-art performance. Tianjun Shi, Jinnan Gong, Jianming Hu, Xiyang Zhi, Guiyi Zhu, Binhuan Yuan, Yu Sun 0028, Wei Zhang 0220 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Complex Optical Remote-Sensing Aircraft Detection Dataset and BenchmarkabstractAircraft detection in remote sensing images is significant in both military and civilian fields, such as air traffic control and battlefield dynamic monitoring. Deep learning methods can achieve promising detection performance with sufficient and labeled samples. However, current aircraft datasets are mainly from a single data source and lack diverse scenes and targets, making it difficult to train a robust and generalized detector. Therefore, we manually label and construct a complex optical remote sensing aircraft target detection dataset (CORS-ADD) from Google Earth and multiple satellites such as WorldView-2, WorldView-3, Pleiades, Jilin-1, and IKONOS. It contains 7,337 images covering typical airports and various rare scenes, including the aircraft carrier, ocean and land with flying aircraft. The dataset consists of 32,285 civil and military aircraft instances, including bombers, fighters, and early warning aircraft. These targets range from 4×4 pixels to 240×240 pixels and are all labeled with both horizontal bounding box (HBB) and oriented bounding box (OBB) annotations. The various scenes and sufficient instances can fully support the training and evaluation of data-driven algorithms. Meanwhile, based on the constructed dataset, we train and evaluate several detectors to provide a benchmark and help promote the development of aircraft detection techniques. Tianjun Shi, Jinnan Gong, Shikai Jiang, Xiyang Zhi, Guangzhen Bao, Yu Sun 0028, Wei Zhang 0220 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Influence of Space Variability on Remote Sensing Image Restoration PerformancesabstractWith the continuous increase in the resolution of optical remote sensing satellites, the influence of space variations on the image quality cannot be ignored, especially in new imaging systems such as thin-film diffraction and rectangular rotating pupils. This paper was conducted to analyze the influence of space variability on restoration performances of different methods, then a new processing strategy of space-variant images is proposed. According to the analytical experiment results, we suggest using the block method when the PSV < 0.20% and otherwise selecting the global method. In order to ensure the final image quality, we also suggest controlling the PSV within 0.28% when designing optical systems. This study can provide a foundation for optimizing the design of front-end optical systems and selecting back-end processing methods in engineering applications. Shikai Jiang, Xiyang Zhi, Tianjun Shi, Jianming Hu, Wei Zhang 0220, Jinnan Gong |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Supervised Multi-Scale Attention-Guided Ship Detection in Optical Remote Sensing ImagesabstractShip detection in optical remote sensing images plays a significant role in a wide range of civilian and military tasks. However, it is still a challenging issue owing to complex environmental interferences and a large variety of target scales and positions. To overcome these limitations, we propose a supervised multi-scale attention-guided detection framework, which can effectively detect ships of different scales both in complex pure ocean and port scenes. Specifically, a multi-scale supervision module is first proposed to adjust the semantic consistency of different feature levels, obtaining extracted features with small semantic gaps. Next, an attention-guided module is utilized to aggregate context information from both spatial and channel dimensions by calculating map correlations, adaptively enhancing the feature representation. Moreover, to preserve the attribute and spatial relationship of the optimized features, we adopt a capsule-based module as the classifier and obtain satisfactory classification performance. Experimental results conducted on two public high-quality datasets demonstrate that the proposed method obtains state-of-the-art performance in comparison with several advanced methods. Jianming Hu, Xiyang Zhi, Shikai Jiang, Hao Tang 0005, Wei Zhang 0220, Lorenzo Bruzzone |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Global Information Transmission Model-Based Multiobjective Image Inversion Restoration Method for Space Diffractive Membrane Imaging SystemsabstractDiffractive membrane imaging systems have been an important development trend for high-orbit satellite cameras owing to their advantages of large aperture, light weight, rapid manufacture, and low cost. However, caused by the cross-coupling effects of diffraction imaging, membrane properties, subaperture stitching, on-orbit disturbances, and other physical factors, lager-aperture space diffractive membrane imaging systems have specific and complex degradation characteristics: the modulation transfer function (MTF) and signal-to-noise ratio (SNR) have more prominent degradation and serious space-variant characteristics over fields of view, with obvious background radiation properties that seriously affect the application of imaging products. To address this problem, this study established a global information transmission model by characterizing the PSF and background radiation in a full field of view to represent the imaging law of an on-orbit system. Aiming at the inverse problem of the information transmission model, we also propose a novel image inversion restoration method for the special degradation characteristics. In particular, the effect of diffraction efficiency is introduced into the inversion restoration method to solve the background radiation problem. Moreover, we innovatively designed matrix regularization parameters to further improve the correction ability of spatial variation. When the diffraction efficiency was experimentally higher than 60% and the mean measured spatial variability was less than 0.2, the proposed method exhibited a satisfactory processing performance, and could improve multiobjective comprehensive processing, such as transfer function compensation, spatial variation correction, and background radiation removal. Shikai Jiang, Xiyang Zhi, Wei Zhang 0220, Dawei Wang 0007, Jianming Hu |
IEEE Trans. Geosci. Remote. Sens. | 2 |