VLDB 2026 Research / reviewers in the wild / expert
Yuanxin Ye
dblp:170/9759
· DBLP profile ↗
37ranked-venue papers
14as first author
28since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 36 · 14 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CIRSM-Net: A Cyclic Registration Network for SAR and Optical ImagesabstractThe registration of synthetic aperture radar (SAR) and optical images is critical in multimodal remote sensing image fusion. In recent years, deep learning-based registration networks have been continuously introduced. However, owing to the significant disparities in viewing angles and radiometric properties between SAR and optical images, current deep learning methods struggle to fully exploit the physical properties of radar imaging. In addition, many existing matching networks typically perform only a forward pass, resulting in suboptimal model performance. This article proposes a cyclic iterative registration SAR mechanism network (termed as CIRSM-Net) for the registration of SAR and optical images. First, we design a learning module that integrates the radar equation with a microwave scattering model to capture deep features from SAR images, and design a corresponding scattering feature loss to aid in better generalization across various radar images. Then, to explore optimization methods for matching networks, this study proposes a strategy of multiple iterative optimizations within the matching network. Specifically, it integrates speeding-up radiation-variation insensitive feature transform (RIFT2) supervision in the backend matching network and iteratively optimizes the final output. Finally, during the iteration process, we propose an innovative matching loss function that combines the rotation invariance supervision of RIFT2 with iterative optimization techniques to enhance feature matching accuracy. Experimental results on both public and our own datasets additionally confirm the effectiveness and superiority of the proposed approach, demonstrating its significant potential for practical applications. Peng Wang 0030, Daiyin Zhu, Xunqiang Gong, Yuanxin Ye, Harry F. Lee, Bo Huang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | RMSO-ConvNeXt: A Lightweight CNN Network for Robust SAR and Optical Image Matching Under Strong Noise InterferenceabstractSynthetic aperture radar (SAR) and optical images provide complementary imaging information, and their joint application holds broad prospects in military reconnaissance and aircraft visual navigation, with the key task of achieving accurate image matching. However, due to significant differences in imaging characteristics and the presence of strong noise interference in complex electromagnetic environments, SAR imaging may face serious challenges in matching with optical images. In addition, limited hardware resources on airborne platforms make it difficult for existing matching algorithms to meet the requirements for online real-time matching. In this article, a lightweight, high-performance, and robust method for robust SAR and optical image matching is proposed. The main contributions include the design of a lightweight convolutional neural network (CNN) and an SAR image random mask noise-resistant feature reconstruction network (named RMSO-ConvNeXt). First, a lightweight pseudo-Siamese dense feature extraction network module was explored for pixel-wise feature extraction of SAR and optical images, which can efficiently extract common features between heterogeneous images. Second, an SAR random noise mask feature reconstruction network was designed to address noise interference in SAR images, and a new joint loss function was constructed to significantly enhance the model’s ability to resist strong noise interference. Furthermore, a high-quality large-scale SAR and optical dataset has been made publicly available to contribute to the advancement and application of image matching. Finally, extensive experiments demonstrate that compared to current state-of-the-art matching methods, the proposed approach exhibits stronger robustness, better matching performance, and smaller model parameters. The code and dataset are shared onhttps://github.com/yeyuanxin110/RMSO-ConvNeXt. Chao Yang 0028, Guoqing Gong, Jiwei Deng, Yuanxin Ye |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Tuple Perturbation-Based Contrastive Learning Framework for Multimodal Remote Sensing Image Semantic SegmentationabstractDeep learning models exhibit promising potential in multimodal remote sensing image semantic segmentation (MRSISS). However, the constrained access to labeled samples for training deep learning networks significantly influences the performance of these models. To address that, self-supervised learning (SSL) methods have garnered significant interest in the remote sensing community. Accordingly, this article proposes a novel multimodal contrastive learning framework based on tuple perturbation, which includes the pretraining and fine-tuning stages. First, a tuple perturbation-based multimodal contrastive learning network (TMCNet) is designed to better explore shared and different feature representations across modalities during the pretraining stage and the tuple perturbation module is introduced to improve the network’s ability to extract multimodal features by generating more complex negative samples. In the fine-tuning stage, we develop a simple and effective multimodal semantic segmentation network (MSSNet), which can reduce noise by using complementary information from various modalities to integrate multimodal features more effectively, resulting in better semantic segmentation performance. Extensive experiments have been carried out on two published multimodal image datasets including optical and synthetic aperture radar (SAR) pairs, and the results show that the proposed framework can obtain more superior performance of semantic segmentation than the current state-of-the-art methods in cases of limited labeled samples. The source code is available athttps://github.com/yeyuanxin110/TMCNet-MSSNet. Yuanxin Ye, Jinkun Dai, Keyi Duan, Ran Tao 0003, Wei Li 0032, Danfeng Hong |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | An Attention-Based Fusion for Handcrafted and Deep Feature to Improve Optical and SAR Image MatchingabstractMatching optical and synthetic aperture radar (SAR) images has always been a challenging task in remote sensing image processing. Recently, optical-SAR image matching using deep learning has outperformed traditional methods that rely on handcrafted features. However, most current deep learning methods emphasize capturing deep semantics while disregarding stable handcrafted features. This approach may restrict the advancement of optical-SAR image matching technology. Accordingly, this study introduces an attention-based feature fusion method for integrating handcrafted and deep features, which enhances the performance of optical-SAR image matching. First, we extract handcrafted and deep features separately. Then, an attention module is designed to integrate them for computing the regions of interest. Finally, we use the calculated weights to fuse the handcrafted and deep features for image matching. Compared with single-type features, fused features can better capture the common characteristics of optical and SAR images. The experimental results demonstrate a significant improvement in matching accuracy with the use of merged features. Zhiqiang Han, Jinkun Dai, Yuanxin Ye |
IGARSS | 4 |
| 2024 | ESAM-CD: Fine-Tuned EfficientSAM Network With LoRA for Weakly Supervised Remote Sensing Image Change DetectionabstractChange detection (CD) has become an attractive research topic in the field of remote sensing imagery in recent years. Despite significant advancements driven by deep learning (DL) techniques, most current methods predominantly rely on fully supervised strategies. These methods require the collection of a large number of pixel-level labels, which is quite time consuming and label intensive. To address that, we propose a weakly supervised CD method with EfficientSAM (ESAM)-CD termed, which leverages multiscale class activation map (CAM) fusion and a fine-tuned EfficientSAM’s image encoder. First, we construct a classification model employing image-level labels with a deep supervision strategy to generate high-quality multiscale CAM. Subsequently, a multiscale CAM fusion module is proposed to refine the boundaries of change targets by harnessing information from various scales. Then, we utilize EfficientSAM with powerful generalization capabilities as the backbone and fine-tune it using a low-rank adaptation (LoRA) strategy to establish a CD network. In such a network, bitemporal images and the generated pseudolabels are fed into the network. In addition, to overcome the reliance of EfficientSAM’s decoder on prompts, we propose a prompt-free decoder based on the general convolutional layers to predict change maps. Finally, we validate the effectiveness of the proposed ESAM-CD using two publicly available CD datasets (i.e., WHU-CD and LEVIR-CD). Comprehensive experiments demonstrate that our method outperforms other weakly supervised CD methods, achieving outstanding performance on both datasets. Mengmeng Wang 0007, Xinghua Li 0002, Yuanxin Ye |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Robust Optical and SAR Image Matching Using Attention-Enhanced Structural FeaturesabstractDue to the complementary nature of optical and SAR images, their alignment is of increasing interest. However, due to the significant radiometric differences between them, precise matching becomes a very challenging problem. Although current advanced structural features and deep learning-based methods have proposed feasible solutions, there is still much potential for improvement. In this paper, we propose a hybrid matching method using attention-enhanced structural features (namely AESF), which combines the advantages of both handcrafted-based and learning-based methods to improve the accuracy of optical and SAR image matching. It mainly consists of two modules: a novel effective multi-branch global attention (MBGA) module and a joint multi-cropping image matching loss function (MCTM) module. The MBGA module is designed to focus on shared information in structural feature descriptors of heterogeneous images across space and channel dimensions, significantly improving the expressive capacity of the classical structural features and generating more refined and robust image features. The MCTM module is constructed to fully exploit the association between global and local information of the input image, which can optimize the triple loss discriminator to discriminate positive and negative samples. To validate the effectiveness of the proposed method, it is compared with five state-of-the-art matching methods by using various optical and SAR datasets. The experimental results show that the matching accuracy at the 1-pixel threshold is improved by about 1.8%-8.7% compared with the most advanced deep learning method (OSMNet) and 6.5%-23% compared with the handcrafted description method (CFOG). Yuanxin Ye, Chao Yang 0028, Guoqing Gong, Peizhen Yang, Dou Quan, Jiayuan Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Optical and SAR Image Fusion Based on Complementary Feature Decomposition and Visual Saliency FeaturesabstractWith the expansion of optical and SAR image fusion application scenarios, it is necessary to integrate their information in land classification, feature recognition, and target tracking. Current methods focus excessively on integrating multimodal feature information to enhance the information richness of the fused images, whereas neglecting the highly corrupted visual perception of the fused results by modal differences and SAR speckle noise. To address that, this paper proposes a novel optical and SAR image fusion framework named Visual Saliency Features Fusion (VSFF), which is based on the extraction and balancing of significant complementary features of optical and SAR images. Firstly, we propose a decomposition algorithm of complementary features to divide the image into main structure features and detail texture features. Then, for the fusion of main structure features, we reconstruct the visual saliency features maps of the pixel and structure that contain significant information from optical and SAR images, and input them into a total variation constraint model to compute the fusion result and achieve the optimal information transfer. Meanwhile, we construct a new feature descriptor based on the Gabor wavelet, that separates meaningful detail texture features from residual noise and selectively preserves features that can improve the interpretability of fusion result. In a comparative analysis with seven state-of-the-art fusion algorithms, VSFF achieved better results in qualitative and quantitative evaluations, and our fused images have a clear and appropriate visual perception. The source code is publicly available at https://github.com/yeyuanxin110/VSFF. Yuanxin Ye, Xiaoyue Ren, Jianwei Fan |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Masked Conditional Variational Autoencoders for Chromosome StraighteningabstractKaryotyping is of importance for detecting chromosomal aberrations in human disease. However, chromosomes easily appear curved in microscopic images, which prevents cytogeneticists from analyzing chromosome types. To address this issue, we propose a framework for chromosome straightening, which comprises a preliminary processing algorithm and a generative model called masked conditional variational autoencoders (MC-VAE). The processing method utilizes patch rearrangement to address the difficulty in erasing low degrees of curvature, providing reasonable preliminary results for the MC-VAE. The MC-VAE further straightens the results by leveraging chromosome patches conditioned on their curvatures to learn the mapping between banding patterns and conditions. During model training, we apply a masking strategy with a high masking ratio to train the MC-VAE with eliminated redundancy. This yields a non-trivial reconstruction task, allowing the model to effectively preserve chromosome banding patterns and structure details in the reconstructed results. Extensive experiments on three public datasets with two stain styles show that our framework surpasses the performance of state-of-the-art methods in retaining banding patterns and structure details. Compared to using real-world bent chromosomes, the use of high-quality straightened chromosomes generated by our proposed method can improve the performance of various deep learning models for chromosome classification by a large margin. Such a straightening approach has the potential to be combined with other karyotyping systems to assist cytogeneticists in chromosome analysis. Jingxiong Li, Sunyi Zheng, Zhongyi Shui, Shichuan Zhang, Linyi Yang, Yuxuan Sun 0002, Honglin Li 0001, Yuanxin Ye, Peter M. A. van Ooijen, Kang Li 0004, Lin Yang 0002 |
IEEE Trans. Medical Imaging | 9 |
| 2023 | Image Template Matching via Dense and Consistent Contrastive LearningabstractImage template matching refers to localizing a small query image as opposed to a large reference image map. The query image a.k.a template has to be screened across every equal-sized region in the reference map to perform inner-product at pixel-level and the resulting similarity indicates the template location. Due to the domain heterogeneity between template and reference images, the matching performance degrades under dramatic appearance changes. More severely, the asymmetric matching easily leads to over-fitting by suggesting excessively false positive regions. To these ends, we propose an effective template matching method based on contrastive learning to perform a dense and consistent InfoNCEloss during matching. This can increase the matching at finer details, and thus effectively regularizes network training to prevent over-fitting. Extensive experiments on the synthetic aperture radar (SAR) and optical datasets, i.e., SEN1-2 and OS datasets demonstrate that our proposed method outperforms state-of-the-art methods by a large margin. Bo Li 0090, Lin Wu 0001, Deyin Liu, Hongyang Chen 0001, Yuanxin Ye, Xianghua Xie |
ICME | 5 |
| 2023 | Modality-Invariant Structural Feature Representation for Multimodal Remote Sensing Image MatchingabstractEstablishing feature correspondences between multimodal remote sensing images is an essential task for realizing diverse applications. Conventional matching methods, which employ gradient or phase congruency (PC) for feature detection and description, produce limited performance when images suffer from strong noises and intensity differences. In this study, we propose a novel modality-invariant structural feature representation (MISFR) method for multimodal remote sensing image matching. First, a maximal/minimal enhanced PC moment (EPCM) representation is designed by incorporating the PC with a multiscale relative total variation (RTV) model for feature detection and description. The EPCM integrates the advantages of these two models to exploit intrinsic structural features while providing robustness against modality variations. Then, to improve the feature stability and repeatability, an aggregation structural feature detector (ASFD) is proposed, in which the corner and edge points are separately extracted on the minimal and maximum EPCMs. Moreover, an adaptive binning-based log-polar descriptor is constructed on the maximal EPCM, named enhanced features of orientated PC (EOPC), to robustly characterize feature points. Various experiments on a series of multimodal remote sensing images demonstrate that our MISFR significantly improves the matching performance compared with four state-of-the-art approaches. Jianwei Fan, Jian Li 0018, Yuanxin Ye |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2023 | Combining Phase Congruency and Self-Similarity Features for Multimodal Remote Sensing Image MatchingabstractThe structural features using self-similarity have become more popular for multimodal remote sensing image matching. However, mostly because of significant geometric distortions and nonlinear intensity differences between images, these methods produce a limited matching performance when directly applied to multimodal remote sensing images. To address that, we propose a novel feature descriptor named pyramid features of orientated self-similarity (POSS) for multimodal remote sensing image matching, which integrates phase congruency (PC) into the self-similarity model for better encoding structural information. Unlike these conventional self-similarity-based descriptors, the POSS is constructed by using PC instead of image intensity in a pyramid manner and thus has better discriminability and robustness to modality variations. In addition, we design a uniform multimoment of PC (UMMPC) detector for feature detection, which can improve the number, distribution, and repeatability of feature points. Experimental results demonstrate its superior applicability of the proposed method over the state-of-the-art methods, showing much better matching performance. Jianwei Fan, Yuanxin Ye, Jian Li 0018 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2023 | Generative Adversarial Network With Transformer for Hyperspectral Image ClassificationabstractIn recent years, generative adversarial networks (GAN) have made great progress in the field of hyperspectral image classification (HIC), which alleviates the problem of insufficient training samples to a large extent. At present, GAN in the field of HIC are all based on Convolutional Neural Network (CNN). But CNN cannot extract sequence information well, and it is difficult to model remote dependencies. However, hyperspectrum is rich in spectral sequence information, and Transformer has been proven to be good at processing sequence information. Therefore, in order to process spectral information and alleviate the problem of insufficient training samples of hyperspectral images (HSI), we put forward a new frame Transformer with residual upscale GAN (TRUG). TRUG includes a generator G and a discriminator D. In the G, we propose residual upscale (RU) to improve the resolution of generated features, while also extracting texture features and capturing context relationships. In addition, we visualized the generated fake images for more intuitive analysis. In the D, we adopt Transformer block with progressively decreasing scale, and use grid self-attention mechanism in the first layer to better extract features for classification. In addition, GAN are prone to the problem of unstable training. In order to solve this problem, we improve the normalization algorithm and add relative position coding. We applied a pure Transformer based GAN to HIC datasets. Experimental results show that the proposed TRUG model has better performance than other models on the three datasets. Siyuan Hao, Yufeng Xia, Yuanxin Ye |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2023 | Adjacent-Level Feature Cross-Fusion With 3-D CNN for Remote Sensing Image Change DetectionabstractDeep learning-based change detection (CD) using remote sensing images has received increasing attention in recent years. However, how to effectively extract and fuse the deep features of bi-temporal images for improving the accuracy of CD is still a challenge. To address that, a novel adjacent-level feature fusion network with 3D convolution (named AFCF3D-Net) is proposed in this article. First, through the inner fusion property of 3D convolution, we design a new feature fusion way that can simultaneously extract and fuse the feature information from bi-temporal images. Then, to alleviate the semantic gap between low-level features and high-level features, we propose an adjacent-level feature cross-fusion (AFCF) module to aggregate complementary feature information between the adjacent levels. Furthermore, the full-scale skip connection strategy is introduced to improve the capability of pixel-wise prediction and the compactness of changed objects in the results. Finally, the proposed AFCF3D-Net has been validated on the three challenging remote sensing CD datasets: the Wuhan building dataset (WHU-CD), the LEVIR building dataset (LEVIR-CD), and the Sun Yat-Sen University dataset (SYSU-CD). The results of quantitative analysis and qualitative comparison demonstrate that the proposed AFCF3D-Net achieves better performance compared to other state-of-the-art methods. The code for this work is available at https://github.com/wm-Githuber/AFCF3D-Net. Yuanxin Ye, Mengmeng Wang 0007, Guangyang Lei, Jianwei Fan, Yao Qin 0002 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | R₂FD₂: Fast and Robust Matching of Multimodal Remote Sensing Images via Repeatable Feature Detector and Rotation-Invariant Feature DescriptorabstractIdentifying feature correspondences between multimodal images is facing enormous challenges because of the significant differences both in radiation and geometry. To address these problems, we propose a novel feature matching method (named R2FD2) that is robust to radiation and rotation differences, which consists of a repeatable feature detector and a rotation-invariant feature descriptor. In the first stage, a repeatable feature detector called the Multi-channel Auto-correlation of the Log-Gabor (MALG) is presented for feature detection, which combines the multi-channel auto-correlation strategy with the Log-Gabor wavelets to detect interest points (IPs) with high repeatability and uniform distribution. In the second stage, a rotation-invariant feature descriptor is constructed, named the Rotation-invariant Maximum index map of the Log-Gabor (RMLG), which includes fast assignment of dominant orientation and construction of feature representation. In the process of fast assignment of dominant orientation, a Rotation-invariant Maximum Index Map (RMIM) is built to address rotation deformations. Then, the proposed RMLG incorporates the rotation-invariant RMIM with the spatial configuration of DAISY to improve RMLG’s resistance to radiation and rotation variances. Finally, we conduct experiments to validate the matching performance of our R2FD2utilizing different types of multimodal image datasets. Experimental results show that the proposed R2FD2outperforms five state-of-the-art feature matching methods. Moreover, our R2FD2achieves the accuracy of matching within two pixels and has a great advantage in matching efficiency over contrastive methods. Bai Zhu, Chao Yang 0028, Jinkun Dai, Jianwei Fan, Yao Qin 0002, Yuanxin Ye |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Comparative Analysis of Pixel Level Fusion Algorithms in High Resolution SAR and Optical Image FusionabstractFusion of Synthetic aperture radar (SAR) and optical images is a significant topic in the field of remote sensing. As a typical category of image fusion methods, pixel level image fusion algorithms have been widely used in SAR-optical image fusion to integrate their complementary information and facilitate the subsequent interpretation and application. The effectiveness of these methods has been demonstrated in different literatures based on the experiment carried on specific, individual datasets, which make a comprehensive comparison of these algorithms difficult to achieve. This paper builds a sub-meter SAR and optical image dataset covering different types of scenes, the performance of 11 pixel level image methods is then investigated based on qualitative and quantitative analysis. Result shows the gradient pyramid (GP) achieve a high quality fusion when dealing with Optical-SAR image fusion task of residents, the non subsampled contourlet transform (NSCT) performs best when fusing images containing farmland and mountains. Yuanxin Ye, Chao Yang 0028, Yangang Zhao |
IGARSS | 2 |
| 2022 | Remote Sensing Object Detection Based on Receptive Field Expansion BlockabstractDue to the rapid development of deep learning techniques and the collection of large-scale remote sensing datasets, convolutional neural networks (CNNs) have made significant progress in remote sensing object detection. However, due to the diversity of objects in remote sensing images, multiscale object detection is still a challenging task. In this letter, a novel object detection framework based on feature pyramid network (FPN) is proposed to improve the detection performance of multiscale objects. First, a receptive field expansion block (RFEB) is designed and added on the top of the backbone to expand the receptive field of FPN adaptively. In this way, the context information around each object is well captured. Then, the features obtained via RFEB are delivered to feature maps at all pyramid levels, remedying the drawback of FPN that semantic information captured by deep layers is gradually diluted when transmitted to lower layers. Third, since the classic backbone of FPN, which produces large receptive fields based on large downsampling factors, may limit the effectiveness of RFEB, the backbone of the original FPN is modified using dilated convolution to ease the resolution drop of feature maps while maintaining a large receptive field. As a feature extractor, the proposed framework can be easily deployed in other FPN-based methods. The experiments on the benchmark for object DetectIon in Optical Remote sensing images (DIOR) dataset demonstrate the proposed method’s superiority over considered state-of-the-art baseline methods in terms of detection accuracy. Xiaohu Dong, Ruigang Fu, Yinghui Gao, Yao Qin 0002, Yuanxin Ye |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | Remote Sensing Object Detection Based on Gated Context-Aware ModuleabstractRecently, deep learning algorithms, especially feature pyramid network (FPN), have achieved significant progress in object detection of natural scene images. However, due to the complex scenes of remote sensing images and the diversity of remote sensing objects, FPN still faces the following drawback when applied to remote sensing object detection. Specifically, in the original FPN, the features of each proposal are extracted by RoIAlign. However, these features have limited effective receptive fields, making FPN lack of crucial contextual information to accurately classify and locate objects, as well as filter some background noises that possess similar appearance with objects. To alleviate the above problem, in this letter, we propose a gated context aware module (G-CAM), and replace the original RoIAlign in FPN with the proposed G-CAM to adaptively incorporate the useful local context surrounding each proposal and the global context of the whole image into FPN, enabling FPN to effectively detect objects in remote sensing images Extensive experiments have been conducted on the DIOR and RSOD datasets, which validates that the proposed method achieves superior performance to the considered state-of-the-art methods in terms of detection accuracy. Xiaohu Dong, Yao Qin 0002, Ruigang Fu, Yinghui Gao, Yuanxin Ye |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2022 | Multiscale Deformable Attention and Multilevel Features Aggregation for Remote Sensing Object DetectionabstractIn this letter, a novel object detection method based on feature pyramid network (FPN) is proposed to improve the detection performance of remote sensing objects. First, since the information in the background regions may interfere with object detection, a novel multi-scale deformable attention module (MSDAM) is designed and added on the top of the backbone of FPN to make the network suppress the background features while highlight the target features. The proposed MSDAM generates attention maps from feature maps with multi-scale deformable receptive fields, thus can fit remote sensing objects of various shapes and sizes better and predict more precise attention maps for remote sensing images. Second, in the original FPN, each proposal is predicted based on feature grids pooled from only one feature level. This process is suboptimal as the information discarded in other feature levels and the global contextual information are also meaningful to object detection. Thus, a multi-level features aggregation module (MLFAM) is proposed to aggregate the multi-level outputs of FPN and the global context of the whole image, generating more powerful pyramidal representations for the subsequent object detection. The experiments conducted on the DIOR and RSOD datasets demonstrate the superiority of the proposed method over the considered state-of-the-art baseline methods in terms of detection accuracy. Xiaohu Dong, Yao Qin 0002, Ruigang Fu, Yinghui Gao, Yuanxin Ye |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2022 | Phase Congruency Order-Based Local Structural Feature for SAR and Optical Image MatchingabstractAutomatic matching of synthetic aperture radar (SAR) and optical images is a fundamental task in many remote sensing applications. However, due to different imaging modalities, conventional matching methods provide limited performances. In this letter, based on the observation that structural features are maintained across different modality images, we propose a novel feature-based method to effectively address SAR and optical image matching. The proposed method is built on the phase congruency (PC) model and consists mainly of two stages. First, a modified version of the uniform nonlinear diffusion-based Harris (MUND-Harris) detector is introduced to extract the local features. Unlike the UND-Harris, MUND-Harris employs the PC instead of image intensity for feature extraction and thus obtain well-distributed and highly repeatable feature points. Second, a local structural descriptor, namely PC order-based local structural (PCOLS), is designed for the extracted points. PCOLS is constructed in a grouping manner and further encodes image structures with an adaptive descriptor structure, which provides robustness against modality variations including significant geometric and intensity differences. Experimental results obtained on several SAR and optical image pairs demonstrate the encouraging performance of the proposed method. Jianwei Fan, Yuanxin Ye, Guichi Liu, Jian Li 0018 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | An Optical Flow SBAS Technique for Glacier Surface Velocity Extraction Using SAR ImagesabstractThe pixel offset-tracking (PO) technique developed from correlation matching for use on synthetic aperture radar (SAR) images has been widely employed to monitor glacier dynamics. However, the decorrelation caused by rapid changes in a glacier surface reduces the integrity of flow velocity extraction. In this letter, we propose a novel method, termed the optical flow (OF) small baseline subset(SBAS), developed from the optical flow algorithm, which is defined as the apparent motion of individual pixels on the image plane. The OF algorithm can compute dense flows at the individual pixel level with a low computational cost and may serve as an alternative to PO for estimating the glacier velocity field. During processing, We selected the image pairs having short spatiotemporal baselines and calculated their offset series according to the OF algorithm. We then used an interval estimation strategy to eliminate outliers to refine stacked offsets. Finally, the least-squares method was used to calculate the displacement series for each pixel. We tested the proposed method utilizing eight ALOS-2/PALSAR-2 images on a temperate debris-covered glacier, the Hailuogou Glacier on the southeastern Tibetan Plateau. Compared with PO-SBAS, our method effectively improves the coverage of glacier flow velocity monitoring from 73.5% to 99.6%. Yin Fu, Bo Zhang 0067, Guoxiang Liu 0001, Rui Zhang 0052, Qiao Liu 0004, Yuanxin Ye |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2022 | Spectral and Spatial Feature Fusion for Hyperspectral Image ClassificationabstractCompared with traditional images, hyperspectral images (HSI) not only have spatial information, but also have rich spectral information. However, the mainstream hyperspectral image classification (HIC) methods are all based on Convolutional Neural Network (CNN), which has great advantages in extracting spatial features, but it has certain limitations in dealing with spectral continuous sequence information. Therefore Transformer which is good at processing sequences, has also been gradually applied to HIC. Besides, Since HSI are typical three-dimensional structures, we believe that the correlation of the three dimensions is also an important information. So in order to fully extract the spectral spatial information, as well as the correlation of the three dimensions. we propose a spectral and spatial feature fusion module (i.e., TransCNN) for HIC. TransCNN consists of CNNs and a Transformer. The former is in charge of mining the spatial and spectral information from different dimensions, while the latter not only undertakes the most critical fusion but also captures the deeper relationship characteristics. We transpose the data to extract features and their correlation through three CNNs branches. we believe that these feature maps still have deep spectral information. Therefore, we have embedded them into one-dimensional vectors and use Transformer’s Encoder to extract features. However, some information will be lost when embedding into one-dimensional vectors. Therefore we use Decoder, which has been ignored in the field of vision, to fuse the features before passing Encoder and the features after extracted by Encoders. Two kinds of features are fused by Decoder, and the obtained information is finally input into the classifier for classification. Experimental results on real HSIs show that the proposed architecture can achieve competitive performance compared with the state-of-the-art methods. Siyuan Hao, Yufeng Xia, Lijian Zhou, Yuanxin Ye, Wei Wang 0108 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | An Unsupervised SAR and Optical Image Fusion Network Based on Structure-Texture DecompositionabstractAlthough the unique advantages of optical and synthetic aperture radar (SAR) images promote their fusion, the integration of complementary features from the two types of data and their effective fusion remains a vital problem. To address that, a novel framework is designed based on the observation that the structure of SAR images and the texture of optical images look complementary. The proposed framework, named SOSTF, is an unsupervised end-to-end fusion network that aims to integrate structural features from SAR images and detailed texture features from optical images into the fusion results. The proposed method adopts the nest connect-based architecture, including an encoder network, a fusion part, and a decoder network. To maintain the structure and texture information of input images, the encoder architecture is utilized to extract multi-scale features from images. Then, we use the densely connected convolutional network (DenseNet) to perform feature fusion. Finally, we reconstruct the fusion image using a decoder network. In the training stage, we introduce a structure-texture decomposition model. In addition, a novel texture-preserving and structure-enhancing loss function are designed to train the DenseNet to enhance the structure and texture features of fusion results. Qualitative and quantitative comparisons of the fusion results with nine advanced methods demonstrate that the proposed method can fuse the complementary features of SAR and optical images more effectively. Yuanxin Ye, Wanchun Liu, Qizhi Xu |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | Optical-to-SAR Image Matching Using Multiscale Masked Structure FeaturesabstractAutomatic and precise matching between optical and synthetic aperture radar (SAR) images is still a challenging task because of significant radiation and texture differences between such images. Recently, structure feature-based methods are popular for the matching of SAR and optical images. However, current structure descriptors include many noninformative features, which degrade their matching performance. To address that, we present a robust matching method by a multiscale masked structure feature representation. We first extract pixelwise gradient structure features on multiple scales of images. Then, a mask is constructed according to large contours of an image, which is used to increase the contribution of the main structure region and alleviate the influence of noninformative regions. Finally, a fast template scheme based on fast Fourier transform (FFT) is employed to obtain correspondences. The proposed method is tested using the optical and SAR images from the Sentinel and GaoFen sensors. Experiment results show that the proposed method significantly improves the matching performance compared with the state-of-the-art methods, especially for the images with poor structure features. Yuanxin Ye, Chao Yang 0028, Jianwei Fan, Yao Qin 0002 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | Robust Matching for SAR and Optical Images Using Multiscale Convolutional Gradient FeaturesabstractImage matching is a key preprocessing step for the integrated application of synthetic aperture radar (SAR) and optical images. Due to significant nonlinear intensity differences between such images, automatic matching for them is still quite challenging. Recently, structure features have been effectively applied to SAR-to-optical image matching because of their robustness to nonlinear intensity differences. However, structure features designed by handcraft are limited to achieve further improvement. Accordingly, this letter employs the deep learning technique to refine structure features for improving image matching. First, we extract multiorientated gradient features to depict the structure properties of images. Then, a shallow pseudo-Siamese network is built to convolve the gradient feature maps in a multiscale manner, which produces the multiscale convolutional gradient features (MCGFs). Finally, MCGF is used to achieve image matching by a fast template scheme. MCGF can capture finer common features between SAR and optical images than traditional handcrafted structure features. Moreover, it also can overcome some limitations of current matching methods based on deep learning, which requires solving a huge number of model parameters by a large number of training samples. Two sets of SAR and optical images with different resolutions are used to evaluate the matching performance of MCGF. The experimental results show its advantage over other state-of-the-art methods. Yuanxin Ye, Tengfeng Tang, Ke Nan, Yao Qin 0002 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | A Novel Multiscale Adaptive Binning Phase Congruency Feature for SAR and Optical Image RegistrationabstractAutomatic registration of synthetic aperture radar (SAR) and optical images is still a challenging problem because of the potential differences in geometric and intensity. In this work, we propose a robust and efficient method for improving the registration performance of SAR and optical images. Our work consists mainly of two steps, including the feature point detection stage and the feature description stage. In the first stage, we present a new feature extraction method, named nonlinear diffusion-based Harris-Laplace (NDHL) detector, which incorporates the nonlinear diffusion and a spatial consistency strategy into the Harris-Laplace detector, which can weaken the modality variations and enhance the structural features of the multimodal images, and can detect many more similar and highly repeatable feature points. In the second stage, we design a novel structural descriptor, named multiscale adaptive binning phase congruency (MABPC). The proposed MABPC descriptor encodes multiscale phase congruency features with an adaptive binning spatial structure, which brings an improvement of the robustness against geometric and nonlinear intensity discrepancies. Experimental results on both simulated and real image pairs show that the proposed registration method achieves encouraging performance improvements over other state-of-the-art methods. In comparison with the ROS-PC and RIFT methods, the proposed method obtains an average 54.6% improvement for the number of correct matches and has an average registration precision of 1.38 pixels. Jianwei Fan, Yuanxin Ye, Jian Li 0018, Guichi Liu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | A Multiscale Framework With Unsupervised Learning for Remote Sensing Image RegistrationabstractRegistration for multisensor or multimodal image pairs with a large degree of distortions is a fundamental task for many remote sensing applications. To achieve accurate and low-cost remote sensing image registration, we propose a multiscale framework with unsupervised learning, named MU-Net. Without costly ground truth labels, MU-Net directly learns the end-to-end mapping from the image pairs to their transformation parameters. MU-Net stacks several deep neural network (DNN) models on multiple scales to generate a coarse-to-fine registration pipeline, which prevents the backpropagation from falling into a local extremum and resists significant image distortions. We design a novel loss function paradigm based on structural similarity, which makes MU-Net suitable for various types of multimodal images. MU-Net is compared with traditional feature-based and area-based methods, as well as supervised and other unsupervised learning methods on the optical-optical, optical-infrared, optical-synthetic aperture radar (SAR), and optical-map datasets. Experimental results show that MU-Net achieves more comprehensive and accurate registration performance between these image pairs with geometric and radiometric distortions. We share the code implemented by Pytorch athttps://github.com/yeyuanxin110/MU-Net. Yuanxin Ye, Tengfeng Tang, Bai Zhu, Chao Yang 0028, Bo Li 0090, Siyuan Hao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | A Fast and Robust Matching System for Multimodal Remote Sensing Image RegistrationabstractThe rapid and explosive growth of remote sensing image dataset (e.g., optical, SAR, LiDAR) promotes the development of the aerospace industry. However, images with complex coverage scenes are usually captured by either different sensors from different perspectives or the same sensor in different periods [1]. These factors have brought a great challenge to precision image co-registration, and it is difficult to identify a fully universal method to cope with all registration cases. Any kind of image registration algorithm needs to consider the imaging principle, radiometric and geometric distortions, noise interference, and so on. To date, numerous efforts have been made to overcome these challenges and improve the performance of multimodal remote sensing image registration, which can be classified into three categories: area-based methods, feature-based methods and a joint of previous two categories [2]. Yuanxin Ye, Bai Zhu, Lorenzo Bruzzone |
IGARSS | 1 |
| 2021 | A Novel Keypoint Detector Combining Corners and Blobs for Remote Sensing Image RegistrationabstractKeypoint detection is a crucial step for feature-based image registration. The traditional detectors only extract one type of keypoint such as a corner or a blob, which is not quite beneficial to image registration. Accordingly, this letter presents a novel keypoint detector that aims to simultaneously extract corners and blobs. The proposed detector is named as Harris-Difference of Gaussian (DoG), which combines the advantages of the Harris-Laplace corner detector and the DoG blob detector. In the definition of Harris-DoG, we first build an image scale space and extract the corners by using the multiscale Harris detector. Then, these corners are screened by an automatic scale selection technique based on their DoG responses. This can make the corners robust to scale changes. Meanwhile, DoG also is used to detect the blobs in the image scale space by a nonmaxima suppression scheme. Finally, the scale invariant feature transform (SIFT) descriptors are computed for the detected corners and blobs, and they are applied together for image registration. The proposed Harris-DoG has been tested by using three pairs of multisensor remote sensing images. The experimental results show that Harris-DoG can effectively increase the number of correct matches and improve the registration accuracy compared with the state-of-the-art keypoint detectors. Yuanxin Ye, Mengmeng Wang 0007, Siyuan Hao, Qing Zhu 0012 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2019 | Cross-Domain Collaborative Learning via Cluster Canonical Correlation Analysis and Random Walker for Hyperspectral Image ClassificationabstractThis paper introduces a novel heterogeneous domain adaptation (HDA) method for hyperspectral image (HSI) classification with a limited amount of labeled samples in both domains. The method is achieved in the way of cross-domain collaborative learning (CDCL), which is addressed via cluster canonical correlation analysis (C-CCA) and random walker (RW) algorithms. To be specific, the proposed CDCL method is an iterative process of three main components, i.e., RW-based pseudolabeling, cross-domain learning via C-CCA, and final classification based on extended RW (ERW) algorithm. First, given the initially labeled target samples as the training set (TS), the RW-based pseudolabeling is employed to update TS and extract target clusters (TCs) by fusing the segmentation results obtained by RW and ERW classifiers. Second, cross-domain learning via C-CCA is applied using labeled source samples and TCs. The unlabeled target samples are then classified with the estimated probability maps using the model trained in the projected correlation subspace. The newly estimated probability map and TS are used for updating TS again via RW-based pseudolabeling. Finally, when the iterative process converges, the result obtained by the ERW classifier using the final TS and estimated probability maps is regarded as the final classification map. Experimental results on four real HSIs demonstrate that the proposed method can achieve better performance compared with the state-of-the-art HDA and ERW methods. Yao Qin 0002, Lorenzo Bruzzone, Yuanxin Ye |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2019 | Fast and Robust Matching for Multimodal Remote Sensing Image RegistrationabstractWhile image matching has been studied in remote sensing community for decades, matching multimodal data [e.g., optical, light detection and ranging (LiDAR), synthetic aperture radar (SAR), and map] remains a challenging problem because of significant nonlinear intensity differences between such data. To address this problem, we present a novel fast and robust template matching framework integrating local descriptors for multimodal images. First, a local descriptor [such as histogram of oriented gradient (HOG) and local self-similarity (LSS) or speeded-up robust feature (SURF)] is extracted at each pixel to form a pixelwise feature representation of an image. Then, we define a fast similarity measure based on the feature representation using the fast Fourier transform (FFT) in the frequency domain. A template matching strategy is employed to detect correspondences between images. In this procedure, we also propose a novel pixelwise feature representation using orientated gradients of images, which is named channel features of orientated gradients (CFOG). This novel feature is an extension of the pixelwise HOG descriptor with superior performance in image matching and computational efficiency. The major advantages of the proposed matching framework include: 1) structural similarity representation using the pixelwise feature description and 2) high computational efficiency due to the use of FFT. The proposed matching framework has been evaluated using many different types of multimodal images, and the results demonstrate its superior matching performance with respect to the state-of-the-art methods. Yuanxin Ye, Lorenzo Bruzzone, Jie Shan, Francesca Bovolo, Qing Zhu 0012 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2018 | A Deep Network Architecture for Super-Resolution-Aided Hyperspectral Image Classification With Classwise LossabstractThe supervised deep networks have shown great potential in improving the classification performance. However, training these supervised deep networks is very challenging for hyperspectral image given the fact that usually only a small amount of labeled samples are available. In order to overcome this problem and enhance the discriminative ability of the network, in this paper, we propose a deep network architecture for a super-resolution (SR)-aided hyperspectral image classification with classwise loss (SRCL). First, a three-layer SR convolutional neural network (SRCNN) is employed to reconstruct a high-resolution image from a low-resolution image. Second, an unsupervised triplet-pipeline CNN (TCNN) with an improved classwise loss is built to encourage intraclass similarity and interclass dissimilarity. Finally, SRCNN, TCNN, and a classification module are integrated to define the SRCL, which can be fine-tuned in an end-to-end manner with a small amount of training data. Experimental results on real hyperspectral images demonstrate that the proposed SRCL approach outperforms other state-of-the-art classification methods, especially for the task in which only a small amount of training data are available. Siyuan Hao, Wei Wang 0108, Yuanxin Ye, Enyu Li, Lorenzo Bruzzone |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2018 | Two-Stream Deep Architecture for Hyperspectral Image ClassificationabstractMost traditional approaches classify hyperspectral image (HSI) pixels relying only on the spectral values of the input channels. However, the spatial context around a pixel is also very important and can enhance the classification performance. In order to effectively exploit and fuse both the spatial context and spectral structure, we propose a novel two-stream deep architecture for HSI classification. The proposed method consists of a two-stream architecture and a novel fusion scheme. In the two-stream architecture, one stream employs the stacked denoising autoencoder to encode the spectral values of each input pixel, and the other stream takes as input the corresponding image patch and deep convolutional neural networks are employed to process the image patch. In the fusion scheme, the prediction probabilities from two streams are fused by adaptive class-specific weights, which can be obtained by a fully connected layer. Finally, a weight regularizer is added to the loss function to alleviate the overfitting of the class-specific fusion weights. Experimental results on real HSIs demonstrate that the proposed two-stream deep architecture can achieve competitive performance compared with the state-of-the-art methods. Siyuan Hao, Wei Wang 0108, Yuanxin Ye, Tingyuan Nie, Lorenzo Bruzzone |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2017 | Fast and robust structure-based multimodal geospatial image matchingabstractThis paper presents a fast and robust framework integrating local features for the matching of multimodal geospatial data (e.g., optical, LiDAR, SAR and map). In the proposed framework, local feature descriptors, such as Histogram of Oriented Gradient (HOG) and Local Self Similarity (LSS), are first extracted for every pixel to form a pixel-wise structural feature representation of an image. Then we define a similarity metric based on the feature representation in frequency domain using the 3 Dimensional Fast Fourier Transform (3DFFT) technique, followed by a template matching scheme to detect control points between multimodal data. The proposed framework is based on the hypothesis that structural similarity between images is preserved across different modalities. The major advantages of this framework include (1) structural similarity representation using pixel-wise feature description and (2) high computational efficiency due to the use of 3DFFT. Experimental results on different types of multimodal geospatial data show more accurate matching performance of the proposed framework than the state-of-the-art methods. Yuanxin Ye, Lorenzo Bruzzone, Jie Shan, Li Shen 0004 |
IGARSS | 1 |
| 2017 | Robust Optical-to-SAR Image Matching Based on Shape PropertiesabstractAlthough image matching techniques have been developed in the last decades, automatic optical-to-synthetic aperture radar (SAR) image matching is still a challenging task due to significant nonlinear intensity differences between such images. This letter addresses this problem by proposing a novel similarity metric for image matching using shape properties. A shape descriptor named dense local self-similarity (DLSS) is first developed based on self-similarities within images. Then a similarity metric (named DLSC) is defined using the normalized cross correlation (NCC) of the DLSS descriptors, followed by a template matching strategy to detect correspondences between images. DLSC is robust against significant nonlinear intensity differences because it captures the shape similarity between images, which is independent of intensity patterns. DLSC has been evaluated with four pairs of optical and SAR images. Experimental results demonstrate its advantage over the state-of-the-art similarity metrics (such as NCC and mutual information), and show the superior matching performance. Yuanxin Ye, Li Shen 0004 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2017 | Robust Registration of Multimodal Remote Sensing Images Based on Structural SimilarityabstractAutomatic registration of multimodal remote sensing data [e.g., optical, light detection and ranging (LiDAR), and synthetic aperture radar (SAR)] is a challenging task due to the significant nonlinear radiometric differences between these data. To address this problem, this paper proposes a novel feature descriptor named the histogram of orientated phase congruency (HOPC), which is based on the structural properties of images. Furthermore, a similarity metric named HOPCncc is defined, which uses the normalized correlation coefficient (NCC) of the HOPC descriptors for multimodal registration. In the definition of the proposed similarity metric, we first extend the phase congruency model to generate its orientation representation and use the extended model to build HOPCncc. Then, a fast template matching scheme for this metric is designed to detect the control points between images. The proposed HOPCncc aims to capture the structural similarity between images and has been tested with a variety of optical, LiDAR, SAR, and map data. The results show that HOPCncc is robust against complex nonlinear radiometric differences and outperforms the state-of-the-art similarities metrics (i.e., NCC and mutual information) in matching performance. Moreover, a robust registration method is also proposed in this paper based on HOPCncc, which is evaluated using six pairs of multimodal remote sensing images. The experimental results demonstrate the effectiveness of the proposed method for multimodal image registration. Yuanxin Ye, Jie Shan, Lorenzo Bruzzone, Li Shen 0004 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2015 | Research on relationship between remote sensing image quality and performance of interest point detectionabstractThe paper researches on the relationship between image quality and the performance of interest point detection. In this paper, we use the image quality metrics and interest point repeatability as the measures of image quality and the performance of interest point detection respectively. Considering the differences of image's scene and quality degradation factor, nine images covering three kinds of classic scenes are selected as the experiment data, and they are respectively contaminated by Gaussian blur and Gaussian noise. The results show that image quality metrics and the repeatability have a certain quantitative relationship, and the repeatability is reduced with the decrease of image quality metrics. Moreover, the relationship can be simulated by some simple functions such as linear and exponential model when the image's scene and degradation factor are fixed. Yuanxin Ye, Li Shen 0004, Songbo Wu |
IGARSS | 2 |
| 2015 | Automatic matching of optical and SAR imagery through shape propertyabstractAutomatic matching of optical and SAR images could be challenging due to the significant non-linear intensity differences caused by radiometric variations among such images. To address this problem, this paper utilize the Shape Property to detect the correspondences between images. A new shape descriptor is first built on the basis of the local self-similarity descriptor. Then, the normalized correlation coefficient of the descriptor is used as the similarity metric (named DLSS), which represents the shape similarity among images. Finally, a template-matching strategy is applied to achieve the correspondences between images. The proposed method has been evaluated with four pairs of optical and SAR images. The experimental results demonstrate that the proposed method is robust to non-linear intensity differences, and improves the matching performance compared with the state-of-the-art methods. Yuanxin Ye, Li Shen 0004 |
IGARSS | 1 |