Yongxiang Yao

dblp:50/5670 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
7since 2021 · last 2025
0000-0001-5492-0564ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 CasP: Improving Semi-Dense Feature Matching Pipeline Leveraging Cascaded Correspondence Priors for Guidance
abstract
Semi-dense feature matching methods have shown strong performance in challenging scenarios. However, the existing pipeline relies on a global search across the entire feature map to establish coarse matches, limiting further improvements in accuracy and efficiency. Motivated by this limitation, we propose a novel pipeline, CasP, which leverages cascaded correspondence priors for guidance. Specifically, the matching stage is decomposed into two progressive phases, bridged by a region-based selective cross-attention mechanism designed to enhance feature discriminability. In the second phase, one-to-one matches are determined by restricting the search range to the one-to-many prior areas identified in the first phase. Additionally, this pipeline benefits from incorporating high-level features, which helps reduce the computational costs of low-level feature extraction. The acceleration gains of CasP increase with higher resolution, and our lite model achieves a speedup of $\sim2.2\times$ at a resolution of 1152 compared to the most efficient method, ELoFTR. Furthermore, extensive experiments demonstrate its superiority in geometric estimation, particularly with impressive cross-domain generalization. These advantages highlight its potential for latency-sensitive and high-robustness applications, such as SLAM and UAV systems. Code is available at https://github.com/pq-chen/CasP.
Peiqi Chen, Lei Yu 0005, Yi Wan 0001, Yingying Pei, Xinyi Liu 0002, Yongxiang Yao, Lixiang Ru, Liheng Zhong, Jingdong Chen, Ming Yang 0007, Yongjun Zhang 0002
ICCV6
2025 Topology-Aware Hierarchical Mamba for Salient Object Detection in Remote Sensing Imagery
Wei Yang 0043, Zhiqi Yi, Andong Huang, Ying Wang 0123, Yongxiang Yao, Yansheng Li 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 DDRL: Domain Distribution Reconstruction Learning for Binary Change Detection in Remote Sensing Images
abstract
Change detection (CD) aims to identify and locate changes in the same observed surface coverage area across bitemporal images. This technique has widespread applications in urban planning, land use, and disaster damage extraction. Deep learning-based CD methods typically use learnable encoders to map bitemporal images to a common domain distribution space, allowing for the discrimination and localization of change and invariant features. However, due to differences in imaging mechanisms, seasons, and shooting angles, a large number of pseudochanges may easily appear, affecting the accurate recognition of the domain distribution space. In addition, binary CD focuses solely on whether scene targets have changed, resulting in change labels that encompass a variety of different objects, thus increasing the significance of intraclass differences. To address the aforementioned issues, we propose a domain distribution reconstruction learning (DDRL) framework for binary CD, which effectively mitigates the problem of pseudochanges by detecting abnormal feature domain distributions. Specifically, DDRL first extracts multiscale features from bitemporal images using a Siamese cross-window self-attention module, achieving feature domain transformation from the original space. Subsequently, it employs a graph attention enhanced (GAE) module to improve the low-level domain distribution, enabling it to focus on change regions. In addition, DDRL utilizes a cross-domain feature contrastive learning (CFCL) module for reconstructive learning of high-level fused features. This process ensures that intraclass features are compact, while interclass features are dispersed within the high-level domain distribution, thereby significantly improving the domain distribution representation to discriminate pseudochanges. Experimental results show that the proposed DDRL performs excellently across multiple public datasets, surpassing mainstream methods and significantly improving CD performance. The source code will be made available athttps://github.com/yzygit1230/DDRL.
Wei Yang 0043, Zhaoyi Ye, Liye Mei, Yongxiang Yao, Yansheng Li 0001
IEEE Trans. Geosci. Remote. Sens.4
2025 Multimodal Remote Sensing Image Robust Matching Based on Second-Order Tensor Orientation Feature Transformation
abstract
Nonrigid deformation (NRD) and image noise in multimodal remote sensing images (MRSI) lead to abrupt changes in feature directions, resulting in sensitivity to rotational variation, sparse correct matches, and high false match rates. In order to address these challenges, this article proposes a second-order tensor orientation feature transformation (SOFT) method to improve the rotational invariance of MRSI matching and increase the number of correct matches (NCMs). The SOFT method has two main contributions: 1) a novel second-order tensor orientation descriptor is constructed by generating a tensor orientation feature map using a designed second-order tensor function, which is then combined with a gradient location and orientation histogram (GLOH)-like descriptor framework to achieve robust rotational invariance in multimodal image matching and 2) an error-removal global-local iterative optimization (EGIO) is introduced, employing a skewness of mixed pixel intensity (SMPI) function to automatically select matching seed points, followed by an iterative partition optimization strategy for refining corresponding points. Experiments on 744 groups of typical MRSIs demonstrate that the SOFT method significantly outperforms nine state-of-the-art methods, achieving an average 97% improvement in the NCMs, an average 25.51% improvement in the rate of correct matches (RCMs), and an average reduction in RMSE of 2.69 pixels. The proposed SOFT method, thus, offers robust MRSI matching with strong rotational invariance and precise identification of corresponding points, proving its effectiveness for complex remote sensing scenarios. Access to experiment-related data and codes will be provided athttps://skyearth.org/research/.
Yongjun Zhang 0002, Peihao Wu, Yongxiang Yao, Yi Wan 0001, Wenfei Zhang, Yansheng Li 0001, Xiaohu Yan
IEEE Trans. Geosci. Remote. Sens.3
2024 Adjacent Self-Similarity 3-D Convolution for Multimodal Image Registration
abstract
Significant challenges exist in the registration of multimodal images (MMIs) due to nonlinear radiation differences, variations in lighting, and interference from image noise. These issues often lead to unreliable similarity measurements and low accuracy in point matching during multimodal registration. To address these challenges, this letter introduces a novel MMI registration method based on adjacent self-similarity 3-D convolution (ASTC). The proposed method consists of three main steps: feature point extraction, where key points are uniformly extracted via the block-FAST method; ASTC salient feature construction, where a local adjacent self-similarity (ASS) model is employed to create multidimensional features; and feature structure enhancement, where a 3-D convolution is used for feature enhancement and finishing the process of image feature description. This letter evaluates the ASTC method against six sets of representative MMIs and compares it with six other algorithms. The results demonstrate that: 1) the ASTC algorithm effectively overcomes radiation distortion, intensity differences, and lighting differences in MMIs, leading to improved accuracy in point matching; and 2) the ASTC algorithm achieves higher matching efficiency and reduces time consumption, making it a practical choice for various data types. In summary, the proposed ASTC algorithm offers a robust solution for reliable registration of MMIs, addressing common challenges related to image differences and improving the overall accuracy of the process. The experimental data and code link used in this letter can be found athttps://github.com/yangwill81/ASTC.
Wei Yang 0043, Liye Mei, Zhaoyi Ye, Ying Wang 0123, Xinglong Hu, Yongxiang Yao
IEEE Geosci. Remote. Sens. Lett.7
2023 CloudViT: A Lightweight Vision Transformer Network for Remote Sensing Cloud Detection
abstract
Clouds inevitably exist in satellite images, which limit the processing and application of satellite images to a certain extent. Therefore, cloud detection is a preprocessing task in satellite image extraction and analysis processing. However, the existing methods are difficult to mine robust features, and the number of parameters and computation are large, which is not conducive to the deployment of the model. In this letter, cloud vision transformer (CloudViT), a lightweight vision transformer network for cloud detection from satellite imagery, is proposed. In detail, to utilize dark channel priors in multispectral imagery to guide the network to learn features, a multiscale dark channel extractor is used to first predict dark channels, and then, the dark channel features and image features are input to the attention mechanism-based dark channel-guided context aggregation module to enhance image features, which in turn makes cloud detection results more accurate. At the same time, to enhance the transfer ability of the network between different satellite sensors, a plug-and-play channel adaptive module is proposed to deal with the inconsistency of the number of different satellite sensor bands. The experimental results on the Landsat7 dataset show that our network CloudViT outperforms the state-of-the-art methods while keeping the number of parameters and computation small. At the same time, the experimental results on transfer to three other datasets show that using the channel adaptation module can greatly improve the transfer ability of the model.
Bin Zhang 0046, Yongjun Zhang 0002, Yansheng Li 0001, Yi Wan 0001, Yongxiang Yao
IEEE Geosci. Remote. Sens. Lett.5
2022 Multi-Modal Remote Sensing Image Matching Considering Co-Occurrence Filter
abstract
Traditional image feature matching methods cannot obtain satisfactory results for multi-modal remote sensing images (MRSIs) in most cases because different imaging mechanisms bring significant nonlinear radiation distortion differences (NRD) and complicated geometric distortion. The key to MRSI matching is trying to weakening or eliminating the NRD and extract more edge features. This paper introduces a new robust MRSI matching method based on co-occurrence filter (CoF) space matching (CoFSM). Our algorithm has three steps: (1) a new co-occurrence scale space based on CoF is constructed, and the feature points in the new scale space are extracted by the optimized image gradient; (2) the gradient location and orientation histogram algorithm is used to construct a 152-dimensional log-polar descriptor, which makes the multi-modal image description more robust; and (3) a position-optimized Euclidean distance function is established, which is used to calculate the displacement error of the feature points in the horizontal and vertical directions to optimize the matching distance function. The optimization results then are rematched, and the outliers are eliminated using a fast sample consensus algorithm. We performed comparison experiments on our CoFSM method with the scale-invariant feature transform (SIFT), upright-SIFT, PSO-SIFT, and radiation-variation insensitive feature transform (RIFT) methods using a multi-modal image dataset. The algorithms of each method were comprehensively evaluated both qualitatively and quantitatively. Our experimental results show that our proposed CoFSM method can obtain satisfactory results both in the number of corresponding points and the accuracy of its root mean square error. The average number of obtained matches is namely 489.52 of CoFSM, and 412.52 of RIFT. As mentioned earlier, the matching effect of the proposed method was significantly greater than the three state-of-art methods. Our proposed CoFSM method achieved good effectiveness and robustness. Executable programs of CoFSM and MRSI datasets are published: https://skyearth.org/publication/project/CoFSM/.
Yongxiang Yao, Yongjun Zhang 0002, Yi Wan 0001, Xinyi Liu 0002, Xiaohu Yan, Jiayuan Li 0001
IEEE Trans. Image Process.1