Xiaoxian Chen

dblp:157/9659 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
7since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2024 CSR-Net++: Rethinking Context Structure Representation Learning for Feature Matching
abstract
Seeking good feature correspondences between two remote sensing (RS) images is an essential and important problem in the fields of RS and photogrammetry. Traditional approaches often necessitate a predefined geometric transformation model or additional manually crafted descriptors, significantly constraining the versatility. In this work, we adopt the recent context structure representation network (CSR-Net), which has shown promising performance in general feature matching problems, and propose modifications, named CSR-Net++, to overcome its main limitations. Specifically, CSR-Net is combined with a PointNet-like geometry estimator, which is sensitive to large deformations, for global preregistration. In addition, CSR-Net learns local consensus representation through a fixed-size grid, leading to limited space-aware capacities due to grid pixelwise max-pooling operations. To tackle the abovementioned limitations, we first introduce a pruning layer for matching guided by global consensus, as opposed to relying on a geometric estimator. In addition, for directly learning consensus representation from points, we propose a modified context structure representation (CSR) learning module including an independent spatial location stream and a stand-alone visual stream (VS). This decomposition separates local consensus into positional consensus and visual consensus. The proposed dual-stream representation learning not only avoids the introduction of grid anchors but also provides visual contextual priors. To demonstrate the robustness and versatility of our CSR-Net++, we conducted comprehensive experiments using diverse sets of real image pairs for general feature matching. The results demonstrate the superiority of our CSR-Net++ in most matching scenarios, achieving a 0.47%–4.70% improvement in F-score for multimodal images over existing leading methods.
Xiaoxian Chen, Jiaxuan Chen 0002
IEEE Trans. Geosci. Remote. Sens.1
2023 StateNet: Deep State Learning for Robust Feature Matching of Remote Sensing Images
abstract
Seeking good correspondences between two images is a fundamental and challenging problem in the remote sensing (RS) community, and it is a critical prerequisite in a wide range of feature-based visual tasks. In this article, we propose a flexible and general deep state learning network for both rigid and nonrigid feature matching, which provides a mechanism to change the state of matches into latent canonical forms, thereby weakening the degree of randomness in matching patterns. Different from the current conventional strategies (i.e., imposing a global geometric constraint or designing additional handcrafted descriptor), the proposed StateNet is designed to perform alternating two steps: 1) recalibrates matchwise feature responses in the spatial domain and 2) leverages the spatially local correlation across two sets of feature points for transformation update. For this purpose, our network contains two novel operations: adaptive dual-aggregation convolution (ADAConv) and point rendering layer (PRL). These two operations are differentiable, so our network can be inserted into the existing classification architecture to reduce the cost of establishing reliable correspondences. To demonstrate the robustness and universality of our approach, extensive experiments on various real image pairs for feature matching are conducted. Experiments reveal the superiority of our StateNet significantly over the state-of-the-art alternatives.
Jiaxuan Chen 0002, Shuang Chen 0008, Xiaoxian Chen, Yang Yang 0032, Yujing Rao
IEEE Trans. Neural Networks Learn. Syst.3
2022 LSV-ANet: Deep Learning on Local Structure Visualization for Feature Matching
abstract
Feature matching is a fundamental and important task in many applications of remote sensing and photogrammetry. Remote sensing images often involve complex spatial relationships due to the ground relief variations and imaging viewpoint changes. Therefore, using a pre-defined geometrical model will probably lead to inferior matching accuracy. In order to find good correspondences, we propose a simple yet efficient deep learning network, which we term the “local structure visualization-attention” network (LSV-ANet). Our main aim is to transform outlier detection into a dynamic visual similarity evaluation. Specifically, we first map the local spatial distribution into a regular grid as descriptor LSV, and then customized a spatial SCale Attention (SCA) module and a spatial STructure Attention (STA) module, which explicitly allows structure manipulation and scale selection of LSV within the network. Finally, the embedded SCA and STA deduce optimal LSV for solving feature matching task by training the LSV-ANet end-to-end. In order to demonstrate the robustness and universality of our LSV-ANet, extensive experiments on various real image pairs for general feature matching are conducted and compared against eight state-of-the-art methods. The experiment results demonstrate the superiority of our method over state of the art.
Jiaxuan Chen 0002, Shuang Chen 0008, Xiaoxian Chen, Yang Yang 0032, Linjie Xing, Yujing Rao
IEEE Trans. Geosci. Remote. Sens.3
2022 IGS-Net: Seeking Good Correspondences via Interactive Generative Structure Learning
abstract
Feature matching, which aims to seek good correspondences from an image pair of the same or similar scene, is one of the important studies on digital remote sensing (RS) image processing. However, alongside the common degradation problems, such as geometric distortion, RS images also often face nonlinear radiation distortions, thereby posing more complex matching patterns. To reduce the cost of establishing reliable correspondences, this article proposes a simple, but effective end-to-end hierarchical learning framework, termed interactive generative structure learning network (IGS-Net). The key thinking of our approach is to offer a structure self-generate learning mechanism, called interactive generative structure learning (IGSL) block, for modeling the local context information of potential correspondences. Specifically, IGSL contains two novel operations: adaptive structure-aware representation (ASR) and physical constraint embedding. Besides, we introduce a coarse-to-fine geometry estimation pipeline aligning two sets of feature points to weaken the degree of randomness in matching patterns, thus improving the generalization ability of representation learning. Overall, this differentiable representation learning architecture can be inserted into existing classification models easily for robust outlier detection and removal. In order to demonstrate that our IGS-Net can boost the baselines, we intensively experiment on both single modality and multimodal RS image datasets. The large amounts of experiment results reveal that the matching performances of IGS-Net are significantly improved over eight state-of-the-art competitors.
Jiaxuan Chen 0002, Shuang Chen 0008, Xiaoxian Chen, Yujing Rao, Chengjiang Zhou, Yang Yang 0032
IEEE Trans. Geosci. Remote. Sens.4
2022 A Hierarchical Consensus Attention Network for Feature Matching of Remote Sensing Images
abstract
Feature matching, referring to establishing high reliable correspondences between two or more scenes with overlapping regions, is of extremely significance to various remote sensing (RS) tasks, such as panorama mosaic and change detection. In this work, we propose an end-to-end deep network for mismatch removal, named hierarchical consensus attention network (HCA-Net), which is one of the critical steps in matching pipeline. Unlike existing practices, our HCA-Net does not rely on global geometric constraints and handcrafted structural representations. The key principle of the proposed HCA-Net is to adaptively enhance neighborhood consensus before evaluating correspondence. To this end, we design a consensus attention mechanism to regularize sparse matches directly. More specifically, consensus attention consists of two novel operations: an encoder–decoder module for calculating compatibility scores and a context-based density representation module. Such attention mechanism can be easily plugged into the existing inlier/outlier classification model in a stacked way to reject outliers. We also propose a hierarchical global-aware network for further improving the accuracy of outlier detection. We compare the proposed HCA-Net with seven state-of-the-art algorithms on several datasets (including various RS images), and the results reveal that our method significantly outperforms the other competitors.
Shuang Chen 0008, Jiaxuan Chen 0002, Yujing Rao, Xiaoxian Chen, Haicheng Bai, Lin Xing, Chengjiang Zhou, Yang Yang 0032
IEEE Trans. Geosci. Remote. Sens.4
2022 CSR-Net: Learning Adaptive Context Structure Representation for Robust Feature Correspondence
abstract
Feature matching, which refers to identifying and then corresponding the same or similar visual pattern from two or more images, is a key technique in any image processing task that requires establishing good correspondences between images. Given potential correspondences (matches) in two scenes, a novel whole-part deep learning framework, termed as Context Structure Representation Network (CSR-Net), is designed to infer the probabilities of arbitrary correspondences being inliers. Traditional approaches commonly build the local relation between correspondences by manually engineered criteria. Different from existing attempts, the main idea of our work is to learn explicitly neighborhood structure of each correspondence, allowing us to formulate the matching problem into a dynamic local structure consensus evaluation in an end-to-end fashion. For this purpose, we propose a permutation-invariant STructure Representation (STR) learning module, which can easily merge different types of networks into a unified architecture to deal with sparse matches directly. By the collaborative use of STR, we introduce a Context-Aware Attention (CAA) mechanism to adaptively re-calibrate structure features via a rotation-invariant context aware encoding and simple feature gating, thus arising the ability of fine-grained patterns recognition. Moreover, to further weaken the cost of establishing reliable correspondences, the CSR-Net is formulated as whole-part consensus learning, where the aim of whole level is compensating rigid transformations. In order to demonstrate our CSR-Net can effectively boost the baselines, we intensively experiment on image matching and other visual tasks. The results of the experiment confirm that the matching performances of CSR-Net have significantly improved over nine state-of-the-art competitors.
Jiaxuan Chen 0002, Shuang Chen 0008, Xiaoxian Chen, Yuan Dai, Yang Yang 0032
IEEE Trans. Image Process.3
2021 Landslide Detection of High-Resolution Satellite Images using Asymmetric Dual-Channel Network
abstract
Landslide detection from high-resolution satellite imagery plays a significant role in disaster management. Recently, deep learning has emerged as one of the most powerful tools for landslide detection. However, the existing deep learning models for landslide detection still have room for improvement. In this paper, we propose a novel deep learning model named DCA-Net for the automatic detection of landslides. This model adopts an asymmetric encoder-decoder structure. The core module designed in the DCA-Net is the dual-channel depthwise block, integrating two parallel channels of asymmetric depthwise separable convolutions and residual connections to enlarge the receptive field and improve the performance. Experiments are conducted on the open-source Bijie landslide dataset. Several state-of-the-art networks are also employed for quantitative and qualitative comparisons. The results indicate that our proposed DCA-Net is superior to other deep learning models, which can be well utilized in landslide detection from aerial images.
Yaohui Liu 0001, Xiaoxian Chen, Mingyang Yu 0007, Yingjun Sun, Xiwei Fan
IGARSS3