EDBT 2026 Demo / reviewers in the wild / expert
Jiaxuan Chen 0002
dblp:64/10174-2
· DBLP profile ↗
10ranked-venue papers
6as first author
10since 2021 · last 2025
0000-0002-4387-9671ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MindGPT: Interpreting What You See With Non-Invasive Brain RecordingsabstractDecoding of seen visual contents with non-invasive brain recordings has important scientific and practical values. Efforts have been made to recover the seen images from brain signals. However, most existing approaches cannot faithfully reflect the visual contents due to insufficient image quality or semantic mismatches. Compared with reconstructing pixel-level visual images, speaking is a more efficient and effective way to explain visual information. Here we introduce a non-invasive neural decoder, termed MindGPT, which interprets perceived visual stimuli into natural languages from functional Magnetic Resonance Imaging (fMRI) signals in an end-to-end manner. Specifically, our model builds upon a visually guided neural encoder with a cross-attention mechanism. By the collaborative use of data augmentation techniques, this architecture permits us to guide latent neural representations towards a desired language semantic direction in a self-supervised fashion. Through doing so, we found that the neural representations of the MindGPT are explainable, which can be used to evaluate the contributions of visual properties to language semantics. Our experiments show that the generated word sequences truthfully represented the visual information (with essential details) conveyed in the seen stimuli. The results also suggested that with respect to language decoding tasks, the higher visual cortex (HVC) is more semantically informative than the lower visual cortex (LVC), and using only the HVC can recover most of the semantic information. The source code for the MindGPT model is publicly available at https://github.com/JxuanC/MindGPT. Jiaxuan Chen 0002, Yueming Wang 0001, Gang Pan 0001 |
IEEE Trans. Image Process. | 1 |
| 2024 | CSR-Net++: Rethinking Context Structure Representation Learning for Feature MatchingabstractSeeking good feature correspondences between two remote sensing (RS) images is an essential and important problem in the fields of RS and photogrammetry. Traditional approaches often necessitate a predefined geometric transformation model or additional manually crafted descriptors, significantly constraining the versatility. In this work, we adopt the recent context structure representation network (CSR-Net), which has shown promising performance in general feature matching problems, and propose modifications, named CSR-Net++, to overcome its main limitations. Specifically, CSR-Net is combined with a PointNet-like geometry estimator, which is sensitive to large deformations, for global preregistration. In addition, CSR-Net learns local consensus representation through a fixed-size grid, leading to limited space-aware capacities due to grid pixelwise max-pooling operations. To tackle the abovementioned limitations, we first introduce a pruning layer for matching guided by global consensus, as opposed to relying on a geometric estimator. In addition, for directly learning consensus representation from points, we propose a modified context structure representation (CSR) learning module including an independent spatial location stream and a stand-alone visual stream (VS). This decomposition separates local consensus into positional consensus and visual consensus. The proposed dual-stream representation learning not only avoids the introduction of grid anchors but also provides visual contextual priors. To demonstrate the robustness and versatility of our CSR-Net++, we conducted comprehensive experiments using diverse sets of real image pairs for general feature matching. The results demonstrate the superiority of our CSR-Net++ in most matching scenarios, achieving a 0.47%–4.70% improvement in F-score for multimodal images over existing leading methods. Xiaoxian Chen, Jiaxuan Chen 0002 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | StateNet: Deep State Learning for Robust Feature Matching of Remote Sensing ImagesabstractSeeking good correspondences between two images is a fundamental and challenging problem in the remote sensing (RS) community, and it is a critical prerequisite in a wide range of feature-based visual tasks. In this article, we propose a flexible and general deep state learning network for both rigid and nonrigid feature matching, which provides a mechanism to change the state of matches into latent canonical forms, thereby weakening the degree of randomness in matching patterns. Different from the current conventional strategies (i.e., imposing a global geometric constraint or designing additional handcrafted descriptor), the proposed StateNet is designed to perform alternating two steps: 1) recalibrates matchwise feature responses in the spatial domain and 2) leverages the spatially local correlation across two sets of feature points for transformation update. For this purpose, our network contains two novel operations: adaptive dual-aggregation convolution (ADAConv) and point rendering layer (PRL). These two operations are differentiable, so our network can be inserted into the existing classification architecture to reduce the cost of establishing reliable correspondences. To demonstrate the robustness and universality of our approach, extensive experiments on various real image pairs for feature matching are conducted. Experiments reveal the superiority of our StateNet significantly over the state-of-the-art alternatives. Jiaxuan Chen 0002, Shuang Chen 0008, Xiaoxian Chen, Yang Yang 0032, Yujing Rao |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | VLSG-SANet: A feature matching algorithm for remote sensing image registration
Linjie Xing, Jiaxuan Chen 0002, Shuang Chen 0008, Haicheng Bai, Lin Xing, Chengjiang Zhou, Yang Yang 0032 |
Knowl. Based Syst. | 3 |
| 2022 | Robust Feature Matching via Hierarchical Local Structure VisualizationabstractFeature matching, which refers to seeking good correspondence between two feature point sets, is a critical prerequisite in many applications of remote sensing and photogrammetry. This work can be viewed as an extension of LSV-ANet. Traditional local structure visualization (LSV) descriptor is very sensitive to outliers existing in the small region around a feature point, which limits the ability of LSV-ANet to recognize fine-grained patterns and its generalizability for complex matching scenes. Thus, in this letter, we propose a hierarchical local structure visualization (HLSV) to solve this issue. Specifically, local structures of feature points are first decoupled by hierarchical tensor and then reconstructed in a learning fashion, which explicitly permits us to learn a nonmutually exclusive relationship for enhancement structure manipulation. In addition, in order to further improve the network generalization ability, we design a permutation-invariant network layer to ensure that the output of model is invariant to the input order. To demonstrate the robustness of the HLSV, extensive experiments on various real remote sensing image pairs for feature matching are conducted. The experiment results reveal that HLSV is superior to the current eight state-of-the-art alternatives. Jiaxuan Chen 0002, Shuang Chen 0008, Yang Yang 0032, Haicheng Bai |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | LSV-ANet: Deep Learning on Local Structure Visualization for Feature MatchingabstractFeature matching is a fundamental and important task in many applications of remote sensing and photogrammetry. Remote sensing images often involve complex spatial relationships due to the ground relief variations and imaging viewpoint changes. Therefore, using a pre-defined geometrical model will probably lead to inferior matching accuracy. In order to find good correspondences, we propose a simple yet efficient deep learning network, which we term the “local structure visualization-attention” network (LSV-ANet). Our main aim is to transform outlier detection into a dynamic visual similarity evaluation. Specifically, we first map the local spatial distribution into a regular grid as descriptor LSV, and then customized a spatial SCale Attention (SCA) module and a spatial STructure Attention (STA) module, which explicitly allows structure manipulation and scale selection of LSV within the network. Finally, the embedded SCA and STA deduce optimal LSV for solving feature matching task by training the LSV-ANet end-to-end. In order to demonstrate the robustness and universality of our LSV-ANet, extensive experiments on various real image pairs for general feature matching are conducted and compared against eight state-of-the-art methods. The experiment results demonstrate the superiority of our method over state of the art. Jiaxuan Chen 0002, Shuang Chen 0008, Xiaoxian Chen, Yang Yang 0032, Linjie Xing, Yujing Rao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | IGS-Net: Seeking Good Correspondences via Interactive Generative Structure LearningabstractFeature matching, which aims to seek good correspondences from an image pair of the same or similar scene, is one of the important studies on digital remote sensing (RS) image processing. However, alongside the common degradation problems, such as geometric distortion, RS images also often face nonlinear radiation distortions, thereby posing more complex matching patterns. To reduce the cost of establishing reliable correspondences, this article proposes a simple, but effective end-to-end hierarchical learning framework, termed interactive generative structure learning network (IGS-Net). The key thinking of our approach is to offer a structure self-generate learning mechanism, called interactive generative structure learning (IGSL) block, for modeling the local context information of potential correspondences. Specifically, IGSL contains two novel operations: adaptive structure-aware representation (ASR) and physical constraint embedding. Besides, we introduce a coarse-to-fine geometry estimation pipeline aligning two sets of feature points to weaken the degree of randomness in matching patterns, thus improving the generalization ability of representation learning. Overall, this differentiable representation learning architecture can be inserted into existing classification models easily for robust outlier detection and removal. In order to demonstrate that our IGS-Net can boost the baselines, we intensively experiment on both single modality and multimodal RS image datasets. The large amounts of experiment results reveal that the matching performances of IGS-Net are significantly improved over eight state-of-the-art competitors. Jiaxuan Chen 0002, Shuang Chen 0008, Xiaoxian Chen, Yujing Rao, Chengjiang Zhou, Yang Yang 0032 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | A Hierarchical Consensus Attention Network for Feature Matching of Remote Sensing ImagesabstractFeature matching, referring to establishing high reliable correspondences between two or more scenes with overlapping regions, is of extremely significance to various remote sensing (RS) tasks, such as panorama mosaic and change detection. In this work, we propose an end-to-end deep network for mismatch removal, named hierarchical consensus attention network (HCA-Net), which is one of the critical steps in matching pipeline. Unlike existing practices, our HCA-Net does not rely on global geometric constraints and handcrafted structural representations. The key principle of the proposed HCA-Net is to adaptively enhance neighborhood consensus before evaluating correspondence. To this end, we design a consensus attention mechanism to regularize sparse matches directly. More specifically, consensus attention consists of two novel operations: an encoder–decoder module for calculating compatibility scores and a context-based density representation module. Such attention mechanism can be easily plugged into the existing inlier/outlier classification model in a stacked way to reject outliers. We also propose a hierarchical global-aware network for further improving the accuracy of outlier detection. We compare the proposed HCA-Net with seven state-of-the-art algorithms on several datasets (including various RS images), and the results reveal that our method significantly outperforms the other competitors. Shuang Chen 0008, Jiaxuan Chen 0002, Yujing Rao, Xiaoxian Chen, Haicheng Bai, Lin Xing, Chengjiang Zhou, Yang Yang 0032 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Learning Relaxed Neighborhood Consistency for Feature MatchingabstractFeature matching is a critical prerequisite in many applications of remote sensing, and its aim is to establish reliable correspondences between two sets of features. Existing attempts typically involve estimating the underlying image transformations to remove false matches in putative matches. However, the image transformation could vary with different application scenarios, which means that using a predefined geometrical model may lead to inferior matching accuracy, especially if the image transformation is nonrigid. This article casts the mismatch removal into a neighborhood consistency evaluation problem under a customized learning framework. With only seven training image pairs involving approximately 8000 putative matches, our method can handle different types of images or transformation models (affine, homography, piecewise-linear transformation, and others). Extensive experiments on feature matching and image registration are conducted to demonstrate the superiority of our method over the eight state-of-the-art competitors. Shuang Chen 0008, Jiaxuan Chen 0002, Zenghui Xiong, Linjie Xing, Yang Yang 0032, Kai Yan 0002, Hao Li 0093 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | CSR-Net: Learning Adaptive Context Structure Representation for Robust Feature CorrespondenceabstractFeature matching, which refers to identifying and then corresponding the same or similar visual pattern from two or more images, is a key technique in any image processing task that requires establishing good correspondences between images. Given potential correspondences (matches) in two scenes, a novel whole-part deep learning framework, termed as Context Structure Representation Network (CSR-Net), is designed to infer the probabilities of arbitrary correspondences being inliers. Traditional approaches commonly build the local relation between correspondences by manually engineered criteria. Different from existing attempts, the main idea of our work is to learn explicitly neighborhood structure of each correspondence, allowing us to formulate the matching problem into a dynamic local structure consensus evaluation in an end-to-end fashion. For this purpose, we propose a permutation-invariant STructure Representation (STR) learning module, which can easily merge different types of networks into a unified architecture to deal with sparse matches directly. By the collaborative use of STR, we introduce a Context-Aware Attention (CAA) mechanism to adaptively re-calibrate structure features via a rotation-invariant context aware encoding and simple feature gating, thus arising the ability of fine-grained patterns recognition. Moreover, to further weaken the cost of establishing reliable correspondences, the CSR-Net is formulated as whole-part consensus learning, where the aim of whole level is compensating rigid transformations. In order to demonstrate our CSR-Net can effectively boost the baselines, we intensively experiment on image matching and other visual tasks. The results of the experiment confirm that the matching performances of CSR-Net have significantly improved over nine state-of-the-art competitors. Jiaxuan Chen 0002, Shuang Chen 0008, Xiaoxian Chen, Yuan Dai, Yang Yang 0032 |
IEEE Trans. Image Process. | 1 |