EDBT 2026 Demo / reviewers in the wild / expert
Xingyu Jiang 0005
dblp:23/4062-5
· DBLP profile ↗
21ranked-venue papers
7as first author
15since 2021 · last 2025
0000-0001-9790-8856ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 3 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MINIMA: Modality Invariant Image MatchingabstractImage matching for both cross-view and cross-modality plays a critical role in multimodal perception. In practice, the modality gap caused by different imaging systems/styles poses great challenges to the matching task. Existing works try to extract invariant features for specific modalities and train on limited datasets, showing poor generalization. In this paper, we present MINIMA, a unified image matching framework for multiple cross-modal cases. Without pursuing fancy modules, our MINIMA aims to enhance universal performance from the perspective of data scaling up. For such purpose, we propose a simple yet effective data engine that can freely produce a large dataset containing multiple modalities, rich scenarios, and accurate matching labels. Specifically, we scale up the modalities from cheap but rich RGB-only matching data, by means of generative models. Under this setting, the matching labels and rich diversity of the RGB dataset are well inherited by the generated multimodal data. Benefiting from this, we construct MD-syn, a new comprehensive dataset that fills the data gap for general multimodal image matching. With MD-syn, we can directly train any advanced matching pipeline on randomly selected modality pairs to obtain cross-modal ability. Extensive experiments on in-domain and zero-shot matching tasks, including 19 cross-modal cases, demonstrate that our MINIMA can significantly outperform the baselines and even surpass modality-specific methods. The dataset and code are available at https://github.com/LSXI7/MINIMA. Jiangwei Ren, Xingyu Jiang 0005, Zizhuo Li, Dingkang Liang, Xin Zhou 0013, Xiang Bai |
CVPR | 2 |
| 2025 | NAUTILUS: A Large Multimodal Model for Underwater Scene UnderstandingabstractUnderwater exploration offers critical insights into our planet and attracts increasing attention for its broader applications in resource exploration, national security, etc. We study the underwater scene understanding methods, which aim to achieve automated underwater exploration. The underwater scene understanding task demands multi-task perceptions from multiple granularities. However, the absence of large-scale underwater multi-task instruction-tuning datasets hinders the progress of this research. To bridge this gap, we construct NautData, a dataset containing 1.45 M image-text pairs supporting eight underwater scene understanding tasks. It enables the development and thorough evaluation of the underwater scene understanding models. Underwater image degradation is a widely recognized challenge that interferes with underwater tasks. To improve the robustness of underwater scene understanding, we introduce physical priors derived from underwater imaging models and propose a plug-and-play vision feature enhancement (VFE) module, which explicitly restores clear underwater information. We integrate this module into renowned baselines LLaVA-1.5 and Qwen2.5-VL and build our underwater LMM, NAUTILUS. Experiments conducted on the NautData and public underwater datasets demonstrate the effectiveness of the VFE module, consistently improving the performance of both baselines on the majority of supported tasks, thus ensuring the superiority of NAUTILUS in the underwater scene understanding area. Data and models are available at https://github.com/H-EmbodVis/NAUTILUS. Wei Xu 0037, Cheng Wang 0048, Dingkang Liang, Zongchuang Zhao, Xingyu Jiang 0005, Xiang Bai |
NeurIPS | 5 |
| 2025 | AVS-Net: Point sampling with adaptive voxel size for 3D scene understanding
Hongcheng Yang, Dingkang Liang, Dingyuan Zhang, Zhe Liu 0033, Zhikang Zou, Xingyu Jiang 0005, Yingying Zhu 0005 |
Neurocomputing | 6 |
| 2025 | Layerlink: Bridging remote sensing object detection and large vision models with efficient fine-tuning
Xingkui Zhu, Dingkang Liang, Xingyu Jiang 0005, Yiran Guan, Yingying Zhu 0005, Xiang Bai |
Pattern Recognit. | 3 |
| 2023 | Robust Model Reasoning and Fitting via Dual Sparsity PursuitabstractIn this paper, we contribute to solving a threefold problem: outlier rejection, true model reasoning and parameter estimation with a unified optimization modeling. To this end, we first pose this task as a sparse subspace recovering problem, to search a maximum of independent bases under an over-embedded data space. Then we convert the objective into a continuous optimization paradigm that estimates sparse solutions for both bases and errors. Wherein a fast and robust solver is proposed to accurately estimate the sparse subspace parameters and error entries, which is implemented by a proximal approximation method under the alternating optimization framework with the ``optimal'' sub-gradient descent. Extensive experiments regarding known and unknown model fitting on synthetic and challenging real datasets have demonstrated the superiority of our method against the state-of-the-art. We also apply our method to multi-class multi-model fitting and loop closure detection, and achieve promising results both in accuracy and efficiency. Code is released at: https://github.com/StaRainJ/DSP. Xingyu Jiang 0005, Jiayi Ma 0001 |
NeurIPS | 1 |
| 2023 | Improving sparse graph attention for feature matching by informative keypoints exploration
Xingyu Jiang 0005, Xiao-Ping Zhang 0002, Jiayi Ma 0001 |
Comput. Vis. Image Underst. | 1 |
| 2023 | Smoothness-Driven Consensus Based on Compact Representation for Robust Feature MatchingabstractFor robust feature matching, a popular and particularly effective method is to recover smooth functions from the data to differentiate the true correspondences (inliers) from false correspondences (outliers). In the existing works, the well-established regularization theory has been extensively studied and exploited to estimate the functions while controlling its complexity to enforce the smoothness constraint, which has shown prominent advantages in this task. However, despite the theoretical optimality properties, the high complexities in both time and space are induced and become the main obstacle of their application. In this article, we propose a novel method for multivariate regression and point matching, which exploits the sparsity structure of smooth functions. Specifically, we use compact Fourier bases for constructing the function, which inherently allows a coarse-to-fine representation. The smoothness constraint can be explicitly imposed by adopting a few low-frequency bases for representation, resulting in reduced computational complexities of the induced multivariate regression algorithm. To cope with potential gross outliers, we formulate the learning problem into a Bayesian framework with latent variables indicating the inliers and outliers and a mixture model accounting for the distribution of data, where a fast expectation-maximization solution can be derived. Extensive experiments are conducted on synthetic data and real-world image matching, and point set registration datasets, which demonstrates the advantages of our method against the current state-of-the-art methods in terms of both scalability and robustness. Aoxiang Fan, Xingyu Jiang 0005, Yong Ma 0001, Xiaoguang Mei, Jiayi Ma 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Feature Matching via Motion-Consistency Driven Probabilistic Graphical Model
Jiayi Ma 0001, Aoxiang Fan, Xingyu Jiang 0005, Guobao Xiao |
Int. J. Comput. Vis. | 3 |
| 2022 | Efficient Deterministic Search With Robust Loss Functions for Geometric Model FittingabstractGeometric model fitting is a fundamental task in computer vision, which serves as the pre-requisite of many downstream applications. While the problem has a simple intrinsic structure where the solution can be parameterized within a few degrees of freedom, the ubiquitously existing outliers are the main challenge. In previous studies, random sampling techniques have been established as the practical choice, since optimization-based methods are usually too time-demanding. This prospective study is intended to design efficient algorithms that benefit from a general optimization-based view. In particular, two important types of loss functions are discussed, \emph{i.e.} truncated and$l_1$losses, and efficient solvers have been derived for both upon specific approximations. Based on this philosophy, a class of algorithms are introduced to perform deterministic search for the inliers or geometric model. Recommendations are made based on theoretical and experimental analyses. Compared with the existing solutions, the proposed methods are both simple in computation and robust to outliers. Extensive experiments are conducted on publicly available datasets for geometric estimation, which demonstrate the superiority of our methods compared with the state-of-the-art ones. Additionally, we apply our method to the recent benchmark for wide-baseline stereo evaluation, leading to a significant improvement of performance. Aoxiang Fan, Jiayi Ma 0001, Xingyu Jiang 0005, Haibin Ling |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Robust image matching via local graph structure consensus
Xingyu Jiang 0005, Xiao-Ping Zhang 0002, Jiayi Ma 0001 |
Pattern Recognit. | 1 |
| 2022 | Appearance-Based Loop Closure Detection via Locality-Driven Accurate Motion Field LearningabstractLoop closure detection (LCD) is of significant importance in simultaneous localization and mapping. It represents the robot’s ability to recognize whether the current surrounding corresponds to a previously observed one. In this paper, we conduct this task in a two-step strategy: candidate frame selection and loop closure verification. The first step aims to search semantically similar images for the query one using features obtained by Key.Net with HardNet. Instead of adopting the traditional Bag-of-Words strategy, we utilize the aggregated selective match kernel to calculate the similarity between images. Subsequently, based on the potential property of motion field in the LCD scene, we propose a novel feature matching method,i.e., exploiting the smoothness prior and learning the motion field for an image pair in a reproducing kernel Hilbert space (RKHS), to implement loop closure verification. Concretely, we formulate the learning problem into a Bayesian framework with latent variables indicating the true/false correspondences and a mixture model accounting for the distribution of data. Furthermore, we propose a locality-driven mechanism to enhance the local relevance of motion vectors and term the algorithm as locality-driven accurate motion field learning (LAL). To satisfy the requirement of efficiency in the LCD task, we use a sparse approximation and search a suboptimal solution for the motion field in the RKHS, termed as LAL*. Extensive experiments are conducted on public datasets for feature matching and LCD tasks. The quantitative results demonstrate the effectiveness of our method over the current state-of-the-art, meanwhile showing its potential for long-term visual localization. The codes of LAL and LAL* are publicly available athttps://github.com/KN-Zhang/LAL. Kaining Zhang, Xingyu Jiang 0005, Jiayi Ma 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Motion Field Consensus with Locality Preservation: A Geometric Confirmation Strategy for Loop Closure DetectionabstractLoop closure detection (LCD), which aims to deal with the drift emerging when robots travel around the route, plays a key role in a simultaneous localization and mapping system. Unlike most current methods which focus on seeking an appropriate representation of images, we propose a novel two-stage pipeline dominated by the estimation of spatial geometric relationship. When a query image occurs, we select semantically similar images based on the SuperPoint network and the aggregated selective match kernel in the first stage, and then conduct robust geometric confirmation to verify true loop-closing pairs in the second stage. Based on the potential property of motion field in the LCD scene, a robust feature matching algorithm, termed as motion field consensus with locality preservation (MFC-LP), is proposed. In particular, we exploit the smoothness prior to guide the learning of the motion field for an image pair in a reproducing kernel Hilbert space (RKHS). Meanwhile, to enhance the local relevance of motion vectors, we design a locality preservation mechanism thus making the learned motion field more accurate. Extensive experiments on several publicly available datasets reveal that MFC-LP has a good performance in the general feature matching task and the proposed pipeline outperforms the current state-of-the-art approaches in the LCD task. Kaining Zhang, Xingyu Jiang 0005, Xiaoguang Mei, Huabing Zhou, Jiayi Ma 0001 |
IROS | 2 |
| 2021 | Image Matching from Handcrafted to Deep Features: A SurveyabstractAbstract As a fundamental and critical task in various visual applications, image matching can identify then correspond the same or similar structure/content from two or more images. Over the past decades, growing amount and diversity of methods have been proposed for image matching, particularly with the development of deep learning techniques over the recent years. However, it may leave several open questions about which method would be a suitable choice for specific applications with respect to different scenarios and task requirements and how to design better image matching methods with superior performance in accuracy, robustness and efficiency. This encourages us to conduct a comprehensive and systematic review and analysis for those classical and latest techniques. Following the feature-based image matching pipeline, we first introduce feature detection, description, and matching techniques from handcrafted methods to trainable ones and provide an analysis of the development of these methods in theory and practice. Secondly, we briefly introduce several typical image matching-based applications for a comprehensive understanding of the significance of image matching. In addition, we also provide a comprehensive and objective comparison of these classical and latest techniques through extensive experiments on representative datasets. Finally, we conclude with the current status of image matching technologies and deliver insightful discussions and prospects for future works. This survey can serve as a reference for (but not limited to) researchers and engineers in image matching and related fields. Jiayi Ma 0001, Xingyu Jiang 0005, Aoxiang Fan, Junjun Jiang, Junchi Yan |
Int. J. Comput. Vis. | 2 |
| 2021 | Ranking list preservation for feature matching
Junjun Jiang, Xingyu Jiang 0005, Jiayi Ma 0001 |
Pattern Recognit. | 3 |
| 2021 | Robust Feature Matching for Remote Sensing Image Registration via Linear Adaptive FilteringabstractAs a fundamental and critical task in feature-based remote sensing image registration, feature matching refers to establishing reliable point correspondences from two images of the same scene. In this article, we propose a simple yet efficient method termed linear adaptive filtering (LAF) for both rigid and nonrigid feature matching of remote sensing images and apply it to the image registration task. Our algorithm starts with establishing putative feature correspondences based on local descriptors and then focuses on removing outliers using geometrical consistency priori together with filtering and denoising theory. Specifically, we first grid the correspondence space into several nonoverlapping cells and calculate a typical motion vector for each one. Subsequently, we remove false matches by checking the consistency between each putative match and the typical motion vector in the corresponding cell, which is achieved by a Gaussian kernel convolution operation. By refining the typical motion vector in an iterative manner, we further introduce a progressive strategy based on the coarse-to-fine theory to promote the matching accuracy gradually. In addition, an adaptive parameter setting strategy and posterior probability estimation based on the expectation-maximization algorithm enhance the robustness of our method to different data. Most importantly, our method is quite efficient where the gridding strategy enables it to achieve linear time complexity. Consequently, some sparse point-based tasks may inspire from our method when they are achieved by deep learning techniques. Extensive feature matching and image registration experiments on several remote sensing data sets demonstrate the superiority of our approach over the state of the art. Xingyu Jiang 0005, Jiayi Ma 0001, Aoxiang Fan, Haiping Xu, Geng Lin, Tao Lu 0001, Xin Tian 0006 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | Geometric Estimation via Robust Subspace Recovery
Aoxiang Fan, Xingyu Jiang 0005, Junjun Jiang, Jiayi Ma 0001 |
ECCV (22) | 2 |
| 2020 | Robust Feature Matching Using Spatial Clustering With Heavy OutliersabstractThis paper focuses on removing mismatches from given putative feature matches created typically based on descriptor similarity. To achieve this goal, existing attempts usually involve estimating the image transformation under a geometrical constraint, where a pre-defined transformation model is demanded. This severely limits the applicability, as the transformation could vary with different data and is complex and hard to model in many real-world tasks. From a novel perspective, this paper casts the feature matching into a spatial clustering problem with outliers. The main idea is to adaptively cluster the putative matches into several motion consistent clusters together with an outlier/mismatch cluster. To implement the spatial clustering, we customize the classic density based spatial clustering method of applications with noise (DBSCAN) in the context of feature matching, which enables our approach to achieve quasi-linear time complexity. We also design an iterative clustering strategy to promote the matching performance in case of severely degraded data. Extensive experiments on several datasets involving different types of image transformations demonstrate the superiority of our approach over state-of-the-art alternatives. Our approach is also applied to near-duplicate image retrieval and co-segmentation and achieves promising performance. Xingyu Jiang 0005, Jiayi Ma 0001, Junjun Jiang, Xiaojie Guo 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | Progressive Filtering for Feature MatchingabstractIn this paper, we propose a simple yet efficient method termed as Progressive Filtering for Feature Matching, which is able to establish accurate correspondences between two images of common or similar scenes. Our algorithm first grids the correspondence space and calculates a typical motion vector for each cell, and then removes false matches by checking the consistency between each putative match and the typical motion vector in the corresponding cell, which is achieved by a convolution operation. By refining the typical motion vector in an iterative manner, we further introduce a progressive matching strategy based on the coarse-to-fine theory to promote the matching accuracy gradually. The density estimation is utilized to address the island samples and accelerate the convergency of the mismatch removal procedure. In addition, our method is quite efficient where the gridding strategy enables it to achieve linear time complexity. Extensive experiments on several representative real images involving different types of geometric transformations demonstrate the superiority of our approach over the state-of-the-art. Xingyu Jiang 0005, Jiayi Ma 0001, Jun Chen 0019 |
ICASSP | 1 |
| 2019 | Feature-guided Gaussian mixture model for image matching
Jiayi Ma 0001, Xingyu Jiang 0005, Junjun Jiang, Yuan Gao 0015 |
Pattern Recognit. | 2 |
| 2019 | Multiscale Locality and Rank Preservation for Robust Feature Matching of Remote Sensing ImagesabstractAs a fundamental and important task in many applications of remote sensing and photogrammetry, feature matching tries to seek correspondences between the two feature sets extracted from an image pair of the same object or scene. This paper focuses on eliminating mismatches from a set of putative feature correspondences constructed according to the similarity of existing well-designed feature descriptors. Considering the stable local topological relationship of the potential true correspondences, we propose a simple yet efficient method named multiscale Top K Rank Preservation (mTopKRP) for robust feature matching. To this end, we first search the K-nearest neighbors of each feature point and generate a ranking list accordingly. Then we design a metric based on the weighted Spearman's footrule distance to describe the similarity of two ranking lists specifically for the matching problem. We build a mathematical optimization model and derive its closed-form solution, enabling our method to establish reliable correspondences in linearithmic time complexity, which requires only tens of milliseconds to handle over 1000 putative matches. We also introduce a multiscale strategy for neighborhood construction, which increases the robustness of our method and can deal with different types of degradation, even when the image pair suffers from a large scale change, rotation, nonrigid deformation, or a large number of mismatches. Extensive experiments on several representative remote sensing image data sets demonstrate the superiority of our method over state of the art. Xingyu Jiang 0005, Junjun Jiang, Aoxiang Fan, Zhongyuan Wang 0001, Jiayi Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | LMR: Learning a Two-Class Classifier for Mismatch RemovalabstractFeature matching, which refers to establishing reliable correspondence between two sets of features, is a critical prerequisite in a wide spectrum of vision-based tasks. Existing attempts typically involve the mismatch removal from a set of putative matches based on estimating the underlying image transformation. However, the transformation could vary with different data. Thus, a pre-defined transformation model is often demanded, which severely limits the applicability. From a novel perspective, this paper casts the mismatch removal into a two-class classification problem, learning a general classifier to determine the correctness of an arbitrary putative match, termed as Learning for Mismatch Removal (LMR). The classifier is trained based on a general match representation associated with each putative match through exploiting the consensus of local neighborhood structures based on a multiple K -nearest neighbors strategy. With only ten training image pairs involving about 8000 putative matches, the learned classifier can generate promising matching results in linearithmic time complexity on arbitrary testing data. The generality and robustness of our approach are verified under several representative supervised learning techniques as well as on different training and testing data. Extensive experiments on feature matching, visual homing, and near-duplicate image retrieval are conducted to reveal the superiority of our LMR over the state-of-the-art competitors. Jiayi Ma 0001, Xingyu Jiang 0005, Junjun Jiang, Ji Zhao 0001, Xiaojie Guo 0001 |
IEEE Trans. Image Process. | 2 |