Xiaoyong Lu

dblp:20/6907 · DBLP profile ↗
← Back
13ranked-venue papers
9as first author
9since 2021 · last 2026
0009-0003-5379-0034ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 6 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author
YearPublicationVenuePosition
2026 Toward Free-Form Local Feature Matching
abstract
Existing feature matching methods are strongly coupled to their pre-defined position priors. For instance, sparse matchers are coupled to keypoints, and semi-dense matchers are coupled to grids. The coupled position prior dictates the distribution of matching points and imposes inherent limitations on the matcher. Consequently, sparse matchers suffer from a reliance on keypoint repeatability, while semi-dense matchers lack texture-based precision. Our preliminary work RCM leverages the keypoint prior in the source image and the grid prior in the target image, ensuring texture-based precision with keypoints while eliminating reliance on repeatability. However, RCM still relies heavily on keypoints in the source image, inheriting limitations such as sparsity and poor distribution in challenging scenes. To address these challenges, we introduce RCM+, which presents a novel free-form matching paradigm. By combining a position-agnostic encoder with a parameter-free decoder, we decouple the matcher from any position prior. As a result, the free-form matcher can match arbitrary input positions in a zero-shot manner, including detected keypoints, lines, edges, grids of any resolution, user-specified points, and more. This paradigm offers exceptional flexibility, allowing users to select position priors based on scene properties without retraining. Thus, RCM+ can leverage the advantages of various position priors without over-relying on any single prior, avoiding limitations in specific scenarios. To better match multiple position priors, we propose the Balancer, which reconciles all input position priors to achieve a more favorable point distribution for downstream tasks. Additionally, we enhance the view switcher and conflict-free matching layer introduced in RCM, further improving matching quality. Comprehensive experiments demonstrate the excellent performance, efficiency, and flexibility of RCM+, underscoring its promising potential for applications.
Xiaoyong Lu, Songlin Du, Yaping Yan, Xiaobo Lu, Takeshi Ikenaga
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 Parallel consensus transformer for local feature matching
Xiaoyong Lu, Bin Kang, Songlin Du
Pattern Recognit.1
2026 SceneGlue: Scene-Aware Transformer for Feature Matching Without Scene-Level Annotation
abstract
Local feature matching plays a critical role in understanding the correspondence between cross-view images. However, traditional methods are constrained by the inherent local nature of feature descriptors, limiting their ability to capture non-local scene information that is essential for accurate cross-view correspondence. In this paper, we introduce SceneGlue, a scene-aware feature matching framework designed to overcome these limitations. SceneGlue leverages a hybridizable matching paradigm that integrates implicit parallel attention and explicit cross-view visibility estimation. The parallel attention mechanism simultaneously exchanges information among local descriptors within and across images, enhancing the scene’s global context. To further enrich the scene awareness, we propose the Visibility Transformer, which explicitly categorizes features into visible and invisible regions, providing an understanding of cross-view scene visibility. By combining explicit and implicit scene-level awareness, SceneGlue effectively compensates for the local descriptor constraints. Notably, SceneGlue is trained using only local feature matches, without requiring scene-level groundtruth annotations. This scene-aware approach not only improves accuracy and robustness but also enhances interpretability compared to traditional methods. Extensive experiments on applications such as homography estimation, pose estimation, image matching, and visual localization validate SceneGlues superior performance. The source code is available at https://github.com/songlindu/ SceneGlue.
Songlin Du, Xiaoyong Lu, Yaping Yan, Guobao Xiao, Xiaobo Lu, Takeshi Ikenaga
IEEE Trans. Circuits Syst. Video Technol.2
2025 JamMa: Ultra-lightweight Local Feature Matching with Joint Mamba
abstract
Existing state-of-the-art feature matchers capture long-range dependencies with Transformers but are hindered by high spatial complexity, leading to demanding training and high-latency inference. Striking a better balance between performance and efficiency remains a challenge in feature matching. Inspired by the linear complexity $\mathcal{O}(N)$ of Mamba, we propose an ultra-lightweight Mamba-based matcher, named JamMa, which converges on a single GPU and achieves an impressive performance-efficiency balance in inference. To unlock the potential of Mamba for feature matching, we propose Joint Mamba with a scan-merge strategy named JEGO, which enables: (1) Joint scan of two images to achieve high-frequency mutual interaction, (2) Efficient scan with skip steps to reduce sequence length, (3) Global receptive field, and (4) Omnidirectional feature representation. With the above properties, the JEGO strategy significantly outperforms the scan-merge strategies proposed in VMamba and EVMamba in the feature matching task. Compared to attention-based sparse and semi-dense matchers, JamMa demonstrates a superior balance between performance and efficiency, delivering better performance with less than 50% of the parameters and FLOPs. Project page: https://leoluxxx.github.io/JamMa-page/.
Xiaoyong Lu, Songlin Du
CVPR1
2025 SemMatcher: Semantic-aware feature matching with neighborhood consensus
Qimin Jiang, Xiaoyong Lu, Dong Liang 0008, Songlin Du
J. Vis. Commun. Image Represent.2
2024 Raising the Ceiling: Conflict-Free Local Feature Matching with Dynamic View Switching
Xiaoyong Lu, Songlin Du
ECCV (42)1
2023 ParaFormer: Parallel Attention Transformer for Efficient Feature Matching
abstract
Heavy computation is a bottleneck limiting deep-learning-based feature matching algorithms to be applied in many real-time applications. However, existing lightweight networks optimized for Euclidean data cannot address classical feature matching tasks, since sparse keypoint based descriptors are expected to be matched. This paper tackles this problem and proposes two concepts: 1) a novel parallel attention model entitled ParaFormer and 2) a graph based U-Net architecture with attentional pooling. First, ParaFormer fuses features and keypoint positions through the concept of amplitude and phase, and integrates self- and cross-attention in a parallel manner which achieves a win-win performance in terms of accuracy and efficiency. Second, with U-Net architecture and proposed attentional pooling, the ParaFormer-U variant significantly reduces computational complexity, and minimize performance loss caused by downsampling. Sufficient experiments on various applications, including homography estimation, pose estimation, and image matching, demonstrate that ParaFormer achieves state-of-the-art performance while maintaining high efficiency. The efficient ParaFormer-U variant achieves comparable performance with less than 50% FLOPs of the existing attention-based models.
Xiaoyong Lu, Yaping Yan, Bin Kang, Songlin Du
AAAI1
2023 Scene-Aware Feature Matching
abstract
Current feature matching methods focus on point-level matching, pursuing better representation learning of individual features, but lacking further understanding of the scene. This results in significant performance degradation when handling challenging scenes such as scenes with large viewpoint and illumination changes. To tackle this problem, we propose a novel model named SAM, which applies attentional grouping to guide Scene-Aware feature Matching. SAM handles multi-level features, i.e., image tokens and group tokens, with attention layers, and groups the image tokens with the proposed token grouping module. Our model can be trained by ground-truth matches only and produce reasonable grouping results. With the sense-aware grouping guidance, SAM is not only more accurate and robust but also more interpretable than conventional feature matching models. Sufficient experiments on various applications, including homography estimation, pose estimation, and image matching, demonstrate that our model achieves state-of-the-art performance.
Xiaoyong Lu, Yaping Yan, Songlin Du
ICCV1
2022 NCTR: Neighborhood Consensus Transformer for Feature Matching
abstract
This paper presents NCTR, a feature matcher that enhances input descriptors and finds the correspondences between them. In NCTR, Transformer is applied to aggregate global context for each descriptor. To solve the lack of neighborhood consensus that Transformer may bring, we propose a novel method to evaluate the neighborhood consensus of each key-point and integrate it into the attentional aggregation. The combination of global and local information greatly enhances the model capability and improves the match quality. The experiments on homography estimation and outdoor pose estimation show that NCTR outperforms other hand-designed or learning-based methods and achieves state-of-the-art results.
Xiaoyong Lu, Songlin Du
ICIP1
2020 Review of Depression Recognition Based on Speech
Xiaoyong Lu, Daimin Shi
ICCE1
2020 A New Method Design for Multi-modal Depression Auxiliary Diagnosis from Perspective of Psychology
Xiaoyong Lu, Daimin Shi
ICCE1
2016 A Novel Graph Partitioning Criterion Based Short Text Clustering Method
Xiaohong Li 0012, Tingnian He, Hongyan Ran, Xiaoyong Lu
ICIC (3)4
2016 Effectively Classifying Short Texts via Improved Lexical Category and Semantic Features
Huifang Ma, Runan Zhou, Xiaoyong Lu
ICIC (1)4