EDBT 2026 Demo / reviewers in the wild / expert
Kaining Zhang
dblp:249/2561
· DBLP profile ↗
23ranked-venue papers
10as first author
23since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 5 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 6 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SGPFeat: Semantic and Geometric Priors for Multi-modal Image MatchingabstractMulti-modal image matching is a fundamental task in multi-view and multi-modal image processing. Its key challenge lies in extracting features that remain consistent despite drastic appearance variations across modalities. However, the learning of the feature is hindered by the scarcity and the inaccurate alignment of existing multi-modal datasets. To address this, we propose a knowledge distillation framework termed SGPFeat that transfers rich prior knowledge from large-scale unimodal tasks to enhance multi-modal representation learning. Specifically, semantic priors from a vision foundation model guide the feature extractor to identify shared semantic structures across modalities, enabling better generalization under large appearance gaps. In parallel, geometric priors derived from accurately aligned visible-light datasets improve detection precision on noisy aligned multi-modal pairs. Furthermore, we introduce a Heterogeneous Feature Aggregation (HFA) module to facilitate effective distillation and feature representation. Extensive experiments demonstrate that semantic and geometric priors bring significant improvement for our SGPFeat across diverse multi-modal image matching benchmarks. Yuxin Deng 0002, Botian Wang, Kaining Zhang, Hao Zhang 0073, Jiayi Ma 0001 |
AAAI | 3 |
| 2026 | MGLD-TLNet: Multigeometric and Long-Distance Representation Network for Transmission Line InspectionabstractEffective transmission line (TL) inspection in complex corridor environments is essential for ensuring reliable power delivery. This work presents a 3D-based perception method for this task. The proposed method is designed by considering two key characteristics of TL inspection. First, the point cloud data are sparse and class distributions are highly imbalanced, which weakens the signals from thin conductors and tower components. To address this issue, we model long-range spatial relations along the corridor to mitigate data sparsity and imbalance. Second, strong structural correlations exist between conductors and towers, which can be leveraged to improve perception performance. To exploit this property, we construct a unified 3-D representation that jointly models towers, conductors, and vegetation, while fusing Cartesian and polar geometries through geometry-aware alignment. Experiments on real-world corridor datasets demonstrate that the proposed method, termed multigeometric and long-distance TL perception Network (MGLD-TLNet), consistently improves stability and accuracy under conditions of sparsity, occlusion, and complex environmental interactions. Hui Zhang 0023, Kaining Zhang, Baheti Biekezat, Hang Zhong, Junfei Yi, Jianxu Mao, Yaonan Wang 0001 |
IEEE Trans. Cybern. | 3 |
| 2026 | FPF: A Focused Perception Framework for Small Defect Identification in Complex Power Scenarios
Hui Zhang 0023, Baheti Biekezat, Yunkang Cao, Kaining Zhang, Tongzhi Niu, Yaonan Wang 0001 |
IEEE Trans. Ind. Informatics | 6 |
| 2026 | SAF: A Structure-Aware Framework for Radial Ice Thickness Detection on Overhead Transmission LinesabstractIce thickness estimation on overhead transmission lines (OHTL) is essential for mitigating icing-induced mechanical failures and ensuring safe grid operation. To address the challenges of detecting radial ice thickness in complex power line corridors, particularly geometric fragmentation of slender conductors and semantic ambiguity near occluded boundaries, this work proposes a structure-aware framework (SAF) based on 3-D point cloud segmentation and geometry-guided modeling. SAF introduces a structure-aware segmentation network, which integrates a cross-level spatial encoding module to preserve geometric continuity and a partition-aware loss to improve boundary localization under vegetation or tower occlusion. Building on accurate segmentation, a geometry-guided module performs centerline fitting and cross-sectional reconstruction to infer slice-level ice thickness. To support evaluation, a large-scale uncrewed aerial vehicle (UAV)-based point cloud dataset covering 32 OHTL is constructed, including six lines with ground-truth ice labels. Experimental results demonstrate that SAF achieves robust and accurate ice estimation across varied voltage levels and terrains, supporting its practical application in intelligent transmission line inspection and icing risk prevention. Hui Zhang 0023, Youyuan Tang, Yihong Cao, Kaining Zhang, Yunkang Cao, Tongzhi Niu, Jianxu Mao, Yaonan Wang 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2025 | The Brain Knows What You Prefer: Using EEG to Decode AR Input Preferences
Kaining Zhang, Theophilus Teo, Eunhee Chang, Xianglin Zheng, Allison Jing, Mark Billinghurst |
CHI | 1 |
| 2025 | Adapting Dense Matching for Homography Estimation with Grid-based AccelerationabstractCurrent deep homography estimation methods are typically constrained to processing low-resolution image pairs due to network architecture and computational limitations. For high-resolution images, downsampling is often required, which can greatly degrade estimation accuracy. In contrast, image matching methods, which match pixels and compute homography from correspondences, provide greater resolution flexibility. So in this work, we revisit the traditional image matching paradigm for homography estimation and propose GFNet, a Grid Flow regression Network that adapts the high-accuracy dense matching framework for homography estimation while enhancing efficiency through a grid-based strategy—estimating flow only over a coarse grid by leveraging homography’s global smoothness. We demonstrate the effectiveness of GFNet on a wide range of experiments on multiple datasets, including the common scene MSCOCO, multimodal datasets VIS-IR and GoogleMap, and the dynamic scene VIRAT. Notably, on 448×448 GoogleMap, GFNet achieves an improvement of +13.5% in auc@3 while reducing MACs by ~47% compared to the SOTA dense matching method. Additionally, it shows a 1.8× improvement in auc@3 over the SOTA deep homography method. Code is available at https://github.com/KN-Zhang/GFNet. Kaining Zhang, Yuxin Deng 0002, Jiayi Ma 0001, Paolo Favaro |
CVPR | 1 |
| 2025 | ArgMatch: Adaptive Refinement Gathering for Efficient Dense Matching
Yuxin Deng 0002, Kaining Zhang, Linfeng Tang, Jiaqi Yang 0002, Jiayi Ma 0001 |
ICCV | 2 |
| 2025 | Headzoom: Hands-Free Zooming and Panning for 2D Image Navigation Using Head MotionabstractWe introduce HeadZoom, a hands-free interaction technique for navigating two-dimensional visual content using head movements. HeadZoom enables fluid zooming and panning using only real-time head tracking. It supports natural control in applications such as map exploration, radiograph inspection, and image browsing, where physical interaction is limited. We evaluated HeadZoom in a withinsubjects study comparing three interaction techniques-Static, Tilt Zoom, and Parallel Zoom-across spatial, error, and subjective metrics. Parallel Zoom significantly reduced total head movement compared to Static and Tilt modes. Users reported significantly lower perceived exertion for Parallel Zoom, confirming its suitability for prolonged or precision-based tasks. By minimizing movement demands while maintaining task effectiveness, HeadZoom advances the design of head-based 2D interaction in VR and creates new opportunities for accessible hands-free systems for image exploration. Kaining Zhang, Catarina Moreira, Pedro Belchior, Gun A. Lee, Mark Billinghurst, Joaquim Jorge 0001 |
ISMAR | 1 |
| 2025 | TITAN: A Trajectory-Informed Technique for Adaptive Parameter Freezing in Large-Scale VQEabstractVariational quantum Eigensolver (VQE) is a leading candidate for harnessing quantum computers to advance quantum chemistry and materials simulations, yet its training efficiency deteriorates rapidly for large Hamiltonians. Two issues underlie this bottleneck: (i) the no-cloning theorem imposes a linear growth in circuit evaluations with the number of parameters per gradient step; and (ii) deeper circuits encounter barren plateaus (BPs), leading to exponentially increasing measurement overheads. To address these challenges, here we propose a deep learning framework, dubbed Titan, which identifies and freezes inactive parameters of a given ansätze at initialization for a specific class of Hamiltonians, reducing the optimization overhead without sacrificing accuracy. The motivation of Titan starts with our empirical findings that a subset of parameters consistently has negligible influence on training dynamics. Its design combines a theoretically grounded data construction strategy, ensuring each training example is informative and BP-resilient, with an adaptive neural architecture that generalizes across ansätze of varying sizes. Across benchmark transverse-field Ising models, Heisenberg models, and multiple molecule systems up to $30$ qubits, Titan achieves up to $3\times$ faster convergence and $40$–$60\%$ fewer circuit evaluations than state-of-the-art baselines, while matching or surpassing their estimation accuracy. By proactively trimming parameter space, Titan lowers hardware demands and offers a scalable path toward utilizing VQE to advance practical quantum chemistry and materials science. Yifeng Peng, Samuel Yen-Chi Chen, Kaining Zhang, Zhiding Liang |
NeurIPS | 4 |
| 2025 | AiDE-Q: Synthetic Labeled Datasets Can Enhance Learning Models for Quantum Property EstimationabstractQuantum many-body problems are central to various scientific disciplines, yet their ground-state properties are intrinsically challenging to estimate. Recent advances in deep learning (DL) offer potential solutions in this field, complementing prior purely classical and quantum approaches. However, existing DL-based models typically assume access to a large-scale and noiseless labeled dataset collected by infinite sampling. This idealization raises fundamental concerns about their practical utility, especially given the limited availability of quantum hardware in the near term. To unleash the power of these DL-based models, we propose AiDE-Q (\underline{a}utomat\underline{i}c \underline{d}ata \underline{e}ngine for \underline{q}uantum property estimation), an effective framework that addresses this challenge by iteratively generating high-quality synthetic labeled datasets. Specifically, AiDE-Q utilizes a confidence-check method to assess the quality of synthetic labels and continuously improves the employed DL models with the identified high-quality synthetic dataset. To verify the effectiveness of AiDE-Q, we conduct extensive numerical simulations on a diverse set of quantum many-body and molecular systems, with up to 50 qubits. The results show that AiDE-Q enhances prediction performance for various reference learning models, with improvements of up to $14.2\\%$. Moreover, we exhibit that a basic supervised learning model integrated with AiDE-Q outperforms advanced reference models, highlighting the importance of a synthetic dataset. Our work paves the way for more efficient and practical applications of DL for quantum property estimation. Xinbiao Wang, Zihan Lou, Kaining Zhang, Yong Luo 0002, Bo Du 0001, Dacheng Tao |
NeurIPS | 5 |
| 2024 | ResMatch: Residual Attention Learning for Feature MatchingabstractAttention-based graph neural networks have made great progress in feature matching. However, the literature lacks a comprehensive understanding of how the attention mechanism operates for feature matching. In this paper, we rethink cross- and self-attention from the viewpoint of traditional feature matching and filtering. To facilitate the learning of matching and filtering, we incorporate the similarity of descriptors into cross-attention and relative positions into self-attention. In this way, the attention can concentrate on learning residual matching and filtering functions with reference to the basic functions of measuring visual and spatial correlation. Moreover, we leverage descriptor similarity and relative positions to extract inter- and intra-neighbors. Then sparse attention for each point can be performed only within its neighborhoods to acquire higher computation efficiency. Extensive experiments, including feature matching, pose estimation and visual localization, confirm the superiority of the proposed method. Our codes are available at https://github.com/ACuOoOoO/ResMatch. Yuxin Deng 0002, Kaining Zhang, Yansheng Li 0001, Jiayi Ma 0001 |
AAAI | 2 |
| 2024 | Sparse-to-dense Multimodal Image Registration via Multi-Task LearningabstractAligning image pairs captured by different sensors or those undergoing significant appearance changes is crucial for various computer vision and robotics applications. Existing approaches cope with this problem via either Sparse feature Matching (SM) or Dense direct Alignment (DA) paradigms. Sparse methods are efficient but lack accuracy in textureless scenes, while dense ones are more accurate in all scenes but demand for good initialization. In this paper, we propose SDME, a Sparse-to-Dense Multimodal feature Extractor based on a novel multi-task network that simultaneously predicts SM and DA features for robust multimodal image registration. We propose the sparse-to-dense registration paradigm: we first perform initial registration via SM and then refine the result via DA. By using the well-designed SDME, the sparse-to-dense approach combines the merits from both SM and DA. Extensive experiments on MSCOCO, GoogleEarth, VIS-NIR and VIS-IR-drone datasets demonstrate that our method achieves remarkable performance on multimodal cases. Furthermore, our approach exhibits robust generalization capabilities, enabling the fine-tuning of models initially trained on single-modal datasets for use with smaller multimodal datasets. Our code is available at https://github.com/KN-Zhang/SDME. Kaining Zhang, Jiayi Ma 0001 |
ICML | 1 |
| 2024 | Quantum Imitation LearningabstractDespite remarkable successes in solving various complex decision-making tasks, training an imitation learning (IL) algorithm with deep neural networks (DNNs) suffers from the high-computational burden. In this work, we propose quantum IL (QIL) with a hope to utilize quantum advantage to speed up IL. Concretely, we develop two QIL algorithms: quantum behavioral cloning (Q-BC) and quantum generative adversarial IL (Q-GAIL). Q-BC is trained with a negative log-likelihood (NLL) loss in an offline manner that suits extensive expert data cases, whereas Q-GAIL works in an inverse reinforcement learning (IRL) scheme, which is online, on-policy, and is suitable for limited expert data cases. For both QIL algorithms, we adopt variational quantum circuits (VQCs) in place of DNNs for representing policies, which are modified with data reuploading and scaling parameters to enhance the expressivity. We first encode classical data into quantum states as inputs, then perform VQCs, and finally measure quantum outputs to obtain control signals of agents. Experiment results demonstrate that both Q-BC and Q-GAIL can achieve comparable performance compared to classical counterparts, with the potential of quantum speedup. To our knowledge, we are the first to propose the concept of QIL and conduct pilot studies, which paves the way for the quantum era. Zhihao Cheng, Kaining Zhang, Li Shen 0008, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Offline Quantum Reinforcement Learning in a Conservative MannerabstractRecently, to reap the quantum advantage, empowering reinforcement learning (RL) with quantum computing has attracted much attention, which is dubbed as quantum RL (QRL). However, current QRL algorithms employ an online learning scheme, i.e., the policy that is run on a quantum computer needs to interact with the environment to collect experiences, which could be expensive and dangerous for practical applications. In this paper, we aim to solve this problem in an offline learning manner. To be more specific, we develop the first offline quantum RL (offline QRL) algorithm named CQ2L (Conservative Quantum Q-learning), which learns from offline samples and does not require any interaction with the environment. CQ2L utilizes variational quantum circuits (VQCs), which are improved with data re-uploading and scaling parameters, to represent Q-value functions of agents. To suppress the overestimation of Q-values resulting from offline data, we first employ a double Q-learning framework to reduce the overestimation bias; then a penalty term that encourages generating conservative Q-values is designed. We conduct abundant experiments to demonstrate that the proposed method CQ2L can successfully solve offline QRL tasks that the online counterpart could not. Zhihao Cheng, Kaining Zhang, Li Shen 0008, Dacheng Tao |
AAAI | 2 |
| 2023 | Approximate k-Nearest Neighbor Query over Spatial Data Federation
Kaining Zhang, Yongxin Tong, Yexuan Shi, Yuxiang Zeng, Yi Xu 0013, Lei Chen 0002, Zimu Zhou, Ke Xu 0001, Weifeng Lv, Zhiming Zheng 0001 |
DASFAA (1) | 1 |
| 2023 | Recent Advances for Quantum Neural Networks in Generative LearningabstractQuantum computers are next-generation devices that hold promise to perform calculations beyond the reach of classical computers. A leading method towards achieving this goal is through quantum machine learning, especially quantum generative learning. Due to the intrinsic probabilistic nature of quantum mechanics, it is reasonable to postulate that quantum generative learning models (QGLMs) may surpass their classical counterparts. As such, QGLMs are receiving growing attention from the quantum physics and computer science communities, where various QGLMs that can be efficiently implemented on near-term quantum machines with potential computational advantages are proposed. In this paper, we review the current progress of QGLMs from the perspective of machine learning. Particularly, we interpret these QGLMs, covering quantum circuit Born machines, quantum generative adversarial networks, quantum Boltzmann machines, and quantum variational autoencoders, as the quantum extension of classical generative learning models. In this context, we explore their intrinsic relations and their fundamental differences. We further summarize the potential applications of QGLMs in both conventional machine learning tasks and quantum physics. Last, we discuss the challenges and further research directions for QGLMs. Jinkai Tian, Shanshan Zhao 0001, Qing Liu 0027, Kaining Zhang, Wanrong Huang, Xingyao Wu, Min-Hsiu Hsieh, Tongliang Liu, Wenjing Yang 0002, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Loop Closure Detection With Bidirectional Manifold Representation ConsensusabstractLoop closure detection (LCD) is an indispensable module in simultaneous localization and mapping. It is responsible to recognize pre-visited areas during the navigation of a robot, providing auxiliary information to revise pose estimation. Unlike most current methods which focus on seeking an appropriate representation of images, we propose a novel two-stage pipeline dominated by the estimation of spatial geometric relationship. Specifically, to avoid unnecessary memory costs, consecutive images are segmented into sequences as per the similarity of their global features. Then the sequence descriptor is incrementally inserted into hierarchical navigable small world for the construction of reference database, from which the most similar image for the query one is searched parallelly. To further identify whether the candidate pair is geometry-consistent, a feature matching method termed as bidirectional manifold representation consensus (BMRC) is proposed. It constructs local neighborhood structures of feature points via manifold representation, and formulates the matching problem into an optimization model, enabling linearithmic time complexity via a closed-form solution. Meanwhile, an accelerated version of it is introduced (BMRC*), which performs about 63% faster than BMRC in an image pair with 352 initial correspondences. Extensive experiments on nine publicly available datasets demonstrate that BMRC and BMRC* perform well in feature matching and the proposed pipeline has remarkable performance in the LCD task. Kaining Zhang, Zizhuo Li, Jiayi Ma 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | Escaping from the Barren Plateau via Gaussian Initializations in Deep Variational Quantum CircuitsabstractVariational quantum circuits have been widely employed in quantum simulation and quantum machine learning in recent years. However, quantum circuits with random structures have poor trainability due to the exponentially vanishing gradient with respect to the circuit depth and the qubit number. This result leads to a general standpoint that deep quantum circuits would not be feasible for practical tasks. In this work, we propose an initialization strategy with theoretical guarantees for the vanishing gradient problem in general deep quantum circuits. Specifically, we prove that under proper Gaussian initialized parameters, the norm of the gradient decays at most polynomially when the qubit number and the circuit depth increase. Our theoretical results hold for both the local and the global observable cases, where the latter was believed to have vanishing gradients even for very shallow circuits. Experimental results verify our theoretical findings in quantum simulation and quantum chemistry. Kaining Zhang, Liu Liu 0014, Min-Hsiu Hsieh, Dacheng Tao |
NeurIPS | 1 |
| 2022 | Cross Fusion Net: A Fast Semantic Segmentation Network for Small-Scale Semantic Information Capturing in Aerial ScenesabstractCapturing accurate multiscale semantic information from the images is of great importance for high-quality semantic segmentation. Over the past years, a large number of methods attempt to improve the multiscale information capturing ability of the networks via various means. However, these methods always suffer unsatisfactory efficiency (e.g., speed or accuracy) on the images that include a large number of small-scale objects, for example, aerial images. In this article, we propose a new network named cross fusion net (CF-Net) for fast and effective extraction of the multiscale semantic information, especially for small-scale semantic information. In particular, the proposed CF-Net can capture more accurate small-scale semantic information from two aspects. On the one hand, we develop a channel attention refinement block to select the informative features. On the other hand, we propose a cross fusion block to enlarge the receptive field of the low-level feature maps. As a result, the network can encode more accurate semantic information from the small-scale objects, and the segmentation accuracy of the small-scale objects is improved accordingly. We have compared the proposed CF-Net with several state-of-the-art semantic segmentation methods on two popular aerial image segmentation data sets. Experimental results reveal that the average$F_{1}$score gain brought by our CF-Net is about 0.43% and the$F_{1}$score gain of the small-scale objects (e.g., cars) is about 2.61%. In addition, our CF-Net has the fastest inference speed, which proves its superiority in the aerial scenes. Our code will be released at:https://github.com/pcl111/CF-Net. Chengli Peng, Kaining Zhang, Yong Ma 0001, Jiayi Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Fast and Robust Loop-Closure Detection via Convolutional Auto-Encoder and Motion ConsensusabstractLoop-closure detection is an indispensable module in the visual simultaneous localization and mapping (vSLAM) system. It typically consists of three main steps: image representation, loop-closure candidate selection, and loop-closure event verification. This article proposes a novel approach for loop-closure detection. In particular, we first introduce a lightweight convolutional auto-encoder network trained by the deep perceptual similarity loss for image representation. We then propose an image-to-sequence selection approach based on place sequence division and distance-weighted voting for loop-closure candidate selection. Furthermore, we propose a motion vector consensus constraint to improve locality preserving matching, which can be used for efficient loop-closure event verification that is robust for various complex environments. Extensive experiments have been conducted on four publicly available datasets. The results demonstrate that our method is able to achieve better recall performance than the state-of-the-art and meet the real-time requirement of vSLAM systems. Jiayi Ma 0001, Shenyue Wang, Kaining Zhang, Zheng He 0001, Jun Huang 0008, Xiaoguang Mei |
IEEE Trans. Ind. Informatics | 3 |
| 2022 | Appearance-Based Loop Closure Detection via Locality-Driven Accurate Motion Field LearningabstractLoop closure detection (LCD) is of significant importance in simultaneous localization and mapping. It represents the robot’s ability to recognize whether the current surrounding corresponds to a previously observed one. In this paper, we conduct this task in a two-step strategy: candidate frame selection and loop closure verification. The first step aims to search semantically similar images for the query one using features obtained by Key.Net with HardNet. Instead of adopting the traditional Bag-of-Words strategy, we utilize the aggregated selective match kernel to calculate the similarity between images. Subsequently, based on the potential property of motion field in the LCD scene, we propose a novel feature matching method,i.e., exploiting the smoothness prior and learning the motion field for an image pair in a reproducing kernel Hilbert space (RKHS), to implement loop closure verification. Concretely, we formulate the learning problem into a Bayesian framework with latent variables indicating the true/false correspondences and a mixture model accounting for the distribution of data. Furthermore, we propose a locality-driven mechanism to enhance the local relevance of motion vectors and term the algorithm as locality-driven accurate motion field learning (LAL). To satisfy the requirement of efficiency in the LCD task, we use a sparse approximation and search a suboptimal solution for the motion field in the RKHS, termed as LAL*. Extensive experiments are conducted on public datasets for feature matching and LCD tasks. The quantitative results demonstrate the effectiveness of our method over the current state-of-the-art, meanwhile showing its potential for long-term visual localization. The codes of LAL and LAL* are publicly available athttps://github.com/KN-Zhang/LAL. Kaining Zhang, Xingyu Jiang 0005, Jiayi Ma 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2021 | Appearance-based Loop Closure Detection via Bidirectional Manifold Representation ConsensusabstractLoop closure detection (LCD), which aims to deal with the drift emerging when robots travel around the route, plays a key role in a simultaneous localization and mapping system. Unlike most current methods which focus on seeking an appropriate representation of images, we propose a novel two-stage pipeline dominated by the estimation of spatial geometric relationship. When a query image occurs, we select candidates on-line according to the similarity of global semantic features in the first stage, and then conduct robust geometric confirmation to verify true loop-closing pairs in the second stage. To this end, a robust feature matching algorithm, termed as bidirectional manifold representation consensus (BMRC), is proposed. In particular, we utilize manifold representation to construct local neighborhood structures of feature points and formulate the matching problem into an optimization model, enabling linearithmic time complexity via a closed-form solution. Furthermore, we propose a dynamic place partition strategy based on BMRC to segment image streams with similar content into a place, which can mine more valid candidate frames, improving the recall rate of the whole system. Extensive experiments on several publicly available datasets reveal that BMRC has a good performance in the general feature matching task and the proposed pipeline outperforms the current state-of-the-art approaches in the LCD task. Kaining Zhang, Zizhuo Li, Jiayi Ma 0001 |
ICRA | 1 |
| 2021 | Motion Field Consensus with Locality Preservation: A Geometric Confirmation Strategy for Loop Closure DetectionabstractLoop closure detection (LCD), which aims to deal with the drift emerging when robots travel around the route, plays a key role in a simultaneous localization and mapping system. Unlike most current methods which focus on seeking an appropriate representation of images, we propose a novel two-stage pipeline dominated by the estimation of spatial geometric relationship. When a query image occurs, we select semantically similar images based on the SuperPoint network and the aggregated selective match kernel in the first stage, and then conduct robust geometric confirmation to verify true loop-closing pairs in the second stage. Based on the potential property of motion field in the LCD scene, a robust feature matching algorithm, termed as motion field consensus with locality preservation (MFC-LP), is proposed. In particular, we exploit the smoothness prior to guide the learning of the motion field for an image pair in a reproducing kernel Hilbert space (RKHS). Meanwhile, to enhance the local relevance of motion vectors, we design a locality preservation mechanism thus making the learned motion field more accurate. Extensive experiments on several publicly available datasets reveal that MFC-LP has a good performance in the general feature matching task and the proposed pipeline outperforms the current state-of-the-art approaches in the LCD task. Kaining Zhang, Xingyu Jiang 0005, Xiaoguang Mei, Huabing Zhou, Jiayi Ma 0001 |
IROS | 1 |