EDBT 2026 Demo / reviewers in the wild / expert
Yanmei Jiao
dblp:227/7803
· DBLP profile ↗
12ranked-venue papers
5as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 5 since 2021Systems, architecture and hardware · 7 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | 3D Model-Free Visual Localization System From Essential Matrix Under Local Planar MotionabstractVisual localization plays a critical role in the functionality of low-cost autonomous mobile robots. Contemporary leading methods for precise visual localization are predominantly 3D scene-specific, necessitating extra computational and memory overhead to construct a 3D scene model in novel environments. An alternative approach of directly using a database of 2D images for visual localization offers more flexibility. However, such methods currently suffer from limited localization accuracy. In this paper, we propose an accurate and robust multiple checking-based 3D model-free visual localization system to address the aforementioned issues. To ensure high accuracy, our focus is on estimating the pose of a query image relative to the retrieved database images using 2D-2D feature matches. Theoretically, by incorporating the local planar motion constraint into both the estimation of the essential matrix and the triangulation stages, we reduce the minimum required feature matches for absolute pose estimation, thereby enhancing the robustness of outlier rejection. Additionally, we introduce a multiple-checking mechanism to ensure the correctness of the solution throughout the solving process. The efficacy of our approach is substantiated through both qualitative and quantitative assessments on simulated and two real-world datasets evidencing significant improvements in accuracy and robustness provided by our 3D model-free visual localization system.Note to Practitioners—The motivation of this article stems from the need to develop an accurate visual localization system with simplicity and flexibility of map construction and easy adaption to new environments. Such a system holds great practical value for a range of applications, including warehouse robots, service robots, and countless others. Existing visual localization systems that achieve high accuracy are dependent on a pre-built accurate 3D scene map, which pose challenges in terms of map construction and consume significant storage resources onboard, particularly for large scenes. And the aforementioned efforts need to be repeated when changing to a new scene. In this article, an accurate and robust 3D model-free visual localization system is proposed to handle this problem. The map construction is simplified to build a set of database images with associated camera poses, which is trivial as it amounts to adding posed images to a database. The core idea for achieving high accuracy and robustness is to model the local planar motion characteristic of general ground-moving robots into both essential matrix estimation and triangulation stages to obtain two minimal solutions. The proposed localization system simplifies the task of switching between different application scenarios for the robot, reducing additional workload and lowering the difficulty of use. Yanmei Jiao, Binxin Zhang, Peng Jiang 0016, Chaoqun Wang 0009, Haojian Lu, Rong Xiong, Yue Wang 0020 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2025 | Robust Vehicle Localization for Spherical Camera Models: Solution, Framework, and VerificationabstractVehicle visual localization uses vision sensors to capture environmental information, enabling precise localization of autonomous vehicles within their surroundings. However, current visual localization methods generally have some shortcomings: on one hand, they are limited by the camera’s field of view, on the other hand, their robustness is often inadequate under challenging conditions such as lighting changes, long-term scene changes, or occlusions. To address these issues, we formulate a general spherical camera model for both fisheye and panoramic cameras and propose a minimal solution for pose estimation using this model based on vehicle motion characteristic. The minimal solution cannot filter outliers, so a robust estimation framework is necessary. For outlier-rejection, we introduce two frameworks: a probabilistic optimal RANSAC and a globally optimal graph-based framework. We conduct a probabilistic analysis of the RANSAC to demonstrate its enhanced robustness given by the proposed minimal solution. To achieve robustness to extreme outliers (higher than 90%), we decouple the rotation and translation space through the minimal solution to construct maximum consensus graph for the two sub-problems. We then employs a maximum clique search algorithm to find the optimal solutions, achieving deterministic convergence while maintaining real-time performance. Extensive experiments with synthetic data, real-world fisheye images, and 360∘panoramic images validate the robustness and efficiency of our proposed algorithms. Yanmei Jiao, Dibin Zhou, Xiumei Li, Rong Xiong, Yue Wang 0020 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | Vertebrae-based Global X-ray to CT Registration for Thoracic SurgeriesabstractX-ray to CT registration is an essential technique to provide on-site guidance for clinicians and medical robots by aligning preoperative information with intraoperative images. Current methods focus on local registration with small capture ranges and necessitate a manual initial alignment before precise registration. Some existing global methods are likely to fail in thoracic surgeries because of the respiratory motion and the nearly colinear nature of vertebrae landmarks. In this study, we propose a vertebrae-based global X-ray to CT registration method with the assistance of clinical setups for thoracic surgeries. Firstly, vertebrae centroids are automatically localized by CNN-based networks in CT and X-ray for establishing 2D/3-D correspondences. Then, inspired by clinical setup, we address the degradation of colinear landmarks of 6-DoF pose estimation by introducing a 4-DoF solver. Considering the inaccurate priori and landmark mislocalization, the solver is embedded into the Adaptive Error-Aware Estimator (AE2) to simultaneously estimate weights and aggregate candidate poses. Finally, the whole method is trained in an end-to-end manner for better performance. Evaluations on both the public LIDC-IDRI dataset and clinical dataset demonstrate that our method outperforms existing optimization-based and learningbased approaches in terms of registration accuracy and success rate. Our code: https://github.com/LiuLiluZJU/2P-AE2 Lilu Liu, Yanmei Jiao, Zhou An, Honghai Ma, Chunlin Zhou, Haojian Lu, Rong Xiong, Yue Wang 0020 |
IROS | 2 |
| 2024 | Fusing Multiple Isolated Maps to Visual Inertial Odometry Online: A Consistent FilterabstractVisual inertial odometry (VIO) is widely used in various kinds of mobile platforms to provide the ego-pose of the platforms. With the help of pre-built map information, the drift of the VIO can be constrained. However, constructing a globally consistent map is a tough job, especially for large scenes. In this paper, we propose a filter-based framework aiming to leverage multiple isolated maps to improve the performance of VIO such that building a globally consistent map can be avoided. In this framework, the relative transformations between the local VIO reference frame and the multiple map reference frames are regarded as 6 degrees of freedom (DoF) pose features to be online estimated. We call these relative transformations asaugmented variables. With theseaugmented variables, the map-based information can be tightly coupled into the VIO system to ease the drift of VIO. To fuse these maps consistently, we first theoretically analyze the observability properties of our proposed framework. Based on the analysis, the Schmidt extended Kalman filter (EKF) and the first-estimate Jacobian (FEJ) are employed to maintain the consistency of the system. Simulation and real-world experiments are conducted to demonstrate the effectiveness and consistency of our framework.Note to Practitioners—Visual inertial odometry (VIO) is widely applied to positioning mobile platforms including autonomous vehicles, robots, and virtual/augmented reality (VR/AR) devices. However, VIO inevitably suffers from drift, which will reduce positioning accuracy. This problem can be solved by fusing prior maps into VIO. Existing works mainly support online fusing one map into VIO. This requires users to offline merge multiple maps into one beforehand, which is complicated and troublesome and sometimes even unrealizable (e.g., the multiple maps have no overlap). According to theoretical analyses, this paper introduces a new system that can online fuse multiple maps into VIO. Our system has the following benefits: 1) It is a light-weighted filter-based system suitable for onboard deployments; 2) It can online fuse multiple maps such that pre-work of merging multiple maps into one can be bypassed; 3) Our system can consistently fuse the multi-map information while keeps the computation at a low level. Experiments show that with our system, the VIO’s drift can be significantly alleviated to benefit downstream tasks like planning, navigation, and control. However, our system needs to fix some linearization points to maintain the correct observability of the system, which will sacrifice some precision. Future works will include investigating more elegant techniques to maintain the observability of the system. Zhuqing Zhang, Yanmei Jiao, Rong Xiong, Yue Wang 0020 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2023 | DAMS-LIO: A Degeneration-Aware and Modular Sensor-Fusion LiDAR-inertial OdometryabstractWith robots being deployed in increasingly complex environments like underground mines and planetary surfaces, the multi-sensor fusion method has gained more and more attention which is a promising solution to state estimation in the such scene. The fusion scheme is a central component of these methods. In this paper, a light-weight iEKF-based LiDAR-inertial odometry system is presented, which utilizes a degeneration-aware and modular sensor-fusion pipeline that takes both LiDAR points and relative pose from another odometry as the measurement in the update process only when degeneration is detected. Both the Cramer-Rao Lower Bound (CRLB) theory and simulation test are used to demonstrate the higher accuracy of our method compared to methods using a single observation. Furthermore, the proposed system is evaluated in perceptually challenging datasets against various state-of-the-art sensor-fusion methods. The results show that the proposed system achieves real-time and high estimation accuracy performance despite the challenging environment and poor observations. Fuzhang Han, Rong Xiong, Yue Wang 0020, Yanmei Jiao |
ICRA | 6 |
| 2022 | Translation Invariant Global Estimation of Heading Angle Using Sinogram of LiDAR Point CloudabstractGlobal point cloud registration is an essential module for localization, of which the main difficulty exists in estimating the rotation globally without initial value. With the aid of gravity alignment, the degree of freedom in point cloud registration could be reduced to 4DoF, in which only the heading angle is required for rotation estimation. In this paper, we propose a fast and accurate global heading angle estimation method for gravity-aligned point clouds. Our key idea is that we generate a translation invariant representation based on Radon Transform, allowing us to solve the decoupled heading angle globally with circular cross-correlation. Besides, for heading angle estimation between point clouds with different distributions, we implement this heading angle estimator as a differentiable module to train a feature extraction network end-to-end. The experimental results validate the effectiveness of the proposed method in heading angle estimation and show better performance compared with other methods. Xiaqing Ding, Xuecheng Xu, Yanmei Jiao, Mengwen Tan, Rong Xiong, Huanjun Deng, Mingyang Li 0001, Yue Wang 0020 |
ICRA | 4 |
| 2022 | FEJ-VIRO: A Consistent First-Estimate Jacobian Visual-Inertial-Ranging OdometryabstractIn recent years, Visual-Inertial Odometry (VIO) has achieved many significant progresses. However, VIO meth-ods suffer from localization drift over long trajectories. In this paper, we propose a First-Estimates Jacobian Visual-Inertial-Ranging Odometry (FEJ-VIRO) to reduce the localization drifts of VIO by incorporating ultra-wideband (UWB) ranging measurements into the VIO framework consistently. Consid-ering that the initial positions of UWB anchors are usually unavailable, we propose a long-short window structure to initialize the UWB anchors' positions as well as the covariance for state augmentation. After initialization, the FEJ - VIRO estimates the UWB anchors' positions simultaneously along with the robot poses. We further analyze the observability of the visual-inertial-ranging estimators and proved that there are four unobservable directions in the ideal case, while one of them vanishes in the actual case due to the gain of spurious information. Based on these analyses, we leverage the FEJ technique to enforce the unobservable directions, hence reducing inconsistency of the estimator. Finally, we validate our analysis and evaluate the proposed FEJ-VIRO with both simulation and real-world experiments. Shenhan Jia, Yanmei Jiao, Zhuqing Zhang, Rong Xiong, Yue Wang 0020 |
IROS | 2 |
| 2022 | Deterministic Optimality for Robust Vehicle Localization Using Visual MeasurementsabstractLocalization is fundamental for autonomous vehicle applications. Compared with widely developed LiDAR-based vehicle localization, vision based localization has attracted considerable attention in recent years owing to the low-cost sensor. One of the challenges for visual localization is the outlier in measurements due to the appearance changes caused by illumination, season, and weather. To address this problem, we present a real-time robust visual localization approach that achieves deterministic optimality with global convergence. The idea is to decouple the rotation and translation estimation by utilizing the fact that the pitch and roll angles of the query pose are similar to those of the map reference pose, since the vehicle motion is locally planar. Based on the decoupled formulation, we first estimate the optimal yaw angle and eliminate the majority of outliers by an efficient inlier voting method, then find the optimal translation by maximum clique search. The two subproblem estimators are embedded into a prioritized search paradigm to guarantee deterministic optimality. In the experiments, the simulation demonstrates that the proposed method can achieve superior robustness even dealing with extreme outlier rates (95%). Results on both public and self-collected real-world vehicle datasets validate the effectiveness of the proposed method in the real application. Yanmei Jiao, Yue Wang 0020, Xiaqing Ding, Minhang Wang, Rong Xiong |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2021 | Robust localization for planar moving robot in changing environment: A perspective on density of correspondence and depthabstractVisual localization for planar moving robot is important to various indoor service robotic applications. To handle the textureless areas and frequent human activities in indoor environments, a novel robust visual localization algorithm which leverages dense correspondence and sparse depth for planar moving robot is proposed. The key component is a minimal solution which computes the absolute camera pose with one 3D-2D correspondence and one 2D-2D correspondence. The advantages are obvious in two aspects. First, the robustness is enhanced as the sample set for pose estimation is maximal by utilizing all correspondences with or without depth. Second, no extra effort for dense map construction is required to exploit dense correspondences for handling textureless and repetitive texture scenes. That is meaningful as building a dense map is computational expensive especially in large scale. Moreover, a probabilistic analysis among different solutions is presented and an automatic solution selection mechanism is designed to maximize the success rate by selecting appropriate solutions in different environmental characteristics. Finally, a complete visual localization pipeline considering situations from the perspective of correspondence and depth density is summarized and validated on both simulation and public real-world indoor localization dataset. Yanmei Jiao, Lilu Liu, Bo Fu 0006, Xiaqing Ding, Minhang Wang, Yue Wang 0020, Rong Xiong |
ICRA | 1 |
| 2020 | Globally optimal consensus maximization for robust visual inertial localization in point and line mapabstractMap based visual inertial localization is a crucial step to reduce the drift in state estimation of mobile robots. The underlying problem for localization is to estimate the pose from a set of 3D-2D feature correspondences, of which the main challenge is the presence of outliers, especially in changing environment. In this paper, we propose a robust solution based on efficient global optimization of the consensus maximization problem, which is insensitive to high percentage of outliers. We first introduce translation invariant measurements (TIMs) for both points and lines to decouple the consensus maximization problem into rotation and translation subproblems, allowing for a two-stage solver with reduced search space. Then we show that (i) the rotation can be estimated by minimizing TIMs using only 1-dimensional branch-and-bound (BnB), (ii) the translation can be estimated by running 1-dimensional search for each of the three axes with prioritized progressive voting. Compared with the popular randomized solver, our solver achieves deterministic global convergence without requiring an initial value. Furthermore, ours is exponentially faster compared with existing BnB based methods. Finally, our experiments on both simulation and real-world datasets demonstrate that the proposed method gives accurate pose estimation even in the presence of 90% outliers (only 2 inliers). Yanmei Jiao, Yue Wang 0020, Bo Fu 0006, Qimeng Tan, Lei Chen 0106, Minhang Wang, Shoudong Huang, Rong Xiong |
IROS | 1 |
| 2019 | 2-Entity RANSAC for robust visual localization in changing environmentabstractVisual localization has attracted considerable attention due to its low-cost and stable sensor, which is desired in many applications, such as autonomous driving, inspection robots and unmanned aerial vehicles. However, current visual localization methods still struggle with environmental changes across weathers and seasons, as there is significant appearance variation between the map and the query image. The crucial challenge in this situation is that the percentage of outliers, i.e. incorrect feature matches, is high. In this paper, we derive minimal closed form solutions for 3D-2D localization with the aid of inertial measurements, using only 2 point matches or 1 point match and 1 line match. These solutions are further utilized in the proposed 2-entity RANSAC, which is more robust to outliers as both line and point features can be used simultaneously and the number of matches required for pose calculation is reduced. Furthermore, we introduce three feature sampling strategies with different advantages, enabling an automatic selection mechanism. With the mechanism, our 2-entity RANSAC can be adaptive to the environments with different distribution of feature types in different segments. Finally, we evaluate the method on both synthetic and real-world datasets, validating its performance and effectiveness in inter-session scenarios. Yanmei Jiao, Yue Wang 0020, Bo Fu 0006, Xiaqing Ding, Qimeng Tan, Lei Chen 0106, Rong Xiong |
IROS | 1 |
| 2018 | MASD: A Multimodal Assembly Skill Decoding System for Robot Programming by DemonstrationabstractProgramming by demonstration (PBD) transforms the robot programming from the code level to automated interface between robot and human, promoting the flexibility of robotized automation. In this paper, we focus on programming the industrial robot for assembly tasks by parsing the human demonstration into a series of assembly skills and compiling the skill to the robot executables. To achieve this goal, an identification system using multimodal information to recognize the assembly skill, called MASD, is proposed including: 1) an initial learning stage using a hierarchical model to recognize the action by considering the features from action-object effect, gesture, and trajectory and 2) a retrospective thinking stage using a segmentation method to cut the continuous demonstrations into multiple assembly skills optimally. Using MASD, the demonstration of assembly tasks can be explained with high accuracy in real time, driving a hypothesis that a PBD system on the top of MASD can be extended to more realistic assembly tasks beyond pure positional moving and picking. In experiments, the skill identification module is used to recognize the five kinds of assembly skills in demonstrations of both single and multiple assembly skills, and outperforms the comparative action identification methods. Besides integrated with the MASD, the PBD system can generate the program based on the demonstration and successfully enable an ABB industrial robotic arm simulator to assemble a flashlight and a switch, verifying the initial hypothesis. Note to Practitioners-In the conventional robotized automation, the key role of the robot mainly owes to its capacity for repeating a wide variety of tasks with high speed and accuracy in long term, with a cost of days to months of programming for deployment. On the other hand, the new trend of customization brings the new characteristics: production in short cycle and small volume. This irreversible momentum urges the robot to switch from task to task efficiently. The biggest bottleneck here is the tedious programming, which also has high prerequisites for most practitioners in manufacturing. This situation motivates the development of a PBD system that can understand the assembly skills performed by the human experts in the demonstration and accordingly generate the program for robot's execution of the taught task. In this paper, we present a skill decoding system to parse the observational raw demonstration into symbolic sequences, which is the crucial bridge to enable the automatic programming. The system achieves high performance in recognition and is tailored for the PBD in assembly tasks by considering both advantages and disadvantages in the background of assembly, such as controllable environment and limited computational resources. It is particularly useful for assembly tasks with modularized actions based on a set of standard parts. At the perspective of industrial application, the PBD upon the proposed system is a promising solution to improve the flexibility of manufacture, which is expected to be true in midterm but an important step toward this goal. Yue Wang 0020, Yanmei Jiao, Rong Xiong, Hongsheng Yu, Jiafan Zhang, Yong Liu 0007 |
IEEE Trans Autom. Sci. Eng. | 2 |