Mingyang Li 0001

dblp:85/7189-1 · DBLP profile ↗
← Back
28ranked-venue papers
5as first author
10since 2021 · last 2025
0000-0002-1854-3076ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 5 first-author · 7 since 2021Systems, architecture and hardware · 16 · 5 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Through the Curved Cover: Synthesizing Cover Aberrated Scenes with Refractive Field
abstract
Recent extended reality headsets and field robots have adopted covers to protect the front-facing cameras from environmental hazards and falls. The surface irregularities on the cover can lead to optical aberrations like blurring and non-parametric distortions. Novel view synthesis methods like NeRF and 3D Gaussian Splatting are ill-equipped to synthesize from sequences with optical aberrations. To address this challenge, we introduce SynthCover to enable novel view synthesis through protective covers for downstream extended reality applications. SynthCover employs a Refractive Field that estimates the cover's geometry, enabling precise analytical calculation of refracted rays. Experiments on synthetic and real-world scenes demonstrate our method's ability to accurately model scenes viewed through protective covers, achieving a significant improvement in rendering quality compared to prior methods. We also show that the model can adjust well to various cover geometries with synthetic sequences captured with covers of different surface curvatures. To motivate further studies on this problem, we provide the benchmarked dataset containing real and synthetic walkable scenes captured with protective cover optical aberrations.
Liuyue Xie, Jiancong Guo, László A. Jeni, Zhiheng Jia, Mingyang Li 0001, Yunwen Zhou
WACV5
2025 FMGS: Foundation Model Embedded 3D Gaussian Splatting for Holistic 3D Scene Understanding
Xingxing Zuo 0001, Pouya Samangouei, Yunwen Zhou, Yan Di, Mingyang Li 0001
Int. J. Comput. Vis.5
2024 NeRF-VINS: A Real-time Neural Radiance Field Map-based Visual-Inertial Navigation System
abstract
Achieving efficient and consistent localization with a prior map remains challenging in robotics. Conventional keyframe-based approaches often suffer from sub-optimal viewpoints due to limited field of view (FOV) and/or constrained motion, thus degrading the localization performance. To address this issue, we design a real-time tightly-coupled Neural Radiance Fields (NeRF)-aided visual-inertial navigation system (VINS). In particular, by effectively leveraging the NeRF’s potential to synthesize novel views, the proposed NeRF-VINS overcomes the limitations of traditional keyframe-based maps (with limited views) and optimally fuses IMU, monocular images, and synthetically rendered images within an efficient filter-based framework. This tightly-coupled fusion enables efficient 3D motion tracking with bounded errors. We extensively validate the proposed NeRF-VINS against the state-of-the-art methods that use prior map information, and demonstrate its ability to perform real-time localization, at 15 Hz, on a resource-constrained Jetson AGX Orin embedded platform.
Saimouli Katragadda, Woosik Lee 0003, Yuxiang Peng 0002, Patrick Geneva, Chuchu Chen, Mingyang Li 0001, Guoquan Huang 0001
ICRA7
2024 Visual-Based Kinematics and Pose Estimation for Skid-Steering Robots
abstract
To build commercial robots, skid-steering mechanical design is of increased popularity due to its manufacturing simplicity and unique mechanism. However, these also cause significant challenges on software and algorithm design, especially for the pose estimation (i.e., determining the robot’s rotation and position) of skid-steering robots, since they change their orientation with an inevitable skid. To tackle this problem, we propose a probabilistic sliding-window estimator dedicated to skid-steering robots, using measurements from a monocular camera, the wheel encoders, and optionally an inertial measurement unit (IMU). Specifically, we explicitly model the kinematics of skid-steering robots by both track instantaneous centers of rotation (ICRs) and correction factors, which are capable of compensating for the complexity of track-to-terrain interaction, the imperfectness of mechanical design, terrain conditions and smoothness, etc. To prevent performance reduction in robots’ long-term missions, the time- and location- varying kinematic parameters are estimated online along with pose estimation states in a tightly-coupled manner. More importantly, we conduct in-depth observability analysis for different sensors and design configurations in this paper, which provides us with theoretical tools in making the correct choice when building real commercial robots. In our experiments, we validate the proposed method by both simulation tests and real-world experiments, which demonstrate that our method outperforms competing methods by wide margins. Note to Practitioners—This paper was motivated by the problem of long-term pose estimation of the commonly commercial-used skid-steering robots with only low-cost sensors. Skid-steering robots change their orientation with a skid, which poses a significant challenge for pose estimation when using the wheel encoders. We propose to online estimate the robot’s kinematics, which succeeds in compensating for the complexity of track-to-terrain interaction, due to the slippage, the imperfectness of mechanical design, terrain conditions and smoothness. It is critical to estimate the kinematics and poses jointly to prevent performance reduction in robots’ long-term missions. We further theoretically analyze whether the kinematics parameters can be estimated under different sensor configurations, and find out the special degrade motions that make the parameters unobservable.
Xingxing Zuo 0001, Mingming Zhang 0008, Mengmeng Wang 0005, Yiming Chen 0001, Guoquan Huang 0001, Yong Liu 0007, Mingyang Li 0001
IEEE Trans Autom. Sci. Eng.7
2022 DuMLP-Pin: A Dual-MLP-Dot-Product Permutation-Invariant Network for Set Feature Extraction
abstract
Existing permutation-invariant methods can be divided into two categories according to the aggregation scope, i.e. global aggregation and local one. Although the global aggregation methods, e. g., PointNet and Deep Sets, get involved in simpler structures, their performance is poorer than the local aggregation ones like PointNet++ and Point Transformer. It remains an open problem whether there exists a global aggregation method with a simple structure, competitive performance, and even much fewer parameters. In this paper, we propose a novel global aggregation permutation-invariant network based on dual MLP dot-product, called DuMLP-Pin, which is capable of being employed to extract features for set inputs, including unordered or unstructured pixel, attribute, and point cloud data sets. We strictly prove that any permutation-invariant function implemented by DuMLP-Pin can be decomposed into two or more permutation-equivariant ones in a dot-product way as the cardinality of the given input set is greater than a threshold. We also show that the DuMLP-Pin can be viewed as Deep Sets with strong constraints under certain conditions. The performance of DuMLP-Pin is evaluated on several different tasks with diverse data sets. The experimental results demonstrate that our DuMLP-Pin achieves the best results on the two classification problems for pixel sets and attribute sets. On both the point cloud classification and the part segmentation, the accuracy of DuMLP-Pin is very close to the so-far best-performing local aggregation method with only a 1-2% difference, while the number of required parameters is significantly reduced by more than 85% in classification and 69% in segmentation, respectively. The code is publicly available on https://github.com/JaronTHU/DuMLP-Pin.
Jiajun Fei, Wenlei Liu, Zhidong Deng, Mingyang Li 0001, Huanjun Deng
AAAI5
2022 SuperLine3D: Self-supervised Line Segmentation and Description for LiDAR Point Cloud
Xiangrui Zhao, Sheng Yang 0007, Tianxin Huang, Jun Chen 0023, Mingyang Li 0001, Yong Liu 0007
ECCV (9)6
2022 Translation Invariant Global Estimation of Heading Angle Using Sinogram of LiDAR Point Cloud
abstract
Global point cloud registration is an essential module for localization, of which the main difficulty exists in estimating the rotation globally without initial value. With the aid of gravity alignment, the degree of freedom in point cloud registration could be reduced to 4DoF, in which only the heading angle is required for rotation estimation. In this paper, we propose a fast and accurate global heading angle estimation method for gravity-aligned point clouds. Our key idea is that we generate a translation invariant representation based on Radon Transform, allowing us to solve the decoupled heading angle globally with circular cross-correlation. Besides, for heading angle estimation between point clouds with different distributions, we implement this heading angle estimator as a differentiable module to train a feature extraction network end-to-end. The experimental results validate the effectiveness of the proposed method in heading angle estimation and show better performance compared with other methods.
Xiaqing Ding, Xuecheng Xu, Yanmei Jiao, Mengwen Tan, Rong Xiong, Huanjun Deng, Mingyang Li 0001, Yue Wang 0020
ICRA8
2022 The Visual-Inertial- Dynamical Multirotor Dataset
abstract
Recently, the community has witnessed numerous datasets built for developing and testing state estimators. However, for some applications such as aerial transportation or search-and-rescue, the contact force or other disturbance must be perceived for robust planning and control, which is beyond the capacity of these datasets. This paper introduces a Visual-Inertial-Dynamical (VID) dataset, not only focusing on traditional six degrees of freedom (6-DOF) pose estimation but also providing dynamical characteristics of the flight platform for external force perception or dynamics-aided estimation. The VID dataset contains hardware synchronized imagery and inertial measurements, with accurate ground truth trajectories for evaluating common visual-inertial estimators. Moreover, the proposed dataset highlights rotor speed and motor current measurements, control inputs, and ground truth 6-axis force data to evaluate external force estimation. To the best of our knowledge, the proposed VID dataset is the first public dataset containing visual-inertial and complete dynamical information in the real world for pose and external force evaluation. The dataset1and related files2are open-sourced.
Kunyi Zhang, Tiankai Yang 0002, Ziming Ding, Sheng Yang 0007, Mingyang Li 0001, Chao Xu 0001, Fei Gao 0011
ICRA6
2021 IMU Data Processing For Inertial Aided Navigation: A Recurrent Neural Network Based Approach
abstract
In this work, we propose a novel method for performing inertial aided navigation, by using deep neural net-works (DNNs). To date, most DNN inertial navigation methods focus on the task of inertial odometry, by taking gyroscope and accelerometer readings as input and regressing for integrated IMU poses (i.e., position and orientation). While this design has been successfully applied on a number of applications, it is not of theoretical performance guarantee unless patterned motion is involved. This inevitably leads to significantly reduced accuracy and robustness in certain use cases. To solve this problem, we design a framework to compute observable IMU integration terms using DNNs, followed by the numerical pose integration and sensor fusion to achieve the performance gain. Specifically, we perform detailed analysis on the motion terms in IMU kinematic equations, propose a dedicated network design, loss functions, and training strategies for the IMU data processing, and conduct extensive experiments. The results show that our method is generally applicable and outperforms both traditional and DNN methods by wide margins.
Mingming Zhang 0008, Yiming Chen 0001, Mingyang Li 0001
ICRA4
2021 Pose Estimation for Ground Robots: On Manifold Representation, Integration, Reparameterization, and Optimization
abstract
In this article, we focus on pose estimation dedicated to nonholonomic ground robots with low-cost sensors, by probabilistically fusing measurements from wheel odometers and an exteroceptive sensor. For ground robots, wheel odometers are widely used in pose estimation tasks, especially in applications in planar scenes. However, since wheel odometer only provides two-dimensional (2D) motion measurements, it is extremely challenging to use that for accurate full 6-D pose (3-D position and 3-D orientation) estimation. Traditional methods for 6-D pose estimation with wheel odometers either approximate motion profiles at the cost of accuracy reduction, or rely on other sensors, e.g., inertial measurement unit, to provide complementary measurements. By contrast, we propose a novel motion-manifold-based method for pose estimation of ground robots, which enables to utilize wheel odometers for high-precision 6-D pose estimation. Specifically, the proposed method, first, formulates the motion manifold of ground robots by a parametric representation, second, performs manifold-based 6-D integration with wheel odometer measurements only, and third, reparameterizes manifold representation periodically for error reduction. To demonstrate the effectiveness and applicability of the proposed algorithmic module, we integrate that into a sliding-window pose estimator by using measurements from wheel odometers and a monocular camera. Extensive simulated and real-world experiments are conducted for evaluation, and the proposed algorithm is shown to outperform competing the state-of-the-art algorithms by a significant margin in pose estimation accuracy, especially when deployed in complex, large-scale real-world environments.
Mingming Zhang 0008, Xingxing Zuo 0001, Yiming Chen 0001, Yong Liu 0007, Mingyang Li 0001
IEEE Trans. Robotics5
2020 MonoPair: Monocular 3D Object Detection Using Pairwise Spatial Relationships
abstract
Monocular 3D object detection is an essential component in autonomous driving while challenging to solve, especially for those occluded samples which are only partially visible. Most detectors consider each 3D object as an independent training target, inevitably resulting in a lack of useful information for occluded samples. To this end, we propose a novel method to improve the monocular 3D object detection by considering the relationship of paired samples. This allows us to encode spatial constraints for partially-occluded objects from their adjacent neighbors. Specifically, the proposed detector computes uncertainty-aware predictions for object locations and 3D distances for the adjacent object pairs, which are subsequently jointly optimized by nonlinear least squares. Finally, the one-stage uncertainty-aware prediction structure and the post-optimization module are dedicatedly integrated for ensuring the run-time efficiency. Experiments demonstrate that our method yields the best performance on KITTI 3D detection benchmark, by outperforming state-of-the-art competitors by wide margins, especially for the hard samples.
Yongjian Chen, Lei Tai, Mingyang Li 0001
CVPR4
2020 Interpretable Foreground Object Search as Knowledge Distillation
Boren Li, Po-Yu Zhuang, Mingyang Li 0001
ECCV (28)4
2020 Overflow Aware Quantization: Accelerating Neural Network Inference by Low-bit Multiply-Accumulate Operations
abstract
The inherent heavy computation of deep neural networks prevents their widespread applications. A widely used method for accelerating model inference is quantization, by replacing the input operands of a network using fixed-point values. Then the majority of computation costs focus on the integer matrix multiplication accumulation. In fact, high-bit accumulator leads to partially wasted computation and low-bit one typically suffers from numerical overflow. To address this problem, we propose an overflow aware quantization method by designing trainable adaptive fixed-point representation, to optimize the number of bits for each input tensor while prohibiting numeric overflow during the computation. With the proposed method, we are able to fully utilize the computing power to minimize the quantization loss and obtain optimized inference performance. To verify the effectiveness of our method, we conduct image classification, object detection, and semantic segmentation tasks on ImageNet, Pascal VOC, and COCO datasets, respectively. Experimental results demonstrate that the proposed method can achieve comparable performance with state-of-the-art quantization methods while accelerating the inference process by about 2 times.
Hongwei Xie, Yafei Song 0002, Mingyang Li 0001
IJCAI4
2019 Seq-SG2SL: Inferring Semantic Layout From Scene Graph Through Sequence to Sequence Learning
abstract
Generating semantic layout from scene graph is a crucial intermediate task connecting text to image. We present a conceptually simple, flexible and general framework using sequence to sequence (seq-to-seq) learning for this task. The framework, called Seq-SG2SL, derives sequence proxies for the two modality and a Transformer-based seq-to-seq model learns to transduce one into the other. A scene graph is decomposed into a sequence of semantic fragments (SF), one for each relationship. A semantic layout is represented as the consequence from a series of brick-action code segments (BACS), dictating the position and scale of each object bounding box in the layout. Viewing the two building blocks, SF and BACS, as corresponding terms in two different vocabularies, a seq-to-seq model is fittingly used to translate. A new metric, semantic layout evaluation understudy (SLEU), is devised to evaluate the task of semantic layout prediction inspired by BLEU. SLEU defines relationships within a layout as unigrams and looks at the spatial distribution for n-grams. Unlike the binary precision of BLEU, SLEU allows for some tolerances spatially through thresholding the Jaccard Index and is consequently more adapted to the task. Experimental results on the challenging Visual Genome dataset show improvement over a non-sequential approach based on graph convolution.
Boren Li, Boyu Zhuang, Mingyang Li 0001
ICCV3
2019 2D LiDAR Map Prediction via Estimating Motion Flow with GRU
abstract
It is a significant problem to predict the 2D LiDAR map at next moment for robotics navigation and path-planning. To tackle this problem, we resort to the motion flow between adjacent maps, as motion flow is a powerful tool to process and analyze the dynamic data, which is named optical flow in video processing. However, unlike video, which contains abundant visual features in each frame, a 2D LiDAR map lacks distinctive local features. To alleviate this challenge, we propose to estimate the motion flow based on deep neural networks inspired by its powerful representation learning ability in estimating the optical flow of the video. To this end, we design a recurrent neural network based on gated recurrent unit, which is named LiDAR-FlowNet. As a recurrent neural network can encode the temporal dynamic information, our LiDAR-FlowNet can estimate motion flow between the current map and the unknown next map only from the current frame and previous frames. A self-supervised strategy is further designed to train the LiDAR-FlowNet model effectively, while no training data need to be manually annotated. With the estimated motion flow, it is straightforward to predict the 2D LiDAR map at the next moment. Experimental results verify the effectiveness of our LiDAR-FlowNet as well as the proposed training strategy. The results of the predicted LiDAR map also show the advantages of our motion flow based method.
Yafei Song 0002, Yonghong Tian 0001, Gang Wang 0012, Mingyang Li 0001
ICRA4
2019 Perception System Design for Low-Cost Commercial Ground Robots: Sensor Configurations, Calibration, Localization and Mapping
abstract
For commercially successful ground robots, high degree of autonomy, low manufacturing and maintenance cost, as well as minimized deployment limitations in different environments are essential attributes. To deliver an `anywhere deployable' product, it is impractical to rely on one single sensor or one single piece of algorithm to overcome all related challenges. Instead, the entire robotic system should be dedicated designed, including the choices of sensors, processors, algorithm integration for various functionality, and so on.This paper presents our design of perception system for commercial ground robots, which is able to operate in most common environments. The designed system is equipped with low-cost sensors and processors. The first key contribution of this paper is the design of the robotic sensory system, which includes a monocular camera, a 2D laser range finder (LRF), wheel encoders, and an inertial measurement unit (IMU). Our sensory system can be built at a cost of as low as $100. Furthermore, the selected sensors provide complementary characteristics for perception of both robot ego-motion and its surrounding environments, which are the prerequisites for `anywhere' deployment. The second key contribution of this paper is that a complete set of technologies is proposed based on our sensor systems, including sensor calibration (factory calibration and online calibration), localization (environmental exploring and re-localization), as well as mapping. The proposed methodology includes both efficient engineering implementation and theoretical novelty for high performance systems. Experimental results from our robotic testing platform and off-the-shelf commercial robots are presented. These results demonstrate that the proposed system can be deployed in various environmental conditions without performance compromise.
Yiming Chen 0001, Mingming Zhang 0008, Dongsheng Hong, Chengcheng Deng, Mingyang Li 0001
IROS5
2019 Learning Local Feature Descriptor with Motion Attribute For Vision-based Localization
abstract
In recent years, camera-based localization has been widely used for robotic applications, and most proposed algorithms rely on local features extracted from recorded images. For better performance, the features used for open-loop localization are required to be short-term globally static, and the ones used for re-localization or loop closure detection need to be long-term static. Therefore, the motion attribute of a local feature point could be exploited to improve localization performance, e.g., the feature points extracted from moving persons or vehicles can be excluded from these systems due to their unsteadiness. In this paper, we design a fully convolutional network (FCN), named MD-Net, to perform motion attribute estimation and feature description simultaneously. MD-Net has a shared backbone network to extract features from the input image and two network branches to complete each sub-task. With MD-Net, we can obtain the motion attribute while avoiding increasing much more computation. Experimental results demonstrate that the proposed method can learn distinct local feature descriptor along with motion attribute only using an FCN, by outperforming competing methods by a wide margin. We also show that the proposed algorithm can be integrated into a vision-based localization algorithm to improve estimation accuracy significantly.
Yafei Song 0002, Jia Li 0003, Yonghong Tian 0001, Mingyang Li 0001
IROS5
2019 Vision-Aided Localization For Ground Robots
abstract
In this paper, we focus on the problem of vision-based localization for ground robotic applications. In recent years, camera only or camera-IMU (inertial measurement unit) based localization methods are widely studied, in terms of theoretical properties, algorithm design, and real-world applications. However, we experimentally find that none of existing methods is able to perform high-precision and robust localization for ground robots in large-scale complicated 3D environments. To this end, in this paper, we propose a novel vision-based localization algorithm dedicatedly designed for ground robots, by fusing measurements from a camera, an IMU, and the wheel odometer. The first contribution of this paper is that we propose a novel algorithm for approximating the motion manifold for ground robots by parametric representation and performing pose integrating via IMU and wheel odometer measurements. Secondly, we propose a complete localization algorithm, by using a sliding-window based estimator. The estimator is designed based on iterative optimization to fuse measurements from multiple sensors on the proposed manifold representation. We show that, based on a variety of real-world experiments, the proposed algorithm outperforms a number of the state-of-the-art vision based localization algorithms by a significant margin, especially when deployed in large-scale complicated environments.
Mingming Zhang 0008, Yiming Chen 0001, Mingyang Li 0001
IROS3
2019 Visual-Inertial Localization for Skid-Steering Robots with Kinematic Constraints
Xingxing Zuo 0001, Mingming Zhang 0008, Yiming Chen 0001, Yong Liu 0007, Guoquan Huang 0001, Mingyang Li 0001
ISRR6
2017 Visual-inertial self-calibration on informative motion segments
abstract
Environmental conditions and external effects, such as shocks, have a significant impact on the calibration parameters of visual-inertial sensor systems. Thus long-term operation of these systems cannot fully rely on factory calibration. Since the observability of certain parameters is highly dependent on the motion of the device, using short data segments at device initialization may yield poor results. When such systems are additionally subject to energy constraints, it is also infeasible to use full-batch approaches on a big dataset and careful selection of the data is of high importance. In this paper, we present a novel approach for resource efficient self-calibration of visual-inertial sensor systems. This is achieved by casting the calibration as a segment-based optimization problem that can be run on a small subset of informative segments. Consequently, the computational burden is limited as only a predefined number of segments is used. We also propose an efficient information-theoretic selection to identify such informative motion segments. In evaluations on a challenging dataset, we show our approach to significantly outperform state-of-the-art in terms of computational burden while maintaining a comparable accuracy.
Thomas Schneider 0007, Mingyang Li 0001, Michael Burri, Juan I. Nieto 0001, Roland Siegwart, Igor Gilitschenski
ICRA2
2017 Photometric patch-based visual-inertial odometry
abstract
In this paper we present a novel direct visual-inertial odometry algorithm, for estimating motion in unknown environments. The algorithm utilizes image patches extracted around image features, and formulates measurement residuals in the image intensity space directly. One key characteristic of the proposed method is that it models the true irradiance at each pixel as a random variable to be estimated and marginalized out. The formulation of the photometric residual explicitly accounts for the camera response function and lens vignetting (which can be calibrated in advance), as well as unknown illumination gains and biases, which are estimated on a per-feature or per-image basis. We present a detailed evaluation of our algorithm on 50 datasets with high-precision ground truth, which amount to approximately 1.5 hours of localization data. Through a direct comparison with a point-feature based method, we demonstrate that the use of photometric residuals results in increased pose estimation accuracy, with approximately 23% lower estimation errors, on average.
Xing Zheng, Zack Moratto, Mingyang Li 0001, Anastasios I. Mourikis
ICRA3
2017 A dense flow-based framework for real-time object registration under compound motion
Songfan Yang, Yinjie Lei, Mingyang Li 0001, Ninad Thakoor, Bir Bhanu, Yiguang Liu
Pattern Recognit.4
2014 High-fidelity sensor modeling and self-calibration in vision-aided inertial navigation
abstract
In this paper, we propose a high-precision pose estimation algorithm for systems equipped with low-cost inertial sensors and rolling-shutter cameras. The key characteristic of the proposed method is that it performs online self-calibration of the camera and the IMU, using detailed models for both sensors and for their relative configuration. Specifically, the estimated parameters include the camera intrinsics (focal length, principal point, and lens distortion), the readout time of the rolling-shutter sensor, the IMU's biases, scale factors, axis misalignment, and g-sensitivity, the spatial configuration between the camera and IMU, as well as the time offset between the timestamps of the camera and IMU. An additional contribution of this work is a novel method for processing the measurements of the rolling-shutter camera, which employs an approximate representation of the estimation errors, instead of the state itself. We demonstrate, in both simulation tests and real-world experiments, that the proposed approach is able to accurately calibrate all the considered parameters in real time, and leads to significantly improved estimation precision compared to existing approaches.
Mingyang Li 0001, Hongsheng Yu, Xing Zheng, Anastasios I. Mourikis
ICRA1
2013 Real-time motion tracking on a cellphone using inertial sensing and a rolling-shutter camera
abstract
All existing methods for vision-aided inertial navigation assume a camera with a global shutter, in which all the pixels in an image are captured simultaneously. However, the vast majority of consumer-grade cameras use rolling-shutter sensors, which capture each row of pixels at a slightly different time instant. The effects of the rolling shutter distortion when a camera is in motion can be very significant, and are not modelled by existing visual-inertial motion-tracking methods. In this paper we describe the first, to the best of our knowledge, method for vision-aided inertial navigation using rolling-shutter cameras. Specifically, we present an extended Kalman filter (EKF)-based method for visual-inertial odometry, which fuses the IMU measurements with observations of visual feature tracks provided by the camera. The key contribution of this work is a computationally tractable approach for taking into account the rolling-shutter effect, incurring only minimal approximations. The experimental results from the application of the method show that it is able to track, in real time, the position of a mobile phone moving in an unknown environment with an error accumulation of approximately 0.8% of the distance travelled, over hundreds of meters.
Mingyang Li 0001, Byung Hyung Kim, Anastasios I. Mourikis
ICRA1
2013 3-D motion estimation and online temporal calibration for camera-IMU systems
abstract
When measurements from multiple sensors are combined for real-time motion estimation, the time instant at which each measurement was recorded must be precisely known. In practice, however, the timestamps of each sensor's measurements are typically affected by a delay, which is different for each sensor. This gives rise to a temporal misalignment (i.e., a time offset) between the sensors' data streams. In this work, we propose an online approach for estimating the time offset between the data obtained from different sensors. Specifically, we focus on the problem of motion estimation using visual and inertial sensors in extended Kalman filter (EKF)-based methods. The key idea proposed here is to explicitly include the time offset between the camera and IMU in the EKF state vector, and estimate it online along with all other variables of interest (the IMU pose, the camera-to-IMU calibration, etc). Our proposed approach is general, and can be employed in several classes of estimation problems, such as motion estimation based on mapped features, EKF-based SLAM, or visual-inertial odometry. Our simulation and experimental results demonstrate that the proposed approach yields high-precision, consistent estimates, in scenarios involving both constant and time-varying offsets.
Mingyang Li 0001, Anastasios I. Mourikis
ICRA1
2012 Improving the accuracy of EKF-based visual-inertial odometry
abstract
In this paper, we perform a rigorous analysis of EKF-based visual-inertial odometry (VIO) and present a method for improving its performance. Specifically, we examine the properties of EKF-based VIO, and show that the standard way of computing Jacobians in the filter inevitably causes inconsistency and loss of accuracy. This result is derived based on an observability analysis of the EKF's linearized system model, which proves that the yaw erroneously appears to be observable. In order to address this problem, we propose modifications to the multi-state constraint Kalman filter (MSCKF) algorithm [1], which ensure the correct observability properties without incurring additional computational cost. Extensive simulation tests and real-world experiments demonstrate that the modified MSCKF algorithm outperforms competing methods, both in terms of consistency and accuracy.
Mingyang Li 0001, Anastasios I. Mourikis
ICRA1
2012 Vision-aided inertial navigation for resource-constrained systems
abstract
In this paper we present a resource-adaptive framework for real-time vision-aided inertial navigation. Specifically, we focus on the problem of visual-inertial odometry (VIO), in which the objective is to track the motion of a mobile platform in an unknown environment. Our primary interest is navigation using miniature devices with limited computational resources, similar for example to a mobile phone. Our proposed estimation framework consists of two main components: (i) a hybrid EKF estimator that integrates two algorithms with complementary computational characteristics, namely a sliding-window EKF and EKF-based SLAM, and (ii) an adaptive image-processing module that adjusts the number of detected image features based on the availability of resources. By combining the hybrid EKF estimator, which optimally utilizes the feature measurements, with the adaptive image-processing algorithm, the proposed estimation architecture fully utilizes the system's computational resources. We present experimental results showing that the proposed estimation framework is capable of real-time processing of image and inertial data on the processor of a mobile phone.
Mingyang Li 0001, Anastasios I. Mourikis
IROS1
2011 A particle filter for monocular vision-aided odometry
abstract
We propose a particle filter-based algorithm for monocular vision-aided odometry for mobile robot localization. The algorithm fuses information from odometry with observations of naturally occurring static point features in the environment. A key contribution of this work is a novel approach for computing the particle weights, which does not require including the feature positions in the state vector. As a result, the computational and sample complexities of the algorithm remain low even in feature-dense environments. We validate the effectiveness of the approach extensively with both simulations as well as real-world data, and compare its performance against that of the extended Kalman filter (EKF) and FastSLAM. Results from the simulation tests show that the particle filter approach is better than these competing approaches in terms of the RMS error. Moreover, the experiments demonstrate that the approach is capable of achieving good localization accuracy in complex environments.
Teddy N. Yap Jr., Mingyang Li 0001, Anastasios I. Mourikis, Christian R. Shelton
ICRA2