EDBT 2026 Demo / reviewers in the wild / expert
Xingxing Zuo 0001
dblp:210/0948
· DBLP profile ↗
28ranked-venue papers
7as first author
22since 2021 · last 2026
0000-0003-4158-3153ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 6 first-author · 13 since 2021Systems, architecture and hardware · 11 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Flying Co-Stereo: Enabling Long-Range Aerial Dense Mapping via Collaborative Stereo Vision of Dynamic-BaselineabstractFor Unmanned Aerial Vehicle (UAV) swarms operating in large-scale unknown environments, lightweight long-range mapping is crucial for enhancing safe navigation. Traditional stereo cameras constrained by a short fixed baseline suffer from limited perception ranges. To overcome this limitation, we present Flying Co-Stereo, a cross-agent collaborative stereo vision system that leverages the wide-baseline spatial configuration of two UAVs for long-range dense mapping. However, realizing this capability presents several challenges. First, the independent motion of each UAV leads to a dynamic and continuously changing stereo baseline, making accurate and robust estimation difficult. Second, efficiently establishing feature correspondences across independently moving viewpoints is constrained by the limited computational capacity of onboard edge devices. To tackle these challenges, we introduce the Flying Co-Stereo system within a novel Collaborative Dynamic-Baseline Stereo Mapping (CDBSM) framework. We first develop a dual-spectrum visual-inertial ranging estimator to achieve robust and precise online estimation of the baseline between the two UAVs. In addition, we propose a hybrid feature association strategy that integrates cross-agent feature matching—based on a computationally intensive yet accurate deep neural network—with intra-agent, optical-flow based lightweight feature tracking. Furthermore, benefiting from the wide baselines between the two UAVs, our system accurately recovers long-range co-visible 3D sparse points. We then employ a monocular depth network to predict up-to-scale dense depth maps, which are refined using accurate metric scales derived from the triangulated sparse points via exponential fitting. Extensive real-world experiments demonstrate that the proposed Flying Co-Stereo system achieves robust and accurate dynamic baseline estimation in complex environments while maintaining efficient feature matching with resource-constrained computers under varying viewpoints. Ultimately, our system achieves dense 3D mapping at distances of up to 70 meters with a relative error between 2.3% and 9.7%. This corresponds to up to a 350% improvement in maximum perception range and up to a 450% increase in coverage area compared to conventional stereo vision systems with fixed compact baselines. Xingxing Zuo 0001, Wei Dong 0008 |
IEEE Trans. Robotics | 2 |
| 2025 | Gaussian-LIC: Real-Time Photo-Realistic SLAM with Gaussian Splatting and LiDAR-Inertial-Camera FusionabstractIn this paper, we present a real-time photo-realistic SLAM method based on marrying Gaussian Splatting with LiDAR-Inertial-Camera SLAM. Most existing radiance-field-based SLAM systems mainly focus on bounded indoor environments, equipped with RGB-D or RGB sensors. However, they are prone to decline when expanding to unbounded scenes or encountering adverse conditions, such as violent motions and changing illumination. In contrast, oriented to general scenarios, our approach additionally tightly fuses LiDAR, IMU, and camera for robust pose estimation and photo-realistic online mapping. To compensate for regions unobserved by the LiDAR, we propose to integrate both the triangulated visual points from images and LiDAR points for initializing 3D Gaussians. In addition, the modeling of the sky and varying camera exposure have been realized for high-quality rendering. Notably, we implement our system purely with C++ and CUDA, and meticulously design a series of strategies to accelerate the online optimization of the Gaussian-based scene representation. Extensive experiments demonstrate that our method outperforms its counterparts while maintaining real-time capability. Impressively, regarding photo-realistic mapping, our method with our estimated poses even surpasses all the compared approaches that utilize privileged ground-truth poses for mapping. Our code will be released on project page https://xingxingzuo.github.io/gaussian_lic. Xiaolei Lang, Laijian Li, Chenming Wu, Chen Zhao 0011, Lina Liu 0010, Yong Liu 0007, Jiajun Lv, Xingxing Zuo 0001 |
ICRA | 8 |
| 2025 | L2COcc: Lightweight Camera-Centric Semantic Scene Completion via Distillation of LiDAR ModelabstractSemantic Scene Completion (SSC) constitutes a pivotal element in autonomous driving perception systems, tasked with inferring the 3D semantic occupancy of a scene from sensory data. To improve accuracy, prior research has implemented various computationally demanding and memory-intensive 3D operations, imposing significant computational requirements on the platform during training and testing. This paper proposes L2COcc, a lightweight camera-centric SSC framework that also accommodates LiDAR inputs. With our proposed efficient voxel transformer (EVT) and cross-modal knowledge modules, including feature similarity distillation (FSD), TPV distillation (TPVD) and prediction alignment distillation (PAD), our method substantially reduce computational burden while maintaining high accuracy. The experimental evaluations demonstrate that our proposed method surpasses the current state-of-the-art vision-based SSC methods regarding accuracy on both the SemanticKITTI and SSCBench-KITTI-360 benchmarks, respectively. Additionally, our method is more lightweight, exhibiting a reduction in both memory consumption and inference time by over 23% compared to the current state-of-the-arts method. Code is available at our project page: https://studyingfufu.github.io/L2COcc/. Yukai Ma, Sheng Tao, Haoang Li, Zongzhi Zhu, Yong Liu 0007, Xingxing Zuo 0001 |
IROS | 8 |
| 2025 | FMGS: Foundation Model Embedded 3D Gaussian Splatting for Holistic 3D Scene Understanding
Xingxing Zuo 0001, Pouya Samangouei, Yunwen Zhou, Yan Di, Mingyang Li 0001 |
Int. J. Comput. Vis. | 1 |
| 2025 | ROEVO: Robust Organized Edge Feature-Based Visual Odometry Using RGB-D CamerasabstractThis work presents a visual odometry (VO) system that leverages image edge features. Edges are spatially expressive cues commonly present across diverse environments, offering rich textural and structural information. However, existing edge-based VO methods often fail to fully exploit this potential. To this end, we introduce a novel feature representation termedorganized edges, which transforms disjoint edge pixels into sequentialized clusters, enabling more effective retention and utilization of the underlying textural and structural information. Another nice property of this formulation is that organized edges can perform edge-level association across multiple frames, enabling the establishment of a co-visibility graph. To achieve precise and efficient pose estimation, we propose a range of particularly designed tracking and joint optimization methods based on the characteristics of organized edges. For tracking, we formulate edge-wise rather than pixel-wise residuals to achieve robust and accurate inter-frame registration. For joint optimization, we introduce a novel shape-preserving edge-fitting method and an organized edge-based Bundle Adjustment (BA) approach, which decomposes the traditional BA problem into fitting and registration to preserve the structural integrity. Based on these novel techniques, we develop a complete VO system that exclusively employs organized edge features, achieving efficient tracking and precise local mapping. Extensive experiments demonstrate its accuracy and robustness in indoor environments, outperforming or achieving comparable performance to state-of-the-art methods. Xingxing Zuo 0001, Renlang Huang, Minglei Zhao, Jiming Chen 0001, Liang Li 0010 |
IEEE Trans. Robotics | 2 |
| 2024 | A Multimodal, Multi-Task Adapting Framework for Video Action RecognitionabstractRecently, the rise of large-scale vision-language pretrained models like CLIP, coupled with the technology of Parameter-Efficient FineTuning (PEFT), has captured substantial attraction in video action recognition. Nevertheless, prevailing approaches tend to prioritize strong supervised performance at the expense of compromising the models' generalization capabilities during transfer. In this paper, we introduce a novel Multimodal, Multi-task CLIP adapting framework named M2-CLIP to address these challenges, preserving both high supervised performance and robust transferability. Firstly, to enhance the individual modality architectures, we introduce multimodal adapters to both the visual and text branches. Specifically, we design a novel visual TED-Adapter, that performs global Temporal Enhancement and local temporal Difference modeling to improve the temporal representation capabilities of the visual encoder. Moreover, we adopt text encoder adapters to strengthen the learning of semantic label information. Secondly, we design a multi-task decoder with a rich set of supervisory signals, including the original contrastive learning head, a cross-modal classification head, a cross-modal masked language modeling head, and a visual classification head. This multi-task decoder adeptly satisfies the need for strong supervised performance within a multimodal framework. Experimental results validate the efficacy of our approach, demonstrating exceptional performance in supervised learning while maintaining strong generalization in zero-shot scenarios. Mengmeng Wang 0005, Jiazheng Xing, Boyuan Jiang, Jun Chen 0023, Jianbiao Mei, Xingxing Zuo 0001, Guang Dai, Jingdong Wang 0001, Yong Liu 0007 |
AAAI | 6 |
| 2024 | Dynamic LiDAR Re-Simulation Using Compositional Neural FieldsabstractWe introduce DyNFL, a novel neural field-based approach for high-fidelity re-simulation of LiDAR scans in dynamic driving scenes. DyNFL processes LiDAR measurements from dynamic environments, accompanied by bounding boxes of moving objects, to construct an editable neural field. This field, comprising separately reconstructed static background and dynamic objects, allows users to modify viewpoints, adjust object positions, and seamlessly add or remove objects in the re-simulated scene. A key innovation of our method is the neural field composition technique, which effectively integrates reconstructed neural assets from various scenes through a ray drop test, accounting for occlusions and transparent surfaces. Our evaluation with both synthetic and real-world environments demonstrates that DyNFL substantially improves dynamic scene LiDAR simulation, offering a combination of physical fidelity and flexible editing capabilities. [project page] Hanfeng Wu, Xingxing Zuo 0001, Stefan Leutenegger, Or Litany, Konrad Schindler |
CVPR | 2 |
| 2024 | Caltech Aerial RGB-Thermal Dataset in the Wild
Connor Lee, Matthew Anderson 0005, Nikhil Ranganathan, Xingxing Zuo 0001, Kevin Do, Georgia Gkioxari, Soon-Jo Chung |
ECCV (63) | 4 |
| 2024 | LaPose: Laplacian Mixture Shape Modeling for RGB-Based Category-Level Object Pose Estimation
Ruida Zhang, Ziqin Huang, Gu Wang 0001, Chenyangguang Zhang, Yan Di, Xingxing Zuo 0001, Jiwen Tang, Xiangyang Ji |
ECCV (25) | 6 |
| 2024 | RadarCam-Depth: Radar-Camera Fusion for Depth Estimation with Learned Metric ScaleabstractWe present a novel approach for metric dense depth estimation based on the fusion of a single-view image and a sparse, noisy Radar point cloud. The direct fusion of heterogeneous Radar and image data, or their encodings, tends to yield dense depth maps with significant artifacts, blurred boundaries, and suboptimal accuracy. To circumvent this issue, we learn to augment versatile and robust monocular depth prediction with the dense metric scale induced from sparse and noisy Radar data. We propose a Radar-Camera framework for highly accurate and fine-detailed dense depth estimation with four stages, including monocular depth prediction, global scale alignment of monocular depth with sparse Radar points, quasi-dense scale estimation through learning the association between Radar points and image patches, and local scale refinement of dense depth using a scale map learner. Our proposed method significantly outperforms the state-of-the-art Radar-Camera depth estimation methods by reducing the mean absolute error (MAE) of depth estimation by 25.6% and 40.2% on the challenging nuScenes dataset and our self-collected ZJU-4DRadarCam dataset, respectively. Our code and dataset will be released at https://github.com/MMOCKING/RadarCam-Depth. Yukai Ma, Yaqing Gu, Kewei Hu, Yong Liu 0007, Xingxing Zuo 0001 |
ICRA | 6 |
| 2024 | Visual-Based Kinematics and Pose Estimation for Skid-Steering RobotsabstractTo build commercial robots, skid-steering mechanical design is of increased popularity due to its manufacturing simplicity and unique mechanism. However, these also cause significant challenges on software and algorithm design, especially for the pose estimation (i.e., determining the robot’s rotation and position) of skid-steering robots, since they change their orientation with an inevitable skid. To tackle this problem, we propose a probabilistic sliding-window estimator dedicated to skid-steering robots, using measurements from a monocular camera, the wheel encoders, and optionally an inertial measurement unit (IMU). Specifically, we explicitly model the kinematics of skid-steering robots by both track instantaneous centers of rotation (ICRs) and correction factors, which are capable of compensating for the complexity of track-to-terrain interaction, the imperfectness of mechanical design, terrain conditions and smoothness, etc. To prevent performance reduction in robots’ long-term missions, the time- and location- varying kinematic parameters are estimated online along with pose estimation states in a tightly-coupled manner. More importantly, we conduct in-depth observability analysis for different sensors and design configurations in this paper, which provides us with theoretical tools in making the correct choice when building real commercial robots. In our experiments, we validate the proposed method by both simulation tests and real-world experiments, which demonstrate that our method outperforms competing methods by wide margins. Note to Practitioners—This paper was motivated by the problem of long-term pose estimation of the commonly commercial-used skid-steering robots with only low-cost sensors. Skid-steering robots change their orientation with a skid, which poses a significant challenge for pose estimation when using the wheel encoders. We propose to online estimate the robot’s kinematics, which succeeds in compensating for the complexity of track-to-terrain interaction, due to the slippage, the imperfectness of mechanical design, terrain conditions and smoothness. It is critical to estimate the kinematics and poses jointly to prevent performance reduction in robots’ long-term missions. We further theoretically analyze whether the kinematics parameters can be estimated under different sensor configurations, and find out the special degrade motions that make the parameters unobservable. Xingxing Zuo 0001, Mingming Zhang 0008, Mengmeng Wang 0005, Yiming Chen 0001, Guoquan Huang 0001, Yong Liu 0007, Mingyang Li 0001 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2024 | RIDERS: Radar-Infrared Depth Estimation for Robust SensingabstractDense depth recovery is crucial in autonomous driving, serving as a foundational element for obstacle avoidance, 3D object detection, and local path planning. Adverse weather conditions, including haze, dust, rain, snow, and darkness, introduce significant challenges to accurate dense depth estimation, thereby posing substantial safety risks in autonomous driving. These challenges are particularly pronounced for traditional depth estimation methods that rely on short electromagnetic wave sensors, such as visible spectrum cameras and near-infrared LiDAR, due to their susceptibility to diffraction noise and occlusion in such environments. To fundamentally overcome this issue, we present a novel approach for robust metric depth estimation by fusing a millimeter-wave radar and a monocular infrared thermal camera, which are capable of penetrating atmospheric particles and unaffected by lighting conditions. Our proposed Radar-Infrared fusion method achieves highly accurate and finely detailed dense depth estimation through three stages, including monocular depth prediction with global scale alignment, quasi-dense radar augmentation by learning radar-pixels correspondences, and local scale refinement of dense depth using a scale map learner. Our method achieves exceptional visual quality and accurate metric estimation by addressing the challenges of ambiguity and misalignment that arise from directly fusing multi-modal long-wave features. We evaluate the performance of our approach on the NTU4DRadLM dataset and our self-collected challenging ZJU-Multispectrum dataset. Especially noteworthy is the unprecedented robustness demonstrated by our proposed method in smoky scenarios. Our code will be released athttps://github.com/MMOCKING/RIDERS. Yukai Ma, Yuehao Huang, Yaqing Gu, Yong Liu 0007, Xingxing Zuo 0001 |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2023 | SimpleMapping: Real-Time Visual-Inertial Dense Mapping with Deep Multi-View StereoabstractWe present a real-time visual-inertial dense mapping method capable of performing incremental 3D mesh reconstruction with high quality using only sequential monocular images and inertial measurement unit (IMU) readings. 6-DoF camera poses are estimated by a robust feature-based visual-inertial odometry (VIO), which also generates noisy sparse 3D map points as a by-product. We propose a sparse point aided multi-view stereo neural network (SPA-MVSNet) that can effectively leverage the informative but noisy sparse points from the VIO system. The sparse depth from VIO is firstly completed by a single-view depth completion network. This dense depth map, although naturally limited in accuracy, is then used as a prior to guide our MVS network in the cost volume generation and regularization for accurate dense depth prediction. Predicted depth maps of keyframe images by the MVS network are incrementally fused into a global map using TSDF-Fusion. We extensively evaluate both the proposed SPA-MVSNet and the entire dense mapping system on several public datasets as well as our own dataset, demonstrating the system’s impressive generalization capabilities and its ability to deliver high-quality 3D reconstruction online. Our proposed dense mapping system achieves a 39.7% improvement in F-score over existing systems when evaluated on the challenging scenarios of the EuRoC dataset. Yingye Xin, Xingxing Zuo 0001, Dongyue Lu, Stefan Leutenegger |
ISMAR | 2 |
| 2023 | High-Quality RGB-D Reconstruction via Multi-View Uncalibrated Photometric Stereo and Gradient-SDFabstractFine-detailed reconstructions are in high demand in many applications. However, most of the existing RGB-D reconstruction methods rely on pre-calculated accurate camera poses to recover the detailed surface geometry, where the representation of a surface needs to be adapted when optimizing different quantities. In this paper, we present a novel multi-view RGB-D based reconstruction method that tackles camera pose, lighting, albedo, and surface normal estimation via the utilization of a gradient signed distance field (gradient-SDF). The proposed method formulates the image rendering process using specific physically-based model(s) and optimizes the surface’s quantities on the actual surface using its volumetric representation, as opposed to other works which estimate surface quantities only near the actual surface. To validate our method, we investigate two physically-based image formation models for natural light and point light source applications. The experimental results on synthetic and real-world datasets demonstrate that the proposed method can recover high-quality geometry of the surface more faithfully than the state-of-the-art and further improves the accuracy of estimated camera poses1. Lu Sang, Bjoern Haefner, Xingxing Zuo 0001, Daniel Cremers |
WACV | 3 |
| 2023 | Online Self-Calibration for Visual-Inertial Navigation: Models, Analysis, and DegeneracyabstractAs sensor calibration plays an important role in visual-inertial sensor fusion, this article performs an in-depth investigation of online self-calibration for robust and accurate visual-inertial state estimation. To this end, we first conduct complete observability analysis for visual-inertial navigation systems (VINS) with full calibration of sensing parameters, including inertial measurement unit (IMU)/camera intrinsics and IMU-camera spatial-temporal extrinsic calibration, along with readout time of rolling shutter (RS) cameras (if used). We study different inertial model variants containing intrinsic parameters that encompass most commonly used models for low-cost inertial sensors. With these models, the observability analysis of linearized VINS with full sensor calibration is performed. Our analysis theoretically proves the intuition commonly assumed in the literature—that is, VINS with full sensor calibration has four unobservable directions, corresponding to the system's global yaw and position, while all sensor calibration parameters are observable given fully excited motions. Moreover, we, for the first time, identify degenerate motion primitives for IMU and camera intrinsic calibration, which, when combined, may produce complex degenerate motions. We compare the proposedonlineself-calibration on commonly used IMUs against the state-of-artofflinecalibration toolbox Kalibr, showing that the proposed system achieves better consistency and repeatability. Based on our analysis and experimental evaluations, we also offer practical guidelines to effectively perform online IMU-camera self-calibration in practice. Patrick Geneva, Xingxing Zuo 0001, Guoquan Huang 0001 |
IEEE Trans. Robotics | 3 |
| 2022 | Symmetry and Uncertainty-Aware Object SLAM for 6DoF Object Pose EstimationabstractWe propose a keypoint-based object-level SLAM framework that can provide globally consistent 6DoF pose estimates for symmetric and asymmetric objects alike. To the best of our knowledge, our system is among the first to utilize the camera pose information from SLAM to provide prior knowledge for tracking keypoints on symmetric objects - ensuring that new measurements are consistent with the current 3D scene. Moreover, our semantic key-point network is trained to predict the Gaussian covariance for the keypoints that captures the true error of the prediction, and thus is not only useful as a weight for the residuals in the system's optimization problems, but also as a means to detect harmful statistical outliers without choosing a manual threshold. Experiments show that our method provides competitive performance to the state of the art in 6DoF object pose estimation, and at a real-time speed. Our code, pre-trained models, and keypoint labels are available https://github.com/rpng/suo_slam. Nathaniel W. Merrill, Yuliang Guo, Xingxing Zuo 0001, Xinyu Huang 0001, Stefan Leutenegger, Liu Ren 0001, Guoquan Huang 0001 |
CVPR | 3 |
| 2022 | Visual-Inertial SLAM with Tightly-Coupled Dropout-Tolerant GPS FusionabstractRobotic applications are continuously striving towards higher levels of autonomy. To achieve that goal, a highly robust and accurate state estimation is indispensable. Combining visual and inertial sensor modalities has proven to yield accurate and locally consistent results in short-term applications. Unfortunately, visual-inertial state estimators suffer from the accumulation of drift for long-term trajectories. To eliminate this drift, global measurements can be fused into the state estimation pipeline. The most known and widely available source of global measurements is the Global Positioning System (GPS). In this paper, we propose a novel approach that fully combines stereo Visual-Inertial Simultaneous Localisation and Mapping (SLAM), including visual loop closures, with the fusion of global sensor modalities in a tightly-coupled and optimisation-based framework. Incorporating measurement uncertainties, we provide a robust criterion to solve the global reference frame initialisation problem. Furthermore, we propose a loop-closure-like optimisation scheme to compensate drift accumulated during outages in receiving GPS signals. Experimental validation on datasets and in a real-world experiment demonstrates the robustness of our approach to GPS dropouts as well as its capability to estimate highly accurate and globally consistent trajectories compared to existing state-of-the-art methods. Simon Boche, Xingxing Zuo 0001, Simon Schaefer, Stefan Leutenegger |
IROS | 2 |
| 2022 | Observability-Aware Intrinsic and Extrinsic Calibration of LiDAR-IMU SystemsabstractAccurate and reliable sensor calibration is essential to fuse LiDAR and inertial measurements, which are usually available in robotic applications. In this article, we propose a novel LiDAR-IMU calibration method within the continuous-time batch-optimization framework, where the intrinsics of both sensors and the spatial-temporal extrinsics between sensors are calibrated without using calibration infrastructure, such as fiducial tags. Compared to discrete-time approaches, the continuous-time formulation has natural advantages for fusing high-rate measurements from LiDAR and IMU sensors. To improve efficiency and address degenerate motions, the following two observability-aware modules are leveraged: first, The information-theoretic data selection policy selectsonlythe most informative segments for calibration during data collection, which significantly improves the calibration efficiency by processing only the selected informative segments. Second, the observability-aware state update mechanism in nonlinear least-squares optimization updatesonlythe identifiable directions in the state space with truncated singular value decomposition, which enables accurate calibration results even under degenerate cases where informative data segments are not available. The proposed LiDAR-IMU calibration approach has been validated extensively in both simulated and real-world experiments with different robot platforms, demonstrating its high accuracy and repeatability in commonly-seen human-made environments. Jiajun Lv, Xingxing Zuo 0001, Kewei Hu, Jinhong Xu, Guoquan Huang 0001, Yong Liu 0007 |
IEEE Trans. Robotics | 2 |
| 2021 | MBA-VO: Motion Blur Aware Visual OdometryabstractMotion blur is one of the major challenges remaining for visual odometry methods. In low-light conditions where longer exposure times are necessary, motion blur can appear even for relatively slow camera motions. In this paper we present a novel hybrid visual odometry pipeline with direct approach that explicitly models and estimates the camera’s local trajectory within the exposure time. This allows us to actively compensate for any motion blur that occurs due to the camera motion. In addition, we also contribute a novel benchmarking dataset for motion blur aware visual odometry. In experiments we show that by directly modeling the image formation process, we are able to improve robustness of the visual odometry, while keeping comparable accuracy as that for images without motion blur. Both the code and the datasets can be found from https://github.com/ethliup/MBA-VO. Peidong Liu 0001, Xingxing Zuo 0001, Viktor Larsson, Marc Pollefeys |
ICCV | 2 |
| 2021 | CodeVIO: Visual-Inertial Odometry with Learned Optimizable Dense DepthabstractIn this work, we present a lightweight, tightly-coupled deep depth network and visual-inertial odometry (VIO) system, which can provide accurate state estimates and dense depth maps of the immediate surroundings. Leveraging the proposed lightweight Conditional Variational Autoencoder (CVAE) for depth inference and encoding, we provide the network with previously marginalized sparse features from VIO to increase the accuracy of initial depth prediction and generalization capability. The compact representation of dense depth, termed depth code, can be updated jointly with navigation states in a sliding window estimator in order to provide the dense local scene geometry. We additionally propose a novel method to obtain the CVAE’s Jacobian which is shown to be more than an order of magnitude faster than previous works, and we additionally leverage First-Estimate Jacobian (FEJ) to avoid recalculation. As opposed to previous works that rely on completely dense residuals, we propose to only provide sparse measurements to update the depth code and show through careful experimentation that our choice of sparse measurements and FEJs can still significantly improve the estimated depth maps. Our full system also exhibits state-of-the-art pose estimation accuracy, and we show that it can run in real-time with single-thread execution while utilizing GPU acceleration only for the network and code Jacobian. Xingxing Zuo 0001, Nathaniel W. Merrill, Wei Li 0111, Yong Liu 0007, Marc Pollefeys, Guoquan Huang 0001 |
ICRA | 1 |
| 2021 | CLINS: Continuous-Time Trajectory Estimation for LiDAR-Inertial SystemabstractIn this paper, we propose a highly accurate continuous-time trajectory estimation framework dedicated to SLAM (Simultaneous Localization and Mapping) applications, which enables fuse high-frequency and asynchronous sensor data effectively. We apply the proposed framework in a 3D LiDAR-inertial system for evaluations. The proposed method adopts a non-rigid registration method for continuous-time trajectory estimation and simultaneously removing the motion distortion in LiDAR scans. Additionally, we propose a two-state continuous-time trajectory correction method to efficiently and efficiently tackle the computationally-intractable global optimization problem when loop closure happens. We examine the accuracy of the proposed approach on several publicly available datasets and the data we collected. The experimental results indicate that the proposed method outperforms the discrete-time methods regarding accuracy especially when aggressive motion occurs. Furthermore, we open source our code at https://github.com/APRIL-ZJU/clins to benefit research community. Jiajun Lv, Kewei Hu, Jinhong Xu, Yong Liu 0007, Xiushui Ma, Xingxing Zuo 0001 |
IROS | 6 |
| 2021 | Pose Estimation for Ground Robots: On Manifold Representation, Integration, Reparameterization, and OptimizationabstractIn this article, we focus on pose estimation dedicated to nonholonomic ground robots with low-cost sensors, by probabilistically fusing measurements from wheel odometers and an exteroceptive sensor. For ground robots, wheel odometers are widely used in pose estimation tasks, especially in applications in planar scenes. However, since wheel odometer only provides two-dimensional (2D) motion measurements, it is extremely challenging to use that for accurate full 6-D pose (3-D position and 3-D orientation) estimation. Traditional methods for 6-D pose estimation with wheel odometers either approximate motion profiles at the cost of accuracy reduction, or rely on other sensors, e.g., inertial measurement unit, to provide complementary measurements. By contrast, we propose a novel motion-manifold-based method for pose estimation of ground robots, which enables to utilize wheel odometers for high-precision 6-D pose estimation. Specifically, the proposed method, first, formulates the motion manifold of ground robots by a parametric representation, second, performs manifold-based 6-D integration with wheel odometer measurements only, and third, reparameterizes manifold representation periodically for error reduction. To demonstrate the effectiveness and applicability of the proposed algorithmic module, we integrate that into a sliding-window pose estimator by using measurements from wheel odometers and a monocular camera. Extensive simulated and real-world experiments are conducted for evaluation, and the proposed algorithm is shown to outperform competing the state-of-the-art algorithms by a significant margin in pose estimation accuracy, especially when deployed in complex, large-scale real-world environments. Mingming Zhang 0008, Xingxing Zuo 0001, Yiming Chen 0001, Yong Liu 0007, Mingyang Li 0001 |
IEEE Trans. Robotics | 2 |
| 2020 | Targetless Calibration of LiDAR-IMU System Based on Continuous-time Batch EstimationabstractSensor calibration is the fundamental block for a multi-sensor fusion system. This paper presents an accurate and repeatable LiDAR-IMU calibration method (termed LI-Calib), to calibrate the 6-DOF extrinsic transformation between the 3D LiDAR and the Inertial Measurement Unit (IMU). Regarding the high data capture rate for LiDAR and IMU sensors, LI-Calib adopts a continuous-time trajectory formulation based on B-Spline, which is more suitable for fusing high-rate or asynchronous measurements than discrete-time based approaches. Additionally, LI-Calib decomposes the space into cells and identifies the planar segments for data association, which renders the calibration problem well-constrained in usual scenarios without any artificial targets. We validate the proposed calibration approach on both simulated and real-world experiments. The results demonstrate the high accuracy and good repeatability of the proposed method in common human-made scenarios. To benefit the research community, we open-source our code at https://github.com/APRIL-ZJU/lidar_IMU_calib. Jiajun Lv, Jinhong Xu, Kewei Hu, Yong Liu 0007, Xingxing Zuo 0001 |
IROS | 5 |
| 2020 | LIC-Fusion 2.0: LiDAR-Inertial-Camera Odometry with Sliding-Window Plane-Feature TrackingabstractMulti-sensor fusion of multi-modal measurements from commodity inertial, visual and LiDAR sensors to provide robust and accurate 6DOF pose estimation holds great potential in robotics and beyond. In this paper, building upon our prior work (i.e., LIC-Fusion), we develop a sliding-window filter based LiDAR-Inertial-Camera odometry with online spatiotemporal calibration (i.e., LIC-Fusion 2.0), which introduces a novel sliding-window plane-feature tracking for efficiently processing 3D LiDAR point clouds. In particular, after motion compensation for LiDAR points by leveraging IMU data, low-curvature planar points are extracted and tracked across the sliding window. A novel outlier rejection criteria is proposed in the plane-feature tracking for high quality data association. Only the tracked planar points belonging to the same plane will be used for plane initialization, which makes the plane extraction efficient and robust. Moreover, we perform the observability analysis for the IMU-LiDAR subsystem under consideration and report the degenerate cases for spatiotemporal calibration using plane features. While the estimation consistency and identified degenerate motions are validated in Monte-Carlo simulations, different real-world experiments are also conducted to show that the proposed LIC-Fusion 2.0 outperforms its predecessor and other state-of-the-art methods. Xingxing Zuo 0001, Patrick Geneva, Jiajun Lv, Yong Liu 0007, Guoquan Huang 0001, Marc Pollefeys |
IROS | 1 |
| 2019 | Tightly-Coupled Aided Inertial Navigation with Point and Plane FeaturesabstractThis paper presents a tightly-coupled aided inertial navigation system (INS) with point and plane features, a general sensor fusion framework applicable to any visual and depth sensor (e.g., RGBD, LiDAR) configuration, in which the camera is used for point feature tracking and depth sensor for plane extraction. The proposed system exploits geometrical structures (planes) of the environments and adopts the closest point (CP) for plane parameterization. Moreover, we distinguish planar point features from non-planar point features in order to enforce point-on-plane constraints which are used in our state estimator, thus further exploiting structural information from the environment. We also introduce a simple but effective plane feature initialization algorithm for feature-based simultaneous localization and mapping (SLAM). In addition, we perform online spatial calibration between the IMU and the depth sensor as it is difficult to obtain this critical calibration parameter in high precision. Both Monte-Carlo simulations and real-world experiments are performed to validate the proposed approach. Patrick Geneva, Xingxing Zuo 0001, Kevin Eckenhoff, Yong Liu 0007, Guoquan Huang 0001 |
ICRA | 3 |
| 2019 | LIC-Fusion: LiDAR-Inertial-Camera OdometryabstractThis paper presents a tightly-coupled multi-sensor fusion algorithm termed LiDAR-inertial-camera fusion (LIC-Fusion), which efficiently fuses IMU measurements, sparse visual features, and extracted LiDAR points. In particular, the proposed LIC-Fusion performs online spatial and temporal sensor calibration between all three asynchronous sensors, in order to compensate for possible calibration variations. The key contribution is the optimal (up to linearization errors) multi-modal sensor fusion of detected and tracked sparse edge/surf feature points from LiDAR scans within an efficient MSCKF-based framework, alongside sparse visual feature observations and IMU readings. We perform extensive experiments in both indoor and outdoor environments, showing that the proposed LIC-Fusion outperforms the state-of-the-art visual-inertial odometry (VIO) and LiDAR odometry methods in terms of estimation accuracy and robustness to aggressive motions. Xingxing Zuo 0001, Patrick Geneva, Woosik Lee 0003, Yong Liu 0007, Guoquan Huang 0001 |
IROS | 1 |
| 2019 | Visual-Inertial Localization for Skid-Steering Robots with Kinematic Constraints
Xingxing Zuo 0001, Mingming Zhang 0008, Yiming Chen 0001, Yong Liu 0007, Guoquan Huang 0001, Mingyang Li 0001 |
ISRR | 1 |
| 2017 | Robust visual SLAM with point and line featuresabstractIn this paper, we develop a robust efficient visual SLAM system that utilizes heterogeneous point and line features. By leveraging ORB-SLAM [1], the proposed system consists of stereo matching, frame tracking, local mapping, loop detection, and bundle adjustment of both point and line features. In particular, as the main theoretical contributions of this paper, we, for the first time, employ the orthonormal representation as the minimal parameterization to model line features along with point features in visual SLAM and analytically derive the Jacobians of the re-projection errors with respect to the line parameters, which significantly improves the SLAM solution. The proposed SLAM has been extensively tested in both synthetic and real-world experiments whose results demonstrate that the proposed system outperforms the state-of-the-art methods in various scenarios. Xingxing Zuo 0001, Xiaojia Xie, Yong Liu 0007, Guoquan Huang 0001 |
IROS | 1 |