Shenghai Yuan 0001

dblp:133/3411 · DBLP profile ↗
← Back
64ranked-venue papers
4as first author
58since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 45 · 4 first-author · 40 since 2021Systems, architecture and hardware · 27 · 2 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 11 since 2021Computer networks · 4 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SplatSSC: Decoupled Depth-Guided Gaussian Splatting for Semantic Scene Completion
abstract
Monocular 3D Semantic Scene Completion (SSC) is a challenging yet promising task that aims to infer dense geometric and semantic descriptions of a scene from a single image. While recent object-centric paradigms significantly improve efficiency by leveraging flexible 3D Gaussian primitives, they still rely heavily on a large number of randomly initialized primitives, which inevitably leads to 1) inefficient primitive initialization and 2) outlier primitives that introduce erroneous artifacts. In this paper, we propose SplatSSC, a novel framework that resolves these limitations with a depth-guided initialization strategy and a principled Gaussian aggregator. Instead of random initialization, SplatSSC utilizes a dedicated depth branch composed of a Group-wise Multi-scale Fusion (GMF) module, which integrates multi-scale image and depth features to generate a sparse yet representative set of initial Gaussian primitives. To mitigate noise from outlier primitives, we develop the Decoupled Gaussian Aggregator (DGA), which enhances robustness by decomposing geometric and semantic predictions during the Gaussian-to-voxel splatting process. Complemented with a specialized Probability Scale Loss, our method achieves state-of-the-art performance on the Occ-ScanNet dataset, outperforming prior approaches by over 6.3% in IoU and 4.1% in mIoU, while reducing both latency and memory cost by more than 9.3%.
Rui Qian 0005, Haozhi Cao, Tianchen Deng, Shenghai Yuan 0001, Lihua Xie 0001
AAAI4
2026 Learning structural consistency and monocular priors for progressive depth completion
Haochen Chai, Yang Lyu, Shenghai Yuan 0001, Meimei Su, Zhunga Liu
Neurocomputing3
2026 Modeling multi-segment acceleration for event camera rotation estimation
Chenyang Shi, Weijun Ding, Ningfang Song, Shenghai Yuan 0001
Pattern Recognit.6
2026 FAM-HRI: Foundation-Model Assisted Multimodal Human-Robot Interaction Combining Gaze and Speech
abstract
Effective Human-Robot Interaction (HRI) is crucial for enhancing accessibility and usability in real-world robotics applications. However, existing solutions often rely on gesture-only or language-only commands, making interaction inefficient and ambiguous, particularly for users with physical impairments. In this paper, we introduce FAM-HRI, an efficient multimodal framework for HRI that integrates language and gaze inputs via foundation models. By leveraging lightweight Meta ARIA glasses, our system captures real-time multimodal signals and utilizes large language models (LLMs) to fuse user intention with scene context, enabling intuitive and precise robot manipulation. Our method accurately determines the gaze fixation time interval, reducing noise caused by the gaze dynamic nature. Experimental evaluations demonstrate that FAM-HRI achieves a high success rate in task execution while maintaining a low interaction time, providing a practical solution for individuals with limited physical mobility or motor impairments. To support the community, we have released our system design, algorithms, and solutions at https://github.com/laiyuzhi/FAM-HRI.
Yuzhi Lai, Shenghai Yuan 0001, Peizheng Li, Benjamin Kiefer, Tianchen Deng, Andreas Zell
IEEE Trans Autom. Sci. Eng.2
2026 NVMS-SLAM: Normal Vector-Based Multi-Session LiDAR SLAM in Indoor Environments
abstract
Multi-session SLAM is essential for long-term robotic operations in indoor environments such as warehouses, office buildings, and industrial facilities. However, the thin walls separating enclosed spaces in such environments introduce a challenge known as the double-sided issue, where point clouds from opposite sides are mistakenly associated as a single surface during single-session mapping, and are prone to being grouped into the same voxel during voxelization in multi-session map fusion, leading to poor voxel planarity, which causes voxel invalidation and reduces the available constraints for global optimization. To address this, we propose NVMS-SLAM, a normal vector-based multi-session LiDAR SLAM system tailored for indoor environments. For single-session mapping, an extended voxel map is designed to preserve normal vector information and to distinguish between primary and secondary surfaces, thereby improving data association. At the multi-session level, a density-encoded indoor scan-context descriptor is introduced for robust loop closure. In addition, a two-stage global map fusion strategy is adopted, combining joint pose graph optimization and normal vector-based bundle adjustment to ensure globally consistent mapping. Experiments on simulated datasets and real-world environments demonstrate that NVMS-SLAM can effectively resolve the double-sided issue at both the single-session and multi-session stages.
Yongxin Ma, Chengwei Zhao 0003, Jie Xu 0066, Xuanxuan Zhang 0002, Shenghai Yuan 0001, Lihua Xie 0001
IEEE Trans Autom. Sci. Eng.6
2026 A Third-Order Gaussian Process Trajectory Representation Framework With Closed-Form Kinematics for Continuous-Time Motion Estimation
abstract
In this paper, we propose a third-order, i.e., white-noise-on-jerk, Gaussian Process (GP) Trajectory Representation (TR) framework for continuous-time (CT) motion estimation (ME) tasks. Our framework features a unified trajectory representation that encapsulates the kinematic models of both SO(3)$times$R3and SE(3) pose representations. This encapsulation strategy allows users to use the same implementation of measurement-based factors for either choice of pose representation, which facilitates experimentation and comparison to make a better choice for the ME task. In addition, unique to our framework, we derive the kinematic models with theclosed-form temporal derivatives of the local variables ofSO(3) and SE(3), which so far has only been approximated based on Taylor expansion in the literature. Our experiments show that these kinematic models can improve the estimation accuracy in high-speed scenarios. All analytical Jacobians of the interpolated states with respect to the support states of the trajectory representation, as well as the motion prior factors, are also provided for accelerated Gauss-Newton (GN) optimization. Our experiments demonstrate the efficacy and efficiency of the framework in various motion estimation tasks such as localization, calibration, and odometry, facilitating fast prototyping for ME researchers. We release the source code for the benefit of the community. Our project is available athttps://github.com/brytsknguyen/gptr.
Thien-Minh Nguyen, Ziyu Cao, Kailai Li 0001, William Talbot, Tongxing Jin, Shenghai Yuan 0001, Tim D. Barfoot, Lihua Xie 0001
IEEE Trans. Robotics6
2025 MNE-SLAM: Multi-Agent Neural SLAM for Mobile Robots
abstract
Neural implicit scene representations have recently shown promising results in dense visual SLAM. However, existing implicit SLAM algorithms are constrained to single-agent scenarios, and fall difficulty in large indoor scenes and long sequences. Existing multi-agent SLAM frameworks cannot meet the constraints of communication bandwidth. To this end, we propose the first distributed multi-agent collaborative SLAM framework with distributed mapping and camera tracking, joint scene representation, intra-to-inter loop closure, and multi-submap fusion. Specifically, our proposed distributed neural mapping and tracking framework only needs peer-to-peer communication, which can greatly improve multi-agent cooperation and communication efficiency. A novel intra-to-inter loop closure method is designed to achieve local (single-agent) and global (multi-agent) consistency. Furthermore, to the best of our knowledge, there is no real-world dataset for NeRF-based/GS-based SLAM that provides both continuous-time trajectories groundtruth and high-accuracy 3D meshes groundtruth. To this end, we propose the first real-world indoor neural slam (INS) dataset covering both single-agent and multi-agent scenarios, ranging from small room to large-scale scenes, with high-accuracy ground truth for both 3D mesh and continuous-time camera trajectory. This dataset can advance the development of the community. Experiments on various datasets demonstrate the superiority of the proposed method in both mapping, tracking, and communication. The dataset and code will be open-source on https://github.com/dtc111111/MNESLAM.
Tianchen Deng, Guole Shen, Chen Xun, Shenghai Yuan 0001, Tongxin Jin, Hongming Shen, Jingchuan Wang, Hesheng Wang 0001, Danwei Wang, Weidong Chen 0001
CVPR4
2025 Audio Array-Based 3D UAV Trajectory Estimation with LiDAR Pseudo-Labeling
abstract
As small unmanned aerial vehicles (UAVs) become increasingly prevalent, there is growing concern regarding their impact on public safety and privacy, highlighting the need for advanced tracking and trajectory estimation solutions. In response, this paper introduces a novel framework that utilizes audio array for 3D UAV trajectory estimation. Our approach incorporates a self-supervised learning model, starting with the conversion of audio data into mel-spectrograms, which are analyzed through an encoder to extract crucial temporal and spectral information. Simultaneously, UAV trajectories are estimated using LiDAR point clouds via unsupervised methods. These LiDAR-based estimations act as pseudo labels, enabling the training of an Audio Perception Network without requiring labeled data. In this architecture, the LiDAR-based system operates as the Teacher Network, guiding the Audio Perception Network, which serves as the Student Network. Once trained, the model can independently predict 3D trajectories using only audio signals, with no need for LiDAR data or external ground truth during deployment. To further enhance precision, we apply Gaussian Process modeling for improved spatiotemporal tracking. Our method delivers toptier performance on the MMAUD dataset, establishing a new benchmark in trajectory estimation using self-supervised learning techniques without reliance on ground truth annotations.
Allen Lei, Tianchen Deng, Han Wang 0001, Jianfei Yang 0001, Shenghai Yuan 0001
ICASSP5
2025 Unsupervised UAV 3D Trajectories Estimation with Sparse Point Clouds
abstract
Compact UAV systems, while advancing delivery and surveillance, pose significant security challenges due to their small size, which hinders detection by traditional methods. This paper presents a cost-effective, unsupervised UAV detection method using spatial-temporal sequence processing to fuse multiple LiDAR scans for accurate UAV tracking in real-world scenarios. Our approach segments point clouds into foreground and background, analyzes spatial-temporal data, and employs a scoring mechanism to enhance detection accuracy. Tested on a public dataset, our solution placed 4th in the CVPR 2024 UG2+ Challenge, demonstrating its practical effectiveness. We plan to open-source all designs, code and sample data for the research community @ github.com/lianghanfang/UnLiDAR-UAV-Est.
Hanfang Liang, Yizhuo Yang 0001, Jinming Hu, Jianfei Yang 0001, Shenghai Yuan 0001
ICASSP6
2025 UAVScenes: A Multi-Modal Dataset for UAVs
Shangshu Yu, Shenghai Yuan 0001, Rui She 0001, Quanjiang Guo, Jinxuan Zheng, Ong Kang Howe, Leonrich Chandra, Shrivarshann Srijeyan, Aditya Sivadas, Toshan Aggarwal, Heyuan Liu, Chujie Chen, Junyu Jiang, Lihua Xie 0001, Wee-Peng Tay
ICCV5
2025 Realm: Real-Time Line-of-Sight Maintenance in Multi-Robot Navigation with Unknown Obstacles
abstract
Multi-robot navigation in complex environments relies on inter-robot communication and mutual observation for situational awareness. This paper studies the multi-robot navigation problem in unknown environments with line-ofsight (LoS) connectivity constraints. While previous works are limited to known environment models to derive the LoS constraints between robots, this paper eliminates such requirements by directly formulating the LoS constraints from realtime LiDAR scans, adopting techniques in point cloud visibility analysis. Based on that, we propose a novel LoS-distance metric to quantify both the urgency and sensitivity of losing LoS between robots considering their potential movements. Moreover, to address the imbalanced urgency of losing LoS between two robots, we design a fusion function to capture the overall urgency while generating gradients that facilitate robots' collaborative behavior to maintain LoS. The team connectivity is guaranteed by encoding the LoS constraints into a potential function that preserves the positivity of the Fiedler eigenvalue of robots' underlying graph. Finally, we establish a LoS-constrained exploration framework integrating the proposed connectivity controller. We showcase its applications in multi-robot exploration in complex unknown environments, where robots can always maintain the LoS connectivity through distributed sensing and communication while collaboratively exploring unknown environments. Our implementations are available at https://github.com/bairuofei/LoS_constrained_navigation.
Ruofei Bai, Shenghai Yuan 0001, Kun Li 0028, Hongliang Guo 0003, Weiyun Yau, Lihua Xie 0001
ICRA2
2025 Swept Volume-Aware Trajectory Planning and MPC Tracking for Multi-Axle Swerve-Drive AMRs
abstract
Multi-axle autonomous mobile robots (AMRs) are set to revolutionize the future of robotics in logistics. As the backbone of next-generation solutions, these robots face a critical challenge: managing and minimizing swept volume during turns while maintaining precise control. Traditional systems designed for standard vehicles often struggle with the complex dynamics of multi-axle configurations, leading to inefficiency and increased safety risk in confined spaces. Our innovative framework overcomes these limitations by combining swept volume minimization with Signed Distance Field (SDF) path planning and model predictive control (MPC) for independent wheel steering. This approach not only plans paths with an awareness of the swept volume, but actively minimizes it in real-time, allowing each axle to follow a precise trajectory while significantly reducing the space the vehicle occupies. By predicting future states and adjusting the turning radius of each wheel, our method enhances both maneuverability and safety, even in the most constrained environments. Unlike previous works, our solution goes beyond basic path calculation and tracking, offering real-time path optimization with minimal swept volume and efficient individual axle control. To our knowledge, this is the first comprehensive approach to tackle these challenges, delivering life-saving improvements in control, efficiency, and safety for multi-axle AMRs. Furthermore, we will open-source our work to foster collaboration and enable others to advance safer and more efficient autonomous systems.
Tianxin Hu, Shenghai Yuan 0001, Ruofei Bai, Xinhang Xu, Yuwen Liao, Lihua Xie 0001
ICRA2
2025 GERA: Geometric Embedding for Efficient Point Registration Analysis
abstract
Point cloud registration aims to provide estimated transformations to align point clouds, which plays a crucial role in pose estimation of various navigation systems, such as surgical guidance systems and autonomous vehicles. Despite the impressive performance of recent models on benchmark datasets, many rely on complex modules like KPConv and Transformers, which impose significant computational and memory demands. These requirements hinder their practical application, particularly in resource-constrained environments such as mobile robotics. In this paper, we propose a novel point cloud registration network that leverages a pure MLP architecture, constructing geometric information offline. This approach eliminates the computational and memory burdens associated with traditional complex feature extractors and significantly reduces inference time and resource consumption. Our method is the first to replace 3D coordinate inputs with offline-constructed geometric encoding, improving generalization and stability, as demonstrated by Maximum Mean Discrepancy (MMD) comparisons. This efficient and accurate geometric representation marks a significant advancement in point cloud analysis, particularly for applications requiring fast and reliability.
Haozhi Cao, Shenghai Yuan 0001, Jianfei Yang 0001
ICRA4
2025 HelmetPoser: A Helmet-Mounted IMU Dataset for Data-Driven Estimation of Human Head Motion in Diverse Conditions
abstract
Helmet-mounted wearable positioning systems are crucial for enhancing safety and facilitating coordination in industrial, construction, and emergency rescue environments. These systems, including LiDAR-Inertial Odometry (LIO) and Visual-Inertial Odometry (VIO), often face challenges in localization due to adverse environmental conditions such as dust, smoke, and limited visual features. To address these limitations, we propose a novel head-mounted Inertial Measurement Unit (IMU) dataset with ground truth, aimed at advancing data-driven IMU pose estimation. Our dataset captures human head motion patterns using a helmet-mounted system, with data from ten participants performing various activities. We explore the application of neural networks, specifically Long Short-Term Memory (LSTM) and Transformer networks, to correct IMU biases and improve localization accuracy. Additionally, we evaluate the performance of these methods across different IMU data window dimensions, motion patterns, and sensor types. We release a publicly available dataset, demonstrate the feasibility of advanced neural network approaches for helmet-based localization, and provide evaluation metrics to establish a baseline for future studies in this field. Data and code can be found at https://lqiutong.github.io/HelmetPoser.github.io/.
Jianping Li 0004, Qiutong Leng, Xinhang Xu, Tongxin Jin, Muqing Cao, Thien-Minh Nguyen, Shenghai Yuan 0001, Kun Cao 0002, Lihua Xie 0001
ICRA8
2025 ULOC: Learning to Localize in Complex Large-Scale Environments with Ultra-Wideband Ranges
abstract
While UWB-based methods can achieve high localization accuracy in small-scale areas, their accuracy and reliability are significantly challenged in large-scale environments. In this paper, we propose a learning-based framework named ULOC for Ultra-Wideband (UWB) based localization in such complex, large-scale environments. First, anchors are deployed in the environment without knowledge of their actual position. Then, UWB observations are collected when the vehicle travels in the environment. At the same time, map-consistent pose estimates are developed from registering onboard self-localization data (from VIO, LIO, and other SLAM methods) with the prior map to provide the training labels. We then propose a network based on MAMBA that learns the ranging patterns of UWBs over a complex, large-scale environment. The experiment demonstrates that our solution can ensure high localization accuracy on a large scale compared to the state-of-the-art. We release our source code to benefit the community at https://github.com/brytsknguyen/uloc.
Thien-Minh Nguyen, Yizhuo Yang 0001, Tien-Dat Nguyen, Shenghai Yuan 0001, Lihua Xie 0001
ICRA4
2025 Large-Scale UWB Anchor Calibration and One-Shot Localization Using Gaussian Process
abstract
Ultra-wideband (UWB) is gaining popularity with devices like AirTags for precise home item localization but faces significant challenges when scaled to large environments like seaports. The main challenges are calibration and localization under obstructed conditions, which are common in logistics environments. Traditional calibration methods, dependent on line-of-sight (LoS), are slow, costly, and unreliable in seaports and warehouses, making large-scale localization a significant pain point in the industry. To overcome these challenges, we propose a one-shot calibration and localization framework based on UWB-LiDAR fusion. Our method uses Gaussian processes to estimate the anchor position from continuous-time LiDAR Inertial Odometry with sampled UWB ranges. This approach ensures accurate and reliable calibration with only one round of sampling in large-scale areas, i.e.,$600 \times 450 ~\mathrm{m}^{2}$. With LoS issues, UWB-only localization can be problematic, even when anchor positions are known. We demonstrate that by applying a UWB-range filter, the search range for LiDAR loop closure descriptors is significantly reduced, improving both accuracy and speed. This concept can be applied to other loop closure detection methods, enabling cost-effective localization in large-scale warehouses and seaports. It significantly improves precision in challenging environments where the UWB-only and LiDAR-Inertial methods fail, as shown in the video https://https://youtu.be/oY8jQKdM7lU. We will open-source our datasets and calibration codes for community use.
Shenghai Yuan 0001, Boyang Lou, Thien-Minh Nguyen, Pengyu Yin, Muqing Cao, Xinghang Xu, Jianping Li 0004, Jie Xu 0066, Siyu Chen 0036, Lihua Xie 0001
ICRA1
2025 BEV-LIO(LC): BEV Image Assisted LiDAR-Inertial Odometry with Loop Closure
abstract
This work introduces BEV-LIO(LC), a novel LiDAR-Inertial Odometry (LIO) framework that combines Bird’s Eye View (BEV) image representations of LiDAR data with geometry-based point cloud registration and incorporates loop closure (LC) through BEV image features. By normalizing point density, we project LiDAR point clouds into BEV images, thereby enabling efficient feature extraction and matching. A lightweight convolutional neural network (CNN) based feature extractor is employed to extract distinctive local and global descriptors from the BEV images. Local descriptors are used to match BEV images with FAST keypoints for reprojection error construction, while global descriptors facilitate loop closure detection. Reprojection error minimization is then integrated with point-to-plane registration within an iterated Extended Kalman Filter (iEKF). In the back-end, global descriptors are used to create a KD-tree-indexed keyframe database for accurate loop closure detection. When a loop closure is detected, Random Sample Consensus (RANSAC) computes a coarse transform from BEV image matching, which serves as the initial estimate for Iterative Closest Point (ICP). The refined transform is subsequently incorporated into a factor graph along with odometry factors, improving the global consistency of localization. Extensive experiments conducted in various scenarios with different LiDAR types demonstrate that BEVLIO(LC) outperforms state-of-the-art methods, achieving competitive localization accuracy. Our code and video can be found at https://github.com/HxCa1/BEV-LIO-LC.
Haoxin Cai, Shenghai Yuan 0001, Jianqi Liu
IROS2
2025 CGS-SLAM: Compact 3D Gaussian Splatting for Dense Visual SLAM
abstract
Recent work has shown that 3D Gaussian-based SLAM enables high-quality reconstruction, accurate pose estimation, and real-time rendering of scenes. However, these approaches are built on a tremendous number of redundant 3D Gaussian ellipsoids, leading to high memory and storage costs and slow training speed. To address this limitation, we propose a compact 3D Gaussian Splatting SLAM system that reduces the number and the parameter size of Gaussian ellipsoids. A sliding window-based masking strategy is first proposed to reduce the redundant ellipsoids. Then, a novel geometry codebook-based quantization method is proposed to further compress 3D Gaussian geometric attributes. Robust and accurate pose estimation is achieved by a local-to-global bundle adjustment method with reprojection loss. Extensive experiments demonstrate that our method achieves faster training, rendering speed, and low memory usage while maintaining the state-of-the-art (SOTA) quality of the scene representation.
Tianchen Deng, Yaohui Chen 0003, Jianfei Yang 0001, Shenghai Yuan 0001, Jiuming Liu, Danwei Wang, Weidong Chen 0001
IROS4
2025 LiMo-Calib: On-Site Fast LiDAR-Motor Calibration for Quadruped Robot-Based Panoramic 3D Sensing System
abstract
Conventional single LiDAR systems are inherently constrained by their limited field of view (FoV), leading to blind spots and incomplete environmental awareness, particularly on robotic platforms with strict payload limitations. Integrating a motorized LiDAR offers a practical solution by significantly expanding the sensor’s FoV and enabling adaptive panoramic 3D sensing. However, the high-frequency vibrations of the quadruped robot introduce calibration challenges: these oscillations continually disturb the LiDAR–motor extrinsics, so parameters calibrated once may drift during operation and degrade sensing accuracy.Existing calibration methods that use artificial targets or dense feature extraction lack feasibility for on-site applications and real-time implementation. To overcome these limitations, we propose LiMo-Calib, an efficient on-site calibration method that eliminates the need for external targets by leveraging geometric features directly from raw LiDAR scans. LiMo-Calib optimizes feature selection based on normal distribution to accelerate convergence while maintaining accuracy and incorporates a reweighting mechanism that evaluates local plane fitting quality to enhance robustness. We integrate and validate the proposed method on a motorized LiDAR system mounted on a quadruped robot, demonstrating significant improvements in calibration efficiency and 3D sensing accuracy, making LiMo-Calib well-suited for real-world robotic applications. We further demonstrate the accuracy improvements of the Lidar Inertial Odometry (LIO) on the panoramic 3D sensing system using the calibrated parameters. The code will be available at: https://github.com/kafeiyin00/LiMo-Calib.
Jianping Li 0004, Zhongyuan Liu, Xinhang Xu, Xiong Qin, Shenghai Yuan 0001, Lihua Xie 0001
IROS6
2025 AirSwarm: Enabling Cost-Effective Multi-UAV Research with COTS drones
abstract
Traditional unmanned aerial vehicle (UAV) swarm missions rely heavily on expensive custom-made drones with onboard perception or external positioning systems, limiting their widespread adoption in research and education. To address this issue, we propose AirSwarm. AirSwarm democratizes multi-drone coordination using low-cost commercially available drones such as Tello or Anafi, enabling affordable swarm aerial robotics research and education. Key innovations include a hierarchical control architecture for reliable multi-UAV coordination, an infrastructure-free visual SLAM system for precise localization without external motion capture, and a ROS-based software framework for simplified swarm development. Experiments demonstrate cm-level tracking accuracy, low-latency control, communication failure resistance, formation flight, and trajectory tracking. By reducing financial and technical barriers, AirSwarm makes multi-robot education and research more accessible. The complete instructions and open source code will be available at https://github.com/vvEverett/tello_ros.
Ruofei Bai, Shenghai Yuan 0001, Lihua Xie 0001
IROS5
2025 Underwater target 6D State Estimation via UUV Attitude Enhance Observability
abstract
Accurate relative state observation of Unmanned Underwater Vehicles (UUVs) for tracking uncooperative targets remains a significant challenge due to the absence of GPS, complex underwater dynamics, and sensor limitations. Existing localization approaches rely on either global positioning infrastructure or multi-UUV collaboration, both of which are impractical for a single UUV operating in large or unknown environments. To address this, we propose a novel persistent relative 6D state estimation framework that enables a single UUV to estimate its relative motion to a non-cooperative target using only successive noisy range measurements from two monostatic sonar sensors. Our key contribution is an observability-enhanced attitude control strategy, which optimally adjusts the UUV’s orientation to improve the observability of relative state estimation using a Kalman filter, effectively mitigating the impact of sensor noise and drift accumulation. Additionally, we introduce a rigorously proven Lyapunov-based tracking control strategy that guarantees long-term stability by ensuring that the UUV maintains an optimal measurement range, preventing localization errors from diverging over time. Through theoretical analysis and simulations, we demonstrate that our method significantly improves 6D relative state estimation accuracy and robustness compared to conventional approaches. This work provides a scalable, infrastructure-free solution for UUVs tracking uncooperative targets underwater.
Chengfeng Jia, Shenghai Yuan 0001, Rong Su 0001
IROS4
2025 Autonomous 3D Moving Target Encirclement and Interception with Range Measurement
abstract
Commercial UAVs are an emerging security threat as they are capable of carrying hazardous payloads or disrupting air traffic. To counter UAVs, we introduce an autonomous 3D target encirclement and interception strategy. Unlike traditional ground-guided systems, this strategy employs autonomous drones to track and engage non-cooperative hostile UAVs, which is effective in non-line-of-sight conditions, GPS denial, and radar jamming, where conventional detection and neutralization from ground guidance fail. Using two noisy real-time distances measured by drones, guardian drones estimate the relative position from their own to the target using observation and velocity compensation methods, based on anti-synchronization (AS) and an X−Y circular motion combined with vertical jitter. An encirclement control mechanism is proposed to enable UAVs to adaptively transition from encircling and protecting a target to encircling and monitoring a hostile target. Upon breaching a warning threshold, the UAVs may even employ a suicide attack to neutralize the hostile target. We validate this strategy through real-world UAV experiments and simulated analysis in MATLAB, demonstrating its effectiveness in detecting, encircling, and intercepting hostile drones. More details: https://youtu.be/5eHW56lPVto.
Shenghai Yuan 0001, Thien-Minh Nguyen, Rong Su 0001
IROS2
2025 QLIO: Quantized LiDAR-Inertial Odometry
abstract
LiDAR-Inertial Odometry (LIO) is widely used for autonomous navigation, but its deployment on Size, Weight, and Power (SWaP)-constrained platforms remains challenging due to the computational cost of processing dense point clouds. Conventional LIO frameworks rely on a single onboard processor, leading to computational bottlenecks and high memory demands, making real-time execution difficult on embedded systems. To address this, we propose QLIO, a multi-processor distributed quantized LIO framework that reduces computational load and bandwidth consumption while maintaining localization accuracy. QLIO introduces a quantized state estimation pipeline, where a co-processor pre-processes LiDAR measurements, compressing point-to-plane residuals before transmitting only essential features to the host processor. Additionally, an rQ-vector-based adaptive resampling strategy intelligently selects and compresses key observations, further reducing computational redundancy. Real-World evaluations demonstrate that QLIO achieves a 14.1× reduction in perobservation residual data while preserving localization accuracy. Furthermore, we release an open-source implementation to facilitate further research and real-world deployment. These results establish QLIO as an efficient and scalable solution for real-time autonomous systems operating under computational and bandwidth constraints.
Boyang Lou, Shenghai Yuan 0001, Jianfei Yang 0001, Wenju Su, Yingjian Zhang, Enwen Hu
IROS2
2025 DPGP: A Hybrid 2D-3D Dual Path Potential Ghost Probe Zone Prediction Framework for Safe Autonomous Driving
abstract
Modern robots must coexist with humans in dense urban environments. A key challenge is the ghost probe problem, where pedestrians or objects unexpectedly rush into traffic paths. This issue affects both autonomous vehicles and human drivers. Existing works propose vehicle-to-everything (V2X) strategies and non-line-of-sight (NLOS) imaging for ghost probe zone detection. However, most require high computational power or specialized hardware, limiting real-world feasibility. Additionally, many methods do not explicitly address this issue. To tackle this, we propose DPGP, a hybrid 2D-3D fusion framework for ghost probe zone prediction using only a monocular camera during training and inference. With unsupervised depth prediction, we observe ghost probe zones align with depth discontinuities, but different depth representations offer varying robustness. To exploit this, we fuse multiple feature embeddings to improve prediction. To validate our approach, we created a 12K-image dataset annotated with ghost probe zones, carefully sourced and cross-checked for accuracy. Experimental results show our framework outperforms existing methods while remaining cost-effective. To our knowledge, this is the first work extending ghost probe zone prediction beyond vehicles, addressing diverse non-vehicle objects. We will open-source our code and dataset for community benefit.
Weiming Qu, Shenghai Yuan 0001, Shengyi Liu, Yuanhao Zhu, Jiayi Rao, Xihong Wu, Dingsheng Luo
IROS3
2025 NVP-HRI: Zero shot natural voice and posture-based human-robot interaction via large language model
abstract
Effective Human–Robot Interaction (HRI) is crucial for future service robots in aging societies . Existing solutions are biased towards only well-trained objects, creating a gap when dealing with new objects. Currently, HRI systems using predefined gestures or language tokens for pretrained objects pose challenges for all individuals, especially elderly ones. These challenges include difficulties in recalling commands, memorizing hand gestures, and learning new names. This paper introduces NVP-HRI, an intuitive multi-modal HRI paradigm that combines voice commands and deictic posture. NVP-HRI utilizes the Segment Anything Model (SAM) to analyze visual cues and depth data, enabling precise structural object representation. Through a pre-trained SAM network, NVP-HRI allows interaction with new objects via zero-shot prediction, even without prior knowledge . NVP-HRI also integrates with a large language model (LLM) for multimodal commands, coordinating them with object selection and scene distribution in real time for collision-free trajectory solutions. We also regulate the action sequence with the essential control syntax to reduce LLM hallucination risks. The evaluation of diverse real-world tasks using a Universal Robot showcased up to 59.2% efficiency improvement over traditional gesture control, as illustrated in the video https://youtu.be/EbC7al2wiAc . Our code and design will be openly available at https://github.com/laiyuzhi/NVP-HRI.git .
Yuzhi Lai, Shenghai Yuan 0001, Youssef Nassar, Mingyu Fan, Matthias Rätsch
Expert Syst. Appl.2
2025 Adaptive-LIO: Enhancing Robustness and Precision Through Environmental Adaptation in LiDAR Inertial Odometry
abstract
The emerging Internet of Things (IoT) applications, such as driverless cars, have a growing demand for high-precision positioning and navigation. Nowadays, LiDAR inertial odometry (LIO) becomes increasingly prevalent in robotics and autonomous driving. However, many current SLAM systems lack sufficient adaptability to various scenarios. Challenges include decreased point cloud accuracy with longer frame intervals under the constant velocity assumption, coupling of erroneous IMU information when IMU saturation occurs, and decreased localization accuracy due to the use of fixed-resolution maps during indoor-outdoor scene transitions. To address these issues, we propose a loosely coupled adaptive LIO named Adaptive-LIO, which incorporates adaptive segmentation to enhance mapping accuracy, adapts motion modality through IMU saturation and fault detection, and adjusts map resolution adaptively using multiresolution voxel maps based on the distance from the LiDAR center. Our proposed method has been tested in various challenging scenarios, demonstrating the effectiveness of the improvements we introduce. The code is open-source on GitHub: Adaptive-LIO.
Chengwei Zhao 0003, Kun Hu 0016, Jie Xu 0066, Lijun Zhao 0003, Baiwen Han, Kaidi Wu, Maoshan Tian, Shenghai Yuan 0001
IEEE Internet Things J.8
2025 Incremental Joint Learning of Depth, Pose, and Implicit Scene Representation on Monocular Camera in Large-Scale Scenes
abstract
Dense scene reconstruction for photo-realistic view synthesis has various applications, such as VR/AR, and robotics navigation. Existing dense reconstruction methods are primarily designed for small room scenarios, but in practice, the scenes encountered by robots are typically large-scale environments. Most existing methods have difficulties in large-scale scenes due to three core challenges:(a) inaccurate depth input. Depth information is crucial for both scene geometry reconstruction and pose estimation. Accurate depth input is impossible to get in real-world large-scale scenes.(b) inaccurate pose estimation. Existing methods are not robust enough with the growth of cumulative errors in large scenes and long sequences.(c) insufficient scene representation capability. A single global radiance field lacks the capacity to scale effectively to large-scale scenes. To this end, we propose an incremental joint learning framework, which can achieve accurate depth, pose estimation, and large-scale dense scene reconstruction. For depth estimation, a vision transformer-based network is adopted as the backbone to enhance performance in scale information estimation. For pose estimation, a feature-metric bundle adjustment (FBA) method is designed for accurate and robust camera tracking in large-scale scenes and eliminates pose drift. In terms of implicit scene representation, we propose an incremental scene representation method to construct the entire large-scale scene as multiple local radiance fields to enhance the scalability of 3D scene representation. In local radiance fields, we propose a tri-plane based scene representation method to further improve the accuracy and efficiency of scene reconstruction. We conduct extensive experiments on various datasets, including our own collected data, to demonstrate the effectiveness and accuracy of our method in depth estimation, pose estimation, and large-scale scene reconstruction. The code has been open-sourced on https://github.com/dtc111111/incre-dpsr.
Tianchen Deng, Nailin Wang, Chongdi Wang, Shenghai Yuan 0001, Jingchuan Wang, Hesheng Wang 0001, Danwei Wang, Weidong Chen 0001
IEEE Trans Autom. Sci. Eng.4
2025 Graph Optimality-Aware Stochastic LiDAR Bundle Adjustment With Progressive Spatial Smoothing
abstract
Large-scale LiDAR Bundle Adjustment (LBA) to refine sensor orientation and point cloud accuracy simultaneously for building navigation maps is a fundamental task in logistics, intelligent transportation, and robotics. In the context of autonomous delivery and smart mobility, the 3D map obtained by accurate and robust LBA plays a pivotal role in enabling reliable localization and navigation across complex, large-scale urban environments. Unlike pose-graph-based methods that rely solely on pairwise relationships between LiDAR frames, LBA leverages raw LiDAR correspondences to achieve more precise results, especially when initial pose estimates are unreliable for low-cost sensors. However, existing LBA methods face challenges such as simplistic planar correspondences, extensive observations, and dense normal matrices in the least-squares problem, which limit robustness, efficiency, and scalability. To address these issues, we propose a Graph Optimality-aware Stochastic Optimization scheme with Progressive Spatial Smoothing, namely PSS-GOSO, to achieverobust,efficient, andscalableLBA. The Progressive Spatial Smoothing (PSS) module extractsrobustLiDAR feature association exploiting the prior structure information obtained by the polynomial smooth kernel. The Graph Optimality-aware Stochastic Optimization (GOSO) module first sparsifies the graph according to optimality for anefficientoptimization. GOSO then utilizes stochastic clustering and graph marginalization to solve the large-scale state estimation problem for ascalableLBA. We validate PSS-GOSO across diverse scenes captured by various platforms, demonstrating its superior performance compared to existing methods. Moreover, the resulting point cloud maps are used for automatic last-mile delivery in large-scale complex scenes, showcasing the practical benefits of our method in modern intelligent transportation systems. The project page can be found at:https://kafeiyin00.github.io/PSS-GOSO/
Jianping Li 0004, Thien-Minh Nguyen, Muqing Cao, Shenghai Yuan 0001, Tzu-Yi Hung, Lihua Xie 0001
IEEE Trans. Intell. Transp. Syst.4
2025 Relative Localizability and Localization for Multirobot Systems
abstract
Inter-robot relative positions are crucial for executing various multirobot missions, such as formation maneuvering and collaborative inspection. However, the current sensing technology usually provides part of relative position information, such as inter-robot distances, bearings and angles. This prompts the study of determining inter-robot relative positions, i.e., relative localization, from these partial measurements. Based on the existing results of static networks' localizability and mobile robots' relative localization, we propose a novel concept,relative localizabilityto describe whether a multirobot system isrelatively localizable. Given each robot's self-displacement measurements and inter-robot partial measurements in$d$($d\leq 4$) sampling instants, we show that a multirobot system's relative localization can be achieved in a purelyalgebraicanddistributedmanner, in which the multirobot system is said to be$d$-step relatively localizable. To make the results more general, we consider that the multirobot system consists of landmarks, leaders, and followers, and that the inter-robot measurements can be distances, bearings or angles. When robots' coordinate frames have different orientations, we show that the given local measurements can be used to determine robots' relative positions and their coordinate frames' relative orientations simultaneously. Simulations and experiments of relative localization for ground robots are conducted to validate the obtained results.
Liangming Chen, Chenyang Liang, Shenghai Yuan 0001, Muqing Cao, Lihua Xie 0001
IEEE Trans. Robotics3
2025 AirSLAM: An Efficient and Illumination-Robust Point-Line Visual SLAM System
abstract
In this article, we present an efficient visual simultaneous localization and mapping (SLAM) system designed to tackle both short-term and long-term illumination challenges. Our system adopts a hybrid approach that combines deep learning techniques for feature detection and matching with traditional back-end optimization methods. Specifically, we propose a unified convolutional neural network that simultaneously extracts keypoints and structural lines. These features are then associated, matched, triangulated, and optimized in a coupled manner. In addition, we introduce a lightweight relocalization pipeline that reuses the built map, where keypoints, lines, and a structure graph are used to match the query frame with the map. To enhance the applicability of the proposed system to real-world robots, we deploy and accelerate the feature detection and matching networks using C++ and NVIDIA TensorRT. Extensive experiments conducted on various datasets demonstrate that our system outperforms other state-of-the-art visual SLAM systems in illumination-challenging environments. Efficiency evaluations show that our system can run at a rate of$73\,\mathrm{Hz}$on a PC and$40\,\mathrm{Hz}$on an embedded platform.
Yuefan Hao, Shenghai Yuan 0001, Chen Wang 0033, Lihua Xie 0001
IEEE Trans. Robotics3
2024 MCD: Diverse Large-Scale Multi-Campus Dataset for Robot Perception
abstract
Perception plays a crucial role in various robot applications. However, existing well-annotated datasets are biased towards autonomous driving scenarios, while unlabelled SLAM datasets are quickly over-fitted, and often lack environment and domain variations. To expand the frontier of these fields, we introduce a comprehensive dataset named MCD (Multi-Campus Dataset), featuring a wide range of sensing modalities, high-accuracy ground truth, and diverse challenging environments across three Eurasian university campuses. MCD comprises both CCS (Classical Cylindrical Spinning) and NRE (Non-Repetitive Epicyclic) lidars, high-quality IMUs (Inertial Measurement Units), cameras, and UWB (Ultra-WideBand) sensors. Further-more, in a pioneering effort, we introduce semantic annotations of 29 classes over 59k sparse NRE lidar scans across three domains, thus providing a novel challenge to existing semantic segmentation research upon this largely unexplored modality. Finally, we propose, for the first time to the best of our knowledge, continuous-time ground truth based on optimization-based registration of lidar-inertial data on three survey-grade prior maps, each several times larger than the next largest publicly available ones. We conduct a rigorous evaluation of numerous state-of-the-art algorithms on MCD, report their performance, and highlight the challenges awaiting solutions from the research community.
Thien-Minh Nguyen, Shenghai Yuan 0001, Thien Hoang Nguyen, Pengyu Yin, Haozhi Cao, Lihua Xie 0001, Maciej Wozniak 0001, Patric Jensfelt, Marko Thiel 0002, Justin Ziegenbein, Noel Blunder
CVPR2
2024 Reliable Spatial-Temporal Voxels For Multi-modal Test-Time Adaptation
Haozhi Cao, Yuecong Xu, Jianfei Yang 0001, Pengyu Yin, Xingyu Ji, Shenghai Yuan 0001, Lihua Xie 0001
ECCV (28)6
2024 MoPA: Multi-Modal Prior Aided Domain Adaptation for 3D Semantic Segmentation
abstract
Multi-modal unsupervised domain adaptation (MM-UDA) for 3D semantic segmentation is a practical solution to embed semantic understanding in autonomous systems without expensive point-wise annotations. While previous MM-UDA methods can achieve overall improvement, they suffer from significant class-imbalanced performance, restricting their adoption in real applications. This imbalanced performance is mainly caused by: 1) self-training with imbalanced data and 2) the lack of pixel-wise 2D supervision signals. In this work, we propose Multi-modal Prior Aided (MoPA) domain adaptation to improve the performance of rare objects. Specifically, we develop Valid Ground-based Insertion (VGI) to rectify the imbalance supervision signals by inserting prior rare objects collected from the wild while avoiding introducing artificial artifacts that lead to trivial solutions. Meanwhile, our SAM consistency loss leverages the 2D prior semantic masks from SAM as pixel-wise supervision signals to encourage consistent predictions for each object in the semantic mask. The knowledge learned from modal-specific prior is then shared across modalities to achieve better rare object segmentation. Extensive experiments show that our method achieves state-of-the-art performance on the challenging MM-UDA benchmark. Code will be available at https://github.com/AronCao49/MoPA.
Haozhi Cao, Yuecong Xu, Jianfei Yang 0001, Pengyu Yin, Shenghai Yuan 0001, Lihua Xie 0001
ICRA5
2024 Jacquard V2: Refining Datasets using the Human In the Loop Data Correction Method
abstract
In the context of rapid advancements in industrial automation, vision-based robotic grasping plays an increasingly crucial role. In order to enhance visual recognition accuracy, the utilization of large-scale datasets is imperative for training models to acquire implicit knowledge related to the handling of various objects. Creating datasets from scratch is a time and labor-intensive process. Moreover, existing datasets often contain errors due to automated annotations aimed at expediency, making the improvement of these datasets a substantial research challenge. Consequently, several issues have been identified in the annotation of grasp bounding boxes within the popular Jacquard Grasp Dataset [1]. We propose utilizing a Human-In-The-Loop(HIL) method to enhance dataset quality. This approach relies on backbone deep learning networks to predict object positions and orientations for robotic grasping. Predictions with Intersection over Union (IOU) values below 0.2 undergo an assessment by human operators. After their evaluation, the data is categorized into False Negatives(FN) and True Negatives(TN). FN are then subcategorized into either missing annotations or catastrophic labeling errors. Images lacking labels are augmented with valid grasp bounding box information, whereas images afflicted by catastrophic labeling errors are completely removed. The open-source tool Labelbee was employed for 53,026 iterations of HIL dataset enhancement, leading to the removal of 2,884 images and the incorporation of ground truth information for 30,292 images. The enhanced dataset, named the Jacquard V2 Grasping Dataset, served as the training data for a range of neural networks. We have empirically demonstrated that these dataset improvements significantly enhance the training and prediction performance of the same network, resulting in an increase of 7.1% across most popular detection architectures for ten iterations. This refined dataset will be accessible on One Drive and Baidu Netdisk, while the associated tools, source code, and benchmarks will be made available on GitHub (https://github.com/lqh12345/Jacquard_V2).
Qiuhao Li, Shenghai Yuan 0001
ICRA2
2024 Outram: One-shot Global Localization via Triangulated Scene Graph and Global Outlier Pruning
abstract
One-shot LiDAR localization refers to the ability to estimate the robot pose from one single point cloud, which yields significant advantages in initialization and relocalization processes. In the point cloud domain, the topic has been extensively studied as a global descriptor retrieval (i.e., loop closure detection) and pose refinement (i.e., point cloud registration) problem both in isolation or combined. However, few have explicitly considered the relationship between candidate retrieval and correspondence generation in pose estimation, leaving them brittle to substructure ambiguities. To this end, we propose a hierarchical one-shot localization algorithm called Outram that leverages substructures of 3D scene graphs for locally consistent correspondence searching and global substructure-wise outlier pruning. Such a hierarchical process couples the feature retrieval and the correspondence extraction to resolve the substructure ambiguities by conducting a local-to-global consistency refinement. We demonstrate the capability of Outram in a variety of scenarios in multiple large-scale outdoor datasets. Our implementation is open-sourced: https://github.com/Pamphlett/Outram.
Pengyu Yin, Haozhi Cao, Thien-Minh Nguyen, Shenghai Yuan 0001, Kangcheng Liu, Lihua Xie 0001
ICRA4
2024 MMAUD: A Comprehensive Multi-Modal Anti-UAV Dataset for Modern Miniature Drone Threats
abstract
In response to the evolving challenges posed by small unmanned aerial vehicles (UAVs), which possess the potential to transport harmful payloads or independently cause damage, we introduce MMAUD: a comprehensive Multi-Modal Anti-UAV Dataset. MMAUD addresses a critical gap in contemporary threat detection methodologies by focusing on drone detection, UAV-type classification, and trajectory estimation. MMAUD stands out by combining diverse sensory inputs, including stereo vision, various Lidars, Radars, and audio arrays. It offers a unique overhead aerial detection vital for addressing real-world scenarios with higher fidelity than datasets captured on specific vantage points using thermal and RGB. Additionally, MMAUD provides accurate Leica-generated ground truth data, enhancing credibility and enabling confident refinement of algorithms and models, which has never been seen in other datasets. Most existing works do not disclose their datasets, making MMAUD an invaluable resource for developing accurate and efficient solutions. Our proposed modalities are cost-effective and highly adaptable, allowing users to experiment and implement new UAV threat detection tools. Our dataset closely simulates real-world scenarios by incorporating ambient heavy machinery sounds. This approach enhances the dataset’s applicability, capturing the exact challenges faced during proximate vehicular operations. It is expected that MMAUD can play a pivotal role in advancing UAV threat detection, classification, trajectory estimation capabilities, and beyond. Our dataset, codes, and designs will be available in https://ntu-aris.github.io/MMAUD.
Shenghai Yuan 0001, Yizhuo Yang 0001, Thien Hoang Nguyen, Thien-Minh Nguyen, Jianfei Yang 0001, Jianping Li 0004, Han Wang 0001, Lihua Xie 0001
ICRA1
2024 PSS-BA: LiDAR Bundle Adjustment with Progressive Spatial Smoothing
abstract
Accurate and consistent construction of point clouds from LiDAR scanning data is fundamental for 3D modeling applications. Current solutions, such as multiview point cloud registration and LiDAR bundle adjustment, predominantly depend on the local plane assumption, which may be inadequate in complex environments lacking of planar geometries or substantial initial pose errors. To mitigate this problem, this paper presents a LiDAR bundle adjustment with progressive spatial smoothing, which is suitable for complex environments and exhibits improved convergence capabilities. The proposed method consists of a spatial smoothing module and a pose adjustment module, which combines the benefits of local consistency and global accuracy. With the spatial smoothing module, we can obtain robust and rich surface constraints employing smoothing kernels across various scales. Then the pose adjustment module corrects all poses utilizing the novel surface constraints. Ultimately, the proposed method simultaneously achieves fine poses and parametric surfaces that can be directly employed for high-quality point cloud reconstruction. The effectiveness and robustness of our proposed approach have been validated on both simulation and real-world datasets. The experimental results demonstrate that the proposed method outperforms the existing methods and achieves better accuracy in complex environments with low planar structures.
Jianping Li 0004, Thien-Minh Nguyen, Shenghai Yuan 0001, Lihua Xie 0001
IROS3
2024 Multi-Robot Active Graph Exploration with Reduced Pose-SLAM Uncertainty via Submodular Optimization
abstract
This paper considers the multi-robot active graph exploration problem, where robots need to collaboratively cover a graph environment while maintaining reliable pose estimation in collaborative Simultaneous Localization and Mapping (SLAM). Considering both objectives presents challenges for multi-robot pathfinding, as it involves the expensive covariance propagation for SLAM uncertainty evaluation, especially when considering various combinations of robots’ paths. To reduce the computational complexity, we propose an efficient two-stage strategy where exploration paths are first generated for quick coverage, and then enhanced by adding informative loop-closing actions along the paths for reliable pose estimation. We formulate the latter problem as a non-monotone submodular maximization problem by relating SLAM uncertainty with pose graph topology, which (1) facilitates a more efficient evaluation of SLAM uncertainty than covariance inference, and (2) allows the employment of approximation algorithms in submodular optimization to provide suboptimality guarantees. We further introduce ordering heuristics to improve the objective values while preserving the optimality bound. Simulation experiments over randomly generated graph environments verify the effectiveness of our methods to achieve quick coverage and enhanced pose graph reliability, and benchmark the performance of the approximation algorithms and the greedy-based algorithm in the loop edge selection problem. Our implementations will be open-source at https://github.com/bairuofei/CGE.
Ruofei Bai, Shenghai Yuan 0001, Hongliang Guo 0003, Pengyu Yin, Weiyun Yau, Lihua Xie 0001
IROS2
2024 I2EKF-LO: A Dual-Iteration Extended Kalman Filter Based LiDAR Odometry
abstract
LiDAR odometry is a pivotal technology in the fields of autonomous driving and autonomous mobile robotics. However, most of the current works focus on nonlinear optimization methods, and still existing many challenges in using the traditional Iterative Extended Kalman Filter (IEKF) framework to tackle the problem: IEKF only iterates over the observation equation, relying on a rough estimate of the initial state, which is insufficient to fully eliminate motion distortion in the input point cloud; the system process noise is difficult to be determined during state estimation of the complex motions; and the varying motion models across different sensor carriers. To address these issues, we propose the Dual-Iteration Extended Kalman Filter (I2EKF) and the LiDAR odometry based on I2EKF (I2EKF-LO). This approach not only iterates over the observation equation but also leverages state updates to iteratively mitigate motion distortion in LiDAR point clouds. Moreover, it dynamically adjusts process noise based on the confidence level of prior predictions during state estimation and establishes motion models for different sensor carriers to achieve accurate and efficient state estimation. Comprehensive experiments demonstrate that I2EKF-LO achieves outstanding levels of accuracy and computational efficiency in the realm of LiDAR odometry. Additionally, to foster community development, our code is open-sourced.1
Wenlu Yu, Jie Xu 0066, Chengwei Zhao 0003, Lijun Zhao 0003, Thien-Minh Nguyen, Shenghai Yuan 0001, Mingming Bai, Lihua Xie 0001
IROS6
2024 M-DIVO: Multiple ToF RGB-D Cameras-Enhanced Depth-Inertial-Visual Odometry
abstract
Time-of-Flight (ToF) RGB-D cameras provide a wealth of information for SLAM systems. However, the limited field of view (FOV) of a single ToF RGB-D camera and the small range of its depth measurement module make it prone to degeneracy when relying solely on visual or depth information for SLAM, a problem typical of unimodal SLAM algorithms. To address this issue, this article presents M-DIVO: an IEKF-based odometry that fuses visual, depth (similar to LiDAR), and inertial modules from multiple ToF RGB-D cameras. It comprises two direct method subsystems: 1) the depth–inertial odometry (DIO) subsystem, which constructs point-to-plane constraints from multiple depth modules and 2) the visual–inertial odometry (VIO) subsystem, which optimizes pose using photometric error constructed by multiple cameras. Additionally, to manage the significant computational load from processing multiple sensors and multimodal information, we introduce a multimodal redundancy scheduling mechanism (MRSM): prioritizing the DIO subsystem with the VIO subsystem as auxiliary, executing the VIO subsystem only when degeneracy occurs in the DIO subsystem. We also propose a “External First, Internal Last” strategy for calibrating multiple external and internal sensors. Experiments demonstrate that compared to unimodal SLAM, our method achieves higher robustness and precision, as well as satisfactory real-time performance. The proposed calibration strategy is demonstrated to be more accurate than the traditional inertial measurement unit-centric approach. The code is open source.
Jie Xu 0066, Wenlu Yu, Shenghai Yuan 0001, Lijun Zhao 0003, Ruifeng Li 0001, Lihua Xie 0001
IEEE Internet Things J.4
2023 Multi-Modal Continual Test-Time Adaptation for 3D Semantic Segmentation
abstract
Continual Test-Time Adaptation (CTTA) generalizes conventional Test-Time Adaptation (TTA) by assuming that the target domain is dynamic over time rather than stationary. In this paper, we explore Multi-Modal Continual Test-Time Adaptation (MM-CTTA) as a new extension of CTTA for 3D semantic segmentation. The key to MMCTTA is to adaptively attend to the reliable modality while avoiding catastrophic forgetting during continual domain shifts, which is out of the capability of previous TTA or CTTA methods. To fulfill this gap, we propose an MM-CTTA method called Continual Cross-Modal Adaptive Clustering (CoMAC) that addresses this task from two perspectives. On one hand, we propose an adaptive dual-stage mechanism to generate reliable cross-modal predictions by attending to the reliable modality based on the class-wise feature-centroid distance in the latent space. On the other hand, to perform test-time adaptation without catastrophic forgetting, we design class-wise momentum queues that capture confident target features for adaptation while stochastically restoring pseudo-source features to revisit source knowledge. We further introduce two new benchmarks to facilitate the exploration of MM-CTTA in the future. Our experimental results show that our method achieves state-of-the-art performance on both benchmarks. Visit our project website at https://sites.google.com/view/mmcotta.
Haozhi Cao, Yuecong Xu, Jianfei Yang 0001, Pengyu Yin, Shenghai Yuan 0001, Lihua Xie 0001
ICCV5
2023 Non-cooperative Stochastic Target Encirclement by Anti-synchronization Control via Range-only Measurement
abstract
This paper investigates the stochastic moving target encirclement problem in a realistic setting. In contrast to typical assumptions in related works, the target in our work is non-cooperative and capable of escaping the circle containment by boosting its speed to maximum for a short duration. In extreme conditions, where GPS signals are not available, weight restrictions are present, and ground guidance is absent, the agents can rely solely on their onboard single-modality perception tools to measure the distances to the target. The distance measurement allows for creating a position estimator by providing a target position-dependent variable. Furthermore, the construction of the unique distributed anti-synchronization controller (DASC) can guarantee that the two agents track and encircle the target swiftly. The convergence of the estimator and controller is rigorously evaluated using the Lyapunov technique. A real-world UAV-based experiment is conducted to illustrate the performance of the proposed methodology in addition to a simulated Matlab numerical sample. Our video demonstration can be found in the URL https://youtu.be/EDVLvP-bk8M.
Shenghai Yuan 0001, Wei Meng 0002, Rong Su 0001, Lihua Xie 0001
ICRA2
2023 Segregator: Global Point Cloud Registration with Semantic and Geometric Cues
abstract
This paper presents Segregator, a global point cloud registration framework that exploits both semantic information and geometric distribution to efficiently build up outlier-robust correspondences and search for inliers. Current state-of-the-art algorithms rely on point features to set up putative correspondences and refine them by employing pair-wise distance consistency checks. However, such a scheme suffers from degenerate cases, where the descriptive capability of local point features downgrades, and unconstrained cases, where length-preserving (1-TRIMs)-based checks cannot sufficiently constrain whether the current observation is consistent with others, resulting in a complexified NP-complete problem to solve. To tackle these problems, on the one hand, we propose a novel degeneracy-robust and efficient corresponding procedure consisting of both instance-level semantic clusters and geometric-level point features. On the other hand, Gaussian distribution-based translation and rotation invariant measurements (G-TRIMs) are proposed to conduct the consistency check and further constrain the problem size. We validated our proposed algorithm on extensive real-world data-based experiments. The code is available: https://github.com/Pamphlett/Segregator.
Pengyu Yin, Shenghai Yuan 0001, Haozhi Cao, Xingyu Ji, Lihua Xie 0001
ICRA2
2023 DoubleBee: A Hybrid Aerial-Ground Robot with Two Active Wheels
abstract
In this paper, we present the dynamic model and control of DoubleBee, a novel hybrid aerial-ground vehicle consisting of two propellers mounted on tilting servo motors and two motor-driven wheels. DoubleBee exploits the high energy efficiency of a bicopter configuration in aerial mode, and enjoys the low power consumption of a two-wheel self-balancing robot on the ground. Furthermore, the propeller thrusts act as additional control inputs on the ground, enabling a novel decoupled control scheme where the attitude of the robot is controlled using thrusts and the translational motion is realized using wheels. A prototype of DoubleBee is constructed using commercially available components. The power efficiency and the control performance of the robot are verified through comprehensive experiments. Challenging tasks in indoor and outdoor environments demonstrate the capability of DoubleBee to traverse unstructured environments, fly over and move under barriers, and climb steep and rough terrains.
Muqing Cao, Xinhang Xu, Shenghai Yuan 0001, Kun Cao 0002, Kangcheng Liu, Lihua Xie 0001
IROS3
2023 AirVO: An Illumination-Robust Point-Line Visual Odometry
abstract
This paper proposes an illumination-robust visual odometry (VO) system that incorporates both accelerated learning-based corner point algorithms and an extended line feature algorithm. To be robust to dynamic illumination, the proposed system employs the convolutional neural network (CNN) and graph neural network (GNN) to detect and match reliable and informative corner points. Then point feature matching results and the distribution of point and line features are utilized to match and triangulate lines. By accelerating CNN and GNN parts and optimizing the pipeline, the proposed system is able to run in real-time on low-power embedded platforms. The proposed VO was evaluated on several datasets with varying illumination conditions, and the results show that it outperforms other state-of-the-art VO systems in terms of accuracy and robustness. The open-source nature of the proposed system allows for easy implementation and customization by the research community, enabling further development and improvement of VO for various applications.
Yuefan Hao, Shenghai Yuan 0001, Chen Wang 0033, Lihua Xie 0001
IROS3
2023 AV-PedAware: Self-Supervised Audio-Visual Fusion for Dynamic Pedestrian Awareness
abstract
In this study, we introduce AV-PedAware, a self-supervised audio-visual fusion system designed to improve dynamic pedestrian awareness for robotics applications. Pedestrian awareness is a critical requirement in many robotics applications. However, traditional approaches that rely on cameras and LIDARs to cover multiple views can be expensive and susceptible to issues such as changes in illumination, occlusion, and weather conditions. Our proposed solution replicates human perception for 3D pedestrian detection using low-cost audio and visual fusion. This study represents the first attempt to employ audio-visual fusion to monitor footstep sounds for the purpose of predicting the movements of pedestrians in the vicinity. The system is trained through self-supervised learning based on LIDAR-generated labels, making it a cost-effective alternative to LIDAR-based pedestrian awareness. AV-PedAware achieves comparable results to LIDAR-based systems at a fraction of the cost. By utilizing an attention mechanism, it can handle dynamic lighting and occlusions, overcoming the limitations of traditional LIDAR and camera-based systems. To evaluate our approach's effectiveness, we collected a new multimodal pedestrian detection dataset and conducted experiments that demonstrate the system's ability to provide reliable 3D detection results using only audio and visual data, even in extreme visual conditions. We will make our collected dataset and source code available online for the community to encourage further development in the field of robotics perception systems.
Yizhuo Yang 0001, Shenghai Yuan 0001, Muqing Cao, Jianfei Yang 0001, Lihua Xie 0001
IROS2
2023 MM-Fi: Multi-Modal Non-Intrusive 4D Human Dataset for Versatile Wireless Sensing
abstract
4D human perception plays an essential role in a myriad of applications, such as home automation and metaverse avatar simulation. However, existing solutions which mainly rely on cameras and wearable devices are either privacy intrusive or inconvenient to use. To address these issues, wireless sensing has emerged as a promising alternative, leveraging LiDAR, mmWave radar, and WiFi signals for device-free human sensing. In this paper, we propose MM-Fi, the first multi-modal non-intrusive 4D human dataset with 27 daily or rehabilitation action categories, to bridge the gap between wireless sensing and high-level human perception tasks. MM-Fi consists of over 320k synchronized frames of five modalities from 40 human subjects. Various annotations are provided to support potential sensing tasks, e.g., human pose estimation and action recognition. Extensive experiments have been conducted to compare the sensing capacity of each or several modalities in terms of multiple tasks. We envision that MM-Fi can contribute to wireless sensing research with respect to action recognition, human pose estimation, multi-modal learning, cross-modal supervision, and interdisciplinary healthcare research.
Jianfei Yang 0001, Yunjiao Zhou, Xinyan Chen 0002, Yuecong Xu, Shenghai Yuan 0001, Han Zou, Xiaoxuan Lu 0001, Lihua Xie 0001
NeurIPS6
2023 GaitFi: Robust Device-Free Human Identification via WiFi and Vision Multimodal Learning
abstract
As an important biomarker for human identification, human gait can be collected at a distance by passive sensors without subject cooperation, which plays an essential role in crime prevention, security detection, and other human identification applications. Presently, most research works are based on cameras and computer vision techniques to perform gait recognition. However, vision-based methods are not reliable when confronting poor illuminations, leading to degrading performances. In this article, we propose a novel multimodal gait recognition method, namely, GaitFi, which leverages WiFi signals and videos for human identification. In GaitFi, channel state information (CSI) that reflects the multipath propagation of WiFi is collected to capture human gaits, while videos are captured by cameras. To learn robust gait information, we propose a lightweight residual convolution network (LRCN) as the backbone network and further propose the two-stream GaitFi by integrating WiFi and vision features for the gait retrieval task. The GaitFi is trained by the triplet loss and classification loss on different levels of features. Extensive experiments are conducted in the real world, which demonstrates that the GaitFi outperforms state-of-the-art gait recognition methods based on single WiFi or camera, achieving 94.2% for human identification tasks of 12 subjects.
Lang Deng, Jianfei Yang 0001, Shenghai Yuan 0001, Han Zou, Xiaoxuan Lu 0001, Lihua Xie 0001
IEEE Internet Things J.3
2023 MetaFi++: WiFi-Enabled Transformer-Based Human Pose Estimation for Metaverse Avatar Simulation
abstract
In the metaverse, digital avatar plays an important role in representing human beings for various interaction with virtual objects and environments, which puts a high demand on effective pose estimation. Though camera-based solutions yield remarkable performance, they encounter privacy issues and degraded performance caused by varying illumination, especially in the smart home. In this article, we propose a WiFi-based Internet of Things-enabled human pose estimation scheme for metaverse avatar simulation, namely, MetaFi++. Specifically, WPFormer is designed with a shared convolutional module and a Transformer block to map the channel state information of WiFi signals to human pose landmarks, effectively exploring spatial information of human pose through self-attention. It is enforced to learn the annotations from the accurate computer vision model, thus achieving cross-modal supervision. Due to the ubiquitous existence of WiFi and robustness to various illumination conditions, WiFi-based human poses are suitable to instruct the movement of digital avatars in the metaverse, promoting avatar applications in smart homes. The experiments are conducted in the real world, and the results show that the MetaFi++ achieves very high performance with a PCK@50 of 97.30%. Our codes are available inhttps://github.com/pridy999/metafi_pose_estimation.
Yunjiao Zhou, Shenghai Yuan 0001, Han Zou, Lihua Xie 0001, Jianfei Yang 0001
IEEE Internet Things J.3
2023 SE-Calib: Semantic Edge-Based LiDAR-Camera Boresight Online Calibration in Urban Scenes
abstract
Rigorous boresight calibration between light detection and ranging (LiDAR) and the camera is crucial for geometry and optical information fusion in earth observation and robotic applications. Although boresight parameters can be obtained through pre-calibration with artificial targets, unforeseen movement of sensors during data collection can lead to significant errors in the boresight parameters. To address this issue, we propose SE-Calib, an automatic and target-free online boresight calibration method for LiDAR-Camera systems. SE-Calib firstly extracts semantic edge features from both point clouds and images simultaneously using the 3D semantic segmentation (3D-SS) and 2D semantic edge detection (2D-SED) methods. The boresight parameters are then optimized with an adaptive solver and maximizing the Soft Semantic Response Consistency Metric (SSRCM) scores iteratively. The SSRCM is designed to evaluate the coherence of cross-modular semantic edge features, and a confidence function is proposed to filter out unreliable optimization results. Experiments conducted on challenging urban datasets show an average boresight error of 0.206 degrees (2.47 pixels in reprojection error), demonstrating the effectiveness and robustness of the proposed method.
Youqi Liao, Jianping Li 0004, Shuhao Kang, Guifang Zhu, Shenghai Yuan 0001, Zhen Dong 0005, Bisheng Yang
IEEE Trans. Geosci. Remote. Sens.6
2023 NEPTUNE: Nonentangling Trajectory Planning for Multiple Tethered Unmanned Vehicles
abstract
Despite recent progress in trajectory planning for multiple robots and a single tethered robot, trajectory planning for multiple tethered robots to reach their individual targets without entanglements remains a challenging problem. In this article, a complete approach is presented to address this problem. First, a multirobot tether-aware representation of homotopy is proposed to efficiently evaluate the feasibility and safety of a potential path in terms of 1) the cable length required to reach a target following the path, and 2) the risk of entanglements with the cables of other robots. Then the proposed representation is applied in a decentralized and online planning framework, which includes a graph-based kinodynamic trajectory finder and an optimization-based trajectory refinement, to generate entanglement-free, collision-free, and dynamically feasible trajectories. The efficiency of the proposed homotopy representation is compared against the existing single and multiple tethered robot planning approaches. Simulations with up to eight UAVs show the effectiveness of the approach in entanglement prevention and its real-time capabilities. Flight experiments using three tethered UAVs verify the practicality of the presented approach. The software implementation is publicly available online.1
Muqing Cao, Kun Cao 0002, Shenghai Yuan 0001, Thien-Minh Nguyen, Lihua Xie 0001
IEEE Trans. Robotics3
2023 Vision-Based Plane Estimation and Following for Building Inspection With Autonomous UAV
abstract
In this article, we focus on enabling the autonomous perception and control of a small unmanned aerial vehicle (UAV) for a façade inspection task. Specifically, we consider the perception as a planar object pose estimation problem by simplifying the building structure as a concatenation of planes, and the control as an optimal reference tracking control problem. First, a vision-based adaptive observer is proposed for plane pose estimation which converges fast and is insensitive to noise under very mild observation conditions. Second, a model predictive controller (MPC) is designed to achieve stable plane following and smooth transition in a multiple-plane scenario, while the persistent excitation (PE) condition of the observer and the maneuver constraints of the UAV are satisfied. The stability of the observer and the MPC controller is also investigated to ensure theoretical completeness. The proposed autonomous plane pose estimation and plane tracking methods are tested in both simulation and practical building façade inspection scenarios, which demonstrate their effectiveness and practicability.
Yang Lyu, Muqing Cao, Shenghai Yuan 0001, Lihua Xie 0001
IEEE Trans. Syst. Man Cybern. Syst.3
2022 Robust RGB-D SLAM in Dynamic Environments for Autonomous Vehicles
abstract
Vision-based SLAM has played an important role in many robotic applications. However, most existing visual SLAM methods are developed under a static world assumption and the robustness in dynamic environments remains a challenging problem. In this paper, we propose a robust RGB-D SLAM system for autonomous vehicles in dynamic scenarios which uses geometry-only information to reduce the impact of moving objects. To achieve this, we introduce an effective and efficient dynamic points detection module in a feature- based SLAM system. Specifically, for each new RGB-D image pair, we first segment the depth image into a few regions using the KMeans algorithm, and then identify the dynamic regions via their reprojection errors. The feature points located in these dynamic regions are then removed and only static ones are used for pose estimation. A dense map that contains only static parts of the environment is also produced by removing dynamic regions in the keyframes. Extensive experiments on public dataset and in real-world scenarios demonstrate that our method provides significant improvement in localization accuracy and mapping quality in dynamic environments.
Tete Ji, Shenghai Yuan 0001, Lihua Xie 0001
ICARCV2
2022 Overcoming Catastrophic Forgetting for Semantic Segmentation Via Incremental Learning
abstract
Deep learning based semantic segmentation models have achieved remarkable results in recent years. However, many deep learning based models encounter the problem of catastrophic forgetting, i.e. when the model is required to learn a new task without labels for old objects, its performance drops significantly for the previous tasks. To solve this problem, an incremental learning method, a Combination of Old Prediction and Modified Label (COPML), is developed in this paper. The proposed method utilizes the prediction results of the old model and the modified labels of the new task to create pseudo labels which are close to the ground truths. By using these pseudo labels for training, the model is expected to preserve the knowledge of old tasks. In addition, knowledge distillation, the replay and parameter freezing strategy are also applied to the proposed method to further assist the model in overcoming catastrophic forgetting. The effectiveness of the proposed method is validated on two semantic segmentation models: Unet and Deeplab3 in Pascal- VOC 2012 dataset and a self-made dataset. The experimental results demonstrate that COPML enables the model to maintain most of the old knowledge while obtaining an excellent performance on a new task.
Yizhuo Yang 0001, Shenghai Yuan 0001, Lihua Xie 0001
ICARCV2
2022 Real-time recognition and warning of mask wearing based on improved YOLOv5 R6.1
abstract
Since the new crown epidemic, mask-wearing has become a new normal in people's work and life. The inspection mechanism for mask-wearing at the entrance and exit of public places is seriously insufficient. The phenomenon of “pick-up on entry” has led to the severe formalization of mask-wearing inspection. Manual detection of mask-wearing in an open and dynamic crowded environment is unrealistic, which is not only time-consuming and labor-intensive but also cannot achieve early warning throughout the entire process. In response to this problem, this paper proposes a real-time recognition and early warning method for mask-wearing in an open, dynamic, complex environment based on improved YOLOv5 R6.1. First, replacing the first Conv structure of the backbone network in the YOLOv5 R6.1 model with an improved Stem structure to minimize the computational overhead while improving the performance. Then by normalizing the data, the random erasure data expansion technique is used to enhance the antiocclusion robustness of the algorithm. Finally, according to the mask-wearing specification in the training data set, optimizing and adjusting the anchor box parameters of the YOLOv5 R6.1 model to improve the model's ability to recognize small targets. The experiments are based on open data sets, and the results show that the mean precision (mAP), precision, and recall of this method reach 92.9%, 94.1%, and 88.5% on average, and the average frames per second (FPS) reaches 117. Moreover, the mAP and FPS are improved by an average of 6.5% and 474% compared with algorithms based on RetinaNet, Attention-Retina, Single Shot multibox Detector, Fast-RCNN, YOLOv4, and YOLOv5.
Shenghai Yuan 0001, Tiancai Liang, Wenchao Jiang, Sui Lin, Zhiming Zhao
Int. J. Intell. Syst.1
2022 Achieving Real-Time Path Planning in Unknown Environments Through Deep Neural Networks
abstract
Real-time path planning is crucial for intelligent vehicles to achieve autonomous navigation. In this paper, we propose a novel deep neural network (DNN) based method for real-time online path planning in unknown cluttered environments. Firstly, an end-to-end DNN architecture named online three-dimensional path planning network (OTDPP-Net) is designed to learn 3D local path planning policies. It determines actions in 3D space based on multiple value iteration computations approximated by recurrent 2D convolutional neural networks. Moreover, a path planning framework is also developed to realize near-optimal real-time online path planning. The effectiveness of the proposed planner is further improved by a switching scheme, and the path quality is optimized by line-of-sight checks. Both virtual and real-world experimental results demonstrate the remarkable performance of the proposed DNN-based path planner in terms of efficiency, success rate and path quality. Different from existing methods, the computational time and effectiveness of the developed DNN-based path planner are both independent of environmental conditions, which reveals its superiority in large-scale complex environments. A video of our experiments can be found at:https://youtu.be/gb4nSG4hd6s.
Keyu Wu 0002, Han Wang 0001, Mahdi Abolfazli Esfahani, Shenghai Yuan 0001
IEEE Trans. Intell. Transp. Syst.4
2022 VIRAL-Fusion: A Visual-Inertial-Ranging-Lidar Sensor Fusion Approach
abstract
In recent years, onboard self-localization (OSL) methods based on cameras or lidar have achieved many significant progresses. However, some issues such as estimation drift and robustness in low-texture environment still remain inherent challenges for OSL methods. On the other hand, infrastructure-based methods can generally overcome these issues, but at the expense of some installation cost. This poses an interesting problem of how to effectively combine these methods, so as to achieve localization with long-term consistency as well as flexibility compared to any single method. To this end, we propose a comprehensive optimization-based estimator for the 15-D state of an unmanned aerial vehicle (UAV), fusing data from an extensive set of sensors: inertial measurement unit (IMU), ultrawideband (UWB) ranging sensors, and multiple onboard visual-inertial and lidar odometry subsystems. In essence, a sliding window is used to formulate a sequence of robot poses, where relative rotational and translational constraints between these poses are observed in the IMU preintegration and OSL observations, while orientation and position are coupled in thebody-offsetUWB range observations. An optimization-based approach is developed to estimate the trajectory of the robot in this sliding window. We evaluate the performance of the proposed scheme in multiple scenarios, including experiments on public datasets, high-fidelity graphical-physical simulation, and field-collected data from UAV flight tests. The result demonstrates that our integrated localization method can effectively resolve the drift issue, while incurring minimal installation requirements.
Thien-Minh Nguyen, Muqing Cao, Shenghai Yuan 0001, Yang Lyu, Thien Hoang Nguyen, Lihua Xie 0001
IEEE Trans. Robotics3
2021 LIRO: Tightly Coupled Lidar-Inertia-Ranging Odometry
abstract
In recent years, thanks to the continuously reduced cost and weight of 3D lidar, the applications of this type of sensor in the community have become increasingly popular. Despite many progresses, estimation drift and tracking loss are still prevalent concerns associated with these systems. However, in theory these issues can be resolved with the use of some observations to fixed landmarks in the operation environments. This motivates us to investigate a sensor fusion scheme of lidar and inertia measurements with Ultra-Wideband (UWB) range measurements to such landmarks, which can be easily deployed in the environments with minimal cost and time. Hence, data from IMU, lidar and UWB are tightly-coupled with the robot's states on a sliding window based on their timestamps. Then, we construct a cost function comprising of factors from UWB, lidar and IMU preintegration measurements. Finally an optimization process is carried out to estimate the robot's position and orientation. It is demonstrated through some real world experiments that the method can effectively resolve the drift issue, while only requiring two or three anchors deployed in the environment.
Thien-Minh Nguyen, Muqing Cao, Shenghai Yuan 0001, Yang Lyu, Thien Hoang Nguyen, Lihua Xie 0001
ICRA3
2020 Unsupervised Scene Categorization, Path Segmentation and Landmark Extraction while Traveling Path
abstract
Segmenting the movement path is an essential requirement of intelligent mobile robots. It assists intelligent systems in gaining a better understanding of the scene and identifying re-visited spaces. Moreover, it helps intelligent robots quantize the wide scene into sub-spaces that visually represent the same content-for instance, distinguishing rooms and kitchen in an indoor environment. This paper proposes an unsupervised approach to understand the transition of the scene while a robot is moving and helps to extract sub-spaces that visually represent a similar environment. The proposed approach benefits from a pre-trained deep network architecture to extract a description (feature representation) for the mobile robot's visual information at each time step. Then, based on the pairwise distance of the feature representations, sub-spaces of the scene and transition points are extracted.
Mahdi Abolfazli Esfahani, Han Wang 0001, Keyu Wu 0002, Shenghai Yuan 0001
ICARCV4
2020 From Local Understanding to Global Regression in Monocular Visual Odometry
abstract
The most significant part of any autonomous intelligent robot is the localization module that gives the robot knowledge about its position and orientation. This knowledge assists the robot to move to the location of its desired goal and complete its task. Visual Odometry (VO) measures the displacement of the robots’ camera in consecutive frames which results in the estimation of the robot position and orientation. Deep Learning, nowadays, helps to learn rich and informative features for the problem of VO to estimate frame-by-frame camera movement. Recent Deep Learning-based VO methods train an end-by-end network to solve VO as a regression problem directly without visualizing and sensing the label of training data in the training procedure. In this paper, a new approach to train Convolutional Neural Networks (CNNs) for the regression problems, such as VO, is proposed. The proposed method first changes the problem to a classification problem to learn different subspaces with similar observations. After solving the classification problem, the problem converts to the original regression problem to solve using the knowledge achieved by solving the classification problem. This approach helps CNN to solve regression problem globally in a local domain learned in the classification step, and improves the performance of the regression module for approximately 10%.
Mahdi Abolfazli Esfahani, Keyu Wu 0002, Shenghai Yuan 0001, Han Wang 0001
Int. J. Pattern Recognit. Artif. Intell.3
2020 AbolDeepIO: A Novel Deep Inertial Odometry Network for Autonomous Vehicles
abstract
Inertial measurement units (IMUs) suffer from bias and measurement noise, which makes it much more complicated to tackle the problem of inertial odometry (IO). Due to the error propagation over time, while estimating robot position, an inaccurate estimation or a small error will cause the odometry and a localization system unreliable and unusable in a split of seconds. This paper presents a novel triple-channel deep IO network architecture based on the physical and mathematical models of IMUs. The proposed method simulates the noise model in the training phase and becomes robust to noise during testing. Besides, the proposed network architecture also considers the time interval between two consecutive IMU readings (sampling time) so that it is robust to the change of IMU frequency and the missing of IMU information. To the best of our knowledge, this paper is the first work reviewing and analyzing the existing IO methods used by the deep-learning-based visual-IO approaches. The proposed network architecture outperforms all the existing solutions on the IMU readings of the challenging Micro Aerial Vehicle dataset and improves the accuracy by approximately 25%.
Mahdi Abolfazli Esfahani, Han Wang 0001, Keyu Wu 0002, Shenghai Yuan 0001
IEEE Trans. Intell. Transp. Syst.4
2019 TDPP-Net: Achieving three-dimensional path planning via a deep neural network architecture
Keyu Wu 0002, Mahdi Abolfazli Esfahani, Shenghai Yuan 0001, Han Wang 0001
Neurocomputing3
2019 DeepDSAIR: Deep 6-DOF camera relocalization using deblurred semantic-aware image representation for large-scale outdoor environments
Mahdi Abolfazli Esfahani, Keyu Wu 0002, Shenghai Yuan 0001, Han Wang 0001
Image Vis. Comput.3
2014 Autonomous object level segmentation
abstract
In this paper we describe a new technique for segment meaningfully objects autonomously. Traditional segmentation scheme tries to find the best segmentation result at some trade off between level of user input and level of meaningful segmentation. Segmentation with user input will ensure better segmentation result but is not applicable to the realtime autonomous robotics system. For most of the commercial software users, they prefer program to know where to segmented before they even touches. No user input usually means that system tends to either over segment or wrongly segmented. Also tuning parameters of the autonomous segmentation scheme is painful. In this paper, we propose a novel idea to find initial seed for the segmentation scheme which need manual input and product object level segments fully autonomously. We also proposed a new evaluation scheme for robotics based autonomous segmentation measurement.
Shenghai Yuan 0001, Han Wang 0001
ICARCV1