Rui Tian 0002

dblp:15/9722-2 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0002-8944-1966ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 UniQuadric: A SLAM Backend for Unknown Rigid Object 3-D Tracking and Light-Weight Modeling
abstract
Tracking and modeling unknown rigid objects in the environment play a crucial role in autonomous uncrewed systems and virtual-real interactive applications. However, many existing Simultaneous Localization, Mapping and Moving Object Tracking (SLAMMOT) methods focus solely on estimating specific object poses and lack estimation of object scales and are unable to effectively track unknown objects. In this paper, we propose a novel SLAM backend that unifies ego-motion tracking, rigid object motion tracking, and modeling within a joint optimization framework. In the perception part, we designed a pixel-level asynchronous object tracker (AOT) based on the Segment Anything Model (SAM) and DeAOT, enabling the tracker to effectively track target unknown objects guided by various predefined tasks and prompts. In the modeling part, we present a novel object-centric quadric parameterization to unify both static and dynamic object initialization and optimization. Subsequently, in the part of object state estimation, we propose a tightly coupled optimization model for object pose and scale estimation, incorporating hybrids constraints into a novel dual sliding window optimization framework for joint estimation. To our knowledge, we are the first to tightly couple object pose tracking with light-weight modeling of dynamic and static objects using quadric. We conduct qualitative and quantitative experiments on simulation datasets and real-world datasets, demonstrating the state-of-the-art robustness and accuracy in motion estimation and modeling. This showcases the significant potential of our method for object perception in complex dynamic environments.
Linghao Yang, Yanmin Wu, Rui Tian 0002, Xinggang Hu
IEEE Trans. Intell. Transp. Syst.4
2024 DynaQuadric: Dynamic Quadric SLAM for Quadric Initialization, Mapping, and Tracking
abstract
Dynamic SLAM is a key technology for autonomous driving and robotics, and accurate pose estimation of surrounding objects is important for semantic perception tasks. Current quadric SLAM methods are based on the assumption of a static environment and can only reconstruct static quadrics in the scene, which limits their applications in complex dynamic scenarios. In this paper, we propose a visual SLAM system that is capable of reconstructing dynamic objects as quadrics, with a unified framework for jointly optimizing pose estimation, multi-object tracking (MOT), and quadric parameters. We propose a robust object-centric quadric initialization algorithm for both static and moving objects, which decouples the prior estimation of the object pose from the quadric parameters. The object is initialized with a coarse sphere, and quadric parameters are further refined. We design a novel factor graph that tightly optimizes camera pose, object pose, map points and quadric parameters within the sliding window-based optimization. To the best of our knowledge, we are the first to propose a dynamic SLAM that combines quadric representations and MOT in a tightly coupled optimization. We perform qualitative and quantitative experiments on both simulated and real-world datasets, and demonstrate the robustness and accuracy in terms of camera localization, dynamic quadric initialization, mapping and tracking. Our system demonstrates the potential application of object perception with quadric representation in complex dynamic scenes.
Rui Tian 0002, Yunzhou Zhang, Linghao Yang, Sonya A. Coleman, Dermot Kerr
IEEE Trans. Intell. Transp. Syst.1
2024 Fast, Robust, Accurate, Multi-Body Motion Aware SLAM
abstract
Simultaneous ego localization and surrounding object motion awareness are significant issues for the navigation capability of unmanned systems and virtual-real interaction applications. Robust and accurate data association at object and feature levels is one of the key factors in solving this problem. However, currently available solutions ignore the complementarity among different cues in the front-end object association and the negative effects of poorly tracked features on the back-end optimization. It makes them not robust enough in practical applications. Motivated by these observations, we make up rigid environment as a unified whole to assist state decoupling by integrating high-level semantic information, ultimately enabling simultaneous multi-states estimation. A filter-based multi-cues fusion object tracker is proposed for establishing more stable object-level data association. Combined with the object’s motion priors, the motion-aided feature tracking algorithm is proposed to improve the feature-level data association performance. Furthermore, a novel state estimation factor graph is designed which integrates a specific feature observation uncertainty model and the intrinsic priors of tracked object, and solved through sliding-window optimization. Our system is evaluated using the KITTI dataset and achieves comparable performance to state-of-the-art object pose estimation systems both quantitatively and qualitatively. We have also validated our system on simulation environment and a real-world dataset to confirm the potential application value in different practical scenarios.
Linghao Yang, Yunzhou Zhang, Rui Tian 0002, Shiwen Liang, You Shen, Sonya A. Coleman, Dermot Kerr
IEEE Trans. Intell. Transp. Syst.3
2024 CapsLoc3D: Point Cloud Retrieval for Large-Scale Place Recognition Based on 3D Capsule Networks
abstract
Point cloud-based place recognition can be used for global localization in large-scale scenes and loop-closure detection in simultaneous localization and mapping (SLAM) systems in the absence of GPS. Current learning-based approaches aim to extract global and local features from 3D point clouds to encode them as descriptors for point cloud retrieval. The key problems are that the occlusion of point clouds by dynamic objects in the scene affects the point cloud structure, a single perceptual field of the network cannot adequately extract point cloud features, and the correlation between features is not fully utilized. To overcome this, we propose a novel network called CapsLoc3D. We first use the static point cloud generation module to remove the occlusion effects of dynamic objects, and then obtain the point cloud descriptors by processing with the CapsLoc3D network which contains the point spatial transformation module, multi-scale feature fusion module, Capsnet module and a GeM Pooling layer. After validation using the Oxford RobotCar, KITTI, and NEU datasets, experiments show that our method performs better and also has good generalization performance and computational efficiency compared with current state-of-the-art algorithms.
Yunzhou Zhang, Ming Liao, Rui Tian 0002, Sonya A. Coleman, Dermot Kerr
IEEE Trans. Intell. Transp. Syst.4
2024 MVSE-Net: A Multi-View Deep Network With Semantic Embedding for LiDAR Place Recognition
abstract
Place recognition is a critical technology in robot navigation and autonomous driving, remains challenging due to inefficient point cloud computation, limited feature representation capability, and poor robustness to long-term environmental changes. We propose MVSE-Net, a feature extraction network with embedded semantic information for multi-view feature fusion. MVSE-Net can convert point cloud data acquired by LiDAR in real time into global descriptors for retrieval. Processing a point cloud by projecting it onto a 2D image can greatly improve computational efficiency. We projected the point cloud into a range-view (RV) image and a bird’s-eye-view (BEV) image in forward and top view, respectively. The semantic segmentation network is then used to process the RV image, and the feature extraction part of the semantic model is connected to the transformer attention module to further refine the features for the place recognition task. The point cloud containing the semantic segmentation results is then converted into a semantic BEV image, and the multi-channel BEV image is processed using a group convolutional network. Finally, the features of the two branches are fused into a global feature representation by post-fusion. Our experiments on three publicly available datasets demonstrate that MVSE-Net exhibits high recall and strong generalization in LiDAR place recognition, outperforming previous state-of-the-art methods.
Yunzhou Zhang, Lei Rong, Rui Tian 0002, Sizhan Wang
IEEE Trans. Intell. Transp. Syst.4
2023 Object SLAM With Robust Quadric Initialization and Mapping for Dynamic Outdoors
abstract
Object SLAM is a popular approach for autonomous driving and robotics, but accurate object perception in outdoor environments remains a challenge. State-of-the-art object SLAM algorithms rely on assumptions and are sensitive to observation noise, limiting their application in real-world scenarios. To address these challenges, we propose a novel object SLAM system that utilizes a quadric initialization algorithm based on constrained quadric optimization, which does not rely on planar assumptions and is robust to partial observations. Additionally, we introduce an automatic object data association algorithm capable of detecting motion states while associating objects across frames. To further enhance the accuracy of the quadric mapping, an extra thread is used to refine the ellipsoid parameters within a local sliding window composed of keyframes. Our system utilizes a joint optimization framework that optimizes camera poses, object landmarks, and point clouds in the local mapping thread for further global optimization while maintaining a consistent map. Experimental results on the real-world KITTI dataset show that the proposed system is more robust and significantly outperforms current state-of-the-art methods in quadric initialization and mapping in outdoor scenarios. Moreover, our system achieves real-time performance, making it suitable for practical applications.
Rui Tian 0002, Yunzhou Zhang, Zhenzhong Cao, Linghao Yang, Sonya A. Coleman, Dermot Kerr
IEEE Trans. Intell. Transp. Syst.1
2022 SemLoc: Accurate and Robust Visual Localization with Semantic and Structural Constraints from Prior Maps
abstract
Semantic information and geometrical structures of a prior map can be leveraged in visual localization to bound drift errors and improve accuracy. In this paper, we propose SemLoc, a pure visual localization system, for accurate localization in a prior semantic map. To tightly couple semantic and structure information from prior maps, a hybrid constraint is presented by using the Dirichlet distribution. Then, with the local landmarks and their semantic states tracked in the frontend, the camera poses and data associations are jointly optimized through Expectation-Maximization (EM) algorithm. We validate the effectiveness of our approach in both monocular and stereo modes on the public KITTI dataset. Experimental results demonstrate that our system can greatly reduce drift errors with an satisfying real-time performance. Compared with several state-of-the-art visual localization systems, the proposed framework achieves a competitive localization performance.
Shiwen Liang, Yunzhou Zhang, Rui Tian 0002, Delong Zhu 0001, Linghao Yang, Zhenzhong Cao
ICRA3
2021 Accurate and Robust Scale Recovery for Monocular Visual Odometry Based on Plane Geometry
abstract
Scale ambiguity is a fundamental problem in monocular visual odometry. Typical solutions include loop closure detection and environment information mining. For applications like self-driving cars, loop closure is not always available, hence mining prior knowledge from the environment becomes a more promising approach. In this paper, with the assumption of a constant height of the camera above the ground, we develop a light-weight scale recovery framework leveraging an accurate and robust estimation of the ground plane. The framework includes a ground point extraction algorithm for selecting high-quality points on the ground plane, and a ground point aggregation algorithm for joining the extracted ground points in a local sliding window. Based on the aggregated data, the scale is finally recovered by solving a least-squares problem using a RANSAC-based optimizer. Sufficient data and robust optimizer enable a highly accurate scale recovery. Experiments on the KITTI dataset show that the proposed framework can achieve state-of-the-art accuracy in terms of translation errors, while maintaining competitive performance on the rotation error. Due to the light-weight design, our framework also demonstrates a high frequency of 20 Hz on the dataset.
Rui Tian 0002, Yunzhou Zhang, Delong Zhu 0001, Shiwen Liang, Sonya A. Coleman, Dermot Kerr
ICRA1