EDBT 2026 Demo / reviewers in the wild / expert
Tong Qin 0001
dblp:58/6849-1
· DBLP profile ↗
27ranked-venue papers
7as first author
17since 2021 · last 2025
0000-0002-0994-9816ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 5 first-author · 13 since 2021Systems, architecture and hardware · 14 · 5 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RL-OGM-Parking: Lidar OGM-Based Hybrid Reinforcement Learning Planner for Autonomous ParkingabstractAutonomous parking has become a critical application in automatic driving research and development. Parking operations often suffer from limited space and complex environments, requiring accurate perception and precise maneuvering. Traditional rule-based parking algorithms struggle to adapt to diverse and unpredictable conditions, while learning-based algorithms lack consistent and stable performance in various scenarios. Therefore, a hybrid approach is necessary that combines the stability of rule-based methods and the generalizability of learning-based methods. Recently, reinforcement learning (RL) based policy has shown robust capability in planning tasks. However, the simulation-to-reality (sim-to-real) transfer gap seriously blocks the real-world deployment. To address these problems, we employ a hybrid policy, consisting of a rule-based Reeds-Shepp (RS) planner and a learningbased reinforcement learning (RL) planner. A real-time LiDARbased Occupancy Grid Map (OGM) representation is adopted to bridge the sim-to-real gap, leading the hybrid policy can be applied to real-world systems seamlessly. We conducted extensive experiments both in the simulation environment and real-world scenarios, and the result demonstrates that the proposed method outperforms pure rule-based and learningbased methods. The real-world experiment further validates the feasibility and efficiency of the proposed method. Zhitao Wang, Mingyang Jiang, Tong Qin 0001, Ming Yang 0002 |
ICRA | 4 |
| 2025 | Direct, Targetless and Automatic Joint Calibration of LiDAR-Camera Intrinsic and ExtrinsicabstractThis paper presents a direct, targetless, and automatic LiDAR-Camera joint calibration method that effectively overcomes the intrinsic precision limitations. We propose an iterative two-stage optimization methodology that leverages 3D LiDAR measurements to simultaneously refine both intrinsic and extrinsic. In the first stage, the intrinsic is optimized using a normalized information distance (NID) metric, an information-theoretic measure that quantifies the statistical alignment between LiDAR and image intensities, while initial extrinsic parameters derived from CAD specifications facilitate the projection of LiDAR point clouds onto the camera image plane. In the second stage, the refined intrinsic guides further optimization of extrinsic using the same NID-based evaluation metrics. This alternating process iteratively enhances both intrinsic and extrinsic through their mutual interdependence. Experiments across multiple datasets demonstrate that our method achieves sub-pixel intrinsic accuracy and extrinsic parameters that closely align with CAD specifications, validating the superior performance of our methodology for sensor fusion applications. Yishu Shen, Shaojie Shen, Tong Qin 0001 |
IROS | 4 |
| 2025 | Embodied Escaping: End-to-End Reinforcement Learning for Robot Navigation in Narrow EnvironmentabstractAutonomous navigation is a fundamental task for robot vacuum cleaners in indoor environments. Since their core function is to clean entire areas, robots inevitably encounter dead zones in cluttered and narrow scenarios. Existing planning methods often fail to escape due to complex environmental constraints, high-dimensional search spaces, and high difficulty maneuvers. To address these challenges, this paper proposes an embodied escaping model that leverages a reinforcement learning-based policy with an efficient action mask for dead zone escaping. To alleviate the issue of the sparse reward in training, we introduce a hybrid training policy that improves learning efficiency. In handling redundant and ineffective action options, we design a novel action representation to reshape the discrete action space with a uniform turning radius. Furthermore, we develop an action mask strategy to select valid actions quickly, balancing precision and efficiency. In real-world experiments, our robot is equipped with a Lidar, IMU, and two-wheel encoders. Extensive quantitative and qualitative experiments across varying difficulty levels demonstrate that our robot can consistently escape from challenging dead zones. Moreover, our approach significantly outperforms compared path planning and reinforcement learning methods in terms of success rate and collision avoidance. A video showcasing our methodology and real-world demonstrations is available at https://youtu.be/kBaaYWGhNuE. Mingyang Jiang, Peiyuan Liu, Tong Qin 0001, Ming Yang 0002 |
IROS | 6 |
| 2025 | CrossGLoc: Cross-Modal Global Localization Leveraging Pretrained Diffusion Models and Semantic Cues for Intelligent VehiclesabstractCross-modal global localization matches visual information with pre-built LiDAR maps, which has attracted more and more attention for its low cost and potential robustness. However, the inherent modality difference between images and point clouds makes it challenging. This paper proposes a novel cross-modal global localization system, named CrossGLoc, which leverages pre-trained diffusion models and semantic cues to address this challenge. The main idea is leveraging the semantic cues shared between different modalities to bridge the modality gap, and utilizing pre-trained diffusion models to extract modality-consistent high-dimensional features guided by these semantic cues. To achieve this, ControlNet is used to generate intermediate feature maps from semantic images and semantic map projections, and a semantic categories-based feature aggregation algorithm is proposed to aggregate these feature maps into global descriptors. Furthermore, a semantic edge key points-based pose estimation algorithm is proposed to estimate the pose of retrieved image and point cloud pairs. Extensive experiments on the KITTI dataset, the KITTI360 dataset and the self-collected dataset demonstrate that the proposed method achieves state-of-the-art performance in cross-modal global localization. Hengwang Zhao, Qiyuan Shen, Hanyang Zhuang, Tong Qin 0001, Ming Yang 0002 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | GS-LIVO: Real-Time LiDAR, Inertial, and Visual Multisensor Fused Odometry With Gaussian MappingabstractIn recent years, 3D Gaussian splatting (3D-GS) has emerged as a novel scene representation approach. However, existing vision-only 3D-GS methods often rely on hand-crafted heuristics for point-cloud densification and face challenges in handling occlusions and high GPU memory and computation consumption [1]. LiDAR-Inertial-Visual (LIV) sensor configuration has demonstrated superior performance in precise localization and dense mapping by leveraging complementary sensing characteristics: rich texture information from cameras, precise geometric measurements from LiDAR, and high-frequency motion data from IMU [2]-[8]. Inspired by this, we propose a novel real-time Gaussian-based simultaneous localization and mapping (SLAM) system. Our map system comprises a global Gaussian map and a sliding window of Gaussians, along with an IESKF-based real-time odometry utilizing Gaussian maps. The structure of the global Gaussian map consists of hash-indexed voxels organized in a recursive octree. This hierarchical structure effectively covers sparse spatial volumes while adapting to different levels of detail and scales in the environment. The Gaussian map is efficiently initialized through multi-sensor fusion and optimized with photometric gradients. Our system incrementally maintains a sliding window of Gaussians with minimal graphics memory usage, significantly reducing GPU computation and memory consumption by only optimizing the map within the sliding window, enabling real-time optimization. Moreover, we implement a tightly coupled multi-sensor fusion odometry with an iterative error state Kalman filter (IESKF), which leverages real-time updating and rendering of the Gaussian map to achieve competitive localization accuracy. Our system represents the first real-time Gaussian-based SLAM framework deployable on resource-constrained embedded systems (all implemented in C++/CUDA for efficiency), demonstrated on theNVIDIA Jetson Orin NXplatform. The framework achieves real-time performance while maintaining robust multi-sensor fusion capabilities. All implementation algorithms, hardware designs, and CAD models and demo video of our GPU-accelerated system will be publicly available athttps://github.com/HKUST-Aerial-Robotics/GS-LIVO. Chunran Zheng, Yishu Shen, Changze Li, Fu Zhang 0002, Tong Qin 0001, Shaojie Shen |
IEEE Trans. Robotics | 6 |
| 2024 | ParkingE2E: Camera-based End-to-end Parking Network, from Images to PlanningabstractAutonomous parking is a crucial task in the intelligent driving field. Traditional parking algorithms are usually implemented using rule-based schemes. However, these methods are less effective in complex parking scenarios due to the intricate design of the algorithms. In contrast, neural-network-based methods tend to be more intuitive and versatile than the rule-based methods. By collecting a large number of expert parking trajectory data and emulating human strategy via learning-based methods, the parking task can be effectively addressed. In this paper, we employ imitation learning to perform end-to-end planning from RGB images to path planning by imitating human driving trajectories. The proposed end-to-end approach utilizes a target query encoder to fuse images and target features, and a transformer-based decoder to autoregressively predict future waypoints. We conduct extensive experiments in real-world scenarios, and the results demonstrate that the proposed method achieved an average parking success rate of 87.8% across four different real-world garages. Real-vehicle experiments further validate the feasibility and effectiveness of the method proposed in this paper. The code can be found at: https://github.com/qintonguav/ParkingE2E. Changze Li, Ziheng Ji, Tong Qin 0001, Ming Yang 0002 |
IROS | 4 |
| 2024 | Cross-Modal Visual Relocalization in Prior LiDAR Maps Utilizing Intensity TexturesabstractCross-modal localization has drawn increasing attention in recent years, while the visual relocalization in prior LiDAR maps is less studied. Related methods usually suffer from inconsistency between the 2D texture and 3D geometry, neglecting the intensity features in the LiDAR point cloud. In this paper, we propose a cross-modal visual relocalization system in prior LiDAR maps utilizing intensity textures, which consists of three main modules: map projection, coarse retrieval, and fine relocalization. In the map projection module, we construct the database of intensity channel map images leveraging the dense characteristic of panoramic projection. The coarse retrieval module retrieves the top-K most similar map images to the query image from the database, and retains the top-K’ results by covisibility clustering. The fine relocalization module applies a two-stage 2D-3D association and a covisibility inlier selection method to obtain robust correspondences for 6DoF pose estimation. The experimental results on our self-collected datasets demonstrate the effectiveness in both place recognition and pose estimation tasks. Qiyuan Shen, Hengwang Zhao, Weihao Yan 0001, Tong Qin 0001, Ming Yang 0002 |
IROS | 5 |
| 2024 | MapLocNet: Coarse-to-Fine Feature Registration for Visual Re-Localization in Navigation MapsabstractRobust localization is the cornerstone of autonomous driving, especially in challenging urban environments where GPS signals suffer from multipath errors. Traditional localization approaches rely on high-definition (HD) maps, which consist of precisely annotated landmarks. However, building HD map is expensive and challenging to scale up. Given these limitations, leveraging navigation maps has emerged as a promising low-cost alternative for localization. Current approaches based on navigation maps can achieve highly accurate localization, but their complex matching strategies lead to unacceptable inference latency that fails to meet the real-time demands. To address these limitations, we introduce MapLocNet, a novel transformer-based neural re-localization method. Inspired by image registration, our approach performs a coarse-to-fine neural feature registration between navigation map features and visual bird’s-eye view features. MapLocNet substantially outperforms the current state-of-the-art methods on both nuScenes and Argoverse datasets, demonstrating significant improvements in localization accuracy and inference speed across both single-view and surround-view input settings. We highlight that our research presents an HD-map-free localization method for autonomous driving, offering a costeffective, reliable, and scalable solution for challenging urban environments. Siyuan Lin, Xiangru Mu, Ming Yang 0002, Tong Qin 0001 |
IROS | 7 |
| 2024 | Pix2Planning: End-to-End Planning by Vision-language Model for Autonomous Driving on Carla SimulatorabstractThe end-to-end neural network has become a hot topic in recent years. Compared with traditional module-based solutions, the end-to-end paradigm is able to reduce the accumulated error and avoid information loss, so that it earns great attention in autonomous driving tasks. However, the current end-to-end network designs easily lose useful information during training due to the complexity of mapping high-dimensional visual observation to navigation waypoints. Since the future navigation point is reasoned from the former one, the planning task is like a sequence generation task. Inspired by the great power of the neural language model, we propose an end-to-end framework, which transfers the planning task as a language sequence generation task conditioned on pixel inputs. The proposed method firstly extracts and transforms the image feature from camera-view to bird-eye-view (BEV). Then the target navigation point is constructed into a text sequence, as the prompt of the visual-language transformer. Finally, the auto-regressive transformer decoder receives the BEV feature and the text sequences to generate sequential waypoints. Overall, our proposed method can make full use of the environmental information and express the planning trajectory as a language sequence to learn the correspondence between trajectory sequences and images. We have conducted extensive experiments on CARLA benchmarks and our model achieves state-of-the-art performance compared with other visual methods. Xiangru Mu, Tong Qin 0001, Songan Zhang, Chunjing Xu, Ming Yang 0002 |
IV | 2 |
| 2024 | 2D-3D Cross-Modality Network for End-to-End Localization with Probabilistic SupervisionabstractAccurate localization ability is a crucial component for autonomous robots. Given existing LiDAR 3D points maps, it is cost-effective to localize the robot only with onboard camera compared to LiDAR. However, matching 2D visual information with 3D point cloud maps presents huge challenges due to different modalities, dimensions, noise and occlusion issues. To overcome it, we propose an end-to-end neural network-based solution, which determines the 6-DoF pose of the camera relative to an existing LiDAR map with centimeter accuracy. Given a query image, a pre-acquired point cloud and an initial pose, the cross-modality network will output a precise pose. By projecting the 3D point cloud onto the image plane, a depth image is acquired as seen from the initial pose. Subsequently, a cross-modality flow network establishes the correspondences of 2D pixels and projected points. Importantly, we leverage a robust probabilistic Perspective-n-Point (PnP) module, which are capable of fine-tuning 2D pairs and learning the pairs weight in an end-to-end manner. A comprehensive evaluation of our proposed algorithm is conducted in KITTI datasets. Furthermore, deploying the algorithm on the real-world parking lot scenario validates its strong practicality of the proposed algorithm. We highlight that this research offers a cost-effective and highly accurate solution that can be readily deployed in low-cost commercial vehicles. Xiangru Mu, Tong Qin 0001, Chunjing Xu, Ming Yang 0002 |
IV | 3 |
| 2024 | BLOS-BEV: Navigation Map Enhanced Lane Segmentation Network, Beyond Line of SightabstractBird’s-eye-view (BEV) representation is crucial for the perception function in autonomous driving tasks. It is difficult to balance the accuracy, efficiency and range of BEV representation. The existing works are restricted to a limited perception range within 50 meters. Extending the BEV representation range can greatly benefit downstream tasks such as topology reasoning, scene understanding, and planning by offering more comprehensive information and reaction time. The Standard-Definition (SD) navigation maps can provide a lightweight representation of road structure topology, characterized by ease of acquisition and low maintenance costs. An intuitive idea is to combine the close-range visual information from onboard cameras with the beyond line-of-sight (BLOS) environmental priors from SD maps to realize expanded perceptual capabilities. In this paper, we propose BLOS-BEV, a novel BEV segmentation model that incorporates SD maps for accurate beyond line-of-sight perception, up to 200m. Our approach is applicable to common BEV architectures and can achieve excellent results by incorporating information derived from SD maps. We explore various feature fusion schemes to effectively integrate the visual BEV representations and semantic features from the SD map, aiming to leverage the complementary information from both sources optimally. Extensive experiments demonstrate that our approach achieves state-of-the-art performance in BEV segmentation on nuScenes and Argoverse benchmark. Through multi-modal inputs, BEV segmentation is significantly enhanced at close ranges below 50m, while also demonstrating superior performance in long-range scenarios, surpassing other methods by over 20% mIoU at distances ranging from 50-200m. Siyuan Lin, Tong Qin 0001, Chunjing Xu, Ming Yang 0002 |
IV | 4 |
| 2024 | E2E Parking: Autonomous Parking by the End-to-end Neural Network on the CARLA SimulatorabstractAutonomous parking is a crucial application for intelligent vehicles, especially in crowded parking lots. The confined space requires highly precise perception, planning, and control. Currently, the traditional Automated Parking Assist (APA) system, which utilizes geometric-based perception and rule-based planning, can assist with parking tasks in simple scenarios. With noisy measurement, the handcrafted rule often lacks flexibility and robustness in various environments, which performs poorly in super crowded and narrow spaces. On the contrary, there are many experienced human drivers, who are good at parking in narrow slots without explicit modeling and planning. Inspired by this, we expect a neural network to learn how to park directly from experts without handcrafted rules. Therefore, in this paper, we present an end-to-end neural network to handle parking tasks. The inputs are the images captured by surrounding cameras and basic vehicle motion state, while the outputs are control signals, including steer angle, acceleration, and gear. The network learns how to control the vehicle by imitating experienced drivers. We conducted closed-loop experiments on the CARLA Simulator to validate the feasibility of controlling the vehicle by the proposed neural network in the parking task. The experiment demonstrated the effectiveness of our end-to-end system in achieving the average position and orientation errors of 0.3 meters and 0.9 degrees with an overall success rate of 91%. The code is available at: https://github.com/qintonguav/e2e-parking-carla Yunfan Yang, Denglong Chen, Tong Qin 0001, Xiangru Mu, Chunjing Xu, Ming Yang 0002 |
IV | 3 |
| 2024 | Crowd-Sourced NeRF: Collecting Data From Production Vehicles for 3D Street View ReconstructionabstractRecently, Neural Radiance Fields (NeRF) achieved impressive results in novel view synthesis. Block-NeRF showed the capability of leveraging NeRF to build large city-scale models. For large-scale modeling, a mass of image data is necessary. Collecting images from specially designed data-collection vehicles can not support large-scale applications. How to acquire massive high-quality data remains an opening problem. Noting that the automotive industry has a huge amount of image data, crowd-sourcing is a convenient way for large-scale data collection. In this paper, we present a crowd-sourced framework, which utilizes substantial data captured by production vehicles to reconstruct the scene with the NeRF model. This approach solves the key problem of large-scale reconstruction, that is where the data comes from and how to use them. Firstly, the crowd-sourced massive data is filtered to remove redundancy and keep a balanced distribution in terms of time and space. Then a structure-from-motion module is performed to refine camera poses. Finally, images, as well as poses, are used to train the NeRF model in a certain block. We highlight that we presents a comprehensive framework that integrates multiple modules, including data selection, sparse 3D reconstruction, sequence appearance embedding, depth supervision of ground surface, and occlusion completion. The complete system is capable of effectively processing and reconstructing high-quality 3D scenes from crowd-sourced data. Extensive quantitative and qualitative experiments were conducted to validate the performance of our system. Moreover, we proposed an application, named first-view navigation, which leveraged the NeRF model to generate 3D street view and guide the driver with a synthesized video. Tong Qin 0001, Changze Li, Haoyang Ye, Shaowei Wan, Minzhen Li, Ming Yang 0002 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | FlowMap: Path Generation for Automated Vehicles in Open Space Using Traffic FlowabstractThere is extensive literature on perceiving road structures by fusing various sensor inputs such as lidar point clouds and camera images using deep neural nets. Leveraging the latest advance of neural architects (such as transformers) and bird-eye-view (BEV) representation, the road cognition accuracy keeps improving. However, how to cognize the “road” for automated vehicles where there is no well-defined “roads” remains an open problem. For example, how to find paths inside intersections without HD maps is hard since there is neither an explicit definition for “roads” nor explicit features such as lane markings. The idea of this paper comes from a proverb: it becomes a way when people walk on it. Although there are no “roads” from sensor readings, there are “roads” from tracks of other vehicles. In this paper, we propose FlowMap, a path generation framework for automated vehicles based on traffic flows. FlowMap is built by extending our previous work RoadMap [1], a light-weight semantic map, with an additional traffic flow layer. A path generation algorithm on traffic flow fields (TFFs) is proposed to generate human-like paths. The proposed framework is validated using real-world driving data and is amenable to generating paths for super complicated intersections without using HD maps. Wenchao Ding 0001, Jieru Zhao, Yubin Chu, Haihui Huang, Tong Qin 0001, Chunjing Xu, Zhongxue Gan 0001 |
ICRA | 5 |
| 2023 | Inverse Perspective Mapping-Based Neural Occupancy Grid Map for Visual ParkingabstractSensing environmental obstacles and establishing an occupancy map of surroundings are critical to achieving automated parking for autonomous vehicles. This paper presents a method to obtain surrounding occupancy information from inverse perspective mapping (IPM) images. This method uses the easily-accessed pseudo-labels from LiDAR to supervise a visual network, which can detect occupied boundaries of obstacles. Fusing this visual occupancy with ego-motion information, we develop a multi-frame fusion approach to build a local OGM to realize online environment mapping. Compared with other learning-based occupancy approaches, our method does not require time-consuming and labor-intensive labeling for the environment due to the ground truth of surrounding occupancy coming from LiDAR easily. The proposed method achieves LiDAR-like performance with pure visual inputs, which greatly decreases the cost of real products. Experiments on driving and parking environments prove that our method can accurately sense surrounding occupancy information and build a robust occupancy map of the environment. Xiangru Mu, Haoyang Ye, Daojun Zhu, Tongqing Chen, Tong Qin 0001 |
ICRA | 5 |
| 2021 | A Light-Weight Semantic Map for Visual Localization towards Autonomous DrivingabstractAccurate localization is of crucial importance for autonomous driving tasks. Nowadays, we have seen a lot of sensor-rich vehicles (e.g. Robo-taxi) driving on the street autonomously, which rely on high-accurate sensors (e.g. Lidar and RTK GPS) and high-resolution map. However, low-cost production cars cannot afford such high expenses on sensors and maps. How to reduce costs? How do sensor-rich vehicles benefit low-cost cars? In this paper, we proposed a light-weight localization solution, which relies on low-cost cameras and compact visual semantic maps. The map is easily produced and updated by sensor-rich vehicles in a crowd-sourced way. Specifically, the map consists of several semantic elements, such as lane line, crosswalk, ground sign, and stop line on the road surface. We introduce the whole framework of on-vehicle mapping, on-cloud maintenance, and user-end localization. The map data is collected and preprocessed on vehicles. Then, the crowd-sourced data is uploaded to a cloud server. The mass data from multiple vehicles are merged on the cloud so that the semantic map is updated in time. Finally, the semantic map is compressed and distributed to production cars, which use this map for localization. We validate the performance of the proposed map in real-world experiments and compare it against other algorithms. The average size of the semantic map is 36 kb/km. We highlight that this framework is a reliable and practical localization solution for autonomous driving. Tong Qin 0001, Tongqing Chen |
ICRA | 1 |
| 2021 | Real-Time Temporal and Rotational Calibration of Heterogeneous Sensors Using Motion Correlation AnalysisabstractAccurate and robust calibration is crucial to a multisensor fusion-based system. The calibration of heterogeneous sensors is particularly challenging because of the huge difference of the captured sensor data. On the other hand, many calibration approaches ignore temporal calibration that is in fact as important as spatial calibration. In this article, we focus on the temporal calibration of heterogeneous sensors, and the corresponding extrinsic rotation is also derived. Most existing methods are specialized for a certain sensor combination, such as an inertial measurement unit (IMU) camera or a camera-Lidar system. However, heterogeneous multisensor fusion is a tendency in the robotics area, so a unified calibration method is desired. To this end, we leverage the 3-D rotational motion feature for calibration, and auxiliary calibration boards are not needed since multiple odometry methods are available to capture 3-D sensor motion. Using a high-frequency IMU as the calibration reference, an IMU-centric scheme is designed to achieve a unified framework that adapts to various target sensors that can independently estimate 3-D rotational motion. By combining independent IMU-centric calibration pairs, an arbitrary pair of sensors can also be calibrated using the same reference IMU. Due to a novel 3-D motion correlation quantification and analysis mechanism, the temporal offset can be first estimated in real time. Given temporally aligned sensor motion, the extrinsic rotation can be derived in closed-form in the same 3-D motion correlation mechanism. Experimental results of certain sensor combinations show the accuracy and robustness of the proposed method through comparison with state-of-the-art calibration approaches, and the calibration result of a heterogeneous multisensor set demonstrates the scalability and versatility of our method. Kejie Qiu, Tong Qin 0001, Jie Pan 0005, Siqi Liu 0022, Shaojie Shen |
IEEE Trans. Robotics | 2 |
| 2020 | AVP-SLAM: Semantic Visual Mapping and Localization for Autonomous Vehicles in the Parking LotabstractAutonomous valet parking is a specific application for autonomous vehicles. In this task, vehicles need to navigate in narrow, crowded and GPS-denied parking lots. Accurate localization ability is of great importance. Traditional visual-based methods suffer from tracking lost due to texture-less regions, repeated structures, and appearance changes. In this paper, we exploit robust semantic features to build the map and localize vehicles in parking lots. Semantic features contain guide signs, parking lines, speed bumps, etc, which typically appear in parking lots. Compared with traditional features, these semantic features are long-term stable and robust to the perspective and illumination change. We adopt four surround-view cameras to increase the perception range. Assisting by an IMU (Inertial Measurement Unit) and wheel encoders, the proposed system generates a global visual semantic map. This map is further used to localize vehicles at the centimeter level. We analyze the accuracy and recall of our system and compare it against other methods in real experiments. Furthermore, we demonstrate the practicability of the proposed system by the autonomous parking application. Tong Qin 0001, Tongqing Chen |
IROS | 1 |
| 2019 | Tracking 3-D Motion of Dynamic Objects Using Monocular Visual-Inertial SensingabstractSix degree-of-freedom (6-DoF) visual tracking of dynamic objects is fundamental to a large variety of robotics and augmented reality (AR) applications. A key to this problem is accurate distance measurement of dynamic objects, which is usually obtained via stereo cameras, RGB-D sensors, or LiDARs. In this paper, however, we address the problem using only a monocular camera rigidly mounted with a low-cost inertial measurement unit. This is a light-weight, small-size, and low-cost solution, which is particularly suitable for tracking dynamic objects on drones or on mobile phones. Starting from a generic image-based two-dimensional tracker, we propose a novel method to resolve the object scale ambiguity in monocular vision in a geometric manner based on correlation analysis. This enables accurate metric three-dimensional tracking of arbitrary objects without requiring any prior knowledge about the object shape or size. We discuss the applicability by analyzing the observability condition and degenerated cases for object scale recovery. Simulation and real-world experimental results with ground truth comparison, along with AR application examples, demonstrate the feasibility of the proposed 6-DoF tracking method. Kejie Qiu, Tong Qin 0001, Wenliang Gao, Shaojie Shen |
IEEE Trans. Robotics | 2 |
| 2018 | Stereo Vision-Based Semantic 3D Object and Ego-Motion Tracking for Autonomous Driving
Peiliang Li 0001, Tong Qin 0001, Shaojie Shen |
ECCV (2) | 2 |
| 2018 | SLAM-based localization of 3D gaze using a mobile eye trackerabstractPast work in eye tracking has focused on estimating gaze targets in two dimensions (2D), e.g. on a computer screen or scene camera image. Three-dimensional (3D) gaze estimates would be extremely useful when humans are mobile and interacting with the real 3D environment. We describe a system for estimating the 3D locations of gaze using a mobile eye tracker. The system integrates estimates of the user's gaze vector from a mobile eye tracker, estimates of the eye tracker pose from a visual-inertial simultaneous localization and mapping (SLAM) algorithm, a 3D point cloud map of the environment from a RGB-D sensor. Experimental results indicate that our system produces accurate estimates of 3D gaze over a much larger range than remote eye trackers. Our system will enable applications, such as the analysis of 3D human attention and more anticipative human robot interfaces. Haofei Wang 0001, Jimin Pi, Tong Qin 0001, Shaojie Shen, Bertram E. Shi |
ETRA | 3 |
| 2018 | Relocalization, Global Optimization and Map Merging for Monocular Visual-Inertial SLAMabstractThe monocular visual-inertial system (VINS), which consists one camera and one low-cost inertial measurement unit (IMU), is a popular approach to achieve accurate 6-DOF state estimation. However, such locally accurate visual-inertial odometry is prone to drift and cannot provide absolute pose estimation. Leveraging history information to relocalize and correct drift has become a hot topic. In this paper, we propose a monocular visual-inertial SLAM system, which can relocalize camera and get the absolute pose in a previous-built map. Then 4-DOF pose graph optimization is performed to correct drifts and achieve global consistent. The 4-DOF contains x, y, z, and yaw angle, which is the actual drifted direction in the visual-inertial system. Furthermore, the proposed system can reuse a map by saving and loading it in an efficient way. Current map and previous map can be merged together by the global pose graph optimization. We validate the accuracy of our system on public datasets and compare against other state-of-the-art algorithms. We also evaluate the map merging ability of our system in the large-scale outdoor environment. The source code of map reuse is integrated into our public code, VINS-Monol11https://github.com/HKUST-Aerial-Robotics/VINS-Mono. Tong Qin 0001, Peiliang Li 0001, Shaojie Shen |
ICRA | 1 |
| 2018 | Online Temporal Calibration for Monocular Visual-Inertial SystemsabstractAccurate state estimation is a fundamental module for various intelligent applications, such as robot navigation, autonomous driving, virtual and augmented reality. Visual and inertial fusion is a popular technology for 6-DOF state estimation in recent years. Time instants at which different sensors' measurements are recorded are of crucial importance to the system's robustness and accuracy. In practice, timestamps of each sensor typically suffer from triggering and transmission delays, leading to temporal misalignment (time offsets) among different sensors. Such temporal offset dramatically influences the performance of sensor fusion. To this end, we propose an online approach for calibrating temporal offset between visual and inertial measurements. Our approach achieves temporal offset calibration by jointly optimizing time offset, camera and IMU states, as well as feature locations in a SLAM system. Furthermore, the approach is a general model, which can be easily employed in several feature-based optimization frameworks. Simulation and experimental results demonstrate the high accuracy of our calibration approach even compared with other state-of-art offline tools. The VIO comparison against other methods proves that the online temporal calibration significantly benefits visual-inertial systems. The source code of temporal calibration is integrated into our public project, VINS-Mono1. Tong Qin 0001, Shaojie Shen |
IROS | 1 |
| 2018 | Estimating Metric Poses of Dynamic Objects Using Monocular Visual-Inertial FusionabstractA monocular 3D object tracking system generally has only up-to-scale pose estimation results without any prior knowledge of the tracked object. In this paper, we propose a novel idea to recover the metric scale of an arbitrary dynamic object by optimizing the trajectory of the objects in the world frame, without motion assumptions. By introducing an additional constraint in the time domain, our monocular visual-inertial tracking system can obtain continuous six degree of freedom (6-DoF) pose estimation without scale ambiguity. Our method requires neither fixed multi-camera nor depth sensor settings for scale observability, instead, the IMU inside the monocular sensing suite provides scale information for both camera itself and the tracked object. We build the proposed system on top of our monocular visual-inertial system (VINS) to obtain accurate state estimation of the monocular camera in the world frame. The whole system consists of a 2D object tracker, an object region-based visual bundle adjustment (BA), VINS and a correlation analysis-based metric scale estimator. Experimental comparisons with ground truth demonstrate the tracking accuracy of our 3D tracking performance while a mobile augmented reality (AR) demo shows the feasibility of potential applications. Kejie Qiu, Tong Qin 0001, Hongwen Xie, Shaojie Shen |
IROS | 2 |
| 2018 | VINS-Mono: A Robust and Versatile Monocular Visual-Inertial State EstimatorabstractOne camera and one low-cost inertial measurement unit (IMU) form a monocular visual-inertial system (VINS), which is the minimum sensor suite (in size, weight, and power) for the metric six degrees-of-freedom (DOF) state estimation. In this paper, we present VINS-Mono: a robust and versatile monocular visual-inertial state estimator. Our approach starts with a robust procedure for estimator initialization. A tightly coupled, nonlinear optimization-based method is used to obtain highly accurate visual-inertial odometry by fusing preintegrated IMU measurements and feature observations. A loop detection module, in combination with our tightly coupled formulation, enables relocalization with minimum computation. We additionally perform 4-DOF pose graph optimization to enforce the global consistency. Furthermore, the proposed system can reuse a map by saving and loading it in an efficient way. The current and previous maps can be merged together by the global pose graph optimization. We validate the performance of our system on public datasets and real-world experiments and compare against other state-of-the-art algorithms. We also perform an onboard closed-loop autonomous flight on the microaerial-vehicle platform and port the algorithm to an iOS-based demonstration. We highlight that the proposed work is a reliable, complete, and versatile system that is applicable for different applications that require high accuracy in localization. We open source our implementations for both PCs (https://github.com/HKUST-Aerial-Robotics/VINS-Mono) and iOS mobile devices (https://github.com/HKUST-Aerial-Robotics/VINS-Mobile). Tong Qin 0001, Peiliang Li 0001, Shaojie Shen |
IEEE Trans. Robotics | 1 |
| 2017 | Robust initialization of monocular visual-inertial estimation on aerial robotsabstractIn this paper, we propose a robust on-the-fly estimator initialization algorithm to provide high-quality initial states for monocular visual-inertial systems (VINS). Due to the non-linearity of VINS, a poor initialization can severely impact the performance of either filtering-based or graph-based methods. Our approach starts with a vision-only structure from motion (SfM) to build the up-to-scale structure of camera poses and feature positions. By loosely aligning this structure with pre-integrated IMU measurements, our approach recovers the metric scale, velocity, gravity vector, and gyroscope bias, which are treated as initial values to bootstrap the nonlinear tightly-coupled optimization framework. We highlight that our approach can perform on-the-fly initialization in various scenarios without using any prior information about system states and movement. The performance of the proposed approach is verified through the public UAV dataset and real-time onboard experiment. We make our implementation open source, which is the initialization part integrated in the VINS-Mono1. Tong Qin 0001, Shaojie Shen |
IROS | 1 |
| 2017 | Monocular Visual-Inertial State Estimation for Mobile Augmented RealityabstractMobile phones equipped with a monocular camera and an inertial measurement unit (IMU) are ideal platforms for augmented reality (AR) applications, but the lack of direct metric distance measurement and the existence of aggressive motions pose significant challenges on the localization of the AR device. In this work, we propose a tightly-coupled, optimization-based, monocular visual-inertial state estimation for robust camera localization in complex indoor and outdoor environments. Our approach does not require any artificial markers, and is able to recover the metric scale using the monocular camera setup. The whole system is capable of online initialization without relying on any assumptions about the environment. Our tightly-coupled formulation makes it naturally robust to aggressive motions. We develop a lightweight loop closure module that is tightly integrated with the state estimator to eliminate drift. The performance of our proposed method is demonstrated via comparison against state-of-the-art visual-inertial state estimators on public datasets and real-time AR applications on mobile devices. We release our implementation on mobile devices as open source software1. Peiliang Li 0001, Tong Qin 0001, Botao Amber Hu, Shaojie Shen |
ISMAR | 2 |