VLDB 2026 Research / reviewers in the wild / expert
Shaojie Shen
dblp:07/9968
· DBLP profile ↗
124ranked-venue papers
6as first author
58since 2021 · last 2026
0000-0002-5573-2909ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 90 · 6 first-author · 36 since 2021Systems, architecture and hardware · 76 · 6 first-author · 27 since 2021Applied, interdisciplinary, general and emerging computing · 29 · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 since 2021Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fast and Scalable Game-Theoretic Trajectory Planning With Intentional UncertaintiesabstractTrajectory planning involving multi-agent interactions has been a long-standing challenge in the field of robotics, primarily burdened by the inherent yet intricate interactions among agents. While game-theoretic methods are widely acknowledged for their effectiveness in managing multi-agent interactions, significant impediments persist when it comes to accommodating the intentional uncertainties of agents. In the context of intentional uncertainties, the heavy computational burdens associated with existing game-theoretic methods are induced, leading to inefficiencies and poor scalability. In this paper, we propose a novel game-theoretic interactive trajectory planning method to effectively address the intentional uncertainties of agents, and it demonstrates both high efficiency and enhanced scalability. As the underpinning basis, we model the interactions between agents under intentional uncertainties as a static Bayesian game, and we show that its agent-form equivalence can be represented as a potential game under certain assumptions. The existence and attainability of the optimal interactive trajectories are illustrated, as the corresponding static Bayesian Nash equilibrium can be attained by optimizing a unified optimization problem. Additionally, we present a distributed algorithm based on the dual consensus alternating direction method of multipliers (ADMM) tailored to the parallel solving of the problem, thereby significantly improving the scalability. The attendant outcomes from simulations and experiments demonstrate that the proposed method is effective across a range of scenarios characterized by general forms of intentional uncertainties. Its scalability surpasses that of existing centralized and decentralized baselines, allowing for real-time interactive trajectory planning in uncertain game settings. The source code will be available onhttps://github.com/zhuangdf/Potential-Bayesian-Game-release. Zhenmin Huang, Yusen Xie, Benshan Ma, Shaojie Shen, Jun Ma 0008 |
IEEE Trans. Robotics | 4 |
| 2025 | Foresight in Motion: Reinforcing Trajectory Prediction with Reward HeuristicsabstractMotion forecasting for on-road traffic agents presents both a significant challenge and a critical necessity for ensuring safety in autonomous driving systems. In contrast to most existing data-driven approaches that directly predict future trajectories, we rethink this task from a planning perspective, advocating a "First Reasoning, Then Forecasting" strategy that explicitly incorporates behavior intentions as spatial guidance for trajectory prediction. To achieve this, we introduce an interpretable, reward-driven intention reasoner grounded in a novel query-centric Inverse Reinforcement Learning (IRL) scheme. Our method first encodes traffic agents and scene elements into a unified vectorized representation, then aggregates contextual features through a query-centric paradigm. This enables the derivation of a reward distribution, a compact yet informative representation of the target agent's behavior within the given scene context via IRL. Guided by this reward heuristic, we perform policy rollouts to reason about multiple plausible intentions, providing valuable priors for subsequent trajectory generation. Finally, we develop a hierarchical DETR-like decoder integrated with bidirectional selective state space models to produce accurate future trajectories along with their associated probabilities. Extensive experiments on the large-scale Argoverse and nuScenes motion forecasting datasets demonstrate that our approach significantly enhances trajectory prediction confidence, achieving highly competitive performance relative to state-of-the-art methods. Muleilan Pei, Shaoshuai Shi, Shaojie Shen |
ICCV | 5 |
| 2025 | GoIRL: Graph-Oriented Inverse Reinforcement Learning for Multimodal Trajectory PredictionabstractTrajectory prediction for surrounding agents is a challenging task in autonomous driving due to its inherent uncertainty and underlying multimodality. Unlike prevailing data-driven methods that primarily rely on supervised learning, in this paper, we introduce a novel Graph-oriented Inverse Reinforcement Learning (GoIRL) framework, which is an IRL-based predictor equipped with vectorized context representations. We develop a feature adaptor to effectively aggregate lane-graph features into grid space, enabling seamless integration with the maximum entropy IRL paradigm to infer the reward distribution and obtain the policy that can be sampled to induce multiple plausible plans. Furthermore, conditioned on the sampled plans, we implement a hierarchical parameterized trajectory generator with a refinement module to enhance prediction accuracy and a probability fusion strategy to boost prediction confidence. Extensive experimental results showcase our approach not only achieves state-of-the-art performance on the large-scale Argoverse & nuScenes motion forecasting benchmarks but also exhibits superior generalization abilities compared to existing supervised models. Muleilan Pei, Shaoshuai Shi, Lu Zhang 0047, Peiliang Li 0001, Shaojie Shen |
ICML | 5 |
| 2025 | SLABIM: A SLAM-BIM Coupled Dataset in HKUST Main BuildingabstractExisting indoor SLAM datasets primarily focus on robot sensing, often lacking building architectures. To address this gap, we design and construct the first dataset to couple the SLAM and BIM, named SLABIM. This dataset provides BIM and SLAM -oriented sensor data, both modeling a university building at HKUST. The as-designed BIM is decomposed and converted for ease of use. We employ a multi-sensor suite for multi-session data collection and mapping to obtain the as-built model. All the related data are timestamped and organized, enabling users to deploy and test effectively. Furthermore, we deploy advanced methods and report the experimental results on three tasks: registration, localization and semantic mapping, demonstrating the effectiveness and practicality of SLAB 1M. We make our dataset open-source at https://github.com/HKUST-Aerial-Robotics/SLABIM. Haoming Huang, Zhijian Qiao, Zehuan Yu, Chuhao Liu, Shaojie Shen, Fumin Zhang 0001, Huan Yin |
ICRA | 5 |
| 2025 | Multimodal Integrated Prediction and Decision-making with Adaptive Interaction Modality ExplorationsabstractNavigating dense and dynamic environments poses a significant challenge for autonomous driving systems, owing to the intricate nature of multimodal interaction, wherein the actions of various traffic participants and the autonomous vehicle are complex and implicitly coupled. In this paper, we propose a novel framework, Multimodal Integrated predictioN and Decision-making (MIND), which addresses the challenges by efficiently generating joint predictions and decisions covering multiple distinctive interaction modalities. Specifically, MIND leverages learning-based scenario predictions to obtain integrated predictions and decisions with socially-consistent interaction modality and utilizes a modality-aware dynamic branching mechanism to generate scenario trees that efficiently capture the evolutions of distinctive interaction modalities with low growth of interaction uncertainty along the planning horizon. The scenario trees are seamlessly utilized by the contingency planning under interaction uncertainty to obtain clear and considerate maneuvers accounting for multimodal evolutions. Comprehensive experimental results in the closed-loop simulation based on the real-world driving dataset showcase superior performance to other strong baselines under various driving contexts. Code is available at: https://github.com/HKUST-Aerial-Robotics/MIND. Lu Zhang 0047, Sikang Liu 0002, Shaojie Shen |
IROS | 4 |
| 2025 | LMMCoDrive: Cooperative Driving with Large Multimodal ModelsabstractTo address the intricate challenges of cooperative scheduling and motion planning in Autonomous Mobility-on-Demand (AMoD) systems, this paper introduces LMMCoDrive, a novel cooperative driving framework that leverages a Large Multimodal Model (LMM) to improve traffic efficiency and passenger experience in dynamic urban environments. This framework seamlessly integrates scheduling and motion planning processes to ensure the effective operation of Cooperative Autonomous Vehicles (CAVs). The spatial relationship between CAVs and passenger requests is abstracted into a Bird’s-Eye View (BEV) image to fully exploit the potential of the multimodal understanding ability of LMMs. Besides, trajectories are cautiously refined for each CAV while ensuring collision avoidance through safety constraints. A decentralized optimization strategy, facilitated by the Alternating Direction Method of Multipliers (ADMM) within the LMM framework, is proposed to drive the graph evolution of CAVs. Simulation results in diverse urban scenarios demonstrate the pivotal role and significant impact of LMM in optimizing CAV scheduling and seamlessly serving a decentralized cooperative optimization process for each CAV. This marks a substantial stride towards practical, efficient, and safe AMoD systems that are poised to revolutionize urban transportation. The code is available at https://github.com/henryhcliu/LMMCoDrive. Haichao Liu 0003, Ruoyu Yao, Zhenmin Huang, Shaojie Shen, Jun Ma 0008 |
IROS | 4 |
| 2025 | Direct, Targetless and Automatic Joint Calibration of LiDAR-Camera Intrinsic and ExtrinsicabstractThis paper presents a direct, targetless, and automatic LiDAR-Camera joint calibration method that effectively overcomes the intrinsic precision limitations. We propose an iterative two-stage optimization methodology that leverages 3D LiDAR measurements to simultaneously refine both intrinsic and extrinsic. In the first stage, the intrinsic is optimized using a normalized information distance (NID) metric, an information-theoretic measure that quantifies the statistical alignment between LiDAR and image intensities, while initial extrinsic parameters derived from CAD specifications facilitate the projection of LiDAR point clouds onto the camera image plane. In the second stage, the refined intrinsic guides further optimization of extrinsic using the same NID-based evaluation metrics. This alternating process iteratively enhances both intrinsic and extrinsic through their mutual interdependence. Experiments across multiple datasets demonstrate that our method achieves sub-pixel intrinsic accuracy and extrinsic parameters that closely align with CAD specifications, validating the superior performance of our methodology for sensor fusion applications. Yishu Shen, Shaojie Shen, Tong Qin 0001 |
IROS | 3 |
| 2025 | From Satellite to Street: Semantic and Depth Information for Enhanced Geo-LocalizationabstractAccurate positioning is essential for autonomous driving, but localization using 2D maps is challenging due to the domain gap between perspective view and 2D map. While GNSS accuracy is often limited by atmospheric effects, multipath, and signal blockages. We propose a novel positioning method that combines perspective view images with satellite images retrieved based on rough GNSS positions to achieve precise three-degree-of-freedom (3-DoF) pose estimation. Our method leverages the Swin Transformer for satellite image processing and semantic completion for monocular image analysis. By extracting depth and semantic information from monocular images, we convert these to overhead projections, effectively bridging the gap between different viewpoints. This cross-view transformation allows for precise alignment of features from monocular images onto semantically enriched satellite images. Additionally, we integrate a robust global position estimator using the semantic information from satellite images to further enhance accuracy and robustness. The experimental results demonstrate that our method excels in various complex scenarios; we successfully improved the positioning accuracy within 1 m to 80.67% and the heading in 1° to 33.78%. However, longitudinal localization remains more challenging, with higher errors than lateral positioning. Yilong Zhu, Jianhao Jiao, Hexiang Wei, Jin Wu 0002, Bohuan Xue, Shaojie Shen |
IROS | 7 |
| 2025 | Speak the Same Language: Global LiDAR Registration on BIM Using Pose Hough TransformabstractLight detection and ranging (LiDAR) point clouds and building information modeling (BIM) represent two distinct data modalities in the fields of robot perception and construction. These modalities originate from different sources and are associated with unique reference frames. The primary goal of this study is to align these modalities within a shared reference frame using a global registration approach, effectively enabling them to “speak the same language”. To achieve this, we propose a cross-modality registration method, spanning from the front end to the back end. At the front end, we extract triangle descriptors by identifying walls and intersected corners, enabling the matching of corner triplets with a complexity independent of the BIM’s size. For the back-end transformation estimation, we utilize the Hough transform to map the matched triplets to the transformation space and introduce a hierarchical voting mechanism to hypothesize multiple pose candidates. The final transformation is then verified using our designed occupancy-aware scoring method. To assess the effectiveness of our approach, we conducted real-world multi-session experiments in a large-scale university building, employing two different types of LiDAR sensors. We make the collected datasets and codes publicly available to benefit the community. Note to Practitioners—Our proposed registration method leverages walls and corners as shared features between LiDAR and BIM data, making it particularly well-suited for scenarios with well-defined structural layouts. Accumulating a larger LiDAR submap provides richer structural information, which further aids in achieving accurate alignment. To optimize computational efficiency, we recommend constructing the descriptor database offline and loading it during runtime, enabling a theoretical retrieval complexity of$O(1)$. Despite its advantages, our approach has certain limitations. First, it primarily focuses on planar structures, which limits its effectiveness in utilizing non-planar features. Second, the method may underperform in cases where significant deviations exist between the as-designed BIM and as-is LiDAR data. Lastly, in ambiguous scenarios, such as long corridors or similar layouts within the same or across different floors, our method may struggle to verify the correct transformation among candidates. To address these challenges, incorporating additional information, particularly semantic cues such as floor numbers, room numbers, and room types, could enhance its robustness and reliability. Zhijian Qiao, Haoming Huang, Chuhao Liu, Zehuan Yu, Shaojie Shen, Fumin Zhang 0001, Huan Yin |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2025 | G3Reg: Pyramid Graph-Based Global Registration Using Gaussian Ellipsoid ModelabstractThis study introduces a novel framework, G3Reg, for fast and robust global registration of LiDAR point clouds. In contrast to conventional complex keypoints and descriptors, we extract fundamental geometric primitives, including planes, clusters, and lines (PCL) from the raw point cloud to obtain low-level semantic segments. Each segment is represented as a unified Gaussian Ellipsoid Model (GEM), using a probability ellipsoid to ensure the ground truth centers are encompassed with a certain degree of probability. Utilizing these GEMs, we present a distrust-and-verify scheme based on a Pyramid Compatibility Graph for Global Registration (PAGOR). Specifically, we establish an upper bound, which can be traversed based on the confidence level for compatibility testing to construct the pyramid graph. Then, we solve multiple maximum cliques (MAC) for each level of the pyramid graph, thus generating the corresponding transformation candidates. In the verification phase, we adopt a precise and efficient metric for point cloud alignment quality, founded on geometric primitives, to identify the optimal candidate. The algorithm’s performance is validated on three publicly available datasets and a self-collected multi-session dataset. Parameter settings remained unchanged during the experiment evaluations. The results exhibit superior robustness and real-time performance of the G3Reg framework compared to state-of-the-art methods. Furthermore, we demonstrate the potential for integrating individual GEM and PAGOR components into other registration frameworks to enhance their efficacy.Note to Practitioners—Our proposed method aims to perform global registration for outdoor LiDAR point clouds. Our methodology, which extracts point cloud segments and utilizes their centers for registration, differs from conventional approaches that rely on keypoints and descriptors. We further propose GEM to model the uncertainty of the centers and embed it into our distrust-and-verify framework. In theory, our method can be applied to any registration task that involves primitives representable as sets of Gaussians or points. Additionally, practitioners should consider the following to enhance applicability. First, practitioners can fine-tune the parameters of the segmentation algorithm to generate more repeatable segmentation results. Second, although our default setting uses four compatibility test thresholds, fewer may suffice, especially when translations between point clouds are minor. Finally, for geometrically uninformative segments such as vegetation, consider extracting descriptors within these segments to increase correspondences. Zhijian Qiao, Zehuan Yu, Binqian Jiang, Huan Yin, Shaojie Shen |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2025 | GS-LIVO: Real-Time LiDAR, Inertial, and Visual Multisensor Fused Odometry With Gaussian MappingabstractIn recent years, 3D Gaussian splatting (3D-GS) has emerged as a novel scene representation approach. However, existing vision-only 3D-GS methods often rely on hand-crafted heuristics for point-cloud densification and face challenges in handling occlusions and high GPU memory and computation consumption [1]. LiDAR-Inertial-Visual (LIV) sensor configuration has demonstrated superior performance in precise localization and dense mapping by leveraging complementary sensing characteristics: rich texture information from cameras, precise geometric measurements from LiDAR, and high-frequency motion data from IMU [2]-[8]. Inspired by this, we propose a novel real-time Gaussian-based simultaneous localization and mapping (SLAM) system. Our map system comprises a global Gaussian map and a sliding window of Gaussians, along with an IESKF-based real-time odometry utilizing Gaussian maps. The structure of the global Gaussian map consists of hash-indexed voxels organized in a recursive octree. This hierarchical structure effectively covers sparse spatial volumes while adapting to different levels of detail and scales in the environment. The Gaussian map is efficiently initialized through multi-sensor fusion and optimized with photometric gradients. Our system incrementally maintains a sliding window of Gaussians with minimal graphics memory usage, significantly reducing GPU computation and memory consumption by only optimizing the map within the sliding window, enabling real-time optimization. Moreover, we implement a tightly coupled multi-sensor fusion odometry with an iterative error state Kalman filter (IESKF), which leverages real-time updating and rendering of the Gaussian map to achieve competitive localization accuracy. Our system represents the first real-time Gaussian-based SLAM framework deployable on resource-constrained embedded systems (all implemented in C++/CUDA for efficiency), demonstrated on theNVIDIA Jetson Orin NXplatform. The framework achieves real-time performance while maintaining robust multi-sensor fusion capabilities. All implementation algorithms, hardware designs, and CAD models and demo video of our GPU-accelerated system will be publicly available athttps://github.com/HKUST-Aerial-Robotics/GS-LIVO. Chunran Zheng, Yishu Shen, Changze Li, Fu Zhang 0002, Tong Qin 0001, Shaojie Shen |
IEEE Trans. Robotics | 7 |
| 2025 | SG-Reg: Generalizable and Efficient Scene Graph RegistrationabstractThis paper addresses the challenges of registering two rigid semantic scene graphs, an essential capability when an autonomous agent needs to register its map against a remote agent, or against a prior map. The hand-crafted descriptors in classical semantic-aided registration, or the ground-truth annotation reliance in learning-based scene graph registration, impede their application in practical real-world environments. To address the challenges, we design a scene graph network to encode multiple modalities of semantic nodes: open-set semantic feature, local topology with spatial awareness, and shape feature. These modalities are fused to create compact semantic node features. The matching layers then search for correspondences in a coarse-to-fine manner. In the back-end, we employ a robust pose estimator to decide transformation according to the correspondences. We manage to maintain a sparse and hierarchical scene representation. Our approach demands fewer GPU resources and fewer communication bandwidth in multi-agent tasks. Moreover, we design a new data generation approach using vision foundation models and a semantic mapping module to reconstruct semantic scene graphs. It differs significantly from previous works, which rely on ground-truth semantic annotations to generate data. We validate our method in a two-agent SLAM benchmark. It significantly outperforms the hand-crafted baseline in terms of registration success rate. Compared to visual loop closure networks, our method achieves a slightly higher registration recall while requiring only 52 KB of communication bandwidth for each query frame. Code available at:http://github.com/HKUST-Aerial-Robotics/SG-Reg Chuhao Liu, Zhijian Qiao, Jieqi Shi, Ke Wang 0058, Peize Liu, Shaojie Shen |
IEEE Trans. Robotics | 6 |
| 2025 | Event-Based Visual-Inertial State Estimation for High-Speed ManeuversabstractNeuromorphic event-based cameras are bio-inspired visual sensors with asynchronous pixels and extremely high temporal resolution. Such favorable properties make them an excellent choice for solving state estimation tasks under high-speed maneuvers. However, failures of camera pose tracking are frequently witnessed in state-of-the-art event-based visual odometry systems when the local map cannot be updated timely or feature matching is unreliable. One of the biggest roadblocks in this field is the absence of efficient and robust methods for data association without imposing any assumptions on the environment. This problem seems, however, unlikely to be addressed as in standard vision because of the motion-dependent nature of event data. To address this, we propose a map-free design for event-based visual-inertial state estimation in this paper. Instead of estimating camera position, we find that recovering the instantaneous linear velocity aligns better with event cameras' differential working principle. The proposed system uses raw data from a stereo event camera and an inertial measurement unit (IMU) as input, and adopts a dual-end architecture. The front-end preprocesses raw events and executes the computation of normal flow and depth information. To handle the temporally non-equispaced event data and establish association with temporally non-aligned IMU's measurements, the back-end employs a continuous-time formulation and a sliding-window scheme that can progressively estimate the linear velocity and IMU's bias. Experiments on synthetic and real data show our method achieves low-latency, metric-scale velocity estimation. To the best of our knowledge, this is the first real-time, purely event-based visual-inertial state estimator for high-speed maneuvers, requiring only sufficient textures and imposing no additional constraints on either the environment or motion pattern. Xiuyuan Lu, Yi Zhou 0010, Jiayao Mai, Kuan Dai, Yang Xu 0083, Shaojie Shen |
IEEE Trans. Robotics | 6 |
| 2025 | ESVO2: Direct Visual-Inertial Odometry With Stereo Event CamerasabstractEvent-based visual odometry is a specific branch of visual simultaneous localization and mapping (SLAM) techniques, which aims at solving tracking and mapping subproblems (typically in parallel), by exploiting the special working principles of neuromorphic (i.e., event-based) cameras. Due to the motion-dependent nature of event data, explicit data association (i.e., feature matching) under large-baseline viewpoint changes is difficult to establish, making direct methods a more rational choice. However, state-of-the-art direct methods are limited by the high computational complexity of the mapping subproblem and the degeneracy of camera pose tracking in certain degrees of freedom (DoF) in rotation. In this article, we tackle these issues by building an event-based stereo visual-inertial odometry system, which is built upon a direct pipeline known as event-based stereo visual odometry (ESVO). Specifically, to speed up the mapping operation, we propose an efficient strategy for sampling contour points according to the local dynamics of events. The mapping performance is also improved in terms of structure completeness and local smoothness by merging the temporal stereo and static stereo results. To circumvent the degeneracy of camera pose tracking in recovering the pitch and yaw components of general 6-DoF motion, we introduce IMU measurements as motion priors via preintegration. To this end, a compact back-end is proposed for continuously updating the IMU bias and predicting the linear velocity, enabling an accurate motion prediction for camera pose tracking. The resulting system scales well with modern high-resolution event cameras and leads to better global positioning accuracy in large-scale outdoor environments. Extensive evaluations on five publicly available datasets featuring different resolutions and scenarios justify the superior performance of the proposed system against five state-of-the-art methods. Compared to ESVO, our new pipeline significantly reduces the camera pose tracking error by 40%–80% and 20%–80% in terms of absolute trajectory error and relative pose error, respectively; at the same time, the mapping efficiency is improved by a factor of five. We release our pipeline as an open-source software for future research in this field. Junkai Niu, Xiuyuan Lu, Shaojie Shen, Guillermo Gallego 0002, Yi Zhou 0010 |
IEEE Trans. Robotics | 4 |
| 2025 | Autonomous Flights Inside Narrow TunnelsabstractMultirotors are usually desired to enter confined narrow tunnels that are barely accessible to humans in various applications including inspection, search and rescue, and so on. This task is extremely challenging since the lack of geometric features and illuminations, together with the limited field of view, cause problems in perception; the restricted space and significant ego airflow disturbances induce control issues. This article introduces an autonomous aerial system designed for navigation through tunnels as narrow as 0.5 m in diameter. The real-time and online system includes a virtual omni-directional perception module tailored for the mission and a novel motion planner that incorporates perception and ego airflow disturbance factors modeled using camera projections and computational fluid dynamics analyses, respectively. Extensive flight experiments on a custom-designed quadrotor are conducted in multiple realistic narrow tunnels to validate the superior performance of the system, even over human pilots, proving its potential for real applications. In addition, a deployment pipeline on other multirotor platforms is outlined and open-source packages are provided for future developments. Yan Ning, Hongming Chen 0005, Peize Liu, Yang Xu 0083, Hao Xu 0032, Ximin Lyu, Shaojie Shen |
IEEE Trans. Robotics | 8 |
| 2025 | SLIM: Scalable and Lightweight LiDAR Mapping in Urban EnvironmentsabstractLight detection and ranging (LiDAR) point cloud maps are extensively utilized on roads for robot navigation due to their high consistency. However, dense point clouds face challenges of high memory consumption and reduced maintainability for long-term operations. In this study, we introduce scalable and lightweight LiDAR mapping (SLIM), a scalable and lightweight mapping system for long-term LiDAR mapping in urban environments. The system begins by parameterizing structural point clouds into lines and planes. These lightweight and structural representations meet the requirements of map merging, pose graph optimization, and bundle adjustment, ensuring incremental management and local consistency. For long-term operations, a map-centric nonlinear factor recovery method is designed to sparsify poses while preserving mapping accuracy. We validate the SLIM system with multisession real-world LiDAR data from classical LiDAR mapping datasets, including KITTI, NCLT, HeLiPR, and M2DGR. The experiments demonstrate its capabilities in mapping accuracy, lightweightness, and scalability. Map reuse is also verified through map-based robot localization. Finally, with multisession LiDAR data, the SLIM system provides a globally consistent map with low memory consumption ($\sim$130 KB/km on KITTI). Zehuan Yu, Zhijian Qiao, Huan Yin, Shaojie Shen |
IEEE Trans. Robotics | 5 |
| 2025 | FALCON: Fast Autonomous Aerial Exploration Using Coverage Path GuidanceabstractIn this article, we introduce a novelFastAutonomous expLoration framework usingCOverage path guidaNce (FALCON), which aims at setting a new performance benchmark in the field of autonomous aerial exploration. Despite recent advancements in the domain, existing exploration planners often suffer from inefficiencies, such as frequent revisitations of previously explored regions. FALCON effectively harnesses the full potential of online generated coverage paths in enhancing exploration efficiency. The framework begins with an incremental connectivity-aware space decomposition and connectivity graph construction, which facilitate efficient coverage path planning. Subsequently, a hierarchical planner generates a coverage path spanning the entire unexplored space, serving as a global guidance. Then, a local planner optimizes the frontier visitation order, minimizing traversal time while consciously incorporating the intention of the global guidance. Finally, minimum-time smooth and safe trajectories are produced to visit the frontier viewpoints. For fair and comprehensive benchmark experiments, we introduce a lightweightexploration planner evaluation environmentthat allows for comparing exploration planners across a variety of testing scenarios using an identical quadrotor simulator. In addition, an in-depth analysis and evaluation is conducted to highlight the significant performance advantages of FALCON in comparison with the state-of-the-art exploration planners based on objective criteria. Extensive ablation studies demonstrate the effectiveness of each component in the proposed framework. Real-world experiments conducted fully onboard further validate FALCON’s practical capability in complex and challenging environments. The source code of both the exploration planner FALCON and the exploration planner evaluation environment has been released to benefit the community. Xinyi Chen 0002, Chen Feng 0006, Boyu Zhou, Shaojie Shen |
IEEE Trans. Robotics | 5 |
| 2024 | GeoWizard: Unleashing the Diffusion Priors for 3D Geometry Estimation from a Single Image
Wei Yin 0006, Mu Hu, Yuexin Ma, Ping Tan 0002, Shaojie Shen, Dahua Lin, Xiaoxiao Long |
ECCV (22) | 7 |
| 2024 | Event-Aided Time-to-Collision Estimation for Autonomous Driving
Bangyan Liao, Xiuyuan Lu, Peidong Liu 0001, Shaojie Shen, Yi Zhou 0010 |
ECCV (54) | 5 |
| 2024 | FC-Planner: A Skeleton-guided Planning Framework for Fast Aerial Coverage of Complex 3D Scenesabstract3D coverage path planning for UAVs is a crucial problem in diverse practical applications. However, existing methods have shown unsatisfactory system simplicity, computation efficiency, and path quality in large and complex scenes. To address these challenges, we propose FC-Planner, a skeleton-guided planning framework that can achieve fast aerial coverage of complex 3D scenes without pre-processing. We decompose the scene into several simple subspaces by a skeleton-based space decomposition (SSD). Additionally, the skeleton guides us to effortlessly determine free space. We utilize the skeleton to efficiently generate a minimal set of specialized and informative viewpoints for complete coverage. Based on SSD, a hierarchical planner effectively divides the large planning problem into independent sub-problems, enabling parallel planning for each subspace. The carefully designed global and local planning strategies are then incorporated to guarantee both high quality and efficiency in path generation. We conduct extensive benchmark and real-world tests, where FC-Planner computes over 10 times faster compared to state-of-the-art methods with shorter path and more complete coverage. The source code will be made publicly available to benefit the community3. Project page: https://hkust-aerial-robotics.github.io/FC-Planner. Chen Feng 0006, Haojia Li, Xinyi Chen 0002, Boyu Zhou, Shaojie Shen |
ICRA | 6 |
| 2024 | APACE: Agile and Perception-aware Trajectory Generation for Quadrotor FlightsabstractVarious perception-aware planning approaches have attempted to enhance the state estimation accuracy during maneuvers, while the feature matchability among frames, a crucial factor influencing estimation accuracy, has often been overlooked. In this paper, we present APACE, an Agile and Perception-Aware trajeCtory gEneration framework for quadrotors aggressive flight, that takes into account feature matchability during trajectory planning. We seek to generate a perception-aware trajectory that reduces the error of visual-based estimator while satisfying the constraints on smoothness, safety, agility and the quadrotor dynamics. The perception objective is achieved by maximizing the number of covisible features while ensuring small enough parallax angles. Additionally, we propose a differentiable and accurate visibility model that allows decomposition of the trajectory planning problem for efficient optimization resolution. Through validations conducted in both a photorealistic simulator and real-world experiments, we demonstrate that the trajectories generated by our method significantly improve state estimation accuracy, with root mean square error (RMSE) reduced by up to an order of magnitude. The source code will be released to benefit the community1. Xinyi Chen 0002, Boyu Zhou, Shaojie Shen |
ICRA | 4 |
| 2024 | Less is More: Physical-Enhanced Radar-Inertial OdometryabstractRadar offers the advantage of providing additional physical properties related to observed objects. In this study, we design a physical-enhanced radar-inertial odometry system that capitalizes on the Doppler velocities and radar cross-section information. The filter for static radar points, correspondence estimation, and residual functions are all strengthened by integrating the physical properties. We conduct experiments on both public datasets and our self-collected data, with different mobile platforms and sensor types. Our quantitative results demonstrate that the proposed radar-inertial odometry system outperforms alternative methods using the physical-enhanced components. Our findings also reveal that using the physical properties results in fewer radar points for odometry estimation, but the performance is still guaranteed and even improved, thus aligning with the "less is more" principle. Qiucan Huang, Zhijian Qiao, Shaojie Shen, Huan Yin |
ICRA | 4 |
| 2024 | Parallel Optimization with Hard Safety Constraints for Cooperative Planning of Connected Autonomous VehiclesabstractThe development of connected autonomous vehicles (CAVs) facilitates the enhancement of traffic efficiency in complicated scenarios. Difficulties remain unsolved in developing an effective and efficient coordination strategy for CAVs. In this paper, we formulate the cooperative autonomous driving task of CAVs as an optimal control problem with safety conditions enforced as hard constraints, and propose a computationally-efficient parallel optimization framework to generate strategies for CAVs with the travel efficiency improved and the hard safety constraints satisfied. Specifically, all constraints involved are addressed appropriately with convex approximation, such that the convexity property of the reformulated optimization problem is exhibited. Then, a parallel optimization algorithm is presented to solve the reformulated optimization problem, with an embodied iterative nearest neighbor search strategy to determine the optimal passing sequence. It is noteworthy that the travel efficiency is enhanced and the computation burden is considerably alleviated with the proposed innovation development. We also examine the proposed method in CARLA simulator and perform thorough comparisons to demonstrate the effectiveness and efficiency of the proposed approach. Zhenmin Huang, Haichao Liu 0003, Shaojie Shen, Jun Ma 0008 |
ICRA | 3 |
| 2024 | A Two-step Nonlinear Factor Sparsification for Scalable Long-term SLAM BackendabstractThis paper proposes a new nonlinear factor sparsification paradigm for general feature-based long-term SLAM backend. Given a pose sparsification policy, we aim to scale the SLAM problem with space explored instead of time in a principled way so that the number of time-indexed poses can be limited. At the same time, their influence and the long-lived landmarks are appropriately maintained. To do this, we propose a new two-step sparsification pipeline. Given a pose node to remove, the first step is performed in the Markov blankets of affected landmarks. It transforms pose-landmark constraints into pose-pose constraints while preserving observability and minimizing information loss in the blanket. Moreover, since landmarks are conditionally independent, we can do this in parallel, disconnecting a pose from all the landmarks. The second step marginalizes the pose of interest with pure pose-wise constraints without affecting landmarks. Our method decouples the management of landmarks from pose-only measurements, making it general for any feature-based SLAM. We also give a practical example of how our backend works by concatenating it to a monocular VIO frontend. In simulation and realworld dataset, our sparsified backend is accurate and efficient. We open-source our backend, along with the VIO+Backend example, to contribute to the community’s betterment. Binqian Jiang, Shaojie Shen |
ICRA | 2 |
| 2024 | OmniNxt: A Fully Open-source and Compact Aerial Robot with Omnidirectional Visual PerceptionabstractAdopting omnidirectional Field of View (FoV) cameras in aerial robots vastly improves perception ability, significantly advancing aerial robotics’s capabilities in inspection, reconstruction, and rescue tasks. However, such sensors also elevate system complexity, e.g., hardware design, and corresponding algorithm, which limits researchers from utilizing aerial robots with omnidirectional FoV in their research. To bridge this gap, we propose OmniNxt, a fully open-source aerial robotics platform with omnidirectional perception. We design a high-performance flight controller Nxt-FC and a multi-fisheye camera set for OmniNxt. Meanwhile, the compatible software is carefully devised, which empowers OmniNxt to achieve accurate localization and real-time dense mapping with limited computation resource occupancy. We conducted extensive real-world experiments to validate the superior performance of OmniNxt in practical applications. All the hardware and software are open-access at3, and we provide docker images of each crucial module in the proposed system. Project page: https://hkust-aerial-robotics.github.io/OmniNxt. Peize Liu, Chen Feng 0006, Yang Xu 0083, Yan Ning, Hao Xu 0032, Shaojie Shen |
IROS | 6 |
| 2024 | Environment Transformer and Policy Optimization for Model-Based Offline Reinforcement LearningabstractInteracting with the actual environment to acquire data is often costly and time-consuming in robotic tasks. Model-based offline reinforcement learning (RL) provides a feasible solution. On the one hand, it eliminates the requirements of interaction with the actual environment. On the other hand, it learns the transition dynamics and reward function from the offline datasets and generates simulated rollouts to accelerate training. Previous model-based offline RL methods adopt probabilistic ensemble neural networks (NN) to model aleatoric uncertainty and epistemic uncertainty. However, this results in a great increase in training time and computing resource requirements. Furthermore, these methods are easily disturbed by the accumulative errors of the environment dynamics models when simulating long-term rollouts. To solve the above problems, we propose an uncertainty-aware sequence modeling architecture called Environment Transformer. It models the probability distribution of the environment dynamics and reward function to capture aleatoric uncertainty and treats epistemic uncertainty as a learnable noise parameter. Benefiting from the accurate modeling of the transition dynamics and reward function, Environment Transformer can be combined with arbitrary planning, dynamics programming, or policy optimization algorithms for offline RL. In this case, we perform Conservative Q-Learning (CQL) to learn a conservative Q-function. Through simulation experiments, we demonstrate that our method achieves or exceeds state-of-the-art performance in widely studied offline RL benchmarks. Moreover, we show that Environment Transformer's simulated rollout quality, sample efficiency, and long-term rollout simulation capability are superior to those of previous model-based offline RL methods. Pengqin Wang, Meixin Zhu, Shaojie Shen |
IROS | 3 |
| 2024 | SOAR: Simultaneous Exploration and Photographing with Heterogeneous UAVs for Fast Autonomous ReconstructionabstractUnmanned Aerial Vehicles (UAVs) have gained significant popularity in scene reconstruction. This paper presents SOAR, a LiDAR-Visual heterogeneous multi-UAV system specifically designed for fast autonomous reconstruction of complex environments. Our system comprises a LiDAR-equipped explorer with a large field-of-view (FoV), alongside photographers equipped with cameras. To ensure rapid acquisition of the scene’s surface geometry, we employ a surface frontier-based exploration strategy for the explorer. As the surface is progressively explored, we identify the uncovered areas and generate viewpoints incrementally. These viewpoints are then assigned to photographers through solving a Consistent Multiple Depot Multiple Traveling Salesman Problem (Consistent-MDMTSP), which optimizes scanning efficiency while ensuring task consistency. Finally, photographers utilize the assigned viewpoints to determine optimal coverage paths for acquiring images. We present extensive benchmarks in the realistic simulator, which validates the performance of SOAR compared with classical and state-of-the-art methods. For more details, please see our project page at sysu-star.github.io/SOAR. Chen Feng 0006, Zengzhi Li, Guiyong Zheng, Zhu Wang 0006, Jinni Zhou, Shaojie Shen, Boyu Zhou |
IROS | 8 |
| 2024 | A Survey on Global LiDAR Localization: Challenges, Advances and Open Problems
Huan Yin, Xuecheng Xu, Xieyuanli Chen, Rong Xiong, Shaojie Shen, Cyrill Stachniss, Yue Wang 0020 |
Int. J. Comput. Vis. | 6 |
| 2024 | Metric3D v2: A Versatile Monocular Geometric Foundation Model for Zero-Shot Metric Depth and Surface Normal EstimationabstractWe introduce Metric3D v2, a geometric foundation model designed for zero-shot metric depth and surface normal estimation from single images, critical for accurate 3D recovery. Depth and normal estimation, though complementary, present distinct challenges. State-of-the-art monocular depth methods achieve zero-shot generalization through affine-invariant depths, but fail to recover real-world metric scale. Conversely, current normal estimation techniques struggle with zero-shot performance due to insufficient labeled data. We propose targeted solutions for both metric depth and normal estimation. For metric depth, we present a canonical camera space transformation module that resolves metric ambiguity across various camera models and large-scale datasets, which can be easily integrated into existing monocular models. For surface normal estimation, we introduce a joint depth-normal optimization module that leverages diverse data from metric depth, allowing normal estimators to improve beyond traditional labels. Our model, trained on over 16 million images from thousands of camera models with varied annotations, excels in zero-shot generalization to new camera settings. As shown in Fig. 1, It ranks the 1st in multiple zero-shot and standard benchmarks for metric depth and surface normal prediction. Our method enables the accurate recovery of metric 3D structures on randomly collected internet images, paving the way for plausible single-image metrology. Our model also relieves the scale drift issues of monocular-SLAM (Fig. 3), leading to high-quality metric scale dense mapping. Such applications highlight the versatility of Metric3D v2 models as geometric foundation models. Mu Hu, Wei Yin 0006, Chi Zhang 0007, Zhipeng Cai 0003, Xiaoxiao Long, Hao Chen 0041, Gang Yu 0002, Chunhua Shen, Shaojie Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 10 |
| 2024 | An Efficient Spatial-Temporal Trajectory Planner for Autonomous Vehicles in Unstructured EnvironmentsabstractAs a fundamental component of autonomous driving systems, motion planning has garnered significant attention from both academia and industry. This paper focuses on efficient and spatial-temporal optimal trajectory optimization in unstructured environments using compact convex approximations of vehicle shapes. Conventional approaches typically model the task as an optimal control problem by discretizing the motion process in state configuration space. However, this often results in a tradeoff between optimality and efficiency since generating high-quality motion trajectories often requires high-precision discretization of the dynamic process, which imposes a substantial computational burden. To address this issue, we leverage the differential flatness property of car-like robots to simplify the trajectory representation and analytically formulate the spatial-temporal joint optimization problem with flat outputs in a compact manner, while ensuring the feasibility of nonholonomic dynamics. Moreover, we achieve efficient obstacle avoidance with a collision-free driving corridor for unmodelled obstacles and signed distance approximations for dynamic moving objects. We present comprehensive benchmarks with State-of-the-Art methods, demonstrating the significance of the proposed method in terms of efficiency and trajectory quality. Real-world experiments verify the practicality of our algorithm. We will release our codes for the research community. Zhichao Han 0002, Yuwei Wu 0005, Lu Zhang 0047, Liuao Pei, Long Xu 0002, Changjia Ma, Chao Xu 0001, Shaojie Shen, Fei Gao 0011 |
IEEE Trans. Intell. Transp. Syst. | 10 |
| 2024 | Adaptive Tracking and Perching for Quadrotor in Dynamic ScenariosabstractPerching on the moving platforms is a promising solution to enhance the endurance and operational range of quadrotors, which could benefit the efficiency of a variety of air ground cooperative tasks. To ensure robust perching, tracking with a steady relative state and reliable perception is a prerequisite. This paper presents an adaptive dynamic tracking and perching scheme for autonomous quadrotors to achieve tight integration with moving platforms. For reliable perception of dynamic targets, we introduce elastic visibility aware planning to actively avoid occlusion and target loss. Additionally, we propose a flexible terminal adjustment method that adapts the changes in flight duration and the couple d terminal states, ensuring full state synchronization with the time varying perching surface at various angles. A relaxation strategy is developed by optimizing the tangential relative speed to address the dynamics and safety violations brought by hard bo undary conditions. Moreover, we take SE(3) motion planning into account to ensure no collision until the contact moment. Furthermore, we propose an efficient spatiotemporal trajectory optimization framework considerin g full state dynamics The proposed method is extensively tested through benchmark comparisons and ablation studies. To facilitate the application of academic research to industry and to validate the efficiency under strictly limited computational resources, we deploy our system on a commercial drone (DJI MAVIC3) with a full size sport utility vehicle (SUV). We conduct extensive real world experiments, where the drone successfully tracks and perches at 30 km/h (8.3 m/ s) on the top of the SUV, and at 3.5∼m/s with 60° inclined into the trunk of the SUV. Yuman Gao, Jialin Ji, Qianhao Wang, Yi Lin 0010, Zhimeng Shang, Yanjun Cao, Shaojie Shen, Chao Xu 0001, Fei Gao 0011 |
IEEE Trans. Robotics | 8 |
| 2024 | Impact-Aware Planning and Control for Aerial Robots With Suspended PayloadsabstractA quadrotor with a cable-suspended payload imposes great challenges in impact-aware planning and control. This joint system has dual motion modes, depending on whether the cable is slack or not, and presents complicated dynamics. Therefore, generating feasible agile flight while preserving the retractable nature of the cable is still a challenging task. In this paper, we propose a novel impact-aware planning and control framework that resolves potential impacts caused by motion mode switching. Our method leverages the augmented Lagrangian method (ALM) to solve an optimization problem with nonlinear complementarity constraints (ONCC), which ensures trajectory feasibility with high accuracy while maintaining efficiency. We further propose a hybrid nonlinear model predictive control method to address the model mismatch issue in agile flight. Our methods have been comprehensively validated in both simulation and experiments, demonstrating superior performance compared to existing approaches. To the best of our knowledge, we are the first to successfully perform automatic multiple motion mode switching for aerial payload systems in real-world experiments. The video supplement is available athttps://sites.google.com/view/suspended-payload/. Haokun Wang 0007, Haojia Li, Boyu Zhou, Fei Gao 0011, Shaojie Shen |
IEEE Trans. Robotics | 5 |
| 2024 | $D^{2}$SLAM: Decentralized and Distributed Collaborative Visual-Inertial SLAM System for Aerial SwarmabstractCollaborative simultaneous localization and mapping (CSLAM) is essential for autonomous aerial swarms, laying the foundation for downstream algorithms, such as planning and control. To address existing CSLAM systems' limitations in relative localization accuracy, crucial for close-range UAV collaboration, this article introduces$D^{2}$SLAM—a novel decentralized and distributed CSLAM system.$D^{2}$SLAM innovatively manages near-field estimation for precise relative state estimation in proximity and far-field estimation for consistent global trajectories. Its adaptable front-end supports both stereo and omnidirectional cameras, catering to various operational needs and overcoming field-of-view challenges in aerial swarms. Experiments demonstrate$D^{2}$SLAM's effectiveness in accurate ego-motion estimation, relative localization, and global consistency. Enhanced by distributed optimization algorithms,$D^{2}$SLAM exhibits remarkable scalability and resilience to network delays, making it well suited for a wide range of real-world aerial swarm applications. We believe the adaptability and proven performance of$D^{2}$SLAM signify a notable advancement in autonomous aerial swarm technology. Hao Xu 0032, Peize Liu, Xinyi Chen 0002, Shaojie Shen |
IEEE Trans. Robotics | 4 |
| 2023 | BiFF: Bi-level Future Fusion with Polyline-based Coordinate for Interactive Trajectory PredictionabstractPredicting future trajectories of surrounding agents is essential for safety-critical autonomous driving. Most existing work focuses on predicting marginal trajectories for each agent independently. However, it has rarely been explored in predicting joint trajectories for interactive agents. In this work, we propose Bi-level Future Fusion (BiFF) to explicitly capture future interactions between interactive agents. Concretely, BiFF fuses the high-level future intentions followed by low-level future behaviors. Then the polyline-based coordinate is specifically designed for multi-agent prediction to ensure data efficiency, frame robustness, and prediction accuracy. Experiments show that BiFF achieves state-of-the-art performance on the interactive prediction benchmark of Waymo Open Motion Dataset. Yiyao Zhu, Di Luan, Shaojie Shen |
ICCV | 3 |
| 2023 | The Devil is in the Wrongly-classified Samples: Towards Unified Open-set Recognition
Jun Cen, Di Luan, Shiwei Zhang 0001, Yixuan Pei, Yingya Zhang, Deli Zhao, Shaojie Shen, Qifeng Chen 0001 |
ICLR | 7 |
| 2023 | PredRecon: A Prediction-boosted Planning Framework for Fast and High-quality Autonomous Aerial ReconstructionabstractAutonomous UAV path planning for 3D reconstruction has been actively studied in various applications for high-quality 3D models. However, most existing works have adopted explore-then-exploit, prior-based or exploration-based strategies, demonstrating inefficiency with repeated flight and low autonomy. In this paper, we propose PredRecon, a prediction-boosted planning framework that can autonomously generate paths for high 3D reconstruction quality. We obtain inspiration from humans can roughly infer the complete construction structure from partial observation. Hence, we devise a surface prediction module (SPM) to predict the coarse complete surfaces of the target from the current partial reconstruction. Then, the uncovered surfaces are produced by online volumetric mapping waiting for observation by UAV. Lastly, a hierarchical planner plans motions for 3D reconstruction, which sequentially finds efficient global coverage paths, plans local paths for maximizing the performance of Multi-View Stereo (MVS), and generates smooth trajectories for image-pose pairs acquisition. We conduct benchmarks in the realistic simulator, which validates the performance of PredRecon compared with the classical and state-of-the-art methods. The open-source code is released at https://github.com/HKUST-Aerial-Robotics/PredRecon. Chen Feng 0006, Haojia Li, Fei Gao 0011, Boyu Zhou, Shaojie Shen |
ICRA | 5 |
| 2023 | Contour Context: Abstract Structural Distribution for 3D LiDAR Loop Detection and Metric Pose EstimationabstractThis paper proposes Contour Context, a simple, effective, and efficient topological loop closure detection pipeline with accurate 3-DoF metric pose estimation, targeting the urban autonomous driving scenario. We interpret the Cartesian bird's eye view (BEV) image projected from 3D LiDAR points as layered distribution of structures. To recover elevation information from BEVs, we slice them at different heights, and connected pixels at each level form contours. Each contour is parameterized by abstract information, e.g., pixel count, center position, covariance, and mean height. The similarity of two BEVs is calculated in sequential discrete and continuous steps. The first step considers the geometric consensus of graph-like constellations formed by contours in particular localities. The second step models the majority of contours as a 2.5D Gaussian mixture model, which is used to calculate correlation and optimize relative transform in continuous space. A retrieval key is designed to accelerate the search of a database indexed by layered KD-trees. We validate the efficacy of our method by comparing it with recent works on public datasets. Binqian Jiang, Shaojie Shen |
ICRA | 2 |
| 2023 | Towards View-invariant and Accurate Loop Detection Based on Scene GraphabstractLoop detection plays a key role in visual Si-multaneous Localization and Mapping (SLAM) by correcting the accumulated pose drift. In indoor scenarios, the richly distributed semantic landmarks are view-point invariant and hold strong descriptive power in loop detection. The current semantic-aided loop detection embeds the topology between semantic instances to search a loop. However, current semantic-aided loop detection methods face challenges in dealing with ambiguous semantic instances and drastic viewpoint differences, which are not fully addressed in the literature. This paper introduces a novel loop detection method based on an incremen-tally created scene graph, targeting the visual SLAM at indoor scenes. It jointly considers the macro-view topology, micro-view topology, and occupancy of semantic instances to find correct correspondences. Experiments using handheld RGB-D sequence show our method is able to accurately detect loops in drastically changed viewpoints. It maintains a high precision in observing objects with similar topology and appearance. Our method also demonstrates that it is robust in changed indoor scenes. Chuhao Liu, Shaojie Shen |
ICRA | 2 |
| 2023 | Are All Point Clouds Suitable for Completion? Weakly Supervised Quality Evaluation Network for Point Cloud CompletionabstractIn the practical application of point cloud completion tasks, real data quality is usually much worse than the CAD datasets used for training. A small amount of noisy data will usually significantly impact the overall system's accuracy. In this paper, we propose a quality evaluation network to score the point clouds and help judge the quality of the point cloud before applying the completion model. We believe our scoring method can help researchers select more appropriate point clouds for subsequent completion and reconstruction and avoid manual parameter adjustment. Moreover, our evaluation model is fast and straightforward and can be directly inserted into any model's training or use process to facilitate the automatic selection and post-processing of point clouds. We propose a complete dataset construction and model evaluation method based on ShapeNet. We verify our network using detection and flow estimation tasks on KITTI, a real-world dataset for autonomous driving. The experimental results show that our model can effectively distinguish the quality of point clouds and help in practical tasks. Jieqi Shi, Peiliang Li 0001, Xiaozhi Chen, Shaojie Shen |
ICRA | 4 |
| 2023 | Rollvox: Real-Time and High-Quality LiDAR Colorization with Rolling Shutter CameraabstractIn this study, we propose a novel system for real-time coloring LiDAR point clouds with a low-cost RS camera. The main challenges are dealing with the motion distortion of the RS camera and the multi-sensor time synchronization. To tackle these challenges, we carefully design a hardware synchronizer to ensure the strict alignment of the LiDAR, inertial measurement unit, and RS camera. With accurate timestamps, we first use LiDAR-inertial odometry (LIO) for pose estimation, and the poses of image line exposure are calculated by forward propagation based on a constant velocity motion model. Then, we propose our method based on the RS constraint for colorizing the LiDAR point cloud. For comparison, we colorize the LiDAR point cloud with conventional rolling shutter image undistortion. In the real-world tests, The results show that our proposed method produces more accurate and efficient colorization of point clouds. Besides, considering the situation of readout time not being provided, we propose a method to calibrate the readout time by minimizing the reprojection error of LIO's inter-frame pose and image optical flows. We release our code and self-collected datasets on Github33https://github.com/sheng00125/Rollvox to benefit the community. Chunran Zheng, Huan Yin, Shaojie Shen |
IROS | 4 |
| 2023 | Online Monocular Lane Mapping Using Catmull-Rom SplineabstractIn this study, we introduce an online monocular lane mapping approach that solely relies on a single camera and odometry for generating spline-based maps. Our proposed technique models the lane association process as an assignment issue utilizing a bipartite graph, and assigns weights to the edges by incorporating Chamfer distance, pose uncertainty, and lateral sequence consistency. Furthermore, we meticulously design control point initialization, spline parameterization, and optimization to progressively create, expand, and refine splines. In contrast to prior research that assessed performance using self-constructed datasets, our experiments are conducted on the openly accessible OpenLane dataset. The experimental outcomes reveal that our suggested approach enhances lane association and odometry precision, as well as overall lane map quality. We have open-sourced out code11https://github.com/HKUST-Aerial-Robotics/MonoLaneMapping for this project. Zhijian Qiao, Zehuan Yu, Huan Yin, Shaojie Shen |
IROS | 4 |
| 2023 | Pyramid Semantic Graph-Based Global Point Cloud Registration with Low OverlapabstractGlobal point cloud registration is essential in many robotics tasks like loop closing and relocalization. Unfortunately, the registration often suffers from the low overlap between point clouds, a frequent occurrence in practical applications due to occlusion and viewpoint change. In this paper, we propose a graph-theoretic framework to address the problem of global point cloud registration with low overlap. To this end, we construct a consistency graph to facilitate robust data association and employ graduated non-convexity (GNC) for reliable pose estimation, following the state-of-the-art (SoTA) methods. Unlike previous approaches, we use semantic cues to scale down the dense point clouds, thus reducing the problem size. Moreover, we address the ambiguity arising from the consistency threshold by constructing a pyramid graph with multi-level consistency thresholds. Then we propose a cascaded gradient ascend method to solve the resulting densest clique problem and obtain multiple pose candidates for every consistency threshold. Finally, fast geometric verification is employed to select the optimal estimation from multiple pose candidates. Our experiments, conducted on a self-collected indoor dataset and the public KITTI dataset, demonstrate that our method achieves the highest success rate despite the low overlap of point clouds and low semantic quality. We have open-sourced our code1for this project. Zhijian Qiao, Zehuan Yu, Huan Yin, Shaojie Shen |
IROS | 4 |
| 2023 | Multi-Session, Localization-Oriented and Lightweight LiDAR Mapping Using Semantic Lines and PlanesabstractIn this paper, we present a centralized framework for multi-session LiDAR mapping in urban environments, by utilizing lightweight line and plane map representations instead of widely used point clouds. The proposed framework achieves consistent mapping in a coarse-to-fine manner. Global place recognition is achieved by associating lines and planes on the Grassmannian manifold, followed by an outlier rejection-aided pose graph optimization for map merging. Then a novel bundle adjustment is also designed to improve the local consistency of lines and planes. In the experimental section, both public and self-collected datasets are used to demonstrate efficiency and effectiveness. Extensive results validate that our LiDAR mapping framework could merge multi-session maps globally, optimize maps incrementally, and is applicable for lightweight robot localization. Zehuan Yu, Zhijian Qiao, Liuyang Qiu, Huan Yin, Shaojie Shen |
IROS | 5 |
| 2023 | Decentralized iLQR for Cooperative Trajectory Planning of Connected Autonomous Vehicles via Dual Consensus ADMMabstractCooperative trajectory planning of connected autonomous vehicles (CAVs) generally admits strong nonlinearity and non-convexity, rendering great difficulties in finding the optimal solution. Existing methods typically suffer from low computational efficiency and poor scalability, which hinder the appropriate applications in large-scale scenarios involving an increasing number of vehicles. To tackle this problem, we propose a novel decentralized iterative linear quadratic regulator (iLQR) algorithm by leveraging the dual consensus alternating direction method of multipliers (ADMM). First, the original non-convex optimization problem is reformulated into a series of convex optimization problems through iterative neighbourhood approximation. Then, the dual of each convex optimization problem is shown to have a consensus structure, which facilitates the use of consensus ADMM to solve for the dual solution in a fully decentralized and parallel architecture. Finally, the primal solution corresponding to the trajectory of each vehicle is recovered by solving a linear quadratic regulator (LQR) problem iteratively, and a novel trajectory update strategy is proposed to ensure the dynamic feasibility of vehicles. With the proposed development, the computation burden is significantly alleviated such that real-time performance is attainable. Two traffic scenarios are presented to validate the proposed algorithm, and thorough comparisons between our proposed method and baseline methods (including centralized iLQR, IPOPT, and SQP) are conducted to demonstrate the scalability of the proposed approach. Zhenmin Huang, Shaojie Shen, Jun Ma 0008 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Event-Based Motion Segmentation With Spatio-Temporal Graph CutsabstractIdentifying independently moving objects is an essential task for dynamic scene understanding. However, traditional cameras used in dynamic scenes may suffer from motion blur or exposure artifacts due to their sampling principle. By contrast, event-based cameras are novel bio-inspired sensors that offer advantages to overcome such limitations. They report pixel-wise intensity changes asynchronously, which enables them to acquire visual information at exactly the same rate as the scene dynamics. We develop a method to identify independently moving objects acquired with an event-based camera, that is, to solve the event-based motion segmentation problem. We cast the problem as an energy minimization one involving the fitting of multiple motion models. We jointly solve two sub-problems, namely event-cluster assignment (labeling) and motion model fitting, in an iterative manner by exploiting the structure of the input event data in the form of a spatio-temporal graph. Experiments on available datasets demonstrate the versatility of the method in scenes with different motion patterns and number of moving objects. The evaluation shows state-of-the-art results without having to predetermine the number of expected moving objects. We release the software and dataset under an open source license to foster research in the emerging topic of event-based motion segmentation. Yi Zhou 0010, Guillermo Gallego 0002, Xiuyuan Lu, Siqi Liu 0022, Shaojie Shen |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | RACER: Rapid Collaborative Exploration With a Decentralized Multi-UAV SystemabstractAlthough the use of multiple unmanned aerial vehicles (UAVs) has great potential for fast autonomous exploration, it has received far too little attention. In this article, we present a RApid Collaborative ExploRation (RACER) approach using a fleet of decentralized UAVs. To effectively dispatch the UAVs, a pairwise interaction based on an online hgrid space decomposition is used. It ensures that all UAVs simultaneously explore distinct regions, using only asynchronous and limited communication. Furthermore, we optimize the coverage paths of unknown space and balance the workloads partitioned to each UAV with a capacitated vehicle routing problem formulation. Given the task allocation, each UAV constantly updates the coverage path and incrementally extracts crucial information to support the exploration planning. A hierarchical planner finds exploration paths, refines local viewpoints, and generates minimum-time trajectories in sequence to explore the unknown space agilely and safely. The proposed approach is evaluated extensively, showing high exploration efficiency, scalability, and robustness to limited communication. Furthermore, for the first time, we achieve fully decentralized collaborative exploration with multiple UAVs in the real world. We will release our implementation as an open-source package. Boyu Zhou, Hao Xu 0032, Shaojie Shen |
IEEE Trans. Robotics | 3 |
| 2022 | Exploration with Global Consistency Using Real-Time Re-integration and Active Loop ClosureabstractDespite recent progress of robotic exploration, most methods assume that drift-free localization is available, which is problematic in reality and causes severe distortion of the reconstructed map. In this work, we present a systematic exploration mapping and planning framework that deals with drifted localization, allowing efficient and globally consistent reconstruction. A real-time re-integration-based mapping approach along with a frame pruning mechanism is proposed, which rectifies map distortion effectively when drifted localization is corrected upon detecting loop-closure. Besides, an exploration planning method considering historical viewpoints is presented to enable active loop closing, which promotes a higher opportunity to correct localization errors and further improves the mapping quality. We evaluate both the mapping and planning methods as well as the entire system comprehensively in simulation and real-world experiments, showing their effectiveness in practice. The implementation of the proposed method will be made open-source for the benefit of the robotics community. Boyu Zhou, Shaojie Shen |
ICRA | 4 |
| 2022 | Fast 3D Sparse Topological Skeleton Graph Generation for Mobile Robot Global PlanningabstractIn recent years, mobile robots are becoming ambitious and deployed in large-scale scenarios. Serving as a high-level understanding of environments, a sparse skeleton graph is beneficial for more efficient global planning. Currently, existing solutions for skeleton graph generation suffer from several major limitations, including poor adaptiveness to different map representations, dependency on robot inspection trajectories and high computational overhead. In this paper, we propose an efficient and flexible algorithm generating a trajectory-independent 3D sparse topological skeleton graph capturing the spatial structure of the free space. In our method, an efficient ray sampling and validating mechanism are adopted to find distinctive free space regions, which contributes to skeleton graph vertices, with traversability between adjacent vertices as edges. A cycle formation scheme is also utilized to maintain skeleton graph compactness. Benchmark comparison with state-of-the-art works demonstrates that our approach generates sparse graphs in a substantially shorter time, giving high-quality global planning paths. Experiments conducted in real-world maps further validate the capability of our method in real-world scenarios. Our method will be made open source to benefit the community. Xinyi Chen 0002, Boyu Zhou, Jiarong Lin, Fu Zhang 0002, Shaojie Shen |
IROS | 6 |
| 2022 | A LiDAR-inertial Odometry with Principled Uncertainty ModelingabstractThis paper proposes a LiDAR-inertial odometry that properly solves the uncertainty estimation problem, guided by the rules of designing a consistent estimator. Our system is built upon an iterated extended Kalman filter, with multiple states in an optimization window. To survive environments without distinctive geometric structures, we do not track features over time. We only extract planar primitives from the local map and use a direct point-to-plane distance metric as the measurement model. The realistic noise parameters are estimated online by modeling point distributions. We use nullspace projection to remove dependency on the feature planes, which is equivalent to transforming the pose-map measurement into relative pose constraints. To avoid reintegrating all the laser points in the local window after every state correction, we use the Schmidt Kalman update to consider the probabilistic effects of past poses while their values are left unaltered. A collection of octrees with an adaptive resolution is designed to manage measurement points and the map efficiently. The consistency and robustness of our system are verified in both simulation and real-world experiments. Binqian Jiang, Shaojie Shen |
IROS | 2 |
| 2022 | Trajectory Prediction with Graph-based Dual-scale Context FusionabstractMotion prediction for traffic participants is essential for a safe and robust automated driving system, especially in cluttered urban environments. However, it is highly challenging due to the complex road topology as well as the uncertain intentions of the other agents. In this paper, we present a graph-based trajectory prediction network named the Dual Scale Predictor (DSP), which encodes both the static and dynamical driving context in a hierarchical manner. Different from methods based on a rasterized map or sparse lane graph, we consider the driving context as a graph with two layers, focusing on both geometrical and topological features. Graph neural networks (GNNs) are applied to extract features with different levels of granularity, and features are subsequently aggregated with attention-based inter-layer networks, realizing better local-global feature fusion. Following the recent goal-driven trajectory prediction pipeline, goal candidates with high likelihood for the target agent are extracted, and predicted trajectories are generated conditioned on these goals. Thanks to the proposed dual-scale context fusion network, our DSP is able to generate accurate and human-like multi-modal trajectories. We evaluate the proposed method on the large-scale Argoverse motion forecasting benchmark, and it achieves promising results, outperforming the recent state-of-the-art methods. We release the code on our project website.11https://github.com/HKUST-Aerial-Robotics/DSP Lu Zhang 0047, Peiliang Li 0001, Jing Chen 0016, Shaojie Shen |
IROS | 4 |
| 2022 | GVINS: Tightly Coupled GNSS-Visual-Inertial Fusion for Smooth and Consistent State EstimationabstractVisual–inertial odometry (VIO) is known to suffer from drifting, especially over long-term runs. In this article, we present GVINS, a nonlinear optimization-based system that tightly fuses global navigation satellite system (GNSS) raw measurements with visual and inertial information for real-time and drift-free stateestimation. Our system aims to provide accurate global six-degree-of-freedom estimation under complex indoor–outdoor environments, where GNSS signals may be intermittent or even inaccessible. To establish the connection between global measurements and local states, a coarse-to-fine initialization procedure is proposed to efficiently calibrate the transformation online and initialize GNSS states from only a short window of measurements. The GNSS code pseudorange and Doppler shift measurements, along with visual and inertial information, are then modeled and used to constrain the system states in a factor graph framework. For complex and GNSS-unfriendly areas, the degenerate cases are discussed and carefully handled to ensure robustness. Thanks to the tightly coupled multisensor approach and system design, our system fully exploits the merits of three types of sensors and is able to seamlessly cope with the transition between indoor and outdoor environments, where satellites are lost and reacquired. We extensively evaluate the proposed system by both simulation and real-world experiments, and the results demonstrate that our system substantially suppresses the drift of the VIO and preserves the local accuracy in spite of noisy GNSS measurements. The versatility and robustness of the system are verified on large-scale data collected in challenging environments. In addition, experiments show that our system can still benefit from the presence of only one satellite, whereas at least four satellites are required for its conventional GNSS counterparts. Shaozu Cao, Xiuyuan Lu, Shaojie Shen |
IEEE Trans. Robotics | 3 |
| 2022 | EPSILON: An Efficient Planning System for Automated Vehicles in Highly Interactive EnvironmentsabstractIn this article, we present an efficient planning system for automated vehicles in highly interactive environments (EPSILON). EPSILON is an efficient interaction-aware planning system for automated driving, and is extensively validated in both simulation and real-world dense city traffic. It follows a hierarchical structure with an interactive behavior planning layer and an optimization-based motion planning layer. The behavior planning is formulated from a partially observable Markov decision process (POMDP), but is much more efficient than naively applying a POMDP to the decision-making problem. The key to efficiency is guided branching in both the action space and observation space, which decomposes the original problem into a limited number of closed-loop policy evaluations. Moreover, we introduce a new driver model with a safety mechanism to overcome the risk induced by the potential imperfectness of prior knowledge. For motion planning, we employ a spatio-temporal semantic corridor (SSC) to model the constraints posed by complex driving environments in a unified way. Based on the SSC, a safe and smooth trajectory is optimized, complying with the decision provided by the behavior planner. We validate our planning system in both simulations and real-world dense traffic, and the experimental results show that our EPSILON achieves human-like driving behaviors in highly interactive traffic flow smoothly and safely without being overconservative compared to the existing planning methods. Wenchao Ding 0001, Lu Zhang 0047, Jing Chen 0016, Shaojie Shen |
IEEE Trans. Robotics | 4 |
| 2022 | Omni-Swarm: A Decentralized Omnidirectional Visual-Inertial-UWB State Estimation System for Aerial SwarmsabstractDecentralized state estimation is one of the most fundamental components of autonomous aerial swarm systems in GPS-denied areas; yet, it remains a highly challenging research topic. Omni-swarm, a decentralized omnidirectional visual–inertial–ultrawideband (UWB) state estimation system for aerial swarms, is proposed in this article to address this research niche. To solve the issues of observability, complicated initialization, insufficient accuracy, and lack of global consistency, we introduce an omnidirectional perception front end in Omni-swarm. It consists of stereo wide-field-of-view cameras and UWB sensors, visual–inertial odometry, multidrone map-based localization, and visual drone tracking algorithms. The measurements from the front end are fused with graph-based optimization in the back end. The proposed method achieves centimeter-level relative state estimation accuracy while guaranteeing global consistency in the aerial swarm, as evidenced by the experimental results. Moreover, supported by Omni-swarm, interdrone collision avoidance can be accomplished without any external devices, demonstrating the potential of Omni-swarm as the foundation of autonomous aerial swarms. Hao Xu 0032, Boyu Zhou, Xinjie Yao, Guotao Meng, Shaojie Shen |
IEEE Trans. Robotics | 7 |
| 2021 | Estimation and Adaption of Indoor Ego Airflow Disturbance with Application to Quadrotor Trajectory PlanningabstractIt is ubiquitously accepted that during the autonomous navigation of the quadrotors, one of the most widely adopted unmanned aerial vehicles (UAVs), safety always has the highest priority. However, it is observed that the ego airflow disturbance can be a significant adverse factor during flights, causing potential safety issues, especially in narrow and confined indoor environments. Therefore, we propose a novel method to estimate and adapt indoor ego airflow disturbance of quadrotors, meanwhile applying it to trajectory planning. Firstly, the hover experiments for different quadrotors are conducted against the proximity effects. Then with the collected acceleration variance, the disturbances are modeled for the quadrotors according to the proposed formulation. The disturbance model is also verified under hover conditions in different reconstructed complex environments. Furthermore, the approximation of Hamilton-Jacobi reachability analysis is performed according to the estimated disturbances to facilitate the safe trajectory planning, which consists of kinodynamic path search as well as B-spline trajectory optimization. The whole planning framework is validated on multiple quadrotor platforms in different indoor environments. Boyu Zhou, Chuhao Liu, Shaojie Shen |
ICRA | 4 |
| 2021 | Event-based Motion Segmentation by Cascaded Two-Level Multi-Model FittingabstractAmong prerequisites for a synthetic agent to inter-act with dynamic scenes, the ability to identify independently moving objects is specifically important. From an application perspective, nevertheless, standard cameras may deteriorate remarkably under aggressive motion and challenging illumination conditions. In contrast, event-based cameras, as a category of novel biologically inspired sensors, deliver advantages to deal with these challenges. Its rapid response and asynchronous nature enables it to capture visual stimuli at exactly the same rate of the scene dynamics. In this paper, we present a cascaded two-level multi-model fitting method for identifying independently moving objects (i.e., the motion segmentation problem) with a monocular event camera. The first level leverages tracking of event features and solves the feature clustering problem under a progressive multi-model fitting scheme. Initialized with the resulting motion model instances, the second level further addresses the event clustering problem using a spatio-temporal graph-cut method. This combination leads to efficient and accurate event-wise motion segmentation that cannot be achieved by any of them alone. Experiments demonstrate the effectiveness and versatility of our method in real-world scenes with different motion patterns and an unknown number of independently moving objects. Xiuyuan Lu, Yi Zhou 0010, Shaojie Shen |
IROS | 3 |
| 2021 | Real-Time Temporal and Rotational Calibration of Heterogeneous Sensors Using Motion Correlation AnalysisabstractAccurate and robust calibration is crucial to a multisensor fusion-based system. The calibration of heterogeneous sensors is particularly challenging because of the huge difference of the captured sensor data. On the other hand, many calibration approaches ignore temporal calibration that is in fact as important as spatial calibration. In this article, we focus on the temporal calibration of heterogeneous sensors, and the corresponding extrinsic rotation is also derived. Most existing methods are specialized for a certain sensor combination, such as an inertial measurement unit (IMU) camera or a camera-Lidar system. However, heterogeneous multisensor fusion is a tendency in the robotics area, so a unified calibration method is desired. To this end, we leverage the 3-D rotational motion feature for calibration, and auxiliary calibration boards are not needed since multiple odometry methods are available to capture 3-D sensor motion. Using a high-frequency IMU as the calibration reference, an IMU-centric scheme is designed to achieve a unified framework that adapts to various target sensors that can independently estimate 3-D rotational motion. By combining independent IMU-centric calibration pairs, an arbitrary pair of sensors can also be calibrated using the same reference IMU. Due to a novel 3-D motion correlation quantification and analysis mechanism, the temporal offset can be first estimated in real time. Given temporally aligned sensor motion, the extrinsic rotation can be derived in closed-form in the same 3-D motion correlation mechanism. Experimental results of certain sensor combinations show the accuracy and robustness of the proposed method through comparison with state-of-the-art calibration approaches, and the calibration result of a heterogeneous multisensor set demonstrates the scalability and versatility of our method. Kejie Qiu, Tong Qin 0001, Jie Pan 0005, Siqi Liu 0022, Shaojie Shen |
IEEE Trans. Robotics | 5 |
| 2021 | Event-Based Stereo Visual OdometryabstractEvent-based cameras are bioinspired vision sensors whose pixels work independently from each other and respond asynchronously to brightness changes, with microsecond resolution. Their advantages make it possible to tackle challenging scenarios in robotics, such as high-speed and high dynamic range scenes. We present a solution to the problem of visual odometry from the data acquired by a stereo event-based camera rig. Our system follows a parallel tracking-and-mapping approach, where novel solutions to each subproblem (three-dimensional (3-D) reconstruction and camera pose estimation) are developed with two objectives in mind: being principled and efficient, for real-time operation with commodity hardware. To this end, we seek to maximize the spatio-temporal consistency of stereo event-based data while using a simple and efficient representation. Specifically, the mapping module builds a semidense 3-D map of the scene by fusing depth estimates from multiple viewpoints (obtained by spatio-temporal consistency) in a probabilistic fashion. The tracking module recovers the pose of the stereo rig by solving a registration problem that naturally arises due to the chosen map and event data representation. Experiments on publicly available datasets and on our own recordings demonstrate the versatility of the proposed method in natural scenes with general 6-DoF motion. The system successfully leverages the advantages of event-based cameras to perform visual odometry in challenging illumination conditions, such as low-light and high dynamic range, while running in real-time on a standard CPU. We release the software and dataset under an open source license to foster research in the emerging topic of event-based simultaneous localization and mapping. Yi Zhou 0010, Guillermo Gallego 0002, Shaojie Shen |
IEEE Trans. Robotics | 3 |
| 2021 | RAPTOR: Robust and Perception-Aware Trajectory Replanning for Quadrotor Fast FlightabstractRecent advances in trajectory replanning have enabled quadrotor to navigate autonomously in unknown environments. However, high-speed navigation still remains a significant challenge. Given very limited time, existing methods have no strong guarantee on the feasibility or quality of the solutions. Moreover, most methods do not consider environment perception, which is the key bottleneck to fast flight. In this article, we present RAPTOR, a robust and perception-aware replanning framework to support fast and safe flight, which addresses these issues systematically. A path-guided optimization approach that incorporates multiple topological paths is devised, to ensure finding feasible and high-quality trajectories in very limited time. We also introduce two perception-aware planning approaches to actively observe and avoid unknown obstacles. A risk-aware trajectory refinement ensures that unknown obstacles which may endanger the quadrotor can be observed earlier and avoid in time. The motion of yaw angle is planned to actively explore the surrounding space that is relevant for safe navigation. The proposed methods are tested extensively through benchmark comparisons and challenging indoor and outdoor aggressive flights. We release our implementation as an open-source package1for the community. Boyu Zhou, Jie Pan 0004, Fei Gao 0011, Shaojie Shen |
IEEE Trans. Robotics | 4 |
| 2020 | Joint Spatial-Temporal Optimization for Stereo 3D Object TrackingabstractDirectly learning multiple 3D objects motion from sequential images is difficult, while the geometric bundle adjustment lacks the ability to localize the invisible object centroid. To benefit from both the powerful object understanding skill from deep neural network meanwhile tackle precise geometry modeling for consistent trajectory estimation, we propose a joint spatial-temporal optimization-based stereo 3D object tracking method. From the network, we detect corresponding 2D bounding boxes on adjacent images and regress an initial 3D bounding box. Dense object cues (local depth and local coordinates) that associating to the object centroid are then predicted using a region-based network. Considering both the instant localization accuracy and motion consistency, our optimization models the relations between the object centroid and observed cues into a joint spatial-temporal error function. All historic cues will be summarized to contribute to the current estimation by a per-frame marginalization strategy without repeated computation. Quantitative evaluation on the KITTI tracking dataset shows our approach outperforms previous image-based 3D tracking methods by significant margins. We also report extensive results on multiple categories and larger datasets (KITTI raw and Argoverse Tracking) for future benchmarking. Peiliang Li 0001, Jieqi Shi, Shaojie Shen |
CVPR | 3 |
| 2020 | PiP: Planning-Informed Trajectory Prediction for Autonomous Driving
Haoran Song, Wenchao Ding 0001, Shaojie Shen, Michael Yu Wang, Qifeng Chen 0001 |
ECCV (21) | 4 |
| 2020 | FP-Stereo: Hardware-Efficient Stereo Vision for Embedded ApplicationsabstractFast and accurate depth estimation, or stereo matching, is essential in embedded stereo vision systems, requiring substantial design effort to achieve an appropriate balance among accuracy, speed and hardware cost. To reduce the design effort and achieve the right balance, we propose FP-Stereo for building high-performance stereo matching pipelines on FPGAs automatically. FP-Stereo consists of an open-source hardware-efficient library, allowing designers to obtain the desired implementation instantly. Diverse methods are supported in our library for each stage of the stereo matching pipeline and a series of techniques are developed to exploit the parallelism and reduce the resource overhead. To improve the usability, FP-Stereo can generate synthesizable C code of the FPGA accelerator with our optimized HLS templates automatically. To guide users for the right design choice meeting specific application requirements, detailed comparisons are performed on various configurations of our library to investigate the accuracy/speed/cost trade-off. Experimental results also show that FP-Stereo outperforms the state-of-the-art FPGA design from all aspects, including 6.08% lower error, 2x faster speed, 30% less resource usage and 40% less energy consumption. Compared to GPU designs, FP-Stereo achieves the same accuracy at a competitive speed while consuming much less energy. Jieru Zhao, Tingyuan Liang, Liang Feng 0001, Wenchao Ding 0001, Sharad Sinha, Wei Zhang 0012, Shaojie Shen |
FPL | 7 |
| 2020 | Geometric Pretraining for Monocular Depth EstimationabstractImageNet-pretrained networks have been widely used in transfer learning for monocular depth estimation. These pretrained networks are trained with classification losses for which only semantic information is exploited while spatial information is ignored. However, both semantic and spatial information is important for per-pixel depth estimation. In this paper, we design a novel self-supervised geometric pretraining task that is tailored for monocular depth estimation using uncalibrated videos. The designed task decouples the structure information from input videos by a simple yet effective conditional autoencoder-decoder structure. Using almost unlimited videos from the internet, networks are pretrained to capture a variety of structures of the scene and can be easily transferred to depth estimation tasks using calibrated images. Extensive experiments are used to demonstrate that the proposed geometric-pretrained networks perform better than ImageNet-pretrained networks in terms of accuracy, few-shot learning and generalization ability. Using existing learning methods, geometric-transferred networks achieve new state-of-the-art results by a large margin. The pretrained networks will be open source soon1. Hengkai Guo, Linfu Wen, Shaojie Shen |
ICRA | 5 |
| 2020 | FlowNorm: A Learning-based Method for Increasing Convergence Range of Direct AlignmentabstractMany approaches have been proposed to estimate camera poses by directly minimizing photometric error. However, due to the non-convex property of direct alignment, proper initialization is still required for these methods. Many robust norms (e.g. Huber norm) have been proposed to deal with the outlier terms caused by incorrect initializations. These robust norms are solely defined on the magnitude of each error term. In this paper, we propose a novel robust norm, named FlowNorm, that exploits the information from both the local error term and the global image registration information. While the local information is defined on patch alignments, the global information is estimated using a learning-based network. Using both the local and global information, we achieve a large convergence range in which images can be aligned given large view angle changes or small overlaps. We further demonstrate the usability of the proposed robust norm by integrating it into the direct methods DSO and BA-Net, and generate more robust and accurate results in real-time. Ke Wang 0058, Shaojie Shen |
ICRA | 3 |
| 2020 | Decentralized Visual-Inertial-UWB Fusion for Relative State Estimation of Aerial SwarmabstractThe collaboration of unmanned aerial vehicles (UAVs) has become a popular research topic for its practicability in multiple scenarios. The collaboration of multiple UAVs, which is also known as aerial swarm is a highly complex system, which still lacks a state-of-art decentralized relative state estimation method. In this paper, we present a novel fully decentralized visual-inertial-UWB fusion framework for relative state estimation and demonstrate the practicability by performing extensive aerial swarm flight experiments. The comparison result with ground truth data from the motion capture system shows the centimeter-level precision which outperforms all the Ultra-WideBand (UWB) and even vision based method. The system is not limited by the field of view (FoV) of the camera or Global Positioning System (GPS), meanwhile on account of its estimation consistency, we believe that the proposed relative state estimation framework has the potential to be prevalently adopted by aerial swarm applications in different scenarios in multiple scales. Hao Xu 0032, Kejie Qiu, Shaojie Shen |
ICRA | 5 |
| 2020 | Efficient Uncertainty-aware Decision-making for Automated Driving Using Guided BranchingabstractDecision-making in dense traffic scenarios is challenging for automated vehicles (AVs) due to potentially stochastic behaviors of other traffic participants and perception uncertainties (e.g., tracking noise and prediction errors, etc.). Although the partially observable Markov decision process (POMDP) provides a systematic way to incorporate these uncertainties, it quickly becomes computationally intractable when scaled to the real-world large-size problem. In this paper, we present an efficient uncertainty-aware decision-making (EUDM) framework, which generates long-term lateral and longitudinal behaviors in complex driving environments in real-time. The computation complexity is controlled to an appropriate level by two novel techniques, namely, the domain-specific closed-loop policy tree (DCP-Tree) structure and conditional focused branching (CFB) mechanism. The key idea is utilizing domain-specific expert knowledge to guide the branching in both action and intention space. The proposed framework is validated using both onboard sensing data captured by a real vehicle and an interactive multi-agent simulation platform. We also release the code of our framework to accommodate benchmarking. Lu Zhang 0047, Wenchao Ding 0001, Jing Chen 0016, Shaojie Shen |
ICRA | 4 |
| 2020 | Robust Real-time UAV Replanning Using Guided Gradient-based Optimization and Topological PathsabstractGradient-based trajectory optimization (GTO) has gained wide popularity for quadrotor trajectory replanning. However, it suffers from local minima, which is not only fatal to safety but also unfavorable for smooth navigation. In this paper, we propose a replanning method based on GTO addressing this issue systematically. A path-guided optimization (PGO) approach is devised to tackle infeasible local minima, which improves the replanning success rate significantly. A topological path searching algorithm is developed to capture a collection of distinct useful paths in 3-D environments, each of which then guides an independent trajectory optimization. It activates a more comprehensive exploration of the solution space and output superior replanned trajectories. Benchmark evaluation shows that our method outplays state-of-the-art methods regarding replanning success rate and optimality. Challenging experiments of aggressive autonomous flight are presented to demonstrate the robustness of our method. We will release our implementation as an open-source package1. Boyu Zhou, Fei Gao 0011, Jie Pan 0004, Shaojie Shen |
ICRA | 4 |
| 2020 | An Augmented Reality Interaction Interface for Autonomous DroneabstractHuman drone interaction in autonomous navigation incorporates spatial interaction tasks, including reconstructed 3D map from the drone and human desired target position. Augmented Reality (AR) devices can be powerful interactive tools for handling these spatial interactions. In this work, we build an AR interface that displays the reconstructed 3D map from the drone on physical surfaces in front of the operator. Spatial target positions can be further set on the 3D map by intuitive head gaze and hand gesture. The AR interface is deployed to interact with an autonomous drone to explore an unknown environment. A user study is further conducted to evaluate the overall interaction performance. Chuhao Liu, Shaojie Shen |
IROS | 2 |
| 2020 | Teach-Repeat-Replan: A Complete and Robust System for Aggressive Flight in Complex EnvironmentsabstractIn this article, we propose a complete and robust system for the aggressive flight of autonomous quadrotors. The proposed system is built upon on the classical teach-and-repeat framework, which is widely adopted in infrastructure inspection, aerial transportation, and search-and-rescue. For these applications, a human's intention is essential for deciding the topological structure of the flight trajectory of the drone. However, poor teaching trajectories and changing environments prevent a simple teach-and-repeat system from being applied flexibly and robustly. In this article, instead of commanding the drone to precisely follow a teaching trajectory, we propose a method to automatically convert a human-piloted trajectory, which can be arbitrarily jerky, to a topologically equivalent one. The generated trajectory is guaranteed to be smooth, safe, and dynamically feasible, with a human preferable aggressiveness. Also, to avoid unmapped or moving obstacles during flights, a fast local perception method and a sliding-windowed replanning method are integrated into our system, to generate safe and dynamically feasible local trajectories onboard. We name our system as teach-repeat-replan. It can capture users' intention of a flight mission, convert an arbitrarily jerky teaching path to a smooth repeating trajectory, and generate safe local replans to avoid unexpected collisions. The proposed planning system is integrated into a complete autonomous quadrotor with global and local perception and localization submodules. Our system is validated by performing aggressive flights in challenging indoor/outdoor environments. We release all components in our quadrotor system as open-source ros packages. Fei Gao 0011, Boyu Zhou, Xin Zhou 0015, Jie Pan 0004, Shaojie Shen |
IEEE Trans. Robotics | 6 |
| 2019 | Stereo R-CNN Based 3D Object Detection for Autonomous DrivingabstractWe propose a 3D object detection method for autonomous driving by fully exploiting the sparse and dense, semantic and geometry information in stereo imagery. Our method, called Stereo R-CNN, extends Faster R-CNN for stereo inputs to simultaneously detect and associate object in left and right images. We add extra branches after stereo Region Proposal Network (RPN) to predict sparse keypoints, viewpoints, and object dimensions, which are combined with 2D left-right boxes to calculate a coarse 3D object bounding box. We then recover the accurate 3D bounding box by a region-based photometric alignment using left and right RoIs. Our method does not require depth input and 3D position supervision, however, outperforms all existing fully supervised image-based methods. Experiments on the challenging KITTI dataset show that our method outperforms the state-of-the-art stereo-based method by around 30% AP on both 3D detection and 3D localization tasks. Code will be made publicly available. Peiliang Li 0001, Xiaozhi Chen, Shaojie Shen |
CVPR | 3 |
| 2019 | Predicting Vehicle Behaviors Over An Extended Horizon Using Behavior Interaction NetworkabstractAnticipating possible behaviors of traffic participants is an essential capability of autonomous vehicles. Many behavior detection and maneuver recognition methods only have a very limited prediction horizon that leaves inadequate time and space for planning. To avoid unsatisfactory reactive decisions, it is essential to count long-term future rewards in planning, which requires extending the prediction horizon. In this paper, we uncover that clues to vehicle behaviors over an extended horizon can be found in vehicle interaction, which makes it possible to anticipate the likelihood of a certain behavior, even in the absence of any clear maneuver pattern. We adopt a recurrent neural network (RNN) for observation encoding, and based on that, we propose a novel vehicle behavior interaction network (VBIN) to capture the vehicle interaction from the hidden states and connection feature of each interaction pair. The output of our method is a probabilistic likelihood of multiple behavior classes, which matches the multimodal and uncertain nature of the distant future. A systematic comparison of our method against two state-of-the-art methods and another two baseline methods on a publicly available real highway dataset is provided, showing that our method has superior accuracy and advanced capability for interaction modeling. Wenchao Ding 0001, Jing Chen 0016, Shaojie Shen |
ICRA | 3 |
| 2019 | Online Vehicle Trajectory Prediction using Policy Anticipation Network and optimization-based Context ReasoningabstractIn this paper, we present an online two-level vehicle trajectory prediction framework for urban autonomous driving where there are complex contextual factors, such as lane geometries, road constructions, traffic regulations and moving agents. Our method combines high-level policy anticipation with low-level context reasoning. We leverage a long short-term memory (LSTM) network to anticipate the vehicle's driving policy (e.g., forward, yield, turn left, turn right, etc.) using its sequential history observations. The policy is then used to guide a low-level optimization-based context reasoning process. We show that it is essential to incorporate the prior policy anticipation due to the multimodal nature of the future trajectory. Moreover, contrary to existing regression-based trajectory prediction methods, our optimization-based reasoning process can cope with complex contextual factors. The final output of the two-level reasoning process is a continuous trajectory that automatically adapts to different traffic configurations and accurately predicts future vehicle motions. The performance of the proposed framework is analyzed and validated in an emerging autonomous driving simulation platform (CARLA). Wenchao Ding 0001, Shaojie Shen |
ICRA | 2 |
| 2019 | Real-time Scalable Dense Surfel MappingabstractIn this paper, we propose a novel dense surfel mapping system that scales well in different environments with only CPU computation. Using a sparse SLAM system to estimate camera poses, the proposed mapping system can fuse intensity images and depth images into a globally consistent model. The system is carefully designed so that it can build from room-scale environments to urban-scale environments using depth images from RGB-D cameras, stereo cameras or even a monocular camera. First, superpixels extracted from both intensity and depth images are used to model surfels in the system. superpixel-based surfels make our method both runtime efficient and memory efficient. Second, surfels are further organized according to the pose graph of the SLAM system to achieve O(1) fusion time regardless of the scale of reconstructed models. Third, a fast map deformation using the optimized pose graph enables the map to achieve global consistency in real-time. The proposed surfel mapping system is compared with other state-of-the-art methods on synthetic datasets. The performances of urban-scale and room-scale reconstruction are demonstrated using the KITTI dataset [1] and autonomous aggressive flights, respectively. The code is available for the benefit of the community. Fei Gao 0011, Shaojie Shen |
ICRA | 3 |
| 2019 | FIESTA: Fast Incremental Euclidean Distance Fields for Online Motion Planning of Aerial RobotsabstractEuclidean Signed Distance Field (ESDF) is useful for online motion planning of aerial robots since it can easily query the distance and gradient information against obstacles. Fast incrementally built ESDF map is the bottleneck for conducting real-time motion planning. In this paper, we investigate this problem and propose a mapping system called FIESTA to build global ESDF map incrementally. By introducing two independent updating queues for inserting and deleting obstacles separately, and using Indexing Data Structures and Doubly Linked Lists for map maintenance, our algorithm updates as few as possible nodes using a BFS framework. Our ESDF map has high computational performance and produces near-optimal results. We show our method outperforms other up-to-date methods in term of performance and accuracy by both theory and experiments. We integrate FIESTA into a completed quadrotor system and validate it by both simulation and onboard experiments. We release our method as open-source software for the community. Luxin Han, Fei Gao 0011, Boyu Zhou, Shaojie Shen |
IROS | 4 |
| 2019 | Flying through a narrow gap using neural network: an end-to-end planning and control approachabstractIn this paper, we investigate the problem of enabling a drone to fly through a tilted narrow gap, without a traditional planning and control pipeline. To this end, we propose an end-to-end policy network, which imitates from the traditional pipeline and is fine-tuned using reinforcement learning. Unlike previous works which plan dynamical feasible trajectories using motion primitives and track the generated trajectory by a geometric controller, our proposed method is an end-to-end approach which takes the flight scenario as input and directly outputs thrust-attitude control commands for the quadrotor. Key contributions of our paper are: 1) presenting an imitate-reinforce training framework. 2) flying through a narrow gap using an end-to-end policy network, showing that learning based method can also address the highly dynamic control problem as the traditional pipeline does (see attached video1). 3) propose a robust imitation of an optimal trajectory generator using multilayer perceptrons. 4) show how reinforcement learning can improve the performance of imitation learning, and the potential to achieve higher performance over the model-based method. Jiarong Lin, Fei Gao 0011, Shaojie Shen, Fu Zhang 0002 |
IROS | 4 |
| 2019 | A GPS-aided Omnidirectional Visual-Inertial State Estimator in Ubiquitous EnvironmentsabstractThe visual-inertial navigation system (VINS) has been a practical approach for state estimation in recent years. In this paper, we propose a general GPS-aided omnidirectional visual-inertial state estimator capable of operating in ubiquitous environments and platforms. Our system consists of two parts: 1) the pre-processing of omnidirectional cameras, IMU, and GPS measurements, and 2) the sliding window based nonlinear optimization for accurate state estimation. We test our system in different conditions including an indoor office, campus roads, and challenging open water surface. Experiment results demonstrate the high accuracy of our approach than state-of-the-art VINSs in all scenarios. The proposed odometry achieves drift ratio less than 0.5% in 1200 m length outdoors campus road in overexposure conditions and 0.65% in open water surface, without a loop closure, compared with a centimeter accuracy GPS reference. Yang Yu 0028, Wenliang Gao, Shaojie Shen, Ming Liu 0001 |
IROS | 4 |
| 2019 | Temporal Scheduling and Optimization for Multi-MAV Planning
William Wu, Fei Gao 0011, Boyu Zhou, Shaojie Shen |
ISRR | 5 |
| 2019 | An Efficient B-Spline-Based Kinodynamic Replanning Framework for QuadrotorsabstractTrajectory replanning for quadrotors is essential to enable fully autonomous flight in unknown environments. Hierarchical motion planning frameworks, which combine path planning with path parameterization, are popular due to their time efficiency. However, the path planning cannot properly deal with nonstatic initial states of the quadrotor, which may result in nonsmooth or even dynamically infeasible trajectories. In this article, we present an efficient kinodynamic replanning framework by exploiting the advantageous properties of the B-spline, which facilitates dealing with the nonstatic state and guarantees safety and dynamical feasibility. Our framework starts with an efficient B-spline-based kinodynamic (EBK) search algorithm, which finds a feasible trajectory with minimum control effort and time. To compensate for the discretization induced by the EBK search, an elastic optimization approach is proposed to refine the control point placement to the optimal location. Systematic comparisons against the state-of-the-art are conducted to validate the performance. Comprehensive onboard experiments using two different vision-based quadrotors are carried out showing the general applicability of the framework. Wenchao Ding 0001, Wenliang Gao, Shaojie Shen |
IEEE Trans. Robotics | 4 |
| 2019 | Tracking 3-D Motion of Dynamic Objects Using Monocular Visual-Inertial SensingabstractSix degree-of-freedom (6-DoF) visual tracking of dynamic objects is fundamental to a large variety of robotics and augmented reality (AR) applications. A key to this problem is accurate distance measurement of dynamic objects, which is usually obtained via stereo cameras, RGB-D sensors, or LiDARs. In this paper, however, we address the problem using only a monocular camera rigidly mounted with a low-cost inertial measurement unit. This is a light-weight, small-size, and low-cost solution, which is particularly suitable for tracking dynamic objects on drones or on mobile phones. Starting from a generic image-based two-dimensional tracker, we propose a novel method to resolve the object scale ambiguity in monocular vision in a geometric manner based on correlation analysis. This enables accurate metric three-dimensional tracking of arbitrary objects without requiring any prior knowledge about the object shape or size. We discuss the applicability by analyzing the observability condition and degenerated cases for object scale recovery. Simulation and real-world experimental results with ground truth comparison, along with AR application examples, demonstrate the feasibility of the proposed 6-DoF tracking method. Kejie Qiu, Tong Qin 0001, Wenliang Gao, Shaojie Shen |
IEEE Trans. Robotics | 4 |
| 2018 | MVDepthNet: Real-Time Multiview Depth Estimation Neural NetworkabstractAlthough deep neural networks have been widely applied to computer vision problems, extending them into multiview depth estimation is non-trivial. In this paper, we present MVDepthNet, a convolutional network to solve the depth estimation problem given several image-pose pairs from a localized monocular camera in neighbor viewpoints. Multiview observations are encoded in a cost volume and then combined with the reference image to estimate the depth map using an encoder-decoder network. By encoding the information from multiview observations into the cost volume, our method achieves real-time performance and the flexibility of traditional methods that can be applied regardless of the camera intrinsic parameters and the number of images. Geometric data augmentation is used to train MVDepthNet. We further apply MVDepthNet in a monocular dense mapping system that continuously estimates depth maps using a single localized moving camera. Experiments show that our method can generate depth maps efficiently and precisely. Shaojie Shen |
3DV | 2 |
| 2018 | Stereo Vision-Based Semantic 3D Object and Ego-Motion Tracking for Autonomous Driving
Peiliang Li 0001, Tong Qin 0001, Shaojie Shen |
ECCV (2) | 3 |
| 2018 | SLAM-based localization of 3D gaze using a mobile eye trackerabstractPast work in eye tracking has focused on estimating gaze targets in two dimensions (2D), e.g. on a computer screen or scene camera image. Three-dimensional (3D) gaze estimates would be extremely useful when humans are mobile and interacting with the real 3D environment. We describe a system for estimating the 3D locations of gaze using a mobile eye tracker. The system integrates estimates of the user's gaze vector from a mobile eye tracker, estimates of the eye tracker pose from a visual-inertial simultaneous localization and mapping (SLAM) algorithm, a 3D point cloud map of the environment from a RGB-D sensor. Experimental results indicate that our system produces accurate estimates of 3D gaze over a much larger range than remote eye trackers. Our system will enable applications, such as the analysis of 3D human attention and more anticipative human robot interfaces. Haofei Wang 0001, Jimin Pi, Tong Qin 0001, Shaojie Shen, Bertram E. Shi |
ETRA | 4 |
| 2018 | Trajectory Replanning for Quadrotors Using Kinodynamic Search and Elastic OptimizationabstractWe focus on a replanning scenario for quadrotors where considering time efficiency, non-static initial state and dynamical feasibility is of great significance. We propose a real-time B-spline based kinodynamic (RBK) search algorithm, which transforms a position-only shortest path search (such as A * and Dijkstra) into an efficient kinodynamic search, by exploring the properties of B-spline parameterization. The RBK search is greedy and produces a dynamically feasible time-parameterized trajectory efficiently, which facilitates non-static initial state of the quadrotor. To cope with the limitation of the greedy search and the discretization induced by a grid structure, we adopt an elastic optimization (EO) approach as a post-optimization process, to refine the control point placement provided by the RBK search. The EO approach finds the optimal control point placement inside an expanded elastic tube which represents the free space, by solving a Quadratically Constrained Quadratic Programming (QCQP) problem. We design a receding horizon replanner based on the local control property of B-spline. A systematic comparison of our method against two state-of-the-art methods is provided. We integrate our replanning system with a monocular vision-based quadrotor and validate our performance onboard. Wenchao Ding 0001, Wenliang Gao, Shaojie Shen |
ICRA | 4 |
| 2018 | Online Safe Trajectory Generation for Quadrotors Using Fast Marching Method and Bernstein Basis PolynomialabstractIn this paper, we propose a framework for online quadrotor motion planning for autonomous navigation in unknown environments. Based on the onboard state estimation and environment perception, we adopt a fast marching-based path searching method to find a path on a velocity field induced by the Euclidean signed distance field (ESDF) of the map, to achieve better time allocation. We generate a flight corridor for the quadrotor to travel through by inflating the path against the environment. We represent the trajectory as piecewise Bézier curves by using Bernstein polynomial basis and formulate the trajectory generation problem as typical convex programs. By using Bézier curves, we are able to bound positions and higher order dynamics of the trajectory entirely within safe regions. The proposed motion planning method is integrated into a customized light-weight quadrotor platform and is validated by presenting fully autonomous navigation in unknown cluttered indoor and outdoor environments. We also release our code for trajectory generation as an open-source package. Fei Gao 0011, William Wu, Yi Lin 0010, Shaojie Shen |
ICRA | 4 |
| 2018 | ACT: An Autonomous Drone Cinematography System for Action ScenesabstractDrones are enabling new forms of cinematography. Aerial filming via drones in action scenes is difficult because it requires users to understand the dynamic scenarios and operate the drone and camera simultaneously. Existing systems allow the user to manually specify the shots and guide the drone to capture footage, while none of them employ aesthetic objectives to automate aerial filming in action scenes. Meanwhile, these drone cinematography systems depend on the external motion capture systems to perceive the human action, which is limited to the indoor environment. In this paper, we propose an Autonomous CinemaTography system “ACT” on the drone platform to address the above the challenges. To our knowledge, this is the first drone camera system which can autonomously capture cinematic shots of action scenes based on limb movements in both indoor and outdoor environments. Our system includes the following novelties. First, we propose an efficient method to extract 3D skeleton points via a stereo camera. Second, we design a real-time dynamical camera planning strategy that fulfills the aesthetic objectives for filming and respects the physical limits of a drone. At the system level, we integrate cameras and GPUs into the limited space of a drone and demonstrate the feasibility of running the entire cinematography system onboard in real-time. Experimental results in both simulation and real-world scenarios demonstrate that our cinematography system “ACT” can capture more expressive video footage of human action than that of a state-of-the-art drone camera system. Chong Huang 0005, Fei Gao 0011, Jie Pan 0004, Weihao Qiu, Peng Chen 0008, Xin Yang 0008, Shaojie Shen, Kwang-Ting Cheng |
ICRA | 8 |
| 2018 | Relocalization, Global Optimization and Map Merging for Monocular Visual-Inertial SLAMabstractThe monocular visual-inertial system (VINS), which consists one camera and one low-cost inertial measurement unit (IMU), is a popular approach to achieve accurate 6-DOF state estimation. However, such locally accurate visual-inertial odometry is prone to drift and cannot provide absolute pose estimation. Leveraging history information to relocalize and correct drift has become a hot topic. In this paper, we propose a monocular visual-inertial SLAM system, which can relocalize camera and get the absolute pose in a previous-built map. Then 4-DOF pose graph optimization is performed to correct drifts and achieve global consistent. The 4-DOF contains x, y, z, and yaw angle, which is the actual drifted direction in the visual-inertial system. Furthermore, the proposed system can reuse a map by saving and loading it in an efficient way. Current map and previous map can be merged together by the global pose graph optimization. We validate the accuracy of our system on public datasets and compare against other state-of-the-art algorithms. We also evaluate the map merging ability of our system in the large-scale outdoor environment. The source code of map reuse is integrated into our public code, VINS-Monol11https://github.com/HKUST-Aerial-Robotics/VINS-Mono. Tong Qin 0001, Peiliang Li 0001, Shaojie Shen |
ICRA | 3 |
| 2018 | Learning Unmanned Aerial Vehicle Control for Autonomous Target FollowingabstractWhile deep reinforcement learning (RL) methods have achieved unprecedented successes in a range of challenging problems, their applicability has been mainly limited to simulation or game domains due to the high sample complexity of the trial-and-error learning process. However, real-world robotic applications often need a data-efficient learning process with safety-critical constraints. In this paper, we consider the challenging problem of learning unmanned aerial vehicle (UAV) control for tracking a moving target. To acquire a strategy that combines perception and control, we represent the policy by a convolutional neural network. We develop a hierarchical approach that combines a model-free policy gradient method with a conventional feedback proportional-integral-derivative (PID) controller to enable stable learning without catastrophic failure. The neural network is trained by a combination of supervised learning from raw images and reinforcement learning from games of self-play. We show that the proposed approach can learn a target following policy in a simulator efficiently and the learned behavior can be successfully transferred to the DJI quadrotor platform for real-world UAV control. Tianbo Liu 0001, Chi Zhang 0067, Dit-Yan Yeung, Shaojie Shen |
IJCAI | 5 |
| 2018 | Optimal Time Allocation for Quadrotor Trajectory GenerationabstractIn this paper, we present a framework to do optimal time allocation for quadrotor trajectory generation. Using this method, we can generate minimum-time piecewise polynomial trajectories for quadrotor flights. We decouple the quadrotor trajectory generation problem into two folds. Firstly we generate a smooth and safe curve which is parameterized by a virtual variable. This curve named spatial trajectory is independent of time and has fixed spatial properties. Then a mapping function which decides how the quadrotor moves along the spatial trajectory respecting kinodynamic limits is found by minimizing total trajectory time. The mapping function maps the virtual variable to time is named temporal trajectory. We formulate the minimum-time temporal trajectory generation problem as a convex program which can be efficiently solved. We show that the proposed method can corporate with various types of previous trajectory generation method to obtain the optimal time allocation. The proposed method is integrated into a customized light-weight quadrotor platform and is validated by presenting autonomous flights in indoor and outdoor environments. We release our code for time optimization as an open-source ros-package. Fei Gao 0011, William Wu, Jie Pan 0004, Boyu Zhou, Shaojie Shen |
IROS | 5 |
| 2018 | Probabilistic Dense Reconstruction from a Moving CameraabstractThis paper presents a probabilistic approach for online dense reconstruction using a single monocular camera moving through the environment. Compared to spatial stereo, depth estimation from motion stereo is challenging due to insufficient parallaxes, visual scale changes, pose errors, etc. We utilize both the spatial and temporal correlations of consecutive depth estimates to increase the robustness and accuracy of monocular depth estimation. An online, recursive, probabilistic scheme to compute depth estimates, with corresponding covariances and inlier probability expectations, is proposed in this work. We integrate the obtained depth hypotheses into dense 3D models in an uncertainty-aware way. We show the effectiveness and efficiency of our proposed approach by comparing it with state-of-the-art methods in the TUM RGB-D SLAM & ICL-NUIM dataset. Online indoor and outdoor experiments are also presented for performance demonstration. Yonggen Ling, Shaojie Shen |
IROS | 3 |
| 2018 | Online Temporal Calibration for Monocular Visual-Inertial SystemsabstractAccurate state estimation is a fundamental module for various intelligent applications, such as robot navigation, autonomous driving, virtual and augmented reality. Visual and inertial fusion is a popular technology for 6-DOF state estimation in recent years. Time instants at which different sensors' measurements are recorded are of crucial importance to the system's robustness and accuracy. In practice, timestamps of each sensor typically suffer from triggering and transmission delays, leading to temporal misalignment (time offsets) among different sensors. Such temporal offset dramatically influences the performance of sensor fusion. To this end, we propose an online approach for calibrating temporal offset between visual and inertial measurements. Our approach achieves temporal offset calibration by jointly optimizing time offset, camera and IMU states, as well as feature locations in a SLAM system. Furthermore, the approach is a general model, which can be easily employed in several feature-based optimization frameworks. Simulation and experimental results demonstrate the high accuracy of our calibration approach even compared with other state-of-art offline tools. The VIO comparison against other methods proves that the online temporal calibration significantly benefits visual-inertial systems. The source code of temporal calibration is integrated into our public project, VINS-Mono1. Tong Qin 0001, Shaojie Shen |
IROS | 2 |
| 2018 | Estimating Metric Poses of Dynamic Objects Using Monocular Visual-Inertial FusionabstractA monocular 3D object tracking system generally has only up-to-scale pose estimation results without any prior knowledge of the tracked object. In this paper, we propose a novel idea to recover the metric scale of an arbitrary dynamic object by optimizing the trajectory of the objects in the world frame, without motion assumptions. By introducing an additional constraint in the time domain, our monocular visual-inertial tracking system can obtain continuous six degree of freedom (6-DoF) pose estimation without scale ambiguity. Our method requires neither fixed multi-camera nor depth sensor settings for scale observability, instead, the IMU inside the monocular sensing suite provides scale information for both camera itself and the tracked object. We build the proposed system on top of our monocular visual-inertial system (VINS) to obtain accurate state estimation of the monocular camera in the world frame. The whole system consists of a 2D object tracker, an object region-based visual bundle adjustment (BA), VINS and a correlation analysis-based metric scale estimator. Experimental comparisons with ground truth demonstrate the tracking accuracy of our 3D tracking performance while a mobile augmented reality (AR) demo shows the feasibility of potential applications. Kejie Qiu, Tong Qin 0001, Hongwen Xie, Shaojie Shen |
IROS | 4 |
| 2018 | Quadtree-Accelerated Real-Time Monocular Dense MappingabstractIn this paper, we propose a novel mapping method for robotic navigation. High-quality dense depth maps are estimated and fused into 3D reconstructions in real-time using a single localized moving camera. The quadtree structure of the intensity image is used to reduce the computation burden by estimating the depth map in multiple resolutions. Both the quadtree-based pixel selection and the dynamic belief propagation are proposed to speed up the mapping process: pixels are selected and optimized with the computation resource according to their levels in the quadtree. Solved depth estimations are further interpolated and fused temporally into full resolution depth maps and fused into dense 3D maps using truncated signed distance function (TSDF). We compare our method with other state-of-the-art methods using the public datasets. Onboard UAV autonomous flight is also used to further prove the usability and efficiency of our method on portable devices. For the benefit of the community, the implementation is also released as open source at https://github.com/HKUST-Aerial-Robotics/open_quadtree_mapping. Wenchao Ding 0001, Shaojie Shen |
IROS | 3 |
| 2018 | Adaptive Baseline Monocular Dense Mapping with Inter-Frame Depth PropagationabstractState-of-the-art monocular dense mapping methods usually divide the image sequence into several separate multi-view stereo problems thus have limited utilization of the information in multi-baseline observations and sequential depth estimations. In this paper, two core contributions are proposed to improve the mapping performance by exploiting the information. The first is an adaptive baseline matching cost computation that uses the sequential input images to provide each pixel with wide-baseline observations. The second is a frame-to-frame propagated depth filter which integrates the sequential depth estimation of the same physical point in a robust probabilistic manner. Two contributions are integrated into a monocular dense mapping system that generates the depth maps in real-time for both pinhole and fisheye cameras. Our system is fully parallelized and can run at more than 25 fps on a Nvidia Jetson TX2. We compare our work with state-of-the-art methods on the public dataset. Onboard UAV mapping and handhold experiments are also used to demonstrate the performance of our method. For the benefit of the community, we make the implementation open source. Shaojie Shen |
IROS | 2 |
| 2018 | Guest Editorial Special Section on Aerial Swarm RoboticsabstractThe papers in this special section present recent advances in aerial swarm robotics, and aims to put together a cohesive set of research goals and visions toward realizing fully autonomous aerial swarm systems. One objective is to emphasize the three-way tradeoff among computational efficiency for large-scale swarms, stability, and robustness under uncertainty, and the optimal system performance. Aerial robotics has been one of the most active areas of research within the robotics community, and recently there have been many reports of promising results in aerial swarm systems. This is partly due to the commoditization of multicopter platforms, and communication, sensing, and processing hardware that has substantially lowered the barriers to entry to the field of aerial swarm robotics. Aerial swarms differ from swarms of ground-based vehicles in two major respects: Aerial robots or unmanned aerial vehicles (UAVs) operate in a three-dimensional space, and the dynamics of individual vehicles add an extra layer of complexity to the problems of path planning and trajectory design. Furthermore, the success of aerial swarms is predicated on the distributed and synergistic capabilities of individual and cooperative control, estimation, and decision making of aerial robots with limited resources, such as modest onboard computation and sensing capabilities and size, weight, and power constraints. Soon-Jo Chung, Aditya A. Paranjape, Philip M. Dames, Shaojie Shen, Vijay Kumar 0001 |
IEEE Trans. Robotics | 4 |
| 2018 | A Survey on Aerial Swarm RoboticsabstractThe use of aerial swarms to solve real-world problems has been increasing steadily, accompanied by falling prices and improving performance of communication, sensing, and processing hardware. The commoditization of hardware has reduced unit costs, thereby lowering the barriers to entry to the field of aerial swarm robotics. A key enabling technology for swarms is the family of algorithms that allow the individual members of the swarm to communicate and allocate tasks amongst themselves, plan their trajectories, and coordinate their flight in such a way that the overall objectives of the swarm are achieved efficiently. These algorithms, often organized in a hierarchical fashion, endow the swarm with autonomy at every level, and the role of a human operator can be reduced, in principle, to interactions at a higher level without direct intervention. This technology depends on the clever and innovative application of theoretical tools from control and estimation. This paper reviews the state of the art of these theoretical tools, specifically focusing on how they have been developed for, and applied to, aerial swarms. Aerial swarms differ from swarms of ground-based vehicles in two respects: they operate in a three-dimensional space and the dynamics of individual vehicles adds an extra layer of complexity. We review dynamic modeling and conditions for stability and controllability that are essential in order to achieve cooperative flight and distributed sensing. The main sections of this paper focus on major results covering trajectory generation, task allocation, adversarial control, distributed sensing, monitoring, and mapping. Wherever possible, we indicate how the physics and subsystem technologies of aerial robots are brought to bear on these individual areas. Soon-Jo Chung, Aditya A. Paranjape, Philip M. Dames, Shaojie Shen, Vijay Kumar 0001 |
IEEE Trans. Robotics | 4 |
| 2018 | VINS-Mono: A Robust and Versatile Monocular Visual-Inertial State EstimatorabstractOne camera and one low-cost inertial measurement unit (IMU) form a monocular visual-inertial system (VINS), which is the minimum sensor suite (in size, weight, and power) for the metric six degrees-of-freedom (DOF) state estimation. In this paper, we present VINS-Mono: a robust and versatile monocular visual-inertial state estimator. Our approach starts with a robust procedure for estimator initialization. A tightly coupled, nonlinear optimization-based method is used to obtain highly accurate visual-inertial odometry by fusing preintegrated IMU measurements and feature observations. A loop detection module, in combination with our tightly coupled formulation, enables relocalization with minimum computation. We additionally perform 4-DOF pose graph optimization to enforce the global consistency. Furthermore, the proposed system can reuse a map by saving and loading it in an efficient way. The current and previous maps can be merged together by the global pose graph optimization. We validate the performance of our system on public datasets and real-world experiments and compare against other state-of-the-art algorithms. We also perform an onboard closed-loop autonomous flight on the microaerial-vehicle platform and port the algorithm to an iOS-based demonstration. We highlight that the proposed work is a reliable, complete, and versatile system that is applicable for different applications that require high accuracy in localization. We open source our implementations for both PCs (https://github.com/HKUST-Aerial-Robotics/VINS-Mono) and iOS mobile devices (https://github.com/HKUST-Aerial-Robotics/VINS-Mobile). Tong Qin 0001, Peiliang Li 0001, Shaojie Shen |
IEEE Trans. Robotics | 3 |
| 2017 | Deep-mapnets : A residual network for 3D environment representationabstractThe ability to localize in the co-ordinate system of a 3D model presents an opportunity for safe trajectory planning. While SLAM-based approaches provide estimates of incremental poses with respect to the first camera frame, they do not provide global localization. With the availability of mobile GPUs like the Nvidia TX1 etc., our method provides a novel, elegant and high performance visual method for model based robot localization. We propose a method to learn an environment representation with deep residual nets for localization in a known 3D model representing a real-world area of 25,000 sq. meters. We use the power of modern GPUs and game engines for rendering training images mimicking a downward looking high flying drone using a photorealistic 3D model. We use these images to drive the learning loop of a 50-layer deep neural network to learn camera positions. We next propose to do data augmentation to accelerate training and to make our trained model robust for cross domain generalization, which has been verified with experiments. We test our trained model with synthetically generated data as well as real data captured from a downward looking drone. It takes about 25 miliseconds of GPU processing to predict camera pose. Unlike previous methods, the proposed method does not do rendering at test time and does independent prediction from a learned environment representation. Manohar Kuse, Sunil Prasad Jaiswal, Shaojie Shen |
ICIP | 3 |
| 2017 | Improving octree-based occupancy maps using environment sparsity with application to aerial robot navigationabstractIn this paper, we present an improved octree-based mapping framework for autonomous navigation of mobile robots. Octree is best known for its memory efficiency for representing large-scale environments. However, existing implementations, including the state-of-the-art OctoMap [1], are computationally too expensive for online applications that require frequent map updates and inquiries. Utilizing the sparse nature of the environment, we propose a ray tracing method with early termination for efficient probabilistic map update. We also propose a divide-and-conquer volume occupancy inquiry method which serves as the core operation for generation of free-space configurations for optimization-based trajectory generation. We experimentally demonstrate that our method maintains the same storage advantage of the original OctoMap, but being computationally more efficient for map update and occupancy inquiry. Finally, by integrating the proposed map structure in a complete navigation pipeline, we show autonomous quadrotor flight through complex environments. Jing Chen 0016, Shaojie Shen |
ICRA | 2 |
| 2017 | Quadrotor trajectory generation in dynamic environments using semi-definite relaxation on nonconvex QCQPabstractIn this paper, we present an optimization-based framework for generating quadrotor trajectories which are free of collision in dynamic environments with both static and moving obstacles. Using the finite-horizon motion prediction of moving obstacles, our method is able to generate safe and smooth trajectories with minimum control efforts. Our method optimizes trajectories globally for all observed moving and static obstacles, such that the avoidance behavior is most unnoticeable. This method first utilizes semi-definite relaxation on a quadratically constrained quadratic programming (QCQP) problem to eliminate the nonconvex constraints in the moving obstacle avoidance problem. A feasible and reasonably good solution to the original nonconvex problem is obtained using a randomization method and convex linear restriction. We detail the trajectory generation formulation and the solving procedure of the nonconvex quadratic program. Our approach is validated by both simulation and experimental results. Fei Gao 0011, Shaojie Shen |
ICRA | 2 |
| 2017 | High altitude monocular visual-inertial state estimation: Initialization and sensor fusionabstractObtaining reliable state estimates at high altitude but GPS-denied environments, such as between high-rise buildings or in the middle of deep canyons, is known to be challenging, due to the lack of direct distance measurements. Monocular visual-inertial systems provide a possible way to recover the metric distance through proper integration of visual and inertial measurements. However, the nonlinear optimization problem for state estimation suffers from poor numerical conditioning or even degeneration, due to difficulties in obtaining observations of visual features with sufficient parallax, and the excessive period of inertial measurement integration. In this paper, we propose a spline-based high altitude estimator initialization method for monocular visual-inertial navigation system (VINS) with special attention to the numerical issues. Our formulation takes only inertial measurements that contain sufficient excitation, and drops uninformative measurements such as those obtained during hovering. In addition, our method explicitly reduces the number of parameters to be estimated in order to achieve earlier convergence. Based on the initialization results, a complete closed-loop system is constructed for high altitude navigation. Extensive experiments are conducted to validate our approach. Tianbo Liu 0001, Shaojie Shen |
ICRA | 2 |
| 2017 | Design and implementation of a quadrotor tail-sitter VTOL UAVabstractWe present the design and implementation of a quadrotor tail-sitter Vertical Take-Off and Landing (VTOL) Unmanned Aerial Vehicle (UAV). The VTOL UAV combines the advantage of a quadrotor, vertical take-off and landing and hovering at a stationary point, with that of a fixed-wing, efficient level flight. We describe our vehicle design with special considerations on fully autonomous operation in a real outdoor environment where the wind is present. The designed quadrotor tail-sitter UAV has insignificant vibration level and achieves stable hovering and landing performance when a cross wind is present. Wind tunnel test is conducted to characterize the full envelope aerodynamics of the aircraft, based on which a flight controller is designed, implemented and tested. MATLAB simulation is presented and shows that our vehicle can achieve a continuous transition from hover flight to level flight. Finally, both indoor and outdoor flight experiments are conducted to verify the performance of our vehicle and the designed controller. Ximin Lyu, Haowei Gu, Zexiang Li 0001, Shaojie Shen, Fu Zhang 0002 |
ICRA | 5 |
| 2017 | Gesture-based piloting of an aerial robot using monocular visionabstractAerial robots are becoming popular among general public, and with the development of artificial intelligence (AI), there is a trend to equip aerial robots with a natural user interface (NUI). Hand/arm gestures are an intuitive way to communicate for humans, and various research works have focused on controlling an aerial robot with natural gestures. However, the techniques in this area are still far from mature. Many issues in this area have been poorly addressed, such as the principles of choosing gestures from the design point of view, hardware requirements from an economic point of view, considerations of data availability, and algorithm complexity from a practical perspective. Our work focuses on building an economical monocular system particularly designed for gesture-based piloting of an aerial robot. Natural arm gestures are mapped to rich target directions and convenient fine adjustment is achieved. Practical piloting scenarios, hardware cost and algorithm applicability are jointly considered in our system design. The entire system is successfully implemented in an aerial robot and various properties of the system are tested. Ting Sun 0001, Shengyi Nie, Dit-Yan Yeung, Shaojie Shen |
ICRA | 4 |
| 2017 | Real-time monocular dense mapping on aerial robots using visual-inertial fusionabstractIn this work, we present a solution to real-time monocular dense mapping. A tightly-coupled visual-inertial localization module is designed to provide metric and high-accuracy odometry. A motion stereo algorithm is proposed to take the video input from one camera to produce local depth measurements with semi-global regularization. The local measurements are then integrated into a global map for noise filtering and map refinement. The global map obtained is able to support navigation and obstacle avoidance for aerial robots through our indoor and outdoor experimental verification. Our system runs at 10Hz on an Nvidia Jetson TX1 by properly distributing computation to CPU and GPU. Through onboard experiments, we demonstrate its ability to close the perception-action loop for autonomous aerial robots. We release our implementation as open-source software1. Zhenfei Yang, Fei Gao 0011, Shaojie Shen |
ICRA | 3 |
| 2017 | Using a quadrotor to track a moving target with arbitrary relative motion patternsabstractWe propose a novel approach for safe tracking of a moving target in cluttered environments using a quadrotor. The key contribution of our work is a formulation that enables the generation of safe and dynamical feasible tracking trajectories that satisfy arbitrary relative motion patterns (circling, parallel tracking, undirectional tracking, etc.) with respect to the target. In our framework, forming the desired relative motion pattern between the quadrotor and the target only requires a generative function that specifies relative positions at different time instants. Our method generates samples to fit a piecewise-polynomial representation of the desired relative motion pattern and embeds it into a cost function for solving valid tracking trajectories via quadratic programming. Collision avoidance is achieved by squeezing the trajectory into a collision-free flight corridor, and dynamical feasibility is achieved by enforcing bounds on corresponding derivatives. Both of which can be written as linear constraints for the quadratic programming. Our approach is lightweight and can be implemented for real-time target tracking. We use a simulated cluttered environment and multiple desired relative motion patterns to demonstrate the performance of the proposed approach. Jing Chen 0016, Shaojie Shen |
IROS | 2 |
| 2017 | Gradient-based online safe trajectory generation for quadrotor flight in complex environmentsabstractIn this paper, we propose a trajectory generation framework for quadrotor autonomous navigation in unknown 3-D complex environments using gradient information. We decouple the trajectory generation problem as front-end path searching and back-end trajectory refinement. Based on the map that is incrementally built onboard, we adopt a sampling-based informed path searching method to find a safe path passing through obstacles. We convert the path consists of line segments to an initial safe trajectory. An optimization-based method which minimizes the penalty of collision cost, smoothness and dynamical feasibility is used to refine the trajectory. Our method shows the ability to online generate smooth and dynamical feasible trajectories with safety guarantee. We integrate the state estimation, dense mapping and motion planning module into a customized light-weight quadrotor platform. We validate our proposed method by presenting fully autonomous navigation in unknown cluttered indoor and outdoor environments. Fei Gao 0011, Yi Lin 0010, Shaojie Shen |
IROS | 3 |
| 2017 | Dual-fisheye omnidirectional stereoabstractWe propose a novel omnidirectional stereo camera setup that is formed by two ultra-wide field-of-view (FOV) fisheye cameras. The proposed configuration is formed by two 245-degree FOV fisheye cameras, facing opposite directions, that are rigidly mounted on two sides of a rod. The overlapping view in the two fisheye images forms a ring-shaped spatial stereo setup. Our system provides stereo observations with full 360-degree FOV in horizontal directions and 65-degree FOV in the vertical direction. In addition, the two fisheye cameras altogether also provide full spherical monocular coverage of the surrounding environment. We address challenges in camera modeling, fisheye intrinsic calibration, stereo self-calibration, and depth estimation. We develop a lens-specific fisheye camera calibration method that uses manufacturer data for the optical lens to assist with the intrinsic calibration and develop an online self-calibration approach for estimating stereo extrinsic parameters. The overlapping camera views are rectified into stereo image pairs, from which a spatial stereo matching pipeline is developed for depth estimation in all horizontal directions. We show both qualitative and quantitative analysis to validate our approach. Wenliang Gao, Shaojie Shen |
IROS | 2 |
| 2017 | Building maps for autonomous navigation using sparse visual SLAM featuresabstractAutonomous navigation, which consists of a systematic integration of localization, mapping, motion planning and control, is the core capability of mobile robotic systems. However, most research considers only isolated technical modules. There exist significant gaps between maps generated by SLAM algorithms and maps required for motion planning. This paper presents a complete online system that consists in three modules: incremental SLAM, real-time dense mapping, and free space extraction. The obtained free-space volume (i.e. a tessellation of tetrahedra) can be served as regular geometric constraints for motion planning. Our system runs in real-time thanks to the engineering decisions proposed to increase the system efficiency. We conduct extensive experiments on the KITTI dataset to demonstrate the run-time performance. Qualitative and quantitative results on mapping accuracy are also shown. For the benefit of the community, we make the source code public. Yonggen Ling, Shaojie Shen |
IROS | 2 |
| 2017 | A hierarchical control approach for a quadrotor tail-sitter VTOL UAV and experimental verificationabstractWe present a hierarchical control approach that can be used to fulfill autonomous flight, including vertical takeoff, landing, hovering, transition, and level flight, of a quadrotor tail-sitter vertical takeoff and landing unmanned aerial vehicle (VTOL UAV). A unified attitude controller, together with a moment allocation scheme between elevons and motor differential thrust, is developed for all flight modes. A comparison study via real flight tests is performed to verify the effectiveness of using elevons in addition to motor differential thrust. With the well-designed switch scheme proposed in this paper, the aircraft can transit between different flight modes with negligible altitude drop or gain. Intensive flight tests have been performed to verify the effectiveness of the proposed control approach in both manual and fully autonomous flight mode. Ximin Lyu, Haowei Gu, Jinni Zhou, Zexiang Li 0001, Shaojie Shen, Fu Zhang 0002 |
IROS | 5 |
| 2017 | Robust initialization of monocular visual-inertial estimation on aerial robotsabstractIn this paper, we propose a robust on-the-fly estimator initialization algorithm to provide high-quality initial states for monocular visual-inertial systems (VINS). Due to the non-linearity of VINS, a poor initialization can severely impact the performance of either filtering-based or graph-based methods. Our approach starts with a vision-only structure from motion (SfM) to build the up-to-scale structure of camera poses and feature positions. By loosely aligning this structure with pre-integrated IMU measurements, our approach recovers the metric scale, velocity, gravity vector, and gyroscope bias, which are treated as initial values to bootstrap the nonlinear tightly-coupled optimization framework. We highlight that our approach can perform on-the-fly initialization in various scenarios without using any prior information about system states and movement. The performance of the proposed approach is verified through the public UAV dataset and real-time onboard experiment. We make our implementation open source, which is the initialization part integrated in the VINS-Mono1. Tong Qin 0001, Shaojie Shen |
IROS | 2 |
| 2017 | Model-aided monocular visual-inertial state estimation and dense mappingabstractRobust state estimation and real-time dense mapping are two core capabilities for autonomous navigation of mobile robots. Global Navigation Satellite System (GNSS) and visual odometry/SLAM are popular methods for state estimation. However, when working between tall buildings or in indoor environments, GNSS fails due to limited sky view or obstruction from buildings. Visual odometry/SLAM are prone to long-term drifting in the absence of reliable loop closure detection. A state estimation method with global-consistent guarantee is desirable for navigation applications. As for real-time mapping, SLAM methods usually get a sparse map that is not good enough for obstacle avoidance and path-planning, and high-quality dense mapping is often computationally too demanding for mobile devices. Realizing the availability of city-scale 3D models, in this work, we improve our previous work on model-based global localization, and propose a model-aided monocular visual-inertial state estimation and dense mapping solution. We first develop a global-consistent state estimator by fusing visual-inertial odometry with the model-based localization results. Utilizing depth prior from the model, we perform motion stereo with semi-global disparity smoothing. Our dense mapping pipeline is capable of online detection of obstacles that are originally not included in the offline 3D model. Our method runs onboard an embedded computer in real-time. We validate both the state estimation and mapping accuracy in real-world experiments. Kejie Qiu, Shaojie Shen |
IROS | 2 |
| 2017 | A unified control method for quadrotor tail-sitter UAVs in all flight modes: Hover, transition, and level flightabstractThis paper presents a unified control framework for controlling a quadrotor tail-sitter UAV. The most salient feature of this framework is its capability of uniformly treating the hovering and forward flight, and enabling continuous transition between these two modes, depending on the commanded velocity. The key part of this framework is a nonlinear solver that solves for the proper attitude and thrust that produces the required acceleration set by the position controller in an online fashion. The planned attitude and thrust are then achieved by an inner attitude controller that is global asymptotically stable. To characterize the aircraft aerodynamics, a full envelope wind tunnel test is performed on the full-scale quadrotor tail-sitter UAV. In addition to planning the attitude and thrust required by the position controller, this framework can also be used to analyze the UAV's equilibrium state (trimmed condition), especially when wind gust is present. Finally, simulation results are presented to verify the controller's capacity, and experiments are conducted to show the attitude controller's performance. Jinni Zhou, Ximin Lyu, Zexiang Li 0001, Shaojie Shen, Fu Zhang 0002 |
IROS | 4 |
| 2017 | Monocular Visual-Inertial State Estimation for Mobile Augmented RealityabstractMobile phones equipped with a monocular camera and an inertial measurement unit (IMU) are ideal platforms for augmented reality (AR) applications, but the lack of direct metric distance measurement and the existence of aggressive motions pose significant challenges on the localization of the AR device. In this work, we propose a tightly-coupled, optimization-based, monocular visual-inertial state estimation for robust camera localization in complex indoor and outdoor environments. Our approach does not require any artificial markers, and is able to recover the metric scale using the monocular camera setup. The whole system is capable of online initialization without relying on any assumptions about the environment. Our tightly-coupled formulation makes it naturally robust to aggressive motions. We develop a lightweight loop closure module that is tightly integrated with the state estimator to eliminate drift. The performance of our proposed method is demonstrated via comparison against state-of-the-art visual-inertial state estimators on public datasets and real-time AR applications on mobile devices. We release our implementation on mobile devices as open source software1. Peiliang Li 0001, Tong Qin 0001, Botao Amber Hu, Shaojie Shen |
ISMAR | 5 |
| 2017 | Monocular Visual-Inertial State Estimation With Online Initialization and Camera-IMU Extrinsic CalibrationabstractThere have been increasing demands for developing microaerial vehicles with vision-based autonomy for search and rescue missions in complex environments. In particular, the monocular visual-inertial system (VINS), which consists of only an inertial measurement unit (IMU) and a camera, forms a great lightweight sensor suite due to its low weight and small footprint. In this paper, we address two challenges for rapid deployment of monocular VINS: 1) the initialization problem and 2) the calibration problem. We propose a methodology that is able to initialize velocity, gravity, visual scale, and camera-IMU extrinsic calibration on the fly. Our approach operates in natural environments and does not use any artificial markers. It also does not require any prior knowledge about the mechanical configuration of the system. It is a significant step toward plug-and-play and highly customizable visual navigation for mobile robots. We show through online experiments that our method leads to accurate calibration of camera-IMU transformation, with errors less than 0.02 m in translation and 1° in rotation. We compare out method with a state-of-the-art marker-based offline calibration method and show superior results. We also demonstrate the performance of the proposed approach in large-scale indoor and outdoor experiments. Zhenfei Yang, Shaojie Shen |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2016 | Online generation of collision-free trajectories for quadrotor flight in unknown cluttered environmentsabstractWe present an online method for generating collision-free trajectories for autonomous quadrotor flight through cluttered environments. We consider the real-world scenario that the quadrotor aerial robot is equipped with limited sensing and operates in initially unknown environments. During flight, an octree-based environment representation is incrementally built using onboard sensors. Utilizing efficient operations in the octree data structure, we are able to generate free-space flight corridors consisting of large overlapping 3-D grids in an online fashion. A novel optimization-based method then generates smooth trajectories that both are bounded entirely within the safe flight corridor and satisfy higher order dynamical constraints. Our method computes valid trajectories within fractions of a second on a moderately fast computer, thus permitting online re-generation of trajectories for reaction to new obstacles. We build a complete quadrotor testbed with onboard sensing, state estimation, mapping, and control, and integrate the proposed method to show online navigation through complex unknown environments. Jing Chen 0016, Tianbo Liu 0001, Shaojie Shen |
ICRA | 3 |
| 2016 | Robust camera motion estimation using direct edge alignment and sub-gradient methodabstractThere has been a paradigm shifting trend towards feature-less methods due to their elegant formulation, accuracy and ever increasing computational power. In this work, we present a direct edge alignment approach for 6-DOF tracking. We argue that photo-consistency based methods are plagued by a much smaller convergence basin and are extremely sensitive to noise, changing illumination and fast motion. We propose to use the Distance Transform in the energy formulation which can significantly extend the influence of the edges for tracking. We address the problem of non-differentiability of our cost function and of the previous methods by use of a sub-gradient method. Through extensive experiments we show that the proposed method gives comparable performance to the previous method under nominal conditions and is able to run at 30 Hz in single threaded mode. In addition, under large motion we demonstrate our method outperforms previous methods using the same runtime configuration for our method. Manohar Kuse, Shaojie Shen |
ICRA | 2 |
| 2016 | Aggressive quadrotor flight using dense visual-inertial fusionabstractIn this work, we address the problem of aggressive flight of a quadrotor aerial vehicle using cameras and IMUs as the only sensing modalities. We present a fully integrated quadrotor system and demonstrate through online experiment the capability of autonomous flight with linear velocities up to 4.2 m/s, linear accelerations up to 9.6 m/s2, and angular velocities up to 245.1 degree/s. Central to our approach is a dense visual-inertial state estimator for reliable tracking of aggressive motions. An uncertainty-aware direct dense visual tracking module provides camera pose tracking that takes inverse depth uncertainty into account and is resistant to motion blur. Measurements from IMU pre-integration and multi-constrained dense visual tracking are fused probabilistically using an optimization-based sensor fusion framework. Extensive statistical analysis and comparison are presented to verify the performance of the proposed approach. We also release our code as open-source ROS packages. Yonggen Ling, Tianbo Liu 0001, Shaojie Shen |
ICRA | 3 |
| 2016 | Tracking a moving target in cluttered environments using a quadrotorabstractWe address the challenging problem of tracking a moving target in cluttered environments using a quadrotor. Our online trajectory planning method generates smooth, dynamically feasible, and collision-free polynomial trajectories that follow a visually-tracked moving target. As visual observations of the target are obtained, the target trajectory can be estimated and used to predict the target motion for a short time horizon. We propose a formulation to embed both limited horizon tracking error and quadrotor control costs in the cost function for a quadratic programming (QP), while encoding both collision avoidance and dynamical feasibility as linear inequality constraints for the QP. Our method generates tracking trajectories in the order of milliseconds and is therefore suitable for online target tracking with a limited sensing range. We implement our approach on-board a quadrotor testbed equipped with cameras, a laser range finder, an IMU, and onboard computing. Statistical analysis, simulation, and real-world experiments are conducted to demonstrate the effectiveness of our approach. Jing Chen 0016, Tianbo Liu 0001, Shaojie Shen |
IROS | 3 |
| 2016 | High-precision online markerless stereo extrinsic calibrationabstractStereo cameras and dense stereo matching algorithms are core components for many robotic applications due to their abilities to directly obtain dense depth measurements and their robustness against changes in lighting conditions. However, the performance of dense depth estimation relies heavily on accurate stereo extrinsic calibration. In this work, we present a real-time markerless approach for obtaining high-precision stereo extrinsic calibration using a novel 5-DOF (degrees-of-freedom) and nonlinear optimization on a manifold, which captures the observability property of vision-only stereo calibration. Our method minimizes epipolar errors between spatial per-frame sparse natural features. It does not require temporal feature correspondences, making it not only invariant to dynamic scenes and illumination changes, but also able to run significantly faster than standard bundle adjustment-based approaches. We introduce a principled method to determine if the calibration converges to the required level of accuracy, and show through online experiments that our approach achieves a level of accuracy that is comparable to offline marker-based calibration methods. Our method refines stereo extrinsic to the accuracy that is sufficient for block matching-based dense disparity computation. It provides a cost-effective way to improve the reliability of stereo vision systems for long-term autonomy. Yonggen Ling, Shaojie Shen |
IROS | 2 |
| 2016 | Self-calibrating multi-camera visual-inertial fusion for autonomous MAVsabstractWe address the important problem of achieving robust and easy-to-deploy visual state estimation for micro aerial vehicles (MAVs) operating in complex environments. We use a sensor suite consisting of multiple cameras and an IMU to maximize perceptual awareness of the surroundings and provide sufficient redundancy against sensor failures. Our approach starts with an online initialization procedure that simultaneously estimates the transformation between each camera and the IMU, as well as the initial velocity and attitude of the platform, without any prior knowledge about the mechanical configuration of the sensor suite. Based on the initial calibrations, a tightly-coupled, optimization-based, generalized multi-camera-inertial fusion method runs onboard the MAV with online camera-IMU calibration refinement and identification of sensor failures. Our approach dynamically configures the system into monocular, stereo, or other multi-camera visual-inertial settings, with their respective perceptual advantages, based on the availability of visual measurements. We show that even under random camera failures, our method can be used for feedback control of the MAVs. We highlight our approach in challenging indoor-outdoor navigation tasks with large variations in vehicle height and speed, scene depth, and illumination. Zhenfei Yang, Tianbo Liu 0001, Shaojie Shen |
IROS | 3 |
| 2015 | Tightly-coupled monocular visual-inertial fusion for autonomous flight of rotorcraft MAVsabstractThere have been increasing interests in the robotics community in building smaller and more agile autonomous micro aerial vehicles (MAVs). In particular, the monocular visual-inertial system (VINS) that consists of only a camera and an inertial measurement unit (IMU) forms a great minimum sensor suite due to its superior size, weight, and power (SWaP) characteristics. In this paper, we present a tightly-coupled nonlinear optimization-based monocular VINS estimator for autonomous rotorcraft MAVs. Our estimator allows the MAV to execute trajectories at 2 m/s with roll and pitch angles up to 30 degrees. We present extensive statistical analysis to verify the performance of our approach in different environments with varying flight speeds. Shaojie Shen, Nathan Michael, Vijay Kumar 0001 |
ICRA | 1 |
| 2014 | Multi-sensor fusion for robust autonomous flight in indoor and outdoor environments with a rotorcraft MAVabstractWe present a modular and extensible approach to integrate noisy measurements from multiple heterogeneous sensors that yield either absolute or relative observations at different and varying time intervals, and to provide smooth and globally consistent estimates of position in real time for autonomous flight. We describe the development of algorithms and software architecture for a new 1.9kg MAV platform equipped with an IMU, laser scanner, stereo cameras, pressure altimeter, magnetometer, and a GPS receiver, in which the state estimation and control are performed onboard on an Intel NUC 3rdgeneration i3 processor. We illustrate the robustness of our framework in large-scale, indoor-outdoor autonomous aerial navigation experiments involving traversals of over 440 meters at average speeds of 1.5 m/s with winds around 10 mph while entering and exiting buildings. Shaojie Shen, Yash Mulgaonkar, Nathan Michael, Vijay Kumar 0001 |
ICRA | 1 |
| 2013 | Vision-based state estimation for autonomous rotorcraft MAVs in complex environmentsabstractIn this paper, we consider the development of a rotorcraft micro aerial vehicle (MAV) system capable of vision-based state estimation in complex environments. We pursue a systems solution for the hardware and software to enable autonomous flight with a small rotorcraft in complex indoor and outdoor environments using only onboard vision and inertial sensors. As rotorcrafts frequently operate in hover or nearhover conditions, we propose a vision-based state estimation approach that does not drift when the vehicle remains stationary. The vision-based estimation approach combines the advantages of monocular vision (range, faster processing) with that of stereo vision (availability of scale and depth information), while overcoming several disadvantages of both. Specifically, our system relies on fisheye camera images at 25 Hz and imagery from a second camera at a much lower frequency for metric scale initialization and failure recovery. This estimate is fused with IMU information to yield state estimates at 100 Hz for feedback control. We show indoor experimental results with performance benchmarking and illustrate the autonomous operation of the system in challenging indoor and outdoor environments. Shaojie Shen, Yash Mulgaonkar, Nathan Michael, Vijay Kumar 0001 |
ICRA | 1 |
| 2012 | Autonomous indoor 3D exploration with a micro-aerial vehicleabstractIn this paper, we propose a stochastic differential equation-based exploration algorithm to enable exploration in three-dimensional indoor environments with a payload constrained micro-aerial vehicle (MAV). We are able to address computation, memory, and sensor limitations by considering only the known occupied space in the current map. We determine regions for further exploration based on the evolution of a stochastic differential equation that simulates the expansion of a system of particles with Newtonian dynamics. The regions of most significant particle expansion correlate to unexplored space. After identifying and processing these regions, the autonomous MAV navigates to these locations to enable fully autonomous exploration. The performance of the approach is demonstrated through numerical simulations and experimental results in single and multi-floor indoor experiments. Shaojie Shen, Nathan Michael, Vijay Kumar 0001 |
ICRA | 1 |
| 2011 | Autonomous multi-floor indoor navigation with a computationally constrained MAVabstractIn this paper, we consider the problem of autonomous navigation with a micro aerial vehicle (MAV) in indoor environments. In particular, we are interested in autonomous navigation in buildings with multiple floors. To ensure that the robot is fully autonomous, we require all computation to occur on the robot without need for external infrastructure, communication, or human interaction beyond high-level commands. Therefore, we pursue a system design and methodology that enables autonomous navigation with real time performance on a mobile processor using only onboard sensors. Specifically, we address multi-floor mapping with loop closure, localization, planning, and autonomous control, including adaptation to aerodynamic effects during traversal through spaces with low vertical clearance or strong external disturbances. We present experimental results with ground truth comparisons and performance analysis. Shaojie Shen, Nathan Michael, Vijay Kumar 0001 |
ICRA | 1 |
| 2011 | Autonomous multi-floor indoor navigation with a computationally constrained micro aerial vehicleabstractIn this paper, we consider the problem of autonomous navigation with a micro aerial vehicle (MAV) in indoor environments. In particular, we are interested in autonomous navigation in buildings with multiple floors. To ensure that the robot is fully autonomous, we require all computation to occur on the robot without need for external infrastructure, communication, or human interaction beyond high-level commands. Therefore, we pursue a system design and methodology that enables autonomous navigation with real time performance on a mobile processor using only onboard sensors. Specifically, we address multi-floor mapping with loop closure, localization, planning, and autonomous control, including adaptation to aerodynamic effects during traversal through spaces with low vertical clearance or strong external disturbances. We present experimental results with ground truth comparisons and performance analysis. Shaojie Shen, Nathan Michael, Vijay Kumar 0001 |
ICRA | 1 |