EDBT 2026 Demo / reviewers in the wild / expert
Yuxiang Sun 0002
dblp:75/1112-2
· DBLP profile ↗
35ranked-venue papers
2as first author
29since 2021 · last 2026
0000-0002-7704-0559ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 18 · 2 first-author · 16 since 2021Artificial intelligence and machine learning · 15 · 11 since 2021Systems, architecture and hardware · 11 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Greedy Priority Inheritance With Backtracking for Multi-Agent Pathfinding ProblemabstractSome real-world transportation systems require moving robots from their start points to goal points without collision, which can be formalized as the multi-agent pathfinding (MAPF) problem. The Priority Inheritance with Backtracking (PIBT) algorithm is an efficient approach for MAPF with low-degree polynomial time complexity and proven completeness. PIBT employs a dynamic prioritization scheme to determine the order of sequential agent planning. However, this scheme may lead to low solution quality, which is measured by the sum of time steps each agent takes to reach its goal for the first time. In this work, we propose two greedy variants of the PIBT algorithm that improve solution quality while preserving completeness: Backflow-based Greedy PIBT (GPIBT-B) and Reduction-based Greedy PIBT (GPIBT-R). GPIBT-B introduces a backflow mechanism that allows a determined agent to adjust its plan during the planning of another agent, achieving higher solution quality while maintaining the same low time complexity as PIBT. GPIBT-R formulates the planning problem as a Mixed-Integer Linear Programming (MILP), enabling the use of efficient MILP solvers to find high-quality solutions. Experimental results show that GPIBT-B substantially improves solution quality with minimal additional computation time, while GPIBT-R achieves even better solution quality at the cost of increased computational time. Mingkai Tang 0002, Yuanhang Li, Lu Gan 0001, Chengxi Zhang, Yuxiang Sun 0002, Jin Wu 0002 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2026 | Spatial Balancing for RGB-Thermal Semantic Segmentation in Autonomous Driving: A Study From Analysis to ImprovementabstractSemantic segmentation based on RGB-Thermal (RGB-T) data fusion has made great progress in the field of autonomous driving. However, we find that most existing RGB-T semantic segmentation methods exhibit inferior performance in image central regions, in which segmentation performance is critical for driving safety. We refer to this phenomenon as spatial bias. To discover the reason for spatial bias, we design a series of experiments. The results challenge the common knowledge that more training data lead to better segmentation performance, and reveal a close causal relationship between segmentation performance and object complexity as well as image quality. We also provide a theoretical interpretation for the causal relationship using information theory and feature space analysis. Based on the findings, we propose a Gaussian-guided regional balancing masking method to balance segmentation performance across different image regions. Moreover, we introduce a spatial-weighted loss to further enhance the overall segmentation performance. Experimental results on two public datasets demonstrate the effectiveness of our method in mitigating spatial bias and improving balanced performance. Henry K. Chu, Yuxiang Sun 0002 |
IEEE Trans. Robotics | 3 |
| 2025 | Dense Semantic Bird-Eye-View Map Generation from Sparse LiDAR Point Clouds via Distribution-aware Feature FusionabstractSemantic scene understanding in bird-eye view (BEV) plays a crucial role in autonomous driving. A common approach to generating BEV maps from LiDAR point-cloud data involves constructing a pillar-level representation by projecting 3D point clouds onto a 2D plane. This process partially discards spatial geometric information, and produces sparse semantic maps. However, downstream tasks (e.g., trajectory planning and prediction), typically require dense grid-like semantic BEV maps rather than sparse segmentation outputs. To bridge this gap, we propose PointDenseBEV, an end-to-end, distribution-aware feature fusion framework. It takes as input sparse LiDAR point clouds and directly generates dense semantic BEV maps. Spatial geometric information and temporal context are embedded as auxiliary semantic cues within the BEV grid representation to enhance semantic density. Extensive experiments on the SemanticKITTI dataset demonstrate that our method achieves competitive performance compared to existing approaches. Kunyu Peng, Yuxiang Sun 0002 |
IROS | 3 |
| 2025 | MMFSeg: Multi-Structure Multi-Feature Fusion for Segmentation of Road Potholes
Yanning Guo, Rui Fan 0001, Yuxiang Sun 0002 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2025 | Boundary-Aware Semantic Bird-Eye-View Map Generation Based on Conditional Diffusion ModelsabstractSemantic bird-eye-view (BEV) map is an efficient data representation for environment perception in autonomous driving. In real driving scenarios, the collected sensory data usually exhibit class imbalance. For example, road layouts are often the majority classes and road objects are the minority. Such imbalanced data could lead to inferior performance in BEV map generation, particularly for minority objects due to insufficient learning samples. This work attempts to mitigate this issue from the perspective of network and loss function design. To this end, a diffusion-guided semantic BEV map generation network with a boundary-aware loss is proposed. The network learns the underlying distribution of the data, including the relationship between majority and minority classes. The boundary-aware loss increases weighting for minority classes during training, making the network focus on these classes. Experimental results on a public dataset demonstrate our superiority over the state-of-the-art methods, and our effectiveness in addressing the class imbalance issue. Qiang Wang 0001, Yuxiang Sun 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Seq-BEV: Semantic Bird-Eye-View Map Generation in Full View Using Sequential Images for Autonomous DrivingabstractSemantic Bird-Eye-View (BEV) map is a straightforward data representation for environment perception. It can be used for downstream tasks, such as motion planning and trajectory prediction. However, taking as input a front-view image from a single camera, most existing methods can only provide V-shaped semantic BEV maps, which limits the field-of-view for the BEV maps. To provide a solution to this problem, we propose a novel end-to-end network to generate semantic BEV maps in full view by taking as input the equidistant sequential images. Specifically, we design a self-adapted sequence fusion module to fuse the features from different images in a distance sequence. In addition, a road-aware view transformation module is introduced to wrap the front-view feature map into BEV based on an attention mechanism. We also create a dataset with semantic labels in full BEV from the public nuScenes data. The experimental results demonstrate the effectiveness of our design and the superiority over the state-of-the-art methods. Qiang Wang 0001, Yuxiang Sun 0002 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2025 | SkyLoc: Cross-Modal Global Localization With a Sky-Looking Fish-Eye Camera and OpenStreetMapabstractGlobal localization can estimate geo-referenced locations (e.g., longitude and latitude), which is a fundamental capability for autonomous vehicles. Most existing solutions rely on the Global Navigation Satellite Systems (GNSS). Their accuracy could be degraded by the multi-path effects or occlusions of GNSS signals in urban environments. Some GNSS-free methods could achieve global localization by comparing the current on-line sensory data with pre-built databases/maps. However, they require tedious human efforts to drive a vehicle to collect and maintain the databases/maps. Moreover, most of these methods use front-looking cameras or LiDARs, so the captured data could be easily contaminated by dynamic objects (e.g., moving vehicles and pedestrians). To provide a solution to these problems, this paper proposes a novel global localization method by comparing an image from a sky-looking fish-eye camera with the publicly available OpenStreetMap (OSM), and using particle filter to achieve real-time metric localization in dynamic traffic environments. To evaluate our method, we extend a public dataset with OSM data, which are retrieved through the given geo-referenced location information. Experimental results demonstrate the effectiveness and efficiency of our method. Weixin Ma, Shoudong Huang, Yuxiang Sun 0002 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2025 | Semantic-MoSeg: Semantics-Assisted Moving-Obstacle Segmentation in Bird-Eye-View for Autonomous DrivingabstractBird-eye-view (BEV) perception for autonomous driving has become popular in recent years. Among various BEV perception tasks, moving-obstacle segmentation is very important, since it can provide necessary information for downstream tasks, such as motion planning and decision making, in dynamic traffic environments. Many existing methods segment moving obstacles with LiDAR point clouds. The point-wise segmentation results can be easily projected into BEV since point clouds are 3-D data. However, these methods could not produce dense 2-D BEV segmentation maps, because LiDAR point clouds are usually sparse. Moreover, 3-D LiDARs are still expensive to vehicles. To provide a solution to these issues, this paper proposes a semantics-assisted moving-obstacle segmentation network using only low-cost visual cameras to produce segmentation results in dense 2-D BEV maps. Our network takes as input visual images from six surrounding cameras as well as the corresponding semantic segmentation maps at the current and previous moments, and directly outputs the BEV map for the current moment. We also propose a movable-obstacle segmentation auxiliary task to provide semantic information to further benefit moving-obstacle segmentation. Extensive experimental results on the public nuScenes and Lyft datasets demonstrate the effectiveness and superiority of our network. Shiyu Meng, Yuxiang Sun 0002 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | Forecasting Semantic Bird-Eye-View Maps for Autonomous DrivingabstractCorrectly understanding surrounding environments is a fundamental capability for autonomous driving. Semantic forecasting of bird-eye-view (BEV) maps can provide semantic perception information in advance, which is important for environment understanding. Currently, the research works on combining semantic forecasting and semantic BEV map generation is limited. Most existing work focuses on individual tasks only. In this work, we attempt to forecast semantic BEV maps in an end-to-end framework for future front-view (FV) images. To this end, we predict depth distributions and context features for FV input images and then forecast depth-context features for the future. The depth-context features are finally converted to the future semantic BEV maps. We conduct ablation studies and create baselines for evaluation and comparison. The results demonstrate that our network achieves superior performance. Qiang Wang 0001, David Navarro-Alarcon, Yuxiang Sun 0002 |
IV | 4 |
| 2024 | Obstacle-sensitive Semantic Bird-Eye-View Map Generation with Boundary-aware Loss for Autonomous drivingabstractDetection of road obstacles is important for autonomous driving. However, road obstacles, like pedestrians, usually account for quite a small portion compared with other semantics, such as road layouts. This leads to the class-imbalance problem in real-world driving datasets and hinders environment perception for autonomous driving. In this paper, we propose an obstacle-sensitive network to improve the semantic Bird-Eye-View (BEV) map generation performance for minority classes. To this end, a context-depth attention module and a boundary-aware loss are introduced. We conduct ablation studies to verify the effectiveness of the proposed network. We also compare our network with other semantic BEV map generation methods. The results demonstrate that our network achieves better performance in terms of semantic BEV map generation, especially for minority classes. Qiang Wang 0001, Yuxiang Sun 0002 |
IV | 3 |
| 2024 | ST-TrackNet: A Multiple-Object Tracking Network Using Spatio-Temporal InformationabstractMultiple-object tracking (MOT) is a crucial component in autonomous driving systems. However, inaccurate object detection is always the bottleneck for MOT. Most detectors are not designed to take the temporal information across consecutive frames into consideration. To take advantage of such information, we design a novel data representation, the spatio-temporal (ST) map, which collects a batch of detection results spatio-temporally, and we train a novel network, ST-TrackNet, to assign predicted track IDs to each positive detection across a sequence. With our ST map detection fed into the tracker, the correlation of objects between adjacent frames becomes prominent, which improves the performance of the tracker in the data association step. Moreover, the long-term trajectory in a sequence also helps to refine the detection results. We train and evaluate our network on the KITTI dataset, a CARLA simulation dataset, and a dataset recorded in a factory environment. Our approach generally achieves superior performance over the state-of-the-art. Note to Practitioners—We investigate the MOT problem in this paper. A spatio-temporal pipeline is proposed to provide a solution to this problem. Object detection results produced by off-the-shelf object detectors are used to form the proposed ST maps. In low signal-to-noise ratio (SNR) situations, our proposed framework can achieve more accurate and robust tracking results with more false-positives. Due to the simplicity and modular design of our framework, it can be applied directly after the detection stage to achieve the online tracking task. The proposed method is evaluated on several datasets, and the experimental results demonstrate its effectiveness. Our method can also be used for other autonomous driving applications, such as path planning and trajectory prediction. Sukai Wang, Yuxiang Sun 0002, Zheng Wang 0002, Ming Liu 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2024 | Learning Semantic Alignment Using Global Features and Multi-Scale ConfidenceabstractSemantic alignment aims to establish pixel correspondences between images based on semantic consistency. It can serve as a fundamental component for various downstream computer vision tasks, such as style transfer and exemplar-based colorization, etc. Many existing methods use local features and their cosine similarities to infer semantic alignment. However, they struggle with significant intra-class variation of objects, such as appearance, size, etc. In other words, contents with the same semantics tend to be significantly different in vision. To address this issue, we propose a novel deep neural network of which the core lies in global feature enhancement and adaptive multi-scale inference. Specifically, two modules are proposed: an enhancement transformer for enhancing semantic features with global awareness; a probabilistic correlation module for adaptively fusing multi-scale information based on the learned confidence scores. We use the unified network architecture to achieve two types of semantic alignment, namely, cross-object semantic alignment and cross-domain semantic alignment. Experimental results demonstrate that our method achieves competitive performance on five standard cross-object semantic alignment benchmarks, and outperforms the state of the arts in cross-domain semantic alignment. Huaiyuan Xu, Jing Liao 0001, Huaping Liu 0001, Yuxiang Sun 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Multimodal-XAD: Explainable Autonomous Driving Based on Multimodal Environment DescriptionsabstractIn recent years, deep learning-based end-to-end autonomous driving has become increasingly popular. However, deep neural networks are like black boxes. Their outputs are generally not explainable, making them not reliable to be used in real-world environments. To provide a solution to this problem, we propose an explainable deep neural network that jointly predicts driving actions and multimodal environment descriptions of traffic scenes, including bird-eye-view (BEV) maps and natural-language environment descriptions. In this network, both the context information from BEV perception and the local information from semantic perception are considered before producing the driving actions and natural-language environment descriptions. To evaluate our network, we build a new dataset with hand-labelled ground truth for driving actions and multimodal environment descriptions. Experimental results show that the combination of context information and local information enhances the prediction performance of driving action and environment description, thereby improving the safety and explainability of our end-to-end autonomous driving network. Yuchao Feng, Wei Hua 0002, Yuxiang Sun 0002 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | SLAMesh: Real-time LiDAR Simultaneous Localization and MeshingabstractMost current LiDAR simultaneous localization and mapping (SLAM) systems build maps in point clouds, which are sparse when zoomed in, even though they seem dense to human eyes. Dense maps are essential for robotic applications, such as map-based navigation. Due to the low memory cost, mesh has become an attractive dense model for mapping in recent years. However, existing methods usually produce mesh maps by using an offline post-processing step to generate mesh maps. This two-step pipeline does not allow these methods to use the built mesh maps online and to enable localization and meshing to benefit each other. To solve this problem, we propose the first CPU-only real-time LiDAR SLAM system that can simultaneously build a mesh map and perform localization against the mesh map. A novel and direct meshing strategy with Gaussian process reconstruction realizes the fast building, registration, and updating of mesh maps. We perform experiments on several public datasets. The results show that our SLAM system can run at around 40Hz. The localization and meshing accuracy also outperforms the state-of-the-art methods, including the TSDF map and Poisson reconstruction. Our code and video demos are available at: https://github.com/lab-sun/SLAMesh. Jianyuan Ruan, Yuxiang Sun 0002 |
ICRA | 4 |
| 2023 | CenterLineDet: CenterLine Graph Detection for Road Lanes with Vehicle-mounted Sensors by Transformer for HD Map GenerationabstractWith the fast development of autonomous driving technologies, there is an increasing demand for high-definition (HD) maps, which provide reliable and robust prior information about the static part of the traffic environments. As one of the important elements in HD maps, road lane centerline is critical for downstream tasks, such as prediction and planning. Manually annotating centerlines for road lanes in HD maps is labor-intensive, expensive and inefficient, severely restricting the wide applications of autonomous driving systems. Previous work seldom explores the lane centerline detection problem due to the complicated topology and severe overlapping issues of lane centerlines. In this paper, we propose a novel method named CenterLineDet to detect lane centerlines for automatic HD map generation. Our CenterLineDet is trained by imitation learning and can effectively detect the graph of centerlines with vehicle-mounted sensors (i.e., six cameras and one LiDAR) through iterations. Due to the use of the DETR-like transformer network, CenterLineDet can handle complicated graph topology, such as lane intersections. The proposed approach is evaluated on the large-scale public dataset NuScenes. The superiority of our CenterLineDet is demonstrated by the comparative results. Our code, supplementary materials, and video demonstrations are available at https://tonyxuqaq.github.io/projects/CenterLineDet/. Zhenhua Xu 0003, Yuxuan Liu 0008, Yuxiang Sun 0002, Ming Liu 0001, Lujia Wang 0001 |
ICRA | 3 |
| 2023 | Adaptive-Mask Fusion Network for Segmentation of Drivable Road and Negative Obstacle With Untrustworthy FeaturesabstractSegmentation of drivable roads and negative obstacles is critical to the safe driving of autonomous vehicles. Currently, many multi-modal fusion methods have been proposed to improve segmentation accuracy, such as fusing RGB and depth images. However, we find that when fusing two modals of data with untrustworthy features, the performance of multi-modal networks could be degraded, even lower than those using a single modality. In this paper, the untrustworthy features refer to those extracted from regions (e.g., far objects that are beyond the depth measurement range) with invalid depth data (i.e., 0 pixel value) in depth images. The untrustworthy features can confuse the segmentation results, and hence lead to inferior results. To provide a solution to this issue, we propose the adaptive-mask fusion Network (AMFNet) by introducing adaptive-weight masks in the fusion module to fuse features from RGB and depth images with inconsistency. In addition, we release a large-scale RGB-depth dataset with manually-labeled ground truth based on the NPO dataset for drivable roads and negative obstacles segmentation. Extensive experimental results demonstrate that our network achieves state-of-the-art performance compared with other networks. Our code and dataset are available at: https://github.com/lab-sun/AMFNet. Yuchao Feng, Yanning Guo, Yuxiang Sun 0002 |
IV | 4 |
| 2023 | A Task-Driven Scene-Aware LiDAR Point Cloud Coding Framework for Autonomous VehiclesabstractLiDAR sensors are almost indispensable for autonomous robots to perceive the surrounding environment. However, the transmission of large-scale LiDAR point clouds is highly bandwidth-intensive, which can easily lead to transmission problems, especially for unstable communication networks. Meanwhile, existing LiDAR data compression is mainly based on rate-distortion optimization, which ignores the semantic information of ordered point clouds and the task requirements of autonomous robots. To address these challenges, this article presents a task-driven Scene-Aware LiDAR Point Clouds Coding (SA-LPCC) framework for autonomous vehicles. Specifically, a semantic segmentation model is developed based on multidimension information, in which both 2-D texture and 3-D topology information are fully utilized to segment movable objects. Furthermore, a prediction-based deep network is explored to remove the spatial–temporal redundancy. The experimental results on the benchmark semantic KITTI dataset validate that our SA-LPCC achieves state-of-the-art performance in terms of the reconstruction quality and storage space for downstream tasks. We believe that SA-LPCC jointly considers the scene-aware characteristics of movable objects and removes the spatial–temporal redundancy from an end-to-end learning mechanism, which will boost the related applications from algorithm optimization to industrial products. Xuebin Sun, Miaohui Wang, Jingxin Du, Yuxiang Sun 0002, Shing Shin Cheng, Wuyuan Xie |
IEEE Trans. Ind. Informatics | 4 |
| 2023 | NLE-DM: Natural-Language Explanations for Decision Making of Autonomous Driving Based on Semantic Scene UnderstandingabstractIn recent years, the advancement of deep-learning technologies has greatly promoted the research progress of autonomous driving. However, deep neural network is like a black box. Given a specific input, it is difficult to explain the output of the network. Without explainable results, it would be unsafe to deploy deep networks in unseen environments or environments with potential unexpected situations. Especially for decision-making networks, inappropriate outputs could lead to severe traffic accidents. To provide a solution to this problem, we propose a deep neural network that jointly predicts the decision-making actions and corresponding natural-language explanations based on semantic scene understanding. Two types of explanations, the reasons of driving actions and the surrounding environment descriptions of the ego-vehicle, are designed. Both the reasons and descriptions are in the form of natural language. The decision-making actions could be explained by the corresponding reasons or the environment descriptions. We also release a large-scale dataset with hand-labelled ground truth including driving actions and environment descriptions. The superiority of our network over other methods is demonstrated on both our dataset and a public dataset. Yuchao Feng, Wei Hua 0002, Yuxiang Sun 0002 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | A Novel Inertial-Aided Visible Light Positioning System Using Modulated LEDs and Unmodulated Lights as LandmarksabstractIndoor localization with high accuracy and efficiency has attracted much attention. Due to visible light communication (VLC), the LED lights in buildings, once modulated, hold great potential to be ubiquitous indoor localization infrastructure. However, this entails retrofitting the lighting system and is hence costly in wide adoption. To alleviate this problem, we propose to exploit modulated LEDs and existing unmodulated lights as landmarks. On this basis, we present a novel inertial-aided visible light positioning (VLP) system for lightweight indoor localization on resource-constrained platforms, such as service robots and mobile devices. With blob detection, tracking, and VLC decoding on rolling-shutter camera images, a visual front end extracts two types of blob features, i.e., mapped landmarks (MLs) and opportunistic features (OFs). These are tightly fused with inertial measurements in a stochastic cloning sliding-window extended Kalman filter (EKF) for localization. We evaluate the system by extensive experiments. The results show that it can provide lightweight, accurate, and robust global pose estimates in real time. Compared with our previous ML-only inertial-aided VLP solution, the proposed system has superior performance in terms of positional accuracy and robustness under challenging light configurations, such as sparse ML/OF distribution. Note to Practitioners—This article is motivated by the problem that many existing visible light positioning (VLP) systems require high-cost environmental modifications, i.e., replacing a large portion of original lights with modulated LEDs as beacons. To reduce costs in wide adoption, we seek to use fewer modulated LEDs if possible. Accordingly, we present a novel inertial-aided VLP system that uses both modulated LEDs and unmodulated lights as landmarks. Like in other VLP systems, the successfully decoded LEDs provide absolute pose measurements for global localization. Unmodulated lights and the LEDs with decoding failures provide relative motion constraints, allowing the reduction of pose drift during the outage of modulated LEDs. Due to the tightly coupled sensor fusion by filtering, the system can provide efficient and accurate localization when modulated LEDs are sparse. The system is lightweight to run on resource-constrained platforms. For practical deployment of our system at scale, creating LED maps accurately and efficiently remains a problem. It is desired to develop automated LED mapping solutions in future work. Yuxiang Sun 0002, Lujia Wang 0001, Ming Liu 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2022 | A Novel Coding Scheme for Large-Scale Point Cloud Sequences Based on Clustering and RegistrationabstractDue to the huge volume of point cloud data, storing and transmitting it is currently difficult and expensive in autonomous driving. Learning from the high-efficiency video coding (HEVC) framework, we propose a novel compression scheme for large-scale point cloud sequences, in which several techniques have been developed to remove the spatial and temporal redundancy. The proposed strategy consists mainly of three parts: intracoding, intercoding, and residual data coding. For intracoding, inspired by the depth modeling modes (DMMs), in 3-D HEVC (3-D-HEVC), a cluster-based prediction method is proposed to remove the spatial redundancy. For intercoding, a point cloud registration algorithm is utilized to transform two adjacent point clouds into the same coordinate system. By calculating the residual map of their corresponding depth image, the temporal redundancy can be removed. Finally, the residual data are compressed either by lossless or lossy methods. Our approach can deal with multiple types of point cloud data, from simple to more complex. The lossless method can compress the point cloud data to 3.63% of its original size by intracoding and 2.99% by intercoding without distance distortion. Experiments on the KITTI dataset also demonstrate that our method yields better performance compared with recent well-known methods.Note to Practitioners—This article deals with the problem of efficient compression of point cloud sequences that come from light detection and ranging (LiDARs) mounted on autonomous mobile robots. The vast amount of point cloud data could be an important bottleneck for transmission and storage. Inspired by the HEVC algorithm, we develop a novel coding architecture for the point cloud sequence. The scans are divided into intraframe and interframe, which are encoded separately using different techniques. Our method can be used for the compression of LiDAR point cloud sequences or dense LiDAR point cloud map and will significantly reduce the transmission bandwidth and storage spaces. We have to admit that although our method is less effective for real-time solutions, it can be highly efficient for off-line applications. Future studies will concentrate on further optimizing the coding algorithm to reduce the computational complexity and trying to find a balance between them. Xuebin Sun, Yuxiang Sun 0002, Weixun Zuo, Shing Shin Cheng, Ming Liu 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2022 | Dynamic Fusion Module Evolves Drivable Area and Road Anomaly Detection: A Benchmark and AlgorithmsabstractJoint detection of drivable areas and road anomalies is very important for mobile robots. Recently, many semantic segmentation approaches based on convolutional neural networks (CNNs) have been proposed for pixelwise drivable area and road anomaly detection. In addition, some benchmark datasets, such as KITTI and Cityscapes, have been widely used. However, the existing benchmarks are mostly designed for self-driving cars. There lacks a benchmark for ground mobile robots, such as robotic wheelchairs. Therefore, in this article, we first build a drivable area and road anomaly detection benchmark for ground mobile robots, evaluating existing state-of-the-art (SOTA) single-modal and data-fusion semantic segmentation CNNs using six modalities of visual features. Furthermore, we propose a novel module, referred to as the dynamic fusion module (DFM), which can be easily deployed in existing data-fusion networks to fuse different types of visual features effectively and efficiently. The experimental results show that the transformed disparity image is the most informative visual feature and the proposed DFM-RTFNet outperforms the SOTAs. In addition, our DFM-RTFNet achieves competitive performance on the KITTI road benchmark. Hengli Wang, Rui Fan 0001, Yuxiang Sun 0002, Ming Liu 0001 |
IEEE Trans. Cybern. | 3 |
| 2022 | RNGDet: Road Network Graph Detection by Transformer in Aerial ImagesabstractRoad network graphs provide critical information for autonomous-vehicle applications, such as drivable areas that can be used for motion planning algorithms. To find road network graphs, manual annotation is usually inefficient and labor-intensive. Automatically detecting road network graphs could alleviate this issue, but existing works still have some limitations. For example, segmentation-based approaches could not ensure satisfactory topology correctness, and graph-based approaches could not present precise enough detection results. To provide a solution to these problems, we propose a novel approach based on transformer and imitation learning in this article. In view of that high-resolution aerial images could be easily accessed all over the world nowadays, we make use of aerial images in our approach. Taken as input an aerial image, our approach iteratively generates road network graphs vertex-by-vertex. Our approach can handle complicated intersection points with various numbers of incident road segments. We evaluate our approach on a publicly available dataset. The superiority of our approach is demonstrated through comparative experiments. Our work is accompanied by a demonstration video which is available athttps://tonyxuqaq.github.io/projects/RNGDet/. Zhenhua Xu 0003, Yuxuan Liu 0008, Lu Gan 0001, Yuxiang Sun 0002, Xinyu Wu 0001, Ming Liu 0001, Lujia Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | DQ-GAT: Towards Safe and Efficient Autonomous Driving With Deep Q-Learning and Graph Attention NetworksabstractAutonomous driving in multi-agent dynamic traffic scenarios is challenging: the behaviors of road users are uncertain and are hard to model explicitly, and the ego-vehicle should apply complicated negotiation skills with them, such as yielding, merging and taking turns, to achieve both safe and efficient driving in various settings. Traditional planning methods are largely rule-based and scale poorly in these complex dynamic scenarios, often leading to reactive or even overly conservative behaviors. Therefore, they require tedious human efforts to maintain workability. Recently, deep learning-based methods have shown promising results with better generalization capability but less hand engineering efforts. However, they are either implemented with supervised imitation learning (IL), which suffers from dataset bias and distribution mismatch issues, or are trained with deep reinforcement learning (DRL) but focus on one specific traffic scenario. In this work, we propose DQ-GAT to achieve scalable and proactive autonomous driving, where graph attention-based networks are used to implicitly model interactions, and deep Q-learning is employed to train the network end-to-end in an unsupervised manner. Extensive experiments in a high-fidelity driving simulator show that our method achieves higher success rates than previous learning-based methods and a traditional rule-based method, and better trades off safety and efficiency in both seen and unseen scenarios. Moreover, qualitative results on a trajectory dataset indicate that our learned policy can be transferred to the real world for practical applications with real-time speeds. Demonstration videos are available athttps://caipeide.github.io/dq-gat/. Peide Cai, Hengli Wang, Yuxiang Sun 0002, Ming Liu 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | Learning Interpretable End-to-End Vision-Based Motion Planning for Autonomous Driving with Optical Flow DistillationabstractRecently, deep-learning based approaches have achieved impressive performance for autonomous driving. However, end-to-end vision-based methods typically have limited interpretability, making the behaviors of the deep networks difficult to explain. Hence, their potential applications could be limited in practice. To address this problem, we propose an interpretable end-to-end vision-based motion planning approach for autonomous driving, referred to as IVMP. Given a set of past surrounding-view images, our IVMP first predicts future egocentric semantic maps in bird’s-eye-view space, which are then employed to plan trajectories for self-driving vehicles. The predicted future semantic maps not only provide useful interpretable information, but also allow our motion planning module to handle objects with low probability, thus improving the safety of autonomous driving. Moreover, we also develop an optical flow distillation paradigm, which can effectively enhance the network while still maintaining its real-time performance. Extensive experiments on the nuScenes dataset and closed-loop simulation show that our IVMP significantly outperforms the state-of-the-art approaches in imitating human drivers with a much higher success rate. Our project page is available at https://sites.google.com/view/ivmp. Hengli Wang, Peide Cai, Yuxiang Sun 0002, Lujia Wang 0001, Ming Liu 0001 |
ICRA | 3 |
| 2021 | S2P2: Self-Supervised Goal-Directed Path Planning Using RGB-D Data for Robotic WheelchairsabstractPath planning is a fundamental capability for autonomous navigation of robotic wheelchairs. With the impressive development of deep-learning technologies, imitation learning-based path planning approaches have achieved effective results in recent years. However, the disadvantages of these approaches are twofold: 1) they may need extensive time and labor to record expert demonstrations as training data; and 2) existing approaches could only receive high-level commands, such as turning left/right. These commands could be less sufficient for the navigation of mobile robots (e.g., robotic wheelchairs), which usually require exact poses of goals. We contribute a solution to this problem by proposing S2P2, a self-supervised goal-directed path planning approach. Specifically, we develop a pipeline to automatically generate planned path labels given as input RGB-D images and poses of goals. Then, we present a best-fit regression plane loss to train our data-driven path planning model based on the generated labels. Our S2P2 does not need pre-built maps, but it can be integrated into existing map-based navigation systems through our framework. Experimental results show that our S2P2 outperforms traditional path planning algorithms, and increases the robustness of existing map-based navigation systems. Our project page is available at https://sites.google.com/view/s2p2. Hengli Wang, Yuxiang Sun 0002, Rui Fan 0001, Ming Liu 0001 |
ICRA | 2 |
| 2021 | DiGNet: Learning Scalable Self-Driving Policies for Generic Traffic Scenarios with Graph Neural NetworksabstractTraditional decision and planning frameworks for self-driving vehicles (SDVs) scale poorly in new scenarios, thus they require tedious hand-tuning of rules and parameters to maintain acceptable performance in all foreseeable cases. Recently, self-driving methods based on deep learning have shown promising results with better generalization capability but less hand engineering effort. However, most of the previous learning-based methods are trained and evaluated in limited driving scenarios with scattered tasks, such as lane-following, autonomous braking, and conditional driving. In this paper, we propose a graph-based deep network to achieve scalable self-driving that can handle massive traffic scenarios. Specifically, more than 7,000 km of evaluation is conducted in a high-fidelity driving simulator, in which our method can obey the traffic rules and safely navigate the vehicle in a large variety of urban, rural, and highway environments, including unprotected left turns, narrow roads, roundabouts, and pedestrian-rich intersections. Demonstration videos are available at https: //caipeide.github.io/dignet/. Peide Cai, Hengli Wang, Yuxiang Sun 0002, Ming Liu 0001 |
IROS | 3 |
| 2021 | CP-loss: Connectivity-preserving Loss for Road Curb Detection in Autonomous Driving with Aerial ImagesabstractRoad curb detection is important for autonomous driving. It can be used to determine road boundaries to constrain vehicles on roads, so that potential accidents could be avoided. Most of the current methods detect road curbs online using vehicle-mounted sensors, such as cameras or 3-D Lidars. However, these methods usually suffer from severe occlusion issues. Especially in highly-dynamic traffic environments, most of the field of view is occupied by dynamic objects. To alleviate this issue, we detect road curbs offline using high-resolution aerial images in this paper. Moreover, the detected road curbs can be used to create high-definition (HD) maps for autonomous vehicles. Specifically, we first predict the pixel-wise segmentation map of road curbs, and then conduct a series of post-processing steps to extract the graph structure of road curbs. To tackle the disconnectivity issue in the segmentation maps, we propose an innovative connectivity-preserving loss (CP-loss) to improve the segmentation performance. The experimental results on a public dataset demonstrate the effectiveness of our proposed loss function. This paper is accompanied with a demonstration video and a supplementary document, which are available at https://sites.google.com/view/cp-loss. Zhenhua Xu 0003, Yuxiang Sun 0002, Lujia Wang 0001, Ming Liu 0001 |
IROS | 2 |
| 2021 | Consensus-Based Cooperative Formation Guidance Strategy for Multiparafoil Airdrop SystemsabstractParafoil airdrop is an important way to deliver goods and materials to area where road vehicles are not easy to reach. However, it is difficult to deliver large quantities of goods and materials to a given location with only one parafoil. Airdropping multiple parafoils is an effective choice for transporting large quantities of goods and materials. To realize the cooperative airdrop of multiple parafoils, a cooperative guidance framework is proposed. First, a trajectory planning algorithm is designed to plan the multiphase trajectory for the parafoil group. Then, a trajectory tracking algorithm is developed for the pilot parafoil in the parafoil group to reliably follow the planned trajectory. Finally, a cooperative formation guidance strategy is designed based on the leader–follower consensus theory. Under this strategy, the position and speed of the follower parafoil can be consensus with those of the leader parafoil. Lyapunov’s theorem proves the stability of this strategy. We evaluate the effectiveness of this framework through simulations. The results demonstrate that our algorithms can realize the precise airdrop of massive goods and materials with upwind landing using multiple parafoils. In addition, the parafoils could be gradually gathered to a desired formation, and safe distances could be maintained between parafoils during the airdrop process.Note to Practitioners—This article was motivated by the problem of airdropping massive goods and materials. Existing methods usually adopt a single heavy parafoil, or use centralized multiparafoil systems. Both these methods have their limitations. For the former, there is an upper limit of the load capacity for a single parafoil. For the latter, the parafoils in the centralized system lack fully autonomous ability. Distributed multiparafoil systems could solve the problem effectively. However, compared to single-parafoil systems, there are still some challenges, for example, the multiparafoil gathering, collision avoidance and cooperative formation, as well as the upwind landing. Fortunately, existing parafoils are equipped with sensors, communication, and control devices, so they could be viewed as agents with autonomous capabilities. In this article, a formation guidance framework for multiple autonomous parafoils is proposed. First, we plan a trajectory for the pilot parafoil. Then, we show how to effectively track the planned trajectory. Finally, we demonstrate how multiple parafoils could coordinate with each other to accomplish airdrop tasks. The simulation results confirm the feasibility of this strategy. Yuxiang Sun 0002, Min Zhao 0011, Ming Liu 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2021 | FuseSeg: Semantic Segmentation of Urban Scenes Based on RGB and Thermal Data FusionabstractSemantic segmentation of urban scenes is an essential component in various applications of autonomous driving. It makes great progress with the rise of deep learning technologies. Most of the current semantic segmentation networks use single-modal sensory data, which are usually the RGB images produced by visible cameras. However, the segmentation performance of these networks is prone to be degraded when lighting conditions are not satisfied, such as dim light or darkness. We find that thermal images produced by thermal imaging cameras are robust to challenging lighting conditions. Therefore, in this article, we propose a novel RGB and thermal data fusion network named FuseSeg to achieve superior performance of semantic segmentation in urban scenes. The experimental results demonstrate that our network outperforms the state-of-the-art networks.Note to Practitioners—This article investigates the problem of semantic segmentation of urban scenes when lighting conditions are not satisfied. We provide a solution to this problem via information fusion with RGB and thermal data. We build an end-to-end deep neural network, which takes as input a pair of RGB and thermal images and outputs pixel-wise semantic labels. Our network could be used for urban scene understanding, which serves as a fundamental component of many autonomous driving tasks, such as environment modeling, obstacle avoidance, motion prediction, and planning. Moreover, the simple design of our network allows it to be easily implemented using various deep learning frameworks, which facilitates the applications on different hardware or software platforms. Yuxiang Sun 0002, Weixun Zuo, Peng Yun, Hengli Wang, Ming Liu 0001 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2020 | Monocular Visual Odometry using Learned Repeatability and DescriptionabstractRobustness and accuracy for monocular visual odometry (VO) under challenging environments are widely concerned. In this paper, we present a monocular VO system leveraging learned repeatability and description. In a hybrid scheme, the camera pose is initially tracked on the predicted repeatability maps in a direct manner and then refined with the patch-wise 3D-2D association. The local feature parameterization and the adapted mapping module further boost different functionalities in the system. Extensive evaluations on challenging public datasets are performed. The competitive performance on camera pose estimation demonstrates the effectiveness of our method. Additional studies on the local reconstruction accuracy and running time exhibit that our system is capable of maintaining a robust and lightweight backend. Huaiyang Huang, Haoyang Ye, Yuxiang Sun 0002, Ming Liu 0001 |
ICRA | 3 |
| 2020 | Applying Surface Normal Information in Drivable Area and Road Anomaly Detection for Ground Mobile RobotsabstractThe joint detection of drivable areas and road anomalies is a crucial task for ground mobile robots. In recent years, many impressive semantic segmentation networks, which can be used for pixel-level drivable area and road anomaly detection, have been developed. However, the detection accuracy still needs improvement. Therefore, we develop a novel module named the Normal Inference Module (NIM), which can generate surface normal information from dense depth images with high accuracy and efficiency. Our NIM can be deployed in existing convolutional neural networks (CNNs) to refine the segmentation performance. To evaluate the effectiveness and robustness of our NIM, we embed it in twelve state-of-the-art CNNs. The experimental results illustrate that our NIM can greatly improve the performance of the CNNs for drivable area and road anomaly detection. Furthermore, our proposed NIM-RTFNet ranks 8th on the KITTI road benchmark and exhibits a real-time inference speed. Hengli Wang, Rui Fan 0001, Yuxiang Sun 0002, Ming Liu 0001 |
IROS | 3 |
| 2019 | Metric Monocular Localization Using Signed Distance FieldsabstractMetric localization plays a critical role in vision-based navigation. For overcoming the degradation of matching photometry under appearance changes, recent research resorted to introducing geometry constraints of the prior scene structure. In this paper, we present a metric localization method for the monocular camera, using the Signed Distance Field (SDF) as a global map representation. Leveraging the volumetric distance information from SDFs, we aim to relax the assumption of an accurate structure from the local Bundle Adjustment (BA) in previous methods. By tightly coupling the distance factor with temporal visual constraints, our system corrects the odometry drift and jointly optimizes global camera poses with the local structure. We validate the proposed approach on both indoor and outdoor public datasets. Compared to the state-of-the-art methods, it achieves a comparable performance with a minimal sensor configuration. Huaiyang Huang, Yuxiang Sun 0002, Haoyang Ye, Ming Liu 0001 |
IROS | 2 |
| 2019 | Active Perception for Foreground Segmentation: An RGB-D Data-Based Background Modeling MethodabstractForeground moving object segmentation is a fundamental problem in many computer vision applications. As a solution for foreground segmentation, background modeling has been intensively studied over past years and many effective algorithms have been developed. However, accurate foreground segmentation is still a difficult problem. Currently, most of the algorithms work solely within the color space, in which the segmentation performance is prone to be degraded by a multitude of challenges, such as illumination changes, shadows, automatic camera adjustments, and color camouflage. RGB-D cameras are active visual sensors that provide depth measurements along with color images. We present in this paper an innovative background modeling method by using both the color and depth information from an RGB-D camera. The proposed method is evaluated using a public RGB-D data set. Various experiments confirm that our method is able to achieve superior performance compared with existing well-known methods. Yuxiang Sun 0002, Ming Liu 0001, Max Q.-H. Meng |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2019 | Autonomous Robotic Exploration by Incremental Road Map ConstructionabstractIn this paper, we propose a novel path planning framework for autonomous exploration in unknown environments using a mobile robot. A graph structure is incrementally constructed along with the exploration process. The structure is the road map that represents the topology of the explored environment. To construct the road map, we design a sampling strategy to get random points in the explored environment uniformly. A global path from the current location of the robot to the target area can be found on this road map efficiently. We utilize a lazy collision checking method that only checks the feasibility of the generated global path to improve the planning efficiency. The feasible global path is further optimized with our proposed trajectory optimization method considering the motion constraints of the robot. This mechanism can facilitate the path cost evaluation for the next best view selection. In order to select the next best target region, we propose a utility function that takes into account both the path cost and the information gain of a candidate target region. Moreover, we present a target reselection mechanism to evaluate the target region and reduce the extra path cost. The efficiency and effectiveness of our approach are demonstrated using a mobile robot in both simulation and real experimental studies. Chaoqun Wang 0009, Wenzheng Chi, Yuxiang Sun 0002, Max Q.-H. Meng |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2018 | A Locomotion Recognition System Using Depth ImagesabstractPowered lower-limb orthoses and prostheses are attracting an increasing amount of attention in assisting daily living activities. To safely and naturally collaborate with human users, the key technology relies on an intelligent controller to accurately decode users' movement intention. In this work, we proposed an innovative locomotion recognition system based on depth images. Composed of a feature extraction subsystem and a finite-state-machine based recognition subsystem, the proposed approach is capable of capturing both the limb movements and the terrains right in front of the user. This makes it possible to anticipate the detection of locomotion modes, especially at transition states, thus enabling the associated wearable robot to deliver a smooth and seamless assistance. Validation experiments were implemented with nine subjects to trace a track that comprised of standing, walking, stair ascending, and stair descending, for three rounds each. The results showed that in steady state, the proposed system could recognize all four locomotion tasks with approximate 100% of accuracy. Out of 216 mode transitions, 82.4% of the intended locomotion tasks can be detected before the transition happened. Thanks to its high accuracy and promising prediction performance, the proposed locomotion recognition system is expected to significantly improve the safety as well as the effectiveness of a lower-limb assistive device. Tingfang Yan, Yuxiang Sun 0002, Chi-Hong Cheung, Max Q.-H. Meng |
ICRA | 2 |