Ming Yang 0002

dblp:98/2604-2 · DBLP profile ↗
← Back
90ranked-venue papers
1as first author
53since 2021 · last 2026
0000-0002-8679-9137ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 44 · 1 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 34 · 23 since 2021Systems, architecture and hardware · 19 · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 8 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DrivingEditor: 4D Composite Gaussian Splatting for Reconstruction and Edition of Dynamic Autonomous Driving Scenes
abstract
In recent years, with the development of autonomous driving, 3D reconstruction for unbounded large-scale scenes has attracted researchers' attention. Existing methods have achieved outstanding reconstruction accuracy in autonomous driving scenes, but most of them lack the ability to edit scenes. Although some methods have the capability to edit scenarios, they are highly dependent on manually annotated 3D bounding boxes, leading to their poor scalability. To address the issues, we introduce a new Gaussian representation, called DrivingEditor, which decouples the scene into two parts and handles them by separate branches to individually model the dynamic foreground objects and the static background during the training process. By proposing a framework for decoupled modeling of scenarios, we can achieve accurate editing of any dynamic target, such as dynamic objects removal, adding and etc, meanwhile improving the reconstruction quality of autonomous driving scenes especially the dynamic foreground objects, without resorting to 3D bounding boxes. Extensive experiments on Waymo Open Dataset and KITTI benchmarks demonstrate the performance in 3D reconstruction for both dynamic and static scenes. Besides, we conduct extra experiments on unstructured large-scale scenarios, which can more convincingly demonstrate the performance and robustness of our proposed model when rendering the unstructured scenes. Our code is available at https://github.com/WangXu-xxx/DrivingEditor.
Yeqiang Qian, Yun-Fu Liu, Lei Tuo, Huiyong Chen, Ming Yang 0002
IEEE Trans. Image Process.6
2026 An Automatic Calibration Method for Low-Overlapped Roadside RGB-D Camera Network
abstract
Roadside multi-sensor networks, such as RGB-D camera networks, play a crucial role in Intelligent Transportation Systems (ITS). These systems rely on accurate extrinsic parameters (i.e., the relative positions and orientations) of each camera in the network. However, achieving fast and accurate large-scale extrinsic parameter calibration is challenging, especially when the overlap between camera views is limited due to cost constraints. To address this issue, we propose an automated, scalable, and marker-free calibration method that requires no human intervention. Our approach leverages dynamic rigid bodies, such as moving vehicles, as bridges to establish associations among all cameras without manual placement or supervision. The proposed method consists of three main stages: 1) calibrating camera height, roll, and pitch using the ground plane; 2) estimating each camera’s 2D position and yaw angle based on trajectories; and 3) refining these estimates by matching features on the road surface using SuperPoint and SuperGlue. Real-world experiments involving 58 roadside RGB-D cameras deployed in a parking lot demonstrate that our method significantly improves calibration efficiency while maintaining high accuracy, making it well-suited for large-scale RGB-D camera network deployment.
Jingda Chen, Hanyang Zhuang, Yuesheng He, Ming Yang 0002
IEEE Trans. Intell. Transp. Syst.5
2026 Learning-Based 3D Reconstruction in Autonomous Driving: A Comprehensive Survey
abstract
Learning-based 3D reconstruction has emerged as a transformative technique in autonomous driving, enabling precise modeling of environments through advanced neural representations. It has inspired pioneering solutions for vital tasks in autonomous driving, such as dense mapping and closed-loop simulation, as well as comprehensive scene feature for driving scene understanding and reasoning. Given the rapid growth in related research, this survey provides a comprehensive review of both technical evolutions and practical applications in autonomous driving. We begin with an introduction to the preliminaries of learning-based 3D reconstruction to provide a solid technical background foundation, then progress to a rigorous, multi-dimensional examination of cutting-edge methodologies, systematically organized according to the distinctive technical requirements and fundamental challenges of autonomous driving. Through analyzing and summarizing development trends and cutting-edge research, we identify existing technical challenges, along with insufficient disclosure of on-board validation and safety verification details in the current literature, and ultimately suggest potential directions to guide future studies.
Liewen Liao, Weihao Yan 0001, Ming Yang 0002, Songan Zhang, H. Eric Tseng
IEEE Trans. Intell. Transp. Syst.4
2026 ParkOcc: A Novel Dataset and Benchmark for Surround-View Fisheye 3-D Semantic Occupancy Prediction in Automated Parking Scenarios
Yeqiang Qian, Ming Yang 0002
IEEE Trans. Intell. Transp. Syst.4
2026 APC-Scheduler: Allocation-Planning-Charging Integrated Scheduling System for Cooperative Automated Valet Parking
abstract
The Automated Valet Parking system (AVP), as one of the promising technologies, offers significant benefits in saving maneuver time and parking cost. With the increase number of vehicles using AVP system in large-scale parking lots, the overall efficiency is limited by the selfish decision-making of each individual; therefore, cooperative-AVP (C-AVP) is developed to achieve global optimization by scheduling the vehicles. Existing C-AVP methods focus only on one of the key processes in AVP, such as space allocation, trajectory planning, and electric vehicle(EV) charging. However, these factors interact with each other and are rarely considered in an entire system. Therefore, this paper aims to model the space allocation, trajectory planning, and EV charging problem in a whole framework, named APC-scheduler, and adopt it to improve the overall efficiency for large-scale parking lot with EV. The APC-scheduler comprises two parts: 1) a zone-based parking space allocation module dynamically assigns parking spaces based on real-time conditions using a hierarchical optimization strategy. It contains a charging priority estimation process to determine the EV’s charging requirement urgency; 2) a conflict-based trajectory planning module is developed to reduce trajectory overlaps and vehicle conflicts. It uses spatial-temporal path planner and speed planner to eliminate conflicts. Experiments under reasonable vehicle arrival and departure statistics have been conducted in a large-scale parking lot with 526 parking spaces. The results demonstrate that the proposed method effectively enhances parking efficiency and minimizes conflicts, particularly under high vehicle arrival frequencies and in dense traffic conditions.
Hanyang Zhuang, Qizhe Xu, Ming Yang 0002
IEEE Trans. Intell. Transp. Syst.5
2026 MPSS: A Model Pruning Method for Semantic Image Segmentation Networks
abstract
This paper proposes a model pruning named MPSS for semantic image segmentation networks, so that semantic image segmentation models can be deployed into embedded devices. Most existing model pruning methods aim at image classification models. Since semantic segmentation is a fine grained task, the directly use of traditional model pruning methods greatly reduces the model accuracy. The core problems of model pruning in the semantic segmentation task are to determine the appropriate pruning kernels and the pruning structure. We propose a new composite index that defines the similarity between convolution kernels to determine the pruning kernels. Furthermore, we propose a structure mending method based on the neural architecture search to determine the pruning structure. Compared with the method of manually defining the pruning rate, the proposed structure mending method obtains a better pruning structure. We conduct experiments based on two semantic segmentation networks, the FCN and the FASSD-Net. The experimental results show that the proposed model pruning method enables the pruned network to obtain higher accuracy under the same compression rate. In addition, we deploy the compressed models on an embedded platform, and the FASSD Net inference speed is twice as fast as the unpruned model on NVIDIA Xavier NX.
Yeqiang Qian, Qihang Su, Ming Yang 0002
IEEE Trans. Multim.4
2026 Structure-Guided Memory-Efficient 3D Gaussians for Large-Scale Reconstruction
abstract
3D reconstruction is a critical technology with significant implications for applications such as urban planning, autonomous driving, and virtual reality. Recent advances in 3D Gaussian Splatting (3DGS) have demonstrated impressive results in small-scale scenes, achieving high-quality reconstructions with real-time rendering capabilities. However, when applied to large-scale scenes, existing 3DGS methods face significant challenges due to the exponential growth of model size, often exceeding the memory capacity of consumer-grade GPUs and making training and rendering infeasible. In this paper, we propose a structure-guided memory-efficient 3DGS framework that uses only half the memory of current large-scale 3DGS methods while maintaining state-of-the-art reconstruction accuracy. Specifically, we introduce a structure-guided density control mechanism that uses a heuristic approach to split Gaussian ellipsoids in challenging regions and optimizes their attributes during densification, significantly reducing memory storage requirements while preserving structural details with fewer ellipsoids. Moreover, we propose a novel structure loss to supervise the learning of scene structural information, enabling the model to better capture and preserve geometric details such as straight lines and edges, further enhancing reconstruction accuracy. We also propose the largest known drone dataset for 3D reconstruction, comprising over 10,000 high-resolution images covering more than 2.5 million square meters. Extensive experiments on multiple benchmark datasets and our proposed dataset demonstrate that our new method is highly memory-efficient with high accuracy. We strongly recommend you to watch our demo at https://lvzinan.github.io/STGS.github.io/.
Zinan Lv, Yeqiang Qian, Ming Yang 0002
IEEE Trans. Vis. Comput. Graph.4
2025 RL-OGM-Parking: Lidar OGM-Based Hybrid Reinforcement Learning Planner for Autonomous Parking
abstract
Autonomous parking has become a critical application in automatic driving research and development. Parking operations often suffer from limited space and complex environments, requiring accurate perception and precise maneuvering. Traditional rule-based parking algorithms struggle to adapt to diverse and unpredictable conditions, while learning-based algorithms lack consistent and stable performance in various scenarios. Therefore, a hybrid approach is necessary that combines the stability of rule-based methods and the generalizability of learning-based methods. Recently, reinforcement learning (RL) based policy has shown robust capability in planning tasks. However, the simulation-to-reality (sim-to-real) transfer gap seriously blocks the real-world deployment. To address these problems, we employ a hybrid policy, consisting of a rule-based Reeds-Shepp (RS) planner and a learningbased reinforcement learning (RL) planner. A real-time LiDARbased Occupancy Grid Map (OGM) representation is adopted to bridge the sim-to-real gap, leading the hybrid policy can be applied to real-world systems seamlessly. We conducted extensive experiments both in the simulation environment and real-world scenarios, and the result demonstrates that the proposed method outperforms pure rule-based and learningbased methods. The real-world experiment further validates the feasibility and efficiency of the proposed method.
Zhitao Wang, Mingyang Jiang, Tong Qin 0001, Ming Yang 0002
ICRA5
2025 Embodied Escaping: End-to-End Reinforcement Learning for Robot Navigation in Narrow Environment
abstract
Autonomous navigation is a fundamental task for robot vacuum cleaners in indoor environments. Since their core function is to clean entire areas, robots inevitably encounter dead zones in cluttered and narrow scenarios. Existing planning methods often fail to escape due to complex environmental constraints, high-dimensional search spaces, and high difficulty maneuvers. To address these challenges, this paper proposes an embodied escaping model that leverages a reinforcement learning-based policy with an efficient action mask for dead zone escaping. To alleviate the issue of the sparse reward in training, we introduce a hybrid training policy that improves learning efficiency. In handling redundant and ineffective action options, we design a novel action representation to reshape the discrete action space with a uniform turning radius. Furthermore, we develop an action mask strategy to select valid actions quickly, balancing precision and efficiency. In real-world experiments, our robot is equipped with a Lidar, IMU, and two-wheel encoders. Extensive quantitative and qualitative experiments across varying difficulty levels demonstrate that our robot can consistently escape from challenging dead zones. Moreover, our approach significantly outperforms compared path planning and reinforcement learning methods in terms of success rate and collision avoidance. A video showcasing our methodology and real-world demonstrations is available at https://youtu.be/kBaaYWGhNuE.
Mingyang Jiang, Peiyuan Liu, Tong Qin 0001, Ming Yang 0002
IROS7
2025 Pose Tracking of Leading Vehicle Using Mass-Produced Sensors
abstract
Vehicle-following presents significant advantages in flexible scenarios such as platooning and valet parking, where a human-driven leader navigates complex environments, and an autonomous follower replicates the leader's trajectory. The key objective is accurate, robust, long-term pose tracking of the leader. This paper proposes a framework using mass-produced sensors, including a front fisheye camera, front millimeter-wave radar, and wireless communication. The framework consists of three modules: 1) a Visual Feature Tracking module using a DiMP tracker for robust taillight feature tracking; 2) a Visual-radar Fusion module to estimate the rear center of the leader vehicle for reference point tracking; 3) a Leader Pose Tracking module combining sensor data and vehicle-to-vehicle communication in a particle filter framework for precise pose tracking in three degrees of freedom. Real-world experiments show lateral and longitudinal errors of 0.056 meters and 0.165 meters, respectively.
Hanyang Zhuang, Ming Yang 0002
IV4
2025 FP-TTC: Fast Prediction of Time-to-Collision Using Monocular Images
abstract
Time-to-Collision (TTC) is a measure of the time until an object collides with the observation plane which is a critical input indicator for obstacle avoidance and other downstream modules. Previous works have utilized deep neural networks to estimate TTC with monocular cameras in an end-to-end manner, which obtain the state-of-the-art (SOTA) accuracy performance. However, these models usually have deep layers and numerous parameters, resulting in long inference time and high computational overhead. Moreover, existing methods use two frames which are the current and future moments as input to calculate the TTC resulting in a delay during the calculation process. To solve these issues, we propose a novel fast TTC prediction model: FP-TTC. We first use an attention-based scale encoder to model the scale-matching process between images, which significantly reduces the computational overhead as well as improves the model’s accuracy. Meanwhile, a simple but powerful trick is introduced to the model, where we built a time-series decoder and predict the current TTC from RGB images in the past, avoiding the computational delay caused by the system time step interval, and further improved the TTC prediction speed. Our model achieves a parameter reduction of 89.1%, a 5.5-fold increase in inference speed, a 19.3% improvement in accuracy. We also provided a lightweight version of FP-TTC, which further optimized the inference speed and parameter count by 15%. Our code is available athttps://github.com/LChanglin/FP-TTC.
Yeqiang Qian, Songan Zhang, Ming Yang 0002
IEEE Trans. Circuits Syst. Video Technol.5
2025 AdaptiveOcc: Adaptive Octree-Based Network for Multi-Camera 3D Semantic Occupancy Prediction in Autonomous Driving
abstract
Multi-camera 3D semantic occupancy prediction is a critical task for autonomous driving, playing a vital role in understanding the environment. Current methods mainly rely on uniform voxel representation to encode space, which greatly limits their resolution scalability. It causes most existing methods to struggle with scaling to finer granularities, as the cubic growth nature of uniform voxel leads to a significant increase in the demand for computational and storage resources when scaling. To address this, we propose a multi-level hierarchical model AdaptiveOcc. Using the octree structure, our model can adaptively represent different parts of space with varying voxel granularity. It can selectively extend resolution only for a small subset of voxels, thus mitigating the substantial computational and storage burden brought by scaling. To endow our model with adaptability, we propose a distance-adaptive octree construction rule for generating supervised labels. Considering that the voxel granularity requirements vary for different distance ranges in environmental perception, such a construction rule results in a higher likelihood of coarser granularity for distant regions and finer granularity for nearby regions. This ensures a more efficient and rational allocation of computational resources, further reducing the inference latency. Extensive experiments on nuScenes, SemanticKITTI and Waymo dataset validate that our method can scale to finer granularities with faster speed, and less training memory compared with other state-of-the-art methods. Our code is available athttps://github.com/yty-sky/AdaptiveOcc.
Yeqiang Qian, Weihao Yan 0001, Ming Yang 0002
IEEE Trans. Circuits Syst. Video Technol.5
2025 HOPE: A Reinforcement Learning-Based Hybrid Policy Path Planner for Diverse Parking Scenarios
abstract
Automated parking stands as a highly anticipated application of autonomous driving technology. However, existing path planning methodologies fall short of addressing this need due to their incapability to handle the diverse and complex parking scenarios in reality. While non-learning methods provide reliable planning results, they are vulnerable to intricate occasions, whereas learning-based ones are good at exploration but unstable in converging to feasible solutions. To leverage the strengths of both approaches, we introduce Hybrid pOlicy Path plannEr (HOPE). This novel solution integrates a reinforcement learning agent with Reeds-Shepp curves, enabling effective planning across diverse scenarios. HOPE guides the exploration of the reinforcement learning agent by applying an action mask mechanism and employs a transformer to integrate the perceived environmental information with the mask. To facilitate the training and evaluation of the proposed planner, we propose a criterion for categorizing the difficulty level of parking scenarios based on space and obstacle distribution. Experimental results demonstrate that our approach outperforms typical rule-based algorithms and traditional reinforcement learning methods, showing higher planning success rates and generalization across various scenarios. We also conduct real-world experiments to verify the practicability of HOPE. The code for our solution is openly available onhttps://github.com/jiamiya/HOPE.
Mingyang Jiang, Yueyuan Li, Songan Zhang, Ming Yang 0002
IEEE Trans. Intell. Transp. Syst.6
2025 TA-TOS: Terrain-Aware Tiny Obstacle Segmentation Based on MRF Road Modeling Using 3-D LiDAR Scans
abstract
Robust obstacle segmentation remains critical for the safety of intelligent transportation systems (ITS), where LiDAR-based perception systems form the cornerstone of vehicle-environment interaction. Although state-of-the-art (SOTA) LiDAR-based approaches have demonstrated high performance in segmenting common obstacles, the results for tiny obstacle segmentation are still unsatisfactory. However, such tiny obstacles, e.g., curbs, gravel, and potholes, pose significant threats to ground vehicles, undermining ITS operational safety and surface transportation traffic efficiency. It is challenging for SOTA methods to distinguish tiny obstacles due to their inability to precisely model road surfaces, particularly bumpy road surfaces. To address this problem, this paper proposes a road modeling method based on the Markov random field (MRF), possessing stronger road surface modeling capability. A novel negative exponential energy function is introduced to simultaneously ensure the smoothness of the road model and the consistency with the road undulation. After the energy minimization of the MRF, the segmentation of obstacles (including both positive and negative obstacles) is achieved by computing the signed distance to the refined road model. Our proposed terrain-aware tiny obstacle segmentation (TA-TOS) method is compatible with different terrains and different LiDARs, without any prior data or pre-training. We evaluate the performance of TA-TOS on the SemanticKITTI dataset, and two self-built datasets containing tiny obstacles from actual urban mobility systems (road scenarios) and mining haulage systems (off-road scenarios), respectively. Our proposed TA-TOS method achieves much better performance than the SOTA LiDAR-based segmentation approaches, particularly on roads with pronounced undulation. The results show that the improvement is more significant for the segmentation of smaller obstacles. Our source code is publicly available at github.com/ryming2001/TA-TOS.
Nan Ming, Yeqiang Qian, Chunyu Feng, Ming Yang 0002
IEEE Trans. Intell. Transp. Syst.5
2025 Local Vectorized High Definition Map Construction for Autonomous Driving: A Comprehensive Review
abstract
With the advancement of autonomous driving technology, high-definition (HD) maps are crucial for accurate vehicle positioning and safe navigation. Traditional HD map construction relies on offline processing and manual annotation, which are costly and inflexible for dynamic road environments. Local vectorized HD map construction (LV-HDMC) has emerged as a key technology to address these limitations. LV-HDMC uses advanced computer vision techniques to generate map elements from vehicle-mounted sensors in real time, meeting the demands of autonomous driving. This review provides a comprehensive analysis of LV-HDMC research, tracing the evolution of HD map generation and offering an overview of the LV-HDMC task. It explores methods for creating ground truth, including local region acquisition and map element representation, and classifies network structures related to LV-HDMC, emphasizing the role of computer vision in feature extraction and decoding. Evaluation metrics and benchmarks are introduced, comparing the performance of existing methods. The review also discusses future research directions, highlighting how advancements in computer vision could enhance the accuracy and efficiency of LV-HDMC. This review aims to offer valuable insights into LV-HDMC tasks and contribute to the advancement of autonomous driving technology.
Yangrong Zhang, Yeqiang Qian, Hongjun Yi, Ming Yang 0002
IEEE Trans. Intell. Transp. Syst.7
2025 CrossGLoc: Cross-Modal Global Localization Leveraging Pretrained Diffusion Models and Semantic Cues for Intelligent Vehicles
abstract
Cross-modal global localization matches visual information with pre-built LiDAR maps, which has attracted more and more attention for its low cost and potential robustness. However, the inherent modality difference between images and point clouds makes it challenging. This paper proposes a novel cross-modal global localization system, named CrossGLoc, which leverages pre-trained diffusion models and semantic cues to address this challenge. The main idea is leveraging the semantic cues shared between different modalities to bridge the modality gap, and utilizing pre-trained diffusion models to extract modality-consistent high-dimensional features guided by these semantic cues. To achieve this, ControlNet is used to generate intermediate feature maps from semantic images and semantic map projections, and a semantic categories-based feature aggregation algorithm is proposed to aggregate these feature maps into global descriptors. Furthermore, a semantic edge key points-based pose estimation algorithm is proposed to estimate the pose of retrieved image and point cloud pairs. Extensive experiments on the KITTI dataset, the KITTI360 dataset and the self-collected dataset demonstrate that the proposed method achieves state-of-the-art performance in cross-modal global localization.
Hengwang Zhao, Qiyuan Shen, Hanyang Zhuang, Tong Qin 0001, Ming Yang 0002
IEEE Trans. Intell. Transp. Syst.6
2025 AMFD: Distillation via Adaptive Multimodal Fusion for Multispectral Pedestrian Detection
abstract
Multispectral pedestrian detection has been shown to be effective in improving performance in complex illumination scenarios. However, prevalent double-stream networks in multispectral detection employ two separate feature extraction branches for multi-modal data, leading to nearly double the inference time compared to single-stream networks utilizing only one feature extraction branch. This increased inference time has hindered the widespread employment of multispectral pedestrian detection in embedded devices for autonomous systems. To efficiently compress multispectral object detection networks, we propose a novel distillation method, the Adaptive Modal Fusion Distillation (AMFD) framework. Unlike traditional distillation methods, the AMFD framework fully leverages the original modal features from the teacher network, thereby significantly enhancing the performance of the student network. Specifically, a Modal Extraction Alignment (MEA) module is utilized to derive learning weights for student networks, integrating focal and global attention mechanisms. This methodology enables the student network to acquire optimal fusion strategies independent from that of teacher network without necessitating an additional feature fusion module. Furthermore, we present the SMOD dataset, a well-aligned challenging multispectral dataset for detection. Extensive experiments on the challenging KAIST, LLVIP, SUNRGB-D and SMOD datasets are conducted to validate the effectiveness of AMFD. The results demonstrate that our method outperforms existing state-of-the-art methods in both reducing log-average Miss Rate and improving mean Average Precision. The code is available athttps://github.com/bigD233/AMFD.git.
Zizhao Chen, Yeqiang Qian, Xiaoxiao Yang, Ming Yang 0002
IEEE Trans. Multim.5
2024 Cross-Modal Registration Using Adaptive Modeling in Infrastructure-based Vehicle Localization
abstract
Infrastructure-based vehicle localization, in comparison to single-agent approaches, offers several advantages including reduced system cost, extended perception range, enhanced data fusion capabilities, and energy savings. Many conventional approaches impose limitations on the types of objects due to the need for specific object-end modifications, such as applying perceptual markers like color-labeled plates and reflective balls. LiDAR presents a solution in terms of object arbitrariness, as it addresses the challenges of feature-free object modeling and continuous registration. However, achieving complete environmental coverage with LiDAR remains prohibitively expensive, particularly in extensive areas. Hence, this study proposes a cross-modal localization approach using adaptive modeling, employing LiDAR for object modeling and cost-effective sensor cameras for object tracking through image-point-cloud registration. Accurate correspondence between the model and observation can be estimated in real-time. The experiments are conducted in a typical scenario that requires adaptive modeling: Autonomous Valet Parking (AVP). Results demonstrate that the proposed system achieves comparable performance with significantly reduced system costs, highlighting its potential for large-scale deployment.
Yuesheng He, Hanyang Zhuang, Chenxi Yang 0002, Ming Yang 0002
ICRA5
2024 MOSFormer: A Transformer-based Multi-Modal Fusion Network for Moving Object Segmentation
abstract
3D moving object segmentation (MOS) is vital for autonomous systems, providing essential information for downstream tasks like mapping and localization. However, current MOS methods face challenges due to the limitation of existing datasets, which are sparse in moving objects and limited in scene diversity. Meanwhile, the prevalent methods are projection-based, struggling with the challenge of blurred boundaries. To tackle the dataset issue, we introduce a nuScenes-based MOS dataset, which provides richer scenes and more dynamic instances. To alleviate the boundary blur-ring issue and further improve accuracy and generalizability, we propose a dual-branch multimodal fusion MOS network, MOSFormer. The Transformer structure is incorporated to extract spatio-temporal information better, while image semantic information is utilized to refine the boundaries of moving objects. Finally, experiments on two datasets show that our method achieves state-of-the-art performance, and a mapping experiment with our method confirms its effectiveness in downstream tasks such as mapping and localization.
Zike Cheng, Hengwang Zhao, Qiyuan Shen, Weihao Yan 0001, Ming Yang 0002
IROS6
2024 ParkingE2E: Camera-based End-to-end Parking Network, from Images to Planning
abstract
Autonomous parking is a crucial task in the intelligent driving field. Traditional parking algorithms are usually implemented using rule-based schemes. However, these methods are less effective in complex parking scenarios due to the intricate design of the algorithms. In contrast, neural-network-based methods tend to be more intuitive and versatile than the rule-based methods. By collecting a large number of expert parking trajectory data and emulating human strategy via learning-based methods, the parking task can be effectively addressed. In this paper, we employ imitation learning to perform end-to-end planning from RGB images to path planning by imitating human driving trajectories. The proposed end-to-end approach utilizes a target query encoder to fuse images and target features, and a transformer-based decoder to autoregressively predict future waypoints. We conduct extensive experiments in real-world scenarios, and the results demonstrate that the proposed method achieved an average parking success rate of 87.8% across four different real-world garages. Real-vehicle experiments further validate the feasibility and effectiveness of the method proposed in this paper. The code can be found at: https://github.com/qintonguav/ParkingE2E.
Changze Li, Ziheng Ji, Tong Qin 0001, Ming Yang 0002
IROS5
2024 Cross-Modal Visual Relocalization in Prior LiDAR Maps Utilizing Intensity Textures
abstract
Cross-modal localization has drawn increasing attention in recent years, while the visual relocalization in prior LiDAR maps is less studied. Related methods usually suffer from inconsistency between the 2D texture and 3D geometry, neglecting the intensity features in the LiDAR point cloud. In this paper, we propose a cross-modal visual relocalization system in prior LiDAR maps utilizing intensity textures, which consists of three main modules: map projection, coarse retrieval, and fine relocalization. In the map projection module, we construct the database of intensity channel map images leveraging the dense characteristic of panoramic projection. The coarse retrieval module retrieves the top-K most similar map images to the query image from the database, and retains the top-K’ results by covisibility clustering. The fine relocalization module applies a two-stage 2D-3D association and a covisibility inlier selection method to obtain robust correspondences for 6DoF pose estimation. The experimental results on our self-collected datasets demonstrate the effectiveness in both place recognition and pose estimation tasks.
Qiyuan Shen, Hengwang Zhao, Weihao Yan 0001, Tong Qin 0001, Ming Yang 0002
IROS6
2024 An Online Automatic Calibration Method for Infrastructure-Based LiDAR-Camera via Cross-modal Object Matching
abstract
In indoor environments where the Global Navigation Satellite System (GNSS) isn’t available, the infrastructure-based LiDAR-camera joint array can provide high-precision localization for mobile robots, such as Autonomous Valet Parking (AVP). The primary challenge in employing the infrastructure-based LiDAR-camera joint array is the extrinsic calibration between the LiDAR and the camera. Moreover, to handle interference deviation caused by vibrations or inadequate mounting stiffness during operation, the calibration’s extrinsic parameters must be automatically updated online, presenting higher demands for infrastructure-based LiDAR-camera extrinsic calibration. This paper proposes an infrastructure LiDAR-camera online automatic calibration method based on prior knowledge of cross-modal target registration. This method requires no manual targets and initial pose guesses and can achieve extrinsic calibration. The object-prior model based on a lightweight object detection algorithm can rapidly detect scenes favorable for extrinsic calibration in sub-images of camera images. This creates favorable conditions for the registration of cross-modal networks and poses optimization of the LiDAR camera. Additionally, because a lightweight algorithm is used, the process does not compromise efficiency or consume excessive computational resources. Experimental results demonstrate that the proposed calibration method is suitable for calibrating infrastructure-based LiDAR-camera, with comparable accuracy and the ability to perform online calibration. Comparative experiments also show that the object-prior model can indeed select better scenes for LiDAR-camera extrinsic calibration, thus improving the accuracy and stability of extrinsic calibration to some extent.
Yuesheng He, Hanyang Zhuang, Ming Yang 0002
IROS4
2024 Active Vehicle Re-localization Based on Non-repetitive LiDAR with Gimbal Motion Strategy
abstract
The installation of a multi-layer 3D LiDAR atop the vehicle is a widely adopted hardware configuration for map-matching-based localization in intelligent driving. By offering a comprehensive 360° horizontal Field of View (FoV), this setup aims to achieve precise matching outcomes through the imposition of substantial geometric constraints against dynamic interferences and structural degradation. However, several factors limit its environmental adaptability, such as sparse point cloud density at distances, insufficient maximum sensing range, and notably, the restricted beam elevation angle, limiting the perception of the environment beyond obstacles. The rapid advancement of non-repetitive scanning LiDARs shows promise in mitigating such limitations. Nevertheless, their narrow FoV remains a challenge to overcome. In this study, we propose a solution by mounting such one single LiDAR on a two-axis rotating gimbal, enabling the vehicle to surpass the ranges and vertical FoV limitations of traditional setups actively. The corresponding gimbal motion strategy has been designed to automatically focus on the environment component with the most robust geometric constraints. Experimental results validate that the proposed method achieves superior robustness under high dynamic interference while delivering sufficient performance under standard conditions.
Xin'Ao Wu, Chenxi Yang 0002, Yiyang Guo, Hanyang Zhuang, Ming Yang 0002
IROS6
2024 MapLocNet: Coarse-to-Fine Feature Registration for Visual Re-Localization in Navigation Maps
abstract
Robust localization is the cornerstone of autonomous driving, especially in challenging urban environments where GPS signals suffer from multipath errors. Traditional localization approaches rely on high-definition (HD) maps, which consist of precisely annotated landmarks. However, building HD map is expensive and challenging to scale up. Given these limitations, leveraging navigation maps has emerged as a promising low-cost alternative for localization. Current approaches based on navigation maps can achieve highly accurate localization, but their complex matching strategies lead to unacceptable inference latency that fails to meet the real-time demands. To address these limitations, we introduce MapLocNet, a novel transformer-based neural re-localization method. Inspired by image registration, our approach performs a coarse-to-fine neural feature registration between navigation map features and visual bird’s-eye view features. MapLocNet substantially outperforms the current state-of-the-art methods on both nuScenes and Argoverse datasets, demonstrating significant improvements in localization accuracy and inference speed across both single-view and surround-view input settings. We highlight that our research presents an HD-map-free localization method for autonomous driving, offering a costeffective, reliable, and scalable solution for challenging urban environments.
Siyuan Lin, Xiangru Mu, Ming Yang 0002, Tong Qin 0001
IROS6
2024 Non-Repetitive: A Promising LiDAR Scanning Pattern
abstract
LiDAR is an essential sensor for intelligent vehicles. Recently, LiDARs used in vehicles produced by different companies have significant differences in their scanning patterns. Some vehicles use mechanical and solid-state (repetitive) LiDARs, while others use prism-based (non-repetitive) LiDARs. The scanning pattern of a LiDAR has a profound impact on its scanning performance. To investigate the influence of LiDAR scanning patterns, we created the "Repetitive-or-not" dataset, which is collected simultaneously by LiDARs with both repetitive and non-repetitive scanning patterns in the CARLA simulation environment. Using this dataset, we conducted a comprehensive statistical analysis of the scanning ability of repetitive and non-repetitive LiDARs. Furthermore, we looked into the effects of these two LiDAR scanning patterns on the performance of various 3D object detection algorithms. Finally, we explored the domain gap in the point cloud data produced by repetitive and non-repetitive LiDARs. Through an in-depth investigation of the "Repetitive-or-not" dataset, we have discovered that non-repetitive LiDAR shows great promise. This conclusion is primarily supported by its superior object scanning capabilities.
Angchen Xie, Yeqiang Qian, Weihao Yan 0001, Ming Yang 0002
IROS5
2024 Pix2Planning: End-to-End Planning by Vision-language Model for Autonomous Driving on Carla Simulator
abstract
The end-to-end neural network has become a hot topic in recent years. Compared with traditional module-based solutions, the end-to-end paradigm is able to reduce the accumulated error and avoid information loss, so that it earns great attention in autonomous driving tasks. However, the current end-to-end network designs easily lose useful information during training due to the complexity of mapping high-dimensional visual observation to navigation waypoints. Since the future navigation point is reasoned from the former one, the planning task is like a sequence generation task. Inspired by the great power of the neural language model, we propose an end-to-end framework, which transfers the planning task as a language sequence generation task conditioned on pixel inputs. The proposed method firstly extracts and transforms the image feature from camera-view to bird-eye-view (BEV). Then the target navigation point is constructed into a text sequence, as the prompt of the visual-language transformer. Finally, the auto-regressive transformer decoder receives the BEV feature and the text sequences to generate sequential waypoints. Overall, our proposed method can make full use of the environmental information and express the planning trajectory as a language sequence to learn the correspondence between trajectory sequences and images. We have conducted extensive experiments on CARLA benchmarks and our model achieves state-of-the-art performance compared with other visual methods.
Xiangru Mu, Tong Qin 0001, Songan Zhang, Chunjing Xu, Ming Yang 0002
IV5
2024 2D-3D Cross-Modality Network for End-to-End Localization with Probabilistic Supervision
abstract
Accurate localization ability is a crucial component for autonomous robots. Given existing LiDAR 3D points maps, it is cost-effective to localize the robot only with onboard camera compared to LiDAR. However, matching 2D visual information with 3D point cloud maps presents huge challenges due to different modalities, dimensions, noise and occlusion issues. To overcome it, we propose an end-to-end neural network-based solution, which determines the 6-DoF pose of the camera relative to an existing LiDAR map with centimeter accuracy. Given a query image, a pre-acquired point cloud and an initial pose, the cross-modality network will output a precise pose. By projecting the 3D point cloud onto the image plane, a depth image is acquired as seen from the initial pose. Subsequently, a cross-modality flow network establishes the correspondences of 2D pixels and projected points. Importantly, we leverage a robust probabilistic Perspective-n-Point (PnP) module, which are capable of fine-tuning 2D pairs and learning the pairs weight in an end-to-end manner. A comprehensive evaluation of our proposed algorithm is conducted in KITTI datasets. Furthermore, deploying the algorithm on the real-world parking lot scenario validates its strong practicality of the proposed algorithm. We highlight that this research offers a cost-effective and highly accurate solution that can be readily deployed in low-cost commercial vehicles.
Xiangru Mu, Tong Qin 0001, Chunjing Xu, Ming Yang 0002
IV5
2024 BLOS-BEV: Navigation Map Enhanced Lane Segmentation Network, Beyond Line of Sight
abstract
Bird’s-eye-view (BEV) representation is crucial for the perception function in autonomous driving tasks. It is difficult to balance the accuracy, efficiency and range of BEV representation. The existing works are restricted to a limited perception range within 50 meters. Extending the BEV representation range can greatly benefit downstream tasks such as topology reasoning, scene understanding, and planning by offering more comprehensive information and reaction time. The Standard-Definition (SD) navigation maps can provide a lightweight representation of road structure topology, characterized by ease of acquisition and low maintenance costs. An intuitive idea is to combine the close-range visual information from onboard cameras with the beyond line-of-sight (BLOS) environmental priors from SD maps to realize expanded perceptual capabilities. In this paper, we propose BLOS-BEV, a novel BEV segmentation model that incorporates SD maps for accurate beyond line-of-sight perception, up to 200m. Our approach is applicable to common BEV architectures and can achieve excellent results by incorporating information derived from SD maps. We explore various feature fusion schemes to effectively integrate the visual BEV representations and semantic features from the SD map, aiming to leverage the complementary information from both sources optimally. Extensive experiments demonstrate that our approach achieves state-of-the-art performance in BEV segmentation on nuScenes and Argoverse benchmark. Through multi-modal inputs, BEV segmentation is significantly enhanced at close ranges below 50m, while also demonstrating superior performance in long-range scenarios, surpassing other methods by over 20% mIoU at distances ranging from 50-200m.
Siyuan Lin, Tong Qin 0001, Chunjing Xu, Ming Yang 0002
IV8
2024 E2E Parking: Autonomous Parking by the End-to-end Neural Network on the CARLA Simulator
abstract
Autonomous parking is a crucial application for intelligent vehicles, especially in crowded parking lots. The confined space requires highly precise perception, planning, and control. Currently, the traditional Automated Parking Assist (APA) system, which utilizes geometric-based perception and rule-based planning, can assist with parking tasks in simple scenarios. With noisy measurement, the handcrafted rule often lacks flexibility and robustness in various environments, which performs poorly in super crowded and narrow spaces. On the contrary, there are many experienced human drivers, who are good at parking in narrow slots without explicit modeling and planning. Inspired by this, we expect a neural network to learn how to park directly from experts without handcrafted rules. Therefore, in this paper, we present an end-to-end neural network to handle parking tasks. The inputs are the images captured by surrounding cameras and basic vehicle motion state, while the outputs are control signals, including steer angle, acceleration, and gear. The network learns how to control the vehicle by imitating experienced drivers. We conducted closed-loop experiments on the CARLA Simulator to validate the feasibility of controlling the vehicle by the proposed neural network in the parking task. The experiment demonstrated the effectiveness of our end-to-end system in achieving the average position and orientation errors of 0.3 meters and 0.9 degrees with an overall success rate of 91%. The code is available at: https://github.com/qintonguav/e2e-parking-carla
Yunfan Yang, Denglong Chen, Tong Qin 0001, Xiangru Mu, Chunjing Xu, Ming Yang 0002
IV6
2024 Crowd-Sourced NeRF: Collecting Data From Production Vehicles for 3D Street View Reconstruction
abstract
Recently, Neural Radiance Fields (NeRF) achieved impressive results in novel view synthesis. Block-NeRF showed the capability of leveraging NeRF to build large city-scale models. For large-scale modeling, a mass of image data is necessary. Collecting images from specially designed data-collection vehicles can not support large-scale applications. How to acquire massive high-quality data remains an opening problem. Noting that the automotive industry has a huge amount of image data, crowd-sourcing is a convenient way for large-scale data collection. In this paper, we present a crowd-sourced framework, which utilizes substantial data captured by production vehicles to reconstruct the scene with the NeRF model. This approach solves the key problem of large-scale reconstruction, that is where the data comes from and how to use them. Firstly, the crowd-sourced massive data is filtered to remove redundancy and keep a balanced distribution in terms of time and space. Then a structure-from-motion module is performed to refine camera poses. Finally, images, as well as poses, are used to train the NeRF model in a certain block. We highlight that we presents a comprehensive framework that integrates multiple modules, including data selection, sparse 3D reconstruction, sequence appearance embedding, depth supervision of ground surface, and occlusion completion. The complete system is capable of effectively processing and reconstructing high-quality 3D scenes from crowd-sourced data. Extensive quantitative and qualitative experiments were conducted to validate the performance of our system. Moreover, we proposed an application, named first-view navigation, which leveraged the NeRF model to generate 3D street view and guide the driver with a synthesized video.
Tong Qin 0001, Changze Li, Haoyang Ye, Shaowei Wan, Minzhen Li, Ming Yang 0002
IEEE Trans. Intell. Transp. Syst.7
2024 Graph Correspondence-Based Point Set Registration
abstract
Point set registration, crucial in computer vision and robotics applications, encounters challenges, such as noise, outliers, and misalignment. Current methods often struggle with these issues, leading to suboptimal registration accuracy. This article proposes a novel graph correspondence-based algorithm to address these challenges in rigid point set registration. We model point sets as graphs, transforming the registration problem into a graph isomorphism problem. This approach is enhanced with probabilistic linear programming heuristics to efficiently establish correspondences between point sets. Our method significantly improves robustness against common registration errors and does not require initial pose estimation, a notable advantage over existing algorithms. Extensive experiments on various datasets, including applications in intelligent vehicle mapping and localization, demonstrate superior performance in correspondence establishment and registration accuracy compared to state-of-the-art methods, particularly under conditions of noise, outliers, and misalignment.
Liang Li 0010, Ming Yang 0002
IEEE Trans. Syst. Man Cybern. Syst.2
2023 Cross-Modal Monocular Localization in Prior LiDAR Maps Utilizing Semantic Consistency
abstract
Visual localization for mobile robots and intelligent vehicles in prior LiDAR maps can achieve high accuracy and low cost. However, algorithms for finding the cross-modal correspondences between images and LiDAR map points are not yet stable. In this paper, we propose a monocular visual localization system in prior LiDAR maps, which is based on the cross-modal registration to optimize the camera pose. To align the point clouds from vision and LiDAR map, a point-to-plane Iterative Closest Point algorithm utilizing semantic consistency is designed, and a decoupling optimization strategy is proposed to compute the affine transformation for the monocular scale ambiguity. Experiments on KITTI dataset show that utilizing the semantic consistency and geometric information of the map makes our system competitive with other methods. On the self-collected dataset, experiments on different light intensities demonstrate the robustness of the system in long-term localization tasks, and the ablation study demonstrates the effectiveness of the proposed algorithms.
Hengwang Zhao, Xuanlai Tang, Ming Yang 0002
ICRA5
2023 TTC4MCP: Monocular Collision Prediction Based on Self-Supervised TTC Estimation
abstract
Vision-based collision prediction for autonomous driving is a challenging task due to the dynamic movement of vehicles and diverse types of obstacles. Most existing methods rely on object detection algorithms, which only predict predefined collision targets, such as vehicles and pedestrians, and cannot anticipate emergencies caused by unknown obstacles. To address this limitation, we propose a novel approach using pixel-wise time-to-collision (TTC) estimation for monocular collision prediction (TTC4MCP). Our approach predicts TTC and optical flow from monocular images and identifies potential collision areas using feature clustering and motion analysis. To overcome the challenge of training TTC estimation models without ground truth data in new scenes, we propose a self-supervised TTC training method, enabling collision prediction in a wider range of scenarios. TTC4MCP is evaluated on multiple road conditions and demonstrates promising results in terms of accuracy and robustness.
Yeqiang Qian, Weihao Yan 0001, Ming Yang 0002
IROS6
2023 Truss Feature Based Robust Localization Method for Vehicles in Dynamic Industrial Scene
abstract
Localization is a key problem for autonomous vehicles in unmanned logistics of industrial scenes. However, due to a large amount of observation noise, the localization robustness of commonly used SLAM methods and traditional sensor configuration schemes is insufficient in dynamic industrial environments. To address the above issue, this paper proposes a system based on the truss feature and solves the key problem that the severe interference of dynamic objects such as industrial equipment, vehicles, and pedestrians causes errors in map matching results. First, the LiDAR with a hemispherical field of view is placed upwards to cover the truss feature. Then, the truss feature is effectively extracted based on the point cloud curvature for point cloud registration. After that, the wheel speed and steering angle information are fused to obtain the final poses. Considerable experiments in real dynamic environments demonstrate the superior robustness of our system, with an average localization error of less than 5.0 cm at 20 Hz.
Wei Yuan 0002, Bing Wang 0006, Chaochun Lian, Yongchun Yao, Yan Cai 0014, Ming Yang 0002
IV7
2023 Cy-CNN: cylinder convolution based rotation-invariant neural network for point cloud registration
Hengwang Zhao, Zhidong Liang, Yuesheng He, Ming Yang 0002
Sci. China Inf. Sci.5
2023 Threshold-Adaptive Unsupervised Focal Loss for Domain Adaptation of Semantic Segmentation
abstract
Semantic segmentation is an important task for intelligent vehicles to understand the environment. Current deep learning based methods require large amounts of labeled data for training. Manual annotation is expensive, while simulators can provide accurate annotations. However, the performance of the semantic segmentation model trained with synthetic datasets will significantly degenerate in the actual scenes. Unsupervised domain adaptation (UDA) for semantic segmentation is used to reduce the domain gap and improve the performance on the target domain. Existing adversarial-based and self-training methods usually involve complex training procedures, while entropy-based methods have recently received attention for their simplicity and effectiveness. However, entropy-based UDA methods have problems that they barely optimize hard samples and lack an explicit semantic connection between the source and target domains. In this paper, we propose a novel two-stage entropy-based UDA method for semantic segmentation. In stage one, we design a threshold-adaptative unsupervised focal loss to regularize the prediction in the target domain. It first introduces unsupervised focal loss into UDA for semantic segmentation, helping to optimize hard samples and avoiding generating unreliable pseudo-labels in the target domain. In stage two, we employ cross-domain image mixing (CIM) to bridge the semantic knowledge between two domains and incorporate long-tail class pasting to alleviate the class imbalance problem. Extensive experiments on synthetic-to-real and cross-city benchmarks demonstrate the effectiveness of our method. It achieves state-of-the-art performance using DeepLabV2, as well as competitive performance using the lightweight BiSeNet with great advantages in training and inference time.
Weihao Yan 0001, Yeqiang Qian, Ming Yang 0002
IEEE Trans. Intell. Transp. Syst.4
2023 Pseudo-Anchors: Robust Semantic Features for Lidar Mapping in Highly Dynamic Scenarios
abstract
Dynamic environments are challenging for anchor-free mapping using lidar in intelligent driving. This study imitates anchor-based approaches such as magnetic nails by applying novel Static Confidence Criteria (SCC) to the point-cloud semantic candidates to ensure their robustness. We name such verified features Pseudo-Anchors (P-A) as they hold similar properties to the anchor nodes: The P-A nodes are improbably formed by dynamic objects, and nodes’ blockage state can be immediately noticed once they are occluded. Another major challenge for mapping is improving large-scale global performance without sacrificing local consistency. Unrecognized GNSS pose drift may deteriorate local trajectory accuracy through post-processing such as graph optimization. In this study, we use the road network to provide the intersection information as a prior so that the GNSS can be better regarded as a reliable anchor factor. Three experiments are designed for this study. The first is ablations to verify the P-A concept; The second proves that the P-A-based lidar odometry outperformed the LOAM-based mainstream methods in highly dynamic scenarios; The third shows that our usage of the GNSS strengthens large-scale maps’ global consistency while causing less deterioration towards the local one. As a knowledge-based method, the P-A concept shows a high deployment efficiency, indicating the potential for migration to other features or even other sensors.
Chenxi Yang 0002, Lei He 0019, Hanyang Zhuang, Ming Yang 0002
IEEE Trans. Intell. Transp. Syst.5
2023 Global Pose Initialization Based on Gridded Gaussian Distribution With Wasserstein Distance
abstract
Feature descriptors, as abstraction of the critical information in the lidar point clouds, are often used in global pose initialization in large-area to provide a pose reference for intelligent driving system. The current state-of-the-art method Scan Context descriptor is generated based on points’ maximum height and, therefore, designed especially for outdoor scenarios without a ceiling. This study proposes a generic descriptor for both outdoor and indoor scenarios based on the point cloud’s Gridded Gaussian Distribution (GGD). Wasserstein distance is introduced to this field to evaluate the proposed GGD descriptors’ matching performance because it not only has a solid mathematical foundation in comparing two Gaussian distributions but also shows excellent time efficiency via a straightforward analytical solution without enumeration or iteration. We construct a multi-step error function to initialize the vehicle pose using conventional cosine similarity, and Wasserstein distance. Two experiments are designed for this study. The first experiment compares the pose initialization performances under various multi-frame superimposition distance in space to find an efficient GGD descriptor extraction setting. The second experiment verifies that the proposed method achieves a better pose initialization success rate than the mainstream methods.
Chenxi Yang 0002, Hanyang Zhuang, Ming Yang 0002
IEEE Trans. Intell. Transp. Syst.5
2022 HR-Planner: A Hierarchical Highway Tactical Planner based on Residual Reinforcement Learning
abstract
Tactical planning is crucial for safe and efficient driving on the highway. However, the problem is complicated by the uncertain intention of surrounding vehicles, as well as observation noise caused by measurement noise and perception errors. Rule-based tactical planning methods are ineffective in handling dynamic scenarios with uncertainty, and susceptible to observation noise. To tackle this problem, we propose a hierarchical tactical planning framework based on residual reinforcement learning. Besides, a new reinforcement learning from demonstrations scheme that views rule-based methods as soft guidance is developed to combine prior knowledge with data-driven methods. Based on the framework and the training scheme, rule-based methods not only can be improved in highway scenarios with uncertainty and observation noise, but also will guide the training procedure for increased sampling efficiency. Additionally, to boost in-depth and consistent exploration in a vehicle system with inertia, we employ noisy networks to explore the optimal policy. The proposed method is validated in a stochastic and uncertain simulation environment, and the results reveal that our method outperforms both rule-based methods and pure data-driven methods in terms of safety and driving efficiency under noisy observations and uncertainty.
Yueyuan Li, Hanyang Zhuang, Ming Yang 0002
ICRA5
2022 BAANet: Learning Bi-directional Adaptive Attention Gates for Multispectral Pedestrian Detection
abstract
Thermal infrared (TIR) image has proven effectiveness in providing temperature cues to the RGB features for multispectral pedestrian detection. Most existing methods directly inject the TIR modality into the RGB-based framework or simply ensemble the results of two modalities. This, however, could lead to inferior detection performance, as the RGB and TIR features generally have modality-specific noise, which might worsen the features along with the propagation of the network. Therefore, this work proposes an effective and efficient cross-modality fusion module called Bi-directional Adaptive Attention Gate (BAA-Gate). Based on the attention mechanism, the BAA-Gate is devised to distill the informative features and recalibrate the representations asymptotically. Concretely, a bi-direction multi-stage fusion strategy is adopted to progressively optimize features of two modalities and retain their specificity during the propagation. Moreover, an adaptive interaction of BAA-Gate is introduced by the illumination-based weighting strategy to adaptively adjust the recalibrating and aggregating strength in the BAA-Gate and enhance the robustness towards illumination changes. Considerable experiments on the challenging KAIST dataset demonstrate the superior performance of our method with satisfactory speed.
Xiaoxiao Yang, Yeqiang Qian, Hui-Jie Zhu, Ming Yang 0002
ICRA5
2022 Robust Localization for Intelligent Vehicles Based on Pole-Like Features Using the Point Cloud
abstract
Localization in the complex urban environment is an open problem for current methods. The occlusion from dynamic objects, such as vehicles and pedestrians, degenerates the precision of the localization result. This article proposes a pole-like feature-based localization framework to solve this problem. Pole-like objects, such as posts of lamps or traffic sign and tree trunks, widely exist in the urban environment and are robust to occlusion, as they are usually higher than the objects on the road. First, this type of feature is extracted from the point cloud by a robust clustering algorithm. Then, the features from different frames of data are stitched to generate a feature map. For online localization, a Monte Carlo localization (MCL) framework is used to fuse the vehicle motion data and the map-matching result. An improved version of iterative closest point (ICP) that is specifically designed for the pole-like feature association is used for map matching based on the state of every particle. With the MCL scheme, localization is robust to the local minimum or robot kidnapping problem. Experimental results in the real urban environment demonstrate the precision and robustness of the proposed method, with mean absolute errors less than 0.20 m and 0.5°. The results also show that the proposed method outperforms some state-of-the-art localization methods in the complex urban environment.Note to Practitioners—There are some works using features from the 3-D point cloud, e.g., corners, planes, and reflectance, for robot localization. Instead of the abstract features, this article presents an object-feature-based localization scheme. We propose a novel pole-like object extraction algorithm based on the spatial distribution of the 3-D points. This algorithm can extract most types of pole-like objects in the urban environment. As these objects are highly distinct from other types of objects in their surroundings, localization is achieved by associating the pole-like features in the map and the features detected in real time through maximizing the likelihood. The whole system is verified with data collected in the real world, which indicates that its accuracy can fulfill the requirements of autonomous driving. The limitation of the proposed method is that it highly depends on one specific type of feature, which may not work well in the rural environment. In future research, we will address this problem by incorporating more types of semantic features for localization.
Liang Li 0010, Ming Yang 0002, Lihong Weng
IEEE Trans Autom. Sci. Eng.2
2022 Pedestrian Graph +: A Fast Pedestrian Crossing Prediction Model Based on Graph Convolutional Networks
abstract
Estimating when pedestrians cross the street is essential for intelligent transportation systems. Accurate, real-time prediction is critical to ensure the safety of the most vulnerable road users while improving passenger comfort. In the present work, we developed a model called Pedestrian Graph +, an improvement of our previous work, Pedestrian Graph, which predicts pedestrian crossing action in urban areas based on a Graph Convolution Network. We integrated two convolutional modules in the new model that provide additional context information (cropped images, cropped segmentation maps, ego-vehicle velocity data) to the main Graph Convolutional module, thus increasing accuracy. Our model is faster and smaller than other state-of-the-art models, achieving equivalent accuracy. Our model is faster than state-of-the-art models, with an inference time of 6 ms (on a GTX 1080) and low memory consumption (0.3 MB). We tested our model on two datasets, Joint Attention in Autonomous Driving (JAAD) and Pedestrian Intention Estimation (PIE), achieving 86% and 89% accuracy, respectively. Another contribution of our work is the ability to dynamically process almost any input size in the time domain without significant loss of accuracy. It is possible due to the fully convolutional property of ConvNets. Our models and results are available athttps://github.com/RodrigoGantier/Pedestrian_graph_plus.
Pablo Rodrigo Gantier Cadena, Yeqiang Qian, Ming Yang 0002
IEEE Trans. Intell. Transp. Syst.4
2022 LiDAR SLAM Based Multivehicle Cooperative Localization Using Iterated Split CIF
abstract
High-precision localization in an unknown environment is the fundamental requirement of autonomous vehicle for safe advanced driving. Cooperative localization has obvious advantages in precision, fault tolerance, and flexibility compared with single-vehicle localization, which estimates the vehicle pose via fusing the multiple sources data from sensors and extra shared neighbor vehicle information available when multiple vehicles operate simultaneously. For cooperative localization, there are more than one type of error data sources that degenerate the state estimation need to be considered, such as inter-estimation correlation and innovation, observation outliers, which may exist simultaneously. Previous works usually evaluate performance under one of the error types at one time, or they do not perform well under some extreme situations where both of the above-mentioned correlation and outliers exist. This paper presents an accurate and robust iterated split covariance intersection filter (Iterated Split CIF) based cooperative localization strategy with a decentralized framework, which can ensure the performance when data sources of various error types exist at the same time. In addition, we adopt an effective point cloud registration method to obtain the cooperative relative pose estimation using mutually shared information from neighbor vehicles. A CARLA simulator based comparative study demonstrates the potential and advantage of the proposed multi-vehicle cooperative localization using Iterated Split CIF in terms of accuracy, robustness and efficiency.
Susu Fang, Hao Li 0024, Ming Yang 0002
IEEE Trans. Intell. Transp. Syst.3
2022 Point Cloud Registration Based on Direct Deep Features With Applications in Intelligent Vehicles
abstract
Point cloud registration is widely used in the research of intelligent vehicles, typical problems include map matching, visual odometer, pose estimation,etc. This paper proposes a deep learning-based registration method that can input point clouds directly, thereby preventing information loss of preprocessing needed by alternative deep-learning approaches. Our network, named DPFNet (Direct Point Feature Net), gradually downsamples the point cloud and aggregates points around determined reference points to formulate local features automatically. This is facilitated by a novel convolution-like operator and a novel loss function. The points in the point cloud are mapped to a high dimensional embedding through the designed deep neural network, where every embedding reflects the local feature of a specific spatial area. Based on the embedding features, correspondences between points can be estimated robustly and the registration between the point clouds can be obtained using an external geometric optimization algorithm. Experimental results on open benchmarks validate the proposed method and show that its performance is favourable over several baseline methods. Specifically, we test the proposed algorithm on KITTI benchmark, which shows its potential in tasks of intelligent vehicles,e.g., map matching, visual or LiDAR odometer.
Liang Li 0010, Ming Yang 0002
IEEE Trans. Intell. Transp. Syst.2
2022 Gated-Residual Block for Semantic Segmentation Using RGB-D Data
abstract
Semantic segmentation is an important technique for scene understanding in the intelligent transportation system. RGB-D data shows great advantages over the unimodal data in this area, and it can be easily obtained from consumer sensors nowadays. How to design effective fusion structures to fuse RGB and depth signals in RGB-D data is a challenging problem. This paper proposes a novel gated-residual block to address this problem. The structure consists of two residual units and one gated fusion unit, the residual unit progressively aggregates modality-specific features from the modality-specific signals and the gate mechanism computes complementary features for them. Based on the gated-residual block, the paper presents the deep multimodal networks, named GRBNet, for RGB-D semantic segmentation. Experiments on ScanNet, Cityscapes and SUN RGB-D datasets verify the effectiveness of the proposed approach and demonstrate that the GRBNet achieved competitive performance.
Yeqiang Qian, Liuyuan Deng, Tianyi Li 0003, Ming Yang 0002
IEEE Trans. Intell. Transp. Syst.5
2022 Survey on Fish-Eye Cameras and Their Applications in Intelligent Vehicles
abstract
Fish-eye cameras have become essential sensors in intelligent vehicles. Due to its unique projection principle, a fish-eye camera can provide a large field of view. Benefiting from this special feature, fish-eye cameras have rich applications in intelligent vehicles. However, dataset and distortion problems are still challenges when applying fish-eye cameras in reality. This work introduces the projection principle of fish-eye cameras, and four classic fish-eye image representation models are presented. Then, the typical fish-eye datasets are presented, including real collected data and virtually generated data. Through the organization and summarization of the relevant studies, we demonstrate various applications of fish-eye cameras in intelligent vehicles, e.g., object detection and tracking, image segmentation, mapping and localization, and around-view monitoring. These works design various strategies to exploit the advantages of fish-eye cameras and prevent image distortion problems, showing the broad application prospects of such cameras. Finally, we discuss the development tendencies of intelligent vehicle applications involving fish-eye cameras.
Yeqiang Qian, Ming Yang 0002, John M. Dolan
IEEE Trans. Intell. Transp. Syst.2
2022 AGBM: An Adaptive Gradient Balanced Mechanism for the End-to-End Steering Estimation
abstract
End-to-end steering estimation is one of the important deep regression tasks. However, driving datasets are always imbalanced on the distribution of the steering value, which makes end-to-end learning models perform bad steering estimation accuracy, especially for the sharp steering value estimation, which finally, leads to an imbalanced training problem. In this paper, the essential reason for the imbalanced training problem for the deep regression task is first revealed as the gradient disharmony. To hedge the gradient disharmony, a general gradient balanced mechanism for the deep regression task is proposed in this paper. Meanwhile, the general gradient balanced mechanism is applied for the steering estimation task. According to the steering distribution, a novel Adaptive Gradient Balanced Mechanism (AGBM) is proposed to hedge the gradient disharmony based on the Laplace distribution with the general gradient balanced mechanism theory. Further, a new loss function embedding with AGBM is proposed to balance the gradient disharmony for the training process. AGBM makes the process of gradient hedging with adaptive style without parameters finetuning. Based on five driving datasets, four experiments are conducted for demonstrations. The experiment results show the AGBM enables the end-to-end models to perform the best accuracy of the steering estimation, including the sharp steering samples. Meanwhile, AGBM is generalized with different end-to-end learning models.
Wei Yuan 0002, Hanyang Zhuang, Ming Yang 0002
IEEE Trans. Intell. Transp. Syst.4
2021 Multi-Vehicle Cooperative SLAM Using Iterated Split Covariance Intersection Filter
abstract
Simultaneous localization and mapping (SLAM) is important to outdoor intelligent vehicle applications. Multivehicle cooperative SLAM which takes advantage of data sharing can outperform single vehicle SLAM. This paper proposes a multi-vehicle cooperative SLAM method using iterated split covariance intersection filter (Iterated Split CIF). In the proposed method, a vehicle can flexibly perform cooperative SLAM with other vehicles in decentralized way, without complicated monitoring and controlling of data flow. Moreover, the innovation and observation outliers caused by modeling errors or abnormal measurements can be solved reliably. A simulation-based comparative study demonstrates the potential and advantage of the proposed multi-vehicle cooperative SLAM using Iterated Split CIF in terms of accuracy and robustness.
Susu Fang, Hao Li 0024, Ming Yang 0002
IV3
2021 SPADE-E2VID: Spatially-Adaptive Denormalization for Event-Based Video Reconstruction
abstract
Event-based cameras have several advantages over traditional cameras that shoot videos in frames. Event cameras have a high temporal resolution, high dynamic range, and almost non-existence of blurriness. The data that is produced by event sensors forms a chain of events when a change in brightness is reported in each pixel. This feature makes it difficult to directly apply existing algorithms and take advantage of the event camera data. Due to the developments in neural networks, important advances were made in event-based image reconstruction. Even though these neural networks achieve precise reconstructions while preserving most of the properties of the event cameras, there is still an initialization time that needs to have the highest possible quality in the reconstructed frames. In this work, we present the SPADE-E2VID neural network model that improves the quality of early frames in an event-based reconstructed video, as well as the overall contrast. The SPADE-E2VID model improves the quality of the first reconstructed frames by 15.87% for MSE error, 4.15% for SSIM, and 2.5% in LPIPS. In addition, the SPADE layer in our model allows training our model to reconstruct videos without a temporal loss function. Another advantage of our model is that it has a faster training time. In a many-to-one training style, we avoid running the loss function at each step, executing the loss function at the end of each loop only once. In the present work, we also carried out experiments with event cameras that do not have polarity data. Our model produces quality video reconstructions with non-polarity events in HD resolution (1200 × 800). The Video, the code, and the datasets will be available at: https://github.com/RodrigoGantier/SPADE_E2VID.
Pablo Rodrigo Gantier Cadena, Yeqiang Qian, Ming Yang 0002
IEEE Trans. Image Process.4
2021 Neutral Cross-Entropy Loss Based Unsupervised Domain Adaptation for Semantic Segmentation
abstract
The generalization performance for semantic segmentation remains a major challenge when the data distributions between the source and target domain mismatch. Unsupervised domain adaptation (UDA) approaches are proposed to mitigate the problem above, among which entropy-minimization-based methods have gained more and more attention. However, the methods merely follow the cluster assumption sharpening the prediction distribution, thus have limited performance improvement. Without additional priors, the entropy loss can easily over-sharpen the prediction distribution, which brings noisy information into the learning process. On the other hand, the gradient of the entropy loss is strongly biased toward easy samples, also leading to limited generalization advances. In this paper, we firstly propose a pixel-level consistency regularization method, which introduces the smoothness prior to the UDA problem. Furthermore, we propose the neutral cross-entropy loss based on the consistency regularization, and reveal that its internal neutralization mechanism mitigates the over-sharpness of entropy minimization via the flatness effect of consistency regularization. We also demonstrate that the gradient bias toward easy samples is inherently tackled via the neutral cross-entropy loss. The experiments show that the proposed method has outperformed state-of-the-art methods in two synthetic-to-real experiments, only using the lightweight network.
Ming Yang 0002, Liuyuan Deng, Yeqiang Qian
IEEE Trans. Image Process.2
2021 Adversarial Training-Based Hard Example Mining for Pedestrian Detection in Fish-Eye Images
abstract
Since fish-eye cameras are popular in the intelligent transportation systems, accurate pedestrian detection in fish-eye images becomes more and more critical in low-speed scenarios. Usually, big data based training is the key for detectors to handle the distortion problem in fish-eye images. Especially, hard examples are more important for detectors. They have more complex features and are hard to recognize. In conventional methods, fish-eye images are collected and labeled manually. These methods are expensive and labor-intense. More importantly, these methods are still hard to collect abundant hard examples since they are rare in reality. This work proposes the Distortion Generation Network, which generates generous fish-eye images automatically using only small samples. Moreover, the Adversarial Distortion Generation Network is proposed to mine hard examples via adversarial training. These hard examples benefit detectors to be more robust to seriously distorted objects in fish-eye images. Experiments with the ETH, the KITTI, the GM-ATCI and real fish-eye datasets demonstrate that the proposed methods achieve higher accuracy than conventional methods in pedestrian detection in fish-eye images.
Yeqiang Qian, Ming Yang 0002, Hao Li 0024, Bing Wang 0006
IEEE Trans. Intell. Transp. Syst.2
2021 SteeringLoss: A Cost-Sensitive Loss Function for the End-to-End Steering Estimation
abstract
Imbalanced training is a challenge in the field of autonomous driving. For the steering estimation task, imbalanced training is the core reason that an end-to-end model cannot estimate sharp steering value well. Inspired by researches on the steering estimation, this paper proposes a novel loss function to train a high-performance end-to-end model for handling the imbalanced training problem, which is named SteeringLoss. Firstly, the imbalanced distribution of the steering value for driving datasets is analyzed, which is similar to the double long-tailed distribution. Secondly, with the feature of distribution, this paper designs a cost-sensitive loss function step by step, the new loss function can improve the impact of the sharp steering value while maintaining the impact of the small steering value. Thirdly, three typical steering estimation models are established with SteeringLoss for demonstration, including the CNN model, the CNN-LSTM model and the 3DCNN-LSTM model. Finally, three experiments are designed for SteeringLoss: Experiment I demonstrates SteeringLoss can avoid imbalanced training problem; Experiment II discusses the finetuning principle of SteeringLoss and gives the basic guideline for using SteeringLoss; Experiment III shows the results on different typical end-to-end steering estimation models, which shows effectiveness of SteeringLoss for all the models and gives different solutions for the steering estimation. Moreover, the SteeringLoss is suitable for the imbalanced training with similar distributions of datasets besides end-to-end steering estimation, which indicates the potential value of the research on SteeringLoss in the future.
Wei Yuan 0002, Ming Yang 0002, Hao Li 0024, Bing Wang 0006
IEEE Trans. Intell. Transp. Syst.2
2021 Joint Localization Based on Split Covariance Intersection on the Lie Group
abstract
This article presents a pose fusion method that accounts for the possible correlations among measurements. The proposed method can handle data fusion problems whose uncertainty has both independent and dependent parts. Different from the existing methods, the uncertainties of the various states or measurements are modeled on the Lie algebra and projected to the manifold through the exponential map, which is more precise than that modeled in the vector space. The correlation is based on the theory of covariance intersection, where the independent and dependent parts are split to yield a more consistent result. In this article, we provide a novel method for the correlated pose fusion algorithm on the manifold. Theoretical derivation and analysis are detailed first, and then, the experimental results are presented to support the proposed theory. The main contributions are threefold: First, we provide a theoretical foundation for the split covariance intersection filter performed on the manifold, where the uncertainty is associated with the Lie algebra. Second, the proposed method gives an explicit fusion formalism on$ \text{SE}(3)$and$ \text{SE}(2)$, which covers the most use cases in the field of robotics. Third, we present a localization framework that can work for both single-robot and multirobot systems, where not only the fusion with possible correlation is derived on the manifold but also the state evolution and relative pose computation are performed on the manifold. The experimental results validate the advantage of this approach over state-of-the-art methods.
Liang Li 0010, Ming Yang 0002
IEEE Trans. Robotics2
2020 ROI-cloud: A Key Region Extraction Method for LiDAR Odometry and Localization
abstract
We present a novel key region extraction method of point cloud, ROI-cloud, for LiDAR odometry and localization with autonomous robots. Traditional methods process massive point cloud data in every region within the field of view. In dense urban environments, however, processing redundant and dynamic regions of point cloud is time-consuming and harmful to the results of matching algorithms. In this paper, a voxelized cube set, ROI-cloud, is proposed to solve this problem by exclusively reserving the regions of interest for better point set registration and pose estimation. 3D space is firstly voxelized into weighted cubes. The key idea is to update their weights continually and extract cubes with high importance as key regions. By extracting geometrical features of a LiDAR scan, the importance of each cube is evaluated as a new measurement. With the help of on-board IMU/odometry data as well as new measurements, the weights of cubes are updated recursively through Bayes filtering. Thus, dynamic and redundant point cloud inside cubes with low importance are discarded by means of Monte Carlo sampling. Our method is validated on various datasets, and results indicate that the ROI-cloud improves the existing method in both accuracy and speed.
Ming Yang 0002, Bing Wang 0006
ICRA2
2020 Split Covariance Intersection Filter based Front-Vehicle Track Estimation for Vehicle Platooning without Communication
abstract
Vehicle platooning is an innovative technology for intelligent transportation systems. Each vehicle in the platoon is required to autonomously follow its front vehicle's path unconditionally and accurately. For platoons without communication, accurate path and velocity estimation of the front vehicle is challenging and crucial. Instead of memorizing the original detection result of the front vehicle to generate a path, the path is estimated by fusing motion prediction and observation in this paper. Since there is a correlation between the two estimates, Split Covariance Intersection Filter is used to guarantee the fusion consistency. Besides, a motion model considering velocity is used in the filter to achieve a precise estimate of the predecessor's velocity simultaneously. Moreover, a path generation approach is designed in a high-frequency loop independent of front vehicle detection to improve continuity of the path. Experimental validations in the real-world environment highlight the remarkable improvement in the accuracy of both path and velocity estimation. Meanwhile, the path generated by the proposed approach is more continuous compared with the commonly used method. Furthermore, complete vehicle platooning demonstrations in diverse environments prove the practicality and robustness of the proposed approach.
Ming Yang 0002, Wei Yuan 0002, Hao Li 0024
IV2
2020 Time-of-Flight Camera Based Indoor Parking Localization Leveraging Manhattan World Regulation
abstract
Localization is a key problem for autonomous driving in indoor parking. There have been some previously proposed methods based on UWB, LiDAR, fisheye cameras, etc. However, most of these methods have some drawbacks such as high cost or dependency on light conditions. To address these challenges, this paper proposes a novel Time-of-Flight (ToF) camera based mapping and localization system for indoor parking lots leveraging Manhattan World Regulation. ToF cameras are low-cost and can actively generate dense point clouds of the environment without external light sources. To overcome the shortcoming of ToF camera small field of view, the proposed system utilizes the structural information of the ceiling of indoor parking lots and Manhattan World Regulation. We track the surface normals on the unit sphere for drift-free rotation estimation. Based on this drift-free rotation, we can effectively calculate 6-DOF pose with decoupled rotation and translation estimation during mapping or global localization. This new system runs in real-time on limited computation resources and is demonstrated on two different challenging indoor parking, achieving real-time performance at 10 Hz and localization error less than 0.1 meter.
Hengwang Zhao, Ming Yang 0002, Yuesheng He
IV2
2020 G2P: a new descriptor for pedestrian detection
Ming Yang 0002, Yeqiang Qian, Linji Xue, Hao Li 0024, Liuyuan Deng
Neural Comput. Appl.1
2020 Monocular pedestrian orientation estimation based on deep 2D-3D feedforward
Chenchen Zhao 0001, Yeqiang Qian, Ming Yang 0002
Pattern Recognit.3
2020 Robust Point Set Registration Using Signature Quadratic Form Distance
abstract
Point set registration is a problem with a long history in many pattern recognition tasks. This paper presents a robust point set registration algorithm based on optimizing the distance between two probability distributions. A major problem in point to point algorithms is defining the correspondence between two point sets. This paper follows the idea of some probability-based point set registration methods by representing the point sets as Gaussian mixture models (GMMs). By optimizing the distance between the two GMMs, rigid transformations (rotation and translation) between two point sets can be obtained without having to find a correspondence. Previous studies have used L2, Kullback Leibler, etc. distance to measure similarity between two GMMs; however, these methods have problems with robustness to noise and outliers, especially when the covariance matrix is large, or a local minimum exists. Therefore, in this paper, the signature quadratic form distance is derived to measure the distribution similarity. The contribution of this paper lies in adopting the signature quadratic form distance for the point set registration algorithm. The experimental results show the precision and robustness of this algorithm and demonstrate that it outperforms other state-of-the-art point set registration algorithms regarding factors, such as noise, outliers, missing partial structures, and initial misalignment.
Liang Li 0010, Ming Yang 0002, Bing Wang 0006
IEEE Trans. Cybern.2
2020 Restricted Deformable Convolution-Based Road Scene Semantic Segmentation Using Surround View Cameras
abstract
Understanding the surrounding environment of the vehicle is still one of the challenges for autonomous driving. This paper addresses 360-degree road scene semantic segmentation using surround view cameras, which are widely equipped in existing production cars. First, in order to address large distortion problem in the fisheye images, Restricted Deformable Convolution (RDC) is proposed for semantic segmentation, which can effectively model geometric transformations by learning the shapes of convolutional filters conditioned on the input feature map. Second, in order to obtain a large-scale training set of surround view images, a novel method called zoom augmentation is proposed to transform conventional images to fisheye images. Finally, an RDC based semantic segmentation model is built; the model is trained for real-world surround view images through a multi-task learning architecture by combining real-world images with transformed images. Experiments demonstrate the effectiveness of the RDC to handle images with large distortions, and that the proposed approach shows a good performance using surround view cameras with the help of the transformed images.
Liuyuan Deng, Ming Yang 0002, Hao Li 0024, Tianyi Li 0003
IEEE Trans. Intell. Transp. Syst.2
2020 DLT-Net: Joint Detection of Drivable Areas, Lane Lines, and Traffic Objects
abstract
Perception is an essential task for self-driving cars, but most perception tasks are usually handled independently. We propose a unified neural network named DLT-Net to detect drivable areas, lane lines, and traffic objects simultaneously. These three tasks are most important for autonomous driving, especially when a high-definition map and accurate localization are unavailable. Instead of separating tasks in the decoder, we construct context tensors between sub-task decoders to share designate influence among tasks. Therefore, each task can benefit from others during multi-task learning. Experiments show that our model outperforms the conventional multi-task network in terms of the task-wise accuracy and the overall computational efficiency, in the challenging BDD dataset.
Yeqiang Qian, John M. Dolan, Ming Yang 0002
IEEE Trans. Intell. Transp. Syst.3
2020 Oriented Spatial Transformer Network for Pedestrian Detection Using Fish-Eye Camera
abstract
Pedestrian detection using fish-eye cameras is a principal research focus in computer vision. Lack of pedestrian datasets of fish-eye images and pedestrian distortion in fish-eye images are two primary challenges. In this paper, two approaches are proposed to deal with these two challenges, respectively. On the one hand, the projective model transformation (PMT) algorithm is proposed, which can transform normal images into fish-eye images. The PMT can be applied to most of the pedestrian datasets and generates corresponding fish-eye image datasets. In this way, enough training data can be provided through the PMT. On the other hand, the oriented spatial transformer network (OSTN) is designed to rectify warped pedestrian features using CNNs, so that pedestrians in fish-eye images are easier for detectors to recognize. The OSTN can be embedded into universal deep learning based detectors easily. Moreover, the new pedestrian detector, where the OSTN is embedded, can be trained end to end. Finally, the OSTN based fish-eye pedestrian detectors can be trained using fish-eye images, which are generated using the PMT. Experiments on ETH, KITTI, Citypersons, and real pedestrian datasets show the effectiveness of the PMT and accuracy improvement of pedestrian detection in fish-eye images using the OSTN.
Yeqiang Qian, Ming Yang 0002, Xu Zhao 0001, Bing Wang 0006
IEEE Trans. Multim.2
2019 Hierarchical Depthwise Graph Convolutional Neural Network for 3D Semantic Segmentation of Point Clouds
abstract
This paper proposes a hierarchical depthwise graph convolutional neural network (HDGCN) for point cloud semantic segmentation. The main chanllenge for learning on point clouds is to capture local structures or relationships. Graph convolution has the strong ability to extract local shape information from neighbors. Inspired by depthwise convolution, we propose a depthwise graph convolution which requires less memory consumption compared with the previous graph convolution. While depthwise graph convolution aggregates features channel-wisely, pointwise convolution is used to learn features across different channels. A customized block called DGConv is specially designed for local feature extraction based on depthwise graph convolution and pointwise convolution. The DGConv block can extract features from points and transfer features to neighbors while being invariant to different point orders. HDGCN is constructed by a series of DGConv blocks using a hierarchical structure which can extract both local and global features of point clouds. Experiments show that HDGCN achieves the state-of-the-art performance in the indoor dataset S3DIS and the outdoor dataset Paris-Lille-3D.
Zhidong Liang, Ming Yang 0002, Liuyuan Deng, Bing Wang 0006
ICRA2
2019 Multi-Vehicle Cooperative Local Mapping Using Split Covariance Intersection Filter
abstract
Local mapping plays an important role in outdoor intelligent vehicle applications and multi-vehicle cooperative local mapping which takes advantage of vehicular communication can bring considerable benefits to this important task. In this paper, a multi-vehicle cooperative local mapping architecture using split covariance intersection filter (Split CIF) is proposed. In the proposed method, a vehicle can flexibly perform cooperative local mapping with other vehicles in decentralized way, without complicated monitoring and controlling of data flow among vehicles; fused maps can be shared freely among vehicles. An efficient and accurate implementation of the Split CIF is also introduced. A simulation-based comparative study demonstrates the potential and advantage of the proposed multi-vehicle cooperative local mapping architecture using Split CIF.
Hao Li 0024, Ming Yang 0002
IROS2
2019 3D Semantic Modelling With Label Correction For Extensive Outdoor Scene
abstract
Semantic point cloud labeling is an important step of 3D urban scene mapping. One doable method is projecting labels from 2D semantic image to 3D point cloud. But the projection accuracy depends on the performance of calibration which cannot be absolutely accurate. In addition, the difference of information dimension between image and point cloud, with the addition of different sensor viewpoints can lead to label error in final 3D semantic results. This paper describes a projection based 3D semantic point cloud building method for extensive outdoor scene, and proposed a segmentation based voting strategy to handle the mislabeling problem in raw projection result. Firstly, based on pre-calibration among sensors, we get semantic point cloud by label projection from semantic image to point cloud. Secondly, we employ a segmentation method to cluster point cloud into different parts which belongs to different objects, this process increases the spatial consistency of semantic label at the object-level. Finally, in each part, we use a voting strategy to filter the label of each point. This strategy was tested in real world urban scene, experiments demonstrate the outstanding performance in decreasing the rate of semantic label error.
Ming Yang 0002, Bing Wang 0006
IV2
2019 Graph Matching Pose SLAM based on Road Network Information
abstract
This paper presents a Graph Matching Pose SLAM to build the perception map which is a prerequisite of localization and environment perception of the intelligent vehicle. Graph-based simultaneous localization and mapping (SLAM) is widely used to build a map with global consistency and requires loop closure to eliminate the accumulative errors. However, loop closure forces the vehicle to re-visit a previously entered area. In this paper, a road network-based graph SLAM method without the requirement of loop closure is proposed to build a consistent map for the real environment. The estimated poses from the original SLAM are linked with the available road network based on the graph matching. Thus the available road network introduces the real topology of the environment to Pose SLAM regularly. In this way, the map is optimized by using the factor graph inference frequently. The proposed method is validated on the KITTI dataset and the real-world experiments. This method outperforms the state-of-the-art methods.
Lei He 0019, Ming Yang 0002, Hao Li 0024, Yuesheng He, Bing Wang 0006
IV2
2019 SteeringLoss: Theory and Application for Steering Prediction
abstract
Imbalanced datasets are deathful for model training. In the field of steering prediction, imbalanced training is the core reason that model is unable to predict sharp steering value well. This paper proposes a new loss framework to train the robust end-to-end model, which is named SteeringLoss. The imbalanced distribution of steering value for datasets is analyzed, which is similar with Gaussian distribution. With the feature of distribution, the gain factor (1+α|y|β)γis added to square loss function. This new SteeringLoss framework is able to improve the impact of sharp steering value while maintain the impact of small steering value. Experiment results show the SteeringLoss based model performs higher performance than traditional square loss based model with higher prediction precision and wider prediction range. Meanwhile, γ is able to control the time of training process, and β can control the model performance, different distribution needs to choose different pair of parameters. What's more, the SteeringLoss framework is suitable for imbalanced training with similar distribution of dataset. The code can be found at: https://github.com/weiy1991/SteeringLoss.
Wei Yuan 0002, Ming Yang 0002, Bing Wang 0006
IV2
2019 A Terrain-Based Vehicle Localization Approach Robust to Braking
abstract
Terrain-based vehicle localization is a valuable alternative to GPS when the signal is blocked or dynamic occlusions exist. However, braking events can induce the slips and vibrations of the vehicle and, thus, increase the errors of terrain-based position estimations. This paper develops a terrain-based localization approach which is robust to braking events. First, the terrain map is generated and stored before localization, which includes distance measurements, terrain feature data, and the geographic locations. The terrain feature data are pitch differences to eliminate the accumulated error. Next, particle-filter-based and acceleration-considered vehicle localization is proposed for position estimation. It takes into account the influence of braking in the vehicle model and the inference process. Experimental results demonstrate the repeatability of pitch difference measurements and the robustness of the proposed approach in the cases of smooth driving and braking.
Tianyi Li 0003, Ming Yang 0002, Hao Li 0024, Liuyuan Deng
IEEE Trans. Intell. Transp. Syst.2
2019 Vehicle Localization at an Intersection Using a Traffic Light Map
abstract
Traditional vehicle localization methods use information from a GPS, inertial navigation, and an odometer, but a GPS is often limited by availability in urban areas for its sensitivity to terrain and interference. Aiming at the above problems, this paper presents a method of intersection localization using a traffic lights' map. In particular, traffic lights with significant visual characteristics in the urban environment are used as landmarks. The localization accuracy of the autonomous vehicle at intersections is improved by combining the location information of traffic lights provided by a high-precision map. The whole scheme of the localization system is designed and the sensors used are determined. First, the coordinates of sensors are established, respectively, and the transformation and rotation are calibrated. Then, the state model and measurement model of the vehicle vision system are established. Combined with the location and height information of the traffic lights provided by the high-precision map, the extended Kalman filter is used to fuse the vision detection results of the traffic lights with the inertial measurement unit information. The experiments demonstrate that the method proposed in this paper improves the lateral localization accuracy and the accuracy of the vehicle's yaw angle.
Hairu Huang, Bing Wang 0006, Ming Yang 0002
IEEE Trans. Intell. Transp. Syst.5
2018 Putting the Anchors Efficiently: Geometric Constrained Pedestrian Detection
Liangji Fang, Xu Zhao 0001, Xiao Song 0002, Shiquan Zhang, Ming Yang 0002
ACCV (5)5
2018 BSN: Boundary Sensitive Network for Temporal Action Proposal Generation
Xu Zhao 0001, Haisheng Su, Chongjing Wang, Ming Yang 0002
ECCV (4)5
2018 Pedestrian Feature Generation in Fish-Eye Images via Adversary
abstract
Pedestrian detection in fish-eye images is always an important problem in advanced driver assistance systems (ADAS). In conventional methods, pedestrian detectors will be trained using fish-eye images. But it is hard to collect and label enough fish-eye images manually. Therefore, a new strategy for training fish-eye pedestrian detectors using images from normal pedestrian datasets is proposed in this work. Concretely, Fish-eye Spatial Transformer Network (FSTN) is designed to generate pedestrian features in fish-eye images. FSTN aims to simulate distorted pedestrian features on the feature maps. Then the entire network is trained via adversary. FSTN is trained to generate examples which are difficult for pedestrian detectors to classify. So that the detectors are more robust to the deformation. FSTN can be embedded into state-of-the-art detectors easily. And the entire pedestrian detector, where the FSTN embedded, can be trained end to end via adversary. Moreover, experiments on ETH and KITTI pedestrian datasets show the slight accuracy improvement of pedestrian detection in fish-eye images using adversarial network compared with conventional methods.
Yeqiang Qian, Ming Yang 0002, Bing Wang 0006
ICRA2
2018 An Efficient Hierarchical Convolutional Neural Network for Traffic Object Detection
abstract
In this paper, we propose a novel hierarchical convolutional neural network for traffic object detection, which is defined as Fusion and Multi-level Alignment CNN (namely FMLA-CNN). The method extends a popular two-stage detector by incorporating a remodified feature fusion module and a multi-level alignment (MLA) strategy such that it is capable of efficiently detecting multi-scale objects in autonomous driving scenario. The feature fusion strategy in proposal generation network improves detection accuracy by inserting high-level semantics to the whole pyramidal feature hierarchy. Subsequently the MLA strategy in the second detection stage can exactly reserve spatial locations from corresponding feature layers determined by hierarchical region-of-interest proposals. In the experiments on KITTI benchmark, our FMLA-CNN achieves an impressively better trade-off between accuracy and efficiency compared with other state-of-the-art methods.
Qianqian Bi, Ming Yang 0002, Bing Wang 0006
Intelligent Vehicles Symposium2
2018 Monocular Visual-Inertial Odometry Based on Sparse Feature Selection with Adaptive Grid
abstract
For sparse feature based visual-inertial odometry, feature selection is vital to the performance. The selected features should be evenly distributed in the image and be appropriate for tracking. This paper presents a visual-inertial odometry approach based on sparse feature selection with the adaptive grid to improve the performance in different environments. In the proposed approach, FAST corner detection is employed in every grid, the size of which is adaptively adjusted. Features in the same grid are ranked and selected based on scores to ensure good feature quality. Subsequently, the selected features are tracked by the KLT sparse optical flow with local intensity normalization. Finally, the inertial and visual measurements are jointly optimized by sliding window based nonlinear optimization to achieve six degrees of freedom motion estimation. The proposed method is validated on the public dataset and real-world experiments, which shows our approach reaches state-of-the-art performance.
Zhiao Cai, Ming Yang 0002, Bing Wang 0006
Intelligent Vehicles Symposium2
2018 Hybrid Filtering Framework Based Robust Localization for Industrial Vehicles
abstract
This paper presents a precise and robust localization framework for autonomous vehicles. In contrast to simultaneous localization and mapping, localization and mapping in our method are separate (i.e., first mapping and then localization within this map). The map used in this paper is a 3-D occupancy map that is generated from a stitched point cloud. For localization, a hybrid filtering framework is proposed to match the live data with the prior map. In the upper layer, odometer data, IMU data, and the map matching result are fused by a cubature Kalman filter, which will limit the predictive pose within a reasonable bound. In the lower layer, the map matching problem is converted into a point set registration problem that is solved using a particle filter, which will make the matching result robust to local minima. This hybrid scheme makes localization more robust to convergence to a local minimum, which is often encountered in the localization task for autonomous vehicles. This method can also guarantee decimeter-level precision in industrial environments. Experiments demonstrate the validity of this method and also show that it outperforms some state-of-the-art methods.
Liang Li 0010, Ming Yang 0002, Bing Wang 0006
IEEE Trans. Ind. Informatics2
2018 Cubature Split Covariance Intersection Filter-Based Point Set Registration
abstract
Point set registration is a basic but still open problem in numerous computer vision tasks. In general, there are more than one type of error sources for registration, for example, noise, outliers and false initialization may exist simultaneously. These errors could influence the registration independently and dependently. Previous works usually test performance under one of the two types of errors at one time, or they do not perform well under some extreme situations with both of the error sources. This work presents a robust point set registration algorithm under a filtering framework, which aims to be robust under various types of errors simultaneously. The point set registration problem can be cast into a non-linear state space model. We use a split covariance intersection filter (SCIF) to capture the correlation between the state transition and the observation (moving point set). The two above-mentioned types of errors can be represented as dependent and independent parts in the SCIF. The covariance of the two types of errors will be updated every iteration. Meanwhile, the non-linearity of the observation model is approximated by a cubature transformation. First, the recursive cubature split covariance intersection filter is derived based on the non-linear state space model. Then, we use this algorithm to solve the point set registration problem. This algorithm can approximate non-linearity by a third-order term and consider correlations between the process model and the observation model. Compared to other filtering-based methods, this algorithm is more robust and precise. Tests on both public datasets and experiments validate the precision and robustness of this algorithm to outliers and noise. Comparison experiments show that this algorithm outperforms state-of-the-art point set registration algorithms in certain respects.
Liang Li 0010, Ming Yang 0002, Bing Wang 0006
IEEE Trans. Image Process.2
2018 Rigid Point Set Registration Based on Cubature Kalman Filter and Its Application in Intelligent Vehicles
abstract
Point set registration is a key problem in intelligent vehicle localization and mapping. This paper presents a rigid point set registration algorithm based on the cubature Kalman filter (CKF). First, the point set registration problem is cast into the state space model assuming that the correspondence between these two point sets is previously unknown. Then, CKF is used to solve this nonlinear filtering problem. At every iterative step, all the points in the moving point set will be considered, and the corresponding points in the model point set will be updated accordingly. In the time update, the scale of the free space that can be explored is significant for this registration algorithm. Herein, continuous simulated annealing (CSA) is adopted to gradually optimize the covariance of the model noise to speed up convergence. Tests on public data sets show that the CKF-based point set registration algorithm is robust to outliers, noise, and initialization misalignment. Then, the application of this algorithm on intelligent vehicles and intelligent transportation systems is demonstrated through mapping and localization experiments. The precision and robustness are both validated compared with the traditional ICP, NDT, and CPD based point set registrations in the localization experiments. Thus, the contributions of this paper are threefold. First, a CKF scheme is utilized in the point set registration problem. Second, CSA serves as the local optimizer to speed up the convergence process and make it more accurate. Third, this filtering-based method is adopted in intelligent vehicles and SLAM applications.
Liang Li 0010, Ming Yang 0002, Bing Wang 0006
IEEE Trans. Intell. Transp. Syst.2
2017 An Online Approach for Gesture Recognition Toward Real-World Applications
Zhaoxuan Fan, Xu Zhao 0001, Wanli Jiang, Ming Yang 0002
ICIG (1)6
2017 Gaussian mixture model-signature quadratic form distance based point set registration
abstract
Point set registration is a long addressed problem in lots of pattern recognition tasks. This paper presents a robust point set registration algorithm based on optimization of distance between two probability distributions. A major problem encountered in the point to point algorithms is the definition of correspondence between two point sets. This paper follows the idea of some probability based point set registration methods and the point set is represented as Gaussian Mixture Models (GMMs). Through optimizing distance between the two GMMs, the rigid transformation (rotation and translation) between two point sets will be obtained while averting the trouble of finding correspondence. Previous studies used L2 distance, KL distance, etc. to measure similarity between two GMMs, the problem therein is the robustness to noise and outliers, especially when the covariance matrix is large or there exists local minimum. So in this work, the signature quadratic form distance is derived for the distribution similarity measurement. The contribution of this paper is as follows. First, we derive the signature quadratic form distance for GMMs similarity measurement. Second, the signature quadratic form distance is adopted to the point set registration algorithm. And performance of the proposed method compared with some existing widely used point set registration algorithms is also presented. Experimental results show precision and robustness of this algorithm. The results also demonstrate this algorithm outperforms some state-of-the-art point set registration algorithms in terms of noise, outliers, partial structures and misalignment initialization, etc.
Liang Li 0010, Ming Yang 0002, Bing Wang 0006
IROS2
2017 CNN based semantic segmentation for urban traffic scenes using fisheye camera
abstract
Semantic segmentation is an important step of visual scene understanding for autonomous driving. Recently, Convolutional Neural Network (CNN) based methods have successfully applied in semantic segmentation using narrow-angle or even wide-angle pinhole camera. However, in urban traffic environments, autonomous vehicles need wider field of view to perceive surrounding things and stuff, especially at intersections. This paper describes a CNN-based semantic segmentation solution using fisheye camera which covers a large field of view. To handle the complex scene in the fisheye image, Overlapping Pyramid Pooling (OPP) module is proposed to explore local, global and pyramid local region context information. Based on the OPP module, a network structure called OPP-net is proposed for semantic segmentation. The net is trained and evaluated on a fisheye image dataset for semantic segmentation which is generated from an existing dataset of urban traffic scenes. In addition, zoom augmentation, a novel data augmentation policy specially designed for fisheye image, is proposed to improve the net's generalization performance. Experiments demonstrate the outstanding performance of the OPP-net for urban traffic scenes and the effectiveness of the zoom augmentation.
Liuyuan Deng, Ming Yang 0002, Yeqiang Qian, Bing Wang 0006
Intelligent Vehicles Symposium2
2017 Self-adapting part-based pedestrian detection using a fish-eye camera
abstract
Nowadays, fish-eye cameras play an increasingly important role in intelligent vehicles because of its wide field of view. Using fish-eye camera, pedestrians around the vehicles could be monitored expediently, but the problem of pedestrian distortion has always existed. This paper creates a new warping pedestrian benchmark using imaging principle of the fish-eye camera based on ETH pedestrian benchmark. With this practical benchmark, warping pedestrians are trained differently according to the position in fish-eye images. A self-adapting part-based algorithm is proposed to detect pedestrian with different degrees of deformation. Moreover, GPU is used to accelerate the whole algorithm to guarantee the real-time performance. Experiments show that the algorithm has competitive accuracy.
Yeqiang Qian, Ming Yang 0002, Bing Wang 0006
Intelligent Vehicles Symposium2
2016 Road DNA based localization for autonomous vehicles
abstract
High-precision and reliable localization is current research focus in the area of autonomous vehicles. Previous studies rely on either high-cost sensors or some specific characteristics, which means that the methods are limited to only a bit given situations. In this paper, a road DNA based localization method is proposed. It could afford high-precision result and does not have the shortcomings of previous methods at the same time. The scenery on both sides of the roads are used to generate the prior-map. The map is presented as grid map by the joint probability of occupation and reflectivity. With this type of map, different environments show different properties, which means that this method is not limited to specific environments and is effective in most cases. It costs much less memory than the previous maps. The map and live road scene flatting are both generated by data collected by low-cost LIDAR. Normalized Information Distance is utilized to align the live road scene flatting with the road DNA. Experiments show the validation and precision of this method.
Liang Li 0010, Ming Yang 0002, Bing Wang 0006
Intelligent Vehicles Symposium2
2016 A robust terrain-based road vehicle localization algorithm
abstract
Terrain-based localization is an alternate to the global positioning system (GPS) in signal blocked areas. However, terrain-based localization technique may suffer from low accuracy or even fail when brake vibration occurs. This paper presents a real-time algorithm for vehicle localization which is robust against brake vibration. The input includes a reference map of pitch difference and measurements from rear wheel encoders and inertial measurement units (IMU). This method consists of two steps. In the first step, terrain map is generated using pitch difference at equidistant intervals. After that, the Bayesian inference and particle filters are adopted in the second step to identify the vehicle location during travel. To enhance system stability, we propose dynamic distributions of filter variances according to acceleration input. Experimental results demonstrate that the localization method with dynamic distributions can localize the vehicle quickly with high accuracy even when a quite severe shuddering happens.
Tianyi Li 0003, Ming Yang 0002, Xujin Zhou
Intelligent Vehicles Symposium2
2016 Curvature Map-Based Magnetic Guidance for Automated Vehicles in an Urban Environment
abstract
Magnetic guidance is commonly used in real applications of automated vehicles for its reliability. However, due to the downward view of magnet detectors, the control of a vehicle based on magnetic guidance is difficult, particularly in an urban environment. This paper proposes the strategy of using a curvature map of magnets to carry out look-ahead control for magnetic guidance-based vehicles. First, a Gauss curve fitting-based lateral offset estimation method is proposed in the magnet localization to improve the robustness to sensors' noise. Then, the curvature of the road is generated based on a magnet-tracking algorithm, which records passed road segments by tracking every magnet detected by the magnetic sensor. Magnet-tracking results describe the relative position of the vehicle with respect to the reference road, whereas the curvature map contains forward road information. The fusion of these two sources of information gives a full description of the local road segment. Furthermore, to suppress noises, a simple local road model with a general form is used to fuse tracking results and curvature map information. Compared with magnetic guidance methods based on accurately positioned magnets, the proposed curvature map-based method boasts a simple map-generation process, in which the vehicle only needs to be driven manually along the reference trajectory once. Experiments on real application scenarios have verified the effectiveness of the proposed methods.
Ming Yang 0002, Hao Li 0024, Bing Wang 0006
IEEE Trans. Intell. Transp. Syst.2
2015 Integrating visual selective attention model with HOG features for traffic light detection and recognition
abstract
Traffic light detection and recognition play a more important role in Advanced Driver Assistance Systems and driverless cars. This paper presents a method of integrating Visual Selective Attention (VSA) model with HOG features to solve the problem of detecting and recognizing traffic lights in complex urban environment. First of all, the VSA model is used to get candidate regions of the traffic lights. Then, the HOG features of the traffic lights and SVM classifier are used in these candidate regions to get precise regions of traffic lights. Within these regions, the color of traffic light is recognized according to the information in the gray-scale image of channel A. Experimental results show that the proposed method has strong robustness and high accuracy.
Ming Yang 0002, Zhengchen Lu
Intelligent Vehicles Symposium2
2014 A new approach for autonomous vehicle navigation in urban scenarios based on roadway Magnets
abstract
Magnetic guidance is a commonly used vehicle navigation solution in real applications due to its reliability. The vehicle control of magnetic guidance, however, is difficult because of the look-down property of road detecting sensors. This paper proposes a curvature map based approach to realize look-ahead control for magnetic guidance used in urban scenarios. The basis of the approach is a magnet tracking algorithm, which makes it possible to calculate the curvature of the passed road. The tracking algorithm is used not only to localize the vehicle but also to build the curvature map of the reference trajectory. Once the magnetic ruler detects a magnet, a magnet tracker is initialized and tracks the magnet in the vehicle coordinate. Then these tracking results combined with the curvature map are used to predict the upcoming road's curvature. The curvature map is obtained by running the tracking algorithm when the vehicle is driving along the magnetic trajectory by hand. Compared with existing methods, the algorithm predominates in implementation and robustness. Experiments on real application scenario have verified the effectiveness of the proposed idea.
Ming Yang 0002, Bing Wang 0006
Intelligent Vehicles Symposium2
2013 Split Covariance Intersection Filter: Theory and Its Application to Vehicle Localization
abstract
Data fusion is an important process in a variety of tasks in the intelligent transportation systems field. Most existing data fusion methods rely on the assumption of conditional independence or known statistics of data correlation. In contrast, the split covariance intersection filter (split CIF) was heuristically presented in literature, which aims at providing a mechanism to reasonably handle both known independent information and unknown correlated information in source data. In this paper, we provide a theoretical foundation for the split CIF. First, we clearly specify the consistency definition (coined as split consistency) for estimates in split form. Second, we provide a theoretical proof for the fusion consistency of the split CIF. Finally, we provide a theoretical derivation of the split CIF for the partial observation case. We also present a general architecture of decentralized vehicle localization, which serves as a concrete application example of the split CIF to demonstrate the advantages of the split CIF and how it can potentially benefit vehicle localization (noncooperative and cooperative). In general, this paper aims at providing a baseline for researchers who might intend to incorporate the split CIF (a useful tool for general data fusion) into their prospective research works.
Hao Li 0024, Fawzi Nashashibi, Ming Yang 0002
IEEE Trans. Intell. Transp. Syst.3
2010 CyberC3: A Prototype Cybernetic Transportation System for Urban Applications
abstract
In this paper, a prototype cybernetic transportation system called cybernetic technologies for cars in Chinese cities (CyberC3) is introduced. This system contains as many modules as necessary to evaluate the feasibility for its potential mass application in cities, including a central control room, five stations, three on-road monitoring cameras, three intelligent vehicles, and a green-energy power system. The entire system is centrally controlled and runs in two different modes—the “shuttle mode” and the “on-demand mode.” The control algorithm is divided into three logical layers: scheduling, planning, and executing. The scheduling layer manages the entire system, the planning layer navigates the vehicle, and the executing layer controls the vehicle in real time. The vehicle is powered by supercapacitor batteries. This system is a demonstration system, as well as a research platform, and has been open to the public in Shanghai, China, since May 2007, which has helped to evaluate the entire system and spread the concept of “cybernetic transportation systems” in China.
Tingkai Xia, Ming Yang 0002, Ruqing Yang
IEEE Trans. Intell. Transp. Syst.2
2009 Ground-Texture-Based Localization for Intelligent Vehicles
abstract
Localization is a critical problem in the research of intelligent vehicles. Although it can be achieved by using a real-time kinematic global positioning system (RTK-GPS, or fused with other methods such as dead reckoning), it may be unfeasible if every vehicle has to be equipped with such an expensive sensor. This paper proposes a ground-texture-based map-matching approach to address the localization problem. To reduce the effect of complicated illumination in outdoor environments, a camera is fixed downward at the bottom of a vehicle, and controllable lights are also equipped around the camera for consistent illumination. The proposed approach includes two steps: 1) mapping and 2) localization. RTK-GPS is only used in the mapping, and other sensor data from camera and odometry are captured with time stamps to create a global ground texture map. A multiple-view registration-based optimization algorithm is applied to improve map accuracy. In the localization step, vehicle pose is estimated by matching the current camera frame with the best submap frame and by fusion strategy. Results with both synthetic and real experiments prove the feasibility and effectiveness of the proposed approach.
Ming Yang 0002, Ruqing Yang
IEEE Trans. Intell. Transp. Syst.3
2009 Conflict-Probability-Estimation-Based Overtaking for Intelligent Vehicles
abstract
Overtaking is a complex and hazardous driving maneuver for intelligent vehicles. When to initiate overtaking and how to complete overtaking are critical issues for an overtaking intelligent vehicle. We propose an overtaking control method based on the estimation of the conflict probability. This method uses the conflict probability as the safety indicator and completes overtaking by tracking a safe conflict probability. The conflict probability is estimated by the future relative position of intelligent vehicles, and the future relative position is estimated by using the dynamics models of the intelligent vehicles. The proposed method uses model predictive control to track a desired safe conflict probability and synthesizes decision making and control of the overtaking maneuver. The effectiveness of this method has been validated in different experimental configurations, and the effects of some parameters in this control method have also been investigated.
Fenghui Wang, Ming Yang 0002, Ruqing Yang
IEEE Trans. Intell. Transp. Syst.2