Hanyang Zhuang

dblp:295/0775 · DBLP profile ↗
← Back
15ranked-venue papers
2as first author
15since 2021 · last 2026
0000-0001-6668-9523ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Systems, architecture and hardware · 4 · 4 since 2021
YearPublicationVenuePosition
2026 Evaluating Roadside LiDAR Placement with a Probability-Based Surrogate Metric
Zhiqi Qi, Xintong Dong, Hanyang Zhuang, Xin Xia 0007
IV3
2026 An Automatic Calibration Method for Low-Overlapped Roadside RGB-D Camera Network
abstract
Roadside multi-sensor networks, such as RGB-D camera networks, play a crucial role in Intelligent Transportation Systems (ITS). These systems rely on accurate extrinsic parameters (i.e., the relative positions and orientations) of each camera in the network. However, achieving fast and accurate large-scale extrinsic parameter calibration is challenging, especially when the overlap between camera views is limited due to cost constraints. To address this issue, we propose an automated, scalable, and marker-free calibration method that requires no human intervention. Our approach leverages dynamic rigid bodies, such as moving vehicles, as bridges to establish associations among all cameras without manual placement or supervision. The proposed method consists of three main stages: 1) calibrating camera height, roll, and pitch using the ground plane; 2) estimating each camera’s 2D position and yaw angle based on trajectories; and 3) refining these estimates by matching features on the road surface using SuperPoint and SuperGlue. Real-world experiments involving 58 roadside RGB-D cameras deployed in a parking lot demonstrate that our method significantly improves calibration efficiency while maintaining high accuracy, making it well-suited for large-scale RGB-D camera network deployment.
Jingda Chen, Hanyang Zhuang, Yuesheng He, Ming Yang 0002
IEEE Trans. Intell. Transp. Syst.2
2026 Enhancing Robotaxi Pick-Up Through Vision-Language-Model-Based Passenger Identification
abstract
Robotaxis are emerging as a key component of urban transportation. However, most current robotaxis rely on preset pick-up points, which cluster robotaxi activity and intensify local traffic congestion, increase deadheading and detour distances, extend roadway occupancy, and raise operating costs. GNSS-based localization may be unreliable especially in urban area. Generally, human drivers can communicate with passengers through text/voice interactions to accurately locate and reach the human’s location. How to align the robotaxi’s visual recognition results with human text/voice description is the key challenge. Therefore, we propose a comprehensive framework, including its architecture, requirment, opertional logic, workflow, to enhance the passenger-robotaxi interaction for accurate passenger pick-up without preset pick-up points. In this framework, the key algorithm is VLMIdentification, a realtime human identification model based on LVLM. VLMIdentification comprises three modules: i) Human Input Processing, which extracts textual features from passenger text/voice and converts human-centric descriptions into third-person, robotaxi-centric attributes; ii)Candidate Searching, which couples conventional detectors with an LVLM to adapt detection thresholds to scene complexity and to translate detections into textual descriptors; and iii) Human Identification, which performs matching between the processed passenger description and candidates to locate the correct human. We defined the human identification with multimodal task, proposed evaluation metrics, and built a new HID (Human Identification with Description) dataset based on existing autonomous driving datasets. Experimental results demonstrate that VLMIdentification outperforms baseline in comprehensive metrics and retains robust, performance under adverse environment and cross-scenario generalization tests, thereby confirming its generalization and robustness. Code is available athttps://github.com/fanwu66/VLMIdentification
Shaojing Song, Qinge Wu, Zhiqing Miao, Hanyang Zhuang
IEEE Trans. Intell. Transp. Syst.6
2026 APC-Scheduler: Allocation-Planning-Charging Integrated Scheduling System for Cooperative Automated Valet Parking
abstract
The Automated Valet Parking system (AVP), as one of the promising technologies, offers significant benefits in saving maneuver time and parking cost. With the increase number of vehicles using AVP system in large-scale parking lots, the overall efficiency is limited by the selfish decision-making of each individual; therefore, cooperative-AVP (C-AVP) is developed to achieve global optimization by scheduling the vehicles. Existing C-AVP methods focus only on one of the key processes in AVP, such as space allocation, trajectory planning, and electric vehicle(EV) charging. However, these factors interact with each other and are rarely considered in an entire system. Therefore, this paper aims to model the space allocation, trajectory planning, and EV charging problem in a whole framework, named APC-scheduler, and adopt it to improve the overall efficiency for large-scale parking lot with EV. The APC-scheduler comprises two parts: 1) a zone-based parking space allocation module dynamically assigns parking spaces based on real-time conditions using a hierarchical optimization strategy. It contains a charging priority estimation process to determine the EV’s charging requirement urgency; 2) a conflict-based trajectory planning module is developed to reduce trajectory overlaps and vehicle conflicts. It uses spatial-temporal path planner and speed planner to eliminate conflicts. Experiments under reasonable vehicle arrival and departure statistics have been conducted in a large-scale parking lot with 526 parking spaces. The results demonstrate that the proposed method effectively enhances parking efficiency and minimizes conflicts, particularly under high vehicle arrival frequencies and in dense traffic conditions.
Hanyang Zhuang, Qizhe Xu, Ming Yang 0002
IEEE Trans. Intell. Transp. Syst.1
2025 Pose Tracking of Leading Vehicle Using Mass-Produced Sensors
abstract
Vehicle-following presents significant advantages in flexible scenarios such as platooning and valet parking, where a human-driven leader navigates complex environments, and an autonomous follower replicates the leader's trajectory. The key objective is accurate, robust, long-term pose tracking of the leader. This paper proposes a framework using mass-produced sensors, including a front fisheye camera, front millimeter-wave radar, and wireless communication. The framework consists of three modules: 1) a Visual Feature Tracking module using a DiMP tracker for robust taillight feature tracking; 2) a Visual-radar Fusion module to estimate the rear center of the leader vehicle for reference point tracking; 3) a Leader Pose Tracking module combining sensor data and vehicle-to-vehicle communication in a particle filter framework for precise pose tracking in three degrees of freedom. Real-world experiments show lateral and longitudinal errors of 0.056 meters and 0.165 meters, respectively.
Hanyang Zhuang, Ming Yang 0002
IV1
2025 SLAM2: Simultaneous Localization and Multimode Mapping for indoor dynamic environments
abstract
Traditional visual Simultaneous Localization and Mapping (SLAM) methods based on point features are often limited by strong static assumptions and texture information, resulting in inaccurate camera pose estimation and object localization. To address these challenges, we present SLAM 2 , a novel semantic RGB-D SLAM system that can obtain accurate estimation of the camera pose and the 6DOF pose of other objects, resulting in complete and clean static 3D model mapping in dynamic environments. Our system makes full use of the point, line, and plane features in space to enhance the camera pose estimation accuracy. It combines the traditional geometric method with a deep learning method to detect both known and unknown dynamic objects in the scene. Moreover, our system is designed with a three-mode mapping method, including dense, semi-dense, and sparse, where the mode can be selected according to the needs of different tasks. This makes our visual SLAM system applicable to diverse application areas. Evaluation in the TUM RGB-D and Bonn RGB-D datasets demonstrates that our SLAM system achieves the most advanced localization accuracy and the cleanest static 3D mapping of the scene in dynamic environments, compared to state-of-the-art methods. Specifically, our system achieves a root mean square error (RMSE) of 0.018 m in the highly dynamic TUM w/half sequence, outperforming ORB-SLAM3 (0.231 m) and DRG-SLAM (0.025 m). In the Bonn dataset, our system demonstrates superior performance in 14 out of 18 sequences, with an average RMSE reduction of 27.3% compared to the next best method. • Semantic vSLAM fusing point, line and plane improves perception and camera pose estimation. • Dense, semi-dense and sparse modes for flexible mapping in dynamic scenes. • 6DOF pose estimation of objects for enriched semantic information in mapping. • Semantic and motion fusion enables advanced pose estimation in dynamic environments. • Depth and plane-based refinement for cleaner global static 3D point cloud maps.
Qi Zhang 0072, Zhen Tian 0002, Peizhuo Yu, Hanyang Zhuang, Jianglin Lan
Pattern Recognit.6
2025 A Conflicts-Free, Speed-Lossless KAN-Based Reinforcement Learning Decision System for Interactive Driving in Roundabouts
abstract
Safety and efficiency are crucial for autonomous driving in roundabouts, especially mixed traffic with both autonomous vehicles (AVs) and human-driven vehicles. This paper presents a learning-based algorithm that promotes safe and efficient driving across varying roundabout traffic conditions. A deep Q-learning network is used to learn optimal strategies in complex multi-vehicle roundabout scenarios, while a Kolmogorov-Arnold Network (KAN) improves the AVs’ environmental understanding. To further enhance safety, an action inspector filters unsafe actions, and a route planner optimizes driving efficiency. Moreover, model predictive control ensures stability and precision in execution. Experimental results demonstrate that the proposed system consistently outperforms state-of-the-art methods, achieving fewer collisions, reduced travel time, and stable training with smooth reward convergence.
Zhen Tian 0002, Jianglin Lan, Qi Zhang 0072, Hanyang Zhuang, Xianxian Zhao
IEEE Trans. Intell. Transp. Syst.6
2025 CrossGLoc: Cross-Modal Global Localization Leveraging Pretrained Diffusion Models and Semantic Cues for Intelligent Vehicles
abstract
Cross-modal global localization matches visual information with pre-built LiDAR maps, which has attracted more and more attention for its low cost and potential robustness. However, the inherent modality difference between images and point clouds makes it challenging. This paper proposes a novel cross-modal global localization system, named CrossGLoc, which leverages pre-trained diffusion models and semantic cues to address this challenge. The main idea is leveraging the semantic cues shared between different modalities to bridge the modality gap, and utilizing pre-trained diffusion models to extract modality-consistent high-dimensional features guided by these semantic cues. To achieve this, ControlNet is used to generate intermediate feature maps from semantic images and semantic map projections, and a semantic categories-based feature aggregation algorithm is proposed to aggregate these feature maps into global descriptors. Furthermore, a semantic edge key points-based pose estimation algorithm is proposed to estimate the pose of retrieved image and point cloud pairs. Extensive experiments on the KITTI dataset, the KITTI360 dataset and the self-collected dataset demonstrate that the proposed method achieves state-of-the-art performance in cross-modal global localization.
Hengwang Zhao, Qiyuan Shen, Hanyang Zhuang, Tong Qin 0001, Ming Yang 0002
IEEE Trans. Intell. Transp. Syst.3
2024 Cross-Modal Registration Using Adaptive Modeling in Infrastructure-based Vehicle Localization
abstract
Infrastructure-based vehicle localization, in comparison to single-agent approaches, offers several advantages including reduced system cost, extended perception range, enhanced data fusion capabilities, and energy savings. Many conventional approaches impose limitations on the types of objects due to the need for specific object-end modifications, such as applying perceptual markers like color-labeled plates and reflective balls. LiDAR presents a solution in terms of object arbitrariness, as it addresses the challenges of feature-free object modeling and continuous registration. However, achieving complete environmental coverage with LiDAR remains prohibitively expensive, particularly in extensive areas. Hence, this study proposes a cross-modal localization approach using adaptive modeling, employing LiDAR for object modeling and cost-effective sensor cameras for object tracking through image-point-cloud registration. Accurate correspondence between the model and observation can be estimated in real-time. The experiments are conducted in a typical scenario that requires adaptive modeling: Autonomous Valet Parking (AVP). Results demonstrate that the proposed system achieves comparable performance with significantly reduced system costs, highlighting its potential for large-scale deployment.
Yuesheng He, Hanyang Zhuang, Chenxi Yang 0002, Ming Yang 0002
ICRA3
2024 An Online Automatic Calibration Method for Infrastructure-Based LiDAR-Camera via Cross-modal Object Matching
abstract
In indoor environments where the Global Navigation Satellite System (GNSS) isn’t available, the infrastructure-based LiDAR-camera joint array can provide high-precision localization for mobile robots, such as Autonomous Valet Parking (AVP). The primary challenge in employing the infrastructure-based LiDAR-camera joint array is the extrinsic calibration between the LiDAR and the camera. Moreover, to handle interference deviation caused by vibrations or inadequate mounting stiffness during operation, the calibration’s extrinsic parameters must be automatically updated online, presenting higher demands for infrastructure-based LiDAR-camera extrinsic calibration. This paper proposes an infrastructure LiDAR-camera online automatic calibration method based on prior knowledge of cross-modal target registration. This method requires no manual targets and initial pose guesses and can achieve extrinsic calibration. The object-prior model based on a lightweight object detection algorithm can rapidly detect scenes favorable for extrinsic calibration in sub-images of camera images. This creates favorable conditions for the registration of cross-modal networks and poses optimization of the LiDAR camera. Additionally, because a lightweight algorithm is used, the process does not compromise efficiency or consume excessive computational resources. Experimental results demonstrate that the proposed calibration method is suitable for calibrating infrastructure-based LiDAR-camera, with comparable accuracy and the ability to perform online calibration. Comparative experiments also show that the object-prior model can indeed select better scenes for LiDAR-camera extrinsic calibration, thus improving the accuracy and stability of extrinsic calibration to some extent.
Yuesheng He, Hanyang Zhuang, Ming Yang 0002
IROS3
2024 Active Vehicle Re-localization Based on Non-repetitive LiDAR with Gimbal Motion Strategy
abstract
The installation of a multi-layer 3D LiDAR atop the vehicle is a widely adopted hardware configuration for map-matching-based localization in intelligent driving. By offering a comprehensive 360° horizontal Field of View (FoV), this setup aims to achieve precise matching outcomes through the imposition of substantial geometric constraints against dynamic interferences and structural degradation. However, several factors limit its environmental adaptability, such as sparse point cloud density at distances, insufficient maximum sensing range, and notably, the restricted beam elevation angle, limiting the perception of the environment beyond obstacles. The rapid advancement of non-repetitive scanning LiDARs shows promise in mitigating such limitations. Nevertheless, their narrow FoV remains a challenge to overcome. In this study, we propose a solution by mounting such one single LiDAR on a two-axis rotating gimbal, enabling the vehicle to surpass the ranges and vertical FoV limitations of traditional setups actively. The corresponding gimbal motion strategy has been designed to automatically focus on the environment component with the most robust geometric constraints. Experimental results validate that the proposed method achieves superior robustness under high dynamic interference while delivering sufficient performance under standard conditions.
Xin'Ao Wu, Chenxi Yang 0002, Yiyang Guo, Hanyang Zhuang, Ming Yang 0002
IROS4
2023 Pseudo-Anchors: Robust Semantic Features for Lidar Mapping in Highly Dynamic Scenarios
abstract
Dynamic environments are challenging for anchor-free mapping using lidar in intelligent driving. This study imitates anchor-based approaches such as magnetic nails by applying novel Static Confidence Criteria (SCC) to the point-cloud semantic candidates to ensure their robustness. We name such verified features Pseudo-Anchors (P-A) as they hold similar properties to the anchor nodes: The P-A nodes are improbably formed by dynamic objects, and nodes’ blockage state can be immediately noticed once they are occluded. Another major challenge for mapping is improving large-scale global performance without sacrificing local consistency. Unrecognized GNSS pose drift may deteriorate local trajectory accuracy through post-processing such as graph optimization. In this study, we use the road network to provide the intersection information as a prior so that the GNSS can be better regarded as a reliable anchor factor. Three experiments are designed for this study. The first is ablations to verify the P-A concept; The second proves that the P-A-based lidar odometry outperformed the LOAM-based mainstream methods in highly dynamic scenarios; The third shows that our usage of the GNSS strengthens large-scale maps’ global consistency while causing less deterioration towards the local one. As a knowledge-based method, the P-A concept shows a high deployment efficiency, indicating the potential for migration to other features or even other sensors.
Chenxi Yang 0002, Lei He 0019, Hanyang Zhuang, Ming Yang 0002
IEEE Trans. Intell. Transp. Syst.3
2023 Global Pose Initialization Based on Gridded Gaussian Distribution With Wasserstein Distance
abstract
Feature descriptors, as abstraction of the critical information in the lidar point clouds, are often used in global pose initialization in large-area to provide a pose reference for intelligent driving system. The current state-of-the-art method Scan Context descriptor is generated based on points’ maximum height and, therefore, designed especially for outdoor scenarios without a ceiling. This study proposes a generic descriptor for both outdoor and indoor scenarios based on the point cloud’s Gridded Gaussian Distribution (GGD). Wasserstein distance is introduced to this field to evaluate the proposed GGD descriptors’ matching performance because it not only has a solid mathematical foundation in comparing two Gaussian distributions but also shows excellent time efficiency via a straightforward analytical solution without enumeration or iteration. We construct a multi-step error function to initialize the vehicle pose using conventional cosine similarity, and Wasserstein distance. Two experiments are designed for this study. The first experiment compares the pose initialization performances under various multi-frame superimposition distance in space to find an efficient GGD descriptor extraction setting. The second experiment verifies that the proposed method achieves a better pose initialization success rate than the mainstream methods.
Chenxi Yang 0002, Hanyang Zhuang, Ming Yang 0002
IEEE Trans. Intell. Transp. Syst.3
2022 HR-Planner: A Hierarchical Highway Tactical Planner based on Residual Reinforcement Learning
abstract
Tactical planning is crucial for safe and efficient driving on the highway. However, the problem is complicated by the uncertain intention of surrounding vehicles, as well as observation noise caused by measurement noise and perception errors. Rule-based tactical planning methods are ineffective in handling dynamic scenarios with uncertainty, and susceptible to observation noise. To tackle this problem, we propose a hierarchical tactical planning framework based on residual reinforcement learning. Besides, a new reinforcement learning from demonstrations scheme that views rule-based methods as soft guidance is developed to combine prior knowledge with data-driven methods. Based on the framework and the training scheme, rule-based methods not only can be improved in highway scenarios with uncertainty and observation noise, but also will guide the training procedure for increased sampling efficiency. Additionally, to boost in-depth and consistent exploration in a vehicle system with inertia, we employ noisy networks to explore the optimal policy. The proposed method is validated in a stochastic and uncertain simulation environment, and the results reveal that our method outperforms both rule-based methods and pure data-driven methods in terms of safety and driving efficiency under noisy observations and uncertainty.
Yueyuan Li, Hanyang Zhuang, Ming Yang 0002
ICRA3
2022 AGBM: An Adaptive Gradient Balanced Mechanism for the End-to-End Steering Estimation
abstract
End-to-end steering estimation is one of the important deep regression tasks. However, driving datasets are always imbalanced on the distribution of the steering value, which makes end-to-end learning models perform bad steering estimation accuracy, especially for the sharp steering value estimation, which finally, leads to an imbalanced training problem. In this paper, the essential reason for the imbalanced training problem for the deep regression task is first revealed as the gradient disharmony. To hedge the gradient disharmony, a general gradient balanced mechanism for the deep regression task is proposed in this paper. Meanwhile, the general gradient balanced mechanism is applied for the steering estimation task. According to the steering distribution, a novel Adaptive Gradient Balanced Mechanism (AGBM) is proposed to hedge the gradient disharmony based on the Laplace distribution with the general gradient balanced mechanism theory. Further, a new loss function embedding with AGBM is proposed to balance the gradient disharmony for the training process. AGBM makes the process of gradient hedging with adaptive style without parameters finetuning. Based on five driving datasets, four experiments are conducted for demonstrations. The experiment results show the AGBM enables the end-to-end models to perform the best accuracy of the steering estimation, including the sharp steering samples. Meanwhile, AGBM is generalized with different end-to-end learning models.
Wei Yuan 0002, Hanyang Zhuang, Ming Yang 0002
IEEE Trans. Intell. Transp. Syst.2