Yizhuo Yang 0001

dblp:37/11052-1 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2025
0009-0000-9139-1542ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Unsupervised UAV 3D Trajectories Estimation with Sparse Point Clouds
abstract
Compact UAV systems, while advancing delivery and surveillance, pose significant security challenges due to their small size, which hinders detection by traditional methods. This paper presents a cost-effective, unsupervised UAV detection method using spatial-temporal sequence processing to fuse multiple LiDAR scans for accurate UAV tracking in real-world scenarios. Our approach segments point clouds into foreground and background, analyzes spatial-temporal data, and employs a scoring mechanism to enhance detection accuracy. Tested on a public dataset, our solution placed 4th in the CVPR 2024 UG2+ Challenge, demonstrating its practical effectiveness. We plan to open-source all designs, code and sample data for the research community @ github.com/lianghanfang/UnLiDAR-UAV-Est.
Hanfang Liang, Yizhuo Yang 0001, Jinming Hu, Jianfei Yang 0001, Shenghai Yuan 0001
ICASSP2
2025 Learning Dynamic Weight Adjustment for Spatial-Temporal Trajectory Planning in Crowd Navigation
abstract
Robot navigation in dense human crowds poses a significant challenge due to the complexity of human behavior in dynamic and obstacle-rich environments. In this work, we propose a dynamic weight adjustment scheme using a neural network to predict the optimal weights of objectives in an optimization-based motion planner. We adopt a spatial-temporal trajectory planner and incorporate diverse objectives to achieve a balance among safety, efficiency, and goal achievement in complex and dynamic environments. We design the network structure, observation encoding, and reward function to effectively train the policy network using reinforcement learning, allowing the robot to adapt its behavior in real time based on environmental and pedestrian information. Simulation results show improved safety compared to the fixed-weight planner and the state-of-the-art learning-based methods, and verify the ability of the learned policy to adaptively adjust the weights based on the observed situations. The feasibility of the approach is demonstrated in a navigation task using an autonomous delivery robot across a crowded corridor over a 300 m distance. Video: https://youtu.be/nSCbNaaF_VM
Muqing Cao, Xinhang Xu, Yizhuo Yang 0001, Jianping Li 0004, Tongxing Jin, Tzu-Yi Hung, Guosheng Lin, Lihua Xie 0001
ICRA3
2025 ULOC: Learning to Localize in Complex Large-Scale Environments with Ultra-Wideband Ranges
abstract
While UWB-based methods can achieve high localization accuracy in small-scale areas, their accuracy and reliability are significantly challenged in large-scale environments. In this paper, we propose a learning-based framework named ULOC for Ultra-Wideband (UWB) based localization in such complex, large-scale environments. First, anchors are deployed in the environment without knowledge of their actual position. Then, UWB observations are collected when the vehicle travels in the environment. At the same time, map-consistent pose estimates are developed from registering onboard self-localization data (from VIO, LIO, and other SLAM methods) with the prior map to provide the training labels. We then propose a network based on MAMBA that learns the ranging patterns of UWBs over a complex, large-scale environment. The experiment demonstrates that our solution can ensure high localization accuracy on a large scale compared to the state-of-the-art. We release our source code to benefit the community at https://github.com/brytsknguyen/uloc.
Thien-Minh Nguyen, Yizhuo Yang 0001, Tien-Dat Nguyen, Shenghai Yuan 0001, Lihua Xie 0001
ICRA2
2024 MMAUD: A Comprehensive Multi-Modal Anti-UAV Dataset for Modern Miniature Drone Threats
abstract
In response to the evolving challenges posed by small unmanned aerial vehicles (UAVs), which possess the potential to transport harmful payloads or independently cause damage, we introduce MMAUD: a comprehensive Multi-Modal Anti-UAV Dataset. MMAUD addresses a critical gap in contemporary threat detection methodologies by focusing on drone detection, UAV-type classification, and trajectory estimation. MMAUD stands out by combining diverse sensory inputs, including stereo vision, various Lidars, Radars, and audio arrays. It offers a unique overhead aerial detection vital for addressing real-world scenarios with higher fidelity than datasets captured on specific vantage points using thermal and RGB. Additionally, MMAUD provides accurate Leica-generated ground truth data, enhancing credibility and enabling confident refinement of algorithms and models, which has never been seen in other datasets. Most existing works do not disclose their datasets, making MMAUD an invaluable resource for developing accurate and efficient solutions. Our proposed modalities are cost-effective and highly adaptable, allowing users to experiment and implement new UAV threat detection tools. Our dataset closely simulates real-world scenarios by incorporating ambient heavy machinery sounds. This approach enhances the dataset’s applicability, capturing the exact challenges faced during proximate vehicular operations. It is expected that MMAUD can play a pivotal role in advancing UAV threat detection, classification, trajectory estimation capabilities, and beyond. Our dataset, codes, and designs will be available in https://ntu-aris.github.io/MMAUD.
Shenghai Yuan 0001, Yizhuo Yang 0001, Thien Hoang Nguyen, Thien-Minh Nguyen, Jianfei Yang 0001, Jianping Li 0004, Han Wang 0001, Lihua Xie 0001
ICRA2
2023 AV-PedAware: Self-Supervised Audio-Visual Fusion for Dynamic Pedestrian Awareness
abstract
In this study, we introduce AV-PedAware, a self-supervised audio-visual fusion system designed to improve dynamic pedestrian awareness for robotics applications. Pedestrian awareness is a critical requirement in many robotics applications. However, traditional approaches that rely on cameras and LIDARs to cover multiple views can be expensive and susceptible to issues such as changes in illumination, occlusion, and weather conditions. Our proposed solution replicates human perception for 3D pedestrian detection using low-cost audio and visual fusion. This study represents the first attempt to employ audio-visual fusion to monitor footstep sounds for the purpose of predicting the movements of pedestrians in the vicinity. The system is trained through self-supervised learning based on LIDAR-generated labels, making it a cost-effective alternative to LIDAR-based pedestrian awareness. AV-PedAware achieves comparable results to LIDAR-based systems at a fraction of the cost. By utilizing an attention mechanism, it can handle dynamic lighting and occlusions, overcoming the limitations of traditional LIDAR and camera-based systems. To evaluate our approach's effectiveness, we collected a new multimodal pedestrian detection dataset and conducted experiments that demonstrate the system's ability to provide reliable 3D detection results using only audio and visual data, even in extreme visual conditions. We will make our collected dataset and source code available online for the community to encourage further development in the field of robotics perception systems.
Yizhuo Yang 0001, Shenghai Yuan 0001, Muqing Cao, Jianfei Yang 0001, Lihua Xie 0001
IROS1
2022 Overcoming Catastrophic Forgetting for Semantic Segmentation Via Incremental Learning
abstract
Deep learning based semantic segmentation models have achieved remarkable results in recent years. However, many deep learning based models encounter the problem of catastrophic forgetting, i.e. when the model is required to learn a new task without labels for old objects, its performance drops significantly for the previous tasks. To solve this problem, an incremental learning method, a Combination of Old Prediction and Modified Label (COPML), is developed in this paper. The proposed method utilizes the prediction results of the old model and the modified labels of the new task to create pseudo labels which are close to the ground truths. By using these pseudo labels for training, the model is expected to preserve the knowledge of old tasks. In addition, knowledge distillation, the replay and parameter freezing strategy are also applied to the proposed method to further assist the model in overcoming catastrophic forgetting. The effectiveness of the proposed method is validated on two semantic segmentation models: Unet and Deeplab3 in Pascal- VOC 2012 dataset and a self-made dataset. The experimental results demonstrate that COPML enables the model to maintain most of the old knowledge while obtaining an excellent performance on a new task.
Yizhuo Yang 0001, Shenghai Yuan 0001, Lihua Xie 0001
ICARCV1