EDBT 2026 Demo / reviewers in the wild / expert
Weihao Yan 0001
dblp:272/6222-1
· DBLP profile ↗
7ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0002-7838-0790ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 4 since 2021Systems, architecture and hardware · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning-Based 3D Reconstruction in Autonomous Driving: A Comprehensive SurveyabstractLearning-based 3D reconstruction has emerged as a transformative technique in autonomous driving, enabling precise modeling of environments through advanced neural representations. It has inspired pioneering solutions for vital tasks in autonomous driving, such as dense mapping and closed-loop simulation, as well as comprehensive scene feature for driving scene understanding and reasoning. Given the rapid growth in related research, this survey provides a comprehensive review of both technical evolutions and practical applications in autonomous driving. We begin with an introduction to the preliminaries of learning-based 3D reconstruction to provide a solid technical background foundation, then progress to a rigorous, multi-dimensional examination of cutting-edge methodologies, systematically organized according to the distinctive technical requirements and fundamental challenges of autonomous driving. Through analyzing and summarizing development trends and cutting-edge research, we identify existing technical challenges, along with insufficient disclosure of on-board validation and safety verification details in the current literature, and ultimately suggest potential directions to guide future studies. Liewen Liao, Weihao Yan 0001, Ming Yang 0002, Songan Zhang, H. Eric Tseng |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | AdaptiveOcc: Adaptive Octree-Based Network for Multi-Camera 3D Semantic Occupancy Prediction in Autonomous DrivingabstractMulti-camera 3D semantic occupancy prediction is a critical task for autonomous driving, playing a vital role in understanding the environment. Current methods mainly rely on uniform voxel representation to encode space, which greatly limits their resolution scalability. It causes most existing methods to struggle with scaling to finer granularities, as the cubic growth nature of uniform voxel leads to a significant increase in the demand for computational and storage resources when scaling. To address this, we propose a multi-level hierarchical model AdaptiveOcc. Using the octree structure, our model can adaptively represent different parts of space with varying voxel granularity. It can selectively extend resolution only for a small subset of voxels, thus mitigating the substantial computational and storage burden brought by scaling. To endow our model with adaptability, we propose a distance-adaptive octree construction rule for generating supervised labels. Considering that the voxel granularity requirements vary for different distance ranges in environmental perception, such a construction rule results in a higher likelihood of coarser granularity for distant regions and finer granularity for nearby regions. This ensures a more efficient and rational allocation of computational resources, further reducing the inference latency. Extensive experiments on nuScenes, SemanticKITTI and Waymo dataset validate that our method can scale to finer granularities with faster speed, and less training memory compared with other state-of-the-art methods. Our code is available athttps://github.com/yty-sky/AdaptiveOcc. Yeqiang Qian, Weihao Yan 0001, Ming Yang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | MOSFormer: A Transformer-based Multi-Modal Fusion Network for Moving Object Segmentationabstract3D moving object segmentation (MOS) is vital for autonomous systems, providing essential information for downstream tasks like mapping and localization. However, current MOS methods face challenges due to the limitation of existing datasets, which are sparse in moving objects and limited in scene diversity. Meanwhile, the prevalent methods are projection-based, struggling with the challenge of blurred boundaries. To tackle the dataset issue, we introduce a nuScenes-based MOS dataset, which provides richer scenes and more dynamic instances. To alleviate the boundary blur-ring issue and further improve accuracy and generalizability, we propose a dual-branch multimodal fusion MOS network, MOSFormer. The Transformer structure is incorporated to extract spatio-temporal information better, while image semantic information is utilized to refine the boundaries of moving objects. Finally, experiments on two datasets show that our method achieves state-of-the-art performance, and a mapping experiment with our method confirms its effectiveness in downstream tasks such as mapping and localization. Zike Cheng, Hengwang Zhao, Qiyuan Shen, Weihao Yan 0001, Ming Yang 0002 |
IROS | 4 |
| 2024 | Cross-Modal Visual Relocalization in Prior LiDAR Maps Utilizing Intensity TexturesabstractCross-modal localization has drawn increasing attention in recent years, while the visual relocalization in prior LiDAR maps is less studied. Related methods usually suffer from inconsistency between the 2D texture and 3D geometry, neglecting the intensity features in the LiDAR point cloud. In this paper, we propose a cross-modal visual relocalization system in prior LiDAR maps utilizing intensity textures, which consists of three main modules: map projection, coarse retrieval, and fine relocalization. In the map projection module, we construct the database of intensity channel map images leveraging the dense characteristic of panoramic projection. The coarse retrieval module retrieves the top-K most similar map images to the query image from the database, and retains the top-K’ results by covisibility clustering. The fine relocalization module applies a two-stage 2D-3D association and a covisibility inlier selection method to obtain robust correspondences for 6DoF pose estimation. The experimental results on our self-collected datasets demonstrate the effectiveness in both place recognition and pose estimation tasks. Qiyuan Shen, Hengwang Zhao, Weihao Yan 0001, Tong Qin 0001, Ming Yang 0002 |
IROS | 3 |
| 2024 | Non-Repetitive: A Promising LiDAR Scanning PatternabstractLiDAR is an essential sensor for intelligent vehicles. Recently, LiDARs used in vehicles produced by different companies have significant differences in their scanning patterns. Some vehicles use mechanical and solid-state (repetitive) LiDARs, while others use prism-based (non-repetitive) LiDARs. The scanning pattern of a LiDAR has a profound impact on its scanning performance. To investigate the influence of LiDAR scanning patterns, we created the "Repetitive-or-not" dataset, which is collected simultaneously by LiDARs with both repetitive and non-repetitive scanning patterns in the CARLA simulation environment. Using this dataset, we conducted a comprehensive statistical analysis of the scanning ability of repetitive and non-repetitive LiDARs. Furthermore, we looked into the effects of these two LiDAR scanning patterns on the performance of various 3D object detection algorithms. Finally, we explored the domain gap in the point cloud data produced by repetitive and non-repetitive LiDARs. Through an in-depth investigation of the "Repetitive-or-not" dataset, we have discovered that non-repetitive LiDAR shows great promise. This conclusion is primarily supported by its superior object scanning capabilities. Angchen Xie, Yeqiang Qian, Weihao Yan 0001, Ming Yang 0002 |
IROS | 3 |
| 2023 | TTC4MCP: Monocular Collision Prediction Based on Self-Supervised TTC EstimationabstractVision-based collision prediction for autonomous driving is a challenging task due to the dynamic movement of vehicles and diverse types of obstacles. Most existing methods rely on object detection algorithms, which only predict predefined collision targets, such as vehicles and pedestrians, and cannot anticipate emergencies caused by unknown obstacles. To address this limitation, we propose a novel approach using pixel-wise time-to-collision (TTC) estimation for monocular collision prediction (TTC4MCP). Our approach predicts TTC and optical flow from monocular images and identifies potential collision areas using feature clustering and motion analysis. To overcome the challenge of training TTC estimation models without ground truth data in new scenes, we propose a self-supervised TTC training method, enabling collision prediction in a wider range of scenarios. TTC4MCP is evaluated on multiple road conditions and demonstrates promising results in terms of accuracy and robustness. Yeqiang Qian, Weihao Yan 0001, Ming Yang 0002 |
IROS | 4 |
| 2023 | Threshold-Adaptive Unsupervised Focal Loss for Domain Adaptation of Semantic SegmentationabstractSemantic segmentation is an important task for intelligent vehicles to understand the environment. Current deep learning based methods require large amounts of labeled data for training. Manual annotation is expensive, while simulators can provide accurate annotations. However, the performance of the semantic segmentation model trained with synthetic datasets will significantly degenerate in the actual scenes. Unsupervised domain adaptation (UDA) for semantic segmentation is used to reduce the domain gap and improve the performance on the target domain. Existing adversarial-based and self-training methods usually involve complex training procedures, while entropy-based methods have recently received attention for their simplicity and effectiveness. However, entropy-based UDA methods have problems that they barely optimize hard samples and lack an explicit semantic connection between the source and target domains. In this paper, we propose a novel two-stage entropy-based UDA method for semantic segmentation. In stage one, we design a threshold-adaptative unsupervised focal loss to regularize the prediction in the target domain. It first introduces unsupervised focal loss into UDA for semantic segmentation, helping to optimize hard samples and avoiding generating unreliable pseudo-labels in the target domain. In stage two, we employ cross-domain image mixing (CIM) to bridge the semantic knowledge between two domains and incorporate long-tail class pasting to alleviate the class imbalance problem. Extensive experiments on synthetic-to-real and cross-city benchmarks demonstrate the effectiveness of our method. It achieves state-of-the-art performance using DeepLabV2, as well as competitive performance using the lightweight BiSeNet with great advantages in training and inference time. Weihao Yan 0001, Yeqiang Qian, Ming Yang 0002 |
IEEE Trans. Intell. Transp. Syst. | 1 |