Weihao Gu

dblp:203/9883 · DBLP profile ↗
← Back
13ranked-venue papers
0as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 7 since 2021Systems, architecture and hardware · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Distilling Future Temporal Knowledge with Masked Feature Reconstruction for 3D Object Detection
abstract
Camera-based temporal 3D object detection has shown impressive results in autonomous driving, with offline models improving accuracy by using future frames. Knowledge distillation (KD) can be an appealing framework for transferring rich information from offline models to online models. However, existing KD methods overlook future frames, as they mainly focus on spatial feature distillation under strict frame alignment or on temporal relational distillation, thereby making it challenging for online models to effectively learn future knowledge. To this end, we propose a sparse query-based approach, Future Temporal Knowledge Distillation (FTKD), which effectively transfers future frame knowledge from an offline teacher model to an online student model. Specifically, we present a future-aware feature reconstruction strategy to encourage the student model to capture future features without strict frame alignment. In addition, we further introduce future-guided logit distillation to leverage the teacher's stable foreground and background context. FTKD is applied to two high-performing 3D object detection baselines, achieving up to 1.3 mAP and 1.3 NDS gains on the nuScenes dataset, as well as the most accurate velocity estimation, without increasing inference cost.
Hu Zhu, Weihao Gu, Yang Yang 0062, Yanyan Liang 0001
AAAI4
2025 Diffusion-Based Planning for Autonomous Driving with Flexible Guidance
abstract
Achieving human-like driving behaviors in complex open-world environments is a critical challenge in autonomous driving. Contemporary learning-based planning approaches such as imitation learning methods often struggle to balance competing objectives and lack of safety assurance,due to limited adaptability and inadequacy in learning complex multi-modal behaviors commonly exhibited in human planning, not to mention their strong reliance on the fallback strategy with predefined rules. We propose a novel transformer-based Diffusion Planner for closed-loop planning, which can effectively model multi-modal driving behavior and ensure trajectory quality without any rule-based refinement. Our model supports joint modeling of both prediction and planning tasks under the same architecture, enabling cooperative behaviors between vehicles. Moreover, by learning the gradient of the trajectory score function and employing a flexible classifier guidance mechanism, Diffusion Planner effectively achieves safe and adaptable planning behaviors. Evaluations on the large-scale real-world autonomous planning benchmark nuPlan and our newly collected 200-hour delivery-vehicle driving dataset demonstrate that Diffusion Planner achieves state-of-the-art closed-loop performance with robust transferability in diverse driving styles.
Yinan Zheng, Ruiming Liang, Kexin Zheng, Jinliang Zheng, Liyuan Mao, Weihao Gu, Rui Ai 0001, Shengbo Eben Li, Xianyuan Zhan
ICLR7
2025 TA-Detector: A GNN-Based Anomaly Detector via Trust Relationship
abstract
With the rise of mobile Internet and AI, social media integrating short messages, images, and videos has developed rapidly. As a guarantee for the stable operation of social media, information security, especially graph anomaly detection (GAD), has become a hot issue inspired by the extensive attention of researchers. Most GAD methods are mainly limited to enhancing the homophily or considering homophily and heterophilic connections. Nevertheless, due to the deceptive nature of homophily connections among anomalies, the discriminative information of the anomalies can be eliminated. To alleviate the issue, we explore a novel method TA-Detector in GAD by introducing the concept of trust into the classification of connections. In particular, the proposed approach adopts a designed trust classier to distinguish trust and distrust connections with the supervision of labeled nodes. Then, we capture the latent factors related to GAD by graph neural networks, which integrate node interaction type information and node representation. Finally, to identify anomalies in the graph, we use the residual network mechanism to extract the deep semantic embedding information related to GAD. Experimental results on two real benchmark datasets verify that our proposed approach boosts the overall GAD performance in comparison to benchmark baselines.
Nan Jiang 0013, Jie Zhou 0001, Yanpei Li, Hualin Zhan, Guang Kou, Weihao Gu
ACM Trans. Multim. Comput. Commun. Appl.8
2024 Cam4DOcc: Benchmark for Camera-Only 4D Occupancy Forecasting in Autonomous Driving Applications
abstract
Understanding how the surrounding environment changes is crucial for performing downstream tasks safely and reliably in autonomous driving applications. Recent occupancy estimation techniques using only camera images as input can provide dense occupancy representations of large-scale scenes based on the current observation. However, they are mostly limited to representing the current 3D space and do not consider the future state of surrounding objects along the time axis. To extend camera-only occupancy estimation into spatiotemporal prediction, we propose Cam4DOcc, a new benchmark for camera-only 4D occupancy forecasting, evaluating the surrounding scene changes in a near future. We build our benchmark based on multiple publicly available datasets, including nuScenes, nuScenes-Occupancy, and Lyft-Level5, which provides sequential occupancy states of general movable and static objects, as well as their 3D backward centripetal flow. To establish this benchmark for future research with comprehensive comparisons, we introduce four baseline types from diverse camera-based perception and prediction implementations, including a static-world occupancy model, voxelization of point cloud prediction, 2D-3D instance-based prediction, and our proposed novel end-to-end 4D occupancy forecasting network. Furthermore, the standardized evaluation protocol for preset multiple tasks is also provided to compare the performance of all the proposed baselines on present and future occupancy estimation with respect to objects of interest in autonomous driving scenarios. The dataset and our implementation of all four baselines in the proposed Cam4DOcc benchmark are released as open source at https://github.com/haomo-ai/Cam4DOcc.
Junyi Ma, Xieyuanli Chen, Jintao Xu 0001, Weihao Gu, Rui Ai 0001, Hesheng Wang 0001
CVPR7
2024 SuperFusion: Multilevel LiDAR-Camera Fusion for Long-Range HD Map Generation
abstract
High-definition (HD) semantic map generation of the environment is an essential component of autonomous driving. Existing methods have achieved good performance in this task by fusing different sensor modalities, such as LiDAR and camera. However, current works are based on raw data or network feature-level fusion and only consider short-range HD map generation, limiting their deployment to realistic autonomous driving applications. In this paper, we focus on the task of building the HD maps in both short ranges, i.e., within 30m, and also predicting long-range HD maps up to 90m, which is required by downstream path planning and control tasks to improve the smoothness and safety of autonomous driving. To this end, we propose a novel network named SuperFusion, exploiting the fusion of LiDAR and camera data at multiple levels. We use LiDAR depth to improve image depth estimation and use image features to guide long-range LiDAR feature prediction. We benchmark our SuperFusion on the nuScenes dataset and a self-recorded dataset and show that it outperforms the state-of-the-art baseline methods with large margins on all intervals. Additionally, we apply the generated HD map to a downstream path planning task, demonstrating that the long-range HD maps predicted by our method can lead to better path planning for autonomous vehicles. Our code and self-recorded dataset have been released at https://github.com/haomo-ai/SuperFusion.
Hao Dong 0011, Weihao Gu, Xianjing Zhang, Jintao Xu 0001, Rui Ai 0001, Huimin Lu 0002, Juho Kannala, Xieyuanli Chen
ICRA2
2024 ModaLink: Unifying Modalities for Efficient Image-to-PointCloud Place Recognition
abstract
Place recognition is an important task for robots and autonomous cars to localize themselves and close loops in pre-built maps. While single-modal sensor-based methods have shown satisfactory performance, cross-modal place recognition that retrieving images from a point-cloud database remains a challenging problem. Current cross-modal methods transform images into 3D points using depth estimation for modality conversion, which are usually computationally intensive and need expensive labeled data for depth supervision. In this work, we introduce a fast and lightweight framework to encode images and point clouds into place-distinctive descriptors. We propose an effective Field of View (FoV) transformation module to convert point clouds into an analogous modality as images. This module eliminates the necessity for depth estimation and helps subsequent modules achieve real-time performance. We further design a non-negative factorization-based encoder to extract mutually consistent semantic features between point clouds and images. This encoder yields more distinctive global descriptors for retrieval. Experimental results on the KITTI dataset show that our proposed methods achieve state-of-the-art performance while running in real time. Additional evaluation on the HAOMO dataset covering a 17 km trajectory further shows the practical generalization capabilities. We have released the implementation of our methods as open source at: https://github.com/haomo-ai/ModaLink.git.
Weidong Xie, Lun Luo, Nanfei Ye, Shaoyi Du, Minhang Wang, Jintao Xu 0001, Rui Ai 0001, Weihao Gu, Xieyuanli Chen
IROS9
2024 PRISM: PRogressive dependency maxImization for Scale-invariant image Matching
abstract
Image matching aims at identifying corresponding points between a pair of images. Currently, detector-free methods have shown impressive performance in challenging scenarios, thanks to their capability of generating dense matches and global receptive field. However, performing feature interaction and proposing matches across the entire image is unnecessary, because not all image regions contribute to the matching process. Interacting and matching in unmatchable areas can introduce errors, reducing matching accuracy and efficiency. Meanwhile, the scale discrepancy issue still troubles existing methods. To address above issues, we propose PRogressive dependency maxImization for Scale-invariant image Matching (PRISM), which jointly prunes irrelevant patch features and tackles the scale discrepancy. To do this, we firstly present a Multi-scale Pruning Module (MPM) to adaptively prune irrelevant features by maximizing the dependency between the two feature sets. Moreover, we design the Scale-Aware Dynamic Pruning Attention (SADPA) to aggregate information from different scales via a hierarchical design. Our method's superior matching performance and generalization capability are confirmed by leading accuracy across various evaluation benchmarks and downstream tasks. The code is publicly available at https://github.com/Master-cai/PRISM.
Yongcai Wang, Lun Luo, Minhang Wang, Deying Li 0001, Jintao Xu 0001, Weihao Gu, Rui Ai 0001
ACM Multimedia7
2024 Joint Scene Flow Estimation and Moving Object Segmentation on Rotational LiDAR Data
abstract
LiDAR-based scene flow estimation (SFE) and moving object segmentation (MOS) are important tasks with broad-ranging applications in autonomous driving, such as traffic surveillance, motion analysis, obstacle avoidance, etc. Most existing works address SFE and MOS separately, ignoring the underlying shared geometric constraints and their inherent correlation. This article rethinks LiDAR-based SFE and MOS tasks, providing our key insight that jointly addressing them can tackle challenges in both tasks, and their solutions can reinforce one another to improve the performance of both. Based on this insight, we introduce a novel framework that exploits shared geometric constraints by explicitly partitioning the scene into static and moving regions and subsequently estimating flow differently for these regions. A lightweight and interpretable neural network dubbed SFEMOS is proposed. It employs an encoder and two specially designed head modules for each task, achieving MOS without relying on prior poses and online point-wise flow estimation for 360-degree point clouds. Due to the absence of public datasets for concurrently evaluating both tasks, we generate ground truth flow data using MOS labels from SemanticKITTI. Additionally, we establish a new dataset using a rotational LiDAR mounted on our own autonomous vehicle. Evaluation results on both datasets validate the superior performance of our proposed SFEMOS. Our dataset and label generation method are released athttps://github.com/nubot-nudt/SFEMOS.
Xieyuanli Chen, Jiafeng Cui, Xianjing Zhang, Jiadai Sun, Rui Ai 0001, Weihao Gu, Jintao Xu 0001, Huimin Lu 0002
IEEE Trans. Intell. Transp. Syst.7
2024 RTrust: toward robust trust evaluation framework for fake news detection in online social networks
Nan Jiang 0013, Ziang Tu, Kanglu Pei, Hualin Zhan, Ximeng Liu, Weihao Gu, Sen Qiu
World Wide Web (WWW)8
2023 I2P-Rec: Recognizing Images on Large-Scale Point Cloud Maps Through Bird's Eye View Projections
abstract
Place recognition is an important technique for autonomous cars to achieve full autonomy since it can provide an initial guess to online localization algorithms. Although current methods based on images or point clouds have achieved satisfactory performance, localizing the images on a large-scale point cloud map remains a fairly unexplored problem. This cross-modal matching task is challenging due to the difficulty in extracting consistent descriptors from images and point clouds. In this paper, we propose the I2P-Rec method to solve the problem by transforming the cross-modal data into the same modality. Specifically, we leverage on the recent success of depth estimation networks to recover point clouds from images. We then project the point clouds into Bird's Eye View (BEV) images. Using the BEV image as an intermediate representation, we extract global features with a Convolutional Neural Network followed by a NetVLAD layer to perform matching. The experimental results evaluated on the KITTI dataset show that, with only a small set of training data, I2P-Rec achieves recall rates at Top-l % over 80% and 90%, when localizing monocular and stereo images on point cloud maps, respectively. We further evaluate I2P-Rec on a 1 km trajectory dataset collected by an autonomous logistics car and show that I2P- Rec can generalize well to previously unseen environments.
Shuhang Zheng, Zhu Yu 0001, Beinan Yu, Si-Yuan Cao, Minhang Wang, Jintao Xu 0001, Rui Ai 0001, Weihao Gu, Lun Luo
IROS9
2023 NAH: neighbor-aware attention-based heterogeneous relation network model in E-commerce recommendation
Nan Jiang 0013, Zihao Hu, Weihao Gu, Ziang Tu, Ximeng Liu, Jianfei Gong, Fengtao Lin
World Wide Web (WWW)5
2022 Efficient Spatial-Temporal Information Fusion for LiDAR-Based 3D Moving Object Segmentation
abstract
Accurate moving object segmentation is an es-sential task for autonomous driving. It can provide effective information for many downstream tasks, such as collision avoidance, path planning, and static map construction. How to effectively exploit the spatial-temporal information is a critical question for 3D LiDAR moving object segmentation (LiDAR-MOS). In this work, we propose a novel deep neural network exploiting both spatial-temporal information and different representation modalities of LiDAR scans to improve LiDAR-MOS performance. Specifically, we first use a range image-based dual-branch structure to separately deal with spatial and temporal information that can be obtained from sequential LiDAR scans, and later combine them using motion-guided attention modules. We also use a point refinement module via 3D sparse convolution to fuse the information from both LiDAR range image and point cloud representations and reduce the artifacts on the borders of the objects. We verify the effectiveness of our proposed approach on the LiDAR-MOS benchmark of SemanticKITTI. Our method outperforms the state-of-the-art methods significantly in terms of LiDAR-MOS IoU. Benefiting from the devised coarse-to-fine architecture, our method operates online at sensor frame rate. Code is available at: https://github.com/haomo-ai/MotionSeg3D.
Jiadai Sun, Yuchao Dai, Xianjing Zhang, Jintao Xu 0001, Rui Ai 0001, Weihao Gu, Xieyuanli Chen
IROS6
2017 Baidu driving dataset and end-to-end reactive control model
abstract
End-to-end autonomous driving system has obtained great progress recently. In this paper, we will introduce our open source dataset: Baidu Driving Dataset(BDD), and our end-to-end reactive control model trained on BDD. The BDD comes from Baidu street view project, which generates millions of kilometers driving data every year. Among them, we publish 10000 kilometers driving data for end-to-end autonomous driving research. The BDD consists of two parts: forward images and vehicle motion attitude. The vehicle motion attitude is derived from real time kinematic GPS location data with standard deviation of 3 centimeters. Our reactive control model consists of lateral control and longitudinal control. We employ curvature instead of steering angle for lateral control, and leverage acceleration, not throttle or brake, for longitudinal control. CNN network is employed for lateral control model, mapping a single image from forward camera directly to corresponding curvature. For longitudinal control, stacked convolutional LSTM is used to extract spatial and temporal features from a sequence of frames, and to map the features with longitudinal control commands. The demo and data are in http://roadhackers.baidu.com. To the best of our knowledge, it is the first time that both lateral and longitudinal control are implemented in an end-to-end style.
Weihao Gu
Intelligent Vehicles Symposium3