Dawei Zhao 0003

dblp:52/266-3 · DBLP profile ↗
← Back
15ranked-venue papers
2as first author
11since 2021 · last 2026
0000-0002-4583-1268ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Autonomous driving · 78% Reinforcement learning · 18% 3D vision · 4%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Autonomous driving
autonomous driving perception
1.322024
DriveWorld: 4D Pre-Trained Scene Understanding via World Models for Autonomous Driving · CVPR 2024
Trajectory Prediction for Autonomous Driving with Topometric Map · ICRA 2022
Machine learning › Reinforcement learning › model-based reinforcement learning
world model
0.812024
DriveWorld: 4D Pre-Trained Scene Understanding via World Models for Autonomous Driving · CVPR 2024
Robotics › Autonomous driving › road detection
off-road freespace detection
0.612022
ORFD: A Dataset and Benchmark for Off-Road Freespace Detection · ICRA 2022
Robotics › Autonomous driving
road detection
0.612022
ORFD: A Dataset and Benchmark for Off-Road Freespace Detection · ICRA 2022
Robotics › Autonomous driving
trajectory prediction
0.612022
Trajectory Prediction for Autonomous Driving with Topometric Map · ICRA 2022
Robotics › Autonomous driving
perception
0.212022
ORFD: A Dataset and Benchmark for Off-Road Freespace Detection · ICRA 2022
Computer vision › 3D vision › depth estimation
visual-LiDAR fusion
0.212022
ORFD: A Dataset and Benchmark for Off-Road Freespace Detection · ICRA 2022

Methods — techniques the papers use, named apart from their topics

pre-training · 0.8memory state-space model · 0.8transformer network · 0.6transformer · 0.6end-to-end learning · 0.6deep learning · 0.6cross-attention · 0.6
YearPublicationVenuePosition
2026 PointSlice: Accurate and efficient slice-based representation for 3D object detection from point clouds
Dawei Zhao 0003, Yabo Dong, Liang Xiao 0007, Juan Wang 0033, Weizhong Jiang, Dongming Lu, Yiming Nie
Pattern Recognit.2
2026 IDSTT: Iterative Dual-Sample-Teacher for Semi-Supervised Visual Object Tracking
Kunlong Zhao, Dawei Zhao 0003, Liang Xiao 0007, Yiming Nie, Yulong Huang 0003, Yonggang Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 TPKD: Teacher-Pruned Knowledge Distillation for Point Cloud-Based 3D Object Detection
Liang Xiao 0007, Dawei Zhao 0003, Qi Zhu 0004, Yiming Nie, Bin Dai 0001
ICIC (22)3
2025 Spatiotemporal Context Adapting Framework for Visual Object Tracking
abstract
ABSTRACT Visual object tracking is widely applied in intelligent transportation systems and visual surveillance systems that serve smart cities, as well as in autonomous vehicles. Existing methods usually utilise a relation‐modelling framework to model the visual object tracking problem, with auxiliary spatial context and temporal information. The spatial context is often extracted by enlarging the target template, which can introduce more background and positional information. The temporal correlation is obtained by associating the search image with previous images. However, due to noise interference, existing methods often partially exploit auxiliary data, leading to underutilisation of spatiotemporal information. To address these issues, we propose a novel and concise tracking framework, uniformly encoding all auxiliary data, including the enlarged target template, previous images, and corresponding target bounding boxes. Specifically, to mitigate the unstable factors introduced by these raw inputs, we propose a spatiotemporal context adaptive encoder, which can adaptively select appropriate information in noisy data. Extensive experiments show that the proposed method achieves state‐of‐the‐art performance on various benchmarks, demonstrating its superiority.
Kunlong Zhao, Dawei Zhao 0003, Xu Wang 0043, Liang Xiao 0007, Yulong Huang 0003, Yiming Nie, Yonggang Zhang 0001, Bin Dai 0001
IET Image Process.2
2025 Efficient Distillation Using Channel Pruning for Point Cloud-Based 3D Object Detection
abstract
Although point cloud-based 3D object detectors have advanced significantly in recent years, they are frequently hindered by substantial computational overheads. Lightweight model techniques, such as knowledge distillation, have recently been proven effective for 3D object detector compression. However, neural network pruning’s complementary role in knowledge distillation is often overlooked. In this paper, we propose an efficient distillation using channel pruning for point cloud-based 3D object detection. Firstly, given the complete teacher model, we introduce random and magnitude channel pruning methods to generate several compact student models and investigate the effects of different combinations on 3D and 2D layers. Secondly, we introduce model compression scores to explore the impact of channel compression ratios and input resolutions, enabling us to select suitable pruned models for distillation from the given set. Furthermore, we employ multi-source knowledge distillation to facilitate more effective spatial and semantic knowledge transfer. To highlight the features of the foreground regions during distillation, we then propose a soft pivotal position selection mask. Extensive evaluations on various datasets using both pillar-and voxel-based 3D detectors validate the efficiency of our method in compressing point cloud-based 3D detectors. Codes are publicly available at https://github.com/lifuyang-1919/Efficient-Distillation.git
Juan Wang 0033, Liang Xiao 0007, Dawei Zhao 0003, Yiming Nie, Bin Dai 0001
IEEE Trans. Intell. Transp. Syst.5
2024 DriveWorld: 4D Pre-Trained Scene Understanding via World Models for Autonomous Driving
abstract
Vision-centric autonomous driving has recently raised wide attention due to its lower cost. Pretraining is essential for extracting a universal representation. However, current vision-centric pretraining typically relies on either 2D or 3D pre-text tasks, overlooking the temporal characteristics of autonomous driving as a 4D scene understanding task. In this paper, we address this challenge by introducing a world model-based autonomous driving 4D representation learning framework, dubbed DriveWorld, which is capable of pretraining from multi-camera driving videos in a spatiotemporal fashion. Specifically, we propose a Memory State-Space Model for spatiotemporal modelling, which consists of a Dynamic Memory Bank module for learning temporal-aware latent dynamics to predict future changes and a Static Scene Propagation module for learning spatial-aware latent statics to offer comprehensive scene contexts. We additionally introduce a Task Prompt to decouple task-aware features for various downstream tasks. The experiments demonstrate that DriveWorld delivers promising results on various autonomous driving tasks. When pretrained with the OpenScene dataset, DriveWorld achieves a 7.5% increase in mAP for 3D object detection, a 3.0% increase in IoU for online mapping, a 5.0% increase in AMOTA for multi-object tracking, a 0.1m decrease in minADE for motionforecasting, a 3.0% increase in IoU for occupancy prediction, and a 0.34m reduction in average L2 error for planning.
Dawei Zhao 0003, Liang Xiao 0007, Jian Zhao 0006, Xinli Xu, Lei Jin 0003, Jianshu Li, Yulan Guo, Junliang Xing, Liping Jing, Yiming Nie, Bin Dai 0001
CVPR2
2024 Pre-pruned Distillation for Point Cloud-based 3D Object Detection
abstract
Knowledge distillation has recently been proven to be effective for model compression and acceleration of point cloud-based 3D object detection. However, the complementary network pruning is often overlooked during knowledge distillation. In this paper, we propose a pre-pruned distillation framework that combines network pruning and knowledge distillation to better transfer knowledge from the teacher to the student. To maintain the feature consistency between the student and the teacher, we train a teacher model and then generate a compact student model by structural channel pruning. Then, we employ multi-source knowledge distillation to transfer both mid-level and high-level information to the student model. Additionally, to improve the object detection performance of the student model, we propose a soft pivotal position selection mask to emphasize the features of the foreground regions during distillation. We conduct experiments on both pillarand voxel-based 3D object detectors on the Waymo datasets, demonstrating the effectiveness of our approach in compressing point cloud-based 3D detectors.
Liang Xiao 0007, Dawei Zhao 0003, Shubin Si, Hanzhang Xue, Yiming Nie, Bin Dai 0001
IV4
2024 A Two-Stage Active Domain Adaptation Framework for Vehicle Re-Identification
Linzhi Shang, Dawei Zhao 0003, Yiming Nie, Kunlong Zhao, Liang Xiao 0007, Bin Dai 0001
PRCV (1)2
2024 Deep Reinforcement Learning: A Survey
abstract
Deep reinforcement learning (DRL) integrates the feature representation ability of deep learning with the decision-making ability of reinforcement learning so that it can achieve powerful end-to-end learning control capabilities. In the past decade, DRL has made substantial advances in many tasks that require perceiving high-dimensional input and making optimal or near-optimal decisions. However, there are still many challenging problems in the theory and applications of DRL, especially in learning control tasks with limited samples, sparse rewards, and multiple agents. Researchers have proposed various solutions and new theories to solve these problems and promote the development of DRL. In addition, deep learning has stimulated the further development of many subfields of reinforcement learning, such as hierarchical reinforcement learning (HRL), multiagent reinforcement learning, and imitation learning. This article gives a comprehensive overview of the fundamental theories, key algorithms, and primary research domains of DRL. In addition to value-based and policy-based DRL algorithms, the advances in maximum entropy-based DRL are summarized. The future research topics of DRL are also analyzed and discussed.
Xu Wang 0043, Xingxing Liang, Dawei Zhao 0003, Jincai Huang 0001, Xin Xu 0001, Bin Dai 0001, Qiguang Miao
IEEE Trans. Neural Networks Learn. Syst.4
2022 ORFD: A Dataset and Benchmark for Off-Road Freespace Detection
abstract
Freespace detection is an essential component of autonomous driving technology and plays an important role in trajectory planning. In the last decade, deep learning based freespace detection methods have been proved feasible. However, these efforts were focused on urban road environments and few deep learning based methods were specifically designed for off-road freespace detection due to the lack of off-road dataset and benchmark. In this paper, we present the ORFD dataset, which, to our knowledge, is the first off-road freespace detection dataset. The dataset was collected in different scenes (woodland, farmland, grassland and countryside), different weather conditions (sunny, rainy, foggy and snowy) and different light conditions (bright light, daylight, twilight, darkness), which totally contains 12,198 LiDAR point cloud and RGB image pairs with the traversable area, non-traversable area and unreachable area annotated in detail. We propose a novel network named OFF-Net, which unifies Transformer architecture to aggregate local and global information, to meet the requirement of large receptive fields for freespace detection task. We also propose the cross-attention to dynamically fuse LiDAR and RGB image information for accurate off-road freespace detection. Dataset and code are publicly available at https://github.com/chaytonmin/OFF-Net.
Weizhong Jiang, Dawei Zhao 0003, Jiaolong Xu, Liang Xiao 0007, Yiming Nie, Bin Dai 0001
ICRA3
2022 Trajectory Prediction for Autonomous Driving with Topometric Map
abstract
State-of-the-art autonomous driving systems rely on high definition (HD) maps for localization and navigation. However, building and maintaining HD maps is time-consuming and expensive. Furthermore, the HD maps assume structured environment such as the existence of major road and lanes, which are not present in rural areas. In this work, we propose an end-to-end transformer networks based approach for map-less autonomous driving. The proposed model takes raw LiDAR data and noisy topometric map as input and produces precise local trajectory for navigation. We demonstrate the effectiveness of our method in real-world driving data, including both urban and rural areas. The experimental results show that the proposed method outperforms state-of-the-art multimodal methods and is robust to the perturbations of the topometric map. The code of the proposed method is publicly available at https://github.com/Jiaolong/trajectory-prediction.
Jiaolong Xu, Liang Xiao 0007, Dawei Zhao 0003, Yiming Nie, Bin Dai 0001
ICRA3
2020 Self-Supervised Domain Adaptation with Consistency Training
abstract
We consider the problem of unsupervised domain adaptation for image classification. To learn target-domain-aware features from the unlabeled data, we create a self-supervised pretext task by augmenting the unlabeled data with a certain type of transformation (specifically, image rotation) and ask the learner to predict the properties of the transformation. However, the obtained feature representation may contain a large amount of irrelevant information with respect to the main task. To provide further guidance, we force the feature representation of the augmented data to be consistent with that of the original data. Intuitively, the consistency introduces additional constraints to representation learning, therefore, the learned representation is more likely to focus on the right information about the main task. Our experimental results validate the proposed method and demonstrate state-of-the-art performance on classical domain adaptation benchmarks. Code is available at https://github.com/Jiaolong/ss-da-consistency.
Liang Xiao 0007, Jiaolong Xu, Dawei Zhao 0003, Yiming Nie, Bin Dai 0001
ICPR3
2020 Drosophila-inspired 3D moving object detection based on point clouds
Dawei Zhao 0003, Tao Wu 0001, Hao Fu 0001, Liang Xiao 0007, Xin Xu 0001, Bin Dai 0001
Inf. Sci.2
2019 Augmenting cascaded correlation filters with spatial-temporal saliency for visual tracking
Dawei Zhao 0003, Liang Xiao 0007, Hao Fu 0001, Tao Wu 0001, Xin Xu 0001, Bin Dai 0001
Inf. Sci.1
2014 Efficient Vehicle Localization Based on Road-Boundary Maps
Dawei Zhao 0003, Tao Wu 0001, Yuqiang Fang, Ruili Wang 0001, Bin Dai 0001
PRICAI1