Liang Xiao 0007

dblp:x/LiangXiao7 · DBLP profile ↗
← Back
19ranked-venue papers
3as first author
13since 2021 · last 2026
0000-0001-6959-4343ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Autonomous driving · 34% Transfer learning and domain adaptation · 28% Vision and language · 19%

Topics — the 16 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Autonomous driving
autonomous driving perception
1.322024
DriveWorld: 4D Pre-Trained Scene Understanding via World Models for Autonomous Driving · CVPR 2024
Trajectory Prediction for Autonomous Driving with Topometric Map · ICRA 2022
Machine learning › Transfer learning and domain adaptation › domain generalization
cross-dataset generalization
0.912025
An Effective Levelling Paradigm for Unlabeled Scenarios · NeurIPS 2025
Machine learning › Transfer learning and domain adaptation
domain generalization
0.912025
An Effective Levelling Paradigm for Unlabeled Scenarios · NeurIPS 2025
Machine learning › Transfer learning and domain adaptation › domain generalization
unsupervised domain generalization
0.912025
An Effective Levelling Paradigm for Unlabeled Scenarios · NeurIPS 2025
Computer vision › Vision and language › vision-language model
prompt learning
0.812024
Advancing Prompt Learning through an External Layer · ACM Multimedia 2024
Computer vision › Vision and language › vision-language model
vision-language model adaptation
0.812024
Advancing Prompt Learning through an External Layer · ACM Multimedia 2024
Machine learning › Reinforcement learning › model-based reinforcement learning
world model
0.812024
DriveWorld: 4D Pre-Trained Scene Understanding via World Models for Autonomous Driving · CVPR 2024
Robotics › Autonomous driving › road detection
off-road freespace detection
0.612022
ORFD: A Dataset and Benchmark for Off-Road Freespace Detection · ICRA 2022
Robotics › Autonomous driving
road detection
0.612022
ORFD: A Dataset and Benchmark for Off-Road Freespace Detection · ICRA 2022
Robotics › Autonomous driving
trajectory prediction
0.612022
Trajectory Prediction for Autonomous Driving with Topometric Map · ICRA 2022
Machine learning › Deep learning architectures and training › recurrent neural network
LSTM
0.312017
Wider and Deeper, Cheaper and Faster: Tensorized LSTMs for Sequence Learning · NIPS 2017
Machine learning › Efficient and distributed learning
parameter sharing
0.312017
Wider and Deeper, Cheaper and Faster: Tensorized LSTMs for Sequence Learning · NIPS 2017
Machine learning › Deep learning architectures and training
recurrent neural network
0.312017
Wider and Deeper, Cheaper and Faster: Tensorized LSTMs for Sequence Learning · NIPS 2017
Computer vision › Vision and language
cross-modal alignment
0.312025
An Effective Levelling Paradigm for Unlabeled Scenarios · NeurIPS 2025
Robotics › Autonomous driving
perception
0.212022
ORFD: A Dataset and Benchmark for Off-Road Freespace Detection · ICRA 2022
Computer vision › 3D vision › depth estimation
visual-LiDAR fusion
0.212022
ORFD: A Dataset and Benchmark for Off-Road Freespace Detection · ICRA 2022

Methods — techniques the papers use, named apart from their topics

visual loss · 0.9multi-objective optimization · 0.9pre-training · 0.8optimal transport · 0.8memory state-space model · 0.8learnable visual embeddings · 0.8external layer · 0.8transformer · 0.6deep learning · 0.6cross-attention · 0.6
YearPublicationVenuePosition
2026 PointSlice: Accurate and efficient slice-based representation for 3D object detection from point clouds
Dawei Zhao 0003, Yabo Dong, Liang Xiao 0007, Juan Wang 0033, Weizhong Jiang, Dongming Lu, Yiming Nie
Pattern Recognit.4
2026 IDSTT: Iterative Dual-Sample-Teacher for Semi-Supervised Visual Object Tracking
Kunlong Zhao, Dawei Zhao 0003, Liang Xiao 0007, Yiming Nie, Yulong Huang 0003, Yonggang Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 TPKD: Teacher-Pruned Knowledge Distillation for Point Cloud-Based 3D Object Detection
Liang Xiao 0007, Dawei Zhao 0003, Qi Zhu 0004, Yiming Nie, Bin Dai 0001
ICIC (22)2
2025 An Effective Levelling Paradigm for Unlabeled Scenarios
abstract
Advancements in direct-integration fine-tuning frameworks have underscored their potential to enhance the performance of labeled scenarios and tasks. To enhance the generalization of different categories in the same dataset, some methods have added visual loss to these frameworks for unlabeled scenarios. However, the performance of these methods through visual loss does not improve significantly in domain generalization and cross-dataset generalization tasks. This may be attributed to the uncoordinated learning of the two-modalities alignment and visual loss. To mitigate this issue of uncoordinated learning, we propose a novel method called Levelling Paradigm (LePa) to improve performance for unlabeled tasks or scenarios. The proposed LePa, designed as a plug-in module, dynamically constrains and coordinates multiple objective functions, thereby improving the generalization of these baseline methods. Comprehensive experiments have shown that our design can effectively address generalized scenarios and tasks.
Fangming Cui, Yuqiang Ren, Liang Xiao 0007, Xinmei Tian 0001
NeurIPS5
2025 Spatiotemporal Context Adapting Framework for Visual Object Tracking
abstract
ABSTRACT Visual object tracking is widely applied in intelligent transportation systems and visual surveillance systems that serve smart cities, as well as in autonomous vehicles. Existing methods usually utilise a relation‐modelling framework to model the visual object tracking problem, with auxiliary spatial context and temporal information. The spatial context is often extracted by enlarging the target template, which can introduce more background and positional information. The temporal correlation is obtained by associating the search image with previous images. However, due to noise interference, existing methods often partially exploit auxiliary data, leading to underutilisation of spatiotemporal information. To address these issues, we propose a novel and concise tracking framework, uniformly encoding all auxiliary data, including the enlarged target template, previous images, and corresponding target bounding boxes. Specifically, to mitigate the unstable factors introduced by these raw inputs, we propose a spatiotemporal context adaptive encoder, which can adaptively select appropriate information in noisy data. Extensive experiments show that the proposed method achieves state‐of‐the‐art performance on various benchmarks, demonstrating its superiority.
Kunlong Zhao, Dawei Zhao 0003, Xu Wang 0043, Liang Xiao 0007, Yulong Huang 0003, Yiming Nie, Yonggang Zhang 0001, Bin Dai 0001
IET Image Process.4
2025 Efficient Distillation Using Channel Pruning for Point Cloud-Based 3D Object Detection
abstract
Although point cloud-based 3D object detectors have advanced significantly in recent years, they are frequently hindered by substantial computational overheads. Lightweight model techniques, such as knowledge distillation, have recently been proven effective for 3D object detector compression. However, neural network pruning’s complementary role in knowledge distillation is often overlooked. In this paper, we propose an efficient distillation using channel pruning for point cloud-based 3D object detection. Firstly, given the complete teacher model, we introduce random and magnitude channel pruning methods to generate several compact student models and investigate the effects of different combinations on 3D and 2D layers. Secondly, we introduce model compression scores to explore the impact of channel compression ratios and input resolutions, enabling us to select suitable pruned models for distillation from the given set. Furthermore, we employ multi-source knowledge distillation to facilitate more effective spatial and semantic knowledge transfer. To highlight the features of the foreground regions during distillation, we then propose a soft pivotal position selection mask. Extensive evaluations on various datasets using both pillar-and voxel-based 3D detectors validate the efficiency of our method in compressing point cloud-based 3D detectors. Codes are publicly available at https://github.com/lifuyang-1919/Efficient-Distillation.git
Juan Wang 0033, Liang Xiao 0007, Dawei Zhao 0003, Yiming Nie, Bin Dai 0001
IEEE Trans. Intell. Transp. Syst.4
2025 Contrastive Label Disambiguation for Self-Supervised Terrain Traversability Learning in Off-Road Environments
abstract
Discriminating terrain traversability stands as a pivotal challenge for autonomous driving in off-road environments. The complexity arises from the diverse and ambiguous nature of off-road conditions, coupled with the specific characteristics of the driving platform. To address this challenge, we introduce a novel self-supervised learning framework for terrain traversability analysis, incorporating a contrastive label disambiguation mechanism. The proposed framework integrates traversability learning with real-time scene reconstruction. By projecting actual driving experience onto the terrain models, weakly labeled training samples with pseudo-labels can be automatically generated. Furthermore, a prototype-based contrastive representation learning method with the aid of a local window-based transformer encoder is designed to learn distinguishable embeddings, facilitating the self-supervised updating of those pseudo labels. Through the iterative interaction between representation learning and pseudo label updating, the inherent ambiguities associated with those pseudo labels are gradually eliminated. This enables the acquisition of fine-grained and platform-specific terrain traversability insights, eliminating the need for any human-provided annotations. Experimental results on the publicly available RELLIS-3D dataset and two self-collected datasets demonstrate the effectiveness of the proposed method.
Hanzhang Xue, Liang Xiao 0007, Xiaochang Hu, Hao Fu 0001, Yiming Nie, Bin Dai 0001
IEEE Trans. Intell. Transp. Syst.2
2024 DriveWorld: 4D Pre-Trained Scene Understanding via World Models for Autonomous Driving
abstract
Vision-centric autonomous driving has recently raised wide attention due to its lower cost. Pretraining is essential for extracting a universal representation. However, current vision-centric pretraining typically relies on either 2D or 3D pre-text tasks, overlooking the temporal characteristics of autonomous driving as a 4D scene understanding task. In this paper, we address this challenge by introducing a world model-based autonomous driving 4D representation learning framework, dubbed DriveWorld, which is capable of pretraining from multi-camera driving videos in a spatiotemporal fashion. Specifically, we propose a Memory State-Space Model for spatiotemporal modelling, which consists of a Dynamic Memory Bank module for learning temporal-aware latent dynamics to predict future changes and a Static Scene Propagation module for learning spatial-aware latent statics to offer comprehensive scene contexts. We additionally introduce a Task Prompt to decouple task-aware features for various downstream tasks. The experiments demonstrate that DriveWorld delivers promising results on various autonomous driving tasks. When pretrained with the OpenScene dataset, DriveWorld achieves a 7.5% increase in mAP for 3D object detection, a 3.0% increase in IoU for online mapping, a 5.0% increase in AMOTA for multi-object tracking, a 0.1m decrease in minADE for motionforecasting, a 3.0% increase in IoU for occupancy prediction, and a 0.34m reduction in average L2 error for planning.
Dawei Zhao 0003, Liang Xiao 0007, Jian Zhao 0006, Xinli Xu, Lei Jin 0003, Jianshu Li, Yulan Guo, Junliang Xing, Liping Jing, Yiming Nie, Bin Dai 0001
CVPR3
2024 Pre-pruned Distillation for Point Cloud-based 3D Object Detection
abstract
Knowledge distillation has recently been proven to be effective for model compression and acceleration of point cloud-based 3D object detection. However, the complementary network pruning is often overlooked during knowledge distillation. In this paper, we propose a pre-pruned distillation framework that combines network pruning and knowledge distillation to better transfer knowledge from the teacher to the student. To maintain the feature consistency between the student and the teacher, we train a teacher model and then generate a compact student model by structural channel pruning. Then, we employ multi-source knowledge distillation to transfer both mid-level and high-level information to the student model. Additionally, to improve the object detection performance of the student model, we propose a soft pivotal position selection mask to emphasize the features of the foreground regions during distillation. We conduct experiments on both pillarand voxel-based 3D object detectors on the Waymo datasets, demonstrating the effectiveness of our approach in compressing point cloud-based 3D detectors.
Liang Xiao 0007, Dawei Zhao 0003, Shubin Si, Hanzhang Xue, Yiming Nie, Bin Dai 0001
IV3
2024 Advancing Prompt Learning through an External Layer
abstract
Prompt learning represents a promising method for adapting pre-trained vision-language models (VLMs) to various downstream tasks by learning a set of text embeddings. One challenge inherent to these methods is the poor generalization performance due to the invalidity of the learned text embeddings for unseen tasks. A straightforward approach to bridge this gap is to freeze the text embeddings in prompts, which results in a lack of capacity to adapt VLMs for downstream tasks. To address this dilemma, we propose a paradigm called EnPrompt with a novel External Layer (EnLa). Specifically, we propose a textual external layer and learnable visual embeddings for adapting VLMs to downstream tasks. The learnable external layer is built upon valid embeddings of pre-trained CLIP. This design considers the balance of learning capabilities between the two branches. To align the textual and visual features, we propose a novel two-pronged approach: i) we introduce the optimal transport as the discrepancy metric to align the vision and text modalities, and ii) we introduce a novel strengthening feature to enhance the interaction between these two modalities. Four representative experiments (i.e., base-to-novel generalization, few-shot learning, cross-dataset generalization, domain shifts generalization) across 15 datasets demonstrate that our method outperforms the existing prompt learning method.
Fangming Cui, Xun Yang 0001, Chao Wu 0001, Liang Xiao 0007, Xinmei Tian 0001
ACM Multimedia4
2024 A Two-Stage Active Domain Adaptation Framework for Vehicle Re-Identification
Linzhi Shang, Dawei Zhao 0003, Yiming Nie, Kunlong Zhao, Liang Xiao 0007, Bin Dai 0001
PRCV (1)5
2022 ORFD: A Dataset and Benchmark for Off-Road Freespace Detection
abstract
Freespace detection is an essential component of autonomous driving technology and plays an important role in trajectory planning. In the last decade, deep learning based freespace detection methods have been proved feasible. However, these efforts were focused on urban road environments and few deep learning based methods were specifically designed for off-road freespace detection due to the lack of off-road dataset and benchmark. In this paper, we present the ORFD dataset, which, to our knowledge, is the first off-road freespace detection dataset. The dataset was collected in different scenes (woodland, farmland, grassland and countryside), different weather conditions (sunny, rainy, foggy and snowy) and different light conditions (bright light, daylight, twilight, darkness), which totally contains 12,198 LiDAR point cloud and RGB image pairs with the traversable area, non-traversable area and unreachable area annotated in detail. We propose a novel network named OFF-Net, which unifies Transformer architecture to aggregate local and global information, to meet the requirement of large receptive fields for freespace detection task. We also propose the cross-attention to dynamically fuse LiDAR and RGB image information for accurate off-road freespace detection. Dataset and code are publicly available at https://github.com/chaytonmin/OFF-Net.
Weizhong Jiang, Dawei Zhao 0003, Jiaolong Xu, Liang Xiao 0007, Yiming Nie, Bin Dai 0001
ICRA5
2022 Trajectory Prediction for Autonomous Driving with Topometric Map
abstract
State-of-the-art autonomous driving systems rely on high definition (HD) maps for localization and navigation. However, building and maintaining HD maps is time-consuming and expensive. Furthermore, the HD maps assume structured environment such as the existence of major road and lanes, which are not present in rural areas. In this work, we propose an end-to-end transformer networks based approach for map-less autonomous driving. The proposed model takes raw LiDAR data and noisy topometric map as input and produces precise local trajectory for navigation. We demonstrate the effectiveness of our method in real-world driving data, including both urban and rural areas. The experimental results show that the proposed method outperforms state-of-the-art multimodal methods and is robust to the perturbations of the topometric map. The code of the proposed method is publicly available at https://github.com/Jiaolong/trajectory-prediction.
Jiaolong Xu, Liang Xiao 0007, Dawei Zhao 0003, Yiming Nie, Bin Dai 0001
ICRA2
2020 Self-Supervised Domain Adaptation with Consistency Training
abstract
We consider the problem of unsupervised domain adaptation for image classification. To learn target-domain-aware features from the unlabeled data, we create a self-supervised pretext task by augmenting the unlabeled data with a certain type of transformation (specifically, image rotation) and ask the learner to predict the properties of the transformation. However, the obtained feature representation may contain a large amount of irrelevant information with respect to the main task. To provide further guidance, we force the feature representation of the augmented data to be consistent with that of the original data. Intuitively, the consistency introduces additional constraints to representation learning, therefore, the learned representation is more likely to focus on the right information about the main task. Our experimental results validate the proposed method and demonstrate state-of-the-art performance on classical domain adaptation benchmarks. Code is available at https://github.com/Jiaolong/ss-da-consistency.
Liang Xiao 0007, Jiaolong Xu, Dawei Zhao 0003, Yiming Nie, Bin Dai 0001
ICPR1
2020 Drosophila-inspired 3D moving object detection based on point clouds
Dawei Zhao 0003, Tao Wu 0001, Hao Fu 0001, Liang Xiao 0007, Xin Xu 0001, Bin Dai 0001
Inf. Sci.6
2019 Augmenting cascaded correlation filters with spatial-temporal saliency for visual tracking
Dawei Zhao 0003, Liang Xiao 0007, Hao Fu 0001, Tao Wu 0001, Xin Xu 0001, Bin Dai 0001
Inf. Sci.2
2018 Hybrid conditional random field based camera-LIDAR fusion for road detection
Liang Xiao 0007, Ruili Wang 0001, Bin Dai 0001, Yuqiang Fang, Daxue Liu, Tao Wu 0001
Inf. Sci.1
2017 Wider and Deeper, Cheaper and Faster: Tensorized LSTMs for Sequence Learning
abstract
Long Short-Term Memory (LSTM) is a popular approach to boosting the ability of Recurrent Neural Networks to store longer term temporal information. The capacity of an LSTM network can be increased by widening and adding layers. However, usually the former introduces additional parameters, while the latter increases the runtime. As an alternative we propose the Tensorized LSTM in which the hidden states are represented by tensors and updated via a cross-layer convolution. By increasing the tensor size, the network can be widened efficiently without additional parameters since the parameters are shared across different locations in the tensor; by delaying the output, the network can be deepened implicitly with little additional runtime since deep computations for each timestep are merged into temporal computations of the sequence. Experiments conducted on five challenging sequence learning tasks show the potential of the proposed model.
Shaobing Gao, Liang Xiao 0007, Daxue Liu, Hangen He, David Barber
NIPS3
2015 CRF based road detection with multi-sensor fusion
abstract
In this paper, we propose to fuse the LIDAR and monocular image in the framework of conditional random field to detect the road robustly in challenging scenarios. LIDAR points are aligned with pixels in image by cross calibration. Then boosted decision tree based classifiers are trained for image and point cloud respectively. The scores of the two kinds of classifiers are treated as the unary potentials of the corresponding pixel nodes of the random field. The fused conditional random field can be solved efficiently with graph cut. Extensive experiments tested on KITTI-Road benchmark show that our method reaches the state-of-the-art.
Liang Xiao 0007, Bin Dai 0001, Daxue Liu, Tingbo Hu, Tao Wu 0001
Intelligent Vehicles Symposium1