VLDB 2026 Research / reviewers in the wild / expert
Jiantao Gao
dblp:265/1310
· DBLP profile ↗
16ranked-venue papers
1as first author
13since 2021 · last 2025
0000-0001-5057-0229ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 1 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | VisionPAD: A Vision-Centric Pre-training Paradigm for Autonomous DrivingabstractThis paper introduces VisionPAD, a novel self-supervised pre-training paradigm designed for vision-centric algorithms in autonomous driving. In contrast to previous approaches that employ neural rendering with explicit depth supervision, VisionPAD utilizes more efficient 3D Gaussian Splatting to reconstruct multi-view representations using only images as supervision. Specifically, we introduce a self-supervised method for voxel velocity estimation. By warping voxels to adjacent frames and supervising the rendered outputs, the model effectively learns motion cues in the sequential data. Furthermore, we adopt a multi-frame photometric consistency approach to enhance geometric perception. It projects adjacent frames to the current frame based on rendered depths and relative poses, boosting the 3D geometric representation through pure image supervision. Extensive experiments on autonomous driving datasets demonstrate that VisionPAD significantly improves performance in 3D object detection, occupancy prediction and map segmentation, surpassing state-of-the-art pre-training strategies by a considerable margin. Haiming Zhang 0001, Wending Zhou, Yiyao Zhu, Xu Yan 0005, Jiantao Gao, Dongfeng Bai, Yingjie Cai, Shuguang Cui, Zhen Li 0026 |
CVPR | 5 |
| 2024 | RadOcc: Learning Cross-Modality Occupancy Knowledge through Rendering Assisted Distillationabstract3D occupancy prediction is an emerging task that aims to estimate the occupancy states and semantics of 3D scenes using multi-view images. However, image-based scene perception encounters significant challenges in achieving accurate prediction due to the absence of geometric priors. In this paper, we address this issue by exploring cross-modal knowledge distillation in this task, i.e., we leverage a stronger multi-modal model to guide the visual model during training. In practice, we observe that directly applying features or logits alignment, proposed and widely used in bird's-eye-view (BEV) perception, does not yield satisfactory results. To overcome this problem, we introduce RadOcc, a Rendering assisted distillation paradigm for 3D Occupancy prediction. By employing differentiable volume rendering, we generate depth and semantic maps in perspective views and propose two novel consistency criteria between the rendered outputs of teacher and student models. Specifically, the depth consistency loss aligns the termination distributions of the rendered rays, while the semantic consistency loss mimics the intra-segment similarity guided by vision foundation models (VLMs). Experimental results on the nuScenes dataset demonstrate the effectiveness of our proposed method in improving various 3D occupancy prediction approaches, e.g., our proposed methodology enhances our baseline by 2.2% in the metric of mIoU and achieves 50% in Occ3D benchmark. Haiming Zhang 0001, Xu Yan 0005, Dongfeng Bai, Jiantao Gao, Shuguang Cui, Zhen Li 0026 |
AAAI | 4 |
| 2024 | Spatial-Aware Learning in Feature Embedding and Classification for One-Stage 3-D Object DetectionabstractOne-stage 3D object detection, known for its simplicity and high-speed inference, is attracting increasing attention in autonomous driving scenarios. However, current one-stage detectors tend to perform sub-optimally compared to two-stage competitors. Our experimental findings suggest that one-stage detectors underperform due to the underutilization of spatial information in feature embedding and classification. Concretely, the spatial context is severely lost during feature propagation, inducing distorted spatial awareness. On the other hand, category recognition relies on the full utilization of spatial information, which is neglected by current detectors. This inadequate spatial awareness of the classification branch can exacerbate misclassification. To address these issues, we propose Spatial-aware Learning in Feature Embedding and Classification for One-stage 3D Object Detection (SLDet). Specifically, to restore the distorted spatial awareness, Category-wise Spatial Augmentation (CSA) is proposed to adaptively bring the network with pre-encoding multi-scale spatial contexts. As for misclassification, Spatial Guiding Classification (SGC) is introduced to guide the classification using explicit scale information. It employs the natural scale divergences among categories to rectify misclassification. Comprehensive experiments demonstrate that SLDet efficiently utilizes spatial information and achieves newly state-of-the-art performance on both the Waymo Open Dataset and the ONCE Dataset. Furthermore, additional experiments demonstrate the excellent generalization capacity of SLDet. Yiqiang Wu, Weiping Xiao, Jiantao Gao, Chang Liu 0082, Yan Peng 0001, Xiaomao Li |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Spatio-Temporal Contextual Learning for Single Object Tracking on Point CloudsabstractSingle object tracking (SOT) is one of the most active research directions in the field of computer vision. Compared with the 2-D image-based SOT which has already been well-studied, SOT on 3-D point clouds is a relatively emerging research field. In this article, a novel approach, namely, the contextual-aware tracker (CAT), is investigated to achieve a superior 3-D SOT through spatially and temporally contextual learning from the LiDAR sequence. More precisely, in contrast to the previous 3-D SOT methods merely exploiting point clouds in the target bounding box as the template, CAT generates templates by adaptively including the surroundings outside the target box to use available ambient cues. This template generation strategy is more effective and rational than the previous area-fixed one, especially when the object has only a small number of points. Moreover, it is deduced that LiDAR point clouds in 3-D scenes are often incomplete and significantly vary from frame to another, which makes the learning process more difficult. To this end, a novel cross-frame aggregation (CFA) module is proposed to enhance the feature representation of the template by aggregating the features from a historical reference frame. Leveraging such schemes enables CAT to achieve a robust performance, even in the case of extremely sparse point clouds. The experiments confirm that the proposed CAT outperforms the state-of-the-art methods on both the KITTI and NuScenes benchmarks, achieving 3.9% and 5.6% improvements in terms of precision. Jiantao Gao, Xu Yan 0005, Weibing Zhao, Zhen Lyu, Yinghong Liao, Chaoda Zheng |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Mitigate the classification ambiguity via localization-classification sequence in object detection
Chang Liu 0082, Shaorong Xie, Xiaomao Li, Jiantao Gao, Weiping Xiao, Baojie Fan, Yan Peng 0001 |
Pattern Recognit. | 4 |
| 2023 | Balanced Sample Assignment and Objective for Single-Model Multi-Class 3D Object DetectionabstractAccurately detecting multi-class objects in a single pass is critical but challenging for real-world autonomous driving scenarios. Several single-class anchor-based methods have recently achieved the state-of-the-art performance in the car category, but when extending to multi-class detection tasks, their performance on small objects (i.e., pedestrians and cyclists) is limited. We find that the core problem that causes this phenomenon lies in the unbalanced sample quality and the classification objective. To address this problem, we proposed a single-model multi-class 3D object detector with balanced sample assignment and objective, named BSAODet. Specifically, the quality-balanced sample assignment (QBSA) is introduced to dynamically collect stable high-quality samples for each class according to the predicted sample performance and geometric constraints. In conjunction with the QBSA, the class-balanced classification objective (CBCO) performs instance-wise label normalization and weighting on positive samples, preventing the model from biasing toward objects with more samples. Extensive experiments on the popular KITTI dataset, the latest large-scale ONCE dataset, and the challenging Waymo Open Dataset show that our method steadily improves the performance of current state-of-the-art detectors by 2–7 mAP in pedestrians and cyclists while maintaining competitiveness in cars. Moreover, our best model achieves 66.31 mAP on three classes, outperforming all published LiDAR-only detectors on the KITTI benchmark. Weiping Xiao, Yan Peng 0001, Chang Liu 0082, Jiantao Gao, Yiqiang Wu, Xiaomao Li |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | 2DPASS: 2D Priors Assisted Semantic Segmentation on LiDAR Point Clouds
Xu Yan 0005, Jiantao Gao, Chaoda Zheng, Chao Zheng 0004, Ruimao Zhang, Shuguang Cui, Zhen Li 0026 |
ECCV (28) | 2 |
| 2022 | Let Images Give You More: Point Cloud Cross-Modal Training for Shape AnalysisabstractAlthough recent point cloud analysis achieves impressive progress, the paradigm of representation learning from single modality gradually meets its bottleneck. In this work, we take a step towards more discriminative 3D point cloud representation using 2D images, which inherently contain richer appearance information, e.g., texture, color, and shade. Specifically, this paper introduces a simple but effective point cloud cross-modality training (PointCMT) strategy, which utilizes view-images, i.e., rendered or projected 2D images of the 3D object, to boost point cloud classification. In practice, to effectively acquire auxiliary knowledge from view-images, we develop a teacher-student framework and formulate the cross-modal learning as a knowledge distillation problem. Through novel feature and classifier enhancement criteria, PointCMT eliminates the distribution discrepancy between different modalities and avoid potential negative transfer effectively. Note that PointCMT efficiently improves the point-only representation without any architecture modification. Sufficient experiments verify significant gains on various datasets based on several backbones, i.e., equipped with PointCMT, PointNet++ and PointMLP achieve state-of-the-art performance on two benchmarks, i.e., 94.4% and 86.7% accuracy on ModelNet40 and ScanObjectNN, respectively. Xu Yan 0005, Heshen Zhan, Chaoda Zheng, Jiantao Gao, Ruimao Zhang, Shuguang Cui, Zhen Li 0026 |
NeurIPS | 4 |
| 2022 | RE-Det3D: RoI-enhanced 3D object detector
Yiqiang Wu, Weiping Xiao, Chang Liu 0082, Jiantao Gao, Guozhu Tan, Xiaomao Li |
Image Vis. Comput. | 4 |
| 2022 | 3D-VDNet: Exploiting the vertical distribution characteristics of point clouds for 3D object detection and augmentation
Weiping Xiao, Xiaomao Li, Chang Liu 0082, Jiantao Gao, Jun Luo 0006, Yan Peng 0001 |
Image Vis. Comput. | 4 |
| 2021 | Sparse Single Sweep LiDAR Point Cloud Segmentation via Learning Contextual Shape Priors from Scene CompletionabstractLiDAR point cloud analysis is a core task for 3D computer vision, especially for autonomous driving. However, due to the severe sparsity and noise interference in the single sweep LiDAR point cloud, the accurate semantic segmentation is non-trivial to achieve. In this paper, we propose a novel sparse LiDAR point cloud semantic segmentation framework assisted by learned contextual shape priors. In practice, an initial semantic segmentation (SS) of a single sweep point cloud can be achieved by any appealing network and then flows into the semantic scene completion (SSC) module as the input. By merging multiple frames in the LiDAR sequence as supervision, the optimized SSC module has learned the contextual shape priors from sequential LiDAR data, completing the sparse single sweep point cloud to the dense one. Thus, it inherently improves SS optimization through fully end-to-end training. Besides, a Point-Voxel Interaction (PVI) module is proposed to further enhance the knowledge fusion between SS and SSC tasks, i.e., promoting the interaction of incomplete local geometry of point cloud and complete voxel-wise global structure. Furthermore, the auxiliary SSC and PVI modules can be discarded during inference without extra burden for SS. Extensive experiments confirm that our JS3C-Net achieves superior performance on both SemanticKITTI and SemanticPOSS benchmarks, i.e., 4% and 3% improvement correspondingly. Xu Yan 0005, Jiantao Gao, Jie Li 0002, Ruimao Zhang, Zhen Li 0026, Shuguang Cui |
AAAI | 2 |
| 2021 | Box-Aware Feature Enhancement for Single Object Tracking on Point CloudsabstractCurrent 3D single object tracking approaches track the target based on a feature comparison between the target template and the search area. However, due to the common occlusion in LiDAR scans, it is non-trivial to conduct accurate feature comparisons on severe sparse and incomplete shapes. In this work, we exploit the ground truth bounding box given in the first frame as a strong cue to enhance the feature description of the target object, enabling a more accurate feature comparison in a simple yet effective way. In particular, we first propose the BoxCloud, an informative and robust representation, to depict an object using the point-to-box relation. We further design an efficient box-aware feature fusion module, which leverages the aforementioned BoxCloud for reliable feature matching and embedding. Integrating the proposed general components into an existing model P2B [27], we construct a superior box-aware tracker (BAT)1. Experiments confirm that our proposed BAT outperforms the previous state-of-the-art by a large margin on both KITTI and NuScenes benchmarks, achieving a 12.8% improvement in terms of precision while running ∼20% faster. Chaoda Zheng, Xu Yan 0005, Jiantao Gao, Weibing Zhao, Wei Zhang 0001, Zhen Li 0026, Shuguang Cui |
ICCV | 3 |
| 2021 | PointLIE: Locally Invertible Embedding for Point Cloud Sampling and RecoveryabstractPoint Cloud Sampling and Recovery (PCSR) is critical for massive real-time point cloud collection and processing since raw data usually requires large storage and computation. This paper addresses a fundamental problem in PCSR: How to downsample the dense point cloud with arbitrary scales while preserving the local topology of discarded points in a case-agnostic manner (i.e., without additional storage for point relationships)? We propose a novel Locally Invertible Embedding (PointLIE) framework to unify the point cloud sampling and upsampling into one single framework through bi-directional learning. Specifically, PointLIE decouples the local geometric relationships between discarded points from the sampled points by progressively encoding the neighboring offsets to a latent variable. Once the latent variable is forced to obey a pre-defined distribution in the forward sampling path, the recovery can be achieved effectively through inverse operations. Taking the recover-pleasing sampled points and a latent embedding randomly drawn from the specified distribution as inputs, PointLIE can theoretically guarantee the fidelity of reconstruction and outperform state-of-the-arts quantitatively and qualitatively. Weibing Zhao, Xu Yan 0005, Jiantao Gao, Ruimao Zhang, Jiayan Zhang, Zhen Li 0026, Shuguang Cui |
IJCAI | 3 |
| 2020 | Characterizing Label Errors: Confident Learning for Noisy-Labeled Image Segmentation
Minqing Zhang, Jiantao Gao, Zhen Lyu, Weibing Zhao, Qin Wang 0011, Weizhen Ding, Sheng Wang 0001, Zhen Li 0026, Shuguang Cui |
MICCAI (1) | 2 |
| 2020 | Diverse receptive field network with context aggregation for fast object detection
Shaorong Xie, Chang Liu 0082, Jiantao Gao, Xiaomao Li, Jun Luo 0006, Baojie Fan, Jiahong Chen, Huayan Pu, Yan Peng 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2020 | An Efficient Entropy-Based Causal Discovery Method for Linear Structural Equation Models With IID Noise VariablesabstractThe discovery of causal relationships from the observational data is an important task. To identify the unique causal structure belonging to a Markov equivalence class, a number of algorithms, such as the linear non-Gaussian acyclic model (LiNGAM), have been proposed. However, two challenges remain to be met: 1) these algorithms fail to work on the data which follow linear structural equation model with Gaussian noise and 2) they misjudge the causal direction when the data contain additional measurement errors. In this paper, we propose an entropy-based two-phase iterative algorithm for arbitrary distribution data with additional measurement errors under some mild assumptions. In the first phase of the algorithm, based on the property that entropy can measure the amount of information behind the data with arbitrary distribution, we design a general approach for the identification of exogenous variable on both Gaussian and non-Gaussian data, and we give the corresponding theoretical derivation. In the second phase, to eliminate the effects of measurement errors, we revise the value of the exogenous variable by removing its measurement error and further use the revised value to remove its effect on the remaining variables. Experimental results on real-world causal structures are presented to demonstrate the effectiveness and stability of our method. We also apply the proposed algorithm on the mobile-base-station data with measurement errors, and the results further prove the effectiveness of our algorithm. Feng Xie 0002, Ruichu Cai, Yan Zeng 0002, Jiantao Gao, Zhifeng Hao 0004 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |