VLDB 2026 Research / reviewers in the wild / expert
Tuo Feng 0001
dblp:203/2907-1
· DBLP profile ↗
7ranked-venue papers
6as first author
5since 2021 · last 2025
0000-0001-5882-3315ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 5 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Gaussian-Based World Model: Gaussian Priors for Voxel-Based Occupancy Prediction and Future Motion Prediction
Tuo Feng 0001, Wenguan Wang, Yi Yang 0001 |
ICCV | 1 |
| 2024 | Interpretable3D: An Ad-Hoc Interpretable Classifier for 3D Point Cloudsabstract3D decision-critical tasks urgently require research on explanations to ensure system reliability and transparency. Extensive explanatory research has been conducted on 2D images, but there is a lack in the 3D field. Furthermore, the existing explanations for 3D models are post-hoc and can be misleading, as they separate explanations from the original model. To address these issues, we propose an ad-hoc interpretable classifier for 3D point clouds (i.e., Interpretable3D). As an intuitive case-based classifier, Interpretable3D can provide reliable ad-hoc explanations without any embarrassing nuances. It allows users to understand how queries are embedded within past observations in prototype sets. Interpretable3D has two iterative training steps: 1) updating one prototype with the mean of the embeddings within the same sub-class in Prototype Estimation, and 2) penalizing or rewarding the estimated prototypes in Prototype Optimization. The mean of embeddings has a clear statistical meaning, i.e., class sub-centers. Moreover, we update prototypes with their most similar observations in the last few epochs. Finally, Interpretable3D classifies new samples according to prototypes. We evaluate the performance of Interpretable3D on four popular point cloud models: DGCNN, PointNet2, PointMLP, and PointNeXt. Our Interpretable3D demonstrates comparable or superior performance compared to softmax-based black-box models in the tasks of 3D shape classification and part segmentation. Our code is released at: github.com/FengZicai/Interpretable3D. Tuo Feng 0001, Ruijie Quan, Wenguan Wang, Yi Yang 0001 |
AAAI | 1 |
| 2024 | LSK3DNet: Towards Effective and Efficient 3D Perception with Large Sparse KernelsabstractAutonomous systems need to process large-scale, sparse, and irregular point clouds with limited compute resources. Consequently, it is essential to develop LiDAR perception methods that are both efficient and effective. Although naive-ly enlarging 3D kernel size can enhance performance, it will also lead to a cubically-increasing overhead. Therefore, it is crucial to develop streamlined 3D large kernel designs that eliminate redundant weights and work effectively with larger kernels. In this paper, we propose an efficient and effective Large Sparse Kernel 3D Neural Network (LSK3DNet) that leverages dynamic pruning to amplify the 3D kernel size. Our method comprises two core components: Spatial-wise Dynamic Sparsity (SDS) and Channel-wise Weight Selection (CWS). SDS dynamically prunes and regrows volumetric weights from the beginning to learn a large sparse 3D kernel. It not only boosts performance but also significantly reduces model size and computational cost. Moreover, CWS selects the most important channels for 3D convolution during training and subsequently prunes the redundant channels to accelerate inference for 3D vision tasks. We demonstrate the effectiveness of LSK3DNet on three benchmark datasets and five tracks compared with classical models and large kernel designs. Notably, LSK3DNet achieves the state-of-the-art performance on SemanticKITTI (i.e., 75.6% on single-scan and 63.4% on multi-scan), with roughly 40% model size re-duction and 60% computing operations reduction compared to the naive large 3D kernel model. Tuo Feng 0001, Wenguan Wang, Fan Ma, Yi Yang 0001 |
CVPR | 1 |
| 2024 | Shape2Scene: 3D Scene Representation Learning Through Pre-training on Shape Data
Tuo Feng 0001, Wenguan Wang, Ruijie Quan, Yi Yang 0001 |
ECCV (55) | 1 |
| 2023 | Clustering based Point Cloud Representation Learning for 3D AnalysisabstractPoint cloud analysis (such as 3D segmentation and detection) is a challenging task, because of not only the irregular geometries of many millions of unordered points, but also the great variations caused by depth, viewpoint, occlusion, etc. Current studies put much focus on the adaption of neural networks to the complex geometries of point clouds, but are blind to a fundamental question: how to learn an appropriate point embedding space that is aware of both discriminative semantics and challenging variations? As a response, we propose a clustering based supervised learning scheme for point cloud analysis. Unlike current de-facto, scene-wise training paradigm, our algorithm conducts within-class clustering on the point embedding space for automatically discovering subclass patterns which are latent yet representative across scenes. The mined patterns are, in turn, used to repaint the embedding space, so as to respect the underlying distribution of the entire training dataset and improve the robustness to the variations. Our algorithm is principled and readily pluggable to modern point cloud segmentation networks during training, without extra overhead during testing. With various 3D network architectures (i.e., voxel-based, point-based, Transformer-based, automatically searched), our algorithm shows notable improvements on famous point cloud segmentation datasets (i.e., 2.0-2.6% on single-scan and 2.0-2.2% multi-scan of SemanticKITTI, 1.8-1.9% on S3DIS, in terms of mIoU). Our algorithm also demonstrates utility in 3D detection, showing 2.0-3.4% mAP gains on KITTI. Tuo Feng 0001, Wenguan Wang, Yi Yang 0001 |
ICCV | 1 |
| 2020 | A Novel Object Re-Track Framework for 3D Point Cloudsabstract3D point cloud data is an important data source for autonomous vehicles to perceive the surroundings. Achieving accurate object tracking of 3D point clouds has become a challenging task. In this paper, we propose a 3D object two-stage re-track framework directly utilizing point clouds as the input, without using the ground truth as the reference box. The framework consists of a coarse stage and a fine stage. By tracking back the previous T frames and expanding the search space for each frame, we add the fine stage to re-track the lost objects of the coarse stage. Moreover, we design a dense AutoEncoder to enhance the discrimination in the latent space and improve shape completion performance, thus improving tracking performance. A Sample Update Strategy is also proposed to aggregate similar model shape samples in different frames, which improves the quality of the model shape. In terms of motion models for the proposed re-track framework, we further compare Kalman Filter with PointLSTM and do an extensive analysis. Finally, we test the re-track framework on the KITTI tracking dataset and outperform the public benchmark by 17.1%/15.5% in Success and Precision, respectively. Our code and model are available at https://github.com/FengZicai/Re-Track. Tuo Feng 0001, Licheng Jiao, Hao Zhu 0009 |
ACM Multimedia | 1 |
| 2019 | A Dense Pointnet++ Architecture for 3D Point Cloud Semantic Segmentationabstract3D point cloud data has been widely used in remote sensing mapping because it is not affected by lighting, shadows and other factors. How to improve the performance of semantic segmentation of 3D point cloud data has attracted more and more attention. Previous works connected shallow features in encoders directly with deep features in decoders, which will lead to semantic gap. In this paper, we propose a Dense PointNet++ architecture, called DPNet, for semantic segmentation of 3D point cloud data. In order to weaken the semantic gap, multiple nested up-sampling layers and a series of cumulative, short and long skip link concatenation are introduced in the network to obtain more abundant point cloud features. Grid map and model fusion are used to further correct the results of network segmentation. The experimental results on US3D data set show that DPNet is superior to existing advanced architectures, especially for the categories with small samples. Moreover, DPNet with grid map and model fusion ranks the first place in 2019 IEEE GRSS Data fusion contest 3D point cloud classification challenge. Yanchao Lian, Tuo Feng 0001, Jinliu Zhou |
IGARSS | 2 |