VLDB 2026 Research / reviewers in the wild / expert
Jie Wang 0097
dblp:29/5259-97
· DBLP profile ↗
12ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0002-4847-3697ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Local grid rendering networks for 3D object detection in point clouds
Jianan Li 0001, Lihe Ding, Jie Wang 0097, Tingfa Xu |
Pattern Recognit. | 3 |
| 2025 | PvNeXt: Rethinking Network Design and Temporal Motion for Point Cloud Video RecognitionabstractPoint cloud video perception has become an essential task for the realm of 3D vision. Current 4D representation learning techniques typically engage in iterative processing coupled with dense query operations. Although effective in capturing temporal features, this approach leads to substantial computational redundancy. In this work, we propose a framework, named as PvNeXt, for effective yet efficient point cloud video recognition, via personalized one-shot query operation. Specially, PvNeXt consists of two key modules, the Motion Imitator and the Single-Step Motion Encoder. The former module, the Motion Imitator, is designed to capture the temporal dynamics inherent in sequences of point clouds, thus generating the virtual motion corresponding to each frame. The Single-Step Motion Encoder performs a one-step query operation, associating point cloud of each frame with its corresponding virtual motion frame, thereby extracting motion cues from point cloud sequences and capturing temporal dynamics across the entire sequence. Through the integration of these two modules, {PvNeXt} enables personalized one-shot queries for each frame, effectively eliminating the need for frame-specific looping and intensive query processes. Extensive experiments on multiple benchmarks demonstrate the effectiveness of our method. Jie Wang 0097, Tingfa Xu, Lihe Ding, Long Bai 0008, Jianan Li 0001 |
ICLR | 1 |
| 2025 | Endo-4DGX: Robust Endoscopic Scene Reconstruction and Illumination Correction with Gaussian Splatting
Yiming Huang 0007, Long Bai 0008, Beilei Cui, Yanheng Li 0002, Tong Chen 0011, Jie Wang 0097, Jinlin Wu, Zhen Lei 0001, Hongbin Liu 0001, Hongliang Ren 0001 |
MICCAI (9) | 6 |
| 2025 | Towards Robust Point Cloud Recognition With Sample-Adaptive Auto-AugmentationabstractRobust 3D perception amidst corruption is a crucial task in the realm of 3D vision. Conventional data augmentation methods aimed at enhancing corruption robustness typically apply random transformations to all point cloud samples offline, neglecting sample structure, which often leads to over- or under-enhancement. In this study, we propose an alternative approach to address this issue by employing sample-adaptive transformations based on sample structure, through an auto-augmentation framework named AdaptPoint++. Central to this framework is an imitator, which initiates with Position-aware Feature Extraction to derive intrinsic structural information from the input sample. Subsequently, a Deformation Controller and a Mask Controller predict per-anchor deformation and per-point masking parameters, respectively, facilitating corruption simulations. In conjunction with the imitator, a discriminator is employed to curb the generation of excessive corruption that deviates from the original data distribution. Moreover, we integrate a perception-guidance feedback mechanism to steer the generation of samples towards an appropriate difficulty level. To effectively train the classifier using the generated augmented samples, we introduce a Structure Reconstruction-assisted learning mechanism, bolstering the classifier's robustness by prioritizing intrinsic structural characteristics over superficial discrepancies induced by corruption. Additionally, to alleviate the scarcity of real-world corrupted point cloud data, we introduce two novel datasets: ScanObjectNN-C and MVPNET-C, closely resembling actual data in real-world scenarios. Experimental results demonstrate that our method attains state-of-the-art performance on multiple corruption benchmarks. Jianan Li 0001, Jie Wang 0097, Tingfa Xu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | PAPooling: Graph-based Position Adaptive Aggregation of Local Geometry in Point CloudsabstractFine-grained geometry, obtained through the assimilation of localized point features, is crucial in the realms of object recognition and scene comprehension within point cloud contexts. Traditional point cloud backbones predominantly utilize max pooling for the amalgamation of local features, a process that tends to overlook spatial interrelations among points, consequently leading to the potential loss of fine-grained geometric details. To overcome this limitation, we introduce an innovative operation termed Position Adaptive Pooling (PAPooling), which is designed to amalgamate local features while sensitively considering the spatial positions of points. This is achieved by employing a graph-based representation to explicitly model the spatial relationships of points. PAPooling involves two principal components: first, the local graph construction , which establishes a local graph for a set of points by linking a central point with its adjacent points, thereby transforming pairwise relative positions into channel-specific attention weights; second, the attentive feature aggregation , which adeptly takes into account the contribution of each node and simulates the inter-node relationships within the local graph, effectively extracting representations of local features through a Graph Convolution Network (GCN). PAPooling’s simplicity and efficacy make it a versatile addition to widely used point-based backbones such as PointNet++ and DGCNN, offering a plug-and-play solution. Comprehensive experimental analysis demonstrates PAPooling’s enhanced capability in capturing local geometry, contributing significantly across a spectrum of applications including 3D shape classification, part segmentation, scene segmentation, and corruption defense, all with minimal computational increase. Code will be public at https://github.com/Roywangj/PAPooling/ . Jie Wang 0097, Tingfa Xu, Liqiang Song, Lihe Ding, Peng Jiang 0013, Yuqi Han, Jianan Li 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Dual-Stage Hyperspectral Image Classification Model with Spectral Supertoken
Peifu Liu, Tingfa Xu, Jie Wang 0097, Huan Chen 0018, Huiyan Bai, Jianan Li 0001 |
ECCV (34) | 3 |
| 2024 | OSSAR: Towards Open-Set Surgical Activity Recognition in Robot-assisted SurgeryabstractIn the realm of automated robotic surgery and computer-assisted interventions, understanding robotic surgical activities stands paramount. Existing algorithms dedicated to surgical activity recognition predominantly cater to pre-defined closed-set paradigms, ignoring the challenges of real-world open-set scenarios. Such algorithms often falter in the presence of test samples originating from classes unseen during training phases. To tackle this problem, we introduce an innovative Open-Set Surgical Activity Recognition (OSSAR) framework. Our solution leverages the hyperspherical reciprocal point strategy to enhance the distinction between known and unknown classes in the feature space. Additionally, we address the issue of over-confidence in the closed set by refining model calibration, avoiding misclassification of unknown classes as known ones. To support our assertions, we establish an open-set surgical activity benchmark utilizing the public JIGSAWS dataset. Besides, we also collect a novel dataset on endoscopic submucosal dissection for surgical activity tasks. Extensive comparisons and ablation experiments on these datasets demonstrate the significant outperformance of our method over existing state-of-the-art approaches. Our proposed solution can effectively address the challenges of real-world surgical scenarios. Our code is publicly accessible at github.com/longbai1006/OSSAR. Long Bai 0008, Guankun Wang, Jie Wang 0097, Xiaoxiao Yang, Huxin Gao, An Wang 0007, Mobarakol Islam, Hongliang Ren 0001 |
ICRA | 3 |
| 2024 | Target-Guided Adversarial Point Cloud Transformer Towards Recognition Against Real-world CorruptionsabstractAchieving robust 3D perception in the face of corrupted data presents an challenging hurdle within 3D vision research. Contemporary transformer-based point cloud recognition models, albeit advanced, tend to overfit to specific patterns, consequently undermining their robustness against corruption. In this work, we introduce the Target-Guided Adversarial Point Cloud Transformer, termed APCT, a novel architecture designed to augment global structure capture through an adversarial feature erasing mechanism predicated on patterns discerned at each step during training. Specifically, APCT integrates an Adversarial Significance Identifier and a Target-guided Promptor. The Adversarial Significance Identifier, is tasked with discerning token significance by integrating global contextual analysis, utilizing a structural salience index algorithm alongside an auxiliary supervisory mechanism. The Target-guided Promptor, is responsible for accentuating the propensity for token discard within the self-attention mechanism, utilizing the value derived above, consequently directing the model attention towards alternative segments in subsequent stages. By iteratively applying this strategy in multiple steps during training, the network progressively identifies and integrates an expanded array of object-associated patterns. Extensive experiments demonstrate that our method achieves state-of-the-art results on multiple corruption benchmarks. Jie Wang 0097, Tingfa Xu, Lihe Ding, Jianan Li 0001 |
NeurIPS | 1 |
| 2024 | PointGL: A Simple Global-Local Framework for Efficient Point Cloud AnalysisabstractEfficient analysis of point clouds holds paramount significance in real-world 3D applications. Currently, prevailing point-based models adhere to the PointNet++ methodology, which involves embedding and abstracting point features within a sequence of spatially overlapping local point sets, resulting in noticeable computational redundancy. Drawing inspiration from the streamlined paradigm of pixel embedding followed by regional pooling in Convolutional Neural Networks (CNNs), we introduce a novel, uncomplicated yet potent architecture known as PointGL, crafted to facilitate efficient point cloud analysis. PointGL employs a hierarchical process of feature acquisition through two recursive steps. First, theGlobal Point Embeddingleverages straightforward residual Multilayer Perceptrons (MLPs) to effectuate feature embedding for each individual point. Second, the novelLocal Graph Poolingtechnique characterizes point-to-point relationships and abstracts regional representations through succinct local graphs. The harmonious fusion of one-time point embedding and parameter-free graph pooling contributes to PointGL's defining attributes of minimized model complexity and heightened efficiency. Our PointGL attains state-of-the-art accuracy on the ScanObjectNN dataset while exhibiting a runtime that is more than 5 times faster and utilizing only approximately 4% of the FLOPs and 30% of the parameters compared to the recent PointMLP model. The code for PointGL is available athttps://github.com/Roywangj/PointGL. Jianan Li 0001, Jie Wang 0097, Tingfa Xu |
IEEE Trans. Multim. | 2 |
| 2023 | Sample-adaptive Augmentation for Point Cloud Recognition Against Real-world CorruptionsabstractRobust 3D perception under corruption has become an essential task for the realm of 3D vision. While current data augmentation techniques usually perform random transformations on all point cloud objects in an offline way and ignore the structure of the samples, resulting in over-or-under enhancement. In this work, we propose an alternative to make sample-adaptive transformations based on the structure of the sample to cope with potential corruption via an auto-augmentation framework, named as Adapt-Point. Specially, we leverage a imitator, consisting of a Deformation Controller and a Mask Controller, respectively in charge of predicting deformation parameters and producing a per-point mask, based on the intrinsic structural information of the input point cloud, and then conduct corruption simulations on top. Then a discriminator is utilized to prevent the generation of excessive corruption that deviates from the original data distribution. In addition, a perception-guidance feedback mechanism is incorporated to guide the generation of samples with appropriate difficulty level. Furthermore, to address the paucity of real-world corrupted point cloud, we also introduce a new dataset ScanObjectNN-C, that exhibits greater similarity to actual data in real-world environments, especially when contrasted with preceding CAD datasets. Experiments show that our method achieves state-of-the-art results on multiple corruption benchmarks, including ModelNet-C, our ScanObjectNN-C, and ShapeNet-C. Jie Wang 0097, Lihe Ding, Tingfa Xu, Shaocong Dong, Xinli Xu, Long Bai 0008, Jianan Li 0001 |
ICCV | 1 |
| 2022 | FH-Net: A Fast Hierarchical Network for Scene Flow Estimation on Real-World Point Clouds
Lihe Ding, Shaocong Dong, Tingfa Xu, Xinli Xu, Jie Wang 0097, Jianan Li 0001 |
ECCV (39) | 5 |
| 2022 | MsSVT: Mixed-scale Sparse Voxel Transformer for 3D Object Detection on Point Cloudsabstract3D object detection from the LiDAR point cloud is fundamental to autonomous driving. Large-scale outdoor scenes usually feature significant variance in instance scales, thus requiring features rich in long-range and fine-grained information to support accurate detection. Recent detectors leverage the power of window-based transformers to model long-range dependencies but tend to blur out fine-grained details. To mitigate this gap, we present a novel Mixed-scale Sparse Voxel Transformer, named MsSVT, which can well capture both types of information simultaneously by the divide-and-conquer philosophy. Specifically, MsSVT explicitly divides attention heads into multiple groups, each in charge of attending to information within a particular range. All groups' output is merged to obtain the final mixed-scale features. Moreover, we provide a novel chessboard sampling strategy to reduce the computational complexity of applying a window-based transformer in 3D voxel space. To improve efficiency, we also implement the voxel sampling and gathering operations sparsely with a hash map. Endowed by the powerful capability and high efficiency of modeling mixed-scale information, our single-stage detector built on top of MsSVT surprisingly outperforms state-of-the-art two-stage detectors on Waymo. Our project page: https://github.com/dscdyc/MsSVT. Shaocong Dong, Lihe Ding, Tingfa Xu, Xinli Xu, Jie Wang 0097, Ziyang Bian, Ying Wang 0064, Jianan Li 0001 |
NeurIPS | 6 |