EDBT 2026 Demo / reviewers in the wild / expert
Sheng Ao
dblp:266/9135
· DBLP profile ↗
16ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0001-6896-1869ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 10 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RCP-LO: A Relative Coordinate Prediction Framework for Generalizable Deep LiDAR OdometryabstractLiDAR odometry is a critical component of SLAM in autonomous driving and robotics. Learning-based methods have shown remarkable performance by regressing relative poses in an end-to-end manner. However, when applying these trained models, originally developed on the widely used KITTI dataset, to other scenes, performance often drops significantly. In other words, existing methods struggle to generalize well to new environments. To address this challenge, we propose RCP-LO, a simple yet effective LiDAR odometry framework. We introduce a novel representation for relative poses, reformulating them as relative coordinates, which can then be solved using geometrical verification. This approach avoids overly simplified pose representations and makes better use of scene geometry, thereby improving generalization. Moreover, to capture the inherent uncertainties in relative pose estimation from occluded LiDAR point clouds from dynamic environments, we adapt our framework to learn a denoising diffusion model, allowing for sampling plausible relative coordinates while improving robustness. We also introduce a differentiable geometric weighted singular value decomposition module, enabling efficient pose estimation through a single forward pass. Extensive experiments demonstrate that RCP-LO, trained exclusively on the KITTI dataset, achieves competitive performance compared to SOTA learning-based methods and generalizes effectively to the KITTI-360, Ford, and Oxford datasets. Wen Li 0005, Yongshu Huang, Minghang Zhu, Yuyang Yang, Dunqiang Liu, Sheng Ao, Cheng Wang 0003 |
AAAI | 7 |
| 2026 | POSITION: Open World 3D Scene CAD Recompositionabstract3D scene CAD recomposition aims to reconstruct a given scene by retrieving and assembling CAD models from a database, so as to accurately simulate the geometric properties and spatial arrangement of the original environment. Recent methods learn this task through training on limited scan-to-CAD annotation data, which hinders their generalization to diverse real-world scenes. In this paper, we propose POSITION, an open-world 3D scene CAD recomposition method to construct the 3D scene with CADs retrieved from an open-set database. POSITION is designed following a divide-and-conquer strategy. Firstly, we extract open-world multi-modal object representations from a captured 3D scene. Secondly, on top of the representations, we propose a coarse-to-fine retrieval method to retrieve CADs that are visually, geometrically and semantically match real objects. Thirdly, we present a physically plausible pose alignment method to adjust retrieved CAD models to maintain consistent geometry and layout with the observation. By decomposing the problem into well-defined subtasks, our approach achieves generalization across various scene types and scalable CAD databases without retraining or fine-tuning. Our approach demonstrates superior CAD recomposition performance on both the Scan2CAD and diverse real-world 3D scene datasets. Our project page: https://yangrongkun.github.io/position/. Rongkun Yang, Hongda Liu 0001, Sheng Ao, Longguang Wang, Shunbo Zhou, Yulan Guo |
IEEE Trans. Image Process. | 4 |
| 2025 | DiffLO: Semantic-Aware LiDAR Odometry with Diffusion-Based RefinementabstractLiDAR odometry is a critical module in autonomous driving systems, responsible for accurate localization by estimating the relative pose transformation between consecutive point cloud frames. However, existing studies frequently encounter challenges with unreliable pose estimation, due to the lack of in-depth understanding of scenario and the presence of noise interference. To address this challenge, we propose DiffLO, a semantic-aware LiDAR odometry network with diffusion-based refinement. To mitigate the impact of challenging cases such as dynamic, repetitive patterns, and low textures, we introduce a semantic distillation method that integrates semantic information into the odometry task. This allows the network to gain a semantic understanding of the scene, enabling it to focus more on the objects that are beneficial for pose estimation. Additionally, to enhance the robustness, we propose a diffusion-based refinement method. This method uses pose-related features as conditional constraints for generative diversity, iteratively refining the pose estimation to achieve greater accuracy. Comparative experiments on the KITTI odometry dataset demonstrate that the proposed method achieves state-of-the-art performance among existing learning-based approaches. Furthermore, the proposed DiffLO method outperforms the classic A-LOAM on most evaluation sequences. The code will be released at https://github.com/hytree7/difflo. Yongshu Huang, Minghang Zhu, Sheng Ao, Chenglu Wen, Cheng Wang 0003 |
CVPR | 4 |
| 2025 | DropoutGS: Dropping Out Gaussians for Better Sparse-view RenderingabstractAlthough 3D Gaussian Splatting (3DGS) has demonstrated promising results in novel view synthesis, its performance degrades dramatically with sparse inputs and generates undesirable artifacts. As the number of training views decreases, the novel view synthesis task degrades to a highly under-determined problem such that existing methods suffer from the notorious overfitting issue. Interestingly, we observe that models with fewer Gaussian primitives exhibit less overfitting under spare inputs. Inspired by this observation, we propose a Random Dropout Regularization (RDR) to exploit the advantages of low-complexity models to alleviate overfitting. In addition, to remedy the lack of high-frequency details for these models, an Edge-guided Splitting Strategy (ESS) is developed. With these two techniques, our method (termed DropoutGS) provides a simple yet effective plug-in approach to improve the generalization performance of existing 3DGS methods. Extensive experiments show that our DropoutGS produces state-of-the-art performance under sparse views on benchmark datasets including Blender, LLFF, and DTU. The project page is at: https://xuyx55.github.io/DropoutGS/. Yexing Xu, Longguang Wang, Minglin Chen, Sheng Ao, Li Li 0100, Yulan Guo |
CVPR | 4 |
| 2025 | Progressive Correspondence Regenerator for Robust 3D RegistrationabstractObtaining enough high-quality correspondences is crucial for robust registration. Existing correspondence refinement methods mostly follow the paradigm of outlier removal, which either fails to correctly identify the accurate correspondences under extreme outlier ratios, or select too few correct correspondences to support robust registration. To address this challenge, we propose a novel approach named Regor, which is a progressive correspondence regenerator that generates higher-quality matches whist sufficiently robust for numerous outliers. In each iteration, we first apply prior-guided local grouping and generalized mutual matching to generate the local region correspondences. A powerful center-aware three-point consistency is then presented to achieve local correspondence correction, instead of removal. Further, we employ global correspondence refinement to obtain accurate correspondences from a global perspective. Through progressive iterations, this process yields a large number of high-quality correspondences. Extensive experiments on both indoor and outdoor datasets demonstrate that the proposed Regor significantly outperforms existing outlier removal techniques. More critically, our approach obtain 10 times more correct correspondences than outlier removal methods. As a result, our method is able to achieve robust registration even with weak features. The code is available at [Regor]. Guiyu Zhao, Sheng Ao, Ye Zhang 0037, Kai Xu 0004, Yulan Guo |
CVPR | 2 |
| 2025 | RALoc: Enhancing Outdoor LiDAR Localization via Rotation Awareness
Yuyang Yang, We Li, Sheng Ao, Shangshu Yu |
ICCV | 3 |
| 2025 | $U^2$ Frame: A Unified and Unsupervised Learning Framework for LiDAR-Based Loop ClosingabstractLoop closing is critically important in Simultaneous Localization and Mapping (SLAM) due to its ability to correct accumulated localization errors. However, existing methods are hindered by the difficulty of acquiring pose labels and the unreliability of ground truth data. In this paper, we propose$U^{2}$Frame, a unified LiDAR-based loop closing framework that handles both loop closure detection and relative pose estimation without any ground truth training data. Specifically, the natural temporal-spatial correlation in point cloud sequences is first leveraged to supervise the network training, where near scans are treated as positives and vice versa as negatives. A new neural architecture is then constructed to jointly learn highly discriminative local and global features for loop closure detection. Additionally, an effective candidate verification module that exploits high-order geometric information is presented to further filter out false loop closures and estimate precise poses. We extensively evaluate$U^{2}$Frame on multiple datasets according to two tasks derived from loop closing: loop closure detection and loop pose estimation. Comparative experiments demonstrate that our method outperforms existing state-of-the-art supervised techniques and has a strong generalization ability across unseen scenarios. Our code is released at https://github.com/yxin-zhang/U2Frame. Sheng Ao, Ye Zhang 0037, Qingyong Hu, Tao Chang, Yulan Guo |
ICRA | 2 |
| 2025 | Unleashing the Power of Data Generation in One-Pass Outdoor LiDAR LocalizationabstractPoint cloud regression localization technology has a wide range of applications in the multimedia field. For example, in virtual reality and augmented reality, accurate point cloud localization can significantly enhance the user experience. Recently, point cloud pose regression algorithms based on APR (Absolute Pose Regression) and SCR (Scene Coordinate Regression) have achieved near sub-meter accuracy, requiring multiple repetitive trajectories for training. The key to their success lies in the diversity of viewpoints, temporal changes, and trajectories, which is resource-consuming. However, due to the errors in GPS/INS, the coupling between trajectories is not ideal, and the stability of re-localization is insufficient. Since LiDAR has covered most of the scene, single-shot localization has the potential to approach or even surpass multi-trajectory localization methods through pose enhancement. Specifically, we present Pose Enhancement Localization (PELoc), which feeds one trajectory, proposing SSDA (Single-shot Data Augmentation) and LTI (LiDAR Trajectories-coupled Interpolation) to simulate different driving poses, and we introduce KP-CL (Key Points Contrastive Learning) through feature perturbation to mitigate the differences in viewpoint/temporal phase transformations in similar scenes across different trajectories. Our algorithm has been tested on the Oxford, QE-Oxford, and NCLT datasets, where single-shot localization accuracy can approach near sub-meter level on QE-Oxford and NCLT. The code will be published in https://github.com/Eaton2022/PELoc. Yidong Chen 0006, Yuyang Yang, Wen Li 0005, Sheng Ao, Cheng Wang 0003 |
ACM Multimedia | 5 |
| 2025 | Enhancing Event-Based Video Reconstruction With Bidirectional Temporal InformationabstractEvent-based video reconstruction has emerged as an appealing research direction to break through the limitations of traditional cameras to better record dynamic scenes. Most existing methods reconstruct each frame from its corresponding event subset in chronological order. Since the temporal information contained in the whole event sequence is not fully exploited, these methods suffer inferior reconstruction quality. In this paper, we propose to enhance event-based video reconstruction by leveraging the bidirectional temporal information in event sequences. The proposed model processes event sequences in a bidirectional fashion, allowing for exploiting bidirectional information in the whole sequence. Furthermore, a transformer-based temporal information fusion module is introduced to aggregate long-range information in both temporal and spatial dimensions. Additionally, we propose a new dataset for the event-based video reconstruction task which contains a variety of objects and movement patterns. Extensive experiments demonstrate that the proposed model outperforms existing state-of-the-art event-based video reconstruction methods both quantitatively and qualitatively. Pinghai Gao, Longguang Wang, Sheng Ao, Ye Zhang 0037, Yulan Guo |
IEEE Trans. Multim. | 3 |
| 2024 | GRLoR: A Unified Global Retrieval and Local Reranking Framework for 3-D Place RecognitionabstractThree-dimensional place recognition aims to search point cloud in a large database that matches the query. It is an essential task in remote sensing applications, such as smart city management and disaster monitoring. The existing methods commonly leverage global descriptors to perform point cloud retrieval for place recognition. However, these methods rely on spatial aggregation to obtain global descriptors, which are neither discriminative nor general. In this letter, we propose a unified global retrieval and local reranking (namely, GRLoR) framework for 3-D place recognition. Specifically, we first utilize a self-attention mechanism to capture the channel dependencies of local features and design a spatial-fusion pooling (SFP) approach to obtain a discriminative global descriptor for retrieval. We then construct a feature correlation module for local reranking, which uses a cross-attention mechanism to determine whether the point cloud pair matches correctly by predicting the similarity of local regions. Experiments conducted on several public benchmarks validate the superiority performance of our method. For instance, it outperforms the strongest model by an average of about 1% on the public datasets in terms of AR@1. Wenshuo Liu, Sheng Ao, Ye Zhang 0037, Hanyun Wang, Yulan Guo |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2023 | BUFFER: Balancing Accuracy, Efficiency, and Generalizability in Point Cloud RegistrationabstractAn ideal point cloud registration framework should have superior accuracy, acceptable efficiency, and strong generalizability: However, this is highly challenging since existing registration techniques are either not accurate enough, far from efficient, or generalized poorly. It remains an open question that how to achieve a satisfying balance between this three key elements. In this paper, we propose BUFFER, a point cloud registration method for balancing accuracy, efficiency, and generalizability. The key to our approach is to take advantage of both point-wise and patch-wise techniques, while overcoming the inherent drawbacks simultaneously. Different from a simple combination of existing methods, each component of our network has been carefully crafted to tackle specific issues. Specifically, a Point-wise Learner is first introduced to enhance computational efficiency by predicting keypoints and improving the representation capacity of features by estimating point orientations, a Patch-wise Embedder which leverages a lightweight local feature learner is then deployed to extract efficient and general patch features. Additionally, an Inliers Generator which combines simple neural layers and general features is presented to search inlier correspondences. Extensive experiments on real-world scenarios demonstrate that our method achieves the best of both worlds in accuracy, efficiency, and generalization. In particular, our method not only reaches the highest success rate on unseen domains, but also is almost 30 times faster than the strong baselines specializing in generalization. Code is available at https://github.com/aosheng1996/BUFFER. Sheng Ao, Qingyong Hu, Hanyun Wang, Kai Xu 0004, Yulan Guo |
CVPR | 1 |
| 2023 | You Only Train Once: Learning General and Distinctive 3D Local DescriptorsabstractExtracting distinctive, robust, and general 3D local features is essential to downstream tasks such as point cloud registration. However, existing methods either rely on noise-sensitive handcrafted features, or depend on rotation-variant neural architectures. It remains challenging to learn robust and general local feature descriptors for surface matching. In this paper, we propose a new, simple yet effective neural network, termed SpinNet, to extract local surface descriptors which are rotation-invariant whilst sufficiently distinctive and general. A Spatial Point Transformer is first introduced to embed the input local surface into an elaborate cylindrical representation (SO(2) rotation-equivariant), further enabling end-to-end optimization of the entire framework. A Neural Feature Extractor, composed of point-based and 3D cylindrical convolutional layers, is then presented to learn representative and general geometric patterns. An invariant layer is finally used to generate rotation-invariant feature descriptors. Extensive experiments on both indoor and outdoor datasets demonstrate that SpinNet outperforms existing state-of-the-art techniques by a large margin. More critically, it has the best generalization ability across unseen scenarios with different sensor modalities. Sheng Ao, Yulan Guo, Qingyong Hu, Bo Yang 0027, Andrew Markham, Zengping Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | SpinNet: Learning a General Surface Descriptor for 3D Point Cloud RegistrationabstractExtracting robust and general 3D local features is key to downstream tasks such as point cloud registration and reconstruction. Existing learning-based local descriptors are either sensitive to rotation transformations, or rely on classical handcrafted features which are neither general nor representative. In this paper, we introduce a new, yet conceptually simple, neural architecture, termed SpinNet, to extract local features which are rotationally invariant whilst sufficiently informative to enable accurate registration. A Spatial Point Transformer is first introduced to map the input local surface into a carefully designed cylindrical space, enabling end-to-end optimization with SO(2) equivariant representation. A Neural Feature Extractor which leverages the powerful point-based and 3D cylindrical convolutional neural layers is then utilized to derive a compact and representative descriptor for matching. Extensive experiments on both indoor and outdoor datasets demonstrate that SpinNet outperforms existing state-of-the-art techniques by a large margin. More critically, it has the best generalization ability across unseen scenarios with different sensor modalities. The code is available at https://github.com/QingyongHu/SpinNet. Sheng Ao, Qingyong Hu, Bo Yang 0027, Andrew Markham, Yulan Guo |
CVPR | 1 |
| 2020 | SurfaceNet: A Surface Focused Network for Pedestrian Detection and Segmentation in 3D Point CloudsabstractPedestrian detection is an important problem for autonomous driving. It is still chanllenging to detect and segment pedestrians from point clouds. In this paper, we propose a method named SurfaceNet to detect and segment pedestrians from point clouds. Specifically, we propose a novel representation, named surface map, to represent a point cloud as a 2D pseudo-image. For pedestrian detection, the proposed method comprises of four modules: 1) a grid feature encoder that can processes arbitrary number of points within each grid; 2) a surface feature convolutional module that employs a set of 2D convolutional layers to extract high level features; 3) a view transform module that transforms features from front view to bird's eye view; and 4) an anchor-free 3D object detection head that produces rotated 3D bounding box predictions. For semantic segmentation, the 2D pseudo-image is used for semantic segmentation and the segmentation results are re-projected to the original point cloud to achieve point cloud segmentation. Experimental results on the KITTI dataset show that our method achieves promising performance on pedestrian detection and segmentation in point clouds. Yongcong Zhang, Minglin Chen, Sheng Ao, Yulan Guo |
ICARCV | 3 |
| 2020 | SGHs for 3D local surface descriptionabstractThis study proposes a distinctive and robust spatial and geometric histograms (SGHs) feature descriptor for three‐dimensional (3D) local surface description. The authors also introduce a new local reference frame for the generation of their SGH descriptor. To fully describe a local surface, the SGH descriptor considers both spatial distribution and geometrical characteristics in its underlying support region. To encode neighbourhood information, the SGH descriptor is constructed using histogram statistics with spatial partition and interpolation strategies. The performance of the SGH descriptor was rigorously tested on six public datasets for applications of both 3D object recognition and registration. Compared to eight state‐of‐the‐art descriptors, experimental results show that SGH achieves the best performance on noise‐free data. It also produces the best results even under different nuisances. The promising descriptiveness and robustness of their SGH descriptor have been fully demonstrated. Sheng Ao, Yulan Guo, Shangtai Gu, Jindong Tian, Dong Li 0050 |
IET Comput. Vis. | 1 |
| 2020 | A repeatable and robust local reference frame for 3D surface matching
Sheng Ao, Yulan Guo, Jindong Tian, Dong Li 0050 |
Pattern Recognit. | 1 |