EDBT 2026 Demo / reviewers in the wild / expert
Shangshu Yu
dblp:256/9581
· DBLP profile ↗
21ranked-venue papers
8as first author
19since 2021 · last 2026
0000-0002-5000-0979ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 5 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NeuSpring: Neural Spring Fields for Reconstruction and Simulation of Deformable Objects from VideosabstractIn this paper, we aim to create physical digital twins of deformable objects under interaction. Existing methods focus more on the physical learning of current state modeling, but generalize worse to future prediction. This is because existing methods ignore the intrinsic physical properties of deformable objects, resulting in the limited physical learning in the current state modeling. To address this, we present NeuSpring, a neural spring field for the reconstruction and simulation of deformable objects from videos. Built upon spring-mass models for realistic physical simulation, our method consists of two major innovations: 1) a piecewise topology solution that efficiently models multi-region spring connection topologies using zero-order optimization, which considers the material heterogeneity of real-world objects. 2) a neural spring field that represents spring physical properties across different frames using a canonical coordinate-based neural network, which effectively leverages the spatial associativity of springs for physical learning. Experiments on real-world datasets demonstrate that our NeuSping achieves superior reconstruction and simulation performance for current state modeling and future prediction, with Chamfer distance improved by 20% and 25%, respectively. Qingshan Xu 0001, Jiao Liu 0006, Shangshu Yu, Yuan Zhou 0016, Junbao Zhou, Jiequan Cui, Yew-Soon Ong, Hanwang Zhang |
AAAI | 3 |
| 2026 | Performance Analysis of Collaborative Vision-based Localization in UAV-assisted Vehicular Networks
Xulun Huang, Xiaoshi Song, Zhengbin Jiao, Shangshu Yu, Liying Tian |
INFOCOM | 5 |
| 2025 | STGC-NeRF: Spatial-Temporal Geometric Consistency for LiDAR Neural Radiance Fields in Dynamic ScenesabstractWhile Neural Radiance Fields (NeRFs) have advanced the frontiers of novel view synthesis (NVS) using LiDAR data, they still struggle in dynamic scenes. Due to the low frequency and sparsity characteristics of LiDAR point clouds, it is challenging to spontaneously learn a dynamic and consistent scene representation from posed scans. In this paper, we propose STGC-NeRF, a novel LiDAR NeRF method that combines spatial-temporal geometry consistency to enhance the reconstruction of dynamic scenes. First, we propose a temporal geometry consistency regularization to enhance the regression of time-varying scene geometries from low-frequency LiDAR sequences. By estimating the pointwise correspondences between synthetic (or real) and real frames at different times, we convert them into various forms of temporal supervision. This alleviates the inconsistency caused by moving objects in dynamic scenes. Second, to improve the reconstruction of sparse LiDAR data, we propose spatial geometric consistency constraints. By computing multiple neighborhood feature descriptors incorporating geometric and contextual information, we capture structural geometry information from sparse LiDAR data. This helps encourage consistent direction, smoothness, and detail of the local surface. Extensive experiments on the KITTI-360 and nuScenes datasets demonstrate that STGC-NeRF outperforms state-of-the-art methods in both geometry and intensity accuracy for dynamic LiDAR scene reconstruction. Shangshu Yu, Xiaotian Sun 0005, Wen Li 0005, Qingshan Xu 0001, Zhimin Yuan, Rui She 0001, Cheng Wang 0003 |
AAAI | 1 |
| 2025 | LightLoc: Learning Outdoor LiDAR Localization at Light SpeedabstractScene coordinate regression achieves impressive results in outdoor LiDAR localization but requires days of training. Since training needs to be repeated for each new scene, long training times make these impractical for applications requiring time-sensitive system upgrades, such as autonomous driving, drones, robotics, etc. We identify large coverage areas and vast amounts of data in large-scale outdoor scenes as key challenges that limit fast training. In this paper, we propose LightLoc, the first method capable of efficiently learning localization in a new scene at light speed. Beyond freezing the scene-agnostic feature backbone and training only the scene-specific prediction heads, we introduce two novel techniques to address these challenges. First, we introduce sample classification guidance to assist regression learning, reducing ambiguity from similar samples and improving training efficiency. Second, we propose redundant sample downsampling to remove well-learned frames during training, reducing training time without compromising accuracy. In addition, the fast training and confidence estimation characteristics of sample classification enable its integration into SLAM, effectively eliminating error accumulation. Extensive experiments on large-scale outdoor datasets demonstrate that LightLoc achieves state-of-the-art performance with just 1 hour of training—50× faster than existing methods. Our Code is available at https://github.com/liw95/LightLoc. Wen Li 0005, Shangshu Yu, Dunqiang Liu, Chenglu Wen, Cheng Wang 0003 |
CVPR | 3 |
| 2025 | Multi-Modal Aerial-Ground Cross-View Place Recognition with Neural ODEsabstractPlace recognition (PR) aims at retrieving the query place from a database and plays a crucial role in various applications, including navigation, autonomous driving, and augmented reality. While previous multi-modal PR works have mainly focused on the same-view scenario in which ground-view descriptors are matched with a database of ground-view descriptors during inference, the multi-modal cross-view scenario, in which ground-view descriptors are matched with aerial-view descriptors in a database, remains underexplored. We propose AGPlace, a model that effectively integrates information from multi-modal ground sensors (cameras and LiDARs) to achieve accurate aerial-ground PR. AGPlace achieves effective aerial-ground cross-view PR by leveraging a manifold-based neural ordinary differential equation (ODE) framework with a multi-domain alignment loss. It outperforms existing state-of-the-art cross-view PR models on large-scale datasets. As most existing PR models are designed for ground-ground PR, we adapt these baselines into our cross-view pipeline. Experiments demonstrate that this direct adaptation performs worse than our overall model architecture AGPlace. AGPlace represents a significant advancement in multi-modal aerial-ground PR, with promising implications for real-world applications. Rui She 0001, Qiyu Kang, Disheng Li, Tianyu Geng, Shangshu Yu, Wee-Peng Tay |
CVPR | 7 |
| 2025 | UAVScenes: A Multi-Modal Dataset for UAVs
Shangshu Yu, Shenghai Yuan 0001, Rui She 0001, Quanjiang Guo, Jinxuan Zheng, Ong Kang Howe, Leonrich Chandra, Shrivarshann Srijeyan, Aditya Sivadas, Toshan Aggarwal, Heyuan Liu, Chujie Chen, Junyu Jiang, Lihua Xie 0001, Wee-Peng Tay |
ICCV | 4 |
| 2025 | RALoc: Enhancing Outdoor LiDAR Localization via Rotation Awareness
Yuyang Yang, We Li, Sheng Ao, Shangshu Yu |
ICCV | 5 |
| 2025 | GTR-Loc: Geospatial Text Regularization Assisted Outdoor LiDAR LocalizationabstractPrevailing scene coordinate regression methods for LiDAR localization suffer from localization ambiguities, as distinct locations can exhibit similar geometric signatures — a challenge that current geometry-based regression approaches have yet to solve. Recent vision–language models show that textual descriptions can enrich scene understanding, supplying potential localization cues missing from point cloud geometries. In this paper, we propose GTR-Loc, a novel text-assisted LiDAR localization framework that effectively generates and integrates geospatial text regularization to enhance localization accuracy. We propose two novel designs: a Geospatial Text Generator that produces discrete pose-aware text descriptions, and a LiDAR-Anchored Text Embedding Refinement module that dynamically constructs view-specific embeddings conditioned on current LiDAR features. The geospatial text embeddings act as regularization to effectively reduce localization ambiguities. Furthermore, we introduce a Modality Reduction Distillation strategy to transfer textual knowledge. It enables high-performance LiDAR-only localization during inference, without requiring runtime text generation. Extensive experiments on challenging large-scale outdoor datasets, including QEOxford, Oxford Radar RobotCar, and NCLT, demonstrate the effectiveness of GTR-Loc. Our method significantly outperforms state-of-the-art approaches, notably achieving a 9.64%/8.04% improvement in position/orientation accuracy on QEOxford. Our code is available at https://github.com/PSYZ1234/GTR-Loc. Shangshu Yu, Wen Li 0005, Xiaotian Sun 0005, Zhimin Yuan, Rui She 0001, Cheng Wang 0003 |
NeurIPS | 1 |
| 2025 | VFM-Depth: Leveraging Vision Foundation Model for Self-Supervised Monocular Depth EstimationabstractSelf-supervised monocular depth estimation has exploited semantics to reduce depth ambiguities in texture-less regions and object boundaries. However, existing methods struggle to obtain universal semantics across scenes for effective depth estimation. This paper proposes VFM-Depth, a novel self-supervised teacher-student framework, that effectively leverages the vision foundation model as semantic regularization to significantly improve the accuracy of monocular depth estimation. Firstly, we propose a novel Geometric-Semantic Aggregation Encoding, integrating universal semantic constraints from the foundation model to reduce ambiguities in the teacher model. Specifically, semantic features from the foundation model and geometric features from the depth model are first encoded and then fused through cross-modal aggregation. Secondly, we introduce a novel Multi-Alignment for Depth Distillation to distill semantic constraints from the teacher, further leveraging knowledge from the foundation model. We obtain a lightweight yet effective student model through an innovative approach that combines distance category alignment with complementary feature and depth imitation. Extensive experiments on KITTI, Cityscapes, and Make3D datasets demonstrate that VFM-Depth (both teacher and student) outperforms state-of-the-art self-supervised methods by a large margin. Shangshu Yu, Meiqing Wu, Siew-Kei Lam |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | EDS-Depth: Enhancing Self-Supervised Monocular Depth Estimation in Dynamic ScenesabstractSelf-supervised monocular depth estimation usually assumes that training samples contain only static objects, which leads to poor performance in real-world environments. The presence of dynamic objects incurs camera motion estimation errors, motion blur, and occlusions, which induce significant challenges for network training. To address these issues, we introduce EDS-Depth, a self-supervised learning framework, that improves monocular depth estimation in dynamic scenes. Firstly, we propose a novel TCE (Temporal Continuity Enhancement) strategy to reduce camera motion estimation errors and motion blur caused by dynamic objects. Video frames are interpolated to generate more continuous frames in order to smooth dynamic changes and enrich motion details. Secondly, we design a novel IPDM (Iterative Pseudo Depth Masking) module to address inaccurate object motion and occlusions in dynamic scenes. The module integrates multiple optical flows from different frames for triangulation, generating optimal depth as pseudo-supervision labels in dynamic regions. Extensive experiments on Cityscapes and KITTI datasets demonstrate the effectiveness of EDS-Depth, which surpasses state-of-the-art self-supervised monocular depth estimation methods, particularly in dynamic scenes. Shangshu Yu, Meiqing Wu, Siew-Kei Lam, Changshuo Wang 0001, Ruiping Wang 0005 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | DiffLoc: Diffusion Model for Outdoor LiDAR LocalizationabstractAbsolute pose regression (APR) estimates global pose in an end-to-end manner, achieving impressive results in learn-based LiDAR localization. However, compared to the top-performing methods reliant on 3D-3D correspondence matching, APR's accuracy still has room for improvement. We recognize APR's lack of robust features learning and iterative denoising process leads to suboptimal results. In this paper, we propose DiffLoc, a novel framework that formulates LiDAR localization as a conditional generation of poses. First, we propose to utilize the foundation model and static-object-aware pool to learn robust features. Second, we incorporate the iterative denoising process into APR via a diffusion model conditioned on the learned geometrically robust features. In addition, due to the unique nature of diffusion models, we propose to adapt our models to two additional applications: (1) using multiple inferences to evaluate pose uncertainty, and (2) seamlessly introducing geometric constraints on denoising steps to improve prediction accuracy. Extensive experiments conducted on the Oxford Radar RobotCar and NCLT datasets demonstrate that DiffLoc outperforms better than the state-of-the-art methods. Especially on the NCLT dataset, we achieve 35% and 34.7% improvement on position and orientation accuracy, respectively. Our code is released at https://github.com/liw95/DiffLoc. Wen Li 0005, Yuyang Yang, Shangshu Yu, Guosheng Hu, Chenglu Wen, Ming Cheng 0002, Cheng Wang 0003 |
CVPR | 3 |
| 2024 | GPSFormer: A Global Perception and Local Structure Fitting-Based Transformer for Point Cloud Understanding
Changshuo Wang 0001, Meiqing Wu, Siew-Kei Lam, Xin Ning 0001, Shangshu Yu, Ruiping Wang 0005, Weijun Li 0002, Thambipillai Srikanthan |
ECCV (8) | 5 |
| 2024 | NIDALoc: Neurobiologically Inspired Deep LiDAR LocalizationabstractAbsolute pose regression has shown great potential in LiDAR localization, which learns to regress 6-DoF LiDAR poses through deep networks. However, recent regression methods suffer from scene ambiguities in challenging scenarios, leading to inaccurate and unstable localization. Inspired by neurobiological localization mechanisms, i.e., the firing mechanism of place cells, head-direction cells, and grid cells in mammalian brains, we propose a novel LiDAR localization framework called NIDALoc to achieve more robust and accurate results. First, we propose a Hebbian memory module, motivated by place cells, to preserve historical information, which helps refine local view features to reduce scene ambiguities. Specifically, the memory module stores scene information and then recalls it when revisiting an old place. Second, we propose a novel pose constrained framework, consisting of an orientation classification task and a grid center regression task, to regularize orientation and position estimation, respectively. The framework based on head-direction cells and grid cells constrains the absolute pose regression to reduce wrong predictions. Extensive experiments on two outdoor datasets demonstrate the effectiveness of NIDALoc, which outperforms state-of-the-art localization methods, especially in large-scale challenging scenes. The source code is available on the project website at https://github.com/PSYZ1234/NIDALoc. Shangshu Yu, Xiaotian Sun 0005, Wen Li 0005, Chenglu Wen, Yunuo Yang, Bailu Si, Guosheng Hu, Cheng Wang 0003 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | SGLoc: Scene Geometry Encoding for Outdoor LiDAR LocalizationabstractLiDAR-based absolute pose regression estimates the global pose through a deep network in an end-to-end manner, achieving impressive results in learning-based localization. However, the accuracy of existing methods still has room to improve due to the difficulty of effectively encoding the scene geometry and the unsatisfactory quality of the data. In this work, we propose a novel LiDAR localization frame-work, SGLoc, which decouples the pose estimation to point cloud correspondence regression and pose estimation via this correspondence. This decoupling effectively encodes the scene geometry because the decoupled correspondence regression step greatly preserves the scene geometry, leading to significant performance improvement. Apart from this decoupling, we also design a tri-scale spatial feature aggregation module and inter-geometric consistency constraint loss to effectively capture scene geometry. Moreover, we empirically find that the ground truth might be noisy due to GPS/INS measuring errors, greatly reducing the pose estimation performance. Thus, we propose a pose quality evaluation and enhancement method to measure and correct the ground truth pose. Extensive experiments on the Oxford Radar RobotCar and NCLT datasets demonstrate the effectiveness of SGLoc, which outperforms state-of-the-art regression-based localization methods by 68.5% and 67.6% on position accuracy, respectively. Wen Li 0005, Shangshu Yu, Cheng Wang 0003, Guosheng Hu, Chenglu Wen |
CVPR | 2 |
| 2023 | Prototype-Guided Multitask Adversarial Network for Cross-Domain LiDAR Point Clouds Semantic SegmentationabstractUnsupervised domain adaptation (UDA) segmentation aims to leverage labeled source data to make accurate predictions on unlabeled target data. The key is to make the segmentation network learn domain-invariant representations. In this work, we propose a prototype-guided multitask adversarial network (PMAN) to achieve this. First, we propose an intensity-aware segmentation network (IAS-Net) that leverages the private intensity information of target data to substantially facilitate feature learning of the target domain. Second, the category-level cross-domain feature alignment strategy is introduced to flee the side effects of global feature alignment. It employs the prototype (class centroid) and includes two essential operations: 1) build an auxiliary nonparametric classifier to evaluate the semantic alignment degree of each point based on the prediction consistency between the main and auxiliary classifiers and 2) introduce two class-conditional point-to-prototype learning objectives for better alignment. One is to explicitly perform category-level feature alignment in a progressive manner, and the other aims to shape the source feature representation to be discriminative. Extensive experiments reveal that our PMAN outperforms state-of-the-art results on two benchmark datasets. Zhimin Yuan, Ming Cheng 0002, Wankang Zeng, Yanfei Su, Weiquan Liu, Shangshu Yu, Cheng Wang 0003 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | STCLoc: Deep LiDAR Localization With Spatio-Temporal ConstraintsabstractLiDAR localization is of great importance to autonomous vehicles and robotics. Absolute pose regression, directly estimating the mapping from a scene to a 6-DoF pose, has achieved impressive results in learning-based localization. Different from traditional map-based methods, it does not need a pre-built 3D map during inference. However, current regression networks typically suffer from scene ambiguities, especially in challenging traffic environments, leading to large wrong predictions (e.g., outliers) and limited applications. To address this problem, a novel LiDAR localization framework with spatio-temporal constraints is proposed, termed STCLoc, to reduce scene ambiguities and achieve more accurate localization. First, we propose to regularize regression in the spatial dimension with a novel classification task to reduce outliers. Specifically, the classification task categorizes the point cloud in terms of position and orientation and then couples it with the regression task to conduct multi-task learning. Second, to learn discriminative features to reduce scene ambiguities, we propose using attention-based feature aggregation to capture the correlation in LiDAR sequences. We conduct extensive experiments on two benchmark datasets, where the localization takes 97ms on each dataset. Results show that our model outperforms state-of-the-art methods by 43.33%/36.76% (position/orientation) on the Oxford Radar RobotCar dataset, verifying the effectiveness of our method. The source code is available on the project website athttps://github.com/PSYZ1234/STCLoc. Shangshu Yu, Cheng Wang 0003, Yitai Lin, Chenglu Wen, Ming Cheng 0002, Guosheng Hu |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | Category-Level Adversaries for Outdoor LiDAR Point Clouds Cross-Domain Semantic SegmentationabstractUnsupervised domain adaptation (UDA) is a low-cost way to deal with the lack of annotations in a new domain. For outdoor point clouds in urban transportation scenes, the mismatch of sampling patterns and the transferability difference between classes make cross-domain segmentation extremely difficult. To overcome these challenges, we propose a category-level adversarial framework. Firstly, we propose a multi-scale domain conditioned block that facilitates to extract the critical low-level domain-dependent knowledge and reduce the domain gap caused by distinct LiDAR sampling patterns. Secondly, we make full use of multiple representation forms (i.e., point-based sets and voxel-based cells) and utilize the prediction consistency between the two forms to measure how well each point is semantically aligned. The model then focuses on the poorly-aligned points without affecting the well-aligned points. Experimental results on three autonomous driving point cloud datasets show that the proposed method outperforms existing methods by a large margin, especially on the low-beam to high-beam cross-domain segmentation task. Zhimin Yuan, Chenglu Wen, Ming Cheng 0002, Yanfei Su, Weiquan Liu, Shangshu Yu, Cheng Wang 0003 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2022 | LiDAR-based localization using universal encoding and memory-aware regression
Shangshu Yu, Cheng Wang 0003, Chenglu Wen, Ming Cheng 0002, Minghao Liu 0007, Zhihong Zhang 0001, Xin Li 0003 |
Pattern Recognit. | 1 |
| 2022 | Corrigendum to "LiDAR-based localization using universal encoding and memory-aware regression" Pattern Recognition Volume 128 (2022) 108685
Shangshu Yu, Cheng Wang 0003, Chenglu Wen, Ming Cheng 0002, Minghao Liu 0007, Zhihong Zhang 0001, Xin Li 0003 |
Pattern Recognit. | 1 |
| 2020 | Learning to Match Ground Camera Image and UAV 3-D Model-Rendered Image Based on Siamese Network With Attention MechanismabstractDifferent domain image sensors or imaging mechanisms provide cross-domain images when sensing the same scene. There is a domain shift between cross-domain images so that the image gap between different domains is the major challenge for measuring the similarity of the feature descriptors extracted from different domain images. Specifically, matching ground camera images and unmanned aerial vehicle (UAV) 3-D model-rendered images, which are two kinds of extremely challenging cross-domain images, is a way to establish indirectly the spatial relationship between 2-D and 3-D spaces. This provides a solution for the virtual-real registration of augmented reality (AR) in outdoor environments. However, during matching, handcrafted descriptors and existing learning-based feature descriptors limit the rendered images. In this letter, first, to learn robust and invariant 128-D local feature descriptors for ground camera and rendered images, we present a novel network structure, SiamAM-Net, which embeds the autoencoders with an attention mechanism into the Siamese network. Then, to narrow the gap between the cross-domain images during the optimizing of SiamAM-Net, we design an adaptive margin for the loss function. Finally, we match the ground camera-rendered images by using the learned local feature descriptors and explore the outdoor AR virtual-real registration. Experiments show that the local feature descriptors, learned by SiamAM-Net, are robust and achieve state-of-the-art retrieval performance on the cross-domain image data set of ground camera and rendered images. In addition, several outdoor AR applications also demonstrate the usefulness of the proposed outdoor AR virtual-real registration. Weiquan Liu, Cheng Wang 0003, Xuesheng Bian, Shangshu Yu, Xiuhong Lin, Shang-Hong Lai, Dongdong Weng, Jonathan Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2020 | Urban 3D modeling with mobile laser scanning: a reviewabstractMobile laser scanning (MLS) systems mainly comprise laser scanners and mobile mapping platforms. Typical MLS systems are able to acquire three-dimensional point clouds with 1-10 centimeter point spacing at a normal driving or walking speed in the street or indoor environments. The MLS' advantages of efficiency and stability make it a quite practical tool for three-dimensional urban modeling. This paper reviews the latest advances in 3D modeling of the LiDAR-based mobile mapping system (MMS) point cloud, including LiDAR Simultaneous Localization and Mapping (SLAM), point cloud registration, feature extraction, object extraction, semantic segmentation, and deep learning processing. Then typical urban modeling applications based on MMS are also discussed. Cheng Wang 0003, Chenglu Wen, Yudi Dai, Shangshu Yu, Minghao Liu 0007 |
Virtual Real. Intell. Hardw. | 4 |