EDBT 2026 Demo / reviewers in the wild / expert
Zhen Dong 0005
dblp:60/1749-5
· DBLP profile ↗
43ranked-venue papers
0as first author
38since 2021 · last 2026
0000-0002-0152-3300ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 20 · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 17 since 2021Artificial intelligence and machine learning · 16 · 15 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GAGS: Granularity-Aware Feature Distillation for Language Gaussian Splattingabstract3D open-vocabulary scene understanding, which accurately perceives complex semantic properties of objects in space, has gained significant attention in recent years. In this paper, we propose GAGS, a framework that distills 2D CLIP features into 3D Gaussian splatting, enabling open-vocabulary queries for renderings on arbitrary viewpoints. The main challenge of distilling 2D features for 3D fields lies in the multiview inconsistency of extracted 2D features, which provides unstable supervision for the 3D feature field. GAGS addresses this challenge with two novel strategies. First, GAGS associates the prompt point density of SAM with the camera distances to scene objects, which significantly improves the multiview consistency of segmentation results. Second, GAGS further decodes a granularity factor to guide the distillation process and this granularity factor can be learned in a unsupervised manner to only select the multiview consistent 2D features in the distillation process. Experimental results on two datasets show that GAGS improves visual grounding accuracy by an average of 10.9% and semantic segmentation accuracy by an average of 7.0%, with an inference speed 2× faster than baseline methods. Yuning Peng, Haiping Wang 0004, Yuan Liu 0025, Chenglu Wen, Zhen Dong 0005, Bisheng Yang |
AAAI | 5 |
| 2026 | AirDC: Adaptive Iterative Depth Refinement Framework for Full-Range Metric Depth CompletionabstractAccurate metric depth completion across wide depth ranges is critical for autonomous systems. However, existing methods often struggle to efficiently capture depth features at both close and long ranges, primarily due to the inadequate modeling of fine-grained depth cues specific to different depth ranges. To address these limitations, we propose AirDC, an adaptive iterative depth refinement framework for full-range metric depth completion. The core contributions of our model lie in the design of two key modules. Specifically, we first construct an adaptive fine-grained stereo-LiDAR feature fusion module to fundamentally strengthen the model's capacity to preserve original full-range depth information. Built upon metric-aligned depth volumes (i.e., a 3D representation composed of cubic voxels uniformly partitioned in real-world metric space), this module employs an adaptive sub-voxel depth attention mechanism to enhance sensitivity to subtle depth variations across the full range, thereby both avoiding long-range accuracy degradation introduced by conventional disparity conversion and alleviating the coarse near-range granularity inherent in metric depth representations. Second, we introduce an iterative hypothesis-guided depth refinement module to improve prediction accuracy while maintaining memory efficiency. By integrating multi-scale multi-modal guidance information from depth hypotheses, this module enables explicit and progressive refinement of the initial depth estimation with a small parameter overhead. Experiments on multiple mainstream real-world and synthetic benchmarks demonstrate that AirDC achieves state-of-the-art performance, providing an effective solution for full-range metric depth completion. The code and data are available at https://github.com/yunqidu/AirDC. Yunqi Du, Hongjuan Zhang, Zhen Dong 0005, Luliang Tang |
IEEE Trans. Image Process. | 5 |
| 2026 | LifelongPR: Lifelong Point Cloud Place Recognition Based on Sample Replay and Prompt LearningabstractPoint cloud place recognition (PCPR) determines the geo-location within a prebuilt map and plays a crucial role in photogrammetry and robotics applications such as autonomous driving, intelligent transportation, and augmented reality. In real-world large-scale deployments of a geographic positioning system, PCPR models must continuously acquire, update, and accumulate knowledge to adapt to diverse and dynamic environments, i.e., the ability known as continual learning (CL). However, existing PCPR models often suffer from catastrophic forgetting, leading to significant performance degradation in previously learned scenes when adapting to new environments or sensor types. This results in poor model scalability, increased maintenance costs, and system deployment difficulties, undermining the practicality of PCPR. To address these issues, we propose LifelongPR, a novel continual learning framework for PCPR, which effectively extracts and fuses knowledge from sequential point cloud data. First, to alleviate the knowledge loss, we propose a replay sample selection method that dynamically allocates sample sizes according to each dataset’s information quantity and selects spatially diverse samples for maximal representativeness. Second, to handle domain shifts, we design a prompt learning-based CL framework with a lightweight continuous prompt module and a two-stage training strategy, enabling domain-specific feature adaptation while minimizing forgetting. Comprehensive experiments on large-scale public and self-collected datasets are conducted to validate the effectiveness of the proposed method. Compared with the state-of-the-art (SOTA) method, our method achieves 6.50% improvement in$mIR\text{@}1$, 7.96% improvement in$mR\text{@}1$, and an 8.95% reduction in$F$. The code and pre-trained models are publicly available athttps://zouxianghong.github.io/LifelongPR Xianghong Zou, Jianping Li 0004, Zhe Chen 0028, Zhen Dong 0005, Qiegen Liu, Bisheng Yang |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | Align3R: Aligned Monocular Depth Estimation for Dynamic VideosabstractRecent developments in monocular depth estimation methods enable high-quality depth estimation of single-view images but fail to estimate consistent video depth across different frames. Very recent works address this problem by applying a video diffusion model to generate video depth conditioned on the input video, which is training-expensive and can only produce scale-invariant depth values without camera poses. In this paper, we propose a novel video-depth estimation method called Align3R to estimate temporally consistent depth maps for a dynamic video. Our key idea is to utilize the recent DUSt3R model to align estimated monocular depth maps of different timesteps. First, we fine-tune the DUSt3R model with additional estimated monocular depth as inputs for the dynamic scenes. Then, we apply optimization to reconstruct both depth maps and camera poses. Extensive experiments demonstrate that Align3R estimates consistent video depth and camera poses for a monocular video with superior performance than baseline methods. Jiahao Lu 0001, Zhiyang Dou, Cheng Lin 0001, Zhiming Cui 0001, Zhen Dong 0005, Sai-Kit Yeung, Wenping Wang 0001, Yuan Liu 0025 |
CVPR | 7 |
| 2025 | Learning to Detect Objects from Multi-Agent LiDAR Scans without Manual LabelsabstractUnsupervised 3D object detection serves as an important solution for offline 3D object annotation. However, due to the data sparsity and limited views, the clustering-based label fitting in unsupervised object detection often generates low-quality pseudo-labels. Multi-agent collaborative dataset, which involves the sharing of complementary observations among agents, holds the potential to break through this bottleneck. In this paper, we introduce a novel unsupervised method that learns to Detect Objects from Multi-Agent LiDAR scans, termed DOtA, without using labels from external. DOtA first uses the internally shared ego-pose and ego-shape of collaborative agents to initialize the detector, leveraging the generalization performance of neural networks to infer preliminary labels. Subsequently, DOtA uses the complementary observations between agents to perform multi-scale encoding on preliminary labels, then decodes high-quality and low-quality labels. These labels are further used as prompts to guide a correct feature learning process, thereby enhancing the performance of the unsupervised object detection task. Extensive experiments on the V2V4Real and OPV2V datasets show that our DOtA outperforms state-of-the-art unsupervised 3D object detection methods. Additionally, we also validate the effectiveness of the DOtA labels under various collaborative perception frameworks. The code is available at https://github.com/xmuqimingxia/DOtA. Qiming Xia, Wenkai Lin, Haoen Xiang, Xun Huang 0003, Siheng Chen, Zhen Dong 0005, Cheng Wang 0003, Chenglu Wen |
CVPR | 6 |
| 2025 | DeepLA-Net: Very Deep Local Aggregation Networks for Point Cloud AnalysisabstractDue to the irregular and disordered data structure in 3D point clouds, prior works have focused on designing more sophisticated local representation methods to capture these complex local patterns. However, the recognition performance has saturated over the past few years, indicating that increasingly complex and redundant designs no longer make improvements to local learning. This phenomenon prompts us to diverge from the trend in 3D vision and instead pursue an alternative and successful solution: deeper neural networks. In this paper, we propose DeepLA-Net, a series of very deep networks for point cloud analysis. The key insight of our approach is to exploit a small but mighty local learning block, which uses 10× fewer FLOPs, enabling the construction of very deep networks. Furthermore, we design a training supervision strategy to ensure smooth gradient backpropagation and optimization in very deep networks. We construct the DeepLA-Net family with a depth of up to 120 blocks — at least 5× deeper than recent methods — trained on a single RTX 3090. An ensemble of the DeepLA-Net achieves state-of-the-art performance on classification and segmentation tasks of S3DIS Area5 (+2.2% mIoU), ScanNet test set (+1.6% mIoU), ScanObjectNN (+2.1% OA), and ShapeNet-Part (+0.9% cls.mIoU). The code are released at https://github.com/zeng-ziyin/DeepLA-Net. Ziyin Zeng, Mingyue Dong, Jian Zhou 0011, Huan Qiu, Zhen Dong 0005 |
CVPR | 5 |
| 2025 | Vistadream: Sampling Multiview Consistent Images for Single-View Scene ReconstructionabstractIn this paper, we propose VistaDream a novel framework to reconstruct a 3D scene from a single-view image. Recent diffusion models enable generating high-quality novel-view images from a single-view input image. Most existing methods only concentrate on building the consistency between the input image and the generated images while losing the consistency between the generated images. VistaDream addresses this problem by a two-stage pipeline. In the first stage, VistaDream begins with building a global coarse 3D scaffold by zooming out a little step with inpainted boundaries and an estimated depth map. Then, on this global scaffold, we use iterative diffusion-based RGB-D inpainting to generate novel-view images to inpaint the holes of the scaffold. In the second stage, we further enhance the consistency between the generated novel-view images by a novel training-free Multiview Consistency Sampling (MCS) that introduces multi-view consistency constraints in the reverse sampling process of diffusion models. Experimental results demonstrate that without training or fine-tuning existing diffusion models, VistaDream achieves consistent and high-quality novel view synthesis using just single-view images and outperforms baseline methods by a large margin. The code, videos, and interactive demos are available at https://vistadream-project-page.github.io/. Haiping Wang 0004, Yuan Liu 0025, Ziwei Liu 0002, Wenping Wang 0001, Zhen Dong 0005, Bisheng Yang |
ICCV | 5 |
| 2025 | CityAnchor: City-scale 3D Visual Grounding with Multi-modality LLMsabstractIn this paper, we present a 3D visual grounding method called CityAnchor for localizing an urban object in a city-scale point cloud. Recent developments in multiview reconstruction enable us to reconstruct city-scale point clouds but how to conduct visual grounding on such a large-scale urban point cloud remains an open problem. Previous 3D visual grounding system mainly concentrates on localizing an object in an image or a small-scale point cloud, which is not accurate and efficient enough to scale up to a city-scale point cloud. We address this problem with a multi-modality LLM which consists of two stages, a coarse localization and a fine-grained matching. Given the text descriptions, the coarse localization stage locates possible regions on a projected 2D map of the point cloud while the fine-grained matching stage accurately determines the most matched object in these possible regions. We conduct experiments on the CityRefer dataset and a new synthetic dataset annotated by us, both of which demonstrate our method can produce accurate 3D visual grounding on a city-scale 3D point cloud. Haiping Wang 0004, Yuan Liu 0025, Zhiyang Dou, Yuexin Ma, Sibei Yang, Wenping Wang 0001, Zhen Dong 0005, Bisheng Yang |
ICLR | 10 |
| 2025 | Layered Denoising and Classification of Photon Point Cloud Data From ICESat-2 in Forest AreaabstractIce, Cloud, and land Elevation Satellite (ICESat-2) carries the Advanced Topographic Laser Altimeter System (ATLAS), which enhancing along-track sampling density but introduces substantial noise in photon point cloud data. Therefore, this study establishes a denoising and classification feature parameter system grounded in the three-dimensional spatial distribution characteristics of photon point clouds. Modeling is conducted in two layers: one layer for upper noise photons and canopy signal photons, and another layer for lower noise photons and ground signal photons. Machine learning and neural network algorithms are utilized to denoise and classify the original photon point clouds from ICESat-2, aiming to obtain a transferable and universally applicable supervised classification model for denoising photon point clouds. Recall, Precision, and the harmonic mean of Recall and Precision (F1 Score) are used as evaluation metrics to verify the accuracy of local, transfer, and global models. The results indicate that under various forest types and external conditions, the proposed photon point cloud Layered Denoising and Classification Model (LDCM) outperforms the Differential Regressive and Gaussian Adaptive Nearest Neighbor (DRAGANN, ICESat-2 ATL08 production algorithm), Ordering Points to Identify the Clustering Structure (OPTICS), and Adaptive Elevation Difference Thresholding (AEDTA) algorithms in terms of accuracy. Compared to the DRAGANN algorithm, the maximum accuracy improvement is 60%, with an average improvement of approximately 20%; compared to the OPTICS algorithm, the maximum accuracy improvement is 36%, with an average improvement of about 28%; compared to the AEDTA algorithm, the maximum accuracy improvement is 27%, with an average improvement of about 14%. The F1 Score for the validation set of the machine learning and neural network algorithms is above 0.94, with the Categorical Boosting (CatBoost) algorithm achieving the best performance. Both the transfer model and the global model have F1 Scores above 0.90. Therefore, the proposed photon point cloud LDCM not only demonstrates excellent classification accuracy but also exhibits good transferability and general applicability. Junfan Bao, Ningning Zhu, Zhen Dong 0005, Sheng Nie, Wenxia Dai, Ruixiong Kou, Bisheng Yang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | WHU-Synthetic: A Synthetic Perception Dataset for 3-D Multitask Model ResearchabstractEnd-to-end models capable of handling multiple subtasks in parallel have become a new trend, thereby presenting significant challenges and opportunities for the integration of multiple tasks within the domain of 3-D vision. The limitations of 3-D data acquisition conditions have not only restricted the exploration of many innovative research problems but have also caused existing 3-D datasets to predominantly focus on single tasks. This has resulted in a lack of systematic approaches and theoretical frameworks for 3-D multitask learning, with most efforts merely serving as auxiliary support to the primary task. In this article, we introduce WHU-Synthetic, a large-scale 3-D synthetic perception dataset designed for multitask learning, from the initial data augmentation (upsampling and depth completion), through scene understanding (segmentation), to macrolevel tasks (place recognition and 3-D reconstruction). Collected in the same environmental domain, we ensure inherent alignment across subtasks to construct multitask models without separate training methods. In addition, we implement several novel settings, making it possible to realize certain ideas that are difficult to achieve in real-world scenarios. This supports more adaptive and robust multitask perception tasks, such as sampling on city-level models, providing point clouds with different densities, and simulating temporal changes. Using our dataset, we conduct several experiments to investigate mutual benefits between subtasks, revealing new observations, challenges, and opportunities for future research. The dataset is accessible at:https://github.com/WHU-USI3DV/WHU-Synthetic. Chen Long, Conglang Zhang, Boheng Li, Haiping Wang 0004, Zhe Chen 0028, Zhen Dong 0005 |
IEEE Trans. Geosci. Remote. Sens. | 9 |
| 2025 | RoCalib: Large-Scale Autonomous Geo-Calibration for Roadside Lidar With High-Definition MapabstractTo address the challenges of low overlap and viewing direction differences in large-scale roadside LiDAR calibration, this paper proposes RoCalib, a novel automatic roadside LiDAR calibration method based on the high-definition (HD) map. This method enables geographic registration of the LiDAR without any specific target and improves the efficiency and safety of the large-scale roadside facility calibration and maintenance. First, a novel virtual reprojection model is designed to construct a virtual mapping from the HD map to LiDAR, reducing representation differences. Based on this, a universal spatial context descriptor is introduced, applicable to various LiDAR systems, facilitating rapid retrieval of LiDAR positions within the HD map. Finally, based on the multi-feature optimization method considering the road structure, the fine registration and parameter calibration of the roadside LiDAR and the HD map are completed. The proposed framework is validated on simulated, public, and self-collected datasets, demonstrating that this method can automatically and accurately achieve multi-LiDAR geographic calibration, yielding superior performance. Cong Duan, Jian Zhou 0011, Zhen Dong 0005, Youchen Tang, Jinsheng Xiao |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | INF-PCA: Implicit Neural Field-Based Interactive Point Cloud Semantic AnnotationabstractPoint cloud semantic segmentation helps Intelligent Transportation Systems understand traffic scenes by assigning semantic label to each point in the point cloud, and it relies on large amounts of annotated training data. Nevertheless, manually annotating large-scale datasets of complex traffic scenes is quite time-consuming and tedious. This paper proposes INF-PCA, an interactive point cloud semantic annotation method based on implicit neural field, which allows users to achieve high-quality, large-scene and fast-response semantic annotation with only a few dozen mouse clicks. Firstly, the appearance, geometry and semantics of the point clouds are jointly represented by an implicit neural field, which maps a 3D spatial coordinate to its corresponding attributes. Secondly, an uncertainty-based semantic entropy loss and a supervoxel-based local consistency loss are designed to force the network to produce deterministic predictions with local consistency, thus generating smoother and more accurate boundaries. Furthermore, an active learning-based strategy for click-free annotation is proposed and analyzed to further reduce annotation pressure. Comprehensive experiments on multiple datasets including the road scene dataset Toronto3D revealed that INF-PCA can achieve more accurate annotations with faster response speed and only half of the clicks employed by the state-of-the-art methods, and that INF-PCA can be directly applied to intelligent transportation applications such as interactive segmentation of road scenes, inventory of transportation infrastructure assets, and production of high-definition map. Chen Long, Wang Wang, Zhen Dong 0005, Bisheng Yang |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | Fade3D: Fast and Deployable 3D Object Detection for Autonomous Drivingabstract3D object detection is an essential scene perception capability for autonomous vehicles. In intelligent transportation systems, autonomous vehicles require minimal inference latency to sense their surroundings in real-time. However, advanced 3D detection methods often suffer from high inference latency. This limits the real-time deployment of 3D detection models in the real world. To address this problem, this paper proposes a fast and deployable 3D object detection method from the LiDAR point cloud for autonomous driving, namedFade3D. Firstly, we propose a Lightweight Input Encoder (LIE) to extract the most critical features from point clouds. Then, we develop a Spatial Feature Enhancement BEV backbone (SFENet) that efficiently encodes geometry features into compact representations. Additionally, we design an IoU-aware Loss Re-weighting (ILR) that enhances performance by shifting more attention to hard samples. Leveraging LIE and SFENet, our approach is independent of point cloud density and number, achieving significant speed advantages in processing large-scale point clouds and being deployment-friendly. Extensive experiments on KITTI and Waymo Open Dataset (WOD) datasets comparing various baseline detectors demonstrate its universality and superiority. Specifically, our method demonstrates impressive real-time inference capabilities, achieving 51.5 Hz on an RTX3090 GPU and 12.4 Hz on a Jetson Orin embedded development board. Code will be available at https://github.com/wayyeah/Fade3D Qiming Xia, Zhen Dong 0005, Ruofei Zhong, Cheng Wang 0003, Chenglu Wen |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | WIN: Variable-View Implicit LiDAR Upsampling NetworkabstractLiDAR upsampling aims to increase the resolution of sparse point sets obtained from low-cost sensors, providing better performance for various downstream tasks (e.g., Autonomous Driving, High Definitation Map). Existing methods transfer LiDAR point cloud into range view, and focus on designing complex encoders or interpolation strategies to improve the resolution of LiDAR range images. However, our analysis shows that using the range view inevitably results in the loss of geometric information. We propose a Variable-View Implicit LiDAR Upsampling network, named WIN to solve this problem. It decouples range views into two novel virtual view representations, Horizontal Range View (HRV) and Vertical Range View (VRV). The key idea behind this is that introducing more perspectives can make up for the geometric information lost in a single perspective. We also prove theoretically that the proposed virtual view representation has a smaller error range compared to the range view representation. In addition, we design two novel strategies (i.e., contrast selection module and selection loss) to fuse the upsampling results of these two virtual representations and stabilize the whole training process. As a result, compared with the current state-of-the art (SOTA) method ILN, WIN introduces only 0.4M additional parameters, yet achieves a +4.53% increase in the MAE and a +7.01% increase in the IoU on the CARLA dataset. Furthermore, our method also outperforms all existing methods in downstream tasks (i.e., Depth Completion and Localization). The code and pre-trained models are available athttps://github.com/WHU-USI3DV/WIN Conglang Zhang, Chen Long, Zhen Dong 0005, Bisheng Yang |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | Explicitly Guided Information Interaction Network for Cross-Modal Point Cloud Completion
Chen Long, Yuan Liu 0025, Zhen Dong 0005, Bisheng Yang |
ECCV (12) | 6 |
| 2024 | FreeReg: Image-to-Point Cloud Registration Leveraging Pretrained Diffusion Models and Monocular Depth EstimatorsabstractMatching cross-modality features between images and point clouds is a fundamental problem for image-to-point cloud registration. However, due to the modality difference between images and points, it is difficult to learn robust and discriminative cross-modality features by existing metric learning methods for feature matching. Instead of applying metric learning on cross-modality data, we propose to unify the modality between images and point clouds by pretrained large-scale models first, and then establish robust correspondence within the same modality. We show that the intermediate features, called diffusion features, extracted by depth-to-image diffusion models are semantically consistent between images and point clouds, which enables the building of coarse but robust cross-modality correspondences. We further extract geometric features on depth maps produced by the monocular depth estimator. By matching such geometric features, we significantly improve the accuracy of the coarse correspondences produced by diffusion features. Extensive experiments demonstrate that without any task-specific training, direct utilization of both features produces accurate image-to-point cloud registration. On three public indoor and outdoor benchmarks, the proposed method averagely achieves a 20.6 percent improvement in Inlier Ratio, a $3.0\times$ higher Inlier Number, and a 48.6 percent improvement in Registration Recall than existing state-of-the-arts. The code and additional results are available at \url{https://whu-usi3dv.github.io/FreeReg/}. Haiping Wang 0004, Yuan Liu 0025, Bing Wang 0013, Yujing Sun 0001, Zhen Dong 0005, Wenping Wang 0001, Bisheng Yang |
ICLR | 5 |
| 2024 | An Adaptive Local Joint-Weighted Method for Continuous Stripping Intensity Correction of Airborne Bathymetric LiDARabstractAlongside geometric information accepted widely, airborne laser bathymetry (ALB) typically records the radiometric properties of sensed targets and assists with strip registration, ground (sediment)-type classification, and geometric modeling. Continuous intensity stripping frequently occurs in the ALB-recorded intensity images due to automatic gain control (AGC), an established circuit to compensate for echo power variations caused by distance changes. This situation is exacerbated by the delayed gain control response and inadequate adaptability to target. To address this issue, an adaptive local joint-weighted method for continuous intensity stripping correction is designed. In this method, an index-sharing mechanism is established between intensity, point cloud, emission angle, and return number for rapid localization and intervention of intensity. Unordered intensities are divided into scan line units based on the signal emission period, and subsequent stripped intensity range identification is achieved by assessing the intensity difference between adjacent scan lines. Neighboring scan lines and neighborhood intensities are weighted to jointly reset the stripped intensity values. In the new method, a bidirectional shifting strategy is implemented to attenuate the degradation of the intensity correction accuracy due to the gradual accumulation of correction defects. From the two measurement missions using the ALB, it is concluded that the stripping intensity is mainly distributed in the region of maximum water depth detectable by the sensor and areas with mixed high and low returns. Compared with the original intensity, the mean absolute percentage error (MAPE) and the root mean square error (RMSE) of the corrected intensity decreased by 35% and 40%, respectively, and the variation coefficient of the stripped intensity and the neighboring intensity decreased from 0.4 to 0.18. Xue Ji, Zhen Dong 0005, Mingchang Wang, Wenxue Xu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Detection and Restoration of Saturated Laser Waveforms Without Prior KnowledgeabstractThe avalanche photodiode equipped in LiDAR heavily relies on the reliability of the gain amplifier to consistently output a sufficient voltage signal, enabling the weak return signal to be captured. However, in certain situations, an increase in voltage will lead to saturation, which distorts the full waveform and results in inaccurate coordinate and intensity readings, causing issues in various applications. To minimize data voids and eliminate artifacts related to the analysis of clipped waveforms, this study focuses on developing an automated framework [Detection and Restoration of Saturated Waveforms without Prior Information (DRSW)] for detecting and recovering saturated waveforms associated with high reflectivity or near-field targets without any prior knowledge. The framework designs four feature descriptors based on topographic, intensity, and waveform characteristics to differentiate between saturated and unsaturated signals, enabling automatic segmentation into two categories. To address the challenges posed by signal saturation-induced peak flattening and broadening, a combined halves-Gaussian model (CHGM) is crucially presented to describe the rising and falling edges of saturated waveform with different halves of normal Gaussian functions (halves-Gaussian functions). In the CHGM, background noise is segmented rather than as a whole to improve modeling accuracy. The effectiveness of CHGM in reconstructing saturated waveforms is evaluated through rigorous comparative experiments with cubic spline, Kriging, and Gaussian mixture model (GMM) methods. The results indicate that the maximum correction error of the CHGM is limited to within 7 DN, and the peak position remains unaltered without any shift. Xue Ji, Zhen Dong 0005, Wenxue Xu, Yanxiong Liu, Mingchang Wang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | A Novel Method for Registration of MLS and Stereo Reconstructed Point CloudsabstractCross-source point cloud registration is a prerequisite for effectively leveraging the complementary information of multiple 3D sensors. However, existing point cloud registration methods have primarily focused on the registration of mono-source point clouds and typically fail to register cross-source data with varying noise patterns and capture characteristics. In this paper, we present a new algorithm for cross-source point cloud registration between MLS point clouds and stereo-reconstructed point clouds. Our method has two key designs. Firstly, we design a novel descriptor with in-plane rotation-equivariance by leveraging the accessible gravity prior, yielding strong descriptiveness, better robustness, and improved efficiency. Secondly, based on the noise pattern of stereo-reconstructed point clouds, a novel disparity-weighted correspondence scoring strategy is proposed to strengthen the registration accuracy. In comparison to existing registration baselines, our method achieves a 32.6% higher Registration Recall on cross-source datasets of KITTI and KITTI-360 and a 23.1% higher Registration Recall on mono-source datasets of KITTI. Notably, our method also outperforms RANSAC-based methods in terms of computational efficiency with a 10× ~ 70× speedup. The source code and datasets have been available at https://github.com/WHU-USI3DV/MSReg. Haiping Wang 0004, Zhen Dong 0005, Yuan Liu 0025, Bisheng Yang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | PointNAT: Large-Scale Point Cloud Semantic Segmentation via Neighbor Aggregation With TransformerabstractGiven the prominence of 3D sensors in recent years, 3D point clouds are worthy to be further investigated for environment perception and scene understanding. Learning accurate local and global contexts in point clouds is pivotal for semantic segmentation, and neighbor aggregation and Transformers have achieved notable success in local and global perception in point cloud analysis, respectively. Nevertheless, studying each independently is far from the optimal solution for comprehensive feature learning. To address this, we take a novel step towards investigating and integrating the structures of neighbor aggregation and Transformers. In this paper, we introduce Point Neighbor Aggregation with Transformer (PointNAT), a conceptually straightforward and effective approach aiming to enhance the performance of 3D point cloud semantic segmentation. PointNAT consists of a Neighbor Aggregation Block (NAB) for local perception, a Point Transformer Block (PTB) for global modeling, and a Hybrid Block to connect NABs and PTBs. NABs effectively learn complex local features at varying scales through an improved neighbor aggregation operation and a multi-head mechanism. PTBs efficiently perform global attention using a small set of learnable key points. Hybrid Blocks serve as high-and-low frequency signal hybridizers, merging the strengths of these two blocks by adaptively assigning hybrid weights to local and global contexts. We have evaluated the performance of PointNAT with state-of-the-art networks on several benchmarks, including S3DIS, Toronto3D, and SensatUrban. PointNAT achieves mIoU scores of 77.8%, 84.7%, and 65.2% in these three dataset, respectively. Furthermore, it outperforms the baseline approach PointNeXt by 3.0%, 1.3%, and 4.2%, respectively, while utilizing only 59.9% of the parameters and 15.2% of the FLOPs. The results demonstrate PointNAT’s superior ability in accurately segmenting large-scale 3D point cloud scenes, emphasizing its potential to advance environment perception and scene understanding. Our code is available at https://github.com/zeng-ziyin/PointNAT. Ziyin Zeng, Huan Qiu, Jian Zhou 0011, Zhen Dong 0005, Jinsheng Xiao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | WHU-Railway3D: A Diverse Dataset and Benchmark for Railway Point Cloud Semantic SegmentationabstractPoint cloud semantic segmentation (PCSS) shows great potential in generating accurate 3D semantic maps for digital twin railways. Deep learning-based methods have seen substantial advancements, driven by numerous PCSS datasets. Nevertheless, existing datasets tend to neglect railway scenes, with limitations in scale, categories, and scene diversity. This motivated us to establish WHU-Railway3D, a diverse PCSS dataset specifically designed for railway scenes. WHU-Railway3D is categorized into urban, rural, and plateau railways based on scene complexity and semantic class distribution. The dataset spans approximately 30 km with 4.6 billion points labeled into 11 classes, such as rails, masts, overhead lines, and fences. In addition to 3D coordinates, WHU-Railway3D provides rich attribute information such as reflected intensity, scanning angle, and number of returns. Cutting-edge methods are extensively evaluated on the dataset, followed by in-depth analysis. Lastly, key challenges and potential future work are identified to stimulate further innovative research. The dataset is accessible athttps://github.com/WHU-USI3DV/WHU-Railway3D. Yuzhou Zhou, Bing Wang 0013, Jianping Li 0004, Zhen Dong 0005, Chenglu Wen, Zhiliang Ma, Bisheng Yang |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2023 | KT-Net: Knowledge Transfer for Unpaired 3D Shape CompletionabstractUnpaired 3D object completion aims to predict a complete 3D shape from an incomplete input without knowing the correspondence between the complete and incomplete shapes. In this paper, we propose the novel KTNet to solve this task from the new perspective of knowledge transfer. KTNet elaborates a teacher-assistant-student network to establish multiple knowledge transfer processes. Specifically, the teacher network takes complete shape as input and learns the knowledge of complete shape. The student network takes the incomplete one as input and restores the corresponding complete shape. And the assistant modules not only help to transfer the knowledge of complete shape from the teacher to the student, but also judge the learning effect of the student network. As a result, KTNet makes use of a more comprehensive understanding to establish the geometric correspondence between complete and incomplete shapes in a perspective of knowledge transfer, which enables more detailed geometric inference for generating high-quality complete shapes. We conduct comprehensive experiments on several datasets, and the results show that our method outperforms previous methods of unpaired point cloud completion by a large margin. Code is available at https://github.com/a4152684/KT-Net. Xin Wen 0003, Zhen Dong 0005, Yu-Shen Liu, Xiongwu Xiao, Bisheng Yang |
AAAI | 4 |
| 2023 | Robust Multiview Point Cloud Registration with Reliable Pose Graph Initialization and History ReweightingabstractIn this paper, we present a new method for the multi-view registration of point cloud. Previous multiview registration methods rely on exhaustive pairwise registration to construct a densely-connected pose graph and apply Iteratively Reweighted Least Square (IRLS) on the pose graph to compute the scan poses. However, constructing a densely-connected graph is time-consuming and contains lots of outlier edges, which makes the subsequent IRLS struggle to find correct poses. To address the above problems, we first propose to use a neural network to estimate the overlap between scan pairs, which enables us to construct a sparse but reliable pose graph. Then, we design a novel history reweighting function in the IRLS scheme, which has strong robustness to outlier edges on the graph. In comparison with existing multiview registration methods, our method achieves 11% higher registration recall on the 3DMatch dataset and ~ 13% lower registration errors on the ScanNet dataset while reducing ~ 70% required pairwise registrations. Comprehensive ablation studies are conducted to demonstrate the effectiveness of our designs. The source code is available at https://github.com/WHU-USI3DV/SGHR. Haiping Wang 0004, Yuan Liu 0025, Zhen Dong 0005, Yulan Guo, Yu-Shen Liu, Wenping Wang 0001, Bisheng Yang |
CVPR | 3 |
| 2023 | RoReg: Pairwise Point Cloud Registration With Oriented Descriptors and Local RotationsabstractWe present RoReg, a novel point cloud registration framework that fully exploits oriented descriptors and estimated local rotations in the whole registration pipeline. Previous methods mainly focus on extracting rotation-invariant descriptors for registration but unanimously neglect the orientations of descriptors. In this paper, we show that the oriented descriptors and the estimated local rotations are very useful in the whole registration pipeline, including feature description, feature detection, feature matching, and transformation estimation. Consequently, we design a novel oriented descriptor RoReg-Desc and apply RoReg-Desc to estimate the local rotations. Such estimated local rotations enable us to develop a rotation-guided detector, a rotation coherence matcher, and a one-shot-estimation RANSAC, all of which greatly improve the registration performance. Extensive experiments demonstrate that RoReg achieves state-of-the-art performance on the widely-used 3DMatch and 3DLoMatch datasets, and also generalizes well to the outdoor ETH dataset. In particular, we also provide in-depth analysis on each component of RoReg, validating the improvements brought by oriented descriptors and the estimated local rotations. Source code and supplementary material are available at https://github.com/HpWang-whu/RoReg. Haiping Wang 0004, Yuan Liu 0025, Qingyong Hu, Bing Wang 0013, Zhen Dong 0005, Yulan Guo, Wenping Wang 0001, Bisheng Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | A Robust Density Estimation Method for Glacier-Height Retrieval From ICESat-2 Photon-Counting DataabstractRetrieval of glacier surface heights from ICESat-2 photon-counting data is significant to monitor glacier surface morphologies (e.g., crevasses) and their changes. However, existing methods are susceptible to complex glacier surfaces and diverse signal-to-noise ratios (SNRs). Therefore, we propose a robust density estimation method for glacier-height retrieval from ICESat-2 photon-counting data. A multi-scale random sample consensus (RANSAC) strategy is firstly employed to obtain the best elliptical neighborhood for each photon, to improve the consistency of neighboring photons. Then, a hybrid weighted density estimation method is developed to robustly describe the differences in the spatial distribution patterns between signal and noise photons. High-quality extraction of signal photons is finally implemented using adaptive thresholds in the local along-track segments. To test its performance, four datasets with strong and weak beams from the Jakobshavn Isbræ Glacier, western Greenland, were selected. Results showed that the proposed method achieved excellent performances in glacier-height retrieval in various glacier surface morphologies and diverse SNRs, due to discriminative densities between noise and signal photons were obtained. The average percentages of along-track segments with signal densities larger than noise densities for weak and strong beams were 92.0% and 99.9%. The averageF1-scoresreached 0.927 and 0.968 for the weak and strong beams, even theF1-scorereached 0.892 for the heavily crevassed complex surface with a low SNR. Moreover, comparison experiments with existing methods demonstrated that the proposed method showed superior performances in density estimation and glacier-height retrieval, and improved more than 17.33% and 16.10% in the most complex surface, respectively. Ruijie Chang, Ronggang Huang, Liming Jiang 0002, Zhen Dong 0005, Hansheng Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | WHU-Helmet: A Helmet-Based Multisensor SLAM Dataset for the Evaluation of Real-Time 3-D Mapping in Large-Scale GNSS-Denied EnvironmentsabstractReal-time 3D mapping of large-scale Global Navigation Satellite System (GNSS)-denied environments plays an important role in forest inventory management, disaster emergency response, and underground facility maintenance. Compact helmet laser scanning (HLS) systems keep the same direction as the user’s line of sight and have the advantage of “what you see is what you get”, providing a promising and efficient solution for 3D geospatial information acquisition. However, the violent motion of the helmet, the limited field of view of the laser scanner, and the repeated symmetrical geometric structures in GNSS-denied environments pose enormous challenges for the existing simultaneous localization and mapping (SLAM) algorithms. To promote the development of HLS and explore its application in large-scale GNSS-denied environments, the first large-scale HLS dataset covering multiple difficult GNSS-denied areas (e.g., forests, mountains, underground spaces) was built in this study. Besides using an additional very high accuracy fiber-optic inertial measurement unit (IMU), a novel post-processing multi-source fusion method—progressive trajectory correction (PTC)—is proposed to generate a reliable ground-truth trajectory for the benchmark, which overcomes the problems of scan matching degradation and non-rigid distortion. The accuracies of the ground truth are controlled and checked by manually surveyed feature points along the trajectory. Finally, the existing state-of-the-art SLAM methods were evaluated on the WHU-Helmet dataset, summarizing the future HLS SLAM research trends. The full dataset is available for download at: https: //github.com/kafeiyin00/WHU-HelmetDataset. Jianping Li 0004, Weitong Wu 0002, Bisheng Yang, Xianghong Zou, Yandi Yang, Xin Zhao 0026, Zhen Dong 0005 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2023 | SE-Calib: Semantic Edge-Based LiDAR-Camera Boresight Online Calibration in Urban ScenesabstractRigorous boresight calibration between light detection and ranging (LiDAR) and the camera is crucial for geometry and optical information fusion in earth observation and robotic applications. Although boresight parameters can be obtained through pre-calibration with artificial targets, unforeseen movement of sensors during data collection can lead to significant errors in the boresight parameters. To address this issue, we propose SE-Calib, an automatic and target-free online boresight calibration method for LiDAR-Camera systems. SE-Calib firstly extracts semantic edge features from both point clouds and images simultaneously using the 3D semantic segmentation (3D-SS) and 2D semantic edge detection (2D-SED) methods. The boresight parameters are then optimized with an adaptive solver and maximizing the Soft Semantic Response Consistency Metric (SSRCM) scores iteratively. The SSRCM is designed to evaluate the coherence of cross-modular semantic edge features, and a confidence function is proposed to filter out unreliable optimization results. Experiments conducted on challenging urban datasets show an average boresight error of 0.206 degrees (2.47 pixels in reprojection error), demonstrating the effectiveness and robustness of the proposed method. Youqi Liao, Jianping Li 0004, Shuhao Kang, Guifang Zhu, Shenghai Yuan 0001, Zhen Dong 0005, Bisheng Yang |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2023 | Multiobjective Optimization-Based Terrestrial Laser Scanning Layout Planning for Landslide MonitoringabstractTerrestrial laser scanning (TLS) is an important means to monitor landslides, and the layout is the key to guarantee captured point clouds with a high quality and low cost. Nevertheless, TLS layout in landslide monitoring is currently determined by user’s subjective experience, and shows a poor performance in reliability and point accuracy. Therefore, we propose multi-objective optimization based TLS layout planning for landslide monitoring. Three-dimensional (3D) scanning simulation of candidate scan positions is first carried out. A multi-objective optimization problem (MOP) is then constructed according to requirements of landslide TLS monitoring in the coverage, cost, and point accuracy. Finally, the combination of an improved non-dominated sorting genetic algorithm Ⅱ (NSGA-II) and the Pareto-Edgeworth-Grierson (PEG) algorithm is introduced to solve the MOP and obtain the optimal layout. To test the proposed method, experiments were performed in two landslide sites in Wuhan, China. Results showed that the planned TLS layouts reached a trade-off among coverage, point accuracy, and the number of TLS scan positions. Comparison experiments based on simulation demonstrated that the constraints fused into NSGA-II dramatically improved the reliability of the planned TLS layout, and the proposed method improved the coverage and point accuracy by 11.62% and 1.52 times compared to the user’s experience. Real scanning demonstrated the proposed method improved the coverage by 10.48%, and reduced the deformation error by 0.78 mm~7.66 mm compared to the user’s experience. In summary, the proposed method is an efficient approach for TLS layout planning, which can improve the ability of landslide TLS monitoring. Wendian Zhang, Ronggang Huang, Zhen Dong 0005, Liming Jiang 0002, Yuanping Xia, Benfu Chen, Hansheng Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Point Cloud Completion Via Skeleton-Detail TransformerabstractPoint cloud shape completion plays a central role in diverse 3D vision and robotics applications. Early methods used to generate global shapes without local detail refinement. Current methods tend to leverage local features to preserve the observed geometric details. However, they usually adopt the convolutional architecture over the incomplete point cloud to extract local features to restore the diverse information of both latent shape skeleton and geometric details, where long-distance correlation among the skeleton and details is ignored. In this work, we present a coarse-to-fine completion framework, which makes full use of both neighboring and long-distance region cues for point cloud completion. Our network leverages a Skeleton-Detail Transformer, which contains cross-attention and self-attention layers, to fully explore the correlation from local patterns to global shape and utilize it to enhance the overall skeleton. Also, we propose a selective attention mechanism to save memory usage in the attention process without significantly affecting performance. We conduct extensive experiments on the ShapeNet dataset and real-scanned datasets. Qualitative and quantitative evaluations demonstrate that our proposed network outperforms current state-of-the-art methods. Huajian Zhou, Zhen Dong 0005, Jun Liu 0036, Qingan Yan, Chunxia Xiao |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2023 | Rank-PointRetrieval: Reranking Point Cloud Retrieval via a Visually Consistent Registration EvaluationabstractPoint cloud-based place recognition is a fundamental part of the localization task, and it can be achieved through a retrieval process. Reranking is a critical step in improving the retrieval accuracy, yet little effort has been devoted to reranking in point cloud retrieval. In this paper, we investigate the versatility of rigid registration in reranking the point cloud retrieval results. Specifically, after obtaining the initial retrieval list based on the global point cloud feature distance, we perform registration between the query and point clouds in the retrieval list. We propose an efficient strategy based on visual consistency to evaluate each registration with a registration score in an unsupervised manner. The final reranked list is computed by considering both the original global feature distance and the registration score. In addition, we find that the registration score between two point clouds can also be used as a pseudo label to judge whether they represent the same place. Thus, we can create a self-supervised training dataset when there is no ground truth of positional information. Moreover, we develop a new probability-based loss to obtain more discriminative descriptors. The proposed reranking approach and the probability-based loss can be easily applied to current point cloud retrieval baselines to improve the retrieval accuracy. Experiments on various benchmark datasets show that both the reranking registration method and probability-based loss can significantly improve the current state-of-the-art baselines. Huajian Zhou, Zhen Dong 0005, Qingan Yan, Chunxia Xiao |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2022 | 3D Shape Reconstruction from 2D Images with Disentangled Attribute FlowabstractReconstructing 3D shape from a single 2D image is a challenging task, which needs to estimate the detailed 3D structures based on the semantic attributes from 2D image. So far, most of the previous methods still struggle to extract semantic attributes for 3D reconstruction task. Since the semantic attributes of a single image are usually implicit and entangled with each other, it is still challenging to reconstruct 3D shape with detailed semantic structures represented by the input image. To address this problem, we propose 3DAttriFlow to disentangle and extract semantic attributes through different semantic levels in the input images. These disentangled semantic attributes will be integrated into the 3D shape reconstruction process, which can provide definite guidance to the reconstruction of specific attribute on 3D shape. As a result, the 3D decoder can explicitly capture high-level semantic features at the bottom of the network, and utilize low-level features at the top of the network, which allows to reconstruct more accurate 3D shapes. Note that the explicit disentangling is learned without extra labels, where the only supervision used in our training is the input image and its corresponding 3D shape. Our comprehensive experiments on ShapeNet dataset demonstrate that 3DAttriFlow outperforms the state-of-the-art shape reconstruction methods, and we also validate its generalization ability on shape completion task. Code is available at https://github.com/junshengzhou/3DAttriFlow. Xin Wen 0003, Junsheng Zhou, Yu-Shen Liu, Zhen Dong 0005, Zhizhong Han |
CVPR | 5 |
| 2022 | PC2-PU: Patch Correlation and Point Correlation for Effective Point Cloud UpsamplingabstractPoint cloud upsampling is to densify a sparse point set acquired from 3D sensors, providing a denser representation for the underlying surface. Existing methods divide the input points into small patches and upsample each patch separately, however, ignoring the global spatial consistency between patches. In this paper, we present a novel method PC$^2$-PU, which explores patch-to-patch and point-to-point correlations for more effective and robust point cloud upsampling. Specifically, our network has two appealing designs: (i) We take adjacent patches as supplementary inputs to compensate the loss structure information within a single patch and introduce a Patch Correlation Module to capture the difference and similarity between patches. (ii) After augmenting each patch's geometry, we further introduce a Point Correlation Module to reveal the relationship of points inside each patch to maintain the local spatial consistency. Extensive experiments on both synthetic and real scanned datasets demonstrate that our method surpasses previous upsampling methods, particularly with the noisy inputs. The code and data are at: https://github.com/chenlongwhu/PC2-PU.git. Chen Long, Ruihui Li, Hao Wang 0057, Zhen Dong 0005, Bisheng Yang |
ACM Multimedia | 5 |
| 2022 | You Only Hypothesize Once: Point Cloud Registration with Rotation-equivariant DescriptorsabstractIn this paper, we propose a novel local descriptor-based framework, called You Only Hypothesize Once (YOHO), for the registration of two unaligned point clouds. In contrast to most existing local descriptors which rely on a fragile local reference frame to gain rotation invariance, the proposed descriptor achieves the rotation invariance by recent technologies of group equivariant feature learning, which brings more robustness to point density and noise. Meanwhile, the descriptor in YOHO also has a rotation-equivariant part, which enables us to estimate the registration from just one correspondence hypothesis. Such property reduces the searching space for feasible transformations, thus greatly improving both the accuracy and the efficiency of YOHO. Extensive experiments show that YOHO achieves superior performances with much fewer needed RANSAC iterations on four widely-used datasets, the 3DMatch/3DLoMatch datasets, the ETH dataset and the WHU-TLS dataset. More details are shown in our project page: https://hpwang-whu.github.io/YOHO/. Haiping Wang 0004, Yuan Liu 0025, Zhen Dong 0005, Wenping Wang 0001 |
ACM Multimedia | 3 |
| 2022 | Automated 3D Road Boundary Extraction and Vectorization Using MLS Point CloudsabstractTo meet the urgent demands in a wide range of geospatial applications, such as road management, intelligent transportation systems, road safety evaluation, and traffic accident analysis, automatic and accurate extraction of 3D roads and associated geometric parameters from point clouds is receiving wide attention. In this paper, we propose an accurate 3D road boundary extraction and vectorization method to bridge the gap from unstructured mobile laser scanning (MLS) point clouds to the vector-based representation of road boundary. Firstly, we propose a supervoxel generation method to extract candidate curbs with fine border preservation and high computation efficiencies. Then the candidate curb supervoxels are recognized and clustered to produce continuous road boundary segments with a contracted distance clustering strategy. Finally, the vectorized road boundary is represented by fitting, tracking, and completion from the extracted road boundary segments, resulting in road geometric parameters including boundary location, road widths, turning radius, and slopes. The performance of the proposed method was evaluated on two large-scale datasets collected in urban and industrial areas. Comprehensive experiments reveal that the proposed method is robust to various road shapes and point densities, in terms of precision of 95.0% and recall of 91.0%, respectively. Xiaoxin Mi, Bisheng Yang, Zhen Dong 0005, Chi Chen 0002, Jianxiang Gu |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | Learnable Motion Coherence for Correspondence PruningabstractMotion coherence is an important clue for distinguishing true correspondences from false ones. Modeling motion coherence on sparse putative correspondences is challenging due to their sparsity and uneven distributions. Existing works on motion coherence are sensitive to parameter settings and have difficulty in dealing with complex motion patterns. In this paper, we introduce a network called Laplacian Motion Coherence Network (LMCNet) to learn motion coherence property for correspondence pruning. We propose a novel formulation of fitting coherent motions with a smooth function on a graph of correspondences and show that this formulation allows a closed-form solution by graph Laplacian. This closed-form solution enables us to design a differentiable layer in a learning framework to capture global motion coherence from putative correspondences. The global motion coherence is further combined with local coherence extracted by another local layer to robustly detect inlier correspondences. Experiments demonstrate that LMCNet has superior performances to the state of the art in relative camera pose estimation and correspondences pruning of dynamic scenes1. Yuan Liu 0025, Lingjie Liu, Cheng Lin 0001, Zhen Dong 0005, Wenping Wang 0001 |
CVPR | 4 |
| 2021 | P2-Net: Joint Description and Detection of Local Features for Pixel and Point MatchingabstractAccurately describing and detecting 2D and 3D key-points is crucial to establishing correspondences across images and point clouds. Despite a plethora of learning-based 2D or 3D local feature descriptors and detectors having been proposed, the derivation of a shared descriptor and joint keypoint detector that directly matches pixels and points remains under-explored by the community. This work takes the initiative to establish fine-grained correspondences between 2D images and 3D point clouds. In order to directly match pixels and points, a dual fully-convolutional framework is presented that maps 2D and 3D inputs into a shared latent representation space to simultaneously describe and detect keypoints. Furthermore, an ultra-wide reception mechanism and a novel loss function are designed to mitigate the intrinsic information variations between pixel and point local regions. Extensive experimental results demonstrate that our framework shows competitive performance in fine-grained matching between images and point clouds and achieves state-of-the-art results for the task of indoor visual localization. Our source code is available at https://github.com/BingCS/P2-Net. Bing Wang 0013, Changhao Chen, Zhaopeng Cui, Jie Qin 0004, Xiaoxuan Lu 0001, Zhengdi Yu, Peijun Zhao, Zhen Dong 0005, Fan Zhu 0001, Agathoniki Trigoni, Andrew Markham |
ICCV | 8 |
| 2021 | AdaFit: Rethinking Learning-based Normal Estimation on Point CloudsabstractThis paper presents a neural network for robust normal estimation on point clouds, named AdaFit, that can deal with point clouds with noise and density variations. Existing works use a network to learn point-wise weights for weighted least squares surface fitting to estimate the normals, which has difficulty in finding accurate normals in complex regions or containing noisy points. By analyzing the step of weighted least squares surface fitting, we find that it is hard to determine the polynomial order of the fitting surface and the fitting surface is sensitive to outliers. To address these problems, we propose a simple yet effective solution that adds an additional offset prediction to improve the quality of normal estimation. Furthermore, in order to take advantage of points from different neighborhood sizes, a novel Cascaded Scale Aggregation layer is proposed to help the network predict more accurate point-wise offsets and weights. Extensive experiments demonstrate that AdaFit achieves state-of-the-art performance on both the synthetic PCPNet dataset and the real-word SceneNN dataset. The code is publicly available at https://github.com/Runsong123/AdaFit. Runsong Zhu, Yuan Liu 0025, Zhen Dong 0005, Yuan Wang 0035, Tengping Jiang, Wenping Wang 0001, Bisheng Yang |
ICCV | 3 |
| 2021 | Leaf and Wood Separation for Individual Trees Using the Intensity and Density Data of Terrestrial Laser ScannersabstractTerrestrial laser scanning (TLS) is a highly effective and noninvasive technology for retrieving the structural and biophysical attributes of trees using 3-D high-accuracy and high-density point clouds. The separation of leaf and wood points in TLS data is a prerequisite for the accurate and reliable derivation of these attributes. In this study, a new method is proposed to separate the leaf and wood points of individual trees by combining the TLS radiometric (intensity) and geometric (density) data. The leaf points are separated from the wood ones through three steps. First, the corrected intensity data are used to separate a part of the leaf points preliminarily given the differences in reflectance characteristics. Second, the density data are adopted for the further separation of another part of the leaf points because the density of the remaining leaf points is smaller than that of the wood points. Finally, a connectivity clustering algorithm is conducted to form several clusters with different sizes (points) and the remaining leaf points are separated in accordance with the cluster sizes. Eight different trees are selected to evaluate the performance of the proposed method. The averaged overall accuracy and kappa coefficient of the eight trees are approximately 95% and 0.81, respectively. The results suggest that the combination of TLS intensity and density data can perform a superior separation of leaf and wood points in terms of efficiency and accuracy, and the proposed separation method can be accurately and robustly used for various trees with different species, sizes, and structures. Kai Tan 0002, Zhen Dong 0005, Xiaojun Cheng |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2018 | Iterative Global Similarity Points: A Robust Coarse-to-Fine Integration Solution for Pairwise 3D Point Cloud RegistrationabstractIn this paper, we propose a coarse-to-fine integration solution inspired by the classical ICP algorithm, to pairwise 3D point cloud registration with two improvements of hybrid metric spaces (e.g., BSC feature and Euclidean geometry spaces) and globally optimal correspondences matching. First, we detect the keypoints of point clouds and use the Binary Shape Context (BSC) descriptor to encode their local features. Then, we formulate the correspondence matching task as an energy function, which models the global similarity of keypoints on the hybrid spaces of BSC feature and Euclidean geometry. Next, we estimate the globally optimal correspondences through optimizing the energy function by the Kuhn-Munkres algorithm and then calculate the transformation based on the correspondences. Finally, we iteratively refine the transformation between two point clouds by conducting optimal correspondences matching and transformation calculation in a mutually reinforcing manner, to achieve the coarse-to-fine registration under an unified framework. The proposed method is evaluated and compared to several state-of-the-art methods on selected challenging datasets with repetitive, symmetric and incomplete structures. Comprehensive experiments demonstrate that the proposed IGSP algorithm obtains good performance and outperforms the state-of-the-art methods in terms of both rotation and translation errors. Yue Pan 0009, Bisheng Yang, Fuxun Liang, Zhen Dong 0005 |
3DV | 4 |
| 2018 | Joint Point Cloud and Image Based Localization for Efficient Inspection in Mixed RealityabstractThis paper introduces a method of structure inspection using mixed-reality headsets to reduce the human effort in reporting accurate inspection information such as fault locations in 3D coordinates. Prior to every inspection, the headset needs to be localized. While external pose estimation and fiducial marker based localization would require setup, maintenance, and manual calibration; marker-free self-localization can be achieved using the onboard depth sensor and camera. However, due to limited depth sensor range of portable mixed-reality headsets like Microsoft HoloLens, localization based on simple point cloud registration (sPCR) would require extensive mapping of the environment. Also, localization based on camera image would face same issues as stereo ambiguities and hence depends on viewpoint. We thus introduce a novel approach to Joint Point Cloud and Image-based Localization (JPIL) for mixed-reality headsets that uses visual cues and headset orientation to register small, partially overlapped point clouds and save significant manual labor and time in environment mapping. Our empirical results compared to sPCR show average 10 fold reduction of required overlap surface area that could potentially save on average 20 minutes per inspection. JPIL is not only restricted to inspection tasks but also can be essential in enabling intuitive human-robot interaction for spatial mapping and scene understanding in conjunction with other agents like autonomous robotic systems that are increasingly being deployed in outdoor environments for applications like structural inspection. Manash Pratim Das, Zhen Dong 0005, Sebastian A. Scherer |
IROS | 2 |
| 2018 | Automatic change detection in lane-level road networks using GPS trajectoriesabstractLane-level road network updating is crucial for urban traffic applications that use geographic information systems contributing to, for example, intelligent driving, route planning and traffic control. Researchers have developed various algorithms to update road networks using sensor data, such as high-definition images or GPS data; however, approaches that involve change detection for road networks at lane level using GPS data are less common. This paper presents a novel method for automatic change detection of lane-level road networks based on GPS trajectories of vehicles. The proposed method includes two steps: map matching at lane level and lane-level change recognition. To integrate the most up-to-date GPS data with a lane-level road network, this research uses a fuzzy logic road network matching method. The proposed map-matching method starts with a confirmation of candidate lane-level road segments that use error ellipses derived from the GPS data, and then computes the membership degree between GPS data and candidate lane-level segments. The GPS trajectory data is classified into successful or unsuccessful matches using a set of defuzzification rules. Any topological and geometrical changes to road networks are detected by analysing the two kinds of matching results and comparing their relationships with the original road network. Change detection results for road networks in Wuhan, China using collected GPS trajectories show that these methods can be successfully applied to detect lane-level road changes including added lanes, closed lanes and lane-changing and turning rules, while achieving a robust detection precision of above 80%. Xue Yang 0002, Luliang Tang, Kathleen Stewart, Zhen Dong 0005, Qingquan Li 0001 |
Int. J. Geogr. Inf. Sci. | 4 |
| 2016 | CLRIC: Collecting Lane-Based Road Information Via CrowdsourcingabstractLane-based road network information, such as the number and locations of traffic lanes on a road, has played an important role in intelligent transportation systems. In this paper, we propose a Collecting Lane-based Road Information via Crowdsourcing (CLRIC) method, which can automatically extract detailed lane structure of roads by using crowdsourcing data collected by vehicles. First, CLRIC filters the high-precision GPS data from the raw trajectories based on region growing clustering with prior knowledge. Second, CLRIC mines the number and locations of traffic lanes through optimized constrained Gaussian mixture model. Experiments are conducted with taxi GPS trajectories in Wuhan, China, and the results show that CLRIC is quantified and displays detailed road networks with the number and locations of traffic lanes comparing with the satellite image and human-interpreted situation. Luliang Tang, Xue Yang 0002, Zhen Dong 0005, Qingquan Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2013 | Automated Extraction of Building Outlines From Airborne Laser Scanning Point CloudsabstractAutomatic extraction of building outlines from airborne laser scanning (ALS) point clouds has been an active topic in the field of photogrammetry, remote sensing, and computer vision. In this letter, a marked point process method is implemented to extract building outlines from ALS point clouds. First, the Gibbs energy model of building objects is defined to describe the building points. Second, the defined Gibbs energy model is sampled within the framework of reversible-jump Markov chain Monte Carlo and optimized to find an optimal energy configuration by simulated annealing. Finally, the detected building objects are refined to eliminate false detections, and the outlines of buildings are derived from the detected building objects by morphological operators. The standard data set provided by ISPRS is used to verify the validity of the proposed method. The method extracted building objects from the standard data sets with an average completeness of 87.3% and correctness of 91.57% at the pixel level, and an average completeness of 77.6% (97.3%) and correctness of 98.1% (97.9%) at the object$(\hbox{object} > 50\ \hbox{m}^{2})$level. Bisheng Yang, Wenxue Xu, Zhen Dong 0005 |
IEEE Geosci. Remote. Sens. Lett. | 3 |