EDBT 2026 Demo / reviewers in the wild / expert
Bisheng Yang
dblp:56/2808
· DBLP profile ↗
48ranked-venue papers
9as first author
31since 2021 · last 2026
0000-0001-7736-0803ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 19 · 2 first-author · 16 since 2021Databases, data management, data science and information retrieval · 12 · 7 first-authorGraphics, computer vision, multimedia, augmented reality and games · 12 · 10 since 2021Artificial intelligence and machine learning · 11 · 11 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GAGS: Granularity-Aware Feature Distillation for Language Gaussian Splattingabstract3D open-vocabulary scene understanding, which accurately perceives complex semantic properties of objects in space, has gained significant attention in recent years. In this paper, we propose GAGS, a framework that distills 2D CLIP features into 3D Gaussian splatting, enabling open-vocabulary queries for renderings on arbitrary viewpoints. The main challenge of distilling 2D features for 3D fields lies in the multiview inconsistency of extracted 2D features, which provides unstable supervision for the 3D feature field. GAGS addresses this challenge with two novel strategies. First, GAGS associates the prompt point density of SAM with the camera distances to scene objects, which significantly improves the multiview consistency of segmentation results. Second, GAGS further decodes a granularity factor to guide the distillation process and this granularity factor can be learned in a unsupervised manner to only select the multiview consistent 2D features in the distillation process. Experimental results on two datasets show that GAGS improves visual grounding accuracy by an average of 10.9% and semantic segmentation accuracy by an average of 7.0%, with an inference speed 2× faster than baseline methods. Yuning Peng, Haiping Wang 0004, Yuan Liu 0025, Chenglu Wen, Zhen Dong 0005, Bisheng Yang |
AAAI | 6 |
| 2026 | LifelongPR: Lifelong Point Cloud Place Recognition Based on Sample Replay and Prompt LearningabstractPoint cloud place recognition (PCPR) determines the geo-location within a prebuilt map and plays a crucial role in photogrammetry and robotics applications such as autonomous driving, intelligent transportation, and augmented reality. In real-world large-scale deployments of a geographic positioning system, PCPR models must continuously acquire, update, and accumulate knowledge to adapt to diverse and dynamic environments, i.e., the ability known as continual learning (CL). However, existing PCPR models often suffer from catastrophic forgetting, leading to significant performance degradation in previously learned scenes when adapting to new environments or sensor types. This results in poor model scalability, increased maintenance costs, and system deployment difficulties, undermining the practicality of PCPR. To address these issues, we propose LifelongPR, a novel continual learning framework for PCPR, which effectively extracts and fuses knowledge from sequential point cloud data. First, to alleviate the knowledge loss, we propose a replay sample selection method that dynamically allocates sample sizes according to each dataset’s information quantity and selects spatially diverse samples for maximal representativeness. Second, to handle domain shifts, we design a prompt learning-based CL framework with a lightweight continuous prompt module and a two-stage training strategy, enabling domain-specific feature adaptation while minimizing forgetting. Comprehensive experiments on large-scale public and self-collected datasets are conducted to validate the effectiveness of the proposed method. Compared with the state-of-the-art (SOTA) method, our method achieves 6.50% improvement in$mIR\text{@}1$, 7.96% improvement in$mR\text{@}1$, and an 8.95% reduction in$F$. The code and pre-trained models are publicly available athttps://zouxianghong.github.io/LifelongPR Xianghong Zou, Jianping Li 0004, Zhe Chen 0028, Zhen Dong 0005, Qiegen Liu, Bisheng Yang |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2025 | V2X-R: Cooperative LiDAR-4D Radar Fusion with Denoising Diffusion for 3D Object DetectionabstractCurrent Vehicle-to-Everything (V2X) systems have significantly enhanced 3D object detection using LiDAR and camera data. However, they face performance degradation in adverse weather. Weather-robust 4D radar, with Doppler velocity and additional geometric information, offers a promising solution to this challenge. To this end, we present V2X-R, the first simulated V2X dataset incorporating LiDAR, camera, and 4D radar modalities. V2X-R contains 12,079 scenarios with 37,727 frames of LiDAR and 4D radar point clouds, 150,908 images, and 170,859 annotated 3D vehicle bounding boxes. Subsequently, we propose a novel cooperative LiDAR-4D radar fusion pipeline for 3D object detection and implement it with multiple fusion strategies. To achieve weather-robust detection, we additionally propose a Multi-modal Denoising Diffusion (MDD) module in our fusion pipeline. MDD utilizes weather-robust 4D radar feature as a condition to guide the diffusion model in denoising noisy LiDAR features. Experiments show that our LiDAR-4D radar fusion pipeline demonstrates superior performance in the V2X-R dataset. Over and above this, our MDD module further improved the foggy/snowy performance of the basic fusion model by up to 5.73%/6.70% and barely disrupting normal performance. The dataset and code will be publicly available at: https://github.com/ylwhxht/V2X-R. Xun Huang 0003, Qiming Xia, Siheng Chen, Bisheng Yang, Xin Li 0003, Cheng Wang 0003, Chenglu Wen |
CVPR | 5 |
| 2025 | Dual-Frequency Spatio-Temporal Phase UnwrappingabstractPhase unwrapping poses a critical challenge in 3D reconstruction, particularly due to the presence of noise and discontinuities that compromise the accuracy of phase extraction. Existing convolutional neural network (CNN)-based methods have struggled to effectively integrate traditional approaches while fully utilizing both multi-frequency, i.e., temporal, and spatial information of the phase. In this paper, we propose a novel phase unwrapping method, STPhaseNet, which incorporates both temporal and spatial phase information into the CNN framework. Specifically, we introduce a temporal feature fusion module and a local attention mechanism to extract and integrate features from different frequency phases. To further leverage spatial phase information, we develop a spatial information extraction module that enlarges the local receptive field of the convolution and assigns weights based on the phase information of horizontal and vertical coordinates. Additionally, we design a globally optimized gradient residual loss function to exploit spatial constraints more effectively. To address the lack of real-world training data, we apply a Random Matrix Enlargement (RME) method to generate high-quality dual-frequency wrapped phase data along with corresponding absolute phases for training purposes. Extensive experiments demonstrate that STPhaseNet outperforms existing methods, achieving superior performance in phase unwrapping tasks. Shuo Du, Qin Zou 0001, Chi Chen 0002, Bisheng Yang |
ICASSP | 4 |
| 2025 | Vistadream: Sampling Multiview Consistent Images for Single-View Scene ReconstructionabstractIn this paper, we propose VistaDream a novel framework to reconstruct a 3D scene from a single-view image. Recent diffusion models enable generating high-quality novel-view images from a single-view input image. Most existing methods only concentrate on building the consistency between the input image and the generated images while losing the consistency between the generated images. VistaDream addresses this problem by a two-stage pipeline. In the first stage, VistaDream begins with building a global coarse 3D scaffold by zooming out a little step with inpainted boundaries and an estimated depth map. Then, on this global scaffold, we use iterative diffusion-based RGB-D inpainting to generate novel-view images to inpaint the holes of the scaffold. In the second stage, we further enhance the consistency between the generated novel-view images by a novel training-free Multiview Consistency Sampling (MCS) that introduces multi-view consistency constraints in the reverse sampling process of diffusion models. Experimental results demonstrate that without training or fine-tuning existing diffusion models, VistaDream achieves consistent and high-quality novel view synthesis using just single-view images and outperforms baseline methods by a large margin. The code, videos, and interactive demos are available at https://vistadream-project-page.github.io/. Haiping Wang 0004, Yuan Liu 0025, Ziwei Liu 0002, Wenping Wang 0001, Zhen Dong 0005, Bisheng Yang |
ICCV | 6 |
| 2025 | CityAnchor: City-scale 3D Visual Grounding with Multi-modality LLMsabstractIn this paper, we present a 3D visual grounding method called CityAnchor for localizing an urban object in a city-scale point cloud. Recent developments in multiview reconstruction enable us to reconstruct city-scale point clouds but how to conduct visual grounding on such a large-scale urban point cloud remains an open problem. Previous 3D visual grounding system mainly concentrates on localizing an object in an image or a small-scale point cloud, which is not accurate and efficient enough to scale up to a city-scale point cloud. We address this problem with a multi-modality LLM which consists of two stages, a coarse localization and a fine-grained matching. Given the text descriptions, the coarse localization stage locates possible regions on a projected 2D map of the point cloud while the fine-grained matching stage accurately determines the most matched object in these possible regions. We conduct experiments on the CityRefer dataset and a new synthetic dataset annotated by us, both of which demonstrate our method can produce accurate 3D visual grounding on a city-scale 3D point cloud. Haiping Wang 0004, Yuan Liu 0025, Zhiyang Dou, Yuexin Ma, Sibei Yang, Wenping Wang 0001, Zhen Dong 0005, Bisheng Yang |
ICLR | 11 |
| 2025 | Layered Denoising and Classification of Photon Point Cloud Data From ICESat-2 in Forest AreaabstractIce, Cloud, and land Elevation Satellite (ICESat-2) carries the Advanced Topographic Laser Altimeter System (ATLAS), which enhancing along-track sampling density but introduces substantial noise in photon point cloud data. Therefore, this study establishes a denoising and classification feature parameter system grounded in the three-dimensional spatial distribution characteristics of photon point clouds. Modeling is conducted in two layers: one layer for upper noise photons and canopy signal photons, and another layer for lower noise photons and ground signal photons. Machine learning and neural network algorithms are utilized to denoise and classify the original photon point clouds from ICESat-2, aiming to obtain a transferable and universally applicable supervised classification model for denoising photon point clouds. Recall, Precision, and the harmonic mean of Recall and Precision (F1 Score) are used as evaluation metrics to verify the accuracy of local, transfer, and global models. The results indicate that under various forest types and external conditions, the proposed photon point cloud Layered Denoising and Classification Model (LDCM) outperforms the Differential Regressive and Gaussian Adaptive Nearest Neighbor (DRAGANN, ICESat-2 ATL08 production algorithm), Ordering Points to Identify the Clustering Structure (OPTICS), and Adaptive Elevation Difference Thresholding (AEDTA) algorithms in terms of accuracy. Compared to the DRAGANN algorithm, the maximum accuracy improvement is 60%, with an average improvement of approximately 20%; compared to the OPTICS algorithm, the maximum accuracy improvement is 36%, with an average improvement of about 28%; compared to the AEDTA algorithm, the maximum accuracy improvement is 27%, with an average improvement of about 14%. The F1 Score for the validation set of the machine learning and neural network algorithms is above 0.94, with the Categorical Boosting (CatBoost) algorithm achieving the best performance. Both the transfer model and the global model have F1 Scores above 0.90. Therefore, the proposed photon point cloud LDCM not only demonstrates excellent classification accuracy but also exhibits good transferability and general applicability. Junfan Bao, Ningning Zhu, Zhen Dong 0005, Sheng Nie, Wenxia Dai, Ruixiong Kou, Bisheng Yang |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2025 | ATCM: Aerial-Terrestrial LiDAR-Based Collaborative Simultaneous Localization and MappingabstractMulti-robot collaborative simultaneous localization and mapping (C-SLAM) offers precise scene reconstruction over single-robot SLAM and enables the data fusion from heterogeneous robots. However, heterogeneous C-SLAM faces challenges in both accurate inter-robot loop closure detection and globally consistent data fusion due to inherent viewpoint disparities and heterogeneous data characteristics. This paper introduces ATCM, an Aerial-Terrestrial LiDAR-based C-SLAM method designed for heterogeneous robots without priori initial relative position. ATCM comprises three modules: single-robot front-end employing diverse SLAM methods, multi-robot loop closure detection, and global pose graph optimization. A novel LiDAR-based cross-view global loop descriptor is proposed for scan-to-scan heterogeneous inter-robot loop closure detection. By uniformly mapping cross-view information into the height domain and integrating dynamic height, the loop descriptor automatically achieves viewpoint correction. Additionally, we introduce a bidirectional loop detection algorithm that validates inter-robot loop closures through both forward and reverse detections. Finally, the two-stage global pose graph optimization integrates multi-source measurements, ensuring globally consistent mapping and localization with cross-view data. We have validated the effectiveness of ATCM on campus scenario datasets and the KITTI dataset, achieving a remarkable 21.95% improvement in trajectory accuracy and a 17.00% enhancement in map precision compared to high-precision point cloud maps, surpassing state-of-the-art LiDAR-based odometry methods. In the ablation experiments, the proposed loop descriptor achieved 97% accuracy in recognizing heterogeneous inter-robot loop closures. Moreover, compared to the traditional unidirectional method, the bidirectional loop detection method demonstrates up to a 31.2% improvement in loop closure accuracy. Chi Chen 0002, Bisheng Yang, Weitong Wu 0002, Shangzhe Sun, Zhiye Wang, Liuchun Li, Qin Zou 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | INF-PCA: Implicit Neural Field-Based Interactive Point Cloud Semantic AnnotationabstractPoint cloud semantic segmentation helps Intelligent Transportation Systems understand traffic scenes by assigning semantic label to each point in the point cloud, and it relies on large amounts of annotated training data. Nevertheless, manually annotating large-scale datasets of complex traffic scenes is quite time-consuming and tedious. This paper proposes INF-PCA, an interactive point cloud semantic annotation method based on implicit neural field, which allows users to achieve high-quality, large-scene and fast-response semantic annotation with only a few dozen mouse clicks. Firstly, the appearance, geometry and semantics of the point clouds are jointly represented by an implicit neural field, which maps a 3D spatial coordinate to its corresponding attributes. Secondly, an uncertainty-based semantic entropy loss and a supervoxel-based local consistency loss are designed to force the network to produce deterministic predictions with local consistency, thus generating smoother and more accurate boundaries. Furthermore, an active learning-based strategy for click-free annotation is proposed and analyzed to further reduce annotation pressure. Comprehensive experiments on multiple datasets including the road scene dataset Toronto3D revealed that INF-PCA can achieve more accurate annotations with faster response speed and only half of the clicks employed by the state-of-the-art methods, and that INF-PCA can be directly applied to intelligent transportation applications such as interactive segmentation of road scenes, inventory of transportation infrastructure assets, and production of high-definition map. Chen Long, Wang Wang, Zhen Dong 0005, Bisheng Yang |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2025 | WIN: Variable-View Implicit LiDAR Upsampling NetworkabstractLiDAR upsampling aims to increase the resolution of sparse point sets obtained from low-cost sensors, providing better performance for various downstream tasks (e.g., Autonomous Driving, High Definitation Map). Existing methods transfer LiDAR point cloud into range view, and focus on designing complex encoders or interpolation strategies to improve the resolution of LiDAR range images. However, our analysis shows that using the range view inevitably results in the loss of geometric information. We propose a Variable-View Implicit LiDAR Upsampling network, named WIN to solve this problem. It decouples range views into two novel virtual view representations, Horizontal Range View (HRV) and Vertical Range View (VRV). The key idea behind this is that introducing more perspectives can make up for the geometric information lost in a single perspective. We also prove theoretically that the proposed virtual view representation has a smaller error range compared to the range view representation. In addition, we design two novel strategies (i.e., contrast selection module and selection loss) to fuse the upsampling results of these two virtual representations and stabilize the whole training process. As a result, compared with the current state-of-the art (SOTA) method ILN, WIN introduces only 0.4M additional parameters, yet achieves a +4.53% increase in the MAE and a +7.01% increase in the IoU on the CARLA dataset. Furthermore, our method also outperforms all existing methods in downstream tasks (i.e., Depth Completion and Localization). The code and pre-trained models are available athttps://github.com/WHU-USI3DV/WIN Conglang Zhang, Chen Long, Zhen Dong 0005, Bisheng Yang |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2024 | Explicitly Guided Information Interaction Network for Cross-Modal Point Cloud Completion
Chen Long, Yuan Liu 0025, Zhen Dong 0005, Bisheng Yang |
ECCV (12) | 7 |
| 2024 | FreeReg: Image-to-Point Cloud Registration Leveraging Pretrained Diffusion Models and Monocular Depth EstimatorsabstractMatching cross-modality features between images and point clouds is a fundamental problem for image-to-point cloud registration. However, due to the modality difference between images and points, it is difficult to learn robust and discriminative cross-modality features by existing metric learning methods for feature matching. Instead of applying metric learning on cross-modality data, we propose to unify the modality between images and point clouds by pretrained large-scale models first, and then establish robust correspondence within the same modality. We show that the intermediate features, called diffusion features, extracted by depth-to-image diffusion models are semantically consistent between images and point clouds, which enables the building of coarse but robust cross-modality correspondences. We further extract geometric features on depth maps produced by the monocular depth estimator. By matching such geometric features, we significantly improve the accuracy of the coarse correspondences produced by diffusion features. Extensive experiments demonstrate that without any task-specific training, direct utilization of both features produces accurate image-to-point cloud registration. On three public indoor and outdoor benchmarks, the proposed method averagely achieves a 20.6 percent improvement in Inlier Ratio, a $3.0\times$ higher Inlier Number, and a 48.6 percent improvement in Registration Recall than existing state-of-the-arts. The code and additional results are available at \url{https://whu-usi3dv.github.io/FreeReg/}. Haiping Wang 0004, Yuan Liu 0025, Bing Wang 0013, Yujing Sun 0001, Zhen Dong 0005, Wenping Wang 0001, Bisheng Yang |
ICLR | 7 |
| 2024 | Semantic Segmentation of Airborne LiDAR Point Clouds With Noisy LabelsabstractHigh-quality point cloud annotation is labor-intensive and time-consuming, but it serves as a critical factor driving the success of LiDAR point cloud semantic segmentation. Leveraging low-quality labels in LiDAR point cloud processing is overlooked, despite the fact that noisy annotation has low labeling costs and abundant cross-modal resources (e.g., labels from images). To this end, we thoroughly investigate the performance of airborne LiDAR point cloud semantic segmentation models using noisy labels for the first time and find that it is closely related to object categories and learning stages. Then we propose a new semantic segmentation framework for LiDAR point cloud noisy learning called adaptive dynamic noise label correction (ADNLC), which consists of weak category priority, dynamic monitoring (DM), and historical choice (HC). With these methods, we can adaptively correct the noise labels of different categories according to their specific learning situations. Finally, we provide a comprehensive process for noise simulation, accuracy evaluation, and comparisons in airborne LiDAR point cloud learning from noisy labels. We conduct experiments on the ISPRS 3-D Labeling Vaihingen and Large-scale ALS data for Semantic Labeling in Dense Urban Areas (LASDU) datasets, and the results show that our ADNLC outperforms baseline methods by 30% and 16%, respectively, verifying the superiority of ADNLC and demonstrating the potential of noise labels in LiDAR data processing. Yuan Gao 0058, Shaobo Xia, Cheng Wang 0016, Xiaohuan Xi, Bisheng Yang, Chou Xie |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | GPR-Former: Detection and Parametric Reconstruction of Hyperbolas in GPR B-Scan Images With TransformersabstractGround Penetrating Radar (GPR) enables the non-invasive detection of various subsurface objects such as pipes, stones, etc. The location and size of the object in the medium could be obtained by fitting the generated hyperbolic signatures within the GPR B-scan and analyzing its parameters. In this paper, GPR-Former is proposed for automatic target detection and hyperbola fitting on GPR B-scan images. We have designed a transformer-based neural network to extract features to directly regress the parameters of hyperbolic signatures in the GPR B-scan data to detect targets beneath the ground automatically. A symmetry-constrained analytical solution for the hyperbolic parameters is proposed to refine the parameters derived from the transformer network, serving the extraction and analysis of buried objects in underground opaque spaces. Experiments are conducted on three datasets for the qualitative and quantitative validation of the GPR-Former, including ground-penetrating radar detection of submarine pipelines and land pipelines. Results show that the proposed method is able to automatically and efficiently extract hyperbolas from GPR B-scan images. True hyperbola-point precision (TP_Pre) and true hyperbola-point recall (TP_Rec) metrics are introduced to evaluate performances in parametric hyperbola extraction and fitting. The results show that the TP_Pre and TP_Rec of the proposed method reach 0.867, 0.402, 0.744 and 0.762, 0.736, 0.723, with an improvement of 6%, 22%, 4% compared with the state-of-the-art methods (C3 algorithm and migration learning-based method proposed by Yang), respectively. Ang Jin, Chi Chen 0002, Bisheng Yang, Qin Zou 0001, Zhiye Wang, Zhengfei Yan, Shaolong Wu, Jian Zhou 0011 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | A Novel Method for Registration of MLS and Stereo Reconstructed Point CloudsabstractCross-source point cloud registration is a prerequisite for effectively leveraging the complementary information of multiple 3D sensors. However, existing point cloud registration methods have primarily focused on the registration of mono-source point clouds and typically fail to register cross-source data with varying noise patterns and capture characteristics. In this paper, we present a new algorithm for cross-source point cloud registration between MLS point clouds and stereo-reconstructed point clouds. Our method has two key designs. Firstly, we design a novel descriptor with in-plane rotation-equivariance by leveraging the accessible gravity prior, yielding strong descriptiveness, better robustness, and improved efficiency. Secondly, based on the noise pattern of stereo-reconstructed point clouds, a novel disparity-weighted correspondence scoring strategy is proposed to strengthen the registration accuracy. In comparison to existing registration baselines, our method achieves a 32.6% higher Registration Recall on cross-source datasets of KITTI and KITTI-360 and a 23.1% higher Registration Recall on mono-source datasets of KITTI. Notably, our method also outperforms RANSAC-based methods in terms of computational efficiency with a 10× ~ 70× speedup. The source code and datasets have been available at https://github.com/WHU-USI3DV/MSReg. Haiping Wang 0004, Zhen Dong 0005, Yuan Liu 0025, Bisheng Yang |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | WHU-Railway3D: A Diverse Dataset and Benchmark for Railway Point Cloud Semantic SegmentationabstractPoint cloud semantic segmentation (PCSS) shows great potential in generating accurate 3D semantic maps for digital twin railways. Deep learning-based methods have seen substantial advancements, driven by numerous PCSS datasets. Nevertheless, existing datasets tend to neglect railway scenes, with limitations in scale, categories, and scene diversity. This motivated us to establish WHU-Railway3D, a diverse PCSS dataset specifically designed for railway scenes. WHU-Railway3D is categorized into urban, rural, and plateau railways based on scene complexity and semantic class distribution. The dataset spans approximately 30 km with 4.6 billion points labeled into 11 classes, such as rails, masts, overhead lines, and fences. In addition to 3D coordinates, WHU-Railway3D provides rich attribute information such as reflected intensity, scanning angle, and number of returns. Cutting-edge methods are extensively evaluated on the dataset, followed by in-depth analysis. Lastly, key challenges and potential future work are identified to stimulate further innovative research. The dataset is accessible athttps://github.com/WHU-USI3DV/WHU-Railway3D. Yuzhou Zhou, Bing Wang 0013, Jianping Li 0004, Zhen Dong 0005, Chenglu Wen, Zhiliang Ma, Bisheng Yang |
IEEE Trans. Intell. Transp. Syst. | 9 |
| 2024 | SGSR-Net: Structure Semantics Guided LiDAR Super-Resolution Network for Indoor LiDAR SLAMabstractMulti-Beam LiDAR (MBL) sensors sample the real-world with discrete 3D point clouds (PC) and have become a major and essential 3D sensing capability for autonomous robots. To ensure an accurate point sampling on surfaces, high-resolution MBL sensors (e.g., Ouster OS0-128) are commonly used to collect dense point clouds for robot tasks, including object detection and tracking, simultaneous localization and mapping (SLAM), in applications such as autonomous driving vehicles (ADVs). However, the high cost and large volume/weight/energy consumption of such sensors limit their usage in broader applications such as UAV/UGV swarms with small-scale agents with limited payload. Existing studies on Super-Resolution (SR) upsampling of the PC from low-resolution MBL have not considered the geometry semantics of the scenes, thus resulting in less optimal SR points for downstream subtasks (e.g., SLAM). Thus, this article proposes SGSR-Net, a structure semantics-guided MBL Super-Resolution network. SGSR-Net takes the low-resolution range images of the MBL sensors as input and produces dense and structure-aware Super-Resolution point cloud from those sparse measurements through a vertical spatial and channel attention-enhanced CNN model coupling with guided Monte Carlo filtering, for indoor LiDAR-SLAM applications. The SGSR-Net is validated using datasets collected by a UGV equipped with multiple MBL sensors. The results demonstrate that the proposed CG-LSR (CASE Attention Guided Encoder-Decoder LiDAR Super-Resolution Network) reduces the MAE of the SR points by 12.4% down to 0.177 m when compared with the state-of-the-art (SOTA) method Shan et al. (2020), Ren et al. (2021), Kwon et al. (2022), Long and Wang (2022). The indoor SLAM results with SR-points produced by SGSR-Net show that the mean and RMSE of the absolute pose error (APE) are decreased by 27% and 30%, down to 0.849 m and 0.902 m, respectively, which significantly improve the indoor-SLAM performance and stability of SOTA LiDAR-SLAM systems (i.e. LeGO-LOAM Shan and Englot (2018), Dellenbach et al. (2022), Vizzo et al. (2023), Zhang and Singh (2014)). Chi Chen 0002, Ang Jin, Zhiye Wang, Yongwei Zheng, Bisheng Yang, Jian Zhou 0011, Zhigang Tu 0001 |
IEEE Trans. Multim. | 5 |
| 2023 | KT-Net: Knowledge Transfer for Unpaired 3D Shape CompletionabstractUnpaired 3D object completion aims to predict a complete 3D shape from an incomplete input without knowing the correspondence between the complete and incomplete shapes. In this paper, we propose the novel KTNet to solve this task from the new perspective of knowledge transfer. KTNet elaborates a teacher-assistant-student network to establish multiple knowledge transfer processes. Specifically, the teacher network takes complete shape as input and learns the knowledge of complete shape. The student network takes the incomplete one as input and restores the corresponding complete shape. And the assistant modules not only help to transfer the knowledge of complete shape from the teacher to the student, but also judge the learning effect of the student network. As a result, KTNet makes use of a more comprehensive understanding to establish the geometric correspondence between complete and incomplete shapes in a perspective of knowledge transfer, which enables more detailed geometric inference for generating high-quality complete shapes. We conduct comprehensive experiments on several datasets, and the results show that our method outperforms previous methods of unpaired point cloud completion by a large margin. Code is available at https://github.com/a4152684/KT-Net. Xin Wen 0003, Zhen Dong 0005, Yu-Shen Liu, Xiongwu Xiao, Bisheng Yang |
AAAI | 7 |
| 2023 | Robust Multiview Point Cloud Registration with Reliable Pose Graph Initialization and History ReweightingabstractIn this paper, we present a new method for the multi-view registration of point cloud. Previous multiview registration methods rely on exhaustive pairwise registration to construct a densely-connected pose graph and apply Iteratively Reweighted Least Square (IRLS) on the pose graph to compute the scan poses. However, constructing a densely-connected graph is time-consuming and contains lots of outlier edges, which makes the subsequent IRLS struggle to find correct poses. To address the above problems, we first propose to use a neural network to estimate the overlap between scan pairs, which enables us to construct a sparse but reliable pose graph. Then, we design a novel history reweighting function in the IRLS scheme, which has strong robustness to outlier edges on the graph. In comparison with existing multiview registration methods, our method achieves 11% higher registration recall on the 3DMatch dataset and ~ 13% lower registration errors on the ScanNet dataset while reducing ~ 70% required pairwise registrations. Comprehensive ablation studies are conducted to demonstrate the effectiveness of our designs. The source code is available at https://github.com/WHU-USI3DV/SGHR. Haiping Wang 0004, Yuan Liu 0025, Zhen Dong 0005, Yulan Guo, Yu-Shen Liu, Wenping Wang 0001, Bisheng Yang |
CVPR | 7 |
| 2023 | DeepWE: A Deep Bayesian Active Learning Waypoint Estimator for Indoor WalkersabstractWaypoint estimation (WE) has a wide range of applications for indoor walkers, such as fire rescue and navigation to find exit doors, lifts, or stairs as examples of waypoints, etc. Data-driven WE has been on the rise with advancements in deep learning algorithms. The current WE methods, however, face two challenges. On the one hand, most waypoint detection approaches rely on visual sensors, hence, their estimation performance is limited by light when collecting visual data. On the other hand, data-driven methods necessitate a large number of labeled data to train a WE model, which significantly increases the time spent manually marking labels. Targeting the above two challenges, our work first proposes a novel deep Bayesian active learning waypoint estimator for indoor walkers (DeepWE) based on human activity recognition (HAR). This estimates six indoor waypoints through walkers’ daily activities due to the strong correlation between human activities and waypoints. First, an initial DeepWE model is developed using a Bayesian ensembled convolutional neural network (B-CNN) using the accelerometer and gyroscope data. Then, active learning is employed to query the most formative samples from pool points with four acquisition functions, and only these queried samples are labeled manually. Finally, the initial DeepWE model is updated from this labeled data using an incremental learning algorithm. Empirical results on two publicly available USC-HAD and OPPORTUNITY data sets show DeepWE performs a considerable accuracy boost for WE, with a substantial amount of acquired pool points reduction (more than 40%). Stefan Poslad, Qingquan Li 0001, Bisheng Yang, Jizhe Xia, Bang Wu 0001, Zhaoliang Luan, Yonglei Fan |
IEEE Internet Things J. | 4 |
| 2023 | SegTrans: Semantic Segmentation With Transfer Learning for MLS Point Cloudsabstract3D point cloud semantic segmentation plays an essential role in fine-grained scene understanding from photogrammetry to autonomous driving. Although recent efforts have been made to push the 3D semantic segmentation forward, many solutions cannot generalize well to new data with different sensor configurations. For example, when transferring the segmentation model learned from terrestrial laser scanning (TLS) data to mobile laser scanning (MLS) data, the performance drops dramatically. Besides, rich-labeled data is usually required. However, labeling point cloud data is time-consuming and label-intensive in practice. In light of this, we propose SegTrans, an unsupervised domain adaption method for the point cloud semantic segmentation task, which largely improves the generalization performance from one labeled dataset (source domain) to another unlabeled dataset (target domain). Specifically, we first introduce a data selection module (DSM) to tackle the discrepancy between different datasets at the data level. Then an adversarial learning module (ALM) with an adversarial loss is iteratively implemented to align the domain-specific feature in both the source and target domains, which only consists of two fully connected layers. Experiments show the overall accuracy of the proposed method achieves 88% OA on the TUM City Campus dataset (MLS dataset) when trained on the Semantic3D dataset (TLS dataset). Shuo Shen 0003, Yan Xia 0003, Andreas Eich, Yusheng Xu, Bisheng Yang, Uwe Stilla |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2023 | Consistent 3D Hand Reconstruction in Video via Self-Supervised LearningabstractWe present a method for reconstructing accurate and consistent 3D hands from a monocular video. We observe that the detected 2D hand keypoints and the image texture provide important cues about the geometry and texture of the 3D hand, which can reduce or even eliminate the requirement on 3D hand annotation. Accordingly, in this work, we propose$\mathrm{{S}^{2}HAND}$, a self-supervised 3D hand reconstruction model, that can jointly estimate pose, shape, texture, and the camera viewpoint from a single RGB input through the supervision of easily accessible 2D detected keypoints. We leverage the continuous hand motion information contained in the unlabeled video data and explore$\mathrm{{S}^{2}HAND(V)}$, which uses a set of weights shared$\mathrm{{S}^{2}HAND}$to process each frame and exploits additional motion, texture, and shape consistency constrains to obtain more accurate hand poses, and more consistent shapes and textures. Experiments on benchmark datasets demonstrate that our self-supervised method produces comparable hand reconstruction performance compared with the recent full-supervised methods in single-frame as input setup, and notably improves the reconstruction accuracy and consistency when using the video training data. Zhigang Tu 0001, Zhisheng Huang, Yujin Chen, Linchao Bao, Bisheng Yang, Junsong Yuan 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | RoReg: Pairwise Point Cloud Registration With Oriented Descriptors and Local RotationsabstractWe present RoReg, a novel point cloud registration framework that fully exploits oriented descriptors and estimated local rotations in the whole registration pipeline. Previous methods mainly focus on extracting rotation-invariant descriptors for registration but unanimously neglect the orientations of descriptors. In this paper, we show that the oriented descriptors and the estimated local rotations are very useful in the whole registration pipeline, including feature description, feature detection, feature matching, and transformation estimation. Consequently, we design a novel oriented descriptor RoReg-Desc and apply RoReg-Desc to estimate the local rotations. Such estimated local rotations enable us to develop a rotation-guided detector, a rotation coherence matcher, and a one-shot-estimation RANSAC, all of which greatly improve the registration performance. Extensive experiments demonstrate that RoReg achieves state-of-the-art performance on the widely-used 3DMatch and 3DLoMatch datasets, and also generalizes well to the outdoor ETH dataset. In particular, we also provide in-depth analysis on each component of RoReg, validating the improvements brought by oriented descriptors and the estimated local rotations. Source code and supplementary material are available at https://github.com/HpWang-whu/RoReg. Haiping Wang 0004, Yuan Liu 0025, Qingyong Hu, Bing Wang 0013, Zhen Dong 0005, Yulan Guo, Wenping Wang 0001, Bisheng Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2023 | RailSeg: Learning Local-Global Feature Aggregation With Contextual Information for Railway Point Cloud Semantic SegmentationabstractIncomplete or outdated inventories of railway infrastructures may disrupt the railway sector’s administration and maintenance of transportation infrastructure, thus posing potential threats to the safety of traffic networks. Previous studies have adopted point clouds to accelerate inventory and inspection automation procedures. However, owing to the complexity of the railway scenes, previous studies reveal an imbalance between semantic richness, segmentation accuracy, and processing efficiency. This study aims to advance our understanding by providing a deep-learning framework for railway point cloud semantic segmentation. The proposed framework, named RailSeg, encompasses point cloud downsampling, integrated local-global feature extraction, spatial context aggregation, and semantic regularization. The proposed method, validated using point clouds collected in suburban and rural scenes, generates a point-level railway furniture inventory of 11 categories and achieves competitive performance in overall accuracy and mean intersection over union. In addition, RailSeg achieves better results than the baseline for additional types of point clouds (i.e., plateau railway mobile laser scanning (MLS) point clouds, street MLS point clouds, and urban-scale photogrammetric point clouds), demonstrating the superior generalization capabilities of RailSeg. This study may contribute to the development of 3D semantic segmentation, digital railway, and intelligent transportation. Tengping Jiang, Bisheng Yang, Qinyu Zhang 0008, Xin Jin 0014, Wenjun Zeng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | WHU-Helmet: A Helmet-Based Multisensor SLAM Dataset for the Evaluation of Real-Time 3-D Mapping in Large-Scale GNSS-Denied EnvironmentsabstractReal-time 3D mapping of large-scale Global Navigation Satellite System (GNSS)-denied environments plays an important role in forest inventory management, disaster emergency response, and underground facility maintenance. Compact helmet laser scanning (HLS) systems keep the same direction as the user’s line of sight and have the advantage of “what you see is what you get”, providing a promising and efficient solution for 3D geospatial information acquisition. However, the violent motion of the helmet, the limited field of view of the laser scanner, and the repeated symmetrical geometric structures in GNSS-denied environments pose enormous challenges for the existing simultaneous localization and mapping (SLAM) algorithms. To promote the development of HLS and explore its application in large-scale GNSS-denied environments, the first large-scale HLS dataset covering multiple difficult GNSS-denied areas (e.g., forests, mountains, underground spaces) was built in this study. Besides using an additional very high accuracy fiber-optic inertial measurement unit (IMU), a novel post-processing multi-source fusion method—progressive trajectory correction (PTC)—is proposed to generate a reliable ground-truth trajectory for the benchmark, which overcomes the problems of scan matching degradation and non-rigid distortion. The accuracies of the ground truth are controlled and checked by manually surveyed feature points along the trajectory. Finally, the existing state-of-the-art SLAM methods were evaluated on the WHU-Helmet dataset, summarizing the future HLS SLAM research trends. The full dataset is available for download at: https: //github.com/kafeiyin00/WHU-HelmetDataset. Jianping Li 0004, Weitong Wu 0002, Bisheng Yang, Xianghong Zou, Yandi Yang, Xin Zhao 0026, Zhen Dong 0005 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | SE-Calib: Semantic Edge-Based LiDAR-Camera Boresight Online Calibration in Urban ScenesabstractRigorous boresight calibration between light detection and ranging (LiDAR) and the camera is crucial for geometry and optical information fusion in earth observation and robotic applications. Although boresight parameters can be obtained through pre-calibration with artificial targets, unforeseen movement of sensors during data collection can lead to significant errors in the boresight parameters. To address this issue, we propose SE-Calib, an automatic and target-free online boresight calibration method for LiDAR-Camera systems. SE-Calib firstly extracts semantic edge features from both point clouds and images simultaneously using the 3D semantic segmentation (3D-SS) and 2D semantic edge detection (2D-SED) methods. The boresight parameters are then optimized with an adaptive solver and maximizing the Soft Semantic Response Consistency Metric (SSRCM) scores iteratively. The SSRCM is designed to evaluate the coherence of cross-modular semantic edge features, and a confidence function is proposed to filter out unreliable optimization results. Experiments conducted on challenging urban datasets show an average boresight error of 0.206 degrees (2.47 pixels in reprojection error), demonstrating the effectiveness and robustness of the proposed method. Youqi Liao, Jianping Li 0004, Shuhao Kang, Guifang Zhu, Shenghai Yuan 0001, Zhen Dong 0005, Bisheng Yang |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2022 | PC2-PU: Patch Correlation and Point Correlation for Effective Point Cloud UpsamplingabstractPoint cloud upsampling is to densify a sparse point set acquired from 3D sensors, providing a denser representation for the underlying surface. Existing methods divide the input points into small patches and upsample each patch separately, however, ignoring the global spatial consistency between patches. In this paper, we present a novel method PC$^2$-PU, which explores patch-to-patch and point-to-point correlations for more effective and robust point cloud upsampling. Specifically, our network has two appealing designs: (i) We take adjacent patches as supplementary inputs to compensate the loss structure information within a single patch and introduce a Patch Correlation Module to capture the difference and similarity between patches. (ii) After augmenting each patch's geometry, we further introduce a Point Correlation Module to reveal the relationship of points inside each patch to maintain the local spatial consistency. Extensive experiments on both synthetic and real scanned datasets demonstrate that our method surpasses previous upsampling methods, particularly with the noisy inputs. The code and data are at: https://github.com/chenlongwhu/PC2-PU.git. Chen Long, Ruihui Li, Hao Wang 0057, Zhen Dong 0005, Bisheng Yang |
ACM Multimedia | 6 |
| 2022 | Full-Waveform Classification and Segmentation-Based Signal Detection of Single-Wavelength Bathymetric LiDARabstractSingle-wavelength bathymetric LiDAR (532 nm) can provide seamless meter- and submeter-scale DEMs of both the terrestrial surface and seafloor. However, the mixed terrestrial and bathymetric surfaces obtained by this sensor are challenging for full-waveform (FW) signal detection. This study addresses the issues in two FW mixed surfaces: accurate classification of terrestrial and non-terrestrial waveforms from the original waveforms without auxiliary information, and flexible detection of peaks based on a new FW theoretical model. A novel FW signal-detection model (FWSD) for single-wavelength bathymetric LiDAR is proposed without complex feature extraction and iterative procedure through waveform classification and segmentation. The raw FW are divided into 5 categories for subsequent signal detection by utilizing a convolutional neural network that merges local descriptors with contextual information. The signal detection task is then split into FW segment recognition and peak extraction using a new FW model, which integrates a leapfrog sliding window FW segmentation, an improved extreme learning machine (ELM) algorithm for FW segment recognition and a flexible signal detection framework. In order to search for the optimal initial parameters for ELM, a self-annealing particle swarm optimization (SAPSO) algorithm is introduced, and the output weight is adjusted by online sequence to improve its generalization. When combined with the Richardson–Lucy deconvolution (RLD) algorithm, FWSD can be adapted to deal with shallow water waveforms. Finally, a test demonstration with an airborne dataset shows that FWSD has higher detection efficiency and higher accuracy than a generalized Gaussian model optimized using the Levenberg–Marquardt algorithm (LM-GGM) and RLD algorithm. Xue Ji, Bisheng Yang, Yuan Wang 0035, Qiuhua Tang, Wenxue Xu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Automated 3D Road Boundary Extraction and Vectorization Using MLS Point CloudsabstractTo meet the urgent demands in a wide range of geospatial applications, such as road management, intelligent transportation systems, road safety evaluation, and traffic accident analysis, automatic and accurate extraction of 3D roads and associated geometric parameters from point clouds is receiving wide attention. In this paper, we propose an accurate 3D road boundary extraction and vectorization method to bridge the gap from unstructured mobile laser scanning (MLS) point clouds to the vector-based representation of road boundary. Firstly, we propose a supervoxel generation method to extract candidate curbs with fine border preservation and high computation efficiencies. Then the candidate curb supervoxels are recognized and clustered to produce continuous road boundary segments with a contracted distance clustering strategy. Finally, the vectorized road boundary is represented by fitting, tracking, and completion from the extracted road boundary segments, resulting in road geometric parameters including boundary location, road widths, turning radius, and slopes. The performance of the proposed method was evaluated on two large-scale datasets collected in urban and industrial areas. Comprehensive experiments reveal that the proposed method is robust to various road shapes and point densities, in terms of precision of 95.0% and recall of 91.0%, respectively. Xiaoxin Mi, Bisheng Yang, Zhen Dong 0005, Chi Chen 0002, Jianxiang Gu |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | AdaFit: Rethinking Learning-based Normal Estimation on Point CloudsabstractThis paper presents a neural network for robust normal estimation on point clouds, named AdaFit, that can deal with point clouds with noise and density variations. Existing works use a network to learn point-wise weights for weighted least squares surface fitting to estimate the normals, which has difficulty in finding accurate normals in complex regions or containing noisy points. By analyzing the step of weighted least squares surface fitting, we find that it is hard to determine the polynomial order of the fitting surface and the fitting surface is sensitive to outliers. To address these problems, we propose a simple yet effective solution that adds an additional offset prediction to improve the quality of normal estimation. Furthermore, in order to take advantage of points from different neighborhood sizes, a novel Cascaded Scale Aggregation layer is proposed to help the network predict more accurate point-wise offsets and weights. Extensive experiments demonstrate that AdaFit achieves state-of-the-art performance on both the synthetic PCPNet dataset and the real-word SceneNN dataset. The code is publicly available at https://github.com/Runsong123/AdaFit. Runsong Zhu, Yuan Liu 0025, Zhen Dong 0005, Yuan Wang 0035, Tengping Jiang, Wenping Wang 0001, Bisheng Yang |
ICCV | 7 |
| 2021 | A Coarse-to-Fine Strip Mosaicing Model for Airborne Bathymetric LiDAR DataabstractThe airborne light detection and ranging (LiDAR) bathymetry (ALB) system is an extension of the ubiquitous topographic LiDAR mapping system and has been most simply characterized as adding a green laser to the infrared laser of topo systems. Due to the low point cloud density and monotonous objects in the scene, it is difficult to mosaicing the ALB strips. Therefore, the existing airborne laser scanning strip stitching algorithm has poor performance for ALB strips. In this article, a coarse-to-fine strip mosaicing model for ALB is proposed. The framework is fast and efficient and can handle large ALB data. An improved alpha shapes algorithm can fast and accurately determine the overlap region of strip is applied. Due to different data accuracy and spatial characteristics, the water area and land area are processed separately. A weight distribution-based coarse-to-fine registration model is designed for underwater areas. The topological constraint term is added to the nonrigid iterative closest point (ICP) cost function to prevent excessive deformation caused by outliers. The implicit B-spline surface fitting algorithm using the 3L algorithm and the least-squares trend surface fitting algorithm are applied separately to assign weights for overlapping strips to solve the limitation of no control or less control. Moreover, a random sample consensus (RANSAC)-ICP registration model characterized by the normal vector and curvature is constructed for land area. Finally, the comparisons with ICP highlight the superiority of the proposed approach in flexibility and accuracy. The root-mean-square error (RMSE) is 0.12 m and the maximum error is 0.36 m. Xue Ji, Bisheng Yang, Qiuhua Tang, Wenxue Xu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | Density-Adaptive and Geometry-Aware Registration of TLS Point Clouds Based on Coherent Point DriftabstractProbabilistic registration algorithms [e.g., coherent point drift, (CPD)] provide effective solutions for point cloud alignment. However, using the original CPD algorithm for automatic registration of terrestrial laser scanner (TLS) point clouds is highly challenging because of density variations caused by scanning acquisition geometry. In this letter, we propose a new global registration method, introducing the use of the CPD framework for TLS point clouds. We first consider the measurement geometry and the intrinsic characteristics of the scene to simplify points. In addition to the Euclidean distance, we incorporate geometric information as well as structural constraints in the probabilistic model to optimize the so-called matching probability matrix. Among the structural constraints, we use a spectral graph to measure the structural similarity between matches at each iteration. The method is tested on three data sets collected by different TLS scanners. Experimental results demonstrate that the proposed method is robust to density variations and can decrease iterations effectively. The average registration errors of the three data sets are 0.05, 0.12, and 0.08 m, respectively. It is also shown that our registration framework is superior to the state-of-the-art methods in terms of both registration errors and efficiency. The experiments demonstrate the effectiveness and efficiency of the proposed probabilistic global registration. Yufu Zang, Roderik C. Lindenbergh, Bisheng Yang, Haiyan Guan |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2020 | Visual perception driven 3D building structure representation from airborne laser scanning point cloudabstractThree-dimensional (3D) building models with unambiguous roof plane geometry parameters, roof structure units, and linked topology are essential data for many applications related to human activities in urban environments. The task of 3D reconstruction from point clouds is still in the phase of development, especially to recognize and interpret roof topological structures. This paper proposes a novel visual perception-based approach to automatically decompose and reconstruct building point clouds into meaningful and simple parametric structures, while these mutual relationships between roof plane geometry and roof structure units are expressed by a hierarchical topology tree. It starts with roof plane extraction performed by a multi-label graph cut energy optimization framework, and then a roof structure graph (RSG) model is constructed to describe the roof topological geometry with common adjacency, symmetry, and convexity rules. Moreover, a progressive roof decomposition and refinement are performed, generating a hierarchical representation of 3D roof structures models. Finally, a visual plane fitted residuals or areas constraint process is adopted to generate the RSG model in different levels of details. Two airborne laser scanning (ALS) datasets with different point densities and roof styles were tested, and the performance evaluation metrics are obtained by the International Society for Photogrammetry and Remote Sensing (ISPRS), achieving correctness and accuracy in terms of 97.7% and 0.29m, respectively. The standardized assessment results demonstrate the effectiveness and robustness of the proposed approach, showing its abilities to generate a variety of structural models, even in the presence of missing data. Pingbo Hu, Bisheng Yang |
Virtual Real. Intell. Hardw. | 2 |
| 2018 | Iterative Global Similarity Points: A Robust Coarse-to-Fine Integration Solution for Pairwise 3D Point Cloud RegistrationabstractIn this paper, we propose a coarse-to-fine integration solution inspired by the classical ICP algorithm, to pairwise 3D point cloud registration with two improvements of hybrid metric spaces (e.g., BSC feature and Euclidean geometry spaces) and globally optimal correspondences matching. First, we detect the keypoints of point clouds and use the Binary Shape Context (BSC) descriptor to encode their local features. Then, we formulate the correspondence matching task as an energy function, which models the global similarity of keypoints on the hybrid spaces of BSC feature and Euclidean geometry. Next, we estimate the globally optimal correspondences through optimizing the energy function by the Kuhn-Munkres algorithm and then calculate the transformation based on the correspondences. Finally, we iteratively refine the transformation between two point clouds by conducting optimal correspondences matching and transformation calculation in a mutually reinforcing manner, to achieve the coarse-to-fine registration under an unified framework. The proposed method is evaluated and compared to several state-of-the-art methods on selected challenging datasets with repetitive, symmetric and incomplete structures. Comprehensive experiments demonstrate that the proposed IGSP algorithm obtains good performance and outperforms the state-of-the-art methods in terms of both rotation and translation errors. Yue Pan 0009, Bisheng Yang, Fuxun Liang, Zhen Dong 0005 |
3DV | 2 |
| 2016 | A polygon-based approach for matching OpenStreetMap road networks with regional transit authority dataabstractMatching road networks is an essential step for data enrichment and data quality assessment, among other processes. Conventionally, road networks from two datasets are matched using a line-based approach that checks for the similarity of properties of line segments. In this article, a polygon-based approach is proposed to match the OpenStreetMap road network with authority data. The algorithm first extracts urban blocks that are central elements of urban planning and are represented by polygons surrounded by their surrounding streets, and it then assigns road lines to edges of urban blocks by checking their topologies. In the matching process, polygons of urban blocks are matched in the first step by checking for overlapping areas. In the second step, edges of a matched urban block pair are further matched with each other. Road lines that are assigned to the same matched pair of urban block edges are then matched with each other. The computational cost is substantially reduced because the proposed approach matches polygons instead of road lines, and thus, the process of matching is accelerated. Experiments on Heidelberg and Shanghai datasets show that the proposed approach achieves good and robust matching results, with a precision higher than 96% and a F1-score better than 90%. Hongchao Fan, Bisheng Yang, Alexander Zipf, Adam Rousell |
Int. J. Geogr. Inf. Sci. | 2 |
| 2015 | Pattern-mining approach for conflating crowdsourcing road networks with POIsabstractCrowdsourcing geospatial data mainly collected by public citizens have brought about a profound transformation on data acquisition and utilization. However, the unpredictable positional accuracies, unstructured semantic descriptions, and invalid spatial relations occur to crowdsourcing geospatial data, causing difficulties for conflating heterogeneous data sets collected by different professional agencies or volunteers. We thus propose a novel pattern-mining approach to conflate crowdsourcing road networks with points of interest (POIs) geometrically and semantically. The proposed method mines the geometric patterns between road networks and POIs respectively and generates the pattern-related skeleton graphs for them. Then, corresponding points are determined between the two skeleton graphs to align POIs and road networks geometrically, and the road-related semantic data between the associated POIs and the road segments are compared to check the data quality of POIs and infer the road names of the road segments. Experimental results show the advantages of our proposed method, demonstrating a functional and promising solution for enriching POIs and road network geometrically and semantically. Bisheng Yang |
Int. J. Geogr. Inf. Sci. | 1 |
| 2014 | Polygon-based approach for extracting multilane roads from OpenStreetMap urban road networksabstractThis study proposes a novel approach for extracting multilane roads from urban road networks in OpenStreetMap (OSM) data sets as functional high-level roads, thereby allowing comparative analyses to determine the differences between this functional hierarchy and other hierarchies. OSM road networks have high levels of detail and complex structures, but they also have large numbers of duplicated lines for the same road features, which leads to difficulties and low efficiency when extracting multilane roads using conventional methods based on the analysis and operations of line segments. To overcome these deficiencies, a polygon-based method is proposed that is based on shape analysis and Gestalt theory, which treats polygons surrounded by roads as operating elements. First, shape descriptors are calculated for each polygon in networks and are used for classification. Second, candidate multilane polygons are classified as seeds based on all the polygons used as shape descriptors by a support vector machine. Finally, based on the seed polygons, a region-growing method is proposed that connects and fills the multilane features according to Gestalt theory. An experiment using OSM data from different urban networks verified the validity of the proposed method. The method achieved good and effective extraction performance, regardless of the complexity and duplication of data sets. Thus, a comparative analysis with high-level roads extracted based on road type attributes and structural analysis was performed to demonstrate the differences between the constructed road levels and other hierarchies. Hongchao Fan, Xuechen Luan, Bisheng Yang, Lin Liu 0005 |
Int. J. Geogr. Inf. Sci. | 4 |
| 2014 | Geometric-based approach for integrating VGI POIs and road networksabstractIntegrating heterogeneous spatial data is a crucial problem for geographical information systems (GIS) applications. Previous studies mainly focus on the matching of heterogeneous road networks or heterogeneous polygonal data sets. Few literatures attempt to approach the problem of integrating the point of interest (POI) from volunteered geographic information (VGI) and professional road networks from official mapping agencies. Hence, the article proposes an approach for integrating VGI POIs and professional road networks. The proposed method first generates a POI connectivity graph by mining the linear cluster patterns from POIs. Secondly, the matching nodes between the POI connectivity graph and the associated road network are fulfilled by probabilistic relaxation and refined by a vector median filtering (VMF). Finally, POIs are aligned to the road network by an affine transformation according to the matching nodes. Experiments demonstrate that the proposed method integrates both the POIs from VGI and the POIs from official mapping agencies with the associated road networks effectively and validly, providing a promising solution for enriching professional road networks by integrating VGI POIs. Bisheng Yang |
Int. J. Geogr. Inf. Sci. | 1 |
| 2013 | A probabilistic relaxation approach for matching road networksabstractGeospatial data matching is an important prerequisite for data integration, change detection and data updating. At present, crowdsourcing geospatial data are attracting considerable attention with its significant potential for timely and cost-effective updating of geospatial data and Geographical Information Science (GIS) applications. To integrate the available and up-to-date information of multi-source geospatial data, this article proposes a heuristic probabilistic relaxation road network matching method. The proposed method starts with an initial probabilistic matrix according to the dissimilarities in the shapes and then integrates the relative compatibility coefficient of neighbouring candidate pairs to iteratively update the initial probabilistic matrix until the probabilistic matrix is globally consistent. Finally, the initial 1:1 matching pairs are selected on the basis of probabilities that are calculated and refined on the basis of the structural similarity of the selected matching pairs. A process of matching is then implemented to find M:N matching pairs. Matching between OpenStreetMap network data and professional road network data shows that our method is independent of matching direction, successfully matches 1:0 (Null), 1:1 and M:N pairs, and achieves a robust matching precision of above 95%. Bisheng Yang, Xuechen Luan |
Int. J. Geogr. Inf. Sci. | 1 |
| 2013 | Semiautomated Building Facade Footprint Extraction From Mobile LiDAR Point CloudsabstractThis letter presents a novel method for automated footprint extraction of building facades from mobile LiDAR point clouds. The proposed method first generates the georeferenced feature image of a mobile LiDAR point cloud and then uses image segmentation to extract contour areas which contain facade points of buildings, points of trees, and points of other objects in the georeferenced feature image. After all the points in each contour area are extracted, a classification based on principal component analysis (PCA) method is adopted to identify building objects from point clouds extracted in contour areas. Then, all the points in a building object are segmented into different planes using the random sample consensus algorithm. For each building, points in facade planes are chosen to calculate the direction, the start point, and the end point of the facade footprints using PCA. Finally, footprints of different facades of building are refined, harmonized, and joined. Two data sets of downtown areas and one data set of a residential area captured by Optech's LYNX mobile mapping system were tested to verify the validities of the proposed method. Experimental results show that the proposed method provides a promising and valid solution for automatically extracting building facade footprints from mobile LiDAR point clouds. Bisheng Yang, Qingquan Li 0001, Jonathan Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2013 | Automated Extraction of Building Outlines From Airborne Laser Scanning Point CloudsabstractAutomatic extraction of building outlines from airborne laser scanning (ALS) point clouds has been an active topic in the field of photogrammetry, remote sensing, and computer vision. In this letter, a marked point process method is implemented to extract building outlines from ALS point clouds. First, the Gibbs energy model of building objects is defined to describe the building points. Second, the defined Gibbs energy model is sampled within the framework of reversible-jump Markov chain Monte Carlo and optimized to find an optimal energy configuration by simulated annealing. Finally, the detected building objects are refined to eliminate false detections, and the outlines of buildings are derived from the detected building objects by morphological operators. The standard data set provided by ISPRS is used to verify the validity of the proposed method. The method extracted building objects from the standard data sets with an average completeness of 87.3% and correctness of 91.57% at the pixel level, and an average completeness of 77.6% (97.3%) and correctness of 98.1% (97.9%) at the object$(\hbox{object} > 50\ \hbox{m}^{2})$level. Bisheng Yang, Wenxue Xu, Zhen Dong 0005 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2011 | Generating hierarchical strokes from urban street networks based on spatial pattern recognitionabstractStrokes are products of a higher-level aggregation of street segments that can reflect functional importance and perceptual significance that is associated with them in human spatial mental conceptualizations, which is of vital importance for network analysis, street selection, and map generalization. Street properties (e.g., street names) and angles between street segments are the two main elements used for generating street strokes according to the continuity principle of perceptual grouping into networks. However, it is difficult to automatically generate strokes with good continuity from street networks with multiple lanes such as dual carriageways or complex street junctions. This article proposes a method for generating street strokes that maintain good continuity across multiple lanes and complex street junctions. The proposed method first detects dual carriageways and complex junctions in street networks and then generates strokes according to the continuity principle of perceptual grouping. Finally, it groups the generated street strokes across the dual carriageways and complex street junctions to maintain good continuity. Moreover, the generated strokes are hierarchically ranked based on stroke length and centrality measurements. Experimental studies demonstrate the validity and effectiveness of the proposed method. The result shows that the generated street strokes maintain good continuity and reflect well the hierarchical structure of the street networks. Bisheng Yang, Xuechen Luan, Qingquan Li 0001 |
Int. J. Geogr. Inf. Sci. | 1 |
| 2008 | A Multi-parameter Approach to Automated Building Grouping and Generalization
Haowen Yan, Robert Weibel, Bisheng Yang |
GeoInformatica | 3 |
| 2008 | Variable-resolution Compression of Vector Data
Bisheng Yang, Ross Purves, Robert Weibel |
GeoInformatica | 1 |
| 2007 | The design and implementation of SPIRIT: a spatially aware search engine for information retrieval on the InternetabstractMuch of the information stored on the web contains geographical context, but current search engines treat such context in the same way as all other content. In this paper we describe the design, implementation and evaluation of a spatially aware search engine which is capable of handling queries in the form of the triplet of ⟨theme⟩⟨spatial relationship⟩⟨location⟩. The process of identifying geographic references in documents and assigning appropriate footprints to documents, to be stored together with document terms in an appropriate indexing structure allowing real‐time search, is described. Methods allowing users to query and explore results which have been relevance‐ranked in terms of both thematic and spatial relevance have been implanted and a usability study indicates that users are happy with the range of spatial relationships available and intuitively understand how to use such a search engine. Normalised precision for 38 queries, containing four types of spatial relationships, is significantly higher (p<0.001) for searches exploiting spatial information than pure text search. Ross Purves, Paul D. Clough, Christopher B. Jones, Avi Arampatzis, Bénédicte Bucher, David Finch, Gaihua Fu, Hideo Joho, Awase Khirni Syed, Subodh Vaid, Bisheng Yang |
Int. J. Geogr. Inf. Sci. | 11 |
| 2007 | Efficient transmission of vector data over the InternetabstractThis paper proposes and implements a new methodology for progressive transmission of vector lines and polygons over the Internet. The methodology generates continuous vector data through constrained remove operations of vertices on the server side, maintains consistent topology via a set of constraint rules, and restores original vector data through a reconstruction operator on the client side. The prototype system was implemented to investigate the performance of the methodology in terms of the preservation of consistent topology, transmission time, and qualities of the resulting visualizations. The method is shown to be scalable to large data sets, to produce graphically acceptable results and to maintain topology. Bisheng Yang, Ross Purves, Robert Weibel |
Int. J. Geogr. Inf. Sci. | 1 |
| 2005 | An integrated TIN and Grid method for constructing multi-resolution digital terrain modelsabstractMulti‐resolution terrain models are an efficient approach to improve the speed of three‐dimensional (3D) visualizations, especially for terrain visualization in Geographical Information Systems (GIS). As a further development to existing algorithms and models, a new model is proposed for the construction of multi‐resolution terrain models in a 3D GIS. The new model represents multi‐resolution terrains using two major methods for terrain representation: Triangulated Irregular Network (TIN) and regular grid (Grid). In this paper, first, the concepts and formal definitions of the new model are presented. Second, the methodology for constructing multi‐resolution terrain models based on the new model is proposed. Third, the error of multi‐resolution terrain models is analysed, and a set of rules is proposed to retain the important features (e.g. boundaries of man‐made objects) within the multi‐resolution terrain models. Finally, several experiments are undertaken to test the performance of the new model. The experimental results demonstrate that the new model can be applied to construct multi‐resolution terrain models with good performance in terms of time cost and maintenance of the important features. Furthermore, a comparison with previous algorithms/models shows that the speed of rendering for 3D walking/flying through has been greatly improved by applying the new model. Bisheng Yang, Wenzhong Shi, Qingquan Li 0001 |
Int. J. Geogr. Inf. Sci. | 1 |
| 2003 | An object-oriented data model for complex objects in three-dimensional geographical information systemsabstractDeveloping a three-dimensional (3D) data model for Geographic Information Systems (GIS) is an essential and complex issue. 3D modelling in GIS is becoming ever more important for the development of cyber cities and digital earth, which have recently become feasible. A competent 3D model forms an efficient foundation for 3D visualization, query and spatial analysis. As a development of the existing 3D models, this study proposes particular improvements in handling complex 3D objects. We present an object-oriented data model for handling complex 3D objects in GIS. First, the conceptual data model is developed based on the principle of object-oriented (OO) data modelling. This model is designed based on the following three basic geometric elements: node, segment and triangle. Accordingly, the abstract geometric objects are defined: including points, lines, surfaces and volumes. Second, the corresponding 3D logical model is designed based on the defined abstract objects and the relationships between them. Third, a formal representation of the 3D spatial objects is described in detail. Fourth, a prototype 3D GIS is developed based on the proposed 3D data model. Finally, we describe the results of an experimental study to reconstruct 3D objects using this 3D GIS and a comparison with the performance of other 3D data models. The proposed model is able to handle complex objects, such as complex buildings and TV towers, which is an essential functionality for building large-scale cyber cities, such as for Hong Kong. The proposed data model proves to be very efficient, particularly in visualization and rendering. The experimental results show which the data volume of the proposed model is compacted and the visualization speed for 3D objects is improved, compared with the existing models. Wenzhong Shi, Bisheng Yang, Qingquan Li 0001 |
Int. J. Geogr. Inf. Sci. | 2 |