VLDB 2026 Research / reviewers in the wild / expert
Ruisheng Wang 0001
dblp:42/5211-1
· DBLP profile ↗
44ranked-venue papers
3as first author
28since 2021 · last 2026
0000-0003-0745-5158ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 26 · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 5 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BuildingWorld: A Structured 3D Building Dataset for Urban Foundation ModelsabstractAs digital twins become central to the transformation of modern cities, accurate and structured 3D building models emerge as a key enabler of high-fidelity, updatable urban representations. These models underpin diverse applications including energy modeling, urban planning, autonomous navigation, and real-time reasoning. Despite recent advances in 3D urban modeling, most learning-based models are trained on building datasets with limited architectural diversity, which significantly undermines their generalizability across heterogeneous urban environments. To address this limitation, we present BuildingWorld, a comprehensive and structured 3D building dataset designed to bridge the gap in stylistic diversity. It encompasses buildings from geographically and architecturally diverse regions—including North America, Europe, Asia, Africa, and Oceania—offering a globally representative dataset for urban-scale foundation modeling and analysis. Specifically, BuildingWorld provides about Five million LOD2 building models collected from diverse sources, accompanied by both real and simulated airborne LiDAR point clouds. This enables comprehensive research on 3D reconstruction, building detection and segmentation, as well as roof structure segmentation. Cyber City, a virtual city model, is introduced to enable the generation of unlimited training data with customized and structurally diverse point cloud distributions. Furthermore, we provide standardized evaluation metrics tailored for building reconstruction, aiming to facilitate the training, evaluation, and comparison of large-scale vision models and foundation models in structured 3D urban environments Shangfeng Huang, Ruisheng Wang 0001 |
AAAI | 2 |
| 2026 | PLAN: Fast and Approximate Gaussian Kernel Density Visualization in Road Networks
Tsz Nam Chan, Hongwei Ye, Bojian Zhu, Leong Hou U, Dingming Wu 0001, Ruisheng Wang 0001, Joshua Zhexue Huang |
ICDE | 6 |
| 2026 | DNA: A Distribution-and-Aggregation Solution for Spatiotemporal K-Function-Based Analysis
Tsz Nam Chan, Bojian Zhu, Dingming Wu 0001, Renchi Yang, Ruisheng Wang 0001 |
ICDE | 5 |
| 2025 | EdgeDiff: Edge-aware Diffusion Network for Building Reconstruction from Point CloudsabstractBuilding reconstruction is a challenging problem at the intersection of computer vision, photogrammetry and computer graphics. 3D wireframe presents a compelling representation for building modeling through its compact structure. Existing wireframe reconstruction methods employing vertex detection and edge regression have achieved promising results. In this paper, we develop an Edge-aware Diffusion network, dubbed EdgeDiff. As a novel paradigm for wireframe reconstruction, the EdgeDiff generates wireframe models from noise using a conditional diffusion model. During the training process, the ground truth wireframes firstly are formulated as a set of parameterized edges and then diffused into a random noise distribution. EdgeDiff learns both the noise reversal process and the network structure simultaneously. During inference, EdgeDiff iteratively refines the generated edge distribution using the denoising diffusion implicit model, enabling flexible single- or multi-step denoising and dynamic adaptation to buildings of varying complexity. Additionally, given the unique structure of wireframes, we introduce an edge attention module to extract point-wise attention from point features, using it as auxiliary information to facilitate learning of edge cues and guide the network toward improved edge awareness. Extensive experiments on the real-world Building3D dataset demonstrate that our approach achieves state-of-the-art performance. Yujun Liu 0005, Ruisheng Wang 0001, Shangfeng Huang, Guo-Rong Cai |
CVPR | 2 |
| 2025 | Complementary Information Guided Occupancy Prediction via Multi-Level Representation FusionabstractCamera-based occupancy prediction is a main-stream approach for 3D perception in autonomous driving, aiming to infer complete 3D scene geometry and semantics from 2D images. Almost existing methods focus on improving performance through structural modifications, such as lightweight backbones and complex cascaded frameworks, with good yet limited performance. Few studies explore from the perspective of representation fusion, leaving the rich diversity of features in 2D images underutilized. Motivated by this, we propose CIGOcc, a two-stage occupancy prediction framework based on multi-level representation fusion. CIGOcc extracts segmentation, graphics, and depth features from an input image and introduces a deformable multi-level fusion mechanism to fuse these three multi-level features. Additionally, CIGOcc incorporates knowledge distilled from SAM to further enhance prediction accuracy. Without increasing training costs, CIGOcc achieves state-of-the-art performance on the SemanticKITTI benchmark. The code is provided in the supplementary material and will be released project page. Rongtao Xu, Jinzhou Lin 0001, Jialei Zhou, Jiahua Dong 0001, Changwei Wang 0001, Ruisheng Wang 0001, Li Guo 0004, Shibiao Xu, Xiaodan Liang |
ICRA | 6 |
| 2025 | A Fast and Accurate Block Compression Solution for Spatiotemporal Kernel Density VisualizationabstractSpatiotemporal Kernel Density Visualization (STKDV) has been widely used across various domains in geospatial analysis, e.g., urban planning, traffic/traffic accident hotspot analysis, crime hotspot analysis, and disease spread modeling.However, STKDV is a computationally expensive tool, which has been complained by many domain experts.Although many recent solutions, including the sliding-window-based solution (SWS) and the prefix-matrix-based solution (PREFIX), have been proposed for improving the efficiency of generating an exact STKDV, these solutions still cannot be scalable to handle large-scale location datasets.To tackle this efficiency issue, we propose the pioneering block compression solution, called COMP, which can compress (or represent) a location dataset by a small amount of blocks.By combining COMP with the existing exact solutions, i.e., SWS and PREFIX, we show that COMP SWS and COMP PREFIX can generate approximate STKDV with an 𝜖-absolute error guarantee based on properly tuning the block size.Experimental results on four large-scale location datasets (up to 6.782 million data points) also verify that COMP SWS and COMP PREFIX can achieve speedups of 4.1x to 677.16x and 1.45x to 143.52x compared with SWS and PREFIX, respectively, without degrading the visualization results.The code of this paper can be found in https://github.com/YovelaZ/COMP. Tsz Nam Chan, Leong Hou U, Dingming Wu 0001, Wei Tu 0001, Ruisheng Wang 0001, Joshua Zhexue Huang |
KDD (2) | 6 |
| 2025 | A Dual-Branch Deep Learning Framework at the Grid Scale for Individual Tree SegmentationabstractIndividual tree segmentation from point clouds is essential for diverse forest applications. A dual-branch segmentation deep learning network operating at the grid scale was proposed, which includes the semantic segmentation branch for partitioning point clouds of tree trunks and the instance segmentation branch for individual trunk extraction. Meanwhile, the network analyzes input forest points at the grid scale instead of pointwise processing to preserve local geometric information of the forest points while reducing computational load. After extraction of each tree trunk in the understory layer using our network, a hierarchical k-nearest neighbors algorithm based on the extracted trunk parts was employed to accomplish individual tree segmentation. For the forest plots, our proposed approach achieves precision, recall,${F}1$-score, and mean intersection over union (MIoU) of 89.66%, 89.13%, 89.40%, and 90.84%, respectively. These results represent a significant improvement in accuracy and rapid execution capability compared to prior methods. Ze Ding, Huaiqing Zhang, Ruisheng Wang 0001, Li Zhang 0057, Hanxiao Jiang 0004, Ting Yun |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2025 | HSPFormer: Hierarchical Spatial Perception Transformer for Semantic SegmentationabstractSemantic perception in driving scenarios plays a crucial role in intelligent transportation systems. However, existing Transformer-based semantic segmentation methods often do not fully exploit their potential in understanding driving scene dynamically. These methods typically lack spatial reasoning, failing to effectively correlate image pixels with their spatial positions, leading to attention drift. To address this issue, we propose a novel architecture, the Hierarchical Spatial Perception Transformer (HSPFormer), which integrates monocular depth estimation and semantic segmentation into a unified framework for the first time. We introduce the Spatial Depth Perception Auxiliary Network (SDPNet), a framework for multiscale feature extraction and multilayer depth map prediction to establish hierarchical spatial coherence. Additionally, we design the Hierarchical Pyramid Transformer Network (HPTNet), which uses depth estimation as learnable position embeddings to form spatially correlated semantic representations and generate global contextual information. Experiments on benchmark datasets such as KITTI-360, Cityscapes, and NYU Depth V2, demonstrate that HSPFormer outperforms several state-of-the-art networks, and achieves promising performance with 66.82% top-1 mIoU on KITTI-360, 83.8% mIoU on Cityscapes, and 57.7% mIoU on NYU Depth V2, respectively. The code will be made publicly available athttps://github.com/SY-Ch/HSPFormer. Siyu Chen 0004, Ting Han 0001, Changshe Zhang, Jinhe Su, Ruisheng Wang 0001, Yiping Chen 0002, Zongyue Wang, Guo-Rong Cai |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | TUC-Net: A Point Cloud Segmentation Network Based on Neighborhood Feature Perception Aggregation for Tunnels Under ConstructionabstractThe high risks in tunnel construction underscore the critical necessity for intelligent tunnel construction. Unmanned tunnel data collection is vital for intelligent construction; additionally, semantic segmentation aids in understanding the environment. However, complex tunnel terrains are challenging for three-dimensional (3D) laser scanning, and diverse interior structures and nontunnel elements complicate accurate segmentation by subsequent networks. Therefore, this paper proposed a tunnel mobile 3D mapping system (TMMS) for complex terrain in construction tunnels using a quadruped robot and simultaneous localization and mapping (SLAM). Additionally, a deep learning-based semantic segmentation network (TUC-Net) is proposed for analysing 3D point clouds in tunnels under construction. The research presented the neighbourhood feature perception enhancement (NFPE) module to enhance the representation of local features, introduce a self-attention (SA) module and improve the loss function to improve network accuracy. The NFPE module enhances feature aggregation for unstructured objects, and the SA module improves the learning of global features, critical for tunnel point cloud segmentation. The TMMS is used to collect point cloud data from tunnels under construction, leading to the creation of the tunnels under construction point clouds (TUCPC) dataset for training and evaluating the TUC-Net network. Compared to other 3D point cloud semantic segmentation methods, the proposed method demonstrated superior performance, achieving an overall accuracy (OA) of 99.45% and a mean intersection over union (mIoU) of 94.06%, surpassing that of other methods by at least 4.41%. In addition, ablation studies were also performed on the NFPE and SA modules to validate their efficacy. Xing Zhang 0003, Xinglin Huang, Kaipeng Hong, Qingquan Li 0001, Ruisheng Wang 0001, Baoding Zhou |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | Handcrafted Local Feature Descriptor-Based Point Cloud Registration and Its Applications: A ReviewabstractPoint cloud registration serves as a fundamental problem across multiple fields including computer vision, computer graphics, and remote sensing. While local feature descriptors (LFDs) have long been established as a cornerstone for point cloud registration and the LFD-based approach has been extensively studied, the field has witnessed significant advancements in recent years. Despite these developments, the research community lacks a systematic review to consolidate these contributions, leaving many researchers unaware of recent progress in LFD-based registration. To address this gap, we present a comprehensive review that critically examines both state-of-the-art and widely referenced methods across all subtasks of LFD-based registration. Our work provides: (1) an extensive survey of existing methodologies, (2) in-depth analysis of their respective strengths and limitations, (3) insightful observations and practical recommendations, and (4) a thorough summary of relevant applications and publicly available datasets. This systematic overview offers valuable guidance for researchers pursuing future investigations in this domain. Wuyong Tao, Ruisheng Wang 0001, Xianghong Hua, Jingbin Liu, Xijiang Chen, Yufu Zang, Dong Chen 0009, Dong Xu 0011 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | PBWR: Parametric-Building-Wireframe Reconstruction from Aerial LiDAR Point CloudsabstractIn this paper, we present an end-to-end 3D-building-wireframe reconstruction method to regress edges directly from aerial light-detection-and-ranging (LiDAR) point clouds. Our method, named parametric-building-wireframe reconstruction (PBWR), takes aerial LiDAR point clouds and initial edge entities as input and fully uses the self-attention mechanism of transformers to regress edge parameters without any intermediate steps such as corner prediction. We propose an edge non-maximum suppression (E-NMS) module based on edge similarity to remove redundant edges. Additionally, a dedicated edge loss function is utilized to guide the PBWR in regressing edges parameters when the simple use of the edge distance loss is not suitable. In our experiments, our proposed method demonstrated state-of-the-art results on the Building3D dataset, achieving an improvement of approximately 36% in Entry-level dataset edge accuracy and around a 42% improvement in the Tallinn dataset. Shangfeng Huang, Ruisheng Wang 0001 |
CVPR | 2 |
| 2024 | Efficient Roof Vertex Clustering for Wireframe Simplification Based on the Extended Multiclass Twin Support Vector MachineabstractThis study introduces an efficient approach for clustering roof wireframe vertices within the realm of model simplification based on a multiclass twin support vector machine (TWSVM) framework. The proposed method first assigns a dynamic label to each point of the input point cloud, and it then iteratively identifies k cluster center 3-D lines by maintaining short distances between wireframe candidate vertices sharing the same corner. In addition, it ensures that these wireframe candidates from one corner are distanced from the wireframe vertices from the other corners in a drafting roof dataset. This study extends the multiclass TWSVM to tackle the clustering problem of roof wireframe vertices, thus facilitating model simplification. Remarkably, this problem can be solved using a straightforward and efficient iterative algorithm. The results demonstrate that our proposed method achieves more accurate clustering results on 20 out of 24 roof wireframe vertex datasets compared with other relevant methods. Furthermore, the proposed method can efficiently and accurately extract the majority of vertices from roof wireframes in real-world Building3D dataset. Shangfeng Huang, Ruisheng Wang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | Self-Supervised Pre-Training for 3-D Roof Reconstruction on LiDAR DataabstractReconstructing building roofs from light detection and ranging (LiDAR) point clouds from aerial perspectives is significantly important in photogrammetry domains. This letter proposes a novel approach for three-dimensional (3D) real-world building roof reconstruction in Estonia, employing a two-stage self-supervised pre-training architecture to transform 3D roof point clouds into wireframe models. We utilize a self-supervised pre-training framework that incorporates a purpose-designed and efficient self-attention mechanism to generate point-wise features. Subsequently, we develop modules for corner detection and edge prediction to classify and regress the coordinates of corner points and determine optimal edge selections, respectively, to construct the final wireframe model. The effectiveness of our approach is evaluated on real-world roof datasets, achieving corner and edge precision accuracies of 83% and 78%, respectively. In addition, fine-tuning our self-supervised pre-training method with varying ratios of labeled data, particularly with only 50% partially labeled data, attains superior performance, achieving 84% and 85% corner and edge precision, respectively. Shangfeng Huang, Ruisheng Wang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | Easy-Net: A Lightweight Building Extraction Network Based on Building FeaturesabstractThe efficient, accurate, and automatic extraction of buildings from remote sensing imagery is a key task in the intelligent extraction of remote sensing information owing to its importance in applications including urban planning, change detection, and unmanned aerial vehicle (UAV) navigation. However, the fast and accurate extraction of buildings from remote sensing images remains difficult owing to the complex, variable nature of geographic information, and variable external appearances of buildings. This is because many existing building extraction networks fail to incorporate building features into their design. Also, generally, simple lightweight networks do not accurately identify buildings, while large complex networks have high operational costs. Therefore, in this article, we proposed a simple, effective feature fusion strategy based on the building features extracted by the lightweight backbone network; also we have improved the feature fusion performance by combining the advantages of a convolutional neural network (CNN) and transformer; and presented the lightweight building extraction network called Easy-Net. We conducted experiments comparing Easy-Net with existing high-performing networks on the public dataset WHU and self-made datasets; results showed the efficiency and accuracy of our method in the task of building extraction from remote sensing images. Thus, Easy-Net was found to be a promising alternative to existing building extraction networks. Code has been released at: github.com/teddy132/EasyNet_for_building_extraction. Huaigang Huang, Ruisheng Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Self-Supervised Pretraining Framework for Extracting Global Structures From Building Point Clouds via CompletionabstractThe exterior structural information of buildings are crucial for advancing smart city initiatives and reconstructing 3-D edifices. However, practical obstacles, such as sparse or incomplete building point clouds—stemming from various scanning angles or sensor limitations—present significant challenges. To mitigate the high costs and labor demands associated with data labeling, we introduce an innovative pretraining framework with self-supervised learning (SSL) that incorporates a modified point cloud completion (PCC) subnetwork to extract building structures. Specifically, the modified PCC subnetwork completes the original partial building point clouds by capturing both fine-grained and high-level semantic information of 3-D shapes. Following this, self-supervised feature extractor using a masked autoencoder (MAE) and a multiscale feature mechanism generates pointwise features from the completed building point clouds. We evaluate the effectiveness of the proposed integrated framework in extracting global structures using both established wireframe construction methods and our newly proposed edge point identification that incorporates a novel edge point regression loss. Extensive experimental results demonstrate that our modified PCC network reaches a 93.5% convergence rate that is higher than the results from competing methods. Our self-supervised pretraining framework extracts more accurate global structures with better loss convergence than traditional edge point identification loss designs. Finally, our combined framework improves performance for subsequent processes (such as wireframe construction and edge point identification) when using completed datasets instead of the original partial datasets. Ruisheng Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Exploring Intrinsic Discrimination and Consistency for Weakly Supervised Object LocalizationabstractWeakly supervised object localization (WSOL) is a challenging and promising task that aims to localize objects solely based on the supervision of image category labels. In the absence of annotated bounding boxes, WSOL methods must employ the intrinsic properties of the image classification task pipeline to generate object localizations. In this work, we propose a WSOL method for exploring the Intrinsic Discrimination and Consistency in the image classification task pipeline, and call it as IDC. First, we develop a Triplet Metrics Based Foreground Modeling (TMFM) framework to directly predict object foreground regions using intrinsic discrimination. Unlike Class Activation Map (CAM) based methods that also rely on intrinsic discrimination, our TMFM framework alleviates the problem of only focusing on the most discriminative parts by optimizing foreground and background regions synergistically. Second, we design a Dual Geometric Transformation Consistency Constraints (DGTC2) training strategy to introduce additional supervision and regularization constraints for WSOL by leveraging intrinsic geometric transformation consistency. The proposed pixel-wise and object-wise consistency constraint losses cost-effectively provide spontaneous supervision for WSOL. Extensive experiments show that our IDC method achieves significant and consistent performance gains compared to existing state-of-the-art WSOL approaches. Code is available at: https://github.com/vignywang/IDC. Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Ruisheng Wang 0001, Xiaopeng Zhang 0001 |
IEEE Trans. Image Process. | 5 |
| 2023 | Building3D: An Urban-Scale Dataset and Benchmarks for Learning Roof Structures from Point CloudsabstractUrban modeling from LiDAR point clouds is an important topic in computer vision, computer graphics, photogrammetry and remote sensing. 3D city models have found a wide range of applications in smart cities, autonomous navigation, urban planning and mapping etc. However, existing datasets for 3D modeling mainly focus on common objects such as furniture or cars. Lack of building datasets has become a major obstacle for applying deep learning technology to specific domains such as urban modeling. In this paper, we present an urban-scale dataset consisting of more than 160 thousands buildings along with corresponding point clouds, mesh and wireframe models, covering 16 cities in Estonia about 998 Km2. We extensively evaluate performance of state-of-the-art algorithms including handcrafted and deep feature based methods. Experimental results indicate that Building3D has challenges of high intra-class variance, data imbalance and large-scale noises. The Building3D is the first and largest urban-scale building modeling benchmark, allowing a comparison of supervised and self-supervised learning methods. We believe that our Building3D will facilitate future research on urban modeling, aerial path planning, mesh simplification, and semantic/part segmentation etc. Ruisheng Wang 0001, Shangfeng Huang |
ICCV | 1 |
| 2023 | MCTNet: Multiscale Cross-Attention-Based Transformer Network for Semantic Segmentation of Large-Scale Point CloudabstractIn this work, we implement a hybrid method to utilize sufficient information by aggregating both fine-grained and globally contextual features for point cloud semantic segmentation with a hierarchical network. By surpassing the defects of convolution operation mainly for extracting low-level features, we combine higher-level cross-attention based Transformer to investigate the importance of long-range relations together with position embedding for multiscale feature representation. Specifically, adding a learnable token to the feature sequence of a layer, a Transformer encoder is first implemented with limited scope to embed these features. Furthermore, instead of performing all-to-all attention, we merely fuse tokens spanning various scales. To improve efficiency, we propose a simple yet efficient token-fusing architecture based on cross-attention, in which the computation of attention maps can be restricted within linear time by only using a token to calculate the query. The cross-attention module can be efficiently aggregated in a multiscale network to further enlarge the scope of the receptive field for attention. Experiments show that our MCTNet achieves promising results on three largest point cloud datasets, DALES, DublinCity and S3DIS datasets. For the DALES benchmark dataset, MCTNet improves the mean intersection-over-union (mIoU) to 83.3% and the overall accuracy (OA) to 98.3%, which outperforms other existing baselines. We also perform abundant ablation studies on various attention and normalization modules and discuss the effect of parameters to validate the descriptive power of cross-attention module and provide an understanding of how long-range dependency can be used to learn fair and unbiased features. Ruisheng Wang 0001, Wenchao Guo, Alex Hayman Ng, Wenfeng Bai |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Toward Accurate and Efficient Road Extraction by Leveraging the Characteristics of Road ShapesabstractAutomatically extracting roads from very high resolution (VHR) remote sensing images is of great importance in a wide range of remote sensing applications. However, complex shapes of roads (i.e., long, geometrically deformed, and thin) always affected the extraction accuracy, which is one of the challenges of road extraction. Based on the insight into road shape characteristics, we propose a novel road shape aware network (RSANet) to achieve efficient and accurate road extraction. First, we introduce the Efficient Strip Transformer Module (ESTM) to efficiently capture the global context to model the long-distance dependence required by the long roads. Second, we design a Geometric Deformation Estimation Module (GDEM) to adaptively extract the context from the shape deformation caused by shooting roads from different perspectives. Third, we provide a simple but effective Road Edge Focal Loss (REF loss) to make the network focus on optimizing the pixels around the road to alleviate the unbalanced distribution of foreground and background pixels caused by the roads being too thin. Finally, we conduct extensive evaluations on public datasets to verify the effectiveness of RSANet and each of the proposed components. Experiments validate that our RSANet outperforms state-of-the-art methods for road extraction in remote sensing images. Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Ruisheng Wang 0001, Jiguang Zhang, Xiaopeng Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Self-Supervised Learning for 3-D Point Clouds Based on a Masked Linear AutoencoderabstractMotivated by the success of a masked autoencoder in three-dimensional (3D) point cloud-based learning, this study proposes an innovative framework for self-supervised learning on 3D point clouds with linear complexity. In the proposed framework, every input point cloud is divided into multiple point patches, which are randomly masked at different ratios. Then, unmasked point patches are then fed to an improved Transformer model, which uses an advanced linear self-attention mechanism autoencoder to learn high-level features. The pre-training objective is to recover the masked patches under the guidance of the unmasked point patches’ features obtained by the designed Transformer. Further, a linear self-attention mechanism is designed to use three projection matrices to decompose the original scaled dot-product attention into smaller parts, using the properties of low-rank and linear decomposition to reduce the time complexity from quadratic to linear. The results of extensive experiments demonstrate that the proposed pre-trained model can achieve high accuracy of 93.6% and 84.77% on the ModelNet40 and ScanObjectNN datasets, respectively, at a masking ratio of 40%. In addition, the results show that the proposed method, which uses a linear self-attention mechanism, can enhance the computational efficiency by significantly reducing inference time and minimizing the storage memory requirements forQ,K, andV(Query, Key, and Value) matrices compared with the existing methods. Finally, the results indicate that the proposed method can achieve state-of-the-art performance on the classification ModelNet40 dataset. Ruisheng Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Attention-Based Multi-Scale Graph Convolution for Point Cloud Semantic SegmentationabstractGeometric deep learning on non-Euclidean data, particularly point clouds, has in recent years been getting a lot of attention and success. Despite this success, the relationship between points in a delineated subgraph have still not been fully explored - like the under exploration of correlations between inter-class and/or intra-class point connections - leading to suboptimal point cloud segmentations. Thus, this study proposes a scale-invariant graph attention convolution network that has the capacity to dynamically adapt to sub-graph structures at varying scales, and reduce noisy local feature propagation due to mixed object class neighborhoods. The efficacy of the proposed framework is evaluated using the Toronto3D benchmark dataset and attained an mI-oU of 61.2%, outperforming all the existing methods it was compared to. Perpetual Hope Akwensi, Ruisheng Wang 0001 |
IGARSS | 2 |
| 2022 | BERT-Enhanced with Context-Aware Embedding for Instance Segmentation in 3D Point CloudsabstractInspired by the successful implementation of transformer network in the Natural Language Processing (NLP), we propose a novel Bidirectional Encoder Representations from Transformers (BERT)-based point cloud segmentation method. Specifically, the whole point cloud is scanned by multiple overlapping windows. We made the first attempt ever to input each window-point-cloud into the BERT model which outputs points' semantic labels and high-dimensional context-aware point embeddings. In the process of training, the Kullback-Leibler (KL)-Divergence-based clustering loss is utilized to optimize the network's parameters by calculating similarity matrices between the point embeddings and the predicted semantic labels. The final instance labels can be obtained by softmax function on these optimized point embeddings. By evaluating on the Stanford 3D Indoor Scene (S3DIS) dataset, our proposed method has reached a micro-mean accuracy (mAcc) of 87.3% on the semantic segmentation task and an Average Precision (mAP) on the instance segmentation task. The results on both tasks have surpassed the traditional point cloud segmentation models. Hailun Yan, Ruisheng Wang 0001 |
IGARSS | 3 |
| 2022 | A Lightweight Network for Building Extraction From Remote Sensing ImagesabstractBuilding extraction is a fundamental research topic in remote sensing image interpretation. Convolutional neural network (CNN)-based building extraction algorithms have achieved high accuracy but require a large account of parameters and calculations, which hinders the practical application of these algorithms. To address the challenge, we propose a lightweight network (RSR-Net) for building extraction from remote sensing images. The network consists of three basic units with only a few parameters, and uses the idea of the fusion of shallow features and deep features, which is proposed by U-Net. Before features fusion, the squeeze-and-excitation (SE) module in RSR-Net assigned channel weights to these deep and shallow features. This operation can effectively reduce the influence of noise caused by shallow features in feature fusion, so as to improve the performance of the model. We estimated our network on datasets and achieved 88.32%, 71.58%, and 77.07% intersection-over-union (IoU) on datasets of aerial image and satellite image in Wuhan University (WHU) dataset, and the self-made building dataset of Guangzhou University Town, with only 2.81 M parameters and 6.91 G floating point operations (FLOPs). In addition, we propose a strategy combining target and background prediction, which makes RSR-Net achieve 0.37% improvement in IoU on WHU aerial image dataset. The effectiveness of RSR-Net is high. It showed that the proposed network is light and fast for the application of convolution neural network algorithm in practice. Huaigang Huang, Yiping Chen 0002, Ruisheng Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Building Instance Mapping From ALS Point Clouds Aided by Polygonal MapsabstractBuilding region extraction from ALS point clouds has been widely studied, whereas instance-level building mapping has been overlooked and remains unsolved. In this study, we present a method to extract individual buildings from ALS point clouds with the help of widely accessible polygonal footprints. The key idea is to merge roof segments to a set of building candidates, from which correct instances are selected by finding optimal matches between polygonal footprints and building candidates. The method has three steps: roof segmentation, building candidate generation, and instance-polygon matching. The method is tested on two large-scale scenes of different building types and can generally achieve high instance-level building mapping accuracy (around 90%) when there are large positioning errors (6.0 m) among polygons. Future work will focus on classification errors in preprocessing, shape inconsistency between point clouds and polygons, and building footprint delineation and updating in postprocessing. Shaobo Xia, Sheng Xu 0003, Ruisheng Wang 0001, Jonathan Li 0001, Guanghui Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | SSA3D: Semantic Segmentation Assisted One-Stage Three-Dimensional Vehicle Object DetectionabstractOne-stage 3D object detection using mobile light detection and ranging (LiDAR) has developed rapidly in recent years. Specifically, one-stage methods have attracted attention because of their high efficiency and light weight compared with two-stage methods. Inspired by this, we present the semantic segmentation assisted one-stage three-dimensional vehicle object detection (SSA3D), a network for the rapid detection of objects that keeps the advantages of the semantic segmentation module in the two-stage methods without increasing redundant computational load. First, we modified the sampling of the farthest point to improve the quality of the sampling points. This helps to reduce sampling outlier points and bad points that are difficult to perceive in the spatial structure information surrounding the point. Second, a neighbor attention group module is devoted to selectively add extra weight to neighbor points because of the different importance of neighbor points for the corresponding sampling point. Correctly increasing the weight is helpful to obtain richer spatial structure information. Finally, a delicate box generation module is included as a voted center point layer based on the generalized Hoff vote method and an anchor-free regression. We used the feature aggregation module as the backbone and the feature propagation module as the auxiliary network to achieve efficiency. At the same time, the auxiliary network retains the ability to extract point-wise features from the state-of-the-art semantic segmentation network. In experiments, we evaluated and tested the SSA3D on a common KITTI dataset and achieved improved performance in the class of car accuracy. Shangfeng Huang, Guo-Rong Cai, Zongyue Wang, Qiming Xia, Ruisheng Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | Robust Lane Extraction From MLS Point Clouds Towards HD Maps Especially in Curve RoadabstractThis article presents a semi-automated method to extract the lane features along the curved roads from mobile laser scanning (MLS) point clouds. The proposed method consists of four steps. After data pre-processing, a road edge detection algorithm is performed to distinguish road curbs and extract road surfaces. Then, textual and directional road markings such as arrows, symbols, and words, to inform drivers in necessary cases, are detected by intensity thresholding and conditional Euclidean clustering algorithms. Furthermore, lane markings are extracted by local intensity analysis and distance thresholding methods according to road design standards, because they are more regular along the road. Finally, centerline points on lanes are estimated based on the coordinates of extracted lane markings. Our method shows strong feasibility and robustness when creating high-definition (HD) maps from MLS data, by increasing the number of blocks in the curve and the distance threshold control in curved lane centerline extraction. Quantitative evaluations show that the average recall, precision, and F1-score obtained from four datasets for road marking extraction are 93.87%, 93.76%, and 93.73%, respectively. The generated lane centerlines are evaluated by overlaying them on manually labeled reference buffers from 4 cm resolution orthoimagery. The comparative study indicates that the proposed methods can achieve higher accuracy and robustness than most state-of-the-art methods. Chengming Ye, He Zhao 0007, Lingfei Ma, Han Jiang 0005, Hongfu Li, Ruisheng Wang 0001, Michael A. Chapman, José Marcato Junior, Jonathan Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2021 | DRB-GAN: A Dynamic ResBlock Generative Adversarial Network for Artistic Style TransferabstractThe paper proposes a Dynamic ResBlock Generative Adversarial Network (DRB-GAN) for artistic style transfer. The style code is modeled as the shared parameters for Dynamic ResBlocks connecting both the style encoding network and the style transfer network. In the style encoding network, a style class-aware attention mechanism is used to attend the style feature representation for generating the style codes. In the style transfer network, multiple Dynamic ResBlocks are designed to integrate the style code and the extracted CNN semantic feature and then feed into the spatial window Layer-Instance Normalization (SW-LIN) decoder, which enables high-quality synthetic images with artistic style transfer. Moreover, the style collection conditional discriminator is designed to equip our DRB-GAN model with abilities for both arbitrary style transfer and collection style transfer during the training stage. No matter for arbitrary style transfer or collection style transfer, extensive experiments strongly demonstrate that our proposed DRB-GAN outperforms state-of-the-art methods and exhibits its superior performance in terms of visual quality and efficiency. Our source code is available at https://github.com/xuwenju123/DRB-GAN. Wenju Xu, Chengjiang Long, Ruisheng Wang 0001, Guanghui Wang 0001 |
ICCV | 3 |
| 2021 | Plane Segmentation Based on the Optimal-Vector-Field in LiDAR Point CloudsabstractOne key challenge in the point cloud segmentation is the detection and split of overlapping regions between different planes. The existing methods depend on the similarity and the dissimilarity in neighbor regions without a global constraint, which brings the 'over-' and 'under-' segmentation in the results. Hence, this paper presents a pipeline of the accurate plane segmentation for point clouds to address the shortcoming in the local optimization. There are two phases included in the proposed segmentation process. One is a local phase to calculate connectivity scores between different planes based on local variations of surface normals. In this phase, a new optimal-vector-field is formulated to detect the plane intersections. The optimal-vector-field is large in magnitude at plane intersections and vanishing at other regions. The other one is a global phase to smooth local segmentation cues to mimic leading eigenvector computation in the graph-cut. Evaluation of two datasets shows that the achieved precision and recall is 94.50 percent and 90.81 percent on the collected mobile LiDAR data and obtains an average accuracy of 75.4 percent on an open benchmark, which outperforms the state-of-the-art methods in terms of completeness and correctness. Sheng Xu 0003, Ruisheng Wang 0001, Ruigang Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | An Optimal Hierarchical Clustering Approach to Mobile LiDAR Point CloudsabstractThis paper aims to propose a new optimal hierarchical clustering approach to 3D mobile light detection and ranging (LiDAR) point clouds. The hierarchical clustering is performed on unorganized point clouds based on a proximity matrix that consists of a distance term and a direction term. In the dissimilarity calculation of two clusters, a pair of points from each of two clusters is selected, respectively, and Euclidean distances between the points are employed to define the distance term. The direction term is obtained by the differences of normal vectors at chosen points. The main contribution is that the cluster combination in the hierarchical clustering is optimized by a point-based graph model. The cluster combination is formulated as a problem of matching, optimized by finding the minimum-cost perfect matching in a bipartite graph. The results show that the proposed hierarchical clustering method succeeds in segmenting object from point clouds without any human-computer interaction and outperforms the state-of-the-art segmentation approaches in terms of completeness and correctness. Sheng Xu 0003, Ruisheng Wang 0001, Han Zheng 0002 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2019 | Incremental Labelling of Voronoi Vertices for Shape ReconstructionabstractAbstract We present an incremental Voronoi vertex labelling algorithm for approximating contours, medial axes and dominant points (high curvature points) from 2D point sets. Though there exist many number of algorithms for reconstructing curves, medial axes or dominant points, a unified framework capable of approximating all the three in one place from points is missing in the literature. Our algorithm estimates the normals at each sample point through poles (farthest Voronoi vertices of a sample point) and uses the estimated normals and the corresponding tangents to determine the spatial locations (inner or outer) of the Voronoi vertices with respect to the original curve. The vertex classification helps to construct a piece‐wise linear approximation to the object boundary. We provide a theoretical analysis of the algorithm for points non‐uniformly (ε‐sampling) sampled from simple, closed, concave and smooth curves. The proposed framework has been thoroughly evaluated for its usefulness using various test data. Results indicate that even sparsely and non‐uniformly sampled curves with outliers or collection of curves are faithfully reconstructed by the proposed algorithm. Jiju Poovvancheri, Amal Dev Parakkat, Andrea Tagliasacchi, Ruisheng Wang 0001, M. Ramanathan 0001 |
Comput. Graph. Forum | 4 |
| 2019 | A Novel Framework for 2.5-D Building Contouring From Large-Scale Residential ScenesabstractThis paper introduces a novel methodology for residential building contouring from large-scale airborne point clouds. Unlike other methods that handle linearization and regularization of the linear primitives separately by imposing rigid constraints, we propose an optimization-based linearization and global regularization to form accurate, topologically error-free, and lightweight polygons. To this end, we enhance the classic density-based spatial clustering of applications with noise algorithm to segment individual building entities at the instance level. The initial contours of each individual building are then delineated and further decomposed by a novel topologically aware propagation process and a global optimization technique. The decomposed linear primitives are fed into the global regularization step, from which the regular shapes are learned and enforced hierarchically by imposing constraints, such as parallelism, homogeneity, orthogonality, and collinearity. Based on the concept of hybrid representation, the regularized and unaltered linear primitives are jointly connected in an esthetic way. Various experiments using representative buildings and large-scale residential scenes from the Dutch AHN3 data set have shown that the proposed methodology generates meaningful building contouring representation in terms of accuracy, compactness, topology, and levels of detail abstraction while being robust and scalable. Jianli Du, Dong Chen 0009, Ruisheng Wang 0001, Jiju Poovvancheri, P. Takis Mathiopoulos, Lei Xie 0010, Ting Yun |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2019 | Semiautomatic Construction of 2-D Façade Footprints From Mobile LiDAR DataabstractAlthough mobile light detection and ranging (LiDAR) technology has excellent potential in mapping street scenes, there is little research in constructing façade footprints from unorganized, uneven, and incomplete mobile LiDAR point clouds. In fact, façade footprint vectorization from mobile LiDAR data still involves a lot of manual work, especially in complex street scenes with various types of buildings. In this paper, we present a new and effective framework for extracting 2-D façade footprints from mobile LiDAR point clouds. The proposed framework consists of three steps: 1) line segment extraction from projected point clouds based on a hypotheses and selection strategy; 2) completion of missing parts between adjacent walls using line intersections; and 3) delineation of footprints through finding the least cost path in the graph of the line segments. We compare our method with several existing ones and discuss its robustness against data missing and noise such as nonwall structures and vegetation. Our proposed method is also tested in two large-scale data sets, a residential data set, and an urban data set. The coverage ratio, i.e., the percentage of outer wall points covered by the generated outlines in the residential data set is 93.4% and 91.7% in the urban data set is achieved. The mean distance between points of ground truth and constructed footprints for the residential data set and urban data set is 0.019 and 0.028 m, respectively. The experimental results demonstrate that the proposed framework is effective in modeling various façade footprints from mobile LiDAR point clouds. Shaobo Xia, Ruisheng Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2018 | Fully Convolutional Networks for Semantic Segmentation of Very High Resolution Remotely Sensed Images Combined With DSMabstractRecently, approaches based on fully convolutional networks (FCN) have achieved state-of-the-art performance in the semantic segmentation of very high resolution (VHR) remotely sensed images. One central issue in this method is the loss of detailed information due to downsampling operations in FCN. To solve this problem, we introduce the maximum fusion strategy that effectively combines semantic information from deep layers and detailed information from shallow layers. Furthermore, this letter develops a powerful backend to enhance the result of FCN by leveraging the digital surface model, which provides height information for VHR images. The proposed semantic segmentation scheme has achieved an overall accuracy of 90.6% on the ISPRS Vaihingen benchmark. Weiwei Sun 0006, Ruisheng Wang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2017 | Enhancing Urban Façades via LiDAR-Based SculptingabstractAbstract Buildings with symmetrical façades are ubiquitous in urban landscapes and detailed models of these buildings enhance the visual realism of digital urban scenes. However, a vast majority of the existing urban building models in web‐based 3D maps such as Google earth are either less detailed or heavily rely on texturing to render the details. We present a new framework for enhancing the details of such coarse models, using the geometry and symmetry inferred from the light detection and ranging (LiDAR) scans and 2D templates. The user‐defined 2D templates, referred to as coded planar meshes (CPMs), encodes the geometry of the smallest repeating 3D structures of the façades via face codes. Our encoding scheme, take into account the directions, type as well as the offset distance of the sculpting to be applied at the respective locations on the coarse model. In our approach, LiDAR scan is registered with the coarse models taken from Google earth 3D or Bing maps 3D and decomposed into dominant planar segments (each representing the frontal or lateral walls of the building). The façade segments are then split into horizontal and vertical tiles using a weighted point count function defined over the window or door boundaries. This is followed by an automatic identification of CPM locations with the help of a template fitting algorithm that respects the alignment regularity as well as the inter‐element spacing on the façade layout. Finally, 3D boolean sculpting operations are applied over the boxes induced by CPMs and the coarse model, and a detailed 3D model is generated. The proposed framework is capable of modelling details even with occluded scans and enhances not only the frontal façades (facing to the streets) but also the lateral façades of the buildings. We demonstrate the potentials of the proposed framework by providing several examples of enhanced Google earth models and highlight the advantages of our method when designing photo‐realistic urban façades. Jiju Poovvancheri, Ruisheng Wang 0001 |
Comput. Graph. Forum | 2 |
| 2017 | A Fast Edge Extraction Method for Mobile Lidar Point CloudsabstractEdges in mobile light detection and ranging (lidar) point clouds are important for many applications but usually overlooked. In this letter, we propose a fast edge extraction method for mobile lidar. First, an edge index based on geometric center is introduced and then gradients in unorganized 3-D point clouds are defined. By analyzing the ratio between eigenvalues, edge candidates can be detected. Finally, an edge linking algorithm named graph snapping is proposed. The method is tested extensively and the experimental results demonstrate that the proposed method is able to quickly extract most of 3-D edges with higher accuracy than the existing methods. Shaobo Xia, Ruisheng Wang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2017 | Topologically Aware Building Rooftop Reconstruction From Airborne Laser Scanning Point CloudsabstractThis paper presents a novel topologically aware 2.5-D building modeling methodology from airborne laser scanning point clouds. The building reconstruction process consists of three main steps: primitive clustering, boundary representation, and geometric modeling. In primitive clustering, we propose an enhanced probability density clustering algorithm to cluster the rooftop primitives by taking into account the topological consistency among primitives. In the second step, we employ a novel Voronoi subgraph-based algorithm to seamlessly trace the primitive boundaries. This algorithm guarantees the production of geometric models without crack defects among adjacent primitives. The primitive boundaries are further divided into multiple linear segments, from which the key points are generated. These key points help to form a hybrid representation of the boundary by combining the projected points with part of the original boundary points. The model representation by the hybrid key points is flexible and well captures the rooftop details to generate lightweight and highly regular building models. Finally, we assemble the primitive boundaries to form the topologically correct entities, which are regarded as the basic units for primitive triangulation. The reconstructed models not only have accurate geometry and correct topology but more importantly have abundant semantics, by which five levels of building models can be generated in real time. The proposed reconstruction method has been comprehensively evaluated on Toronto data set in terms of model compactness, multilevel model representation, and geometric accuracy. Dong Chen 0009, Ruisheng Wang 0001, Jiju Poovvancheri |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2017 | Road Curb Extraction From Mobile LiDAR Point CloudsabstractAutomatic extraction of road curbs from uneven, unorganized, noisy, and massive 3-D point clouds is a challenging task. Existing methods often project 3-D point clouds onto 2-D planes to extract curbs. However, the projection causes loss of 3-D information, which degrades the performance of the detection. This paper presents a robust, accurate, and efficient method to extract road curbs from 3-D mobile LiDAR point clouds. Our method consists of two steps: 1) extracting candidate points of curbs based on the proposed novel energy function and 2) refining candidate points using the proposed least cost path model. We evaluated the method on a large scale of residential area (16.7 GB, 300 million points) and an urban area (1.07 GB, 20 million points) mobile LiDAR point clouds. Results indicate that the proposed method is superior to the state-of-the-art methods in terms of robustness, accuracy, and efficiency. The proposed curb extraction method achieved a completeness of 78.62% and a correctness of 83.29%. Experiments demonstrate that our method is a promising solution to extract road curbs from mobile LiDAR point clouds. Sheng Xu 0003, Ruisheng Wang 0001, Han Zheng 0002 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2017 | Recognizing Street Lighting Poles From Mobile LiDAR DataabstractIn this paper, a novel segmentation and recognition approach to automatically extract street lighting poles from mobile LiDAR data is proposed. First, points on or around the ground are extracted and removed through a piecewise elevation histogram segmentation method. Then, a new graph-cut-based segmentation method is introduced to extract the street lighting poles from each cluster obtained through a Euclidean distance clustering algorithm. In addition to the spatial information, the street lighting pole's shape and the point's intensity information are also considered to formulate the energy function. Finally, a Gaussian-mixture-model-based method is introduced to recognize the street lighting poles from the candidate clusters. The proposed approach is tested on several point clouds collected by different mobile LiDAR systems. Experimental results show that the proposed method is robust to noises and achieves an overall performance of 90% in terms of true positive rate. Han Zheng 0002, Ruisheng Wang 0001, Sheng Xu 0003 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2016 | Enhancing scene parsing by transferring structures via efficient low-rank graph matchingabstractScene parsing has attracted significant attention for its practical and theoretical value in computer vision. A typical scene parsing algorithm seeks to densely label pixels or 3-dimensional points from a scene. Traditionally, this procedure relies on a pre-trained classifier to identify the label information, and a smoothing step via Markov Random Field to enhance the consistency. LabelTranfer is a category of scene parsing algorithms to enhance traditional scene parsing framework, by finding dense correspondence and transferring labels across scenes. In this paper, we present a novel scene parsing algorithm which matches maximal similar structures between scenes via efficient low-rank graph matching. The inputs of the algorithm are images, and well- aligned point clouds if available. The images and the point clouds are processed in separate pipelines. The pipeline of images is to learn a reliable classifier and to match local structures via graph matching. The pipeline of point clouds is to conduct preliminary segmentation and to generate feasible label sets. The two pipelines are merged at inference step, in which we elaborate effective and efficient potential functions. We propose a new graph matching model incorporating low-rank and Frobenius regularization, which not only guarantees an accurate solution, but also provides high optimization efficiency via an eigen-decomposition strategy. Several challenging experiments are conducted, showing competitive performance of the proposed method compared to state-of-the-art LabelTransfer algorithm. Further, with point clouds, the performance can be significantly enhanced. Tianshu Yu 0001, Ruisheng Wang 0001 |
SIGSPATIAL/GIS | 2 |
| 2016 | Graph matching with low-rank regularizationabstractGraph matching is a widely researched topic which has been utilized in various applications of computer vision. Due to the combinatorial nature of graph matching, it is NP-hard to find an exact solution. So exact graph matching is always relaxed to inexact graph matching which seeks to find an approximate solution for the original problem. For a matching problem in quadratic form, semidefinite programming (SDP) relaxation is proven to be effective. However, previous SDP relaxation methods discard the constraint that the solution matrix is rank one, because the rank of a matrix is non-convex. In this paper, we explore some good properties of the solution matrix. By relaxing the rank into convex form using the properties, we propose to reformulate the graph matching with low rank constraint into a standard SDP, which can be easily solved. We test our method on both synthetic and real world data. The experimental results demonstrate that our method effectively handles low rank constraint and achieves competitive performance on robustness test against state-of-the-art counterparts. Tianshu Yu 0001, Ruisheng Wang 0001 |
WACV | 2 |
| 2016 | Scene parsing using graph matching on street-view data
Tianshu Yu 0001, Ruisheng Wang 0001 |
Comput. Vis. Image Underst. | 2 |
| 2012 | A new upsampling method for mobile LiDAR dataabstractWe present a novel method to upsample mobile LiDAR data using panoramic images collected in urban environments. Our method differs from existing methods in the following aspects: First, we consider point visibility with respect to a given viewpoint, and use only visible points for interpolation; second, we present a multi-resolution depth map based visibility computation method; third, we present ray casting methods for upsampling mobile LiDAR data incorporating constraints from color information of spherical images. The experiments show the effectiveness of the proposed approach. Ruisheng Wang 0001, Jeff Bach, Jane MacFarlane, Frank P. Ferrie |
WACV | 1 |
| 2011 | Window detection from mobile LiDAR dataabstractWe present an automatic approach to window and façade detection from LiDAR (Light Detection And Ranging) data collected from a moving vehicle along streets in urban environments. The proposed method combines bottom-up with top-down strategies to extract façade planes from noisy LiDAR point clouds. The window detection is achieved through a two-step approach: potential window point detection and window localization. The facade pattern is automatically inferred to enhance the robustness of the window detection. Experimental results on six datasets result in 71.2% and 88.9% in the first two datasets, 100% for the rest four datasets in terms of completeness rate, and 100% correctness rate for all the tested datasets, which demonstrate the effectiveness of the proposed solution. The application potential includes generation of building facade models with street-level details and texture synthesis for producing realistic occlusion-free façade texture. Ruisheng Wang 0001, Jeff Bach, Frank P. Ferrie |
WACV | 1 |
| 2009 | Next generation map making: geo-referenced ground-level LIDAR point clouds for automatic retro-reflective road feature extractionabstractThis paper presents a novel method to process large scale, ground level Light Detection and Ranging (LIDAR) data to automatically detect geo-referenced navigation attributes (traffic signs and lane markings) corresponding to a collection travel path. A mobile data collection device is introduced. Both the intensity of the LIDAR light return and 3-D information of the point clouds are used to find retroreflective, painted objects. Panoramic and high definition images are registered with 3-D point clouds so that the content of the sign and color can subsequently be extracted. Brad Kohlmeyer, Matei Stroila, Narayanan Alwar, Ruisheng Wang 0001, Jeff Bach |
GIS | 5 |