EDBT 2026 Demo / reviewers in the wild / expert
Yiping Chen 0002
dblp:40/1106-2
· DBLP profile ↗
58ranked-venue papers
4as first author
37since 2021 · last 2026
0000-0003-1465-6599ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 47 · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SGS-3D: High-Fidelity 3D Instance Segmentation via Reliable Semantic Mask Splitting and GrowingabstractAccurate 3D instance segmentation is crucial for high-quality scene understanding in the 3D vision domain. However, 3D instance segmentation based on 2D-to-3D lifting approaches struggle to produce precise instance-level segmentation, due to accumulated errors introduced during the lifting process from ambiguous semantic guidance and insufficient depth constraints. To tackle these challenges, we propose Splitting and Growing reliable Semantic mask for high-fidelity 3D instance segmentation (SGS-3D), a novel "split-then-grow" framework that first purifies and splits ambiguous lifted masks using geometric primitives, and then grows them into complete instances within the scene. Unlike existing approaches that directly rely on raw lifted masks and sacrifice segmentation accuracy, SGS-3D serves as a training-free refinement method that jointly fuses semantic and geometric information, enabling effective cooperation between the two levels of representation. Specifically, for semantic guidance, we introduce a mask filtering strategy that leverages the co-occurrence of 3D geometry primitives to identify and remove ambiguous masks, thereby ensuring more reliable semantic consistency with the 3D object instances. For the geometric refinement, we construct fine-grained object instances by exploiting both spatial continuity and high-level features, particularly in the case of semantic ambiguity between distinct objects. Experimental results on ScanNet200, ScanNet++, and KITTI-360 demonstrate that SGS-3D substantially improves segmentation accuracy and robustness against inaccurate masks from pre-trained models, yielding high-fidelity object instances while maintaining strong generalization across diverse indoor and outdoor environments. Chaolei Wang, Yang Luo 0002, Siyu Chen 0004, Yiping Chen 0002, Ting Han 0001 |
AAAI | 5 |
| 2026 | A Policy-Driven Black-Box Adversarial Example With Location Optimization Against 3D Object DetectionabstractAdversarial attack strategies for 3D object detection have highlighted the critical importance of addressing security concerns in this domain. However, white-box methods require full access to the victim model in large-scale point cloud applications. To this end, we propose a novel Policy-Driven Black-box Attack (BAT) that is designed to optimize attack locations without necessitating detailed knowledge of the victim models. First, we introduce a density-aware pattern generator that creates scene-adaptive attack clusters. Second, we leverage the deep deterministic policy gradient in deep reinforcement learning to train an attack agent capable of targeting the victim model. Ultimately, the attack agent is iteratively directed towards optimal attack locations through the joint application of critic loss and actor loss. To the best of our knowledge, this represents the first reinforcement learning-based black-box attack applied to practical 3D object detection. Experimental results on the KITTI, nuScenes, and Waymo datasets demonstrate that BAT effectively diminishes the accuracy of notable models. Importantly, BAT significantly enhances the attack success rate (surpassing state-of-the-art both white-box and black-box methods) and increases transferability (by 20 times) through simple deep deterministic policy gradient, thus establishing a new baseline for adversarial attacks in 3D object detection. Ting Han 0001, Xiaobin Wu, Chaolei Wang, Huan Luo 0001, Xiaochun Cao, Li Liu 0002, Yiping Chen 0002 |
IEEE Trans. Image Process. | 8 |
| 2025 | Scene4U: Hierarchical Layered 3D Scene Reconstruction from Single Panoramic Image for Your Immerse ExplorationabstractThe reconstruction of immersive and realistic 3D scenes holds significant practical importance in various fields of computer vision and computer graphics. Typically, immersive and realistic scenes should be free from obstructions by dynamic objects, maintain global texture consistency, and allow for unrestricted exploration. The current mainstream methods for image-driven scene construction involves iteratively refining the initial image using a moving virtual camera to generate the scene. However, previous methods struggle with visual discontinuities due to global texture inconsistencies under varying camera poses, and they frequently exhibit scene voids caused by foreground-background occlusions. To this end, we propose a novel layered 3D scene reconstruction framework from panoramic image, named Scene4U. Specifically, Scene4U integrates an open-vocabulary segmentation model with a large language model to decompose a real panorama into multiple layers. Then, we employs a layered repair module based on diffusion model to restore occluded regions using visual cues and depth information, generating a hierarchical representation of the scene. The multi-layer panorama is then initialized as a 3D Gaussian Splatting representation, followed by layered optimization, which ultimately produces an immersive 3D scene with semantic and structural consistency that supports free exploration. Scene4U outperforms state-of-the-art method, improving by 24.24% in LPIPS and 24.40% in BRISQUE, while also achieving the fastest training speed. Additionally, to demonstrate the robustness of Scene4U and allow users to experience immersive scenes from various landmarks, we build WorldVista3D dataset for 3D scene reconstruction, which contains panoramic images of globally renowned sites. The implementation code and dataset will be made publicly available. Junyan Ye, Lihan Jiang, Yiping Chen 0002, Ting Han 0001 |
CVPR | 6 |
| 2025 | VoxelFlow: 2D Semantic Mask-Guided Voxel Flow for Open-Vocabulary 3D Instance SegmentationabstractOpen-vocabulary 3D instance segmentation (OV3IS) has emerged as a promising field, successfully bridging point cloud data and text descriptions through intermediate image modalities. However, existing approaches often rely on strong 3D clustering priors or a single, isolated 2D mask branch, which typically results in insufficient and inefficient feature interaction between the point cloud and image domains. To address these limitations, we propose VoxelFlow, a novel framework that effectively combines 2D mask semantics with 3D geometric features. Our core insight is to leverage the implicit semantic and spatial information embedded in 2D masks and leverage them to replace rigid 3D priors. Specifically, voxel masks, initially derived from 2D projections, are meticulously refined through a unique two-stage process: growth and cross-frame merging, ensuring precise alignment with true instance masks under both semantic and geometric constraints. The methodology unfolds in three key steps. Firstly, we project 2D masks onto a voxelized 3D scene to establish initial seed voxel masks in 3D. Secondly, these voxel masks are grown based on local 3D geometric features and progressively merged across frames using a dynamic threshold, yielding robust super voxel masks. Finally, these semantically consistent and geometrically aligned super voxel masks are integrated with a visual language model and text embeddings to achieve comprehensive open-vocabulary instance segmentation. Extensive experiments on ScanNet200 and ScanNet++ demonstrate that VoxelFlow not only significantly outperforms robust 3D prior baselines but also achieves superior performance in both class-agnostic and open-vocabulary 3D instance segmentation. Chaolei Wang, Huan Chen 0025, Ting Han 0001, Yiping Chen 0002 |
CW | 5 |
| 2025 | CityInsight: Incorporating Dual-Condition-Based Diffusion Model Into Building Footprint Segmentation From Remote Sensing ImageryabstractAccurately identifying urban building layouts plays a crucial role in understanding the complexity of urban construction and the level of economic development. Previous footprint segmentation methods have struggled to adapt to the diverse morphology of buildings, limiting the accurate extraction of building footprints from remote sensing imagery and impeding insights and understanding of urban areas. To this end, we propose a framework named CityInsight for analyzing urban building morphology from remote sensing imagery. First, we establish a semantic segmentation network, dual-condition diffusion network (DC-Net), based on a diffusion model to accurately identify building footprints from remote sensing images. Second, we use uncertainty attention and condition attention to generate spatial and semantic priors. Finally, we design a condition injection module to incorporate spatial and semantic information into the diffusion learning. Comprehensive experiments demonstrate the accuracy, robustness, and generalization of the proposed method. The$F_{1}$-scores of the DC-Net on the large-scale remote sensing datasets SpaceNet, WHU Building, Inria, and Massachusetts are 92.05%, 96.59%, 92.17%, and 92.86%, respectively. Furthermore, the footprint segmentation is utilized for subsequent urban function identification and urban analysis of the Zona Oeste of Rio de Janeiro, to emphasize the application value of CityInsight. Our code is available athttps://github.com/Ting-Devin-Han/CityInsight Ting Han 0001, Chaolei Wang, Yang Luo 0002, Hongchao Fan, José Marcato Junior, Xinchang Zhang 0002, Yiping Chen 0002 |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2025 | CSFNet: Cross-Modal Semantic Focus Network for Semantic Segmentation of Large-Scale Point CloudsabstractSemantic segmentation of large-scale point clouds is an indispensable component of outdoor scene perception, providing essential 3-D semantic insights for applications in scene reconstruction, urban planning, autonomous driving, and more. However, the discriminative capability of point clouds features declines with increasing distance from the sensor, causing current methods to usually perform poorly in segmenting distant objects. To overcome this challenge and improve the differentiation between classes with similar geometric features, we propose the cross-modal semantic focus network (CSFNet). Firstly, we design a multiscale feature dynamic fusion (MDF) module to leverage multiscale image features, thereby enriching the feature representation of point clouds with additional images color and texture information. Then, in order to extract the distinguishing features of distant and different categories of objects more efficiently, we propose a semantic focus module (SFM) that employs a multiclass contrastive learning strategy to enhance feature discrimination. Finally, we introduce cross-modal knowledge distillation (KD) to augment the model’s comprehension of point clouds. Extensive experiments conducted on the SemanticKITTI and nuScenes datasets demonstrate the effectiveness of our method. Notably, our method achieves superior segmentation accuracy across multiple classes at various distances compared to current methods. Yang Luo 0002, Ting Han 0001, Yujun Liu 0005, Jinhe Su, Yiping Chen 0002, Yun-Dong Wu, Guo-Rong Cai |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | RSD-BiasEval: A Framework for Remote Sensing Dataset Bias Analysis and Evaluation
Xinrui Xie, ZiYue Lin, Haifeng Li 0007, Ji Qi 0001, Yibo Wang 0014, Yiping Chen 0002, Xinchang Zhang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | HSPFormer: Hierarchical Spatial Perception Transformer for Semantic SegmentationabstractSemantic perception in driving scenarios plays a crucial role in intelligent transportation systems. However, existing Transformer-based semantic segmentation methods often do not fully exploit their potential in understanding driving scene dynamically. These methods typically lack spatial reasoning, failing to effectively correlate image pixels with their spatial positions, leading to attention drift. To address this issue, we propose a novel architecture, the Hierarchical Spatial Perception Transformer (HSPFormer), which integrates monocular depth estimation and semantic segmentation into a unified framework for the first time. We introduce the Spatial Depth Perception Auxiliary Network (SDPNet), a framework for multiscale feature extraction and multilayer depth map prediction to establish hierarchical spatial coherence. Additionally, we design the Hierarchical Pyramid Transformer Network (HPTNet), which uses depth estimation as learnable position embeddings to form spatially correlated semantic representations and generate global contextual information. Experiments on benchmark datasets such as KITTI-360, Cityscapes, and NYU Depth V2, demonstrate that HSPFormer outperforms several state-of-the-art networks, and achieves promising performance with 66.82% top-1 mIoU on KITTI-360, 83.8% mIoU on Cityscapes, and 57.7% mIoU on NYU Depth V2, respectively. The code will be made publicly available athttps://github.com/SY-Ch/HSPFormer. Siyu Chen 0004, Ting Han 0001, Changshe Zhang, Jinhe Su, Ruisheng Wang 0001, Yiping Chen 0002, Zongyue Wang, Guo-Rong Cai |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | Supervised Contrastive Learning for Indoor Point Cloud OversegmentationabstractPoint cloud oversegmentation method can obtain a series of superpoints by grouping points that are semantically and geometrically consistent. The generated superpoints can be treated as the basic processing units in various downstream tasks to improve task performance and processing efficiency. However, due to the high semantic and geometric complexity of point cloud scenes, obtaining high-quality superpoints is still challenging. Aiming to generate high-quality indoor superpoints, we propose an end-to-end supervised contrastive learning framework SCL-OverSeg for indoor point cloud oversegmentation. Firstly, to solve the challenge of balancing the importance of geometric similarity and spatial proximity constraint between points and superpoints in indoor scenes, we integrate the geometric similarity and spatial proximity constraint into the supervision signal by generating the superpoint ground truth. To solve the challenge of superpoints crossing objects, we propose to utilize instance labels rather than semantic labels to generate the ideal superpoint ground truth as the object-level supervision signal. Secondly, to construct the distinguishable embedding space facilitating to the assignments of points to superpoints, we propose point-superpoint contrastive learning to compel the network to project each point to be closer to the reasonable superpoint in embedding space. Besides, with the instance labels, to improve the superpoint performance on object boundaries, we propose the object boundary contrastive learning to enhance the feature distinguishability between tough points across the object boundaries. Extensive experiments demonstrate that SCL-OverSeg can effectively improve indoor oversegmentation performance, especially on object boundaries. The relevant codes will be available onhttps://github.com/sssssyf/SCL-OverSeg. Yifan Sun 0008, Chenguang Dai, Wenke Li, Song Ji, Anzhu Yu, Yiping Chen 0002, Hanyun Wang |
IEEE Trans. Multim. | 7 |
| 2024 | Incorporating Building Information into Road Extraction from Remote Sensing ImagesabstractRoad extraction is a critical task in the field of remote sensing, and is widely applied in various tasks. However, significant challenges arise in road extraction due to occlusions from surrounding objects and chaotic backgrounds. Signed Distance Map (SDM) of buildings contains distance information from pixels to building edges, which can serve as auxiliary information for road extraction tasks. To address these challenges, a Road-SDM Fusion Network (RSDMNet) is proposed to explore the latent relationship between roads and buildings to enhance contextual reasoning capabilities. By incorporating the distribution information of buildings, the proposed method provides additional prior features and more global information for road extraction. Experimental results demonstrate that the proposed method achieves a higher IoU and outperforms state-of-the-art approaches. Xiangyi Xie, Ting Han 0001, Yumeng Du, Yiping Chen 0002 |
IGARSS | 4 |
| 2024 | SPTNet: Sparse Convolution and Transformer Network for Woody and Foliage Components Separation From Point CloudsabstractThe separation of woody and foliage components is beneficial in estimating the physical parameters of forests. However, many current methods incur high computational costs and rely on extensive prior knowledge. These methods display weak abilities in generalization for component separation from various LiDAR sensors and tree species. In this paper, a network that combines sparse convolution and transform blocks is proposed for the separation of woody and foliage components in tree point clouds called SPTNet. The sparse convolution block facilitates efficient and effective local feature extraction, while the transformer block offers a solution for the inadequate global feature extraction in sparse convolution blocks. Point feature extraction blocks, called Morphological Detection Coefficient (MDC) and Normal Difference Operator (NDO), were specifically developed to aid in the segmentation task. Distinct adaptive radius strategies are implemented for each geometric feature block to minimize the need for a priori knowledge. Eight different tree species datasets were used to improve methods, including a simulated larch dataset. The other datasets consist of actual trees and comprise seven distinct tree species along with a large tropical tree dataset. Our experimental results demonstrate that our method attains state-of-the-art performance across all datasets. It’s worth mentioning that SPTNet obtains an OA of 94.69% and 89.96% mIoU on the large tropical dataset, which encompasses 15 tree species. Moreover, SPTNet outperforms FWCNN, the current leading branch and leaf separation approach, by 0.43% OA and 0.72% mIoU. Shuai Zhang 0043, Yiping Chen 0002, Dong Pan 0003, Wuming Zhang, Aiguang Li |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Crack-U2Net: Multiscale Feature Learning Network for Pavement Crack Detection From Large-Scale MLS Point CloudsabstractDeep learning-based algorithms detect pavement cracks in an end-to-end manner from Mobile Laser Scanning (MLS) point clouds, achieving impressive results. However, the accuracy of existing methods still has room to improve due to the difficulty of effectively encoding multiscale features and the limited training data. In this paper, we propose a novel pavement crack detection framework, Crack-U2Net, which innovatively incorporates a two-level nested U-Net architecture for feature learning. This design enables the learning of intra-stage multiscale features without introducing significant memory and computation costs, resulting in substantial improvements in accuracy. Moreover, to solve the challenge of insufficient training data, we propose a Geometry-based Data Augmentation (GDA) strategy, aiming to expand the pavement dataset while preserving the pavement geometry. Extensive experiments on the Qinghai-Tibet Highway point cloud dataset demonstrate the higher accuracy and efficiency of Crack-U2Net over the state-of-the-art methods, achieving an average precision, recall, F$1\text - $score, and accuracy of 83.8%, 77.6%, 80.1%, and 95.8%, respectively. Huifang Feng 0002, Wen Li 0005, Lingfei Ma, Yiping Chen 0002, Haiyan Guan, Yongtao Yu, José Marcato Junior, Jonathan Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | eViTBins: Edge-Enhanced Vision-Transformer Bins for Monocular Depth Estimation on Edge DevicesabstractMonocular depth estimation (MDE) remains a fundamental yet not well-solved problem in computer vision. Current wisdom of MDE often achieves blurred or even indistinct depth boundaries, degenerating the quality of vision-based intelligent transportation systems. This paper presents an edge-enhanced vision transformer bins network for monocular depth estimation, termed eViTBins. eViTBins has three core modules to predict monocular depth maps with exceptional smoothness, accuracy, and fidelity to scene structures and object edges. First, a multi-scale feature fusion module is proposed to circumvent the loss of depth information at various levels during depth regression. Second, an image-guided edge-enhancement module is proposed to accurately infer depth values around image boundaries. Third, a vision transformer-based depth discretization module is introduced to comprehend the global depth distribution. Meanwhile, unlike most MDE models that rely on high-performance GPUs, eViTBins is optimized for seamless deployment on edge devices, such as NVIDIA Jetson Nano and Google Coral SBC, making it ideal for real-time intelligent transportation systems applications. Extensive experimental evaluations corroborate the superiority of eViTBins over competing methods, notably in terms of preserving depth edges and global depth representations. Yutong She, Peng Li 0064, Mingqiang Wei, Dong Liang 0008, Yiping Chen 0002, Haoran Xie 0001, Fu Lee Wang |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | Joint Structure Detection and Multi-Scale Clustering Filtering for Tunnel Lining Extraction From Point CloudsabstractAccurate extraction of tunnel lining is crucial for deformation monitoring, damage detection, and tunnel modeling. However, the extraction of tunnel lining is still challenging due to the different tunnel shapes and complicated construction scenes. As for tunnel point clouds, inconsistent local density and non-structural features also pose challenges. Previous research has shown that utilizing the centerline of a tunnel is the preferred method for extracting tunnel lining. However, imprecision often arises due to noise points and point cloud hole. To address these challenges, we propose a novel extraction framework for various tunnel shapes that seldom depends on the centerline. First, the structure detection module is designed to detect the lining structure from the raw point cloud. Following that, multi-scale clustering filtering is applied for fine extraction. The global clustering filtering procedure is then implemented to process the projected point cloud. Finally, the local filtering approach is created for confusion clustering. Experimental results show that the highest Kappa coefficient of the proposed method is 95.5%. Compared with the elliptic cylinder method, the angle threshold method, and the template method, the extraction accuracy of our method is improved by 6.2%, 10.9%, and 9.7%, respectively. Yipeng Zhao, Aiguang Li, Zhigang Du, Yiping Chen 0002, Haili Sun, Zhiyang Zhi |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | Tree Species Classfifcation Using Deep Learning Based 3d Point Cloud Transformer on Airborne Lidar DataabstractThis paper applied a transformer based deep learning model 3D Point Cloud Transformer (3DPCT) to conduct a tree species classification of Airborne LiDAR data. There are a total 1291 single tree point clouds of 11 different species from coniferous and deciduous used in this paper. The model integrated the local and global feature learning modules from both pointwise and channel-wise, which provide promising results of tree species classification. We also investigate by adding more channels the classification results can be improved. Different number of points per each sample as the model input also deliver different accuracy. The highest overall accuracy of 11 categories classification achieved 86.1%, and precision and recall of each category provide more directions of future study. Dening Lu, Weikai Tan, Yiping Chen 0002, Jonathan Li 0001 |
IGARSS | 4 |
| 2023 | Feature Graph Convolution Network With Attentive Fusion for Large-Scale Point Clouds Semantic SegmentationabstractUnstructured nature of 3D point clouds in large scenes is a challenging problem to effectively learn local geometric structures for point cloud semantic segmentation. To address this issue, we proposed a Feature Graph Convolution Network with Attentive Fusion (FGC-AFNet) in this letter. Our method takes large point clouds as input and uses the Feature Graph Convoluton (FGC) module to construct a graph of the central point with its neighboring points to extract local features. Then, we reduced the number of points using Random Sampling (RS) to expand the receptive field gradually to obtain multi-level features. The network also employs a dual Attention Fusion (AF) mechanism for efficient feature aggregation. One is at different levels for semantic feature fusion, another is for narrowing the semantic feature gap between the encoder and decoder. Compared to state-of-the-art methods on the S3DIS and Toronto3D datasets, our method obtained competitive results, with an overall accuracy of 88.6% and 96.58%, and a mean intersection over union of 71.2% and 81.92% on S3DIS and Toronto3D, respectively. Jun Chen 0005, Yiping Chen 0002, Cheng Wang 0003 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2023 | BrGAN: Blur Resist Generative Adversarial Network With Multiple Joint Dilated Residual Convolutions for Chlorophyll Color Image RestorationabstractThis paper presents a Blur Resist Generative Adversarial Network (GAN) (BrGAN) with multiple joint dilated residual convolutions for chlorophyll image restoration of the Geostationary Ocean Color Imager (GOCI). First, a publicly available dataset was built to support this study. Second, a multiple attention perception mechanism and a multiple joint dilated residual convolution module was proposed to cope with the challenge of large missing areas in GOCI chlorophyll images. Third, a patch GAN based discrimination module was proposed to avoid the restored areas with generating mosaic and shadows. Our experimental results demonstrate that the BrGAN can reach 37.06 in the peak signal-to-noise ratio (PSNR) and 0.0485 in the Learned Perceptual Image Patch Similarity (LPIPS), respectively. The comparative study shows that the BrGAN achieves the highest effectiveness and advancement among other seven state-of-the-art methods. Ziyi Chen 0001, Yuhua Luo, Yiping Chen 0002, Jing Wang 0049, Dilong Li, Kyle Gao, Cheng Wang 0003, Jonathan Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | A Click-Based Interactive Segmentation Network for Point CloudsabstractInteractive segmentation plays an essential role in several tasks involving point clouds. However, existing methods suffer from low segmentation accuracy and cannot adjust the segmentation results according to the user’s personal demands. This paper presents a novel deep learning-based interactive segmentation method, named Click Rough Segmentation Network (CRSNet), designed to handle point clouds. The method allows users to iteratively click to segment interesting objects. CRSNet consists of two key parts: a CRS module and a feature extraction module. First, the CRS module transforms click operations into an appropriate representation to input into the feature extraction module. The CRS module takes raw point clouds and click operations as input and outputs 3D Gaussian vectors and roughly segmented blocks, which adapt to different-sized and densely-distributed objects in complex environments. Second, the feature extraction module, which uses a novel mix loss-based analysis algorithm, extracts deep features and obtains instance segmentation results. The module is highly compatible because its backbones can be replaced by different deep learning architectures. Experimental results on the KITTI, Apolloscape, Roadmarking, Scannet, and SemanticKITTI datasets show that our method outperforms state-of-the-art semantic segmentation methods with one click. Moreover, our method can generalize well to unseen objects and datasets. Wentao Sun, Yiping Chen 0002, Huxiong Li, José Marcato Junior, Wesley Nunes Gonçalves, Jonathan Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Automated Detection of Oil/Gas Well Sites Detection from Multi-Source High Spatial Resolution ImagesabstractWith the development of oil/gas production, its adverse impact has drawn much attention. Therefore, automated detecting oil/gas well sites become important. Current research has been focused on detection using sole source RGB images, which was not the common case in remote sensing. In this study, we explored the use of a combination of the Residual Channel Attention Network (RCAN) and the state-of-the-art object detection method, You Only Look Once (YOLO) v4, to detect oil/gas well sites from multi-sensor images. To testify the feasibility of the combination, we selected 18 RapidEye and 15 WorldView images which cover the oil sands area in Alberta, Canada. We applied a pre-trained RCAN to unify the spatial resolution of different images to 2m/pixel. To maximize the feature space, we preserved 5 bands which were available in both images. YOLO v4 was applied on cropped image patches to detect oil/gas well sites. The experiment results showed that using the framework proposed in this study, oil/gas well sites can be localized accurately although the bounding boxes of the sites may not perfectly align with the objects. Hongjie He 0003, Hongzhang Xu, Michael A. Chapman, Yiping Chen 0002, Jonathan Li 0001 |
IGARSS | 5 |
| 2022 | Assessing the Impact of Covid-19 on Human Activities in the Greater Toronto Area by Nighttime Light Images and Active Covid-19 CasesabstractThis paper explores the effect of COVID-19 outbreaks on human activity through nighttime light images of Greater Toronto Area (GTA), Canada. The methods used in this paper include image preprocessing, image classification, and spatial analysis. By using the nighttime light radiance data from VIIRS/NPP data products and COVID-19 cases and comparing this data from the pre-pandemic year, the impact of COVID-19 was analyzed. The result shows that during the pandemic year the monthly average nighttime light radiance has decreased about 4.3-5.0% compared to the pre-pandemic year. The classification results shows that the average percentage of changes in residential areas, public facilities, and commercial areas are 0.3%, −0.7%, and −1.2%, respectively of each corresponding month. Meanwhile, the spatial analysis results show population distribution patterns in GTA during the pandemic year. Overall, the nighttime lights (NTL) images can be used for a preliminary understanding of how COVID-19 affected human activities and is corroborated with other forms data collection used for the pandemic analysis. Jianshen Wang, Sarah Narges Fatholahi, Michael A. Chapman, Yiping Chen 0002, Jonathan Li 0001 |
IGARSS | 5 |
| 2022 | Foreground-Background Segmentation of Sequential Point CloudsabstractPoint clouds are receiving increasing attention in the field of computer vision. Hitherto, segmentation tasks for point clouds have dealt with single-frame data without considering the temporal information. In this paper, we propose a point cloud segmentation network based on a neuron-like model that can exploit the time information of point cloud data to improve network performance. The network's encoder adopts the structure of the U-Net network and sparse convolution. The features extracted by the encoder and the last moment are used as input to the time integration module. The time integration module is a timing-based neuron model which can fuse features from two moments and use them to enhance the features of the current moment. Experiments on the public autopilot dataset NuScenes validate that our proposed point cloud segmentation network can achieve a 97.7 % Intersection over Union (IoU) for segmenting static scenes and dynamic objects. The experiments further validate that fusing temporal information improves the performance of point cloud foreground-background segmentation compared to considering only data from a single frame point cloud data. Chengzhe Yang, Yiping Chen 0002, Cheng Wang 0003, Jonathan Li 0001 |
IGARSS | 2 |
| 2022 | Adaptive Pyramid Context Fusion for Point Cloud PerceptionabstractDeep learning for 3-D point cloud perception has been a very active research topic in recent years. A current trend is toward the combination of the semantically strong and the fine-grained information from different scales of intermediate representations to boost network generalization power and robustness against scale variation. One prominent challenge is how to effectively conduct the allocation of multiple scales of information. In this letter, we propose a module, named adaptive pyramid context fusion (APCF), to adaptively capture scales of contextual information from a multiscale feature pyramid for the point cloud. The APCF module reweights and aggregates the features from different levels in the feature pyramid via a softmax attention strategy. The allocation of information is adaptively conducted level by level from bottom to up first and then from top to bottom. To ensure both effectiveness and efficiency, we propose a multiscale context-aware network APCF-Net through applying our proposed APCF to the PointConv architecture. Experiments demonstrate that APCF-Net surpasses its vanilla counterpart by a large margin both in effectiveness and efficiency. Especially, APCF-Net outperforms state-of-the-art approaches on 3-D object classification and semantic segmentation task, with the overall accuracy of 93.3% on ModelNet40 and mIoU of 63.1% on ScanNet V2 online test. Haojia Lin, Wen Li 0005, Yiping Chen 0002, Cheng Wang 0003, Jonathan Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Building Instance Extraction Method Based on Improved Hybrid Task CascadeabstractAutomatic building extraction from remote sensing imagery is crucial to urban construction and management. To address the main challenges of diverse building scale and appearance, this letter proposes an automatic building instance extraction method based on an improved hybrid task cascade (HTC). Our method consists of three components by obtaining high-resolution representation, defining guided anchor, and forming focal loss to boost the adaptability of automatic building instance extraction. Comprehensive experimental results on WHU aerial building data set demonstrated that compared with the mainstream Mask R-CNN method, our method increased AP and AR in bounding box branch and mask branch by 9.8%–6.5% and 10.7%–8.0% respectively, especially AP$_{S}$and AP$_{L}$in the two branches by 10.1%–6.9% and 3.4%–2.4%, respectively. We evaluated the effectiveness and complexity of these components separately and discussed the universality and practicability of deep learning method in automatic building extraction. Yiping Chen 0002, Mingqiang Wei, Cheng Wang 0003, Wesley Nunes Gonçalves, José Marcato Junior, Jonathan Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | STN: Saliency-Guided Transformer Network for Point-Wise Semantic Segmentation of Urban ScenesabstractAccurate and effective road object semantic segmentation plays a significant role in supporting extensive intelligent transportation system (ITS)-related applications. However, most existing image-based methods and point-based methods cannot deliver promising solutions with respect to segmentation accuracy and robustness, especially in complex urban road scenes. Thus, we design a saliency-guided transformer architecture (STN) in this letter for point-wise semantic segmentation from mobile laser scanning (MLS) point clouds. First, four types of feature saliency maps are constructed to obtain more compact feature spaces for enhancing the feature encoding semantics. Then, integrated with offset attention mechanisms and edge convolutions, an effective point-wise transformer network is proposed to extract high-level features for point-wise label assignment of road objects. The STN model is evaluated on the Pairs-Lille-3D dataset and achieves satisfactory experimental results with 87.2% overall accuracy and 81.7% mean IoU, respectively. Comparative studies with five deep learning-based methods also prove the superior performance of the STN model for large-scale semantic segmentation tasks. Lingfei Ma, Jonathan Li 0001, Haiyan Guan, Yongtao Yu, Yiping Chen 0002 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | A Supervoxel Approach to Road Boundary Enhancement From 3-D LiDAR Point CloudsabstractRapid and accurate enhancement of road boundaries from terrestrial laser scanning (TLS) 3-D point clouds has been a challenging task in road infrastructure inventory. To address the challenge with a lack of ability to enhance object boundaries when the supervoxel number is less, this letter proposes a novel supervoxel segmentation algorithm framework for enhancing road boundaries from 3-D point clouds. First, we utilize radius$k$nearest-neighbor search method to obtain the neighborhood information after partitioning points on octrees with seed points. Second, the iterative weighted least square algorithm and spatial structure judgment are used to segment point clouds based on seed points. Finally, an update method to adjust the supervoxel centroids is applied with surrounding information in the first part. To verify the excellent performance, we tested the proposed method on two publicly large-scale point clouds benchmarks—IQmulus and TerraMobilita (IQTM) and Semantic 3-D. The experimental results demonstrate that our approach achieved approximately 48.98% and 68.41% boundary recall higher than two existing classical methods in the street scene, and our running time is feasible and effective. Zhengchuan Sha, Yiping Chen 0002, Yangbin Lin, Cheng Wang 0003, José Marcato Junior, Jonathan Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Dense Point Cloud Completion Based on Generative Adversarial NetworkabstractPoint cloud completion aims to reconstruct complete point clouds from partial point clouds, which is widely used in various fields such as autonomous driving and robotics. Most existing methods are sparse point cloud completion, where the number of point clouds after completion is relatively small and the details are insufficient. This article proposes a novel end-to-end generative adversarial network-based dense point cloud completion architecture (DPCG-Net). We design two generative adversarial network (GAN)-based modules that translate point cloud completion into mapping between global feature distributions obtained by encoding partial point clouds and ground truth, respectively. The first designed generator module proposes skip connections to fully connected layer-based network for regenerating global feature and changing the global feature distribution derived from the encoder module to approximate the ground truth global feature distribution. The second proposed discriminator module divides high-dimensional global feature vectors into several smaller batches for judgment to guarantee the similarity between the regenerated global feature and the ground truth. We perform quantitative and qualitative experiments on the ShapeNet and KITTI datasets. Experiments on ShapeNet demonstrate that our model outperforms other models in cases where the lack of a large proportion of point clouds results in a large loss of spatial structure, especially when 80% of point clouds are missing. Moreover, KITTI experiments reveal that it is also valid for realistic situations. In addition, application in classification shows that the classification accuracy of point clouds completed with DPCG-Net is as high as 86.5% under the condition of 80% missing point clouds. Ming Cheng 0002, Guoyan Li, Yiping Chen 0002, Jun Chen 0005, Cheng Wang 0003, Jonathan Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | A Lightweight Network for Building Extraction From Remote Sensing ImagesabstractBuilding extraction is a fundamental research topic in remote sensing image interpretation. Convolutional neural network (CNN)-based building extraction algorithms have achieved high accuracy but require a large account of parameters and calculations, which hinders the practical application of these algorithms. To address the challenge, we propose a lightweight network (RSR-Net) for building extraction from remote sensing images. The network consists of three basic units with only a few parameters, and uses the idea of the fusion of shallow features and deep features, which is proposed by U-Net. Before features fusion, the squeeze-and-excitation (SE) module in RSR-Net assigned channel weights to these deep and shallow features. This operation can effectively reduce the influence of noise caused by shallow features in feature fusion, so as to improve the performance of the model. We estimated our network on datasets and achieved 88.32%, 71.58%, and 77.07% intersection-over-union (IoU) on datasets of aerial image and satellite image in Wuhan University (WHU) dataset, and the self-made building dataset of Guangzhou University Town, with only 2.81 M parameters and 6.91 G floating point operations (FLOPs). In addition, we propose a strategy combining target and background prediction, which makes RSR-Net achieve 0.37% improvement in IoU on WHU aerial image dataset. The effectiveness of RSR-Net is high. It showed that the proposed network is light and fast for the application of convolution neural network algorithm in practice. Huaigang Huang, Yiping Chen 0002, Ruisheng Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | A GCN-Based Method for Extracting Power Lines and Pylons From Airborne LiDAR DataabstractExtracting the power lines and pylons automatically and accurately from airborne LiDAR data is a critical step in inspecting the routine power line, especially in the remote mountainous areas. However, challenges arise in using existing methods to extract the targets from large scenarios of remote mountainous areas since the terrain is undulating, and the features are difficult to distinguish. In this article, to overcome these challenges, we propose a graph convolutional network (GCN)-based method to extract power lines and pylons from Airborne LiDAR point clouds. First, data augmentation and near-ground filtering methods are developed to overcome the problems of insufficient and imbalanced samples in the LiDAR data. Then, a GCN-based framework is proposed to extract the power lines and pylons, which consist of two main modules, i.e., the neighborhood dimension information (NDI) module and the neighborhood geometry information aggregation (NGIA) module. These two modules are designed to strengthen the model’s ability to portray local geometric details. Besides, an attention fusion module is investigated to further improve the NDI and NGIA features. Finally, a line structure constraint algorithm is proposed to identify individual power lines, where the power corridor is reconstructed using a polynomial-based algorithm. Numerical experiments are conducted based on two different power line scenarios acquired in mountainous areas. The results demonstrate the superior performances of the proposed method over several existing algorithms, where the$F_{1}$score and quality of the power line are 99.3% and 98.6%, and the results of the pylon are 96% and 92.4%, respectively. The identification rate of power line identification is above 98%. Wen Li 0005, Zhenlong Xiao, Yiping Chen 0002, Cheng Wang 0003, Jonathan Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Detection of Individual Trees in UAV LiDAR Point Clouds Using a Deep Learning Framework Based on Multichannel RepresentationabstractIndividual tree detection is critical for forest investigation and monitoring. Several existing methods have difficulties to detect trees in complex forest environments due to insufficiently mining descriptive features. This study proposes a deep learning (DL) framework based on a designed multichannel information complementarity representation for detecting trees in complex forest using UAV laser scanning point clouds. The proposed method consists of two main stages: ground filtering and tree detection. In the first stage, a modified graph convolution network with a local topological information layer is designed to separate the ground points. Unlike most existing parametric methods, our ground filtering method avoids the optimal parameters selection to adapt to different kinds of environments. For tree detection, a top-down slice (TDS) module is first designed to mine the vertical structure information in a top-down way. Then, a special multichannel representation (MCR) is developed to preserve different distribution patterns of points from complementary perspectives. Finally, a multibranch network (MBNet) is proposed for individual tree detection by fusing multichannel features, which can provide discriminative information for MBNet to detect trees more accurately. MBNet was evaluated on seven forest areas [UAV light detection and ranging (LiDAR) data with the mean size of 14$000~\text {m}^{2}$and point density of 250 points/$\text {m}^{2}$]. Experimental results showed that the proposed framework achieves excellent performance. Our method obtains promising performance with a mean recall of 89.23% and a mean F1-score of 87.04%. Wen Li 0005, Yiping Chen 0002, Cheng Wang 0003, Abdul Nurunnabi, Jonathan Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | DFAN: Dual-Branch Feature Alignment Network for Domain Adaptation on Point CloudsabstractUnsupervised domain adaptation (UDA) significantly reduces the gap between the source domain and the target domain in machine learning and computer vision tasks. Most UDA approaches are applied to images and videos, and only a few methods implement domain adaptation on 3-D computer vision problems. The existing UDA approaches operating on point clouds try to extract domain-invariant features in different domains for feature alignment. However, higher commonality brings less diversity and results a loss of detailed information. In this article, we propose a novel dual-branch feature alignment network (DFAN) architecture for domain adaptation on point cloud visual tasks to better exploit the respective characteristics of local and global features. Our approach specializes in the extraction and alignment of global and local features with different strategies in each branch to complement each other. We also introduce a hierarchical alignment strategy for local feature alignment and a distribution alignment strategy for global feature alignment. Experiments on the PointDA-10 and PointSegDA datasets show that our approach achieves state-of-the-art performance on the UDA of point cloud classification and segmentation tasks. The ablation study demonstrates the effectiveness of the dual-branch design and the feature alignment strategies. Liangwei Shi, Zhimin Yuan, Ming Cheng 0002, Yiping Chen 0002, Cheng Wang 0003 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Detecting Occluded and Dense Trees in Urban Terrestrial Views With a High-Quality Tree Detection DatasetabstractUrban trees are often densely planted along the two sides of a street. When observing these trees from a fixed view, they are inevitably occluded with each other and the passing vehicles. The high density and occlusion of urban tree scenes significantly degrade the performance of object detectors. This paper raises an intriguing learning-related question – if a module is developed to enable the network to adaptively cope with occluded and un-occluded regions while enhancing its feature extraction capabilities, can the performance of a cutting-edge detection model be improved? To answer it, a lightweight yet effective object detection network is proposed for discerning occluded and dense urban trees, called OD-UTDNet. The main contribution is a newly-designed Dilated Attention Cross Stage Partial (DACSP) module. DACSP can expand the fields-of-view of OD-UTDNet for paying more attention to the un-occluded region, while enhancing the network’s feature extraction ability in the occluded region. This work further explores both the self-calibrated convolution module and GFocal loss, which enhance the OD-UTDNet’s ability to resolve the challenging problem of high densities and occlusions. Finally, to facilitate the detection task of urban trees, a high-quality urban tree detection dataset is established, named UTD; to our knowledge, this is the first time. Extensive experiments show clear improvements of the proposed OD-UTDNet over twelve representative object detectors on UTD. The code and dataset are available at https://github.com/yzwang/OD-UTDNet. Yongzhen Wang 0001, Xuefeng Yan 0001, Hexiang Bao, Yiping Chen 0002, Lina Gong, Mingqiang Wei, Jonathan Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | GCN-Based Pavement Crack Detection Using Mobile LiDAR Point CloudsabstractMobile Laser Scanning (MLS) system can provide high-density and accurate 3D point clouds that enable rapid pavement crack detection for road maintenance tasks. Supervised learning-based algorithms have been proved pretty effective for handling such a large amount of inhomogeneous and unstructured point clouds. However, these algorithms often rely on a lot of annotated data, which is labor-intensive and time-consuming. This paper presents a semi-supervised point-level approach to overcome this challenge. We propose a graph-widen module to construct a reasonable graph structure for point clouds, increasing the detection performance of graph convolutional networks (GCN). The constructed graph characterizes the local features from a small amount of annotated data, avoiding information loss and dramatically reduces the dependence on annotated data. The MLS point clouds acquired by a commercial RIEGL VMX-450 system are used in this study. The experimental results demonstrate that our method outperforms the state-of-the-art point-level methods in terms of recall, F1 score, and efficiency while achieving comparable accuracy. Huifang Feng 0002, Wen Li 0005, Yiping Chen 0002, Sarah Narges Fatholahi, Ming Cheng 0002, Cheng Wang 0003, José Marcato Junior, Jonathan Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | Rapid Extraction of Urban Road Guardrails From Mobile LiDAR Point CloudsabstractMobile Laser Scanning (MLS) systems provide highly dense 3D point clouds that enable the acquisition of accurate traffic facilities information for intelligent transportation system. Road guardrails with safety features that can separate traffic and define moving spaces for pedestrians and vehicles face challenges such as diverse guardrail types and continuous slopes in point clouds data. This paper proposes a novel approach for rapidly extracting urban road guardrails from MLS point clouds, combining a proposed multi-level filtering with a modified Density-Based Spatial Clustering of Applications with Noise (DBSCAN) clustering, and adapting for most types of guardrails and rough slope roads. We develop a multi-level filter to detect the road surface and remove the undesirable points. Through a proposed modified DBSCAN clustering, the guardrails are extracted after a four-step screening, which includes the limits based on the number of points, the fitting error, the bounding box size and the average reflection intensity for each cluster. The proposed method achieves high precisions of 97.2% and 96.4% respectively for the lane-separating guardrails and the anti-fall guardrails on the dataset. Extensive experiments with test dataset captured by a RIEGL VMX-450 MLS, show that our method outperforms the state-of-the-art method to extract 3D guardrails from point clouds. Jianlan Gao, Yiping Chen 0002, José Marcato Junior, Cheng Wang 0003, Jonathan Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Cycle-SNSPGAN: Towards Real-World Image Dehazing via Cycle Spectral Normalized Soft Likelihood Estimation Patch GANabstractImage dehazing is a common operation in autonomous driving, traffic monitoring and surveillance. Learning-based image dehazing has achieved excellent performance recently. However, it is nearly impossible to capture pairs of hazy/clean images from the real world to train an image dehazing network. Most of existing dehazing models that are learnt from synthetically generated hazy images generalize poorly on real-world hazy scenarios due to the obvious domain shift. To deal with this unpaired problem arisen by real-world hazy images, we present Cycle Spectral Normalized Soft likelihood estimation Patch Generative Adversarial Network (Cycle-SNSPGAN) for image dehazing. Cycle-SNSPGAN is an unsupervised dehazing framework to boost the generalization ability on real-world hazy images. To leverage unpaired samples of real-world hazy images without relying on their clean counterparts, we design an SN-Soft-Patch GAN and exploit a new cyclic self-perceptual loss which avoids using the ground-truth image to compute the perceptual similarity. Moreover, a significant color loss is adopted to brighten the dehazed images as human expects. Both visual and numerical results show clear improvements of the proposed Cycle-SNSPGAN over state-of-the-arts in terms of hazy-robustness and image detail recovery, with even only a small dataset training our Cycle-SNSPGAN. Code has been available athttps://github.com/yz-wang/Cycle-SNSPGAN. Yongzhen Wang 0001, Xuefeng Yan 0001, Donghai Guan, Mingqiang Wei, Yiping Chen 0002, Xiao-Ping Zhang 0002, Jonathan Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2021 | A Local Topological Information Aware Based Deep Learning Method for Ground Filtering from Airborne Lidar DataabstractAs a foundational preprocessing step for a lot of downstream tasks, ground filtering from airborne LiDAR data is designed to separate the ground points and preserve the off-ground points with complete shape information. However, because of the undulating terrain, it is still a challenge work to filter the ground under complex mountain regions. In this paper, we provide a deep learning based model to improve the ground filtering performance in abrupt slope using airborne LiDAR point clouds. Specifically, we first design a local topological information mining module to extract the local features. Then a modified graph convolutional networks (GCNs) is developed to fusion the local features and global features. Compared with most existing methods, our model not only enjoys the parameter-free advantage, which means it can be applied easily in various areas, but also obtains better ground filtering performance and can preserve more complete information contained in off-ground points. Experiments was implemented on seven forest areas. The proposed method obtains promising ground filtering results with mean total error of 6.46% and the mean kappa coefficient of 86.01%. Wen Li 0005, Haojia Lin, Yiping Chen 0002, Cheng Wang 0003, Jonathan Li 0001 |
IGARSS | 5 |
| 2021 | Semantic Segmentation of UAV Lidar Point Clouds of a Stack Interchange with Deep Neural NetworksabstractStack interchanges are essential components of transportation systems. Mobile laser scanning (MLS) systems have been widely used in road infrastructure mapping, but accurate mapping of complicated multi-layer stack interchanges are still challenging. This study examined the point clouds collected by a new Unmanned Aerial Vehicle (UAV) Light Detection and Ranging (LiDAR) system to perform the semantic segmentation task of a stack interchange. An end-to-end supervised 3D deep learning framework was proposed to classify the point clouds. The proposed method has proven to capture 3D features in complicated interchange scenarios with stacked convolution and the result achieved over 93% classification accuracy. In addition, the new low-cost semi-solid-state LiDAR sensor Livox Mid-40 featuring a incommensurable rosette scanning pattern has demonstrated its potential in high-definition urban mapping. Weikai Tan, Dedong Zhang, Lingfei Ma, Nannan Qin, Yiping Chen 0002, Jonathan Li 0001 |
IGARSS | 6 |
| 2021 | Direction-aware Feature-level Frequency Decomposition for Single Image DerainingabstractWe present a novel direction-aware feature-level frequency decomposition network for single image deraining. Compared with existing solutions, the proposed network has three compelling characteristics. First, unlike previous algorithms, we propose to perform frequency decomposition at feature-level instead of image-level, allowing both low-frequency maps containing structures and high-frequency maps containing details to be continuously refined during the training procedure. Second, we further establish communication channels between low-frequency maps and high-frequency maps to interactively capture structures from high-frequency maps and add them back to low-frequency maps and, simultaneously, extract details from low-frequency maps and send them back to high-frequency maps, thereby removing rain streaks while preserving more delicate features in the input image. Third, different from existing algorithms using convolutional filters consistent in all directions, we propose a direction-aware filter to capture the direction of rain streaks in order to more effectively and thoroughly purge the input images of rain streaks. We extensively evaluate the proposed approach in three representative datasets and experimental results corroborate our approach consistently outperforms state-of-the-art deraining algorithms. Yidan Feng, Mingqiang Wei, Haoran Xie 0001, Yiping Chen 0002, Jonathan Li 0001, Xiao-Ping Zhang 0002, Harry Qin |
IJCAI | 5 |
| 2020 | A Boundary-Enhanced Supervoxel Method for 3D Point CloudsabstractThis paper presents a boundary-enhanced supervoxel method to solve over-segmentation problems in supervoxel generation of Voxel Cloud Connectivity Segmentation (VCCS). First, we use different searching methods to obtain the neighborhood of each point. Second, three variants of neighbor points are clustered by the local k-means clustering method on points directly instead of on voxels. Finally, a scale metric is used to measure the difference between two points that considers underlying 3D spatial structure of the points. Our proposed is tested on two publicly available benchmark point cloud datasets acquired by mobile laser scanning (MLS) and terrestrial laser scanning (TLS) systems, respectively. Results of the experiments show that the boundary recall approximately 7 and 4 times higher than VCCS for the best results, which our proposed methods are effective, and the cost time is feasible and effective. Zhengchuan Sha, Qing Zhu 0012, Yiping Chen 0002, Cheng Wang 0003, Abdul Nurunnabi, Jonathan Li 0001 |
IGARSS | 3 |
| 2019 | Geometric Multi-Model Fitting by Deep Reinforcement LearningabstractThis paper deals with the geometric multi-model fitting from noisy, unstructured point set data (e.g., laser scanned point clouds). We formulate multi-model fitting problem as a sequential decision making process. We then use a deep reinforcement learning algorithm to learn the optimal decisions towards the best fitting result. In this paper, we have compared our method against the state-of-the-art on simulated data. The results demonstrated that our approach significantly reduced the number of fitting iterations. Zongliang Zhang, Hongbin Zeng, Jonathan Li 0001, Yiping Chen 0002, Chenhui Yang, Cheng Wang 0003 |
AAAI | 4 |
| 2019 | Reconstruction of 3D Zebra Crossings from Mobile Laser Scanning Point CloudsabstractThis paper presents a novel method for reconstruction of three-dimensional (3D) zebra crossings from mobile laser scanning (MLS) point clouds. Firstly, we extract the zebra crossings from the 3D point cloud data in data preprocessing. Secondly, the fitting model is generated by seven parameters to determine one plane commonly and then calculating similarity for fitting the zebra crossings point clouds. Finally, the cuckoo search algorithm is used to adjust the parameters and optimize the results to obtain the optimal model and the geometric information of the zebra crossings can be acquired simultaneously. The proposed algorithm is tested on a set of point-clouds acquired by a RIEGL VMX-450 LiDAR system. The experimental results show the feasibility and stability of our method and the zebra crossings of urban road can be automatically and effectively reconstructed. Hongbin Zeng, Yiping Chen 0002, Zongliang Zhang, Cheng Wang 0003, Jonathan Li 0001 |
IGARSS | 2 |
| 2019 | Slam-Based Multi-Sensor Backpack Lidar Systems in Gnss-Denied EnvironmentsabstractThe backpack LiDAR system is a highly efficient device for indoor positioning and navigation. With the use of multi-sensor interactions inside of the backpack, it can solve the problem of undetermined trajectories or locations in a GNSS-denied environment. This paper presents a multi-sensor-based backpack LiDAR system for mapping in GNSS-denied environments. With this device, we solve the problem of inaccurate positioning caused by the inability to receive GNSS signals in a small-sized lidar portable device, also reduces the cumulative error caused by the IMU module through the analysis of the vibration pattern. Through the demonstration of the results, the method is of great significance for 3D reconstruction of the indoor environment and object detection. Dedong Zhang, Yiping Chen 0002, John S. Zelek, Jonathan Li 0001 |
IGARSS | 3 |
| 2018 | LiDAR-Video Driving Dataset: Learning Driving Policies EffectivelyabstractLearning autonomous-driving policies is one of the most challenging but promising tasks for computer vision. Most researchers believe that future research and applications should combine cameras, video recorders and laser scanners to obtain comprehensive semantic understanding of real traffic. However, current approaches only learn from large-scale videos, due to the lack of benchmarks that consist of precise laser-scanner data. In this paper, we are the first to propose a LiDAR-Video dataset, which provides large-scale high-quality point clouds scanned by a Velodyne laser, videos recorded by a dashboard camera and standard drivers' behaviors. Extensive experiments demonstrate that extra depth information help networks to determine driving policies indeed. Yiping Chen 0002, Jingkang Wang, Jonathan Li 0001, Cewu Lu, Cheng Wang 0003 |
CVPR | 1 |
| 2018 | Rural Road Networks Matching Via Extending LineabstractRoad network matching has played an important role in road network extraction and update, yet has got extensive researching during the recent decades. Differ from previous road matching methods focus mainly on the city areas, which have accurate and regular road networks, this paper aim to address the matching between incomplete ground survey road network and extracted road network from remote sensing images. Specifically, we propose an extending line based matching scheme to calculate the road primitive similarity by taking into account the surrounding connections and contextual information. The experimental results show that the proposed method is able to provide high quality matching results, even the ground survey data are very different from the extracted road network of the satellite. Thus makes it possible to implement the road network update for the wide rural regions without interested ground survey road network data. Yiping Chen 0002, Cheng Wang 0003, Jonathan Li 0001 |
IGARSS | 3 |
| 2018 | Estimation of Forest Trees Diameter from Terrestrial Laser Scanning Point Clouds Based on a Circle Fitting MethodabstractIn forest monitoring and management, any rational decision needs to be based on forest parameters. The diameter at breast height (DBH) of a tree is considered to be the most significant parameter among them. This paper presents a novel method for extracting tree stems and estimating DBH of trees in a forest environment from 3D point clouds data acquired by a terrestrial laser scanning (TLS) system. In the proposed method, a downward-growing algorithm is used to extract individual tree stems and DBH of trees are estimated by the circle fitting algorithm. This proposed method can avoid errors caused from tilted trees by estimating a plane perpendicular to the tree stem. With this method, 17 trees were extracted from single-scan point cloud data consisting of 21 trees. The estimated DBH had a bias of 0.38 cm and a root mean squared error of 1.76 cm, These experiment results show the feasibility of the proposed method. Rongren Wu, Yiping Chen 0002, Cheng Wang 0003, Jonathan Li 0001 |
IGARSS | 2 |
| 2017 | Use of ground penetrating radar for detecting underground holes in urban areas: XMU's experienceabstractThis paper presents a cosine-based back-projection(CBP) algorithm for ground penetrating radar(GPR) imaging to detect underground holes. Compared with the classic back-projection imaging algorithm, the CBP algorithm provides a better performance due to better effect of clutter suppression. In the proposed algorithm, a cosine-based measure function is established to describe the similar feature between every two different echo signals to achieve excellent artifact suppression. Then numerical simulation and experimental data set were used to test the SBP algorithm. The experimental data set was gained in Double Han Road, Siming District of Xiamen, China, and acquired by the the Latvia radar system, zond-12e model, which has a 115 kHz transmitting frequency, 40/80/160/320 Hz scanning frequency, ±40dB receive gain and 2000ns time-window width. The results fully demonstrate the effectiveness and superiority of CBP algorithm. Zhiyou Hong, Jonathan Li 0001, Zhenmiao Deng, Yiping Chen 0002 |
IGARSS | 5 |
| 2017 | Quality evaluation of point cloud model for interior structure of a common buildingabstractThis paper presents a standardized quality criteria to evaluate the 3D point cloud model of the indoor building which is based on point cloud's data accuracy, the prior characteristics of the building and the coincidence errors of the point cloud model. Our assessment framework involves three steps: the point cloud data acquisition, model generation and quality evaluation. In model generation progress, incapacity of scanning the whole building information one time since its multi-storied spatial structure, the building model need to be registered and merged. In evaluation step, taking into account the need of mapping, indoor location and navigation, the establishment of an interior spatial building model requires accurate measurement. Therefore, we adopt data noise analysis to give a judgement. Then, since the geometric characteristics of the building model are varying, the geometric analysis is proposed to evaluate acquisition errors and registration error. Comparative experiments demonstrate our method give integrate, realistic and reliable quality framework for the indoor building point cloud model. Yiping Chen 0002, Chenglu Wen, Cheng Wang 0003, Jonathan Li 0001, JinYong Chen |
IGARSS | 2 |
| 2017 | Automated extraction of urban roadside trees from mobile laser scanning point clouds based on a voxel growing methodabstractThis paper presents a new method for extracting urban roadside trees automatically from mobile laser scanning point clouds. This method mainly includes three steps. First, ground point clouds are removed by voxel-based upward growing method. Second, Euclidean distance segment method is used to cluster non-ground point clouds into certain individual objects. Then crown seeds of the initial layer is found by comparing the number of points in each layer after using the voxel modeling algorithm. Crown seeds in other layers can thereafter be detected via the upward inter-sectional analysis. Third, a crown voxel growing algorithm is used to make the crown grow in horizontal. The experimental results show that the voxel models of the individual roadside trees can be automatically and effectively extracted with our method. Zhenlong Xiao, Yiping Chen 0002, Pengdi Huang, Rongren Wu, Jonathan Li 0001 |
IGARSS | 3 |
| 2017 | Broken and degraded document images binarization
Yiping Chen 0002, Liansheng Wang 0002 |
Neurocomputing | 1 |
| 2017 | Traffic Sign Occlusion Detection Using Mobile Laser Scanning Point CloudsabstractFor survey and maintenance of traffic signs, this paper presents a novel traffic sign occlusion detection method using 3-D point clouds and trajectory data acquired by a mobile laser scanning system. To produce a maintenance guide, our method aims to obtain the degree of occlusion by analyzing the spatial relationship between traffic signs, surroundings, and drivers on the road. First, a detection method considering both reflectance and geometric features is developed to capture traffic signs. Next, to simulate the driver's view, a trajectory-based method is proposed to determine driver's observation location and the corresponding observed traffic sign. Finally, to determine whether a traffic sign is in occlusion, a hidden point removal algorithm is adopted and carried out. Furthermore, we develop two indices to evaluate the degree of occlusion. The proposed method is tested using two point cloud data sets collected by an RIEGL VMX-450 system along a 23.68-km-long urban road. The obtained results illustrate the feasibility of the proposed occlusion detection method. Pengdi Huang, Ming Cheng 0002, Yiping Chen 0002, Huan Luo 0001, Cheng Wang 0003, Jonathan Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2016 | Vehicle Detection in High-Resolution Aerial Images via Sparse Representation and SuperpixelsabstractThis paper presents a study of vehicle detection from high-resolution aerial images. In this paper, a superpixel segmentation method designed for aerial images is proposed to control the segmentation with a low breakage rate. To make the training and detection more efficient, we extract meaningful patches based on the centers of the segmented superpixels. After the segmentation, through a training sample selection iteration strategy that is based on the sparse representation, we obtain a complete and small training subset from the original entire training set. With the selected training subset, we obtain a dictionary with high discrimination ability for vehicle detection. During training and detection, the grids of histogram of oriented gradient descriptor are used for feature extraction. To further improve the training and detection efficiency, a method is proposed for the defined main direction estimation of each patch. By rotating each patch to its main direction, we give the patches consistent directions. Comprehensive analyses and comparisons on two data sets illustrate the satisfactory performance of the proposed algorithm. Ziyi Chen 0001, Cheng Wang 0003, Chenglu Wen, Xiuhua Teng, Yiping Chen 0002, Haiyan Guan, Huan Luo 0001, Liujuan Cao, Jonathan Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2016 | Vehicle Detection in High-Resolution Aerial Images Based on Fast Sparse Representation Classification and Multiorder FeatureabstractThis paper presents an algorithm for vehicle detection in high-resolution aerial images through a fast sparse representation classification method and a multiorder feature descriptor that contains information of texture, color, and high-order context. To speed up computation of sparse representation, a set of small dictionaries, instead of a large dictionary containing all training items, is used for classification. To extract the context information of a patch, we proposed a high-order context information extraction method based on the proposed fast sparse representation classification method. To effectively extract the color information, the RGB color space is transformed into color name space. Then, the color name information is embedded into the grids of histogram of oriented gradient feature to represent the low-order feature of vehicles. By combining low- and high-order features together, a multiorder feature is used to describe vehicles. We also proposed a sample selection strategy based on our fast sparse representation classification method to construct a complete training subset. Finally, a set of dictionaries, which are trained by the multiorder features of the selected training subset, is used to detect vehicles based on superpixel segmentation results of aerial images. Experimental results illustrate the satisfactory performance of our algorithm. Ziyi Chen 0001, Cheng Wang 0003, Huan Luo 0001, Hanyun Wang, Yiping Chen 0002, Chenglu Wen, Yongtao Yu, Liujuan Cao, Jonathan Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2015 | Extraction of street trees from mobile laser scanning point clouds based on subdivided dimensional featuresabstractThis paper proposes a method for automated extraction of street trees in a typical urban environment from 3D point cloud data acquired by the mobile laser scanning system. First, the algorithm utilizes the voxel-based method to remove the ground points from the scene. Second, the Euclidean distance clustering is adopted to cluster points into individual objects. The eigenvalues of neighborhood covariance matrix and the corresponding normalized centroid distance are computed for each point to obtain the subdivided dimensional features. Finally, the statistical component features and horizontal information are calculated for object detection. The experiment results show the feasibility of the proposed algorithm. Pengdi Huang, Yiping Chen 0002, Jonathan Li 0001, Yongtao Yu, Cheng Wang 0003, Hongshan Nie |
IGARSS | 2 |
| 2015 | Using mobile LiDAR point clouds for traffic sign detection and sign visibility estimationabstractThis paper presents a novel method for traffic sign detection and visibility evaluation from mobile Light Detection and Ranging (LiDAR) point clouds and the corresponding images. Our algorithm involves two steps. Firstly, a detection algorithm based on high retro-reflectivity of the traffic sign from the MLS point clouds is designed for sign detection in complicated road scenes. To solve the spatial features of traffic signs, we also create geo-referenced relations between traffic signs and roads according to the normal of ground. Secondly, we propose a visibility estimation method to evaluate the visibility level of the traffic sign based on a combination of visual appearance and spatial-related features. The proposed algorithm is validated on a set of transportation-related point-clouds acquired by a RIEGL VMX-450 LiDAR system. The experiment results demonstrate that the efficiency and reliability of the proposed algorithm in detection traffic signs are robust, and also prove the potential of using mobile LiDAR data for traffic sign visibility evaluation. Chenglu Wen, Huan Luo 0001, Yiping Chen 0002, Cheng Wang 0003, Jonathan Li 0001 |
IGARSS | 4 |
| 2015 | Inventory of 3D street lighting poles using mobile laser scanning point cloudsabstractThis paper presents a novel approach for extracting street lighting poles directly from MLS point clouds. The approach includes four stages: 1) elevation filtering to remove ground points, 2) Euclidean distance clustering to cluster points, 3) voxel-based normalized cut (Ncut) segmentation to separate overlapping objects, and 4) statistical analysis of geometric properties to extract 3D street lighting poles. A Dataset acquired by a RIEGL VMX-450 MLS system are tested with the proposed approach. The results demonstrate the efficiency and reliability of the proposed approach to extract 3D street lighting poles. Dawei Zai, Yiping Chen 0002, Jonathan Li 0001, Yongtao Yu, Cheng Wang 0003, Hongshan Nie |
IGARSS | 2 |
| 2015 | Application of L0-Norm Regularization to Epicardial Potential Reconstruction
Liansheng Wang 0002, Yiping Chen 0002, Harry Qin |
MICCAI (2) | 3 |
| 2015 | Hidden target detection from the multi-echo small-footprint LiDAR point cloudsabstractWe propose a new approach for hidden or potential object detection behind the vegetation based on the multi-echo small-footprint of light detection and ranging(LiDAR) point cloud. According to the specific characteristic that laser beam is able to penetrate foliage gaps, which offers the opportunity to perceive and detect the object invisible to the naked eye. First, the waveform sample data of the small footprint multi-echo liDAR uses a Gaussian fitting tool to curve fitting. Second, the peak detection of the waveform can be classified in statistical results of its wave numbers. Decomposition and correction are the processing for classifying based on the statistic data. After selecting the point clouds which are contained in the multi-peaks echoes, we obtain the tree and the embedded target behind it as well as ground elimination. The range of the waveform component is used to separate the penetrations material and the target by distance discriminant function. Experiments are implemented on the waveforms acquired by small-footprint LiDAR system VZ-1000 Sensor. The results indicate that the algorithm could provide an optimal solution for LiDAR waveform hidden target detection. Yiping Chen 0002, Cheng Wang 0003, Jonathan Li 0001 |
MMSP | 1 |
| 2015 | Road Boundaries Detection Based on Local Normal Saliency From Mobile Laser Scanning DataabstractThe accurate extraction of roads is a prerequisite for the automatic extraction of other road features. This letter describes a method for detecting road boundaries from mobile laser scanning (MLS) point clouds in an urban environment. The key idea of our method is directly constructing a saliency map on 3-D unorganized point clouds to extract road boundaries. The method consists of four major steps, i.e., road partition with the assistance of the vehicle trajectory, salient map construction and salient points extraction, curb detection and curb lowest points extraction, and road boundaries fitting. The performance of the proposed method is evaluated on the point clouds of an urban scene collected by a RIEGL VMX-450 MLS system. The completeness, correctness, and quality of the extracted road boundaries are 95.41%, 99.35%, and 94.81%, respectively. Experimental results demonstrate that our method is feasible for detecting road boundaries in MLS point clouds. Hanyun Wang, Huan Luo 0001, Chenglu Wen, Jun Cheng 0002, Peng Li 0064, Yiping Chen 0002, Cheng Wang 0003, Jonathan Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2010 | Structure-preserving multiscale vessel enhancing diffusion filterabstractEnhancement of vessels in medical images is still an unsolved problem. Multiscale approaches were proposed to improve the vessel enhancement effect based on the structure size and image resolution. Vessel enhancing diffusion (VED) filter is one of the multiscale approaches, which was based on the scale space theory. VED performs well on enhancing vessel structures but cannot preserve complex structures such as the vessel junctions. In this paper, a structure-preserving diffusion tensor is defined in the diffusion equation, which brings a structure-preserving vessel enhancing diffusion filter. Through the multiscale framework, the proposed method enhances the vessel structures especially the complex structure such as junctions. Experimental evaluation performed on various vessel data sets demonstrated the effectiveness of the proposed method. Yiping Chen 0002, Liansheng Wang 0002, Lin Shi 0001, Defeng Wang, Pheng-Ann Heng, Tien-Tsin Wong, Xiang Li 0014 |
ICIP | 1 |