Jonathan Li 0001

dblp:85/6906-1 · DBLP profile ↗
← Back
239ranked-venue papers
0as first author
89since 2021 · last 2026
0000-0001-7899-0049ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 191 · 74 since 2021Artificial intelligence and machine learning · 30 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 6 since 2021Databases, data management, data science and information retrieval · 6 · 4 since 2021Systems, architecture and hardware · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2Computer networks · 1Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 BCWildfire: A Long-term Multi-factor Dataset and Deep Learning Benchmark for Boreal Wildfire Risk Prediction
abstract
Wildfire risk prediction remains a critical yet challenging task due to the complex interactions among fuel conditions, meteorology, topography, and human activity. Despite growing interest in data-driven approaches, publicly available benchmark datasets that support long-term temporal modeling, large-scale spatial coverage, and multimodal drivers remain scarce. To address this gap, we present a 25-year, daily-resolution wildfire dataset covering 240 million hectares across British Columbia and surrounding regions. The dataset includes 38 covariates, encompassing active fire detections, weather variables, fuel conditions, terrain features, and anthropogenic factors. Using this benchmark, we evaluate a diverse set of time-series forecasting models, including CNN-based, linear-based, Transformer-based, and Mamba-based architectures. We also investigate effectiveness of position embedding and the relative importance of different fire-driving factors.
Zhengsen Xu, Sibo Cheng, Hongjie He 0003, Wentao Sun, Jonathan Li 0001, Lincoln Linlin Xu
AAAI6
2026 Co-MixPL: An optimized semi-supervised learning method for tunnel water leakage detection
Xujie Long, Jing Teng, Shaobo Zhao, Mengyang Pu, Ruifeng Shi, Jonathan Li 0001, Guoqing Jing
Adv. Eng. Informatics8
2026 SFE-CapsNet: Spatial Feature Enhanced Capsule Networks for Remote Sensing Object Detection
abstract
Remote sensing imagery often involves complex backgrounds and multi-scale targets, while variations such as rotation and scaling significantly degrade the performance of existing object detection algorithms and hinder effective modeling of spatial relationships between objects. To address these challenges, we propose SFE-CapsNet. First, we fuse scalar features extracted by convolutional neural networks (CNNs) with vector features from capsule networks through structural reorganization to form virtual capsules, approximating the dynamic routing process via a single fully connected layer. This design preserves object pose and texture information while streamlining information flow. Second, we introduce a capsule attention module that generates attention masks to dynamically enhance target-relevant features and suppress background noise, strengthening multi-level feature representations. Integrated with a feature pyramid network (FPN) architecture, our approach achieves precise detection of targets at varying scales. Experimental results demonstrate that multi-level feature fusion and the capsule attention mechanism significantly improve detection accuracy and robustness, achieving 77.65% mean Average Precision (mAP) on the DOTA dataset and 97.63% mAP on HRSC2016, highlighting its effectiveness and efficiency in complex scenes.
Ziyi Chen 0001, Wenhui Qiu, Huayou Wang, Dilong Li, Jin Gou, Cheng Wang 0003, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.8
2026 Visible and Infrared Image Fusion Based on Adaptive Weighted Multimodal Features Extraction and Bidirectional Guidance Structure
abstract
Efficient fusion of infrared and visible images is of critical importance for real-time applications such as autonomous driving. While deep learning-based fusion methods have demonstrated significant improvements in fusion quality in recent years, current network architectures still exhibit unsatisfactory computational complexity and processing speed. To reduce computational complexity and improve fusion efficiency, mask-based methods or approaches driven by downstream tasks often prioritize key regions. However, such methods tend to overemphasize target objects, potentially overlooking contextually significant elements. To address this limitation and achieve more effective fusion, we propose BEFuse, a decoupled two-stage training strategy with an end-to-end inference framework. BEFuse extracts shallow features and gradient information from images in the first stage via a cross-modal image segmentation subnetwork. In the fusion stage, we use the Hadamard product to map features to an implicit quadratic feature space, combining feature similarity and gradient mask information, allowing automatic adjustment of loss weights and improving fusion accuracy. Experiments on four datasets (MSRS, TNO, RoadScene, and M3FD) show that BEFuse outperforms existing methods in both fusion quality and computational speed.
Ziyi Chen 0001, Gaosheng Cai, Dilong Li, Jing Wang 0049, Jin Gou, Cheng Wang 0003, Jonathan Li 0001
IEEE Trans. Circuits Syst. Video Technol.8
2026 Leveraging Multi-View Images to Learn Domain-Invariant Discriminative Embeddings for Cross-View Geo-Localization
abstract
Cross-view geo-localization (CVGL) aims to match images of the same location captured from different viewpoints, such as those captured by Unmanned Aerial Vehicles (UAVs) and satellite platforms. The task is particularly challenging due to significant variations in scale, viewpoint, and illumination. Most existing methods employ symmetric sampling strategy to construct drone–satellite image pairs for deep metric learning, but neglect the potential of incorporating multi-view drone images to enhance the viewpoint robustness of features. To address this, we propose leveraging multi-view images to learn Domain-Invariant Discriminative Embeddings (DIDE) for CVGL. DIDE introduces an Inter-view Feature Aggregation Module (IFAM), which dynamically integrates multi-view drone information into robust embeddings. These are used in contrastive learning with satellite embeddings within batches to learn view-invariant discriminative features, while representation learning further improves scene discrimination across batches. To reduce the domain gap, DIDE constructs and aligns drone and satellite prototypes for effective cross-domain feature alignment. Furthermore, we adopt a parameter-efficient transfer learning strategy that leverages the capabilities of pre-trained foundation models while fine-tuning only dual adapters, significantly reducing the trainable parameters. DIDE achieves the state-of-the-art on University-1652 and University-160k, competitive results on SUES-200, and demonstrates strong cross-dataset transferability, with fewer training parameters and lower computational cost.
Ziyi Chen 0001, Dilong Li, Jin Gou, Cheng Wang 0003, Kyle Gao, Jonathan Li 0001
IEEE Trans. Circuits Syst. Video Technol.7
2026 Boundary-Guided Real-Time Semantic Segmentation and Pixel-Level Quantification of Pavement Cracks
abstract
Timely and accurately extracting and assessing pavement cracks is crucial for intelligent transportation systems (ITS) to improve road maintenance and safety. In this paper, we present an automated framework for crack semantic segmentation and quantification using optical images. First, a unique boundary-guided real-time high-resolution network is proposed, termed as BulletNet, for crack semantic segmentation. BulletNet is a bullet-head structure that can retain crack details while ensuring real-time inference speed, in which a Cross-Scale Global Attention (CSGA) module is designed to enhance global feature representation and pixel-level relations, as well as a Boundary-Guided Fusion (BGF) module proposed to utilize boundary features to guide the fusion of crack details and contextual information. Second, a Pixel-level Crack Quantification (PCQ) algorithm is proposed for complex cracks, incorporating an Improved Discrete Skeleton Evolution (IDSE) method to optimize skeleton pruning for accurate crack length and a normal vector correction method to adjust propagation direction for precise crack width. Comprehensive experiments on three datasets showed that the proposed BulletNet surpassed the comparative models in terms of efficiency and performance, with average F1-score, mIoU, and Frames per second (FPS) of 87.20%, 88.70%, and 125.53, respectively. In addition, tested on 200 images, the PCQ calculated the crack maximum widths and lengths with an average relative error of 6.96% and 4.62%, respectively. Finally, BulletNet was deployed on edge devices for field testing, and a system based on the PCQ algorithm was developed to validate the effectiveness of the entire framework.
Haiyan Guan, Lingfei Ma, Yongtao Yu, Sangning Li, Jonathan Li 0001
IEEE Trans. Intell. Transp. Syst.7
2025 L4DR: LiDAR-4DRadar Fusion for Weather-Robust 3D Object Detection
abstract
LiDAR-based 3D object detection is crucial for autonomous driving. However, due to the quality deterioration of LiDAR point clouds, it suffers from performance degradation in adverse weather conditions. Fusing LiDAR with the weatherrobust 4D radar sensor is expected to solve this problem; however, it faces challenges of significant differences in terms of data quality and the degree of degradation in adverse weather. To address these issues, we introduce L4DR, a weather-robust 3D object detection method that effectively achieves LiDAR and 4D Radar fusion. Our L4DR proposes Multi-Modal Encoding (MME) and Foreground-Aware Denoising (FAD) modules to reconcile sensor gaps, which is the first exploration of the complementarity of early fusion between LiDAR and 4D radar. Additionally, we design an Inter-Modal and IntraModal ({IM}2) parallel feature extraction backbone coupled with a Multi-Scale Gated Fusion (MSGF) module to counteract the varying degrees of sensor degradation under adverse weather conditions. Experimental evaluation on a VoD dataset with simulated fog proves that L4DR is more adaptable to changing weather conditions. It delivers a significant performance increase under different fog levels, improving the 3D mAP by up to 20.0% over the traditional LiDAR-only approach. Moreover, the results on the K-Radar dataset validate the consistent performance improvement of L4DR in realworld adverse weather conditions.
Xun Huang 0003, Ziyu Xu 0002, Qiming Xia, Yan Xia 0003, Jonathan Li 0001, Kyle Gao, Chenglu Wen, Cheng Wang 0003
AAAI7
2025 MTCloud: Multi-type convolutional linkage network for point cloud instance segmentation
Jing Du 0007, Guo-Rong Cai, Zongyue Wang, Jinhe Su, Min Huang 0004, John S. Zelek, José Marcato Junior, Jonathan Li 0001
Expert Syst. Appl.8
2025 Digital Buildings Analysis: 3-D Modeling, GIS Integration, and Visual Descriptions Using Gaussian Splatting, ChatGPT/Deepseek, and Google Maps Platform
abstract
We propose a Digital Building Analysis (DBA), a digital system for building-scale cloud-based data integration and data analytics. By connecting to cloud mapping platforms such as Google Map Platforms APIs, by leveraging state-of-the-art multi-agent Large Language Models data analysis using ChatGPT(4o) and Deepseek-V3/R1, and by using our Gaussian Splatting-based mesh extraction pipeline, our framework can retrieve a building’s 3D model, visual descriptions, and achieve cloud-based mapping integration with large language model-based data analytics using a building’s address, postal code, or geographic coordinates, and be easily extended to perform data analysis on other cloud-based data streams.
Kyle Gao, Dening Lu, Liangzhi Li 0002, Hongjie He 0003, Linlin Xu, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.7
2025 Optimizing Relative Radiometric Normalization: Minimizing Residual Distortions in Multispectral Bitemporal Images Using Trust-Region Reflective and Laplacian Pyramid Fusion
abstract
Accurate relative radiometric normalization (RRN) is important for reliable multitemporal remote sensing image analysis. Traditional methods often depend on coregistered image pairs, limiting their applicability with unregistered data. Keypoint-based RRN (KRRN) relaxes this constraint but remains affected by residual radiometric errors due to normalization inaccuracies and nonlinear effects. This letter introduces a refinement strategy that leverages the Trust-Region Reflective (TRR) algorithm to optimize normalization parameters, coupled with Laplacian Pyramid (LP) fusion for seamless image integration. Evaluation on four multispectral image pairs from different sensors (e.g., Landsat 8 and Sentinel-2, IRS and Landsat 5, Landsat 7 and SPOT-5, and UK-DMC2 and Landsat 5) and one pair from the same sensor (Sentinel-2) showed that our method reduces residual radiometric discrepancies, achieving up to 29% lower RMSE than some well-known models. The source code and datasets are available on GitHub: https://github.com/ArminMoghimi/Tensor-based-keypoint-detection.
Armin Moghimi, Turgay Celik, Ali Mohammadzadeh, Saied Pirasteh, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.5
2025 SECBNet: Semantic Segmentation-Enhanced Color Balance Network for Optical Satellite Images
abstract
Earth observation satellites can capture optical images under different temporal, climatic conditions, and platforms exhibit substantial differences in color and brightness, leading to poor visual experiences when synthesizing large-area optical satellite images. The related issue of color balancing has attracted considerable attention from researchers, yet challenges such as a lack of research data and sensitivity to model parameters persist. To address these problems, this article publishes a publicly open dataset and presents a semantic segmentation-enhanced color balance network (SECBNet). First, to mitigate the scarcity of research data, we develop a publicly available remote sensing image color balance dataset, Zhu Hai color balance image (ZHCBI), to support related research activities. Second, to improve semantic consistency between the color-balanced images and the target images, we design a dual-branch U-Net architecture guided by segmentation results and propose a novel segmentation feature loss function. Finally, to address issues of seams and unnatural transitions between blocks in segmented processing, we introduce a postprocessing module based on weighted averaging. We conducted comparative experiments and analyses with existing mainstream color balancing algorithms on the ZHCBI dataset. The results demonstrate that our proposed method achieves state-of-the-art color balancing quality, with significant improvement in visual effects and a higher peak signal-to-noise ratio (PSNR) (23.64 dB) compared with other mainstream methods.
Ziyi Chen 0001, Hanhuang Chen, Lujuan Gao, Dilong Li, Cheng Wang 0003, Linlin Xu, Somayeh Mollaee, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.8
2025 LVP: Leverage Virtual Points in Multimodal Early Fusion for 3-D Object Detection
abstract
Due to the sparsity and occlusion of point clouds, pure point cloud detection has limited effectiveness in detecting such samples. Researchers have been actively exploring the fusion of multimodal data, attempting to address the bottleneck issue based on LiDAR. In particular, virtual points, generated through depth completion from front-view RGB image, offer the potential for better integration with point clouds. Nevertheless, recent approaches fuse these two modalities in the region of interest (RoI), which limits the fusion effectiveness due to the inaccurate RoI region issue in the point cloud’s branch, especially in hard samples. To overcome it and unleash the potential of virtual points, while combining late fusion, we present leverage virtual point (LVP), a high-performance 3-D object detector which LVPs in early fusion to enhance the quality of RoI generation. LVP consists of three early fusion modules: virtual points painting (VPP), virtual points auxiliary (VPA), and virtual points completion (VPC) to achieve point-level fusion and global-level fusion. The integration of these modules effectively improves occlusion handling and improves the detection of distant small objects. In the KITTI benchmark, LVP achieves 85.45% 3-D mAP. As for large dataset nuScenes, we could improve the detection accuracy of large objects by compensating for errors in depth estimation. Without whistles and bells, these results establish LVP as an impressive solution for a 3-D outdoor object detection algorithm.
Yidong Chen 0006, Guo-Rong Cai, Ziying Song, Zhaoliang Liu, Binghui Zeng, Jonathan Li 0001, Zongyue Wang
IEEE Trans. Geosci. Remote. Sens.6
2025 Enhanced 3-D Urban Scene Reconstruction and Point Cloud Densification Using Gaussian Splatting and Google Earth Imagery
abstract
Three-dimensional urban scene reconstruction and modeling is a crucial research area. From a technical perspective, it is an interdisciplinary research area spanning computer vision, computer graphics, and photogrammetry. Its applications span across multiple disciplines including autonomous navigation with 3-D scene understanding, remote sensing/photogrammetry for the creation of 3-D maps from aerial/drone/satellite images, geographic information systems with urban digital twins, augmented and virtual reality with photorealistic scene reconstructions. Using Google Earth imagery, we create a 3-D Gaussian splatting (3DGS) model of the Waterloo region centered on the University of Waterloo, and are able to achieve view-synthesis results far exceeding previous 3-D view-synthesis results based on neural radiance fields (NeRFs)which we demonstrate in our benchmark. We also retrieve the 3-D geometry of the scene using the 3-D point cloud extracted from the 3DGS model, thereby reconstructing both the 3-D geometry and photorealistic lighting of the large-scale urban scene.
Kyle Gao, Dening Lu, Hongjie He 0003, Linlin Xu, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 Point-SCT: A Multiscale Spatial Convolution-Swin Transformer Network for Point Cloud Ground Filtering in Complex Mountainous Terrains
abstract
Deep learning-based point cloud segmentation methods have been extensively explored, but the majority focus either on local or global feature learning, with few integrating both. These integrated approaches have not been sufficiently explored in complex mountainous scenes with low feature heterogeneity. To address this gap, we propose a novel point-based multi-scale spatial Convolution-Swin Transformer network (Point-SCT). Point-SCT combines convolutional local geometric detail capture with global relationship modeling via dynamic window interactions in Transformer, enhancing ground filtering accuracy in challenging mountainous scenes. The encoder incorporates convolution-based Multi-Scale Local Feature Aggregation (MLFA) approach, integrating Local Geometric Feature Encoding (LGSE) and Diluent Pooling (DP) strategies to effectively aggregate local detailed geometric features while suppressing irrelevant feature vectors and enhancing the representation of low-heterogeneity feature. Additionally, the dynamic spatial window strategy within the Transformer facilitates the capture of long-range feature dependencies. To mitigate noise introduced by RGB in point cloud overlays and sharpen geometric distinctions between the ground and low-lying vegetation, we introduce Boundary Detector, Curvature, and Average Elevation (BCE) as prior inputs, replacing RGB. Finally, quantitative and qualitative analyses of Point-SCT are conducted on an airborne laser scanning (ALS) dataset from a mountainous area, with ablation studies validating the effectiveness of LGSE, DP and BCE. The comprehensive experiments demonstrate that Point-SCT robustly segments ground points in complex mountainous scenes, achieving state-of-the-art levels of accuracy and generalization.
Fuquan Tang, Lingfei Ma, Nur Intan Raihana Ruhaiyem, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.7
2025 Exploring Token Serialization for Mamba-Based LiDAR Point Cloud Segmentation
abstract
LiDAR point cloud segmentation has increasingly benefited from the application of Mamba-based models. However, unordered and irregular natures of point clouds necessitates serialization, which significantly impacts the performance of Mamba-based methods. This paper explores the critical role of token serialization in Mamba-based point cloud processing, using the pure Mamba network, PointMamba, as the baseline. We systematically investigated existing point cloud serialization methods, evaluating their performance on two challenging LiDAR datasets: the airborne MultiSpectral LiDAR (MS-LiDAR) dataset and the aerial DALES dataset. To explore the inherent factors of serialization contributing to Mamba’s performance, we design novel indicators for serialization quality, focusing on spatial and semantic proximity. These indicators are validated across all datasets, offering a valuable reference and guidance for advancing token serialization in Mamba-based point cloud processing. Guided by these indicators, we proposed a new point cloud serialization method that integrates spatial and semantic features through a weighted comprehensive distance matrix. The proposed method achieves superior accuracy on both LiDAR datasets, surpassing existing approaches, and establishes a strong foundation for advancing Mamba-based point cloud processing.
Dening Lu, Kyle Gao, Jonathan Li 0001, Dedong Zhang, Linlin Xu
IEEE Trans. Geosci. Remote. Sens.3
2025 SLAM-TSM: Enhanced Indoor LiDAR SLAM With Total Station Measurements for Accurate Trajectory Estimation
abstract
Simultaneous Localization and Mapping (SLAM) is a crucial task in various domains, including intelligent robotics, computer vision, and indoor navigation. Accurate and robust trajectory estimation is especially challenging in indoor environments due to the presence of feature-poor or repetitive scenes, limited visibility, and dynamic objects. Obtaining highly accurate machine platform odometry is also an important basis for solving the “last mile” problem in intelligent transportation. This paper proposes a novel algorithm that combines LiDAR-based SLAM, total station measurements, and graph optimization to optimize the robot’s trajectory in indoor environments. By integrating highly accurate positional data from total station measurements as additional constraints, the proposed method enhances the performance of indoor LiDAR SLAM, effectively addressing the challenges of drift and trajectory offsets. Moreover, the proposed algorithm can provide trajectory optimization even in the absence of loop closure detection, making it more robust and suitable for a broader range of indoor environments. Experimental results validated the effectiveness of the proposed approach in reducing drift and improving trajectory estimation for low-cost indoor LiDAR devices, demonstrating its potential in various applications such as autonomous navigation, facility management, and augmented reality. It provides targeted ground truth for autonomous driving of machine platforms in intelligent transportation scenarios.
Dedong Zhang, Weikai Tan, John S. Zelek, Lingfei Ma, Jonathan Li 0001
IEEE Trans. Intell. Transp. Syst.5
2024 Detection of Small Objects from UAV Imagery via an Improved Swin Transformer
abstract
Automated detection of small objects such as vehicles in images of complex urban environments taken by unmanned aerial vehicles (UAVs) is one of the most challenging tasks in computer vision and remote sensing communities. Convolutional neural networks (CNNs)-based deep learning models have been widely used to automatically detect objects in UAV images given their high performance. However, their detection accuracy is still unsatisfactory, particularly when it comes to small objects, due to the shortcomings of CNNs. Therefore, in this study, we propose a Swin Transformer-based model that incorporates convolutions with the Swin Transformer to extract more local information, mitigating the problem of small object detection from complex backgrounds in UAV images and further improving the detection accuracy. By using the Swin Transformer, our model leverages both the local feature extraction of convolutions and the global feature modeling of transformers. The framework comprises two primary modules: a Local Context Enhancement (LCE) module and a Residual U-Feature Pyramid Network (RSU-FPN) module. Additionally, it incorporates a loss function that combines L1 loss with Normalized Gaussian Wasserstein Distance. Our experimental results obtained on the UAV Detection and Tracking (UAVDT) dataset indicated that our proposed method increased the average precision (AP) by 21.6%, 22.3% and 25.5% over Cascade Region-based CNN (R-CNN), Faster Region based CNN (R-CNN) with ResNet-50 and with Pyramid Vision Transformer (PVT) B0, and Dynamic R-CNN detectors, respectively, indicating its effectiveness and reliability on small object detection from UAV images.
Weidong Liang, Jingtian Tan, Hongjie He 0003, Hongzhang Xu, Jonathan Li 0001
IGARSS5
2024 UnderstAnding Bag of Tricks of Deep Learning-Based Semantic Segmentation in Pavement Crack Detection
abstract
The rapid development of deep learning has significantly enhanced the performance of models in the detection of pavement cracks, thereby facilitating the deployment of deep learning-based approaches into real-world applications. Nevertheless, it is worth noting that deep learning-based crack detection models represent complex amalgamations of deep learning networks and model training strategies, with the latter frequently being overlooked. Therefore, in this paper, we focus on various techniques in data augmentation, and model deployment stages that are commonly employed in deep learning-based semantic segmentation models. Through extensive experiments, the effectiveness of these techniques in crack detection is evaluated, aiming to provide guidance for subsequent crack detection experiments and project implementations. Consequently, the experiments demonstrate data augmentation methods such as color jittering and CutMix can effectively improve model performance by altering the distribution of the training dataset. Additionally, in case of crack datasets with limited samples and severe class imbalance, loss function selection and pre-training weights can be crucial in model deployment.
Zhengsen Xu, Hiayan Guan, Hongjie He 0003, Jonathan Li 0001
IGARSS4
2024 DBARCT: Road Extraction Based on Double-Branch Architecture and Random Block Coding Transformer
abstract
Although transformer models are main network architectures for the delineation of roads from remote sensing imagery, they have critical limitations due to their regular patch mechanism and inefficiency in local information learning. To address these limitations for enhanced road extraction, this letter presents a novel double-branch architecture and random block coding transformer (DBARCT), with the following contributions. First, to improve local spatial details’ learning, we integrate transformer with convolutional neural network (CNN) into a novel dual-branch encoder-decoder architecture, such that the resulting model is efficient at learning both the local edge information and the global context information that are highly complementary for accurate road extraction. Second, to additionally augment the learning of global contextual information, we integrate the regular patching approach in traditional transformer models with a new irregular patching approach, such that it can better capture the global spatial information correlations that might be ignored by the regular patching approach. Third, an array of tests was carried out to meticulously scrutinize the efficacy of the fundamental elements of the suggested model. The empirical findings reveal that the intersection over union (IoU) metric attained by the proposed methodology on the LRSNY dataset stands at 88.53%, thereby corroborating the efficacy and preeminence of our approach in tasks related to road extraction.
Ziyi Chen 0001, Yucai Chen, Lujuan Gao, Dilong Li, Linlin Xu, Jonathan Li 0001, Cheng Wang 0003, Yewang Chen
IEEE Geosci. Remote. Sens. Lett.6
2024 Neighborhood Attention Makes the Encoder of ResUNet Stronger for Accurate Road Extraction
abstract
In the domain of remote sensing image interpretation, road extraction from high-resolution aerial imagery has already been a hot research topic. Although deep CNNs have presented excellent results for semantic segmentation, the efficiency and capabilities of vision transformers are yet to be fully researched. As such, for accurate road extraction, a deep semantic segmentation neural network that utilizes the abilities of residual learning, HetConvs, UNet, and vision transformers, which is called ResUNetFormer, is proposed in this letter. The developed ResUNetFormer is evaluated on various cutting-edge deep learning-based road extraction techniques on the public Massachusetts road dataset. Statistical and visual results demonstrate the superiority of the ResUNetFormer over the state-of-the-art CNNs and vision transformers for segmentation. The code will be made available publicly at https://github.com/aj1365/ResUNetFormer.
Ali Jamali, Swalpa Kumar Roy, Jonathan Li 0001, Pedram Ghamisi
IEEE Geosci. Remote. Sens. Lett.3
2024 Local Enhanced Transformer Networks for Land Cover Classification With Airborne Multispectral LiDAR Data
abstract
Transformer networks have demonstrated remarkable performance in point cloud processing tasks. However, balancing local feature aggregation with long-range dependency modeling remains a challenging issue. In this work we present a local enhanced Transformer network (LETNet) for land cover classification with multispectral LiDAR data. Specifically, we first rethink position encoding in 3D Transformers and design a novel feature encoding module that embeds comprehensive geometric and semantic information, serving a similar purpose. Then, the proposed local enhanced Transformer module is used to capture the accurate global attention weights and refine the features. Finally, to effectively extract and integrate global features across various scales, an attention-based pooling module is introduced. This module extracts global features from each encoder and decoder layer and constructs a feature pyramid to fuse these multi-scale global features. Both quantitative assessments and comparative analyses demonstrate the competitive capability and advanced performance of the LETNet in land cover classification task.
Dilong Li, Shenghong Zheng, Ziyi Chen 0001, Jonathan Li 0001, Jixiang Du
IEEE Geosci. Remote. Sens. Lett.4
2024 HigherNet-DST: Higher-Resolution Network With Dynamic Scale Training for Rooftop Delineation
abstract
High-definition (HD) maps of building rooftops or footprints are important for urban application and disaster management. Rapid creation of such HD maps through rooftop delineation at the city scale using high-resolution satellite and aerial images with deep learning methods has become feasible and drawn much attention. However, the scale variance issue in rooftop delineation limited the overall performance. Existing methods exhibit considerably poor performance in rooftop delineation of small buildings. In this paper, we propose a new method, namely the Higher Resolution Network with Dynamic Scale Training (HigherNet-DST) to overcome the scale variance problem in rooftop delineation. Specifically, the DST is applied in the model training phase to reduce the negative impact of scale variance. Then, a scale-aware backbone, namely the Higher Resolution Network, is adopted to enhance the feature representation. Finally, the high-resolution supervision targets are used to further boost the delineation performance. Our method was tested on four publicly accessible building datasets and the results demonstrated that our method achieved the highest performance in rooftop delineation among the existing methods. Extensive experiments showed the superior performance of our method with an AP of 68.5% on the AICrowd Building Dataset and an IoU of 82.6% On the Inria Building Dataset, respectively, which surpassed many state-of-the-art (SOTA) methods. On the WHU Building Dataset and the Waterloo Building Dataset, our method also achieved the highest performance among the benchmarked methods, showing the high performance of our method for building boundary delineation.
Hongjie He 0003, Lingfei Ma, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.3
2024 High-Precision Intracity Temperature Estimation Based on Generated Point Clouds
abstract
Estimation of urban surface temperature is crucial for urban planning and emergency management. Due to the complexity of intracity structures, it is very challenging to acquire satisfied prediction errors of the land surface temperature (LST) at very high resolution, like 60-by-60 m. Considering this, we propose a low-cost method for generating urban point clouds via readily accessible city data. Then we design an efficient descriptor, geofeature distribution matrix (GFDM) to describe the complex intracity structure. Using GFDM, we introduce a 3-D urban structure guided temperature prediction network (3D-UP Net) to capture the complex relationship between urban structure, upper atmospheric conditions, and surface temperature. The proposed 3D-UP Net is generalizable, capable of predicting future surface temperature for existing cities and even for those that are planned. Experiments conducted in multiple regions of China demonstrate that our method’s error is less than 1.5 K (in most cases) at a high resolution (60-by-60 m).
Guanjie Huang, KaiQun He, Xiaoyue Lyu, Hang Shu, Lei Zhao 0023, Jonathan Li 0001, Cheng Wang 0003
IEEE Trans. Geosci. Remote. Sens.7
2024 A Novel Weighted Ensemble Transferred U-Net Based Model (WETUM) for Postearthquake Building Damage Assessment From UAV Data: A Comparison of Deep Learning- and Machine Learning-Based Approaches
abstract
Nowadays, unmanned aerial vehicle (UAV) remote sensing data are key operational sources used to produce a reliable building damage map (BDM), which is of great importance in instant response and rescue operations after earthquakes. The present study proposes a novel weighted ensemble transferred U-Net-based model (WETUM) consisting of two major steps to create a reliable binary BDM using UAV data. In the first step of the proposed approach, three individual initial BDMs are predicted by three pre-trained U-Net-based composite networks. In the second step, these three individual predictions are linearly integrated through a proposed grid search technique so that an optimized hybrid BDM (OHBDM) incorporating complementary damage information is made. The proposed WETUM was then compared with several conventional deep learning (DL) and machine learning (ML) models. The models were compared across two pivotal scenarios, addressing the impact of diverse feature sets on model performance and generalizability. Specifically, the first scenario focused solely on spectral features, while the second incorporated both spectral and geometrical features. To make the comparisons, this study conducted empirical analyses using UAV spectral and geometrical data acquired over Sarpol-e Zahab, Iran. The experimental findings showed that the synergic use of spectral and geometrical data boosted both DL- and ML-based approaches in damage detection. Moreover, the proposed WETUM with DDR values of 65.22 and 78.26 (%), respectively, for the first and second scenarios, outperformed all the compared methods. Notably, WETUM with only spectral data outperformed the random forest (RF) classifier equipped with many hand-crafted spectral and geometrical features, indicating the highest potential and generalizability of the proposed WETUM for building damage evaluation in a new unseen earthquake-affected area.
Ehsan Khankeshizadeh, Ali Mohammadzadeh, Hossein Arefi, Amin Mohsenifar, Saied Pirasteh, En Fan, Huxiong Li, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.8
2024 A Multitask CNN-Transformer Network for Semantic Change Detection From Bitemporal Remote Sensing Images
abstract
Bitemporal remote sensing (RS) semantic change detection (SCD) involves discerning and categorizing changes in the same geographical area across two RS images taken at different times. High-performance SCD approaches typically address this task using multitask networks that simultaneously handle binary change detection (BCD) and semantic segmentation (SS). Despite significant advancements in SCD research, constructing a multitask network that fully explores the correlation between BCD and SS remains challenging. To address this, we propose a novel approach called the multitask CNN-transformer network (MCTNet), tailored for SCD using bitemporal RS images. Our Siamese network simultaneously tackles SS and BCD via three subnetworks: two for SS and one for BCD. The methodology begins with a multiscale convolutional neural network (CNN) extracting local features from input images, and converting them into tokens. A Transformer module with an encoder-decoder architecture then captures long-range dependencies among these visual tokens. The extracted features are subsequently passed to multitask heads, generating predicted outputs. To ensure that the BCD results remain consistent regardless of the order of images in the input pair, we introduce spatiotemporal feature learning (SFL), enabling the acquisition of temporal-symmetric representations for BCD. Extensive experimental validation on the WHU-CD, SECOND, and HRSCD datasets demonstrates the effectiveness and efficiency of MCTNet for both SS and BCD tasks. The source code for this article will be published on GitHub in the futurehttps://github.com/kangziwen1/MCTNet.
Ziwen Kang, Yiyuan Lin, Yongtao Yu, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.6
2024 3DGTN: 3-D Dual-Attention GLocal Transformer Network for Point Cloud Classification and Segmentation
abstract
Although the application of Transformers to 3-D point cloud processing has achieved significant progress and success, it is still challenging for existing 3-D Transformer methods to efficiently and accurately learn both valuable global and local features for improved applications. This article presents a novel point cloud representational learning network, called 3-D Dual Self-attention global local (GLocal) Transformer Network (3DGTN), for improved feature learning in both classification and segmentation tasks, with the following key contributions. First, a GLocal feature learning (GFL) block with the dual self-attention mechanism [i.e., a novel point-patch self-attention, called PPSA, and a channel-wise self-attention (CSA)] is designed to efficiently learn the global and local context information. Second, the GFL block is integrated with a multiscale Graph Convolution-based local feature aggregation (LFA) block, leading to a GLocal information extraction module that can efficiently capture critical information. Third, a series of GLocal modules are used to construct a new hierarchical encoder–decoder structure to enable the learning of information in different scales in a hierarchical manner. The proposed framework is evaluated on both classification and segmentation datasets, demonstrating that the proposed method is capable of outperforming many state-of-the-art methods on both synthetic and LiDAR data. Our code has been released athttps://github.com/d62lu/3DGTN.
Dening Lu, Kyle Gao, Qian Xie 0001, Linlin Xu, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 Point Transformer-Based Salient Object Detection Network for 3-D Measurement Point Clouds
abstract
While salient object detection (SOD) on 2D images has been extensively studied, there is very little SOD work on 3D measurement surfaces. We propose an effective point transformer-based SOD network for 3D measurement point clouds, termed PSOD-Net. PSOD-Net is an encoder-decoder network that takes full advantage of transformers to model the contextual information in both multi-scale point- and scene-wise manners. In the encoder, we develop a Point Context Transformer (PCT) module to capture region contextual features at the point level; PCT contains two different transformers to excavate the relationship among points. In the decoder, we develop a Scene Context Transformer (SCT) module to learn context representations at the scene level; SCT contains both Upsampling-and-Transformer blocks and Multi-context Aggregation units to integrate the global semantic and multi-level features from the encoder into the global scene context. Experiments show clear improvements of PSOD-Net over its competitors and validate that PSOD-Net is more robust to challenging cases such as small objects, multiple objects, and objects with complex structures. Code is available at: https://github.com/ZeyongWei/PSOD-Net.
Zeyong Wei, Baian Chen, Weiming Wang 0002, Honghua Chen, Mingqiang Wei, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.6
2024 Hybrid Network-Based Automatic Seamline Detection for Orthophoto Mosaicking
abstract
Seamline detection is a crucial procedure for orthophoto mosaicking. To eliminate the seam effect caused by geometric discontinuities, seamlines must avoid crossing areas containing the obvious ground object, for which manual processing is usually required. Many existing seamline detection methods can generate seamlines bypassing most of obvious ground objects but always take pixel-level computation which may consume much time. To address this problem, this paper presents a seamline detection approach based on a hybrid network search. First, without auxiliary data, the semi-global block matching (SGBM) algorithm was adopted to generate a disparity map for pairwise orthophoto overlapping area. By using adaptive threshold segmentation, a binary cost map containing the ground objects was obtained. Subsequently, a hybrid network was constructed by edge points and uniform points extracted on the cost map. Finally, seamlines were detected by search on this network-based graph. The essential contribution of the proposed method is that the seamline is searched on a sparse hybrid network instead of a raster cost map. Thus, computational complexity can be significantly decreased and produce fine-tuned seamlines. A Series of comparison experiments were carried out between the proposed and well-established methods, using two benchmark datasets with different characteristics. The comparison results demonstrated that the proposed method could generate high-quality seamlines in terms of visual comparison and statistical evaluation. Moreover, the processing speed has a nearly tenfold improvement compared with the control group methods.
Wei Yuan 0004, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.3
2024 RdmkNet & Toronto-RDMK: Large-Scale Datasets for Road Marking Classification and Segmentation
abstract
Effective road marking classification and segmentation play a pivotal role in advancing vehicle-to-everything (V2X) applications and refining road inventory databases. However, the irregular data formats and unordered permutation modes of 3D point clouds, along with the limited availability of large-scale datasets with point-level annotations, remain significant obstacles to designing deep learning-based networks with superior performance. To address these challenges, this paper proposes a novel multi-level feature optimization network structure, named MFPNet, and introduces two point cloud benchmarks, RdmkNet and Toronto-Rdmk, for road marking classification and segmentation in intricate urban environments. MFPNet is composed of three integral modules. First, the M-transformer module, consisting of three transformers obtained from different channels, fully captures rich point cloud background information and long-distance dependencies between objects. Then, the feature pooling aggregation module uses parallel structured pooling attention mechanisms to aggregate features captured by the M-transformer module, while the prediction refinement module further enhances the acquisition of semantic features. Comparative studies indicate that MFPNet can be embedded into general deep learning networks without changing their original network structures, significantly improving the accuracy of multiple baseline networks. Furthermore, extensive experiments demonstrate that the two newly-developed point cloud datasets are meaningful for road marking classification and segmentation tasks, contributing to the development of autonomous driving.
Jing Du 0007, Lingfei Ma, Jing Li 0040, Nannan Qin, John S. Zelek, Haiyan Guan, Jonathan Li 0001
IEEE Trans. Intell. Transp. Syst.7
2024 Crack-U2Net: Multiscale Feature Learning Network for Pavement Crack Detection From Large-Scale MLS Point Clouds
abstract
Deep learning-based algorithms detect pavement cracks in an end-to-end manner from Mobile Laser Scanning (MLS) point clouds, achieving impressive results. However, the accuracy of existing methods still has room to improve due to the difficulty of effectively encoding multiscale features and the limited training data. In this paper, we propose a novel pavement crack detection framework, Crack-U2Net, which innovatively incorporates a two-level nested U-Net architecture for feature learning. This design enables the learning of intra-stage multiscale features without introducing significant memory and computation costs, resulting in substantial improvements in accuracy. Moreover, to solve the challenge of insufficient training data, we propose a Geometry-based Data Augmentation (GDA) strategy, aiming to expand the pavement dataset while preserving the pavement geometry. Extensive experiments on the Qinghai-Tibet Highway point cloud dataset demonstrate the higher accuracy and efficiency of Crack-U2Net over the state-of-the-art methods, achieving an average precision, recall, F$1\text - $score, and accuracy of 83.8%, 77.6%, 80.1%, and 95.8%, respectively.
Huifang Feng 0002, Wen Li 0005, Lingfei Ma, Yiping Chen 0002, Haiyan Guan, Yongtao Yu, José Marcato Junior, Jonathan Li 0001
IEEE Trans. Intell. Transp. Syst.8
2023 Enhancing Spatial Resolution of Building Datasets Using Transformer-Based Single-Image Super-Resolution
abstract
The spatial resolution of Earth Observation (EO) images plays a key role in building footprint extraction. For the spatial resolution enhancement, deep learning-based image super-resolution methods have been widely used due to their remarkable performance. Transformer-based networks are effective and has drawn much attention in computer vision but underutilized in remote sensing, especially for super-resolving building datasets. Therefore, in this paper, we developed a novel transformer-based Single-Image Super-Resolution (SISR) method, named Pyramid Vision Transformer-Residual Feature Aggregation Network (PVT_RFANet), to improve the spatial resolution of building datasets. Specifically, the PVT v2 network was embedded into our Momentum Spatial-Channel Attention Residual Feature Aggregation Network (MSCA-RFANet). Moreover we conducted a comparative study to compare our method with Bicubic interpolation (BI), Super-resolution Convolutional Neural Network (SRCNN), Deep Recursive Residual Network (DRRN), SRResNet, and MSCA-RFANet. Using Peak Signal-Noise Ratio (PSNR) and Similarity Structure Index Measurement (SSIM) as the evaluation metrics, our method showed highest performance with the PSNR of 22.01 dB and the SSIM of 0.50 on the WHU Building Dataset, which demonstrated the superior performance of the proposed method.
Yuwei Cai, Hongjie He 0003, Zhimeng He, Michael A. Chapman, Jing Li 0040, Lingfei Ma, Jonathan Li 0001
IGARSS7
2023 Research on Fast Detection Method of Wind Turbine in Remote Sensing Image Land Area Based on Yolo
abstract
With the development of the social economy, wind turbines are taking up a larger and larger share of new energy sources. The detection of the number and spatial distribution of wind turbines in remotely sensed images holds great scientific significance. Wind turbines are difficult to identify in remote sensing images therefore, a fast detection method based on deep learning is proposed. First, we extract potential wind turbine candidate regions from wind speed, slope, and land use data. Second, the YOLO v5 model was trained using our labeled wind turbine detection dataset. Finally, the images of the candidate regions were used for wind turbine detection using the trained optimal model. The proposed method was demonstrated to have a recall of 94.87% and an accuracy of 82.04% through the experimental results. The proposed method for wind turbine detection is not only reasonable and effective but also offers a heightened level of efficiency.
Deliang Chen, Taotao Cheng, Kyle Gao, Sarah Narges Fatholahi, Jonathan Li 0001
IGARSS6
2023 Natural Language Aided Remote Sensing Image Few-Shot Classification
abstract
We aim to improve the efficiency of traditional deep learning methods for remote sensing by reducing the reliance on annotated data and minimizing training time. Instead of using large-scale unimodal remote sensing image datasets for pre-training, we propose the use of multimodal data (text-image pairs), which we believe to be more effective. To enhance the model's generalization performance in the remote sensing domain and achieve accurate remote sensing image scene classification, we employ the Feature Adaptive Embedding Module. For this purpose, we introduce a cross-modal comparison learning network that is based on openly accessible generalized datasets. This network is capable of recognizing specific photo scenarios from remote sensing photographs, maximizing the accuracy of classification.
Deliang Chen, Jianbo Xiao, Kyle Gao, Sarah Narges Fatholahi, Jonathan Li 0001
IGARSS6
2023 Nighttime Light Missing Data Retrieval Using Modis Version 6 Satellite Data and Mask Dilated Partial Convolutional Neural Network
abstract
Nighttime Lights (NTLs) remote sensing imagery contains tremendous information and has been shown to accurately predict a region’s human dynamics, economic health and energy consumption. Despite its usefulness, NTLs imagery is less widely available than other remote sensing data modalities. Several challenges appear when attempting to reconstruct NTLs data, either from other data modalities or existing NTLs data. These include complex non-linear relationships between NTLs and multispectral bands, non-matching spatial and temporal coverage, and different atmospheric and cloud conditions. This study attempts to create an out-of-the-box model that compensates for missing NTLs data using widely available daytime data in a broadly generalizable manner. The proposed project has two objectives: the construction of an image-to-image dataset mapping daytime multispectral images (MODIS V6 Land Surface Reflectance, MODIS V6 Land Cover, MODIS V6 Vegetation Indices) to NTLs images, and the reconstruction of NTLs data using deep learning techniques by researching, creating, and employing the state-of-the-art architecture of the Mask Partial Convolutional Neural Network in conjunction with dilated convolutions. The project will facilitate the training of new models for predicting missing NTLs and make NTLs data more accessible for future remote sensing research.
Xuanchen Liu, Shuxin Qiao, Kyle Gao, Hongjie He 0003, Lingfei Ma, Jonathan Li 0001
IGARSS6
2023 Tree Species Classfifcation Using Deep Learning Based 3d Point Cloud Transformer on Airborne Lidar Data
abstract
This paper applied a transformer based deep learning model 3D Point Cloud Transformer (3DPCT) to conduct a tree species classification of Airborne LiDAR data. There are a total 1291 single tree point clouds of 11 different species from coniferous and deciduous used in this paper. The model integrated the local and global feature learning modules from both pointwise and channel-wise, which provide promising results of tree species classification. We also investigate by adding more channels the classification results can be improved. Different number of points per each sample as the model input also deliver different accuracy. The highest overall accuracy of 11 categories classification achieved 86.1%, and precision and recall of each category provide more directions of future study.
Dening Lu, Weikai Tan, Yiping Chen 0002, Jonathan Li 0001
IGARSS5
2023 Automatic Image-to-Color Point Cloud Cross-modal Registration Based on Graph Neural Networks and Iterative Reprojection
abstract
Image-to-Color point cloud registration establishes a connection between two-dimensional image data and three-dimensional point cloud data, and plays a vital role in the intelligent city, autonomous driving, and robotics field. However, it is still challenging to automatically register an image and its surrounding color point clouds together, due to the special application scenario, unknown camera intrinsic parameters, lack of train data, and a large amount of noise. To overcome those issues, we propose an iterative reprojection architecture that automatically acquires the matched 2D-3D keypoints pairs between the image and the color point clouds by graph optimization method and mapping transfer first, then completes registration by Alternating Direction Method of Multipliers (ADMM). Experiments results show that the proposed method is more accurate than the manual way.
Shanxin Zhang, Jiande Sun 0001, Cheng Wang 0003, Jonathan Li 0001
ISCAS6
2023 BrGAN: Blur Resist Generative Adversarial Network With Multiple Joint Dilated Residual Convolutions for Chlorophyll Color Image Restoration
abstract
This paper presents a Blur Resist Generative Adversarial Network (GAN) (BrGAN) with multiple joint dilated residual convolutions for chlorophyll image restoration of the Geostationary Ocean Color Imager (GOCI). First, a publicly available dataset was built to support this study. Second, a multiple attention perception mechanism and a multiple joint dilated residual convolution module was proposed to cope with the challenge of large missing areas in GOCI chlorophyll images. Third, a patch GAN based discrimination module was proposed to avoid the restored areas with generating mosaic and shadows. Our experimental results demonstrate that the BrGAN can reach 37.06 in the peak signal-to-noise ratio (PSNR) and 0.0485 in the Learned Perceptual Image Patch Similarity (LPIPS), respectively. The comparative study shows that the BrGAN achieves the highest effectiveness and advancement among other seven state-of-the-art methods.
Ziyi Chen 0001, Yuhua Luo, Yiping Chen 0002, Jing Wang 0049, Dilong Li, Kyle Gao, Cheng Wang 0003, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.8
2023 SAR-Optical Image Matching With Semantic Position Probability Distribution
abstract
We propose a deep learning framework of Semantic Position Probability Distribution for SAR-optical image matching, termed as SPPD. Unlike the pixel-by-pixel searching matching method, a correspondence is directly obtained by an outputted matching position probability distribution. First, multiscale pyramidal features are created for each pixel in the SAR and optical images by using two weight-sharing ResNet-50 + Feature Pyramid Network (FPN) networks. The features containing high-level semantic information are then embedded into the proposed image Position Attention Module to obtain the spatial position dependencies between two images. Then, we present a loss function for semantic position matching to optimize the network from both semantic information and pixel alignment perspectives, converting the probability distribution of semantic matching positions into a point-to-point matching problem. In this paper, the SAR and optical images are set as the sensed and reference images. The effects of different image sizes, training label types, and loss function weights on matching accuracy are explored to obtain the optimal parameter settings for matching. The experimental results show that the proposed method is insensitive to image deformation and achieves cross-modal matching for SAR-optical images with high accuracy compared with the best matching method on different scene images, with several orders of magnitude faster inferences time.
Liangzhi Li 0002, Kyle Gao, Hongjie He 0003, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.7
2023 A Click-Based Interactive Segmentation Network for Point Clouds
abstract
Interactive segmentation plays an essential role in several tasks involving point clouds. However, existing methods suffer from low segmentation accuracy and cannot adjust the segmentation results according to the user’s personal demands. This paper presents a novel deep learning-based interactive segmentation method, named Click Rough Segmentation Network (CRSNet), designed to handle point clouds. The method allows users to iteratively click to segment interesting objects. CRSNet consists of two key parts: a CRS module and a feature extraction module. First, the CRS module transforms click operations into an appropriate representation to input into the feature extraction module. The CRS module takes raw point clouds and click operations as input and outputs 3D Gaussian vectors and roughly segmented blocks, which adapt to different-sized and densely-distributed objects in complex environments. Second, the feature extraction module, which uses a novel mix loss-based analysis algorithm, extracts deep features and obtains instance segmentation results. The module is highly compatible because its backbones can be replaced by different deep learning architectures. Experimental results on the KITTI, Apolloscape, Roadmarking, Scannet, and SemanticKITTI datasets show that our method outperforms state-of-the-art semantic segmentation methods with one click. Moreover, our method can generalize well to unseen objects and datasets.
Wentao Sun, Yiping Chen 0002, Huxiong Li, José Marcato Junior, Wesley Nunes Gonçalves, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.7
2023 LCE-NET: Contour Extraction for Large-Scale 3-D Point Clouds
abstract
The contours, one of the most significant human perceptual features, have a significant impact on point cloud processing. In urban scenes, contour extraction is quite challenging due to the enormous number of unstructured and irregular points (typically greater than 107points). In this paper, we propose a Large-scale 3D point cloud Contour Extraction Network (LCE-NET) to generate contours consistent with human perception of outdoor scenes. To our knowledge, it is the first time that an end-to-end learning-based framework has been proposed for contour extraction on point cloud at this scale. The proposed LCE-Net is essentially a two-phase system. In the first phase, potential vertexes are detected from the input point cloud by the vertex detection module, then in the second phase, a designed overcomplete line proposal set is generated, and invalid line segments are further suppressed by the line proposal discrimination module. The two phases are jointly trained by a uniform loss function to promote the information interchange, thus leading to extraction results with satisfied accurate and false alarm ratings. Since there is hardly any available dataset with labeled contours for the large-scale outdoor scene, we open sourced SemanticLine, the first dataset for large-scale point clouds with labeled contour information, based on re-annotation of previous mapping level point cloud dataset semantic3D. Experimental results demonstrate that LCE-NET can effectively extract parametric contour lines from large-scale point clouds of urban scenes. Additionally, it outperforms the state-of-the-art approaches. The code will be open source on GitHub soon.
Binjie Chen, Yunzhou Xia, Hanyun Guo, Yunuo Yang, Weiquan Liu, Cheng Wang 0003, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.8
2023 CF-YOLO: Cross Fusion YOLO for Object Detection in Adverse Weather With a High-Quality Real Snow Dataset
abstract
Snow is one of the toughest adverse weather conditions for object detection (OD). Currently, not only there is a lack of snowy OD datasets to train cutting-edge detectors, but also these detectors have difficulties of learning latent information beneficial for detection in snow. To alleviate the two above problems, we first establish a real-world snowy OD dataset, named RSOD. Besides, we develop an unsupervised training strategy with a distinctive activation function, called$Peak Act$, to quantitatively evaluate the effect of snow on each object. Peak Act helps grade the images in RSOD into four-difficulty levels. To our knowledge, RSOD is the first quantitatively evaluated and graded real-world snowy OD dataset. Then, we propose a novel Cross Fusion (CF) block to construct a lightweight OD network based on YOLOv5s (called CF-YOLO). CF is a plug-and-play feature aggregation module, which integrates the advantages of Feature Pyramid Network and Path Aggregation Network in a simpler yet more flexible form. Both RSOD and CF lead our CF-YOLO to possess an optimization ability for OD in real-world snow. That is, CF-YOLO can handle unfavorable detection problems of vagueness, distortion and covering of snow. Experiments show that our CF-YOLO achieves better detection results on RSOD, compared to SOTAs. The code and dataset are available athttps://github.com/qqding77/CF-YOLO-and-RSOD.
Qiqi Ding, Peng Li 0064, Xuefeng Yan 0001, Ding Shi, Luming Liang, Weiming Wang 0002, Haoran Xie 0001, Jonathan Li 0001, Mingqiang Wei
IEEE Trans. Intell. Transp. Syst.8
2022 Automated Detection of Oil/Gas Well Sites Detection from Multi-Source High Spatial Resolution Images
abstract
With the development of oil/gas production, its adverse impact has drawn much attention. Therefore, automated detecting oil/gas well sites become important. Current research has been focused on detection using sole source RGB images, which was not the common case in remote sensing. In this study, we explored the use of a combination of the Residual Channel Attention Network (RCAN) and the state-of-the-art object detection method, You Only Look Once (YOLO) v4, to detect oil/gas well sites from multi-sensor images. To testify the feasibility of the combination, we selected 18 RapidEye and 15 WorldView images which cover the oil sands area in Alberta, Canada. We applied a pre-trained RCAN to unify the spatial resolution of different images to 2m/pixel. To maximize the feature space, we preserved 5 bands which were available in both images. YOLO v4 was applied on cropped image patches to detect oil/gas well sites. The experiment results showed that using the framework proposed in this study, oil/gas well sites can be localized accurately although the bounding boxes of the sites may not perfectly align with the objects.
Hongjie He 0003, Hongzhang Xu, Michael A. Chapman, Yiping Chen 0002, Jonathan Li 0001
IGARSS6
2022 Object Classification in Point Cloud Via Conditional Adversarial Domain Adaptation for Forest Inventory
abstract
Recently, laser scanning system is widely used to accurately predict forest inventory attributes. In this work, we propose an efficient framework including Feature Extractor Module and Conditional Adversarial Module for object classification in 3D heterogeneous point clouds. The Feature Extractor Module can aggregate features of objects in point clouds. For domain adaptation task, the Conditional Adversarial Module is proposed to minimize the discrepancy of source and target domains. Since there is no common evaluation benchmark for object classification in 3D point cloud, we build a bench-mark of six tasks from three 3D objects datasets. Evaluations on six tasks with four categories have demonstrated the effectiveness of our proposed framework of object classification. The framework can be applied potentially in forest management.
Lingkai Li, Huan Luo 0001, Cheng Wang 0003, Wenzhong Guo, Jonathan Li 0001
IGARSS6
2022 A LiDAR-Based 3D Indoor Mapping Framework with Mismatch Detection and Optimization
abstract
In this paper, we propose a novel LiDAR-based mapping framework for geometry-featureless scenarios, combining learning-based mismatch detection and intensity-assisted registration optimization. The mismatch detection method hybridizes point cloud features and temporal features of pose to detect the mismatch position, and the optimization method uses the intensity of captured road markers by LiDAR to optimize mismatch pose. Experiments on underground garages data demonstrate that the proposed method performs better localization and mapping accuracy than the LOAM.
Weiquan Liu, Chenglu Wen, Yongfei Shi, Xiaocheng Yan, Jinbin Tan, Cheng Wang 0003, Jonathan Li 0001
IGARSS8
2022 Assessing the Impact of Covid-19 on Human Activities in the Greater Toronto Area by Nighttime Light Images and Active Covid-19 Cases
abstract
This paper explores the effect of COVID-19 outbreaks on human activity through nighttime light images of Greater Toronto Area (GTA), Canada. The methods used in this paper include image preprocessing, image classification, and spatial analysis. By using the nighttime light radiance data from VIIRS/NPP data products and COVID-19 cases and comparing this data from the pre-pandemic year, the impact of COVID-19 was analyzed. The result shows that during the pandemic year the monthly average nighttime light radiance has decreased about 4.3-5.0% compared to the pre-pandemic year. The classification results shows that the average percentage of changes in residential areas, public facilities, and commercial areas are 0.3%, −0.7%, and −1.2%, respectively of each corresponding month. Meanwhile, the spatial analysis results show population distribution patterns in GTA during the pandemic year. Overall, the nighttime lights (NTL) images can be used for a preliminary understanding of how COVID-19 affected human activities and is corroborated with other forms data collection used for the pandemic analysis.
Jianshen Wang, Sarah Narges Fatholahi, Michael A. Chapman, Yiping Chen 0002, Jonathan Li 0001
IGARSS6
2022 Foreground-Background Segmentation of Sequential Point Clouds
abstract
Point clouds are receiving increasing attention in the field of computer vision. Hitherto, segmentation tasks for point clouds have dealt with single-frame data without considering the temporal information. In this paper, we propose a point cloud segmentation network based on a neuron-like model that can exploit the time information of point cloud data to improve network performance. The network's encoder adopts the structure of the U-Net network and sparse convolution. The features extracted by the encoder and the last moment are used as input to the time integration module. The time integration module is a timing-based neuron model which can fuse features from two moments and use them to enhance the features of the current moment. Experiments on the public autopilot dataset NuScenes validate that our proposed point cloud segmentation network can achieve a 97.7 % Intersection over Union (IoU) for segmenting static scenes and dynamic objects. The experiments further validate that fusing temporal information improves the performance of point cloud foreground-background segmentation compared to considering only data from a single frame point cloud data.
Chengzhe Yang, Yiping Chen 0002, Cheng Wang 0003, Jonathan Li 0001
IGARSS4
2022 TopoSeg: Topology-aware Segmentation for Point Clouds
abstract
Point cloud segmentation plays an important role in AI applications such as autonomous driving, AR, and VR. However, previous point cloud segmentation neural networks rarely pay attention to the topological correctness of the segmentation results. In this paper, focusing on the perspective of topology awareness. First, to optimize the distribution of segmented predictions from the perspective of topology, we introduce the persistent homology theory in topology into a 3D point cloud deep learning framework. Second, we propose a topology-aware 3D point cloud segmentation module, TopoSeg. Specifically, we design a topological loss function embedded in TopoSeg module, which imposes topological constraints on the segmentation of 3D point clouds. Experiments show that our proposed TopoSeg module can be easily embedded into the point cloud segmentation network and improve the segmentation performance. In addition, based on the constructed topology loss function, we propose a topology-aware point cloud edge extraction algorithm, which is demonstrated that has strong robustness.
Weiquan Liu, Hanyun Guo, Weini Zhang, Cheng Wang 0003, Jonathan Li 0001
IJCAI6
2022 2D3D-MVPNet: Learning cross-domain feature descriptors for 2D-3D matching based on multi-view projections of point clouds
Baiqi Lai, Weiquan Liu, Cheng Wang 0003, Xiaoliang Fan, Yangbin Lin, Xuesheng Bian, Shangbin Wu, Ming Cheng 0002, Jonathan Li 0001
Appl. Intell.9
2022 Counting and locating high-density objects using convolutional neural network
Mauro dos Santos de Arruda, Lucas Prado Osco, Plabiany Rodrigo Acosta, Diogo Nunes Gonçalves, José Marcato Junior, Ana Paula Marques Ramos, Edson Takashi Matsubara, Jonathan Li 0001, Jonathan de Andrade Silva, Wesley Nunes Gonçalves
Expert Syst. Appl.9
2022 Mesh Oversegmentation with Segmentation-Aware Loss
Jibril Muhammad Adam, Muhammad Kamran Afzal, Saifullahi Aminu Bello, Cheng Wang 0003, Jonathan Li 0001
Inf. Sci.6
2022 Discriminative feature abstraction by deep L2 hypersphere embedding for 3D mesh CNNs
Muhammad Kamran Afzal, Jibril Muhammad Adam, Hafiz Muhammad Rehan Afzal, Saifullahi Aminu Bello, Cheng Wang 0003, Jonathan Li 0001
Inf. Sci.7
2022 Adaptive Pyramid Context Fusion for Point Cloud Perception
abstract
Deep learning for 3-D point cloud perception has been a very active research topic in recent years. A current trend is toward the combination of the semantically strong and the fine-grained information from different scales of intermediate representations to boost network generalization power and robustness against scale variation. One prominent challenge is how to effectively conduct the allocation of multiple scales of information. In this letter, we propose a module, named adaptive pyramid context fusion (APCF), to adaptively capture scales of contextual information from a multiscale feature pyramid for the point cloud. The APCF module reweights and aggregates the features from different levels in the feature pyramid via a softmax attention strategy. The allocation of information is adaptively conducted level by level from bottom to up first and then from top to bottom. To ensure both effectiveness and efficiency, we propose a multiscale context-aware network APCF-Net through applying our proposed APCF to the PointConv architecture. Experiments demonstrate that APCF-Net surpasses its vanilla counterpart by a large margin both in effectiveness and efficiency. Especially, APCF-Net outperforms state-of-the-art approaches on 3-D object classification and semantic segmentation task, with the overall accuracy of 93.3% on ModelNet40 and mIoU of 63.1% on ScanNet V2 online test.
Haojia Lin, Wen Li 0005, Yiping Chen 0002, Cheng Wang 0003, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.6
2022 Building Instance Extraction Method Based on Improved Hybrid Task Cascade
abstract
Automatic building extraction from remote sensing imagery is crucial to urban construction and management. To address the main challenges of diverse building scale and appearance, this letter proposes an automatic building instance extraction method based on an improved hybrid task cascade (HTC). Our method consists of three components by obtaining high-resolution representation, defining guided anchor, and forming focal loss to boost the adaptability of automatic building instance extraction. Comprehensive experimental results on WHU aerial building data set demonstrated that compared with the mainstream Mask R-CNN method, our method increased AP and AR in bounding box branch and mask branch by 9.8%–6.5% and 10.7%–8.0% respectively, especially AP$_{S}$and AP$_{L}$in the two branches by 10.1%–6.9% and 3.4%–2.4%, respectively. We evaluated the effectiveness and complexity of these components separately and discussed the universality and practicability of deep learning method in automatic building extraction.
Yiping Chen 0002, Mingqiang Wei, Cheng Wang 0003, Wesley Nunes Gonçalves, José Marcato Junior, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.7
2022 Domain Adaptation for Object Classification in Point Clouds via Asymmetrical Siamese and Conditional Adversarial Network
abstract
Nowadays, researchers have developed various deep neural networks for processing point clouds effectively. Due to the enormous parameters in deep learning-based models, a lot of manual efforts have to be invested into annotating sufficient training samples. To mitigate such manual efforts of annotating samples for a new scanning device, this letter focuses on proposing a new neural network to achieve domain adaptation in 3D object classification. Specifically, to minimize the data discrepancy of intra-class objects in different domains, an Asymmetrical Siamese module is designed to align the intra-class features. To preserve the discriminative information for distinguishing inter-class objects in different domains, a Conditional Adversarial module is leveraged to consider the classification information conveyed from the classifier. To verify the effectiveness of the proposed method on object classification in heterogeneous point clouds, evaluations are conducted on three point cloud datasets, which are collected in different scenarios by different laser scanning devices. Furthermore, the comparative experiments also demonstrate the superior performance of the proposed method on the classification accuracy.
Huan Luo 0001, Lingkai Li, Lina Fang, Hanyun Wang, Cheng Wang 0003, Wenzhong Guo, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.7
2022 STN: Saliency-Guided Transformer Network for Point-Wise Semantic Segmentation of Urban Scenes
abstract
Accurate and effective road object semantic segmentation plays a significant role in supporting extensive intelligent transportation system (ITS)-related applications. However, most existing image-based methods and point-based methods cannot deliver promising solutions with respect to segmentation accuracy and robustness, especially in complex urban road scenes. Thus, we design a saliency-guided transformer architecture (STN) in this letter for point-wise semantic segmentation from mobile laser scanning (MLS) point clouds. First, four types of feature saliency maps are constructed to obtain more compact feature spaces for enhancing the feature encoding semantics. Then, integrated with offset attention mechanisms and edge convolutions, an effective point-wise transformer network is proposed to extract high-level features for point-wise label assignment of road objects. The STN model is evaluated on the Pairs-Lille-3D dataset and achieves satisfactory experimental results with 87.2% overall accuracy and 81.7% mean IoU, respectively. Comparative studies with five deep learning-based methods also prove the superior performance of the STN model for large-scale semantic segmentation tasks.
Lingfei Ma, Jonathan Li 0001, Haiyan Guan, Yongtao Yu, Yiping Chen 0002
IEEE Geosci. Remote. Sens. Lett.2
2022 Semantic Segmentation of Coastal Zone on Airborne Lidar Bathymetry Point Clouds
abstract
Large-scale semantic segmentation point cloud is an ongoing research topic for on-land environments. However, there is a rare deep learning research study for the sub-surface environment. Although, PointNet and its successor PointNet++ have become the cornerstone of point cloud segmentation. However, these techniques handle a relatively small number of points. This poses a natural difficulty in a large spatial scene with millions of possible points. In particular, for shallow water of coastal zone, the small number of points where the seabed and water surface meet, close points may belong to different classes. In our work, we present the semantic segmentation on a large-scale airborne Lidar bathymetry (ALB) point cloud containing millions of sample points into two classes of water surface and seabed with the voxel sampling pre-processing (VSP) approach. The proposed approach will allow us to capture the complicated outdoor natural scene components of water surface and seabed more accurately and more realistic through nonuniform voxelization in the mixture of dense and sparse points of the ALB point cloud. The performance of validation results show improvement in a per-point accuracy of 72.45% compared with other state-of-the-art deep learning-based methods.
Sajjad Roshandel, Weiquan Liu, Cheng Wang 0003, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.4
2022 A Supervoxel Approach to Road Boundary Enhancement From 3-D LiDAR Point Clouds
abstract
Rapid and accurate enhancement of road boundaries from terrestrial laser scanning (TLS) 3-D point clouds has been a challenging task in road infrastructure inventory. To address the challenge with a lack of ability to enhance object boundaries when the supervoxel number is less, this letter proposes a novel supervoxel segmentation algorithm framework for enhancing road boundaries from 3-D point clouds. First, we utilize radius$k$nearest-neighbor search method to obtain the neighborhood information after partitioning points on octrees with seed points. Second, the iterative weighted least square algorithm and spatial structure judgment are used to segment point clouds based on seed points. Finally, an update method to adjust the supervoxel centroids is applied with surrounding information in the first part. To verify the excellent performance, we tested the proposed method on two publicly large-scale point clouds benchmarks—IQmulus and TerraMobilita (IQTM) and Semantic 3-D. The experimental results demonstrate that our approach achieved approximately 48.98% and 68.41% boundary recall higher than two existing classical methods in the street scene, and our running time is feasible and effective.
Zhengchuan Sha, Yiping Chen 0002, Yangbin Lin, Cheng Wang 0003, José Marcato Junior, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.6
2022 Land Cover Classification of Multispectral LiDAR Data With an Efficient Self-Attention Capsule Network
abstract
Periodically conducting land cover mapping plays a vital role in monitoring the status and changes of the land use. The up-to-date and accurate land use database serves importantly for a wide range of applications. This letter constructs an efficient self-attention capsule network (ESA-CapsNet) for land cover classification of multispectral light detection and ranging (LiDAR) data. First, formulated with a novel capsule encoder–decoder architecture, the ESA-CapsNet performs promisingly in extracting high-level, informative, and strong feature semantics for pixel-wise land cover classification by using the five types of rasterized feature images. Furthermore, designed with a novel capsule-based attention module, the channel and spatial feature encodings are comprehensively exploited to boost the feature saliency and robustness. The ESA-CapsNet is evaluated on two multispectral LiDAR data sets and achieves an advantageous performance with the overall accuracy, average accuracy, and kappa coefficient of over 98.42%, 95.15%, and 0.9776, respectively. Comparative experiments with the existing methods also demonstrate the effectiveness and applicability of the ESA-CapsNet in land cover classification tasks.
Yongtao Yu, Chao Liu 0040, Haiyan Guan, Lanfang Wang, Shangbing Gao, Yahong Zhang, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.8
2022 Dense Point Cloud Completion Based on Generative Adversarial Network
abstract
Point cloud completion aims to reconstruct complete point clouds from partial point clouds, which is widely used in various fields such as autonomous driving and robotics. Most existing methods are sparse point cloud completion, where the number of point clouds after completion is relatively small and the details are insufficient. This article proposes a novel end-to-end generative adversarial network-based dense point cloud completion architecture (DPCG-Net). We design two generative adversarial network (GAN)-based modules that translate point cloud completion into mapping between global feature distributions obtained by encoding partial point clouds and ground truth, respectively. The first designed generator module proposes skip connections to fully connected layer-based network for regenerating global feature and changing the global feature distribution derived from the encoder module to approximate the ground truth global feature distribution. The second proposed discriminator module divides high-dimensional global feature vectors into several smaller batches for judgment to guarantee the similarity between the regenerated global feature and the ground truth. We perform quantitative and qualitative experiments on the ShapeNet and KITTI datasets. Experiments on ShapeNet demonstrate that our model outperforms other models in cases where the lack of a large proportion of point clouds results in a large loss of spatial structure, especially when 80% of point clouds are missing. Moreover, KITTI experiments reveal that it is also valid for realistic situations. In addition, application in classification shows that the classification accuracy of point clouds completed with DPCG-Net is as high as 86.5% under the condition of 80% missing point clouds.
Ming Cheng 0002, Guoyan Li, Yiping Chen 0002, Jun Chen 0005, Cheng Wang 0003, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.6
2022 A GCN-Based Method for Extracting Power Lines and Pylons From Airborne LiDAR Data
abstract
Extracting the power lines and pylons automatically and accurately from airborne LiDAR data is a critical step in inspecting the routine power line, especially in the remote mountainous areas. However, challenges arise in using existing methods to extract the targets from large scenarios of remote mountainous areas since the terrain is undulating, and the features are difficult to distinguish. In this article, to overcome these challenges, we propose a graph convolutional network (GCN)-based method to extract power lines and pylons from Airborne LiDAR point clouds. First, data augmentation and near-ground filtering methods are developed to overcome the problems of insufficient and imbalanced samples in the LiDAR data. Then, a GCN-based framework is proposed to extract the power lines and pylons, which consist of two main modules, i.e., the neighborhood dimension information (NDI) module and the neighborhood geometry information aggregation (NGIA) module. These two modules are designed to strengthen the model’s ability to portray local geometric details. Besides, an attention fusion module is investigated to further improve the NDI and NGIA features. Finally, a line structure constraint algorithm is proposed to identify individual power lines, where the power corridor is reconstructed using a polynomial-based algorithm. Numerical experiments are conducted based on two different power line scenarios acquired in mountainous areas. The results demonstrate the superior performances of the proposed method over several existing algorithms, where the$F_{1}$score and quality of the power line are 99.3% and 98.6%, and the results of the pylon are 96% and 92.4%, respectively. The identification rate of power line identification is above 98%.
Wen Li 0005, Zhenlong Xiao, Yiping Chen 0002, Cheng Wang 0003, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.6
2022 Detection of Individual Trees in UAV LiDAR Point Clouds Using a Deep Learning Framework Based on Multichannel Representation
abstract
Individual tree detection is critical for forest investigation and monitoring. Several existing methods have difficulties to detect trees in complex forest environments due to insufficiently mining descriptive features. This study proposes a deep learning (DL) framework based on a designed multichannel information complementarity representation for detecting trees in complex forest using UAV laser scanning point clouds. The proposed method consists of two main stages: ground filtering and tree detection. In the first stage, a modified graph convolution network with a local topological information layer is designed to separate the ground points. Unlike most existing parametric methods, our ground filtering method avoids the optimal parameters selection to adapt to different kinds of environments. For tree detection, a top-down slice (TDS) module is first designed to mine the vertical structure information in a top-down way. Then, a special multichannel representation (MCR) is developed to preserve different distribution patterns of points from complementary perspectives. Finally, a multibranch network (MBNet) is proposed for individual tree detection by fusing multichannel features, which can provide discriminative information for MBNet to detect trees more accurately. MBNet was evaluated on seven forest areas [UAV light detection and ranging (LiDAR) data with the mean size of 14$000~\text {m}^{2}$and point density of 250 points/$\text {m}^{2}$]. Experimental results showed that the proposed framework achieves excellent performance. Our method obtains promising performance with a mean recall of 89.23% and a mean F1-score of 87.04%.
Wen Li 0005, Yiping Chen 0002, Cheng Wang 0003, Abdul Nurunnabi, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.7
2022 A Two-Step Descriptor-Based Keypoint Filtering Algorithm for Robust Image Matching
abstract
Finding robust and correct keypoints in images remains a challenge, especially when repetitive patterns are present. In this article, we propose a universal two-step filtering method to solve the mismatch problem in repetitive patterns. Having applied a mean-shift clustering algorithm to remove obvious mismatches, the proposed confusion reduction (CR) method uses a novel confusion index (CI) in a gridding schema to identify and filter out the remaining confusing keypoints. In both steps, the descriptors’ statistical properties are evaluated using kernel density estimation. Various synthetic and real stereo pairs, along with multiview image blocks, were used to assess the performance of the presented algorithm. The results were also compared with those obtained by several state-of-the-art mismatch removal methods. The experiments showed that, on average, the proposed strategy improves the accuracy of matching by 10% and the accuracy of photogrammetric blocks by 20%–30%.
Vahid Mousavi, Masood Varshosaz, Fabio Remondino, Saied Pirasteh, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 Detecting Occluded and Dense Trees in Urban Terrestrial Views With a High-Quality Tree Detection Dataset
abstract
Urban trees are often densely planted along the two sides of a street. When observing these trees from a fixed view, they are inevitably occluded with each other and the passing vehicles. The high density and occlusion of urban tree scenes significantly degrade the performance of object detectors. This paper raises an intriguing learning-related question – if a module is developed to enable the network to adaptively cope with occluded and un-occluded regions while enhancing its feature extraction capabilities, can the performance of a cutting-edge detection model be improved? To answer it, a lightweight yet effective object detection network is proposed for discerning occluded and dense urban trees, called OD-UTDNet. The main contribution is a newly-designed Dilated Attention Cross Stage Partial (DACSP) module. DACSP can expand the fields-of-view of OD-UTDNet for paying more attention to the un-occluded region, while enhancing the network’s feature extraction ability in the occluded region. This work further explores both the self-calibrated convolution module and GFocal loss, which enhance the OD-UTDNet’s ability to resolve the challenging problem of high densities and occlusions. Finally, to facilitate the detection task of urban trees, a high-quality urban tree detection dataset is established, named UTD; to our knowledge, this is the first time. Extensive experiments show clear improvements of the proposed OD-UTDNet over twelve representative object detectors on UTD. The code and dataset are available at https://github.com/yzwang/OD-UTDNet.
Yongzhen Wang 0001, Xuefeng Yan 0001, Hexiang Bao, Yiping Chen 0002, Lina Gong, Mingqiang Wei, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.7
2022 CasA: A Cascade Attention Network for 3-D Object Detection From LiDAR Point Clouds
abstract
3D object detection from LiDAR point clouds has gained great attention in recent years due to its wide applications in smart cities and autonomous driving. Cascade framework shows its advancement in 2D object detection but is less investigated in 3D space. Conventional cascade structures use multipleseparatesub-networks to sequentially refine region proposals. Such methods, however, have limited ability to measure proposal quality in all stages, and hard to achieve a desirable performance improvement in 3D space. This paper proposes a new cascade framework, termed CasA, for 3D object detection from LiDAR point clouds. CasA consists of a Region Proposal Network (RPN) and a Cascade Refinement Network (CRN). In CRN, we designed a new Cascade Attention Module that uses multiple sub-networks and attention modules to aggregate the object features from different stages and progressively refine region proposals. CasA can be integrated into various two-stage 3D detectors and improve their performance. Extensive experiments on KITTI and Waymo datasets with various baseline detectors demonstrate the universality and superiority of our CasA. In particular, based on one variant of Voxel-RCNN, we achieve state-of-the-art results on the KITTI dataset. On the KITTI online 3D object detection leaderboard, we achieve a high detection performance of 83.06%, 47.09%, and 73.47% Average Precision (AP) in the moderate Car, Pedestrian, and Cyclist classes, respectively. Code is available at https://github.com/hailanyi/CasA.
Jinhao Deng, Chenglu Wen, Xin Li 0003, Cheng Wang 0003, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.6
2022 Intracity Temperature Estimation by Physics Informed Neural Network Using Modeled Forcing Meteorology and Multispectral Satellite Imagery
abstract
Estimating urban surface temperature at high resolution is crucial for effective urban planning for climate-driven risks. This high-resolution surface temperature over broader scales can usually be obtained via satellite remote sensing for historical period. However, it can be hard for future predictions. This paper presents a Physics Informed Hierarchical Perception (PIHP) network, a novel approach for accurate, high-resolution and generalizable urban surface temperature estimation. The key to our approach is leveraging the implied temperature-related physics information of the land surface structure from high-resolution multi-spectral satellite images, thus achieving precise estimation or prediction for high spatial resolution urban surface temperature. Specifically, a semantic category histogram is first designed to describe the land surface structures. Based on this, a hierarchical urban surface perception network is proposed to capture the complex relationship between the underlying land surface features, upper atmosphere conditions and the intracity temperature. The proposed PIHP-Net makes it possible to generate models that can generalize across different cities, thus to estimating or predicting high-resolution urban surface temperature when the satellite land surface temperature (LST) observation is not available. Experiments over various cities in different climate regions in China show, for the first time, errors less than 2 Kelvin (for most of the cases) at the high resolution (60-by-60 meters grids), thus making it possible to predict futureintracity temperaturefrom forcing meteorology and multi-spectral satellite imagery.
Donghang Wu, Weiquan Liu, Lei Zhao 0023, Shenlong Wang, Cheng Wang 0003, José Marcato Junior, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.10
2022 Building Instance Mapping From ALS Point Clouds Aided by Polygonal Maps
abstract
Building region extraction from ALS point clouds has been widely studied, whereas instance-level building mapping has been overlooked and remains unsolved. In this study, we present a method to extract individual buildings from ALS point clouds with the help of widely accessible polygonal footprints. The key idea is to merge roof segments to a set of building candidates, from which correct instances are selected by finding optimal matches between polygonal footprints and building candidates. The method has three steps: roof segmentation, building candidate generation, and instance-polygon matching. The method is tested on two large-scale scenes of different building types and can generally achieve high instance-level building mapping accuracy (around 90%) when there are large positioning errors (6.0 m) among polygons. Future work will focus on classification errors in preprocessing, shape inconsistency between point clouds and polygons, and building footprint delineation and updating in postprocessing.
Shaobo Xia, Sheng Xu 0003, Ruisheng Wang 0001, Jonathan Li 0001, Guanghui Wang 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 Spectral-Spatial Transformer Network for Hyperspectral Image Classification: A Factorized Architecture Search Framework
abstract
Neural networks have dominated the research of hyperspectral image classification, attributing to the feature learning capacity of convolution operations. However, the fixed geometric structure of convolution kernels hinders long-range interaction between features from distant locations. In this article, we propose a novel spectral–spatial transformer network (SSTN), which consists of spatial attention and spectral association modules, to overcome the constraints of convolution kernels. Also, we design a factorized architecture search (FAS) framework that involves two independent subprocedures to determine the layer-level operation choices and block-level orders of SSTN. Unlike conventional neural architecture search (NAS) that requires a bilevel optimization of both network parameters and architecture settings, the FAS focuses only on finding out optimal architecture settings to enable a stable and fast architecture search. Extensive experiments conducted on five popular HSI benchmarks demonstrate the versatility of SSTNs over other state-of-the-art (SOTA) methods and justify the FAS strategy. On the University of Houston dataset, SSTN obtains comparable overall accuracy to SOTA methods with a small fraction (1.2%) of multiply-and-accumulate operations compared to a strong baseline spectral–spatial residual network (SSRN). Most importantly, SSTNs outperform other SOTA networks using only 1.2% or fewer MACs of SSRNs on the Indian Pines, the Kennedy Space Center, the University of Pavia, and the Pavia Center datasets.
Zilong Zhong, Ying Li 0036, Lingfei Ma, Jonathan Li 0001, Wei-Shi Zheng 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 A New Method for Automated Monitoring of Road Pavement Aging Conditions Based on Recurrent Neural Network
abstract
The automated monitoring of road pavement conditions is a challenging subject in intelligent transportation. However, the existing studies mostly focus on extracting pavement damages such as cracks, while the pavement aging conditions are still less investigated. In this paper, a novel method based on a modified recurrent neural network is designed for automated monitoring of asphalt pavement aging phenomena from fine-resolution satellite imagery. A spectral augmentation method is proposed to enhance the spectral details of the road pavements. A novel loss function is also proposed to improve the bi-directional gated recurrent unit (Bi-GRU) network in order to better classify different degrees of road pavement aging and non-pavement objects. In order to demonstrate the outperformance of the modified network Bi-GRU+, the Worldview-2 satellite image (16360*7728) covering 16 asphalt roads in the southwestern suburb of Beijing City is used. The results show that the proposed approach has better performance than existing machine learning methods, with an overall accuracy of 98.16% and a Kappa coefficient of 0.97. The overall processing time of the proposed method is 7836 seconds in our case study. The proposed method is efficient for large-scale monitoring of road health conditions from fine-resolution satellite imagery. It can become a part of intelligent transportation and provide a new foundation for large-range automated monitoring of road pavement aging conditions.
Xiao Chen 0020, Xianfeng Zhang, Jonathan Li 0001, Miao Ren, Bo Zhou 0023
IEEE Trans. Intell. Transp. Syst.3
2022 GCN-Based Pavement Crack Detection Using Mobile LiDAR Point Clouds
abstract
Mobile Laser Scanning (MLS) system can provide high-density and accurate 3D point clouds that enable rapid pavement crack detection for road maintenance tasks. Supervised learning-based algorithms have been proved pretty effective for handling such a large amount of inhomogeneous and unstructured point clouds. However, these algorithms often rely on a lot of annotated data, which is labor-intensive and time-consuming. This paper presents a semi-supervised point-level approach to overcome this challenge. We propose a graph-widen module to construct a reasonable graph structure for point clouds, increasing the detection performance of graph convolutional networks (GCN). The constructed graph characterizes the local features from a small amount of annotated data, avoiding information loss and dramatically reduces the dependence on annotated data. The MLS point clouds acquired by a commercial RIEGL VMX-450 system are used in this study. The experimental results demonstrate that our method outperforms the state-of-the-art point-level methods in terms of recall, F1 score, and efficiency while achieving comparable accuracy.
Huifang Feng 0002, Wen Li 0005, Yiping Chen 0002, Sarah Narges Fatholahi, Ming Cheng 0002, Cheng Wang 0003, José Marcato Junior, Jonathan Li 0001
IEEE Trans. Intell. Transp. Syst.9
2022 Rapid Extraction of Urban Road Guardrails From Mobile LiDAR Point Clouds
abstract
Mobile Laser Scanning (MLS) systems provide highly dense 3D point clouds that enable the acquisition of accurate traffic facilities information for intelligent transportation system. Road guardrails with safety features that can separate traffic and define moving spaces for pedestrians and vehicles face challenges such as diverse guardrail types and continuous slopes in point clouds data. This paper proposes a novel approach for rapidly extracting urban road guardrails from MLS point clouds, combining a proposed multi-level filtering with a modified Density-Based Spatial Clustering of Applications with Noise (DBSCAN) clustering, and adapting for most types of guardrails and rough slope roads. We develop a multi-level filter to detect the road surface and remove the undesirable points. Through a proposed modified DBSCAN clustering, the guardrails are extracted after a four-step screening, which includes the limits based on the number of points, the fitting error, the bounding box size and the average reflection intensity for each cluster. The proposed method achieves high precisions of 97.2% and 96.4% respectively for the lane-separating guardrails and the anti-fall guardrails on the dataset. Extensive experiments with test dataset captured by a RIEGL VMX-450 MLS, show that our method outperforms the state-of-the-art method to extract 3D guardrails from point clouds.
Jianlan Gao, Yiping Chen 0002, José Marcato Junior, Cheng Wang 0003, Jonathan Li 0001
IEEE Trans. Intell. Transp. Syst.5
2022 3DCTN: 3D Convolution-Transformer Network for Point Cloud Classification
abstract
Point cloud classification is a fundamental task in 3D applications. However, it is challenging to achieve effective feature learning due to the irregularity and unordered nature of point clouds. Lately, 3D Transformers have been adopted to improve point cloud processing. Nevertheless, massive Transformer layers tend to incur huge computational and memory costs. This paper presented a novel hierarchical framework that incorporated convolutions with Transformers for point cloud classification, named 3D Convolution-Transformer Network (3DCTN). It combined the strong local feature learning ability of convolutions with the remarkable global context modeling capability of Transformers. Our method had two main modules operating on the downsampling point sets. Each module consisted of a multi-scale local feature aggregating (LFA) block and a global feature learning (GFL) block, which were implemented by using the Graph Convolution and Transformer respectively. We also conducted a detailed investigation on a series of self-attention variants to explore better performance for our network. Various experiments on ModelNet40 and ScanObjectNN datasets demonstrated that our method achieves state-of-the-art classification performance with a lightweight design. The code is publicly available athttps://github.com/d62lu/3DCTN.
Dening Lu, Qian Xie 0001, Kyle Gao, Linlin Xu, Jonathan Li 0001
IEEE Trans. Intell. Transp. Syst.5
2022 BoundaryNet: Extraction and Completion of Road Boundaries With Deep Learning Using Mobile Laser Scanning Point Clouds and Satellite Imagery
abstract
Robust road boundary extraction and completion play an important role in providing guidance to all road users and supporting high-definition (HD) maps. The significant challenges remain in remarkable and accurate road boundary recovery from poor road boundary conditions. This paper presents a novel deep learning framework, named BoundaryNet, to extract and complete road boundaries by using both mobile laser scanning (MLS) point clouds and high-resolution satellite imagery. First, road boundaries are extracted by conducting a curb-based extraction method. Such extracted 3D road boundary lines are used as inputs to feed into a U-shaped network for erroneous boundary denoising. Then, a convolutional neural network (CNN) model is proposed to complete the road boundaries. Next, to achieve more complete and accurate road boundaries, a conditional deep convolutional generative adversarial network (c-DCGAN) with the assistance of road centerlines extracted from satellite images is developed. Finally, according to the completed road boundaries, the inherent road geometries are calculated. The proposed methods were evaluated using satellite imagery and four MLS point cloud datasets with varying densities and road conditions in urban environments. The quality evaluation metrics of 82.88%, 82.43%, 88.86%, and 84.89% were achieved for four data sets. The experimental results indicate that the BoundaryNet model can provide a promising solution for road boundary completion and road geometry estimation.
Lingfei Ma, Ying Li 0036, Jonathan Li 0001, José Marcato Junior, Wesley Nunes Gonçalves, Michael A. Chapman
IEEE Trans. Intell. Transp. Syst.3
2022 Cycle-SNSPGAN: Towards Real-World Image Dehazing via Cycle Spectral Normalized Soft Likelihood Estimation Patch GAN
abstract
Image dehazing is a common operation in autonomous driving, traffic monitoring and surveillance. Learning-based image dehazing has achieved excellent performance recently. However, it is nearly impossible to capture pairs of hazy/clean images from the real world to train an image dehazing network. Most of existing dehazing models that are learnt from synthetically generated hazy images generalize poorly on real-world hazy scenarios due to the obvious domain shift. To deal with this unpaired problem arisen by real-world hazy images, we present Cycle Spectral Normalized Soft likelihood estimation Patch Generative Adversarial Network (Cycle-SNSPGAN) for image dehazing. Cycle-SNSPGAN is an unsupervised dehazing framework to boost the generalization ability on real-world hazy images. To leverage unpaired samples of real-world hazy images without relying on their clean counterparts, we design an SN-Soft-Patch GAN and exploit a new cyclic self-perceptual loss which avoids using the ground-truth image to compute the perceptual similarity. Moreover, a significant color loss is adopted to brighten the dehazed images as human expects. Both visual and numerical results show clear improvements of the proposed Cycle-SNSPGAN over state-of-the-arts in terms of hazy-robustness and image detail recovery, with even only a small dataset training our Cycle-SNSPGAN. Code has been available athttps://github.com/yz-wang/Cycle-SNSPGAN.
Yongzhen Wang 0001, Xuefeng Yan 0001, Donghai Guan, Mingqiang Wei, Yiping Chen 0002, Xiao-Ping Zhang 0002, Jonathan Li 0001
IEEE Trans. Intell. Transp. Syst.7
2022 Robust Lane Extraction From MLS Point Clouds Towards HD Maps Especially in Curve Road
abstract
This article presents a semi-automated method to extract the lane features along the curved roads from mobile laser scanning (MLS) point clouds. The proposed method consists of four steps. After data pre-processing, a road edge detection algorithm is performed to distinguish road curbs and extract road surfaces. Then, textual and directional road markings such as arrows, symbols, and words, to inform drivers in necessary cases, are detected by intensity thresholding and conditional Euclidean clustering algorithms. Furthermore, lane markings are extracted by local intensity analysis and distance thresholding methods according to road design standards, because they are more regular along the road. Finally, centerline points on lanes are estimated based on the coordinates of extracted lane markings. Our method shows strong feasibility and robustness when creating high-definition (HD) maps from MLS data, by increasing the number of blocks in the curve and the distance threshold control in curved lane centerline extraction. Quantitative evaluations show that the average recall, precision, and F1-score obtained from four datasets for road marking extraction are 93.87%, 93.76%, and 93.73%, respectively. The generated lane centerlines are evaluated by overlaying them on manually labeled reference buffers from 4 cm resolution orthoimagery. The comparative study indicates that the proposed methods can achieve higher accuracy and robustness than most state-of-the-art methods.
Chengming Ye, He Zhao 0007, Lingfei Ma, Han Jiang 0005, Hongfu Li, Ruisheng Wang 0001, Michael A. Chapman, José Marcato Junior, Jonathan Li 0001
IEEE Trans. Intell. Transp. Syst.9
2022 3D Vehicle Detection Using Multi-Level Fusion From Point Clouds and Images
abstract
3D vehicle detectors based on point clouds generally have higher detection performance than detectors based on multi-sensors. However, with the lack of texture information, point-based methods get many missing detection of occluded and distant vehicles, and false detection with high-confidence of similarly shaped objects, which is a potential threat to traffic safety. Therefore, in the long run, fusion-based methods have more potential. This paper presents a multi-level fusion network for 3D vehicle detection from point clouds and images. The fusion network includes three stages: data-level fusion of point clouds and images, feature-level fusion of voxel and Bird’s Eye View (BEV) in the point cloud branch, and feature-level fusion of point clouds and images. Besides, a novel coarse-fine detection header is proposed, which simulates the two-stage detectors, generating coarse proposals on the encoder, and refining them on the decoder. Extensive experiments show that the proposed network has better detection performance on occluded and distant vehicles, and reduces the false detection of similarly shaped objects, proving its superiority over some state-of-the-art detectors on the challenging KITTI benchmark. Ablation studies have also demonstrated the effectiveness of each designed module.
Lingfei Ma, José Marcato Junior, Wesley Nunes Gonçalves, Jonathan Li 0001
IEEE Trans. Intell. Transp. Syst.8
2021 Monitoring Surface Deformation Over Oilfield Using MT-Insar and Production Well Data
abstract
Surface displacements associated with the average subsidence due to hydrocarbon exploitation in southwest of Iran which has a long history in oil production, can lead to significant damages to surface and subsurface structures, and requires serious consideration. In this study, the Small BAseline Subset (SBAS) approach, which is a multitemporal Interferometric Synthetic Aperture Radar (InSAR) algorithm was employed to resolve ground deformation in the Marun region, Iran. A total of 22 interferograms were generated using 10 Envisat ASAR images. The mean velocity map obtained in the Line-Of-Sight (LOS) direction of satellite to the ground reveals the maximum subsidence on order of 13.5 mm per year over the field due to both tectonic and non-tectonic features. In order to assess the effect of non-tectonic features such as petroleum extraction on ground surface displacement, the results of InSAR have been compared with the oil production rate, which have shown a good agreement.
Sarah Narges Fatholahi, Hongjie He 0003, Awase Syed, Jonathan Li 0001
IGARSS5
2021 The Impact of Data Volume on Performance of Depp Learning Based Building Rooftop Extraction Using Very High Spatial Resolution Aerial Images
abstract
Building rooftop data are of importance in several urban applications and in natural disaster management. In contrast to traditional surveying and mapping, by using high spatial resolution aerial images, deep learning-based building rooftops extraction methods are efficient and accurate. Although more training data is preferred in deep learning-based tasks, the effect of data volume on building extraction models is underexplored. Therefore, the paper explores the impact of data volume on the performance of building rooftop extraction from very-high-spatial-resolution (VHSR) images using deep learning-based methods. To do so, we manually labelled 0.12m spatial resolution aerial images and perform a comparative analysis of models trained on datasets of different sizes using popular deep learning architectures for segmentation tasks, including Fully Convolutional Networks (FCN)-8s, U-Net and DeepLabv3+. The experiments showed that with more training data, algorithms converged faster and achieved higher accuracy, while better algorithms were able to better mitigate the lack of training data.
Hongjie He 0003, Yuwei Cai, Zijian Jiang, Qiutong Yu, Sarah Narges Fatholahi, Yan Liu 0043, Hasti Andon Petrosians, Bingxu Hu, Liyuan Qing, Zhehan Zhang, Hongzhang Xu, Kyle Gao, Linlin Xu, Jonathan Li 0001
IGARSS18
2021 Integration of Photogrammetry and Deep Learning in Earth Observation Applications
abstract
The integration of photogrammetry and deep learning methods can be powerful for Earth observation applications. Photogrammetry techniques allow the achievement of detailed geospatial products with em-level positional accuracy. Deep learning enables automatic image classification, segmentation, and object detection. For instance, when dealing with a large data set, photogrammetric processing steps, such as image orientation and dense point cloud generation, results in high computational costs. In contrast, deep learning methods are fast in the inference step. Here, we explore the complementarity of deep learning and photogrammetry, aiming to generate accurate and fast geospatial information. The main aim is to discuss the possibilities of using deep learning in the photogrammetric process. We conduct experiments to present the potential of the Mask R-CNN method trained on the COCO dataset to generate masks, essential to remove image observations from moving objects during the orientation (alignment) step.
José Marcato Junior, Pedro Zamboni, Mariana Batista Campos, Ana Paula Marques Ramos, Lucas Prado Osco, Jonathan de Andrade Silva, Wesley Nunes Gonçalves, Jonathan Li 0001
IGARSS8
2021 Metric Learning for 2D Image Patch and 3D Point Cloud Volume Matching
abstract
Similarity measure of cross-domain descriptors (2D descriptors and 3D descriptors) between 2D image patches and 3D point cloud volumes provides stable retrieval performance and establishes the spatial relationship between 2D and 3D space, which plays the potential applications in geospatial space, such as 2D and 3D interaction of remote sensing, Augmented Reality (AR) and robot navigation. However, the mature handcrafted descriptors of 2D image patches and 3D point cloud volumes are extremely different, resulting in the huge challenge for 2D image patch and 3D point cloud volume matching. In this paper, we propose a novel network which combines both unified descriptor training and descriptor comparison function training for 2D image patch and 3D point cloud volume matching. First, two feature extraction networks are applied for jointly learning the local descriptors for 2D image patches and 3D point cloud volumes, respectively. Second, a fully connected network is introduced to compute the similarity between 2D descriptors and 3D descriptors. Motivated by the successful indicator system on evaluating 2D patch feature representation, we use the false positive rate at 95% recall (FPR95) and precision based on cross-domain descriptors as the measured metric. The experimental results show that our proposed network achieve state-of-the-art performance in the matching of 2D image patches and 3D point cloud volumes.
Baiqi Lai, Weiquan Liu, Cheng Wang 0003, Xuesheng Bian, Xiuhong Lin, Chenglu Wen, Jonathan Li 0001
IGARSS8
2021 A Local Topological Information Aware Based Deep Learning Method for Ground Filtering from Airborne Lidar Data
abstract
As a foundational preprocessing step for a lot of downstream tasks, ground filtering from airborne LiDAR data is designed to separate the ground points and preserve the off-ground points with complete shape information. However, because of the undulating terrain, it is still a challenge work to filter the ground under complex mountain regions. In this paper, we provide a deep learning based model to improve the ground filtering performance in abrupt slope using airborne LiDAR point clouds. Specifically, we first design a local topological information mining module to extract the local features. Then a modified graph convolutional networks (GCNs) is developed to fusion the local features and global features. Compared with most existing methods, our model not only enjoys the parameter-free advantage, which means it can be applied easily in various areas, but also obtains better ground filtering performance and can preserve more complete information contained in off-ground points. Experiments was implemented on seven forest areas. The proposed method obtains promising ground filtering results with mean total error of 6.46% and the mean kappa coefficient of 86.01%.
Wen Li 0005, Haojia Lin, Yiping Chen 0002, Cheng Wang 0003, Jonathan Li 0001
IGARSS7
2021 Semantic Segmentation of UAV Lidar Point Clouds of a Stack Interchange with Deep Neural Networks
abstract
Stack interchanges are essential components of transportation systems. Mobile laser scanning (MLS) systems have been widely used in road infrastructure mapping, but accurate mapping of complicated multi-layer stack interchanges are still challenging. This study examined the point clouds collected by a new Unmanned Aerial Vehicle (UAV) Light Detection and Ranging (LiDAR) system to perform the semantic segmentation task of a stack interchange. An end-to-end supervised 3D deep learning framework was proposed to classify the point clouds. The proposed method has proven to capture 3D features in complicated interchange scenarios with stacked convolution and the result achieved over 93% classification accuracy. In addition, the new low-cost semi-solid-state LiDAR sensor Livox Mid-40 featuring a incommensurable rosette scanning pattern has demonstrated its potential in high-definition urban mapping.
Weikai Tan, Dedong Zhang, Lingfei Ma, Nannan Qin, Yiping Chen 0002, Jonathan Li 0001
IGARSS7
2021 Direction-aware Feature-level Frequency Decomposition for Single Image Deraining
abstract
We present a novel direction-aware feature-level frequency decomposition network for single image deraining. Compared with existing solutions, the proposed network has three compelling characteristics. First, unlike previous algorithms, we propose to perform frequency decomposition at feature-level instead of image-level, allowing both low-frequency maps containing structures and high-frequency maps containing details to be continuously refined during the training procedure. Second, we further establish communication channels between low-frequency maps and high-frequency maps to interactively capture structures from high-frequency maps and add them back to low-frequency maps and, simultaneously, extract details from low-frequency maps and send them back to high-frequency maps, thereby removing rain streaks while preserving more delicate features in the input image. Third, different from existing algorithms using convolutional filters consistent in all directions, we propose a direction-aware filter to capture the direction of rain streaks in order to more effectively and thoroughly purge the input images of rain streaks. We extensively evaluate the proposed approach in three representative datasets and experimental results corroborate our approach consistently outperforms state-of-the-art deraining algorithms.
Yidan Feng, Mingqiang Wei, Haoran Xie 0001, Yiping Chen 0002, Jonathan Li 0001, Xiao-Ping Zhang 0002, Harry Qin
IJCAI6
2021 Y-Net: Learning Domain Robust Feature Representation for ground camera image and large-scale image-based point cloud registration
Weiquan Liu, Cheng Wang 0003, Xuesheng Bian, Baiqi Lai, Xuelun Shen, Ming Cheng 0002, Shang-Hong Lai, Dongdong Weng, Jonathan Li 0001
Inf. Sci.10
2021 Estimating Carbon Sequestration Potential in Vegetation by Distance-Constrained Zonal Analysis
abstract
This letter proposes A DISTANCE-CONSTRAIned (DC) zonal analysis approach to quantify how much more carbon could be further sequestrated by vegetation in mainland China based on multiple data sources. Our approach first segments the area into homogeneous landform-vegetation-soil (LVS) zones. Good land management practice (GLMP) corresponding to high sequestrated carbon (target carbon level) is identified at the locations in the same LVS zone. The target carbon level is set as the 90th percentile of the historically sequestrated carbon using the proxy of net primary productivity (NPP) at the locations within the LVS zone. When GLMP is realized over the entire LVS zone, more carbon could be sequestrated. Our results show that on average about 1/4 of more carbon could be added to the existing amount given the selected “good” land management practices are adopted by neighboring locations where lower carbon sequestration levels exist. The carbon sequestration potential for different land cover types differs significantly.
Zongyao Sha, Ruren Li, Jonathan Li 0001, Yichun Xie
IEEE Geosci. Remote. Sens. Lett.3
2021 Mapping and Semantic Modeling of Underground Parking Lots Using a Backpack LiDAR System
abstract
Presented in this paper is a novel method for the mapping and semantic modeling of an underground parking lot using 3D point clouds collected by a low-cost Backpack Laser Scanning (BLS) or LiDAR system. Our method consists of two parts: a Simultaneous Localization and Mapping (SLAM) algorithm based on Sparse Point Clouds (SPC) and a semantic modeling algorithm based on a modified PointNet model. The main contributions of this paper are as follows: (1) a probability frontend framework for the alignment of point clouds using the local point cloud surface variance as the weight of registration, which modifies registration failure caused by the lack of features in sparse point clouds, (2) a robust submap-based strategy for loop closure detection and back-end optimization under sparse point clouds, and (3) a modified PointNet model for classifying the point clouds of underground parking lots into four categories: ceiling, floor, wall, others. Experimental results show that our SPC-SLAM algorithm achieves centimeter-level accuracy (0.09% trajectory error rate) after closed loop processing in a Global Navigation Satellite System (GNSS)-denied underground parking lot, and precision of 84.8% in semantic segmentation.
Jonathan Li 0001, Chenglu Wen, Cheng Wang 0003, John S. Zelek
IEEE Trans. Intell. Transp. Syst.2
2021 Capsule-Based Networks for Road Marking Extraction and Classification From Mobile LiDAR Point Clouds
abstract
Accurate road marking extraction and classification play a significant role in the development of autonomous vehicles (AVs) and high-definition (HD) maps. Due to point density and intensity variations from mobile laser scanning (MLS) systems, most of the existing thresholding-based extraction methods and rule-based classification methods cannot deliver high efficiency and remarkable robustness. To address this, we propose a capsule-based deep learning framework for road marking extraction and classification from massive and unordered MLS point clouds. This framework mainly contains three modules. Module I is first implemented to segment road surfaces from 3D MLS point clouds, followed by an inverse distance weighting (IDW) interpolation method for 2D georeferenced image generation. Then, in Module II, a U-shaped capsule-based network is constructed to extract road markings based on the convolutional and deconvolutional capsule operations. Finally, a hybrid capsule-based network is developed to classify different types of road markings by using a revised dynamic routing algorithm and large-margin Softmax loss function. A road marking dataset containing both 3D point clouds and manually labeled reference data is built from three types of road scenes, including urban roads, highways, and underground garages. The proposed networks were accordingly evaluated by estimating robustness and efficiency using this dataset. Quantitative evaluations indicate the proposed extraction method can deliver 94.11% in precision, 90.52% in recall, and 92.43% in F1-score, respectively, while the classification network achieves an average of 3.42% misclassification rate in different road scenes.
Lingfei Ma, Ying Li 0036, Jonathan Li 0001, Yongtao Yu, José Marcato Junior, Wesley Nunes Gonçalves, Michael A. Chapman
IEEE Trans. Intell. Transp. Syst.3
2021 Multi-Scale Point-Wise Convolutional Neural Networks for 3D Object Segmentation From LiDAR Point Clouds in Large-Scale Environments
abstract
Although significant improvement has been achieved in fully autonomous driving and semantic high-definition map (HD) domains, most of the existing 3D point cloud segmentation methods cannot provide high representativeness and remarkable robustness. The principally increasing challenges remain in completely and efficiently extracting high-level 3D point cloud features, specifically in large-scale road environments. This paper provides an end-to-end feature extraction framework for 3D point cloud segmentation by using dynamic point-wise convolutional operations in multiple scales. Compared to existing point cloud segmentation methods that are commonly based on traditional convolutional neural networks (CNNs), our proposed method is less sensitive to data distribution and computational powers. This framework mainly includes four modules. Module I is first designed to construct a revised 3D point-wise convolutional operation. Then, a U-shaped downsampling-upsampling architecture is proposed to leverage both global and local features in multiple scales in Module II. Next, in Module III, high-level local edge features in 3D point neighborhoods are further extracted by using an adaptive graph convolutional neural network based on the K-Nearest Neighbor (KNN) algorithm. Finally, in Module IV, a conditional random field (CRF) algorithm is developed for postprocessing and segmentation result refinement. The proposed method was evaluated on three large-scale LiDAR point cloud datasets in both urban and indoor environments. The experimental results acquired by using different point cloud scenarios indicate our method can achieve state-of-the-art semantic segmentation performance in feature representativeness, segmentation accuracy, and technical robustness.
Lingfei Ma, Ying Li 0036, Jonathan Li 0001, Weikai Tan, Yongtao Yu, Michael A. Chapman
IEEE Trans. Intell. Transp. Syst.3
2021 Special Issue on 3D Sensing in Intelligent Transportation
abstract
High-Accuracy and high-efficiency 3-D sensing and associated data processing techniques are urgently needed for today’s roadway inventory, infrastructure health monitoring, autonomous driving, connected vehicles, urban modeling, and smart cities. 3D geospatial data acquired by digital photogrammetry or laser scanning or LiDAR systems have become one of the most critical data sources to support the above-mentioned applications. While progress has been made to applying 3D sensory data to those applications related to intelligent transportation systems (ITS), such as road network extraction, platform localization, obstacle avoidance, high-definition map generation, and transportation infrastructure inventory, many essential questions remain regarding the processing and understanding such massive 3D datasets in ITS-related applications. The authors have selected four articles for review in this Special issue. A summary of these articles is outlined below.
Chenglu Wen, Ayman Habib 0001, Jonathan Li 0001, Charles K. Toth, Cheng Wang 0003, Hongchao Fan
IEEE Trans. Intell. Transp. Syst.3
2021 Deep Learning for LiDAR Point Clouds in Autonomous Driving: A Review
abstract
Recently, the advancement of deep learning (DL) in discriminative feature learning from 3-D LiDAR data has led to rapid development in the field of autonomous driving. However, automated processing uneven, unstructured, noisy, and massive 3-D point clouds are a challenging and tedious task. In this article, we provide a systematic review of existing compelling DL architectures applied in LiDAR point clouds, detailing for specific tasks in autonomous driving, such as segmentation, detection, and classification. Although several published research articles focus on specific topics in computer vision for autonomous vehicles, to date, no general survey on DL applied in LiDAR point clouds for autonomous vehicles exists. Thus, the goal of this article is to narrow the gap in this topic. More than 140 key contributions in the recent five years are summarized in this survey, including the milestone 3-D deep architectures, the remarkable DL applications in 3-D semantic segmentation, object detection, and classification; specific data sets, evaluation metrics, and the state-of-the-art performance. Finally, we conclude the remaining challenges and future researches.
Ying Li 0036, Lingfei Ma, Zilong Zhong, Michael A. Chapman, Dongpu Cao, Jonathan Li 0001
IEEE Trans. Neural Networks Learn. Syst.7
2020 Squeeze-and-Attention Networks for Semantic Segmentation
abstract
The recent integration of attention mechanisms into segmentation networks improves their representational capabilities through a great emphasis on more informative features. However, these attention mechanisms ignore an implicit sub-task of semantic segmentation and are constrained by the grid structure of convolution kernels. In this paper, we propose a novel squeeze-and-attention network (SANet) architecture that leverages an effective squeeze-and-attention (SA) module to account for two distinctive characteristics of segmentation: i) pixel-group attention, and ii) pixel-wise prediction. Specifically, the proposed SA modules impose pixel-group attention on conventional convolution by introducing an 'attention' convolutional channel, thus taking into account spatial-channel inter-dependencies in an efficient manner. The final segmentation results are produced by merging outputs from four hierarchical stages of a SANet to integrate multi-scale contexts for obtaining an enhanced pixel-wise prediction. Empirical experiments on two challenging public datasets validate the effectiveness of the proposed SANets, which achieves 83.2 % mIoU (without COCO pre-training) on PASCAL VOC and a state-of-the-art mIoU of 54.4 % on PASCAL Context.
Zilong Zhong, Zhong Qiu Lin, Rene Bidart, Xiaodan Hu, Ibrahim Ben Daya, Wei-Shi Zheng 0001, Jonathan Li 0001, Alexander Wong
CVPR8
2020 Extraction of Power Lines and Pylons from LiDAR Point Clouds Using a GCN-Based Method
abstract
The routine power line inspection is critical to maintain the reliability, availability, and sustainability of electricity supply. As a key part of inspection, power lines and pylons extraction is essential for resource management and power corridor safety, especially in the mountain regions. In this paper, we proposed a deep learning based method to extract power lines and pylons using ALS point clouds. First, a structure information preserved module is designed to mine the relationship of local neighborhood points. Then, a graph convolutional network (GCN) is used as basic module to extract point features. Finally, three categories, power lines, pylons and other objects are segmented from input point clouds. In addition, we provide an effective data enhancement strategy to generate enough samples to train the proposed model. We evaluated our method using a dataset acquired by our ALS scanning system. Experimental results demonstrate that our method is superior to the state-of-the-art methods on descriptiveness and efficiency. The overall accuracy and mean time are 99.1% and 9.3 seconds, respectively.
Wen Li 0005, Zhenlong Xiao, Cheng Wang 0003, Jonathan Li 0001
IGARSS6
2020 Extracting Vehicles in Point Clouds of Underground Parking Lots Based on Graph Convolution
abstract
Three-dimensional point clouds can describe the shape and position of objects more accurately when compared with 2D images, thereby providing richer information for object recognition, detection, and reconstruction tasks. Extracting vehicles in point clouds of underground parking lots can help autonomous vehicles achieve automatic parking. Camera-based perception algorithms will fail in complicate environments, so it is necessary to study algorithms for extracting targets using point cloud data. In this paper, we designed an effective method to extract the vehicles in the underground parking lot. First, the point clouds belonging to the vehicle will be segmented using a neural network based on graph convolution, and then different vehicles will be separated based on clustering. Finally, the minimum bounding box for each car is calculated. The proposed approach achieved much better results on the point cloud dataset than other state-of-the-art methods. Our method achieves 99.6% in Overall Accuracy and 98.5% in Mean IOU (Intersection over Union).
Zhenlong Xiao, Jonathan Li 0001
IGARSS4
2020 Automated Detection of Manhole Covers in MLS Point Clouds Using a Deep Learning Approach
abstract
Road manhole cover works as an important part of road construction. Timely detection can make a great progress in the development of road management. This paper proposes a rapid road manhole detection method using mobile LiDAR with state-of-the-art computer vision and deep learning techniques. Firstly, the road surface data is extracted from mobile laser scanning system(MLS). Then, the 2D geographic reference feature(GRF) images are formed from 3D point cloud. Finally, the object detector using deep learning technology was applied to locate and annotate the road manholes. Also, we adjusted the training model to present the better result with high confidence over 0.90. Compared with the previous method, the proposed method can correctly detect the manhole cover with higher rate of precision and FI-feature at 0.952 and 0.975 respectively, especially in the complex road situation.
Liyuan Qing, Weikai Tan, Jonathan Li 0001
IGARSS4
2020 A Boundary-Enhanced Supervoxel Method for 3D Point Clouds
abstract
This paper presents a boundary-enhanced supervoxel method to solve over-segmentation problems in supervoxel generation of Voxel Cloud Connectivity Segmentation (VCCS). First, we use different searching methods to obtain the neighborhood of each point. Second, three variants of neighbor points are clustered by the local k-means clustering method on points directly instead of on voxels. Finally, a scale metric is used to measure the difference between two points that considers underlying 3D spatial structure of the points. Our proposed is tested on two publicly available benchmark point cloud datasets acquired by mobile laser scanning (MLS) and terrestrial laser scanning (TLS) systems, respectively. Results of the experiments show that the boundary recall approximately 7 and 4 times higher than VCCS for the best results, which our proposed methods are effective, and the cost time is feasible and effective.
Zhengchuan Sha, Qing Zhu 0012, Yiping Chen 0002, Cheng Wang 0003, Abdul Nurunnabi, Jonathan Li 0001
IGARSS6
2020 Early-Season Crop Classification with Radarsat-2 Polarimetric Synthetic Aperture Radar Imagery
abstract
Timely crop classification maps are essential for the agriculture sector to ensure food security and understand the state and trend of crop growth. Though there are several crop monitoring systems in operation, early-season crop classification is still in demand. We developed a robust crop growth estimation technology previously with synthetic aperture radar (SAR) imagery for canola in Canadian Prairies, and we are extending the procedure to enable accurate early-season crop classification. Here we present a dynamic crop classification technique with RADARSAT-2 (RS2) polarimetric SAR (Pol-SAR) imagery for the classification of canola, corn, soybean and wheat, the four major crop types in Canadian Prairies. The procedure achieved over 90% classification accuracy of the major four crop types in the testing area by the end of July.
Weikai Tan, Abhijit Sinha, Yifeng Li 0003, Lingfei Ma, Jonathan Li 0001
IGARSS5
2020 Understanding urban structures and crowd dynamics leveraging large-scale vehicle mobility data
Zhihan Jiang 0001, Yan Liu 0043, Xiaoliang Fan, Cheng Wang 0003, Jonathan Li 0001, Longbiao Chen
Frontiers Comput. Sci.5
2020 Pairwise registration of TLS point clouds by deep multi-scale local features
Wei Li 0151, Cheng Wang 0003, Chenglu Wen, Congren Lin, Jonathan Li 0001
Neurocomputing6
2020 A Convolutional Capsule Network for Traffic-Sign Recognition Using Mobile LiDAR Data With Digital Images
abstract
Traffic-sign recognition plays an important role in road transportation systems. This letter presents a novel two-stage method for detecting and recognizing traffic signs from mobile Light Detection and Ranging (LiDAR) point clouds and digital images. First, traffic signs are detected from mobile LiDAR point cloud data according to their geometrical and spectral properties, which have been fully studied in our previous work. Afterward, the traffic-sign patches are obtained by projecting the detected points onto the registered digital images. To improve the performance of traffic-sign recognition, we apply a convolutional capsule network to the traffic-sign patches to classify them into different types. We have evaluated the proposed framework on data sets acquired by a RIEGL VMX-450 system. Quantitative evaluations show that a recognition rate of 0.957 is achieved. Comparative studies with the convolutional neural network (CNN) and our previous supervised Gaussian-Bernoulli deep Boltzmann machine (GB-DBM) classifier also confirm that the proposed method performs effectively and robustly in recognizing traffic signs of various types and conditions.
Haiyan Guan, Yongtao Yu, Daifeng Peng, Yufu Zang, JianYong Lu, Aixia Li, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.7
2020 Learning to Match Ground Camera Image and UAV 3-D Model-Rendered Image Based on Siamese Network With Attention Mechanism
abstract
Different domain image sensors or imaging mechanisms provide cross-domain images when sensing the same scene. There is a domain shift between cross-domain images so that the image gap between different domains is the major challenge for measuring the similarity of the feature descriptors extracted from different domain images. Specifically, matching ground camera images and unmanned aerial vehicle (UAV) 3-D model-rendered images, which are two kinds of extremely challenging cross-domain images, is a way to establish indirectly the spatial relationship between 2-D and 3-D spaces. This provides a solution for the virtual-real registration of augmented reality (AR) in outdoor environments. However, during matching, handcrafted descriptors and existing learning-based feature descriptors limit the rendered images. In this letter, first, to learn robust and invariant 128-D local feature descriptors for ground camera and rendered images, we present a novel network structure, SiamAM-Net, which embeds the autoencoders with an attention mechanism into the Siamese network. Then, to narrow the gap between the cross-domain images during the optimizing of SiamAM-Net, we design an adaptive margin for the loss function. Finally, we match the ground camera-rendered images by using the learned local feature descriptors and explore the outdoor AR virtual-real registration. Experiments show that the local feature descriptors, learned by SiamAM-Net, are robust and achieve state-of-the-art retrieval performance on the cross-domain image data set of ground camera and rendered images. In addition, several outdoor AR applications also demonstrate the usefulness of the proposed outdoor AR virtual-real registration.
Weiquan Liu, Cheng Wang 0003, Xuesheng Bian, Shangshu Yu, Xiuhong Lin, Shang-Hong Lai, Dongdong Weng, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.9
2020 Toward Efficient 3-D Colored Mapping in GPS-/GNSS-Denied Environments
abstract
Efficient 3-D mapping provides useful and detailed 3-D data for many applications. In this letter, we present a multisensor calibration and mapping method, to provide highly efficient and relatively accurate colored mapping for GPS-/global navigation satellite system-denied environments. The sensor data include 3-D laser scanning point clouds and camera images. A simultaneous localization and mapping (SLAM)-assisted calibration method is first proposed for multiple multibeam light detection and ranging (LiDAR) and multiple camera calibration. An improved SLAM method with loop closure is proposed for 3-D mapping. With the proposed calibration and mapping methods, centimeter-level colored point clouds can be obtained efficiently. The proposed method was tested with both backpacked and car-mounted systems on indoor and outdoor scenes. Experimental results show the effectiveness and efficiency of the proposed calibration and mapping methods.
Chenglu Wen, Yudi Dai, Yan Xia 0003, Yuhan Lian, Jinbin Tan, Cheng Wang 0003, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.7
2020 A Hybrid Capsule Network for Land Cover Classification Using Multispectral LiDAR Data
abstract
Land cover mapping is an effective way to quantify land resources and monitor their changes. It plays an important role in a wide range of applications. This letter proposes a hybrid capsule network for land cover classification using multispectral light detection and ranging (LiDAR) data. First, the multispectral LiDAR data were rasterized into a set of feature images to exploit the geometrical and spectral properties of different types of land covers. Then, a hybrid capsule network composed of an encoder network and a decoder network is trained to extract both high-level local and global entity-oriented capsule features for accurate land cover classification. Quantitative classification evaluations on two data sets show that the overall accuracy, average accuracy, and kappa coefficient of over 97.89%, 94.54%, and 0.9713, respectively, are obtained. Comparative studies with five existing methods confirm that the proposed method performs robustly and accurately in land cover classification using the multispectral LiDAR data.
Yongtao Yu, Haiyan Guan, Dilong Li, Tiannan Gu, Lanfang Wang, Lingfei Ma, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.7
2020 3-D Feature Matching for Point Cloud Object Extraction
abstract
Effective object extraction plays an important role in many point cloud-based applications. This letter proposes a 3-D feature matching framework for point cloud object extraction. To determine the optimal affine transformation parameters for each template feature point, a convex dissimilarity function and the locally affine-invariant geometric constraints are designed to construct the overall objective function. The 3-D feature matching framework is integrated into a point cloud object extraction workflow. Extraction results on six test data sets show that average completeness, correctness, quality, and F1-measure of 0.96, 0.97, 0.93, and 0.96, respectively, are obtained in extracting light poles, vehicles, and palm trees. Comparative studies also confirm that the proposed method performs effectively and robustly, and exhibits superior or compatible performance over the other compared methods.
Yongtao Yu, Haiyan Guan, Dilong Li, Shenghua Jin, Taiyue Chen, Cheng Wang 0003, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.7
2020 Road Manhole Cover Delineation Using Mobile Laser Scanning Point Cloud Data
abstract
Periodical road manhole cover measurement is extremely important to ensure road safety and reduce traffic disasters. This letter proposes an effective method for delineating road manhole covers from mobile laser scanning point cloud data. To improve processing efficiency, first, road surface points are segmented and rasterized into georeferenced intensity images. Then, object-oriented patches are generated through superpixel segmentation and further fed to a convolutional capsule network classifier for manhole cover detection. Finally, manhole covers are accurately delineated through a marked point process of disks. Quantitative evaluations on three data sets show that an average completeness, correctness, quality, and F1-measure of 0.965, 0.961, 0.929, and 0.963, respectively, are obtained. Comparative studies with three existing methods confirm that the proposed method performs superiorly in delineating manhole covers of varying conditions and on complex road surface environments.
Yongtao Yu, Haiyan Guan, Dilong Li, Cheng Wang 0003, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.6
2020 Generative Adversarial Networks and Conditional Random Fields for Hyperspectral Image Classification
abstract
In this paper, we address the hyperspectral image (HSI) classification task with a generative adversarial network and conditional random field (GAN-CRF)-based framework, which integrates a semisupervised deep learning and a probabilistic graphical model, and make three contributions. First, we design four types of convolutional and transposed convolutional layers that consider the characteristics of HSIs to help with extracting discriminative features from limited numbers of labeled HSI samples. Second, we construct semisupervised generative adversarial networks (GANs) to alleviate the shortage of training samples by adding labels to them and implicitly reconstructing real HSI data distribution through adversarial training. Third, we build dense conditional random fields (CRFs) on top of the random variables that are initialized to the softmax predictions of the trained GANs and are conditioned on HSIs to refine classification maps. This semisupervised framework leverages the merits of discriminative and generative models through a game-theoretical approach. Moreover, even though we used very small numbers of labeled training HSI samples from the two most challenging and extensively studied datasets, the experimental results demonstrated that spectral-spatial GAN-CRF (SS-GAN-CRF) models achieved top-ranking accuracy for semisupervised HSI classification.
Zilong Zhong, Jonathan Li 0001, David A. Clausi, Alexander Wong
IEEE Trans. Cybern.2
2020 Characterization of MSS Channel Reflectance and Derived Spectral Indices for Building Consistent Landsat 1-5 Data Record
abstract
The Landsat 1-5 multispectral scanner system (MSS) collected records of land surface mainly during 1972-1992. Investigations on MSS have been relatively limited compared with the numerous investigations on its successors, such as Thematic Mapper (TM) and Enhanced TM Plus (ETM+). The benefits of the Landsat program are not fully accomplished without the inclusion of MSS archives. Investigations on the Landsat 1-5 MSS channel reflectance characteristics wereperformed followed by derived vegetation spectral indices and the Tasseled Cap (TC) transformed features mainly using a collection of synthesized records. On average, the Landsat 4 MSS is generally comparable to the Landsat 5 MSS. The Landsat 1-3 MSSs show disagreement in channel reflectance compared with the Landsat 5 MSS, especially for the red channel (600-700 nm) and the near-infrared channel (700-800 nm). Meanwhile, the relative differences for vegetation spectral indices of the Landsat 3 MSS are mainly from -16% to -5% with the median about -11.5%, while those of the Landsat 2 MSS are mainly from -15% to -7%. Cross-validation tests and two case applications suggested that between-sensor consistency was improved generally through the transformation models generated by ordinary least-squares regression. To improve the consistency of the vegetation indices and the TC greenness, direct strategy employing respective transformation models was more effective than calculations based on the transformed channel reflectance. Considering the shortages of the Landsat MSS archives, further efforts are needed to improve its comparability with observations by other successive Landsat sensors.
Feng Chen 0022, Qiancong Fan, Shenlong Lou, Martin Claverie, Cheng Wang 0003, José Marcato Junior, Wesley Nunes Gonçalves, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.10
2020 TGNet: Geometric Graph CNN on 3-D Point Cloud Segmentation
abstract
Recent geometric deep learning works define convolution operations in local regions and have enjoyed remarkable success on non-Euclidean data, including graph and point clouds. However, the high-level geometric correlations between the input and its neighboring coordinates or features are not fully exploited, resulting in suboptimal segmentation performance. In this article, we propose a novel graph convolution architecture, which we term as Taylor Gaussian mixture model (GMM) network (TGNet), to efficiently learn expressive and compositional local geometric features from point clouds. The TGNet is composed of basic geometric units, TGConv, that conduct local convolution on irregular point sets and are parametrized by a family of filters. Specifically, these filters are defined as the products of the local point features and the neighboring geometric features extracted from local coordinates. These geometric features are expressed by Gaussian weighted Taylor kernels. Then, a parametric pooling layer aggregates TGConv features to generate new feature vectors for each point. TGNet employs TGConv on multiscale neighborhoods to extract coarse-to-fine semantic deep features while improving its scale invariance. Additionally, a conditional random field (CRF) is adopted within the output layer to further improve the segmentation results. Using three point cloud data sets, qualitative and quantitative experimental results demonstrate that the proposed method achieves 62.2% average accuracy on ScanNet, 57.8% and 68.17% mean intersection over union (mIoU) on Stanford Large-Scale 3D Indoor Spaces (S3DIS) and Paris-Lille-3D data sets, respectively.
Ying Li 0036, Lingfei Ma, Zilong Zhong, Dongpu Cao, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.5
2020 Corse-to-Fine Road Extraction Based on Local Dirichlet Mixture Models and Multiscale-High-Order Deep Learning
abstract
Road extraction from remote sensing images is an attractive but difficult task. Gray-value distribution and structure feature information are both crucial for road extraction task. However, existing methods mainly focus on structure feature information which contains morphological shape features and machine learning features, suffering from lots of false positives which are generated at positions having similar structure features but different gray-value distribution with roads. To effectively fuse the two complementary gray-value distribution and structure feature information, we propose a coarse-to-fine road extraction algorithm from remote sensing images. First, at the coarse level, we introduce a local Dirichlet mixture models (LDMM) which utilizing gray-value distribution information to pre-segment images into potential roads and backgrounds. Thus, most backgrounds having different gray-value distribution with roads can be removed firstly. Compared with original Dirichlet mixture models, the LDMM is much faster and more accurate. Next, at the fine level, we introduce a multiscale-high-order deep learning strategy based on ResNet model which can learn robust structure context features for final road extraction step. Based on the results of LDMM, the multiscale-high-order strategy can further remove false positives which have different structure features with roads. Compared with a single scanning size ResNet, our multiscale-high-order strategy can learn higher-order context information, leading to better performances. We test our algorithm on Shaoshan dataset. Experiments illustrate our better performance compared with other six state-of-the-art methods.
Ziyi Chen 0001, Wentao Fan 0001, Bineng Zhong 0001, Jonathan Li 0001, Jixiang Du, Cheng Wang 0003
IEEE Trans. Intell. Transp. Syst.4
2020 Semi-Automated Generation of Road Transition Lines Using Mobile Laser Scanning Data
abstract
This paper recognizes the research gaps and difficulties in generating transition lines (the paths that pass through a road intersection) in road intersections from mobile laser scanning (MLS) point clouds. The proposed method contains three modules: road surface detection, lane marking extraction, and transition line generation. First, the points covering the road surface are extracted using the voxel-based upward growing and the improved region growing. Then, lane markings are extracted and identified according to the multi-thresholding and the geometric filtering. Finally, transition lines are generated through a combination of the lane node structure generation algorithm and the cubic Catmull-Rom spline algorithm. The experimental results demonstrate that transition lines can be successfully generated for both T- and cross-intersections with promising accuracy. In the validation of lane marking extraction using the manually interpreted lane marking points, the method can achieve average precision, recall, and F1-score of 90.80%, 92.07%, and 91.43%, respectively. The success rate of transition line generation is 96.5%. Furthermore, the buffer-overlay-statistics (BOS) method validates that the proposed method can generate lane centerlines and transition lines within 20-cm-level localization accuracy from the MLS point clouds.
Chengming Ye, Jonathan Li 0001, Han Jiang 0005, He Zhao 0007, Lingfei Ma, Michael A. Chapman
IEEE Trans. Intell. Transp. Syst.2
2020 3D Highway Curve Reconstruction From Mobile Laser Scanning Point Clouds
abstract
The point clouds acquired by a vehicle-borne mobile laser scanning (MLS) system have shown great potential for many applications such as intelligent transportation systems, road infrastructure inventories, and high-definition (HD) maps to support the advanced driver-assistance systems (ADAS) and autonomous vehicles (AVs). This paper presents a novel two-step approach to automated detection and reconstruction of three-dimensional (3D) highway curves from MLS point clouds. However, when dealing with noisy, unstructured, dense point clouds, we often face some challenges, most notably in handling of the outliers introduced during road marking detection and in recognition of curve types during 3D curve reconstruction. Our approach is formed by two main algorithms: a detector based on intensity variance and a robust model fitting estimator. The experimental results obtained using both a virtual scan dataset and a real MLS dataset demonstrated that our approach is very promising in handling of the outliers and reconstruction of 3D road curves. Specifically, a relative accuracy of 0.6% has been achieved in estimation of circle radii based on the virtual scan dataset. A comparative study also showed that our road marking detection approach is more effective and more stable than state-of-the-art approaches.
Zongliang Zhang, Jonathan Li 0001, Yulan Guo, Chenhui Yang, Cheng Wang 0003
IEEE Trans. Intell. Transp. Syst.2
2020 DeepSTD: Mining Spatio-Temporal Disturbances of Multiple Context Factors for Citywide Traffic Flow Prediction
abstract
Deep learning techniques have been widely applied to traffic flow prediction, considering underlying routine patterns, and multiple context factors (e.g., time and weather). However, the complex spatio-temporal dependencies between inherent traffic patterns and multiple disturbances have not been fully addressed. In this paper, we propose a two-phase end-to-end deep learning framework, namely DeepSTD to uncover the spatio-temporal disturbances (STD) to predict the citywide traffic flow. In the STD Modeling phase, we propose an STD modeling method to model both the different regional disturbances caused by various region functions and the spatio-temporal propagating effects. In the Prediction phase, we eliminate the STD from the historical traffic flow to enhance the leaning of inherent traffic patterns and combine the STD at the prediction time interval to consider the future disturbances. The experimental results on two real-world datasets demonstrate that DeepSTD outperforms the state-of-the-art methods.
Chuanpan Zheng, Xiaoliang Fan, Chenglu Wen, Longbiao Chen, Cheng Wang 0003, Jonathan Li 0001
IEEE Trans. Intell. Transp. Syst.6
2019 Geometric Multi-Model Fitting by Deep Reinforcement Learning
abstract
This paper deals with the geometric multi-model fitting from noisy, unstructured point set data (e.g., laser scanned point clouds). We formulate multi-model fitting problem as a sequential decision making process. We then use a deep reinforcement learning algorithm to learn the optimal decisions towards the best fitting result. In this paper, we have compared our method against the state-of-the-art on simulated data. The results demonstrated that our approach significantly reduced the number of fitting iterations.
Zongliang Zhang, Hongbin Zeng, Jonathan Li 0001, Yiping Chen 0002, Chenhui Yang, Cheng Wang 0003
AAAI3
2019 LO-Net: Deep Real-Time Lidar Odometry
abstract
We present a novel deep convolutional network pipeline, LO-Net, for real-time lidar odometry estimation. Unlike most existing lidar odometry (LO) estimations that go through individually designed feature selection, feature matching, and pose estimation pipeline, LO-Net can be trained in an end-to-end manner. With a new mask-weighted geometric constraint loss, LO-Net can effectively learn feature representation for LO estimation, and can implicitly exploit the sequential dependencies and dynamics in the data. We also design a scan-to-map module, which uses the geometric and semantic information learned in LO-Net, to improve the estimation accuracy. Experiments on benchmark datasets demonstrate that LO-Net outperforms existing learning based approaches and has similar accuracy with the state-of-the-art geometry-based approach, LOAM.
Qing Li 0032, Shaoyang Chen, Cheng Wang 0003, Xin Li 0003, Chenglu Wen, Ming Cheng 0002, Jonathan Li 0001
CVPR7
2019 RF-Net: An End-To-End Image Matching Network Based on Receptive Field
abstract
This paper proposes a new end-to-end trainable matching network based on receptive field, RF-Net, to compute sparse correspondence between images. Building end-to-end trainable matching framework is desirable and challenging. The very recent approach, LF-Net, successfully embeds the entire feature extraction pipeline into a jointly trainable pipeline, and produces the state-of-the-art matching results. This paper introduces two modifications to the structure of LF-Net. First, we propose to construct receptive feature maps, which lead to more effective keypoint detection. Second, we introduce a general loss function term, neighbor mask, to facilitate training patch selection. This results in improved stability in descriptor training. We trained RF-Net on the open dataset HPatches, and compared it with other methods on multiple benchmark datasets. Experiments show that RF-Net outperforms existing state-of-the-art methods.
Xuelun Shen, Cheng Wang 0003, Xin Li 0003, Zenglei Yu, Jonathan Li 0001, Chenglu Wen, Ming Cheng 0002
CVPR5
2019 Partial 3D Object Retrieval and Completeness Evaluation for Urban Street Scene
abstract
3D objects detected from real-world data are usually incomplete in different degrees. Objects with different degrees of incompleteness should be treated and processed separately. This paper proposes a framework for partial 3D object retrieval and completeness evaluation in an urban street scene based on mobile laser scanning (MLS) point cloud data. The framework consists of three parts. A deep learning method is first used to detect objects from 3D point cloud data. Then, for each detected object, the most similar object in the reference dataset, which contains complete objects, is obtained by a partial 3D shape retrieval method. Last, a completeness evaluation of the detected object is conducted by calculating the completeness index that reflects the integrity of the detected object, and a missing part prediction is given to guide further completion. The proposed framework is validated on the public dataset KITTI and our own point cloud dataset. The experiment includes 3D detection, the partial 3D shape retrieval, and the completeness evaluation. Results show the good performance of the object detection and partial shape retrieval, also a reasonable evaluation of objects completeness.
Chenglu Wen, Xiaotian Sun 0005, Cheng Wang 0003, Jonathan Li 0001
IGARSS5
2019 Non-Reference Quality Evaluation for Indoor 3d Point Clouds
abstract
This paper proposes a novel approach for indoor point clouds quality evaluation, which works well without reference point clouds. In this paper, we mainly evaluate indoor point clouds quality in two aspects: the smoothness of the walls and the degree of occlusion of the walls and the floor. Our approach involves three steps. Firstly, with the S3DIS dataset, we use a deep learning method to train a detector to label walls and floor from indoor scenes. Next, we calculate the normal vector of the wall and the normal vector of each point on the wall. The degree of smoothness of the wall is judged according to the angle between the normal vector of the wall and the normal vectors of the points. Finally, according to the cause of occlusion (objects on the floor or in front of the wall), the occlusion degree of the wall and the floor is obtained. The experimental results demonstrate that the proposed method is suitable for non-reference indoor point cloud quality evaluation.
Yuhan Lian, Chenglu Wen, Cheng Wang 0003, Jonathan Li 0001
IGARSS4
2019 Long-Term Trend of Ground-Level PM2.5 Concentrations Over 2012-2017 in China
abstract
Ambient suspended fine particulate matter (PM2.5) is a greatest environmental risk factor for premature mortality. We adopted aerosol optical depth (AOD) retrieved from the Moderate Resolution Imaging Spectroradiometer (MODIS) instrument to produce annual-mean PM2.5 concentrations from 2012 to 2017 with a spatial resolution of 3km. A geographically weighted regression model was conducted using vertical- and hydroscopic-corrected AOD and meteorological data. The PM2.5 estimates were validated by the ground measurements, with R2and RMSE (MPE) of 0.79 and 18.26 (12.03) μg/m3. The results show that national average of PM2.5 concentration represented a 31% decline over five years, from 69.37 μg/m3in 2013 to 43.85 μg/m3in 2017, after a slightly rise (6%) during 2012-2013. Significant reduction was revealed in the Beijing-Tianjin-Hebei region, decreasing by 37.31% from 2013 to 2017. Despite low decline in some southeastern provinces, the national-mean PM2.5 concentration has decreased by 31%, indicating the effectiveness of the control policies issued by Chinese government in 2013. Nevertheless, efforts to improve air quality are still required to further reduce the mass concentration in China.
Gaoxiang Zhou, Rebecca K. Saari, Jonathan Li 0001
IGARSS4
2019 Reconstruction of 3D Zebra Crossings from Mobile Laser Scanning Point Clouds
abstract
This paper presents a novel method for reconstruction of three-dimensional (3D) zebra crossings from mobile laser scanning (MLS) point clouds. Firstly, we extract the zebra crossings from the 3D point cloud data in data preprocessing. Secondly, the fitting model is generated by seven parameters to determine one plane commonly and then calculating similarity for fitting the zebra crossings point clouds. Finally, the cuckoo search algorithm is used to adjust the parameters and optimize the results to obtain the optimal model and the geometric information of the zebra crossings can be acquired simultaneously. The proposed algorithm is tested on a set of point-clouds acquired by a RIEGL VMX-450 LiDAR system. The experimental results show the feasibility and stability of our method and the zebra crossings of urban road can be automatically and effectively reconstructed.
Hongbin Zeng, Yiping Chen 0002, Zongliang Zhang, Cheng Wang 0003, Jonathan Li 0001
IGARSS5
2019 Slam-Based Multi-Sensor Backpack Lidar Systems in Gnss-Denied Environments
abstract
The backpack LiDAR system is a highly efficient device for indoor positioning and navigation. With the use of multi-sensor interactions inside of the backpack, it can solve the problem of undetermined trajectories or locations in a GNSS-denied environment. This paper presents a multi-sensor-based backpack LiDAR system for mapping in GNSS-denied environments. With this device, we solve the problem of inaccurate positioning caused by the inability to receive GNSS signals in a small-sized lidar portable device, also reduces the cumulative error caused by the IMU module through the analysis of the vibration pattern. Through the demonstration of the results, the method is of great significance for 3D reconstruction of the indoor environment and object detection.
Dedong Zhang, Yiping Chen 0002, John S. Zelek, Jonathan Li 0001
IGARSS5
2019 Generalized Zero-Shot Vehicle Detection in Remote Sensing Imagery via Coarse-to-Fine Framework
abstract
Vehicle detection and recognition in remote sensing images are challenging, especially when only limited training data are available to accommodate various target categories. In this paper, we introduce a novel coarse-to-fine framework, which decomposes vehicle detection into segmentation-based vehicle localization and generalized zero-shot vehicle classification. Particularly, the proposed framework can well handle the problem of generalized zero-shot vehicle detection, which is challenging due to the requirement of recognizing vehicles that are even unseen during training. Specifically, a hierarchical DeepLab v3 model is proposed in the framework, which fully exploits fine-grained features to locate the target on a pixel-wise level, then recognizes vehicles in a coarse-grained manner. Additionally, the hierarchical DeepLab v3 model is beneficially compatible to combine the generalized zero-shot recognition. To the best of our knowledge, there is no publically available dataset to test comparative methods, we therefore construct a new dataset to fill this gap of evaluation. The experimental results show that the proposed framework yields promising results on the imperative yet difficult task of zero-shot vehicle detection and recognition.
Yongtan Luo, Liujuan Cao, Baochang Zhang 0001, Guodong Guo, Cheng Wang 0003, Jonathan Li 0001, Rongrong Ji
IJCAI7
2019 Ground Camera Images and UAV 3D Model Registration for Outdoor Augmented Reality
abstract
This paper presents a novel virtual-real registration approach for augmented reality (AR) in large-scale outdoor environments. Essentially, it is a pose estimation for the mobile camera images (ground camera images) in 3D model recovered by Unmanned Aerial Vehicle (UAV) image sequence via Structure-From-Motion (SFM) technology. The approach considers to indirectly establish the spatial relationship between 2D and 3D space by inferring the transformation relationship between the ground camera images and the UAV 3D model rendered images. Specifically, the proposed approach can overcome the positioning errors, which are deterioration and drift in the GPS, and deviation of orientation. The experimental results demonstrate the possibility of the proposed virtual-real registration approach, and show that the approach is robust, efficient and intuitive for AR in large-scale outdoor environments.
Weiquan Liu, Cheng Wang 0003, Shang-Hong Lai, Dongdong Weng, Xuesheng Bian, Xiuhong Lin, Xuelun Shen, Jonathan Li 0001
VR9
2019 NormalNet: A voxel-based CNN for 3D object classification and retrieval
Cheng Wang 0003, Ming Cheng 0002, Ferdous Sohel, Mohammed Bennamoun, Jonathan Li 0001
Neurocomputing5
2019 Mini-batch algorithms with online step size
Cheng Wang 0003, Zhemin Zhang, Jonathan Li 0001
Knowl. Based Syst.4
2019 Robust procedural model fitting with a new geometric similarity estimator
Zongliang Zhang, Jonathan Li 0001, Yulan Guo, Xin Li 0003, Yangbin Lin, Guobao Xiao, Cheng Wang 0003
Pattern Recognit.2
2019 Accelerated stochastic gradient descent with step size selection rules
Cheng Wang 0003, Zhemin Zhang, Jonathan Li 0001
Signal Process.4
2019 Joint 2-D-3-D Traffic Sign Landmark Data Set for Geo-Localization Using Mobile Laser Scanning Data
abstract
This paper presents a framework to build a joint 2-D-3-D traffic sign landmark data set for geo-localization using mobile laser scanning (MLS) data. The MLS data include 3-D point clouds and corresponding multi-view images. First, an integrated method, based on a deep learning network and the retro-reflective properties of traffic signs, is developed to accurately extract traffic signs from MLS point clouds. Next, the semantic and spatial properties of the traffic signs (type, location, position, and geometric characteristics) are obtained. Then, a joint 2-D-3-D traffic sign landmark data set is built, and a semantic-spatial organization graph is used to organize the traffic sign data set. Last, based on the traffic sign landmark data set, a geo-localization method for a driving car is proposed to estimate the driving trajectory. It can be used for auxiliary positioning of autonomous vehicles. Experimental results demonstrate the reliability of our proposed method for traffic sign detection and the potential of building 2-D-3-D traffic sign landmark data set for driving trajectory estimation from MLS data.
Changbin You, Chenglu Wen, Cheng Wang 0003, Jonathan Li 0001, Ayman Habib 0001
IEEE Trans. Intell. Transp. Syst.4
2018 Generative Adversarial Networks and Probabilistic Graph Models for Hyperspectral Image Classification
abstract
High spectral dimensionality and the shortage of annotations make hyperspectral image (HSI) classification a challenging problem. Recent studies suggest that convolutional neural networks can learn discriminative spatial features, which play a paramount role in HSI interpretation. However, most of these methods ignore the distinctive spectral-spatial characteristic of hyperspectral data. In addition, a large amount of unlabeled data remains an unexploited gold mine for efficient data use. Therefore, we proposed an integration of generative adversarial networks (GANs) and probabilistic graphical models for HSI classification. Specifically, we used a spectral-spatial generator and a discriminator to identify land cover categories of hyperspectral cubes. Moreover, to take advantage of a large amount of unlabeled data, we adopted a conditional random field to refine the preliminary classification results generated by GANs. Experimental results obtained using two commonly studied datasets demonstrate that the proposed framework achieved encouraging classification accuracy using a small number of data for training.
Zilong Zhong, Jonathan Li 0001
AAAI2
2018 LiDAR-Video Driving Dataset: Learning Driving Policies Effectively
abstract
Learning autonomous-driving policies is one of the most challenging but promising tasks for computer vision. Most researchers believe that future research and applications should combine cameras, video recorders and laser scanners to obtain comprehensive semantic understanding of real traffic. However, current approaches only learn from large-scale videos, due to the lack of benchmarks that consist of precise laser-scanner data. In this paper, we are the first to propose a LiDAR-Video dataset, which provides large-scale high-quality point clouds scanned by a Velodyne laser, videos recorded by a dashboard camera and standard drivers' behaviors. Extensive experiments demonstrate that extra depth information help networks to determine driving policies indeed.
Yiping Chen 0002, Jingkang Wang, Jonathan Li 0001, Cewu Lu, Cheng Wang 0003
CVPR3
2018 Sensing Urban Structures and Crowd Dynamics with Mobility Big Data
Yan Liu 0043, Longbiao Chen, Linjin Liu, Xiaoliang Fan, Cheng Wang 0003, Jonathan Li 0001
GPC7
2018 Joint Denoising and Super-Resolution via Generative Adversarial Training
abstract
Single image denoising and super-resolution are sitting in the core of various image processing and pattern recognition applications. Typically, these two tasks are handled separately, without regarding to joint reinforcement and learning. The former deals with equal-size pixel-to-pixel translation, while the latter deals with scaling up amount of input pixels. In this paper, we propose a Generative Adversarial Network(GAN) towards joint learning of single image denoising and super-resolution. In principle, our design allows both tasks to share several common building blocks, with the linking between both outputs to reinforce each other. Such a reinforcement is accomplished via designing a novel generative network through optimizing a novel loss function to achieve both denoising and super-resolution. Quantitatively comparing to a set of alternative approaches and baselines, the experiment demonstrated superior performance our method in denoising and super-resolution with high upscaling factors.
Li Chen 0007, Wen Dan, Liujuan Cao, Cheng Wang 0003, Jonathan Li 0001
ICPR5
2018 Weakly Supervised Vehicle Detection in Satellite Images via Multiple Instance Ranking
abstract
Given the difficulty in labeling sufficient amount of instances across different resolutions and imaging environment of satellite images, weakly supervised vehicle detection is with great importance for satellite images analysis and processing. To prevent such cumbersome and meticulous manual annotation, naturally we have introduced the weakly supervised detection that has recently explosively prevalent in ordinary viewing angle images. Our program merely stands in need of region-level group annotation, i.e., whether this district convers vehicle(s) without plainly pointing out the coordinates of vehicles. There are two major problems are often encountered for Weakly Supervised Object Detection. One is that it is often chooses only a most expressive instance contains multiple target objects which often have a bigger probability when selecting a target block. For this problem, the number of vehicles can be estimated based on the object counting, a combinatorial selection algorithm can be used to select patch which contains at most one vehicle instance. Another problem is that precise object positioning becomes more difficult due to the lack of instance-level supervision. This problem can be optimized by a progressive learning strategy. Experiments was carried on wide-ranging remote sensing dataset and achieved better results compared to the state-of-the-art weakly supervised vehicle detection schemes.
Yihan Sheng, Liujuan Cao, Cheng Wang 0003, Jonathan Li 0001
ICPR4
2018 Estimating PM 2.5 in British Columbia Before and After Wildfires using 3 KM Modis AOD Products from February to August 2017
abstract
Particulate Matter, representing fine particles with diameters smaller than 2.5μm (PM 2.5), are considered to be harmful for both public health and environment. This paper estimated PM 2.5 concentrations in British Columbia (BC), Canada from February to August 2017 in order to compare PM 2.5 levels before and after wildfires by using ground-level PM 2.5 measurements, 3 km Moderate Resolution Imaging Spectroradiometer (MODIS) aerosol optical depth (AOD) products and the Geographically Weighted Regression (GWR) model. Additional meteorological data was also utilized to perfect the model. The results showed that the variation of AOD highly represents the distribution of PM 2.5, and PM 2.5 concentrations in July and August were much higher than previous month, which indicated that there was a rapid increase in PM 2.5 levels after the wildfires. In addition, the results also demonstrated that the 3 km MODIS AOD products have high accuracy on providing spatial information when estimating PM 2.5 levels.
Mengge Chen, Jonathan Li 0001
IGARSS4
2018 A Modified Framework for Ship Detection from Compact Polarization SAR Image
abstract
In recent years, the compact polarimetric SAR (CP SAR) imaging mode has received much attention because of its advantages in swath width compared to the quad-polarization mode and in information of scattering targets compared to the linear dual-polarization mode, respectively. For object detection (e.g., ship, sea ice, and drilling platform) in maritime monitoring which has been widely discussed, it is still a challenge to eliminate the false alarms caused by ocean clutter. Accordingly, a modified feature-based framework for ship detection using CP SAR data was proposed in this paper. In particular, a Guard-filter was used after feature extraction to reduce the effect of ocean clutter. Several simulated CP SAR imageries based on Gaofen-3 quad-polarization SAR image were selected for further investigations on ship detection. A SVM classifier was included in the framework to do a pixel-wise classification. The result shows that the proposed method has a better performance compared to traditional methods (e.g., constant false alarm rate and polarimetric whitening filter) and the feature-based method without filtering. Overall false alarm rates are decreased by 35% with the proposed method.
Qiancong Fan, Feng Chen 0022, Ming Cheng 0002, Cheng Wang 0003, Jonathan Li 0001
IGARSS5
2018 Dection and Health Analysis of Individual Tree in Urban Environment with Multi-Sensor Platform
abstract
With the technology enhanced, 3D mobile light detection and ranging (LiDAR) can produce more accurate 3D information for the objects. Meanwhile, hyperspectral remote sensing has more number of wavelengths and provides a higher resolution spectrum of objects. This paper proposes a multi-sensor platform to provide these two data for health detection at the individual tree level in urban environments. We firstly locate and segment the suspected tree objects by ground removal and Euclidean distance clustering. Then we take use of spectrum to remove non-tree objects, e.g., buildings, light poles. After that, we use LiDAR data to compute the geometric parameters of each tree and hyperspectral data to analyze its health situation.
Yunhe Feng, Chenglu Wen, Pengdi Huang, Cheng Wang 0003, Jonathan Li 0001
IGARSS5
2018 Segment-Based Traffic Sign Detection from Mobile Laser Scanning Data
abstract
This paper presents a segment-based traffic sign detection method using vehicle-borne mobile laser scanning (MLS) data. This method has three steps: road scene segmentation, clustering and traffic sign detection. The non-ground points are firstly segmented from raw MLS data by estimating road ranges based on vehicle trajectory and geometric features of roads (e.g., surface normals and planarity). The ground points are then removed followed by obtaining non-ground points where traffic signs are contained. Secondly, clustering is conducted to detect the traffic sign segments (or candidates) from the non-ground points. Finally, these segments are classified to specified classes. Shape, elevation, intensity, 2D and 3D geometric and structural features of traffic sign patches are learned by the support vector machine (SVM) algorithm to detect traffic signs among segments. The proposed algorithm has been tested on a MLS point cloud dataset acquired by a Leador system in the urban environment. The results demonstrate the applicability of the proposed algorithm for detecting traffic signs in MLS point clouds.
Ying Li 0036, Lingfei Ma, Yuchun Huang, Jonathan Li 0001
IGARSS4
2018 Discriminative Learning of Point Cloud Feature Descriptors Based on Siamese Network
abstract
It is challenging to direct extract the feature descriptors of the object in the point cloud, although deep learning has been widely used with the classification and detection in the point cloud, those methods hidden feature presentation in the network. Since the point cloud scanned by the Laser Scanner usually have different point density, unordered and even the different occlusion, which go beyond the reach of handcrafted descriptors, e.g. FPH, FPFH, VFH, ROPS. In this paper, we aim to direct extract the feature descriptors of the point cloud object through the raw point cloud. Inspired by the recent success of the Siamese networks[6], PointNet[7] and PointNet++[8], we propose a novel network to direct extract the feature descriptors of the whole point cloud object. We train our network with the Euclidean distance as the loss function which reflects feature descriptors similarity. The experiment object datasets were acquired by Mobile Laser Scanning (MLS) system which contains 6 categories. Experiment result shows that our network has a robust generalization, which can well direct extract the feature descriptors of the whole point cloud object.
Xuelun Shen, Cheng Wang 0003, Chenglu Wen, Weiquan Liu, Xiaotian Sun 0005, Jonathan Li 0001
IGARSS6
2018 Rural Road Networks Matching Via Extending Line
abstract
Road network matching has played an important role in road network extraction and update, yet has got extensive researching during the recent decades. Differ from previous road matching methods focus mainly on the city areas, which have accurate and regular road networks, this paper aim to address the matching between incomplete ground survey road network and extracted road network from remote sensing images. Specifically, we propose an extending line based matching scheme to calculate the road primitive similarity by taking into account the surrounding connections and contextual information. The experimental results show that the proposed method is able to provide high quality matching results, even the ground survey data are very different from the extracted road network of the satellite. Thus makes it possible to implement the road network update for the wide rural regions without interested ground survey road network data.
Yiping Chen 0002, Cheng Wang 0003, Jonathan Li 0001
IGARSS5
2018 Estimation of Forest Trees Diameter from Terrestrial Laser Scanning Point Clouds Based on a Circle Fitting Method
abstract
In forest monitoring and management, any rational decision needs to be based on forest parameters. The diameter at breast height (DBH) of a tree is considered to be the most significant parameter among them. This paper presents a novel method for extracting tree stems and estimating DBH of trees in a forest environment from 3D point clouds data acquired by a terrestrial laser scanning (TLS) system. In the proposed method, a downward-growing algorithm is used to extract individual tree stems and DBH of trees are estimated by the circle fitting algorithm. This proposed method can avoid errors caused from tilted trees by estimating a plane perpendicular to the tree stem. With this method, 17 trees were extracted from single-scan point cloud data consisting of 21 trees. The estimated DBH had a bias of 0.38 cm and a root mean squared error of 1.76 cm, These experiment results show the feasibility of the proposed method.
Rongren Wu, Yiping Chen 0002, Cheng Wang 0003, Jonathan Li 0001
IGARSS4
2018 A Comparison of Ice Freeboard Measurement by Icesat and Envisat Altimeters
abstract
Remote sensing plays an essential role in the continuous monitoring of sea ice in polar regions. This study fills the gap in literature by conducting a cross-mission analysis between ICESat and EnviSat from 2005 to 2009. The methodology consists of two parts: 1) Compare ocean surface elevation measurement. 2) Deriving freeboard from ICESat and EnviSat independently. Results show that: 1) There is a strong correlation between ICESat and EnviSat measured surface elevation. 2) Despite the penetration ability of Ku band radar altimeter, EnviSat measured elevation is consistently above ICESat. 3) According to the estimated freeboard, Southern Beaufort Sea is dominated by first year ice.
Weiya Ye, Jonathan Li 0001
IGARSS2
2018 Traffic Flow Prediction Based on Cascaded Artificial Neural Network
abstract
The prediction of traffic flow is of great significance for the prevention of accidents, the avoidance of congestion and the dispatch of command center. Considering the complexity of traffic data in reality, it is an extraordinarily challenging task to forecast accurately from historical patterns. In this paper, we propose a method based on the cascaded artificial neural network (CANN) to predict traffic flow at positions. In order to express the spatial correlation of traffic data, the actual road network distance is introduced in our model. The realworld data derived from video surveillance cameras in Xiamen is used in the experiment which is compared with five baselines. To the best of our knowledge, this is the first time that CANN is applied to forecast traffic flow. The experimental results demonstrate that the CANN method has superior performance. In addition, We also discuss the impact of some external factors such as temperature, weather and holidays on the prediction results.
Zejian Kang, Zhiyou Hong, Zhemin Zhang, Cheng Wang 0003, Jonathan Li 0001
IGARSS6
2018 Combination of Crop Growth Model and Radiation Transfer Model with Remote Sensing Data Assimilation for Fapar Estimation
abstract
Accurate assessment of Fraction of Absorbed Photosynthetically Active Radiation (FAPAR) in large scale is significant for crop productivity estimation and climate change analysis. The object of study is to simulate FAPAR in the rice growth period for exploring photosynthetic capacity of rice in large-scale. The daily FAPAR is calculated based on a coupled model consisting of the leaf-canopy radiative transfer model (PROSAIL) and the World Food Study Model (WOFOST). Due to the limitation of the PROSAIL and WOFOST model, we introduced the remote sensing data assimilation method, which assimilated the Normalized Difference Vegetation Index (NDVI) into the coupled model, to improve the prediction accuracy and carry out the large-scale application. The results show high correlation between the simulated FAPAR and the measured data, with the determinate coefficient$(R^{2})$of 0.75 in the study area. The spatial distribution of FAPAR is uniform in flat area, which indicates that the rice in the whole study area has well growth condition and photosynthetic capacity. This study suggest that the coupled model (PROSAIL + WOFOST) assimilated with remote sensing data could accurately simulate daily FAPAR during the crop growth period.
Gaoxiang Zhou, Xiangnan Liu, Jonathan Li 0001
IGARSS4
2018 Extraction of Building Windows from Mobile Laser Scanning Point Clouds
abstract
This study recognizes the significance and considerable commercial applications in creating Level of Detail (LoD) building models for 3D city models generation. Accordingly, this paper proposes a novel method to identify and extract window frames on building facades from Mobile Laser Scanning (MLS) point clouds. The proposed method can typically be regarded as a stepwise procedure. Firstly, a voxel-based upward-growing method is applied to distinguish non-ground points from ground points. Next, outliers are filtered out from non-ground points by statistical analysis. Then, all the remaining non-ground points are clustered based on the conditional Euclidean clustering algorithm to segment out building facades. A volumetric box is afterward created to store façade points so that neighbors of each point can be operated. Finally, a manipulator is applied according to the structural characteristics of window frames to extract the potential window points. Quantitative evaluations based on 2D validation and 3D validation were both conducted. In the 2D validation, the lowest F1-measure of the test datasets is 0.740, and the highest can be 0.977. While in the 3D validation, the lowest precision of the test dataset is 79.58%, and the highest can be 97.96%. The results demonstrate the proposed method can successfully extract the rectangular or curved windows in the test datasets with promising accuracies to support the generation of LoD3 building models.
Menglan Zhou, Lingfei Ma, Ying Li 0036, Jonathan Li 0001
IGARSS4
2018 H-Net: Neural Network for Cross-domain Image Patch Matching
abstract
Describing the same scene with different imaging style or rendering image from its 3D model gives us different domain images. Different domain images tend to have a gap and different local appearances, which raise the main challenge on the cross-domain image patch matching. In this paper, we propose to incorporate AutoEncoder into the Siamese network, named as H-Net, of which the structural shape resembles the letter H. The H-Net achieves state-of-the-art performance on the cross-domain image patch matching. Furthermore, we improved H-Net to H-Net++. The H-Net++ extracts invariant feature descriptors in cross-domain image patches and achieves state-of-the-art performance by feature retrieval in Euclidean space. As there is no benchmark dataset including cross-domain images, we made a cross-domain image dataset which consists of camera images, rendering images from UAV 3D model, and images generated by CycleGAN algorithm. Experiments show that the proposed H-Net and H-Net++ outperform the existing algorithms. Our code and cross-domain image dataset are available at https://github.com/Xylon-Sean/H-Net.
Weiquan Liu, Xuelun Shen, Cheng Wang 0003, Zhihong Zhang 0001, Chenglu Wen, Jonathan Li 0001
IJCAI6
2018 Random Barzilai-Borwein step size for mini-batch algorithms
Cheng Wang 0003, Zhemin Zhang, Jonathan Li 0001
Eng. Appl. Artif. Intell.4
2018 Mini-batch algorithms with Barzilai-Borwein update step
Cheng Wang 0003, Jonathan Li 0001
Neurocomputing4
2018 Deep mobile traffic forecast and complementary base station clustering for C-RAN optimization
Longbiao Chen, Dingqi Yang, Daqing Zhang 0001, Cheng Wang 0003, Jonathan Li 0001, Thi Mai Trang Nguyen
J. Netw. Comput. Appl.5
2018 Line Structure-Based Indoor and Outdoor Integration Using Backpacked and TLS Point Cloud Data
abstract
This letter presents a line structure-based method for integration of centimeter-level indoor backpacked scanning point clouds and millimeter-level outdoor terrestrial laser scanning point clouds. Using 3-D lines for registration, instead of matching points directly, can improve the robustness of the method and adapt to multisource point cloud data of different qualities. Considering the limited overlapping between indoor and outdoor scenes, line structures are extracted from overlapped wall areas that may be included in interior and exterior data. Here, a patch-based method labels a point cloud into wall, ceiling, floor categories, as well as assigning the candidate overlapping walls. Then, lines structures are extracted from the wall plane point cloud. Potential door and window line structures are detected and refined for point cloud registration. Last, an iterative closest point-based method is used to fine tune the registration results. Our results show that the proposed method effectively integrates a promising map of indoor and outdoor scenes.
Chenglu Wen, Xiaotian Sun 0005, Shiwei Hou, Jinbin Tan, Yudi Dai, Cheng Wang 0003, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.7
2018 Semantic Labeling of Mobile LiDAR Point Clouds via Active Learning and Higher Order MRF
abstract
Using mobile Light Detection and Ranging point clouds to accomplish road scene labeling tasks shows promise for a variety of applications. Most existing methods for semantic labeling of point clouds require a huge number of fully supervised point cloud scenes, where each point needs to be manually annotated with a specific category. Manually annotating each point in point cloud scenes is labor intensive and hinders practical usage of those methods. To alleviate such a huge burden of manual annotation, in this paper, we introduce an active learning method that avoids annotating the whole point cloud scenes by iteratively annotating a small portion of unlabeled supervoxels and creating a minimal manually annotated training set. In order to avoid the biased sampling existing in traditional active learning methods, a neighbor-consistency prior is exploited to select the potentially misclassified samples into the training set to improve the accuracy of the statistical model. Furthermore, lots of methods only consider short-range contextual information to conduct semantic labeling tasks, but ignore the long-range contexts among local variables. In this paper, we use a higher order Markov random field model to take into account more contexts for refining the labeling results, despite of lacking fully supervised scenes. Evaluations on three data sets show that our proposed framework achieves a high accuracy in labeling point clouds although only a small portion of labels is provided. Moreover, comparative experiments demonstrate that our proposed framework is superior to traditional sampling methods and exhibits comparable performance to those fully supervised models.
Huan Luo 0001, Cheng Wang 0003, Chenglu Wen, Ziyi Chen 0001, Dawei Zai, Yongtao Yu, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.7
2018 Spectral-Spatial Residual Network for Hyperspectral Image Classification: A 3-D Deep Learning Framework
abstract
In this paper, we designed an end-to-end spectral-spatial residual network (SSRN) that takes raw 3-D cubes as input data without feature engineering for hyperspectral image classification. In this network, the spectral and spatial residual blocks consecutively learn discriminative features from abundant spectral signatures and spatial contexts in hyperspectral imagery (HSI). The proposed SSRN is a supervised deep learning framework that alleviates the declining-accuracy phenomenon of other deep learning models. Specifically, the residual blocks connect every other 3-D convolutional layer through identity mapping, which facilitates the backpropagation of gradients. Furthermore, we impose batch normalization on every convolutional layer to regularize the learning process and improve the classification performance of trained models. Quantitative and qualitative results demonstrate that the SSRN achieved the state-of-the-art HSI classification accuracy in agricultural, rural-urban, and urban data sets: Indian Pines, Kennedy Space Center, and University of Pavia.
Zilong Zhong, Jonathan Li 0001, Zhiming Luo, Michael A. Chapman
IEEE Trans. Geosci. Remote. Sens.2
2018 3-D Road Boundary Extraction From Mobile Laser Scanning Data via Supervoxels and Graph Cuts
abstract
Effective extraction of road boundaries plays a significant role in intelligent transportation applications, including autonomous driving, vehicle navigation, and mapping. This paper presents a new method to automatically extract 3-D road boundaries from mobile laser scanning (MLS) data. The proposed method includes two main stages: supervoxel generation and 3-D road boundary extraction. Supervoxels are generated by selecting smooth points as seeds and assigning points into facets centered on these seeds using several attributes (e.g., geometric, intensity, and spatial distance). 3-D road boundaries are then extracted using the α-shape algorithm and the graph cuts-based energy minimization algorithm. The proposed method was tested on two data sets acquired by a RIEGL VMX-450 MLS system. Experimental results show that road boundaries can be robustly extracted with an average completeness over 95%, an average correctness over 98%, and an average quality over 94% on two data sets. The effectiveness and superiority of the proposed method over the state-of-the-art methods is demonstrated.
Dawei Zai, Jonathan Li 0001, Yulan Guo, Ming Cheng 0002, Yangbin Lin, Huan Luo 0001, Cheng Wang 0003
IEEE Trans. Intell. Transp. Syst.2
2017 Auto-Annotation of 3D Objects via ImageNet
abstract
Automatic annotation of 3D objects in cluttered scenes shows its great importance to a variety of applications. Nowadays, 3D point clouds, a new 3D representation of real-world objects, can be easily and rapidly collected by mobile LiDAR systems, e.g. RIEGL VMX-450 system. Moreover, the mobile LiDAR system can also provide a series of consecutive multi-view images which are calibrated with 3D point clouds. This paper proposes to automatically annotate 3D objects of interest in point clouds of road scenes by exploiting a multitude of annotated images in image databases, such as LabelMe and ImageNet. In the proposed method, an object detector trained on the annotated images is used to locate the object regions in acquired multi-view images. Then, based on the correspondences between multi-view images and 3D point clouds, a probabilistic graphical model is used to model the temporal, spatial and geometric constraints to extract the 3D objects automatically. A new dataset was built for evaluation and the experimental results demonstrate a satisfied performance on 3D object extraction.
Huan Luo 0001, Cheng Wang 0003, Jonathan Li 0001
AAAI3
2017 Unlicensed Taxis Detection Service Based on Large-Scale Vehicles Mobility Data
abstract
Unlicensed taxis are widely considered as major obstacles to city traffic regulation and public safety. Thus, many governments have issued restrictions for car-hailing services and alleged that the use of unlicensed vehicles was illegal. However, it is very challenging that traffic administrative enforcements face limited manpower to prohibit unlicensed taxis, due to costly and time-consuming procedure of on-site evidence collection. In this paper, we propose an effective service to incorporate human mobility mechanism into unlicensed taxis detection from massive city-wide vehicles. We first extract 276 spatio-temporal features, which are grouped into two categories, including daily behaviors and sustainable behaviors to capture the mobility characteristics of unlicensed taxis. Second, we investigate the detection accuracy of three machine learning techniques, viz. support vector machines, decision tree, and logical regression. We illustrate our approach using real-world vehicle license plate recognition dataset in Xiamen, China, which contains 336 million passing records for 6.2 million vehicles filmed by 439 devices in August 2016. Experimental results reveal that LR outperforms SVM and DT in prediction accuracy and F-score measurement, while SVM is capable of identifying the largest number of unlicensed taxis.
Xiaoliang Fan, Xiao Liu 0004, Chuanpan Zheng, Longbiao Chen, Cheng Wang 0003, Jonathan Li 0001
ICWS7
2017 Post-typhoon assessment of surface greenness disturbance using Landsat series observations
abstract
Extreme climate events are projected to increase under the context of global warming associated with the increase in greenhouse gas emissions. In particular, coastal regions which are also characterized with significant urbanization will be vulnerable to the extreme events, such as severe typhoon. Timely assessment and accurate information on the extent and severity of the damage caused by extreme event is necessary to better facilitate decision-making for disaster alleviation and post-recovery. Satellite remote sensing has an important advantage in assessing the impacts caused by extreme events on natural environment and socio-economic dimension at a variety of spatial and temporal scales. Surface greenness disturbance mainly over Xiamen, China, caused by Typhoon Meranti (2016) was investigated using a pair of observations by Landsat 7 ETM+ and Landsat 8 OLI. Significant decreases in greenness were detected, which mainly located along the Maluan Bay and the settlements in suburban and rural areas surrounding the track of typhoon.
Feng Chen 0022, Jonathan Li 0001, Cheng Wang 0003
IGARSS2
2017 Use of ground penetrating radar for detecting underground holes in urban areas: XMU's experience
abstract
This paper presents a cosine-based back-projection(CBP) algorithm for ground penetrating radar(GPR) imaging to detect underground holes. Compared with the classic back-projection imaging algorithm, the CBP algorithm provides a better performance due to better effect of clutter suppression. In the proposed algorithm, a cosine-based measure function is established to describe the similar feature between every two different echo signals to achieve excellent artifact suppression. Then numerical simulation and experimental data set were used to test the SBP algorithm. The experimental data set was gained in Double Han Road, Siming District of Xiamen, China, and acquired by the the Latvia radar system, zond-12e model, which has a 115 kHz transmitting frequency, 40/80/160/320 Hz scanning frequency, ±40dB receive gain and 2000ns time-window width. The results fully demonstrate the effectiveness and superiority of CBP algorithm.
Zhiyou Hong, Jonathan Li 0001, Zhenmiao Deng, Yiping Chen 0002
IGARSS3
2017 Combining optial-thermal remote sensing data and topographic slope for the identification of debris-covered glaciers
abstract
Debris-covered glaciers are an important component of glacier systems on Earth. In this paper, using Landsat Thematic Mapper (TM) images, MOD05 products and digital elevation model (DEM) as data source. Land surface reflectance, land surface temperature (LST) and topographic slope are retrieved from optical bands, thermal band and DEM, respectively. A rule set that combines optical reflectance, LST and slope is proposed to identify debris-covered glaciers. The experimental results indicate that when introducing LST, the overall accuracy of identification increase by approximately 2.38% compared to that without using LST, and the majority of misclassifications of debris-covered glaciers are corrected.
Cheng Kou, Jonathan Li 0001
IGARSS4
2017 How many people died due to PM2.5 and where the mortality risks increased? A case study in Beijing
abstract
This study used MODIS 3 KM Aerosol Optical Depth (AOD) products, ground-level PM2.5 measurements in Beijing and the Public Health knowledge to estimate the number of death attributed to the long-term exposure to a harmful level of PM2.5 concentrations. The study results demonstrated that 2015 population-weighted averaged PM2.5 in Beijing was 70.46 μg/m3, 369.73% exceeding the China's yearly standard. Additionally, it was estimated that in 2015, 4172 non-accidental deaths in Beijing may attribute to the long-term exposure of excessive PM2.5 concentrations.
Jonathan Li 0001
IGARSS3
2017 Quality evaluation of point cloud model for interior structure of a common building
abstract
This paper presents a standardized quality criteria to evaluate the 3D point cloud model of the indoor building which is based on point cloud's data accuracy, the prior characteristics of the building and the coincidence errors of the point cloud model. Our assessment framework involves three steps: the point cloud data acquisition, model generation and quality evaluation. In model generation progress, incapacity of scanning the whole building information one time since its multi-storied spatial structure, the building model need to be registered and merged. In evaluation step, taking into account the need of mapping, indoor location and navigation, the establishment of an interior spatial building model requires accurate measurement. Therefore, we adopt data noise analysis to give a judgement. Then, since the geometric characteristics of the building model are varying, the geometric analysis is proposed to evaluate acquisition errors and registration error. Comparative experiments demonstrate our method give integrate, realistic and reliable quality framework for the indoor building point cloud model.
Yiping Chen 0002, Chenglu Wen, Cheng Wang 0003, Jonathan Li 0001, JinYong Chen
IGARSS5
2017 An efficiently volumetric fusing method for structure-frome-motion and terrestrial point cloud
abstract
Airborne acquisition and ground-view 3D point cloud provide complementary 3D information at city scale. A complete but lacks ground-view details, while the latter is incomplete for higher floors and severe occlusion. In this paper, First, a volumetric fusion method based on graph cuts were applied for fusing of airborne and terrestrial 3D LiDAR data. Second, we propose a method of constraints based on the local centroid of point cloud to eliminate the gap of fusion boundary. Finally, the experiments show that the improved fusion algorithm implement blending effectively.
Wei Li 0151, Cheng Wang 0003, Dawei Zai, Pengdi Huang, Weiquan Liu, Jonathan Li 0001
IGARSS6
2017 Automated extraction of urban roadside trees from mobile laser scanning point clouds based on a voxel growing method
abstract
This paper presents a new method for extracting urban roadside trees automatically from mobile laser scanning point clouds. This method mainly includes three steps. First, ground point clouds are removed by voxel-based upward growing method. Second, Euclidean distance segment method is used to cluster non-ground point clouds into certain individual objects. Then crown seeds of the initial layer is found by comparing the number of points in each layer after using the voxel modeling algorithm. Crown seeds in other layers can thereafter be detected via the upward inter-sectional analysis. Third, a crown voxel growing algorithm is used to make the crown grow in horizontal. The experimental results show that the voxel models of the individual roadside trees can be automatically and effectively extracted with our method.
Zhenlong Xiao, Yiping Chen 0002, Pengdi Huang, Rongren Wu, Jonathan Li 0001
IGARSS6
2017 Road width measurement from remote sensing images
abstract
In this paper, we propose a novel approach for road width measurement from high resolution satellite or aerial images. The proposed approach has three main steps. First, we extract line segments and road center lines on the given remote sensing images. Second, we could obtain many pairs of parallel lines with width information by computing the positional relationship between each other. Then K-means is performed to cluster these parallel lines into several clusters by the width information of them. Finally, an energy function is introduced to assign the width range of a cluster to each pixel on road center lines, the width range is viewed as the width of the corresponding road segment. Attribute to parallel lines extraction, parallel lines clustering and our energy function, the proposed road width measurement method is able to provide high quality results on road width measurement.
Zhichao Xia, Cheng Wang 0003, Jonathan Li 0001
IGARSS4
2017 Rapid traffic sign damage inspection in natural scenes using mobile laser scanning data
abstract
This paper proposes a novel approach for traffic sign detection and rapid damage inspection in natural scenes based on mobile laser scanning (MLS) data, including images and point clouds. The inspection results assist traffic management departments to take immediate measures to update and maintain traffic signs after natural disasters leading to many damaged traffic signs. Our approach involves four steps: Firstly, we use a deep learning network, Fast regions with convolutional neural network (Fast R-CNN), to train a traffic sign detector in an open benchmark, where the images are more variable and have a higher resolution. Then, traffic signs in images are detected by using the trained detector. Next, the area of the traffic sign, based on the sign area in the image, is roughly detected in MLS point clouds. Then, an accurate traffic sign is detected. Finally, some placement parameters of the traffic sign are measured for damage inspection and further inventory. Our proposed approach is validated on a set of point-clouds acquired by a RIEGL VMX-450 MLS system. Experimental results demonstrate that the rapidity and reliability of our proposed approach in traffic sign detection and damage inspection are robust.
Changbin You, Chenglu Wen, Huan Luo 0001, Cheng Wang 0003, Jonathan Li 0001
IGARSS5
2017 Deep residual networks for hyperspectral image classification
abstract
Deep neural networks can learn deep feature representation for hyperspectral image (HSI) interpretation and achieve high classification accuracy in different datasets. However, counterintuitively, the classification performance of deep learning models degrades as their depth increases. Therefore, we add identity mappings to convolutional neural networks for every two convolutional layers to build deep residual networks (ResNets). To study the influence of deep learning model size on HSI classification accuracy, this paper applied ResNets and CNNs with different depth and width using two challenging datasets. Moreover, we tested the effectiveness of batch normalization as a regularization method with different model settings. The experimental results demonstrate that ResNets mitigate the declining-accuracy effect and achieved promising classification performance with 10% and 5% training sample percentages for the University of Pavia and Indian Pines datasets, respectively. In addition, t-Distributed Stochastic Neighbor Embedding (t-SNE) provides a direct view of the extracted features through dimensionality reduction.
Zilong Zhong, Jonathan Li 0001, Lingfei Ma, Han Jiang 0005, He Zhao 0007
IGARSS2
2017 Tree Classification in Complex Forest Point Clouds Based on Deep Learning
abstract
Recently, the classification of tree species using 3-D point clouds has drawn wide attention in surveys and forestry investigations. This letter proposes a new voxel-based deep learning method to classify tree species in 3-D point clouds collected from complex forest scenes. The proposed method includes three steps: 1) individual tree extraction based on the density of the point clouds; 2) low-level feature representation through voxel-based rasterization; and 3) classification of tree species by a deep learning model. Two data sets of 3-D forest point clouds acquired by terrestrial laser scanning systems are used to evaluate the proposed method. The method achieves an average classification accuracy of 93.1% and 95.6% on the two data sets. Furthermore, in comparative experiments, the proposed method exhibits performance superior to that of the other 3-D tree species classification methods.
Xinhuai Zou, Ming Cheng 0002, Cheng Wang 0003, Yan Xia 0003, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.5
2017 Facet Segmentation-Based Line Segment Extraction for Large-Scale Point Clouds
abstract
As one of the most common features in the man-made environments, straight lines play an important role in many applications. In this paper, we present a new framework to extract line segments from large-scale point clouds. The proposed method is fast to produce results, easy for implementation and understanding, and suitable for various point cloud data. The key idea is to segment the input point cloud into a collection of facets efficiently. These facets provide sufficient information for determining linear features in the local planar region and make line segment extraction become relatively convenient. Moreover, we introduce the concept “number of false alarms” into 3-D point cloud context to filter the false positive line segment detections. We test our approach on various types of point clouds acquired from different ways. We also compared the proposed method with several other methods and provide both quantitative and visual comparison results. The experimental results show that our algorithm is efficient and effective, and produce more accurate and complete line segments than the comparative methods. To further verify the accuracy of the line segments extracted by the proposed method, we also present a line-based registration framework, which employs these line segments on point clouds registration.
Yangbin Lin, Cheng Wang 0003, Bili Chen, Dawei Zai, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.5
2017 Joint Enhancing Filtering for Road Network Extraction
abstract
In this paper, we propose a task-oriented enhancing technique for extracting road networks from satellite images. By exploiting an approximate estimation of the potential road edges for guidance, we developed a joint enhancing filtering framework to generate a version of the input image that facilitates road network extraction. First, an adaptive smoothing scheme is designed to suppress the interference of noise or heavy textures, such as residential areas or terrain boundaries. By combining this scheme with the proposed novel anisotropic shock filter, the edges of the potential road regions can be kept sharp and clear. Through abundant experimental comparisons with state-of-the-art filtering techniques and quantitative evaluations using data from various satellite sensors, the performance of the proposed approach is comprehensively evaluated. The experimental results demonstrate that our system can address heavy high contrast textures and provide a meaningful improvement in the feature detection for road extraction.
Cheng Wang 0003, Lun Luo, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.6
2017 Traffic Sign Occlusion Detection Using Mobile Laser Scanning Point Clouds
abstract
For survey and maintenance of traffic signs, this paper presents a novel traffic sign occlusion detection method using 3-D point clouds and trajectory data acquired by a mobile laser scanning system. To produce a maintenance guide, our method aims to obtain the degree of occlusion by analyzing the spatial relationship between traffic signs, surroundings, and drivers on the road. First, a detection method considering both reflectance and geometric features is developed to capture traffic signs. Next, to simulate the driver's view, a trajectory-based method is proposed to determine driver's observation location and the corresponding observed traffic sign. Finally, to determine whether a traffic sign is in occlusion, a hidden point removal algorithm is adopted and carried out. Furthermore, we develop two indices to evaluate the degree of occlusion. The proposed method is tested using two point cloud data sets collected by an RIEGL VMX-450 system along a 23.68-km-long urban road. The obtained results illustrate the feasibility of the proposed occlusion detection method.
Pengdi Huang, Ming Cheng 0002, Yiping Chen 0002, Huan Luo 0001, Cheng Wang 0003, Jonathan Li 0001
IEEE Trans. Intell. Transp. Syst.6
2017 Rapid Localization and Extraction of Street Light Poles in Mobile LiDAR Point Clouds: A Supervoxel-Based Approach
abstract
This paper presents a supervoxel-based approach for automated localization and extraction of street light poles in point clouds acquired by a mobile LiDAR system. The method consists of five steps: preprocessing, localization, segmentation, feature extraction, and classification. First, the raw point clouds are divided into segments along the trajectory, the ground points are removed, and the remaining points are segmented into supervoxels. Then, a robust localization method is proposed to accurately identify the pole-like objects. Next, a localization-guided segmentation method is proposed to obtain pole-like objects. Subsequently, the pole features are classified using the support vector machine and random forests. The proposed approach was evaluated on three datasets with 1,055 street light poles and 701 million points. Experimental results show that our localization method achieved an average recall value of 98.8%. A comparative study proved that our method is more robust and efficient than other existing methods for localization and extraction of street light poles.
Chenglu Wen, Yulan Guo, Yongtao Yu, Cheng Wang 0003, Jonathan Li 0001
IEEE Trans. Intell. Transp. Syst.7
2016 Towards Domain Adaptive Vehicle Detection in Satellite Image by Supervised Super-Resolution Transfer
abstract
Vehicle detection in satellite image has attracted extensive research attentions with various emerging applications.However, the detector performance has been significantly degenerated due to the low resolutions of satellite images, as well as the limited training data.In this paper, a robust domain-adaptive vehicle detection framework is proposed to bypass both problems.Our innovation is to transfer the detector learning to the high-resolution aerial image domain,where rich supervision exists and robust detectors can be trained.To this end, we first propose a super-resolution algorithm using coupled dictionary learning to ``augment'' the satellite image region being tested into the aerial domain.Notably, linear detection loss is embedded into the dictionary learning, which enforces the augmented region to be sensitive to the subsequent detector training.Second, to cope with the domain changes, we propose an instance-wised detection using Exemplar Support Vector Machines (E-SVMs), which well handles the intra-class and imaging variations like scales, rotations, and occlusions.With comprehensive experiments on large-scale satellite image collections, we demonstrate that the proposed framework can significantly boost the detection accuracy over several state-of-the-arts.
Liujuan Cao, Rongrong Ji, Cheng Wang 0003, Jonathan Li 0001
AAAI4
2016 A low cost indoor mapping robot based on TinySLAM algorithm
abstract
It is important for a robot or a smart device to locate itself and create a map of its indoor environment. A number of approaches and techniques exist for indoor mapping, many of which require expensive devices and highly complex computational algorithms. In our work, we introduce a low cost robot architecture based on a cheap LiDAR system and a NVIDIA Jetson tk1 platform and perform real-time indoor mapping based on the tiny SLAM algorithm.
Jonathan Li 0001, Wei Li 0151
IGARSS2
2016 Determining the effect of spatial resolution in land use classification using optical aerial imagery
abstract
Unmanned Aircraft System (UAS), a low cost and time efficient remote sensing platform, can help public agencies obtain useful urban aerial imagery of cities and update land use information frequently, especially in natural colour bands. In this paper, a decision tree based land use classification approach using optical aerial imagery is proposed. First, land cover information is extracted through the Maximum Likelihood Classifier and tabulated with an Ownership parcel map. Second, a decision tree is generated to establish the relationship between land cover and land use. Taking advantage of the geometric characteristics of parcels, an organized land use parcel map is produced. Afterwards, by resampling the aerial imagery from 20 cm, 50 cm and 100 cm resolution, effects of spatial resolution in this classification approach are discussed and determined. This land use classification method is flexible and can be widely used in urban planning and landscape monitoring.
Han Jiang 0005, Zilong Zhong, Jonathan Li 0001
IGARSS4
2016 Road network extraction via deep learning and line integral convolution
abstract
In this paper, we propose a learning-based road network extraction scheme from high resolution satellite. First, the convolutional neural network (CNN), which is able to capture large context of local structures, are applied to predict the probability of a pixel belonging to road regions, and assign labels to each pixel to describe whether it is road. Then, a line integral convolution based algorithm is developed to smooth the rough map to connect small gaps. Finally, by combining with some common image processing operators, road centerlines are able to be acquired. Attribute to the learning capacity of CNN, and the line integral convolution based connection scheme, the proposed road extraction method is able to provide high quality results comparing to current state-of-art road extraction methods.
Peikang Li, Cheng Wang 0003, Jonathan Li 0001, Ming Cheng 0002, Lun Luo
IGARSS4
2016 Superpixel-based coastline extraction in SAR images with speckle noise removal
abstract
Coastline extraction in Synthetic aperture radar (SAR) images is a fundamental and challenging task due to the speckle noise. In this paper, we propose a new method for automatic coastline extraction in SAR images. In our method, we combine K-means and speckle noise removal methods together to increase the dissimilarity between sea and land. To enhance the robustness to speckle noise, and preserve the targets boundaries, we treat superpixels as basic regions instead of pixels in traditional pixel-based methods. Finally, an adaptive threshold is applied to classify these regions into sea or land. Based on the classifications, a canny detector is employed to detect the coastline. We evaluate our proposed method on SAR images and the improved coastline extraction method superpixel-based is verified on remote sensing images with RGB channels. The experimental results demonstrate its superior performance on coastline extraction.
Xiaofang Liu, Hong Jia, Liujuan Cao, Cheng Wang 0003, Jonathan Li 0001, Ming Cheng 0002
IGARSS5
2016 Exploiting location information to detect light pole in mobile LiDAR point clouds
abstract
With rapid development of light detection and ranging (LiDAR) technologies, three dimensional point clouds increasingly become a new approach to sense the world. In our previous work, light poles were detected from mobile LiDAR point clouds without using their locations. In this paper, we improve our previous work by considering location information between two neighboring light poles to reduce false alarm. In the proposed method, the potential light poles are first detected by the extended Hough Forest Framework. Then, a gaussian distribution is exploited to model the distance between two light poles by using locations of those detected light poles. Finally, inaccurately detected light poles are removed by considering the distance between two adjacent objects. We evaluate our proposed method on mobile LiDAR point clouds acquired by RIEGL VMX-450 system. On the basis of the experimental test instances, we demonstrate improved accuracy on light pole detection.
Huan Luo 0001, Cheng Wang 0003, Hanyun Wang, Ziyi Chen 0001, Dawei Zai, Shanxin Zhang, Jonathan Li 0001
IGARSS7
2016 Automated segmentation of LiDAR point clouds for building rooftop extraction
abstract
LIDAR (Light Detection and Ranging) is an especially effective tool for acquiring geo-referenced point clouds of urban site. Accurate extraction of elevated features such as building rooftops is vitally important in various applications. However, it is still challenging to determine an accurate rooftop contour from the irregularly distributed LIDAR point clouds. In this paper an efficient LIDAR segmentation method is presented in order to achieve automated rooftop extraction. First, we apply a voxel-based upward growing algorithm that filters out the ground points from the raw point cloud scenes. Second, we employ a Euclidean based clustering method on non-ground points by making use of nearest neighbors. Then we introduce RANSAC (RANdom SAmple Consensus) technique to estimate primitive planes for fitting rooftop facets. Finally, we use concave hull and L0regularization to determine the rooftop contour. Accurate experimental results demonstrate the validity of our segmentation method for rooftop extraction.
Cheng Wang 0003, Jonathan Li 0001, Zongliang Zhang, Dawei Zai, Pengdi Huang, Chenglu Wen
IGARSS3
2016 3D road surface extraction from mobile laser scanning point clouds
abstract
This paper presents a new algorithm to directly extract 3D road boundaries from mobile laser scanning (MLS) point clouds. The algorithm includes two stages: 1) non-ground point removal by a voxel-based elevation filter, and 2) 3D road surface extraction by curb-line detection based on energy minimization and graph cuts. The proposed algorithm was tested on a dataset acquired by a RIEGL VMX-450 MLS system. The results fully demonstrate the effectiveness and superiority of the proposed algorithm.
Dawei Zai, Yulan Guo, Jonathan Li 0001, Huan Luo 0001, Yangbin Lin, Pengdi Huang, Cheng Wang 0003
IGARSS3
2016 Examining urban expansion using multi-temporal landsat imagery: A case study of montreal census metropolitan area
abstract
Greater Montreal is the most populous metropolitan area in Quebec, and the second most populous in Canada after Greater Toronto. In the 1970s, the economic center of Canada shifted from Montreal to Toronto. Since some previous studies have focused on the urbanization process in the Greater Toronto Area, it is important to conduct research on its counterpart. This study uses Landsat images as the data source, combined with census data to detect urban changes in the Montreal census metropolitan area (CMA) from 1975 to 2015. We analyzed spatial patterns and annual urban growth rate by applying four supervised classification algorithms. Also, we mapped temporal land cover categories and evaluated major driving forces that contribute to the urban changes. Our results show that Montreal CMA has experienced a rapid development over the past 40 years, with 442 km2urban growth. Urban expansion in Montreal CMA mainly has two modes: radiated and ribbon.
He Zhao 0007, Lingfei Ma, Jonathan Li 0001
IGARSS4
2016 Fully convolutional networks for building and road extraction: Preliminary results
abstract
Available big geoscientific data and modern powerful computation hardware have laid a solid foundation for the prevailing deep learning models in the field of image classification, detection and segmentation. In these models, fully convolutional networks achieve unprecedented success in image segmentation tasks [6]. In this paper, we apply the contemporary image segmentation models in the context of extracting buildings and roads from high spatial resolution imagery. We estimate the influence of filter stride, learning rate, input data size, training epoch and fine-tuning on model performance. Selected Massachusetts road and building datasets are used for training, validation, and testing the performance of the models with different parameters. As a result of combining shallow fine-grained pooling layer outputs with the deep final-score layer or abandoning coarse-grained pooling layers, the extraction precision rate of the best modified model improves significantly to over 78%.
Zilong Zhong, Jonathan Li 0001, Weihong Cui, Han Jiang 0005
IGARSS2
2016 Local quality assessment of point clouds for indoor mobile mapping
Fangfang Huang, Chenglu Wen, Huan Luo 0001, Ming Cheng 0002, Cheng Wang 0003, Jonathan Li 0001
Neurocomputing6
2016 Vehicle detection from highway satellite images via transfer learning
Liujuan Cao, Cheng Wang 0003, Jonathan Li 0001
Inf. Sci.3
2016 Pole-Like Road Object Detection in Mobile LiDAR Data via Supervoxel and Bag-of-Contextual-Visual-Words Representation
abstract
This letter addresses the problem of detecting pole-like road objects (including light poles and traffic signposts) from mobile light detection and ranging (LiDAR) data for transportation-related applications. The method consists of two consecutive stages: training and pole-like object detection. At the training stage, a contextual visual vocabulary is created from the feature regions generated from a training data set by supervoxel segmentation. At the pole-like object detection stage, a bag-of-contextual-visual-words representation is generated for each semantic object segmented from mobile LiDAR data. The experimental results show that the proposed method achieves correctness, omission, and commission of 88.9%, 11.1%, and 2.8%, respectively, in detecting pole-like road objects. Computational complexity analysis demonstrates that our method provides a promising and effective solution to rapid and accurate detection of pole-like objects from large volumes of mobile LiDAR data.
Haiyan Guan, Yongtao Yu, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.3
2016 An Indoor Backpack System for 2-D and 3-D Mapping of Building Interiors
abstract
This letter presents a backpack mapping system for creating indoor 2-D and 3-D maps of building interiors. For many applications, indoor mobile mapping provides a 3-D structure via an indoor map. Because there are significant roll and pitch motions of the indoor mobile mapping system, the need arises for a moving mobile system with 6 degrees of freedom (DOFs) (x, y, and z positions and roll, yaw, and pitch angles). First, we present a 6-DOF pose estimation algorithm by fusing 2-D laser scanner data with inertial sensor data using an extended Kalman filter-based method. The estimated 6-DOF pose is used as the initialized transformation for consecutive map alignment in 3-D map building. The 6-DOF pose gives a full 3-D estimation of the system pose and is used to accelerate the map alignment process and also align the two maps directly when there are few or no overlapping areas between the maps. Our results show that the proposed system effectively builds a consistent 2-D grid map and a 3-D point cloud map of an indoor environment.
Chenglu Wen, Cheng Wang 0003, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.4
2016 Vehicle Detection in High-Resolution Aerial Images via Sparse Representation and Superpixels
abstract
This paper presents a study of vehicle detection from high-resolution aerial images. In this paper, a superpixel segmentation method designed for aerial images is proposed to control the segmentation with a low breakage rate. To make the training and detection more efficient, we extract meaningful patches based on the centers of the segmented superpixels. After the segmentation, through a training sample selection iteration strategy that is based on the sparse representation, we obtain a complete and small training subset from the original entire training set. With the selected training subset, we obtain a dictionary with high discrimination ability for vehicle detection. During training and detection, the grids of histogram of oriented gradient descriptor are used for feature extraction. To further improve the training and detection efficiency, a method is proposed for the defined main direction estimation of each patch. By rotating each patch to its main direction, we give the patches consistent directions. Comprehensive analyses and comparisons on two data sets illustrate the satisfactory performance of the proposed algorithm.
Ziyi Chen 0001, Cheng Wang 0003, Chenglu Wen, Xiuhua Teng, Yiping Chen 0002, Haiyan Guan, Huan Luo 0001, Liujuan Cao, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.9
2016 Automated Detection of Three-Dimensional Cars in Mobile Laser Scanning Point Clouds Using DBM-Hough-Forests
abstract
This paper presents an automated algorithm for rapidly and effectively detecting cars directly from large-volume 3-D point clouds. Rather than using low-order descriptors, a multilayer feature generation model is created to obtain high-order feature representations for 3-D local patches through deep learning techniques. To handle cars with different levels of incompleteness caused by data acquisition ways and occlusions, a hierarchical visibility estimation model is developed to augment Hough voting. Considering scale and orientation variations in the azimuth direction, a set of multiscale Hough forests is constructed to rotationally cast votes to estimate cars' centroids. Quantitative assessments show that the proposed algorithm achieves average completeness, correctness, quality, and F1-measure of 0.94, 0.96, 0.90, and 0.95, respectively, in detecting 3-D cars. Comparative studies also demonstrate that the proposed algorithm outperforms the other four existing algorithms in accurately and completely detecting 3-D cars from large-scale 3-D point clouds.
Yongtao Yu, Jonathan Li 0001, Haiyan Guan, Cheng Wang 0003
IEEE Trans. Geosci. Remote. Sens.2
2016 Road Network Extraction via Aperiodic Directional Structure Measurement
abstract
In this paper, we present a novel aperiodic directional structure measurement (ADSM) toward road network extraction. Based on the observations from Cognitive Psychology regarding the aperiodicity and local directionality, ADSM can well characterize roadlike structures independent of the spectral character and contrast. By exploiting such measurement as guidance, we construct a mask to denote potential road regions. Then, by combining with some common morphology operators, our approach is able to provide robust road centerlines efficiently. We evaluate our approach with data from various satellite sensors and make comprehensive comparisons with previous state-of-the-art methods. Experimental results demonstrate the merit using our ADSM as a metric to identify potential road structures, as well as the effectiveness and efficiency of our road network extraction system.
Cheng Wang 0003, Liujuan Cao, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.5
2016 Vehicle Detection in High-Resolution Aerial Images Based on Fast Sparse Representation Classification and Multiorder Feature
abstract
This paper presents an algorithm for vehicle detection in high-resolution aerial images through a fast sparse representation classification method and a multiorder feature descriptor that contains information of texture, color, and high-order context. To speed up computation of sparse representation, a set of small dictionaries, instead of a large dictionary containing all training items, is used for classification. To extract the context information of a patch, we proposed a high-order context information extraction method based on the proposed fast sparse representation classification method. To effectively extract the color information, the RGB color space is transformed into color name space. Then, the color name information is embedded into the grids of histogram of oriented gradient feature to represent the low-order feature of vehicles. By combining low- and high-order features together, a multiorder feature is used to describe vehicles. We also proposed a sample selection strategy based on our fast sparse representation classification method to construct a complete training subset. Finally, a set of dictionaries, which are trained by the multiorder features of the selected training subset, is used to detect vehicles based on superpixel segmentation results of aerial images. Experimental results illustrate the satisfactory performance of our algorithm.
Ziyi Chen 0001, Cheng Wang 0003, Huan Luo 0001, Hanyun Wang, Yiping Chen 0002, Chenglu Wen, Yongtao Yu, Liujuan Cao, Jonathan Li 0001
IEEE Trans. Intell. Transp. Syst.9
2016 Patch-Based Semantic Labeling of Road Scene Using Colorized Mobile LiDAR Point Clouds
abstract
Semantic labeling of road scenes using colorized mobile LiDAR point clouds is of great significance in a variety of applications, particularly intelligent transportation systems. However, many challenges, such as incompleteness of objects caused by occlusion, overlapping between neighboring objects, interclass local similarities, and computational burden brought by a huge number of points, make it an ongoing open research area. In this paper, we propose a novel patch-based framework for labeling road scenes of colorized mobile LiDAR point clouds. In the proposed framework, first, three-dimensional (3-D) patches extracted from point clouds are used to construct a 3-D patch-based match graph structure (3D-PMG), which transfers category labels from labeled to unlabeled point cloud road scenes efficiently. Then, to rectify the transferring errors caused by local patch similarities in different categories, contextual information among 3-D patches is exploited by combining 3D-PMG with Markov random fields. In the experiments, the proposed framework is validated on colorized mobile LiDAR point clouds acquired by the RIEGL VMX-450 mobile LiDAR system. Comparative experiments show the superior performance of the proposed framework for accurate semantic labeling of road scenes.
Huan Luo 0001, Cheng Wang 0003, Chenglu Wen, Zhipeng Cai 0003, Ziyi Chen 0001, Hanyun Wang, Yongtao Yu, Jonathan Li 0001
IEEE Trans. Intell. Transp. Syst.8
2016 Spatial-Related Traffic Sign Inspection for Inventory Purposes Using Mobile Laser Scanning Data
abstract
This paper presents a spatial-related traffic sign inspection process for sign type, position, and placement using mobile laser scanning (MLS) data acquired by a RIEGL VMX-450 system and presents its potential for traffic sign inventory applications. First, the paper describes an algorithm for traffic sign detection in complicated road scenes based on the retroreflectivity properties of traffic signs in MLS point clouds. Then, a point cloud-to-image registration process is proposed to project the traffic sign point clouds onto a 2-D image plane. Third, based on the extracted traffic sign points, we propose a traffic sign position and placement inspection process by creating geospatial relations between the traffic signs and road environment. For further inventory applications, we acquire several spatial-related inventory measurements. Finally, a traffic sign recognition process is conducted to assign sign type. With the acquired sign type, position, and placement data, a spatial-associated sign network is built. Experimental results indicate satisfactory performance of the proposed detection, recognition, position, and placement inspection algorithms. The experimental results also prove the potential of MLS data for automatic traffic sign inventory applications.
Chenglu Wen, Jonathan Li 0001, Huan Luo 0001, Yongtao Yu, Zhipeng Cai 0003, Hanyun Wang, Cheng Wang 0003
IEEE Trans. Intell. Transp. Syst.2
2016 Bag of Contextual-Visual Words for Road Scene Object Detection From Mobile Laser Scanning Data
abstract
This paper proposes a novel algorithm for detecting road scene objects (e.g., light poles, traffic signposts, and cars) from 3-D mobile-laser-scanning point cloud data for transportation-related applications. To describe local abstract features of point cloud objects, a contextual visual vocabulary is generated by integrating spatial contextual information of feature regions. Objects of interest are detected based on the similarity measures of the bag of contextual-visual words between the query object and the segmented semantic objects. Quantitative evaluations on two selected data sets show that the proposed algorithm achieves an average recall, precision, quality, and F-score of 0.949, 0.970, 0.922, and 0.959, respectively, in detecting light poles, traffic signposts, and cars. Comparative studies demonstrate the superior performance of the proposed algorithm over other existing methods.
Yongtao Yu, Jonathan Li 0001, Haiyan Guan, Cheng Wang 0003, Chenglu Wen
IEEE Trans. Intell. Transp. Syst.2
2015 Big Data Analytics and Visualization with Spatio-Temporal Correlations for Traffic Accidents
Xiaoliang Fan, Baoqin He, Cheng Wang 0003, Jonathan Li 0001, Ming Cheng 0002, Huaqiang Huang, Xiao Liu 0004
ICA3PP (2)4
2015 A tensor voting approach to dark spot detection in RADARSAT-1 intensity imagery
abstract
This paper presents a tensor voting approach to automated detection of dark spots in RADARSAT-1 ScanSAR Narrow Beam mode images. First, a thresholding algorithm that well maximizes the ratio of between-class variance to within-class variance is used to detect potential dark spot candidates. Next, a tensor voting framework integrated with sparse and dense ball votings is carried out to suppress noise while maintaining dark spots. Then, a saliency map that reflects the probability of a pixel being located within a dark spot is generated using the saliencies of ball tensors. Finally, a segmentation method is applied to ascertain dark spots based on the saliency map. The proposed approach has been tested on a set of RADARSAT-1 ScanSAR Narrow Beam intensity images. Quantitative evaluations demonstrate that the proposed approach achieves an average commission error, omission error, and quality of 0.003, 0.037, and 0.956, respectively, for detecting dark spots in SAR intensity imagery.
Haiyan Guan, Yongtao Yu, Jonathan Li 0001
IGARSS3
2015 Extraction of street trees from mobile laser scanning point clouds based on subdivided dimensional features
abstract
This paper proposes a method for automated extraction of street trees in a typical urban environment from 3D point cloud data acquired by the mobile laser scanning system. First, the algorithm utilizes the voxel-based method to remove the ground points from the scene. Second, the Euclidean distance clustering is adopted to cluster points into individual objects. The eigenvalues of neighborhood covariance matrix and the corresponding normalized centroid distance are computed for each point to obtain the subdivided dimensional features. Finally, the statistical component features and horizontal information are calculated for object detection. The experiment results show the feasibility of the proposed algorithm.
Pengdi Huang, Yiping Chen 0002, Jonathan Li 0001, Yongtao Yu, Cheng Wang 0003, Hongshan Nie
IGARSS3
2015 Improving urban impervious surface classification by combining Landsat and PolSAR images: A case study in Kitchener-Waterloo, Ontario, Canada
abstract
Urban impervious surface mapping using moderate-resolution optical images such as Landsat images could be challenging due to the complexity of urban land cover. The study aims to combine optical and PolSAR images to improve accuracy of impervious surface classification. A scene of Landsat-5 TM image and a scene of RADARSAT-2 full-polarized imagery of Kitchener-Waterloo were used. The classification accuracies of Landsat image with the combination of different polarizations were compared. The results demonstrated the improvement of impervious surface classification with the combination of RADARSAT-2 PolSAR imagery with Landsat imagery. The major improvement was distinguishing between dark and bright impervious surface. In addition, generally more polarizations generated better results, and HV had the most contributions compared to the rest three polarizations. The results of the study may serve as a reference for further application for combining PolSAR and optical images.
Weikai Tan, Renfang Liao, Yikang Du, Jun Lu 0008, Jonathan Li 0001
IGARSS5
2015 Evaluation of regional-scale snow albedo characteristics during winter season from 2003 to 2014
abstract
Snow is a very important component of the climate system. It can influence the energy budget of the atmosphere and hydrological system significantly. The main goal of this paper is to use remote sensing and geographical information system techniques to analyze the spatial and temporal variations in regional scale and to find the relations between meteorological parameters and the snow albedo for the future modeling in snow albedo study. The results revealed spatial and temporal variation throughout different months during the winter season. In addition, the Pearson correlation coefficient analysis showed partial correlation between snow albedo and meteorological variables, which can be used to model snow albedo in some hydrological studies.
Jonathan Li 0001, Claude R. Duguay, Dilong Li
IGARSS2
2015 Using mobile LiDAR point clouds for traffic sign detection and sign visibility estimation
abstract
This paper presents a novel method for traffic sign detection and visibility evaluation from mobile Light Detection and Ranging (LiDAR) point clouds and the corresponding images. Our algorithm involves two steps. Firstly, a detection algorithm based on high retro-reflectivity of the traffic sign from the MLS point clouds is designed for sign detection in complicated road scenes. To solve the spatial features of traffic signs, we also create geo-referenced relations between traffic signs and roads according to the normal of ground. Secondly, we propose a visibility estimation method to evaluate the visibility level of the traffic sign based on a combination of visual appearance and spatial-related features. The proposed algorithm is validated on a set of transportation-related point-clouds acquired by a RIEGL VMX-450 LiDAR system. The experiment results demonstrate that the efficiency and reliability of the proposed algorithm in detection traffic signs are robust, and also prove the potential of using mobile LiDAR data for traffic sign visibility evaluation.
Chenglu Wen, Huan Luo 0001, Yiping Chen 0002, Cheng Wang 0003, Jonathan Li 0001
IGARSS6
2015 Inventory of 3D street lighting poles using mobile laser scanning point clouds
abstract
This paper presents a novel approach for extracting street lighting poles directly from MLS point clouds. The approach includes four stages: 1) elevation filtering to remove ground points, 2) Euclidean distance clustering to cluster points, 3) voxel-based normalized cut (Ncut) segmentation to separate overlapping objects, and 4) statistical analysis of geometric properties to extract 3D street lighting poles. A Dataset acquired by a RIEGL VMX-450 MLS system are tested with the proposed approach. The results demonstrate the efficiency and reliability of the proposed approach to extract 3D street lighting poles.
Dawei Zai, Yiping Chen 0002, Jonathan Li 0001, Yongtao Yu, Cheng Wang 0003, Hongshan Nie
IGARSS3
2015 Turning mobile laser scanning points into 2D/3D on-road object models: Current status
abstract
Traditional road surveying methods rely largely on in-situ measurements, which are time consuming and labor intensive. Recent Mobile Laser Scanning (MLS) techniques enable collection of road data at a normal driving speed. However, extracting required information from collected MLS data remains a challenging task. This paper focuses on examining the current status of automated on-road object extraction techniques from 3D MLS points over the last five years. Several kinds of on-road objects are included in this paper: curbs and road surfaces, road markings, pavement cracks, as well as manhole and sewer well covers. We evaluate the extraction techniques according to their method design, degree of automation, precision, and computational efficiency. Given the large volume of MLS data, to date most MLS object extraction techniques are aiming to improve their precision and efficiency.
Zongliang Zhang, Ming Cheng 0002, Xinqu Chen, Menglan Zhou, Jonathan Li 0001, Hongshan Nie
IGARSS6
2015 Hidden target detection from the multi-echo small-footprint LiDAR point clouds
abstract
We propose a new approach for hidden or potential object detection behind the vegetation based on the multi-echo small-footprint of light detection and ranging(LiDAR) point cloud. According to the specific characteristic that laser beam is able to penetrate foliage gaps, which offers the opportunity to perceive and detect the object invisible to the naked eye. First, the waveform sample data of the small footprint multi-echo liDAR uses a Gaussian fitting tool to curve fitting. Second, the peak detection of the waveform can be classified in statistical results of its wave numbers. Decomposition and correction are the processing for classifying based on the statistic data. After selecting the point clouds which are contained in the multi-peaks echoes, we obtain the tree and the embedded target behind it as well as ground elimination. The range of the waveform component is used to separate the penetrations material and the target by distance discriminant function. Experiments are implemented on the waveforms acquired by small-footprint LiDAR system VZ-1000 Sensor. The results indicate that the algorithm could provide an optimal solution for LiDAR waveform hidden target detection.
Yiping Chen 0002, Cheng Wang 0003, Jonathan Li 0001
MMSP3
2015 Single-image super-resolution in RGB space via group sparse representation
abstract
Super‐resolution (SR) is the problem of generating a high‐resolution (HR) image from one or more low‐resolution (LR) images. This study presents a new approach to single‐image super‐resolution based on group sparse representation. Two dictionaries are constructed corresponding to the LR and HR image patches, respectively. The sparse coefficients of an input LR image patch in terms of the LR dictionary are used to recover the HR patch from the HR dictionary. When constructing the dictionaries, the three colour channels in a training image patch are considered a group composed of three atoms. The whole group is selected simultaneously when representing an image patch so that the correlations between the colour channels can be retained. A dictionary training method is also designed in which the two dictionaries are trained jointly to ensure that the corresponding LR and HR patches have the same sparse coefficients. Experimental results demonstrate the effectiveness of the proposed method and its robustness to noise.
Ming Cheng 0002, Cheng Wang 0003, Jonathan Li 0001
IET Image Process.3
2015 Combinative hypergraph learning for semi-supervised image classification
Binghui Wei, Ming Cheng 0002, Cheng Wang 0003, Jonathan Li 0001
Neurocomputing4
2015 Occluded Boundary Detection for Small-Footprint Groundborne LIDAR Point Cloud Guided by Last Echo
abstract
Occluded boundary detection in a 3-D point cloud is an indispensable preprocessing step for many applications, such as point cloud completion. Meanwhile, existing methods do not have the ability of distinguishing occluded and complete surface borders, such as the border of a sign board. To solve this problem, this letter presents an occluded boundary detection method for small-footprint LIDAR point clouds. The main novelty of this letter is using the last-echo information for occluded boundary detection. Seed boundary (SB) points are subsequently detected using this last-echo information. Finally, the SB points are grown into neighboring points using an occluded boundary growth algorithm. To the best of our knowledge, this method is the first method that uses the last-echo information to detect occluded boundaries. Experimental results with comparisons indicate that the proposed method can accurately and efficiently detect an occluded boundary without contamination from a complete surface border. These advantages allow the proposed method to benefit further applications such as point cloud completion, as demonstrated in the application section.
Zhipeng Cai 0003, Cheng Wang 0003, Chenglu Wen, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.4
2015 Road Boundaries Detection Based on Local Normal Saliency From Mobile Laser Scanning Data
abstract
The accurate extraction of roads is a prerequisite for the automatic extraction of other road features. This letter describes a method for detecting road boundaries from mobile laser scanning (MLS) point clouds in an urban environment. The key idea of our method is directly constructing a saliency map on 3-D unorganized point clouds to extract road boundaries. The method consists of four major steps, i.e., road partition with the assistance of the vehicle trajectory, salient map construction and salient points extraction, curb detection and curb lowest points extraction, and road boundaries fitting. The performance of the proposed method is evaluated on the point clouds of an urban scene collected by a RIEGL VMX-450 MLS system. The completeness, correctness, and quality of the extracted road boundaries are 95.41%, 99.35%, and 94.81%, respectively. Experimental results demonstrate that our method is feasible for detecting road boundaries in MLS point clouds.
Hanyun Wang, Huan Luo 0001, Chenglu Wen, Jun Cheng 0002, Peng Li 0064, Yiping Chen 0002, Cheng Wang 0003, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.8
2015 Three-Dimensional Object Matching in Mobile Laser Scanning Point Clouds
abstract
This letter presents a 3-D object matching framework to support information extraction directly from 3-D point clouds. The problem of 3-D object matching is to match a template, represented by a group of 3-D points, to a point cloud scene containing an instance of that object. A locally affine-invariant geometric constraint is proposed to effectively handle affine transformations, occlusions, incompleteness, and scales in 3-D point clouds. The 3-D object matching framework is integrated into 3-D correspondence computation, 3-D object detection, and point cloud object classification in mobile laser scanning (MLS) point clouds. Experimental results obtained using the 3-D point clouds acquired by a RIEGL VMX-450 system showed that completeness, correctness, and quality of over 0.96, 0.94, and 0.91 are achieved, respectively, with the proposed framework in 3-D object detection. Comparative studies demonstrate that the proposed method outperforms the two existing methods for detecting 3-D objects directly from large-volume MLS point clouds.
Yongtao Yu, Jonathan Li 0001, Haiyan Guan, Fukai Jia, Cheng Wang 0003
IEEE Geosci. Remote. Sens. Lett.2
2015 Robust depth-based object tracking from a moving binocular camera
Liujuan Cao, Cheng Wang 0003, Jonathan Li 0001
Signal Process.3
2015 Iterative Tensor Voting for Pavement Crack Extraction Using Mobile Laser Scanning Data
abstract
The assessment of pavement cracks is one of the essential tasks for road maintenance. This paper presents a novel framework, called ITVCrack, for automated crack extraction based on iterative tensor voting (ITV), from high-density point clouds collected by a mobile laser scanning system. The proposed ITVCrack comprises the following: 1) the preprocessing involving the separation of road points from nonroad points using vehicle trajectory data; 2) the generation of the georeferenced feature (GRF) image from the road points; and 3) the ITV-based crack extraction from the noisy GRF image, followed by an accurate delineation of the curvilinear cracks. Qualitatively, the method is applicable for pavement cracks with low contrast, low signal-to-noise ratio, and bad continuity. Besides the application to GRF images, the proposed framework demonstrates much better crack extraction performance when quantitatively compared to existing methods on synthetic data and pavement images.
Haiyan Guan, Jonathan Li 0001, Yongtao Yu, Michael A. Chapman, Hanyun Wang, Cheng Wang 0003, Ruifang Zhai
IEEE Trans. Geosci. Remote. Sens.2
2015 Semiautomated Extraction of Street Light Poles From Mobile LiDAR Point-Clouds
abstract
This paper proposes a novel algorithm for extracting street light poles from vehicleborne mobile light detection and ranging (LiDAR) point-clouds. First, the algorithm rapidly detects curb-lines and segments a point-cloud into road and nonroad surface points based on trajectory data recorded by the integrated position and orientation system onboard the vehicle. Second, the algorithm accurately extracts street light poles from the segmented nonroad surface points using a novel pairwise 3-D shape context. The proposed algorithm is tested on a set of point-clouds acquired by a RIEGL VMX-450 mobile LiDAR system. The results show that road surfaces are correctly segmented, and street light poles are robustly extracted with a completeness exceeding 99%, a correctness exceeding 97%, and a quality exceeding 96%, thereby demonstrating the efficiency and feasibility of the proposed algorithm to segment road surfaces and extract street light poles from huge volumes of mobile LiDAR point-clouds.
Yongtao Yu, Jonathan Li 0001, Haiyan Guan, Cheng Wang 0003, Jun Yu 0002
IEEE Trans. Geosci. Remote. Sens.2
2015 Automated Road Information Extraction From Mobile Laser Scanning Data
abstract
This paper presents a survey of literature about road feature extraction, giving a detailed description of a Mobile Laser Scanning (MLS) system (RIEGL VMX-450) for transportation-related applications. This paper describes the development of automated algorithms for extracting road features (road surfaces, road markings, and pavement cracks) from MLS point cloud data. The proposed road surface extraction algorithm detects road curbs from a set of profiles that are sliced along vehicle trajectory data. Based on segmented road surface points, we create Geo-Referenced Feature (GRF) images and develop two algorithms, respectively, for extracting the following: 1) road markings with high retroreflectivity and 2) cracks containing low contrast with their surroundings, low signal-to-noise ratio, and poor continuity. A comprehensive comparison illustrates satisfactory performance of the proposed algorithms and concludes that MLS is a reliable and cost-effective alternative for rapid road inspection.
Haiyan Guan, Jonathan Li 0001, Yongtao Yu, Michael A. Chapman, Cheng Wang 0003
IEEE Trans. Intell. Transp. Syst.2
2015 Using Mobile LiDAR Data for Rapidly Updating Road Markings
abstract
Updating road markings is one of the routine tasks of transportation agencies. Compared with traditional road inventory mapping techniques, vehicle-borne mobile light detection and ranging (LiDAR) systems can undertake the job safely and efficiently. However, current hurdles include software and computing challenges when handling huge volumes of highly dense and irregularly distributed 3-D mobile LiDAR point clouds. This paper presents the development and implementation aspects of an automated object extraction strategy for rapid and accurate road marking inventory. The proposed road marking extraction method is based on 2-D georeferenced feature (GRF) images, which are interpolated from 3-D road surface points through a modified inverse distance weighted (IDW) interpolation. Weighted neighboring difference histogram (WNDH)-based dynamic thresholding and multiscale tensor voting (MSTV) are proposed to segment and extract road markings from the noisy corrupted GRF images. The results obtained using 3-D point clouds acquired by a RIEGL VMX-450 mobile LiDAR system in a subtropical urban environment are encouraging.
Haiyan Guan, Jonathan Li 0001, Yongtao Yu, Zheng Ji, Cheng Wang 0003
IEEE Trans. Intell. Transp. Syst.2
2015 Automated Extraction of Urban Road Facilities Using Mobile Laser Scanning Data
abstract
This paper proposes a novel, automated algorithm for rapidly extracting urban road facilities, including street light poles, traffic signposts, and bus stations, for transportation-related applications. A detailed description and implementation of the proposed algorithm is provided using mobile laser scanning data collected by a state-of-the-art RIEGL VMX-450 system. First, to reduce the quantity of data to be handled, a fast voxel-based upward growing method is developed to remove ground points. Then, off-ground points are clustered and segmented into individual objects via Euclidean distance clustering and voxel-based normalized cut segmentation, respectively. Finally, a 3-D object matching framework, benefiting from a locally affine-invariant geometric constraint, is developed to achieve the extraction of 3-D objects. Quantitative evaluations show that the proposed algorithm attains an average completeness, correctness, quality, and F1-measure of 0.949, 0.971, 0.922, and 0.960, respectively, in extracting 3-D light poles, traffic signposts, and bus stations. Comparative studies demonstrate the efficiency and feasibility of the proposed algorithm for automated and rapid extraction of urban road facilities.
Yongtao Yu, Jonathan Li 0001, Haiyan Guan, Cheng Wang 0003
IEEE Trans. Intell. Transp. Syst.2
2014 Oil spill detection based on a superpixel segmentation method for SAR image
abstract
In this paper, a rapid oil spill detection approach which still maintains high detection accuracy is presented. The major contribution of the approach is using a superpixel segmentation method to subdivide the target SAR image into many approximate uniform scale pieces and preserves the boundaries well. Furthermore, a novel approach combine space distance, intensity deviation and size information together (SIS) is presented to eliminate the potential false positive, which is convenient and effective meanwhile. The proposed approach performs well and fast in both the synthetic data and RAD ARS AT-1 ScanSAR data which contain verified oil spills. The processing time is about 6s for a 512×512 image.
Ziyi Chen 0001, Cheng Wang 0003, Xiuhua Teng, Liujuan Cao, Jonathan Li 0001
IGARSS5
2014 Automatic extraction of power lines from mobile laser scanning data
abstract
This paper presents a stepwise algorithm for extracting power-lines from mobile laser scanning (MLS) data. This algorithm first extracts non-road points from MLS data by estimating road ranges with regard to scanning mechanism and applying elevation-difference and slope criteria to the road ranges scan-line by scan-line. Then, three filters, in terms of height, spatial density, and size-and-shape, are proposed to extract power-line points in the identified non-road points, followed by Hough transform and Euclidean distance clustering. Finally, a 3D power line is modelled as a horizontal line in X-Y plane and a vertical catenary curve defined by a hyperbolic cosine function in X-Z plane. The proposed algorithm has been tested on a sample of point clouds acquired by a RIEGL VMX-450 MLS system. The results demonstrate the applicability of the proposed algorithm in extracting power transmission lines.
Haiyan Guan, Jonathan Li 0001, Yongjun Zhou, Yongtao Yu, Cheng Wang 0003, Chenglu Wen
IGARSS2
2014 Earthwork volumes estimation in asphalt pavement reconstruction using a mobile laser scanning systerm
abstract
This paper presents a novel method for estimating earthwork volumes in asphalt pavement reconstruction using a mobile laser scanning (MLS) system. First, based on the static targets, this method registers two point cloud datasets into the same coordinate system, which respectively are acquired in the reconstructing road before and after asphalting. Next, road surface points are detected from each point cloud using a curb-based method, and further divided into a set of blocks. Afterwards, the blocks are perpendicularly partitioned into grids, where two surface features are extracted using the RANSAC. Finally, the volume of each grid is calculated according to these two surface features. The proposed algorithm has been tested on two sets of point clouds acquired by a RIEGL VMX-450 MLS system in the reconstructing road before and after asphalting. The results demonstrate the accuracy and efficiency of the proposed algorithm in estimating earthwork volumes.
Fukai Jia, Jonathan Li 0001, Cheng Wang 0003, Yongtao Yu, Ming Cheng 0002, Dawei Zai
IGARSS2
2014 Object based building extraction by QuickBird image for population estimation: A case study of the City of Waterloo
abstract
This paper used QuickBird high resolution image to estimate the population of the city of Waterloo, ON, Canada. Two approaches of object based classification were compared to extract buildings from the original image. One is rule based classification and the other is example based classification. We chose two districts which are Lakeshore and Columbia as our testing areas. Rule based result is better than example based. The overall accuracy of rule based classification in Lakeshore District and Columbia District are 92.5% and 85.5%. With census data, the average area per person is about 38.8 m2and the estimated population of the city of Waterloo is about 109589.
Wei Li 0151, Shiqian Wang, Jonathan Li 0001
IGARSS3
2014 Automatic markerless registration of mobile LiDAR point-clouds
abstract
Point-cloud registration plays a significant role in the area of mobile LiDAR data processing. This paper proposes an automatic markerless registration algorithm for lidar point-clouds. It first introduces a local feature for point-cloud representation. The feature is invariant to rotations and translations of a point-cloud. It then presents a point-cloud registration method using geometric consistency check and the Iterative Closest Points (ICP) algorithm. Comparative experiments were performed on a publicly available dataset. Experimental results show that our algorithm is very accurate and outperforms the spin image and SHOT based algorithms.
Min Lu 0001, Yulan Guo, Jun Zhang 0044, Jianwei Wan, Jonathan Li 0001
IGARSS5
2014 Detecting floods accurately in SAR images without the interference of the reference image
abstract
Most of methods for change detection of floods in SAR images are based on the difference image (DI), which is from comparing the flood image and the reference image straightforwardly. DI, however, is often distorted by the interferences of the complicated reference image. Meanwhile, the spatial contextual information of floods is difficult to be effectively used. In response to these problems, a novel change detection method by removing the interference of the reference image is proposed. Firstly, the difference image is segmented to derive the initial change mask, i.e. a part of flood areas. Secondly, the whole flood areas are gained by region growing directly in the flood image. Experimental results on the real ERS SAR dataset validate its effectiveness on both the quantitative and subjective aspects.
Jun Lu 0008, Jonathan Li 0001, Min Lu 0001, Zhiyong Li 0008, Boli Xiong, Gangyao Kuang
IGARSS2
2014 Using high resolution remote sensing image to help population estimation in small cities
abstract
This paper introduces how high resolution remotely sensed images and GIS data can help population estimation in small cities through a case study in the city of Waterloo. The foundation of this study lies on the relationship between living space and population. This relationship has been valid in recent studies and using Remote Sensing (RS) technologies for population estimation has gained much attention in RS studies [1]. However, the methods vary a lot due to different landscape and dwelling types. In this paper, a typical small city is chosen as a case study to exam the reliability of the proposed population estimation methods. First, two different classification methods are compared to generate residential area and dwelling count: rule-based feature extraction and sample-based feature extraction. Also, Geographic Information System (GIS) techniques are used to enhance the classification accuracy. Then, two population estimation models are compared: dwelling count model and residential area model. At last, results and discussion are included.
Shiqian Wang, Wei Li 0151, Jonathan Li 0001
IGARSS3
2014 Automated mosaicking of UAV images based on SFM method
abstract
Optical sensors onboard an unmanned aerial vehicle (UAV) can collect high resolution images with small dimensions. Image mosaicking is necessary to cover a larger geographic area. This paper presents a novel approach to mosaicking UAV images automatically. The "Orthophoto Map" is based on Structure From Motion (SFM). This method can fully automatic mosaic generate a wide range, and with well visual effects, and no evident deformation. This method can not only get a panoramic image of wide range of areas, and can get the corresponding three-dimensional terrain model.
Jonathan Li 0001, Liyong Wang, Haiyan Guan, Zexun Geng
IGARSS2
2014 Monitoring urban expansion of the Greater Toronto area from 1985 to 2013 using Landsat images
abstract
This study aims to detect the urban expansion in the Greater Toronto Area (GTA) in the period of 29 years lasting from 1985 to 2013 using the optical remote sensing data. A time series study is carried out and the change of the urban area and non-urban area is analyzed bi-temporally and multi-temporally by using the post-classification comparison method. Landsat images can been used to examine the land use and land cover changes of GTA in a long time period. The extent and spatial patterns of urban expansion are both analyzed quantitatively in the study, and it is effective to integrate GIS technology into remote sensing applications like this paper.
Shiqian Wang, Wei Li 0151, Jonathan Li 0001
IGARSS4
2014 Assimilation of SMOS soil moisture in the MESH model with the ensemble Kalman filter
abstract
Over the past decade, satellite soil moisture retrievals have showed great potential to improve land surface and hydrologic modeling, especially through an advanced data assimilation system. Data assimilation can be viewed as a process to optimally merge the model estimate and the observed information based upon some estimate of their error characteristics. This paper presents a case study of assimilating the Soil Moisture and Ocean Salinity (SMOS) satellite soil moisture retrievals (2010-2013) into a coupled land-surface and hydrological model MESH with an ensemble Kalman filter (EnKF). The assimilation experiment is conducted over the Great Lakes basin. The assimilation is validated against in situ soil moisture measurements (53 sites) from the Michigan Automated Weather Network, the Soil Climate Analysis Network, and the Fluxnet-Canada, in terms of the daily-spaced anomaly time series correlation coefficient (soil moisture skill). Results indicate that the assimilation of SMOS retrievals enhances the MESH model's soil moisture skill.
Xiaoyong Xu, Jonathan Li 0001, Bryan A. Tolson, Ralf M. Staebler, Frank Seglenieks, Bruce Davison, Amin Haghnegahdar, Eric D. Soulis
IGARSS2
2014 3D crack skeleton extraction from mobile LiDAR point clouds
abstract
This paper presents a novel algorithm for extracting 3D crack skeletons from 3D point clouds acquired by a mobile Light Detection and Ranging (LiDAR) system. This algorithm uses intensity information of cloud clouds to identify pavement cracks that usually exhibit lower intensities compared to their surroundings. First, crack candidates are extracted by applying the Otsu thresholding algorithm. Then, a spatial density filter is used to remove outliers. Next, crack points are grouped into crack-lines using a Euclidean distance clustering method. Finally, crack skeletons are extracted based on an L1-medial skeleton extraction method. The proposed algorithm has been tested on a set of mobile LiDAR point clouds acquired by a state-of-the-art RIEGL VMX-450 mobile LiDAR system. The results demonstrate the efficiency and reliability of the proposed algorithm in extracting 3D crack skeletons.
Yongtao Yu, Jonathan Li 0001, Haiyan Guan, Cheng Wang 0003
IGARSS2
2014 Semantic preserving distance metric learning and applications
Jun Yu 0002, Dapeng Tao, Jonathan Li 0001, Jun Cheng 0002
Inf. Sci.3
2014 Sparse Representation Based Pansharpening Using Trained Dictionary
abstract
Sparse representation has been used to fuse high-resolution panchromatic (HRP) and low-resolution multispectral (LRM) images. However, the approach faces the difficulty that the dictionary is generated from the high-resolution multispectral (HRM) images, which are unknown. In this letter, a two-step method is proposed to train the dictionary from the HRP and LRM images. In the first step, coarse HRM images are obtained by additive wavelet fusion method. The initial dictionary is composed of randomly sampled patches from the coarse HRM images. In the second step, a linear constraint K-SVD method is designed to train the dictionary to improve its representation ability. Experimental results using QuickBird and IKONOS data indicate that the trained dictionary yields comparable fusion products with raw-patch-dictionary sampled from HRM images.
Ming Cheng 0002, Cheng Wang 0003, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.3
2014 Semisupervised Classification for Hyperspectral Imagery With Transductive Multiple-Kernel Learning
abstract
The classification of hyperspectral imagery is a challenging problem because few labeled pixels are available. In this letter, we propose a new semisupervised learning algorithm to combine both cluster and manifold assumptions to increase classification reliability and accuracy. The new method uses a concave-convex procedure and sequential minimization optimization technologies for transductive multiple-kernel learning (TMKL). Then, a one-against-all strategy is adopted to generalize the binary TMKL classifiers to solve the multiclass problem of remote sensing images. Experimental results on two real data sets indicate that the proposed method exhibits both high accuracy and good computational performance.
Cheng Wang 0003, Dilong Li, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.4
2014 Object Detection in Terrestrial Laser Scanning Point Clouds Based on Hough Forest
abstract
This letter presents a novel rotation-invariant method for object detection from terrestrial 3-D laser scanning point clouds acquired in complex urban environments. We utilize the Implicit Shape Model to describe object categories, and extend the Hough Forest framework for object detection in 3-D point clouds. A 3-D local patch is described by structure and reflectance features and then mapped to the probabilistic vote about the possible location of the object center. Objects are detected at the peak points in the 3-D Hough voting space. To deal with the arbitrary azimuths of objects in real world, circular voting strategy is introduced by rotating the offset vector. To deal with the interference of adjacent objects, distance weighted voting is proposed. Large-scale real-world point cloud data collected by terrestrial mobile laser scanning systems are used to evaluate the performance. Experimental results demonstrate that the proposed method outperforms the state-of-the-art 3-D object detection methods.
Hanyun Wang, Cheng Wang 0003, Huan Luo 0001, Peng Li 0064, Ming Cheng 0002, Chenglu Wen, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.7
2014 Three-Dimensional Indoor Mobile Mapping With Fusion of Two-Dimensional Laser Scanner and RGB-D Camera Data
abstract
Three-dimensional mobile mapping in indoor environment, mostly global navigation satellite system-denied space, is to consecutively align the frames to build a global 3-D map of an indoor environment. One of the major difficulties of the current solutions is the failure at the insufficient overlapping between the frames, which is the reality of a lack of correspondences between the frames. To overcome this problem, a 3-D indoor mobile mapping system that integrates a 2-D laser scanner, and an RGB-Depth camera is presented in this letter. In this system, a fusion-iterative closest point (ICP) method, which combines the 2-D mobile platform pose from a Rao-Blackwellized particle filter estimation, an ICP, and a generalized-ICP method, is proposed for the consecutive frame alignment. Fusion-ICP achieves effective frame alignment, particularly in solving the insufficient overlapping frame alignment problem. Comparative experiments were conducted to evaluate the mapping system. The experimental results demonstrate the effectiveness and efficiency of our system for 3-D indoor mobile mapping.
Chenglu Wen, Qingyuan Zhu, Cheng Wang 0003, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.5
2014 Bayesian Classification of Hyperspectral Imagery Based on Probabilistic Sparse Representation and Markov Random Field
abstract
This letter presents a Bayesian method for hyperspectral image classification based on the sparse representation (SR) of spectral information and the Markov random field modeling of spatial information. We introduce a probabilistic SR approach to estimate the class conditional distribution, which proved to be a powerful feature extraction technique to be combined with the label prior distribution in a Bayesian framework. The resulting maximum a priori problem is estimated by a graph-cut-based α-expansion technique. The capabilities of the proposed method are proven in several benchmark hyperspectral images of both agricultural and urban areas.
Linlin Xu, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.2
2014 K-P-Means: A Clustering Algorithm of K "Purified" Means for Hyperspectral Endmember Estimation
abstract
This letter presents K-P-Means, a novel approach for hyperspectral endmember estimation. Spectral unmixing is formulated as a clustering problem, with the goal of K-P-Means to obtain a set of “purified” hyperspectral pixels to estimate endmembers. The K-P-Means algorithm alternates iteratively between two main steps (abundance estimation and endmember update) until convergence to yield final endmember estimates. Experiments using both simulated and real hyperspectral images show that the proposed K-P-Means method provides strong endmember and abundance estimation results compared with existing approaches.
Linlin Xu, Jonathan Li 0001, Alexander Wong, Junhuan Peng
IEEE Geosci. Remote. Sens. Lett.2
2014 Automated Detection of Road Manhole and Sewer Well Covers From Mobile LiDAR Point Clouds
abstract
A novel object detection algorithm is developed for automatically detecting road manhole and sewer well covers from mobile light detection and ranging point clouds. This algorithm takes advantage of a marked point process of disks and rectangles to model the locations of manhole and sewer well covers and their geometric dimensions. A reversible jump Markov chain Monte Carlo algorithm is implemented for simulating the posterior distribution obtained using a Bayesian paradigm. The detection results obtained from the road surface point clouds acquired by a RIEGL VMX-450 system show that the manhole and sewer well covers can be detected automatically and accurately. The performance achieved using the proposed algorithm is much more accurate and effective than those of the other three existing algorithms.
Yongtao Yu, Jonathan Li 0001, Haiyan Guan, Cheng Wang 0003, Jun Yu 0002
IEEE Geosci. Remote. Sens. Lett.2
2014 Pairwise Three-Dimensional Shape Context for Partial Object Matching and Retrieval on Mobile Laser Scanning Data
abstract
A novel pairwise 3-D shape context for partial object matching and retrieval is developed for extracting 3-D light poles and trees from mobile laser scanning (MLS) point clouds in a typical urban street scene. Unlike the single-point shape context describing only the local topology of a shape, the pairwise 3-D shape context can simultaneously model the local and global geometric structures of a shape in manifold space. By using histogram descriptors, the pairwise 3-D shape context has such characteristics as invariance to scale, invariance to orientation, and partial insensitivity to topological changes. Our results show that 3-D light poles and individual trees can be extracted from the RIEGL VMX-450 MLS point clouds and the performance achieved using our algorithm is much more accurate and effective than those of the other two existing algorithms.
Yongtao Yu, Jonathan Li 0001, Jun Yu 0002, Haiyan Guan, Cheng Wang 0003
IEEE Geosci. Remote. Sens. Lett.2
2014 SAR Image Denoising via Clustering-Based Principal Component Analysis
abstract
The combination of nonlocal grouping and transformed domain filtering has led to the state-of-the-art denoising techniques. In this paper, we extend this line of study to the denoising of synthetic aperture radar (SAR) images based on clustering the noisy image into disjoint local regions with similar spatial structure and denoising each region by the linear minimum mean-square error (LMMSE) filtering in principal component analysis (PCA) domain. Both clustering and denoising are performed on image patches. For clustering, to reduce dimensionality and resist the influence of noise, several leading principal components identified by the minimum description length criterion are used to feed the K-means clustering algorithm. For denoising, to avoid the limitations of the homomorphic approach, we build our denoising scheme on additive signal-dependent noise model and derive a PCA-based LMMSE denoising model for multiplicative noise. Denoised patches of all clusters are finally used to reconstruct the noise-free image. The experiments demonstrate that the proposed algorithm achieved better performance than the referenced state-of-the-art methods in terms of both noise reduction and image detail preservation.
Linlin Xu, Jonathan Li 0001, Yuanming Shu, Junhuan Peng
IEEE Trans. Geosci. Remote. Sens.2
2013 Hyperspectral image classification using Primal Laplacian SVM in preconditioned conjugate gradient solution
abstract
With the introduction of manifold assumption, Laplacian Support Vector Machine (LapSVM) has advantages over the traditional SVM classifiers. However the dual solution of LapSVM is still a major barrier on the further application of LapSVM. Primal optimization is a promising solution to this problem. In this paper, we introduce a novel primal Laplacian Support Vector Machine with Precondition Conjugate Gradient method (PCG) to the problem of hyperspectral images classification which is one type of primal optimization solution. To prove the effectiveness of the proposed method, we apply it into the hyperspectral image data set Indian Pine. The experiment results show higher accuracy and better generalization ability than dual strategy.
Xiaoli Ma, Cheng Wang 0003, Chenglu Wen, Jonathan Li 0001
IGARSS5
2013 An IconMap-based exploratory analytical approach for multivariate geospatial data
Xianfeng Zhang, Chunhua Liao, Jonathan Li 0001
Sci. China Inf. Sci.4
2013 Multi-view hypergraph learning by patch alignment framework
Jun Yu 0002, Jonathan Li 0001
Neurocomputing3
2013 Learn Multiple-Kernel SVMs for Domain Adaptation in Hyperspectral Data
abstract
This letter presents a novel semisupervised method for addressing a domain adaptation problem in the classification of hyperspectral data. To overcome the influence of distribution bias between the source and target domains, we introduce the domain transfer multiple-kernel learning to simultaneously minimize the maximum mean discrepancy criterion and the structural risk functional of support vector machines. Then, the pairwise binary classifiers are merged as the multiclass classifier for solving the classification problem in hyperspectral data. Both bias and nonbias sampling strategies are introduced to evaluate the robustness of the proposed method against the spectral distribution bias. The results obtained from real data sets show that the proposed method can achieve higher classification accuracy even with cross-domain distribution bias and provide robust solutions with different labeled and unlabeled data sizes.
Cheng Wang 0003, Hanyun Wang, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.4
2013 Semiautomated Building Facade Footprint Extraction From Mobile LiDAR Point Clouds
abstract
This letter presents a novel method for automated footprint extraction of building facades from mobile LiDAR point clouds. The proposed method first generates the georeferenced feature image of a mobile LiDAR point cloud and then uses image segmentation to extract contour areas which contain facade points of buildings, points of trees, and points of other objects in the georeferenced feature image. After all the points in each contour area are extracted, a classification based on principal component analysis (PCA) method is adopted to identify building objects from point clouds extracted in contour areas. Then, all the points in a building object are segmented into different planes using the random sample consensus algorithm. For each building, points in facade planes are chosen to calculate the direction, the start point, and the end point of the facade footprints using PCA. Finally, footprints of different facades of building are refined, harmonized, and joined. Two data sets of downtown areas and one data set of a residential area captured by Optech's LYNX mobile mapping system were tested to verify the validities of the proposed method. Experimental results show that the proposed method provides a promising and valid solution for automatically extracting building facade footprints from mobile LiDAR point clouds.
Bisheng Yang, Qingquan Li 0001, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.4
2013 Unsupervised Land Cover/Land Use Classification Using PolSAR Imagery Based on Scattering Similarity
abstract
This paper presents a new unsupervised land cover/land use classification scheme using polarimetric synthetic aperture radar (PolSAR) imagery based on polarimetric scattering similarity. Compared with theH/alpha classification scheme based on a dominant “average” scattering mechanism, the proposed scheme has such advantages as the following: (1) The major scattering mechanism represents a target scattering in the low-entropy case; (2) it also represents both the major and minor scattering mechanisms in the medium-entropy case; and (3) all the scattering mechanisms in the high-entropy case can be represented. The major and minor scattering mechanisms have been identified automatically based on the relative magnitude of multiple-scattering similarities. The canonical scattering corresponding to maximum scattering similarity is regarded as the major scattering mechanism. The result obtained using the National Aeronautics and Space Administration/Jet Propulsion Laboratory's AIRSAR L-band PolSAR imagery reveals that the proposed scheme is more effective as compared to the existing models and promises to increase the accuracy of the classification and interpretation.
Qiang Chen 0015, Gangyao Kuang, Jonathan Li 0001, Lichun Sui, Diangang Li
IEEE Trans. Geosci. Remote. Sens.3
2013 ETVOS: An Enhanced Total Variation Optimization Segmentation Approach for SAR Sea-Ice Image Segmentation
abstract
This paper presents a novel enhanced total variation optimization segmentation (ETVOS) approach consisting of two phases to segmentation of various sea-ice types. In the total variation optimization phase, the Rudin-Osher-Fatemi total variation model was modified and implemented iteratively to estimate the piecewise constant state from a nonpiecewise constant state (the original noisy imagery) by minimizing the total variation constraints. In the finite mixture model classification phase, based on the pixel distribution, an expectation maximization method was performed to estimate the final class likelihood using a Gaussian mixture model. Then, a maximum likelihood classification technique was utilized to estimate the final class of each pixel that appeared in the product of the total variation optimization phase. The proposed method was tested on a synthetic image and various subsets of RADARSAT-2 imagery, and the results were compared with other well-established approaches. With the advantage of a short processing time, the visual inspection and quantitative analysis of segmentation results confirm the superiority of the proposed ETVOS method over other existing methods.
Tae J. Kwon, Jonathan Li 0001, Alexander Wong
IEEE Trans. Geosci. Remote. Sens.2
2012 Edge-Guided Multiscale Segmentation of Satellite Multispectral Imagery
abstract
This paper presents a new approach to multiscale segmentation of satellite multispectral imagery using edge information. The Canny edge detector is applied to perform multispectral edge detection. The detected edge features are then utilized in a multiscale segmentation loop, and the merge procedure for adjacent image objects is controlled by a separability criterion that combines edge information with segmentation scale. The significance of the edge is measured by adjacent partitioned regions to perform edge assessment. The present method is based on a half-partition structure, which is composed of three steps: single edge detection, separated pixel grouping, and significant feature calculation. The spectral distance of the half-partitions separated by the edge is calculated, compared, and integrated into the edge information. The results show that the proposed approach works well on satellite multispectral images of a coastal area.
Jonathan Li 0001, Delu Pan, Qiankun Zhu, Zhihua Mao
IEEE Trans. Geosci. Remote. Sens.2
2010 Segmentation of SAR Intensity Imagery With a Voronoi Tessellation, Bayesian Inference, and Reversible Jump MCMC Algorithm
abstract
This paper presents a region-based approach to segmentation of the satellite synthetic aperture radar (SAR) intensity imagery. The approach is based on a Voronoi tessellation, the Bayesian inference, and the reversible jump Markov chain Monte Carlo (RJMCMC) algorithm. By Voronoi tessellation, the approach partitions a SAR image into a set of polygons corresponding to the components of the segmented homogenous regions. Each polygon is assigned a label to indicate a homogeneous region. The labels for all the polygons form a label field, which is characterized by an improved Potts model. The intensities of pixels in each polygon are assumed to satisfy identical and independent gamma distributions in terms of their label. Following the Bayesian paradigm, the posterior distribution that characterizes the SAR image segmentation can be obtained up to the integration constant. Then, a RJMCMC scheme is designed to simulate the posterior distribution and estimate its parameters. Finally, an optimal segmentation is obtained by the maximum a posteriori algorithm. The results obtained on both real Radarsat-1/2 and simulated SAR intensity images show that our approach works well and is very promising.
Yu Li 0002, Jonathan Li 0001, Michael A. Chapman
IEEE Trans. Geosci. Remote. Sens.2
2009 Automatic Generation of Seamline Network Using Area Voronoi Diagrams With Overlap
abstract
The mosaicking of orthoimages has been used to cover a large geographic region for various applications ranging from environmental monitoring to disaster management. However, existing mosaicking methods mainly focus on the generation of seamlines between two adjacent orthoimages. In this paper, we present a novel approach based on the use of a seamline network formed by a novel area Voronoi diagrams with overlap and the use of effective mosaic polygons (EMPs) to define the pixels of each orthoimage for the final mosaic. The generated seamline network is global based and is also optimized after refinement. It gives an effective partitioning for the regions of all orthoimages to form EMPs. The partitioning is unique, seamless, and has no redundancy. The algorithm is parallel, and the EMP of each orthoimage only has relation to orthoimages which have overlaps with it. It can ensure the flexibility and efficiency of mosaicking, without an intermediate process and independent of the sequence of the image composite. The experimental results obtained from the mosaicking of 40 color orthoimages demonstrate considerable potential for generating a seamline network automatically and effectively. This is extremely useful when a seamless mosaic is required to cover a large geographic region.
Jun Pan 0001, Mi Wang, DeRen Li, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.4
2002 Combination of vegetation indices with pesticide canopy emission model
abstract
Application of vegetation indices on a pesticide canopy emission model was carried out in this study. A new approach was developed to calculate LAI using TNDVI data. A case study in Indian Head using Landsat-7 TM and IKONOS image data showed the feasibility of the proposed model.
Bing Chen 0003, Gordon H. Huang, Jonathan Li 0001, Baiyu Zhang, H. L. Li
IGARSS3