Lingfei Ma

dblp:69/7771 · DBLP profile ↗
← Back
35ranked-venue papers
4as first author
27since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 31 · 4 first-author · 23 since 2021Computer networks · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Boundary-Guided Real-Time Semantic Segmentation and Pixel-Level Quantification of Pavement Cracks
abstract
Timely and accurately extracting and assessing pavement cracks is crucial for intelligent transportation systems (ITS) to improve road maintenance and safety. In this paper, we present an automated framework for crack semantic segmentation and quantification using optical images. First, a unique boundary-guided real-time high-resolution network is proposed, termed as BulletNet, for crack semantic segmentation. BulletNet is a bullet-head structure that can retain crack details while ensuring real-time inference speed, in which a Cross-Scale Global Attention (CSGA) module is designed to enhance global feature representation and pixel-level relations, as well as a Boundary-Guided Fusion (BGF) module proposed to utilize boundary features to guide the fusion of crack details and contextual information. Second, a Pixel-level Crack Quantification (PCQ) algorithm is proposed for complex cracks, incorporating an Improved Discrete Skeleton Evolution (IDSE) method to optimize skeleton pruning for accurate crack length and a normal vector correction method to adjust propagation direction for precise crack width. Comprehensive experiments on three datasets showed that the proposed BulletNet surpassed the comparative models in terms of efficiency and performance, with average F1-score, mIoU, and Frames per second (FPS) of 87.20%, 88.70%, and 125.53, respectively. In addition, tested on 200 images, the PCQ calculated the crack maximum widths and lengths with an average relative error of 6.96% and 4.62%, respectively. Finally, BulletNet was deployed on edge devices for field testing, and a system based on the PCQ algorithm was developed to validate the effectiveness of the entire framework.
Haiyan Guan, Lingfei Ma, Yongtao Yu, Sangning Li, Jonathan Li 0001
IEEE Trans. Intell. Transp. Syst.4
2025 "Pilot" to "Embodier": Brain-Controlled Robotic Arm With the E-VEP Paradigm in 3-D Manufacturing Scenarios for IoT
abstract
Robotic arm operation based on human–machine collaboration in manufacturing scenarios for the Internet of Things (IoT) has become an important research direction, especially in three-dimensional (3-D) scenarios that require high precision and flexible operation. However, owing to the complexity of operating robotic arms in 3-D scenarios, it is challenging for humans to perform tasks in pilot mode, leading to unnatural human–machine interactions. In this study, an embodied visual evoked potential (E-VEP) paradigm is proposed that can be used to control robotic arms in manufacturing scenarios in embodier mode. In addition, an incremental self-learning intention decoding (ISLID) algorithm is established to address the temporal variability in electroencephalography (EEG) signals. A brain-controlled robotic arm system was developed on the basis of the E-VEP paradigm and the ISLID algorithm. Online free grasping experiments revealed that the task time cost, output delay, and intention output ratio of the proposed system were 89.04 s, 2.22 s, and 46.59%, respectively. Compared with those of brain-controlled robotic arm systems based on the dynamic visual evoked potential (D-VEP) and SSVEP paradigms, the system based on the E-VEP paradigm achieved reductions in the average task time cost of 13.44% and 24.54%, respectively, and reductions in the average intention output ratio of 17.01% and 26.65%, respectively. The proposed brain-controlled robotic arm system holds significant application value in intelligent manufacturing scenarios for the IoT, advancing the integration of brain–machine interfaces and IoT technologies. The video (https://youtu.be/WtRHew4WGyo) demonstrates the utilization process of the proposed brain-controlled robotic arm.
Zhiyuan Ming, Mengxin Liu, Lingfei Ma, Tianyi Yan
IEEE Internet Things J.9
2025 Graph Representation Learning for Infrared and Visible Image Fusion
abstract
Infrared and visible image fusion aims to extract complementary features to synthesize a single fused image. In our method, we covert the regular image format into the graph space and conduct graph convolutional networks (GCNs) to extract NLss for the reliable infrared and visible image fusion. More specifically, GCNs are first performed on each intra-modal set to aggregate the features and propagate the inherent information, thereby extracting independent intra-modal NLss. Then, such intra-modal non-local self-similarity (NLss) features of infrared and visible images are concatenated to explore cross-domain NLss inter-modally and reconstruct the fused images. Extensive experiments show the superior performance of our method with the qualitative and quantitative analysis on the TNO, RoadScene and M3FD datasets, respectively, outperforming many state-of-the-art (SOTA) methods for the robust and effective infrared and visible image fusion.
Jing Li 0040, Lu Bai 0001, Bin Yang 0008, Chang Li 0001, Lingfei Ma
IEEE Trans Autom. Sci. Eng.5
2025 Point-SCT: A Multiscale Spatial Convolution-Swin Transformer Network for Point Cloud Ground Filtering in Complex Mountainous Terrains
abstract
Deep learning-based point cloud segmentation methods have been extensively explored, but the majority focus either on local or global feature learning, with few integrating both. These integrated approaches have not been sufficiently explored in complex mountainous scenes with low feature heterogeneity. To address this gap, we propose a novel point-based multi-scale spatial Convolution-Swin Transformer network (Point-SCT). Point-SCT combines convolutional local geometric detail capture with global relationship modeling via dynamic window interactions in Transformer, enhancing ground filtering accuracy in challenging mountainous scenes. The encoder incorporates convolution-based Multi-Scale Local Feature Aggregation (MLFA) approach, integrating Local Geometric Feature Encoding (LGSE) and Diluent Pooling (DP) strategies to effectively aggregate local detailed geometric features while suppressing irrelevant feature vectors and enhancing the representation of low-heterogeneity feature. Additionally, the dynamic spatial window strategy within the Transformer facilitates the capture of long-range feature dependencies. To mitigate noise introduced by RGB in point cloud overlays and sharpen geometric distinctions between the ground and low-lying vegetation, we introduce Boundary Detector, Curvature, and Average Elevation (BCE) as prior inputs, replacing RGB. Finally, quantitative and qualitative analyses of Point-SCT are conducted on an airborne laser scanning (ALS) dataset from a mountainous area, with ablation studies validating the effectiveness of LGSE, DP and BCE. The comprehensive experiments demonstrate that Point-SCT robustly segments ground points in complex mountainous scenes, achieving state-of-the-art levels of accuracy and generalization.
Fuquan Tang, Lingfei Ma, Nur Intan Raihana Ruhaiyem, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.3
2025 Brain-Controlled Hand Exoskeleton Based on Augmented Reality-Fused Stimulus Paradigm
abstract
Advancements in brain-machine interfaces (BMIs) have led to the development of novel rehabilitation training methods for people with impaired hand function. However, contemporary hand exoskeleton systems predominantly adopt passive control methods, leading to low system performance. In this work, an active brain-controlled hand exoskeleton system is proposed that uses a novel augmented reality-fused stimulus (AR-FS) paradigm as a human-machine interface, which enables users to actively control their fingers to move. Considering that the proposed AR-FS paradigm generates movement artifacts during hand movements, an enhanced decoding algorithm is designed to improve the decoding accuracy and robustness of the system. In online experiments, participants performed online control tasks using the proposed system, with an average task time cost of 16.27 s, an average output latency of 1.54 s, and an average correlation instantaneous rate (CIR) of 0.0321. The proposed system shows 35.37% better efficiency, 8.03% reduced system delay, and 35.28% better stability than the traditional system. This study not only provides an efficient rehabilitation solution for people with impaired hand function but also expands the application prospects of brain-control technology in areas such as human augmentation, patient monitoring, and remote robotic interaction. The video in Graphical Abstract Video demonstrates the user's process of operating the proposed brain-controlled hand exoskeleton system.
Zhiyuan Ming, Lingfei Ma, Dingjie Suo, Guangying Pei, Tianyi Yan
IEEE J. Biomed. Health Informatics7
2025 AI-Powered LiDAR Point Cloud Understanding and Processing: An Updated Survey
abstract
Current advances have enhanced the efficiency and availability of 3D data processing and scene understanding technologies, which confirms the pivotal status of point cloud data structure in 3D data transmission and storage. However, the intrinsic defects of point cloud data structure have always been a considerable challenge for model complexity and accuracy. As substantial researches have been conducted over the years, multitudinous compelling architectures applied in LiDAR point clouds are proposed in succession. To facilitate further research, we present a systematic integrated survey focusing explicitly on more than 200 key contributions to deep learning-based 3D LiDAR point cloud processing over the recent five years, detailing the revolution of feature extraction techniques of point cloud and specific deep learning-based tasks. Based on an introduction of the hardware sensors and devices of several LiDAR systems, the working mechanism and principle of 3D LiDAR point cloud acquisition and storage are interpreted. Moreover, about 30 publicly accessible datasets, including classical and latest outputs, are summarized and collated in accordance with different tasks. Comprehensive insights into the research challenges and opportunities in this topic are suggested. The contribution of this paper is to offer an up-to-date and all-sided overview of this realm, inspiring innovative explorations and further achievements in the AI-powered LiDAR point cloud understanding and processing.
Shanghui Jia, Xinghan Gong, Lingfei Ma
IEEE Trans. Intell. Transp. Syst.4
2025 Dual-Modal Prior Semantic Guided Infrared and Visible Image Fusion for Intelligent Transportation System
abstract
Infrared and visible image fusion (IVF) plays an important role in intelligent transportation system (ITS). The early works predominantly focus on boosting visual appeal of the fused result, although several recent approaches have tried to combine high-level vision task with IVF, they prioritize the design of cascaded structure to seek unified suitable features and fit different tasks. Thus, they tend to bias toward reconstructing raw pixels without considering the significance of semantic features. Therefore, we propose a novel prior semantic guided image fusion method based on the dual-modality strategy, improving the performance of IVF in ITS. Specifically, to explore the independent significant semantic of each modality, we first design two parallel semantic segmentation branches with a refined feature adaptive-modulation (RFaM) mechanism. RFaM can perceive the features that are semantically distinct enough in each semantic segmentation branch. Then, two pilot experiments based on the two branches are conducted to capture the significant prior semantic of source images, which is then applied to guide the fusion task in the integration of semantic segmentation branches and fusion branch. In addition, to aggregate both high-level semantics and impressive visual effects, we further investigate the frequency response of the prior semantics, and propose a multi-level representation-adaptive fusion (MRaF) module to explicitly integrate low-frequency prior semantic with high-frequency details. Extensive experiments on two public datasets demonstrate the superiority of our method over state-of-the-art fusion approaches. Our method has better performance on four quantitative metrics in fusion task and achieves the highest mIoU in semantic segmentation task.
Jing Li 0040, Lu Bai 0001, Bin Yang 0008, Chang Li 0001, Lingfei Ma, Lixin Cui, Edwin R. Hancock
IEEE Trans. Intell. Transp. Syst.5
2025 A Novel Nonlinear Smooth Controller for a Brain-Controlled Driving System in Complex Driving Scenarios
abstract
With the rapid advancement of technology, brain-controlled driving (BCD) has emerged as a contemporary focal point of research in academia and industry. BCD refers to the application of brain-machine interface (BMI) technology to driving, where control commands from the human brain are decoded by BMI technology and used to assist in the control of vehicles. However, existing BCD systems display inadequate performance in joint lateral and longitudinal control, and BCD systems in complex driving scenarios with other vehicles have not been studied. In this study, a nonlinear smooth controller is proposed, and a BCD system for complex driving scenarios is developed based on it. First, the BCD system is built from three modules, namely, the vehicle module, the BMI module and the controller module. Subsequently, the nonlinear smooth controller is developed based on the BMI controller, the proximal policy optimization (PPO) controller, and the self-adaptive collaborative (SAC) controller. The SAC controller is designed based on a sigmoid function to achieve nonlinear smoothness in the process of allocating control authority between the PPO controller and the BMI controller. The results of online driving experiments demonstrate that the proposed controller is better equipped to handle complex driving scenarios, exhibiting superior performance, heightened safety, and improved user experience compared to the PPO controller and BMI controller. This study holds significant value in advancing the practicality of BCD and providing a foundation for future research on BMI control.
Tianyi Yan, Zhiyuan Ming, Lingfei Ma, Dingjie Suo
IEEE Trans. Intell. Transp. Syst.6
2025 SLAM-TSM: Enhanced Indoor LiDAR SLAM With Total Station Measurements for Accurate Trajectory Estimation
abstract
Simultaneous Localization and Mapping (SLAM) is a crucial task in various domains, including intelligent robotics, computer vision, and indoor navigation. Accurate and robust trajectory estimation is especially challenging in indoor environments due to the presence of feature-poor or repetitive scenes, limited visibility, and dynamic objects. Obtaining highly accurate machine platform odometry is also an important basis for solving the “last mile” problem in intelligent transportation. This paper proposes a novel algorithm that combines LiDAR-based SLAM, total station measurements, and graph optimization to optimize the robot’s trajectory in indoor environments. By integrating highly accurate positional data from total station measurements as additional constraints, the proposed method enhances the performance of indoor LiDAR SLAM, effectively addressing the challenges of drift and trajectory offsets. Moreover, the proposed algorithm can provide trajectory optimization even in the absence of loop closure detection, making it more robust and suitable for a broader range of indoor environments. Experimental results validated the effectiveness of the proposed approach in reducing drift and improving trajectory estimation for low-cost indoor LiDAR devices, demonstrating its potential in various applications such as autonomous navigation, facility management, and augmented reality. It provides targeted ground truth for autonomous driving of machine platforms in intelligent transportation scenarios.
Dedong Zhang, Weikai Tan, John S. Zelek, Lingfei Ma, Jonathan Li 0001
IEEE Trans. Intell. Transp. Syst.4
2024 Multi-Granularity Feature Fusion For Point Cloud Semantic Segmentation Under Urban Scenes
abstract
Point cloud semantic segmentation plays a key role in scene understanding and digital twin cities tasks. This article proposed a multi-granularity feature fusion network (MGF-Net) for point cloud semantic segmentation. The model first used a cluster relation aggregation module to extract fine-grained point features and a 3D convolution module to extract coarse-grained voxel features, followed by feature aggregation via a multi-granularity feature adaptive fusion module. Finally, to further improve the model performance, MGF-Net used a global feature attention module to capture long-distance context information. The performance of MGF-Net was evaluated on three point cloud datasets of urban scenes, i.e., Toronto3D, WHU-MLS, and SensatUrban. The quantitative results showed that MGF-Net achieved 80.16%, 51.27%, and 54.20% of mIoU on these datasets, respectively. Moreover, the comparative results showed that the proposed MGF-Net outperformed the baseline for complex urban scenes, and obtained better point cloud semantic segmentation results.
Huchen Li, Lingfei Ma, Haiyan Guan, Nannan Qin, Yufu Zang
IGARSS2
2024 Weakly Supervised Point Cloud Segmentation by Combining Active Learning Annotation and Multi-Consistency Mechanism
abstract
In recent years, fully supervised learning based semantic segmentation algorithms for point clouds have achieved significant advancements. However, a major limitation of these traditional algorithms is their reliance on extensive labeled datasets. This impedes their practical applicability. To overcome this obstacle, this paper proposes a novel point cloud semantic segmentation framework based on weakly supervised learning. This framework is designed to segment point cloud data both efficiently and accurately, even with a limited budget (0.1%) of labeled data. The proposed approach initiates with an active learning annotation strategy. This strategy involves computing the uncertainty scores of each point and ranking them, consequently selecting the top-K points for labeling based on the labeling budget. Furthermore, this paper developed a weakly supervised learning network. This network is enhanced by the calculation of multiple consistency losses to enhance the network's performance. Experimental results demonstrate that with a mere 0.1% labeling ratio, the proposed framework achieves a mean Intersection over Union (mIoU) of 70.3% on the NPM3D dataset.
Haiyan Guan, Lingfei Ma, Nannan Qin, Yufu Zang
IGARSS3
2024 Remote-Oriented Brain-Controlled Unmanned Aerial Vehicle for IoT
abstract
With the rapid development of the internet of things (IoT) systems, the application potential of remote-oriented unmanned aerial vehicle (UAV) in IoT systems is becoming increasingly prominent. Brain-computer interface (BCI)-based remote-oriented UAV systems can not only leverage the natural advantages of the human brain in cognition and response, but also contribute to safer and more efficient operations in certain special environments. However, remote-oriented BCI systems still face challenges in spatial perception and control capabilities. In this study, a compressed-perceptual visual evoked potentials (CPVEP) paradigm and a human-machine closed-loop (HMCL) controller are proposed for a remote-oriented brain-controlled unmanned aerial vehicle (BCUAV). A BCVAV system for remote application scenarios is constructed based on the CPVEP paradigm and the HMCL controller. Online experiments demonstrates that all subjects have completed the navigation task by the proposed remote-oriented BCUAV system. Human-in-the-loop experiments show that the proposed system can significantly improve the system performance and adaptability of BCUAV to different environments, while significantly reducing the user’s workload. In the future, the proposed remote-oriented BCUAV system can be applied to various scenarios such as remote-controlled search and rescue, traffic monitoring and power line inspection.
Zhiyuan Ming, Lingfei Ma, Dingjie Suo, Tianyi Yan
IEEE Internet Things J.7
2024 HigherNet-DST: Higher-Resolution Network With Dynamic Scale Training for Rooftop Delineation
abstract
High-definition (HD) maps of building rooftops or footprints are important for urban application and disaster management. Rapid creation of such HD maps through rooftop delineation at the city scale using high-resolution satellite and aerial images with deep learning methods has become feasible and drawn much attention. However, the scale variance issue in rooftop delineation limited the overall performance. Existing methods exhibit considerably poor performance in rooftop delineation of small buildings. In this paper, we propose a new method, namely the Higher Resolution Network with Dynamic Scale Training (HigherNet-DST) to overcome the scale variance problem in rooftop delineation. Specifically, the DST is applied in the model training phase to reduce the negative impact of scale variance. Then, a scale-aware backbone, namely the Higher Resolution Network, is adopted to enhance the feature representation. Finally, the high-resolution supervision targets are used to further boost the delineation performance. Our method was tested on four publicly accessible building datasets and the results demonstrated that our method achieved the highest performance in rooftop delineation among the existing methods. Extensive experiments showed the superior performance of our method with an AP of 68.5% on the AICrowd Building Dataset and an IoU of 82.6% On the Inria Building Dataset, respectively, which surpassed many state-of-the-art (SOTA) methods. On the WHU Building Dataset and the Waterloo Building Dataset, our method also achieved the highest performance among the benchmarked methods, showing the high performance of our method for building boundary delineation.
Hongjie He 0003, Lingfei Ma, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 RdmkNet & Toronto-RDMK: Large-Scale Datasets for Road Marking Classification and Segmentation
abstract
Effective road marking classification and segmentation play a pivotal role in advancing vehicle-to-everything (V2X) applications and refining road inventory databases. However, the irregular data formats and unordered permutation modes of 3D point clouds, along with the limited availability of large-scale datasets with point-level annotations, remain significant obstacles to designing deep learning-based networks with superior performance. To address these challenges, this paper proposes a novel multi-level feature optimization network structure, named MFPNet, and introduces two point cloud benchmarks, RdmkNet and Toronto-Rdmk, for road marking classification and segmentation in intricate urban environments. MFPNet is composed of three integral modules. First, the M-transformer module, consisting of three transformers obtained from different channels, fully captures rich point cloud background information and long-distance dependencies between objects. Then, the feature pooling aggregation module uses parallel structured pooling attention mechanisms to aggregate features captured by the M-transformer module, while the prediction refinement module further enhances the acquisition of semantic features. Comparative studies indicate that MFPNet can be embedded into general deep learning networks without changing their original network structures, significantly improving the accuracy of multiple baseline networks. Furthermore, extensive experiments demonstrate that the two newly-developed point cloud datasets are meaningful for road marking classification and segmentation tasks, contributing to the development of autonomous driving.
Jing Du 0007, Lingfei Ma, Jing Li 0040, Nannan Qin, John S. Zelek, Haiyan Guan, Jonathan Li 0001
IEEE Trans. Intell. Transp. Syst.2
2024 Crack-U2Net: Multiscale Feature Learning Network for Pavement Crack Detection From Large-Scale MLS Point Clouds
abstract
Deep learning-based algorithms detect pavement cracks in an end-to-end manner from Mobile Laser Scanning (MLS) point clouds, achieving impressive results. However, the accuracy of existing methods still has room to improve due to the difficulty of effectively encoding multiscale features and the limited training data. In this paper, we propose a novel pavement crack detection framework, Crack-U2Net, which innovatively incorporates a two-level nested U-Net architecture for feature learning. This design enables the learning of intra-stage multiscale features without introducing significant memory and computation costs, resulting in substantial improvements in accuracy. Moreover, to solve the challenge of insufficient training data, we propose a Geometry-based Data Augmentation (GDA) strategy, aiming to expand the pavement dataset while preserving the pavement geometry. Extensive experiments on the Qinghai-Tibet Highway point cloud dataset demonstrate the higher accuracy and efficiency of Crack-U2Net over the state-of-the-art methods, achieving an average precision, recall, F$1\text - $score, and accuracy of 83.8%, 77.6%, 80.1%, and 95.8%, respectively.
Huifang Feng 0002, Wen Li 0005, Lingfei Ma, Yiping Chen 0002, Haiyan Guan, Yongtao Yu, José Marcato Junior, Jonathan Li 0001
IEEE Trans. Intell. Transp. Syst.3
2023 Enhancing Spatial Resolution of Building Datasets Using Transformer-Based Single-Image Super-Resolution
abstract
The spatial resolution of Earth Observation (EO) images plays a key role in building footprint extraction. For the spatial resolution enhancement, deep learning-based image super-resolution methods have been widely used due to their remarkable performance. Transformer-based networks are effective and has drawn much attention in computer vision but underutilized in remote sensing, especially for super-resolving building datasets. Therefore, in this paper, we developed a novel transformer-based Single-Image Super-Resolution (SISR) method, named Pyramid Vision Transformer-Residual Feature Aggregation Network (PVT_RFANet), to improve the spatial resolution of building datasets. Specifically, the PVT v2 network was embedded into our Momentum Spatial-Channel Attention Residual Feature Aggregation Network (MSCA-RFANet). Moreover we conducted a comparative study to compare our method with Bicubic interpolation (BI), Super-resolution Convolutional Neural Network (SRCNN), Deep Recursive Residual Network (DRRN), SRResNet, and MSCA-RFANet. Using Peak Signal-Noise Ratio (PSNR) and Similarity Structure Index Measurement (SSIM) as the evaluation metrics, our method showed highest performance with the PSNR of 22.01 dB and the SSIM of 0.50 on the WHU Building Dataset, which demonstrated the superior performance of the proposed method.
Yuwei Cai, Hongjie He 0003, Zhimeng He, Michael A. Chapman, Jing Li 0040, Lingfei Ma, Jonathan Li 0001
IGARSS6
2023 Nighttime Light Missing Data Retrieval Using Modis Version 6 Satellite Data and Mask Dilated Partial Convolutional Neural Network
abstract
Nighttime Lights (NTLs) remote sensing imagery contains tremendous information and has been shown to accurately predict a region’s human dynamics, economic health and energy consumption. Despite its usefulness, NTLs imagery is less widely available than other remote sensing data modalities. Several challenges appear when attempting to reconstruct NTLs data, either from other data modalities or existing NTLs data. These include complex non-linear relationships between NTLs and multispectral bands, non-matching spatial and temporal coverage, and different atmospheric and cloud conditions. This study attempts to create an out-of-the-box model that compensates for missing NTLs data using widely available daytime data in a broadly generalizable manner. The proposed project has two objectives: the construction of an image-to-image dataset mapping daytime multispectral images (MODIS V6 Land Surface Reflectance, MODIS V6 Land Cover, MODIS V6 Vegetation Indices) to NTLs images, and the reconstruction of NTLs data using deep learning techniques by researching, creating, and employing the state-of-the-art architecture of the Mask Partial Convolutional Neural Network in conjunction with dilated convolutions. The project will facilitate the training of new models for predicting missing NTLs and make NTLs data more accessible for future remote sensing research.
Xuanchen Liu, Shuxin Qiao, Kyle Gao, Hongjie He 0003, Lingfei Ma, Jonathan Li 0001
IGARSS5
2023 A systematic mapping study for graphical user interface testing on mobile apps
abstract
Abstract Mobile apps with tested Graphical User Interface (GUI) tend to have higher downloads in the apps store. In recent years, few efforts were made to analyse the research community and research status of the literature for GUI testing on mobile apps, which brings an obstacle to characterise and understand this field. In this study, the authors propose a systematic mapping study to gain insights into the field. First, the authors conduct an extensive search of relevant literature over seven popular digital libraries. From 4427 candidate studies, 114 primary studies published between January 2011 and September 2022 were selected. Next, the authors analyse these primary studies from the perspectives of bibliometric and qualitative analysis. For the bibliometric analysis, first, the authors analyse the popular research topics and their relationships. Second, the authors study the authors' community. For the qualitative analysis, the authors analyse the objectives, approaches and evaluation metrics employed in these primary studies. Their investigation reports several major findings: (1) there are relatively more studies on two topics, that is, test case generation and the automated test; (2) the most productive authors tend to collaborate and often have relatively broad research interests; (3) the functionality is the main objective of GUI testing; the model‐based approach is the most widely used.
Liming Nie, Kabir S. Said, Lingfei Ma, Yaowen Zheng
IET Softw.3
2022 STN: Saliency-Guided Transformer Network for Point-Wise Semantic Segmentation of Urban Scenes
abstract
Accurate and effective road object semantic segmentation plays a significant role in supporting extensive intelligent transportation system (ITS)-related applications. However, most existing image-based methods and point-based methods cannot deliver promising solutions with respect to segmentation accuracy and robustness, especially in complex urban road scenes. Thus, we design a saliency-guided transformer architecture (STN) in this letter for point-wise semantic segmentation from mobile laser scanning (MLS) point clouds. First, four types of feature saliency maps are constructed to obtain more compact feature spaces for enhancing the feature encoding semantics. Then, integrated with offset attention mechanisms and edge convolutions, an effective point-wise transformer network is proposed to extract high-level features for point-wise label assignment of road objects. The STN model is evaluated on the Pairs-Lille-3D dataset and achieves satisfactory experimental results with 87.2% overall accuracy and 81.7% mean IoU, respectively. Comparative studies with five deep learning-based methods also prove the superior performance of the STN model for large-scale semantic segmentation tasks.
Lingfei Ma, Jonathan Li 0001, Haiyan Guan, Yongtao Yu, Yiping Chen 0002
IEEE Geosci. Remote. Sens. Lett.1
2022 Spectral-Spatial Transformer Network for Hyperspectral Image Classification: A Factorized Architecture Search Framework
abstract
Neural networks have dominated the research of hyperspectral image classification, attributing to the feature learning capacity of convolution operations. However, the fixed geometric structure of convolution kernels hinders long-range interaction between features from distant locations. In this article, we propose a novel spectral–spatial transformer network (SSTN), which consists of spatial attention and spectral association modules, to overcome the constraints of convolution kernels. Also, we design a factorized architecture search (FAS) framework that involves two independent subprocedures to determine the layer-level operation choices and block-level orders of SSTN. Unlike conventional neural architecture search (NAS) that requires a bilevel optimization of both network parameters and architecture settings, the FAS focuses only on finding out optimal architecture settings to enable a stable and fast architecture search. Extensive experiments conducted on five popular HSI benchmarks demonstrate the versatility of SSTNs over other state-of-the-art (SOTA) methods and justify the FAS strategy. On the University of Houston dataset, SSTN obtains comparable overall accuracy to SOTA methods with a small fraction (1.2%) of multiply-and-accumulate operations compared to a strong baseline spectral–spatial residual network (SSRN). Most importantly, SSTNs outperform other SOTA networks using only 1.2% or fewer MACs of SSRNs on the Indian Pines, the Kennedy Space Center, the University of Pavia, and the Pavia Center datasets.
Zilong Zhong, Ying Li 0036, Lingfei Ma, Jonathan Li 0001, Wei-Shi Zheng 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 BoundaryNet: Extraction and Completion of Road Boundaries With Deep Learning Using Mobile Laser Scanning Point Clouds and Satellite Imagery
abstract
Robust road boundary extraction and completion play an important role in providing guidance to all road users and supporting high-definition (HD) maps. The significant challenges remain in remarkable and accurate road boundary recovery from poor road boundary conditions. This paper presents a novel deep learning framework, named BoundaryNet, to extract and complete road boundaries by using both mobile laser scanning (MLS) point clouds and high-resolution satellite imagery. First, road boundaries are extracted by conducting a curb-based extraction method. Such extracted 3D road boundary lines are used as inputs to feed into a U-shaped network for erroneous boundary denoising. Then, a convolutional neural network (CNN) model is proposed to complete the road boundaries. Next, to achieve more complete and accurate road boundaries, a conditional deep convolutional generative adversarial network (c-DCGAN) with the assistance of road centerlines extracted from satellite images is developed. Finally, according to the completed road boundaries, the inherent road geometries are calculated. The proposed methods were evaluated using satellite imagery and four MLS point cloud datasets with varying densities and road conditions in urban environments. The quality evaluation metrics of 82.88%, 82.43%, 88.86%, and 84.89% were achieved for four data sets. The experimental results indicate that the BoundaryNet model can provide a promising solution for road boundary completion and road geometry estimation.
Lingfei Ma, Ying Li 0036, Jonathan Li 0001, José Marcato Junior, Wesley Nunes Gonçalves, Michael A. Chapman
IEEE Trans. Intell. Transp. Syst.1
2022 Robust Lane Extraction From MLS Point Clouds Towards HD Maps Especially in Curve Road
abstract
This article presents a semi-automated method to extract the lane features along the curved roads from mobile laser scanning (MLS) point clouds. The proposed method consists of four steps. After data pre-processing, a road edge detection algorithm is performed to distinguish road curbs and extract road surfaces. Then, textual and directional road markings such as arrows, symbols, and words, to inform drivers in necessary cases, are detected by intensity thresholding and conditional Euclidean clustering algorithms. Furthermore, lane markings are extracted by local intensity analysis and distance thresholding methods according to road design standards, because they are more regular along the road. Finally, centerline points on lanes are estimated based on the coordinates of extracted lane markings. Our method shows strong feasibility and robustness when creating high-definition (HD) maps from MLS data, by increasing the number of blocks in the curve and the distance threshold control in curved lane centerline extraction. Quantitative evaluations show that the average recall, precision, and F1-score obtained from four datasets for road marking extraction are 93.87%, 93.76%, and 93.73%, respectively. The generated lane centerlines are evaluated by overlaying them on manually labeled reference buffers from 4 cm resolution orthoimagery. The comparative study indicates that the proposed methods can achieve higher accuracy and robustness than most state-of-the-art methods.
Chengming Ye, He Zhao 0007, Lingfei Ma, Han Jiang 0005, Hongfu Li, Ruisheng Wang 0001, Michael A. Chapman, José Marcato Junior, Jonathan Li 0001
IEEE Trans. Intell. Transp. Syst.3
2022 3D Vehicle Detection Using Multi-Level Fusion From Point Clouds and Images
abstract
3D vehicle detectors based on point clouds generally have higher detection performance than detectors based on multi-sensors. However, with the lack of texture information, point-based methods get many missing detection of occluded and distant vehicles, and false detection with high-confidence of similarly shaped objects, which is a potential threat to traffic safety. Therefore, in the long run, fusion-based methods have more potential. This paper presents a multi-level fusion network for 3D vehicle detection from point clouds and images. The fusion network includes three stages: data-level fusion of point clouds and images, feature-level fusion of voxel and Bird’s Eye View (BEV) in the point cloud branch, and feature-level fusion of point clouds and images. Besides, a novel coarse-fine detection header is proposed, which simulates the two-stage detectors, generating coarse proposals on the encoder, and refining them on the decoder. Extensive experiments show that the proposed network has better detection performance on occluded and distant vehicles, and reduces the false detection of similarly shaped objects, proving its superiority over some state-of-the-art detectors on the challenging KITTI benchmark. Ablation studies have also demonstrated the effectiveness of each designed module.
Lingfei Ma, José Marcato Junior, Wesley Nunes Gonçalves, Jonathan Li 0001
IEEE Trans. Intell. Transp. Syst.2
2021 Semantic Segmentation of UAV Lidar Point Clouds of a Stack Interchange with Deep Neural Networks
abstract
Stack interchanges are essential components of transportation systems. Mobile laser scanning (MLS) systems have been widely used in road infrastructure mapping, but accurate mapping of complicated multi-layer stack interchanges are still challenging. This study examined the point clouds collected by a new Unmanned Aerial Vehicle (UAV) Light Detection and Ranging (LiDAR) system to perform the semantic segmentation task of a stack interchange. An end-to-end supervised 3D deep learning framework was proposed to classify the point clouds. The proposed method has proven to capture 3D features in complicated interchange scenarios with stacked convolution and the result achieved over 93% classification accuracy. In addition, the new low-cost semi-solid-state LiDAR sensor Livox Mid-40 featuring a incommensurable rosette scanning pattern has demonstrated its potential in high-definition urban mapping.
Weikai Tan, Dedong Zhang, Lingfei Ma, Nannan Qin, Yiping Chen 0002, Jonathan Li 0001
IGARSS3
2021 Capsule-Based Networks for Road Marking Extraction and Classification From Mobile LiDAR Point Clouds
abstract
Accurate road marking extraction and classification play a significant role in the development of autonomous vehicles (AVs) and high-definition (HD) maps. Due to point density and intensity variations from mobile laser scanning (MLS) systems, most of the existing thresholding-based extraction methods and rule-based classification methods cannot deliver high efficiency and remarkable robustness. To address this, we propose a capsule-based deep learning framework for road marking extraction and classification from massive and unordered MLS point clouds. This framework mainly contains three modules. Module I is first implemented to segment road surfaces from 3D MLS point clouds, followed by an inverse distance weighting (IDW) interpolation method for 2D georeferenced image generation. Then, in Module II, a U-shaped capsule-based network is constructed to extract road markings based on the convolutional and deconvolutional capsule operations. Finally, a hybrid capsule-based network is developed to classify different types of road markings by using a revised dynamic routing algorithm and large-margin Softmax loss function. A road marking dataset containing both 3D point clouds and manually labeled reference data is built from three types of road scenes, including urban roads, highways, and underground garages. The proposed networks were accordingly evaluated by estimating robustness and efficiency using this dataset. Quantitative evaluations indicate the proposed extraction method can deliver 94.11% in precision, 90.52% in recall, and 92.43% in F1-score, respectively, while the classification network achieves an average of 3.42% misclassification rate in different road scenes.
Lingfei Ma, Ying Li 0036, Jonathan Li 0001, Yongtao Yu, José Marcato Junior, Wesley Nunes Gonçalves, Michael A. Chapman
IEEE Trans. Intell. Transp. Syst.1
2021 Multi-Scale Point-Wise Convolutional Neural Networks for 3D Object Segmentation From LiDAR Point Clouds in Large-Scale Environments
abstract
Although significant improvement has been achieved in fully autonomous driving and semantic high-definition map (HD) domains, most of the existing 3D point cloud segmentation methods cannot provide high representativeness and remarkable robustness. The principally increasing challenges remain in completely and efficiently extracting high-level 3D point cloud features, specifically in large-scale road environments. This paper provides an end-to-end feature extraction framework for 3D point cloud segmentation by using dynamic point-wise convolutional operations in multiple scales. Compared to existing point cloud segmentation methods that are commonly based on traditional convolutional neural networks (CNNs), our proposed method is less sensitive to data distribution and computational powers. This framework mainly includes four modules. Module I is first designed to construct a revised 3D point-wise convolutional operation. Then, a U-shaped downsampling-upsampling architecture is proposed to leverage both global and local features in multiple scales in Module II. Next, in Module III, high-level local edge features in 3D point neighborhoods are further extracted by using an adaptive graph convolutional neural network based on the K-Nearest Neighbor (KNN) algorithm. Finally, in Module IV, a conditional random field (CRF) algorithm is developed for postprocessing and segmentation result refinement. The proposed method was evaluated on three large-scale LiDAR point cloud datasets in both urban and indoor environments. The experimental results acquired by using different point cloud scenarios indicate our method can achieve state-of-the-art semantic segmentation performance in feature representativeness, segmentation accuracy, and technical robustness.
Lingfei Ma, Ying Li 0036, Jonathan Li 0001, Weikai Tan, Yongtao Yu, Michael A. Chapman
IEEE Trans. Intell. Transp. Syst.1
2021 Deep Learning for LiDAR Point Clouds in Autonomous Driving: A Review
abstract
Recently, the advancement of deep learning (DL) in discriminative feature learning from 3-D LiDAR data has led to rapid development in the field of autonomous driving. However, automated processing uneven, unstructured, noisy, and massive 3-D point clouds are a challenging and tedious task. In this article, we provide a systematic review of existing compelling DL architectures applied in LiDAR point clouds, detailing for specific tasks in autonomous driving, such as segmentation, detection, and classification. Although several published research articles focus on specific topics in computer vision for autonomous vehicles, to date, no general survey on DL applied in LiDAR point clouds for autonomous vehicles exists. Thus, the goal of this article is to narrow the gap in this topic. More than 140 key contributions in the recent five years are summarized in this survey, including the milestone 3-D deep architectures, the remarkable DL applications in 3-D semantic segmentation, object detection, and classification; specific data sets, evaluation metrics, and the state-of-the-art performance. Finally, we conclude the remaining challenges and future researches.
Ying Li 0036, Lingfei Ma, Zilong Zhong, Michael A. Chapman, Dongpu Cao, Jonathan Li 0001
IEEE Trans. Neural Networks Learn. Syst.2
2020 Early-Season Crop Classification with Radarsat-2 Polarimetric Synthetic Aperture Radar Imagery
abstract
Timely crop classification maps are essential for the agriculture sector to ensure food security and understand the state and trend of crop growth. Though there are several crop monitoring systems in operation, early-season crop classification is still in demand. We developed a robust crop growth estimation technology previously with synthetic aperture radar (SAR) imagery for canola in Canadian Prairies, and we are extending the procedure to enable accurate early-season crop classification. Here we present a dynamic crop classification technique with RADARSAT-2 (RS2) polarimetric SAR (Pol-SAR) imagery for the classification of canola, corn, soybean and wheat, the four major crop types in Canadian Prairies. The procedure achieved over 90% classification accuracy of the major four crop types in the testing area by the end of July.
Weikai Tan, Abhijit Sinha, Yifeng Li 0003, Lingfei Ma, Jonathan Li 0001
IGARSS4
2020 A Hybrid Capsule Network for Land Cover Classification Using Multispectral LiDAR Data
abstract
Land cover mapping is an effective way to quantify land resources and monitor their changes. It plays an important role in a wide range of applications. This letter proposes a hybrid capsule network for land cover classification using multispectral light detection and ranging (LiDAR) data. First, the multispectral LiDAR data were rasterized into a set of feature images to exploit the geometrical and spectral properties of different types of land covers. Then, a hybrid capsule network composed of an encoder network and a decoder network is trained to extract both high-level local and global entity-oriented capsule features for accurate land cover classification. Quantitative classification evaluations on two data sets show that the overall accuracy, average accuracy, and kappa coefficient of over 97.89%, 94.54%, and 0.9713, respectively, are obtained. Comparative studies with five existing methods confirm that the proposed method performs robustly and accurately in land cover classification using the multispectral LiDAR data.
Yongtao Yu, Haiyan Guan, Dilong Li, Tiannan Gu, Lanfang Wang, Lingfei Ma, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.6
2020 TGNet: Geometric Graph CNN on 3-D Point Cloud Segmentation
abstract
Recent geometric deep learning works define convolution operations in local regions and have enjoyed remarkable success on non-Euclidean data, including graph and point clouds. However, the high-level geometric correlations between the input and its neighboring coordinates or features are not fully exploited, resulting in suboptimal segmentation performance. In this article, we propose a novel graph convolution architecture, which we term as Taylor Gaussian mixture model (GMM) network (TGNet), to efficiently learn expressive and compositional local geometric features from point clouds. The TGNet is composed of basic geometric units, TGConv, that conduct local convolution on irregular point sets and are parametrized by a family of filters. Specifically, these filters are defined as the products of the local point features and the neighboring geometric features extracted from local coordinates. These geometric features are expressed by Gaussian weighted Taylor kernels. Then, a parametric pooling layer aggregates TGConv features to generate new feature vectors for each point. TGNet employs TGConv on multiscale neighborhoods to extract coarse-to-fine semantic deep features while improving its scale invariance. Additionally, a conditional random field (CRF) is adopted within the output layer to further improve the segmentation results. Using three point cloud data sets, qualitative and quantitative experimental results demonstrate that the proposed method achieves 62.2% average accuracy on ScanNet, 57.8% and 68.17% mean intersection over union (mIoU) on Stanford Large-Scale 3D Indoor Spaces (S3DIS) and Paris-Lille-3D data sets, respectively.
Ying Li 0036, Lingfei Ma, Zilong Zhong, Dongpu Cao, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.2
2020 Semi-Automated Generation of Road Transition Lines Using Mobile Laser Scanning Data
abstract
This paper recognizes the research gaps and difficulties in generating transition lines (the paths that pass through a road intersection) in road intersections from mobile laser scanning (MLS) point clouds. The proposed method contains three modules: road surface detection, lane marking extraction, and transition line generation. First, the points covering the road surface are extracted using the voxel-based upward growing and the improved region growing. Then, lane markings are extracted and identified according to the multi-thresholding and the geometric filtering. Finally, transition lines are generated through a combination of the lane node structure generation algorithm and the cubic Catmull-Rom spline algorithm. The experimental results demonstrate that transition lines can be successfully generated for both T- and cross-intersections with promising accuracy. In the validation of lane marking extraction using the manually interpreted lane marking points, the method can achieve average precision, recall, and F1-score of 90.80%, 92.07%, and 91.43%, respectively. The success rate of transition line generation is 96.5%. Furthermore, the buffer-overlay-statistics (BOS) method validates that the proposed method can generate lane centerlines and transition lines within 20-cm-level localization accuracy from the MLS point clouds.
Chengming Ye, Jonathan Li 0001, Han Jiang 0005, He Zhao 0007, Lingfei Ma, Michael A. Chapman
IEEE Trans. Intell. Transp. Syst.5
2018 Segment-Based Traffic Sign Detection from Mobile Laser Scanning Data
abstract
This paper presents a segment-based traffic sign detection method using vehicle-borne mobile laser scanning (MLS) data. This method has three steps: road scene segmentation, clustering and traffic sign detection. The non-ground points are firstly segmented from raw MLS data by estimating road ranges based on vehicle trajectory and geometric features of roads (e.g., surface normals and planarity). The ground points are then removed followed by obtaining non-ground points where traffic signs are contained. Secondly, clustering is conducted to detect the traffic sign segments (or candidates) from the non-ground points. Finally, these segments are classified to specified classes. Shape, elevation, intensity, 2D and 3D geometric and structural features of traffic sign patches are learned by the support vector machine (SVM) algorithm to detect traffic signs among segments. The proposed algorithm has been tested on a MLS point cloud dataset acquired by a Leador system in the urban environment. The results demonstrate the applicability of the proposed algorithm for detecting traffic signs in MLS point clouds.
Ying Li 0036, Lingfei Ma, Yuchun Huang, Jonathan Li 0001
IGARSS2
2018 Extraction of Building Windows from Mobile Laser Scanning Point Clouds
abstract
This study recognizes the significance and considerable commercial applications in creating Level of Detail (LoD) building models for 3D city models generation. Accordingly, this paper proposes a novel method to identify and extract window frames on building facades from Mobile Laser Scanning (MLS) point clouds. The proposed method can typically be regarded as a stepwise procedure. Firstly, a voxel-based upward-growing method is applied to distinguish non-ground points from ground points. Next, outliers are filtered out from non-ground points by statistical analysis. Then, all the remaining non-ground points are clustered based on the conditional Euclidean clustering algorithm to segment out building facades. A volumetric box is afterward created to store façade points so that neighbors of each point can be operated. Finally, a manipulator is applied according to the structural characteristics of window frames to extract the potential window points. Quantitative evaluations based on 2D validation and 3D validation were both conducted. In the 2D validation, the lowest F1-measure of the test datasets is 0.740, and the highest can be 0.977. While in the 3D validation, the lowest precision of the test dataset is 79.58%, and the highest can be 97.96%. The results demonstrate the proposed method can successfully extract the rectangular or curved windows in the test datasets with promising accuracies to support the generation of LoD3 building models.
Menglan Zhou, Lingfei Ma, Ying Li 0036, Jonathan Li 0001
IGARSS2
2017 Deep residual networks for hyperspectral image classification
abstract
Deep neural networks can learn deep feature representation for hyperspectral image (HSI) interpretation and achieve high classification accuracy in different datasets. However, counterintuitively, the classification performance of deep learning models degrades as their depth increases. Therefore, we add identity mappings to convolutional neural networks for every two convolutional layers to build deep residual networks (ResNets). To study the influence of deep learning model size on HSI classification accuracy, this paper applied ResNets and CNNs with different depth and width using two challenging datasets. Moreover, we tested the effectiveness of batch normalization as a regularization method with different model settings. The experimental results demonstrate that ResNets mitigate the declining-accuracy effect and achieved promising classification performance with 10% and 5% training sample percentages for the University of Pavia and Indian Pines datasets, respectively. In addition, t-Distributed Stochastic Neighbor Embedding (t-SNE) provides a direct view of the extracted features through dimensionality reduction.
Zilong Zhong, Jonathan Li 0001, Lingfei Ma, Han Jiang 0005, He Zhao 0007
IGARSS3
2016 Examining urban expansion using multi-temporal landsat imagery: A case study of montreal census metropolitan area
abstract
Greater Montreal is the most populous metropolitan area in Quebec, and the second most populous in Canada after Greater Toronto. In the 1970s, the economic center of Canada shifted from Montreal to Toronto. Since some previous studies have focused on the urbanization process in the Greater Toronto Area, it is important to conduct research on its counterpart. This study uses Landsat images as the data source, combined with census data to detect urban changes in the Montreal census metropolitan area (CMA) from 1975 to 2015. We analyzed spatial patterns and annual urban growth rate by applying four supervised classification algorithms. Also, we mapped temporal land cover categories and evaluated major driving forces that contribute to the urban changes. Our results show that Montreal CMA has experienced a rapid development over the past 40 years, with 442 km2urban growth. Urban expansion in Montreal CMA mainly has two modes: radiated and ribbon.
He Zhao 0007, Lingfei Ma, Jonathan Li 0001
IGARSS2