VLDB 2026 Research / reviewers in the wild / expert
Jie Mei 0004
dblp:47/4497-4
· DBLP profile ↗
22ranked-venue papers
6as first author
12since 2021 · last 2025
0000-0002-9789-4025ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Information sparsity guided transformer for multi-modal medical image super-resolution
Haotian Lu 0003, Jie Mei 0004, Fangwei Hao, Jing Xu 0008 |
Expert Syst. Appl. | 2 |
| 2025 | ECINFusion: A Novel Explicit Channel-Wise Interaction Network for Unified Multi-Modal Medical Image FusionabstractMulti-modal medical image fusion enhance the representation, aggregation and comprehension of functional and structural information, improving accuracy and efficiency for subsequent analysis. However, lacking explicit cross channel modeling and interaction among modalities results in the loss of details and artifacts. To this end, we propose a novelExplicitChannel-wiseInteractionNetwork for unified multi-modal medical imageFusion, namely ECINFusion. ECINFusion encompasses two components: multi-scale adaptive feature modeling (MAFM) and explicit channel-wise interaction mechanism (ECIM). MAFM leverages adaptive parallel convolution and transformer in multi-scale manner to achieve the global context-aware feature representation. ECIM utilizes the designed multi-head channel-attention mechanism for explicit modeling in channel dimension to accomplish the cross-modal interaction. Besides, we introduce a novel adaptive L-Norm loss, preserving fine-grained details. Experiments demonstrate ECINFusion outperforms state-of-the-art approaches in various medical fusion sub-tasks on different metrics. Furthermore, extended experiments reveal the robust generalization of the proposed in different fusion tasks. In breif, the proposed explicit channel-wise interaction mechanism provides new insight for multi-modal interaction. Xinjian Wei, Xiaoxuan Xu, Jing Xu 0008, Jie Mei 0004, Jun Zhang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Salient Object Detection in Traffic Scene Through the TSOD10K DatasetabstractTraffic Salient Object Detection (TSOD) aims to segment the objects critical to driving safety by combining semantic (e.g., collision risks) and visual saliency. Unlike SOD in natural scene images (NSI-SOD), which prioritizes visually distinctive regions, TSOD emphasizes the objects that demand immediate driver attention due to their semantic impact, even with low visual contrast. This dual criterion, i.e., bridging perception and contextual risk, re-defines saliency for autonomous and assisted driving systems. To address the lack of task-specific benchmarks, we collect the first large-scale TSOD dataset with pixel-wise saliency annotations, named TSOD10K. TSOD10K covers the diverse object categories in various real-world traffic scenes under various challenging weather/illumination variations (e.g., fog, snowstorms, low-contrast, and low-light). Methodologically, we propose a Mamba-based TSOD model, termed Tramba. Considering the challenge of distinguishing inconspicuous visual information from complex traffic backgrounds, Tramba introduces a novel Dual-Frequency Visual State Space module equipped with shifted window partitioning and dilated scanning to enhance the perception of fine details and global structure by hierarchically decomposing high/low-frequency components. To emphasize critical regions in traffic scenes, we propose a traffic-oriented Helix 2D-Selective-Scan (Helix-SS2D) mechanism that injects driving attention priors while effectively capturing global multi-direction spatial dependencies. We establish a comprehensive benchmark by evaluating Tramba and 25 existing NSI-SOD models on TSOD10K, demonstrating Tramba's superiority. Our research establishes the first foundation for safety-aware saliency analysis in intelligent transportation systems. The dataset and code will be made publicly available at https://github.com/mj129/Tramba. Jie Mei 0004, Lin Xiao 0002, Jing Xu 0008 |
IEEE Trans. Image Process. | 3 |
| 2024 | Towards semi-supervised multi-modal rectal cancer segmentation: A large-scale dataset and a multi-teacher uncertainty-aware network
Haotian Lu 0003, Jie Mei 0004, Sixu Bao, Jing Xu 0008 |
Expert Syst. Appl. | 3 |
| 2024 | Superpixel-wise contrast exploration for salient object detection
Jie Mei 0004, Jing Xu 0008 |
Knowl. Based Syst. | 2 |
| 2024 | Lightweight Cross-Modal Information Measure and Propagation for Road Extraction From Remote Sensing Image and Trajectory/LiDARabstractRecent studies have confirmed that GPS trajectory can effectively assist in achieving more accurate road extraction from remote sensing images. Therefore, lots of efforts focus on designing effective multi-modal fusion strategies for GPS trajectory and remote sensing image modalities. However, there are still some limitations,e.g., the fusion structures are complex and hinder further improvements. Moreover, the negative impact of redundant information in various modalities is commonly ignored. This paper aims to design a simple yet effective fusion strategy for GPS trajectory and remote sensing image modalities to address the above issues. Inspired by the network pruning algorithm, we design a Cross-Modal Information Propagation (CMIP) mechanism. CMIP utilizes the scaling factors and sparse constraint to distinguish the redundant information of a certain modality that is directly replaced with corresponding information of another modality. We improve the widely-usedL1sparse constraint and propose a novel information balanced constraint which is added on the scaling factors to better identify and prune redundant channels. Embedded in the CMIP mechanism, a multi-modal information propagation network (CMIPNet) is proposed, which can fully explore the complementarities between different modalities to accurately locate roads, especially roads with noise or incomplete information of a certain modality. Since the CMIP is parameter-free and self-adaptive, CMIPNet is lightweight and easy to deploy. The parameter number of CMIPNet can be comparable to single-modal models, which is about 1/3 of the existing multi-modal models. Extensive experiments are performed to demonstrate that CMIPNet outperforms the previous single- and multi-modal road extraction methods. Chenyu Lin, Jie Mei 0004, Haotian Lu 0003, Jing Xu 0008 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Deeply Hybrid Contrastive Learning Based on Semantic Pseudo-Label for Salient Object Detection in Optical Remote Sensing ImagesabstractSalient object detection in natural scene images (NSI-SOD) has undergone remarkable advancements in recent years. However, compared to those of natural images, the properties of remote sensing images (ORSIs), such as diverse spatial resolutions, complex background structures, and varying visual attributes of objects, are more complicated. Hence, how to explore the multiscale structural perceptual information of ORSIs to accurately detect salient objects is more challenging. In this paper, inspired by the superiority of contrastive learning, we propose a novel training paradigm for ORSI-SOD, named Deeply Hybrid Contrastive Learning Based on Semantic Pseudo-Label (DHCont), to force the network to extract rich structural perceptual information and further learn the better-structured feature embedding spaces. Specifically, DHCont first splits the ORSI into several local subregions composed of color- and texture-similar pixels, which act as semantic pseudo-labels. This strategy can effectively explore the underdeveloped semantic categories in ORSI-SOD. To delve deeper into multiscale structure-aware optimization, DHCont incorporates a hybrid contrast strategy that integrates “pixel-to-pixel”, “region-to-region”, “pixel-to-region”, and “region-to-pixel” contrasts at multiple scales. Additionally, to enhance the edge details of salient regions, we develop a hard edge contrast strategy that focuses on improving the detection accuracy of hard pixels near the object boundary. Moreover, we introduce a deep contrast algorithm that adds additional deep-level constraints to the feature spaces of multiple stages. Extensive experiments on two popular ORSI-SOD datasets demonstrate that simply integrating our DHCont into the existing ORSI-SOD models can significantly improve the performance. Jie Mei 0004, Jing Xu 0008 |
IEEE Trans. Multim. | 3 |
| 2023 | D2ANet: Difference-aware attention network for multi-level change detection from satellite imageryabstractRecognizing dynamic variations on the ground, especially changes caused by various natural disasters, is critical for assessing the severity of the damage and directing the disaster response. However, current workflows for disaster assessment usually require human analysts to observe and identify damaged buildings, which is labor-intensive and unsuitable for large-scale disaster areas. In this paper, we propose a difference-aware attention network (D2ANet) for simultaneous building localization and multi-level change detection from the dual-temporal satellite imagery. Considering the differences in different channels in the features of pre- and post-disaster images, we develop a dual-temporal aggregation module using paired features to excite change-sensitive channels of the features and learn the global change pattern. Since the nature of building damage caused by disasters is diverse in complex environments, we design a difference-attention module to exploit local correlations among the multi-level changes, which improves the ability to identify damage on different scales. Extensive experiments on the large-scale building damage assessment dataset xBD demonstrate that our approach provides new state-of-the-art results. Source code is publicly available at https://github.com/mj129/D2ANet . Jie Mei 0004, Yibo Zheng, Ming-Ming Cheng |
Comput. Vis. Media | 1 |
| 2022 | SANet: A Slice-Aware Network for Pulmonary Nodule DetectionabstractLung cancer is the most common cause of cancer death worldwide. A timely diagnosis of the pulmonary nodules makes it possible to detect lung cancer in the early stage, and thoracic computed tomography (CT) provides a convenient way to diagnose nodules. However, it is hard even for experienced doctors to distinguish them from the massive CT slices. The currently existing nodule datasets are limited in both scale and category, which is insufficient and greatly restricts its applications. In this paper, we collect the largest and most diverse dataset named PN9 for pulmonary nodule detection by far. Specifically, it contains 8,798 CT scans and 40,439 annotated nodules from 9 common classes. We further propose a slice-aware network (SANet) for pulmonary nodule detection. A slice grouped non-local (SGNL) module is developed to capture long-range dependencies among any positions and any channels of one slice group in the feature map. And we introduce a 3D region proposal network to generate pulmonary nodule candidates with high sensitivity, while this detection stage usually comes with many false positives. Subsequently, a false positive reduction module (FPR) is proposed by using the multi-scale feature maps. To verify the performance of SANet and the significance of PN9, we perform extensive experiments compared with several state-of-the-art 2D CNN-based and 3D CNN-based detection methods. Promising evaluation results on PN9 prove the effectiveness of our proposed SANet. The dataset and source code is available at https://mmcheng.net/SANet/. Jie Mei 0004, Ming-Ming Cheng, Lan-Ruo Wan |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | SLCRF: Subspace Learning With Conditional Random Field for Hyperspectral Image ClassificationabstractSubspace learning (SL) plays an essential role in hyperspectral image (HSI) classification since it can provide an effective solution to reduce the redundant information in the image pixels of HSIs. Previous works about SL aim to improve the accuracy of HSI recognition. Using a large number of labeled samples, related methods can train the parameters of the proposed solutions to obtain better representations of HSI pixels. However, the data instances may not be sufficient to learn a precise model for HSI classification in real applications. Moreover, it is well known that it takes much time, labor, and human expertise to label HSI images. To avoid the abovementioned problems, a novel SL method that includes the probability assumption called SL with the conditional random field (SLCRF) is developed. In SLCRF, the 3-D convolutional autoencoder (3DCAE) is first introduced to remove the redundant information in HSI pixels. Besides, the relationships are also constructed using spectral-spatial information among the adjacent pixels. Then, the conditional random field (CRF) framework can be constructed and further embedded into the HSI SL procedure with the semisupervised approach. Through the linearized alternating direction method termed LADMAP, the objective function of SLCRF is optimized using a defined iterative algorithm. The proposed method is comprehensively evaluated using the challenging public HSI data sets. We can achieve state-of-the-art performance using these HSI sets. Jie Mei 0004, Yuebin Wang, Liqiang Zhang 0001, Junhuan Peng, Bing Zhang 0001, Yibo Zheng |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | CoANet: Connectivity Attention Network for Road Extraction From Satellite ImageryabstractExtracting roads from satellite imagery is a promising approach to update the dynamic changes of road networks efficiently and timely. However, it is challenging due to the occlusions caused by other objects and the complex traffic environment, the pixel-based methods often generate fragmented roads and fail to predict topological correctness. In this paper, motivated by the road shapes and connections in the graph network, we propose a connectivity attention network (CoANet) to jointly learn the segmentation and pair-wise dependencies. Since the strip convolution is more aligned with the shape of roads, which are long-span, narrow, and distributed continuously. We develop a strip convolution module (SCM) that leverages four strip convolutions to capture long-range context information from different directions and avoid interference from irrelevant regions. Besides, considering the occlusions in road regions caused by buildings and trees, a connectivity attention module (CoA) is proposed to explore the relationship between neighboring pixels. The CoA module incorporates the graphical information and enables the connectivity of roads are better preserved. Extensive experiments on the popular benchmarks (SpaceNet and DeepGlobe datasets) demonstrate that our proposed CoANet establishes new state-of-the-art results. The source code will be made publicly available at: https://mmcheng.net/coanet/. Jie Mei 0004, Roujing Li, Wang Gao 0001, Ming-Ming Cheng |
IEEE Trans. Image Process. | 1 |
| 2021 | JCS: An Explainable COVID-19 Diagnosis System by Joint Classification and SegmentationabstractRecently, the coronavirus disease 2019 (COVID-19) has caused a pandemic disease in over 200 countries, influencing billions of humans. To control the infection, identifying and separating the infected people is the most crucial step. The main diagnostic tool is the Reverse Transcription Polymerase Chain Reaction (RT-PCR) test. Still, the sensitivity of the RT-PCR test is not high enough to effectively prevent the pandemic. The chest CT scan test provides a valuable complementary tool to the RT-PCR test, and it can identify the patients in the early-stage with high sensitivity. However, the chest CT scan test is usually time-consuming, requiring about 21.5 minutes per case. This paper develops a novel Joint Classification and Segmentation (JCS) system to perform real-time and explainable COVID- 19 chest CT diagnosis. To train our JCS system, we construct a large scale COVID- 19 Classification and Segmentation (COVID-CS) dataset, with 144,167 chest CT images of 400 COVID- 19 patients and 350 uninfected cases. 3,855 chest CT images of 200 patients are annotated with fine-grained pixel-level labels of opacifications, which are increased attenuation of the lung parenchyma. We also have annotated lesion counts, opacification areas, and locations and thus benefit various diagnosis aspects. Extensive experiments demonstrate that the proposed JCS diagnosis system is very efficient for COVID-19 classification and segmentation. It obtains an average sensitivity of 95.0% and a specificity of 93.0% on the classification test set, and 78.5% Dice score on the segmentation test set of our COVID-CS dataset. The COVID-CS dataset and code are available at https://github.com/yuhuan-wu/JCS. Yu-Huan Wu, Shanghua Gao, Jie Mei 0004, Jun Xu 0019, Deng-Ping Fan, Rongguo Zhang, Ming-Ming Cheng |
IEEE Trans. Image Process. | 3 |
| 2020 | Topology-Enhanced Urban Road Extraction via a Geographic Feature-Enhanced NetworkabstractUrban road extraction has wide applications in public transportation systems and unmanned vehicle navigation. The high-resolution remote sensing images contain background clutter and the roads have large appearance differences and complex connectivities, which makes it a very challenging task for road extraction. In this article, we propose a novel end-to-end deep learning model for road area extraction from remote sensing images. Road features are learned from three levels, which can remove the distraction of the background and enhance feature representation. A direction-aware attention block is introduced to the deep learning model for keeping road topologies. We compare our method on public remote sensing data sets with other related methods. The experimental results show the superiority of our method in terms of road extraction and connectivity preservation. Yuebin Wang, Liqiang Zhang 0001, Suhong Liu, Jie Mei 0004, Yang Li 0061 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2020 | Deep Learning for Multilabel Remote Sensing Image Annotation With Dual-Level Semantic ConceptsabstractMultilabel remote sensing (RS) image annotation is a challenging and time-consuming task that requires a considerable amount of expert knowledge. Most existing RS image annotation methods are based on handcrafted features and require multistage processes that are not sufficiently efficient and effective. An RS image can be assigned with a single label at the scene level to depict the overall understanding of the scene and with multiple labels at the object level to represent the major components. The multiple labels can be used as supervised information for annotation, whereas the single label can be used as additional information to exploit the scene-level similarity relationships. By exploiting the dual-level semantic concepts, we propose an end-to-end deep learning framework for object-level multilabel annotation of RS images. The proposed framework consists of a shared convolutional neural network for discriminative feature learning, a classification branch for multilabel annotation and an embedding branch for preserving the scene-level similarity relationships. In the classification branch, an attention mechanism is introduced to generate attention-aware features, and skip-layer connections are incorporated to combine information from multiple layers. The philosophy of the embedding branch is that images with the same scene-level semantic concepts should have similar visual representations. The proposed method adopts the binary cross-entropy loss for classification and the triplet loss for image embedding learning. The evaluations on three multilabel RS image data sets demonstrate the effectiveness and superiority of the proposed method in comparison with the state-of-the-art methods. Panpan Zhu, Yumin Tan, Liqiang Zhang 0001, Yuebin Wang, Jie Mei 0004 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2019 | PSASL: Pixel-Level and Superpixel-Level Aware Subspace Learning for Hyperspectral Image ClassificationabstractThe performance of hyperspectral image (HSI) classification relies on the pixel information obtained from hundreds of contiguous and narrow spectral bands. Existing approaches, however, are limited to exploit an appropriate latent subspace for data representation within the pixel-level or superpixel-level. To utilize spectral information and spatial correlation among pixels in HSI and avoid the “salt-and-pepper” problem generated in the pixel-based HSI classification, a novel pixel-level and superpixel-level aware subspace learning method called PSASL is developed. The PSASL constructs the subspace learning framework based on the reconstruction independent component analysis algorithm. The spectral–spatial graph regularization and label space regularization are developed as the pixel-level constraints. To avoid the “salt-and-pepper” problem generated in the pixel-based classification methods, superpixel-level constraints are introduced for integrating the data representations defined in the subspace and class probabilities of the pixels in the same superpixel. The subspace learning and the pixel-level regularization are combined with the superpixel-level regularization to form a unified objective function. The solution to the objective function is efficiently achieved by employing a customized iterative algorithm, and it converges very fast. A discriminative data representation and a universal multiclass classifier are learned simultaneously. We test the PSASL on three widely used HSI data sets. Experimental results demonstrate the superior performance of our method over many recently proposed methods in HSI classification. Jie Mei 0004, Yuebin Wang, Liqiang Zhang 0001, Bing Zhang 0001, Suhong Liu, Panpan Zhu, Yingchao Ren |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | Self-Supervised Feature Learning With CRF Embedding for Hyperspectral Image ClassificationabstractThe challenges in hyperspectral image (HSI) classification lie in the existence of noisy spectral information and lack of contextual information among pixels. Considering the three different levels in HSIs, i.e., subpixel, pixel, and superpixel, offer complementary information, we develop a novel HSI feature learning network (HSINet) to learn consistent features by self-supervision for HSI classification. HSINet contains a three-layer deep neural network and a multifeature convolutional neural network. It automatically extracts the features such as spatial, spectral, color, and boundary as well as context information. To boost the performance of self-supervised feature learning with the likelihood maximization, the conditional random field (CRF) framework is embedded into HSINet. The potential terms of unary, pairwise, and higher order in CRF are constructed by the corresponding subpixel, pixel, and superpixel. Furthermore, the feedback information derived from these terms are also fused into the different-level feature learning process, which makes the HSINet-CRF be a trainable end-to-end deep learning model with the back-propagation algorithm. Comprehensive evaluations are performed on three widely used HSI data sets and our method outperforms the state-of-the-art methods. Yuebin Wang, Jie Mei 0004, Liqiang Zhang 0001, Bing Zhang 0001, Panpan Zhu, Yang Li 0061 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2018 | Joint Margin, Cograph, and Label Constraints for Semisupervised Scene Parsing From Point CloudsabstractTo parse large-scale urban scenes using the supervised methods, a large amount of training data that can account for the vast visual and structural variance of urban environment is necessary. Unfortunately, such training data are mostly obtained by tedious and time-consuming manual work. To overcome the drawback, we propose a semisupervised learning framework that combines the margin, cograph, and label constraints into an objective function for point cloud parsing. Mathematically, the margin constraint is presented to learn a novel distance criterion that can effectively recognize points of different classes. The graph regularization is then employed to characterize the intrinsic geometry structure of the data manifold and explore relationships among points. The label consistency regularization is introduced to ensure the category consistency of the clustered points and single point. To classify the out-of-sample data, the framework successfully transforms the semisupervised classification results into the linear classifier by adopting a linear regression. An iterative algorithm is utilized to efficiently and effectively optimize the objective function with characteristics of multiple variables and highly nonlinear. The point clouds of four urban scenes are used to validate our method. The experimental results show that our method outperforms the state-of-the-art algorithms. Jie Mei 0004, Liqiang Zhang 0001, Yuebin Wang, Zidong Zhu, Huiqian Ding |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2018 | Self-Supervised Low-Rank Representation (SSLRR) for Hyperspectral Image ClassificationabstractLow-rank representation (LRR) can construct the relationships among pixels for hyperspectral image (HSI) classification with a given dictionary and a noise term. However, the accuracy of HSI classification based on LRR methods is degraded with the redundant and noise information existed in pixels. The neglect of semantic information around pixels in the LRR methods may cause “salt-and-pepper” problem in HSI classification. To avoid the aforementioned problems, a novel self-supervised low-rank representation method called SSLRR is developed. In SSLRR, the LRR and spectral–spatial graph regularization are developed as the pixel-level constraints to remove the redundant and noise information in HSIs. Superpixel constraints including data structure and relationship construction are further utilized to provide supervised feedback information to the subspace learning to avoid the “salt-and-pepper” problem generated in the pixel-based classification methods, and simultaneously enhance the performance of LRR. The pixel-level and superpixel-level regularizations are explicitly integrated into a unified objective function for LRR. By means of the linearized alternating direction method with adaptive penalty, the solution to the objective function is achieved by employing a customized iterative algorithm. We perform comprehensive evaluation of the proposed method on three challenging public HSI data sets. We obtain new state-of-the-art performance on these data sets, and achieve improvements of 44.3%, 13.4%, and 30.1% in overall accuracy compared to the best LRR method. Yuebin Wang, Jie Mei 0004, Liqiang Zhang 0001, Bing Zhang 0001, Anjian Li, Yibo Zheng, Panpan Zhu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2018 | LRAGE: Learning Latent Relationships With Adaptive Graph Embedding for Aerial Scene ClassificationabstractThe performance of scene classification relies heavily on the spatial and structural features that are extracted from high spatial resolution remote-sensing images. Existing approaches, however, are limited in adequately exploiting latent relationships between scene images. Aiming to decrease the distances between intraclass images and increase the distances between interclass images, we propose a latent relationship learning framework that integrates an adaptive graph with the constraints of the feature space and label propagation for high-resolution aerial image classification. To describe the latent relationships among scene images in the framework, we construct an adaptive graph that is embedded into the constrained joint space for features and labels. To remove redundant information and improve the computational efficiency, subspace learning is introduced to assist in the latent relationship learning. To address out-of-sample data, linear regression is adopted to project the semisupervised classification results onto a linear classifier. Learning efficiency is improved by minimizing the objective function via the linearized alternating direction method with an adaptive penalty. We test our method on three widely used aerial scene image data sets. The experimental results demonstrate the superior performance of our method over the state-of-the-art algorithms in aerial scene image classification. Yuebin Wang, Liqiang Zhang 0001, Xiaohua Tong, Feiping Nie 0001, Haiyang Huang 0001, Jie Mei 0004 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2017 | 3D tree modeling from incomplete point clouds via optimization and L1-MSTabstractReconstruction of 3D trees from incomplete point clouds is a challenging issue due to their large variety and natural geometric complexity. In this paper, we develop a novel method to effectively model trees from a single laser scan. First, coarse tree skeletons are extracted by utilizing the L1-median skeleton to compute the dominant direction of each point and the local point density of the point cloud. Then we propose a data completion scheme that guides the compensation for missing data. It is an iterative optimization process based on the dominant direction of each point and local point density. Finally, we present a L1-minimum spanning tree (MST) algorithm to refine tree skeletons from the optimized point cloud, which integrates the advantages of both L1-median skeleton and MST algorithms. The proposed method has been validated on various point clouds captured from single laser scans. The experiment results demonstrate the effectiveness and robustness of our method for coping with complex shapes of branching structures and occlusions. Jie Mei 0004, Liqiang Zhang 0001, Zhen Wang 0032, Liang Zhang 0023 |
Int. J. Geogr. Inf. Sci. | 1 |
| 2016 | A Three-Step Approach for TLS Point Cloud ClassificationabstractThe ability to classify urban objects in large urban scenes from point clouds efficiently and accurately still remains a challenging task today. A new methodology for the effective and accurate classification of terrestrial laser scanning (TLS) point clouds is presented in this paper. First, in order to efficiently obtain the complementary characteristics of each 3-D point, a set of point-based descriptors for recognizing urban point clouds is constructed. This includes the 3-D geometry captured using the spin-image descriptor computed on three different scales, the mean RGB colors of the point in the camera images, the LAB values of that mean RGB, and the normal at each 3-D point. The initial 3-D labeling of the categories in urban environments is generated by utilizing a linear support vector machine classifier on the descriptors. These initial classification results are then first globally optimized by the multilabel graph-cut approach. These results are further refined automatically by a local optimization approach based upon the object-oriented decision tree that uses weak priors among urban categories which significantly improves the final classification accuracy. The proposed method has been validated on three urban TLS point clouds, and the experimental results demonstrate that it outperforms the state-of-the-art method in classification accuracy for buildings, trees, pedestrians, and cars. Zhuqiang Li, Liqiang Zhang 0001, Xiaohua Tong, Bo Du 0001, Yuebin Wang, Liang Zhang 0023, Zhenxin Zhang, Jie Mei 0004, Xiaoyue Xing, P. Takis Mathiopoulos |
IEEE Trans. Geosci. Remote. Sens. | 9 |
| 2016 | A Local Structure and Direction-Aware Optimization Approach for Three-Dimensional Tree ModelingabstractModeling 3-D trees from terrestrial laser scanning (TLS) point clouds remains a challenging task for several well-known reasons, including their complex structure and severe occlusions. In order to accurately reconstruct 3-D tree models from TLS point clouds that typically suffer from significant occlusions, in this paper, a novel local structure and direction-aware approach is presented to successfully complete missing structures of trees. In this method, we first extract the coarse tree skeleton from the input point cloud, and thus, the branch dominant direction and the point density of each branch are obtained. By a skeleton-based Laplacian algorithm, the point cloud is further shrunk into a skeleton point cloud to highlight the branch dominant direction of each branch. For obtaining even more accurate point densities, a dictionary-based algorithm is utilized to learn and reconstruct the local structure. Finally, the branch dominant direction and point density are integrated into an iterative optimization process to recover the missing data. Extensive experimental results have shown that the proposed method is very robust to incomplete data sets, and it is capable of accurately reconstructing 3-D trees, which are partially, or even to a large extent, missing from the input point cloud. Zhen Wang 0032, Liqiang Zhang 0001, Tian Fang, Xiaohua Tong, P. Takis Mathiopoulos, Liang Zhang 0023, Jie Mei 0004 |
IEEE Trans. Geosci. Remote. Sens. | 7 |