EDBT 2026 Demo / reviewers in the wild / expert
Yuebin Wang
dblp:145/0744
· DBLP profile ↗
39ranked-venue papers
9as first author
20since 2021 · last 2026
0000-0002-6978-4558ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 35 · 6 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Weakly Supervised Semantic Segmentation of Remote Sensing Scenes With Cross-Image Class Token ConstraintsabstractWeakly supervised semantic segmentation (WSSS) based on image-level labels significantly reduces the labeling burden. However, current mainstream approaches optimise solely using single image information, neglecting the rich semantic correlation among images and struggling to dynamically suppress interfering information. When confronted with complex backgrounds and multi-category remote sensing (RS) images, intra-class consistency and inter-class discrimination pose significant challenges. To address these challenges, this paper proposes the cross-image class token constraints network (CICTC-Net). CICTC-Net establishes semantic correlations across multi-category RS images and implements two modules for targeted optimisation. Specifically, the cross-image token enhancement (CITE) module constructs intra-class token relationship graphs and applies cross-image consistency constraints to enhance semantic consistency among objects of the same category. The class-patch interaction refinement (CPIR) module constructs a directed graph of class-patch relationships and employs a neighbourhood selection mechanism to refine class tokens, thereby enhancing inter-class discriminability. Experiments on two RS datasets demonstrate that this approach significantly outperforms existing state-of-the-art solutions. Zhen Wang 0032, Junhuan Peng, Yuebin Wang, Yasong Mi, Dengxiang Wu |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2025 | SCIIENet: Shared and Complementary Information Interaction Enhancement Network for Self-Supervised Multimodal Remote Sensing Image ClassificationabstractEmploying multimodal remote sensing images (MRSIs) enhances ground object identification, yet the heterogeneity of MRSIs can increase the risk of model overfitting, particularly when training samples are limited. Self-supervised methods, such as masked autoencoder (MAE) reconstruction or contrastive learning, efficiently extract features from MRSIs with minimal reliance on labeled samples. However, current methods have not adequately considered both the shared and complementary features of MRSIs during the pretraining process, limiting feature representation and classification performance. To address these limitations, we propose a self-supervised pretraining framework called shared and complementary information interaction enhancement network (SCIIENet) for hyperspectral images (HSIs) and light detection and ranging (LiDAR)/synthetic aperture radar (SAR) data joint classification. Specifically, the framework incorporates an decoupling-interaction-enhancement approach through a triple-branch MAE architecture. The triple-branch architecture seperately captures the spectral features of HSIs that provide complementary information relative to LiDAR/SAR data, as well as the spatial features of HSIs and LiDAR/SAR data. We also introduce an interaction enhancement strategy that utilizes hierarchical contrastive learning (HCL) to bridge significant gaps in MRSI to emphasize the shared spatial information of MRSIs and incorporates a two-stage cross-modal fusion encoder (TCFE) to integrate the shared and complementary spatial information of different modalities. Moreover, we propose multi-stage skip connections (MSCs) for the reconstruction, which effectively preserve complementary spatial information. Extensive experiments on four datasets show that our method significantly improves classification accuracy compared to several state-of-the-art methods. Junhuan Peng, Yuebin Wang, Zhen Wang 0032, Huiwei Su, Yasong Mi |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | PMA²-Net: Progressive Fusion of Multiscale Axial Attention Network for Hyperspectral and Multispectral ImagesabstractIntegrating low-resolution hyperspectral images (LR-HSIs) with corresponding high-resolution multispectral images (HR-MSIs) for the reconstruction of HR-HSI using deep learning techniques represents a critical area of research. Although convolutional neural networks (CNNs) are widely used for HR-HSI reconstruction, their small receptive fields hinder effective global feature extraction, which limits their potential. Fusion methods that rely on traditional attention mechanisms also lack feature interaction, failing to integrate and harmonize feature information from hyperspectral image (HSI) and MSI effectively. To address these issues, this article develops a novel progressive fusion of multiscale axial attention network (PMA2-Net), which combines multiscale convolutions with axial attention (AA) and employs a progressive interaction approach to reconstruct high-resolution images. Specifically, PMA2-Net extracts the spatial and spectral information from HSI through spatial feature extraction (Spatial-FE) and spectral feature extraction (Spectral-FE). Concurrently, a feature injection module (FIM) is introduced, employing multiscale convolution to capture local features and integrating AA to enhance global feature association. Moreover, a progressive fusion module (PFM) is employed to enhance multidimensional feature collaboration and hierarchical integration. Extensive studies conducted using five significant HSI datasets verify the effectiveness of PMA2-Net, demonstrating its superior performance compared to current state-of-the-art (SOTA) fusion techniques. Shunhui Wang, Yuebin Wang, Danfeng Hong, Liqiang Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | DA&MTSS: An End-to-End Remote Sensing Image Domain Adaptive Semantic Segmentation Framework Combining Data Augmentation and Mobile Threshold Self-SupervisionabstractThe application of deep learning-based semantic segmentation in remote sensing (RS) images has achieved considerable success. However, many supervised methods still heavily rely on a large amount of labeled data, requiring time-consuming and labor-intensive manual annotations. Besides, networks trained on labeled source domain data often perform poorly in inference tasks with target domain data due to the domain shift phenomenon. To address these challenges, we construct a novel end-to-end unsupervised domain adaptation (UDA) framework, named data augmentation and mobile threshold self-supervision (DA&MTSS), which integrates data augmentation with self-supervision. Specifically, we analyze the common factors that cause domain shifts in RS images and adopt different data augmentation techniques to attenuate the domain shift and enhance the generalization ability and robustness of the network in cross-domain inference. In the self-supervision phase, we design a new sample-based mobile threshold method to dynamically control the thresholds of both dominant and long-tail classes during training and generate stable pseudo-labels. Therefore, our method eliminates the need for additional training or expert knowledge and achieves the co-evolution of network parameters and pseudo-label quality in the training process. The results of comprehensive experiments on five tasks across the ISPRS Vaihingen, Potsdam, and LoveDA datasets demonstrate that this method consistently achieves higher mIoU scores, showcasing the performance advantage of DA&MTSS in UDA for the semantic segmentation of RS image. Dong Chen 0009, Yuebin Wang, Liqiang Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Enhanced Local Feature Learning With Simple Offset Attention for Semantic Segmentation of Large-Scale Point CloudsabstractThe semantic segmentation network performance of large-scale outdoor point clouds is usually limited by the number of input point clouds. In the application of most methods, the point cloud is cut into small pieces as the training input, which will not only lead to heavy preprocessing burdens but also destroy the overall geometric structure of the scene. The transformer network has demonstrated remarkable advantages of attention mechanisms in focusing on crucial features and improving model performance. However, its training and inference on large-scale data are confined by computational complexity. To eliminate these challenges, the attention mechanisms are improved to enhance their performance in processing large-scale input data, while removing limitations imposed by computational complexity. Moreover, a novel local feature enhancement (LFE) module is developed to construct the LFE-Net, which can accurately and efficiently extract spatial and attribute features from local point clouds. In particular, an improved attention module, which is called simple offset attention (SOA), is adopted for local point cloud spatial feature learning. Compared with self-attention, SOA requires less memory and can better capture the fine-grained local features. Furthermore, to effectively avoid the destruction of object geometry and diminish the impact of sample imbalance, a training sample collection method based on the number of different classes is designed. To validate the effectiveness of this method, some experiments are conducted based on publicly accessible outdoor point cloud datasets. The results demonstrate that the LFE-Net can achieve substantial improvements compared with other cutting-edge network models. Dong Chen 0009, Yuebin Wang, Liqiang Zhang 0001, Zhizhong Kang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | GLR-CNN: CNN-Based Framework With Global Latent Relationship Embedding for High-Resolution Remote Sensing Image Scene ClassificationabstractHigh-resolution remote sensing image (HRSI) scene classification often faces challenges; for example, the intraclass similarity is low, but the interclass similarity is high due to complex backgrounds and variable scene scales. Convolutional neural networks (CNNs), the leading methods for HRSI scene classification, offer excellent performance. However, traditional CNNs require fixed-size inputs, which are a limitation when dealing with HRSI that represent large image domains, potentially degrading classification performance. To overcome these problems, we propose a CNN-based model named GLR-CNN in this article. First, to capitalize on the information from large-scale scenes adequately, VGG16 is utilized to extract the deep representative features, fine-tuned by the target HRSI of any size. Furthermore, a multilayer feature fusion block based on the channel–spatial attention algorithm is integrated into the CNN to capture more discriminative features from arbitrary-size images. Finally, to enhance the consistency between image features and similarities, a global latent relationship is used to measure the similarities among image features, then embed it into the fully connected layers (FCLs), and construct the latent relationship constraint. The model is optimized by the joint objective function including the latent relationship constraint and cross-entropy loss with label smoothing. Extensive experiments on three HRSI datasets obtained improvements of 3.93%, 6.5%, and 2.2% in overall accuracy compared to the finetuned VGG16 model, proving the effectiveness of the GLR-CNN method. Li Liu 0055, Yuebin Wang, Junhuan Peng, Liqiang Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Patch-Based Transformer Network Construction With Adaptive Feature-Interaction for Hyperspectral Image ClassificationabstractHyperspectral image classification (HSIC) is widely used in such fields as vegetation classification and fine agriculture. Nowadays, numerous classification models based on Transformer networks are inputting data typically in the form of patches. However, such patch-based input format may weaken the spatial relationships between neighbor pixels, thus limiting the ability to obtain contextual semantic information. Additionally, hyperspectral images (HSIs) generally contain hundreds of bands, which in turn include some redundant information. The original spectral vectors may lead to redundancy of information, thus reducing the classification accuracy remarkably. Therefore, in this article, a patch-based transformer network construction with adaptive feature-interaction (AFi) for HSIC called AFinet is developed for HSIC. As an end-to-end network, AFinet consists of a feature extraction module and an AFi module. For the developed model, the feed-forward neural network (FNN) in the standard Transformer framework suffers from a limited ability in exploiting local contexts. In this article, an expression-enhanced FNN (E2FNN) is introduced, which can incorporate depthwise convolution layers to capture local contextual information, so as to enhance the correlations between neighbor pixels. Moreover, this study designs an AFi module to facilitate the sharing of interaction features among Transformer-extracted features. Subsequently, global spatial feature information can be integrated into each spectral channel. The AFi module also includes a channel attention mechanism to help focus more attention on channels that contain critical information, while paying less attention to channels with little critical information. After the approach proposed in this study was tested on three widely used datasets, the experimental results demonstrate that the AFinet developed in this study can effectively improve classification accuracy. Yunbo Li, Yuebin Wang, Haiyan Gu, Liqiang Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | A Spatiotemporal Interpolation Graph Convolutional Network for Estimating PM₂.₅ Concentrations Based on Urban Functional ZonesabstractUrban functional zones (UFZs) contain abundant landscape information that can be adopted to better understand the surroundings. Various landscape compositions and configurations reflect different human activities, which may affect the particulate matter (PM2.5) concentrations. The very high-resolution (VHR) image features can reflect the physical and spatial structures of the UFZs. However, the existing PM2.5 estimation methods neither have been based on the scale of UFZs, nor have the VHR image features of UFZs as independent variables. Hence, this article proposes a spatiotemporal interpolation graph convolutional network (STI-GCN) model and introduces VHR image features to achieve PM2.5 estimation in UFZs. First, UFZs are split, and VHR image features are extracted by the visual geometry group 16 (VGG16). Subsequently, meteorological factors, aerosol optical depth (AOD), and VHR image features are used to estimate the PM2.5 concentrations at the scale of the UFZs. The two metropolises, Beijing and Shanghai, are chosen to assess the validity of the STI-GCN model. As for Beijing and Shanghai, the overall accuracy${R^{2}}$of the STI-GCN model can reach 0.96 and 0.89, the root-mean-square errors (RMSEs) are 8.15 and 6.40$\mu \text {g}/{\text {m}^{3}}$, the mean absolute errors (MAEs) are 5.51 and 4.78$\mu \text {g}/{\text {m}^{3}}$, and the relative prediction errors (RPEs) are 18.53% and 17.38%, respectively. Experiments show that the STI-GCN consistently outperforms other models. What’s more, the PM2.5 values are relatively high in commercial and official zones (COZs) and relatively low in urban green zones (UGZs). Xinya Chen, Yuebin Wang, Liqiang Zhang 0001, Zhiyu Yi, Hanchao Zhang, P. Takis Mathiopoulos |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Adaptive Context Transformer for Semisupervised Remote Sensing Image SegmentationabstractCurrent deep learning methods for semantic segmentation in remote sensing heavily depend on a substantial amount of labeled data. However, obtaining pixel-level labeled data in this field is both time-consuming and laborious. To address this challenge, semi-supervised learning methods have been introduced. Pseudo supervision is one of the most effective methods, which can be adopted to enhance the performance of semi-supervised semantic segmentation of remote sensing images [1]. But incorrect pseudo labels can cause substantially distortions to the segmentation model in semi-supervised learning. Moreover, it is difficult for conventional semantic segmentation methods to deal with global-local features of the remote sensing image without adaptive context feature. In this paper, we propose a novel learning approach based on an adaptive context transformer and pseudo labeling, called Adaptive Context Transformer for semi-supervised (ACTSS) remote sensing image segmentation. We propose an adaptive context attention model with adjustable sliding windows. A small window is used to capture Query (Q) for local feature and bigger windows are used to capture Key (K) and Value (V) for global feature. Then we combine them and get the global-local feature. And we propose a point-line-plane pseudo label filter (PLP) mechanism based on clustering and boundary extraction, which can filter unreliable pseudo labels from three angles: point, line and plane. To validate the effectiveness of the model, we carried out extensive experiments on the LOVEDA, Potsdam and Vaihingen datasets, and compared ACTSS with other methods. These experiments demonstrate that ACTSS achieves state-of-the-art performance for semi-supervised semantic segmentation on all tested datasets. Yunbo Li, Zhiyu Yi, Yuebin Wang, Liqiang Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | A Trajectory Inference Driven Matched Filtering Method for Maneuvering Target RefocusingabstractThe imaging of maneuvering targets may be defocused by the range cell and Doppler frequency migrations initiated by the targets’ complex motions in a long Coherent Processing Interval (CPI). To settle these problems, a generalized matched filtering method driven by the Bayesian motion trajectory inference (BMTI) is presented in this paper. Analyses show that any motion trajectory evolutions caused by targets’ movements can be described by the state equations derived from a polynomial prediction model. It is then demonstrated the uncertainties of the targets entering/leaving the radar beam coverage area will make the state equations into a state labeled multi-Bernoulli random finite set (LMBRFS). With the scale-transformation-based processing pipeline, a measurement random finite set (RFS) taking both the statistical properties of the maneuvering target with random appearance and disappearance and the noise/clutter into consideration is coined for the state LMBRFS. In this way, a state space model for describing the target instantaneous motions in a CPI is established. Once the target motion trajectories are inferred by a modified Bayesian filter, it is shown that the matched filter bank designed by the inferred trajectories can refocus the echo energies. Numerical simulation results verify the correctness of our analytical ones, while it is illustrated that our model and method are universal for indicating high-speed and maneuvering targets with any complex and/or even unknown motion forms in a CPI. Yuebin Wang, Dan Li 0004, Jian Qiu Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Revolutionizing Remote Sensing Image Analysis With BESSL-Net: A Boundary-Enhanced Semi-Supervised Learning NetworkabstractDeep learning (DL) become increasingly popular in remote sensing (RS) change detection (CD), leading to the development of massive networks that surpass traditional methods in accuracy and automation. However, the need for enormous amounts of annotated data remains a major concern and accurate boundary segmentation in RS images is challenging due to their complexity and heterogeneity. Moreover, properly aggregating the bi-temporal feature pairs and creating highly discriminative change features are crucial for detection performance. This article proposes a boundary-enhanced semi-supervised network (BESSL-Net) to tackle these issues for CD tasks. The network adopts dual encoders and one decoder architecture for segmentation and incorporates pseudo-labeling, contrastive learning, along with a teacher-student scheme to leverage unlabeled data. A boundary extraction module (BEM) is used to conduct boundary segmentation, while a change segmentation feature learning module (SFLM) is employed to create discriminative change features in both channel and spatial domains by integrating multi-level features. Three publicly available CD datasets are utilized to validate the proposed BESSL-Net. Compared to the current state-of-the-art networks, the semi-supervised network demonstrates advanced performance metrics, especially regarding Intersection over Union (IoUc) of change-class, showing improvements ranging from 1% to 9%. Zhiyu Yi, Yuebin Wang, Liqiang Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | A novel time-frequency model, analysis and parameter estimation approach: Towards multiple close and crossed chirp modes
Yuebin Wang, Dan Li 0004, Jian Qiu Zhang 0001 |
Signal Process. | 1 |
| 2022 | DSL-BC: Deep Subspace Learning With Boundary Consistency for Hyperspectral Image ClassificationabstractDeep subspace learning (DSL) plays an essential role in hyperspectral image classification, providing an effective solution tool to reduce the redundant information of hyperspectral image (HSI) pixels. Semi-supervised convolutional neural network (CNN)-based DSL methods can extract a more representative representation of latent subspace with the help of the labeled and unlabeled data. However, CNN-based DSL methods may lose the information of the class boundaries leading to misclassifications within regular input. We develop the deep subspace learning method with boundary consistency (DSL-BC) for the HSI classification to address this problem. In DSL-BC, the convolutional autoencoder (CAE) is first applied to extract the deep subspace representation (DSR). The DSR is used to model the boundary consistency. The graph convolutional network (GCN) is further adapted to enforce the boundary consistency by conducting the graph convolution on arbitrarily structured non-Euclidean data and irregular image regions. In addition, the adaptive entropy rate (ER) superpixel segmentation algorithm is applied to generate superpixels, and superpixel constraint is employed to improve the ability of DSL and GCN construction. DSL-BC integrates the DSL, the GCN, and the superpixel constraint into a unified objective function. A customized iterative algorithm is used to solve the objective function of the DSL-BC. The experimental results on three challenging public HSI datasets demonstrate that the DSL-BC can outperform the related state-of-the-art HSI classification methods. Yuebin Wang, Junhuan Peng, Liqiang Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | DRFL-VAT: Deep Representative Feature Learning With Virtual Adversarial Training for Semisupervised Classification of Hyperspectral ImageabstractWhile deep learning algorithms have achieved good results in hyperspectral image (HSI) classification, several supervised classification algorithms rely on a large number of labeled samples to get adequate performance. Collecting a large number of labeled samples is expensive in many real applications. To address this issue, a novel semisupervised HSI classification framework called deep representative feature learning (DRFL) with virtual adversarial training (DRFL-VAT) is developed in this article. By embedding the local manifold learning (LML) into the fully connected layers of a convolutional neural network (CNN), our newly developed DRFL can learn representative features. The VAT regularization is adopted to exploit the prediction label distribution of training samples and addresses the overfitting problem. Finally, the objective function of DRFL-VAT is solved by a customized algorithm. We test our method on three widely public HSI datasets and our results show that our method is competitive when compared to other state-of-the-art approaches. Yuebin Wang, Liqiang Zhang 0001, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Looking Outside the Window: Wide-Context Transformer for the Semantic Segmentation of High-Resolution Remote Sensing ImagesabstractLong-range contextual information is crucial for the semantic segmentation of high-resolution (HR) remote sensing images (RSIs). However, image cropping operations, commonly used for training neural networks, limit the perception of long-range contexts in large RSIs. To overcome this limitation, we propose a wide-context network (WiCoNet) for the semantic segmentation of HR RSIs. Apart from extracting local features with a conventional convolutional neural network (CNN), the WiCoNet has an extra context branch to aggregate information from a larger image area. Moreover, we introduce a context transformer to embed contextual information from the context branch and selectively project it onto the local features. The context transformer extends the vision transformer, an emerging kind of neural networks, to model the dual-branch semantic correlations. It overcomes the locality limitation of CNNs and enables the WiCoNet to see the bigger picture before segmenting the land-cover/land-use (LCLU) classes. Ablation studies and comparative experiments conducted on several benchmark datasets demonstrate the effectiveness of the proposed method. In addition, we present a new Beijing Land-Use (BLU) dataset. This is a large-scale HR satellite dataset with high-quality and fine-grained reference labels, which can facilitate future studies in this field. Lei Ding 0008, Dong Lin, Shaofu Lin, Jing Zhang 0023, Xiaojie Cui, Yuebin Wang, Hao Tang 0005, Lorenzo Bruzzone |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | DFLLR: Deep Feature Learning With Latent Relationship Embedding for Remote Sensing Image RetrievalabstractFor deep networks, accurate image similarities cannot be well characterized with limited iterations, so the latent relationships between images can be embedded to enhance image retrieval performance. In this article, we propose a method named DFLLR to learn deep image features and accurate image similarities for remote sensing image retrieval (RSIR) simultaneously. First, the AlexNet is employed to extract high-level semantic features. Second, to obtain accurate image similarities, latent relationships between images are constructed with manifold learning and embedded in the AlexNet model with fully connected layers; in this way, the latent relationships and image features can be jointly learned. Third, to boost the RSIR performance further, the constraints of central and margin for jointly learning latent relationships and image features are integrated into our DFLLR. The central constraint is used to reduce the discrepancy of the latent relationships at the intraclass level and enhance the accuracies of image features. Moreover, the margin constraint is designed to enhance the accuracies of the latent relationships by maximizing the manifold margin between the latent relationships at the intraclass and interclass levels. To validate our method, we perform comprehensive experiments on three publicly available remote sensing image datasets, and the results demonstrate that it significantly outperforms other state-of-the-art methods. Li Liu 0055, Yuebin Wang, Junhuan Peng, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | DGCC-EB: Deep Global Context Construction With an Enabled Boundary for Land Use Mapping of CSMAabstractLand use mapping (LUM) of a coal mining subsidence area (CMSA) is a significant task. The application of convolutional neural networks (CNNs) has become prevalent in LUM, which can achieve promising performances. However, CNNs cannot process irregular data; as a result, the boundary information is overlooked. The graph convolutional network (GCN) flexibly operates with irregular regions to capture the contextual relations among neighbors. However, the global context is not considered in the GCN. In this paper, we develop the deep global context construction with enabled boundary (DGCC-EB) for the LUM of the CMSA. An original Google Earth image is partitioned into nonoverlapping processing units. The DGCC-EB extracts preliminary features from the processing unit that are further divided into nonoverlapping superpixels with irregular edges. The superpixel features are generated and then embedded into the GCN and vision transformer (ViT). In the GCN, the graph convolution is applied to superpixel features; therefore, the boundary information of objects can be preserved. In the ViT, the multihead attention blocks and positional encoding build the global context among the superpixel features. The feature constraint is calculated to fuse the advantages of the features extracted from the GCN and ViT. To improve the LUM accuracy, the cross-entropy (CE) loss is calculated. The DGCC-EB integrates all modules into a whole end-to-end framework and is then optimized by a customized algorithm. The results of case studies show that the proposed DGCC-EB obtained acceptable OA (89.06%/88.68%) and Kappa (0.86/0.87) values for Shouzhou city and Zezhou city, respectively. Hanchao Zhang, Ning Zang, Yuebin Wang, Liqiang Zhang 0001, Bo Huang 0001, P. Takis Mathiopoulos |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | DS4L: Deep Semisupervised Shared Subspace Learning for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification is essential in remote sensing image analysis. The classification methods based on deep learning have attracted more and more attention. However, classification accuracy is seriously affected by the quantity of labeled data and redundant information. Therefore, a deep semisupervised shared subspace learning (DS4L) model is developed to overcome these problems in this article. DS4L is composed of two parts. First, the basic feature extraction (BFE) network is constructed to preliminary extract high-dimensional-space features of multiscale data and fusion them to one shared subspace. Then, a deep shared subspace learning (DSSL) network is proposed to obtain a deeper and more representative low-dimensional subspace. Moreover, to obtain a more representative subspace and alleviate dependence on labeled samples, the regular, irregular constraint, and cross-entropy (CE) loss are integrated into the model. The regular constraint is adopted to reconstruct the multiscale patches to ensure the quality of the subspace in an unsupervised manner. The irregular constraint can well embed labeled and unlabeled samples into the procedure of subspace learning (SL). Then, the CE loss is used to extract more discriminative subspace using the limited labeled samples. Finally, we perform experiments on three widely used HSI datasets. Compared with the basic SL model, the DS4L’s classification accuracy on the popular Salinas, Indian Pines, and PaviaU datasets are increased by 4.68%, 4.25%, and 1.57%, respectively. Li Liu 0055, Yuebin Wang, Liqiang Zhang 0001, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | SLCRF: Subspace Learning With Conditional Random Field for Hyperspectral Image ClassificationabstractSubspace learning (SL) plays an essential role in hyperspectral image (HSI) classification since it can provide an effective solution to reduce the redundant information in the image pixels of HSIs. Previous works about SL aim to improve the accuracy of HSI recognition. Using a large number of labeled samples, related methods can train the parameters of the proposed solutions to obtain better representations of HSI pixels. However, the data instances may not be sufficient to learn a precise model for HSI classification in real applications. Moreover, it is well known that it takes much time, labor, and human expertise to label HSI images. To avoid the abovementioned problems, a novel SL method that includes the probability assumption called SL with the conditional random field (SLCRF) is developed. In SLCRF, the 3-D convolutional autoencoder (3DCAE) is first introduced to remove the redundant information in HSI pixels. Besides, the relationships are also constructed using spectral-spatial information among the adjacent pixels. Then, the conditional random field (CRF) framework can be constructed and further embedded into the HSI SL procedure with the semisupervised approach. Through the linearized alternating direction method termed LADMAP, the objective function of SLCRF is optimized using a defined iterative algorithm. The proposed method is comprehensively evaluated using the challenging public HSI data sets. We can achieve state-of-the-art performance using these HSI sets. Jie Mei 0004, Yuebin Wang, Liqiang Zhang 0001, Junhuan Peng, Bing Zhang 0001, Yibo Zheng |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | SDFL-FC: Semisupervised Deep Feature Learning With Feature Consistency for Hyperspectral Image ClassificationabstractSemisupervised deep learning methods (DLMs) can mitigate the dependence on large amounts of labeled samples using a small number of labeled samples. However, for semisupervised deep feature learning (SDFL), the quality of extracted features cannot be well ensured without a certain amount of labeled samples. To address this issue, we develop the SDFL method with feature consistency (SDFL-FC) for the hyperspectral image (HSI) classification. The SDFL-FC first adopts the convolutional neural network (CNN) to extract spectral–spatial features of HSI and then uses the fully connected layers (FCLs) to model the feature consistency. Moreover, two constraints that enforce both the feature consistency of single pixel (FCS) and feature consistency of group pixels (FCG) are introduced to obtain the representative and discriminative features. The FCS is achieved by the generative adversarial network (GAN) regularization, which can reconstruct the original data from extracted features. The FCG is based on the assumption that the features of group pixels should have similar characteristics within a superpixel, which is embedded in each FCL. The final FCL outputs the class labels, and the cross-entropy (CE) loss is calculated with the labeled samples, while the two losses of FCS and FCG are calculated with all the training samples (both labeled and unlabeled). SDFL-FC integrates the FCS, FCG, and CE loss into a unified objective function and uses a customized iterative optimization algorithm to optimize it. Experiments demonstrate that the SDFL-FC can outperform the related state-of-the-art HSI classification methods. Yuebin Wang, Junhuan Peng, Chunping Qiu, Lei Ding 0008, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | Combined Nonlocal Spatial Information and Spatial Group Sparsity in NMF for Hyperspectral UnmixingabstractUnmixing is a key but difficult issue in hyperspectral image (HSI) processing, and many unmixing methods have been proposed. However, an effective introduction of the spatial context in unmixing remains a challenge but is a necessary condition for many real scene applications. In this letter, a new nonnegative matrix factorization (NMF) method that combines nonlocal spatial information with spatial group sparsity (NLNMF) is proposed. Each superpixel generated by the simple linear iterative clustering (SLIC) segmentation method was used as a group. The search region of the nonlocal means method was adaptively set using a superpixel label from each spectrum to find the similar spectra to reestimate the reference spectrum. Additionally, the sparsity of spectra in the same superpixel was considered to be the same. Experiment results for synthetic and real HSI showed that the proposed method not only can more accurately estimate the endmember and abundance compared with other unmixing methods but also has good performance regarding antinoise. Longshan Yang, Junhuan Peng, Huiwei Su, Linlin Xu, Yuebin Wang |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2020 | DML-GANR: Deep Metric Learning With Generative Adversarial Network Regularization for High Spatial Resolution Remote Sensing Image RetrievalabstractWith a small number of labeled samples for training, it can save considerable manpower and material resources, especially when the amount of high spatial resolution remote sensing images (HSR-RSIs) increases considerably. However, many deep models face the problem of overfitting when using a small number of labeled samples. This might degrade HSR-RSI retrieval accuracy. Aiming at obtaining more accurate HSR-RSI retrieval performance with small training samples, we develop a deep metric learning approach with generative adversarial network regularization (DML-GANR) for HSR-RSI retrieval. The DML-GANR starts from a high-level feature extraction (HFE) to extract high-level features, which includes convolutional layers and fully connected (FC) layers. Each of the FC layers is constructed by deep metric learning (DML) to maximize the interclass variations and minimize the intraclass variations. The generative adversarial network (GAN) is adopted to mitigate the overfitting problem and validate the qualities of extracted high-level features. DML-GANR is optimized through a customized approach, and the optimal parameters are obtained. The experimental results on the three data sets demonstrate the superior performance of DML-GANR over state-of-the-art techniques in HSR-RSI retrieval. Yuebin Wang, Junhuan Peng, Liqiang Zhang 0001, Linlin Xu, Kai Yan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | Topology-Enhanced Urban Road Extraction via a Geographic Feature-Enhanced NetworkabstractUrban road extraction has wide applications in public transportation systems and unmanned vehicle navigation. The high-resolution remote sensing images contain background clutter and the roads have large appearance differences and complex connectivities, which makes it a very challenging task for road extraction. In this article, we propose a novel end-to-end deep learning model for road area extraction from remote sensing images. Road features are learned from three levels, which can remove the distraction of the background and enhance feature representation. A direction-aware attention block is introduced to the deep learning model for keeping road topologies. We compare our method on public remote sensing data sets with other related methods. The experimental results show the superiority of our method in terms of road extraction and connectivity preservation. Yuebin Wang, Liqiang Zhang 0001, Suhong Liu, Jie Mei 0004, Yang Li 0061 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | Latent Relationship Guided Stacked Sparse Autoencoder for Hyperspectral Imagery ClassificationabstractClassification is an important application of hyperspectral image (HSI). However, it is also a challenging research topic due to the spatial variability of spectral signature and limited training samples. To address these problems, a novel unsupervised feature learning method called latent relationship guided the stacked sparse autoencoder (LRSSAE) is developed in this article, which can effectively exploit the latent relationship under feature space to improve the ability of feature learning. Moreover, the superpixels constraint is employed on the feature representation to avoid the “salt-and-pepper” problem, and it is enforced on the latent relationship to enhance the latent relationship learning additionally. In LRSSAE, combining the stacked sparse autoencoder (SSAE) with the graph regularizations of latent relationship in each hidden layer and the superpixel constraints in the top layer, we extract feature representation in an unsupervised manner. And then, we present a customized iterative algorithm to optimize the LRSSAE. We evaluate the proposed method on three widely used HSI data sets comprehensively. The results demonstrate that our method achieves promising classification performance on these data sets and obtains improvements of 5.06%, 5.77%, and 2.11% in overall accuracy compared to the best SSAE method. Li Liu 0055, Yuebin Wang, Junhuan Peng, Liqiang Zhang 0001, Bing Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | Supervised High-Level Feature Learning With Label Consistencies for Object RecognitionabstractDue to the large intraclass variances and complicated object distribution, recognizing objects with complex appearances and arbitrary orientations has been an active research topic and a challenging task in remote sensing fields. In this article, we formulate object recognition as a high-level feature-learning problem, and a novel supervised method is proposed to learn high-level feature representations from high-resolution remote sensing images for object recognition. Our method simultaneously and coherently achieves high-level feature learning and classifier training, which improves the recognition performance. Two constraints that enforce the label consistencies of group images and label consistencies of single images are introduced in a deep learning framework to obtain the high-level feature space. The high-level feature and a multiclass linear classifier are finally learned by an effective optimization algorithm. Experimental results demonstrate the superior performance of the proposed method over many state-of-the-art techniques in object recognition. Yuebin Wang, Honglei Yang, Liqiang Zhang 0001, Suhong Liu, P. Takis Mathiopoulos |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | Deep Learning for Multilabel Remote Sensing Image Annotation With Dual-Level Semantic ConceptsabstractMultilabel remote sensing (RS) image annotation is a challenging and time-consuming task that requires a considerable amount of expert knowledge. Most existing RS image annotation methods are based on handcrafted features and require multistage processes that are not sufficiently efficient and effective. An RS image can be assigned with a single label at the scene level to depict the overall understanding of the scene and with multiple labels at the object level to represent the major components. The multiple labels can be used as supervised information for annotation, whereas the single label can be used as additional information to exploit the scene-level similarity relationships. By exploiting the dual-level semantic concepts, we propose an end-to-end deep learning framework for object-level multilabel annotation of RS images. The proposed framework consists of a shared convolutional neural network for discriminative feature learning, a classification branch for multilabel annotation and an embedding branch for preserving the scene-level similarity relationships. In the classification branch, an attention mechanism is introduced to generate attention-aware features, and skip-layer connections are incorporated to combine information from multiple layers. The philosophy of the embedding branch is that images with the same scene-level semantic concepts should have similar visual representations. The proposed method adopts the binary cross-entropy loss for classification and the triplet loss for image embedding learning. The evaluations on three multilabel RS image data sets demonstrate the effectiveness and superiority of the proposed method in comparison with the state-of-the-art methods. Panpan Zhu, Yumin Tan, Liqiang Zhang 0001, Yuebin Wang, Jie Mei 0004 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2020 | WeGAN: Deep Image Hashing With Weighted Generative Adversarial NetworksabstractImage hashing has been widely used in image retrieval tasks. Many existing methods generate hashing codes based on image feature representations. They rarely consider the rich information such as image clustering information contained in the image set as well as uncertain relationships between images and tags simultaneously. In this paper, we develop a Weighted Generative Adversarial Networks (WeGAN) to transfer the clustering information of images to construct the hashing code. WeGAN consists three modules: 1) a hashing learning process for transferring knowledge of the image set to hashing codes of single images; 2) by means of hashing codes, a module to generate image content, tag representation, and their joint information which reflects the correlation between the image and the corresponding tags; 3) a discriminator to distinguish the generated data from the original source, and then formulating three loss functions. Different weights are assigned to these loss functions in order to deal with the uncertainties between images and tags. Through introducing the image set to process the image hashing with different tags, WeGAN can naturally provide the information of clustering results, which is useful for image hashing with multi-tags. The generated hashing code has the ability to dynamically process the uncertain relationships between images and tags. Experiments on three challenging datasets show that WeGAN outperforms the state-of-the-art methods. Yuebin Wang, Liqiang Zhang 0001, Feiping Nie 0001 |
IEEE Trans. Multim. | 1 |
| 2019 | PSASL: Pixel-Level and Superpixel-Level Aware Subspace Learning for Hyperspectral Image ClassificationabstractThe performance of hyperspectral image (HSI) classification relies on the pixel information obtained from hundreds of contiguous and narrow spectral bands. Existing approaches, however, are limited to exploit an appropriate latent subspace for data representation within the pixel-level or superpixel-level. To utilize spectral information and spatial correlation among pixels in HSI and avoid the “salt-and-pepper” problem generated in the pixel-based HSI classification, a novel pixel-level and superpixel-level aware subspace learning method called PSASL is developed. The PSASL constructs the subspace learning framework based on the reconstruction independent component analysis algorithm. The spectral–spatial graph regularization and label space regularization are developed as the pixel-level constraints. To avoid the “salt-and-pepper” problem generated in the pixel-based classification methods, superpixel-level constraints are introduced for integrating the data representations defined in the subspace and class probabilities of the pixels in the same superpixel. The subspace learning and the pixel-level regularization are combined with the superpixel-level regularization to form a unified objective function. The solution to the objective function is efficiently achieved by employing a customized iterative algorithm, and it converges very fast. A discriminative data representation and a universal multiclass classifier are learned simultaneously. We test the PSASL on three widely used HSI data sets. Experimental results demonstrate the superior performance of our method over many recently proposed methods in HSI classification. Jie Mei 0004, Yuebin Wang, Liqiang Zhang 0001, Bing Zhang 0001, Suhong Liu, Panpan Zhu, Yingchao Ren |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | Self-Supervised Feature Learning With CRF Embedding for Hyperspectral Image ClassificationabstractThe challenges in hyperspectral image (HSI) classification lie in the existence of noisy spectral information and lack of contextual information among pixels. Considering the three different levels in HSIs, i.e., subpixel, pixel, and superpixel, offer complementary information, we develop a novel HSI feature learning network (HSINet) to learn consistent features by self-supervision for HSI classification. HSINet contains a three-layer deep neural network and a multifeature convolutional neural network. It automatically extracts the features such as spatial, spectral, color, and boundary as well as context information. To boost the performance of self-supervised feature learning with the likelihood maximization, the conditional random field (CRF) framework is embedded into HSINet. The potential terms of unary, pairwise, and higher order in CRF are constructed by the corresponding subpixel, pixel, and superpixel. Furthermore, the feedback information derived from these terms are also fused into the different-level feature learning process, which makes the HSINet-CRF be a trainable end-to-end deep learning model with the back-propagation algorithm. Comprehensive evaluations are performed on three widely used HSI data sets and our method outperforms the state-of-the-art methods. Yuebin Wang, Jie Mei 0004, Liqiang Zhang 0001, Bing Zhang 0001, Panpan Zhu, Yang Li 0061 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2018 | Joint Margin, Cograph, and Label Constraints for Semisupervised Scene Parsing From Point CloudsabstractTo parse large-scale urban scenes using the supervised methods, a large amount of training data that can account for the vast visual and structural variance of urban environment is necessary. Unfortunately, such training data are mostly obtained by tedious and time-consuming manual work. To overcome the drawback, we propose a semisupervised learning framework that combines the margin, cograph, and label constraints into an objective function for point cloud parsing. Mathematically, the margin constraint is presented to learn a novel distance criterion that can effectively recognize points of different classes. The graph regularization is then employed to characterize the intrinsic geometry structure of the data manifold and explore relationships among points. The label consistency regularization is introduced to ensure the category consistency of the clustered points and single point. To classify the out-of-sample data, the framework successfully transforms the semisupervised classification results into the linear classifier by adopting a linear regression. An iterative algorithm is utilized to efficiently and effectively optimize the objective function with characteristics of multiple variables and highly nonlinear. The point clouds of four urban scenes are used to validate our method. The experimental results show that our method outperforms the state-of-the-art algorithms. Jie Mei 0004, Liqiang Zhang 0001, Yuebin Wang, Zidong Zhu, Huiqian Ding |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2018 | Self-Supervised Low-Rank Representation (SSLRR) for Hyperspectral Image ClassificationabstractLow-rank representation (LRR) can construct the relationships among pixels for hyperspectral image (HSI) classification with a given dictionary and a noise term. However, the accuracy of HSI classification based on LRR methods is degraded with the redundant and noise information existed in pixels. The neglect of semantic information around pixels in the LRR methods may cause “salt-and-pepper” problem in HSI classification. To avoid the aforementioned problems, a novel self-supervised low-rank representation method called SSLRR is developed. In SSLRR, the LRR and spectral–spatial graph regularization are developed as the pixel-level constraints to remove the redundant and noise information in HSIs. Superpixel constraints including data structure and relationship construction are further utilized to provide supervised feedback information to the subspace learning to avoid the “salt-and-pepper” problem generated in the pixel-based classification methods, and simultaneously enhance the performance of LRR. The pixel-level and superpixel-level regularizations are explicitly integrated into a unified objective function for LRR. By means of the linearized alternating direction method with adaptive penalty, the solution to the objective function is achieved by employing a customized iterative algorithm. We perform comprehensive evaluation of the proposed method on three challenging public HSI data sets. We obtain new state-of-the-art performance on these data sets, and achieve improvements of 44.3%, 13.4%, and 30.1% in overall accuracy compared to the best LRR method. Yuebin Wang, Jie Mei 0004, Liqiang Zhang 0001, Bing Zhang 0001, Anjian Li, Yibo Zheng, Panpan Zhu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2018 | LRAGE: Learning Latent Relationships With Adaptive Graph Embedding for Aerial Scene ClassificationabstractThe performance of scene classification relies heavily on the spatial and structural features that are extracted from high spatial resolution remote-sensing images. Existing approaches, however, are limited in adequately exploiting latent relationships between scene images. Aiming to decrease the distances between intraclass images and increase the distances between interclass images, we propose a latent relationship learning framework that integrates an adaptive graph with the constraints of the feature space and label propagation for high-resolution aerial image classification. To describe the latent relationships among scene images in the framework, we construct an adaptive graph that is embedded into the constrained joint space for features and labels. To remove redundant information and improve the computational efficiency, subspace learning is introduced to assist in the latent relationship learning. To address out-of-sample data, linear regression is adopted to project the semisupervised classification results onto a linear classifier. Learning efficiency is improved by minimizing the objective function via the linearized alternating direction method with an adaptive penalty. We test our method on three widely used aerial scene image data sets. The experimental results demonstrate the superior performance of our method over the state-of-the-art algorithms in aerial scene image classification. Yuebin Wang, Liqiang Zhang 0001, Xiaohua Tong, Feiping Nie 0001, Haiyang Huang 0001, Jie Mei 0004 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2017 | 3DCNN-DQN-RNN: A Deep Reinforcement Learning Framework for Semantic Parsing of Large-Scale 3D Point CloudsabstractSemantic parsing of large-scale 3D point clouds is an important research topic in computer vision and remote sensing fields. Most existing approaches utilize hand-crafted features for each modality independently and combine them in a heuristic manner. They often fail to consider the consistency and complementary information among features adequately, which makes them difficult to capture high-level semantic structures. The features learned by most of the current deep learning methods can obtain high-quality image classification results. However, these methods are hard to be applied to recognize 3D point clouds due to unorganized distribution and various point density of data. In this paper, we propose a 3DCNN-DQN-RNN method which fuses the 3D convolutional neural network (CNN), Deep Q-Network (DQN) and Residual recurrent neural network (RNN)for an efficient semantic parsing of large-scale 3D point clouds. In our method, an eye window under control of the 3D CNN and DQN can localize and segment the points of the object's class efficiently. The 3D CNN and Residual RNN further extract robust and discriminative features of the points in the eye window, and thus greatly enhance the parsing accuracy of large-scale point clouds. Our method provides an automatic process that maps the raw data to the classification results. It also integrates object localization, segmentation and classification into one framework. Experimental results demonstrate that the proposed method outperforms the state-of-the-art point cloud classification methods. Fangyu Liu 0001, Shuaipeng Li, Liqiang Zhang 0001, Chenghu Zhou, Rongtian Ye, Yuebin Wang, Jiwen Lu |
ICCV | 6 |
| 2017 | A feature extraction and similarity metric-learning framework for urban model retrievalabstractUrban model retrieval has wide applications in the geoscience field, and it is also a very challenging research topic due to the blur and background clutter in query images and the large spatial inconsistencies between query and database images. In this study, a feature extraction and similarity metric-learning framework for urban model retrieval is proposed. In the method, the selective search voting algorithm is presented to automatically localize and segment a query object from an input image with the help of the top-ranked retrieved database images. Then, the local features of object images are extracted via sparse coding, and the global features are learned using the spatial constrained convolutional neural network. We utilize a new similarity metric to match the database images with a query object image. Finally, similar 3D models are retrieved. Both qualitative and quantitative experimental results indicate that the proposed framework can localize and segment a query object from an input image precisely and that the retrieval results are better than those of other related approaches. Yuebin Wang, Liqiang Zhang 0001, Xiaohua Tong, Suhong Liu, Tian Fang |
Int. J. Geogr. Inf. Sci. | 1 |
| 2017 | Learning a Discriminative Distance Metric With Label Consistency for Scene ClassificationabstractTo achieve high scene classification performance of high spatial resolution remote sensing images (HSR-RSIs), it is important to learn a discriminative space in which the distance metric can precisely measure both similarity and dissimilarity of features and labels between images. While the traditional metric learning methods focus on preserving interclass separability, label consistency (LC) is less involved, and this might degrade scene images classification accuracy. Aiming at considering intraclass compactness in HSR-RSIs, we propose a discriminative distance metric learning method with LC (DDML-LC). The DDML-LC starts from the dense scale invariant feature transformation features extracted from HSR-RSIs, and then uses spatial pyramid maximum pooling with sparse coding to encode the features. In the learning process, the intraclass compactness and interclass separability are enforced while the global and local LC after the feature transformation is constrained, leading to a joint optimization of feature manifold, distance metric, and label distribution. The learned metric space can scale to discriminate out-of-sample HSR-RSIs that do not appear in the metric learning process. Experimental results on three data sets demonstrate the superior performance of the DDML-LC over state-of-the-art techniques in HSR-RSI classification. Yuebin Wang, Liqiang Zhang 0001, Hao Deng 0004, Jiwen Lu, Haiyang Huang 0001, Liang Zhang 0023, Jun Liu 0029, Xiaoyue Xing |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2016 | A Three-Step Approach for TLS Point Cloud ClassificationabstractThe ability to classify urban objects in large urban scenes from point clouds efficiently and accurately still remains a challenging task today. A new methodology for the effective and accurate classification of terrestrial laser scanning (TLS) point clouds is presented in this paper. First, in order to efficiently obtain the complementary characteristics of each 3-D point, a set of point-based descriptors for recognizing urban point clouds is constructed. This includes the 3-D geometry captured using the spin-image descriptor computed on three different scales, the mean RGB colors of the point in the camera images, the LAB values of that mean RGB, and the normal at each 3-D point. The initial 3-D labeling of the categories in urban environments is generated by utilizing a linear support vector machine classifier on the descriptors. These initial classification results are then first globally optimized by the multilabel graph-cut approach. These results are further refined automatically by a local optimization approach based upon the object-oriented decision tree that uses weak priors among urban categories which significantly improves the final classification accuracy. The proposed method has been validated on three urban TLS point clouds, and the experimental results demonstrate that it outperforms the state-of-the-art method in classification accuracy for buildings, trees, pedestrians, and cars. Zhuqiang Li, Liqiang Zhang 0001, Xiaohua Tong, Bo Du 0001, Yuebin Wang, Liang Zhang 0023, Zhenxin Zhang, Jie Mei 0004, Xiaoyue Xing, P. Takis Mathiopoulos |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2016 | A Three-Layered Graph-Based Learning Approach for Remote Sensing Image RetrievalabstractWith the emergence of huge volumes of high-resolution remote sensing images produced by all sorts of satellites and airborne sensors, processing and analysis of these images require effective retrieval techniques. To alleviate the dramatic variation of the retrieval accuracy among queries caused by the single image feature algorithms, we developed a novel graph-based learning method for effectively retrieving remote sensing images. The method utilizes a three-layer framework that integrates the strengths of query expansion and fusion of holistic and local features. In the first layer, two retrieval image sets are obtained by, respectively, using the retrieval methods based on holistic and local features, and the top-ranked and common images from both of the top candidate lists subsequently form graph anchors. In the second layer, the graph anchors as an expansion query retrieve six image sets from the image database using each individual feature. In the third layer, the images in the six image sets are evaluated for generating positive and negative data, and SimpleMKL is applied to learn suitable query-dependent fusion weights for achieving the final image retrieval result. Extensive experiments were performed on the UC Merced Land Use-Land Cover data set. The source code has been available at our website. Compared with other related methods, the retrieval precision is significantly enhanced without sacrificing the scalability of our approach. Yuebin Wang, Liqiang Zhang 0001, Xiaohua Tong, Liang Zhang 0023, Zhenxin Zhang, Xiaoyue Xing, P. Takis Mathiopoulos |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2016 | A Multilevel Point-Cluster-Based Discriminative Feature for ALS Point Cloud ClassificationabstractPoint cloud classification plays a critical role in point cloud processing and analysis. Accurately classifying objects on the ground in urban environments from airborne laser scanning (ALS) point clouds is a challenge because of their large variety, complex geometries, and visual appearances. In this paper, a novel framework is presented for effectively extracting the shape features of objects from an ALS point cloud, and then, it is used to classify large and small objects in a point cloud. In the framework, the point cloud is split into hierarchical clusters of different sizes based on a natural exponential function threshold. Then, to take advantage of hierarchical point cluster correlations, latent Dirichlet allocation and sparse coding are jointly performed to extract and encode the shape features of the multilevel point clusters. The features at different levels are used to capture information on the shapes of objects of different sizes. This way, robust and discriminative shape features of the objects can be identified, and thus, the precision of the classification is significantly improved, particularly for small objects. Zhenxin Zhang, Liqiang Zhang 0001, Xiaohua Tong, P. Takis Mathiopoulos, Zhen Wang 0032, Yuebin Wang |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2014 | A Structure-Aware Global Optimization Method for Reconstructing 3-D Tree Models From Terrestrial Laser Scanning DataabstractA 3-D tree structure plays an important role in many scientific fields, including forestry and agriculture. For example, terrestrial laser scanning (TLS) can efficiently capture high-precision 3-D spatial arrangements and structure of trees as a point cloud. In the past, several methods to reconstruct 3-D trees from the TLS point cloud were proposed. However, in general, they fail to process incomplete TLS data. To address such incomplete TLS data sets, a new method that is based on a structure-aware global optimization approach (SAGO) is proposed. The SAGO first obtains the approximate tree skeleton from a distance minimum spanning tree (DMst) and then defines the stretching directions of the branches on the tree skeleton. Based on these stretching directions, the SAGO recovers missing data in the incomplete TLS point cloud. The DMst is applied again to obtain the refined tree skeleton from the optimized data, and the tree skeleton is smoothed by employing a Laplacian function. To reconstruct 3-D tree models, the radius of each branch section is estimated, and leaves are added to form the crown geometry. The developed methodology has been extensively evaluated by employing a dozen TLS point clouds of various types of trees. Both qualitative and quantitative performance evaluation results have indicated that the SAGO is capable of effectively reconstructing 3-D tree models from grossly incomplete TLS point clouds with significant amounts of missing data. Zhen Wang 0032, Liqiang Zhang 0001, Tian Fang, P. Takis Mathiopoulos, Huamin Qu, Dong Chen 0009, Yuebin Wang |
IEEE Trans. Geosci. Remote. Sens. | 7 |