EDBT 2026 Demo / reviewers in the wild / expert
Xiangtao Zheng
dblp:160/1446
· DBLP profile ↗
54ranked-venue papers
16as first author
32since 2021 · last 2025
0000-0002-8398-6324ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 33 · 9 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 10 · 3 first-author · 4 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Context-Aware Local-Global Semantic Alignment for Remote Sensing Image-Text RetrievalabstractRemote sensing image-text retrieval (RSITR) is a cross-modal task that integrates visual and textual information, attracting significant attention in remote sensing research. Remote sensing images typically contain complex scenes with abundant details, presenting significant challenges for accurate semantic alignment between images and texts. Despite advances in the field, achieving precise alignment in such intricate contexts remains a major hurdle. To address this challenge, this article introduces a novel context-aware local-global semantic alignment (CLGSA) method. The proposed method consists of two key modules: the local key feature alignment (LKFA) module and the cross-sample global semantic alignment (CGSA) module. The LKFA module incorporates a local image masking and reconstruction task to improve the alignment between image and text features. Specifically, this module masks certain regions of the image and uses text context information to guide the reconstruction of the masked areas, enhancing the alignment of local semantics and ensuring more accurate retrieval of region-specific content. The CGSA module employs a hard sample triplet loss to improve global semantic consistency. By prioritizing difficult samples during training, this module refines feature space distributions, helping the model better capture global semantics across the entire image-text pair. A series of extensive experiments demonstrates the effectiveness of the proposed method. The method achieves an mR score of 32.07% on the RSICD dataset and 46.63% on the RSITMD dataset, outperforming baseline methods and confirming the robustness and accuracy of the approach. Xiumei Chen, Xiangtao Zheng, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Relevance-Guided Adaptive Learning for Remote Sensing Image-Text RetrievalabstractThe remote sensing image–text retrieval (RSITR) aims to establish semantic alignment between images and texts to enable accurate cross-modal retrieval. Existing methods usually extract features from images and texts independently, aligning them in a shared embedding space to achieve cross-modal retrieval. However, these methods often assume complete alignment between image and text pairs, overlooking the inherent disparities between the rich visual details in remote sensing (RS) images and the abstract nature of textual descriptions. These disparities result in image–text pairs only sharing partial semantic correlations, rather than one-to-one complete alignment. Such incomplete alignment adversely affects model training and retrieval accuracy. To address this problem, a relevance-guided adaptive learning (RGAL) method is proposed, which quantifies and leverages the relevance of image–text pairs to refine the training process while enhancing retrieval performance. First, the proposed method introduces an image–text relevance measurement mechanism that integrates global and local feature distances to accurately evaluate the degree of semantic relevance between images and texts. Second, a relevance-based sample division (RBSD) strategy is proposed, utilizing a Gaussian mixture model to dynamically redivide samples into positive and negative pairs according to the measured image–text relevance. This strategy refines the training dataset, reduces noise, and enhances the effectiveness of model learning. Finally, a relevance-weighted triplet loss (RWTL) is designed to adaptively adjust the contribution of sample pairs to the loss function based on their relevance, further optimizing model training and enhancing retrieval accuracy. Experimental results on multiple RSITR datasets demonstrate that the proposed method significantly improves retrieval accuracy and performance. Xiumei Chen, Xiangtao Zheng, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Reconstruct Multiscale Features for Lightweight Small Object Detection in Remote Sensing ImagesabstractMost small objects are missed when object detection algorithms are transferred from natural images to remote sensing images. Constructing multi-scale features has been proven to be an effective approach for detecting small objects. However, existing methods for multi-scale features have two limitations: insufficient discriminative capability and sparse semantic-spatial information, which fail to fully leverage the potential of multi-scale features. To overcome these limitations, we propose the Multiscale Feature Reconstruction Network (MRN), which introduces three novel modules during feature extraction, fusion, and enhancement: the Composite Multi-scale Feature Extraction Module (CEM), Interlayer Feature Joint Module (IJM), and Spatial-Semantic Information Cross Module (SSM). First, CEM utilizes a multi-branch structure to aggregate scale information. Dilated convolution and asymmetric convolution are extensively used in the branches, which expand the receptive field and capture information of rectangular instances, respectively. Second, the IJM leverages a gating mechanism to achieve pixel-level feature enhancement for feature maps at different hierarchical depths. Finally, the SSM alleviates high-level semantic feature information imbalance through dual-branch information interaction. Furthermore, to utilize the limited computational resources, we propose a lightweight version called MRN_Lite. We evaluate MRN and MRN_Lite on three existing public datasets: AI-TOD, VEDAI, and VisDrone2019. Extensive experiments demonstrate the effectiveness of our method. In comparison experiments, both versions of the model outperform the state of the art (SOTA). And MRN_Lite has less than 50% of the FLOPs and parameters of MRN, which has comparable performance to the original version. Yuancheng Huang, Renwei Qin, Xiangtao Zheng, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Efficient Prompt Tuning of Large Vision-Language Model for Fine-Grained Ship ClassificationabstractRemote-sensing fine-grained ship classification (RS-FGSC) poses a significant challenge due to the high similarity between classes and the limited availability of labeled data, limiting the effectiveness of traditional supervised classification methods. Recent advancements in large pretrained vision-language models (VLMs) have demonstrated impressive capabilities in few-shot or zero-shot learning, particularly in understanding image content. This study delves into harnessing the potential of VLMs to enhance classification accuracy for unseen ship categories, which holds considerable significance in scenarios with restricted data due to cost or privacy constraints. Directly fine-tuning VLMs for RS-FGSC often encounters the challenge of overfitting the seen classes, resulting in suboptimal generalization to unseen classes, which highlights the difficulty in differentiating complex backgrounds and capturing distinct ship features. To address these issues, we introduce a novel prompt tuning technique that employs a hierarchical, multigranularity prompt design. Our approach integrates remote sensing ship priors through bias terms, learned from a small trainable network. This strategy enhances the model’s generalization capabilities while improving its ability to discern intricate backgrounds and learn discriminative ship features. Furthermore, we contribute to the field by introducing a comprehensive dataset, FGSCM-52, significantly expanding existing datasets with more extensive data and detailed annotations for less common ship classes. Extensive experimental evaluations demonstrate the superiority of our proposed method over current state-of-the-art techniques. The source code will be made publicly available. Long Lan, Fengxiang Wang 0004, Xiangtao Zheng, Zengmao Wang, Xinwang Liu 0002 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Integrating Local-Global Structural Interaction Using Siamese Graph Neural Network for Urban Land Use Change Detection From VHR Satellite ImagesabstractDetecting land use changes in urban areas from very-high-resolution (VHR) satellite images presents two primary challenges: 1) traditional methods focus mainly on comparing changes in land cover-related features, which are insufficient for detecting changes in land use and are prone to pseudo-changes caused by illumination differences, seasonal variations, and subtle structural changes and 2) spatial structural information, which is characterized by topological relationships among land cover objects, is crucial for urban land use classification but remains underexplored in change detection. To address these challenges, this study developed a local-global structural interaction network (LGSI-Net) based on a Siamese graph neural network (SGNN) that integrates high-level structural and semantic information to detect urban land use changes from bitemporal VHR images. We developed both local structural feature interaction module (LSIM) and global structural feature interaction module (GSIM) to enhance the representation of bitemporal structural features at the global scene graph and local object node levels. Experiments on the publicly available MtS-WH dataset and two generated datasets, LUCD-FZ and LUCD-HF, show that the proposed method outperforms the existing bag of visual word (BoVW)-based method and CorrFusionNet. Furthermore, we evaluated the detection performance for different semantic feature extraction strategies and structural feature extraction backbones. The results demonstrate that the proposed method, which integrates high-level semantic and graph isomorphism network (GIN)-derived structural features achieves the best performance. The method trained on the LUCD-FZ dataset was successfully transferred to the LUCD-HF dataset with different urban landscapes, indicating its effectiveness in detecting land use changes from VHR satellite images, even in areas with relatively large imbalances between changed and unchanged samples. Kangkai Lou, Mengmeng Li 0002, Fashuai Li, Xiangtao Zheng |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Domain Mapping Network for Remote Sensing Cross-Domain Few-Shot ClassificationabstractIt is a challenging task to recognize novel categories with only a few labeled remote sensing images. Currently, meta-learning solves the problem by learning prior knowledge from another dataset where the classes are disjoint. However, the existing methods assume the training dataset comes from the same domain as the test dataset. For remote sensing images, test dataset may come from different domains. It is impossible to collect a training dataset for each domain. Meta-learning and transfer learning are widely used to tackle the few-shot classification and the cross-domain classification, respectively. However, it is difficult to recognize novel categories from various domains with only a few images. In this paper, a Domain Mapping Network (DMN) is proposed to cope with the few-shot classification under domain shift. DMN trains an efficient few-shot classification model on the source domain and then adapts the model to the target domain. Specifically, dual autoencoders are exploited to fit the source and target domain distribution. First, DMN learns an autoencoder on the source domain to fit the source domain distribution. Then, a target autoencoder is initiated from the source domain autoencoder and further updated with a few target images. To ensure the distribution alignment, cycle-consistency losses are proposed to jointly train the source autoencoder and target autoencoder. Extensive experiments are conducted to validate the generalizable and superiority of the proposed method. Xiaoqiang Lu, Tengfei Gong, Xiangtao Zheng |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Unsupervised real image super-resolution via knowledge distillation network
Nianzeng Yuan, Bangyong Sun, Xiangtao Zheng |
Comput. Vis. Image Underst. | 3 |
| 2023 | Deep Feature Reconstruction Learning for Open-Set Classification of Remote-Sensing ImageryabstractExisting remote sensing scene image (RSSI) classification methods usually rely on static closed-set assumption that testing samples do not belong to unknown classes. However, practical applications are usually the open-set classification problem, which means that RSSIs from unknown classes will appear in the testing set. Most existing methods are prone to forcibly misclassify RSSIs of unknown classes into known classes, resulting in poor practical performance. In this letter, a deep feature reconstruction learning (DFRL) framework is proposed for open-set classification of RSSIs. The proposed DFRL unifies discriminative feature learning and feature reconstruction into an end-to-end network. Firstly, a feature extraction module is utilized to project raw input data from the image space to the feature space to extract deep features. Then, the deep features are fed to a deep feature reconstruction module for distinguishing known and unknown classes based on feature-level reconstruction errors. The feature-level reconstruction can effectively suppress the interference of complex backgrounds. In addition, a sparse regularization is introduced to improve the discrimination of image representation. Experiments on three RSSI datasets demonstrate the effectiveness of DFRL for open-set classification of RSSIs. Hao Sun 0014, Jie Yu 0003, Dongbo Zhou, Wenjing Chen 0003, Xiangtao Zheng, Xiaoqiang Lu |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2023 | Multiple Source Domain Adaptation for Multiple Object Tracking in Satellite VideoabstractSatellite videos capture the dynamic changes in a large observed sense, which provides an opportunity to track the object trajectories. However, existing multiple object tracking methods require massive video annotations, which is time-consuming and fallible. To alleviate this problem, this paper proposes a Cross-Domain multiple object Tracker (CDTrack) to learn knowledge from multiple source domains. First, a cross-domain object detector with multi-level domain alignment is constructed to learn domain-invariant knowledge between remote sensing images and satellite videos. Second, the proposed method adopts a bidirectional teacher-student framework to fuse multiple source domains. Two teacher-student models learn different domain knowledge and teach mutually each other. With mutual learning, the proposed method alleviates the discrepancies between different domains. Finally, a simple weakly supervised re-identification model is proposed for long-term association. Experimental results on the satellite video datasets demonstrate that the proposed method can achieve great performance without satellite video annotations. The code is available at https://github.com/XiangtaoZheng/CDTrack. Xiangtao Zheng, Haowen Cui, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Dual Teacher: A Semisupervised Cotraining Framework for Cross-Domain Ship DetectionabstractCross-domain ship detection tries to identify Synthetic Aperture Radar (SAR) ship by adapting knowledge from labeled optical images, without labor-intensive annotations. In practical applications, a few (e.g., one or three samples) labeled SAR samples are available, which provides an additional supervision for SAR ships. However, the existing cross-domain methods ignore the SAR supervision (a few labeled and unlabeled SAR images), which limits their performances in a practical and under-investigated task: semi-supervised cross-domain ship detection. In this paper, a Dual Teacher framework is proposed to address the mutual interference between the optical supervision and the SAR supervision. First, both optical and SAR supervision are decomposed into two sub-tasks: cross-domain task and semi-supervised task. Then, both cross-domain task and semi-supervised task can be learned interactively in two individual teacher-student models. The teacher-student models generate pseudo-labels on unlabeled SAR images by a teacher network and fine-tune the student network. Finally, the Dual Teacher framework retrains two teacher-student models in co-training strategies. Both cross-domain dataset and semi-supervised dataset are exploited to jointly improve the pseudo-label quality. The effectiveness of the Dual Teacher framework has been fully experimentally demonstrated. The code is available at https://github.com/XiangtaoZheng/DualTeacher. Xiangtao Zheng, Haowen Cui, Chujie Xu, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | NCSiam: Reliable Matching via Neighborhood Consensus for Siamese-Based Object TrackingabstractAn essential need for accurate visual object tracking is to capture better correlations between the tracking target and the search region. However, the dominant Siamese-based trackers are limited to producing dense similarity maps at once via a cross-correlations operation, ignoring to remedy the contamination caused by erroneous or ambiguous matches. In this paper, we propose a novel tracker, termed neighborhood consensus constraint-based siamese tracker (NCSiam), which takes the idea of neighborhood consensus constraint to refine the produced correlation maps. The intuition behind our approach is that we can support the nearby erroneous or ambiguous matches by analyzing a larger context of the scene that contains a unique match. Specifically, we devise a 4D convolution-based multi-level similarity refinement (MLSR) strategy. Taking the primary similarity maps obtained from a cross-correlation as input, MLSR acquires reliable matches by analyzing neighborhood consensus patterns in 4D space, thus enhancing the discriminability between the tracking target and the distractors. Besides, traditional Siamese-based trackers directly perform classification and regression on similarity response maps which discard appearance or semantic information. Therefore, an appearance affinity decoder (AAD) is developed to take full advantage of the semantic information of the search region. To further improve performance, we design a task-specific disentanglement (TSD) module to decouple the learned representations into classification-specific and regression-specific embeddings. Extensive experiments are conducted on six challenging benchmarks, including GOT-10k, TrackingNet, LaSOT, UAV123, OTB2015, and VOT2020. The results demonstrate the effectiveness of our method. The code will be available at https://github.com/laybebe/NCSiam. Pujian Lai, Gong Cheng 0003, Meili Zhang, Jifeng Ning, Xiangtao Zheng, Junwei Han 0001 |
IEEE Trans. Image Process. | 5 |
| 2023 | Identity Feature Disentanglement for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) task aims to retrieve persons from different spectrum cameras (i.e., visible and infrared images). The biggest challenge of VI-ReID is the huge cross-modal discrepancy caused by different imaging mechanisms. Many VI-ReID methods have been proposed by embedding different modal person images into a shared feature space to narrow the cross-modal discrepancy. However, these methods ignore the purification of identity features, which results in identity features containing different modal information and failing to align well. In this article, an identity feature disentanglement method is proposed to disentangle the identity features from identity-irrelevant information, such as pose and modality. Specifically, images of different modalities are first processed to extract shared features that reduce the cross-modal discrepancy preliminarily. Then the extracted feature of each image is disentangled into a latent identity variable and an identity-irrelevant variable. In order to enforce the latent identity variable to contain as much identity information as possible and as little identity-irrelevant information, an ID-discriminative loss and an ID-swapping reconstruction process are additionally designed. Extensive quantitative and qualitative experiments on two popular public VI-ReID datasets, RegDB and SYSU-MM01, demonstrate the efficacy and superiority of the proposed method. Xiumei Chen, Xiangtao Zheng, Xiaoqiang Lu |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | Human action recognition by multiple spatial clues network
Xiangtao Zheng, Tengfei Gong, Xiaoqiang Lu, Xuelong Li 0001 |
Neurocomputing | 1 |
| 2022 | Semisupervised Spectral Degradation Constrained Network for Spectral Super-ResolutionabstractRecently, various deep learning-based methods have been designed to improve the spectral resolution of the multispectral image (MSI) to obtain the hyperspectral image (HSI). These methods usually rely on sufficient MSI/HSI pairs for supervised training. However, collecting plentiful HSIs is time-consuming. In this letter, a semisupervised spectral degradation constrained network (SSDCN) is proposed to improve the spectral resolution of MSI. SSDCN is an autoencoder-like network that is composed of an encoder subnetwork for estimating HSI from input MSI and a decoder subnetwork for reconstructing MSI from the estimated HSI. A semisupervised training method is proposed to explore both MSI/HSI pairs and MSIs without ground-truth HSIs to optimize SSDCN. Simulated and two real databases are employed to demonstrate the effectiveness of SSDCN. Wenjing Chen 0003, Xiangtao Zheng, Xiaoqiang Lu |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Remote Sensing Scene Classification by Local-Global Mutual LearningabstractRemote sensing scene classification (RSSC) attempts to label an image with a specific scene category. Recently, convolutional neural networks (CNNs) have shown the powerful feature extraction capability to combine local and global features. However, both the local and global features are extracted independently, which ignore the complementary representation. In this letter, a local–global mutual learning (LML) method is proposed to capture both the global and local features. Specifically, local regions are first generated by highlighting the semantic areas in the corresponding original image. Then, a two-branch architecture is used to extract features for the local regions and global image, respectively. Both the classification loss and mutual learning loss are exploited to train the local–global branches simultaneously, which constrain the two branches to promote each other. Experiments on two popular datasets demonstrate the effectiveness of the proposed method. Xiumei Chen, Xiangtao Zheng, Yue Zhang 0053, Xiaoqiang Lu |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Meta Self-Supervised Learning for Distribution Shifted Few-Shot Scene ClassificationabstractFew-shot classification tries to recognize novel remote sensing image categories with a few shot samples. However, current methods assume that the test dataset shares the same domain with the labeled training dataset where prior knowledge is learned. It is infeasible to collect a training dataset for each domain, since remote sensing images may come from various domains. Exploiting the existing labeled dataset from another domain (source domain) to help the target dataset (target domain) classification would be efficient. In this paper, both meta-learning and self-supervised learning are jointly conducted for few-shot classification. Specifically, meta-learning is executed over a pre-trained network for few-shot classification. Furthermore, self-supervised learning is exploited to fit the target domain distribution by training on unlabeled target domain images. Experiments are conducted on NWPU, EuroSAT and Merced datasets to validate the effectiveness. Tengfei Gong, Xiangtao Zheng, Xiaoqiang Lu |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Pairwise Comparison Network for Remote-Sensing Scene ClassificationabstractRemote-sensing scene classification aims to assign a specific semantic label to a remote-sensing image. Recently, convolutional neural networks (CNNs) have greatly improved the performance of remote-sensing scene classification. However, some confused images may be easily recognized as the incorrect category, which generally degrade the performance. The differences between image pairs can be used to distinguish image categories. This letter proposed a pairwise comparison network (PCNet), which contains two main steps: pairwise selection and pairwise representation. The proposed network first selects similar image pairs and then represents the image pairs with pairwise representations. The self-representation is introduced to highlight the informative parts of each image itself, while the mutual representation is proposed to capture the subtle differences between image pairs. Comprehensive experimental results on two challenging datasets (AID, NWPU-RESISC45) demonstrate the effectiveness of the proposed network. Yue Zhang 0053, Xiangtao Zheng, Xiaoqiang Lu |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Spectral Super-Resolution of Multispectral Images Using Spatial-Spectral Residual Attention NetworkabstractThe spectral super-resolution of multispectral image (MSI) refers to improving the spectral resolution of the MSI to obtain the hyperspectral image (HSI). Most recent works are based on the sparse representation to unfold the MSI into the 2-D matrix in advance for subsequent operations, which results in that the spatial information of MSI cannot be fully explored. In this article, a spatial–spectral residual attention network (SSRAN) is proposed to simultaneously explore the spatial and spectral information of MSI for reconstructing the HSI. The proposed SSRAN is composed of the feature extraction part, the nonlinear mapping part, and the reconstruction part. Firstly, the multispectral features of the input MSI are extracted in the feature extraction part. Second, in the nonlinear mapping part, the spatial–spectral residual blocks are proposed to explore spatial and spectral information of MSI for mapping the multispectral features to the hyperspectral features. Finally, in the reconstruction part, a 2-D convolution is used to reconstruct the HSI from the hyperspectral features. Also, a neighboring spectral attention module is specially designed to explicitly constrain the reconstructed HSI to maintain the correlation among neighboring spectral bands. The proposed SSRAN outperforms the state-of-the-art methods on both simulated and real databases. Xiangtao Zheng, Wenjing Chen 0003, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Unsupervised Change Detection by Cross-Resolution Difference LearningabstractChange detection (CD) aims to identify the differences between multitemporal images acquired over the same geographical area at different times. With the advantages of requiring no cumbersome labeled change information, unsupervised CD has attracted extensive attention of researchers. Multitemporal images tend to have different resolutions as they are usually captured at different times with different sensor properties. It is difficult to directly obtain one pixelwise change map for two images with different resolutions, so current methods usually resize multitemporal images to a unified size. However, resizing operations change the original information of pixels, which limits the final CD performance. This article aims to detect changes from multitemporal images in the originally different resolutions without resizing operations. To achieve this, a cross-resolution difference learning method is proposed. Specifically, two cross-resolution pixelwise difference maps are generated for the two different resolution images and fused to produce the final change map. First, the two input images are segmented into individual homogeneous regions separately due to different resolutions. Second, each pixelwise difference map is produced according to two measure distances, the mutual information distance and the deep feature distance, between image regions in which the pixel lies. Third, the final binary change map is generated by fusing and binarizing the two cross-resolution difference maps. Extensive experiments on four datasets demonstrate the effectiveness of the proposed method for detecting changes from different resolution images. Xiangtao Zheng, Xiumei Chen, Xiaoqiang Lu, Bangyong Sun |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Generalized Scene Classification From Small-Scale Datasets With Multitask LearningabstractRemote sensing images contain a wealth of spatial information. Efficient scene classification is a necessary precedent step for further application. Despite the great practical value, the mainstream methods using deep convolutional neural networks (CNNs) are generally pretrained on other large datasets (such as ImageNet) and thus fail to capture the specific visual characteristics of remote sensing images. For another, it lacks the generalization ability to new tasks when training a new CNN from scratch with an existing remote sensing dataset. This article addresses the dilemma and uses multiple small-scale datasets to learn a generalized model for efficient scene classification. Since the existing datasets are heterogeneous and cannot be directly combined to train a network, a multitask learning network (MTLN) is developed. The MTLN treats each small-scale dataset as an individual task and uses complementary information contained in multiple tasks to improve generalization. Concretely, the MTLN consists of a shared branch for all tasks and multiple task-specific branches with each for one task. The shared branch extracts shared features for all tasks to achieve information sharing among tasks. The task-specific branch distills the shared features into task-specific features toward the optimal estimation of each specific task. By jointly learning shared features and task-specific features, the MTLN maintains both generalization and discrimination abilities. Two types of MTL scenarios are explored to validate the effectiveness of the proposed method: one is to complete multiple scene classification tasks and the other is to jointly perform scene classification and semantic segmentation. Xiangtao Zheng, Tengfei Gong, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Mutual Attention Inception Network for Remote Sensing Visual Question AnsweringabstractRemote sensing images (RSIs) containing various ground objects have been applied in many fields. To make semantic understanding of RSIs objective and interactive, the task remote sensingvisual question answering(VQA) has appeared. Given an RSI, the goal of remote sensing VQA is to make an intelligent agent answer a question about the remote sensing scene. Existing remote sensing VQA methods utilized a nonspatial fusion strategy to fuse the image features and question features, which ignores the spatial information of images and word-level information of questions. A novel method is proposed to complete the task considering these two aspects. First, convolutional features of the image are included to represent spatial information, and the word vectors of questions are adopted to present semantic word information. Second, attention mechanism and bilinear technique are introduced to enhance the feature considering the alignments between spatial positions and words. Finally, a fully connected layer with softmax is utilized to output an answer from the perspective of the multiclass classification task. To benchmark this task, aRSIVQAdataset is introduced in this article. For each of more than 37 000 RSIs, the proposed dataset contains at least one or more questions, plus corresponding answers. Experimental results demonstrate that the proposed method can capture the alignments between images and questions. The code and dataset are available athttps://github.com/spectralpublic/RSIVQA. Xiangtao Zheng, Binqiang Wang, Xingqian Du, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Visible-Infrared Person Re-Identification via Partially Interactive CollaborationabstractVisible-infrared person re-identification (VI-ReID) task aims to retrieve the same person between visible and infrared images. VI-ReID is challenging as the images captured by different spectra present large cross-modality discrepancy. Many methods adopt a two-stream network and design additional constraint conditions to extract shared features for different modalities. However, the interaction between the feature extraction processes of different modalities is rarely considered. In this paper, a partially interactive collaboration method is proposed to exploit the complementary information of different modalities to reduce the modality gap for VI-ReID. Specifically, the proposed method is achieved in a partially interactive-shared architecture: collaborative shallow layers and shared deep layers. The collaborative shallow layers consider the interaction between modality-specific features of different modalities, encouraging the feature extraction processes of different modalities constrain each other to enhance feature representations. The shared deep layers further embed the modality-specific features to a common space to endow them the same identity discriminability. To ensure the interactive collaborative learning implement effectively, the conventional loss and collaborative loss are utilized jointly to train the whole network. Extensive experiments on two publicly available VI-ReID datasets verify the superiority of the proposed PIC method. Specifically, the proposed method achieves a rank-1 accuracy of 83.6% and 57.5% on RegDB and SYSU-MM01 datasets, respectively. Xiangtao Zheng, Xiumei Chen, Xiaoqiang Lu |
IEEE Trans. Image Process. | 1 |
| 2022 | Rotation-Invariant Attention Network for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification refers to identifying land-cover categories of pixels based on spectral signatures and spatial information of HSIs. In recent deep learning-based methods, to explore the spatial information of HSIs, the HSI patch is usually cropped from original HSI as the input. And 3 ×3 convolution is utilized as a key component to capture spatial features for HSI classification. However, the 3 ×3 convolution is sensitive to the spatial rotation of inputs, which results in that recent methods perform worse in rotated HSIs. To alleviate this problem, a rotation-invariant attention network (RIAN) is proposed for HSI classification. First, a center spectral attention (CSpeA) module is designed to avoid the influence of other categories of pixels to suppress redundant spectral bands. Then, a rectified spatial attention (RSpaA) module is proposed to replace 3 ×3 convolution for extracting rotation-invariant spectral-spatial features from HSI patches. The CSpeA module, the 1 ×1 convolution and the RSpaA module are utilized to build the proposed RIAN for HSI classification. Experimental results demonstrate that RIAN is invariant to the spatial rotation of HSIs and has superior performance, e.g., achieving an overall accuracy of 86.53% (1.04% improvement) on the Houston database. The codes of this work are available at https://github.com/spectralpublic/RIAN. Xiangtao Zheng, Hao Sun 0014, Xiaoqiang Lu, Wei Xie 0008 |
IEEE Trans. Image Process. | 1 |
| 2022 | Disentangled Representation Learning for Cross-Modal Biometric MatchingabstractCross-modal biometric matching (CMBM) aims to determine the corresponding voice from a face, or identify the corresponding face from a voice. Recently, many CMBM methods have been proposed by forcing the distance between two modal features to be narrowed. However, these methods ignore the alignability between the two modal features. Because the feature is extracted under the supervision of identity information from single modal data, it can only reflect the identity information of single modal data. In order to address this problem, a disentangled representation learning method is proposed to disentangle the alignable latent identity factors and nonalignable the modality-dependent factors for CMBM. The proposed method consists of two main steps: 1) feature extraction and 2) disentangled representation learning. Firstly, an image feature extraction network is adopted to obtain face features, and a voice feature extraction network is applied to learn voice features. Secondly, a disentangled latent variable is explored to disentangle the latent identity factors that are shared across the modalities from the modality-dependent factors. The modality-dependent factors are filtered out, while the latent identity factors from the two modalities are enforced to be narrowed to align the same identity information. Then, the disentangled latent identity factors are considered as pure identity information to bridge the two modalities for cross-modal verification, 1:$N$matching, and retrieval. Note that the proposed method learns the identity information from the input face images and voice segments with only identity label as supervised information. Extensive experiments on the challenging VoxCeleb dataset demonstrate the proposed method outperforms the state-of-the-art methods. Hailong Ning, Xiangtao Zheng, Xiaoqiang Lu, Yuan Yuan 0001 |
IEEE Trans. Multim. | 2 |
| 2021 | Audio description from image by modal translation network
Hailong Ning, Xiangtao Zheng, Yuan Yuan 0001, Xiaoqiang Lu |
Neurocomputing | 2 |
| 2021 | Local and correlation attention learning for subtle facial expression recognition
Yuan Yuan 0001, Xiangtao Zheng, Xiaoqiang Lu |
Neurocomputing | 3 |
| 2021 | Remote Sensing Image Generation From AudioabstractGenerating image from other modal data has attracted much attention in cross-modal studies, since the generated image offers intuitive vision information. Unlike the previous works which generate an image from text, a novel task is introduced, generating an image from audio. However, semantic gap intrinsically exists in cross-modal data, which disturbs the generative results. In order to explore the relevance between the audio and image, a novel reranking audio-image translation method is proposed. The proposed method: 1) maps the audio and image into a uniform feature space; 2) designs an audio-audio matching network to match the related audio; and 3) adopts an audio-image matching network for every matched audio to generate a related image, and the most frequent image is voted as the final result. Extensive experiments on two remote sensing cross-modal data sets demonstrate that the proposed method can visualize the content of audio. Jun Chen 0001, Xiangtao Zheng, Xiaoqiang Lu |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2021 | Multisource Remote Sensing Data Classification With Graph Fusion NetworkabstractThe land cover classification has been an important task in remote sensing. With the development of various sensors technologies, carrying out classification work with multisource remote sensing (MSRS) data has shown an advantage over using a single type of data. Hyperspectral images (HSIs) are able to represent the spectral properties of land cover, which is quite common for land cover understanding. Light detection and ranging (LiDAR) images contain altitude information of the ground, which is greatly helpful with urban scene analysis. Current HSI and LiDAR fusion methods perform feature extraction and feature fusion separately, which cannot well exploit the correlation of data sources. In order to make full use of the correlation of multisource data, an unsupervised feature extraction-fusion network for HSI and LiDAR, which utilizes feature fusion to guide the feature extraction procedure, is proposed in this article. More specifically, the network takes multisource data as input and directly output the unified fused feature. A multimodal graph is constructed for feature fusion, and graph-based loss functions including Laplacian loss and t-distributed stochastic neighbor embedding (t-SNE) loss are utilized to constrain the feature extraction network. Experimental results on several data sets demonstrate the proposed network can achieve more excellent classification performance than some state-of-the-art methods. Xingqian Du, Xiangtao Zheng, Xiaoqiang Lu, Alexander A. Doudkin |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Cross-Domain Scene Classification by Integrating Multiple Incomplete SourcesabstractCross-domain scene classification identifies scene categories by learning knowledge from a labeled data set (source domain) to an unlabeled data set (target domain), where the source data and the target data are sampled from different distributions. A lot of domain adaptation methods are used to reduce the distribution shift across domains, and most existing methods assume that the source domain shares the same categories with the target domain. It is usually hard to find a source domain that covers all categories in the target domain. Some works exploit multiple incomplete source domains to cover the target domain. However, in such setting, the categories of each source domain are a subset of the target-domain categories, and the target domain contains “unknown” categories for each source domain. The existence of unknown categories results in the conventional domain adaptation unsuitable. Known and unknown categories should be treated separately. Therefore, a separation mechanism is proposed to separate the known and unknown categories in this article. First, multiple-source classifiers trained on the multiple source domains are used to coarsely separate the known/unknown categories in the target domain. The target images with high similarities to source images are selected as known categories, and the target images with low similarities are selected as unknown categories. Then, a binary classifier trained using the selected images is used to finely separate all target-domain images. Finally, only the known categories are implemented in the cross-domain alignment and classification. The target images get labels by integrating the hypotheses of multiple-source classifiers on the known categories. Experiments are conducted on three cross-domain data sets to demonstrate the effectiveness of the proposed method. Tengfei Gong, Xiangtao Zheng, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Bidirectional Interaction Network for Person Re-IdentificationabstractPerson re-identification (ReID) task aims to retrieve the same person across multiple spatially disjoint camera views. Due to huge image changes caused by various factors such as posture variation and illumination transformation, images of different persons may share the more similar appearances than images of the same one. Learning discriminative representations to distinguish details of different persons is significant for person ReID. Many existing methods learn discriminative representations resorting to a human body part location branch which requires cumbersome expert human annotations or complex network designs. In this article, a novel bidirectional interaction network is proposed to explore discriminative representations for person ReID without any human body part detection. The proposed method regards multiple convolutional features as responses to various body part properties and exploits the inter-layer interaction to mine discriminative representations for person identities. Firstly, an inter-layer bilinear pooling strategy is proposed to feasibly exploit the pairwise feature relations between two convolution layers. Secondly, to explore interaction of multiple layers, an effective bidirectional integration strategy consisting of two different multi-layer interaction processes is designed to aggregate bilinear pooling interaction of multiple convolution layers. The interaction of multiple layers is implemented in a layer-by-layer nesting policy to ensure the two interaction processes are different and complementary. Extensive experiments validate the superiority of the proposed method on four popular person ReID datasets including Market-1501, DukeMTMC-ReID, CUHK03-NP and MSMT17. Specifically, the proposed method achieves a rank-1 accuracy of 95.1% and 88.2% on Market-1501 and DukeMTMC-ReID, respectively. Xiumei Chen, Xiangtao Zheng, Xiaoqiang Lu |
IEEE Trans. Image Process. | 2 |
| 2021 | A Supervised Segmentation Network for Hyperspectral Image ClassificationabstractRecently, deep learning has drawn broad attention in the hyperspectral image (HSI) classification task. Many works have focused on elaborately designing various spectral-spatial networks, where convolutional neural network (CNN) is one of the most popular structures. To explore the spatial information for HSI classification, pixels with its adjacent pixels are usually directly cropped from hyperspectral data to form HSI cubes in CNN-based methods. However, the spatial land-cover distributions of cropped HSI cubes are usually complicated. The land-cover label of a cropped HSI cube cannot simply be determined by its center pixel. In addition, the spatial land-cover distribution of a cropped HSI cube is fixed and has less diversity. For CNN-based methods, training with cropped HSI cubes will result in poor generalization to the changes of spatial land-cover distributions. In this paper, an end-to-end fully convolutional segmentation network (FCSN) is proposed to simultaneously identify land-cover labels of all pixels in a HSI cube. First, several experiments are conducted to demonstrate that recent CNN-based methods show the weak generalization capabilities. Second, a fine label style is proposed to label all pixels of HSI cubes to provide detailed spatial land-cover distributions of HSI cubes. Third, a HSI cube generation method is proposed to generate plentiful HSI cubes with fine labels to improve the diversity of spatial land-cover distributions. Finally, a FCSN is proposed to explore spectral-spatial features from finely labeled HSI cubes for HSI classification. Experimental results show that FCSN has the superior generalization capability to the changes of spatial land-cover distributions. Hao Sun 0014, Xiangtao Zheng, Xiaoqiang Lu |
IEEE Trans. Image Process. | 2 |
| 2021 | Fine-Grained Visual Categorization by Localizing Object Parts With Single ImageabstractFine-grained visual categorization (FGVC) refers to assigning fine-grained labels to images which belong to the same base category. Due to the high inter-class similarity, it is challenging to distinguish fine-grained images under different subcategories. Recently, researchers have proposed to firstly localize key object parts within images and then find discriminative clues on object parts. To localize object parts, existing methods train detectors for different kinds of object parts. However, due to the fact that the same kind of object part in different images often changes intensely in appearance, the existing methods face two shortages: 1) Training part detector for object parts with diverse appearance is laborious; 2) Discriminative parts with unusual appearance may be neglected by the trained part detectors. To localize the key object parts efficiently and accurately, a novel FGVC method is proposed in the paper. The main novelty is that the proposed method localizes the key object parts within each image only depending on a single image and hence avoid the influence of diversity between parts in different images. The proposed FGVC method consists of two key steps. Firstly, the proposed method localizes the key parts in each image independently. To this end, potential object parts in each image are identified and then these potential parts are merged to generate the final representative object parts. Secondly, two kinds of features are extracted for simultaneously describing the discriminative clues within each part and the relationship between object parts. In addition, a part based dropout learning technique is adopted to boost the classification performance further in the paper. The proposed method is evaluated in comparison experiments and the experiment results show that the proposed method can achieve comparable or better performance than state-of-the-art methods. Xiangtao Zheng, Lei Qi 0004, Yutao Ren, Xiaoqiang Lu |
IEEE Trans. Multim. | 1 |
| 2020 | Deep balanced discrete hashing for image retrieval
Xiangtao Zheng, Yichao Zhang 0005, Xiaoqiang Lu |
Neurocomputing | 1 |
| 2020 | Spatial attention based visual semantic learning for action recognition in still images
Yunpeng Zheng, Xiangtao Zheng, Xiaoqiang Lu |
Neurocomputing | 2 |
| 2020 | Multisource Compensation Network for Remote Sensing Cross-Domain Scene ClassificationabstractCross-domain scene classification refers to the scene classification task in which the training set (termed source domain) and the test set (termed target domain) come from different distributions. Various domain adaptation methods have been developed to reduce the distribution discrepancy between different domains. However, current domain adaptation methods assume that the source domain and target domain share the same categories. In reality, it is hard to find a source domain that can completely cover all the categories of target domain. In this article, we propose to use multiple complementary source domains to form the categories of target domain. A multisource compensation network (MSCN) is proposed to tackle these challenges: distribution discrepancy and category incompleteness. First, a pretrained convolutional neural network (CNN) is exploited to learn the feature representation for each domain. Second, a cross-domain alignment module is developed to reduce the domain shift between source and target domains. Domain shift is reduced by mapping the two domain features into a common feature space. Finally, a classifier complement module is proposed to align categories in multiple sources and learn a target classifier. Two cross-domain classification data sets are constructed using four heterogeneous remote sensing scene classification data sets. Extensive experiments are conducted on these datasets to validate the effectiveness of the proposed method. The proposed method can achieve 81.23% and 81.97% average accuracies on two-source-complementary data set and three-source-complementary data set, respectively. Xiaoqiang Lu, Tengfei Gong, Xiangtao Zheng |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | Sound Active Attention Framework for Remote Sensing Image CaptioningabstractAttention mechanism-based image captioning methods have achieved good results in the remote sensing field, but are driven by tagged sentences, which is called passive attention. However, different observers may give different levels of attention to the same image. The attention of observers during testing, then, may not be consistent with the attention during training. As a direct and natural human-machine interaction, speech is much faster than typing sentences. Sound can represent the attention of different observers. This is called active attention. Active attention can be more targeted to describe the image; for example, in disaster assessments, the situation can be obtained quickly and the corresponding disaster areas can be located related to the specific disaster. A novel sound active attention framework is proposed for more specific caption generation according to the interest of the observer. First, sound is modeled by mel-frequency cepstral coefficients (MFCCs) and the image is encoded by convolutional neural networks (CNNs). Then, to handle the continuity characteristic of sound, a sound module and an attention module are designed based on the gated recurrent units (GRUs). Finally, the sound-guided image feature processed by the attention module is imported into the output module to generate descriptive sentence. Experiments based on both fake and real sound data sets show that the proposed method can generate sentences that can capture the focus of human. Xiaoqiang Lu, Binqiang Wang, Xiangtao Zheng |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | Remote Sensing Scene Classification by Gated Bidirectional NetworkabstractRemote sensing (RS) scene classification is a challenging task due to various land covers contained in RS scenes. Recent RS classification methods demonstrate that aggregating the multilayer convolutional features, which are extracted from different hierarchical layers of a convolutional neural network, can effectively improve classification accuracy. However, these methods treat the multilayer convolutional features as equally important and ignore the hierarchical structure of multilayer convolutional features. Multilayer convolutional features not only provide complementary information for classification but also bring some interference information (e.g., redundancy and mutual exclusion). In this paper, a gated bidirectional network is proposed to integrate the hierarchical feature aggregation and the interference information elimination into an end-to-end network. First, the performance of each convolutional feature is quantitatively analyzed and a superior combination of convolutional features is selected. Then, a bidirectional connection is proposed to hierarchically aggregate multilayer convolutional features. Both the top–down direction and the bottom–up direction are considered to aggregate multilayer convolutional features into the semantic-assist feature and appearance-assist feature, respectively, and a gated function is utilized to eliminate interference information in the bidirectional connection. Finally, the semantic-assist feature and appearance-assist feature are merged for classification. The proposed method can compete with the state-of-the-art methods on four RS scene classification data sets (AID, UC-Merced, WHU-RS19, and OPTIMAL-31). Hao Sun 0014, Xiangtao Zheng, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | Spectral-Spatial Attention Network for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification aims to assign each hyperspectral pixel with a proper land-cover label. Recently, convolutional neural networks (CNNs) have shown superior performance. To identify the land-cover label, CNN-based methods exploit the adjacent pixels as an input HSI cube, which simultaneously contains spectral signatures and spatial information. However, at the edge of each land-cover area, an HSI cube often contains several pixels whose land-cover labels are different from that of the center pixel. These pixels, named interfering pixels, will weaken the discrimination of spectral-spatial features and reduce classification accuracy. In this article, a spectral-spatial attention network (SSAN) is proposed to capture discriminative spectral-spatial features from attention areas of HSI cubes. First, a simple spectral-spatial network (SSN) is built to extract spectral-spatial features from HSI cubes. The SSN is composed of a spectral module and a spatial module. Each module consists of only a few 3-D convolution and activation operations, which make the proposed method easy to converge with a small number of training samples. Second, an attention module is introduced to suppress the effects of interfering pixels. The attention module is embedded into the SSN to obtain the SSAN. The experiments on several public HSI databases demonstrate that the proposed SSAN outperforms several state-of-the-art methods. Hao Sun 0014, Xiangtao Zheng, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | Attribute-Cooperated Convolutional Neural Network for Remote Sensing Image ClassificationabstractRemote sensing image (RSI) classification is one of the most important fields in RSI processing. It is well known that RSIs are very complicated due to its various kinds of contents. Therefore, it is very difficult to distinguish different scene categories with similar visual contents, like desert and bare land. To address hard negative categories, an attribute-cooperated convolutional neural network (ACCNN) is proposed to exploit attributes as additional guiding information. First, the classification branch extracts convolutional neural network feature, which is then utilized to recognize the RSI scene categories. Second, the attribute branch is proposed to make the network distinguish scene categories efficiently. The proposed attribute branch shares feature extraction layers with the classification branch and makes the classification branch aware of extra attribute information. Finally, the relationship branch constraints the relationship between the classification branch and the attribute branch. To exploit the attribute information, three attribute-classification data sets are generated (AC-AID, AC-UCM, and AC-Sydney). Experimental results show that the proposed method is competitive to state-of-the-art methods. The data sets are available at https://github.com/CrazyStoneonRoad/Attribute-Cooperated-Classification-Data sets. Yuanlin Zhang 0003, Xiangtao Zheng, Yuan Yuan 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | A Joint Relationship Aware Neural Network for Single-Image 3D Human Pose EstimationabstractThis paper studies the task of 3D human pose estimation from a single RGB image, which is challenging without depth information. Recently many deep learning methods are proposed and achieve great improvements due to their strong representation learning. However, most existing methods ignore the relationship between joint features. In this paper, a joint relationship aware neural network is proposed to take both global and local joint relationship into consideration. First, a whole feature block representing all human body joints is extracted by a convolutional neural network. A Dual Attention Module (DAM) is applied on the whole feature block to generate attention weights. By exploiting the attention module, the global relationship between the whole joints is encoded. Second, the weighted whole feature block is divided into some individual joint features. To capture salient joint feature, the individual joint features are refined by individual DAMs. Finally, a joint angle prediction constraint is proposed to consider local joint relationship. Quantitative and qualitative experiments on 3D human pose estimation benchmarks demonstrate the effectiveness of the proposed method. Xiangtao Zheng, Xiumei Chen, Xiaoqiang Lu |
IEEE Trans. Image Process. | 1 |
| 2019 | Bidirectional adaptive feature fusion for remote sensing scene classification
Xiaoqiang Lu, Weijun Ji, Xuelong Li 0001, Xiangtao Zheng |
Neurocomputing | 4 |
| 2019 | Semantic Descriptions of High-Resolution Remote Sensing ImagesabstractImage captioning has attracted more and more attention in remote sensing filed since it provides more specific information than traditional tasks, such as classification. Though image captioning has gained some developments in recent years, it is difficult to describe the image in one simple sentence. To relieve the limitation, a novel captioning task is proposed and a novel framework is proposed to solve the novel task. The proposed framework uses semantic embedding to measure the image representation and the sentence representation. The captioning performance is improved by a proposed sentence representation (collective representation). Experimental results and human evaluations on three captioning data sets in remote sensing field demonstrate that the proposed framework can lead to advancement in image captioning results. Binqiang Wang, Xiaoqiang Lu, Xiangtao Zheng, Xuelong Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2019 | A Feature Aggregation Convolutional Neural Network for Remote Sensing Scene ClassificationabstractRemote sensing scene classification (RSSC) refers to inferring semantic labels based on the content of the remote sensing scenes. Recently, most works take the pretrained convolutional neural network (CNN) as the feature extractor to build a scene representation for RSSC. The activations in different layers of CNN (named intermediate features) contain different spatial and semantic information. Recent works demonstrate that aggregating intermediate features into a scene representation can significantly improve the classification accuracy for RSSC. However, the intermediate features are aggregated by some unsupervised feature encoding methods (e.g., Bag-of-Visual-Words). Little attention has been paid to explore the information of semantic labels for the feature aggregation. In this paper, in order to explore the semantic label information, an end-to-end feature aggregation CNN (FACNN) is proposed to learn a scene representation for RSSC. In FACNN, a supervised convolutional features' encoding module and a progressive aggregation strategy are proposed to leverage the semantic label information to aggregate the intermediate features. The FACNN integrates the feature learning, feature aggregation, and classifier into a unified end-to-end framework for joint training. In FACNN, the scene representation is learned by considering the information of semantic labels, which can result in better performance for RSSC. Extensive experiments on AID, UC-Merged, and WHU-RS19 databases demonstrate that FACNN performs better than several state-of-the-art methods. Xiaoqiang Lu, Hao Sun 0014, Xiangtao Zheng |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2019 | Hyperspectral Image Denoising by Fusing the Selected Related BandsabstractHyperspectral images (HSIs) convey more useful information than RGB or gray images, which are widely used in many remote sensing tasks. In real scenarios, HSIs are inevitably corrupted by noise because of sensors' imperfectness or atmospheric influence. Recently, many HSI denoising methods have been proposed to utilize the interband information between different spectral bands. However, these methods regard the HSI as a whole and treat the different spectral bands with the same noise level. In fact, the noise levels in different bands are different. Especially, only few certain bands are corrupted by noise, named the target noised bands. Under this circumstance, an HSI denoising method is proposed by considering the band relationship and different noise levels. The target noised bands are adaptively denoised by fusing some selected bands. Specifically, some related but quality superior bands are selected according to the target noised bands. Then, the target noised bands can be denoised by fusing the selected related bands. Experimental results show that the proposed method achieves considerable performances in comparison with several state-of-the-art hyperspectral denoising methods. Xiangtao Zheng, Yuan Yuan 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | A Deep Scene Representation for Aerial Scene ClassificationabstractAs a fundamental problem in earth observation, aerial scene classification tries to assign a specific semantic label to an aerial image. In recent years, the deep convolutional neural networks (CNNs) have shown advanced performances in aerial scene classification. The successful pretrained CNNs can be transferable to aerial images. However, global CNN activations may lack geometric invariance and, therefore, limit the improvement of aerial scene classification. To address this problem, this paper proposes a deep scene representation to achieve the invariance of CNN features and further enhance the discriminative power. The proposed method: 1) extracts CNN activations from the last convolutional layer of pretrained CNN; 2) performs multiscale pooling (MSP) on these activations; and 3) builds a holistic representation by the Fisher vector method. MSP is a simple and effective multiscale strategy, which enriches multiscale spatial information in affordable computational time. The proposed representation is particularly suited at aerial scenes and consistently outperforms global CNN activations without requiring feature adaptation. Extensive experiments on five aerial scene data sets indicate that the proposed method, even with a simple linear classifier, can achieve the state-of-the-art performance. Xiangtao Zheng, Yuan Yuan 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2018 | Exploring Models and Data for Remote Sensing Image Caption GenerationabstractInspired by recent development of artificial satellite, remote sensing images have attracted extensive attention. Recently, notable progress has been made in scene classification and target detection. However, it is still not clear how to describe the remote sensing image content with accurate and concise sentences. In this paper, we investigate to describe the remote sensing images with accurate and flexible sentences. First, some annotated instructions are presented to better describe the remote sensing images considering the special characteristics of remote sensing images. Second, in order to exhaustively exploit the contents of remote sensing images, a large-scale aerial image data set is constructed for remote sensing image caption. Finally, a comprehensive review is presented on the proposed data set to fully advance the task of remote sensing caption. Extensive experiments on the proposed data set demonstrate that the content of the remote sensing image can be completely described by generating language descriptions. The data set is available at https://github.com/201528014227051/RSICD_optimal. Xiaoqiang Lu, Binqiang Wang, Xiangtao Zheng, Xuelong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2017 | Joint Dictionary Learning for Multispectral Change DetectionabstractChange detection is one of the most important applications of remote sensing technology. It is a challenging task due to the obvious variations in the radiometric value of spectral signature and the limited capability of utilizing spectral information. In this paper, an improved sparse coding method for change detection is proposed. The intuition of the proposed method is that unchanged pixels in different images can be well reconstructed by the joint dictionary, which corresponds to knowledge of unchanged pixels, while changed pixels cannot. First, a query image pair is projected onto the joint dictionary to constitute the knowledge of unchanged pixels. Then reconstruction error is obtained to discriminate between the changed and unchanged pixels in the different images. To select the proper thresholds for determining changed regions, an automatic threshold selection strategy is presented by minimizing the reconstruction errors of the changed pixels. Adequate experiments on multispectral data have been tested, and the experimental results compared with the state-of-the-art methods prove the superiority of the proposed method. Contributions of the proposed method can be summarized as follows: 1) joint dictionary learning is proposed to explore the intrinsic information of different images for change detection. In this case, change detection can be transformed as a sparse representation problem. To the authors' knowledge, few publications utilize joint learning dictionary in change detection; 2) an automatic threshold selection strategy is presented, which minimizes the reconstruction errors of the changed pixels without the prior assumption of the spectral signature. As a result, the threshold value provided by the proposed method can adapt to different data due to the characteristic of joint dictionary learning; and 3) the proposed method makes no prior assumption of the modeling and the handling of the spectral signature, which can be adapted to different data. Xiaoqiang Lu, Yuan Yuan 0001, Xiangtao Zheng |
IEEE Trans. Cybern. | 3 |
| 2017 | Remote Sensing Scene Classification by Unsupervised Representation LearningabstractWith the rapid development of the satellite sensor technology, high spatial resolution remote sensing (HSR) data have attracted extensive attention in military and civilian applications. In order to make full use of these data, remote sensing scene classification becomes an important and necessary precedent task. In this paper, an unsupervised representation learning method is proposed to investigate deconvolution networks for remote sensing scene classification. First, a shallow weighted deconvolution network is utilized to learn a set of feature maps and filters for each image by minimizing the reconstruction error between the input image and the convolution result. The learned feature maps can capture the abundant edge and texture information of high spatial resolution images, which is definitely important for remote sensing images. After that, the spatial pyramid model (SPM) is used to aggregate features at different scales to maintain the spatial layout of HSR image scene. A discriminative representation for HSR image is obtained by combining the proposed weighted deconvolution model and SPM. Finally, the representation vector is input into a support vector machine to finish classification. We apply our method on two challenging HSR image data sets: the UCMerced data set with 21 scene categories and the Sydney data set with seven land-use categories. All the experimental results achieved by the proposed method outperform most state of the arts, which demonstrates the effectiveness of the proposed method. Xiaoqiang Lu, Xiangtao Zheng, Yuan Yuan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2017 | Dimensionality Reduction by Spatial-Spectral Preservation in Selected BandsabstractDimensionality reduction (DR) has attracted extensive attention since it provides discriminative information of hyperspectral images (HSI) and reduces the computational burden. Though DR has gained rapid development in recent years, it is difficult to achieve higher classification accuracy while preserving the relevant original information of the spectral bands. To relieve this limitation, in this paper, a different DR framework is proposed to perform feature extraction on the selected bands. The proposed method uses determinantal point process to select the representative bands and to preserve the relevant original information of the spectral bands. The performance of classification is further improved by performing multiple Laplacian eigenmaps (LEs) on the selected bands. Different from the traditional LEs, multiple Laplacian matrices in this paper are defined by encoding spatial-spectral proximity on each band. A common low-dimensional representation is generated to capture the joint manifold structure from multiple Laplacian matrices. Experimental results on three real-world HSIs demonstrate that the proposed framework can lead to a significant advancement in HSI classification compared with the state-of-the-art methods. Xiangtao Zheng, Yuan Yuan 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2017 | Latent Semantic Minimal Hashing for Image RetrievalabstractHashing-based similarity search is an important technique for large-scale query-by-example image retrieval system, since it provides fast search with computation and memory efficiency. However, it is a challenge work to design compact codes to represent original features with good performance. Recently, a lot of unsupervised hashing methods have been proposed to focus on preserving geometric structure similarity of the data in the original feature space, but they have not yet fully refined image features and explored the latent semantic feature embedding in the data simultaneously. To address the problem, in this paper, a novel joint binary codes learning method is proposed to combine image feature to latent semantic feature with minimum encoding loss, which is referred as latent semantic minimal hashing. The latent semantic feature is learned based on matrix decomposition to refine original feature, thereby it makes the learned feature more discriminative. Moreover, a minimum encoding loss is combined with latent semantic feature learning process simultaneously, so as to guarantee the obtained binary codes are discriminative as well. Extensive experiments on several well-known large databases demonstrate that the proposed method outperforms most state-of-the-art hashing methods. Xiaoqiang Lu, Xiangtao Zheng, Xuelong Li 0001 |
IEEE Trans. Image Process. | 2 |
| 2017 | Discovering Diverse Subset for Unsupervised Hyperspectral Band SelectionabstractBand selection, as a special case of the feature selection problem, tries to remove redundant bands and select a few important bands to represent the whole image cube. This has attracted much attention, since the selected bands provide discriminative information for further applications and reduce the computational burden. Though hyperspectral band selection has gained rapid development in recent years, it is still a challenging task because of the following requirements: 1) an effective model can capture the underlying relations between different high-dimensional spectral bands; 2) a fast and robust measure function can adapt to general hyperspectral tasks; and 3) an efficient search strategy can find the desired selected bands in reasonable computational time. To satisfy these requirements, a multigraph determinantal point process (MDPP) model is proposed to capture the full structure between different bands and efficiently find the optimal band subset in extensive hyperspectral applications. There are three main contributions: 1) graphical model is naturally transferred to address band selection problem by the proposed MDPP; 2) multiple graphs are designed to capture the intrinsic relationships between hyperspectral bands; and 3) mixture DPP is proposed to model the multiple dependencies in the proposed multiple graphs, and offers an efficient search strategy to select the optimal bands. To verify the superiority of the proposed method, experiments have been conducted on three hyperspectral applications, such as hyperspectral classification, anomaly detection, and target detection. The reliability of the proposed method in generic hyperspectral tasks is experimentally proved on four real-world hyperspectral data sets. Yuan Yuan 0001, Xiangtao Zheng, Xiaoqiang Lu |
IEEE Trans. Image Process. | 2 |
| 2016 | A target detection method for hyperspectral image based on mixture noise model
Xiangtao Zheng, Yuan Yuan 0001, Xiaoqiang Lu |
Neurocomputing | 1 |
| 2016 | A discriminative representation for human action recognition
Yuan Yuan 0001, Xiangtao Zheng, Xiaoqiang Lu |
Pattern Recognit. | 2 |
| 2015 | Spectral-Spatial Kernel Regularized for Hyperspectral Image DenoisingabstractNoise contamination is a ubiquitous problem in hyperspectral images (HSIs), which is a challenging and promising theme in many remote sensing applications. A large number of methods have been proposed to remove noise. Unfortunately, most denoising methods fail to take full advantages of the high spectral correlation and to simultaneously consider the specific noise distributions in HSIs. Recently, a spectral-spatial adaptive hyperspectral total variation (SSAHTV) was proposed and obtained promising results. However, the SSAHTV model is insensitive to the image details, which makes the edges blur. To overcome all of these drawbacks, a spectral-spatial kernel method for HSI denoising is proposed in this paper. The proposed method is inspired by the observation that the spectral-spatial information is highly redundant in HSIs, which is sufficient to estimate the clear images. In this paper, a spectral-spatial kernel regularization is proposed to maintain the spectral correlations in spectral dimension and to match the original structure between two spatial dimensions. Moreover, an adaptive mechanism is developed to balance the fidelity term according to different noise distributions in each band. Therefore, it cannot only suppress noise in the high-noise band but also preserve information in the low-noise band. The reliability of the proposed method in removing noise is experimentally proved on both simulated data and real data. Yuan Yuan 0001, Xiangtao Zheng, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 2 |