Xiaoqiang Lu

dblp:20/8032 · DBLP profile ↗
← Back
208ranked-venue papers
46as first author
96since 2021 · last 2026
0000-0002-7037-5188ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 97 · 19 first-author · 64 since 2021Artificial intelligence and machine learning · 66 · 19 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 46 · 10 first-author · 12 since 2021Computer networks · 6 · 4 since 2021Systems, architecture and hardware · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 ThermoSplat: Cross-modal 3D Gaussian splatting with feature modulation and geometry decoupling
Zhaoqi Su, Shihai Chen, Xinyan Lin, Liqin Huang, Zhipeng Su, Xiaoqiang Lu
Neurocomputing6
2026 Like Human Rethinking: Contour Transformer AutoRegression for Referring Remote Sensing Interpretation
abstract
Referring remote sensing interpretation holds significant application value in various scenarios such as ecological protection, resource exploration, and emergency management. However, referring remote sensing expression comprehension and segmentation (RRSECS) faces critical challenges, including micro-target localization drift problem caused by insufficient extraction of boundary features in existing paradigms. Moreover, when transferred to remote sensing domains, polygon-based methods encounter issues such as contour-boundary misalignment and multi-task co-optimization conflicts problems. In this paper, we propose SeeFormer, a novel contour autoregressive paradigm specifically designed for RRSECS, which accurately locates and segments micro, irregular targets in remote sensing imagery. We first introduce a brain-inspired feature refocus learning (BIFRL) module that progressively attends to effective object features via a coarse-to-fine scheme, significantly boosting small-object localization and segmentation. Next, we present a language-contour enhancer (LCE) that injects shape-aware contour priors, and a corner-based contour sampler (CBCS) to improve mask-polygon reconstruction fidelity. Finally, we develop an autoregressive dual-decoder paradigm (ARDDP) that preserves sequence consistency while alleviating multi-task optimization conflicts. Extensive experiments on RefDIOR, RRSISD, and OPTRSVG datasets under varying scenarios, scales, and task paradigms demonstrate transformative performance gains: compared to the baseline PolyFormer, our proposed SeeFormer improves oIoU and mIoU by 27.58% and 39.37% for referring image segmentation and by 18.94% and 28.90% for visual grounding on the RefDIOR dataset.
Jinming Chai, Licheng Jiao, Xiaoqiang Lu, Lingling Li 0002, Fang Liu 0001, Xu Liu 0006, Wenping Ma 0001, Weibin Li 0002
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 PC2F: Language-guided Progressive Calibration and Cascade Filtering for Remote Sensing Visual Grounding
Licheng Jiao, Xu Liu 0006, Xiaoqiang Lu, Shuo Li 0010, Xiaolin Tian 0002
Pattern Recognit.6
2026 ReCoTR: Reducing Semantic Cognitive Shift via Dual-Consensus Token Compression for Remote Sensing Image-Text Retrieval
abstract
With the rapid advancement of vision-language models (VLMs) in general-purpose settings, their application to cross-modal retrieval and semantic understanding of large-scale multimodal remote sensing (RS) data is emerging as a key enabler for urban governance, environmental monitoring, and disaster response. However, the pervasive issue of semantic shift in RS image poses a significant challenge to the transferability of pre-trained VLMs. To address this limitation, we propose ReCoTR, an enhanced CLIP-based cross-modal retrieval framework tailored for remote sensing applications. ReCoTR tackles region-level granularity bias and contextual semantic drift through a Dual Consensus Token Evaluation (DCTE) module, which leverages a mixture-of-experts strategy to fuse inter-modal semantic consensus with intra-modal structural consistency, enabling fine-grained estimation of semantic confidence for visual tokens. Moreover, to mitigate representational contamination caused by background noise, we introduce the Semantic Confidence Token Compression (SCTC) module. This module selectively filters and aggregates tokens with high semantic relevance, thus reducing redundancy and alleviating the noise amplification inherent in CLIP's average pooling. Experimental results on three benchmark RS cross-modal retrieval datasets demonstrate that ReCoTR consistently outperforms existing methods on bidirectional image-text retrieval tasks, validating its effectiveness and robustness in remote sensing semantic alignment scenarios. Our source codes are available at: https://github.com/Jerry710/ReCoTR.git.
Jirui Huang, Yaxiong Chen, Chuang Du, Shengwu Xiong 0001, Xiaoqiang Lu
IEEE Trans. Image Process.5
2026 Reconstruction-Contrast Coupling Learning for Open-Set Semi-Supervised Hyperspectral Image Classification
abstract
Although numerous semi-supervised learning methods have been elaborately designed for hyperspectral image (HSI) classification, most existing semi-supervised learning paradigms still rely on a closed-set assumption. These methods implicitly assume that the category spaces of labeled and unlabeled samples are completely aligned, that is, all unlabeled samples must belong to a pre-defined known category set. However, the closed-set assumption is particularly problematic in practical remote sensing scenarios because partial unlabeled data inevitably belong to unknown categories. To address this challenge, this paper proposes a reconstruction-contrast coupling learning (ReCo2L) method for open-set semi-supervised HSI classification, fully leveraging the complementarity between masked feature reconstruction learning and contrastive learning to enhance the encoder’s local detail sensitivity and global discriminative ability. Specifically, we first apply a masked feature reconstruction learning with an adaptive masking strategy to enhance the encoder’s ability to capture local details by high-quality spectral-spatial feature reconstruction. Then, we employ contrastive learning to strengthen the encoder’s capability to extract global characteristics by pulling semantically similar samples closer and pushing dissimilar ones farther apart in the feature space. Finally, a pixel-prototype deviation loss is proposed to further improve both inter-category distinguishability and intra-category compactness by reducing the distances between labeled sample features and their corresponding class anchors. Extensive experiments on three benchmark datasets demonstrate that our proposed ReCo2L achieves superior classification performance in both known and unknown categories and significantly surpasses 10 state-of-the-art HSI classification methods. The code will be available at https://github.com/repository-AI-chen/ReCo2L.
Hao Sun 0014, Renyi Chen, Yong Chen 0024, Wenjing Chen 0003, Wei Xie 0008, Xiaoqiang Lu
IEEE Trans. Image Process.6
2025 Multi-HAP-Assisted Computation Offloading in Space-Air-Ground-Sea Integrated Network
abstract
The growth of maritime activities has boosted the demand for efficient marine communications and computation offloading. Mobile edge computing (MEC)-powered space-air–ground-sea integrated network (SAGSIN) is a promising solution to satisfy these demands. To the best of our knowledge, the existing research on high altitude platforms (HAPs) in SAGSIN remains open. To exploit the HAPs’ advantages of facilitating communication at shorter distances and more stable computing services than terrestrial base stations, this work considers multi-HAP-assisted computation offloading in SAGSIN and formulates an optimization problem, which turns out to be a mixed integer nonlinear programming (MINLP) due to the joint optimization among task-HAP association, computational resource allocation of HAPs and task-satellite association. To this end, the problem is decoupled into three single-variable subproblems by relaxing the delay constraint, and the subproblems are solved by graph theory, the Lagrange multiplier method with Karush-Kuhn–Tucker (KKT) constraints, and the alternating optimization, respectively. Simulation results demonstrate the efficiency of the proposed scheme with superior performances compared to benchmark schemes.
Wei Feng 0001, Yi Fang 0005, Zhijian Lin, Xiaoqiang Lu
IEEE Internet Things J.5
2025 AoI Energy-Efficient Edge Caching in AAV-Assisted Vehicular Networks
abstract
Mobile edge caching (MEC) has grown substantially with the rapid development in scale and complexity of data traffic. By exploiting the expansive coverage of autonomous aerial vehicles (AAVs), MEC enables services for massive vehicle users (VUs) simultaneously, which is promising for enhancing network transmission efficiency. Nonetheless, due to challenges arising from the timeliness and freshness of content services caused by AAVs’ limited endurance and airborne capacity, caching strategy considering the real-time of content in large-scale dynamic Internet of Vehicles (IoV) environments remains open. With the above consideration, in this article, the cache refreshing cycle and content placement are jointly optimized in the cache-enabled AAV-assisted vehicular integrated networks (CAVINs) to minimize the content Age of Information (AoI) and energy consumption of the macro AAV. Since the joint optimization problem is variational coupled with nonconvex binary constraints, it is decoupled and solved by a double-iteration method. Specifically, the optimal cache refreshing cycle is derived in semi-closed form with the Karush-Kuhn-Tucker (KKT) conditions. The locally optimal solution of the content placement is obtained through successive convex approximation (SCA). Simulation results corroborate the effectiveness and superiority of the proposed scheme.
Yang Xiao 0014, Zhijian Lin, Xiaoxiao Cao, Youjia Chen, Xiaoqiang Lu
IEEE Internet Things J.5
2025 Analysis and Design of a Discrete-Time 3-0 MASH Delta-Sigma ADC With 100.2 dB Dynamic Range
abstract
This paper describes the analysis and design of a discrete-time (DT) fully dynamic 3-0 multi-stage noise-shaping (MASH) delta-sigma ($\Delta \Sigma $) analog-to-digital converter (ADC). Through system-level analysis, error source analysis, nonlinearity analysis and modeling of the integrators, and detailed considerations for circuit implementation, the trade-offs between design parameters in the 3-0 MASH$\Delta \Sigma $ADC were evaluated. The proposed ADC is fabricated and measured in a 180 nm CMOS process, achieving a DR, peak SNDR, and SFDR of 100.2 dB, 98.5 dB, and 116.7 dB, respectively, within a 2.56 kHz bandwidth, consuming only$20.1~\mu $W. As a result, the Schreier figure-of-merit (FoM) for SNDR and DR are 179.6 dB and 181.3 dB, respectively. The measurement results of the prototype 3-0 MASH$\Delta \Sigma $ADC closely matched the theoretical predictions. This consistency between the measurements and the theoretical analysis confirms the reliability of the design approach in achieving the expected performance.
Cong Wei 0002, Rongshan Wei, Xiaoqiang Lu, Zhichao Tan
IEEE Trans. Circuits Syst. I Regul. Pap.4
2025 Relation-Aware Multiprototype Learning for Semi-Supervised Hyperspectral Image Classification
Wenjing Chen 0003, Renyi Chen, Zhiwei Ye, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.4
2025 Bow Direction Detection Based on Angular Coding With Heading Intersection Over Union Loss
abstract
Accurate bow direction detection is essential for ship trajectory prediction and port monitoring. Existing ship detection networks typically output angles within 180°, while extending to 360° introduces cyclic issues affecting rotation intersection over union (RIoU) accuracy. This study proposes a novel bow direction detection algorithm that extends network output to 360° and integrates a heading intersection over union (HIoU) loss to enhance detection accuracy and robustness. Additionally, an HIoU loss function is designed to improve bow direction identification and reduce quantization errors in hash codes. The algorithm is evaluated on three datasets: FGSD, OHD-SJTU-S, and OHD-SJTU-L. On FGSD, it achieves mean average precision (mAP) of 91.14%. On OHD-SJTU-S, it attains an$\text {mAP}_{50:95}$of 63.3% and a bow direction prediction accuracy of 90.7%. On OHD-SJTU-L, the$\text {mAP}_{50:95}$is 29.2%, with an accuracy of 80.2%.
Yaxiong Chen, Qiangqiang Huang, Hao Sun 0014, Shengwu Xiong 0001, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.6
2025 Bilinear Parallel Fourier Transformer for Multimodal Remote Sensing Classification
abstract
Vision Transformers (ViTs) have shown promise in multimodal fusion image classification, yet face performance challenges in complex remote sensing scenarios. Single fusion frameworks often fail to fully utilize multimodal diversity, and the uneven distribution of image categories complicates the accurate construction of spatial structures by Transformers. Additionally, traditional cross-entropy tends to favor majority classes, neglecting minority classes, resulting in suboptimal predictions and reduced overall accuracy (OA). To solve these challenges, we propose a novel deep neural network, a bilinear parallel Fourier Transformer (BPFT). We propose a novel dual-fusion feature interaction (DFFI) module that utilizes two distinct types of fused features for learning, namely the spatial-spectral fusion feature and the global fusion feature. Besides, we introduce a dual-feature interaction (DFI) module to improve the utilization of fused feature information. To enable the Transformer to better establish spatial structural relationships, we employ the Fourier transform in place of the self-attention mechanism. To address the focus on minority class labels, we propose an exponential label smoothing cross-entropy loss function. This loss function comprises two components: exponential cross-entropy and label smoothing. The exponential cross-entropy component applies a strong penalty to misclassified samples, thereby increasing attention on minority class labels. To validate the efficacy of our approach, extensive experiments are conducted across two multimodal remote sensing datasets: Augsburg and Berlin, encompassing hyperspectral imaging (HSI) data and synthetic aperture radar (SAR) data. The results of these experiments affirm the superior performance of our proposed BPFT model compared to existing state-of-the-art models in multimodal remote sensing image classification tasks.
Yaxiong Chen, Qicong Wang, Yichen Zhao, Shengwu Xiong 0001, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.5
2025 Global-Local Fusion With Semantic Information Guidance for Accurate Small Object Detection in UAV Aerial Images
abstract
In recent years, the rapid development of the unmanned aerial vehicle (UAV) technology has generated a large number of aerial photography images captured by UAV. Consequently, the object detection in UAV aerial images has emerged as a recent research focus. However, due to the flexible flight heights and diverse shooting angles of UAV, two significant challenges have arisen in UAV aerial images: extreme variation in target scale and the presence of numerous small targets. To address these challenges, this article introduces a semantic information-guided fusion module specifically tailored for small targets. This module utilizes high-level semantic information to guide and align the underlying texture information, thereby enhancing the semantic representation of small targets at the feature level and subsequently improving the model’s ability to detect them. In addition, this article introduces a novel global–local fusion detection strategy to strengthen the detection of small targets. We have redesigned the foreground region assembly method to address the drawbacks of previous methods that involved multiple inferences. Extensive experiments conducted on the VisDrone and UAVDT datasets demonstrate that our two self-designed modules can significantly enhance the detection capability of small targets compared with the YOLOX-M model. Our code is publicly available at:https://github.com/LearnYZZ/GLSDet.
Yaxiong Chen, Zhengze Ye, Haokai Sun 0001, Tengfei Gong, Shengwu Xiong 0001, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.6
2025 Context-Aware Local-Global Semantic Alignment for Remote Sensing Image-Text Retrieval
abstract
Remote sensing image-text retrieval (RSITR) is a cross-modal task that integrates visual and textual information, attracting significant attention in remote sensing research. Remote sensing images typically contain complex scenes with abundant details, presenting significant challenges for accurate semantic alignment between images and texts. Despite advances in the field, achieving precise alignment in such intricate contexts remains a major hurdle. To address this challenge, this article introduces a novel context-aware local-global semantic alignment (CLGSA) method. The proposed method consists of two key modules: the local key feature alignment (LKFA) module and the cross-sample global semantic alignment (CGSA) module. The LKFA module incorporates a local image masking and reconstruction task to improve the alignment between image and text features. Specifically, this module masks certain regions of the image and uses text context information to guide the reconstruction of the masked areas, enhancing the alignment of local semantics and ensuring more accurate retrieval of region-specific content. The CGSA module employs a hard sample triplet loss to improve global semantic consistency. By prioritizing difficult samples during training, this module refines feature space distributions, helping the model better capture global semantics across the entire image-text pair. A series of extensive experiments demonstrates the effectiveness of the proposed method. The method achieves an mR score of 32.07% on the RSICD dataset and 46.63% on the RSITMD dataset, outperforming baseline methods and confirming the robustness and accuracy of the approach.
Xiumei Chen, Xiangtao Zheng, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.3
2025 Relevance-Guided Adaptive Learning for Remote Sensing Image-Text Retrieval
abstract
The remote sensing image–text retrieval (RSITR) aims to establish semantic alignment between images and texts to enable accurate cross-modal retrieval. Existing methods usually extract features from images and texts independently, aligning them in a shared embedding space to achieve cross-modal retrieval. However, these methods often assume complete alignment between image and text pairs, overlooking the inherent disparities between the rich visual details in remote sensing (RS) images and the abstract nature of textual descriptions. These disparities result in image–text pairs only sharing partial semantic correlations, rather than one-to-one complete alignment. Such incomplete alignment adversely affects model training and retrieval accuracy. To address this problem, a relevance-guided adaptive learning (RGAL) method is proposed, which quantifies and leverages the relevance of image–text pairs to refine the training process while enhancing retrieval performance. First, the proposed method introduces an image–text relevance measurement mechanism that integrates global and local feature distances to accurately evaluate the degree of semantic relevance between images and texts. Second, a relevance-based sample division (RBSD) strategy is proposed, utilizing a Gaussian mixture model to dynamically redivide samples into positive and negative pairs according to the measured image–text relevance. This strategy refines the training dataset, reduces noise, and enhances the effectiveness of model learning. Finally, a relevance-weighted triplet loss (RWTL) is designed to adaptively adjust the contribution of sample pairs to the loss function based on their relevance, further optimizing model training and enhancing retrieval accuracy. Experimental results on multiple RSITR datasets demonstrate that the proposed method significantly improves retrieval accuracy and performance.
Xiumei Chen, Xiangtao Zheng, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.3
2025 VGRSS: Datasets and Models for Visual Grounding in Remote Sensing Ship Images
abstract
This paper introduces a task named Visual Grounding of Remote Sensing Ship Images (VGRSS). The goal of VGRSS is to locate ship objects in remote sensing images guided by natural language. Extensive research has been conducted on multimodal processing of remote sensing images and text to retrieve rich information from remote sensing images using natural language. However, due to the unique characteristics of remote sensing ship images, ship localization using natural language remains a challenge. Therefore, in this work, we construct datasets for the VGRSS task and explore deep learning models. Specifically, our contributions can be summarized as follows: First, we construct two remote sensing ship datasets for visual grounding. One is based on the optical remote sensing dataset, named RSSVG, while the other is based on the synthetic aperture radar (SAR) dataset, named SARVG. Second, we propose a Language-Guided Visual Feature Enhancement (LVFE) module. This module enhances visual features through language guidance before Visual-Linguistic Fusion. Third, we propose a Visual-Linguistic Fusion (VLF) module based on multimodal feature stacking. This module inputs the stacked language and visual features, and then performs feature fusion using a Transformer, enabling effective cross-modal interaction and integration. Fourth, we introduce a novel loss calculation method by incorporating Enhanced Intersection over Union (EIoU) into the loss function. Finally, we benchmark extensive state-of-the-art (SOTA) natural image visual grounding methods on the constructed RSSVG and SARVG datasets, then provide insightful analysis based on the results. This work offers valuable insights for developing better VGRSS models.
Yaxiong Chen, Liwen Zhan, Yichen Zhao, Shengwu Xiong 0001, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.5
2025 A Dual-Stage Wavelet and Linear Attention Enhancement Network for Agricultural Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) classification faces unique challenges in agricultural scenario due to spectral-spatial feature similarity caused by complex planting structures and high spectral similarity. Existing spatial-spectral joint feature extraction methods fail to fully exploit the advantages of spatial and spectral information, thus have certain limitations and cannot effectively distinguish similar crops in agricultural scenarios. To address these limitations, we proposed a dual-stage wavelet and linear attention enhancement network (DSW-LAN) for agricultural HSI classification, addressing the challenges of complex spatial-spectral information and high redundancy. We integrates a direction factorized deformable 3D Convolution (DFDWConv3D) module to capture multi-scale spatial-spectral features through adaptive kernel adjustments, while wavelet transform decomposes spatial features into low-frequency (structural) and high-frequency (textural) components for targeted enhancement. Additionally, a spectral probe-guided linear attention mechanism efficiently models long-range spectral dependencies with reduced computational complexity by prioritizing discriminative bands. Experimental results demonstrate superior performance on three challenging agricultural HSI datasets, achieving enhanced classification accuracy with reduced computational complexity.
Yaxiong Chen, Bo Zhang 0069, Shili Xiong, Shengwu Xiong 0001, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.6
2025 Discover the Unknown Ones in Fine-Grained Ship Detection
abstract
Remote sensing image-based ship identification technology has great applications in areas such as national defense and fishery management. However, existing remote sensing ship studies mainly focus on a closed environment and overlook actual sea conditions, while new military ships will be encountered. These unknown categories of ships will be ignored or misclassified by existing models, dramatically affecting the accurate assessment of the maritime situation. Furthermore, existing unknown detection methods for natural images fail to tackle the remote sensing ship detection problem for the property of high similarity in overall appearance. To cope with this problem, this paper proposes a fine-grained unknown ship detection network. Firstly, we explore a class-balanced proposal sampler to avoid inefficient information learning. Secondly, we propose a finegrained memory bank-based contrastive learning strategy to separate different categories. Finally, to further separate unknown classes, we adopt an uncertainty-aware unknown learner with logit to reduce the uncertainty of fine-grained predictions. Experiments conducted in three public ship detection datasets ShipRSImageNet, DOSR, and HRSC2016 show that the method not only achieves good detection on unknown class ships, but also improves the detection accuracy on known classes. The code is available at https://github.com/FoRGEU/DUONet.
Tengfei Gong, Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.5
2025 Multibranch Fusion-Based Feature Enhance for Remote-Sensing Scene Classification
abstract
Remote-sensing (RS) scene classification is a fundamental and significant task in RS image interpretation, involving the annotation of semantic content. RS scene images are characterized by complex backgrounds, rich content, and multiscale targets, exhibiting both intraclass separation and interclass convergence. Therefore, extracting features that effectively express the intrinsic attributes of images and possess high discriminative is crucial for RS scene classification. Existing global-based methods often lack the ability to capture significant detailed information in similar scenes. Conversely, methods based on local discriminative features tend to overlook the interrelationships of objects within the same scene. To address these issues, this article proposes a unified framework named MBFNet to align and fuse features of different scales and levels for accurate RS scene classification. We utilize a multibranch feature-extracting network structure with parallel convolution and Transformer modules. Simultaneously, a kernel-selected multiscale aggregation (KSMSA) module is designed to efficiently process the diverse scale features emanating from these parallel branches. By selecting different convolution kernels, a dynamic receptive field is established to adaptively process features of different scales, reducing semantic differences to achieve effective aggregation of multiscale features. Moreover, a learnable multilevel aggregation (LMLA) module is designed to integrate shallow features, such as shape information, into deep features for more comprehensive feature fusion. Benefiting from KSMSA and LMLA, the proposed MBFNet improves the discriminability of features, thereby enhancing classification performance. Comprehensive experiments on three benchmark datasets demonstrate that the proposed method outperforms state-of-the-art RS scene classification methods in terms of performance.
Xiongbo Lu, Meng Yang 0034, Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.5
2025 LSCF: Long-Term Semantic-Guidance ConvFormer for Referring Remote Sensing Image Segmentation
abstract
Referring Remote Sensing Image Segmentation (RRSIS) task aims to generate segmentation masks for target objects based on language descriptions. It requires precise localization while distinguishing between visually similar yet semantically distinct objects. Fusing vision-language features only during extraction causes information loss and semantic forgetting in the decoder, harming similar target distinction. Additionally, high-resolution remote sensing images present challenges, including complex backgrounds, diverse object scales, and intricate boundaries, limiting the effectiveness of previous methods. To address these issues, we propose the Long-term Semantic-guidance ConvFormer (LSCF) Network. First, we fuse multi-receptive-field local features extracted by the Multi-scale CoordConv (MCC) module with language-aware global features from the Cross-modal Attention (CA) module to obtain multi-modal representations. Second, the Sampling Attention (SA) module enables fine-grained vision context alignment under semantic guidance. Finally, the Global Language Fusion (GLF) module is incorporated in the decoder to maintain long-term vision-language alignment and mitigate semantic degradation. Experimental validation on the RefSegRS, RRSIS-D, and RISBench datasets demonstrates that LSCF achieves oIoU scores of 83.27%, 77.42%, and 74.88%, and mIoU scores of 77.44%, 64.25%, and 68.53%, respectively. On RefSegRS, LSCF surpasses the SOTA method FIANet by 5.53% (oIoU) and 9.58% (mIoU), while delivering competitive performance on RRSIS-D and RISBench. Code and experimental configurations will be released.
Lingling Li 0002, Xiaoqiang Lu, Licheng Jiao, Fang Liu 0001, Wenping Ma 0001, Xu Liu 0006
IEEE Trans. Geosci. Remote. Sens.3
2025 Learning Positive-Negative Prompts for Open-Set Remote Sensing Scene Classification
Hao Sun 0014, Hanlizi Chen, Wenjing Chen 0003, Chengji Wang, Wei Xie 0008, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.6
2025 Class-Aware Consistency Learning for Open-Set Semi-Supervised Hyperspectral Image Classification
abstract
Semi-supervised hyperspectral image (HSI) classification methods focus on exploring the spectral and spatial information of unlabeled samples. However, existing methods generally follow the closed-set setting, assuming that unlabeled samples do not contain novel classes, which is hard to hold in practical applications. This paper aims to study semi-supervised HSI classification in the open-set setting, i.e., unlabeled samples fall into novel classes, and proposes a class-aware consistency learning (CACL) method. First, to explore discriminative spectral-spatial features, a position-aware transformer is developed, which effectively models spatial position priors between the center pixel and its neighboring pixels via a symmetric position-aware encoding. Then, to reduce the interference from novel class samples on the model’s discrimination, a prototype-driven consistency learning is proposed, which accurately selects unlabeled samples belonging to known classes via a known class sampler, and efficiently utilizes their spectral-spatial information by modeling consistent predictions across different views. Finally, to further improve the distinguishability between known classes, a prototype contrastive optimization is proposed to decrease the distance between samples from the same class and increase the distances between those from different classes in the feature domain. Furthermore, an adaptive segmentation threshold is designed to accurately predict known classes and reject novel classes. Extensive experiments verify that our CACL outperforms the state-of-the-art methods, achieving the overall accuracy of 81.65%, 88.61%, and 92.88% on the Indian Pines, Salinas, and Pavia University datasets, with 10 labeled samples in each known class. The code is available at https://github.com/rock-in/CACL-main.
Hao Sun 0014, Renyi Chen, Huaxiong Yao, Yaxiong Chen, Wei Xie 0008, Guirong Feng, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.8
2025 Novel Class Discovery for Hyperspectral Image via Class-Relation Perceptive Distillation With Prototype-Level Clustering Prediction
abstract
Confronted with the increasing emergency of hyperspectral remote sensing categories in the dynamic environment, traditional classification models that depend on fixed-category labeled data encounter difficulties on new classes recognition. Novel class discovery (NCD) aims to discover unknown class-disjoint novel classes in an unlabeled dataset with the pre-existing knowledge of known classes. Notably, the critical goal of NCD is to ensure recognition accuracy of known classes while identifying new ones. In this paper, we propose a class-relation perceptive distillation with prototype-level clustering prediction network (CRPD-PCP) for NCD of hyperspectral image (HSI). The proposed framework comprises an initial training stage (ITS) and a novel class discovery stage (NCDS) with two essential modules. Specifically, we present the class relation perceptive distillation (CRPD) module, which imposes a similarity constraint on the prediction of the distribution of new class data over the models of two stages. With the CRPD operated on the NCDS, our model effectively captures class relation information in spectral-spatial domain between known and novel classes of HSI to avoid forgetting old knowledge. Besides, we establish the prototype-level clustering prediction (PCP) module to generate high-confidence pseudo-labels for unlabeled novel classes. To be specific, we progressively cluster samples with the same spectral angular distance from the perspective of prototypes, and the self-supervised prototype-level knowledge distillation strategy in PCP facilitates effective identification of new categories. Experiments conducted on four datasets demonstrate that the CRPD-PCP model generates superior performance compared to other NCD methods for HSI. Our code will be released at https://github.com/Chirsycy/CRPD-PCP.git.
Chunyan Yu, Xiaowen Zhao, Yulei Wang 0002, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.4
2025 Concern With Center-Pixel Labeling: Center-Specific Perception Transformer Network for Hyperspectral Image Classification
abstract
Self-attention-based approaches that leverage global context information for hyperspectral image (HSI) classification have gained increasing prominence. Nevertheless, due to the assignment of equivalent attention weight to all the tokens (pixels or patches), the existing self-attention mechanism inadvertently prioritizes the non-label-specified information over the instinct label-specified information, which generates attention shifts and redundancy in HSI classification. To alleviate the mentioned barrier, we propose the center-specific perception transformer network (CP-Transformer), which is the first attempt to perform class-guided attention and filter interference factors for HSI classification feature representation. Specifically, the central-pixel focus attention module (CFA) is presented to compute the label-related attention between the center and other pixels. In this manner, CFA reduces computational complexity and closely aligns with the center-pixel labeling strategy. Besides, the spectral saliency focus attention module (SSFA) is developed to capture the spectral correlation by focusing salient bands to provide a beneficial supplement for spatial features. Moreover, the hierarchical integration network (HIN) constructs the inference network to integrate and rectify spatial-spectral features for HSI classification. The experiment results on four popular HSI datasets demonstrate that the proposed method achieves robust performance compared to other state-of-the-art methods. Our code will be released at https://github.com/Chirsycy/CP-Transformer.
Chunyan Yu, Yuanchen Zhu, Yulei Wang 0002, Enyu Zhao, Qiang Zhang 0011, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.6
2025 SSPNet: Spatial-Spectral Perception Network for Mineral Hyperspectral Image Classification
abstract
Unlike general scenes, mineral hyperspectral images often exhibit similar spatial and spectral characteristics across different mines, making traditional classification methods less effective due to compromised robustness. To address this, we propose a Spatial-Spectral Perception Network for mineral hyperspectral image classification. This approach divides spatial-spectral feature extraction into two stages. In the spatial feature perception stage, we introduce a Spatial Frequency Perceptron that maps three-dimensional spatial features into low-frequency and high-frequency domains. We then apply Triple-Cross-Attention to each frequency domain to better differentiate spatial features of similar mines. In the spectral perception stage, we design a Spectral Linear Perceptron using Absolute Linear Attention, which captures fine-grained spectral differences by establishing internal relationships between spectral features through Absolute Positional Weighting. This enables effective separation of similar spectra for final classification. Extensive experiments on three publicly available mineral hyperspectral image datasets and one agricultural hyperspectral dataset show that our method outperforms popular alternatives in both effectiveness and robustness. The open-source code can be accessed at https://github.com/WUTCM-Lab/SSPNet.
Bo Zhang 0069, Yaxiong Chen, Ruilin Yao, Shili Xiong, Shengwu Xiong 0001, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.6
2025 Hyperspectral Image Classification via Cascaded Spatial Cross-Attention Network
abstract
In hyperspectral images (HSIs), different land cover (LC) classes have distinct reflective characteristics at various wavelengths. Therefore, relying on only a few bands to distinguish all LC classes often leads to information loss, resulting in poor average accuracy. To address this problem, we propose a method called Cascaded Spatial Cross-Attention Network (CSCANet) for HSI classification. We design a cascaded spatial cross-attention module, which first performs cross-attention on local and global features in the spatial context, then uses a group cascade structure to sequentially propagate important spatial regions within the different channels, and finally obtains joint attention features to improve the robustness of the network. Moreover, we also design a two-branch feature separation structure based on spatial-spectral features to separate different LC Tokens as much as possible, thereby improving the distinguishability of different LC classes. Extensive experiments demonstrate that our method achieves excellent performance in enhancing classification accuracy and robustness. The source code can be obtained from https://github.com/WUTCM-Lab/CSCANet.
Bo Zhang 0069, Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu
IEEE Trans. Image Process.4
2025 HybridRDN: Delay-Optimal Computation Offloading for Autonomous Vehicle Fleets Based on RSMA
abstract
Rate-splitting multiple access (RSMA), space division multiple access (SDMA), and non-orthogonal multiple access (NOMA) have gained significant popularity and are extensively utilized across various domains. However, it is still unclear whether hybridRSMA-SDMA-NOMA (HybridRDN) would seamlessly combine the advantages of RSMA, SDMA, and NOMA to contribute to the computation offloading of autonomous vehicle systems. To address the above issue, this paper introduces a novel HybridRDN-assisted computation offloading fleet (COF) scheme tailored for autonomous vehicle systems. First, we propose a stochastic-geometry-aided method to model the offloading framework. Afterwards, the task vehicles (TVs) ingeniously employ the proposed HybridRDN scheme to offload tasks to the resource vehicles (RVs) in each COF to relieve their computational burden. Diverging from the sole optimization of the task segmentation ratio or the transmission rate, a joint optimization problem involving the transmission weighting factor, the HybridRDN precoding matrix, the common rate, and the task segmentation ratio, is formulated, which aims to minimize the average delay of the COF system while approaching the rate performance of the ideal HybridRDN. Furthermore, a delay-optimal alternating optimization algorithm (DOAOA) is developed to obtain the solution for the optimization problem. Experimental results validate the plausibility and superiority of the proposed framework compared to the state-of-the-art schemes.
Zhijian Lin, Yang Xiao 0014, Yi Fang 0005, Hongbing Chen, Xiaoqiang Lu
IEEE Trans. Mob. Comput.5
2025 Uncertainty-Aware Semi-Supervised Learning Segmentation for Remote Sensing Images
abstract
Deep learning based remote sensing (RS) image segmentation significantly impacts several real application scenarios. Behind its success, massive labeled data plays an important role. However, annotating high-resolution RS images requires time-consuming and relevant expertise efforts. To address it, many works dive into semi-supervised learning which utilizes raw information embedded in unlabeled data to improve the segmentation model. Nevertheless, previous studies ignore the integrity and effectiveness of the potential context information hidden in RS data. In this work, we propose an uncertainty-aware masked consistency learning (U-MCL) framework that contains an uncertainty-aware masked denoising (U-MD) module and an uncertainty-aware masked image consistency (U-MIC) module. U-MCL initially generates a patch-wise uncertainty map for each unlabeled image during each training iteration, which is then used to derive an adaptive mask ratio for pseudo-label denoising in U-MD. Simultaneously, the uncertainty map is adopted to model a masked unlabeled image for reasoning unseen areas in U-MIC. Consequently, U-MCL is capable of enhancing model performance by engaging in accurate and stable consistency learning while preserving the integrity of the context and employing the context to infer the predictions of the masked regions safely. Extensive experiments on six RS datasets, i.e., ISPRS Vaihingen, FloodNet, MiniFrance, LoveDA, MER, and MSL, demonstrate the superiority of our U-MCL over recent most advanced methods, achieving new state-of-the-art performance under all benchmarks.
Xiaoqiang Lu, Lingling Li 0002, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001
IEEE Trans. Multim.1
2024 Multimodal Segformer for Flood Rapid Mapping with Sentinel-2 Data
abstract
Flood rapid mapping products play an important role in informing flood emergency response and management. To this end, the 2024 IEEE GRSS Data Fusion Contest Track 2 (DFC24-T2) establishes a multimodal benchmark for the segmentation of flood areas from Sentinel-2 multispectral images. However, the problems of imbalanced data distribution, data scarsity, and inter-modal differences severely inhibit the performance of deep-learning-based segmentation networks. In this work, we propose an end-to-end Multimodal Transformer-based Segmentation Network (MTSN) for accurate flood rapid mapping. MTSN first employs two Siamese encoders with shared parameters to accept multimodal inputs and output their respective hierarchical multiscale features, which are then enriched by several channel attention blocks. Subsequently, a Cross-modal Feature Fusion Module (CFFM) based on a gated mechanism is proposed to efficiently integrate the benefits of multimodal features, and generate informative representations. Finally, the fused features are decoded by a lightweight pure multilayer perception decoder to quickly generate mapping results of flood areas. Moreover, we introduce offline data augmentation, semi-supervised learning, test-time augmentation, and multimodal post-process to further boost the performance and generalization of our MTSN. Experimental results and extensive ablations show the effectiveness of our method. Code is available at https://github.com/xiaoqiang-lu/MMSegFormer.
Xiaoqiang Lu, Tong Gou, Zhongjian Huang, Yuting Yang 0008, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001
IGARSS1
2024 Game-based computation offloading and resource allocation in stochastic geometry-modeling vehicular networks
Jianjie Yang, Zhijian Lin, Yingyang Chen, Xiaoqiang Lu, Yi Fang 0005
Sci. China Inf. Sci.4
2024 Domain Adaptation of Anchor-Free object detection for urban traffic
Xiaoyong Yu, Xiaoqiang Lu
Neurocomputing2
2024 Query-aware multi-scale proposal network for weakly supervised temporal sentence grounding in videos
Mingyao Zhou, Wenjing Chen 0003, Hao Sun 0014, Wei Xie 0008, Ming Dong 0004, Xiaoqiang Lu
Knowl. Based Syst.6
2024 Prototype-Based Pseudo-Label Refinement for Semi-Supervised Hyperspectral Image Classification
abstract
Pseudo-label learning-based methods usually regard class confidence above a certain threshold for unlabeled samples as pseudo-labels, which may result in pseudo-labels still containing wrong labels. In this letter, we propose a prototype-based pseudo-label refinement (PPLR) for semi-supervised hyperspectral image classification. The proposed PPLR filters wrong labels from pseudo-labels using class prototypes, which can improve the discrimination of the network. First, PPLR uses multi-head attentions to extract the spectral-spatial features, and designs an adaptive threshold that can be dynamically adjusted to generate high-confidence pseudo-labels. Then, PPLR constructs class prototypes for different categories using labeled sample features and unlabeled sample features with refined pseudo-labels to improve the quality of pseudo-labels by filtering wrong labels. Finally, PPLR further assigns reliable weights to these pseudo-labels in calculating their supervised loss, and introduces a center loss to improve the discrimination of features. When 10 labeled samples per category are utilized for training, PPLR achieves the overall accuracies of 82.11%, 86.70% and 92.50% on the Indian Pines, Houston2013 and Salinas datasets, respectively.
Renyi Chen, Huaxiong Yao, Wenjing Chen 0003, Hao Sun 0014, Wei Xie 0008, Xiaoqiang Lu
IEEE Geosci. Remote. Sens. Lett.7
2024 Dual-Intervention-Constrained Mask-Adversary Framework for Unsupervised Domain Adaptation of Hyperspectral Image Classification
abstract
To mitigate the domain shift and enhance the alignment of the spatial-spectral features, this letter proposes a novel dual-intervention-constrained mask-adversary (DICMA) framework for unsupervised domain adaptation (UDA) of hyperspectral image classification (HSIC). Innovatively, DICMA integrates a generator, masker, and bi-classifier within an adversarial framework constrained by a dual intervention mechanism. Specifically, the correlation intervention module (CIM) ensures the preservation and independence of causal spatial-spectral variables, while the knowledge distillation intervention module completes the spatial-spectral generalization with constrained distillation information. Besides, with the collaborative adversarial training strategy, the proposed approach transfers effective knowledge for spatial-spectral feature alignment. Experimental results and analyses demonstrate the effectiveness of the proposed DICMA model, which yields an accuracy of 91.15% in the Pavia University (PaviaU)$\to $Pavia Center (PaviaC). Our code will be released athttps://github.com/Chirsycy/DICMA.
Chunyan Yu, Mingyang Xu, Qiang Zhang 0011, Xiaoqiang Lu
IEEE Geosci. Remote. Sens. Lett.4
2024 Self Pseudo Entropy Knowledge Distillation for Semi-Supervised Semantic Segmentation
abstract
Recently, semi-supervised semantic segmentation methods based on weak-to-strong consistency learning have achieved the most advanced performance. The key to such a technique lies in strong perturbations and multi-objective co-training. However, CutMix, the most commonly used data augmentation in this field, limits the strength of perturbations as it only focuses on single random local context. Besides, complex optimization targets also reduce computational efficiency. In this work, we propose an efficient consistency learning based framework. Specifically, a novel unsupervised data augmentation strategy, EntropyMix, is present for semi-supervised semantic segmentation. Patches of unlabeled data from multi-view augmentations are combined into new training samples based on their prediction entropy, which provides more informative and powerful perturbations for consistency regularization and impels the model to focus on cross-view local context. On this basis, we further propose Self Pseudo Entropy Knowledge Distillation (SPEED) to learn global pixel relations from multi- and cross-view perturbations by optimizing a linear combination of feature-and logit-level distillation loss, enhancing model performance without additional auxiliary segmentation heads or a complex pre-trained teacher model. The collocation of the two ideas above is a plug-and-play technique without additional modification. Extensive experimental results on PASCAL VOC and Cityscapes datasets under various training settings demonstrate the superiority of the proposed data augmentation strategy and self-distillation loss, achieving new state-of-the-art performance. Remarkably, our method reaches mIoU of 75.16% using only 0.87% labeled data on PASCAL VOC and mIoU of 76.98% using only 6.25% labeled data on Cityscapes. The code is available at https://github.com/xiaoqiang-lu/SPEED.
Xiaoqiang Lu, Licheng Jiao, Lingling Li 0002, Fang Liu 0001, Xu Liu 0006, Shuyuan Yang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2024 Scale-Aware Adaptive Refinement and Cross-Interaction for Remote Sensing Audio-Visual Cross-Modal Retrieval
Yaxiong Chen, Chuang Du, Yunfei Zi, Shengwu Xiong 0001, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.5
2024 Multiscale Salient Alignment Learning for Remote-Sensing Image-Text Retrieval
abstract
Remote-sensing image–text (RSIT) retrieval involves the use of either textual descriptions or remote-sensing images (RSI) as queries to retrieve relevant RSIs or corresponding text descriptions. Many traditional cross-modal RSIT retrieval methods tend to overlook the importance of capturing salient information and establishing the prior similarity between RSIs and texts, leading to a decline in cross-modal retrieval performance. In this article, we address these challenges by introducing a novel approach known as multiscale salient image-guided text alignment (MSITA). This approach is designed to learn salient information by aligning text with images for effective cross-modal RSIT retrieval. The MSITA approach first incorporates a multiscale fusion module and a salient learning module to facilitate the extraction of salient information. In addition, it introduces an image-guided text alignment (IGTA) mechanism that uses image information to guide the alignment of texts, enabling the effective capture of fine-grained correspondences between RSI regions and textual descriptions. In addition to these components, a novel loss function is devised to enhance the similarity across different modalities and reinforce the prior similarity between RSIs and texts. Extensive experiments conducted on four widely adopted RSIT datasets affirm that the MSITA approach significantly enhances cross-modal RSIT retrieval performance in comparison to other state-of-the-art methods.
Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.5
2024 Thread the Needle: Cues-Driven Multiassociation for Remote Sensing Cross-Modal Retrieval
abstract
Rapid advances in Earth observation technologies have yielded numerous remotely sensed images and corresponding text data, enabling cross-modal image–text retrieval to extract valuable clues. However, current methods often focus on learning global semantic information from text and remote sensing (RS) images, while neglecting fine-grained semantic alignment and correlation. In addition, contrastive learning between modalities is often insufficient. To address these issues, we propose an innovative cues-driven multiassociation feature matching network (CDMAN) for cross-modal RS image retrieval. The proposed method primarily involves two key steps: 1) aligning positive samples and enhancing fusion for negative samples based on modal cues. To achieve precise alignment between RS images and text and facilitate the learning process for negative samples in contrastive learning, we have developed a novel fine-grained cues injection module that aligns and guides modalities using fine-grained cues; and 2) establishing multigranularity associative learning. To address the issue of insufficient association between RS images and text, we have implemented multigranularity collaborative associative learning, focusing on general and fine-grained modal associations. By fully leveraging modal cues, our method maintains both detailed associations and overall consistency in global associations. Experiments demonstrate that, compared to baseline methods, this approach achieves more accurate cross-modal retrieval (MCR) by combining fine-grained alignment and multigranularity associations.
Yaxiong Chen, Jirui Huang, Zhaoyang Sun, Shengwu Xiong 0001, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.5
2024 Integrating Multisubspace Joint Learning With Multilevel Guidance for Cross-Modal Retrieval of Remote Sensing Images
abstract
In recent years, with the continuous advancement of remote sensing technology and text processing techniques, there has been a growing abundance of remote sensing images and associated textual data. Combining remote sensing images with their corresponding textual data allows for integrated analysis and retrieval, which holds significant practical implications across multiple application domains, including geographic information systems (GIS), environmental monitoring, and agricultural management. Remote sensing images have the characteristics of multi-targets and multi-scales, and the textual descriptions of these targets are not fully utilized, leading to a decrease in retrieval accuracy. Previous methods have struggled to balance inter-modality information interaction and intra-modality feature fusion, and they have paid little attention to the consistency of distribution within modalities. In light of this, this paper proposes a symmetric multi-level guidance network (SMLGN) for cross-modal retrieval in remote sensing. SMLGN first introduces fusion guidance between local and global within modalities and fine-grained bidirectional guidance between modalities, allowing for the learning of a common semantic space. Furthermore, to address the distribution differences of different modalities within the common semantic space, we design an adversarial joint learning framework and a multi-objective loss function to optimize the SMLGN method and achieve consistency in data distribution. The experimental results demonstrate that the SMLGN method performs well in the task of cross-modal retrieval between remote sensing images and textual data. It effectively integrates the information from both modalities, improving the accuracy and reliability of the retrieval process.
Yaxiong Chen, Jirui Huang, Shengwu Xiong 0001, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.4
2024 Integrating Detailed Features and Global Contexts for Semantic Segmentation in Ultrahigh-Resolution Remote Sensing Images
abstract
Semantic segmentation of ultrahigh-resolution (UHR) remote sensing images is a fundamental task for many downstream applications. Achieving precise pixel-level classification is paramount for obtaining exceptional segmentation results. This challenge becomes even more complex due to the need to address intricate segmentation boundaries and accurately delineate small objects within the remote sensing imagery. To meet these demands effectively, it is critical to integrate two crucial components: global contextual information and spatial detail feature information. In response to this imperative, the multilevel context-aware segmentation network (MCSNet) emerges as a promising solution. MCSNet is engineered to not only model the overarching global context but also extract intricate spatial detail features, thereby optimizing segmentation outcomes. The strength of MCSNet lies in its two pivotal modules, the spatial detail feature extraction (SDFE) module and the refined multiscale feature fusion (RMFF) module. Moreover, to further harness the potential of MCSNet, a multitask learning approach is employed. This approach integrates boundary detection and semantic segmentation, ensuring that the network is well-rounded in its segmentation capabilities. The efficacy of MCSNet is rigorously demonstrated through comprehensive experiments conducted on two established international society for photogrammetry and remote sensing (ISPRS) 2-D semantic labeling datasets: Potsdam and Vaihingen. These experiments unequivocally establish MCSNet stands as a pioneering solution, that delivers state-of-the-art performance, as evidenced by its outstanding mean intersection over union (mIoU) and mean$F1$-score (mF1) metrics. The code is available at:https://github.com/WUTCM-Lab/MCSNet.
Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu, Xiao Xiang Zhu 0001, Lichao Mou
IEEE Trans. Geosci. Remote. Sens.4
2024 A Joint Saliency Temporal-Spatial-Spectral Information Network for Hyperspectral Image Change Detection
abstract
Hyperspectral image change detection (HSI-CD) is a fundamental task in the field of remote sensing (RS) observation, which utilizes the rich spectral and spatial information in bitemporal HSIs to detect subtle changes on the Earth’s surface. However, modern deep learning (DL)-based HSI-CD methods mostly rely on patch-based methods, which leads to spectral band redundancy and spatial information noise in limited receiving domains, thus ignoring the extraction and utilization of saliency information and limiting the improvement of CD performance. To address these issues, this article proposes a joint saliency temporal–spatial–spectral information network (STSS-Net) for HSI-CD. The principal contributions of this article can be summarized: 1) we have designed a spatial saliency information extraction (SSIE) module for denoising based on distance from center pixels and spectral similarity of the substance, which increases the attention to spatial differences between similar spectral substances and different spectral substances; 2) we have designed a compact high-level spectral information tokenizer (CHLSIT) for spectral saliency information, where the high-level conceptual information of changes in spectral interest can be represented by nonlinear combinations of spectral bands, and redundancy can be removed by extracting high-level spectral conceptual features; and 3) utilizing the advantages of CNN and transformer architectures to combine temporal–spatial–spectral information. The experimental results on three real HSI-CD datasets show that STSS-Net can improve the accuracy of CD and has a certain improvement in the detection of edge information and complex information.
Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.5
2024 Learning Starts From Optimizing the Composition of Temporal Information for Hyperspectral Change Detection
abstract
Hyperspectral image change detection (HSI-CD) is a task that utilizes both spectral and spatial features to more effectively detect changes. The features of HSI captured at different times are often influenced by external factors. The annoying variability is not useful for CD. Current methods extract effective information through complex learning together with it. They did not focus on whether different categories (changed and unchanged) of sample pairs have the same information composition. They lack a fundamental analysis and optimization based on the composition of information pairs to address current issues such as insufficient recognition of changed samples and poor feature fusion. To address these issues, we rethink the information composition of bitemporal sample pairs in different categories for HSI-CD, and design a frequency-domain information exchange and generation network with Siamese U-shaped structure (FDIEG-UNet) for HSI-CD. Our contribution can be summarized as follows: 1) by designing the FDGJLS module for learning and separating global-joint spectral-spatial features, to provide better basic frequency-domain information for subsequent feature-domain optimization; 2) based on the information from FDGJLS, we use the ULDT and SSTDFE modules to remove the influence of temporal-correlated invalid style information in feature fusion. Two modules work together to improve the equality of feature description between changed and unchanged samples in CD; and 3) the experimental results on three real HSI-CD datasets demonstrate the effectiveness of our proposed method in solving the above problems. The code is available athttps://github.com/WUTCM-Lab/FDIEG-UNet.
Yaxiong Chen, Jirui Huang, Shengwu Xiong 0001, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.5
2024 Cross-Modal Remote Sensing Image-Audio Retrieval With Adaptive Learning for Aligning Correlation
abstract
An important challenge that existing work has yet to address is the relatively small differences in audio representations compared with the rich content provided by remote sensing (RS) images, making it easy to overlook certain details in the images. This imbalance in information between modalities poses a challenge in maintaining consistent representations. In response to this challenge, we propose a novel cross-modal RS image-audio (RSIA) retrieval method called adaptive learning for aligning correlation (ALAC). ALAC integrates region-level learning into image annotation through a region-enhanced learning attention (RELA) module. By collaboratively suppressing features at different region levels, ALAC is able to provide a more comprehensive visual feature representation. In addition, a novel adaptive knowledge transfer (AKT) strategy has been proposed, which guides the learning process of the frontend network using aligned feature vectors. This approach allows the model to adaptively acquire alignment information during the learning process, thereby facilitating better alignment between the two modalities. Finally, to better use mutual information between different modalities, we introduce a plug-and-play result rerank module. This module optimizes the similarity matrix using retrieval mutual information between modalities as weights, significantly improving retrieval accuracy. Experimental results on four RSIA datasets demonstrate that ALAC outperforms other methods in retrieval performance. Compared with state-of-the-art methods, improvements of 1.49%, 2.25%, 4.24%, and 1.33% were, respectively, achieved by ALAC. The codes are accessible athttps://github.com/huangjh98/ALAC.
Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.4
2024 Visual Contextual Semantic Reasoning for Cross-Modal Drone Image-Text Retrieval
abstract
The cross-modal drone image-text (DIT) retrieval task involves using either text or drone images as queries to retrieve relevant drone images or corresponding text. The primary challenge stems from the diverse and intricate nature of drone images, making effective alignment between image and text challenging. In response, we propose an innovative approach called visual contextual semantic reasoning (VCSR), aimed at precisely aligning information across different modalities. VCSR employs textual cues to guide rich semantic reasoning within the visual context, reducing redundancy in visual information. Furthermore, the method captures drone image information relevant to the text, revealing subtle correspondences between drone image regions and textual content. To enhance visual semantic learning, context region learning (CRL) term and consistency semantic alignment (CSA) terms are introduced for stronger guidance, further intensifying the cross-modal interaction between textual and visual data, resulting in more robust feature representation. Extensive experiments conducted on two self-constructed DIT datasets demonstrate that VCSR outperforms alternative methods in terms of DIT retrieval performance. The codes are accessible athttps://github.com/huangjh98/VCSR.
Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.4
2024 Oriented Object Detector With Gaussian Distribution Cost Label Assignment and Task-Decoupled Head
abstract
Recently, oriented object detection in remote sensing images has garnered significant attention due to its broad range of applications. Early oriented object detection adhered to the established general object detection frameworks, utilizing the label assignment strategy based on the horizontal bounding box annotations or rotation-agnostic cost function. Such strategy may not reflect the large aspect ratio and rotation of arbitrary-oriented objects in remota sensing images and require high parameter-tuning efforts in training process, which will eventually harm the detector performance. Furthermore, the localization quality of oriented object depends on precise rotation angle prediction, exacerbating the inconsistency between classification and regression tasks in oriented object detection. To address these issues, we propose the Gaussian Distribution Cost Optimal Transport Assignment (GCOTA) and Decoupled Layer Attention Angle Head (DLAAH). Specifically, GCOTA utilize Gaussian distribution based cost function for the optimal transport label assignment in training process, alleviating the impact of rotation angle and large aspect ratio in remote sensing images. DLAAH predicts rotation angle independently and incorporates layer attention to obtain the task-specific features based on the shared FPN features, enhancing the angle prediction and improving consistency across different tasks. Based on these proposed components, we present an anchor-free oriented detector, namely Gaussian Distribution and Task-Decoupled head oriented Detector(GTDet) and a a multi-class ship detection dataset in real scenarios (CGWX), which provides a benchmark for fine-grained object recognition in remote sensing images. Comprehensive experiments are conducted on CGWX and several public challenging datasets, including DOTAv1.0, HRSC2016, to demonstrate that our method achieves superior performance on oriented object detection task. The code is available at https://github.com/WUTCM-Lab/GTDet.
Qiangqiang Huang, Ruilin Yao, Xiaoqiang Lu, Jishuai Zhu, Shengwu Xiong 0001, Yaxiong Chen
IEEE Trans. Geosci. Remote. Sens.3
2024 Domain Mapping Network for Remote Sensing Cross-Domain Few-Shot Classification
abstract
It is a challenging task to recognize novel categories with only a few labeled remote sensing images. Currently, meta-learning solves the problem by learning prior knowledge from another dataset where the classes are disjoint. However, the existing methods assume the training dataset comes from the same domain as the test dataset. For remote sensing images, test dataset may come from different domains. It is impossible to collect a training dataset for each domain. Meta-learning and transfer learning are widely used to tackle the few-shot classification and the cross-domain classification, respectively. However, it is difficult to recognize novel categories from various domains with only a few images. In this paper, a Domain Mapping Network (DMN) is proposed to cope with the few-shot classification under domain shift. DMN trains an efficient few-shot classification model on the source domain and then adapts the model to the target domain. Specifically, dual autoencoders are exploited to fit the source and target domain distribution. First, DMN learns an autoencoder on the source domain to fit the source domain distribution. Then, a target autoencoder is initiated from the source domain autoencoder and further updated with a few target images. To ensure the distribution alignment, cycle-consistency losses are proposed to jointly train the source autoencoder and target autoencoder. Extensive experiments are conducted to validate the generalizable and superiority of the proposed method.
Xiaoqiang Lu, Tengfei Gong, Xiangtao Zheng
IEEE Trans. Geosci. Remote. Sens.1
2024 SANet: A Self-Attention Network for Agricultural Hyperspectral Image Classification
abstract
Unlike conventional hyperspectral image (HSI) classification in general scenes, agricultural HSI classification poses greater challenges due to the increased occurrence of “same spectrum different object” and “different spectrum same object” phenomena caused by class similarities. Furthermore, the dense spatial distribution of land cover categories in agricultural scenes and the mixing of spatial–spectral features at crop boundaries add to the complexity of agricultural HSIs. To tackle these issues, we propose SANet, a network designed to enhance crop classification. SANet integrates spectral and contextual information while emphasizing self-correlation within the HSIs. It combines the spatial–spectral nonlocal block structure and the multiscale spectral self-attention (SSA) structure, allocating more attention resources to spatial and spectral dimensions and modeling the existing correlations within the spectral–spatial domain. Additionally, we introduce a two-branch spatial–spectral semantic extraction and fusion structure that can adaptively learn results from both branches. Experimental results demonstrate the promising performance of SANet in agricultural HSI classification by effectively utilizing spectral data, contextual information, and self-attention mechanisms.
Bo Zhang 0069, Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.5
2024 Spectrum-Induced Transformer-Based Feature Learning for Multiple Change Detection in Hyperspectral Images
abstract
The multiple change detection (MCD) of hyperspectral images (HSIs) is the process of detecting change areas and providing “from–to” change information of HSIs obtained from the same area at different times. HSIs have hundreds of spectral bands and contain a large amount of spectral information. However, current deep-learning-based MCD methods do not pay special attention to the interspectral dependency and the effective spectral bands of various land covers, which limits the improvement of HSIs’ change detection (CD) performance. To address the above problems, we propose a spectrum-induced transformer-based feature learning (STFL) method for HSIs. The STFL method includes a spectrum-induced transformer-based feature extraction module (STFEM) and an attention-based detection module (ADM). First, the 3D-2D convolutional neural networks (CNNs) are used to extract deep features, and the transformer encoder (TE) is used to calculate self-attention matrices along the spectral dimension in STFEM. Then, the extracted deep features and the learned self-attention matrices are dot-multiplied to generate more discriminative features that take the long-range dependency of the spectrum into account. Finally, ADM mines the effective spectral bands of the difference features learned from STFEM by the attention block (AB) to explore the discrepancy of difference features and uses the softmax function to identify multiple changes. The proposed STFL method is validated on two hyperspectral datasets, and their experiments illustrate the superiority of the proposed STFL method over the currently existing MCD methods.
Wuxia Zhang, Shiwen Gao, Xiaoqiang Lu, Yi Tang 0003, Shihu Liu
IEEE Trans. Geosci. Remote. Sens.4
2024 Global-Group Attention Network With Focal Attention Loss for Aerial Scene Classification
abstract
Aerial scene classification, aiming at assigning a specific semantic class to each aerial image, is a fundamental task in the remote sensing community. Aerial scene images have more diverse and complex geological features. While some statistics of images can be well fit using convolution, it limits such models to capturing the global context hidden in aerial scenes. Furthermore, to optimize the feature space, many methods add class information to the feature embedding space. However, they seldom combine model structure with class information to obtain more separable feature representations. In this article, we propose to address these limitations in a unified framework (i.e., CGFNet) from two aspects: focusing on the key information of input images and optimizing the feature space. Specifically, we propose a global-group attention module (GGAM) to adaptively learn and selectively focus on important information from input images. GGAM consists of two parallel branches: the adaptive global attention branch (AGAB) and the region-aware attention branch (RAAB). AGAB utilizes an adaptive pooling operation to better model the global context in aerial scenes. As a supplement to AGAB, RAAB combines grouping features with spatial attention to spatially enhance the semantic distribution of features (i.e., selectively focus on effective regions of features and ignore irrelevant semantic regions). In parallel, a focal attention loss (FA-Loss) is exploited to introduce class information into attention vector space, which can improve intraclass consistency and interclass separability. Experimental results on four publicly available and challenging datasets demonstrate the effectiveness of our method. The source code will be released at:https://github.com/zoecheno/CGFNet.
Yichen Zhao, Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.5
2024 Co-Enhanced Global-Part Integration for Remote-Sensing Scene Classification
abstract
Remote sensing (RS) scene classification aims to classify remote sensing images with similar scene characteristics into one category. Plenty of RS images are complex in background, rich in content, and multi-scale in target, exhibiting the characteristics of both intra-class separation and inter-class convergence. Therefore, discriminative feature representations designed to highlight the differences between classes are the key to RS scene classification. Existing methods represent scene images by extracting either global context or discriminative part features from RS images. However, global-based methods often lack salient details in similar RS scenes, while part-based methods tend to ignore the relationships between local ground objects, thus weakening the discriminative feature representation. In this paper, we propose to combine global context and part-level discriminative features within a unified framework called CGINet for accurate RS scene classification. To be specific, we develop a light context-aware attention block (LCAB) to explicitly model the global context to obtain larger receptive fields and contextual information. A co-enhanced loss module (CELM) is also devised to encourage the model to actively locate discriminative parts for feature enhancement. In particular, CELM is only used during training and not activated during inference, which introduces less computational cost. Benefiting from LCAB and CELM, our proposed CGINet improves the discriminability of features, thereby improving classification performance. Comprehensive experiments over four benchmark datasets show that the proposed method achieves consistent performance gains over state-of-the-art RS scene classification methods.
Yichen Zhao, Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu, Xiao Xiang Zhu 0001, Lichao Mou
IEEE Trans. Geosci. Remote. Sens.4
2023 A Strong Vision Transformer Adapter with Adaptive Thresholding for fine-Grained Building Classification
abstract
Fine-grained building classification provides a solid basis for the comparison of city morphologies and the investigation of urban planning. To this aim, the DFC23 establishes a large-scale and multi-modal benchmark for the classification of building roof types. However, the problems of long-tailed distribution, data insufficient, inter-class similarity, and intra-class difference severely inhibit the performance of the detector. In this work, we build a strong vision transformer adapter fine-tuned on the cropped building instances to enhance the capacity of feature extraction and design a cross-modal fusion (CMF) module to effectively aggregate features from RGB and SAR data. When transferring to building instance segmentation, we construct a robust training pipeline and a two-stage test-time results ensemble scheme. Furthermore, we introduce self-training with two key denoising techniques, global average filtering (GAF) and intra-class adaptive thresholding (IAT), to boost the generalization of the model. Experimental results show the effectiveness of our method, ranking 2nd in the test phase of the contest.
Xiaoqiang Lu, Licheng Jiao, Lingling Li 0002, Fang Liu 0001, Xu Liu 0006, Yuting Yang 0008
IGARSS1
2023 Trident Cooperation Network for Building Extraction and Height Estimation
abstract
Building extraction and height estimation provide solid fundamentals for reconstructing city morphologies and investigating urban planning. To this aim, the DFC23 establishes a large-scale and multi-modal benchmark for multi-task learning of building reconstruction. However, the problems of data limitation and fore-background confusion severely inhibit the performance of the model. In this work, we propose a novel trident cooperation network (TCNet) to perform end-to-end building extraction and height estimation using RGB and SAR data. Specifically, to enrich the feature representation and generalization of the shared backbone, we introduce a vision transformer adapter to inject vision-specific inductive biases and design a cross-modal fusion (CMF) module to effectively aggregate features from multi-modal data. For downstream visual tasks, we construct trident decoders including a detector, a lightweight MLP segmentation head, and a pixel-wise regression head. Moreover, to highlight the foreground object, we use the binary mask predicted by the MLP head to cooperate with the height estimation map predicted by the estimator. And the weighted sub-task losses are gathered to optimize our TCNet. Experimental results show the effectiveness of our method, ranking 2nd in the test phase of the contest.
Xiaoqiang Lu, Licheng Jiao, Lingling Li 0002, Fang Liu 0001, Xu Liu 0006, Yuting Yang 0008
IGARSS1
2023 Deep Feature Reconstruction Learning for Open-Set Classification of Remote-Sensing Imagery
abstract
Existing remote sensing scene image (RSSI) classification methods usually rely on static closed-set assumption that testing samples do not belong to unknown classes. However, practical applications are usually the open-set classification problem, which means that RSSIs from unknown classes will appear in the testing set. Most existing methods are prone to forcibly misclassify RSSIs of unknown classes into known classes, resulting in poor practical performance. In this letter, a deep feature reconstruction learning (DFRL) framework is proposed for open-set classification of RSSIs. The proposed DFRL unifies discriminative feature learning and feature reconstruction into an end-to-end network. Firstly, a feature extraction module is utilized to project raw input data from the image space to the feature space to extract deep features. Then, the deep features are fed to a deep feature reconstruction module for distinguishing known and unknown classes based on feature-level reconstruction errors. The feature-level reconstruction can effectively suppress the interference of complex backgrounds. In addition, a sparse regularization is introduced to improve the discrimination of image representation. Experiments on three RSSI datasets demonstrate the effectiveness of DFRL for open-set classification of RSSIs.
Hao Sun 0014, Jie Yu 0003, Dongbo Zhou, Wenjing Chen 0003, Xiangtao Zheng, Xiaoqiang Lu
IEEE Geosci. Remote. Sens. Lett.7
2023 Difference-Enhancement Triplet Network for Change Detection in Multispectral Images
abstract
Change detection is the process of detecting and evaluating differences from bitemporal remote sensing images. Deep-learning-based change detection methods have become the mainstream approaches due to their discriminative features and good change detection performance. However, most of the existing deep-learning-based change detection methods did not perform well in detecting subtle changes and did not fully explore the underlying information of features learned by deep neural networks. To address the above-mentioned problems, we propose an end-to-end deep neural network for multispectral change detection, named difference-enhancement triplet network (DETNet). DETNet mainly includes two modules: the triplet feature extraction module and the difference feature learning module. First, the triplet feature extraction module uses the triple CNN as the backbone to extract representative spatial–spectral features. Second, the difference feature learning module mines the underlying information of difference representations of learned spatial–spectral features to detect subtle changes. Finally, the model uses a compound loss function, which includes triplet loss, contrastive loss, and cross-entropy loss, to guide DETNet toward learning more discriminative features. Extensive experimental results of the proposed DETNet and other state-of-the-art methods on four datasets demonstrate its superiority.
Wuxia Zhang, Liangxu Su, Xiaoqiang Lu
IEEE Geosci. Remote. Sens. Lett.5
2023 Fine Aligned Discriminative Hashing for Remote Sensing Image-Audio Retrieval
abstract
For cross-modal remote sensing image-audio retrieval task, hashing technology has attracted much attention in recent works. Most of them focus on mappingRemote Sensing(RS) images and audios into a Hamming space, whilst neglecting discriminative information of RS images and fine alignment for RS images and audios. In this paper, we tackle these dilemmas with a novelFine Aligned Discriminative Hashing(FADH) approach, which can learn hash codes to capture discriminative information of RS images and learn the corresponding detailed information between RS images and audios simultaneously. We first develop a new discriminative information learning module to learn discriminative information of RS images. Meanwhile, a fine alignment module is proposed to unearth the fine correspondence for RS image regions and audios, which can effectively improve the retrieval performance. On top of the two paths, we design a new objective function, which can maintain the similarity of hash codes, preserve the semantic information of RS image features and audio features and eliminate cross-modal differences. The reliability and significance of the designed framework are effectively demonstrated by diverse experiments on three remote sensing image-audio datasets.
Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.4
2023 Self-Supervision Interactive Alignment for Remote Sensing Image-Audio Retrieval
abstract
Cross-modal remote sensing image-audio retrieval aims to use audio or remote sensing images as queries to retrieve relevant remote sensing images or corresponding audios. Although many approaches leverage labeled samples to achieve good performance, the performance cost of labeled samples is high, because cross-modal remote sensing labeled samples usually requires huge labor resources. Therefore, unsupervised cross-modal learning is very important in real-world applications. In this paper, we propose a novel unsupervised cross-modal remote sensing image-audio retrieval approach, namedSelf-Supervision Interactive Alignment(SSIA), which can take advantage of large amounts of unlabeled samples to learn the salient information, cross-modal alignment and the similarity between remote sensing images and audios. Since self-supervised learning lacks the supervision of label information, we leverage the similarity between the input remote sensing image information and audio information as the supervision information. Besides, to perform cross-modal alignment, a novel interactive alignment module is designed to explore fine correspondence relation for remote sensing images and audios. Moreover, we design an audio guided image de-redundant module to reduce the redundant information of visual information, which can capture salient information of remote sensing images. Extensive experiments on four widely-used remote sensing image-audio datasets testify that the SSIA perform gain better remote sensing image-audio retrieval performance than other compared approaches.
Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.4
2023 Weak-to-Strong Consistency Learning for Semisupervised Image Segmentation
abstract
Supervised remote sensing (RS) image segmentation has achieved remarkable success with large amounts of manually labeled data, which may be difficult to acquire in some practical application scenarios. Semisupervised RS image segmentation can efficiently utilize the knowledge embedded in unlabeled data to improve recognition performance, which is of great significance for the generalization application of segmentation models. In this work, we propose an end-to-end semisupervised RS image segmentation method based on weak-to-strong consistency learning, denoted as WSCL. Specifically, a common strong data augmentation technique for image segmentation is introduced to provide powerful input perturbation to decouple self-biased cognition. By forcing weakly augmented, and strongly augmented perspectives from the same sample to be consistent, WSCL not only enables the model to steadily learn knowledge contained in unlabeled data but also alleviates overfitting. In addition, a novel sparse dual-view cross-sample image generation method is presented to generate new training samples, which helps provide a more comprehensive diversity of perturbations. Furthermore, an adaptive re-weighting strategy based on the entropy maps of the outputs of strongly perturbed samples is proposed to suppress noise, guiding the training process in a positive direction. Extensive experiments demonstrate the significant advantage of WSCL over other advanced methods, achieving new state-of-the-art under several evaluation metrics on DFC22, iSAID, MER, MSL, Vaihingen, and GID-15 datasets. The source code is open-sourced at https://github.com/xiaoqiang-lu/WSCL.
Xiaoqiang Lu, Licheng Jiao, Lingling Li 0002, Fang Liu 0001, Xu Liu 0006, Shuyuan Yang 0001, Zhixi Feng, Puhua Chen
IEEE Trans. Geosci. Remote. Sens.1
2023 Pseudolabel-Based Unreliable Sample Learning for Semi-Supervised Hyperspectral Image Classification
abstract
Recently, pseudo-label-based deep learning methods have shown excellent performance in semi-supervised hyperspectral image (HSI) classification. These methods usually select high-confidence unlabeled samples to help optimize backbone classification networks. However, a large number of remaining low-confidence unlabeled samples, which contain rich land-covers information, are underutilized. In this paper, we propose a pseudo-label-based unreliable sample learning (PUSL) method to fully exploit low-confidence unlabeled samples for semi-supervised HSI classification. Firstly, to avoid overfitting the spatial distribution of labeled samples, we build a position-free transformer (PFT) as the backbone classification network. Secondly, PFT is initially trained with labeled samples in a supervised learning manner to obtain an initial classifier, which is then used to split unlabeled samples into reliable and unreliable unlabeled samples based on the predicted confidence. Thirdly, reliable unlabeled samples participate in training along with labeled samples. Finally, unreliable unlabeled samples are treated as negative samples for corresponding categories to improve the discrimination of PFT in a contrastive learning paradigm. Extensive experiments on three HSI datasets demonstrate that PUSL outperforms compared methods.
Huaxiong Yao, Renyi Chen, Wenjing Chen 0003, Hao Sun 0014, Wei Xie 0008, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.6
2023 MATNet: A Combining Multi-Attention and Transformer Network for Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) has rich spatial-spectral information, high spectral correlation and large redundancy between information. Due to the sparse background distribution of HSI, existing methods generally perform poorly for the classification of class pixels located in the boundary areas of land cover categories. This is largely because the network is vulnerable to surrounding redundant information during the training stage, leading to inaccurate feature extraction and thus poor generalization ability of the model. Based on previous work, we propose a HSI classification network called MATNet which combines multi-attention and Transformer. The network first uses spatial attention and channel attention to pay more attention to the more significant information parts, then uses tokenizer module to make a semantic level representation of different categories of ground objects, and then performs deep semantic feature extraction using the transformer encoder module. Finally, we design a loss function called Lpoly, which adds a polynomial to the label smoothing loss to tune the original first polynomial to accommodate different datasets and tasks. We perform experiments in several well-known HSI datasets as well as for visualization. The results show that our proposed MATNet performs well in extracting spatial-spectral features of HSIs as well as understanding semantic degrees of semantic degrees.
Bo Zhang 0069, Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.5
2023 A Spectrum-Aware Transformer Network for Change Detection in Hyperspectral Imagery
abstract
Change detection in the HyperSpectral Imagery (HSI) detects the changed pixels or areas in bi-temporal images. HSIs contain hundreds of spectral bands, including a large amount of spectral information. However, most of deep learning-based change detection methods did not focus on the spectral dependency of spectral information in the spectral dimension and just adopted the difference strategy to represent the correlation of learned features, which limited the improvement of the change detection performance. To address the above-mentioned problems, we propose an end-to-end change detection network for HSIs, named Spectrum-Aware Transformer Network (SATNet), which includes SETrans feature extraction module, the transformer-based correlation representation module and the detection module. First, SETrans feature extraction module is employed to extract deep features of HSIs. Then, the transformer-based correlation representation module is presented to explore the spectral dependency of spectral information and capture the correlation of learned features of bi-temporal HSIs from both the perspective of difference and dot-product operations, so as to obtain more discriminative features. Finally, the decision fusion strategy in the detection module is utilized to the learned discriminative features to generate the final change map for better change detection performance. Experimental results on three hyperspectral datasets show that the proposed SATNet is superior to the existing change detection methods.
Wuxia Zhang, Liangxu Su, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.4
2023 Multiple Source Domain Adaptation for Multiple Object Tracking in Satellite Video
abstract
Satellite videos capture the dynamic changes in a large observed sense, which provides an opportunity to track the object trajectories. However, existing multiple object tracking methods require massive video annotations, which is time-consuming and fallible. To alleviate this problem, this paper proposes a Cross-Domain multiple object Tracker (CDTrack) to learn knowledge from multiple source domains. First, a cross-domain object detector with multi-level domain alignment is constructed to learn domain-invariant knowledge between remote sensing images and satellite videos. Second, the proposed method adopts a bidirectional teacher-student framework to fuse multiple source domains. Two teacher-student models learn different domain knowledge and teach mutually each other. With mutual learning, the proposed method alleviates the discrepancies between different domains. Finally, a simple weakly supervised re-identification model is proposed for long-term association. Experimental results on the satellite video datasets demonstrate that the proposed method can achieve great performance without satellite video annotations. The code is available at https://github.com/XiangtaoZheng/CDTrack.
Xiangtao Zheng, Haowen Cui, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.3
2023 Dual Teacher: A Semisupervised Cotraining Framework for Cross-Domain Ship Detection
abstract
Cross-domain ship detection tries to identify Synthetic Aperture Radar (SAR) ship by adapting knowledge from labeled optical images, without labor-intensive annotations. In practical applications, a few (e.g., one or three samples) labeled SAR samples are available, which provides an additional supervision for SAR ships. However, the existing cross-domain methods ignore the SAR supervision (a few labeled and unlabeled SAR images), which limits their performances in a practical and under-investigated task: semi-supervised cross-domain ship detection. In this paper, a Dual Teacher framework is proposed to address the mutual interference between the optical supervision and the SAR supervision. First, both optical and SAR supervision are decomposed into two sub-tasks: cross-domain task and semi-supervised task. Then, both cross-domain task and semi-supervised task can be learned interactively in two individual teacher-student models. The teacher-student models generate pseudo-labels on unlabeled SAR images by a teacher network and fine-tune the student network. Finally, the Dual Teacher framework retrains two teacher-student models in co-training strategies. Both cross-domain dataset and semi-supervised dataset are exploited to jointly improve the pseudo-label quality. The effectiveness of the Dual Teacher framework has been fully experimentally demonstrated. The code is available at https://github.com/XiangtaoZheng/DualTeacher.
Xiangtao Zheng, Haowen Cui, Chujie Xu, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.4
2023 Enhanced Bayesian Factorization With Variant Scale Partitioning for Multivariate Time Series Analysis
abstract
Multivariate time series data (Mv-TSD) portray the evolving processes of the system(s) under examination in a “multi-view” manner. Factorization methods are salient for Mv-TSD analysis with the potentials of structural feature construction correlating various data attributes. However, research challenges remain in the derivation of factors due to highly scattered data distribution of Mv-TSD and intensive interferences/outliers embedded in the source data. The proposed Enhanced Bayesian Factorization approach (Enhanced-BF) addresses the challenges in three phases: (1) variant scale partitioning applies to Mv-TSD according to degree of amplitude and obtains the blocks of variant scales; (2) hierarchical Bayesian model for tensor factorization automatically derives the factors of each block with interferences suppressed; (3) Bayesian unification model merges those block factors to construct the final structural features.Enhanced-BFhas been evaluated using a case study of brain data engineering with multivariate electroencephalogram (EEG). Experimental results indicate that the proposed method manifests robustness to the interferences and outperforms the counterparts in terms of operation efficiency and error when factorizing EEG tensor. Besides,Enhanced-BFexcels in factorization-based analysis of ongoing autism spectrum disorder (ASD) EEG: 3 times speed-up in factorization and$87.35\%$accuracy in ASD discrimination. The latent factors (“biomarkers”) can distinctly interpret the typical EEG characteristics of ASD subjects.
Yunbo Tang, Dan Chen 0001, Yiping Zuo, Xiaoqiang Lu, Rajiv Ranjan 0001, Albert Y. Zomaya, Quanming Yao, Xiaoli Li 0002
IEEE Trans. Knowl. Data Eng.4
2023 Identity Feature Disentanglement for Visible-Infrared Person Re-Identification
abstract
Visible-infrared person re-identification (VI-ReID) task aims to retrieve persons from different spectrum cameras (i.e., visible and infrared images). The biggest challenge of VI-ReID is the huge cross-modal discrepancy caused by different imaging mechanisms. Many VI-ReID methods have been proposed by embedding different modal person images into a shared feature space to narrow the cross-modal discrepancy. However, these methods ignore the purification of identity features, which results in identity features containing different modal information and failing to align well. In this article, an identity feature disentanglement method is proposed to disentangle the identity features from identity-irrelevant information, such as pose and modality. Specifically, images of different modalities are first processed to extract shared features that reduce the cross-modal discrepancy preliminarily. Then the extracted feature of each image is disentangled into a latent identity variable and an identity-irrelevant variable. In order to enforce the latent identity variable to contain as much identity information as possible and as little identity-irrelevant information, an ID-discriminative loss and an ID-swapping reconstruction process are additionally designed. Extensive quantitative and qualitative experiments on two popular public VI-ReID datasets, RegDB and SYSU-MM01, demonstrate the efficacy and superiority of the proposed method.
Xiumei Chen, Xiangtao Zheng, Xiaoqiang Lu
ACM Trans. Multim. Comput. Commun. Appl.3
2022 Semi-Supervised Landcover Classification with Adaptive Pixel-Rebalancing Self-Training
abstract
Semi-supervised learning methods can assist in making use of extensive existing data resources while lowering the cost of manual labeling, which is significant for visual scene under-standing. Previous works on semi-supervised semantic segmentation have paid little attention to class imbalance, leading to the aggravation of the long-tail effect already present in the training data. To address this problem, we propose a novel method adaptive pixel-rebalancing self-training (APRST), which rebalances training data by adaptively sampling at the pixel level, thereby alleviating the class imbalance. Based on APRST, we further design a multi-models with cross pseudo supervision scheme, denoted as APRST+, which alleviates the confirmation bias problem and improves the quality of pseudo annotations. In addition, the segmentation performance is further improved by simply utilizing the DEM data for threshold filtering. Experiment results show that our approach achieves the state-of-the-art in the DFC2022 Track SLM test phase (mIOU of 0.5335).
Xiaoqiang Lu, Guojin Cao, Tong Gou
IGARSS1
2022 Human action recognition by multiple spatial clues network
Xiangtao Zheng, Tengfei Gong, Xiaoqiang Lu, Xuelong Li 0001
Neurocomputing3
2022 Semisupervised Spectral Degradation Constrained Network for Spectral Super-Resolution
abstract
Recently, various deep learning-based methods have been designed to improve the spectral resolution of the multispectral image (MSI) to obtain the hyperspectral image (HSI). These methods usually rely on sufficient MSI/HSI pairs for supervised training. However, collecting plentiful HSIs is time-consuming. In this letter, a semisupervised spectral degradation constrained network (SSDCN) is proposed to improve the spectral resolution of MSI. SSDCN is an autoencoder-like network that is composed of an encoder subnetwork for estimating HSI from input MSI and a decoder subnetwork for reconstructing MSI from the estimated HSI. A semisupervised training method is proposed to explore both MSI/HSI pairs and MSIs without ground-truth HSIs to optimize SSDCN. Simulated and two real databases are employed to demonstrate the effectiveness of SSDCN.
Wenjing Chen 0003, Xiangtao Zheng, Xiaoqiang Lu
IEEE Geosci. Remote. Sens. Lett.3
2022 Remote Sensing Scene Classification by Local-Global Mutual Learning
abstract
Remote sensing scene classification (RSSC) attempts to label an image with a specific scene category. Recently, convolutional neural networks (CNNs) have shown the powerful feature extraction capability to combine local and global features. However, both the local and global features are extracted independently, which ignore the complementary representation. In this letter, a local–global mutual learning (LML) method is proposed to capture both the global and local features. Specifically, local regions are first generated by highlighting the semantic areas in the corresponding original image. Then, a two-branch architecture is used to extract features for the local regions and global image, respectively. Both the classification loss and mutual learning loss are exploited to train the local–global branches simultaneously, which constrain the two branches to promote each other. Experiments on two popular datasets demonstrate the effectiveness of the proposed method.
Xiumei Chen, Xiangtao Zheng, Yue Zhang 0053, Xiaoqiang Lu
IEEE Geosci. Remote. Sens. Lett.4
2022 Meta Self-Supervised Learning for Distribution Shifted Few-Shot Scene Classification
abstract
Few-shot classification tries to recognize novel remote sensing image categories with a few shot samples. However, current methods assume that the test dataset shares the same domain with the labeled training dataset where prior knowledge is learned. It is infeasible to collect a training dataset for each domain, since remote sensing images may come from various domains. Exploiting the existing labeled dataset from another domain (source domain) to help the target dataset (target domain) classification would be efficient. In this paper, both meta-learning and self-supervised learning are jointly conducted for few-shot classification. Specifically, meta-learning is executed over a pre-trained network for few-shot classification. Furthermore, self-supervised learning is exploited to fit the target domain distribution by training on unlabeled target domain images. Experiments are conducted on NWPU, EuroSAT and Merced datasets to validate the effectiveness.
Tengfei Gong, Xiangtao Zheng, Xiaoqiang Lu
IEEE Geosci. Remote. Sens. Lett.3
2022 Pairwise Comparison Network for Remote-Sensing Scene Classification
abstract
Remote-sensing scene classification aims to assign a specific semantic label to a remote-sensing image. Recently, convolutional neural networks (CNNs) have greatly improved the performance of remote-sensing scene classification. However, some confused images may be easily recognized as the incorrect category, which generally degrade the performance. The differences between image pairs can be used to distinguish image categories. This letter proposed a pairwise comparison network (PCNet), which contains two main steps: pairwise selection and pairwise representation. The proposed network first selects similar image pairs and then represents the image pairs with pairwise representations. The self-representation is introduced to highlight the informative parts of each image itself, while the mutual representation is proposed to capture the subtle differences between image pairs. Comprehensive experimental results on two challenging datasets (AID, NWPU-RESISC45) demonstrate the effectiveness of the proposed network.
Yue Zhang 0053, Xiangtao Zheng, Xiaoqiang Lu
IEEE Geosci. Remote. Sens. Lett.3
2022 Reconstructive Sequence-Graph Network for Video Summarization
abstract
Exploiting the inner-shot and inter-shot dependencies is essential for key-shot based video summarization. Current approaches mainly devote to modeling the video as a frame sequence by recurrent neural networks. However, one potential limitation of the sequence models is that they focus on capturing local neighborhood dependencies while the high-order dependencies in long distance are not fully exploited. In general, the frames in each shot record a certain activity and vary smoothly over time, but the multi-hop relationships occur frequently among shots. In this case, both the local and global dependencies are important for understanding the video content. Motivated by this point, we propose a reconstructive sequence-graph network (RSGN) to encode the frames and shots as sequence and graph hierarchically, where the frame-level dependencies are encoded by long short-term memory (LSTM), and the shot-level dependencies are captured by the graph convolutional network (GCN). Then, the videos are summarized by exploiting both the local and global dependencies among shots. Besides, a reconstructor is developed to reward the summary generator, so that the generator can be optimized in an unsupervised manner, which can avert the lack of annotated data in video summarization. Furthermore, under the guidance of reconstruction loss, the predicted summary can better preserve the main video content and shot-level dependencies. Practically, the experimental results on three popular datasets (i.e., SumMe, TVsum and VTW) have demonstrated the superiority of our proposed approach to the summarization task.
Bin Zhao 0001, Haopeng Li 0001, Xiaoqiang Lu, Xuelong Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Semisupervised Consistent Projection Metric Learning for Person Reidentification
abstract
Person reidentification is a hot topic in the computer vision field. Many efforts have been paid on modeling a discriminative distance metric. However, existing metric-learning-based methods are a lack of generalization. In this article, the poor generalization of the metric model is argued as the biased estimation problem that the independent identical distribution hypothesis is not valid. The verification experimental result shows that there is a sharp difference between the training and test samples in the metric subspace. A semisupervised consistent projection metric-learning method is proposed to ease the biased estimation problem by learning a consistent constrained metric subspace in which the identified pairs are forced to follow the distribution of the positive training pairs. First, a semisupervised method is proposed to generate potential matching pairs from the k -nearest neighbors of test samples. The potential matching pairs are used to estimate the distances' distribution center of the positive test pairs. Second, the metric subspace is improved by forcing this estimation to be close to the center of the positive training pairs. Finally, extensive experiments are conducted on five datasets and the results demonstrate that the proposed method reaches the best performance, especially on the rank-1 identification rate.
Bangyong Sun, Yutao Ren, Xiaoqiang Lu
IEEE Trans. Cybern.3
2022 A Novel NMF Guided for Hyperspectral Unmixing From Incomplete and Noisy Data
abstract
The nonnegative matrix factorization (NMF)-combined spatial–spectral information has been widely applied in the unmixing of hyperspectral images (HSIs). However, how to select the appropriate similarity pixels and explore the spatial information and how to adapt the unmixing algorithm to complex data are both great challenges. In this article, we propose a novel unmixing method named spatial–spectral neighborhood preserving NMF (SSNPNMF) for incomplete and noisy HSI data. First, a spatial–spectral kernel regularizer is introduced to preprocess the HSI, which can reduce noise and complete missing elements. Second, a distance metric SSD based on spatial–spectral information is designed to select similar pixels in the image. Subsequently, the spatial–spectral relationship of the selected first$k$similar pixels is used to reconstruct the image and obtain the reconstruction matrix. Finally, the reconstruction matrix is used to constrain the abundances and improve the unmixing performance. Experimental results on synthetic data and Cuprite data indicate that SSNPNMF has a more effective unmixing performance compared with the state-of-the-art methods.
Xiaoqiang Lu, Ganchao Liu, Yuan Yuan 0001
IEEE Trans. Geosci. Remote. Sens.2
2022 Simple and Efficient: A Semisupervised Learning Framework for Remote Sensing Image Semantic Segmentation
abstract
Semantic segmentation based on deep learning has achieved impressive results in recent years, but these results are supported by a large amount of labeled data which requires intensive annotation at the pixel level, particularly for high-resolution remote sensing (RS) images. In this work, we propose a simple yet efficient semisupervised learning framework based on linear sampling self-training, named LSST, to improve the performance of RS image semantic segmentation. Specifically, the classical pseudo-labeling-based self-training paradigm is enhanced by injecting strong data augmentations (SDA) applicable to RS images, based on which a powerful baseline is constructed. Nevertheless, the problem of insufficient data training to generate pseudo-labels with a high level of noise persists, and the noisy pseudo-labels will continue to accumulate and impede model improvement during the re-training phase. Previous works commonly employ a pre-defined threshold to remove noise, but it will lead to overfitting the model to easily identified classes. To address it, a method using linear sampling (LS) is presented for assigning thresholds to different classes in an adaptive manner, which provides noiseless regions for re-training. Experiments prove that the proposed pixel-wise selection is more available for segmentation than image-level selection in RS images. Finally, LSST achieves state-of-the-art on several datasets and different evaluation metrics. The source code of the this paper is available at https://github.com/xiaoqiang-lu/LSST.
Xiaoqiang Lu, Licheng Jiao, Fang Liu 0001, Shuyuan Yang 0001, Xu Liu 0006, Zhixi Feng, Lingling Li 0002, Puhua Chen
IEEE Trans. Geosci. Remote. Sens.1
2022 Cross-Attention Spectral-Spatial Network for Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) classification aims to identify categories of hyperspectral pixels. Recently, many convolutional neural networks (CNNs) have been designed to explore the spectrums and spatial information of HSI for classification. In recent CNN-based methods, 2-D or 3-D convolutions are inevitably utilized as basic operations to extract the spatial or spectral–spatial features. However, 2-D and 3-D convolutions are sensitive to the image rotation, which may result in that recent CNN-based methods are not robust to the HSI rotation. In this article, a cross-attention spectral–spatial network (CASSN) is proposed to alleviate the problem of HSI rotation. First, a cross-spectral attention component is proposed to exploit the local and global spectrums of the pixel to generate band weight for suppressing redundant bands. Second, a spectral feature extraction component is utilized to capture spectral features. Then, a cross-spatial attention component is proposed to generate spectral–spatial features from the HSI patch under the guidance of the pixel to be classified. Finally, the spectral–spatial feature is fed to a softmax classifier to obtain the category. The effectiveness of CASSN is demonstrated on three public databases.
Hao Sun 0014, Chunbo Zou, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.4
2022 Spectral Super-Resolution of Multispectral Images Using Spatial-Spectral Residual Attention Network
abstract
The spectral super-resolution of multispectral image (MSI) refers to improving the spectral resolution of the MSI to obtain the hyperspectral image (HSI). Most recent works are based on the sparse representation to unfold the MSI into the 2-D matrix in advance for subsequent operations, which results in that the spatial information of MSI cannot be fully explored. In this article, a spatial–spectral residual attention network (SSRAN) is proposed to simultaneously explore the spatial and spectral information of MSI for reconstructing the HSI. The proposed SSRAN is composed of the feature extraction part, the nonlinear mapping part, and the reconstruction part. Firstly, the multispectral features of the input MSI are extracted in the feature extraction part. Second, in the nonlinear mapping part, the spatial–spectral residual blocks are proposed to explore spatial and spectral information of MSI for mapping the multispectral features to the hyperspectral features. Finally, in the reconstruction part, a 2-D convolution is used to reconstruct the HSI from the hyperspectral features. Also, a neighboring spectral attention module is specially designed to explicitly constrain the reconstructed HSI to maintain the correlation among neighboring spectral bands. The proposed SSRAN outperforms the state-of-the-art methods on both simulated and real databases.
Xiangtao Zheng, Wenjing Chen 0003, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.3
2022 Unsupervised Change Detection by Cross-Resolution Difference Learning
abstract
Change detection (CD) aims to identify the differences between multitemporal images acquired over the same geographical area at different times. With the advantages of requiring no cumbersome labeled change information, unsupervised CD has attracted extensive attention of researchers. Multitemporal images tend to have different resolutions as they are usually captured at different times with different sensor properties. It is difficult to directly obtain one pixelwise change map for two images with different resolutions, so current methods usually resize multitemporal images to a unified size. However, resizing operations change the original information of pixels, which limits the final CD performance. This article aims to detect changes from multitemporal images in the originally different resolutions without resizing operations. To achieve this, a cross-resolution difference learning method is proposed. Specifically, two cross-resolution pixelwise difference maps are generated for the two different resolution images and fused to produce the final change map. First, the two input images are segmented into individual homogeneous regions separately due to different resolutions. Second, each pixelwise difference map is produced according to two measure distances, the mutual information distance and the deep feature distance, between image regions in which the pixel lies. Third, the final binary change map is generated by fusing and binarizing the two cross-resolution difference maps. Extensive experiments on four datasets demonstrate the effectiveness of the proposed method for detecting changes from different resolution images.
Xiangtao Zheng, Xiumei Chen, Xiaoqiang Lu, Bangyong Sun
IEEE Trans. Geosci. Remote. Sens.3
2022 Generalized Scene Classification From Small-Scale Datasets With Multitask Learning
abstract
Remote sensing images contain a wealth of spatial information. Efficient scene classification is a necessary precedent step for further application. Despite the great practical value, the mainstream methods using deep convolutional neural networks (CNNs) are generally pretrained on other large datasets (such as ImageNet) and thus fail to capture the specific visual characteristics of remote sensing images. For another, it lacks the generalization ability to new tasks when training a new CNN from scratch with an existing remote sensing dataset. This article addresses the dilemma and uses multiple small-scale datasets to learn a generalized model for efficient scene classification. Since the existing datasets are heterogeneous and cannot be directly combined to train a network, a multitask learning network (MTLN) is developed. The MTLN treats each small-scale dataset as an individual task and uses complementary information contained in multiple tasks to improve generalization. Concretely, the MTLN consists of a shared branch for all tasks and multiple task-specific branches with each for one task. The shared branch extracts shared features for all tasks to achieve information sharing among tasks. The task-specific branch distills the shared features into task-specific features toward the optimal estimation of each specific task. By jointly learning shared features and task-specific features, the MTLN maintains both generalization and discrimination abilities. Two types of MTL scenarios are explored to validate the effectiveness of the proposed method: one is to complete multiple scene classification tasks and the other is to jointly perform scene classification and semantic segmentation.
Xiangtao Zheng, Tengfei Gong, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.4
2022 Mutual Attention Inception Network for Remote Sensing Visual Question Answering
abstract
Remote sensing images (RSIs) containing various ground objects have been applied in many fields. To make semantic understanding of RSIs objective and interactive, the task remote sensingvisual question answering(VQA) has appeared. Given an RSI, the goal of remote sensing VQA is to make an intelligent agent answer a question about the remote sensing scene. Existing remote sensing VQA methods utilized a nonspatial fusion strategy to fuse the image features and question features, which ignores the spatial information of images and word-level information of questions. A novel method is proposed to complete the task considering these two aspects. First, convolutional features of the image are included to represent spatial information, and the word vectors of questions are adopted to present semantic word information. Second, attention mechanism and bilinear technique are introduced to enhance the feature considering the alignments between spatial positions and words. Finally, a fully connected layer with softmax is utilized to output an answer from the perspective of the multiclass classification task. To benchmark this task, aRSIVQAdataset is introduced in this article. For each of more than 37 000 RSIs, the proposed dataset contains at least one or more questions, plus corresponding answers. Experimental results demonstrate that the proposed method can capture the alignments between images and questions. The code and dataset are available athttps://github.com/spectralpublic/RSIVQA.
Xiangtao Zheng, Binqiang Wang, Xingqian Du, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.4
2022 Visible-Infrared Person Re-Identification via Partially Interactive Collaboration
abstract
Visible-infrared person re-identification (VI-ReID) task aims to retrieve the same person between visible and infrared images. VI-ReID is challenging as the images captured by different spectra present large cross-modality discrepancy. Many methods adopt a two-stream network and design additional constraint conditions to extract shared features for different modalities. However, the interaction between the feature extraction processes of different modalities is rarely considered. In this paper, a partially interactive collaboration method is proposed to exploit the complementary information of different modalities to reduce the modality gap for VI-ReID. Specifically, the proposed method is achieved in a partially interactive-shared architecture: collaborative shallow layers and shared deep layers. The collaborative shallow layers consider the interaction between modality-specific features of different modalities, encouraging the feature extraction processes of different modalities constrain each other to enhance feature representations. The shared deep layers further embed the modality-specific features to a common space to endow them the same identity discriminability. To ensure the interactive collaborative learning implement effectively, the conventional loss and collaborative loss are utilized jointly to train the whole network. Extensive experiments on two publicly available VI-ReID datasets verify the superiority of the proposed PIC method. Specifically, the proposed method achieves a rank-1 accuracy of 83.6% and 57.5% on RegDB and SYSU-MM01 datasets, respectively.
Xiangtao Zheng, Xiumei Chen, Xiaoqiang Lu
IEEE Trans. Image Process.3
2022 Rotation-Invariant Attention Network for Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) classification refers to identifying land-cover categories of pixels based on spectral signatures and spatial information of HSIs. In recent deep learning-based methods, to explore the spatial information of HSIs, the HSI patch is usually cropped from original HSI as the input. And 3 ×3 convolution is utilized as a key component to capture spatial features for HSI classification. However, the 3 ×3 convolution is sensitive to the spatial rotation of inputs, which results in that recent methods perform worse in rotated HSIs. To alleviate this problem, a rotation-invariant attention network (RIAN) is proposed for HSI classification. First, a center spectral attention (CSpeA) module is designed to avoid the influence of other categories of pixels to suppress redundant spectral bands. Then, a rectified spatial attention (RSpaA) module is proposed to replace 3 ×3 convolution for extracting rotation-invariant spectral-spatial features from HSI patches. The CSpeA module, the 1 ×1 convolution and the RSpaA module are utilized to build the proposed RIAN for HSI classification. Experimental results demonstrate that RIAN is invariant to the spatial rotation of HSIs and has superior performance, e.g., achieving an overall accuracy of 86.53% (1.04% improvement) on the Houston database. The codes of this work are available at https://github.com/spectralpublic/RIAN.
Xiangtao Zheng, Hao Sun 0014, Xiaoqiang Lu, Wei Xie 0008
IEEE Trans. Image Process.3
2022 Disentangled Representation Learning for Cross-Modal Biometric Matching
abstract
Cross-modal biometric matching (CMBM) aims to determine the corresponding voice from a face, or identify the corresponding face from a voice. Recently, many CMBM methods have been proposed by forcing the distance between two modal features to be narrowed. However, these methods ignore the alignability between the two modal features. Because the feature is extracted under the supervision of identity information from single modal data, it can only reflect the identity information of single modal data. In order to address this problem, a disentangled representation learning method is proposed to disentangle the alignable latent identity factors and nonalignable the modality-dependent factors for CMBM. The proposed method consists of two main steps: 1) feature extraction and 2) disentangled representation learning. Firstly, an image feature extraction network is adopted to obtain face features, and a voice feature extraction network is applied to learn voice features. Secondly, a disentangled latent variable is explored to disentangle the latent identity factors that are shared across the modalities from the modality-dependent factors. The modality-dependent factors are filtered out, while the latent identity factors from the two modalities are enforced to be narrowed to align the same identity information. Then, the disentangled latent identity factors are considered as pure identity information to bridge the two modalities for cross-modal verification, 1:$N$matching, and retrieval. Note that the proposed method learns the identity information from the input face images and voice segments with only identity label as supervised information. Extensive experiments on the challenging VoxCeleb dataset demonstrate the proposed method outperforms the state-of-the-art methods.
Hailong Ning, Xiangtao Zheng, Xiaoqiang Lu, Yuan Yuan 0001
IEEE Trans. Multim.3
2021 Audio description from image by modal translation network
Hailong Ning, Xiangtao Zheng, Yuan Yuan 0001, Xiaoqiang Lu
Neurocomputing4
2021 Local and correlation attention learning for subtle facial expression recognition
Yuan Yuan 0001, Xiangtao Zheng, Xiaoqiang Lu
Neurocomputing4
2021 Adaptive multiscale feature for object detection
Xiaoyong Yu, Xiaoqiang Lu, Guilong Gao
Neurocomputing3
2021 Generative Adversarial Capsule Network With ConvLSTM for Hyperspectral Image Classification
abstract
Recently, deep learning has been widely applied in hyperspectral image (HSI) classification since it can extract high-level spatial-spectral features. However, deep learning methods are restricted due to the lack of sufficient annotated samples. To address this problem, this letter proposes a novel generative adversarial network (GAN) for HSI classification that can generate artificial samples for data augmentation to improve the HSI classification performance with few training samples. In the proposed network, a new discriminator is designed by exploiting capsule network (CapsNet) and convolutional long short-term memory (ConvLSTM), which extracts the low-level features and combines them together with local space sequence information to form the high-level contextual features. In addition, a structured sparse L2,1constraint is imposed on sample generation to control the modes of data being generated and achieve more stable training. The experimental results on two real HSI data sets show that the proposed method can obtain better classification performance than the several state-of-the-art deep classification methods.
Wei-Ye Wang, Heng-Chao Li 0001, Yangjun Deng, Li-Yang Shao, Xiaoqiang Lu, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.5
2021 Remote Sensing Image Generation From Audio
abstract
Generating image from other modal data has attracted much attention in cross-modal studies, since the generated image offers intuitive vision information. Unlike the previous works which generate an image from text, a novel task is introduced, generating an image from audio. However, semantic gap intrinsically exists in cross-modal data, which disturbs the generative results. In order to explore the relevance between the audio and image, a novel reranking audio-image translation method is proposed. The proposed method: 1) maps the audio and image into a uniform feature space; 2) designs an audio-audio matching network to match the related audio; and 3) adopts an audio-image matching network for every matched audio to generate a related image, and the most frequent image is voted as the final result. Extensive experiments on two remote sensing cross-modal data sets demonstrate that the proposed method can visualize the content of audio.
Jun Chen 0001, Xiangtao Zheng, Xiaoqiang Lu
IEEE Geosci. Remote. Sens. Lett.4
2021 Deep Category-Level and Regularized Hashing With Global Semantic Similarity Learning
abstract
The hashing technique has been extensively used in large-scale image retrieval applications due to its low storage and fast computing speed. Most existing deep hashing approaches cannot fully consider the global semantic similarity and category-level semantic information, which result in the insufficient utilization of the global semantic similarity for hash codes learning and the semantic information loss of hash codes. To tackle these issues, we propose a novel deep hashing approach with triplet labels, namely, deep category-level and regularized hashing (DCRH), to leverage the global semantic similarity of deep feature and category-level semantic information to enhance the semantic similarity of hash codes. There are four contributions in this article. First, we design a novel global semantic similarity constraint about the deep feature to make the anchor deep feature more similar to the positive deep feature than to the negative deep feature. Second, we leverage label information to enhance category-level semantics of hash codes for hash codes learning. Third, we develop a new triplet construction module to select good image triplets for effective hash functions learning. Finally, we propose a new triplet regularized loss (Reg-L) term, which can force binary-like codes to approximate binary codes and eventually minimize the information loss between binary-like codes and binary codes. Extensive experimental results in three image retrieval benchmark datasets show that the proposed DCRH approach achieves superior performance over other state-of-the-art hashing approaches.
Yaxiong Chen, Xiaoqiang Lu
IEEE Trans. Cybern.2
2021 Person Reidentification via Unsupervised Cross-View Metric Learning
abstract
Person reidentification (Re-ID) aims to match observations of individuals across multiple nonoverlapping camera views. Recently, metric learning-based methods have played important roles in addressing this task. However, metrics are mostly learned in supervised manners, of which the performance relies heavily on the quantity and quality of manual annotations. Meanwhile, metric learning-based algorithms generally project person features into a common subspace, in which the extracted features are shared by all views. However, it may result in information loss since these algorithms neglect the view-specific features. Besides, they assume person samples of different views are taken from the same distribution. Conversely, these samples are more likely to obey different distributions due to view condition changes. To this end, this paper proposes an unsupervised cross-view metric learning method based on the properties of data distributions. Specifically, person samples in each view are taken from a mixture of two distributions: one models common prosperities among camera views and the other focuses on view-specific properties. Based on this, we introduce a shared mapping to explore the shared features. Meanwhile, we construct view-specific mappings to extract and project view-related features into a common subspace. As a result, samples in the transformed subspace follow the same distribution and are equipped with comprehensive representations. In this paper, these mappings are learned in an unsupervised manner by clustering samples in the projected space. Experimental results on five cross-view datasets validate the effectiveness of the proposed method.
Yachuang Feng, Yuan Yuan 0001, Xiaoqiang Lu
IEEE Trans. Cybern.3
2021 Vision-to-Language Tasks Based on Attributes and Attention Mechanism
abstract
Vision-to-language tasks aim to integrate computer vision and natural language processing together, which has attracted the attention of many researchers. For typical approaches, they encode image into feature representations and decode it into natural language sentences. While they neglect high-level semantic concepts and subtle relationships between image regions and natural language elements. To make full use of these information, this paper attempt to exploit the text-guided attention and semantic-guided attention (SA) to find the more correlated spatial information and reduce the semantic gap between vision and language. Our method includes two-level attention networks. One is the text-guided attention network which is used to select the text-related regions. The other is SA network which is used to highlight the concept-related regions and the region-related concepts. At last, all these information are incorporated to generate captions or answers. Practically, image captioning and visual question answering experiments have been carried out, and the experimental results have shown the excellent performance of the proposed approach.
Xuelong Li 0001, Aihong Yuan, Xiaoqiang Lu
IEEE Trans. Cybern.3
2021 Bio-Inspired Representation Learning for Visual Attention Prediction
abstract
Visual attention prediction (VAP) is a significant and imperative issue in the field of computer vision. Most of the existing VAP methods are based on deep learning. However, they do not fully take advantage of the low-level contrast features while generating the visual attention map. In this article, a novel VAP method is proposed to generate the visual attention map via bio-inspired representation learning. The bio-inspired representation learning combines both low-level contrast and high-level semantic features simultaneously, which are developed by the fact that the human eye is sensitive to the patches with high contrast and objects with high semantics. The proposed method is composed of three main steps: 1) feature extraction; 2) bio-inspired representation learning; and 3) visual attention map generation. First, the high-level semantic feature is extracted from the refined VGG16, while the low-level contrast feature is extracted by the proposed contrast feature extraction block in a deep network. Second, during bio-inspired representation learning, both the extracted low-level contrast and high-level semantic features are combined by the designed densely connected block, which is proposed to concatenate various features scale by scale. Finally, the weighted-fusion layer is exploited to generate the ultimate visual attention map based on the obtained representations after bio-inspired representation learning. Extensive experiments are performed to demonstrate the effectiveness of the proposed method.
Yuan Yuan 0001, Hailong Ning, Xiaoqiang Lu
IEEE Trans. Cybern.3
2021 Spectral-Spatial Joint Sparse NMF for Hyperspectral Unmixing
abstract
The nonnegative matrix factorization (NMF) combining with spatial-spectral contextual information is an important technique for extracting endmembers and abundances of hyperspectral image (HSI). Most methods constrain unmixing by the local spatial position relationship of pixels or search spectral correlation globally by treating pixels as an independent point in HSI. Unfortunately, they ignore the complex distribution of substance and rich contextual information, which makes them effective in limited cases. In this article, we propose a novel unmixing method via two types of self-similarity to constrain sparse NMF. First, we explore the spatial similarity patch structure of data on the whole image to construct the spatial global self-similarity group between pixels. And according to the regional continuity of the feature distribution, the spectral local self-similarity group of pixels is created inside the superpixel. Then based on the sparse expression of the pixel in the subspace, we sparsely encode the pixels in the same spatial group and spectral group respectively. Finally, the abundance of pixels within each group is forced to be similar to constrain the NMF unmixing framework. Experiments on synthetic and real data fully demonstrate the superiority of our method over other existing methods.
Yuan Yuan 0001, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.3
2021 Multisource Remote Sensing Data Classification With Graph Fusion Network
abstract
The land cover classification has been an important task in remote sensing. With the development of various sensors technologies, carrying out classification work with multisource remote sensing (MSRS) data has shown an advantage over using a single type of data. Hyperspectral images (HSIs) are able to represent the spectral properties of land cover, which is quite common for land cover understanding. Light detection and ranging (LiDAR) images contain altitude information of the ground, which is greatly helpful with urban scene analysis. Current HSI and LiDAR fusion methods perform feature extraction and feature fusion separately, which cannot well exploit the correlation of data sources. In order to make full use of the correlation of multisource data, an unsupervised feature extraction-fusion network for HSI and LiDAR, which utilizes feature fusion to guide the feature extraction procedure, is proposed in this article. More specifically, the network takes multisource data as input and directly output the unified fused feature. A multimodal graph is constructed for feature fusion, and graph-based loss functions including Laplacian loss and t-distributed stochastic neighbor embedding (t-SNE) loss are utilized to constrain the feature extraction network. Experimental results on several data sets demonstrate the proposed network can achieve more excellent classification performance than some state-of-the-art methods.
Xingqian Du, Xiangtao Zheng, Xiaoqiang Lu, Alexander A. Doudkin
IEEE Trans. Geosci. Remote. Sens.3
2021 Cross-Domain Scene Classification by Integrating Multiple Incomplete Sources
abstract
Cross-domain scene classification identifies scene categories by learning knowledge from a labeled data set (source domain) to an unlabeled data set (target domain), where the source data and the target data are sampled from different distributions. A lot of domain adaptation methods are used to reduce the distribution shift across domains, and most existing methods assume that the source domain shares the same categories with the target domain. It is usually hard to find a source domain that covers all categories in the target domain. Some works exploit multiple incomplete source domains to cover the target domain. However, in such setting, the categories of each source domain are a subset of the target-domain categories, and the target domain contains “unknown” categories for each source domain. The existence of unknown categories results in the conventional domain adaptation unsuitable. Known and unknown categories should be treated separately. Therefore, a separation mechanism is proposed to separate the known and unknown categories in this article. First, multiple-source classifiers trained on the multiple source domains are used to coarsely separate the known/unknown categories in the target domain. The target images with high similarities to source images are selected as known categories, and the target images with low similarities are selected as unknown categories. Then, a binary classifier trained using the selected images is used to finely separate all target-domain images. Finally, only the known categories are implemented in the cross-domain alignment and classification. The target images get labels by integrating the hypotheses of multiple-source classifiers on the known categories. Experiments are conducted on three cross-domain data sets to demonstrate the effectiveness of the proposed method.
Tengfei Gong, Xiangtao Zheng, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.3
2021 Bidirectional Interaction Network for Person Re-Identification
abstract
Person re-identification (ReID) task aims to retrieve the same person across multiple spatially disjoint camera views. Due to huge image changes caused by various factors such as posture variation and illumination transformation, images of different persons may share the more similar appearances than images of the same one. Learning discriminative representations to distinguish details of different persons is significant for person ReID. Many existing methods learn discriminative representations resorting to a human body part location branch which requires cumbersome expert human annotations or complex network designs. In this article, a novel bidirectional interaction network is proposed to explore discriminative representations for person ReID without any human body part detection. The proposed method regards multiple convolutional features as responses to various body part properties and exploits the inter-layer interaction to mine discriminative representations for person identities. Firstly, an inter-layer bilinear pooling strategy is proposed to feasibly exploit the pairwise feature relations between two convolution layers. Secondly, to explore interaction of multiple layers, an effective bidirectional integration strategy consisting of two different multi-layer interaction processes is designed to aggregate bilinear pooling interaction of multiple convolution layers. The interaction of multiple layers is implemented in a layer-by-layer nesting policy to ensure the two interaction processes are different and complementary. Extensive experiments validate the superiority of the proposed method on four popular person ReID datasets including Market-1501, DukeMTMC-ReID, CUHK03-NP and MSMT17. Specifically, the proposed method achieves a rank-1 accuracy of 95.1% and 88.2% on Market-1501 and DukeMTMC-ReID, respectively.
Xiumei Chen, Xiangtao Zheng, Xiaoqiang Lu
IEEE Trans. Image Process.3
2021 A Supervised Segmentation Network for Hyperspectral Image Classification
abstract
Recently, deep learning has drawn broad attention in the hyperspectral image (HSI) classification task. Many works have focused on elaborately designing various spectral-spatial networks, where convolutional neural network (CNN) is one of the most popular structures. To explore the spatial information for HSI classification, pixels with its adjacent pixels are usually directly cropped from hyperspectral data to form HSI cubes in CNN-based methods. However, the spatial land-cover distributions of cropped HSI cubes are usually complicated. The land-cover label of a cropped HSI cube cannot simply be determined by its center pixel. In addition, the spatial land-cover distribution of a cropped HSI cube is fixed and has less diversity. For CNN-based methods, training with cropped HSI cubes will result in poor generalization to the changes of spatial land-cover distributions. In this paper, an end-to-end fully convolutional segmentation network (FCSN) is proposed to simultaneously identify land-cover labels of all pixels in a HSI cube. First, several experiments are conducted to demonstrate that recent CNN-based methods show the weak generalization capabilities. Second, a fine label style is proposed to label all pixels of HSI cubes to provide detailed spatial land-cover distributions of HSI cubes. Third, a HSI cube generation method is proposed to generate plentiful HSI cubes with fine labels to improve the diversity of spatial land-cover distributions. Finally, a FCSN is proposed to explore spectral-spatial features from finely labeled HSI cubes for HSI classification. Experimental results show that FCSN has the superior generalization capability to the changes of spatial land-cover distributions.
Hao Sun 0014, Xiangtao Zheng, Xiaoqiang Lu
IEEE Trans. Image Process.3
2021 Fine-Grained Visual Categorization by Localizing Object Parts With Single Image
abstract
Fine-grained visual categorization (FGVC) refers to assigning fine-grained labels to images which belong to the same base category. Due to the high inter-class similarity, it is challenging to distinguish fine-grained images under different subcategories. Recently, researchers have proposed to firstly localize key object parts within images and then find discriminative clues on object parts. To localize object parts, existing methods train detectors for different kinds of object parts. However, due to the fact that the same kind of object part in different images often changes intensely in appearance, the existing methods face two shortages: 1) Training part detector for object parts with diverse appearance is laborious; 2) Discriminative parts with unusual appearance may be neglected by the trained part detectors. To localize the key object parts efficiently and accurately, a novel FGVC method is proposed in the paper. The main novelty is that the proposed method localizes the key object parts within each image only depending on a single image and hence avoid the influence of diversity between parts in different images. The proposed FGVC method consists of two key steps. Firstly, the proposed method localizes the key parts in each image independently. To this end, potential object parts in each image are identified and then these potential parts are merged to generate the final representative object parts. Secondly, two kinds of features are extracted for simultaneously describing the discriminative clues within each part and the relationship between object parts. In addition, a part based dropout learning technique is adopted to boost the classification performance further in the paper. The proposed method is evaluated in comparison experiments and the experiment results show that the proposed method can achieve comparable or better performance than state-of-the-art methods.
Xiangtao Zheng, Lei Qi 0004, Yutao Ren, Xiaoqiang Lu
IEEE Trans. Multim.4
2020 Unregistered Hyperspectral and Multispectral Image Fusion with Synchronous Nonnegative Matrix Factorization
Wenjing Chen 0003, Xiaoqiang Lu
PRCV (1)2
2020 Mobile person re-identification with a lightweight trident CNN
Mingfu Xiong, Dan Chen 0001, Xiaoqiang Lu
Sci. China Inf. Sci.3
2020 Deep discrete hashing with pairwise correlation learning
Yaxiong Chen, Xiaoqiang Lu
Neurocomputing2
2020 Deep balanced discrete hashing for image retrieval
Xiangtao Zheng, Yichao Zhang 0005, Xiaoqiang Lu
Neurocomputing3
2020 Spatial attention based visual semantic learning for action recognition in still images
Yunpeng Zheng, Xiangtao Zheng, Xiaoqiang Lu
Neurocomputing3
2020 Supervised deep hashing with a joint deep network
Yaxiong Chen, Xiaoqiang Lu, Xuelong Li 0001
Pattern Recognit.2
2020 Deep Cross-Modal Image-Voice Retrieval in Remote Sensing
abstract
With the rapid progress of satellite and aircraft technologies, cross-modal remote sensing image-voice retrieval has been studied in geography recently. However, there still exist some bottlenecks: how to consider the characteristics of remote sensing data adequately and how to reduce the memory and improve the retrieval efficiency in large-scale remote sensing data. In this article, we propose a novel deep cross-modal remote sensing image-voice retrieval approach, namely, deep image-voice retrieval (DIVR), to capture more information of remote sensing data to generate hash codes with low memory and fast retrieval properties. Especially, the DIVR approach proposes inception dilated convolution module to capture multiscale contextual information of remote sensing images and voices. Moreover, in order to enhance cross-modal similarity, the deep features' similarity term is designed to make paired similar deep features as close as possible and paired dissimilar deep features as mutually far as possible. In addition, the quantization error term is designed to drive hash-like codes to approximate hash codes, which can effectively reduce the quantization error for hash codes' learning. Extensive experimental results on three remote sensing image-voice data sets show that the proposed DIVR approach can outperform other cross-modal retrieval approaches.
Yaxiong Chen, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.2
2020 Subspace Clustering Constrained Sparse NMF for Hyperspectral Unmixing
abstract
As one of the most important information of hyperspectral images (HSI), spatial information is usually simulated with the similarity among pixels to enhance the unmixing performance of nonnegative matrix factorization (NMF). Nevertheless, the similarity is generally calculated based on the Euclidean distance between pairwise pixels, which is sensitive to noise and fails in capturing subspace information of hyperspectral data. In addition, it is independent of the NMF framework. In this article, we propose a novel unmixing method called subspace clustering constrained sparse NMF (SC-NMF) for hyperspectral unmixing to more accurately extract endmembers and correspond abundances. First, the nonnegative subspace clustering is embedded into the NMF framework to learn a similar graph, which takes full advantage of the characteristics of the reconstructed data itself to extract the spatial correlation of pixels for unmixing. It is noteworthy that the similar graph and NMF will be simultaneously updated. Second, to mitigate the influence of noise in HSI, only the k largest values are retained in each self-expression vector. Finally, we use the idea of subspace clustering to extract endmembers by linearly combining of all pixels in spectral subspace, aiming at giving a reasonable physical significance to the endmembers. We evaluate the proposed SC-NMF on both synthetic and real hyperspectral data, and experimental results demonstrate that the proposed method is effective and superior by comparing with the state-of-the-art methods.
Xiaoqiang Lu, Yuan Yuan 0001
IEEE Trans. Geosci. Remote. Sens.1
2020 Multisource Compensation Network for Remote Sensing Cross-Domain Scene Classification
abstract
Cross-domain scene classification refers to the scene classification task in which the training set (termed source domain) and the test set (termed target domain) come from different distributions. Various domain adaptation methods have been developed to reduce the distribution discrepancy between different domains. However, current domain adaptation methods assume that the source domain and target domain share the same categories. In reality, it is hard to find a source domain that can completely cover all the categories of target domain. In this article, we propose to use multiple complementary source domains to form the categories of target domain. A multisource compensation network (MSCN) is proposed to tackle these challenges: distribution discrepancy and category incompleteness. First, a pretrained convolutional neural network (CNN) is exploited to learn the feature representation for each domain. Second, a cross-domain alignment module is developed to reduce the domain shift between source and target domains. Domain shift is reduced by mapping the two domain features into a common feature space. Finally, a classifier complement module is proposed to align categories in multiple sources and learn a target classifier. Two cross-domain classification data sets are constructed using four heterogeneous remote sensing scene classification data sets. Extensive experiments are conducted on these datasets to validate the effectiveness of the proposed method. The proposed method can achieve 81.23% and 81.97% average accuracies on two-source-complementary data set and three-source-complementary data set, respectively.
Xiaoqiang Lu, Tengfei Gong, Xiangtao Zheng
IEEE Trans. Geosci. Remote. Sens.1
2020 Sound Active Attention Framework for Remote Sensing Image Captioning
abstract
Attention mechanism-based image captioning methods have achieved good results in the remote sensing field, but are driven by tagged sentences, which is called passive attention. However, different observers may give different levels of attention to the same image. The attention of observers during testing, then, may not be consistent with the attention during training. As a direct and natural human-machine interaction, speech is much faster than typing sentences. Sound can represent the attention of different observers. This is called active attention. Active attention can be more targeted to describe the image; for example, in disaster assessments, the situation can be obtained quickly and the corresponding disaster areas can be located related to the specific disaster. A novel sound active attention framework is proposed for more specific caption generation according to the interest of the observer. First, sound is modeled by mel-frequency cepstral coefficients (MFCCs) and the image is encoded by convolutional neural networks (CNNs). Then, to handle the continuity characteristic of sound, a sound module and an attention module are designed based on the gated recurrent units (GRUs). Finally, the sound-guided image feature processed by the attention module is imported into the output module to generate descriptive sentence. Experiments based on both fake and real sound data sets show that the proposed method can generate sentences that can capture the focus of human.
Xiaoqiang Lu, Binqiang Wang, Xiangtao Zheng
IEEE Trans. Geosci. Remote. Sens.1
2020 Exploiting Embedding Manifold of Autoencoders for Hyperspectral Anomaly Detection
abstract
Hyperspectral anomaly detection is an important task in the remote sensing domain. Recently, researchers have shown great interest in deep learning-based methods because they can learn hierarchical, abstract, and high-level representations. However, the latent features learned from the autoencoder (AE) are not always able to reflect the intrinsic structure of hyperspectral data because the locality property is not considered during the learning process. In order to address this problem, a novel manifold constrained AE network (MC-AEN)-based hyperspectral anomaly detection method is proposed in this article. First, the manifold learning method is employed to learn the embedding manifold. Then, the latent representations are learned by an AE network with the learned embedding manifold constraints to preserve the intrinsic structure of hyperspectral data. Finally, the reconstruction errors are calculated to detect anomalies. The global reconstruction error from MC-AEN and the local reconstruction error from the learned latent representations are combined to fully utilize the learned knowledge for better detection performance. We test our proposed algorithm on three different real data sets. Experimental results on these three data sets show the superiority of our proposed method.
Xiaoqiang Lu, Wuxia Zhang, Ju Huang
IEEE Trans. Geosci. Remote. Sens.1
2020 Gated and Axis-Concentrated Localization Network for Remote Sensing Object Detection
abstract
In the multicategory object detection task of high-resolution remote sensing images, small objects are always difficult to detect. This happens because the influence of location deviation on small object detection is greater than on large object detection. The reason is that, with the same intersection decrease between a predicted box and a true box, Intersection over Union (IoU) of small objects drops more than those of large objects. In order to address this challenge, we propose a new localization model to improve the location accuracy of small objects. This model is composed of two parts. First, a global feature gating process is proposed to implement a channel attention mechanism on local feature learning. This process takes full advantages of global features’ abundant semantics and local features’ spatial details. In this case, more effective information is selected for small object detection. Second, an axis-concentrated prediction (ACP) process is adopted to project convolutional feature maps into different spatial directions, so as to avoid interference between coordinate axes and improve the location accuracy. Then, coordinate prediction is implemented with a regression layer using the learned object representation. In our experiments, we explore the relationship between the detection accuracy and the object scale, and the results show that the performance improvements of small objects are distinct using our method. Compared with the classical deep learning detection models, the proposed gated axis-concentrated localization network (GACL Net) has the characteristic of focusing on small objects.
Xiaoqiang Lu, Yuanlin Zhang 0003, Yuan Yuan 0001, Yachuang Feng
IEEE Trans. Geosci. Remote. Sens.1
2020 Nonlocal Graph Convolutional Networks for Hyperspectral Image Classification
abstract
Over the past few years making use of deep networks, including convolutional neural networks (CNNs) and recurrent neural networks (RNNs), classifying hyperspectral images has progressed significantly and gained increasing attention. In spite of being successful, these networks need an adequate supply of labeled training instances for supervised learning, which, however, is quite costly to collect. On the other hand, unlabeled data can be accessed in almost arbitrary amounts. Hence it would be conceptually of great interest to explore networks that are able to exploit labeled and unlabeled data simultaneously for hyperspectral image classification. In this article, we propose a novel graph-based semisupervised network called nonlocal graph convolutional network (nonlocal GCN). Unlike existing CNNs and RNNs that receive pixels or patches of a hyperspectral image as inputs, this network takes the whole image (including both labeled and unlabeled data) in. More specifically, a nonlocal graph is first calculated. Given this graph representation, a couple of graph convolutional layers are used to extract features. Finally, the semisupervised learning of the network is done by using a cross-entropy error over all labeled instances. Note that the nonlocal GCN is end-to-end trainable. We demonstrate in extensive experiments that compared with state-of-the-art spectral classifiers and spectral-spatial classification networks, the nonlocal GCN is able to offer competitive results and high-quality classification maps (with fine boundaries and without noisy scattered points of misclassification).
Lichao Mou, Xiaoqiang Lu, Xuelong Li 0001, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.2
2020 Remote Sensing Scene Classification by Gated Bidirectional Network
abstract
Remote sensing (RS) scene classification is a challenging task due to various land covers contained in RS scenes. Recent RS classification methods demonstrate that aggregating the multilayer convolutional features, which are extracted from different hierarchical layers of a convolutional neural network, can effectively improve classification accuracy. However, these methods treat the multilayer convolutional features as equally important and ignore the hierarchical structure of multilayer convolutional features. Multilayer convolutional features not only provide complementary information for classification but also bring some interference information (e.g., redundancy and mutual exclusion). In this paper, a gated bidirectional network is proposed to integrate the hierarchical feature aggregation and the interference information elimination into an end-to-end network. First, the performance of each convolutional feature is quantitatively analyzed and a superior combination of convolutional features is selected. Then, a bidirectional connection is proposed to hierarchically aggregate multilayer convolutional features. Both the top–down direction and the bottom–up direction are considered to aggregate multilayer convolutional features into the semantic-assist feature and appearance-assist feature, respectively, and a gated function is utilized to eliminate interference information in the bidirectional connection. Finally, the semantic-assist feature and appearance-assist feature are merged for classification. The proposed method can compete with the state-of-the-art methods on four RS scene classification data sets (AID, UC-Merced, WHU-RS19, and OPTIMAL-31).
Hao Sun 0014, Xiangtao Zheng, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.4
2020 Spectral-Spatial Attention Network for Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) classification aims to assign each hyperspectral pixel with a proper land-cover label. Recently, convolutional neural networks (CNNs) have shown superior performance. To identify the land-cover label, CNN-based methods exploit the adjacent pixels as an input HSI cube, which simultaneously contains spectral signatures and spatial information. However, at the edge of each land-cover area, an HSI cube often contains several pixels whose land-cover labels are different from that of the center pixel. These pixels, named interfering pixels, will weaken the discrimination of spectral-spatial features and reduce classification accuracy. In this article, a spectral-spatial attention network (SSAN) is proposed to capture discriminative spectral-spatial features from attention areas of HSI cubes. First, a simple spectral-spatial network (SSN) is built to extract spectral-spatial features from HSI cubes. The SSN is composed of a spectral module and a spatial module. Each module consists of only a few 3-D convolution and activation operations, which make the proposed method easy to converge with a small number of training samples. Second, an attention module is introduced to suppress the effects of interfering pixels. The attention module is embedded into the SSN to obtain the SSAN. The experiments on several public HSI databases demonstrate that the proposed SSAN outperforms several state-of-the-art methods.
Hao Sun 0014, Xiangtao Zheng, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.3
2020 Attribute-Cooperated Convolutional Neural Network for Remote Sensing Image Classification
abstract
Remote sensing image (RSI) classification is one of the most important fields in RSI processing. It is well known that RSIs are very complicated due to its various kinds of contents. Therefore, it is very difficult to distinguish different scene categories with similar visual contents, like desert and bare land. To address hard negative categories, an attribute-cooperated convolutional neural network (ACCNN) is proposed to exploit attributes as additional guiding information. First, the classification branch extracts convolutional neural network feature, which is then utilized to recognize the RSI scene categories. Second, the attribute branch is proposed to make the network distinguish scene categories efficiently. The proposed attribute branch shares feature extraction layers with the classification branch and makes the classification branch aware of extra attribute information. Finally, the relationship branch constraints the relationship between the classification branch and the attribute branch. To exploit the attribute information, three attribute-classification data sets are generated (AC-AID, AC-UCM, and AC-Sydney). Experimental results show that the proposed method is competitive to state-of-the-art methods. The data sets are available at https://github.com/CrazyStoneonRoad/Attribute-Cooperated-Classification-Data sets.
Yuanlin Zhang 0003, Xiangtao Zheng, Yuan Yuan 0001, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.4
2020 A Joint Relationship Aware Neural Network for Single-Image 3D Human Pose Estimation
abstract
This paper studies the task of 3D human pose estimation from a single RGB image, which is challenging without depth information. Recently many deep learning methods are proposed and achieve great improvements due to their strong representation learning. However, most existing methods ignore the relationship between joint features. In this paper, a joint relationship aware neural network is proposed to take both global and local joint relationship into consideration. First, a whole feature block representing all human body joints is extracted by a convolutional neural network. A Dual Attention Module (DAM) is applied on the whole feature block to generate attention weights. By exploiting the attention module, the global relationship between the whole joints is encoded. Second, the weighted whole feature block is divided into some individual joint features. To capture salient joint feature, the individual joint features are refined by individual DAMs. Finally, a joint angle prediction constraint is proposed to consider local joint relationship. Quantitative and qualitative experiments on 3D human pose estimation benchmarks demonstrate the effectiveness of the proposed method.
Xiangtao Zheng, Xiumei Chen, Xiaoqiang Lu
IEEE Trans. Image Process.3
2020 Discrete Deep Hashing With Ranking Optimization for Image Retrieval
abstract
For large-scale image retrieval task, a hashing technique has attracted extensive attention due to its efficient computing and applying. By using the hashing technique in image retrieval, it is crucial to generate discrete hash codes and preserve the neighborhood ranking information simultaneously. However, both related steps are treated independently in most of the existing deep hashing methods, which lead to the loss of key category-level information in the discretization process and the decrease in discriminative ranking relationship. In order to generate discrete hash codes with notable discriminative information, we integrate the discretization process and the ranking process into one architecture. Motivated by this idea, a novel ranking optimization discrete hashing (RODH) method is proposed, which directly generates discrete hash codes (e.g., +1/-1) from raw images by balancing the effective category-level information of discretization and the discrimination of ranking information. The proposed method integrates convolutional neural network, discrete hash function learning, and ranking function optimizing into a unified framework. Meanwhile, a novel loss function based on label information and mean average precision (MAP) is proposed to preserve the label consistency and optimize the ranking information of hash codes simultaneously. Experimental results on four benchmark data sets demonstrate that RODH can achieve superior performance over the state-of-the-art hashing methods.
Xiaoqiang Lu, Yaxiong Chen, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.1
2020 Siamese Dilated Inception Hashing With Intra-Group Correlation Enhancement for Image Retrieval
abstract
For large-scale image retrieval, hashing has been extensively explored in approximate nearest neighbor search methods due to its low storage and high computational efficiency. With the development of deep learning, deep hashing methods have made great progress in image retrieval. Most existing deep hashing methods cannot fully consider the intra-group correlation of hash codes, which leads to the correlation decrease problem of similar hash codes and ultimately affects the retrieval results. In this article, we propose an end-to-end siamese dilated inception hashing (SDIH) method that takes full advantage of multi-scale contextual information and category-level semantics to enhance the intra-group correlation of hash codes for hash codes learning. First, a novel siamese inception dilated network architecture is presented to generate hash codes with the intra-group correlation enhancement by exploiting multi-scale contextual information and category-level semantics simultaneously. Second, we propose a new regularized term, which can force the continuous values to approximate discrete values in hash codes learning and eventually reduces the discrepancy between the Hamming distance and the Euclidean distance. Finally, experimental results in five public data sets demonstrate that SDIH can outperform other state-of-the-art hashing algorithms.
Xiaoqiang Lu, Yaxiong Chen, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.1
2020 Property-Constrained Dual Learning for Video Summarization
abstract
Video summarization is the technique to condense large-scale videos into summaries composed of key-frames or key-shots so that the viewers can browse the video content efficiently. Recently, supervised approaches have achieved great success by taking advantages of recurrent neural networks (RNNs). Most of them focus on generating summaries by maximizing the overlap between the generated summary and the ground truth. However, they neglect the most critical principle, i.e., whether the viewer can infer the original video content from the summary. As a result, existing approaches cannot preserve the summary quality well and usually demand large amounts of training data to reduce overfitting. In our view, video summarization has two tasks, i.e., generating summaries from videos and inferring the original content from summaries. Motivated by this, we propose a dual learning framework by integrating the summary generation (primal task) and video reconstruction (dual task) together, which targets to reward the summary generator under the assistance of the video reconstructor. Moreover, to provide more guidance to the summary generator, two property models are developed to measure the representativeness and diversity of the generated summary. Practically, experiments on four popular data sets (SumMe, TVsum, OVP, and YouTube) have demonstrated that our approach, with compact RNNs as the summary generator, using less training data, and even in the unsupervised setting, can get comparable performance with those supervised ones adopting more complex summary generators and trained on more annotated data.
Bin Zhao 0001, Xuelong Li 0001, Xiaoqiang Lu
IEEE Trans. Neural Networks Learn. Syst.3
2020 Unsupervised Learning of Human Action Categories in Still Images with Deep Representations
abstract
In this article, we propose a novel method for unsupervised learning of human action categories in still images. In contrast to previous methods, the proposed method explores distinctive information of actions directly from unlabeled image databases, attempting to learn discriminative deep representations in an unsupervised manner to distinguish different actions. In the proposed method, action image collections can be used without manual annotations. Specifically, (i) to deal with the problem that unsupervised discriminative deep representations are difficult to learn, the proposed method builds a training dataset with surrogate labels from the unlabeled dataset, then learns discriminative representations by alternately updating convolutional neural network (CNN) parameters and the surrogate training dataset in an iterative manner; (ii) to explore the discriminatory information among different action categories, training batches for updating the CNN parameters are built with triplet groups and the triplet loss function is introduced to update the CNN parameters; and (iii) to learn more discriminative deep representations, a Random Forest classifier is adopted to update the surrogate training dataset, and more beneficial triplet groups then can be built with the updated surrogate training dataset. Extensive experiments on four benchmark datasets demonstrate the effectiveness of the proposed method.
Yunpeng Zheng, Xuelong Li 0001, Xiaoqiang Lu
ACM Trans. Multim. Comput. Commun. Appl.3
2019 Deep Voice-Visual Cross-Modal Retrieval with Deep Feature Similarity Learning
Yaxiong Chen, Xiaoqiang Lu, Yachuang Feng
PRCV (3)2
2019 Muti-stage learning for gender and age prediction
Jie Fang 0001, Yuan Yuan 0001, Xiaoqiang Lu, Yachuang Feng
Neurocomputing3
2019 Bidirectional adaptive feature fusion for remote sensing scene classification
Xiaoqiang Lu, Weijun Ji, Xuelong Li 0001, Xiangtao Zheng
Neurocomputing1
2019 3G structure for image caption generation
Aihong Yuan, Xuelong Li 0001, Xiaoqiang Lu
Neurocomputing3
2019 Detection of ships in inland river using high-resolution optical satellite imagery based on mixture of deformable part models
Lei Qi 0004, Xueming Qian, Xiaoqiang Lu
J. Parallel Distributed Comput.4
2019 Semantic Descriptions of High-Resolution Remote Sensing Images
abstract
Image captioning has attracted more and more attention in remote sensing filed since it provides more specific information than traditional tasks, such as classification. Though image captioning has gained some developments in recent years, it is difficult to describe the image in one simple sentence. To relieve the limitation, a novel captioning task is proposed and a novel framework is proposed to solve the novel task. The proposed framework uses semantic embedding to measure the image representation and the sentence representation. The captioning performance is improved by a proposed sentence representation (collective representation). Experimental results and human evaluations on three captioning data sets in remote sensing field demonstrate that the proposed framework can lead to advancement in image captioning results.
Binqiang Wang, Xiaoqiang Lu, Xiangtao Zheng, Xuelong Li 0001
IEEE Geosci. Remote. Sens. Lett.2
2019 Exploiting spatial relation for fine-grained image classification
Lei Qi 0004, Xiaoqiang Lu, Xuelong Li 0001
Pattern Recognit.2
2019 Weather recognition via classification labels and weather-cue maps
Bin Zhao 0001, Lulu Hua, Xuelong Li 0001, Xiaoqiang Lu, Zhigang Wang 0002
Pattern Recognit.4
2019 Robust Space-Frequency Joint Representation for Remote Sensing Image Scene Classification
abstract
Remote sensing image scene classification is a fundamental problem, which aims to label an image with a specific semantic category automatically. Recent progress on remote sensing image scene classification is substantial, benefitting mostly from the powerful feature extraction capability of convolutional neural networks (CNNs). Even though these CNN-based methods have achieved competitive performances, they only construct the representation of the image in location-sensitive space-domain. As a result, their representations are not robust to rotation-variant remote sensing images, which influence the classification accuracy. In this paper, we propose a novel feature representation method by introducing a frequency-domain branch to the traditional only-space-domain architecture. Our framework takes full advantages of discriminative features from space domain and location-robust features from the frequency domain, providing more advanced representations through an additional joint learning module, a property that is critically needed to perform remote sensing image scene classification. Additionally, our method produces satisfactory performances on four public and challenging remote sensing image scene data sets, Sydney, UC-Merced, WHU-RS19, and AID.
Jie Fang 0001, Yuan Yuan 0001, Xiaoqiang Lu, Yachuang Feng
IEEE Trans. Geosci. Remote. Sens.3
2019 A Feature Aggregation Convolutional Neural Network for Remote Sensing Scene Classification
abstract
Remote sensing scene classification (RSSC) refers to inferring semantic labels based on the content of the remote sensing scenes. Recently, most works take the pretrained convolutional neural network (CNN) as the feature extractor to build a scene representation for RSSC. The activations in different layers of CNN (named intermediate features) contain different spatial and semantic information. Recent works demonstrate that aggregating intermediate features into a scene representation can significantly improve the classification accuracy for RSSC. However, the intermediate features are aggregated by some unsupervised feature encoding methods (e.g., Bag-of-Visual-Words). Little attention has been paid to explore the information of semantic labels for the feature aggregation. In this paper, in order to explore the semantic label information, an end-to-end feature aggregation CNN (FACNN) is proposed to learn a scene representation for RSSC. In FACNN, a supervised convolutional features' encoding module and a progressive aggregation strategy are proposed to leverage the semantic label information to aggregate the intermediate features. The FACNN integrates the feature learning, feature aggregation, and classifier into a unified end-to-end framework for joint training. In FACNN, the scene representation is learned by considering the information of semantic labels, which can result in better performance for RSSC. Extensive experiments on AID, UC-Merged, and WHU-RS19 databases demonstrate that FACNN performs better than several state-of-the-art methods.
Xiaoqiang Lu, Hao Sun 0014, Xiangtao Zheng
IEEE Trans. Geosci. Remote. Sens.1
2019 Remote Sensing Image Scene Classification Using Rearranged Local Features
abstract
Remote sensing image scene classification is a fundamental problem, which aims to label an image with a specific semantic category automatically. Recently, deep learning methods have achieved competitive performance for remote sensing image scene classification, especially the methods based on a convolutional neural network (CNN). However, most of the existing CNN methods only use feature vectors of the last fully connected layer. They give more importance to global information and ignore local information of images. It is common that some images belong to different categories, although they own similar global features. The reason is that the category of an image may be highly related to local features, other than the global feature. To address this problem, a method based on rearranged local features is proposed in this paper. First, outputs of the last convolutional layer and the last fully connected layer are employed to depict the local and global information, respectively. After that, the remote sensing images are clustered to several collections using their global features. For each collection, local features of an image are rearranged according to their similarities with local features of the cluster center. In addition, a fusion strategy is proposed to combine global and local features for enhancing the image representation. The proposed method surpasses the state of the arts on four public and challenging data sets: UC-Merced, WHU-RS19, Sydney, and AID.
Yuan Yuan 0001, Jie Fang 0001, Xiaoqiang Lu, Yachuang Feng
IEEE Trans. Geosci. Remote. Sens.3
2019 Similarity Constrained Convex Nonnegative Matrix Factorization for Hyperspectral Anomaly Detection
abstract
Hyperspectral anomaly detection is very important in the remote sensing domain. The representation-based anomaly method is one of the most important hyperspectral anomaly detection methods, which uses reconstruction errors (REs) to detect anomalies. REs are affected by the basis matrix and its corresponding coefficient matrix. Mixed pixels exist because of the low-spatial resolution of hyperspectral images. The RE is not large enough to correctly distinguish the pixel difficult to classify when the basis matrix is composed of pixels. Moreover, its corresponding coefficients cannot indicate whether pixels are pure or mixed and the abundances of mixed pixels. To address the above-mentioned problems, endmembers referring to pure or relatively pure spectral signatures are explored to build the basis matrix. The RE based on the basis matrix of endmembers is much larger for the anomalous pixel difficult to correctly classify. Furthermore, its corresponding coefficient matrix of endmembers has physical meanings. Hence, a novel hyperspectral anomaly detection based on similarity constrained convex nonnegative matrix factorization is proposed from the perspective of endmembers for the first time. First, convex nonnegative matrix factorization (CNMF) is employed to obtain endmembers of background. Then, CNMF is constrained by the similarity regularization that considers different contributions of endmembers to the pixel under test to acquire the more accurate and meaningful coefficient matrix. Finally, anomalies are detected by calculating REs. The proposed algorithm is verified on both simulated and real data sets. Experimental results show that our proposed algorithm outperforms other state-of-the-art algorithms.
Wuxia Zhang, Xiaoqiang Lu, Xuelong Li 0001
IEEE Trans. Geosci. Remote. Sens.2
2019 Hierarchical and Robust Convolutional Neural Network for Very High-Resolution Remote Sensing Object Detection
abstract
Object detection is a basic issue of very high-resolution remote sensing images (RSIs) for automatically labeling objects. At present, deep learning has gradually gained the competitive advantage for remote sensing object detection, especially based on convolutional neural networks (CNNs). Most of the existing methods use the global information in the fully connected feature vector and ignore the local information in the convolutional feature cubes. However, the local information can provide spatial information, which is helpful for accurate localization. In addition, there are variable factors, such as rotation and scaling, which affect the object detection accuracy in RSIs. In order to solve these problems, this paper presents a hierarchical robust CNN. First, multiscale convolutional features are extracted to represent the hierarchical spatial semantic information. Second, multiple fully connected layer features are stacked together so as to improve the rotation and scaling robustness. Experiments on two data sets have shown the effectiveness of our method. In addition, a large-scale high-resolution remote sensing object detection data set is established to make up for the current situation that the existing data set is insufficient or too small. The data set is available athttps://github.com/CrazyStoneonRoad/TGRS-HRRSD-Dataset.
Yuanlin Zhang 0003, Yuan Yuan 0001, Yachuang Feng, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.4
2019 Hyperspectral Image Denoising by Fusing the Selected Related Bands
abstract
Hyperspectral images (HSIs) convey more useful information than RGB or gray images, which are widely used in many remote sensing tasks. In real scenarios, HSIs are inevitably corrupted by noise because of sensors' imperfectness or atmospheric influence. Recently, many HSI denoising methods have been proposed to utilize the interband information between different spectral bands. However, these methods regard the HSI as a whole and treat the different spectral bands with the same noise level. In fact, the noise levels in different bands are different. Especially, only few certain bands are corrupted by noise, named the target noised bands. Under this circumstance, an HSI denoising method is proposed by considering the band relationship and different noise levels. The target noised bands are adaptively denoised by fusing some selected bands. Specifically, some related but quality superior bands are selected according to the target noised bands. Then, the target noised bands can be denoised by fusing the selected related bands. Experimental results show that the proposed method achieves considerable performances in comparison with several state-of-the-art hyperspectral denoising methods.
Xiangtao Zheng, Yuan Yuan 0001, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.3
2019 A Deep Scene Representation for Aerial Scene Classification
abstract
As a fundamental problem in earth observation, aerial scene classification tries to assign a specific semantic label to an aerial image. In recent years, the deep convolutional neural networks (CNNs) have shown advanced performances in aerial scene classification. The successful pretrained CNNs can be transferable to aerial images. However, global CNN activations may lack geometric invariance and, therefore, limit the improvement of aerial scene classification. To address this problem, this paper proposes a deep scene representation to achieve the invariance of CNN features and further enhance the discriminative power. The proposed method: 1) extracts CNN activations from the last convolutional layer of pretrained CNN; 2) performs multiscale pooling (MSP) on these activations; and 3) builds a holistic representation by the Fisher vector method. MSP is a simple and effective multiscale strategy, which enriches multiscale spatial information in affordable computational time. The proposed representation is particularly suited at aerial scenes and consistently outperforms global CNN activations without requiring feature adaptation. Extensive experiments on five aerial scene data sets indicate that the proposed method, even with a simple linear classifier, can achieve the state-of-the-art performance.
Xiangtao Zheng, Yuan Yuan 0001, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.3
2019 CAM-RNN: Co-Attention Model Based RNN for Video Captioning
abstract
Video captioning is a technique that bridges vision and language together, for which both visual information and text information are quite important. Typical approaches are based on the recurrent neural network (RNN), where the video caption is generated word by word, and the current word is predicted based on the visual content and previously generated words. However, in the prediction of the current word, there is much uncorrelated visual content, and some of the previously generated words provide little information, which may cause interference in generating a correct caption. Based on this point, we attempt to exploit the visual and text features that are most correlated with the caption. In this paper, a co-attention model based recurrent neural network (CAM-RNN) is proposed, where the CAM is utilized to encode the visual and text features, and the RNN works as the decoder to generate the video caption. Specifically, the CAM is composed of a visual attention module, a text attention module, and a balancing gate. During the generation procedure, the visual attention module is able to adaptively attend to the salient regions in each frame and the frames most correlated with the caption. The text attention module can automatically focus on the most relevant previously generated words or phrases. Moreover, between the two attention modules, a balancing gate is designed to regulate the influence of visual features and text features when generating the caption. In practice, the extensive experiments are conducted on four popular datasets, including MSVD, Charades, MSR-VTT, and MPII-MD, which have demonstrated the effectiveness of the proposed approach.
Bin Zhao 0001, Xuelong Li 0001, Xiaoqiang Lu
IEEE Trans. Image Process.3
2019 Face Sketch Synthesis by Multidomain Adversarial Learning
abstract
Given a training set of face photo-sketch pairs, face sketch synthesis targets at learning a mapping from the photo domain to the sketch domain. Despite the exciting progresses made in the literature, it retains as an open problem to synthesize high-quality sketches against blurs and deformations. Recent advances in generative adversarial training provide a new insight into face sketch synthesis, from which perspective the existing synthesis pipelines can be fundamentally revisited. In this paper, we present a novel face sketch synthesis method by multidomain adversarial learning (termed MDAL), which overcomes the defects of blurs and deformations toward high-quality synthesis. The principle of our scheme relies on the concept of "interpretation through synthesis." In particular, we first interpret face photographs in the photodomain and face sketches in the sketch domain by reconstructing themselves respectively via adversarial learning. We define the intermediate products in the reconstruction process as latent variables, which form a latent domain. Second, via adversarial learning, we make the distributions of latent variables being indistinguishable between the reconstruction process of the face photograph and that of the face sketch. Finally, given an input face photograph, the latent variable obtained by reconstructing this face photograph is applied for synthesizing the corresponding sketch. Quantitative comparisons to the state-of-the-art methods demonstrate the superiority of the proposed MDAL method.
Shengchuan Zhang, Rongrong Ji, Jie Hu 0018, Xiaoqiang Lu, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.4
2019 Spatial Structure Preserving Feature Pyramid Network for Semantic Image Segmentation
abstract
Recently, progress on semantic image segmentation is substantial, benefiting from the rapid development of Convolutional Neural Networks. Semantic image segmentation approaches proposed lately have been mostly based on Fully convolutional Networks (FCNs). However, these FCN-based methods use large receptive fields and too many pooling layers to depict the discriminative semantic information of the images. Specifically, on one hand, convolutional kernel with large receptive field smooth the detailed edges, since too much contexture information is used to depict the “center pixel.” However, the pooling layer increases the receptive field through zooming out the latest feature maps, which loses many detailed information of the image, especially in the deeper layers of the network. These operations often cause low spatial resolution inside deep layers, which leads to spatially fragmented prediction. To address this problem, we exploit the inherent multi-scale and pyramidal hierarchy of deep convolutional networks to extract the feature maps with different resolutions and take full advantages of these feature maps via a gradually stacked fusing way. Specifically, for two adjacent convolutional layers, we upsample the features from deeper layer with stride of 2 and then stack them on the features from shallower layer. Then, a convolutional layer with kernels of 1× 1 is followed to fuse these stacked features. The fused feature preserves the spatial structure information of the image; meanwhile, it owns strong discriminative capability for pixel classification. Additionally, to further preserve the spatial structure information and regional connectivity of the predicted category label map, we propose a novel loss term for the network. In detail, two graph model-based spatial affinity matrixes are proposed, which are used to depict the pixel-level relationships in the input image and predicted category label map respectively, and then their cosine distance is backward propagated to the network. The proposed architecture, called spatial structure preserving feature pyramid network, significantly improves the spatial resolution of the predicted category label map for semantic image segmentation. The proposed method achieves state-of-the-art results on three public and challenging datasets for semantic image segmentation.
Yuan Yuan 0001, Jie Fang 0001, Xiaoqiang Lu, Yachuang Feng
ACM Trans. Multim. Comput. Commun. Appl.3
2018 HSA-RNN: Hierarchical Structure-Adaptive RNN for Video Summarization
abstract
Although video summarization has achieved great success in recent years, few approaches have realized the influence of video structure on the summarization results. As we know, the video data follow a hierarchical structure, i.e., a video is composed of shots, and a shot is composed of several frames. Generally, shots provide the activity-level information for people to understand the video content. While few existing summarization approaches pay attention to the shot segmentation procedure. They generate shots by some trivial strategies, such as fixed length segmentation, which may destroy the underlying hierarchical structure of video data and further reduce the quality of generated summaries. To address this problem, we propose a structure-adaptive video summarization approach that integrates shot segmentation and video summarization into a Hierarchical Structure-Adaptive RNN, denoted as HSA-RNN. We evaluate the proposed approach on four popular datasets, i.e., SumMe, TVsum, CoSum and VTW. The experimental results have demonstrated the effectiveness of HSA-RNN in the video summarization task.
Bin Zhao 0001, Xuelong Li 0001, Xiaoqiang Lu
CVPR3
2018 Video Captioning with Tube Features
abstract
Visual feature plays an important role in the video captioning task. Considering that the video content is mainly composed of the activities of salient objects, it has restricted the caption quality of current approaches which just focus on global frame features while paying less attention to the salient objects. To tackle this problem, in this paper, we design an object-aware feature for video captioning, denoted as tube feature. Firstly, Faster-RCNN is employed to extract object regions in frames, and a tube generation method is developed to connect the regions from different frames but belonging to the same object. After that, an encoder-decoder architecture is constructed for video caption generation. Specifically, the encoder is a bi-directional LSTM, which is utilized to capture the dynamic information of each tube. The decoder is a single LSTM extended with an attention model, which enables our approach to adaptively attend to the most correlated tubes when generating the caption. We evaluate our approach on two benchmark datasets: MSVD and Charades. The experimental results have demonstrated the effectiveness of tube feature in the video captioning task.
Bin Zhao 0001, Xuelong Li 0001, Xiaoqiang Lu
IJCAI3
2018 Hyperspectral Band Selection with Convolutional Neural Network
Rui Cai 0002, Yuan Yuan 0001, Xiaoqiang Lu
PRCV (4)3
2018 A CNN-RNN architecture for multi-label weather recognition
Bin Zhao 0001, Xuelong Li 0001, Xiaoqiang Lu, Zhigang Wang 0002
Neurocomputing3
2018 Multi-modal gated recurrent units for image description
Xuelong Li 0001, Aihong Yuan, Xiaoqiang Lu
Multim. Tools Appl.3
2018 Hyperspectral image classification based on joint spectrum of spatial space and spectral space
Zhibin Pan, Xiaoqiang Lu, Bingliang Hu
Multim. Tools Appl.3
2018 Feedback mechanism based iterative metric learning for person re-identification
abstract
Person re-identification problem is targeting to match people in the views of non-overlapped camera networks. It is an important task in the fields of computer vision and video surveillance. It shows great value in applications like surveillance and action recognition. Existing metric learning based methods measure the similarity of sample pairs by learning a metric space in which the positive pairs are closer than negative pairs. However, the appearance features undergo with drastic variation. Person re-identification is a typical small sample problem. It is hard to learn a robust projection of metric subspace that takes all the situations into consideration. The learned metric subspace is usually over-fitting to training dataset due to the strict metric learning constraint . And the hard pairs in training dataset will weaken the discrimination of matching pairs’ similarity. To address these problems, a feedback mechanism based iterative metric learning method is proposed. The proposed method introduces a mean distance of multi-metric subspace to deal with the over-fitting problem. The joint discriminant optimal model on feedback top ranks matching pairs will enhance the discrimination of matching pairs’ similarity. It is a robust and discriminative distance metric which measures the matching pairs similarity with distances of multiple metric projections learned by a set of training datasets. Aiming to learn the multi-metric subspace, the proposed method gives a feedback mechanism based approach which back propagates the top ranks identification results as pseudo training datasets. The effectiveness of proposed mean distance of multi-metric projection is analyzed and proved theoretically. And extensive experiments on three challenging datasets, VIPeR, GRID and CUHK01 are conducted. The results show that the proposed method achieves the best performance and improves the state-of-the-art rank-1 identification rates by 18.48%, 2.00% and 5.41% on three datasets respectively.
Yutao Ren, Xuelong Li 0001, Xiaoqiang Lu
Pattern Recognit.3
2018 Structured dictionary learning for abnormal event detection in crowded scenes
Yuan Yuan 0001, Yachuang Feng, Xiaoqiang Lu
Pattern Recognit.3
2018 Key Frame Extraction in the Summary Space
abstract
Key frame extraction is an efficient way to create the video summary which helps users obtain a quick comprehension of the video content. Generally, the key frames should be representative of the video content, meanwhile, diverse to reduce the redundancy. Based on the assumption that the video data are near a subspace of a high-dimensional space, a new approach, named as key frame extraction in the summary space, is proposed for key frame extraction in this paper. The proposed approach aims to find the representative frames of the video and filter out similar frames from the representative frame set. First of all, the video data are mapped to a high-dimensional space, named as summary space. Then, a new representation is learned for each frame by analyzing the intrinsic structure of the summary space. Specifically, the learned representation can reflect the representativeness of the frame, and is utilized to select representative frames. Next, the perceptual hash algorithm is employed to measure the similarity of representative frames. As a result, the key frame set is obtained after filtering out similar frames from the representative frame set. Finally, the video summary is constructed by assigning the key frames in temporal order. Additionally, the ground truth, created by filtering out similar frames from human-created summaries, is utilized to evaluate the quality of the video summary. Compared with several traditional approaches, the experimental results on 80 videos from two datasets indicate the superior performance of our approach.
Xuelong Li 0001, Bin Zhao 0001, Xiaoqiang Lu
IEEE Trans. Cybern.3
2018 Exploring Models and Data for Remote Sensing Image Caption Generation
abstract
Inspired by recent development of artificial satellite, remote sensing images have attracted extensive attention. Recently, notable progress has been made in scene classification and target detection. However, it is still not clear how to describe the remote sensing image content with accurate and concise sentences. In this paper, we investigate to describe the remote sensing images with accurate and flexible sentences. First, some annotated instructions are presented to better describe the remote sensing images considering the special characteristics of remote sensing images. Second, in order to exhaustively exploit the contents of remote sensing images, a large-scale aerial image data set is constructed for remote sensing image caption. Finally, a comprehensive review is presented on the proposed data set to fully advance the task of remote sensing caption. Extensive experiments on the proposed data set demonstrate that the content of the remote sensing image can be completely described by generating language descriptions. The data set is available at https://github.com/201528014227051/RSICD_optimal.
Xiaoqiang Lu, Binqiang Wang, Xiangtao Zheng, Xuelong Li 0001
IEEE Trans. Geosci. Remote. Sens.1
2018 A Hybrid Sparsity and Distance-Based Discrimination Detector for Hyperspectral Images
abstract
Hyperspectral target detection is an approach which tries to locate targets in a hyperspectral image on the condition of given targets spectrum. Many classical target detectors are based on the linear mixing model (LMM) and sparsity model. The LMM has a poor performance in dealing with the spectral variability. Therefore, more studies focus on the sparsity-based detectors, most of which are based on residual reconstruction. Owing to the fact that the impure dictionary for the test pixel weakens the detection performance and the discrimination ability of residual function has direct influence on the detecting accuracy, the dictionary purity and discriminative residual function are two most important factors affecting the accuracy of sparsity-based target detectors. In order to obtain more purified dictionary and discriminative residual function, this paper proposes a novel sparsity-based detector named the hybrid sparsity and distance-based discrimination (HSDD) detector for target detection in hyperspectral imagery. The residual function is constrained by the discrimination information during the dictionary construction, which enhances the dictionary purification. Only background samples are used to construct the dictionary because it is easier to remove the target pixel than to select it on the condition that majority of pixels are the background pixels. Hence, a purification process is applied for background training samples in order to construct an effective competition between the residual term and discriminative term. Extensive experimental results with four hyperspectral data sets demonstrate that the proposed HSDD algorithm has a better performance than the state-of-the-art algorithms.
Xiaoqiang Lu, Wuxia Zhang, Xuelong Li 0001
IEEE Trans. Geosci. Remote. Sens.1
2018 A Coarse-to-Fine Semi-Supervised Change Detection for Multispectral Images
abstract
Change detection is an important technique providing insights to urban planning, resources monitoring, and environmental studies. For multispectral images, most semi-supervised change detection methods focus on improving the contribution of training samples hard to be classified to the trained classifier. However, hard training samples will weaken the discrimination of the training model for multispectral change detection. Besides, these methods only use the spectral information, while the limited spectral information cannot represent objects very well. In this paper, a method named as coarse-to-fine semi-supervised change detection is proposed to solve the aforementioned problems. First, a novel multiscale feature is exploited by concatenating the spectral vector of the pixel to be detected and its adjacent pixels by different scales. Second, the enhanced metric learning is proposed to acquire more discriminant metric by strengthening the contribution of training samples easy to be classified and weakening the contribution of training samples hard to be classified to the trained model. Finally, a coarse-to-fine strategy is adopted to detect testing samples from the viewpoint of distance metric and label information of neighborhood in spatial space. The coarse detection result obtained from the enhanced metric learning is used to guide the final detection. The effectiveness of our proposed method is verified on two real-life operating scenarios, Taizhou and Kunshan data sets. Extensive experimental results demonstrate that our proposed algorithm has better performance than those of other state-of-the-art algorithms.
Wuxia Zhang, Xiaoqiang Lu, Xuelong Li 0001
IEEE Trans. Geosci. Remote. Sens.2
2018 Video Synopsis in Complex Situations
abstract
Video synopsis is an effective technique for surveillance video browsing and storage. However, most of the existing video synopsis approaches are not suitable for complex situations, especially crowded scenes. This is because these approaches heavily depend on the preprocessing results of foreground segmentation and multiple objects tracking, but the preprocessing techniques usually achieve poor performance in crowded scenes. To address this problem, we propose a comprehensive video synopsis approach which can be applied to scenes with drastically varying crowdedness. The proposed approach differs significantly from the existing methods, and has several appealing properties. First, we propose to detect the crowdedness of a given video, then, extract object tubes in sparse periods and extract video clips in crowded periods, respectively. Through such a solution, the poor performance of preprocessing techniques in crowded scenes can be avoided by extracting the whole video frames. Second, we propose a group-partition algorithm which can discovers the relationships among moving objects and alleviates several segmentation and tracking errors. Third, a group-based greedy optimization algorithm is proposed to automatically determine the length of a synopsis video. Besides, we present extensive experiments that demonstrate the effectiveness and efficiency of the proposed approach.
Xuelong Li 0001, Zhigang Wang 0002, Xiaoqiang Lu
IEEE Trans. Image Process.3
2018 Hierarchical Recurrent Neural Hashing for Image Retrieval With Hierarchical Convolutional Features
abstract
Hashing has been an important and effective technology in image retrieval due to its computational efficiency and fast search speed. The traditional hashing methods usually learn hash functions to obtain binary codes by exploiting hand-crafted features, which cannot optimally represent the information of the sample. Recently, deep learning methods can achieve better performance, since deep learning architectures can learn more effective image representation features. However, these methods only use semantic features to generate hash codes by shallow projection but ignore texture details. In this paper, we proposed a novel hashing method, namely hierarchical recurrent neural hashing (HRNH), to exploit hierarchical recurrent neural network to generate effective hash codes. There are three contributions of this paper. First, a deep hashing method is proposed to extensively exploit both spatial details and semantic information, in which, we leverage hierarchical convolutional features to construct image pyramid representation. Second, our proposed deep network can exploit directly convolutional feature maps as input to preserve the spatial structure of convolutional feature maps. Finally, we propose a new loss function that considers the quantization error of binarizing the continuous embeddings into the discrete binary codes, and simultaneously maintains the semantic similarity and balanceable property of hash codes. Experimental results on four widely used data sets demonstrate that the proposed HRNH can achieve superior performance over other state-of-the-art hashing methods.Hashing has been an important and effective technology in image retrieval due to its computational efficiency and fast search speed. The traditional hashing methods usually learn hash functions to obtain binary codes by exploiting hand-crafted features, which cannot optimally represent the information of the sample. Recently, deep learning methods can achieve better performance, since deep learning architectures can learn more effective image representation features. However, these methods only use semantic features to generate hash codes by shallow projection but ignore texture details. In this paper, we proposed a novel hashing method, namely hierarchical recurrent neural hashing (HRNH), to exploit hierarchical recurrent neural network to generate effective hash codes. There are three contributions of this paper. First, a deep hashing method is proposed to extensively exploit both spatial details and semantic information, in which, we leverage hierarchical convolutional features to construct image pyramid representation. Second, our proposed deep network can exploit directly convolutional feature maps as input to preserve the spatial structure of convolutional feature maps. Finally, we propose a new loss function that considers the quantization error of binarizing the continuous embeddings into the discrete binary codes, and simultaneously maintains the semantic similarity and balanceable property of hash codes. Experimental results on four widely used data sets demonstrate that the proposed HRNH can achieve superior performance over other state-of-the-art hashing methods.
Xiaoqiang Lu, Yaxiong Chen, Xuelong Li 0001
IEEE Trans. Image Process.1
2018 Person Reidentification Based on Elastic Projections
abstract
Person reidentification usually refers to matching people in different camera views in nonoverlapping multicamera networks. Many existing methods learn a similarity measure by projecting the raw feature to a latent subspace to make the same target's distance smaller than different targets' distances. However, the same targets captured in different camera views should hold the same intrinsic attributes while different targets should hold different intrinsic attributes. Projecting all the data to the same subspace would cause loss of such an information and comparably poor discriminability. To address this problem, in this paper, a method based on elastic projections is proposed to learn a pairwise similarity measure for person reidentification. The proposed model learns two projections, positive projection and negative projection, which are both representative and discriminative. The representability refers to: for the same targets captured in two camera views, the positive projection can bridge the corresponding appearance variation and represent the intrinsic attributes of the same targets, while for the different targets captured in two camera views, the negative projection can explore and utilize the different attributes of different targets. The discriminability means that the intraclass distance should become smaller than its original distance after projection, while the interclass distance becomes larger on the contrary, which is the elastic property of the proposed model. In this case, prior information of the original data space is used to give guidance for the learning phase; more importantly, similar targets (but not the same) are effectively reduced by forcing the same targets to become more similar and different targets to become more distinct. The proposed model is evaluated on three benchmark data sets, including VIPeR, GRID, and CUHK, and achieves better performance than other methods.
Xuelong Li 0001, Lina Liu 0006, Xiaoqiang Lu
IEEE Trans. Neural Networks Learn. Syst.3
2017 Image2song: Song Retrieval via Bridging Image Content and Lyric Words
abstract
Image is usually taken for expressing some kinds of emotions or purposes, such as love, celebrating Christmas. There is another better way that combines the image and relevant song to amplify the expression, which has drawn much attention in the social network recently. Hence, the automatic selection of songs should be expected. In this paper, we propose to retrieve semantic relevant songs just by an image query, which is named as the image2song problem. Motivated by the requirements of establishing correlation in semantic/content, we build a semantic-based song retrieval framework, which learns the correlation between image content and lyric words. This model uses a convolutional neural network to generate rich tags from image regions, a recurrent neural network to model lyric, and then establishes correlation via a multi-layer perceptron. To reduce the content gap between image and lyric, we propose to make the lyric modeling focus on the main image content via a tag attention. We collect a dataset from the social-sharing multimodal data to study the proposed problem, which consists of (image, music clip, lyric) triplets. We demonstrate that our proposed model shows noticeable results in the image2song retrieval task and provides suitable songs. Besides, the song2image task is also performed.
Xuelong Li 0001, Di Hu 0001, Xiaoqiang Lu
ICCV3
2017 MAM-RNN: Multi-level Attention Model Based RNN for Video Captioning
abstract
Visual information is quite important for the task of video captioning. However, in the video, there are a lot of uncorrelated content, which may cause interference to generate a correct caption. Based on this point, we attempt to exploit the visual features which are most correlated to the caption. In this paper, a Multi-level Attention Model based Recurrent Neural Network (MAM-RNN) is proposed, where MAM is utilized to encode the visual feature and RNN works as the decoder to generate the video caption. During generation, the proposed approach is able to adaptively attend to the salient regions in the frame and the frames correlated to the caption. Practically, the experimental results on two benchmark datasets, i.e., MSVD and Charades, have shown the excellent performance of the proposed approach.
Xuelong Li 0001, Bin Zhao 0001, Xiaoqiang Lu
IJCAI3
2017 JM-Net and Cluster-SVM for Aerial Scene Classification
abstract
Aerial scene classification, which is a fundamental problem for remote sensing imagery, can automatically label an aerial image with a specific semantic category. Although deep learning has achieved competitive performance for aerial scene classification, training the conventional neural networks with aerial datasets will easily stick in overtting and local minimum. Because the aerial datasets only contain a few hundreds or thousands images, meanwhile the conventional networks usually contain millions of parameters to be trained. To address the problem, a novel convolutional neural network named JM-Net is proposed in this paper, which has different size of convolution kernels in same layer and ignores the fully convolytion layer, so it has fewer parameters and can be trained well on aerial datasets. Additionally, Cluster-SVM, a strategy to improve the accuracy and speed up the classification is used in the specific task. Finally, our method suparssed the state-of-art result on the challenging AID dataset while cost shorter time and used smaller storage space.
Xiaoqiang Lu, Yuan Yuan 0001, Jie Fang 0001
IJCAI1
2017 A Multi-Task Framework for Weather Recognition
abstract
Weather recognition is important in practice, while this task has not been thoroughly explored so far. The current trend of dealing with this task is treating it as a single classification problem, i.e., determining whether a given image belongs to a certain weather category or not. However, weather recognition differs significantly from traditional image classification, since several weather features may appear simultaneously. In this case, a simple classification result is insufficient to describe the weather condition. To address this issue, we propose to provide auxiliary weather related information for comprehensive weather description. Specifically, semantic segmentation of weather-cues, such as blue sky and white clouds, is exploited as an auxiliary task in this paper. Moreover, a convolutional neural network (CNN) based multi-task framework is developed which aims to concurrently tackle weather category classification task and weather-cues segmentation task. Due to the intrinsic relationships between these two tasks, exploring auxiliary semantic segmentation of weather-cues can also help to learn discriminative features for the classification task, and thus obtain superior accuracy. To verify the effectiveness of the proposed approach, extra segmentation masks of weather-cues are generated manually on an existing weather image dataset. Experimental results have demonstrated the superior performance of our approach. The enhanced dataset, source codes and pre-trained models are available at https://github.com/wzgwzg/Multitask_Weather.
Xuelong Li 0001, Zhigang Wang 0002, Xiaoqiang Lu
ACM Multimedia3
2017 Hierarchical Recurrent Neural Network for Video Summarization
abstract
Exploiting the temporal dependency among video frames or subshots is very important for the task of video summarization. Practically, RNN is good at temporal dependency modeling, and has achieved overwhelming performance in many video-based tasks, such as video captioning and classification. However, RNN is not capable enough to handle the video summarization task, since traditional RNNs, including LSTM, can only deal with short videos, while the videos in the summarization task are usually in longer duration. To address this problem, we propose a hierarchical recurrent neural network for video summarization, called H-RNN in this paper. Specifically, it has two layers, where the first layer is utilized to encode short video subshots cut from the original video, and the final hidden state of each subshot is input to the second layer for calculating its confidence to be a key subshot. Compared to traditional RNNs, H-RNN is more suitable to video summarization, since it can exploit long temporal dependency among frames, meanwhile, the computation operations are significantly lessened. The results on two popular datasets, including the Combined dataset and VTW dataset, have demonstrated that the proposed H-RNN outperforms the state-of-the-arts.
Bin Zhao 0001, Xuelong Li 0001, Xiaoqiang Lu
ACM Multimedia3
2017 Learning deep event models for crowd anomaly detection
Yachuang Feng, Yuan Yuan 0001, Xiaoqiang Lu
Neurocomputing3
2017 Penalized Linear Discriminant Analysis of Hyperspectral Imagery for Noise Removal
abstract
The existence of noise in hyperspectral imagery (HSI) seriously affects image quality. Noise removal is one of the most important and challenging tasks to complete before hyperspectral information extraction. Though many advances have been made in alleviating the effect of noise, problems, including a high correlation among bands and predefined structure of noise covariance, still prevent us from the effective implementation of hyperspectral denoising. In this letter, a new algorithm named the penalized linear discriminant analysis (PLDA) and noise adjusted principal components transformation (NAPCT) was proposed. PLDA was applied to search for the best noise covariance structure, while the NAPCT was employed to remove the noise. The results of the tests with both HJ-1A HSI and EO-1 Hyperion showed that the proposed PLDA-NAPCT method could remove the noise effectively and that it could preserve the spectral fidelity of the restored hyperspectral images. Specifically, the recovered spectral curves using the proposed method are visually more similar to the original image compared with the control methods; quantitative matrices, including the noise reduction ration and mean relative deviation, also showed that the PLDA-NAPCT produced less bias than the control methods. Furthermore, the PLDA-NAPCT method is sensor-independent, and it could be easily adapted for removing the noise from different sensors.
Luojia Hu, Tianxiang Yue, Bin Chen 0005, Xiaoqiang Lu, Bing Xu 0001
IEEE Geosci. Remote. Sens. Lett.6
2017 Remote Sensing Image Scene Classification: Benchmark and State of the Art
abstract
Remote sensing image scene classification plays an important role in a wide range of applications and hence has been receiving remarkable attention. During the past years, significant efforts have been made to develop various data sets or present a variety of approaches for scene classification from remote sensing images. However, a systematic review of the literature concerning data sets and methods for scene classification is still lacking. In addition, almost all existing data sets have a number of limitations, including the small scale of scene classes and the image numbers, the lack of image variations and diversity, and the saturation of accuracy. These limitations severely limit the development of new approaches especially deep learning-based methods. This paper first provides a comprehensive review of the recent progress. Then, we propose a large-scale data set, termed “NWPU-RESISC45,” which is a publicly available benchmark for REmote Sensing Image Scene Classification (RESISC), created by Northwestern Polytechnical University (NWPU). This data set contains 31 500 images, covering 45 scene classes with 700 images in each class. The proposed NWPU-RESISC45 1) is large-scale on the scene classes and the total image number; 2) holds big variations in translation, spatial resolution, viewpoint, object pose, illumination, background, and occlusion; and 3) has high within-class diversity and between-class similarity. The creation of this data set will enable the community to develop and evaluate various data-driven algorithms. Finally, several representative methods are evaluated using the proposed data set, and the results are reported as a useful baseline for future research.
Gong Cheng 0003, Junwei Han 0001, Xiaoqiang Lu
Proc. IEEE3
2017 On Combining Social Media and Spatial Technology for POI Cognition and Image Localization
abstract
With fast development of information engineering and social network, people's locations can be conveniently sensed by spatial technology, such as global positioning systems (GPS), base stations, Wi-Fi access points and even from the appearances of the photos they have taken. The social networks and the online shopping platforms have been gathering billions of users, who share a large amount of images taken in places they live in and visit. We can leverage the social networks to express our opinions about the services and places of interest (POIs). The interactions among users, and user and POIs or services generate big social media data, which have rich information for user, location, and service cognition. Many real-time network applications rely heavily on the accurate social users' locations. How to sense the locations from multisource social media data is very important and challenging. Thus, in this paper, we give a systematic review of the works that combine social media and spatial technology for POI cognition and image localization.
Xueming Qian, Xiaoqiang Lu, Junwei Han 0001, Bo Du 0001, Xuelong Li 0001
Proc. IEEE2
2017 Joint Dictionary Learning for Multispectral Change Detection
abstract
Change detection is one of the most important applications of remote sensing technology. It is a challenging task due to the obvious variations in the radiometric value of spectral signature and the limited capability of utilizing spectral information. In this paper, an improved sparse coding method for change detection is proposed. The intuition of the proposed method is that unchanged pixels in different images can be well reconstructed by the joint dictionary, which corresponds to knowledge of unchanged pixels, while changed pixels cannot. First, a query image pair is projected onto the joint dictionary to constitute the knowledge of unchanged pixels. Then reconstruction error is obtained to discriminate between the changed and unchanged pixels in the different images. To select the proper thresholds for determining changed regions, an automatic threshold selection strategy is presented by minimizing the reconstruction errors of the changed pixels. Adequate experiments on multispectral data have been tested, and the experimental results compared with the state-of-the-art methods prove the superiority of the proposed method. Contributions of the proposed method can be summarized as follows: 1) joint dictionary learning is proposed to explore the intrinsic information of different images for change detection. In this case, change detection can be transformed as a sparse representation problem. To the authors' knowledge, few publications utilize joint learning dictionary in change detection; 2) an automatic threshold selection strategy is presented, which minimizes the reconstruction errors of the changed pixels without the prior assumption of the spectral signature. As a result, the threshold value provided by the proposed method can adapt to different data due to the characteristic of joint dictionary learning; and 3) the proposed method makes no prior assumption of the modeling and the handling of the spectral signature, which can be adapted to different data.
Xiaoqiang Lu, Yuan Yuan 0001, Xiangtao Zheng
IEEE Trans. Cybern.1
2017 Statistical Hypothesis Detector for Abnormal Event Detection in Crowded Scenes
abstract
Abnormal event detection is now a challenging task, especially for crowded scenes. Many existing methods learn a normal event model in the training phase, and events which cannot be well represented are treated as abnormalities. However, they fail to make use of abnormal event patterns, which are elements to comprise abnormal events. Moreover, normal patterns in testing videos may be divergent from training ones, due to the existence of abnormalities. To address these problems, in this paper, an abnormality detector is proposed to detect abnormal events based on a statistical hypothesis test. The proposed detector treats each sample as a combination of a set of event patterns. Due to the unavailability of labeled abnormalities for training, abnormal patterns are adaptively extracted from incoming unlabeled testing samples. Contributions of this paper are listed as follows: 1) we introduce the idea of a statistical hypothesis test into the framework of abnormality detection, and abnormal events are identified as ones containing abnormal event patterns while possessing high abnormality detector scores; 2) due to the complexity of video events, noise seldom follows a simple distribution. For this reason, we approximate the complex noise distribution by employing a mixture of Gaussian. This benefits the modeling of video events and improves abnormality detection accuracies; and 3) because of the existence of abnormalities, there are always some unusually occurring normal events in the testing videos, which differ from the training ones. To represent normal events precisely, an online updating strategy is proposed to cover these cases in the normal event patterns. As a result, false detections are eliminated mostly. Extensive experiments and comparisons with state-of-the-art methods verify the effectiveness of the proposed algorithm.
Yuan Yuan 0001, Yachuang Feng, Xiaoqiang Lu
IEEE Trans. Cybern.3
2017 Remote Sensing Scene Classification by Unsupervised Representation Learning
abstract
With the rapid development of the satellite sensor technology, high spatial resolution remote sensing (HSR) data have attracted extensive attention in military and civilian applications. In order to make full use of these data, remote sensing scene classification becomes an important and necessary precedent task. In this paper, an unsupervised representation learning method is proposed to investigate deconvolution networks for remote sensing scene classification. First, a shallow weighted deconvolution network is utilized to learn a set of feature maps and filters for each image by minimizing the reconstruction error between the input image and the convolution result. The learned feature maps can capture the abundant edge and texture information of high spatial resolution images, which is definitely important for remote sensing images. After that, the spatial pyramid model (SPM) is used to aggregate features at different scales to maintain the spatial layout of HSR image scene. A discriminative representation for HSR image is obtained by combining the proposed weighted deconvolution model and SPM. Finally, the representation vector is input into a support vector machine to finish classification. We apply our method on two challenging HSR image data sets: the UCMerced data set with 21 scene categories and the Sydney data set with seven land-use categories. All the experimental results achieved by the proposed method outperform most state of the arts, which demonstrates the effectiveness of the proposed method.
Xiaoqiang Lu, Xiangtao Zheng, Yuan Yuan 0001
IEEE Trans. Geosci. Remote. Sens.1
2017 AID: A Benchmark Data Set for Performance Evaluation of Aerial Scene Classification
abstract
Aerial scene classification, which aims to automatically label an aerial image with a specific semantic category, is a fundamental problem for understanding high-resolution remote sensing imagery. In recent years, it has become an active task in the remote sensing area, and numerous algorithms have been proposed for this task, including many machine learning and data-driven approaches. However, the existing data sets for aerial scene classification, such as UC-Merced data set and WHU-RS19, contain relatively small sizes, and the results on them are already saturated. This largely limits the development of scene classification algorithms. This paper describes the Aerial Image data set (AID): a large-scale data set for aerial scene classification. The goal of AID is to advance the state of the arts in scene classification of remote sensing images. For creating AID, we collect and annotate more than 10000 aerial scene images. In addition, a comprehensive review of the existing aerial scene classification techniques as well as recent widely used deep learning methods is given. Finally, we provide a performance analysis of typical aerial scene classification and deep learning approaches on AID, which can be served as the baseline results on this benchmark.
Gui-Song Xia, Jingwen Hu 0001, Baoguang Shi, Xiang Bai, Yanfei Zhong, Liangpei Zhang 0001, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.8
2017 Dimensionality Reduction by Spatial-Spectral Preservation in Selected Bands
abstract
Dimensionality reduction (DR) has attracted extensive attention since it provides discriminative information of hyperspectral images (HSI) and reduces the computational burden. Though DR has gained rapid development in recent years, it is difficult to achieve higher classification accuracy while preserving the relevant original information of the spectral bands. To relieve this limitation, in this paper, a different DR framework is proposed to perform feature extraction on the selected bands. The proposed method uses determinantal point process to select the representative bands and to preserve the relevant original information of the spectral bands. The performance of classification is further improved by performing multiple Laplacian eigenmaps (LEs) on the selected bands. Different from the traditional LEs, multiple Laplacian matrices in this paper are defined by encoding spatial-spectral proximity on each band. A common low-dimensional representation is generated to capture the joint manifold structure from multiple Laplacian matrices. Experimental results on three real-world HSIs demonstrate that the proposed framework can lead to a significant advancement in HSI classification compared with the state-of-the-art methods.
Xiangtao Zheng, Yuan Yuan 0001, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.3
2017 A General Framework for Edited Video and Raw Video Summarization
abstract
In this paper, we build a general summarization framework for both of edited video and raw video summarization. Overall, our work can be divided into three folds. 1) Four models are designed to capture the properties of video summaries, i.e., containing important people and objects (importance), representative to the video content (representativeness), no similar key-shots (diversity), and smoothness of the storyline (storyness). Specifically, these models are applicable to both edited videos and raw videos. 2) A comprehensive score function is built with the weighted combination of the aforementioned four models. Note that the weights of the four models in the score function, denoted as property-weight, are learned in a supervised manner. Besides, the property-weights are learned for edited videos and raw videos, respectively. 3) The training set is constructed with both edited videos and raw videos in order to make up the lack of training data. Particularly, each training video is equipped with a pair of mixing-coefficients, which can reduce the structure mess in the training set caused by the rough mixture. We test our framework on three data sets, including edited videos, short raw videos, and long raw videos. Experimental results have verified the effectiveness of the proposed framework.
Xuelong Li 0001, Bin Zhao 0001, Xiaoqiang Lu
IEEE Trans. Image Process.3
2017 Latent Semantic Minimal Hashing for Image Retrieval
abstract
Hashing-based similarity search is an important technique for large-scale query-by-example image retrieval system, since it provides fast search with computation and memory efficiency. However, it is a challenge work to design compact codes to represent original features with good performance. Recently, a lot of unsupervised hashing methods have been proposed to focus on preserving geometric structure similarity of the data in the original feature space, but they have not yet fully refined image features and explored the latent semantic feature embedding in the data simultaneously. To address the problem, in this paper, a novel joint binary codes learning method is proposed to combine image feature to latent semantic feature with minimum encoding loss, which is referred as latent semantic minimal hashing. The latent semantic feature is learned based on matrix decomposition to refine original feature, thereby it makes the learned feature more discriminative. Moreover, a minimum encoding loss is combined with latent semantic feature learning process simultaneously, so as to guarantee the obtained binary codes are discriminative as well. Extensive experiments on several well-known large databases demonstrate that the proposed method outperforms most state-of-the-art hashing methods.
Xiaoqiang Lu, Xiangtao Zheng, Xuelong Li 0001
IEEE Trans. Image Process.1
2017 Discovering Diverse Subset for Unsupervised Hyperspectral Band Selection
abstract
Band selection, as a special case of the feature selection problem, tries to remove redundant bands and select a few important bands to represent the whole image cube. This has attracted much attention, since the selected bands provide discriminative information for further applications and reduce the computational burden. Though hyperspectral band selection has gained rapid development in recent years, it is still a challenging task because of the following requirements: 1) an effective model can capture the underlying relations between different high-dimensional spectral bands; 2) a fast and robust measure function can adapt to general hyperspectral tasks; and 3) an efficient search strategy can find the desired selected bands in reasonable computational time. To satisfy these requirements, a multigraph determinantal point process (MDPP) model is proposed to capture the full structure between different bands and efficiently find the optimal band subset in extensive hyperspectral applications. There are three main contributions: 1) graphical model is naturally transferred to address band selection problem by the proposed MDPP; 2) multiple graphs are designed to capture the intrinsic relationships between hyperspectral bands; and 3) mixture DPP is proposed to model the multiple dependencies in the proposed multiple graphs, and offers an efficient search strategy to select the optimal bands. To verify the superiority of the proposed method, experiments have been conducted on three hyperspectral applications, such as hyperspectral classification, anomaly detection, and target detection. The reliability of the proposed method in generic hyperspectral tasks is experimentally proved on four real-world hyperspectral data sets.
Yuan Yuan 0001, Xiangtao Zheng, Xiaoqiang Lu
IEEE Trans. Image Process.3
2016 Temporal Multimodal Learning in Audiovisual Speech Recognition
abstract
In view of the advantages of deep networks in producing useful representation, the generated features of different modality data (such as image, audio) can be jointly learned using Multimodal Restricted Boltzmann Machines (MRB-M). Recently, audiovisual speech recognition based the M-RBM has attracted much attention, and the MRBM shows its effectiveness in learning the joint representation across audiovisual modalities. However, the built networks have weakness in modeling the multimodal sequence which is the natural property of speech signal. In this paper, we will introduce a novel temporal multimodal deep learning architecture, named as Recurrent Temporal Multimodal RB-M (RTMRBM), that models multimodal sequences by transforming the sequence of connected MRBMs into a probabilistic series model. Compared with existing multimodal networks, it's simple and efficient in learning temporal joint representation. We evaluate our model on audiovisual speech datasets, two public (AVLetters and AVLetters2) and one self-build. The experimental results demonstrate that our approach can obviously improve the accuracy of recognition compared with standard MRBM and the temporal model based on conditional RBM. In addition, RTMRBM still outperforms non-temporal multimodal deep networks in the presence of the weakness of long-term dependencies.
Di Hu 0001, Xuelong Li 0001, Xiaoqiang Lu
CVPR3
2016 Deep Representation for Abnormal Event Detection in Crowded Scenes
abstract
Abnormal event detection is extremely important, especially for video surveillance. Nowadays, many detectors have been proposed based on hand-crafted features. However, it remains challenging to effectively distinguish abnormal events from normal ones. This paper proposes a deep representation based algorithm which extracts features in an unsupervised fashion. Specially, appearance, texture, and short-term motion features are automatically learned and fused with stacked denoising autoencoders. Subsequently, long-term temporal clues are modeled with a long short-term memory (LSTM) recurrent network, in order to discover meaningful regularities of video events. The abnormal events are identified as samples which disobey these regularities. Moreover, this paper proposes a spatial anomaly detection strategy via manifold ranking, aiming at excluding false alarms. Experiments and comparisons on real world datasets show that the proposed algorithm outperforms state of the arts for the abnormal event detection problem in crowded scenes.
Yachuang Feng, Yuan Yuan 0001, Xiaoqiang Lu
ACM Multimedia3
2016 Multimodal Learning via Exploring Deep Semantic Similarity
abstract
Deep learning is skilled at learning representation from raw data, which are embedded in the semantic space. Traditional multimodal networks take advantage of this, and maximize the joint distribution over the representations of different modalities. However, the similarity among the representations are not emphasized, which is an important property for multimodal data. In this paper, we will introduce a novel learning method for multimodal networks, named as Semantic Similarity Learning (SSL), which aims at training the model via enhancing the similarity between the high-level features of different modalities. Sets of experiments are conducted for evaluating the method on different multimodal networks and multiple tasks. The experimental results demonstrate the effectiveness of SSL in keeping the shared information and improving the discrimination. Particularly, SSL shows its ability in encouraging each modality to learn transferred knowledge from the other one when faced with missing data.
Di Hu 0001, Xiaoqiang Lu, Xuelong Li 0001
ACM Multimedia2
2016 Local structure learning in high resolution remote sensing image retrieval
Zhongxiang Du, Xuelong Li 0001, Xiaoqiang Lu
Neurocomputing3
2016 A target detection method for hyperspectral image based on mixture noise model
Xiangtao Zheng, Yuan Yuan 0001, Xiaoqiang Lu
Neurocomputing3
2016 Action recognition by joint learning
Yuan Yuan 0001, Lei Qi 0004, Xiaoqiang Lu
Image Vis. Comput.3
2016 Video parsing via spatiotemporally analysis with images
Xuelong Li 0001, Lichao Mou, Xiaoqiang Lu
Multim. Tools Appl.3
2016 A discriminative representation for human action recognition
Yuan Yuan 0001, Xiangtao Zheng, Xiaoqiang Lu
Pattern Recognit.3
2016 Spatiotemporal Statistics for Video Quality Assessment
abstract
It is an important task to design models for universal no-reference video quality assessment (NR-VQA) in multiple video processing and computer vision applications. However, most existing NR-VQA metrics are designed for specific distortion types, which are not often aware in practical applications. A further deficiency is that the spatial and temporal information of videos is hardly considered simultaneously. In this paper, we propose a new NR-VQA metric based on the spatiotemporal natural video statistics in 3D discrete cosine transform (3D-DCT) domain. In the proposed method, a set of features are first extracted based on the statistical analysis of 3D-DCT coefficients to characterize the spatiotemporal statistics of videos in different views. These features are used to predict the perceived video quality via the efficient linear support vector regression model afterward. The contributions of this paper are: 1) we explore the spatiotemporal statistics of videos in the 3D-DCT domain that has the inherent spatiotemporal encoding advantage over other widely used 2D transformations; 2) we extract a small set of simple but effective statistical features for video visual quality prediction; and 3) the proposed method is universal for multiple types of distortions and robust to different databases. The proposed method is tested on four widely used video databases. Extensive experimental results demonstrate that the proposed method is competitive with the state-of-art NR-VQA metrics and the top-performing full-reference VQA and reduced-reference VQA metrics.
Xuelong Li 0001, Xiaoqiang Lu
IEEE Trans. Image Process.3
2016 Surveillance Video Synopsis via Scaling Down Objects
abstract
Video synopsis is an effective technique to provide a compact representation of the original video by removing spatiotemporal redundancies and by preserving the essential activities. Most current approaches for video synopsis will cause collisions among objects, especially when the video is condensed much. In this paper, we present an approach for video synopsis to reduce the collisions. Our approach first shifts active objects along the time axis to compact the original video. Then, the sizes of the objects are reduced when collisions occur. Meanwhile, the geometric centroids of the objects will be kept unchanged to preserve the location information. Our contributions are threefold. First, an approach is proposed to decrease collisions in the synopsis video through reducing the sizes of the objects. Second, an optimization framework is developed to indicate the optimal time position and the appropriate reduction coefficient for each object. Finally, some metrics are proposed, and several experiments are carried out to evaluate the proposed approach. The experiments have demonstrated that the synopsis video produced by our approach has much fewer collisions while the compression ratio is high.
Xuelong Li 0001, Zhigang Wang 0002, Xiaoqiang Lu
IEEE Trans. Image Process.3
2015 MR image super-resolution via manifold regularized sparse learning
Xiaoqiang Lu, Zihan Huang, Yuan Yuan 0001
Neurocomputing1
2015 Image quality assessment: A sparse learning way
Yuan Yuan 0001, Xiaoqiang Lu
Neurocomputing3
2015 Semi-supervised change detection method for multi-temporal hyperspectral images
Yuan Yuan 0001, Haobo Lv, Xiaoqiang Lu
Neurocomputing3
2015 Low-rank representation for 3D hyperspectral images analysis from map perspective
Yuan Yuan 0001, Xiaoqiang Lu
Signal Process.3
2015 Multi-spectral pedestrian detection
Yuan Yuan 0001, Xiaoqiang Lu
Signal Process.2
2015 Scene Parsing From an MAP Perspective
abstract
Scene parsing is an important problem in the field of computer vision. Though many existing scene parsing approaches have obtained encouraging results, they fail to overcome within-category inconsistency and intercategory similarity of superpixels. To reduce the aforementioned problem, a novel method is proposed in this paper. The proposed approach consists of three main steps: 1) posterior category probability density function (PDF) is learned by an efficient low-rank representation classifier (LRRC); 2) prior contextual constraint PDF on the map of pixel categories is learned by Markov random fields; and 3) final parsing results are yielded up to the maximum a posterior process based on the two learned PDFs. In this case, the nature of being both dense for within-category affinities and almost zeros for intercategory affinities is integrated into our approach by using LRRC to model the posterior category PDF. Meanwhile, the contextual priori generated by modeling the prior contextual constraint PDF helps to promote the performance of scene parsing. Experiments on benchmark datasets show that the proposed approach outperforms the state-of-the-art approaches for scene parsing.
Xuelong Li 0001, Lichao Mou, Xiaoqiang Lu
IEEE Trans. Cybern.3
2015 Semi-Supervised Multitask Learning for Scene Recognition
abstract
Scene recognition has been widely studied to understand visual information from the level of objects and their relationships. Toward scene recognition, many methods have been proposed. They, however, encounter difficulty to improve the accuracy, mainly due to two limitations: 1) lack of analysis of intrinsic relationships across different scales, say, the initial input and its down-sampled versions and 2) existence of redundant features. This paper develops a semi-supervised learning mechanism to reduce the above two limitations. To address the first limitation, we propose a multitask model to integrate scene images of different resolutions. For the second limitation, we build a model of sparse feature selection-based manifold regularization (SFSMR) to select the optimal information and preserve the underlying manifold structure of data. SFSMR coordinates the advantages of sparse feature selection and manifold regulation. Finally, we link the multitask model and SFSMR, and propose the semi-supervised learning method to reduce the two limitations. Experimental results report the improvements of the accuracy in scene recognition.
Xiaoqiang Lu, Xuelong Li 0001, Lichao Mou
IEEE Trans. Cybern.1
2015 Substance Dependence Constrained Sparse NMF for Hyperspectral Unmixing
abstract
Hyperspectral unmixing is one of the most important problems in analyzing remote sensing images, which aims to decompose a mixed pixel into a collection of constituent materials named endmembers and their corresponding fractional abundances. Recently, various methods have been proposed to incorporate sparse constraints into hyperspectral unmixing and achieve advanced performance. However, most of them ignore the complex distribution of substances in hyperspectral data so that they are only effective in limited cases. In this paper, the concept of substance dependence is introduced to help hyperspectral unmixing. Generally, substance dependence can be considered in a local region by K-nearest neighbors method. However, since substances of hyperspectral images are complicatedly distributed, number K of the most similar substances to each substance is difficult to decide. In this case, substance dependence should be considered in the whole data space, and the number of the K most similar substances to each substance can be adaptively determined by searching from the whole space. Through maintaining the substance dependence during unmixing, the abundances resulted from the proposed method are closer to the real fractions, which lead to better unmixing performance. The following contributions can be summarized. 1) The concept of substance dependence is proposed to describe the complicated relationship between substances in the hyperspectral image. 2) We propose substance dependence constrained sparse nonnegative matrix factorization (SDSNMF) for hyperspectral unmixing. Using SDSNMF, we meet or exceed state-of-the-art unmixing performance. 3) Adequate experiments on both synthetic and real hyperspectral data have been tested. Compared with the state-of-the-art methods, the experimental results prove the superiority of the proposed method.
Yuan Yuan 0001, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.3
2015 Spectral-Spatial Kernel Regularized for Hyperspectral Image Denoising
abstract
Noise contamination is a ubiquitous problem in hyperspectral images (HSIs), which is a challenging and promising theme in many remote sensing applications. A large number of methods have been proposed to remove noise. Unfortunately, most denoising methods fail to take full advantages of the high spectral correlation and to simultaneously consider the specific noise distributions in HSIs. Recently, a spectral-spatial adaptive hyperspectral total variation (SSAHTV) was proposed and obtained promising results. However, the SSAHTV model is insensitive to the image details, which makes the edges blur. To overcome all of these drawbacks, a spectral-spatial kernel method for HSI denoising is proposed in this paper. The proposed method is inspired by the observation that the spectral-spatial information is highly redundant in HSIs, which is sufficient to estimate the clear images. In this paper, a spectral-spatial kernel regularization is proposed to maintain the spectral correlations in spectral dimension and to match the original structure between two spatial dimensions. Moreover, an adaptive mechanism is developed to balance the fidelity term according to different noise distributions in each band. Therefore, it cannot only suppress noise in the high-noise band but also preserve information in the low-noise band. The reliability of the proposed method in removing noise is experimentally proved on both simulated data and real data.
Yuan Yuan 0001, Xiangtao Zheng, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.3
2015 Scene Recognition by Manifold Regularized Deep Learning Architecture
abstract
Scene recognition is an important problem in the field of computer vision, because it helps to narrow the gap between the computer and the human beings on scene understanding. Semantic modeling is a popular technique used to fill the semantic gap in scene recognition. However, most of the semantic modeling approaches learn shallow, one-layer representations for scene recognition, while ignoring the structural information related between images, often resulting in poor performance. Modeled after our own human visual system, as it is intended to inherit humanlike judgment, a manifold regularized deep architecture is proposed for scene recognition. The proposed deep architecture exploits the structural information of the data, making for a mapping between visible layer and hidden layer. By the proposed approach, a deep architecture could be designed to learn the high-level features for scene recognition in an unsupervised fashion. Experiments on standard data sets show that our method outperforms the state-of-the-art used for scene recognition.
Yuan Yuan 0001, Lichao Mou, Xiaoqiang Lu
IEEE Trans. Neural Networks Learn. Syst.3
2014 Group sparse reconstruction for image segmentation
Xiaoqiang Lu, Xuelong Li 0001
Neurocomputing1
2014 Hybrid structure for robust dimensionality reduction
Xiaoqiang Lu, Yuan Yuan 0001
Neurocomputing1
2014 Refraction angle extracting strategy for fan-beam differential phase contrast CT
Renzhen Ye, Yi Tang 0003, Xiaoqiang Lu
Neurocomputing3
2014 Multiresolution Imaging
abstract
Imaging resolution has been standing as a core parameter in various applications of vision. Mostly, high resolutions are desirable or essential for many applications, e.g., in most remote sensing systems, and therefore much has been done to achieve a higher resolution of an image based on one or a series of images of relatively lower resolutions. On the other hand, lower resolutions are also preferred in some cases, e.g., for displaying images in a very small screen or interface. Accordingly, algorithms for image upsampling or downsampling have also been proposed. In the above algorithms, the downsampled or upsampled (super-resolution) versions of the original image are often taken as test images to evaluate the performance of the algorithms. However, there is one important question left unanswered: whether the downsampled or upsampled versions of the original image can represent the low-resolution or high-resolution real images from a camera? To tackle this point, the following works are carried out: 1) a multiresolution camera is designed to simultaneously capture images in three different resolutions; 2) at a given resolution (i.e., image size), the relationship between a pair of images is studied, one gained via either downsampling or super-resolution, and the other is directly captured at this given resolution by an imaging device; and 3) the performance of the algorithms of super-resolution and image downsampling is evaluated by using the given image pairs. The key reason why we can effectively tackle the aforementioned issues is that the designed multiresolution imaging camera can provide us with real images of different resolutions, which builds a solid foundation for evaluating various algorithms and analyzing the images with different resolutions, which is very important for vision.
Xiaoqiang Lu, Xuelong Li 0001
IEEE Trans. Cybern.1
2014 Alternatively Constrained Dictionary Learning For Image Superresolution
abstract
Dictionaries are crucial in sparse coding-based algorithm for image superresolution. Sparse coding is a typical unsupervised learning method to study the relationship between the patches of high-and low-resolution images. However, most of the sparse coding methods for image superresolution fail to simultaneously consider the geometrical structure of the dictionary and the corresponding coefficients, which may result in noticeable superresolution reconstruction artifacts. In other words, when a low-resolution image and its corresponding high-resolution image are represented in their feature spaces, the two sets of dictionaries and the obtained coefficients have intrinsic links, which has not yet been well studied. Motivated by the development on nonlocal self-similarity and manifold learning, a novel sparse coding method is reported to preserve the geometrical structure of the dictionary and the sparse coefficients of the data. Moreover, the proposed method can preserve the incoherence of dictionary entries and provide the sparse coefficients and learned dictionary from a new perspective, which have both reconstruction and discrimination properties to enhance the learning performance. Furthermore, to utilize the model of the proposed method more effectively for single-image superresolution, this paper also proposes a novel dictionary-pair learning method, which is named as two-stage dictionary training. Extensive experiments are carried out on a large set of images comparing with other popular algorithms for the same purpose, and the results clearly demonstrate the effectiveness of the proposed sparse representation model and the corresponding dictionary learning algorithm.
Xiaoqiang Lu, Yuan Yuan 0001, Pingkun Yan
IEEE Trans. Cybern.1
2014 Double Constrained NMF for Hyperspectral Unmixing
abstract
Given only the collected hyperspectral data, unmixing aims at obtaining the latent constituent materials and their corresponding fractional abundances. Recently, manynonnegative matrix factorization(NMF)-based algorithms have been developed to deal with this issue. Considering that the abundances of most materials may be sparse, the sparseness constraint is intuitively introduced into NMF. Although sparse NMF algorithms have achieved advanced performance in unmixing, the result is still susceptible to unstable decomposition and noise corruption. To reduce the aforementioned drawbacks, the structural information of the data is exploited to guide the unmixing. Since similar pixel spectra often imply similar substance constructions, clustering can explicitly characterize this similarity. Through maintaining the structural information during the unmixing, the resulting fractional abundances by the proposed algorithm can well coincide with the real distributions of constituent materials. Moreover, the additional clustering-based regularization term also lessens the interference of noise to some extent. The experimental results on synthetic and real hyperspectral data both illustrate the superiority of the proposed method compared with other state-of-the-art algorithms.
Xiaoqiang Lu, Hao Wu 0098, Yuan Yuan 0001
IEEE Trans. Geosci. Remote. Sens.1
2013 Sparse coding for image denoising using spike and slab prior
Xiaoqiang Lu, Yuan Yuan 0001, Pingkun Yan
Neurocomputing1
2013 Local tomography based on grey model
Renzhen Ye, Xiaoqiang Lu, Haihua Liu
Neurocomputing2
2013 Robust visual tracking with discriminative sparse learning
Xiaoqiang Lu, Yuan Yuan 0001, Pingkun Yan
Pattern Recognit.1
2013 Image Super-Resolution Via Double Sparsity Regularized Manifold Learning
abstract
Over the past few years, high resolutions have been desirable or essential, e.g., in online video systems, and therefore, much has been done to achieve an image of higher resolution from the corresponding low-resolution ones. This procedure of recovering/rebuilding is called single-image super-resolution (SR). Performance of image SR has been significantly improved via methods of sparse coding. That is to say, the image frame patch can be sparse linear combinations of basis elements. However, most of these existing methods fail to consider the local geometrical structure in the space of the training data. To take this crucial issue into account, this paper proposes a method named double sparsity regularized manifold learning (DSRML). DSRML can preserve the properties of the aforementioned local geometrical structure by employing manifold learning, e.g., locally linear embedding. Based on a large amount of experimental results, DSRML is demonstrated to be more robust and more effective than previous efforts in the task of single-image SR.
Xiaoqiang Lu, Yuan Yuan 0001, Pingkun Yan
IEEE Trans. Circuits Syst. Video Technol.1
2013 Graph-Regularized Low-Rank Representation for Destriping of Hyperspectral Images
abstract
Hyperspectral image destriping is a challenging and promising theme in remote sensing. Striping noise is a ubiquitous phenomenon in hyperspectral imagery, which may severely degrade the visual quality. A variety of methods have been proposed to effectively alleviate the effects of the striping noise. However, most of them fail to take full advantage of the high spectral correlation between the observation subimages in distinct bands and consider the local manifold structure of the hyperspectral data space. In order to remedy this drawback, in this paper, a novel graph-regularized low-rank representation (LRR) destriping algorithm is proposed by incorporating the LRR technique. To obtain desired destriping performance, two sides of performing destriping are included: 1) To exploit the high spectral correlation between the observation subimages in distinct bands, the technique of LRR is first utilized for destriping, and 2) to preserve the intrinsic local structure of the original hyperspectral data, the graph regularizer is incorporated in the objective function. The experimental results and quantitative analysis demonstrate that the proposed method can both remove striping noise and achieve cleaner and higher contrast reconstructed results.
Xiaoqiang Lu, Yuan Yuan 0001
IEEE Trans. Geosci. Remote. Sens.1
2013 Manifold Regularized Sparse NMF for Hyperspectral Unmixing
abstract
Hyperspectral unmixing is one of the most important techniques in analyzing hyperspectral images, which decomposes a mixed pixel into a collection of constituent materials weighted by their proportions. Recently, many sparse nonnegative matrix factorization (NMF) algorithms have achieved advanced performance for hyperspectral unmixing because they overcome the difficulty of absence of pure pixels and sufficiently utilize the sparse characteristic of the data. However, most existing sparse NMF algorithms for hyperspectral unmixing only consider the Euclidean structure of the hyperspectral data space. In fact, hyperspectral data are more likely to lie on a low-dimensional submanifold embedded in the high-dimensional ambient space. Thus, it is necessary to consider the intrinsic manifold structure for hyperspectral unmixing. In order to exploit the latent manifold structure of the data during the decomposition, manifold regularization is incorporated into sparsity-constrained NMF for unmixing in this paper. Since the additional manifold regularization term can keep the close link between the original image and the material abundance maps, the proposed approach leads to a more desired unmixing performance. The experimental results on synthetic and real hyperspectral data both illustrate the superiority of the proposed method compared with other state-of-the-art approaches.
Xiaoqiang Lu, Hao Wu 0098, Yuan Yuan 0001, Pingkun Yan, Xuelong Li 0001
IEEE Trans. Geosci. Remote. Sens.1
2013 Sparse Coding From a Bayesian Perspective
abstract
Sparse coding is a promising theme in computer vision. Most of the existing sparse coding methods are based on either l0 or l1 penalty, which often leads to unstable solution or biased estimation. This is because of the nonconvexity and discontinuity of the l0 penalty and the over-penalization on the true large coefficients of the l1 penalty. In this paper, sparse coding is interpreted from a novel Bayesian perspective, which results in a new objective function through maximum a posteriori estimation. The obtained solution of the objective function can generate more stable results than the l0 penalty and smaller reconstruction errors than the l1 penalty. In addition, the convergence property of the proposed algorithm for sparse coding is also established. The experiments on applications in single image super-resolution and visual tracking demonstrate that the proposed method is more effective than other state-of-the-art methods.
Xiaoqiang Lu, Yuan Yuan 0001
IEEE Trans. Neural Networks Learn. Syst.1
2012 Geometry constrained sparse coding for single image super-resolution
abstract
The choice of the over-complete dictionary that sparsely represents data is of prime importance for sparse coding-based image super-resolution. Sparse coding is a typical unsupervised learning method to generate an over-complete dictionary. However, most of the sparse coding methods for image super-resolution fail to simultaneously consider the geometrical structure of the dictionary and corresponding coefficients, which may result in noticeable super-resolution reconstruction artifacts. In this paper, a novel sparse coding method is proposed to preserve the geometrical structure of the dictionary and the sparse coefficients of the data. Moreover, the proposed method can preserve the incoherence of dictionary entries, which is critical for sparse representation. Inspired by the development on non-local self-similarity and manifold learning, the proposed sparse coding method can provide the sparse coefficients and learned dictionary from a new perspective, which have both reconstruction and discrimination properties to enhance the learning performance. Extensive experimental results on image super-resolution have demonstrated the effectiveness of the proposed method.
Xiaoqiang Lu, Pingkun Yan, Yuan Yuan 0001, Xuelong Li 0001
CVPR1
2012 Robust Alternative Minimization for Matrix Completion
abstract
Recently, much attention has been drawn to the problem of matrix completion, which arises in a number of fields, including computer vision, pattern recognition, sensor network, and recommendation systems. This paper proposes a novel algorithm, named robust alternative minimization (RAM), which is based on the constraint of low rank to complete an unknown matrix. The proposed RAM algorithm can effectively reduce the relative reconstruction error of the recovered matrix. It is numerically easier to minimize the objective function and more stable for large-scale matrix completion compared with other existing methods. It is robust and efficient for low-rank matrix completion, and the convergence of the RAM algorithm is also established. Numerical results showed that both the recovery accuracy and running time of the RAM algorithm are competitive with other reported methods. Moreover, the applications of the RAM algorithm to low-rank image recovery demonstrated that it achieves satisfactory performance.
Xiaoqiang Lu, Tieliang Gong, Pingkun Yan, Yuan Yuan 0001, Xuelong Li 0001
IEEE Trans. Syst. Man Cybern. Part B1
2011 Image Denoising via Improved Sparse Coding
abstract
This paper presents a novel dictionary learning method for image denoising, which removes zero-mean independent identically distributed additive noise from a given image. Choosing noisy image itself to train an over-complete dictionary, the dictionary trained by traditional sparse coding methods contains noise information. Through mathematical derivation of equation, we found that a lower bound of dictionary is related with the level of noise in dictionary learning. The proposed idea is to take advantage of the noise information for designing a sparse coding algorithm called improved sparse coding (ISC), which effectively suppresses the noise influence for training a dictionary. This denoising framework utilizes the effective \nmethod, which is based on sparse representations over trained dictionaries. Acquiring an over-complete dictionary by ISC mainly includes three stages. Firstly, we utilize \nK-means method to group the noisy image patches. Secondly, each dictionary is trained by ISC in corresponding class. Finally, an over-complete dictionary is merged \nby these dictionaries. Theory analysis and experimental results both demonstrate that the proposed method yields excellent performance.
Xiaoqiang Lu, Pingkun Yan, Luoqing Li, Xuelong Li 0001
BMVC1
2011 A novel alternative algorithm for limited angle tomography
abstract
This paper studies incomplete data problems of circular cone-beam computed tomography, which occur frequently in medical imaging and industrial imaging. The incomplete data problems in which projection data are only available in an angular range can be attributed to the limited angle tomography. Limited angle tomography is a severely ill-posed inverse problem. In recent years, image reconstruction based on total variation (TV) was employed to reduce the problem and gave better performance on edge-preserving reconstruction. However, the artificial parameter can only be determined through considerable experimentation. In this paper, an alternating minimization method based on TV is proposed to reduce the data insufficiency in tomographic imaging. This novel alternating minimization method provides a robust and effective reconstruction without any artificial parameter in the iterative processes, by using the TV as a multiplicative constraint. The results demonstrate that this new reconstruction method brings satisfactory performance.
Xiaoqiang Lu, Yuan Yuan 0001, Pingkun Yan, Xuelong Li 0001
ICIP1
2011 Local learning-based image super-resolution
abstract
Local learning algorithm has been widely used in single-frame super-resolution reconstruction algorithm, such as neighbor embedding algorithm [1] and locality preserving constraints algorithm [2]. Neighbor embedding algorithm is based on manifold assumption, which defines that the embedded neighbor patches are contained in a single manifold. While manifold assumption does not always hold. In this paper, we present a novel local learning-based image single-frame SR reconstruction algorithm with kernel ridge regression (KRR). Firstly, Gabor filter is adopted to extract texture information from low-resolution patches as the feature. Secondly, each input low-resolution feature patch utilizes K nearest neighbor algorithm to generate a local structure. Finally, KRR is employed to learn a map from input low-resolution (LR) feature patches to high-resolution (HR) feature patches in the corresponding local structure. Experimental results show the effectiveness of our method.
Xiaoqiang Lu, Yuan Yuan 0001, Pingkun Yan, Luoqing Li, Xuelong Li 0001
MMSP1
2011 Image reconstruction by an alternating minimisation
Xiaoqiang Lu, Yi Sun 0009, Yuan Yuan 0001
Neurocomputing1
2011 Optimization for limited angle tomography in medical image processing
Xiaoqiang Lu, Yi Sun 0009, Yuan Yuan 0001
Pattern Recognit.1
2010 Adaptive wavelet-Galerkin methods for limited angle tomography
Xiaoqiang Lu, Yi Sun 0009, Gangfeng Bai
Image Vis. Comput.1