Yuntao Qian

dblp:25/6037 · DBLP profile ↗
← Back
134ranked-venue papers
16as first author
60since 2021 · last 2026
0000-0002-7418-5891ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 81 · 11 first-author · 38 since 2021Graphics, computer vision, multimedia, augmented reality and games · 40 · 5 first-author · 16 since 2021Artificial intelligence and machine learning · 20 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021
YearPublicationVenuePosition
2026 Bridging the resolution gap: Semantic-aware alignment for cross-resolution change detection
Wang Hao, Fengchao Xiong, Jianfeng Lu 0003, Jingzhou Chen, Yuntao Qian
Pattern Recognit.6
2025 Label Relationship Graph-Enhanced Class Hierarchy for Incremental Classification of Remote Sensing Images
abstract
Incremental learning is a strategy that continuously incorporates new data to tackle emerging tasks without the need for retraining the model. While effective, it encounters the challenge of catastrophic forgetting. Hierarchical Classification (HC) enhances classification accuracy and efficiency by assigning objects to multiple labels within a hierarchical structure. This paper introduces LRGIC, a novel approach specifically designed for incremental hierarchical classification of remote sensing images. LRGIC combines class hierarchy (CH), a Feature Pyramid Network (FPN), and a Learning Without Forgetting (LWF) strategy, while also incorporating a HEX graph to constrain labels and encode hierarchical knowledge. These elements, integrated within a hierarchical residual network, significantly boost classification performance. The FPN captures multi-scale features, and the LWF strategy facilitates the learning of new categories without reusing old samples. Experimental results demonstrate that LRGIC effectively classifies new categories and their hierarchical relationships, while preserving the performance of existing categories, underscoring its substantial research and practical value.
Yang Chu 0001, Yuntao Qian
ICASSP2
2025 Incremental classification of remote sensing images using feature pyramid and class hierarchy enhanced by label relationship graphs
Yang Chu 0001, Yuntao Qian
Appl. Intell.2
2025 Multi-domain universal representation learning for hyperspectral object tracking
Zhuanfeng Li, Fengchao Xiong, Jianfeng Lu 0003, Jing Wang 0062, Diqi Chen, Jun Zhou 0001, Yuntao Qian
Pattern Recognit.7
2025 Hierarchical Contrastive Learning for Multigranularity Ship Classification With Learnable Class Queries
abstract
Ship targets in remote sensing images can be categorized at various granularities due to variations in image quality, ranging from general ship categories to fine-grained classes like Nimitz-class carriers. Traditional studies mainly focus on fine-grained ship classification, often neglecting samples observed at coarser-grained levels. Samples distributed across multiple granularity levels exhibit semantic relationships among their annotated classes, enabling hierarchical knowledge transfer during model training. This paper incorporates two semantic relationships into deep learning-based representation learning and class prediction: parent-child relationships across levels and mutual exclusivity among sibling categories. For hierarchical representation learning, the proposed hierarchical contrastive learning algorithm extracts category-specific representations from input images and aligns them with their semantic relationships, ensuring that parent and child categories share similarities while sibling categories remain distinct. For hierarchical class predictions, a novel consistency loss ensures coherence in probability distributions between parent and child categories. Specially, cross-entropy loss is employed to impose mutual exclusivity among sibling categories. In this paper, a multi-modal dataset is also designedly developed for hierarchical classification, which integrates optical and synthetic aperture radar (SAR) images across multiple hierarchical levels. Experiments on two popular datasets and a multi-modal dataset demonstrate that the proposed method outperforms state-of-the-art approaches in hierarchical multi-granularity ship classification.
Jingzhou Chen, Fengchao Xiong, Yuntao Qian, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.3
2025 Spatial-Spectral-Temporal Correlation Filter for Hyperspectral Object Tracking
abstract
Object tracking with hyperspectral videos (HSVs) offers significant advantages due to the captured spectral fingerprint information, which provides detailed physical material characteristics. While correlation filter (CF)-based tracking methods align well with the high-dimensional nature of HSVs, they often fall short of fully utilizing the spatial–spectral–temporal structure inherent in these data. In this article, we introduce a spatial–spectral–temporal CF (SSTCF) framework to address these limitations. SSTCF employs the spatial-spectral histogram of gradients and fractional abundances as features to characterize the spatial-spectral structure of the object. A low-rank constraint is integrated into the CF framework to enhance the global spectral semantic dependencies among learned filters. In addition, a temporal constraint is incorporated to ensure filter consistency across consecutive frames, further improving tracking continuity between nearby frames. Extensive experiments demonstrate that our SSTCF tracker achieves more accurate and stable performance. The source code will be publicly available athttps://github.com/bearshng/SSTCF
Fengchao Xiong, Yongle Sun, Jun Zhou 0001, Jianfeng Lu 0003, Yuntao Qian
IEEE Trans. Geosci. Remote. Sens.5
2025 Frequency- and Spatial-Domain Saliency Network for Remote Sensing Cross-Modal Retrieval
abstract
Recently, remote sensing text-image retrieval has gained increasing attention among researchers due to its capability to provide abundant, inclusive, and multi-perspective information. To extract the salient representation of the two modalities and realize information alignment, existing methods usually apply convolution calculation, matrix multiplication, and attention mechanisms to model features in the spatial domain. However, unlike natural images, remote sensing data contain a considerable amount of noise, and the spatial domain calculation methods will lead to information smoothing and noise exchange, ultimately reducing the robustness of feature representation. In this article, a Frequency and Spatial-domain Saliency Network (FSSN) is proposed by extracting saliency features in both frequency and spatial domains to further enhance the efficiency of retrieval networks. Specifically, the FSSN first designs a Frequency-domain Intra-modal Low-pass Filter (FILF) by using the Fourier Transform to convert the spatial domain representation into the frequency domain representation, and leverage low-pass filtering to filter out the noise information of the image. Afterwards, the Space-domain Inter-modal Boundary-based Saliency module (SIBS) is devised, which fully utilizes the positive and negative boundaries and draws an innovative assessment mechanism to automatically recognize the effective regions and words without noise shifting in the spatial domain. Finally, the Frequency and Spatial-domain Saliency Fusion (FSSF) is designed to realize the effective integration of features obtained from the frequency and spatial domain according to the consistency of the identity masks. Quantitative and qualitative experiments are implemented on four remote sensing benchmarks to showcase the notable effectiveness achieved by the joint modeling network of multiple domains. Specifically, the proposed FSSN outperforms the best model RemoteClip by 7.58% of the mR evaluation matrix on the UCM dataset.
Jie Nie, Xiu Li 0006, Yuntao Qian, Zhiqiang Wei 0002
IEEE Trans. Geosci. Remote. Sens.5
2025 Whole Semantic Sparse Coding Network for Remote Sensing Image-Text Retrieval
abstract
In recent years, cross-modal text-image retrieval in remote sensing has gained prominence as a research focus due to its potential to provide abundant, inclusive, and multi-perspective information. However, existing methods usually focus on salient features, but these salient features cannot describe the image or text completely, resulting in the loss of some important details and discriminable information. In this article, a Whole Semantic Sparse Coding Network (WSSCN) is proposed for remote sensing image-text retrieval to build a complete and reliable features description for further improving the performance of the retrieval model. Specifically, the WSSCN first designs a Whole Semantic Sparse Representation Coding (WSSRC) module by constructing a robust semantic library to transform the dense features matrix of the image and text into a whole semantic sparse matrix that enables multiple semantic decoupling and leads to a more precise and detailed features expression. Afterwards, the Intra- and Inter-Modal Consistency (IIMC) module is devised to improve the intra-modal and inter-modal consistency of the whole semantic sparse representation from different models. Finally, the Salient and Whole Semantic Adaptive Learning (SWSAL) loss is proposed to focus on salient information or whole semantic information by calculating an adjusted parameter. Quantitative and qualitative experiments are performed on four extensive datasets for cross-modal retrieval in remote sensing to showcase the notable effectiveness achieved by implementing a whole semantic sparse coding network.
Xiu Li 0006, Chenxue Yang, Jie Nie, Yiyun Guo, Yuntao Qian, Zhiqiang Wei 0002
IEEE Trans. Geosci. Remote. Sens.7
2025 SUIT: Spatial-Spectral Union-Intersection Interaction Network for Hyperspectral Object Tracking
abstract
Hyperspectral videos (HSVs), with their inherent spatial-spectral-temporal structure, offer distinct advantages in challenging tracking scenarios such as cluttered backgrounds and small objects. However, existing methods primarily focus on spatial interactions between the template and search regions, often overlooking spectral interactions, leading to suboptimal performance. To address this issue, this paper investigates spectral interactions from both the architectural and training perspectives. At the architectural level, we first establish band-wise long-range spatial relationships between the template and search regions using Transformers. We then model spectral interactions using the inclusion-exclusion principle from set theory, treating them as the union of spatial interactions across all bands. This enables the effective integration of both shared and band-specific spatial cues. At the training level, we introduce a spectral loss to enforce material distribution alignment between the template and predicted regions, enhancing robustness to shape deformation and appearance variations. Extensive experiments demonstrate that our tracker achieves state-of-the-art tracking performance. The source code, trained models and results will be publicly available via https://github.com/bearshng/suit to support reproducibility.
Fengchao Xiong, Zhenxing Wu, Jun Zhou 0001, Sen Jia 0001, Yuntao Qian
IEEE Trans. Image Process.5
2024 RGB Images Enhancing Hyperspectral Image Denoising with Diffusion Model
abstract
Deep learning (DL)-based hyperspectral image (HSI) denoising has achieved remarkable accomplishments but is still limited by insufficient training data of noisy-clean pairs. This paper proposes a novel approach to enhance the diffusion model (DM)-based HSI denoising by leveraging abundant RGB images. Specifically, an RGB-DM is pre-trained on the RGB images to capture comprehensive spatial information, and then it is integrated with the HSI-DM using a fusion operator during the reverse diffusion process, yielding improved denoising results. To optimize this fusion process, we derive the optimal fusion weight by minimizing signal distortion. Experimental results on CAVE and ICVL datasets demonstrate the effectiveness of our approach by outperforming state-of-the-art methods.
Keli Deng, Yuntao Qian
ICASSP3
2024 Incremental Classification of Remote Sensing Images with Feature Pyramid and Class Hierarchy
abstract
Incremental learning is a methodology aimed at addressing novel tasks by continuously acquiring new data while retaining knowledge obtained from previous learning tasks. Hierarchical classification (HC) assigns multiple labels to each object, establishing a hierarchical relationship among these labels. The objective of incremental HC for remote sensing images is to accurately differentiate recently added categories while preserving the HC structure. In this paper, we propose a novel incremental learning method called HC-FPNSP for HC, leveraging class hierarchy (CH), feature pyramid network (FPN), and learning without forgetting (LWF) strategy. This method introduces the information of CH into incremental learning for remote sensing image classification. Experimental results demonstrate the efficacy of the proposed method in accurately capturing the hierarchical relationships of newly added categories while maintaining satisfactory classification performance for existing categories.
Yang Chu 0001, Yuntao Qian
IGARSS2
2024 Semantic-Aware Alignment Network for Cross-Resolution Change Detection
abstract
Cross-resolution change detection (CRCD) is of significant practical importance in disaster assessment, rapid urban transitions, and various applications. Conventional change detection methods are primarily tailored for bitemporal images with consistent spatial resolution, rendering them unsuitable for direct application to CRCD tasks. This limitation stems from the substantial scale differences and pixel-wise misalignment prevalent in cross-resolution remote sensing images. In response to these challenges, we introduce a semantic-aware alignment network (SA-Net). SA-Net utilizes cross-attention to map bitemporal images into a shared semantic space, effectively alleviating the difficulties of the subsequent alignment arising from semantic mismatches. Furthermore, a joint transformer featuring an encoder-decoder architecture is employed to extract global information and learn the geometric parameters for spatial alignment between bitemporal images. Experimental evaluations on two real-collected datasets, HTCD and MRCDD, showcase the superior performance of our proposed SA-Net in CRCD tasks.
Fengchao Xiong, Jianfeng Lu 0003, Minchao Ye, Jun Zhou 0001, Yuntao Qian
IGARSS6
2024 Efficiently Exploiting Spatially Variant Knowledge for Video Deblurring
abstract
Video deblurring is a challenging task as the blur is often spatially variant. Existing methods mainly engage in building the spatial-temporal correspondence among the frames. As one of the widely-used frameworks, the long-range temporal propagation usually suffers from the expensive computation cost and error accumulation caused by the numerous connections among temporal frames. Meanwhile, the exploration of spatial-variant information from the neighbor frames is often ignored in video deblurring. To tackle these issues, we tailor an efficient short-range multi-scale framework slimming the long-range propagation and exploiting the most relevant neighbor temporal knowledge. For capturing spatial knowledge, we further propose a spatial feature extractor, named the spatially variant adaptive block, to adaptively generate the location-wise kernel to cater to the spatially variant character of blur. For efficient temporal exploitation, a simple inter-frame shift as a motion compensation is developed to avoid expensive long temporal relevance modeling. Both quantitative and qualitative evaluation results on benchmark datasets demonstrate that the proposed algorithm performs favorably against state-of-the-art methods.
Qian Xu 0014, Xiaobin Hu, Donghao Luo 0001, Ying Tai, Chengjie Wang 0001, Yuntao Qian
IEEE Trans. Circuits Syst. Video Technol.6
2024 Diffusion-Model-Based Hyperspectral Unmixing Using Spectral Prior Distribution
abstract
Hyperspectral unmixing is a crucial task for identifying the constituent materials and their respective distributions in a scene. Utilizing known spectral libraries as prior information, semi-blind unmixing methods (also known as spectral-library-based methods) have been proven advantageous over unblind and blind approaches. However, such methods encounter two main challenges: difficulty in handling large-scale spectral libraries and vulnerability to variabilities stemming from differences between the underlying signatures and those in the spectral library. To address these challenges, a novel approach named DiffUn, based on a diffusion model, is proposed in this article for semi-blind hyperspectral unmixing. DiffUn considers hyperspectral unmixing as a sampling process from a posterior distribution, where the prior distribution is learned from a spectral library, and the likelihood distribution is estimated from the observed data by the linear spectral mixture model. Specifically, the approach first learns the spectral prior distribution from a spectral library through an unconditional diffusion model, then integrates this prior knowledge into the reverse process of the diffusion model, and finally samples the underlying endmembers and corresponding abundances from the posterior distribution. Since spectral prior distribution estimation is not sensitive to library size, DiffUn exhibits superior unmixing performance even in a large-scale library. Furthermore, DiffUn permits sampling spectral signatures from a continuous probabilistic distribution, whereas conventional semi-blind unmixing methods only allow endmembers selected from the library, which is a discrete space. Thus, DiffUn shows greater robustness to spectral variations. Experimental results on synthetic and real-world datasets demonstrate DiffUn outperforming the state-of-the-art semi-blind unmixing methods. The code is available at https://github.com/Dmsw/DiffUn.git.
Keli Deng, Yuntao Qian, Jie Nie, Jun Zhou 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 Hyperspectral Image Denoising via Spatial-Spectral Recurrent Transformer
abstract
Hyperspectral images (HSIs) often suffer from noise arising from both intra-imaging mechanisms and environmental factors. Leveraging domain knowledge specific to HSIs, such as global spectral correlation (GSC) and non-local spatial self-similarity (NSS), is crucial for effective denoising. Existing methods tend to independently utilize each of these knowledge components with multiple blocks, overlooking the inherent 3D nature of HSIs where domain knowledge is strongly interlinked, resulting in suboptimal performance. To address this challenge, this paper introduces a spatial-spectral recurrent transformer U-Net (SSRT-UNet) for HSI denoising. The proposed SSRT-UNet integrates NSS and GSC properties within a single SSRT block. This block consists of a spatial branch and a spectral branch. The spectral branch employs a combination of transformer and recurrent neural network to perform recurrent computations across bands, allowing for GSC exploitation beyond a fixed number of bands. Concurrently, the spatial branch encodes NSS for each band by sharingkeysandvalueswith the spectral branch under the guidance of GSC. The interaction between the two branches enables the joint utilization of NSS and GSC, avoiding their independent treatment. Experimental results demonstrate that our method outperforms several alternative approaches. The source code will be available at https://github.com/lronkitty/SSRT.
Guanyiman Fu, Fengchao Xiong, Jianfeng Lu 0003, Jun Zhou 0001, Jiantao Zhou 0001, Yuntao Qian
IEEE Trans. Geosci. Remote. Sens.6
2024 Material-Guided Multiview Fusion Network for Hyperspectral Object Tracking
abstract
Hyperspectral videos (HSVs) have more potential in object tracking than color videos thanks to their material identification ability. Nevertheless, previous works have not fully explored the benefits of the material information, resulting in limited representation ability and tracking accuracy. To address this issue, this paper introduces a material-guided multi-view fusion network for improved tracking. Specifically, we combine false-color information, hyperspectral information, and material information obtained by hyperspectral unmixing to provide a rich multi-view representation of the object. Cross-material attention is employed to capture the interaction among materials, enabling the network to focus on the most relevant materials for the target. Furthermore, leveraging the discriminative ability of material view, a novel material-guided multi-view fusion module is proposed to capture both intra-view and cross-view long-range spatial dependencies for effective feature aggregation. Thanks to the enhanced representation ability of each view and the integration of the complementary advantages of all views, our network is more capable of suppressing the tracking drift in various challenging scenes and achieving accurate object localization. Extensive experiments show that our tracker achieves state-of-the-art tracking performance. The source code will be available at https://github.com/hscv/MMF-Net.
Zhuanfeng Li, Fengchao Xiong, Jun Zhou 0001, Jianfeng Lu 0003, Yuntao Qian
IEEE Trans. Geosci. Remote. Sens.6
2024 Wavelet Siamese Network With Semi-Supervised Domain Adaptation for Remote Sensing Image Change Detection
abstract
Change detection is a crucial technique in remote sensing image analysis and faces challenges, such as background complexity and appearance shift, resulting in incomplete change boundaries and pseudochanges. This article introduces a novel wavelet Siamese network with semi-supervised domain adaptation (DA) to address these issues, named WS-Net++. WS-Net++ establishes spatial–frequency interactions between bitemporal images to enhance the completeness of the change boundaries. The spatial-domain interaction highlights the pixelwise differences. The frequency-domain interaction first adaptively adjusts the contributions from different frequency components based on image context. Within-frequency and between-frequency interactions are further constructed to capture the frequency-domain differences, enabling the adaptive and effective handling of both overall and subtle changes. In addition, WS-Net++ employs a semi-supervised DA strategy to mitigate the appearance shifts between bitemporal images. By categorizing regions into changed, unchanged, and regions of no interest in a semi-supervised manner, the network minimizes intraclass discrepancies within unchanged regions and maximizes interclass discrepancies between changed regions, reducing the domain gap. Experimental results on the LEVIR-CD, WHU-CD, and CLCD datasets demonstrate that our WS-Net++ outperforms alternative methods, achieving the$F1$scores of 91.31%, 94.52%, and 79.77%, respectively. The code and models will be publicly available athttps://github.com/JiTaiTai/WS-Net_Plusfor reproducible research.
Fengchao Xiong, Tianhan Li, Yi Yang 0071, Jun Zhou 0001, Jianfeng Lu 0003, Yuntao Qian
IEEE Trans. Geosci. Remote. Sens.6
2024 Adaptive Graph Modeling With Self-Training for Heterogeneous Cross-Scene Hyperspectral Image Classification
abstract
The small-sample-size problem of hyperspectral image (HSI) classification has recently gained considerable attention. Cross-scene HSI classification has emerged as an effective solution to this problem. In real-world applications, different HSI scenes are often captured by diverse sensors, resulting in variations between scenes. Graph modeling, as a method to represent relationships, leverages semantic information to establish connections between scenes, thereby facilitating transfer learning by aligning their features. However, in scenarios with only a few labeled target samples, the resulting graph is typically sparse and can only capture weak cross-scene relationships. Studies have shown that a dense and fault-tolerant graph is beneficial for transfer learning in small-sample-size cases. Consequently, we propose a novel heterogeneous transfer learning approach called adaptive graph modeling with self-training (AGM-ST). Unlike conventional graph modeling methods that employ predefined graph weights, adaptive graph modeling (AGM) employs a learnable network to generate graph weights based on the similarities of spectral–spatial features. Additionally, an adaptive cutoff threshold is trained to eliminate weak relationships between samples that may be potentially incorrect. Subsequently, a cross-scene graph loss is designed based on the generated graph to align the feature spaces of the source and target scenes. Furthermore, the unlabeled samples from the target scene are gradually updated with pseudo labels using the self-training (ST) technique, which enhances semantic information and improves graph modeling. Experimental evaluations conducted on three cross-scene HSI datasets have demonstrated the effectiveness of the proposed AGM-ST approach.
Minchao Ye, Junbin Chen, Fengchao Xiong, Yuntao Qian
IEEE Trans. Geosci. Remote. Sens.4
2024 Discriminative Vision Transformer for Heterogeneous Cross-Domain Hyperspectral Image Classification
abstract
The transformer has been introduced in the hyperspectral image (HSI) classification, demonstrating outstanding capability in capturing global features compared to the convolutional neural network (CNN). However, the small-sample-size problem poses a significant challenge in practical HSI classification, especially in training the transformer. To tackle this issue, cross-domain transfer learning is adopted as a practical solution, which transfers the information from a source domain with abundant labeled samples to a target domain with limited labeled samples. This article proposes a novel transfer learning method for heterogeneous cross-domain HSI classification called cross-domain discriminative vision transformer (CD-DViT). This algorithm primarily contains three key contributions. First, source samples are mapped to the target domain through an encoder-decoder architecture, and the mapped source samples can be used to train the target classifier. Second, the cross-attention mechanism is utilized to construct two blocks for achieving the domainwise and classwise feature alignments (FAs), respectively. Specifically, the combination of the cross-attention mechanism with the domain discriminator aims to learn domain-invariant features, thereby facilitating domainwise alignment and alleviating domain shift. Third, knowledge distillation (KD) is adopted to learn more information from the target domain and assist in classifying target samples. Our experiments on three real-world cross-domain HSI datasets demonstrate the effectiveness of the proposed approach.
Minchao Ye, Jiawei Ling, Wanli Huo, Zhaojuan Zhang, Fengchao Xiong, Yuntao Qian
IEEE Trans. Geosci. Remote. Sens.6
2024 Building Cross-Domain Mapping Chains From Multi-CycleGAN for Hyperspectral Image Classification
abstract
The small-sample-size issue in hyperspectral image (HSI) classification remains a significant challenge. To improve the classification accuracy of a dataset with a few labeled samples (target domain), we can use knowledge from another dataset with sufficient labeled samples (source domain). This is called cross-domain HSI classification. Transfer learning enables knowledge transfer between source and target domains. Different HSI datasets are often acquired by different sensors, resulting in different characteristics. Consequently, different HSIs possess different feature spaces, and knowledge transfer between them becomes difficult due to heterogeneities. CycleGAN, based on adversarial learning, can help solve heterogeneous transfer learning tasks by establishing the two-way mapping between two different feature spaces. However, CycleGAN contains only one cycle for data, leading to large mapping errors. This article proposes a novel CycleGAN-based transfer learning method for cross-domain HSI classification. The proposed method extends the two-way mapping of CycleGAN. It incorporates multiple mapping cycles to construct a multi-CycleGAN, which is then unfolded to derive the Cross-Domain Mapping Chain (CDMC) model. The generators in our proposed CDMC provide accurate mappings between domains. Moreover, we calculate and accumulate errors in each cycle, and the backpropagation of accumulated errors through the chains improves the model’s performance. Besides, auxiliary classifiers are introduced to account for class-conditional distributions in the mapping process. Experimental results on three real-world heterogeneous cross-domain HSI datasets show the effectiveness of the proposed method.
Minchao Ye, Zhihao Meng, Yuntao Qian
IEEE Trans. Geosci. Remote. Sens.3
2024 Iterative Low-Rank Network for Hyperspectral Image Denoising
abstract
Hyperspectral image (HSI) denoising is a crucial preprocessing step for subsequent tasks. The clean HSI usually reside in a low-dimensional subspace, which can be captured by low-rank and sparse representation, known as the physical prior of HSI. It is generally challenging to adequately use such physical properties for effective denoising while preserving image details. This article introduces a novel iterative low-rank network (ILRNet) to address these challenges. ILRNet integrates the strengths of model-driven and data-driven approaches by embedding a rank minimization module (RMM) within a U-Net architecture. This module transforms feature maps into the wavelet domain and applies singular value thresholding (SVT) to the low-frequency components during the forward pass, leveraging the spectral low-rankness of HSIs in the feature domain. The parameter, closely related to the hyperparameter of the singular vector thresholding algorithm, is adaptively learned from the data, allowing for flexible and effective capture of low-rankness across different scenarios. Additionally, ILRNet features an iterative refinement process that adaptively combines intermediate denoised HSIs with noisy inputs. This manner ensures progressive enhancement and superior preservation of image details. Experimental results demonstrate that ILRNet achieves state-of-the-art performance in both synthetic and real-world noise removal tasks.
Jin Ye 0008, Fengchao Xiong, Jun Zhou 0001, Yuntao Qian
IEEE Trans. Geosci. Remote. Sens.4
2023 Iterative Refinement Network for Hyperspectral Image Denoising
abstract
Hyperspectral image (HSI) denoising is an important pre-processing procedure for subsequent tasks. Learning a direct mapping from the observed noisy HSI to its clean counterpart is challenging, especially in the case of very severe noise. The learning difficulty can be greatly reduced with the iterative refinement, combining the denoising results with the noisy HSI to produce a cleaner HSI for further denoising. To this end, we introduce an iterative refinement denoising network (IRDNet) for HSIs. The network consists of three key components, i.e., a coarse estimation module, a multi-stage refinement module, and a λ(•) module. The coarse estimation module provides the initial estimate for the starting point of further refinement. The lightweight refinement module progressively performs noise reduction on the weighted combination of noisy inputs and denoising results from the previous layer. Instead of fixed weights, the λ(•) module adaptively provides the layer-wise weight for each band for the aforementioned combinations. Extensive experiments on synthetic and real-world datasets show that our IRDNet favorably outperforms alternative methods.
Fengchao Xiong, Jun Zhou 0001, Yuntao Qian
ICME4
2023 Incremental Learning of Remote Sensing Target Classification with Class Hierarchy
abstract
Incremental learning can continuously learn to address new tasks from new data while preserving knowledge learned from previously learned tasks. Incremental remote sensing target classification aims to accurately classify newly added classes while maintaining the classification performance of the old classes. In this paper, we propose a new incremental learning method using the information of class hierarchy (CH) and the strategy of Learning without Forgetting (LWF), named CH-LWF. As we know, the target classes in remote sensing images always have hierarchical relationships and are organized as hierarchical tree structures, and CH has been proven to benefit target classification. Our work in this paper may be the first to introduce CH into incremental learning of remote sensing images. In addition, LWF can learn new classes only with new training samples, and the training samples for old classes are not required to be reused, which is convenient in real applications.
Yang Chu 0001, Yuntao Qian
IGARSS3
2023 A Noise-Model-Free Hyperspectral Image Denoising Method Based on Diffusion Model
abstract
Hyperspectral Image (HSI) denoising is a crucial preprocessing step to ensure the accuracy of the subsequent HSI analysis and interpretation. Neural network methods recently achieve state-of-the-art performance in HSI denoising. Nevertheless, these methods are typically trained on specific noise models, which could limit their performance as noise models in HSI may vary across different spectral bands. To mitigate this problem, we introduce an HSI denoising method based on the Diffusion Model (DM) whose training process is independent of the noise model, noted as a noise-model-free method. In this method, we first introduce a DM for HSI by increasing the input and output dimensions to incorporate spectral and spatial information. We then derive a DM-based HSI denoising process for a common noise model. Moreover, to address the issue of overfitting during training, we have introduced a stochastic sampling method to more effectively balance the significance of both spatial and spectral information. Experimental results on in-distribution and out-of-distribution samples demonstrate the efficacy of our approach.
Keli Deng, Zhongshun Jiang, Qipeng Qian, Yuntao Qian
IGARSS5
2023 Hierarchical Superpixel Relation Graph Combined with Convolutional Sparse Coding for Self-Supervised Hyperspectral Image Denoising
abstract
Self-supervised methods have recently been widely used for hyperspectral image (HSI) denoising. As only a single noisy HSI required to be restored is used for learning, effectively exploring the spatial-spectral joint information for HSI representation becomes to be a more critical problem. In this paper, we propose a self-supervised denoiser based on a blind-spot network and a new feature extraction method that combines convolutional neural network (CNN) and graph neural network (GNN). The features from the convolutional backbone network are fed into the GNN as graph nodes. With hierarchical superpixel segmentation of HSI, the feature nodes with contextual and hierarchical relationships are feature-enhanced by the powerful feature interaction capability of GNN. Experimental results show that our method is competitive compared with state-of-the-art methods.
Zhongshun Jiang, Qipeng Qian, Yuntao Qian
IGARSS4
2023 Wavelet Siamese Network for Change Detection in Remote Sensing Images
abstract
Change detection is a technique used to identify semantic differences between co-registered images of the same area captured at different times. However, current methods often overlook the fact that the low-frequency and high-frequency components of these images play distinct roles in change detection. Our method decomposes each feature map into its low-frequency and high-frequency components and then uses an attention mechanism to adjust the contribution of each component to handle different types of changes. Low-frequency information can help detect overall changes, and high-frequency information can enhance the integrity of the change boundaries. Experiments on the LEVIR-CD, WHU-CD and CLCD datasets show that our model outperforms the state-of-the-art method and the ablation study demonstrates that this approach improve the accuracy of the change detection.
Tianhan Li, Fengchao Xiong, Zhuanfeng Li, Jun Zhou 0001, Yuntao Qian
IGARSS6
2023 Cross-Domain Hyperspectral Image Classification Based on Graph Convolutional Networks
abstract
A major challenge in hyperspectral image (HSI) classification is the small-sample-size problem. Cross-domain information can help solve the problem. In cross-domain HSI classification, the source domain has many samples, while the target domain has fewer samples. Transfer learning can transfer knowledge from the source domain to the target domain. The source and target domains are mostly captured by different sensors and thus come from different feature spaces. Heterogeneous transfer learning can solve this problem. This paper proposes a transfer learning method based on a crossdomain graph convolutional network (CD-GCN). A class co-occurrence semantic graph is built between heterogeneous spaces of source and target domains. Then graph convolutional network (GCN) is adopted to learn the features of graphs. To handle the different feature dimensions, a feature alignment subnet is proposed. By combining a feature alignment subnet and a GCN feature extraction subnet, the proposed model CD-GCN transfers knowledge between heterogeneous domains. Experiments on two cross-domain HSI datasets prove that CD-GCN overperforms many transfer learning methods.
Minchao Ye, Yuntao Qian, Qipeng Qian
IGARSS3
2023 Cross-Domain Hyperspectral Image Classification Based on Transformer
abstract
Small-sample-size problem is a big challenge in hyperspectral image (HSI) classification. Deep learning-based methods, especially Transformer, may need more training samples to train a satisfactory model. Cross-domain classification has been proven to be effective in handling the small-sample-size problem. In two HSI scenes sharing the same land-cover classes, one with sufficient labeled samples is called the source domain, while the other with limited labeled samples is called the target domain. Thus, the information on the source domain could help the target domain improve classification performance. This paper proposes a cross-domain Vision Transformer (CD-ViT) method for heterogeneous HSI classification. CD-ViT maps the source samples to the target domain for supplementing training samples. In addition, cross-attention is used to align the source and target features. Moreover, knowledge distillation is employed to learn more transferable information. Experiments on three different cross-domain HSI datasets demonstrate the effectiveness of the proposed approach.
Jiawei Ling, Minchao Ye, Yuntao Qian, Qipeng Qian
IGARSS3
2023 Domain-invariant attention network for transfer learning between cross-scene hyperspectral images
abstract
Abstract Small‐sample‐size problem is always a challenge for hyperspectral image (HSI) classification. Considering the co‐occurrence of land‐cover classes between similar scenes, transfer learning can be performed, and cross‐scene classification is deemed a feasible approach proposed in recent years. In cross‐scene classification, the source scene which possesses sufficient labelled samples is used for assisting the classification of the target scene that has a few labelled samples. In most situations, different HSI scenes are imaged by different sensors resulting in their various input feature dimensions (i.e. number of bands), hence heterogeneous transfer learning is desired. An end‐to‐end heterogeneous transfer learning algorithm namely domain‐invariant attention network (DIAN) is proposed to solve the cross‐scene classification problem. The DIAN mainly contains two modules. (1) A feature‐alignment CNN (FACNN) is applied to extract features from source and target scenes, respectively, aiming at projecting the heterogeneous features from two scenes into a shared low‐dimensional subspace. (2) A domain‐invariant attention block is developed to gain cross‐domain consistency with a specially designed class‐specific domain‐invariance loss, thus further eliminating the domain shift. The experiments on two different cross‐scene HSI datasets show that the proposed DIAN achieves satisfying classification results.
Minchao Ye, Zhihao Meng, Fengchao Xiong, Yuntao Qian
IET Comput. Vis.5
2023 Multitask Sparse Representation Model-Inspired Network for Hyperspectral Image Denoising
abstract
Hyperspectral images (HSIs) are prone to noise because of the imaging mechanism and environment. This paper proposes a multitask sparse representation (SR) model inspired neural network for HSI denoising. Unlike other deep learning-based methods, our network is interpretable, whose network architecture is induced by unfolding the iterative optimization of a multitask sparse representation model. On the one hand, the model globally represents the common structure among bands, such as image edges, with the shared sparse coefficients. On the other hand, it separately encodes the unique structure of individual bands with unshared ones to capture image details. Accordingly, our network has three modules: the shared SR module, the unshared SR module, and the image reconstruction (IR) module. All the modules are connected with a specific operation of the iterative optimization algorithm, equipping the network with clear physical interpretation. Experimental results on both synthetic and real-world datasets demonstrate the superior performance of our method, visually and quantitatively. The codes will be publicly available at https://github.com/bearshng/mtsrnn for reproducible research.
Fengchao Xiong, Jiantao Zhou 0001, Jun Zhou 0001, Jianfeng Lu 0003, Yuntao Qian
IEEE Trans. Geosci. Remote. Sens.5
2023 Deep Parameterized Neural Networks for Hyperspectral Image Denoising
abstract
Sparse representation (SR)-based hyperspectral image (HSI) denoising methods normally average the local denoising results of multiple overlapped cubes to recover the whole HSI. Though interpretable, they rely on cumbersome hyperparameter settings and ignore the relationship between overlapped cubes, leading to poor denoising performance. This article combines SR and convolutional neural networks and introduces a deep parameterized sparse neural network (DPNet-S) to address the above issues. DPNet-S parameterizes the SR-based HSI denoising model with two modules: 1) sparse optimizer to extract sparse feature maps from noisy HSIs via recurrent usage of convolution, deconvolution, and soft shrinkage operations; and 2) image reconstructor to recover the denoised HSI from its sparse feature maps via deconvolution operations. We further replace the soft shrinkage operator with U-Net architecture to account for general HSI priors and more effectively capture the complex structures of HSIs, resulting in DPNet-U. Both networks directly learn the parameters from data and perform denoising on the whole HSI, which overcomes the limitations of SR-based methods. Moreover, our networks are generated from the denoising model and optimization procedures, thus leveraging the knowledge embedded and relying less on the number of training samples. Extensive experiments on both synthetic and real-world HSIs show that our DPNet-S and DPNet-U achieve remarkable results when compared with state-of-the-art methods. The codes will be publicly available athttps://github.com/bearshng/dpnetsfor reproducible research.
Fengchao Xiong, Jun Zhou 0001, Jiantao Zhou 0001, Jianfeng Lu 0003, Yuntao Qian
IEEE Trans. Geosci. Remote. Sens.5
2023 Ca-STANet: Spatiotemporal Attention Network for Chlorophyll-a Prediction With Gap-Filled Remote Sensing Data
abstract
Long-term chlorophyll-a (Chl-a) prediction has the potential to provide an early warning of red tide, and support fishery management and marine ecosystem health. The existing learning-based Chl-a prediction methods mostly predict a single point or multiple points with monitoring data. However, the monitoring data are subject to sparse sampling and difficult to be measured in a large-scale and synchronous way. Moreover, the advanced learning-based models for point Chl-a prediction, such as long short-term memory (LSTM) and convolutional neural network (CNN)-LSTM, are unable to fully mining the spatio-temporal correlation of Chl-a variations. Therefore, by using the satellite remote sensing data with extensive coverage, we design a framework, namely Ca-STANet, to simultaneously predict the Chl-a of all the locations in a large-scale area from the perspective of spatio-temporal field. Specifically in our method, the original data are firstly divided into multiple sub-regions to capture the spatial heterogeneity of large-scale area. Then, two modules are respectively operated to mine the spatial correlation and long-term dependency features. Finally, the outputs from the two modules are integrated by a fusion module to fully mine the spatio-temporal correlations, which are exploited to attain the final Chl-a prediction. In this paper, the proposed Ca-STANet is comprehensively evaluated and compared with the legacy methods based on the OC-CCI Chl-a 5.0 data of the Bohai Sea. The results demonstrate that the proposed Ca-STANet is highly effective for Chl-a prediction and achieves higher prediction accuracy than the baseline methods. Moreover, as the OC-CCI Chl-a 5.0 data have many missing areas, we introduce DINEOF method to fill the data gaps before using them for prediction.
Bohan Li 0005, Jie Nie, Yuntao Qian, Lie-Liang Yang
IEEE Trans. Geosci. Remote. Sens.4
2023 Learning a Deep Ensemble Network With Band Importance for Hyperspectral Object Tracking
abstract
Attributing to material identification ability powered by a large number of spectral bands, hyperspectral videos (HSVs) have great potential for object tracking. Most hyperspectral trackers employ manually designed features rather than deeply learned features to describe objects due to limited available HSVs for training, leaving a huge gap to improve the tracking performance. In this paper, we propose an end-to-end deep ensemble network (SEE-Net) to address this challenge. Specifically, we first establish a spectral self-expressive model to learn the band correlation, indicating the importance of a single band in forming hyperspectral data. We parameterize the optimization of the model with a spectral self-expressive module to learn the nonlinear mapping from input hyperspectral frames to band importance. In this way, the prior knowledge of bands is transformed into a learnable network architecture, which has high computational efficiency and can fast adapt to the changes of target appearance because of no iterative optimization. The band importance is further exploited from two aspects. On the one hand, according to the band importance, each frame of HSVs is divided into several three-channel false-color images which are then used for deep feature extraction and location. On the other hand, based on the band importance, the importance of each false-color image is computed, which is then used to assemble the tracking results from individual false-color images. In this way, the unreliable tracking caused by false-color images of low importance can be suppressed to a large extent. Extensive experimental results show that SEE-Net performs favorably against the state-of-the-art approaches. The source code will be available at https://github.com/hscv/SEE-Net.
Zhuanfeng Li, Fengchao Xiong, Jun Zhou 0001, Jianfeng Lu 0003, Yuntao Qian
IEEE Trans. Image Process.5
2022 Label Relation Graphs Enhanced Hierarchical Residual Network for Hierarchical Multi-Granularity Classification
abstract
Hierarchical multi-granularity classification (HMC) assigns hierarchical multi-granularity labels to each object and focuses on encoding the label hierarchy, e.g., [“Albatross”, “Laysan Albatross”] from coarse-to-fine levels. However, the definition of what is fine-grained is subjective, and the image quality may affect the identification. Thus, samples could be observed at any level of the hierarchy, e.g., [“Albatross”] or [“Albatross”, “Laysan Albatross”], and examples discerned at coarse categories are often neglected in the conventional setting of HMC. In this paper, we study the HMC problem in which objects are labeled at any level of the hierarchy. The essential designs of the proposed method are derived from two motivations: (1) learning with objects labeled at various levels should transfer hierarchical knowledge between levels; (2) lower-level classes should inherit attributes related to upper-level superclasses. The proposed combinatorial loss maximizes the marginal probability of the observed ground truth label by aggregating information from related labels defined in the tree hierarchy. If the observed label is at the leaf level, the combinatorial loss further imposes the multi-class cross-entropy loss to increase the weight of fine-grained classification loss. Considering the hierarchical feature interaction, we propose a hierarchical residual network (HRN), in which granularity-specific features from parent levels acting as residual connections are added to features of children levels. Experiments on three commonly used datasets demonstrate the effectiveness of our approach compared to the state-of-the-art HMC approaches. The code will be available at https://github.com/MonsterZhZh/HRN.
Jingzhou Chen, Peng Wang 0100, Yuntao Qian
CVPR4
2022 Material-Guided Siamese Fusion Network for Hyperspectral Object Tracking
abstract
Hyperspectral videos (HSVs) have more potential in target tracking than color videos thanks to the material identification capability provided by abundant spectral bands. Due to limited HSVs for training, most current hyperspectral trackers are based on hand-crafted features rather than deeply learned ones, resulting in poor tracking performance. This paper introduces a material-guided Siamese fusion network (SiamF) for hyperspectral object tracking to make up this gap. Belonging to the Siamese tracker family and SiamF aims to model the appearance of hyperspectral objects using backbone networks trained on color images. Specifically, SiamF splits each hyperspectral frame into multiple groups of false-color images according to their band importance. Then SiamF employs a hyperspectral feature fusion (HFF) module with a dense connection architecture to integrate the extracted features from different layers and band groups, producing a multi-scale multilevel spatial-spectral representation of the targets. Instead of direct addition or concatenation, HFF employs global-local channel attention for feature fusion, so that yielded features capture the global and local structure of a specific object. Moreover, online spatial and material classifiers are developed to inject spatial and material appearance changes information into SiamF for adaptively online tracking. Experimental results demonstrate our tracker outperforms alternative methods.
Zhuanfeng Li, Fengchao Xiong, Jianfeng Lu 0003, Jun Zhou 0001, Yuntao Qian
ICASSP5
2022 Multitask Sparse Neural Network for Hyperspectral Image Denoising
abstract
Data-driven deep learning (DL)-based methods directly learn the nonlinear mapping between noisy hyperspectral images (HSIs) and corresponding clean ones. However, DLbased methods neglect the prior knowledge of HSIs embodied by physical models. Consequently, they require complex network architectures and a large number of training samples. To address the above issues, this paper introduces a multitask sparse neural network (MTSNN) which bridges the sparsity prior of HSIs with data-driven deep learning for HSI denoising. Specifically, we first build a multitask sparse (MTS) denoising model which shares sparse coefficients among bands to exploit the spectral-spatial correlation and learns a dictionary for each band to depict the distinct spatial structure among bands. The iterative optimization of the MTS model is then unfolded to yield our MTSNN by introducing some learnable parameters. MTSNN is a multi-branch network. Each branch performs a single denoising task for an individual band. All branches are connected by shared coefficients, forming multitask denoising for all bands. The hybrid advantages of the MTS model and data-driven learning equip MTSNN with strong denoising ability, preferable learning capability, superior interpretability, and higher generalization capacity. Experimental results demonstrate that our method achieves state-of-the-art denoising performance compared with several alternative approaches.
Fengchao Xiong, Minchao Ye, Jun Zhou 0001, Jianfeng Lu 0003, Yuntao Qian
ICASSP5
2022 A Self-Supervised Hyperspectral Image Restoration Method Based on Convolutional Sparse Coding and Superpixel Segmentation
abstract
Denoising is a common preprocessing step before hyperspectral image (HSI) analysis and interpretation. Since noise-free training samples are not required, self-supervised deep learning is becoming a new research trend in image restoration. In this paper, we propose a self-supervised denoising method based on noise2noise learning scheme that assumes two noisy signals in a pair being randomly generated from the same clean signal. To obtain noisy-noisy training pairs that meet the requirement of noise2noise learning, the target noisy HSI is divided into several homogeneous image patches by superpixel segmentation. Pixels located in the same superpixel have strong spectral similarity and spatial neighborhood, so they are used to construct noisy-noisy training pairs and the spectral-spatial correlations in an HSI can be embedded in our denoiser via training samples. Moreover, the convolutional sparse coding is introduced and unfolded as the backbone network for spectral signature representation and restoration. Experimental results show that our method is competitive compared with the state-of-the-art methods.
Zhongshun Jiang, Yuntao Qian
IGARSS4
2022 Cross-Scene Hyperspectral Image Classification Based on Cycle-Consistent Adversarial Networks
abstract
Lack of labeled training samples is a challenge in hyperspectral image (HSI) classification. Cross-scene classification is a valid solution to few-shot learning problem. In cross-scene classification, two strongly related HSI scenes are considered, one with sufficient labeled samples is called source scene, while the other one containing limited labeled samples is called target scene. By establishing connections between two scenes, abundant labeled samples in source scene can benefit the classification of target scene. In this paper, a novel model named cycle auxiliary classifier generative adversarial network (Cycle-AC-GAN) is proposed for heterogeneous transfer learning across source and target scenes. In Cycle-AC-GAN, a source-to-target generator and a target-to-source generator are simultaneously built. Thus, a two-way mapping can be effectively established between source and target scenes with the adversarial training. In addition, different from existing CycleGAN, in Cycle-AC-GAN, each discriminator contains a binary domain classifier and an auxiliary land-cover classifier. The auxiliary classifiers can align the class-conditional distributions between source and target HSIs. Inspiring experimental results on two real-world cross-scene HSI datasets demonstrate the effectiveness of the proposed approach.
Zhihao Meng, Minchao Ye, Futian Yao, Fengchao Xiong, Yuntao Qian
IGARSS5
2022 Cross-Domain Attention Network for Hyperspectral Image Classification
abstract
Expensive cost of labeling leads to few-shot learning problem in hyperspectral image (HSI) classification. Cross-scene classification is a novel approach to solve this problem. In this work, we propose an end-to-end heterogeneous transfer learning algorithm namely cross-domain attention network (CDAN) to settle the cross-scene classification problem. CDAN mainly contains two modules. 1) A two-stream HybirdSN architecture is designed for extracting features from source and target scenes, aiming at projecting the features into a shared low-dimensional subspace. 2) Cross-domain attention mechanism is adopted based on the consistency of features between different scenes. A cross-domain updating rule is proposed for training the subnet. CDAN is proved to be effective according to the experiments on two different cross-scene HSI datasets.
Minchao Ye, Ling Lei 0002, Fengchao Xiong, Yuntao Qian
IGARSS5
2022 Spatial-Spectral Convolutional Sparse Neural Network for Hyperspectral Image Denoising
abstract
Sparse representation (SR) is a widely accepted hyper-spectral image (HSI) denoising model. Because of the curse of dimensionality and the desire to better fit the data, the SR models are typically deployed on small and fully overlapping blocks whose results are averaged to produce the global de-noised HSI. This “local-global” denoising mechanism ignores the dependencies between blocks, resulting in visual artifacts. This paper describes the underlying clean HSI with a 3D con-volutional sparse coding (CSC) model, representing the HSI with a linear combination of few shift-invariant 3D spatial-spectral filters in a global dictionary. Instead of operating on patches, the CSC model sees the clean HSI is generated from a sum of local atoms that appear in a small number of locations throughout the image, naturally retaining the relationship between pixels. Moreover, we unfold the optimization process of the model into a spatial-spectral convolutional sparse neural network which absorbs the interpretation ability of the model while supporting discriminative learning from data. Experimental results on both synthetic and real-world datasets show that our network achieves competitive denoising performances, qualitatively and quantitatively.
Fengchao Xiong, Minchao Ye, Jun Zhou 0001, Yuntao Qian
IGARSS4
2022 Self-Supervised Learning Hyperspectral Image Denoiser with Separated Spectral-Spatial Feature Extraction
abstract
Deep learning-based methods have achieved remarkable results in the field of hyperspectral image (HSI) denoising, and these methods are typically trained on pairs of noisy input and clean target images. How to deal with the noise in real-world HSIs when clean targets are unavailable is still a challenging problem. In this paper, we propose a self-supervised HSI denoiser in which only a single noisy HSI is utilized. We exploit the blind-spot network and extend the method in the spatial-spectral space to accomplish self-supervised learning. In order to better extract spatial-spectral features with limited training samples, we use separable feature extraction modules to extract spectral-spatial joint information of HSI separately and finally fuse these features. Experimental results on both simulated and real hyperspectral datasets show that our proposed method outperforms some state-of-the-art denoising approaches.
Minchao Ye, Yuntao Qian
IGARSS4
2022 RLPath: a knowledge graph link prediction method using reinforcement learning based attentive relation path searching and representation learning
Ling Chen 0001, Jun Cui 0003, Xing Tang 0006, Yuntao Qian, Yansheng Li 0001, Yongjun Zhang 0002
Appl. Intell.4
2022 Learning an Occlusion-Aware Network for Video Deblurring
abstract
Video deblurring is a challenging task since the blur is caused by camera shake, object motions, etc. The success of the state-of-the-art methods stems mainly from exploiting the temporal information of neighboring frames through alignment. When there exists occlusion among the sequence, these approaches become less effective for inaccurate alignment. In this paper, we propose an effective occlusion-aware network to handle the occlusion for video deblurring. The proposed module first generates a coarse pixel-wise alignment filter to explore the temporal information and then learns an adaptive affine transformation to deal with the occluded areas. In addition, a self-attention mechanism is developed to better model the occluded pixels. To further improve the performance, we progress a multi-scale strategy and train the network in an end-to-end manner. Both quantitative and qualitative experimental results show that the proposed method achieves favorable performance against state-of-the-art methods on the benchmark datasets. The code and trained models are available at:https://github.com/XQLuck/code.git
Qian Xu 0014, Jinshan Pan, Yuntao Qian
IEEE Trans. Circuits Syst. Video Technol.3
2022 Bidirectional Transformer for Video Deblurring
abstract
We present a bidirectional transformer network to exploit the long-range informative dependency both in temporal and spatial domains for video deblurring. Motivated by the fact that the optical flow is related to the latent frames rather than blurry ones in the degradation process, we first develop a pre-deblur module to generate initial latent frames so that we can use these initial latent frames to estimate optical flow for better exploring the temporal information. Then we propose an effective bidirectional transformer to explore the long-range information, where it first aggregates the temporal information of the whole sequence through backward and forward propagation with the estimated optical flow, then a recurrent pixel-wise neighbor transformer (PWNT) block is developed at the end of the module to extract the useful spatial information. We embed our bidirectional transformer into a deep convolutional neural network and evaluate it on the publicly available video deblurring benchmarks. Extensive experimental results show that the proposed method performs favorably against the state-of-the-art methods. The implementation is available in this website:https://github.com/Rebeccaxq/pwnt
Qian Xu 0014, Yuntao Qian
IEEE Trans. Circuits Syst. Video Technol.2
2022 Hierarchical Multilabel Ship Classification in Remote Sensing Images Using Label Relation Graphs
abstract
Hierarchical multilabel classification (HMC) assigns multiple labels to each instance with the labels organized under hierarchical relations. In ship classification in remote sensing images, depending on the expert knowledge and image quality, the same type of ships in different remote sensing images may be annotated with different class labels from coarse to fine levels such as merchant ship (MS) or container ship (CTS). In this article, we propose a novel deep network with two output channels and their associated loss functions to learn an HMC classifier using samples labeled at different levels in the hierarchy. In the proposed network, a hierarchy and exclusion (HEX) graph is introduced to model the label hierarchy, which satisfies hierarchical constraints by encoding semantic relations between any two labels. The output nodes of the first channel are organized according to the HEX graph, and its corresponding probabilistic classification loss is built to reflect the hierarchical structure of the HEX graph. On the other hand, the output nodes of the second channel only represent the finest grained (last level in the hierarchy) classes, and its multiclass cross-entropy loss is designed to enhance the discriminative power of the HMC classifier on the last level labels, which is also compatible with constraints in the HEX graph. The combination of these two losses from two output channels can effectively transfer the hierarchical information of ship taxonomy during network training. Experimental results on two commonly used ship datasets demonstrate that the proposed method outperforms the state-of-the-art HMC approaches, and is especially advantageous when trained with fewer fine-grained samples.
Jingzhou Chen, Yuntao Qian
IEEE Trans. Geosci. Remote. Sens.2
2022 Incremental Detection of Remote Sensing Objects With Feature Pyramid and Knowledge Distillation
abstract
When a detection model that has been well-trained on a set of classes faces new classes, incremental learning is always necessary to adapt the model to detect the new classes. In most scenarios, it is required to preserve the learned knowledge of the old classes during incremental learning rather than reusing the training data from the old classes. Since the objects in remote sensing images often appear in various sizes, arbitrary directions, and dense distribution, it further makes incremental learning-based object detection more difficult. In this article, a new architecture for incremental object detection is proposed based on feature pyramid and knowledge distillation. Especially, by means of a feature pyramid network (FPN), the objects with various scales are detected in the different layers of the feature pyramid. Motivated by Learning without Forgetting (LwF), a new branch is expended in the last layer of FPN, and knowledge distillation is applied to the outputs of the old branch to maintain the old learning capability for the old classes. Multitask learning is adopted to jointly optimize the losses from two branches. Experiments on two widely used remote sensing data sets show our promising performance compared with state-of-the-art incremental object detection methods.
Jingzhou Chen, Ling Chen 0001, Haibin Cai, Yuntao Qian
IEEE Trans. Geosci. Remote. Sens.5
2022 Nonlocal Spatial-Spectral Neural Network for Hyperspectral Image Denoising
abstract
Hyperspectral image (HSI) denoising is an essential preprocessing step to improve the quality of HSIs. The difficulty of HSI denoising lies in effectively modeling the intrinsic characteristics of HSIs, such as spatial-spectral correlation, global spectral correlation, and nonlocal spatial correlation. This paper introduces a nonlocal spatial-spectral neural network (NSSNN) for HSI denoising by considering the above three factors in a unified network. More specifically, NSSNN is based on the residual U-Net and embedded with the introduced spatial-spectral recurrent (SSR) blocks and nonlocal self-similarity (NSS) blocks. The SSR block comprises 3D convolutions, one light recurrence, and one highway network. 3D convolution helps exploit the spatial-spectral correlation. The light recurrence and highway network make up the recurrent computation component and refined component, respectively, to model the global spectral correlation. NSS block is based on crisscross attention and can exploit the long-range spatial contexts effectively and efficiently. Attributing to effective modeling of the spatial-spectral correlation, the global spectral correlation, and the nonlocal spatial correlation, our NSSNN has a strong denoising ability. Extensive experiments show the superior denoising effectiveness of our method on synthetic and real-world datasets when compared to alternative methods. The source code will be available at https://github.com/lronkitty/NSSNN.
Guanyiman Fu, Fengchao Xiong, Jianfeng Lu 0003, Jun Zhou 0001, Yuntao Qian
IEEE Trans. Geosci. Remote. Sens.5
2022 Hyperspectral Image Restoration With Self-Supervised Learning: A Two-Stage Training Approach
abstract
Hyperspectral image (HSI) denoising is a crucial preprocessing task to improve the performance of the subsequent HSI interpretation and applications. With recent progress in deep learning, HSI denoising methods based on deep neural networks have attracted increasing interest and achieved the state-of-the-art performance. Nevertheless, most of these methods are based on network structures originally developed for grayscale and color images and require a change of network structure to be applicable to HSIs. The new network architectures often lead to complicated models and limited flexibility, which, in turn, result in difficulty in learning and demand of a large number of training samples. In this article, we propose an innovative two-stage learning method including pretraining and fine-tuning procedures. In the first stage, a denoising convolutional neural network can be pretrained with pairs of corrupted and clean images. In the second stage, the pretrained network is fine-tuned via a self-supervised learning strategy to capture the spectral correlation in HSIs. The training pairs in the second stage are constructed from the neighboring band images in the target noisy HSI, leading to a novel idea of embedding spectral information into denoiser through the target image, rather than the change of the network architecture. This model has strong adaptability such that many image denoising networks can be easily adopted for HSIs, while the external hyperspectral training set is optional but not mandatory. Experimental results show that our method has competitive performance compared with the state-of-the-art approaches.
Yuntao Qian, Ling Chen 0001, Jun Zhou 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 MAC-Net: Model-Aided Nonlocal Neural Network for Hyperspectral Image Denoising
abstract
Hyperspectral image (HSI) denoising is an ill-posed inverse problem. The underlying physical model is always important to tackle this problem, which is unfortunately ignored by most of the current deep learning (DL)-based methods, producing poor denoising performance. To address this issue, this article introduces an end-to-end model-aided nonlocal neural network (MAC-Net) which simultaneously takes the spectral low-rank model and spatial deep prior into account for HSI noise reduction. Specifically, motivated by the success of the spectral low-rank model in depicting the strong spectral correlations and the nonlocal similarity prior in capturing spatial long-range dependencies, we first build a spectral low-rank model and then integrate a nonlocal U-Net into the model. In this way, we obtain a hybrid model-based and DL-based HSI denoising method where the spatial local and nonlocal multi-scale and spectral low-rank structures are effectively exploited. After that, we cast the optimization and denoising procedure of the hybrid method as a forward process of a neural network and introduce a set of learnable modules to yield our MAC-Net. Compared with traditional model-based methods, our MAC-Net overcomes the difficulties of accurate modeling, thanks to the strong learning and representation ability of DL. Unlike most “black-box” DL-based methods, the spectral low-rank model is beneficial to increase the generalization ability of the network and decrease the requirement of training samples. Experimental results on the natural and remote-sensing HSIs show that MAC-Net achieves state-of-the-art performance over both model-based and DL-based methods. The source code and data of this article will be made publicly available athttps://github.com/bearshng/mac-netfor reproducible research.
Fengchao Xiong, Jun Zhou 0001, Qinling Zhao, Jianfeng Lu 0003, Yuntao Qian
IEEE Trans. Geosci. Remote. Sens.5
2022 SNMF-Net: Learning a Deep Alternating Neural Network for Hyperspectral Unmixing
abstract
Hyperspectral unmixing is recognized as an important tool to learn the constituent materials and corresponding distribution in a scene. The physical spectral mixture model is always important to tackle this problem because of its highly ill-posed nature. In this article, we introduce a linear spectral mixture model (LMM)-based end-to-end deep neural network named SNMF-Net for hyperspectral unmixing. SNMF-Net shares an alternating architecture and benefits from both model-based methods and learning-based methods. On the one hand, SNMF-Net is of high physical interpretability as it is built by unrolling$L_{p}$sparsity constrained nonnegative matrix factorization ($L_{p}$-NMF) model belonging to LMM families. On the other hand, all the parameters and submodules of SNMF-Net can be seamlessly linked with the alternating optimization algorithm of$L_{p}$-NMF and unmixing problem. This enables us to reasonably integrate the prior knowledge on unmixing, the optimization algorithm, and the sparse representation theory into the network for robust learning, so as to improve unmixing. Experimental results on the synthetic and real-world data show the advantages of the proposed SNMF-Net over many state-of-the-art methods.
Fengchao Xiong, Jun Zhou 0001, Shuyin Tao, Jianfeng Lu 0003, Yuntao Qian
IEEE Trans. Geosci. Remote. Sens.5
2022 Learning a Deep Structural Subspace Across Hyperspectral Scenes With Cross-Domain VAE
abstract
Hyperspectral image (HSI) classification is a small-sample-size problem due to the expensive cost of labeling. As a novel approach to this problem, cross-scene HSI classification has become a hot research topic in recent years. In cross-scene HSI classification, the scene containing enough labeled samples (called source scene) is used to benefit the classification in another scene containing a small number of training samples (called target scene). Transfer learning is a typical solution for cross-scene classification. However, many transfer learning algorithms assume an identical feature space for source and target scenes, which violates the fact that source and target scenes often lie in different feature spaces with various dimensions due to different HSI sensors. Aiming at the different feature spaces between the two scenes, we propose an end-to-end heterogeneous deep transfer learning algorithm, namely, cross-domain variational autoencoder (CDVAE). This algorithm is mainly composed of two key parts: 1) the features of the two scenes are embedded into the shared feature subspace through the two-stream variational autoencoder (VAE) to ensure that the output feature dimensions of the two scenes are identical and 2) graph regularization is used to establish the manifold constraints between source and target scenes in the shared subspace, so as to align the feature spaces. Experiments on two different cross-scene HSI datasets have proved the superior performance of the proposed CDVAE algorithm.
Minchao Ye, Junbin Chen, Fengchao Xiong, Yuntao Qian
IEEE Trans. Geosci. Remote. Sens.4
2022 Spectral-Spatial Boundary Detection in Hyperspectral Images
abstract
In this paper, we propose a novel method for boundary detection in close-range hyperspectral images. This method can effectively predict the boundaries of objects of similar colour but different materials. To effectively extract the material information in the image, the spatial distribution of the spectral responses of different materials or endmembers is first estimated by hyperspectral unmixing. The resulting abundance map represents the fraction of each endmember spectra at each pixel. The abundance map is used as a supportive feature such that the spectral signature and the abundance vector for each pixel are fused to form a new spectral feature vector. Then different spectral similarity measures are adopted to construct a sparse spectral-spatial affinity matrix that characterizes the similarity between the spectral feature vectors of neighbouring pixels within a local neighborhood. After that, a spectral clustering method is adopted to produce eigenimages. Finally, the boundary map is constructed from the most informative eigenimages. We created a new HSI dataset and use it to compare the proposed method with four alternative methods, one for hyperspectral image and three for RGB image. The results exhibit that our method outperforms the alternatives and can cope with several scenarios that methods based on colour images cannot handle.
Suhad Lateef Al-Khafaji, Jun Zhou 0001, Xiao Bai 0001, Yuntao Qian, Alan Wee-Chung Liew
IEEE Trans. Image Process.4
2022 Contrast-Reconstruction Representation Learning for Self-Supervised Skeleton-Based Action Recognition
abstract
Skeleton-based action recognition is widely used in varied areas, e.g., surveillance and human-machine interaction. Existing models are mainly learned in a supervised manner, thus heavily depending on large-scale labeled data, which could be infeasible when labels are prohibitively expensive. In this paper, we propose a novel Contrast-Reconstruction Representation Learning network (CRRL) that simultaneously captures postures and motion dynamics for unsupervised skeleton-based action recognition. It consists of three parts: Sequence Reconstructor (SER), Contrastive Motion Learner (CML), and Information Fuser (INF). SER learns representation from skeleton coordinate sequence via reconstruction. However the learned representation tends to focus on trivial postural coordinates and be hesitant in motion learning. To enhance the learning of motions, CML performs contrastive learning between the representation learned from coordinate sequences and additional velocity sequences, respectively. Finally, in the INF module, we explore varied strategies to combine SER and CML, and propose to couple postures and motions via a knowledge-distillation based fusion strategy which transfers the motion learning from CML to SER. Experimental results on several benchmarks, i.e., NTU RGB+D 60/120, PKU-MMD, CMU, and NW-UCLA, demonstrate the promise of the our method by outperforming state-of-the-art approaches.
Peng Wang 0100, Jun Wen 0001, Chenyang Si, Yuntao Qian, Liang Wang 0001
IEEE Trans. Image Process.4
2022 SMDS-Net: Model Guided Spectral-Spatial Network for Hyperspectral Image Denoising
abstract
Deep learning (DL) based hyperspectral images (HSIs) denoising approaches directly learn the nonlinear mapping between noisy and clean HSI pairs. They usually do not consider the physical characteristics of HSIs. This drawback makes the models lack interpretability that is key to understanding their denoising mechanism and limits their denoising ability. In this paper, we introduce a novel model-guided interpretable network for HSI denoising to tackle this problem. Fully considering the spatial redundancy, spectral low-rankness, and spectral-spatial correlations of HSIs, we first establish a subspace-based multidimensional sparse (SMDS) model under the umbrella of tensor notation. After that, the model is unfolded into an end-to-end network named SMDS-Net, whose fundamental modules are seamlessly connected with the denoising procedure and optimization of the SMDS model. This makes SMDS-Net convey clear physical meanings, i.e., learning the low-rankness and sparsity of HSIs. Finally, all key variables are obtained by discriminative training. Extensive experiments and comprehensive analysis on synthetic and real-world HSIs confirm the strong denoising ability, strong learning capability, promising generalization ability, and high interpretability of SMDS-Net against the state-of-the-art HSI denoising methods. The source code and data of this article will be made publicly available at https://github.com/bearshng/smds-net for reproducible research.
Fengchao Xiong, Jun Zhou 0001, Shuyin Tao, Jianfeng Lu 0003, Jiantao Zhou 0001, Yuntao Qian
IEEE Trans. Image Process.6
2022 DACHA: A Dual Graph Convolution Based Temporal Knowledge Graph Representation Learning Method Using Historical Relation
abstract
Temporal knowledge graph (TKG) representation learning embeds relations and entities into a continuous low-dimensional vector space by incorporating temporal information. Latest studies mainly aim at learning entity representations by modeling entity interactions from the neighbor structure of the graph. However, the interactions of relations from the neighbor structure of the graph are neglected, which are also of significance for learning informative representations. In addition, there still lacks an effective historical relation encoder to model the multi-range temporal dependencies. In this article, we propose a d ual gr a ph c onvolution network based TKG representation learning method using h istorical rel a tions (DACHA). Specifically, we first construct the primal graph according to historical relations, as well as the edge graph by regarding historical relations as nodes. Then, we employ the dual graph convolution network to capture the interactions of both entities and historical relations from the neighbor structure of the graph. In addition, the temporal self-attentive historical relation encoder is proposed to explicitly model both local and global temporal dependencies. Extensive experiments on two event based TKG datasets demonstrate that DACHA achieves the state-of-the-art results.
Ling Chen 0001, Xing Tang 0006, Yuntao Qian, Yansheng Li 0001, Yongjun Zhang 0002
ACM Trans. Knowl. Discov. Data4
2021 NMF-SAE: An Interpretable Sparse Autoencoder for Hyperspectral Unmixing
abstract
Hyperspectral unmixing is an important tool to learn the material constitution and distribution of a scene. Model-based unmixing methods depend on well-designed iterative optimization algorithms, which is usually time consuming. Learning-based methods perform unmixing in a data-driven manner but heavily rely on the quality and quantity of the training samples due to the lack of physical interpretability. In this paper, we combine the advantages of both model-based and learning-based methods and propose a nonnegative matrix factorization (NMF) inspired sparse autoencoder (NMF-SAE) for hyperspectral unmixing. NMF-SAE consists of an encoder and a decoder, both of which are constructed by unrolling the iterative optimization rules of L1sparsity-constrained NMF for the linear spectral mixture model. All parameters in our method are obtained by end-to-end training in a data-driven manner. Our network is not only physically interpretable and flexible but also has higher learning capacity with fewer parameters. Experimental results on both synthetic and real-world data demonstrate that our method is capable of producing desirable unmixing results when compared against several alternative approaches.
Fengchao Xiong, Jun Zhou 0001, Minchao Ye, Jianfeng Lu 0003, Yuntao Qian
ICASSP5
2021 Hierarchical Multi-Label Ship Recognition in Remote Sensing Images Using Label Relation Graphs
abstract
Hierarchical multi-label classification (HMC) aims to assign multiple labels to every instance with the labels organized under hierarchical relations. In the application of ship recognition in remote sensing images, a ship can own coarse-to-fine hierarchical labels, e.g., the military ship, aircraft carrier, and nimitz class aircraft carrier. In this paper, we propose to combine two forms of loss functions to solve the HMC problem based on the neural network. The first probabilistic classification loss is to encode the hierarchical knowledge by introducing hierarchy and exclusion (HEX) graphs to impose constraints on hierarchical labels. The second cross-entropy loss imposes the softmax normalization on leaf nodes in the hierarchy to discriminate fine-grained classes. We evaluate our method on the high resolution satellite image dataset for ship recognition (HRSC), in which hierarchical labels are organized as the three-level tree. The proposed method shows comparative results compared to state-of-art HMC models.
Jingzhou Chen, Yuntao Qian
IGARSS2
2021 Graph Regularized Autoencoder Based Feature Extraction for Hyperspectral Image Classification
abstract
We present a novel stacked autoencoder framework for feature extraction to improve classification of hyperspectral image, leveraging graph regularization to address the shortcomings of classical autoencoder that mainly focuses on learning spectral features. In the proposed method, we firstly construct a graph to represent the spectral-spatial similarity between pixels in a hyperspectral image by measuring their spatial and spectral distances. And then the graph regularized autoencoder is learned to transform the original spectral signatures of pixels into a new feature space used for the downstream pixel classification or other tasks. Our feature extraction method can preserve the intrinsic spectral-spatial distribution in a hyperspectral image and obtain more discriminative and robust features. The experiments on pixel classification show the competitive performance compared with classical autoencoder based and manifold learning based feature extraction approaches.
Xiaotian Fan, Jingzhou Chen, Yuntao Qian
IGARSS3
2021 Learning a Model-Based Deep Hyperspectral Denoiser from a Single Noisy Hyperspectral Image
abstract
Hyperspectral image (HSI) denoising is a crucial preprocessing procedure to improve the quality of HSI. Model-based methods take the degradation model and the structure of underlying clean HSI into account for denoising but require a large number of numerical iterations and exhausting parameter tuning. Deep-learning-based (DL-based) methods directly learn the nonlinear transformation of clean and noisy image HSI pairs, but rely on large-scale high-quality training samples because of its “black box” denoising mechanism. In this paper, we propose a model-based DL method for HSI denoising to combine the advantages of model-based methods and DL-based methods. Specifically, we first build a HSI denoising model based on sparse representation. Then, we unfold the iterative optimization under the framework of gradient descent with momentum to yield a Gradient Momentum Sparse Coding Network (GMSC-Net) for denoising. In order to overcome the unavailability of noisy-clean HSI pairs for training, we directly learn GMSC-Net from a single HSI. The observed noisy HSI is grouped into a number of clusters containing local cubes. The cluster centers are treated as “clean” cubes and are polluted by noises, yielding a set of “noisy-clean” pairs for training. Extensive experiments show the effectiveness of our method on both synthetic and real-world datasets.
Guanyiman Fu, Fengchao Xiong, Shuyin Tao, Jianfeng Lu 0003, Jun Zhou 0001, Yuntao Qian
IGARSS6
2021 Multi-stage stochastic gradient method with momentum acceleration
Zhijian Luo, Siyu Chen 0002, Yuntao Qian, Yueen Hou
Signal Process.3
2020 BAE-Net: A Band Attention Aware Ensemble Network for Hyperspectral Object Tracking
abstract
Hyperspectral videos contain images with a large number of light wavelength indexed bands that can facilitate material identification for object tracking. Most hyperspectral trackers use hand-crafted features rather than deep learning generated features for image representation due to limited training samples. To fill this gap, this paper introduces a band attention aware ensemble network (BAE-Net) for deep hyperspectral object tracking, which takes advantages of deep models trained on color videos for feature representation. Specifically, an autoencoder-like band attention block is introduced to learn the dependencies among bands and generate band-wise weights. Guided by these weights, hyperspectral images are then divided into a number of three-channel images. These three-channel images are fed into a deep color tracking network, producing several weak trackers. Finally, weak trackers are fused using ensemble learning for target location. Experimental results on hyperspectral datasets show the effectiveness and advantages of the proposed deep hyperspectral tracker.
Zhuanfeng Li, Fengchao Xiong, Jun Zhou 0001, Jing Wang 0062, Jianfeng Lu 0003, Yuntao Qian
ICIP6
2020 Few-Shot Learning for Remote Sensing Image Retrieval With MAML
abstract
Few-shot remote sensing image retrieval is devoted to add new retrieval categories with a small number of labeled samples, and simultaneously achieve favorable retrieval performance for new categories and keep the primary retrieval performance for the original categories as far as possible. Few-shot learning has received considerable attention, however its applications in image retrieval, especially for remote sensing image retrieval, are still very few. In this paper, we redefine the few-shot image retrieval problem formally and further propose a few-shot retrieval method under model-agnostic meta-learning (MAML) framework, combined with ResNet and GeM as feature extraction module. Moreover, the optimal mean average precision (mAP) is used as ranking loss for defining the loss function of learning model, and a histogram binning approximation of mAP, which is differential, is thus employed so that the whole few-shot retrieval model can be end-to-end trained. The retrieval effectiveness and efficiency of our method are verified on UC-Merced and AID data sets.
Qian Zhong, Yuntao Qian
ICIP3
2020 Multi-Label Remote Sensing Image Classification with Deformable Convolutions and Graph Neural Networks
abstract
Multi-label remote sensing image classification is a significant yet difficult task due to intra-class variations and label dependencies among land-cover classes. In this paper, we propose a novel multi -label classification model based on deformable convolutions and graph neural networks. Specifically, we first use deformable convolutions to learn image features with geometric transformation invariance and adaptive receptive field. Then we adopt attention mechanism to extract label-related image features. After that, a directed graph is constructed to model the label dependencies, and the label-related features are fused through graph propagation mechanisms. Experiments on UC-Merced and DOTA data sets demonstrate its effectiveness.
Yingyu Diao, Jingzhou Chen, Yuntao Qian
IGARSS3
2020 Nonlocal Low-Rank Nonnegative Tensor Factorization for Hyperspectral Unmixing
abstract
Hyperspectral unmixing decomposes hyperspectral images (HSI) into a collection of constituent materials or end-members and their fractions, i.e., abundances. Nonnegative tensor factorization (NTF) has been utilized thanks to its ability of preserving all the information in HSI. However, NTF based unmixing only makes use of global spatial-spectral information without considering detailed local/non-local spatial information, making it vulnerable to real-world disturbance such as noises. To this end, in this paper, we extend NTF by introducing non-local low-rank constraint to abundance maps. The additional regularization on abundances facilities tensor factorization avoid being trapped into a large number of suspicious solutions, so as to preserve the non-local spatial structure on abundance maps. Experimental results on synthetic data and real-world data show that the proposed method outperforms the state-of-the-art methods.
Fengchao Xiong, Kun Qian 0015, Jianfeng Lu 0003, Jun Zhou 0001, Yuntao Qian
IGARSS5
2020 Residual deep PCA-based feature extraction for hyperspectral image classification
Minchao Ye, Chenxi Ji, Ling Lei 0002, Huijuan Lu, Yuntao Qian
Neural Comput. Appl.6
2020 Spectral Mixture Model Inspired Network Architectures for Hyperspectral Unmixing
abstract
In many statistical hyperspectral unmixing approaches, the unmixing task is essentially an optimization problem given a defined linear or nonlinear spectral mixture model. However, most of the model inference algorithms require a time-consuming iterative procedure. On the other hand, neural networks have been recently used to estimate abundances given some training samples, or directly estimate endmembers and abundances simultaneously in an unsupervised setting. However, their disadvantages are clear: lack of interpretability and reliance on the large training set. Model-inspired neural networks are constructed by the problem model and its corresponding inference algorithm. It incorporates the prior knowledge of physical model and algorithm into network architecture, combining the advantages of model-based and learning-based methods. This article deeply unfolds the linear mixture model and the corresponding iterative shrinkage-thresholding algorithm (ISTA) to build two unmixing network architectures. The first assumes that the set of endmembers are known, and the deep unfolded ISTA model is only for abundance estimation; and the second is used for blind unmixing to estimate both endmembers and abundances at the same time. The networks can be trained by supervised and unsupervised schemes, respectively, with a small-size training set, and then, unmixing becomes a feedforward process, which is very fast since no iteration is required. The experimental results show their competitive performance compared with the state-of-the-art unmixing approaches.
Yuntao Qian, Fengchao Xiong, Qipeng Qian, Jun Zhou 0001
IEEE Trans. Geosci. Remote. Sens.1
2020 Material Based Object Tracking in Hyperspectral Videos
abstract
Traditional color images only depict color intensities in red, green and blue channels, often making object trackers fail in challenging scenarios, e.g., background clutter and rapid changes of target appearance. Alternatively, material information of targets contained in large amount of bands of hyperspectral images (HSI) is more robust to these difficult conditions. In this paper, we conduct a comprehensive study on how material information can be utilized to boost object tracking from three aspects: dataset, material feature representation and material based tracking. In terms of dataset, we construct a dataset of fully-annotated videos, which contain both hyperspectral and color sequences of the same scene. Material information is represented by spectral-spatial histogram of multidimensional gradients, which describes the 3D local spectral-spatial structure in an HSI, and fractional abundances of constituted material components which encode the underlying material distribution. These two types of features are embedded into correlation filters, yielding material based tracking. Experimental results on the collected dataset show the potentials and advantages of material based object tracking.
Fengchao Xiong, Jun Zhou 0001, Yuntao Qian
IEEE Trans. Image Process.3
2019 Dual Dictionary Learning for Mining a Unified Feature Subspace between Different Hyperspectral Image Scenes
abstract
In real-world applications of hyperspectral images (HSIs), we may frequently face the following situation: two HSI scenes (named source and target scenes, respectively) contain similar land over objects, but they are captured in different spots or at different time. Even if they are captured by the same hyper-spectral sensor, there exist spectral shift between them. In our previous work, we tried to solve the spectral shift by dictionary sharing or domain-invariant feature selection. However, a more regular case is that two similar HSI scenes are captured by different hyperspectral sensors. How to mine the relationship between such HSIs is a more challenging problem, since the feature spaces are totally different. A natural approach is to learn a unified low-dimensional feature subspace which can bridge the two HSI scenes. In this work, we propose a dual dictionary nonnegative matrix factorization (DDNMF) algorithm for the aforementioned goal. In details, an individual domain-specific dictionary is learned for each scene, and two dictionary learning tasks (for source and target scenes) are coupled by manifold regularization, ensuring that pixels belonging to a same land cover class have similar representations over the learned dictionaries, even if they come from different scenes. Experimental results show that the proposed algorithm can indeed mine a unified feature subspace shared between two different HSI scenes.
Minchao Ye, Huijuan Lu, Ling Lei 0002, Yuntao Qian
IGARSS5
2019 Feature Extraction of Hyperspectral Imagery Based on Deep NMF
abstract
Feature extraction is an important research topic in hyper-spectral image (HSI) classification. However, most of feature extraction methods only extract low-level features, which makes them not perform well in the applications of HSI. In this paper, we have proposed a non-negative matrix factorization (NMF) based deep feature extraction algorithm, namely deep NMF. Deep NMF tries to construct a deep feature representation by cascading multiple NMFs. Reconstruction residual of NMF is passed layer by layer to reduce information loss. Meanwhile, passing residuals between layers can construct a feature hierarchy from coarse to fine. Furthermore, activation functions are applied between adjacent layers to enhance the ability of non-linear feature extraction. Experimental results have also shown that our algorithm is computationally efficient and effective for HSI classification.
Chenxi Ji, Minchao Ye, Huijuan Lu, Futian Yao, Yuntao Qian
IGARSS5
2019 Spectral-Spatial Joint Noise Estimation for Hyperspectral Images
abstract
Hyperspectral images (HSIs) are always corrupted by noise, which will strongly affect the applications. Various denoising algorithms have been proposed for HSIs. Most existing denoising methods have parameters related to intensity of noise. So noise estimation is an essential step in HSI denoising. In our previous work, we have proposed a homogeneous region based noise estimation algorithm. However, we find it often fails on severely corrupted bands. To solve the problem, two improvements are made in this work: 1) depending on strong correlations between bands, a regression based signal-noise separation is adopted; 2) utilizing the identical spatial structure of different bands, a unified homogeneous region segmentation is performed across all bands via clustering of spectral vectors. Then, noise estimation is done using curve fitting according to the segmented homogeneous regions with separated signal and noise components. By this spectral-spatial joint approach, we have significantly improved the accuracy of noise estimation.
Minchao Ye, Chenxi Ji, Ling Lei 0002, Yuntao Qian
IGARSS5
2019 Paper evolution graph: multi-view structural retrieval for academic literature
abstract
Academic literature retrieval concerns about the selection of papers that are most likely to match a user’s information needs. Most of the retrieval systems are limited to list-output models, in which the retrieval results are isolated from each other. In this paper, we aim to uncover the relationships between the retrieval results and propose a method to build structural retrieval results for academic literature, which we call a paper evolution graph (PEG). The PEG describes the evolution of diverse aspects of input queries through several evolution chains of papers. By using the author, citation, and content information, PEGs can uncover various underlying relationships among the papers and present the evolution of articles from multiple viewpoints. Our system supports three types of input queries: keyword query, single-paper query, and two-paper query. The construction of a PEG consists mainly of three steps. First, the papers are soft-clustered into communities via metagraph factorization, during which the topic distribution of each paper is obtained. Second, topically cohesive evolution chains are extracted from the communities that are relevant to the query. Each chain focuses on one aspect of the query. Finally, the extracted chains are combined to generate a PEG, which fully covers all the topics of the query. Experimental results on a real-world dataset demonstrate that the proposed method can construct meaningful PEGs.
Danping Liao, Yuntao Qian
Frontiers Inf. Technol. Electron. Eng.2
2019 Hyperspectral Unmixing via Total Variation Regularized Nonnegative Tensor Factorization
abstract
Hyperspectral unmixing decomposes a hyperspectral imagery (HSI) into a number of constituent materials and associated proportions. Recently, nonnegative tensor factorization (NTF)-based methods have been proposed for hyperspectral unmixing thanks to their capability in representing an HSI without any information loss. However, tensor factorization-based HSI processing approaches often suffer from low-signal-to-noise ratio condition of HSI and nonuniqueness of the solution. This problem can be effectively alleviated by introducing various spatial constraints into tensor factorization to suppress the noise and decrease the number of extreme, stationary, and saddle points. On the other hand, total variation (TV) adaptively promotes piecewise smoothness while preserving edges. In this paper, we propose a TV regularized matrix-vector NTF method. It takes advantage of tensor factorization in preserving global spectral-spatial information and the merits of TV in exploiting local spatial information, thus generating smooth abundance maps with preserved edges. Experimental results on synthetic and real-world data show that the proposed method outperforms the state-of-the-art methods.
Fengchao Xiong, Yuntao Qian, Jun Zhou 0001, Yuan Yan Tang
IEEE Trans. Geosci. Remote. Sens.2
2019 Hyperspectral Restoration via L0 Gradient Regularized Low-Rank Tensor Factorization
abstract
Due to the mechanism of the data acquisition process, hyperspectral imagery (HSI) are usually contaminated by various noises, e.g., Gaussian noise, impulse noise, strips, and dead lines. In this article, a spectral-spatial L0gradient regularized low-rank tensor factorization (LRTFL0) method is proposed for hyperspectral denoising, in which the restored HSI is approximated by low-rank block term decomposition (BTD). BTD factorizes a tensor into the sum of a series of component tensors, each of which is represented by the outer product of a matrix and a vector. From subspace learning point of view, the vector and matrix can be considered as a spectral atom and its corresponding coding coefficients. In the proposed method, the correlations in both spectral and spatial domains are taken into account via the small size of atom set and low-rankness of coding matrices. In addition, HSIs also have the local structure of piecewise smoothness in both spectral and spatial domains. Motivated by the supreme virtues of L0gradient regularization in image structure exploitation, we develop a spectral-spatial L0gradient regularization and embed it into BTD to explore the spectral-spatial texture information. The proposed method can simultaneously remove various types of noises, and the experimental results on both synthetic data and real-world data show its superiority when compared with several state-of-the-art approaches.
Fengchao Xiong, Jun Zhou 0001, Yuntao Qian
IEEE Trans. Geosci. Remote. Sens.3
2018 Hyperspectral Imagery Denoising via Reweighed Sparse Low-Rank Nonnegative Tensor Factorization
abstract
Hyperspectral imagery (HSI) denoising is an important preprocessing step for real-world applications. Recently, sparse representation and low-rank representation based methods are proven effective in HSI denoising. However, most of these approaches only consider the low-rankness in the spectral domain and the sparsity in coding matrix. They have ignored the property that the coding matrix of each atom is also low-rank, i.e., low-rankness also exists in the spatial domain. In this paper, a reweighed sparse low-rank nonnegative tensor factorization (RSLRNTF) method is proposed to restore an HSI. It takes an HSI as a third-order tensor and factorizes it into the combination of a few component tensors where each one is the outer product of a low-rank matrix (coding matrix) and a vector (atom). Additionally, a reweighed L1 norm is added to coding matrices to enforce their sparsity. The low-rankness in both the spatial domain and the spectral domain as well as sparsity in the spatial domain improve the denoising performance. Furthermore, the nonnegativity in both coding matrices and dictionary leads to parts-based representation of HSI, which facilitates preserving local fine structure information. Experimental results on synthetic data and real-world data demonstrate the superiority of proposed method.
Fengchao Xiong, Jun Zhou 0001, Yuntao Qian
ICIP3
2018 Visualization of Hyperspectral Images Using Moving Least Squares
abstract
Displaying the large number of bands in a hyperspectral image (HSI) on a trichromatic monitor has been an active research topic. The visualized image shall convey as much information as possible from the original data and facilitate image interpretation. Most existing methods display HSIs in false colors, which contradict with human's experience and expectation. In this paper, we propose a nonlinear approach to visualize an input HSI with natural colors by taking advantage of a corresponding RGB image. Our approach is based on Moving Least Squares (MLS), an interpolation scheme for reconstructing a surface from a set of control points, which in our case is a set of matching pixels between the HSI and the corresponding RGB image. Based on MLS, the proposed method solves for each spectral signature a unique transformation so that the nonlinear structure of the HSI can be preserved. The matching pixels between a pair of HSI and RGB image can be reused to display other HSIs captured by the same imaging sensor with natural colors. Experiments show that the output images of the proposed method not only have natural colors but also maintain the visual information necessary for human analysis.
Danping Liao, Siyu Chen 0002, Yuntao Qian
ICPR3
2018 Deep Tensor Factorization for Hyperspectral Image Classification
abstract
High-dimensional spectral feature and limited training samples have caused a range of difficulties for hyperspectral image (HSI) classification. Feature extraction is effective to tackle this problem. Specifically, tensor factorization is superior to some prominent methods such as principle component analysis (PCA) and non-negative matrix factorization (NMF) because it takes spatial information into consideration. Recently, deep learning has gotten more and more attention for efficiently extracting hierarchical features for various tasks. In this paper, we propose a novel feature extraction method, deep tensor factorization (DTF), to extract hierarchical and meaningful features from observed HSI. This method takes advantage of tensor in representing HSI and the merits of convolutional neural network (CNN) in hierarchical feature extraction. Specifically, a convolution operation is firstly applied in the spectral dimension of HSI to suppress the effect of noise. Then, the convolved HSI is fed into tensor factorization to learn a low rank representation of data. After that, the above two process are repeated to learn a hierarchical representation of HSI. Experimental results on two real hyperspectral data sets show the superiority of the proposed method.
Jingzhou Chen, Yuntao Qian, Minchao Ye
IGARSS3
2018 Superpixel-Based Nonnegative Tensor Factorization for Hyperspectral Unmixing
abstract
Hyperspectral unmixing aims at decomposing a hyperspectral image (HSI) into a number of constituted materials and associated proportions. Recently, nonnegative tensor factorization (NTF) based methods have been proved effective and natural for hyperspectral unmixing owing to their virtue of representing an HSI without any information loss. However, these methods take an HSI as a whole, partly ignoring the local information in distinct local regions. In addition, HSIs are high likely to be disturbed by various noise, making the global information unnecessarily reliable. To alleviate these drawbacks, we propose a superpixel-based matrix-vector nonnegative tensor factorization (S-MV-NTF) method for hyperspectral unmixing, where both the global information and local information are taken into consideration. In this method, the HSI is firstly partitioned into numerous superpixels, homogeneous regions with adaptive sizes and compact boundaries, representing the local spatial structure information. Then, such local information is integrated to the tensor factorization to make the pixels lying in the same superpixel share similar abundances. Experimental results on synthetic data and real-world data show that the proposed method dominates the state-of-the-art methods.
Fengchao Xiong, Jingzhou Chen, Jun Zhou 0001, Yuntao Qian
IGARSS4
2018 Cross-Scene Feature Selection for Hyperspectral Images Based on Cross-Domain Information Gain
abstract
Feature selection is an important research topic for hyperspectral images (HSIs). It helps to remove the noisy or redundant features. Traditional feature selection algorithms are mostly performed within a single HSI scene (dataset). However, appearance of massive HSIs requires the feature selection problems to be considered across different HSI scenes, e.g., two HSI scenes obtained from different spots or at different time. In this case, the features are not identically distributed within two scenes due to spectral shift. To solve this problem, a cross-scene feature selection algorithm is proposed in this work for HSIs, which is based on cross-domain information gain (CDIG). The main motivation includes two factors, one is the discriminant of selected features to separate different land-cover classes, while the other is the consistency of the selected features between different scenes. Consequently, the proposed CDIG reaches a compromise between aforementioned two factors. Experimental results on two cross-scene HSI datasets show the advantages of the proposed CDIG in cross-scene feature selection problems.
Minchao Ye, Yongqiu Xu, Huijuan Lu, Ke Yan 0001, Yuntao Qian
IGARSS5
2018 Deconv R-CNN for Small Object Detection on Remote Sensing Images
abstract
Small object detection has drawn increasing interest in computer vision and remote sensing image processing. The Region Proposal Network (RPN) methods (e.g., Faster R-CNN) have obtained promising detection accuracy with several hundred proposals. However, due to the pooling layers in the network structure of the deep model, precise localization of small-size object is still a hard problem. In this paper, we design a network with a deconvolution layer after the last convolution layer of base network for small target detection. We call our model Deconv R-CNN. In the experiment on a remote sensing image dataset, Deconv R-CNN reaches a much higher mean average precision (mAP) than Faster R-CNN.
Sophanyouly Thachan, Jingzhou Chen, Yuntao Qian
IGARSS5
2018 Spectral Image Visualization Using Generative Adversarial Networks
Siyu Chen 0002, Danping Liao, Yuntao Qian
PRICAI (1)3
2018 Corrections to "Dictionary Learning-Based Feature-Level Domain Adaptation for Cross-Scene Hyperspectral Image Classification"
abstract
In the above paper[1], there is an error inFig. 14.Fig. 14should include$3\times3$matrices rather than$7\times7$, since the Shanghai-Hangzhou dataset has three land-cover classes. The corrected figure appears here.
Minchao Ye, Yuntao Qian, Jun Zhou 0001, Yuan Yan Tang
IEEE Trans. Geosci. Remote. Sens.2
2017 On the Sampling Strategy for Evaluation of Spectral-Spatial Methods in Hyperspectral Image Classification
abstract
Spectral-spatial processing has been increasingly explored in remote sensing hyperspectral image classification. While extensive studies have focused on developing methods to improve the classification accuracy, experimental setting and design for method evaluation have drawn little attention. In the scope of supervised classification, we find that traditional experimental designs for spectral processing are often improperly used in the spectral-spatial processing context, leading to unfair or biased performance evaluation. This is especially the case when training and testing samples are randomly drawn from the same image - a practice that has been commonly adopted in the experiments. Under such setting, the dependence caused by overlap between the training and testing samples may be artificially enhanced by some spatial information processing methods, such as spatial filtering and morphological operation. Such enhancement of dependence in return amplifies the classification accuracy, leading to an improper evaluation of spectral-spatial classification techniques. Therefore, the widely adopted pixel-based random sampling strategy is not always suitable to evaluate spectral-spatial classification algorithms, because it is difficult to determine whether the improvement of classification accuracy is caused by incorporating spatial information into classifier or by increasing the overlap between training and testing samples. To tackle this problem, we propose a novel controlled random sampling strategy for spectral-spatial methods. It can greatly reduce the overlap between training and testing samples and provides more objective and accurate evaluation.
Jie Liang 0003, Jun Zhou 0001, Yuntao Qian, Lian Wen, Xiao Bai 0001, Yongsheng Gao 0001
IEEE Trans. Geosci. Remote. Sens.3
2017 Matrix-Vector Nonnegative Tensor Factorization for Blind Unmixing of Hyperspectral Imagery
abstract
Many spectral unmixing approaches ranging from geometry, algebra to statistics have been proposed, in which nonnegative matrix factorization (NMF)-based ones form an important family. The original NMF-based unmixing algorithm loses the spectral and spatial information between mixed pixels when stacking the spectral responses of the pixels into an observed matrix. Therefore, various constrained NMF methods are developed to impose spectral structure, spatial structure, and spectral-spatial joint structure into NMF to enforce the estimated endmembers and abundances preserve these structures. Compared with matrix format, the third-order tensor is more natural to represent a hyperspectral data cube as a whole, by which the intrinsic structure of hyperspectral imagery can be losslessly retained. Extended from NMF-based methods, a matrix-vector nonnegative tensor factorization (NTF) model is proposed in this paper for spectral unmixing. Different from widely used tensor factorization models, such as canonical polyadic decomposition CPD) and Tucker decomposition, the proposed method is derived from block term decomposition, which is a combination of CPD and Tucker decomposition. This leads to a more flexible frame to model various application-dependent problems. The matrix-vector NTF decomposes a third-order tensor into the sum of several component tensors, with each component tensor being the outer product of a vector (endmember) and a matrix (corresponding abundances). From a formal perspective, this tensor decomposition is consistent with linear spectral mixture model. From an informative perspective, the structures within spatial domain, within spectral domain, and cross spectral-spatial domain are retreated interdependently. Experiments demonstrate that the proposed method has outperformed several state-of-the-art NMF-based unmixing methods.
Yuntao Qian, Fengchao Xiong, Shan Zeng, Jun Zhou 0001, Yuan Yan Tang
IEEE Trans. Geosci. Remote. Sens.1
2017 Dictionary Learning-Based Feature-Level Domain Adaptation for Cross-Scene Hyperspectral Image Classification
abstract
A big challenge of hyperspectral image (HSI) classification is the small size of labeled pixels for training classifier. In real remote sensing applications, we always face the situation that an HSI scene is not labeled at all, or is with very limited number of labeled pixels, but we have sufficient labeled pixels in another HSI scene with the similar land cover classes. In this paper, we try to classify an HSI scene containing no labeled sample or only a few labeled samples with the help of a similar HSI scene having a relative large size of labeled samples. The former scene is defined as the target scene, while the latter one is the source scene. We name this classification problem as cross-scene classification. The main challenge of cross-scene classification is spectral shift, i.e., even for the same class in different scenes, their spectral distributions maybe have significant deviation. As all or most training samples are drawn from the source scene, while the prediction is performed in the target scene, the difference in spectral distribution would greatly deteriorate the classification performance. To solve this problem, we propose a dictionary learning-based feature-level domain adaptation technique, which aligns the spectral distributions between source and target scenes by projecting their spectral features into a shared low-dimensional embedding space by multitask dictionary learning. The basis atoms in the learned dictionary represent the common spectral components, which span a cross-scene feature space to minimize the effect of spectral shift. After the HSIs of two scenes are transformed into the shared space, any traditional HSI classification approach can be used. In this paper, sparse logistic regression (SRL) is selected as the classifier. Especially, if there are a few labeled pixels in the target domain, multitask SRL is used to further promote the classification performance. The experimental results on synthetic and real HSIs show the advantages of the proposed method for cross-scene classification.
Minchao Ye, Yuntao Qian, Jun Zhou 0001, Yuan Yan Tang
IEEE Trans. Geosci. Remote. Sens.2
2016 Semisupervised manifold learning for color transfer between multiview images
abstract
In multiview image stitching, the colors of images in a scene might vary when images are taken under different illumination or camera settings. A common way to produce a seamless stitched image is to transform the colors of a target image to match that of a source image. In this paper we present a color transfer method based on two premises: first, pixels in the generated image should have similar colors with their corresponding pixels in the source image. Second, pixels with similar colors should still have similar colors after color transfer. Our method can be considered as a semisupervised manifold learning approach, where the corresponding pixels of the input images serve as the labeled data. Our goal is to learn a final image which not only shares the same colors with the source image but also has the same image structure with the target image. While manifold learning methods aim to find an embedded space to represent the data with minimum structure loss, the proposed method further constrains the solution space using the labeled data. This paper introduces a parametric linear method and a nonparametric nonlinear method to tackle different types of color changes. Experimental results show the effectiveness of our methods both quantitatively and qualitatively.
Danping Liao, Yuntao Qian, Ze-Nian Li
ICPR2
2016 Bound analysis of natural gradient descent in stochastic optimization setting
abstract
Natural gradient descent is a metric aware optimization algorithm which utilizes an underlying Riemannian parameter space, and has successfully improved performance in statistical asymptotic and experimental point of view. In this paper, we investigate the bound property of natural gradient descent in stochastic optimization setting. The bound property is analyzed in both direct and indirect ways. Substituting natural gradient for vanilla gradient is considered as the direct analysis. In this way, we analyze the bound of natural gradient descent method by convergence analysis technique. Afterwards, the bound is analyzed in an indirect way by introducing mirror gradient according to its equivalence to natural gradient. Employing mirror gradient in bound analysis makes the procedure of parameter update more intuitive. We finally present experimental results to support our theoretical findings.
Zhijian Luo, Danping Liao, Yuntao Qian
ICPR3
2016 Parallel adaptive sparsity-constrained NMF algorithm for hyperspectral unmixing
abstract
Sparsity-constrained Nonnegative matrix factorization (NMF) has been proved to be an effective method for hyperspectral unmixing. However, the optimization procedure of sparsity-constrained NMF is computational demanding, which may limit its application in time-constrained conditions. In this paper, a parallel L1/2sparsity-constrained NMF unmixing method on Graphics Processing Units (GPUs) is proposed, and implemented using the Compute Unified Device Architecture (CUDA). It mainly involves the parallelization of multiplicative update rule for endmembers extraction and half thresholding update rule with adaptive regularization parameter strategy for abundance estimation. In particular, the concurrent kernel computation power of modern GPUs is employed to overlap the separated subtasks. The experiment results on the synthetic and real hyperspectral data demonstrate the effectiveness of our implementation.
Wenhong Wang 0004, Yuntao Qian
IGARSS2
2016 Kernel based sparse NMF algorithm for hyperspectral unmixing
abstract
Nonlinear unmixing methods for hyperspectral image (HSI) have attracted increasing interests since they can overcome the inherent limitations of the linear ones. In this paper, we regard the spectral unmixing as a blind source separation problem in the feature space, and develop a novel kernel based nonnegative matrix factorization method to estimate the endmembers and the abundances simultaneously. Moreover, the abundance sparsity property of HSI data in the feature space is exploited to promote more accurate unmixing results. Experiments are conducted on real HSI data with intimate mixture of minerals and the unmixing performance is compared with the state-of-the-art methods to demonstrate the effectiveness of the proposed method.
Wenhong Wang 0004, Yuntao Qian
IGARSS2
2016 Group sparse nonnegative matrix factorization for hyperspectral image denoising
abstract
Hyperspectral image (HSI) denoising is a significant preprocessing step to improve the performance of subsequent applications. Recently, HSI denoising methods using low rank representation and sparse coding have attracted much attention. In the HSI, there exists strong local correlations between spectral signatures within each full-band patch (FBP), i.e., the subcube containing the same area of all spectral bands, which suggests that spectral signatures within a clean FBP can be represented by a small number of bases. Denoising nonlocal similar FBPs jointly is beneficial as extra structure information is brought by the spatial self-similarity. However, there may exist variations among nonlocal similar FBPs and these variations need to be considered. Therefore, we propose a novel HSI denoising method based on group sparse nonnegative matrix factorization (GSNMF). In GSNMF, spectral signatures from nonlocal similar FBPs are assumed to be represented by a small number of bases and the coefficients of linear combination are sparse in nature. With the group s-parse regularization term, spectral signatures within an FBP share a common set of bases for reconstruction, indicating the strong local correlation. Spectral signatures across different nonlocal similar FBPs partially share a set of bases, which means that each of them may remain some non-shared bases. Thus, both nonlocal correlation and variation is considered. The effectiveness of the proposed method is validated in both synthetic and real hyperspectral datasets.
Yuntao Qian
IGARSS2
2016 Preference transfer model in collaborative filtering for implicit data
abstract
Generally, predicting whether an item will be liked or disliked by active users, and how much an item will be liked, is a main task of collaborative filtering systems or recommender systems. Recently, predicting most likely bought items for a target user, which is a subproblem of the rank problem of collaborative filtering, became an important task in collaborative filtering. Traditionally, the prediction uses the user item co-occurrence data based on users’ buying behaviors. However, it is challenging to achieve good prediction performance using traditional methods based on single domain information due to the extreme sparsity of the buying matrix. In this paper, we propose a novel method called the preference transfer model for effective cross-domain collaborative filtering. Based on the preference transfer model, a common basis item-factor matrix and different user-factor matrices are factorized. Each user-factor matrix can be viewed as user preference in terms of browsing behavior or buying behavior. Then, two factor-user matrices can be used to construct a so-called ‘preference dictionary’ that can discover in advance the consistent preference of users, from their browsing behaviors to their buying behaviors. Experimental results demonstrate that the proposed preference transfer model outperforms the other methods on the Alibaba Tmall data set provided by the Alibaba Group.
Bin Ju, Yuntao Qian, Minchao Ye
Frontiers Inf. Technol. Electron. Eng.2
2016 A Manifold Alignment Approach for Hyperspectral Image Visualization With Natural Color
abstract
The trichromatic visualization of hundreds of bands in a hyperspectral image (HSI) has been an active research topic. The visualized image shall convey as much information as possible from the original data and facilitate easy image interpretation. However, most existing methods display HSIs in false color, which contradicts with user experience and expectation. In this paper, we propose a new framework for visualizing an HSI with natural color by the fusion of an HSI and a high-resolution color image via manifold alignment. Manifold alignment projects several data sets to a shared embedding space where the matching points between them are pairwise aligned. The embedding space bridges the gap between the high-dimensional spectral space of the HSI and the RGB space of the color image, making it possible to transfer natural color and spatial information in the color image to the HSI. In this way, a visualized image with natural color distribution and fine spatial details can be generated. Another advantage of the proposed method is its flexible data setting for various scenarios. As our approach only needs to search a limited number of matching pixel pairs that present the same object, the HSI and the color image can be captured from the same or semantically similar sites. Moreover, the learned projection function from the hyperspectral data space to the RGB space can be directly applied to other HSIs acquired by the same sensor to achieve a quick overview. Our method is also able to visualize user-specified bands as natural color images, which is very helpful for users to scan bands of interest.
Danping Liao, Yuntao Qian, Jun Zhou 0001, Yuan Yan Tang
IEEE Trans. Geosci. Remote. Sens.2
2016 Nonnegative-Matrix-Factorization-Based Hyperspectral Unmixing With Partially Known Endmembers
abstract
Hyperspectral unmixing is an important technique for estimating fractions of various materials from remote sensing imagery. Most unmixing methods make the assumption that no prior knowledge of endmembers is available before the estimation. This is, however, not true for some unmixing tasks for which part of the endmember signatures may be known in advance. In this paper, we address the hyperspectral unmixing problem with partially known endmembers. We extend nonnegative-matrix-factorization-based unmixing algorithms to incorporate prior information into their models. The proposed approach uses the spectral signature of known endmembers as a constraint, among others, in the unmixing model, and propagates the knowledge by an optimization process which minimizes the difference between the image data and the prior knowledge. Results on both synthetic and real data have validated the effectiveness of the proposed method and have shown that it has outperformed several state-of-the-art methods that use or do not use prior knowledge of endmembers.
Jun Zhou 0001, Yuntao Qian, Xiao Bai 0001, Yongsheng Gao 0001
IEEE Trans. Geosci. Remote. Sens.3
2015 Local detail enhanced hyperspectral image visualization
abstract
A new method for hyperspectral image (HSI) visualization is proposed in this paper, which emphasizes on pairwise distance preservation and detail enhancement. It includes two sequential steps. The first reduces the high dimensional spectral space of HSI to 3-D color space by distance preservation method, and then enhancesthe detailed information by Laplacian pyramid. Distance preservation is an optimization problem that minimizes the difference between the pairwise pixel distances in the original spectral space and the corresponding color space. In general, solving this optimization problem is always very time and storage consuming. A multi-resolution multidimensional scaling algorithm is proposed in this paper to mitigate this hardness. Obviously the loss of some local details is not avoided in multidimensional scaling. In order to enhance the spatial distinction of different scene objects, Laplacian pyramid is used to draw the locally details from the original HSI, and embed it into the color image. The proposed HSI visualization method takes the global and local information of spectral and spatial distribution in HSI into account for visualization, which makes the color display of HSI carry as much original information as possible.
Jialu Fang, Yuntao Qian
IGARSS2
2015 Multitask Sparse Nonnegative Matrix Factorization for Joint Spectral-Spatial Hyperspectral Imagery Denoising
abstract
Hyperspectral imagery (HSI) denoising is a challenging problem because of the difficulty in preserving both spectral and spatial structures simultaneously. In recent years, sparse coding, among many methods dedicated to the problem, has attracted much attention and showed state-of-the-art performance. Due to the low-rank property of natural images, an assumption can be made that the latent clean signal is a linear combination of a minority of basis atoms in a dictionary, while the noise component is not. Based on this assumption, denoising can be explored as a sparse signal recovery task with the support of a dictionary. In this paper, we propose to solve the HSI denoising problem by sparse nonnegative matrix factorization (SNMF), which is an integrated model that combines parts-based dictionary learning and sparse coding. The noisy image is used as the training data to learn a dictionary, and sparse coding is used to recover the image based on this dictionary. Unlike most HSI denoising approaches, which treat each band image separately, we take the joint spectral-spatial structure of HSI into account. Inspired by multitask learning, a multitask SNMF (MTSNMF) method is developed, in which bandwise denoising is linked across the spectral domain by sharing a common coefficient matrix. The intrinsic image structures are treated differently but interdependently within the spatial and spectral domains, which allows the physical properties of the image in both spatial and spectral domains to be reflected in the denoising model. The experimental results show that MTSNMF has superior performance on both synthetic and real-world data compared with several other denoising methods.
Minchao Ye, Yuntao Qian, Jun Zhou 0001
IEEE Trans. Geosci. Remote. Sens.2
2015 Collaborative work with linear classifier and extreme learning machine for fast text categorization
Yuntao Qian
World Wide Web3
2014 Visualization of Hyperspectral Imaging Data Based on Manifold Alignment
abstract
Tristimulus display of the abundant information contained in a hyper spectral image is a challenging task. Previous visualization approaches focused on preserving as much information as possible in the reduced spectral space, but ended up with displaying hyper spectral images as false color images, which contradicts with human experience and expectation. This paper proposes a new framework to tackle this problem. It is based on the fusion of a hyper spectral image and a high-resolution color image via manifold alignment technique. Manifold learning is an important tool for dimension reduction. Manifold alignment projects a pair of two data sets into a common embedding space so that the pairs of corresponding points in these two data sets are pair wise aligned in this new space. Hyper spectral image and high-resolution color image have strong complementary properties due to the high spectral resolution in the former and the high spatial resolution in the latter. The embedding space produced by manifold alignment bridges a gap between the high dimensional spectral space of hyper spectral image and RGB space of color image, making it possible to transfer the natural color and spatial information of a high-resolution color image to a hyper spectral image to generate a visualized image with natural color distribution and finer details.
Danping Liao, Yuntao Qian, Jun Zhou 0001
ICPR2
2014 Speech enhancement with a GSC-like structure employing sparse coding
abstract
Speech communication is often influenced by various types of interfering signals. To improve the quality of the desired signal, a generalized sidelobe canceller (GSC), which uses a reference signal to estimate the interfering signal, is attracting attention of researchers. However, the interference suppression of GSC is limited since a little residual desired signal leaks into the reference signal. To overcome this problem, we use sparse coding to suppress the residual desired signal while preserving the reference signal. Sparse coding with the learned dictionary is usually used to reconstruct the desired signal. As the training samples of a desired signal for dictionary learning are not observable in the real environment, the reconstructed desired signal may contain a lot of residual interfering signal. In contrast, the training samples of the interfering signal during the absence of the desired signal for interferer dictionary learning can be achieved through voice activity detection (VAD). Since the reference signal of an interfering signal is coherent to the interferer dictionary, it can be well restructured by sparse coding, while the residual desired signal will be removed. The performance of GSC will be improved since the estimate of the interfering signal with the proposed reference signal is more accurate than ever. Simulation and experiments on a real acoustic environment show that our proposed method is effective in suppressing interfering signals.
Li-chun Yang, Yuntao Qian
J. Zhejiang Univ. Sci. C2
2013 Salient object detection in hyperspectral imagery
abstract
Object detection in hyperspectral images is an important task for many applications. While most traditional methods are pixel-based, many recent efforts have been put on extracting spatial-spectral features. In this paper, we introduce Itti's visual saliency model into the spectral domain for object detection. This enables the extraction of salient spectral features, which is related to the material property and spatial layout of objects, in the scale space. To our knowledge, this is the first attempt to combine hyperspectral data with salient object detection. Three methods have been implemented and compared to show how color component in the traditional saliency model can be replaced by spectral information. We have performed experiments on selected images from three online hyperspectral datasets, and show the effectiveness of the proposed methods.
Jie Liang 0003, Jun Zhou 0001, Xiao Bai 0001, Yuntao Qian
ICIP4
2013 Manifold alignment based color transfer for multiview image stitching
abstract
In multiview image stitching, color transfer removes all color inconsistences between different views under different illumination conditions and camera settings to make the stitching more seamless or visually acceptable. This paper presents a manifold alignment method to perform color transfer by exploring manifold structures of partially overlapped source and target images. Manifold alignment projects a pair of source and target images into a common embedding space in which not only the local geometries of color distribution in the respective images are preserved, but also the corresponding pixels in overlapped area across two images are pairwise aligned. Under this new space, color transfer can be considered as a matching problem between different manifolds, i.e. the color of each target pixel is replaced by the color of a source pixel that is nearest to this target pixel in this new space. Compared with other techniques in the literature, the proposed method makes full use of both the correspondences in overlapped area and the intrinsic color structures in the whole stitching scene so that a favorable performance is achieved.
Yuntao Qian, Danping Liao, Jun Zhou 0001
ICIP1
2013 Noise reduction of hyperspectral imagery based on nonlocal tensor factorization
abstract
Noise reduction for hyperspectral imagery (HSI) is an indispensable step before further processes such as object detection and classification. In this paper, we propose a noise reduction method for HSI based on non-local strategy and tensor factorization. Based on the observation that natural images are always locally self-repetitive, we divide the whole HSI into small sub-blocks and cluster similar blocks into groups. Since similar blocks share the same underlying structure, the redundancy can be utilized to remove noise of the blocks jointly. We stack the similar blocks to construct a fourth-order tensor from each group. Noise is reduced by finding the lower dimensional approximation of each of the fourth-order tensors via Tucker factorization. The experimental results indicate that the proposed method has a good quality of restoring the true signal from the noisy observation.
Danping Liao, Minchao Ye, Sen Jia 0001, Yuntao Qian
IGARSS4
2013 Visualization of hyperspectral imagery based on manifold learning
abstract
Displaying the abundant information contained in a hyperspectral image is a challenging task. Previous visualization approach focused only on preserving the structure in the original images. They ended up with presenting pseudo-color images and stopped short of adjusting the color of the images to retrieve more desirable visual effects. In this paper, a new visualization algorithm is proposed. It can be modeled as a two stage approach. At the first stage, Laplacian Eigenmaps algorithm is applied to reduce the dimension of the hyperspectral image. In this way we obtain a three dimensional image with pseudo-color. At the second stage, we transfer the natural color of a panchromatic image to the image obtained by the first step via manifold alignment. Experimental results show that the visualized image not only retains the structure of the hyperspectral image but also possesses natural colors.
Danping Liao, Minchao Ye, Sen Jia 0001, Yuntao Qian
IGARSS4
2013 MT-OMP for hyperspectral imagery denoising with model parameter estimation
abstract
It is extensively accepted that much noise is included in hyperspectral imagery (HSI). Noise removal for HSI is an important but challenging task. Most denoising methods have one or more model parameters. For many algorithms, the denoising performance strongly depends on the values of parameters. In many cases, empirically selected parameters are not adaptive to various noise levels. Another challenge is the computational complexity. Since HSI has numerous bands, band by band HSI denoising is relatively time-consuming when compared to RGB or gray image. So a fast algorithm is preferred in practice. In this work, a multi-task orthogonal matching pursuit (MT-OMP) algorithm is proposed for ℓ2,0non-local sparse denoising. This greedy scheme is a multi-task extension of the famous OMP algorithm. The only parameter of MT-OMP is the sparse reconstruction error, which can be derived via noise variance. Furthermore, it is time-efficient and easy to implement. The experimental results show advantages of the proposed MT-OMP algorithm.
Minchao Ye, Yuntao Qian
IGARSS2
2013 Panchromatic image based dictionary learning for hyperspectral imagery denoising
abstract
Sparse coding based noise reduction algorithms have been extensively applied on hyperspectral imagery (HSI) denoising. Dictionary learning schemes are strongly suggested for sparse reconstruction in many researches, aiming at a smaller error between the underlying clean image and the reconstruction result. In previous researches, the training samples (patches) are selected from either unrelated clean images or the noised image itself. The dictionaries learned form unrelated clean images can not perfectly represent the underlying clean target image, while the dictionaries learned form the noised image itself may be affected by the noise existing in training samples. In this paper, we propose a novel dictionary learning scheme that depends on a panchromatic image from the same or similar scene with HSI. Considering the fact that the noise level of a panchromatic image is always much lower than HSI, we take the patches from panchromatic image as training samples. Taking the multi-scale image representation into consideration, we construct the dictionary from different scales via Gaussian pyramid. The proposed dictionary shows its good denoising performance in our experiments.
Minchao Ye, Yuntao Qian
IGARSS2
2013 Text categorization based on regularization extreme learning machine
Yuntao Qian, Huijuan Lu
Neural Comput. Appl.2
2013 Hyperspectral Image Classification Based on Structured Sparse Logistic Regression and Three-Dimensional Wavelet Texture Features
abstract
Hyperspectral remote sensing imagery contains rich information on spectral and spatial distributions of distinct surface materials. Owing to its numerous and continuous spectral bands, hyperspectral data enable more accurate and reliable material classification than using panchromatic or multispectral imagery. However, high-dimensional spectral features and limited number of available training samples have caused some difficulties in the classification, such as overfitting in learning, noise sensitiveness, overloaded computation, and lack of meaningful physical interpretability. In this paper, we propose a hyperspectral feature extraction and pixel classification method based on structured sparse logistic regression and 3-D discrete wavelet transform (3D-DWT) texture features. The 3D-DWT decomposes a hyperspectral data cube at different scales, frequencies, and orientations, during which the hyperspectral data cube is considered as a whole tensor instead of adapting the data to a vector or matrix. This allows the capture of geometrical and statistical spectral-spatial structures. After the feature extraction step, sparse representation/modeling is applied for data analysis and processing via sparse regularized optimization, which selects a small subset of the original feature variables to model the data for regression and classification purpose. A linear structured sparse logistic regression model is proposed to simultaneously select the discriminant features from the pool of 3D-DWT texture features and learn the coefficients of the linear classifier, in which the prior knowledge about feature structure can be mapped into the various sparsity-inducing norms such as lasso, group, and sparse group lasso. Furthermore, to overcome the limitation of linear models, we extended the linear sparse model to nonlinear classification by partitioning the feature space into subspaces of linearly separable samples. The advantages of our methods are validated on the real hyperspectral remote sensing data sets.
Yuntao Qian, Minchao Ye, Jun Zhou 0001
IEEE Trans. Geosci. Remote. Sens.1
2012 Non-negative Sparse Semantic Coding for text categorization
Yuntao Qian
ICPR2
2012 3-D nonlocal means filter with noise estimation for hyperspectral imagery denoising
abstract
Noise reduction is one of important processing tasks for hyperspectral imagery (HSI). In this paper, a three-dimensional (3-D) nonlocal means filter is proposed for noise reduction of HSI. Recently, non-local means method attracts many attentions due to its global and local integrated property. Nonlocal algorithm searches the similar image patches in the whole scene to build the mean filter, so that it overcomes the disadvantage of local filter that only local pixels within a small neighbor is used, and the disadvantage of global filter that local structure is ignored. In order to explore the spectral-spatial correlation of HSI, nonlocal means method is extended from 2-D to 3-D. Furthermore, as HSI contains both of signal-independent and signal-dependent noises, variance-stabilizing transformation based on noise estimation is used to make noise reduction under the additive Gaussian noise model. Experiments with the real hyperspectral data set indicate that the proposed strategy can work well in both of detail preservation and noise removal.
Yuntao Qian, Yanhao Shen, Minchao Ye
IGARSS1
2012 Noise reduction of hyperspectral imagery using nonlocal sparse representation with spectral-spatial structure
abstract
Noise reduction is always an active research area in image processing due to its importance for the sequential tasks such as object classification and detection. In this paper, we develop a sparse representation based noise reduction method for hyperspectral imagery, which is dependent on the assumption that the non-noise component in the signal can be approximated by only a small number of atoms in a dictionary while noise component has not this property. The main contribution of the paper is in introducing nonlocal similarity and spectral-spatial structure of hyperspectral imagery into sparse representation. Non-locality means the self-similarity of image, by which the whole image can be partitioned into some groups containing similar patches. The similar patches in each group is sparsely represented with shared atoms making the signal and noise more easily separated. Sparse representation with spectral-spatial structure can exploit spectral and spatial joint correlations of hyperspectral imagery also making the signal and noise more distinguished, in which 3-D blocks are instead of 2-D patches for sparse coding. The experimental results indicate that the proposed method has a good quality of restoring the true signal from the noisy observation.
Yuntao Qian, Minchao Ye
IGARSS1
2011 Fast and accurate iris segmentation based on linear basis function and RANSAC
abstract
Iris segmentation is a component of iris recognition system, and noncircular iris is hard to segment accurately. This paper presents an iris segmentation algorithm using linear basis function and RANSAC (Random SAmple Consensus) which iterately derives fine iris boundary curves from coarse iris boundary points. The algorithm consists of three steps. In step 1, coarse center and radius of iris are found using IDO (Integro Differential Operators); in step 2, coarse iris boundary points are located, and then a linear basis function model is constructed to derive coarse iris boundary curves from the boundary points; and in step 3, a RANSAC method is applied to refine the iris boundary curves. The proposed algorithm is tested on two datasets CASIA-Iris V3-Interval and IITD v1.0 and shows the effectiveness comparing with some popular algorithms.
Yuntao Qian
ICIP2
2011 Dimension reduction of hyperspectral images with sparse linear discriminant analysis
abstract
Hyperspectral imagery generally contains enormous amounts of data due to hundreds of spectral bands. Classification for these high-dimensional data often requires a large set of training samples and enormous processing time. Therefore, dimension reduction methods for hyperspectral data are catching the attention of researchers lately. In this paper, a dimension reduction method based on sparse penalty regularized linear discriminant analysis was experimented on hyperspectral data. Through imposing sparsity regularization penalty on the Fisher's discriminant analysis projection matrix via the optimal scoring technique, sparse linear discriminant vectors can be achieved. Therefore, interpretability of the spectral bands' physical meaning and effective low dimensional data transforming can be achieved simultaneously in the same model. Experimental analysis on the sparsity and efficacy of low dimensional outputs showed that, sparse linear discriminant analysis can yield good classification results and interpretability in spectral domain.
Jiming Li, Yuntao Qian
IGARSS2
2011 Structured sparse model based feature selection and classification for hyperspectral imagery
abstract
Sparse modeling is a powerful framework for data analysis and processing. It is especially useful for high-dimensional regression and classification problems in which a large number of feature variables exist but the amount of training samples is limited. In this paper, we address the problems of feature description, feature selection and classifier design for hyperspectral images using structured sparse models. A linear sparse logistic regression model is proposed to combine feature selection and pixel classification into a regularized optimization problem with the constraint of sparsity. To explore the structured features, three-dimensional discrete wavelet transform (3D-DWT) is employed, which processes the hyperspectral data cube as a whole tensor instead of adapting the data to a vector or matrix. This allows more effective capturing of the spatial and spectral structure. The structure of the 3D-DWT features is imposed on the sparse model by group LASSO which selects the features on the group level. The advantages of our method are validated on the real hyperspectral data.
Yuntao Qian, Jun Zhou 0001, Minchao Ye
IGARSS1
2011 Hyperspectral Unmixing via $L_{1/2}$ Sparsity-Constrained Nonnegative Matrix Factorization
abstract
Hyperspectral unmixing is a crucial preprocessing step for material classification and recognition. In the last decade, nonnegative matrix factorization (NMF) and its extensions have been intensively studied to unmix hyperspectral imagery and recover the material end-members. As an important constraint for NMF, sparsity has been modeled making use of the$L_{1}$regularizer. Unfortunately, the$L_{1}$regularizer cannot enforce further sparsity when the full additivity constraint of material abundances is used, hence limiting the practical efficacy of NMF methods in hyperspectral unmixing. In this paper, we extend the NMF method by incorporating the$L_{1/2}$sparsity constraint, which we name$L_{1/2}$-NMF. The$L_{1/2}$regularizer not only induces sparsity but is also a better choice among$L_{q}(0 < q < 1)$regularizers. We propose an iterative estimation algorithm for$L_{1/2}$-NMF, which provides sparser and more accurate results than those delivered using the$L_{1}$norm. We illustrate the utility of our method on synthetic and real hyperspectral data and compare our results to those yielded by other state-of-the-art methods.
Yuntao Qian, Sen Jia 0001, Jun Zhou 0001, Antonio Robles-Kelly
IEEE Trans. Geosci. Remote. Sens.1
2010 Hierarchical alternating least squares algorithm with Sparsity Constraint for hyperspectral unmixing
abstract
In this paper, we not only extend the temporal hierarchical alternating least squares (HALS) to spatial domain, but also incorporate two necessary characteristics of material abundances, full additivity and sparsity, to unmix hyperspectral data. The new algorithm is abbreviated as HALSSC (HALS with Sparsity Constraint). Different from the other endmember extraction approaches, the proposed algorithm does not need the existence assumption of pure pixel of each endmember in the scene. Experimental results on highly mixed synthetic data and real hyperspectral data from Washington DC mall confirm the accuracy of the developed algorithm.
Sen Jia 0001, Yuntao Qian, Jiming Li, Yan Li 0066, Zhong Ming 0001
ICIP2
2010 Regularized logistic regression method for change detection in multispectral data via Pathwise Coordinate optimization
abstract
Remotely sensed data by sensors on satellite or airborne platform, is becoming more and more important in monitoring the local, regional and global resources and environment. In this paper, we utilize the regularized logistic regression model for change detection of large scale remotely sensed bi-temporal multispectral images. Change detection methods base on classification schemes under this kind of condition should put more emphasis on the model's simplicity and efficiency in addition to the detection accuracy. The simple linear classifier is solved by recent proposed “Pathwise Coordinate Descent”. When applied on the L1-regularized regression problem, the algorithm can handle large problems in a comparatively very low timing cost. Through computing the solutions for a decreasing sequence of regularization parameters, the algorithm also combines model selection procedure into itself. We experiment the logistic regression with elastic-net convex penalty. Experimental results from a real data set demonstrate that, models obtained by Pathwise Coordinate Descent algorithm only need very low computational costs. The achieved remarkable efficiency indicates that regularized logistic regression via Pathwise Coordinate Descent is a promising method for large scale change detection problem in remote sensing.
Jiming Li, Yuntao Qian, Sen Jia 0001
ICIP2
2010 Feature extraction and selection hybrid algorithm for hyperspectral imagery classification
abstract
Due to the enormous amounts of data contained in hyperspectral imagery, the main challenge for hyperspectral image classification is to improve the accuracy with less computation complexity. Hence, dimensionality reduction (DR) is often adopted, which includes two different kinds of methods, feature extraction and feature selection. In this paper, discrete wavelet transform (DWT) and affinity propagation (AP), which belong to feature extraction and feature selection respectively, are combined together to accomplish the DR task. Firstly, DWT-based features are extracted from the original hyperspectral data; secondly, AP is applied to select representative features from the obtained ones. Experimental results demonstrate that, compared with some other DR methods which only make use of feature extraction or feature selection, the features acquired by the hybrid technique make the classification results more accurate.
Sen Jia 0001, Yuntao Qian, Jiming Li, Weixiang Liu, Zhen Ji
IGARSS2
2010 Noise-robust subband decomposition blind signal separation for hyperspectral unmixing
abstract
Hyperspectral unmixing can be considered as a blind source separation (BSS) and/or independent component analysis (ICA) problem. This paper presents a new noise-resistant subband decomposition BSS/ICA approach for hyperspectral unmixing. Subband decomposition BSS relaxes the assumption that the source signals are mutual independent, which has been proved successful in some BSS applications. However, the existing subband decomposition and subband selection methods emphasize the “independence” of sub-components, but ignore the impact of their “noise”. It is well known that most subband decomposition such as wavelet and fourier transforms have been successfully used for noise removal, so simultaneously considering independence and noise through subband decomposition is possible. In this paper, we propose wavelet package transform for subband decomposition, independence-and-noise joint measure based ranking method for subband selection. The experimental results indicate that the proposed methods are promising in hyperspectral unmixing.
Yuntao Qian
IGARSS1
2009 Hyperspectral data classification using Margin Infused Relaxed Algorithm
abstract
Obtaining training sets for special hyperspectral data sets or applications seems so time consuming and expensive especially for relatively inaccessible locations. Moreover, current techniques of image processing and pattern recognition are not robust enough to make automated remote sensing interpretation feasible. The Margin Infused Relaxed Algorithm (MIRA) is a new perceptron-like online algorithm with a margin-dependent learning rate; meanwhile, it's also a specific online algorithm that seeks a set of prototypes to represent each class. In this paper, we put emphasis on building an online framework by MIRA, which can naturally combine inputs from human and learn as few labeled data points as possible. Experimental results have proved that the MIRA applied in our method is effective in classification problem and economical of the computation time cost.
Jiming Li, Zhenfang Hu, Yuntao Qian
ICIP3
2009 Band selection based gaussian processes for hyperspectral remote sensing images classification
abstract
Classification of Hyperspectral remote sensing images is an important research direction. Hyperspectral remote sensing images have high dimension and nonlinear property. Band selection is often adopted firstly to reduce computational cost and accelerate knowledge discovery of subsequent classification and analysis. Furthermore, Hyperspectral images often contain some uncertainty brought by mixed pixels. We proposed a new band selection based Gaussian processes method to solve these problems. Our method is a Bayesian kernel-based nonlinear method, so it is suitable for nonlinear data classification and it can reduce the uncertainty by computation of posterior label probabilities. Experiment results show that our method is very good at classification of Hyperspectral remote sensing images with respect to classification accuracy and stability.
Futian Yao, Yuntao Qian
ICIP2
2009 Constrained Nonnegative Matrix Factorization for Hyperspectral Unmixing
abstract
Hyperspectral unmixing is a process to identify the constituent materials and estimate the corresponding fractions from the mixture. During the last few years, nonnegative matrix factorization (NMF), as a suitable candidate for the linear spectral mixture model, has been applied to unmix hyperspectral data. Unfortunately, the local minima caused by the nonconvexity of the objective function makes the solution nonunique, thus only the nonnegativity constraint is not sufficient enough to lead to a well-defined problem. Therefore, in this paper, two inherent characteristics of hyperspectral data, piecewise smoothness (both temporal and spatial) of spectral data and sparseness of abundance fraction of every material, are introduced to NMF. The adaptive potential function from discontinuity adaptive Markov random field model is used to describe the smoothness constraint while preserving discontinuities in spectral data. At the same time, two NMF algorithms, nonsmooth NMF and NMF with sparseness constraint, are used to quantify the degree of sparseness of material abundances. A gradient-based optimization algorithm is presented, and the monotonic convergence of the algorithm is proved. Three important facts are exploited in our method: First, both the spectra and abundances are nonnegative; second, the variation of the material spectra and abundance images is piecewise smooth in wavelength and spatial spaces, respectively; third, the abundance distribution of each material is almost sparse in the scene. Experiments using synthetic and real data demonstrate that the proposed algorithm provides an effective unsupervised technique for hyperspectral unmixing.
Sen Jia 0001, Yuntao Qian
IEEE Trans. Geosci. Remote. Sens.2
2008 Improved recognition of figures containing fluorescence microscope images in online journal articles using graphical models
abstract
MOTIVATION: There is extensive interest in automating the collection, organization and analysis of biological data. Data in the form of images in online literature present special challenges for such efforts. The first steps in understanding the contents of a figure are decomposing it into panels and determining the type of each panel. In biological literature, panel types include many kinds of images collected by different techniques, such as photographs of gels or images from microscopes. We have previously described the SLIF system (http://slif.cbi.cmu.edu) that identifies panels containing fluorescence microscope images among figures in online journal articles as a prelude to further analysis of the subcellular patterns in such images. This system contains a pretrained classifier that uses image features to assign a type (class) to each separate panel. However, the types of panels in a figure are often correlated, so that we can consider the class of a panel to be dependent not only on its own features but also on the types of the other panels in a figure. RESULTS: In this article, we introduce the use of a type of probabilistic graphical model, a factor graph, to represent the structured information about the images in a figure, and permit more robust and accurate inference about their types. We obtain significant improvement over results for considering panels separately. AVAILABILITY: The code and data used for the experiments described here are available from http://murphylab.web.cmu.edu/software.
Yuntao Qian, Robert F. Murphy
Bioinform.1
2007 Face recognition using a kernel fractional-step discriminant analysis algorithm
Guang Dai, Dit-Yan Yeung, Yuntao Qian
Pattern Recognit.3
2007 Spectral and Spatial Complexity-Based Hyperspectral Unmixing
abstract
Hyperspectral unmixing, which decomposes pixel spectra into a collection of constituent spectra, is a preprocessing step for hyperspectral applications like target detection and classification. It can be considered as a blind source separation (BSS) problem. Independent component analysis, which is a widely used method for performing BSS, models a mixed pixel as a linear mixture of its constituent spectra weighted by the correspondent abundance fractions (sources). The sources are assumed to be independent and stationary. However, in many instances, this assumption is not valid. In this paper, a complexity-based BSS algorithm is introduced, which studies the complexity of sources instead of the independence. We extend the 1-D temporal complexity, which is called complexity pursuit that was proposed by Stone, to the 2-D spatial complexity, which is named spatial complexity BSS (SCBSS), to describe the spatial autocorrelation of each abundance fraction. Further, the temporal complexity of spectrum is combined into SCBSS to account for the spectral smoothness, which is termed spectral and spatial complexity BSS. More importantly, a strict theoretic interpretation is given, showing that the complexity-based BSS is very suitable for hyperspectral unmixing. Experimental results on synthetic and real hyperspectral data demonstrate the advantages of the proposed two algorithms with respect to other methods.
Sen Jia 0001, Yuntao Qian
IEEE Trans. Geosci. Remote. Sens.2
2006 Semi-supervised Dynamic Counter Propagation Network
Yuntao Qian
ADMA2
2006 Segmental Semi-Markov Model Based Online Series Pattern Detection Under Arbitrary Time Scaling
Guangjie Ling, Yuntao Qian, Sen Jia 0001
ADMA2
2006 A Network Event Correlation Algorithm Based on Fault Filtration
Qiuhua Zheng, Yuntao Qian
PRICAI2
2005 A semi-supervised color image segmentation method
abstract
A new color image segmentation algorithm based on semi-supervised clustering is proposed, which integrates limited human assistance, a user indicates the relationship of some different regions in an image by mouse, to get the final accurate segmentation result which satisfies the prior segmentation constraints. The algorithm first has the image quantified and then clusters in the quantified color space with prior segmentation information. Experiment results show that the proposed algorithm is effective and has high value of utility.
Yuntao Qian, Wenwu Si
ICIP (2)1
2004 Face Recognition Using Novel LDA-Based Algorithms
Guang Dai, Yuntao Qian
ECAI2
2004 Modified kernel-based nonlinear feature extraction [face recognition example]
abstract
Feature extraction techniques are widely used in many applications to pre-process data in order to reduce the complexity of subsequent processes. A group of kernel-based Fisher discriminant analysis (KFDA) algorithms has attracted much attention due to their high performance. In this paper, the inherent limitations of those KFDA algorithms have been discussed and a novel algorithm is proposed to effectively overcome those limitations. Experimental results on face recognition suggest that this proposed algorithm is superior to the existing methods in terms of correct classification rate.
Guang Dai, Yuntao Qian, Sen Jia 0001
ICASSP (5)2
2004 Kernel generalized nonlinear discriminant analysis algorithm for pattern recognition
Guang Dai, Yuntao Qian
ICIP2
2004 A Gabor direct fractional-step LDA algorithm for face recognition
abstract
Recently, a direct fractional-step linear discriminant analysis (DF-LDA) algorithm was proposed and successfully applied to face recognition (FR). However, the classification performance of DFT-LDA is degraded by the limitations of the direct linear discriminant analysis (D-LDA) used in DF-LDA. We describe a novel DF-LDA to solve this problem, and based on this novel DF-LDA, a novel Gabor DF-LDA (GDF-LDA), which directly applies the novel DF-LDA to the high-dimensional augmented Gabor feature vectors (AGFV) derived from the Gabor wavelet representation of face images, has been proposed for FR. The GDF-LDA not only is robust to facial variations, but also overcomes the limitations of the previous DF-LDA. The comparative results on the ORL database show that the GDF-LDA is more effective than existing FR methods.
Guang Dai, Yuntao Qian
ICME2
2002 Sequential Combination Methods for Data Clustering Analysis
Yuntao Qian, Ching Y. Suen, Yuan Yan Tang
J. Comput. Sci. Technol.1
2000 Clustering Combination Method
abstract
Clustering combination uses more than one clustering method with identical pattern features to improve the clustering performance. In general clustering is an optimization procedure based on a specific clustering criterion, so clustering combination can be regarded as a technique that constructs and processes multiple clustering criteria rather than a single criterion. We propose two methods of combining objective function clustering and graph theory clustering. One incorporates multiple criteria into an objective function according to their importance, and solves this problem with constrained nonlinear optimization programming. The other method consists of two sequential procedures: (a) a traditional objective function clustering for generating the initial result, and (b) an autoassociative additive system based on graph theory clustering for modifying the initial result.
Yuntao Qian, Ching Y. Suen
ICPR1
1999 Image Interpretation with Fuzzy-Graph Based Genetic Algorithm
abstract
Image interpretation is a very challenging problem in the scope of computer vision. A novel method based on fuzzy graph and genetic algorithm is presented in this paper. It consists of two parts. At first, the fuzzy classification memberships are obtained by fuzzy classifier based on the statistic/geometric unary features of segmented regions, and a fuzzy graph used for describing interpretation information is built through the spatial binary features and the prior rule-base concerning spatial relations that is acquired from statistics or experience. Second, genetic searching algorithm is used to combine unary and binary features, realize uncertain analysis on the graph with loop, and achieve an optimistic interpretation result. Moreover this method can deal with the case where the image segmentation result is not precise. The simulation results are promising for this novel image interpretation method. It is an improvement over the one-directional reasoning method based image interpretation such as probabilistic, evidence and fuzzy reasoning.
Yuntao Qian
ICIP (1)1
1997 Image Segmentation Based on Combination of Global and Local Information
abstract
An image segmentation approach based on modified fuzzy c-mean clustering algorithm is developed. This method deals with the global and local image information at a gross scene level, which incorporates the local information including the edge map and the spatial relationship of the pixels into the parameters of its objective function. But the current clustering based segmentation methods usually incorporate the local information into the feature space, or integrate the global and local information at a local level. In addition, we also propose a fuzzy Gaussian basis function neural network to complete fuzzy clustering on the grey-histogram of image as the initial solution, which can automatically determine the number of clusters, and is strong and robust.
Yuntao Qian, Rongchun Zhao
ICIP (1)1