EDBT 2026 Demo / reviewers in the wild / expert
Fengchao Xiong
dblp:196/0256
· DBLP profile ↗
49ranked-venue papers
19as first author
39since 2021 · last 2026
0000-0002-9753-4919ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 32 · 12 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 7 first-author · 9 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bridging the resolution gap: Semantic-aware alignment for cross-resolution change detection
Wang Hao, Fengchao Xiong, Jianfeng Lu 0003, Jingzhou Chen, Yuntao Qian |
Pattern Recognit. | 2 |
| 2025 | Single-frame multi-exposure image fusion via narrowband filter decoupled imaging
Xin Ke, Jing Han 0009, Jun Lu 0006, Lianfa Bai, Shuaifeng Gong, Fengchao Xiong, Duan Wei |
Neurocomputing | 10 |
| 2025 | Multi-domain universal representation learning for hyperspectral object tracking
Zhuanfeng Li, Fengchao Xiong, Jianfeng Lu 0003, Jing Wang 0062, Diqi Chen, Jun Zhou 0001, Yuntao Qian |
Pattern Recognit. | 2 |
| 2025 | Hierarchical Contrastive Learning for Multigranularity Ship Classification With Learnable Class QueriesabstractShip targets in remote sensing images can be categorized at various granularities due to variations in image quality, ranging from general ship categories to fine-grained classes like Nimitz-class carriers. Traditional studies mainly focus on fine-grained ship classification, often neglecting samples observed at coarser-grained levels. Samples distributed across multiple granularity levels exhibit semantic relationships among their annotated classes, enabling hierarchical knowledge transfer during model training. This paper incorporates two semantic relationships into deep learning-based representation learning and class prediction: parent-child relationships across levels and mutual exclusivity among sibling categories. For hierarchical representation learning, the proposed hierarchical contrastive learning algorithm extracts category-specific representations from input images and aligns them with their semantic relationships, ensuring that parent and child categories share similarities while sibling categories remain distinct. For hierarchical class predictions, a novel consistency loss ensures coherence in probability distributions between parent and child categories. Specially, cross-entropy loss is employed to impose mutual exclusivity among sibling categories. In this paper, a multi-modal dataset is also designedly developed for hierarchical classification, which integrates optical and synthetic aperture radar (SAR) images across multiple hierarchical levels. Experiments on two popular datasets and a multi-modal dataset demonstrate that the proposed method outperforms state-of-the-art approaches in hierarchical multi-granularity ship classification. Jingzhou Chen, Fengchao Xiong, Yuntao Qian, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Transformer-Based Cross-Domain Few-Shot Learning for Hyperspectral Target DetectionabstractDeep learning-based methods have made significant progress in hyperspectral target detection (HTD). Unfortunately, limited target prior information and imbalance class resulting from the low occurrence probability of target leaves deep learning-based methods to confront bottlenecks. To ameliorate the abovementioned issues, a Transformer-based cross-domain few-shot learning (TCFSL) method is proposed for HTD. First, the TCFSL leverages cross-domain few-shot learning (FSL) to establish FSL tasks in both the source domain (SD) and the target domain (TD). This allows the TCFSL to learn transferable knowledge of the SD and distinguishable feature embedding model for the TD, to address the problems of target priori lacking and imbalance class. Second, feature-level and distribution-level domain adaptation (DA) is used to tackle the problem of domain shift in cross-domain FSL. The feature-level DA extracts intradomain information of the SD and TD to learn their common features to alleviate domain shift. The distribution-level DA based on cross-Transformer present interdomain distribution-level information aggregation and captures domain similarities of two data domains. By pursuing similarities between two data domains, the distribution-level DA block prompts specific FSL tasks in each domain, facilitating the target detection task. Finally, cross-domain FSL and DA blocks are trained in a unitary manner, which facilitates real-time information interaction and parameter adjustment between different blocks to achieve the optimal model. Experiments conducted on six HSI datasets indicate that the TCFSL outperforms 12 compared methods. Shou Feng, Fengchao Xiong, Chunhui Zhao 0003, Wei Li 0032, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Mid-Range Convolutional Modulated Transformer Network for Hyperspectral Image Classification
Mingzhu Tai, Zigao Liu, Zhenqiu Shu, Fengchao Xiong, Zhengtao Yu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Spatial-Spectral-Temporal Correlation Filter for Hyperspectral Object TrackingabstractObject tracking with hyperspectral videos (HSVs) offers significant advantages due to the captured spectral fingerprint information, which provides detailed physical material characteristics. While correlation filter (CF)-based tracking methods align well with the high-dimensional nature of HSVs, they often fall short of fully utilizing the spatial–spectral–temporal structure inherent in these data. In this article, we introduce a spatial–spectral–temporal CF (SSTCF) framework to address these limitations. SSTCF employs the spatial-spectral histogram of gradients and fractional abundances as features to characterize the spatial-spectral structure of the object. A low-rank constraint is integrated into the CF framework to enhance the global spectral semantic dependencies among learned filters. In addition, a temporal constraint is incorporated to ensure filter consistency across consecutive frames, further improving tracking continuity between nearby frames. Extensive experiments demonstrate that our SSTCF tracker achieves more accurate and stable performance. The source code will be publicly available athttps://github.com/bearshng/SSTCF Fengchao Xiong, Yongle Sun, Jun Zhou 0001, Jianfeng Lu 0003, Yuntao Qian |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | SUIT: Spatial-Spectral Union-Intersection Interaction Network for Hyperspectral Object TrackingabstractHyperspectral videos (HSVs), with their inherent spatial-spectral-temporal structure, offer distinct advantages in challenging tracking scenarios such as cluttered backgrounds and small objects. However, existing methods primarily focus on spatial interactions between the template and search regions, often overlooking spectral interactions, leading to suboptimal performance. To address this issue, this paper investigates spectral interactions from both the architectural and training perspectives. At the architectural level, we first establish band-wise long-range spatial relationships between the template and search regions using Transformers. We then model spectral interactions using the inclusion-exclusion principle from set theory, treating them as the union of spatial interactions across all bands. This enables the effective integration of both shared and band-specific spatial cues. At the training level, we introduce a spectral loss to enforce material distribution alignment between the template and predicted regions, enhancing robustness to shape deformation and appearance variations. Extensive experiments demonstrate that our tracker achieves state-of-the-art tracking performance. The source code, trained models and results will be publicly available via https://github.com/bearshng/suit to support reproducibility. Fengchao Xiong, Zhenxing Wu, Jun Zhou 0001, Sen Jia 0001, Yuntao Qian |
IEEE Trans. Image Process. | 1 |
| 2024 | Semantic-Aware Alignment Network for Cross-Resolution Change DetectionabstractCross-resolution change detection (CRCD) is of significant practical importance in disaster assessment, rapid urban transitions, and various applications. Conventional change detection methods are primarily tailored for bitemporal images with consistent spatial resolution, rendering them unsuitable for direct application to CRCD tasks. This limitation stems from the substantial scale differences and pixel-wise misalignment prevalent in cross-resolution remote sensing images. In response to these challenges, we introduce a semantic-aware alignment network (SA-Net). SA-Net utilizes cross-attention to map bitemporal images into a shared semantic space, effectively alleviating the difficulties of the subsequent alignment arising from semantic mismatches. Furthermore, a joint transformer featuring an encoder-decoder architecture is employed to extract global information and learn the geometric parameters for spatial alignment between bitemporal images. Experimental evaluations on two real-collected datasets, HTCD and MRCDD, showcase the superior performance of our proposed SA-Net in CRCD tasks. Fengchao Xiong, Jianfeng Lu 0003, Minchao Ye, Jun Zhou 0001, Yuntao Qian |
IGARSS | 2 |
| 2024 | SSUMamba: Spatial-Spectral Selective State Space Model for Hyperspectral Image DenoisingabstractDenoising is a crucial preprocessing step for hyperspectral images (HSIs) due to noise arising from intraimaging mechanisms and environmental factors. Long-range spatial-spectral correlation modeling is beneficial for HSI denoising but often comes with high complexity. Based on the state space model (SSM), Mamba is known for its remarkable long-range dependency modeling capabilities and computational efficiency. Building on this, we introduce a memory-efficient spatial-spectral UMamba (SSUMamba) for HSI denoising, with the spatial-spectral continuous scan (SSCS) Mamba being the core component. SSCS Mamba alternates the row, column, and band in six different orders to generate the sequence and uses the bidirectional SSM to exploit long-range spatial-spectral dependencies. In each order, the images are rearranged between adjacent scans to ensure spatial-spectral continuity. In addition, 3-D convolutions are embedded into the SSCS Mamba to enhance local spatial-spectral modeling. Experiments demonstrate that SSUMamba achieves superior denoising results with lower memory consumption per batch compared with transformer-based methods. The source code is available at:https://github.com/lronkitty/SSUMamba. Guanyiman Fu, Fengchao Xiong, Jianfeng Lu 0003, Jun Zhou 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Hyperspectral Image Denoising via Spatial-Spectral Recurrent TransformerabstractHyperspectral images (HSIs) often suffer from noise arising from both intra-imaging mechanisms and environmental factors. Leveraging domain knowledge specific to HSIs, such as global spectral correlation (GSC) and non-local spatial self-similarity (NSS), is crucial for effective denoising. Existing methods tend to independently utilize each of these knowledge components with multiple blocks, overlooking the inherent 3D nature of HSIs where domain knowledge is strongly interlinked, resulting in suboptimal performance. To address this challenge, this paper introduces a spatial-spectral recurrent transformer U-Net (SSRT-UNet) for HSI denoising. The proposed SSRT-UNet integrates NSS and GSC properties within a single SSRT block. This block consists of a spatial branch and a spectral branch. The spectral branch employs a combination of transformer and recurrent neural network to perform recurrent computations across bands, allowing for GSC exploitation beyond a fixed number of bands. Concurrently, the spatial branch encodes NSS for each band by sharingkeysandvalueswith the spectral branch under the guidance of GSC. The interaction between the two branches enables the joint utilization of NSS and GSC, avoiding their independent treatment. Experimental results demonstrate that our method outperforms several alternative approaches. The source code will be available at https://github.com/lronkitty/SSRT. Guanyiman Fu, Fengchao Xiong, Jianfeng Lu 0003, Jun Zhou 0001, Jiantao Zhou 0001, Yuntao Qian |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Material-Guided Multiview Fusion Network for Hyperspectral Object TrackingabstractHyperspectral videos (HSVs) have more potential in object tracking than color videos thanks to their material identification ability. Nevertheless, previous works have not fully explored the benefits of the material information, resulting in limited representation ability and tracking accuracy. To address this issue, this paper introduces a material-guided multi-view fusion network for improved tracking. Specifically, we combine false-color information, hyperspectral information, and material information obtained by hyperspectral unmixing to provide a rich multi-view representation of the object. Cross-material attention is employed to capture the interaction among materials, enabling the network to focus on the most relevant materials for the target. Furthermore, leveraging the discriminative ability of material view, a novel material-guided multi-view fusion module is proposed to capture both intra-view and cross-view long-range spatial dependencies for effective feature aggregation. Thanks to the enhanced representation ability of each view and the integration of the complementary advantages of all views, our network is more capable of suppressing the tracking drift in various challenging scenes and achieving accurate object localization. Extensive experiments show that our tracker achieves state-of-the-art tracking performance. The source code will be available at https://github.com/hscv/MMF-Net. Zhuanfeng Li, Fengchao Xiong, Jun Zhou 0001, Jianfeng Lu 0003, Yuntao Qian |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Wavelet Siamese Network With Semi-Supervised Domain Adaptation for Remote Sensing Image Change DetectionabstractChange detection is a crucial technique in remote sensing image analysis and faces challenges, such as background complexity and appearance shift, resulting in incomplete change boundaries and pseudochanges. This article introduces a novel wavelet Siamese network with semi-supervised domain adaptation (DA) to address these issues, named WS-Net++. WS-Net++ establishes spatial–frequency interactions between bitemporal images to enhance the completeness of the change boundaries. The spatial-domain interaction highlights the pixelwise differences. The frequency-domain interaction first adaptively adjusts the contributions from different frequency components based on image context. Within-frequency and between-frequency interactions are further constructed to capture the frequency-domain differences, enabling the adaptive and effective handling of both overall and subtle changes. In addition, WS-Net++ employs a semi-supervised DA strategy to mitigate the appearance shifts between bitemporal images. By categorizing regions into changed, unchanged, and regions of no interest in a semi-supervised manner, the network minimizes intraclass discrepancies within unchanged regions and maximizes interclass discrepancies between changed regions, reducing the domain gap. Experimental results on the LEVIR-CD, WHU-CD, and CLCD datasets demonstrate that our WS-Net++ outperforms alternative methods, achieving the$F1$scores of 91.31%, 94.52%, and 79.77%, respectively. The code and models will be publicly available athttps://github.com/JiTaiTai/WS-Net_Plusfor reproducible research. Fengchao Xiong, Tianhan Li, Yi Yang 0071, Jun Zhou 0001, Jianfeng Lu 0003, Yuntao Qian |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Adaptive Graph Modeling With Self-Training for Heterogeneous Cross-Scene Hyperspectral Image ClassificationabstractThe small-sample-size problem of hyperspectral image (HSI) classification has recently gained considerable attention. Cross-scene HSI classification has emerged as an effective solution to this problem. In real-world applications, different HSI scenes are often captured by diverse sensors, resulting in variations between scenes. Graph modeling, as a method to represent relationships, leverages semantic information to establish connections between scenes, thereby facilitating transfer learning by aligning their features. However, in scenarios with only a few labeled target samples, the resulting graph is typically sparse and can only capture weak cross-scene relationships. Studies have shown that a dense and fault-tolerant graph is beneficial for transfer learning in small-sample-size cases. Consequently, we propose a novel heterogeneous transfer learning approach called adaptive graph modeling with self-training (AGM-ST). Unlike conventional graph modeling methods that employ predefined graph weights, adaptive graph modeling (AGM) employs a learnable network to generate graph weights based on the similarities of spectral–spatial features. Additionally, an adaptive cutoff threshold is trained to eliminate weak relationships between samples that may be potentially incorrect. Subsequently, a cross-scene graph loss is designed based on the generated graph to align the feature spaces of the source and target scenes. Furthermore, the unlabeled samples from the target scene are gradually updated with pseudo labels using the self-training (ST) technique, which enhances semantic information and improves graph modeling. Experimental evaluations conducted on three cross-scene HSI datasets have demonstrated the effectiveness of the proposed AGM-ST approach. Minchao Ye, Junbin Chen, Fengchao Xiong, Yuntao Qian |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Discriminative Vision Transformer for Heterogeneous Cross-Domain Hyperspectral Image ClassificationabstractThe transformer has been introduced in the hyperspectral image (HSI) classification, demonstrating outstanding capability in capturing global features compared to the convolutional neural network (CNN). However, the small-sample-size problem poses a significant challenge in practical HSI classification, especially in training the transformer. To tackle this issue, cross-domain transfer learning is adopted as a practical solution, which transfers the information from a source domain with abundant labeled samples to a target domain with limited labeled samples. This article proposes a novel transfer learning method for heterogeneous cross-domain HSI classification called cross-domain discriminative vision transformer (CD-DViT). This algorithm primarily contains three key contributions. First, source samples are mapped to the target domain through an encoder-decoder architecture, and the mapped source samples can be used to train the target classifier. Second, the cross-attention mechanism is utilized to construct two blocks for achieving the domainwise and classwise feature alignments (FAs), respectively. Specifically, the combination of the cross-attention mechanism with the domain discriminator aims to learn domain-invariant features, thereby facilitating domainwise alignment and alleviating domain shift. Third, knowledge distillation (KD) is adopted to learn more information from the target domain and assist in classifying target samples. Our experiments on three real-world cross-domain HSI datasets demonstrate the effectiveness of the proposed approach. Minchao Ye, Jiawei Ling, Wanli Huo, Zhaojuan Zhang, Fengchao Xiong, Yuntao Qian |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Iterative Low-Rank Network for Hyperspectral Image DenoisingabstractHyperspectral image (HSI) denoising is a crucial preprocessing step for subsequent tasks. The clean HSI usually reside in a low-dimensional subspace, which can be captured by low-rank and sparse representation, known as the physical prior of HSI. It is generally challenging to adequately use such physical properties for effective denoising while preserving image details. This article introduces a novel iterative low-rank network (ILRNet) to address these challenges. ILRNet integrates the strengths of model-driven and data-driven approaches by embedding a rank minimization module (RMM) within a U-Net architecture. This module transforms feature maps into the wavelet domain and applies singular value thresholding (SVT) to the low-frequency components during the forward pass, leveraging the spectral low-rankness of HSIs in the feature domain. The parameter, closely related to the hyperparameter of the singular vector thresholding algorithm, is adaptively learned from the data, allowing for flexible and effective capture of low-rankness across different scenarios. Additionally, ILRNet features an iterative refinement process that adaptively combines intermediate denoised HSIs with noisy inputs. This manner ensures progressive enhancement and superior preservation of image details. Experimental results demonstrate that ILRNet achieves state-of-the-art performance in both synthetic and real-world noise removal tasks. Jin Ye 0008, Fengchao Xiong, Jun Zhou 0001, Yuntao Qian |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Iterative Refinement Network for Hyperspectral Image DenoisingabstractHyperspectral image (HSI) denoising is an important pre-processing procedure for subsequent tasks. Learning a direct mapping from the observed noisy HSI to its clean counterpart is challenging, especially in the case of very severe noise. The learning difficulty can be greatly reduced with the iterative refinement, combining the denoising results with the noisy HSI to produce a cleaner HSI for further denoising. To this end, we introduce an iterative refinement denoising network (IRDNet) for HSIs. The network consists of three key components, i.e., a coarse estimation module, a multi-stage refinement module, and a λ(•) module. The coarse estimation module provides the initial estimate for the starting point of further refinement. The lightweight refinement module progressively performs noise reduction on the weighted combination of noisy inputs and denoising results from the previous layer. Instead of fixed weights, the λ(•) module adaptively provides the layer-wise weight for each band for the aforementioned combinations. Extensive experiments on synthetic and real-world datasets show that our IRDNet favorably outperforms alternative methods. Fengchao Xiong, Jun Zhou 0001, Yuntao Qian |
ICME | 1 |
| 2023 | Wavelet Siamese Network for Change Detection in Remote Sensing ImagesabstractChange detection is a technique used to identify semantic differences between co-registered images of the same area captured at different times. However, current methods often overlook the fact that the low-frequency and high-frequency components of these images play distinct roles in change detection. Our method decomposes each feature map into its low-frequency and high-frequency components and then uses an attention mechanism to adjust the contribution of each component to handle different types of changes. Low-frequency information can help detect overall changes, and high-frequency information can enhance the integrity of the change boundaries. Experiments on the LEVIR-CD, WHU-CD and CLCD datasets show that our model outperforms the state-of-the-art method and the ablation study demonstrates that this approach improve the accuracy of the change detection. Tianhan Li, Fengchao Xiong, Zhuanfeng Li, Jun Zhou 0001, Yuntao Qian |
IGARSS | 2 |
| 2023 | Multi-Task Attentional U-Net for Hyperspectral Image DenoisingabstractHyperspectral image (HSI) denoising is a critical preprocessing step for ensuring the usability of HSIs. However, current deep learning-based approaches still struggle with modeling the global spectral correlation among bands, which is crucial for high-quality denoising. To address this issue, we introduce a multitask attention module and embed it into a U-Net architecture to yield a multitask attentional U-Net (MTA-Net) for HSI denoising. The module enforces all bands to focus on the same region by sharing the same attention map across all bands. This ensures that all bands capture the same image structure, effectively modeling the global spectral correlation. Experimental results demonstrate that the proposed MTA-Net achieves state-of-the-art performance on both synthetic and real-world data. Fengchao Xiong, Zhongyi Gu, Tianhan Li, Jun Zhou 0001 |
IGARSS | 1 |
| 2023 | Domain-invariant attention network for transfer learning between cross-scene hyperspectral imagesabstractAbstract Small‐sample‐size problem is always a challenge for hyperspectral image (HSI) classification. Considering the co‐occurrence of land‐cover classes between similar scenes, transfer learning can be performed, and cross‐scene classification is deemed a feasible approach proposed in recent years. In cross‐scene classification, the source scene which possesses sufficient labelled samples is used for assisting the classification of the target scene that has a few labelled samples. In most situations, different HSI scenes are imaged by different sensors resulting in their various input feature dimensions (i.e. number of bands), hence heterogeneous transfer learning is desired. An end‐to‐end heterogeneous transfer learning algorithm namely domain‐invariant attention network (DIAN) is proposed to solve the cross‐scene classification problem. The DIAN mainly contains two modules. (1) A feature‐alignment CNN (FACNN) is applied to extract features from source and target scenes, respectively, aiming at projecting the heterogeneous features from two scenes into a shared low‐dimensional subspace. (2) A domain‐invariant attention block is developed to gain cross‐domain consistency with a specially designed class‐specific domain‐invariance loss, thus further eliminating the domain shift. The experiments on two different cross‐scene HSI datasets show that the proposed DIAN achieves satisfying classification results. Minchao Ye, Zhihao Meng, Fengchao Xiong, Yuntao Qian |
IET Comput. Vis. | 4 |
| 2023 | Guest Editorial: Spectral imaging powered computer visionabstractThe increasing accessibility and affordability of spectral imaging technology have revolutionised computer vision, allowing for data capture across various wavelengths beyond the visual spectrum.This advancement has greatly enhanced the capabilities of computers and AI systems in observing, understanding, and interacting with the world.Consequently, new datasets in various modalities, such as infrared, ultraviolet, fluorescent, multispectral, and hyperspectral, have been constructed, presenting fresh opportunities for computer vision research and applications.Although significant progress has been made in processing, learning, and utilising data obtained through spectral imaging technology, several challenges persist in the field of computer vision.These challenges include the presence of low-quality images, sparse input, high-dimensional data, expensive data labelling processes, and a lack of methods to effectively analyse and utilise data considering their unique properties.Many mid-level and high-level computer vision tasks, such as object segmentation, detection and recognition, image retrieval and classification, and video tracking and understanding, still have not leveraged the advantages offered by spectral information.Additionally, the problem of effectively and efficiently fusing data in different modalities to create robust vision systems remains unresolved.Therefore, there is a pressing need for novel computer vision methods and applications to advance this research area.This special issue aims to provide a venue for researchers to present innovative computer vision methods driven by the spectral imaging technology. Jun Zhou 0001, Fengchao Xiong, Naoto Yokoya, Pedram Ghamisi |
IET Comput. Vis. | 2 |
| 2023 | An Attention-Based Multiscale Spectral-Spatial Network for Hyperspectral Target DetectionabstractDeep learning-based methods have made great progress in hyperspectral target detection. Unfortunately, the insufficient utilization of spatial information in most methods leaves deep learning-based methods to confront ineffectiveness. To ameliorate this issue, an attention-based multiscale spectral-spatial detector (AMSSD) for hyperspectral target detection is proposed. Firstly, the AMSSD leverages the Siamese structure to establish a similarity discrimination network, which can enlarge intraclass similarity and interclass dissimilarity to facilitate better discrimination between the target and the background. Secondly, 1D CNN and vision Transformer are used combinedly to extract spectral-spatial features more feasibly and adaptively. The joint use of spectral-spatial information can obtain more comprehensive features, which promotes subsequent similarity measurement. Finally, a multiscale spectral-spatial difference feature fusion module is devised to integrate spectral-spatial difference features of different scales to obtain more distinguishable representation and boost detection competence. Experiments conducted on two HSI datasets indicate that the AMSSD outperforms seven compared methods. Shou Feng, Chunhui Zhao 0003, Fengchao Xiong, Lifu Zhang 0002 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2023 | Rapid coded aperture spectrometer based on energy concentration characteristic
Jiutao Mu, Fengchao Xiong, Jun Lu 0006, Jing Han 0009 |
Signal Process. | 4 |
| 2023 | Multitask Sparse Representation Model-Inspired Network for Hyperspectral Image DenoisingabstractHyperspectral images (HSIs) are prone to noise because of the imaging mechanism and environment. This paper proposes a multitask sparse representation (SR) model inspired neural network for HSI denoising. Unlike other deep learning-based methods, our network is interpretable, whose network architecture is induced by unfolding the iterative optimization of a multitask sparse representation model. On the one hand, the model globally represents the common structure among bands, such as image edges, with the shared sparse coefficients. On the other hand, it separately encodes the unique structure of individual bands with unshared ones to capture image details. Accordingly, our network has three modules: the shared SR module, the unshared SR module, and the image reconstruction (IR) module. All the modules are connected with a specific operation of the iterative optimization algorithm, equipping the network with clear physical interpretation. Experimental results on both synthetic and real-world datasets demonstrate the superior performance of our method, visually and quantitatively. The codes will be publicly available at https://github.com/bearshng/mtsrnn for reproducible research. Fengchao Xiong, Jiantao Zhou 0001, Jun Zhou 0001, Jianfeng Lu 0003, Yuntao Qian |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Deep Parameterized Neural Networks for Hyperspectral Image DenoisingabstractSparse representation (SR)-based hyperspectral image (HSI) denoising methods normally average the local denoising results of multiple overlapped cubes to recover the whole HSI. Though interpretable, they rely on cumbersome hyperparameter settings and ignore the relationship between overlapped cubes, leading to poor denoising performance. This article combines SR and convolutional neural networks and introduces a deep parameterized sparse neural network (DPNet-S) to address the above issues. DPNet-S parameterizes the SR-based HSI denoising model with two modules: 1) sparse optimizer to extract sparse feature maps from noisy HSIs via recurrent usage of convolution, deconvolution, and soft shrinkage operations; and 2) image reconstructor to recover the denoised HSI from its sparse feature maps via deconvolution operations. We further replace the soft shrinkage operator with U-Net architecture to account for general HSI priors and more effectively capture the complex structures of HSIs, resulting in DPNet-U. Both networks directly learn the parameters from data and perform denoising on the whole HSI, which overcomes the limitations of SR-based methods. Moreover, our networks are generated from the denoising model and optimization procedures, thus leveraging the knowledge embedded and relying less on the number of training samples. Extensive experiments on both synthetic and real-world HSIs show that our DPNet-S and DPNet-U achieve remarkable results when compared with state-of-the-art methods. The codes will be publicly available athttps://github.com/bearshng/dpnetsfor reproducible research. Fengchao Xiong, Jun Zhou 0001, Jiantao Zhou 0001, Jianfeng Lu 0003, Yuntao Qian |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Learning a Deep Ensemble Network With Band Importance for Hyperspectral Object TrackingabstractAttributing to material identification ability powered by a large number of spectral bands, hyperspectral videos (HSVs) have great potential for object tracking. Most hyperspectral trackers employ manually designed features rather than deeply learned features to describe objects due to limited available HSVs for training, leaving a huge gap to improve the tracking performance. In this paper, we propose an end-to-end deep ensemble network (SEE-Net) to address this challenge. Specifically, we first establish a spectral self-expressive model to learn the band correlation, indicating the importance of a single band in forming hyperspectral data. We parameterize the optimization of the model with a spectral self-expressive module to learn the nonlinear mapping from input hyperspectral frames to band importance. In this way, the prior knowledge of bands is transformed into a learnable network architecture, which has high computational efficiency and can fast adapt to the changes of target appearance because of no iterative optimization. The band importance is further exploited from two aspects. On the one hand, according to the band importance, each frame of HSVs is divided into several three-channel false-color images which are then used for deep feature extraction and location. On the other hand, based on the band importance, the importance of each false-color image is computed, which is then used to assemble the tracking results from individual false-color images. In this way, the unreliable tracking caused by false-color images of low importance can be suppressed to a large extent. Extensive experimental results show that SEE-Net performs favorably against the state-of-the-art approaches. The source code will be available at https://github.com/hscv/SEE-Net. Zhuanfeng Li, Fengchao Xiong, Jun Zhou 0001, Jianfeng Lu 0003, Yuntao Qian |
IEEE Trans. Image Process. | 2 |
| 2022 | Material-Guided Siamese Fusion Network for Hyperspectral Object TrackingabstractHyperspectral videos (HSVs) have more potential in target tracking than color videos thanks to the material identification capability provided by abundant spectral bands. Due to limited HSVs for training, most current hyperspectral trackers are based on hand-crafted features rather than deeply learned ones, resulting in poor tracking performance. This paper introduces a material-guided Siamese fusion network (SiamF) for hyperspectral object tracking to make up this gap. Belonging to the Siamese tracker family and SiamF aims to model the appearance of hyperspectral objects using backbone networks trained on color images. Specifically, SiamF splits each hyperspectral frame into multiple groups of false-color images according to their band importance. Then SiamF employs a hyperspectral feature fusion (HFF) module with a dense connection architecture to integrate the extracted features from different layers and band groups, producing a multi-scale multilevel spatial-spectral representation of the targets. Instead of direct addition or concatenation, HFF employs global-local channel attention for feature fusion, so that yielded features capture the global and local structure of a specific object. Moreover, online spatial and material classifiers are developed to inject spatial and material appearance changes information into SiamF for adaptively online tracking. Experimental results demonstrate our tracker outperforms alternative methods. Zhuanfeng Li, Fengchao Xiong, Jianfeng Lu 0003, Jun Zhou 0001, Yuntao Qian |
ICASSP | 2 |
| 2022 | Multitask Sparse Neural Network for Hyperspectral Image DenoisingabstractData-driven deep learning (DL)-based methods directly learn the nonlinear mapping between noisy hyperspectral images (HSIs) and corresponding clean ones. However, DLbased methods neglect the prior knowledge of HSIs embodied by physical models. Consequently, they require complex network architectures and a large number of training samples. To address the above issues, this paper introduces a multitask sparse neural network (MTSNN) which bridges the sparsity prior of HSIs with data-driven deep learning for HSI denoising. Specifically, we first build a multitask sparse (MTS) denoising model which shares sparse coefficients among bands to exploit the spectral-spatial correlation and learns a dictionary for each band to depict the distinct spatial structure among bands. The iterative optimization of the MTS model is then unfolded to yield our MTSNN by introducing some learnable parameters. MTSNN is a multi-branch network. Each branch performs a single denoising task for an individual band. All branches are connected by shared coefficients, forming multitask denoising for all bands. The hybrid advantages of the MTS model and data-driven learning equip MTSNN with strong denoising ability, preferable learning capability, superior interpretability, and higher generalization capacity. Experimental results demonstrate that our method achieves state-of-the-art denoising performance compared with several alternative approaches. Fengchao Xiong, Minchao Ye, Jun Zhou 0001, Jianfeng Lu 0003, Yuntao Qian |
ICASSP | 1 |
| 2022 | Ques-to-Visual Guided Visual Question AnsweringabstractVisual question answering (VQA) answers text-based questions about images. The difficulty of VQA lies in the accurate localization of the region related to the question. In this paper, we introduce the ques-to-visual (q2v) feature as the additional input of VQA to tackle this problem. The q2v feature is generated according to the semantics of the question, containing visual semantics that is helpful to locate the region related to the question. We then use self-attention to model the intra-relationship in each modality to enhance different features, i.e., q2v, image, and text features. The enhanced features are then fused by spatial guided-attention and multi-scale channel attention modules for the answer prediction. Experimental results on the VQA2.0 benchmark dataset show that our method achieves higher performance when compared with other methods. Jianfeng Lu 0003, Zhuanfeng Li, Fengchao Xiong |
ICIP | 4 |
| 2022 | Cross-Scene Hyperspectral Image Classification Based on Cycle-Consistent Adversarial NetworksabstractLack of labeled training samples is a challenge in hyperspectral image (HSI) classification. Cross-scene classification is a valid solution to few-shot learning problem. In cross-scene classification, two strongly related HSI scenes are considered, one with sufficient labeled samples is called source scene, while the other one containing limited labeled samples is called target scene. By establishing connections between two scenes, abundant labeled samples in source scene can benefit the classification of target scene. In this paper, a novel model named cycle auxiliary classifier generative adversarial network (Cycle-AC-GAN) is proposed for heterogeneous transfer learning across source and target scenes. In Cycle-AC-GAN, a source-to-target generator and a target-to-source generator are simultaneously built. Thus, a two-way mapping can be effectively established between source and target scenes with the adversarial training. In addition, different from existing CycleGAN, in Cycle-AC-GAN, each discriminator contains a binary domain classifier and an auxiliary land-cover classifier. The auxiliary classifiers can align the class-conditional distributions between source and target HSIs. Inspiring experimental results on two real-world cross-scene HSI datasets demonstrate the effectiveness of the proposed approach. Zhihao Meng, Minchao Ye, Futian Yao, Fengchao Xiong, Yuntao Qian |
IGARSS | 4 |
| 2022 | Cross-Domain Attention Network for Hyperspectral Image ClassificationabstractExpensive cost of labeling leads to few-shot learning problem in hyperspectral image (HSI) classification. Cross-scene classification is a novel approach to solve this problem. In this work, we propose an end-to-end heterogeneous transfer learning algorithm namely cross-domain attention network (CDAN) to settle the cross-scene classification problem. CDAN mainly contains two modules. 1) A two-stream HybirdSN architecture is designed for extracting features from source and target scenes, aiming at projecting the features into a shared low-dimensional subspace. 2) Cross-domain attention mechanism is adopted based on the consistency of features between different scenes. A cross-domain updating rule is proposed for training the subnet. CDAN is proved to be effective according to the experiments on two different cross-scene HSI datasets. Minchao Ye, Ling Lei 0002, Fengchao Xiong, Yuntao Qian |
IGARSS | 4 |
| 2022 | Spatial-Spectral Convolutional Sparse Neural Network for Hyperspectral Image DenoisingabstractSparse representation (SR) is a widely accepted hyper-spectral image (HSI) denoising model. Because of the curse of dimensionality and the desire to better fit the data, the SR models are typically deployed on small and fully overlapping blocks whose results are averaged to produce the global de-noised HSI. This “local-global” denoising mechanism ignores the dependencies between blocks, resulting in visual artifacts. This paper describes the underlying clean HSI with a 3D con-volutional sparse coding (CSC) model, representing the HSI with a linear combination of few shift-invariant 3D spatial-spectral filters in a global dictionary. Instead of operating on patches, the CSC model sees the clean HSI is generated from a sum of local atoms that appear in a small number of locations throughout the image, naturally retaining the relationship between pixels. Moreover, we unfold the optimization process of the model into a spatial-spectral convolutional sparse neural network which absorbs the interpretation ability of the model while supporting discriminative learning from data. Experimental results on both synthetic and real-world datasets show that our network achieves competitive denoising performances, qualitatively and quantitatively. Fengchao Xiong, Minchao Ye, Jun Zhou 0001, Yuntao Qian |
IGARSS | 1 |
| 2022 | Nonlocal Spatial-Spectral Neural Network for Hyperspectral Image DenoisingabstractHyperspectral image (HSI) denoising is an essential preprocessing step to improve the quality of HSIs. The difficulty of HSI denoising lies in effectively modeling the intrinsic characteristics of HSIs, such as spatial-spectral correlation, global spectral correlation, and nonlocal spatial correlation. This paper introduces a nonlocal spatial-spectral neural network (NSSNN) for HSI denoising by considering the above three factors in a unified network. More specifically, NSSNN is based on the residual U-Net and embedded with the introduced spatial-spectral recurrent (SSR) blocks and nonlocal self-similarity (NSS) blocks. The SSR block comprises 3D convolutions, one light recurrence, and one highway network. 3D convolution helps exploit the spatial-spectral correlation. The light recurrence and highway network make up the recurrent computation component and refined component, respectively, to model the global spectral correlation. NSS block is based on crisscross attention and can exploit the long-range spatial contexts effectively and efficiently. Attributing to effective modeling of the spatial-spectral correlation, the global spectral correlation, and the nonlocal spatial correlation, our NSSNN has a strong denoising ability. Extensive experiments show the superior denoising effectiveness of our method on synthetic and real-world datasets when compared to alternative methods. The source code will be available at https://github.com/lronkitty/NSSNN. Guanyiman Fu, Fengchao Xiong, Jianfeng Lu 0003, Jun Zhou 0001, Yuntao Qian |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | MAC-Net: Model-Aided Nonlocal Neural Network for Hyperspectral Image DenoisingabstractHyperspectral image (HSI) denoising is an ill-posed inverse problem. The underlying physical model is always important to tackle this problem, which is unfortunately ignored by most of the current deep learning (DL)-based methods, producing poor denoising performance. To address this issue, this article introduces an end-to-end model-aided nonlocal neural network (MAC-Net) which simultaneously takes the spectral low-rank model and spatial deep prior into account for HSI noise reduction. Specifically, motivated by the success of the spectral low-rank model in depicting the strong spectral correlations and the nonlocal similarity prior in capturing spatial long-range dependencies, we first build a spectral low-rank model and then integrate a nonlocal U-Net into the model. In this way, we obtain a hybrid model-based and DL-based HSI denoising method where the spatial local and nonlocal multi-scale and spectral low-rank structures are effectively exploited. After that, we cast the optimization and denoising procedure of the hybrid method as a forward process of a neural network and introduce a set of learnable modules to yield our MAC-Net. Compared with traditional model-based methods, our MAC-Net overcomes the difficulties of accurate modeling, thanks to the strong learning and representation ability of DL. Unlike most “black-box” DL-based methods, the spectral low-rank model is beneficial to increase the generalization ability of the network and decrease the requirement of training samples. Experimental results on the natural and remote-sensing HSIs show that MAC-Net achieves state-of-the-art performance over both model-based and DL-based methods. The source code and data of this article will be made publicly available athttps://github.com/bearshng/mac-netfor reproducible research. Fengchao Xiong, Jun Zhou 0001, Qinling Zhao, Jianfeng Lu 0003, Yuntao Qian |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | SNMF-Net: Learning a Deep Alternating Neural Network for Hyperspectral UnmixingabstractHyperspectral unmixing is recognized as an important tool to learn the constituent materials and corresponding distribution in a scene. The physical spectral mixture model is always important to tackle this problem because of its highly ill-posed nature. In this article, we introduce a linear spectral mixture model (LMM)-based end-to-end deep neural network named SNMF-Net for hyperspectral unmixing. SNMF-Net shares an alternating architecture and benefits from both model-based methods and learning-based methods. On the one hand, SNMF-Net is of high physical interpretability as it is built by unrolling$L_{p}$sparsity constrained nonnegative matrix factorization ($L_{p}$-NMF) model belonging to LMM families. On the other hand, all the parameters and submodules of SNMF-Net can be seamlessly linked with the alternating optimization algorithm of$L_{p}$-NMF and unmixing problem. This enables us to reasonably integrate the prior knowledge on unmixing, the optimization algorithm, and the sparse representation theory into the network for robust learning, so as to improve unmixing. Experimental results on the synthetic and real-world data show the advantages of the proposed SNMF-Net over many state-of-the-art methods. Fengchao Xiong, Jun Zhou 0001, Shuyin Tao, Jianfeng Lu 0003, Yuntao Qian |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Learning a Deep Structural Subspace Across Hyperspectral Scenes With Cross-Domain VAEabstractHyperspectral image (HSI) classification is a small-sample-size problem due to the expensive cost of labeling. As a novel approach to this problem, cross-scene HSI classification has become a hot research topic in recent years. In cross-scene HSI classification, the scene containing enough labeled samples (called source scene) is used to benefit the classification in another scene containing a small number of training samples (called target scene). Transfer learning is a typical solution for cross-scene classification. However, many transfer learning algorithms assume an identical feature space for source and target scenes, which violates the fact that source and target scenes often lie in different feature spaces with various dimensions due to different HSI sensors. Aiming at the different feature spaces between the two scenes, we propose an end-to-end heterogeneous deep transfer learning algorithm, namely, cross-domain variational autoencoder (CDVAE). This algorithm is mainly composed of two key parts: 1) the features of the two scenes are embedded into the shared feature subspace through the two-stream variational autoencoder (VAE) to ensure that the output feature dimensions of the two scenes are identical and 2) graph regularization is used to establish the manifold constraints between source and target scenes in the shared subspace, so as to align the feature spaces. Experiments on two different cross-scene HSI datasets have proved the superior performance of the proposed CDVAE algorithm. Minchao Ye, Junbin Chen, Fengchao Xiong, Yuntao Qian |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | SMDS-Net: Model Guided Spectral-Spatial Network for Hyperspectral Image DenoisingabstractDeep learning (DL) based hyperspectral images (HSIs) denoising approaches directly learn the nonlinear mapping between noisy and clean HSI pairs. They usually do not consider the physical characteristics of HSIs. This drawback makes the models lack interpretability that is key to understanding their denoising mechanism and limits their denoising ability. In this paper, we introduce a novel model-guided interpretable network for HSI denoising to tackle this problem. Fully considering the spatial redundancy, spectral low-rankness, and spectral-spatial correlations of HSIs, we first establish a subspace-based multidimensional sparse (SMDS) model under the umbrella of tensor notation. After that, the model is unfolded into an end-to-end network named SMDS-Net, whose fundamental modules are seamlessly connected with the denoising procedure and optimization of the SMDS model. This makes SMDS-Net convey clear physical meanings, i.e., learning the low-rankness and sparsity of HSIs. Finally, all key variables are obtained by discriminative training. Extensive experiments and comprehensive analysis on synthetic and real-world HSIs confirm the strong denoising ability, strong learning capability, promising generalization ability, and high interpretability of SMDS-Net against the state-of-the-art HSI denoising methods. The source code and data of this article will be made publicly available at https://github.com/bearshng/smds-net for reproducible research. Fengchao Xiong, Jun Zhou 0001, Shuyin Tao, Jianfeng Lu 0003, Jiantao Zhou 0001, Yuntao Qian |
IEEE Trans. Image Process. | 1 |
| 2021 | NMF-SAE: An Interpretable Sparse Autoencoder for Hyperspectral UnmixingabstractHyperspectral unmixing is an important tool to learn the material constitution and distribution of a scene. Model-based unmixing methods depend on well-designed iterative optimization algorithms, which is usually time consuming. Learning-based methods perform unmixing in a data-driven manner but heavily rely on the quality and quantity of the training samples due to the lack of physical interpretability. In this paper, we combine the advantages of both model-based and learning-based methods and propose a nonnegative matrix factorization (NMF) inspired sparse autoencoder (NMF-SAE) for hyperspectral unmixing. NMF-SAE consists of an encoder and a decoder, both of which are constructed by unrolling the iterative optimization rules of L1sparsity-constrained NMF for the linear spectral mixture model. All parameters in our method are obtained by end-to-end training in a data-driven manner. Our network is not only physically interpretable and flexible but also has higher learning capacity with fewer parameters. Experimental results on both synthetic and real-world data demonstrate that our method is capable of producing desirable unmixing results when compared against several alternative approaches. Fengchao Xiong, Jun Zhou 0001, Minchao Ye, Jianfeng Lu 0003, Yuntao Qian |
ICASSP | 1 |
| 2021 | Learning a Model-Based Deep Hyperspectral Denoiser from a Single Noisy Hyperspectral ImageabstractHyperspectral image (HSI) denoising is a crucial preprocessing procedure to improve the quality of HSI. Model-based methods take the degradation model and the structure of underlying clean HSI into account for denoising but require a large number of numerical iterations and exhausting parameter tuning. Deep-learning-based (DL-based) methods directly learn the nonlinear transformation of clean and noisy image HSI pairs, but rely on large-scale high-quality training samples because of its “black box” denoising mechanism. In this paper, we propose a model-based DL method for HSI denoising to combine the advantages of model-based methods and DL-based methods. Specifically, we first build a HSI denoising model based on sparse representation. Then, we unfold the iterative optimization under the framework of gradient descent with momentum to yield a Gradient Momentum Sparse Coding Network (GMSC-Net) for denoising. In order to overcome the unavailability of noisy-clean HSI pairs for training, we directly learn GMSC-Net from a single HSI. The observed noisy HSI is grouped into a number of clusters containing local cubes. The cluster centers are treated as “clean” cubes and are polluted by noises, yielding a set of “noisy-clean” pairs for training. Extensive experiments show the effectiveness of our method on both synthetic and real-world datasets. Guanyiman Fu, Fengchao Xiong, Shuyin Tao, Jianfeng Lu 0003, Jun Zhou 0001, Yuntao Qian |
IGARSS | 2 |
| 2020 | BAE-Net: A Band Attention Aware Ensemble Network for Hyperspectral Object TrackingabstractHyperspectral videos contain images with a large number of light wavelength indexed bands that can facilitate material identification for object tracking. Most hyperspectral trackers use hand-crafted features rather than deep learning generated features for image representation due to limited training samples. To fill this gap, this paper introduces a band attention aware ensemble network (BAE-Net) for deep hyperspectral object tracking, which takes advantages of deep models trained on color videos for feature representation. Specifically, an autoencoder-like band attention block is introduced to learn the dependencies among bands and generate band-wise weights. Guided by these weights, hyperspectral images are then divided into a number of three-channel images. These three-channel images are fed into a deep color tracking network, producing several weak trackers. Finally, weak trackers are fused using ensemble learning for target location. Experimental results on hyperspectral datasets show the effectiveness and advantages of the proposed deep hyperspectral tracker. Zhuanfeng Li, Fengchao Xiong, Jun Zhou 0001, Jing Wang 0062, Jianfeng Lu 0003, Yuntao Qian |
ICIP | 2 |
| 2020 | Nonlocal Low-Rank Nonnegative Tensor Factorization for Hyperspectral UnmixingabstractHyperspectral unmixing decomposes hyperspectral images (HSI) into a collection of constituent materials or end-members and their fractions, i.e., abundances. Nonnegative tensor factorization (NTF) has been utilized thanks to its ability of preserving all the information in HSI. However, NTF based unmixing only makes use of global spatial-spectral information without considering detailed local/non-local spatial information, making it vulnerable to real-world disturbance such as noises. To this end, in this paper, we extend NTF by introducing non-local low-rank constraint to abundance maps. The additional regularization on abundances facilities tensor factorization avoid being trapped into a large number of suspicious solutions, so as to preserve the non-local spatial structure on abundance maps. Experimental results on synthetic data and real-world data show that the proposed method outperforms the state-of-the-art methods. Fengchao Xiong, Kun Qian 0015, Jianfeng Lu 0003, Jun Zhou 0001, Yuntao Qian |
IGARSS | 1 |
| 2020 | Spectral Mixture Model Inspired Network Architectures for Hyperspectral UnmixingabstractIn many statistical hyperspectral unmixing approaches, the unmixing task is essentially an optimization problem given a defined linear or nonlinear spectral mixture model. However, most of the model inference algorithms require a time-consuming iterative procedure. On the other hand, neural networks have been recently used to estimate abundances given some training samples, or directly estimate endmembers and abundances simultaneously in an unsupervised setting. However, their disadvantages are clear: lack of interpretability and reliance on the large training set. Model-inspired neural networks are constructed by the problem model and its corresponding inference algorithm. It incorporates the prior knowledge of physical model and algorithm into network architecture, combining the advantages of model-based and learning-based methods. This article deeply unfolds the linear mixture model and the corresponding iterative shrinkage-thresholding algorithm (ISTA) to build two unmixing network architectures. The first assumes that the set of endmembers are known, and the deep unfolded ISTA model is only for abundance estimation; and the second is used for blind unmixing to estimate both endmembers and abundances at the same time. The networks can be trained by supervised and unsupervised schemes, respectively, with a small-size training set, and then, unmixing becomes a feedforward process, which is very fast since no iteration is required. The experimental results show their competitive performance compared with the state-of-the-art unmixing approaches. Yuntao Qian, Fengchao Xiong, Qipeng Qian, Jun Zhou 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | Material Based Object Tracking in Hyperspectral VideosabstractTraditional color images only depict color intensities in red, green and blue channels, often making object trackers fail in challenging scenarios, e.g., background clutter and rapid changes of target appearance. Alternatively, material information of targets contained in large amount of bands of hyperspectral images (HSI) is more robust to these difficult conditions. In this paper, we conduct a comprehensive study on how material information can be utilized to boost object tracking from three aspects: dataset, material feature representation and material based tracking. In terms of dataset, we construct a dataset of fully-annotated videos, which contain both hyperspectral and color sequences of the same scene. Material information is represented by spectral-spatial histogram of multidimensional gradients, which describes the 3D local spectral-spatial structure in an HSI, and fractional abundances of constituted material components which encode the underlying material distribution. These two types of features are embedded into correlation filters, yielding material based tracking. Experimental results on the collected dataset show the potentials and advantages of material based object tracking. Fengchao Xiong, Jun Zhou 0001, Yuntao Qian |
IEEE Trans. Image Process. | 1 |
| 2019 | Deep Unfolded Iterative Shrinkage-Thresholding Model for Hyperspectral UnmixingabstractIn this paper, we propose a novel approach for spectral unmixing by unfolding the iterative shrinkage-thresholding algorithm (ISTA) into a deep neural network architecture. Spectral unmixing aims at identifying the endmembers and their fractional abundances in the mixed pixels. Once the endmembers are obtained as a dictionary, abundance estimation can be defined as a sparse coding problem with nonnegativity constraint. There are a number of iterative optimization algorithms for solving this problem, including ISTA, however, they always require hundreds and even thousands iterations, which is too slow for time-sensitive applications. In contrast, deep neural networks can approximate a finite closed-form expression to direct estimate abundances by learning from training samples, but they are closer to black-box mechanism rather than problem-level formulations. Deep unfolding constructs a deep neural network architecture inspired by the problem model and its corresponding optimization algorithm, which incorporates the prior knowledge of physical model and algorithm into network architecture. In this paper, the deep unfolded ISTA model is adopted for abundance estimation. It uses only a small training set to learn the model parameters, and then the abundance estimation become to be a feed-forward process in this model, which is very fast since no iteration is required. Qipeng Qian, Fengchao Xiong, Jun Zhou 0001 |
IGARSS | 2 |
| 2019 | Hyperspectral Unmixing via Total Variation Regularized Nonnegative Tensor FactorizationabstractHyperspectral unmixing decomposes a hyperspectral imagery (HSI) into a number of constituent materials and associated proportions. Recently, nonnegative tensor factorization (NTF)-based methods have been proposed for hyperspectral unmixing thanks to their capability in representing an HSI without any information loss. However, tensor factorization-based HSI processing approaches often suffer from low-signal-to-noise ratio condition of HSI and nonuniqueness of the solution. This problem can be effectively alleviated by introducing various spatial constraints into tensor factorization to suppress the noise and decrease the number of extreme, stationary, and saddle points. On the other hand, total variation (TV) adaptively promotes piecewise smoothness while preserving edges. In this paper, we propose a TV regularized matrix-vector NTF method. It takes advantage of tensor factorization in preserving global spectral-spatial information and the merits of TV in exploiting local spatial information, thus generating smooth abundance maps with preserved edges. Experimental results on synthetic and real-world data show that the proposed method outperforms the state-of-the-art methods. Fengchao Xiong, Yuntao Qian, Jun Zhou 0001, Yuan Yan Tang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | Hyperspectral Restoration via L0 Gradient Regularized Low-Rank Tensor FactorizationabstractDue to the mechanism of the data acquisition process, hyperspectral imagery (HSI) are usually contaminated by various noises, e.g., Gaussian noise, impulse noise, strips, and dead lines. In this article, a spectral-spatial L0gradient regularized low-rank tensor factorization (LRTFL0) method is proposed for hyperspectral denoising, in which the restored HSI is approximated by low-rank block term decomposition (BTD). BTD factorizes a tensor into the sum of a series of component tensors, each of which is represented by the outer product of a matrix and a vector. From subspace learning point of view, the vector and matrix can be considered as a spectral atom and its corresponding coding coefficients. In the proposed method, the correlations in both spectral and spatial domains are taken into account via the small size of atom set and low-rankness of coding matrices. In addition, HSIs also have the local structure of piecewise smoothness in both spectral and spatial domains. Motivated by the supreme virtues of L0gradient regularization in image structure exploitation, we develop a spectral-spatial L0gradient regularization and embed it into BTD to explore the spectral-spatial texture information. The proposed method can simultaneously remove various types of noises, and the experimental results on both synthetic data and real-world data show its superiority when compared with several state-of-the-art approaches. Fengchao Xiong, Jun Zhou 0001, Yuntao Qian |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2018 | Hyperspectral Imagery Denoising via Reweighed Sparse Low-Rank Nonnegative Tensor FactorizationabstractHyperspectral imagery (HSI) denoising is an important preprocessing step for real-world applications. Recently, sparse representation and low-rank representation based methods are proven effective in HSI denoising. However, most of these approaches only consider the low-rankness in the spectral domain and the sparsity in coding matrix. They have ignored the property that the coding matrix of each atom is also low-rank, i.e., low-rankness also exists in the spatial domain. In this paper, a reweighed sparse low-rank nonnegative tensor factorization (RSLRNTF) method is proposed to restore an HSI. It takes an HSI as a third-order tensor and factorizes it into the combination of a few component tensors where each one is the outer product of a low-rank matrix (coding matrix) and a vector (atom). Additionally, a reweighed L1 norm is added to coding matrices to enforce their sparsity. The low-rankness in both the spatial domain and the spectral domain as well as sparsity in the spatial domain improve the denoising performance. Furthermore, the nonnegativity in both coding matrices and dictionary leads to parts-based representation of HSI, which facilitates preserving local fine structure information. Experimental results on synthetic data and real-world data demonstrate the superiority of proposed method. Fengchao Xiong, Jun Zhou 0001, Yuntao Qian |
ICIP | 1 |
| 2018 | Superpixel-Based Nonnegative Tensor Factorization for Hyperspectral UnmixingabstractHyperspectral unmixing aims at decomposing a hyperspectral image (HSI) into a number of constituted materials and associated proportions. Recently, nonnegative tensor factorization (NTF) based methods have been proved effective and natural for hyperspectral unmixing owing to their virtue of representing an HSI without any information loss. However, these methods take an HSI as a whole, partly ignoring the local information in distinct local regions. In addition, HSIs are high likely to be disturbed by various noise, making the global information unnecessarily reliable. To alleviate these drawbacks, we propose a superpixel-based matrix-vector nonnegative tensor factorization (S-MV-NTF) method for hyperspectral unmixing, where both the global information and local information are taken into consideration. In this method, the HSI is firstly partitioned into numerous superpixels, homogeneous regions with adaptive sizes and compact boundaries, representing the local spatial structure information. Then, such local information is integrated to the tensor factorization to make the pixels lying in the same superpixel share similar abundances. Experimental results on synthetic data and real-world data show that the proposed method dominates the state-of-the-art methods. Fengchao Xiong, Jingzhou Chen, Jun Zhou 0001, Yuntao Qian |
IGARSS | 1 |
| 2017 | Matrix-Vector Nonnegative Tensor Factorization for Blind Unmixing of Hyperspectral ImageryabstractMany spectral unmixing approaches ranging from geometry, algebra to statistics have been proposed, in which nonnegative matrix factorization (NMF)-based ones form an important family. The original NMF-based unmixing algorithm loses the spectral and spatial information between mixed pixels when stacking the spectral responses of the pixels into an observed matrix. Therefore, various constrained NMF methods are developed to impose spectral structure, spatial structure, and spectral-spatial joint structure into NMF to enforce the estimated endmembers and abundances preserve these structures. Compared with matrix format, the third-order tensor is more natural to represent a hyperspectral data cube as a whole, by which the intrinsic structure of hyperspectral imagery can be losslessly retained. Extended from NMF-based methods, a matrix-vector nonnegative tensor factorization (NTF) model is proposed in this paper for spectral unmixing. Different from widely used tensor factorization models, such as canonical polyadic decomposition CPD) and Tucker decomposition, the proposed method is derived from block term decomposition, which is a combination of CPD and Tucker decomposition. This leads to a more flexible frame to model various application-dependent problems. The matrix-vector NTF decomposes a third-order tensor into the sum of several component tensors, with each component tensor being the outer product of a vector (endmember) and a matrix (corresponding abundances). From a formal perspective, this tensor decomposition is consistent with linear spectral mixture model. From an informative perspective, the structures within spatial domain, within spectral domain, and cross spectral-spatial domain are retreated interdependently. Experiments demonstrate that the proposed method has outperformed several state-of-the-art NMF-based unmixing methods. Yuntao Qian, Fengchao Xiong, Shan Zeng, Jun Zhou 0001, Yuan Yan Tang |
IEEE Trans. Geosci. Remote. Sens. | 2 |