EDBT 2026 Demo / reviewers in the wild / expert
Kai Jiang 0001
dblp:22/2361-1
· DBLP profile ↗
21ranked-venue papers
5as first author
20since 2021 · last 2026
0000-0001-9921-2043ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TOP-RL: Task-Optimized Progressive Token Pruning with Reinforcement Learning for Vision Language ModelsabstractIn recent years, Large Vision-Language Models (LVLMs) have significantly advanced multimodal tasks. However, their inference requires intensive processing of numerous visual tokens and incurs substantial computational overhead. Existing methods typically compress visual tokens either at the input stage or in early model layers, ignoring variations across tasks and depths. To address these limitations, we introduce TOP-RL, a Task-Optimized Progressive token pruning framework based on Reinforcement Learning. TOP-RL formulates visual token pruning as a multi-stage Markov Decision Process (MDP). It employs an agent trained with dense and fine-grained reward signals to progressively generate differentiable binary masks. This enables TOP-RL to adaptively select crucial visual tokens tailored to each task, effectively balancing accuracy and computational efficiency. Extensive experiments on leading multimodal datasets and advanced LVLMs validate that TOP-RL effectively learns task-optimized pruning policies, significantly boosting inference efficiency while preserving robust performance. For instance, LLaVA-NeXT equipped with TOP-RL achieves a 1.9x speedup in inference time and a 9.3x reduction in FLOPs, with 96% performance preserved. Hengyi Wang, Weiying Xie, Yaotao Wei, Kai Jiang 0001, Mingxiang Cao, Chenhe Hao, Leyuan Fang |
AAAI | 5 |
| 2025 | DiffCLIP: Few-shot Language-driven Multimodal ClassifierabstractVisual language models like Contrastive Language-Image Pretraining (CLIP) have shown impressive performance in analyzing natural images with language information. However, these models often encounter challenges when applied to specialized domains such as remote sensing due to the limited availability of image-text pairs for training. To tackle this issue, we introduce DiffCLIP, a novel framework that extends CLIP to effectively convey comprehensive language-driven semantic information for accurate classification of high-dimensional multimodal remote sensing images. DiffCLIP is a few-shot learning method that leverages unlabeled images for pretraining. It employs unsupervised mask diffusion learning to capture the distribution of diverse modalities without requiring labels. The modality-shared image encoder maps multimodal data into a unified subspace, extracting shared features with consistent parameters across modalities. A well-trained image encoder further enhances learning by aligning visual representations with class-label text information from CLIP. By integrating these approaches, DiffCLIP significantly boosts CLIP performance using a minimal number of image-text pairs. We evaluate DiffCLIP on widely used high-dimensional multimodal datasets, demonstrating its effectiveness in addressing few-shot annotated classification tasks. DiffCLIP achieves an overall accuracy improvement of 10.65% across three remote sensing datasets compared with CLIP, while utilizing only 2-shot image-text pairs. Mingxiang Cao, Kai Jiang 0001, Yunsong Li 0001 |
AAAI | 4 |
| 2025 | Visual Instruction Tuning towards General-Purpose Multimodal Large Language Model: A Survey
Jiaxing Huang 0001, Jingyi Zhang 0005, Kai Jiang 0001, Han Qiu 0008, Xiaoqin Zhang 0002, Ling Shao 0001, Shijian Lu, Dacheng Tao |
Int. J. Comput. Vis. | 3 |
| 2025 | Hyperspectral anomaly detection with self-supervised anomaly prior
Yidan Liu, Kai Jiang 0001, Weiying Xie, Yunsong Li 0001, Leyuan Fang |
Neural Networks | 2 |
| 2025 | M³amba: CLIP-Driven Mamba Model for Multi-Modal Remote Sensing ClassificationabstractMulti-modal fusion holds great promise for integrating information from different modalities. However, due to a lack of consideration for modal consistency, existing multi-modal fusion methods in the field of remote sensing still face challenges of incomplete semantic information and low computational efficiency in their fusion designs. Inspired by the observation that the visual language pre-training model CLIP can effectively extract strong semantic information from visual features, we propose M3amba, a novel end-to-end CLIP-driven Mamba model for multi-modal fusion to address these challenges. Specifically, we introduce CLIP-driven modality-specific adapters in the fusion architecture to avoid the bias of understanding specific domains caused by direct inference, making the original CLIP encoder modality-specific perception. This unified framework enables minimal training to achieve a comprehensive semantic understanding of different modalities, thereby guiding cross-modal feature fusion. To further enhance the consistent association between modality mappings, a multi-modal Mamba fusion architecture with linear complexity and a cross-attention module Cross-SS2D are designed, which fully considers effective and efficient information interaction to achieve complete fusion. Extensive experiments have shown that M3amba has an average performance improvement of at least 5.98% compared with the state-of-the-art methods in multi-modal hyperspectral image classification tasks in the remote sensing field, while also demonstrating excellent training efficiency, achieving a double improvement in accuracy and efficiency. The code is released athttps://github.com/kaka-Cao/M3amba. Mingxiang Cao, Weiying Xie, Xin Zhang 0092, Kai Jiang 0001, Jie Lei 0001, Yunsong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | SHAA: Spatial Hybrid Attention Network With Adaptive Cross-Entropy Loss Function for UAV-View Geo-LocalizationabstractCross-view geo-localization provides an offline visual positioning strategy for unmanned aerial vehicles (UAVs) in Global Navigation Satellite System (GNSS)-denied environments. However, it still faces the following challenges, leading to suboptimal localization performance: 1) Existing methods primarily focus on extracting global features or local features by partitioning feature maps, neglecting the exploration of spatial information, which is essential for extracting consistent feature representations and aligning images of identical targets across different views. 2) Cross-view geo-localization encounters the challenge of data imbalance between UAV and satellite images. To address these challenges, the Spatial Hybrid Attention Network with Adaptive Cross-Entropy Loss Function (SHAA) is proposed. To tackle the first issue, the Spatial Hybrid Attention (SHA) method employs a Spatial Shift-MLP (SSM) to focus on the spatial geometric correspondences in feature maps across different views, extracting both global features and fine-grained features. Additionally, the SHA method utilizes a Hybrid Attention (HA) mechanism to enhance feature extraction diversity and robustness by capturing interactions between spatial and channel dimensions, thereby extracting consistent cross-view features and aligning images. For the second challenge, the Adaptive Cross-Entropy (ACE) loss function incorporates adaptive weights to emphasize hard samples, alleviating data imbalance issues and improving training effectiveness. Extensive experiments on widely recognized benchmarks, including University-1652, SUES-200, and DenseUAV, demonstrate that SHAA achieves state-of-the-art performance, outperforming existing methods by over 3.92%. Nanhua Chen, Dongshuo Zhang, Kai Jiang 0001, Yeqing Zhu, Tai-Shan Lou, Liangyu Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Hyperspectral Target Detection Based on Generative Self-Supervised Learning With Wavelet TransformabstractRecently, generative self-supervised learning (GSSL) has gained extensive attention in hyperspectral remote sensing. For the hyperspectral target detection (HTD) task, traditional GSSL-based algorithms usually require hyperspectral images (HSIs) as additional datasets for pretraining, which are relatively resource-intensive and time-consuming. To better interpret the spectral-spatial information of HSIs while alleviating the dependence on large-scale hyperspectral datasets, we develop a novel two-stage framework for HTD based on GSSL in this article. In the preprocessing for the input HSI, a dimensional transformation (DT) module and a coarse detection reference (CDR) module are constructed to produce feature patches as training samples for subsequent pretraining and fine-tuning. In the pretraining stage for spectral-spatial reconstruction, we construct an asymmetric autoencoder (AE) architecture which leverages the transformer blocks with long-range perception to extract generalized features and explore discriminative feature representations of the input HSI. Specifically, a dual-stream wavelet patch embedding (DWPE) module is proposed to integrate the wavelet transform (WT) mechanism with the convolutional neural networks (CNNs), which extracts robust spectral-spatial features by performing convolutional operations with different frequency components of WT. In the fine-tuning stage, a novel signature-constrained cross-entropy (SC-CE) loss function is proposed to constrain the network optimization. For the final detection, a pixel-level fusion based on coarse detection based pixel-level fusion (CDPF) module is employed after inference to further suppress the interference from background. Experimental results on six real HSIs demonstrate that the proposed method achieves superior detection performance while maintaining the generalization of the pretrained model. Shuai Wang 0057, Yunsong Li 0001, Weiying Xie, Kai Jiang 0001, Kailang Cao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | A Signature-Constrained Two-Stage Framework for Hyperspectral Target Detection Based on Generative Self-Supervised LearningabstractRecently, hyperspectral target detection (HTD) technique based on deep learning (DL) has been developed rapidly. However, existing algorithms show poor generalization across different hyperspectral images (HSIs), where repeated training and inference based on limited prior information are necessary. To liberate HTD from dependence on the quantity and the quality of training samples, this article proposes a signature-constrained two-stage framework for HTD (HTD-STF) based on generative self-supervised learning (GSSL). In the first stage for pre-training, to realize spectral-spatial reconstruction, we build an asymmetric autoencoder (AE) employing transformer blocks with long-range perception for generalized feature extraction. During pre-training, the spectral-spatial similarity loss is designed to improve the effect of reconstruction. In the second stage for fine-tuning and detection, the signature is utilized in preprocessing, training and inference, respectively. Specifically, the coarse sample mining and tiling strategy in preprocessing not only facilitates the framework in flexible input dimension, but also provides pseudo labels for end-to-end training. During training, we adopt the signature as guidance for feature-level fusion, which alleviates the impact of sample imbalance. After training, the final inference based on pixel-level fusion refines the original output. For ideal GSSL, the HyperMix-10K, a new large-scale hyperspectral dataset, has been constructed in this work, which contains numerous unlabeled HSIs captured in various scenes. Experimental results and analysis on real HSIs verify the effectiveness and generalization ability of the HTD-STF. Shuai Wang 0057, Yunsong Li 0001, Weiying Xie, Kai Jiang 0001, Kailang Cao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | DA-BEV: Unsupervised Domain Adaptation for Bird's Eye View Perception
Kai Jiang 0001, Jiaxing Huang 0001, Weiying Xie, Jie Lei 0001, Yunsong Li 0001, Ling Shao 0001, Shijian Lu |
ECCV (82) | 1 |
| 2024 | Open-Vocabulary Object Detection via Language HierarchyabstractRecent studies on generalizable object detection have attracted increasing attention with additional weak supervision from large-scale datasets with image-level labels.
However, weakly-supervised detection learning often suffers from image-to-box label mismatch, i.e., image-level
labels do not convey precise object information.
We design Language Hierarchical Self-training (LHST) that introduces language hierarchy into weakly-supervised detector training for learning more generalizable detectors.
LHST expands the image-level labels with language hierarchy and enables co-regularization between the expanded labels and self-training. Specifically, the expanded labels regularize self-training by providing richer supervision and mitigating the image-to-box label mismatch, while self-training allows assessing and selecting the expanded labels according to the predicted reliability.
In addition, we design language hierarchical prompt generation that introduces language hierarchy into prompt generation which helps bridge the vocabulary gaps between training and testing.
Extensive experiments show that the proposed techniques achieve superior generalization performance consistently across 14 widely studied object detection datasets. Jiaxing Huang 0001, Jingyi Zhang 0005, Kai Jiang 0001, Shijian Lu |
NeurIPS | 3 |
| 2024 | Domain Adaptation for Large-Vocabulary Object DetectorsabstractLarge-vocabulary object detectors (LVDs) aim to detect objects of many categories, which learn super objectness features and can locate objects accurately while applied to various downstream data. However, LVDs often struggle in recognizing the located objects due to domain discrepancy in data distribution and object vocabulary. At the other end, recent vision-language foundation models such as CLIP demonstrate superior open-vocabulary recognition capability.
This paper presents KGD, a Knowledge Graph Distillation technique that exploits the implicit knowledge graphs (KG) in CLIP for effectively adapting LVDs to various downstream domains.
KGD consists of two consecutive stages: 1) KG extraction that employs CLIP to encode downstream domain data as nodes and their feature distances as edges, constructing KG that inherits the rich semantic relations in CLIP explicitly;
and 2) KG encapsulation that transfers the extracted KG into LVDs to enable accurate cross-domain object classification.
In addition, KGD can extract both visual and textual KG independently, providing complementary vision and language knowledge for object localization and object classification in detection tasks over various downstream domains.
Experiments over multiple widely adopted detection benchmarks show that KGD outperforms the state-of-the-art consistently by large margins.
Codes will be released. Kai Jiang 0001, Jiaxing Huang 0001, Weiying Xie, Jie Lei 0001, Yunsong Li 0001, Ling Shao 0001, Shijian Lu |
NeurIPS | 1 |
| 2024 | Distribution-Aware Interactive Attention Network and Large-Scale Cloud Recognition Benchmark on FY-4A Satellite ImageabstractAccurate cloud recognition and warning are crucial for various applications, including in-flight support, weather forecasting, and climate research. However, recent deep learning algorithms have predominantly focused on detecting cloud regions in satellite imagery, with insufficient attention to the specificity required for accurate cloud recognition. This limitation inspired us to develop the novel FY-4A-Himawari-8 (FYH) dataset, which includes nine distinct cloud categories and uses precise domain adaptation methods to align 70419 image-label pairs (including 110000 train/5500 test$100\times 100$size images) in terms of projection, temporal resolution, and spatial resolution, thereby facilitating the training of supervised deep learning networks. Given the complexity and diversity of cloud formations, we have thoroughly analyzed the challenges inherent to cloud recognition tasks, examining the intricate characteristics and distribution of the data. To effectively address these challenges, we designed a distribution-aware interactive-attention network (DIAnet), which preserves pixel-level details through a high-resolution branch and a parallel multiresolution cross-branch. We also integrated a distribution-aware loss (DAL) to mitigate the imbalance across cloud categories. An interactive attention module (IAM) further enhances the robustness of feature extraction combined with spatial and channel information. Empirical evaluations on the FYH dataset demonstrate that our method outperforms other cloud recognition networks, achieving superior performance in terms of mean intersection over union (mIoU). The code for implementing DIAnet is available athttps://github.com/icey-zhang/DIAnet. Jie Lei 0001, Weiying Xie, Kai Jiang 0001, Xin Zhang 0092, Mingxiang Cao, Yunsong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Block-Wise Partner Learning for Model CompressionabstractDespite the great potential of convolutional neural networks (CNNs) in various tasks, the resource-hungry nature greatly hinders their wide deployment in cost-sensitive and low-powered scenarios, especially applications in remote sensing. Existing model pruning approaches, implemented by a "subtraction" operation, impose a performance ceiling on the slimmed model. Self-knowledge distillation (Self-KD) resorts to auxiliary networks that are only active in the training phase for performance improvement. However, the knowledge is holistic and crude, and the learning-based knowledge transfer is mediate and lossy. Here, we propose a novel model-compression method, termed block-wise partner learning (BPL), which comprises "extension" and "fusion" operations and liberates the compressed model from the bondage of baseline. Different from the Self-KD, the proposed BPL creates a partner for each block for performance enhancement in training. For the model to absorb more diverse information, a diversity loss (DL) is designed to evaluate the difference between the original block and the partner. Besides, the partner is fused equivalently instead of being discarded directly. After training, we can simply adopt the fused compressed model that contains the enhancement information of partners but with fewer parameters and less inference cost. As validated using the UC Merced land-use, NWPU-RESISC45, and RSD46-WHU datasets, the BPL demonstrates superiority over other compared model-compression approaches. For example, it attains a substantial floating-point operations (FLOPs) reduction of 73.97% with only 0.24 accuracy (ACC.) loss for ResNet-50 on the UC Merced land-use dataset. The code is available at https://github.com/zhangxin-xd/BPL. Xin Zhang 0092, Weiying Xie, Yunsong Li 0001, Jie Lei 0001, Kai Jiang 0001, Leyuan Fang, Qian Du 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Toward Stable, Interpretable, and Lightweight Hyperspectral Super-ResolutionabstractFor real applications, existing HSI-SR methods are not only limited to unstable performance under unknown scenarios but also suffer from high computation consumption. In this paper, we develop a new coordination optimization framework for stable, interpretable, and lightweight HSI-SR. Specifically, we create a positive cycle between fusion and degradation estimation under a new probabilistic framework. The estimated degradation is applied to fusion as guidance for a degradation-aware HSI-SR. Under the framework, we establish an explicit degradation estimation method to tackle the indeterminacy and unstable performance caused by the black-box simulation in previous methods. Considering the interpretability in fusion, we integrate spectral mixing prior into the fusion process, which can be easily realized by a tiny autoencoder, leading to a dramatic release of the computation burden. Based on the spectral mixing prior, we then develop a partial fine-tune strategy to reduce the computation cost further. Comprehensive experiments demonstrate the superiority of our method against the state-of-the-arts under synthetic and real datasets. For instance, we achieve a 2.3 dB promotion on PSNR with$120\times$model size reduction and$4300 \times$FLOPs reduction under the CAVE dataset. Code is available in https://github.com/WenjinGuo/DAEM. Wen-jin Guo, Weiying Xie, Kai Jiang 0001, Yunsong Li 0001, Jie Lei 0001, Leyuan Fang |
CVPR | 3 |
| 2023 | Weakly supervised adversarial learning via latent space for hyperspectral target detection
Weiying Xie, Yunsong Li 0001, Kai Jiang 0001, Jie Lei 0001, Qian Du 0001 |
Pattern Recognit. | 4 |
| 2023 | A Model-Driven Deep Mixture Network for Robust Hyperspectral Anomaly DetectionabstractHyperspectral anomaly detection (HAD) aims to identify samples with unknown atypical spectra from the background. Deep learning (DL)-based methods, particularly autoencoders (AEs), have proven effective in uncovering the underlying profiles for HAD. However, in real-world applications of hyperspectral images (HSIs), complex background land-covers and anomaly corruptions are common, leading to two issues: 1) A low-dimensional manifold characterized by DL-based HAD methods can only reveal a few underlying variation factors of the background distribution and cannot capture the complex structures behind land-covers of all categories. 2) DL-based HAD methods trained on anomaly-contaminated HSIs tend to overfit specific anomalies, resulting in poor background characterization. To tackle these issues, this study presents a novel and robust framework for HAD called Model-Driven Deep Mixture Network (MDMN) that combines the strengths of model-driven and data-driven approaches while emphasizing interpretability. By assuming that the background, consisting of various land-covers, arises from a mixture of low-dimensional manifolds, the MDMN incorporates a novel deep mixture module to comprehensively characterize the background. This module utilizes a low-dimensional manifold learned by an AE to represent a specific category of background land-covers. To mitigate the impact of anomaly corruptions, the MDMN incorporates a convex relaxation of a sparse constraint, which helps prevent overfitting anomalies. Extensive experimental results demonstrate that the proposed MDMN offers more satisfactory and robust detection performance. Yunsong Li 0001, Kai Jiang 0001, Weiying Xie, Jie Lei 0001, Xin Zhang 0092, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | REAF: Remembering Enhancement and Entropy-Based Asymptotic Forgetting for Filter PruningabstractNeurologically, filter pruning is a procedure of forgetting and remembering recovering. Prevailing methods directly forget less important information from an unrobust baseline at first and expect to minimize the performance sacrifice. However, unsaturated base remembering imposes a ceiling on the slimmed model leading to suboptimal performance. And significantly forgetting at first would cause unrecoverable information loss. Here, we design a novel filter pruning paradigm termed Remembering Enhancement and Entropy-based Asymptotic Forgetting (REAF). Inspired by robustness theory, we first enhance remembering by over-parameterizing baseline with fusible compensatory convolutions which liberates pruned model from the bondage of baseline at no inference cost. Then the collateral implication between original and compensatory filters necessitates a bilateral-collaborated pruning criterion. Specifically, only when the filter has the largest intra-branch distance and its compensatory counterpart has the strongest remembering enhancement power, they are preserved. Further, Ebbinghaus curve-based asymptotic forgetting is proposed to protect the pruned model from unstable learning. The number of pruned filters is increasing asymptotically in the training procedure, which enables the remembering of pretrained weights gradually to be concentrated in the remaining filters. Extensive experiments demonstrate the superiority of REAF over many state-of-the-art (SOTA) methods. For example, REAF removes 47.55% FLOPs and 42.98% parameters of ResNet-50 only with 0.98% TOP-1 accuracy loss on ImageNet. The code is available at https://github.com/zhangxin-xd/REAF. Xin Zhang 0092, Weiying Xie, Yunsong Li 0001, Kai Jiang 0001, Leyuan Fang |
IEEE Trans. Image Process. | 4 |
| 2022 | E2E-LIADE: End-to-End Local Invariant Autoencoding Density Estimation Model for Anomaly Target Detection in Hyperspectral ImageabstractHyperspectral anomaly target detection (also known as hyperspectral anomaly detection (HAD)] is a technique aiming to identify samples with atypical spectra. Although some density estimation-based methods have been developed, they may suffer from two issues: 1) separated two-stage optimization with inconsistent objective functions makes the representation learning model fail to dig out characterization customized for HAD and 2) incapability of learning a low-dimensional representation that preserves the inherent information from the original high-dimensional spectral space. To address these problems, we propose a novel end-to-end local invariant autoencoding density estimation (E2E-LIADE) model. To satisfy the assumption on the manifold, the E2E-LIADE introduces a local invariant autoencoder (LIA) to capture the intrinsic low-dimensional manifold embedded in the original space. Augmented low-dimensional representation (ALDR) can be generated by concatenating the local invariant constrained by a graph regularizer and the reconstruction error. In particular, an end-to-end (E2E) multidistance measure, including mean-squared error (MSE) and orthogonal projection divergence (OPD), is imposed on the LIA with respect to hyperspectral data. More important, E2E-LIADE simultaneously optimizes the ALDR of the LIA and a density estimation network in an E2E manner to avoid the model being trapped in a local optimum, resulting in an energy map in which each pixel represents a negative log likelihood for the spectrum. Finally, a postprocessing procedure is conducted on the energy map to suppress the background. The experimental results demonstrate that compared to the state of the art, the proposed E2E-LIADE offers more satisfactory performance. Kai Jiang 0001, Weiying Xie, Jie Lei 0001, Zan Li 0001, Yunsong Li 0001, Tao Jiang 0031, Qian Du 0001 |
IEEE Trans. Cybern. | 1 |
| 2021 | LREN: Low-Rank Embedded Network for Sample-Free Hyperspectral Anomaly DetectionabstractHyperspectral anomaly detection (HAD) is a challenging task because it explores the intrinsic structure of complex high-dimensional signals without any samples at training time. Deep neural networks (DNNs) can dig out the underlying distribution of hyperspectral data but are limited by the labeling of large-scale hyperspectral datasets, especially the low spatial resolution of hyperspectral data, which makes labeling more difficult. To tackle this problem while ensuring the detection performance, we present an unsupervised low-rank embedded network (LREN) in this paper. LREN is a joint learning network in which the latent representation is specifically designed for HAD, rather than merely as a feature input for the detector. And it searches the lowest rank representation based on a representative and discriminative dictionary in the deep latent space to estimate the residual efficiently. Considering the physically mixing properties in hyperspectral imaging, we develop a trainable density estimation module based on Gaussian mixture model (GMM) in the deep latent space to construct a dictionary that can better characterize the complex hyperspectral images (HSIs). The closed-form solution of the proposed low-rank learner surpasses existing approaches on four real hyperspectral datasets with different anomalies. We argue that this unified framework paves a novel way to combine feature extraction and anomaly estimation-based methods for HAD, which intends to learn the underlying representation tailored for HAD without the prerequisite of manually labeled data. Code available at https://github.com/xdjiangkai/LREN. Kai Jiang 0001, Weiying Xie, Jie Lei 0001, Tao Jiang 0031, Yunsong Li 0001 |
AAAI | 1 |
| 2021 | PTGAN: A Proposal-Weighted Two-Stage GAN with Attention for Hyperspectral Target DetectionabstractIn this paper, a proposal-weighted two-stage generative adversarial network (GAN) with attention mechanism is proposed for hyperspectral target detection (HTD). PTGAN leverages GAN to estimate spectral background distribution and realize mapping from the latent space to the spectral space. Meanwhile, PTGAN conducts the reversed mapping through latent-spectral-latent and spectral-latent-spectral learning. On this basis, PTGAN implements accurate reconstruction of background spectrum via latent space. Therefore, targets of interest can be detected through larger pixel-level reconstruction error. In particular, the variance attention module is designed to make full use of global information among spectral bands to selectively emphasize channel-wise spectral features. Furthermore, a proposal-weighted strategy in a two-stage manner reduces the false alarm of detection by refining the previous detection proposal. Finally, exponential nonlinear fusion combines the discriminative feature from two stages to suppress the background. Extensive experiments on two real hyperspectral images (HSIs) verify the effectiveness of PTGAN. Weiying Xie, Yunsong Li 0001, Kai Jiang 0001, Jie Lei 0001, Qian Du 0001 |
IGARSS | 4 |
| 2020 | Semisupervised Spectral Learning With Generative Adversarial Network for Hyperspectral Anomaly DetectionabstractLimited by the anomalous spectral vectors in unlabeled hyperspectral images (HSIs), anomaly detection methods based on background distribution estimation often suffer from the contamination of anomalies, which decreases the estimation accuracy and, thus, weakens the detection performance. To address this problem, we proposed a novel semisupervised spectral learning (SSL) for the hyperspectral anomaly detection framework based on the generative adversarial network (GAN). GAN is applied and developed to estimate the background distribution in a semisupervised manner and obtain an initial spectral feature because of its strong representational capability and adversarial training advantage. In the proposed framework, an initial spatial feature is generated via morphological attribute filtering. Finally, an exponential constrained nonlinear suppression fusion technique is adopted to suppress the background and combine the complementary information in different features to obtain a fused detection map. The performance of the proposed anomaly detection technique is evaluated on a series of HSIs. Experimental results demonstrate that our method can outperform state-of-the-art anomaly detection methods. Kai Jiang 0001, Weiying Xie, Yunsong Li 0001, Jie Lei 0001, Gang He 0002, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |