Jun Yue 0004

dblp:39/2914-4 · DBLP profile ↗
← Back
31ranked-venue papers
4as first author
30since 2021 · last 2026
0000-0002-6465-5052ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 21 · 2 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021
YearPublicationVenuePosition
2026 Generating Any Changes in the Noise Domain
abstract
Change detection is essential in Earth observation, yet current models heavily rely on large-scale annotated datasets. Generative models offer a promising alternative by synthesizing training data, but generating temporally coherent image pairs with realistic, semantically meaningful changes remains a significant challenge. Existing approaches typically simulate changes by generating pre- and post-change label maps using either heuristic rules (e.g., copy-pasting) or text prompts. However, the former offers limited change diversity, while the latter often fails to maintain spatial consistency between image pairs. We observe that the noise space of diffusion models encodes strong generative capacity and spatial controllability: localized perturbations in the noise can yield meaningful, interpretable changes in corresponding image regions. Motivated by this, we propose Noise2Change, a framework for simulating change directly in the noise domain. The key idea is to manipulate the semantic composition of the initial noise sampled from the noise domain, such that the diffusion process generates structurally consistent pre- and post-change images reflecting realistic transformations. Since the unperturbed noise is shared between both images, the resulting pairs exhibit strong temporal alignment and semantic coherence, effectively addressing the trade-off between realism and consistency. Concretely, we employ a discrete diffusion model to extract high-level semantics from the initial noise. Guided by these semantics, we introduce a change simulation strategy that optimizes the noise to encode intended changes. The modified noise is then used to drive the diffusion process, yielding pre- and post-change label maps with natural structural transitions. These maps are passed through a unified framework for image generation and label refinement, producing highly aligned image-label pairs. Our framework supports diverse change types across a wide range of scenarios. Extensive experiments on multiple change detection tasks demonstrate that our method achieves superior performance compared to existing generative approaches.
Jun Yue 0004, Pedram Ghamisi, Weiying Xie, Leyuan Fang
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Beyond dimensionality explosion: A latent diffusion framework for hyperspectral image classification
Meiyun Lu, Xia Yue, Yicong Li 0015, Jun Yue 0004, Leyuan Fang
Neurocomputing5
2025 HyperEDL: Spectral-Spatial Evidence Deep Learning for Cross-Scene Hyperspectral Image Classification
abstract
Cross-scene hyperspectral image (HSI) classification presents significant challenges due to domain shifts, which amplify epistemic uncertainty and lead to substantial performance drops in unseen scenes. While evidence deep learning (EDL) has shown promise in modeling uncertainty, existing methods fall short, as they do not explicitly account for the epistemic uncertainty arising from spatial-spectral feature interactions. To address these challenges, we propose the spectral-spatial evidence deep learning for cross-scene hyperspectral image classification (HyperEDL) framework, which introduces the spatial-spectral multiorder aggregation module (SS-Moga). This module effectively captures and adaptively encodes multiorder contextual interactions from both spatial and spectral perspectives. By combining multiorder contextual encoding with spatial-spectral confidence, our approach fully aggregates multiorder evidence to mitigate epistemic uncertainty arising from knowledge gaps between seen and unseen scenes. Specifically, it uses Dirichlet distribution to capture correlation between spatial-spectral knowledge about different scenes, which can be generalized to unseen scenes. Extensive experiments on three benchmark datasets demonstrate that HyperEDL outperforms state-of-the-art methods, showcasing its effectiveness and strong generalization ability.
Yangbo Feng, Shuhe Wang, Jun Yue 0004, Shaobo Xia, Leyuan Fang
IEEE Trans. Geosci. Remote. Sens.4
2025 Exemplar-Free Lifelong Hyperspectral Image Classification With Spectral Consistency
abstract
Hyperspectral image (HSI) classification suffers from severe catastrophic forgetting in exemplar-free lifelong learning, where models must continuously learn new land cover categories without accessing historical training samples. This challenge persists due to high-dimensional spectral-volumetric complexity and cross-task spectral drift, which current methods inadequately address. We propose HyperSC, a novel framework that synergizes spectral-consistent auxiliary samples synthesis with stability-plasticity fused learning. The framework consists of three key components: a Spectral Consistency Model Inversion (SCMI) module, a Spectral Progressive Enhancement (SPE) module, and a Fusion Distillation Learning (FDL) module. The SCMI module synthesizes class-conditional auxiliary samples through spectral moment matching, in which the mean and variance of each spectral band are constrained to match class-specific real historical data distributions, thereby achieving spectral consistency. The SPE module injects class-specific Gaussian noise and applies momentum-based updating to enhance sample diversity while preserving spectral fidelity. The FDL module jointly trains on fused real and auxiliary samples by coordinating cross-entropy classification, output-layer knowledge distillation, and intermediate-layer feature alignment, thereby enabling plasticity for new class learning while maintaining stability against catastrophic forgetting for previous tasks. Extensive experiments on three HSI datasets (Indian Pines, Houston, Salinas) demonstrate HyperSC’s superiority compared to previous exemplar-free lifelong learning methods. The code is available at https://github.com/lzlsxs/hypersc.
Zhenlin Li, Shaobo Xia, Shuhe Wang, Jun Yue 0004, Leyuan Fang
IEEE Trans. Geosci. Remote. Sens.4
2025 HyperKD: Lifelong Hyperspectral Image Classification With Cross-Spectral-Spatial Knowledge Distillation
abstract
Hyperspectral image (HSI) classification models suffer from a phenomenon known as catastrophic forgetting, which refers to the sharp decline in performance on previously learned tasks after learning a new one when continuously acquiring new knowledge from a sequence of tasks. In recent years, some lifelong learning approaches have been proposed for HSI classification. Despite some progress, the challenge of catastrophic forgetting in lifelong learning remains significant and unresolved. In this article, we propose a novel lifelong learning framework for HSI classification, which is based on exemplar replay and cross-spectral–spatial feature knowledge distillation (KD), termed HyperKD. Specifically, the proposed framework incorporates a min-max cross-selection (MMCS) module tailored to HSI characteristics with a cross-spectral-spatial knowledge distillation (CSSKD) module. The MMCS module selects the most representative or diverse samples as exemplars from previous tasks for replay. Additionally, the CSSKD module not only transfers the prediction logit distribution from the previous network to the current network but also transfers the spectral-spatial feature distribution via cross-network KD, without directly assessing the similarity of feature distributions, thereby retaining more knowledge and mitigating forgetting. Through experiments conducted on a series of tasks, including the Pavia, Indian Pines, Salinas, and Houston datasets, our approach demonstrates superior performance compared to previous lifelong learning methods for HSI classification, effectively mitigating catastrophic forgetting. The code implementation of our approach will be publicly available athttps://github.com/lzlsxs/hyperkd.
Zhenlin Li, Shaobo Xia, Jun Yue 0004, Leyuan Fang
IEEE Trans. Geosci. Remote. Sens.3
2024 Open Set Recognition in Real World
Zhen Yang 0026, Jun Yue 0004, Pedram Ghamisi, Shiliang Zhang, Jiayi Ma 0001, Leyuan Fang
Int. J. Comput. Vis.2
2024 TAKD: Target-Aware Knowledge Distillation for Remote Sensing Scene Classification
abstract
Remote sensing (RS) scene classification based on deep neural networks (DNNs) has recently drawn remarkable attention. However, the DNNs contain a great number of parameters and require a huge amount of computational costs, which are hard to deploy on edge devices such as onboard embedded systems. To address this issue, in this paper, we propose a target-aware knowledge distillation (TAKD) method for RS scene classification. By considering the characteristics among the target and background regions of the RS images, the TAKD can adaptively distill the knowledge from the teacher model to create a lightweight student model. Specifically, we first introduce a target extraction module that utilizes heatmaps to highlight target regions on the teacher’s feature maps. Next, we propose an adaptive fusion module that aggregates these heatmaps to capture objects with varying scales. Finally, we design a target-aware loss that enables the transfer of knowledge in the target regions from the teacher model to the student model, greatly reducing background disturbance. Our distillation scheme that does not require extra learning parameters is both simple and effective, significantly improving the accuracy of the student model without any additional computational or resource costs. Our experiments on three benchmark datasets demonstrate that our proposed TAKD outperforms the existing state-of-the-art distillation methods.
Jie Wu 0035, Leyuan Fang, Jun Yue 0004
IEEE Trans. Circuits Syst. Video Technol.3
2024 Spectral Query Spatial: Revisiting the Role of Center Pixel in Transformer for Hyperspectral Image Classification
abstract
Recently, there have been significant advancements in Hyperspectral Image (HSI) classification methods employing Transformer architectures. However, these methods, while extracting spectral-spatial features, may introduce irrelevant spatial information that interferes with HSI classification. To address this issue, this paper proposes a Spectral Query Spatial Transformer (SQSFormer) framework. The proposed framework utilizes the center pixel (i.e., pixel to be classified) to adaptively query relevant spatial information from neighboring pixels, thereby preserving spectral features while reducing the introduction of irrelevant spatial information. Specifically, this paper introduces a Rotation-Invariant Position Embedding module to integrate random central rotation and center relative position embedding, mitigating the interference of absolute position and orientation information on spatial feature extraction. Moreover, a Spectral-Spatial Center Attention module is designed to enable the network to focus on the center pixel by adaptively extracting spatial features from neighboring pixels at multiple scales. The pivotal characteristic of the proposed framework achieves adaptive spectral-spatial information fusion using the Spectral Query Spatial paradigm, reducing the introduction of irrelevant information and effectively improving classification performance. Experimental results on multiple public datasets demonstrate that our framework outperforms previous state-of-the-art methods. For the sake of reproducibility, the source code of SQSFormer will be publicly available at https://github.com/chenning0115/SQSFormer.
Leyuan Fang, Shaobo Xia, Hui Liu 0041, Jun Yue 0004
IEEE Trans. Geosci. Remote. Sens.6
2024 SVAFormer: Integrating Random and Hierarchical Spectral View Attention for Hyperspectral Image Classification
abstract
Recently, hyperspectral image (HSI) classification methods based on Transformers have developed rapidly. However, these methods still face challenges in handling the widely varying scales and diverse spatial distribution patterns commonly found in HSIs. To address these issues, this article proposes a simple, yet novel HSI classification framework named the spectral view attention Transformer (SVAFormer). Built on the Transformer mechanism, this framework enhances the integration of spectral and spatial features by allowing the spectral token, corresponding to the pixel to be classified, to access spatial neighborhood information from multiple perspectives and levels. Specifically, the framework employs random masking techniques to provide spectral tokens with spatial neighborhood information from different viewpoints, enabling the model to handle diverse land-cover distribution patterns. Additionally, the framework introduces a spectral token-aware pooling layer between adjacent Transformer blocks, which preserves the central role of spectral tokens while progressively expanding the spatial scale represented by each token. This reduces the Transformer’s focus on spatially fragmented information and enables spectral tokens to concentrate on spatial neighborhood information at various levels and scales. The key characteristic of this framework is its ability to effectively handle land-cover features of different scales and shapes by strengthening the fusion of spectral and spatial characteristics. Experimental results on multiple public datasets demonstrate that our framework outperforms previous state-of-the-art methods. For the sake of reproducibility, the source code of SVAFormer will be publicly available athttps://github.com/chenning0115/SVAFormer.
Zhou Huang 0002, Xia Yue, Anfeng Liu, Meiyun Lu, Jun Yue 0004, Leyuan Fang
IEEE Trans. Geosci. Remote. Sens.6
2024 Point Label Meets Remote Sensing Change Detection: A Consistency-Aligned Regional Growth Network
abstract
The acquisition of a substantial volume of precisely dense pixel-annotated samples plays a crucial role in the effective training of deep learning-based change detection models. Nevertheless, in real-world scenarios, pairwise labeling of massive bitemporal remote sensing images is often laborious and time-consuming, resulting in the lack of labeled samples. In this article, we propose a novel point-based weakly supervised learning approach, called as the consistency-aligned regional growth network (CARGNet), for remote sensing change detection. Unlike pixel-level labels, point labels are easy to label and usually sparse, which leads to a lack of boundary information, making it difficult for the model to accurately capture the details of the changed objects. Therefore, learning directly from them may mislead the training of the network. To address these problems, we introduce a point-based changed regional growth (PCRG) module and consistency alignment (CA) constraint into CARGNet, which breaks the limitation of point labels in losing important target details. Specifically, our CARGNet contains two branches: a base decoder branch and an expanded decoder branch. First, we utilize the PCRG module to generate the expanded annotations from the point annotations. Then, the base decoder is supervised by the original point annotations, while the expanded decoder is supervised by the expanded annotations. Finally, the CA constraint is thereby achieved by minimizing the discrepancy between the predictions from both the base and the expanded decoders, which greatly improves the performance of the model. Experimental results on LEVIR-CD-Point and DSIFN-CD-Point datasets demonstrate that our proposed CARGNet can achieve highly competitive results compared with state-of-the-art fully-supervised methods. Code and datasets are available athttps://github.com/Wanderlust717/CARGNet.
Leyuan Fang, Yiqi Jiang, Jun Yue 0004
IEEE Trans. Geosci. Remote. Sens.5
2024 Enhancing Hyperspectral Image Classification: Leveraging Unsupervised Information With Guided Group Contrastive Learning
abstract
Deep learning (DL) has demonstrated remarkable performance in the classification of hyperspectral images (HSIs) by leveraging its powerful ability to automatically learn deep spectral–spatial features over the years. Nevertheless, the limited supervisory signals along with a vast number of parameters in deep models still pose critical challenges when utilizing a restricted number of samples for training deep networks. To better handle this issue, this article proposes an end-to-end framework called guided group contrastive learning (GGCL) that adaptively integrates unsupervised information into a supervised contrastive learning framework. The proposed method employs a similarity-guided module that measures the spectral–spatial similarity of unsupervised samples based on supervised signals and effectively groups them. Then, the similarity signals of both supervised and unsupervised data are combined with contrastive learning to achieve intragroup feature aggregation and intergroup feature separation with guided group contrastive loss (GGCLoss). The pivotal characteristic of the proposed method lies in the end-to-end incorporation of unsupervised information with supervised signals for contrastive learning. Experiments on three public HSI datasets demonstrate that the proposed method can achieve better performance than existing state-of-the-art (SOTA) methods. For ease of reproducibility, the code of the proposed GGCL will be publicly available athttps://github.com/fanerlight/GGCL_HSI.
Leyuan Fang, Jitong Kang, Jun Yue 0004
IEEE Trans. Geosci. Remote. Sens.5
2024 HyperMamba: A Spectral-Spatial Adaptive Mamba for Hyperspectral Image Classification
abstract
Transformers have significantly advanced hyperspectral image (HSI) classification through their proficiency in modeling long sequences. However, the high dimensionality of HSIs poses a particular challenge for Transformers due to their quadratic computational complexity. In natural language processing, state-space models (SSMs) such as Mamba hold great promise for handling long sequence tasks with significantly reduced computational overhead. However, the original Mamba lacks consideration for the spectral and spatial information inherent in HSIs. Inspired by this, we propose the HyperMamba, a novel spectral-spatial adaptive Mamba for HSI classification. The core idea of HyperMamba involves adaptively scanning spatial neighborhood pixels and dynamically enhancing spectral bands for spectral scanning based on acquired spatial neighborhood information. Specifically, HyperMamba consists of two core modules: the spatial neighborhood adaptive scanning (SNAS) module and the spectral adaptive enhancement scanning (SAES) module. Initially, the SNAS module analyzes the spectral characteristics of classified pixels, adaptively selecting the optimal neighborhood for spatial scanning by balancing spatial neighborhood information and local spatial structure. Subsequently, the SAES module dynamically enhances the spectral features of classified pixels using neighborhood spectral information and conducts spectral scanning. Finally, the spectral features of the target pixels are fed into a single fully connected layer classifier, achieving high-precision HSI classification. Extensive experiments demonstrate the effectiveness of HyperMamba, surpassing state-of-the-art methods across three widely used HSI datasets. The code will be available athttps://github.com/chiangliu/HyperMamba.
Jun Yue 0004, Shaobo Xia, Leyuan Fang
IEEE Trans. Geosci. Remote. Sens.2
2024 Diffusion Models Meet Remote Sensing: Principles, Methods, and Perspectives
abstract
As a newly emerging advance in deep generative models, diffusion models have achieved state-of-the-art results in many fields, including computer vision, natural language processing, and molecule design. The remote sensing (RS) community has also noticed the powerful ability of diffusion models and quickly applied them to a variety of tasks for image processing. Given the rapid increase in research on diffusion models in the field of RS, it is necessary to conduct a comprehensive review of existing diffusion model-based RS papers, to help researchers recognize the potential of diffusion models and provide some directions for further exploration. Specifically, this article first introduces the theoretical background of diffusion models, and then systematically reviews the applications of diffusion models in RS, including image generation, enhancement, and interpretation. Finally, the limitations of existing RS diffusion models and worthy research directions for further exploration are discussed and summarized.
Yidan Liu, Jun Yue 0004, Shaobo Xia, Pedram Ghamisi, Weiying Xie, Leyuan Fang
IEEE Trans. Geosci. Remote. Sens.2
2024 When Vectorization Meets Change Detection
abstract
In long-term Earth observation, change detection (CD) is a crucial and intricate task with applications spanning diverse fields, including land resource planning and natural disaster monitoring. Most existing CD approaches typically output segmentation results in raster format. However, raster format results suffer from higher memory usage, poorer shape accuracy, magnified distortions, and challenges in topological editing. To address the issues of raster format, we propose a novel end-to-end change vectorization network (CVNet), which is the first attempt to extract changes using vector format. The CVNet directly learns the vector components of changed objects and uses them to construct vectors. Specifically, since the vectorization of CD faces the inherent imbalance between changed and unchanged samples, we first introduce the Change-Collector to collect the changed regions and combine them into more compact samples. Next, the vector components learning model (VCLM) is introduced to capture the fundamental components for constructing the vectors, including change maps, junction positions, and segmentation masks. Finally, the changed instances obtained from the masks are used to divide and connect junctions to generate the vector output. To verify the effectiveness of the proposed framework, we construct two building change vectorization datasets by modifying the WHU-CD and LEVIR-CD benchmarks. Experimental results demonstrate that the CVNet outperforms the existing postprocess vectorization methods in terms of the visual effect and all evaluation metrics. The dataset and source code will be made publicly available athttps://github.com/yyyyll0ss/CVNet.
Yinglong Yan, Jun Yue 0004, Jiaxing Lin, Weiying Xie, Leyuan Fang
IEEE Trans. Geosci. Remote. Sens.2
2024 HyperMLL: Toward Robust Hyperspectral Image Classification With Multisource Label Learning
Xia Yue, Anfeng Liu, Shaobo Xia, Jun Yue 0004, Leyuan Fang
IEEE Trans. Geosci. Remote. Sens.5
2024 SemiRS-COC: Semi-Supervised Classification for Complex Remote Sensing Scenes With Cross-Object Consistency
abstract
Semi-supervised learning (SSL), which aims to learn with limited labeled data and massive amounts of unlabeled data, offers a promising approach to exploit the massive amounts of satellite Earth observation images. The fundamental concept underlying most state-of-the-art SSL methods involves generating pseudo-labels for unlabeled data based on image-level predictions. However, complex remote sensing (RS) scene images frequently encounter challenges, such as interference from multiple background objects and significant intra-class differences, resulting in unreliable pseudo-labels. In this paper, we propose the SemiRS-COC, a novel semi-supervised classification method for complex RS scenes. Inspired by the idea that neighboring objects in feature space should share consistent semantic labels, SemiRS-COC utilizes the similarity between foreground objects in RS images to generate reliable object-level pseudo-labels, effectively addressing the issues of multiple background objects and significant intra-class differences in complex RS images. Specifically, we first design a Local Self-Learning Object Perception (LSLOP) mechanism, which transforms multiple background objects interference of RS images into usable annotation information, enhancing the model's object perception capability. Furthermore, we present a Cross-Object Consistency Pseudo-Labeling (COCPL) strategy, which generates reliable object-level pseudo-labels by comparing the similarity of foreground objects across different RS images, effectively handling significant intra-class differences. Extensive experiments demonstrate that our proposed method achieves excellent performance compared to state-of-the-art methods on three widely-adopted RS datasets.
Jun Yue 0004, Weiying Xie, Leyuan Fang
IEEE Trans. Image Process.2
2023 SpectralDiff: A Generative Framework for Hyperspectral Image Classification With Diffusion Models
abstract
Hyperspectral Image (HSI) classification is an important issue in remote sensing field with extensive applications in earth science. In recent years, a large number of deep learning-based HSI classification methods have been proposed. However, existing methods have limited ability to handle high-dimensional, highly redundant, and complex data, making it challenging to capture the spectral-spatial distributions of data and relationships between samples. To address this issue, we propose a generative framework for HSI classification with diffusion models (SpectralDiff) that effectively mines the distribution information of high-dimensional and highly redundant data by iteratively denoising and explicitly constructing the data generation process, thus better reflecting the relationships between samples. The framework consists of a spectral-spatial diffusion module, and an attention-based classification module. The spectral-spatial diffusion module adopts forward and reverse spectral-spatial diffusion processes to achieve adaptive construction of sample relationships without requiring prior knowledge of graphical structure or neighborhood information. It captures spectral-spatial distribution and contextual information of objects in HSI and mines unsupervised spectral-spatial diffusion features within the reverse diffusion process. Finally, these features are fed into the attention-based classification module for per-pixel classification. The diffusion features can facilitate cross-sample perception via reconstruction distribution, leading to improved classification performance. Experiments on three public HSI datasets demonstrate that the proposed method can achieve better performance than state-of-the-art methods. For the sake of reproducibility, the source code of SpectralDiff will be publicly available at https://github.com/chenning0115/SpectralDiff.
Jun Yue 0004, Leyuan Fang, Shaobo Xia
IEEE Trans. Geosci. Remote. Sens.2
2023 Hyperspectral Image Instance Segmentation Using Spectral-Spatial Feature Pyramid Network
abstract
In recent years, hyperspectral image (HSI) classification and detection techniques based on deep learning have been widely applied to various aspects, such as environmental monitoring, urban planning, and energy surveys. As an important image content analysis method, instance segmentation can provide important support for the extraction of ground object information and monomeric application of HSI. This article introduces instance segmentation into HSI interpretation for the first time. In this article, we create the hyperspectral instance segmentation dataset (HS-ISD), which contains a total of 56 images, each with a size of$298\times301$and a number of channels of 48. More than 1000 architectural examples are annotated to apply to the research of HSI instance segmentation. In addition, considering that HSI contains rich spectral and spatial information, and the traditional instance segmentation network model cannot well utilize both types of information effectively, we propose the spectral–spatial feature pyramid network (Spectral–Spatial FPN). The Spectral–Spatial FPN can integrate multiscale spectral information and multiscale spatial information in the feature extraction stage through attention mechanism and bidirectional feature pyramid structure, so as to better improve the performance of the network model by spectral information and spatial information and realize the end-to-end instance segmentation of HSI. The experimental results conducted on the HS-ISD show that the proposed Spectral–Spatial FPN can achieve state-of-the-art results.
Leyuan Fang, Yinglong Yan, Jun Yue 0004, Yue Deng 0001
IEEE Trans. Geosci. Remote. Sens.4
2023 Rethinking Remote Sensing Pretrained Model: Instance-Aware Visual Prompting for Remote Sensing Scene Classification
abstract
Large-scale pre-trained models, such as vision transformers, have made significant progress in remote sensing (RS) scene classification tasks. For a new scene classification task, it is popular to fully fine-tune the pre-trained model parameters to avoid training from scratch. Although such an approach achieves satisfactory results, it will lead to heavy computation and storage burden, which limits the transferability of large pre-trained models to different RS scene classification tasks. To address this challenge, we propose a parameter-efficient tuning approach called as the Instance-Aware Visual Prompting (IVP), which is the first work to explore the prompting in the field of RS scene classification. The proposed IVP adaptively generates prompts based on the complex background and highly variable characteristics of RS images, and updates only a few parameters to transfer the pre-trained RS Transformer model to different scene classification tasks. Specifically, instead of adapting the entire model parameters, we introduce some instance-specific prompt vectors into the input space. Then, considering the significant variability in RS images, we introduce an instance-level prompt generation module to generate specific prompts for each RS image by aggregating contextual information from the input. Finally, these prompt vectors will calibrate the pre-trained features to encode instance-specific information. Extensive experiments on three RS scene classification datasets demonstrate the superiority of IVP over other fine-tuning methods. For example, when updating just 1.1% parameters, the Swin Transformer model achieves about 1.83% and 1.42% improvement compared to the full fine-tuning method on NWPU-19 and NWPU-28, respectively.
Leyuan Fang, Jun Yue 0004
IEEE Trans. Geosci. Remote. Sens.5
2023 Toward the Vectorization of Hyperspectral Imagery
abstract
Hyperspectral images (HSIs) can provide rich spectral-spatial information that has been widely utilized in many fields, such as national defense, mineralogy and agriculture. Most of the recent HSI interpretation methods are conducted in the raster pattern, which results in high memory costs, amplification distortion, and difficulties in topological editing. To address this issue, a novel end-to-end vectorization framework is proposed, called as the HSI Vectorization Network (HSI-VecNet), which learns a vector representation from spectral-spatial information through cross-level interactions. Specifically, this framework integrates low-level geometry information and high-level semantic instance information, which consists of two branches: the HSI Semantic Instance Segmentation (HSIS) and the Spectral-Spatial Junction Prediction (SSJP). The HSIS conducts the raster-based classification and extracts the semantic information of each object in the HSI. In addition, the SSJP exploits spectral-spatial information to predict the positions of junctions in the HSI. The instance information of each object and the relations of junctions are then fused to vectorize the HSI. To verify the effectiveness of the proposed method, four hyperspectral datasets are vectorially labeled. Experimental results on these datasets demonstrate that the proposed end-to-end HSI-VecNet outperforms existing post-process vectorization methods. Our model and datasets will be made publicly available at https://github.com/yyyyll0ss/HSI-VecNet.
Leyuan Fang, Yinglong Yan, Jun Yue 0004, Yue Deng 0001
IEEE Trans. Geosci. Remote. Sens.3
2023 A Multi-Level Label-Aware Semi-Supervised Framework for Remote Sensing Scene Classification
abstract
Semi-supervised learning (SSL) is a promising approach to reduce the labeling burden in remote sensing scene classification tasks. However, most semi-supervised methods typically exploit the single-level semantic information of unlabeled data, ignoring the multi-level semantic structure prevalent in remote sensing data. The multi-level semantic structure, which contains the correlation of different categories and the multi-granularity semantic information, can help the scene classification model to more accurately measure the feature distance between different categories and more effectively utilize unlabeled data. Therefore, this paper proposes a multi-level label-aware semi-supervised scene classification framework, MLLA, which extends the semantic information captured in unlabeled data from single-level to multi-level to improve the scene classification performance. Specifically, we first propose a multi-level prototype awareness module to capture the multi-level semantic structure underlying remote sensing data. Then, based on this structure, a multi-level pseudo-label generation module is designed to assign multi-level pseudo-labels to the unlabeled data. Finally, by combining the labeled samples and the multi-level pseudo-labeled samples, the scene classification model is progressively trained. The experimental results on three benchmark datasets show that the proposed MLLA achieves excellent performance compared to other semi-supervised classification methods.
Linshan Wu, Jun Yue 0004, Leyuan Fang
IEEE Trans. Geosci. Remote. Sens.5
2023 PCLDet: Prototypical Contrastive Learning for Fine-Grained Object Detection in Remote Sensing Images
abstract
The capacity of satellites to supply high-resolution imaging has promoted the fine-grained object detection task in remote sensing images. However, this type of object detection is challenging due to low interclass feature differences in objects. To address this issue, we propose a prototypical contrastive learning-based detector (PCLDet) for fine-grained object detection in remote sensing images. The PCLDet first introduces the prototype to learn the fine-grained objects’ features, and then adopts contrastive learning to compare the target and the learned features, thus improving the differentiability of the fine-grained object. Specifically, we first introduce the prototype, which represents the feature centers of each class, and then construct a prototype bank to store the feature prototypes of each class. Then, we introduce contrastive learning to extract the discriminative features by maximizing the interclass distance and minimizing the intraclass distance. Furthermore, we propose the ProtoCL loss as a part of the model optimization, which enables more representative prototypes to be learned. Finally, to address the long-tail problem in the remote sensing fine-grained object detection dataset, we propose a new proposal sampler, the class-balanced sampler (CBS) that can sample each class equally. Extensive experiments demonstrate that our method can achieve state-of-the-art performance on a commonly used aerial fine-grained object dataset (Fair1M) and aerial fine-grained ship dataset (OFSD) while maintaining high efficiency. The code will be available at https://github.com/G-Naughty/PCLDet.
Lihan Ouyang, Guangmiao Guo, Leyuan Fang, Pedram Ghamisi, Jun Yue 0004
IEEE Trans. Geosci. Remote. Sens.5
2023 Dif-Fusion: Toward High Color Fidelity in Infrared and Visible Image Fusion With Diffusion Models
abstract
Color plays an important role in human visual perception, reflecting the spectrum of objects. However, the existing infrared and visible image fusion methods rarely explore how to handle multi-spectral/channel data directly and achieve high color fidelity. This paper addresses the above issue by proposing a novel method with diffusion models, termed as Dif-Fusion, to generate the distribution of the multi-channel input data, which increases the ability of multi-source information aggregation and the fidelity of colors. In specific, instead of converting multi-channel images into single-channel data in existing fusion methods, we create the multi-channel data distribution with a denoising network in a latent space with forward and reverse diffusion process. Then, we use the the denoising network to extract the multi-channel diffusion features with both visible and infrared information. Finally, we feed the multi-channel diffusion features to the multi-channel fusion module to directly generate the three-channel fused image. To retain the texture and intensity information, we propose multi-channel gradient loss and intensity loss. Along with the current evaluation metrics for measuring texture and intensity fidelity, we introduce Delta E as a new evaluation metric to quantify color fidelity. Extensive experiments indicate that our method is more effective than other state-of-the-art image fusion methods, especially in color fidelity. The source code is available at https://github.com/GeoVectorMatrix/Dif-Fusion.
Jun Yue 0004, Leyuan Fang, Shaobo Xia, Yue Deng 0001, Jiayi Ma 0001
IEEE Trans. Image Process.1
2022 Discrepant multiple instance learning for weakly supervised object detection
Wei Gao 0050, Fang Wan 0001, Jun Yue 0004, Songcen Xu, Qixiang Ye
Pattern Recognit.3
2022 Adaptive Regional Multiple Features for Large-Scale High-Resolution Remote Sensing Image Registration
abstract
The efficient and accurate registration of multitemporal images is essential for many remote sensing applications. With the increase in the imaging resolution and field in satellites, the acquired large-scale remote sensing images have brought serious challenges, since there exist parallax shifts and large background variations among different local regions of the acquired images. To address these issues, this article proposes an adaptive regional multiple features (ARMF) matching method for the registration of multitemporal large-scale high-resolution remote sensing images. Specifically, since large background variations in fixed-size regions of multitemporal images will cause insufficient features and the failure of features matching, the ARMF introduces an adaptive regions searching strategy, which utilizes the pyramid amplification technique to adaptively select the regions that can find the sufficient matched features. Then, the ARMF extracts multiple types of features (i.e., gradient feature, phase feature, and line feature) from the adaptive searched region that can more effectively represent the characteristics of the large regions. Finally, we utilize the feature matching error as the rule to adaptively select the suitable features as the descriptors of the region. The experimental results on large-scale multitemporal image data obtained from Google Earth demonstrated the proposed method can outperform several state-of-the-arts remote sensing registration approaches.
Zezhou Li, Jun Yue 0004, Leyuan Fang
IEEE Trans. Geosci. Remote. Sens.2
2022 Self-Supervised Learning With Adaptive Distillation for Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) classification is an important topic in the community of remote sensing, which has a wide range of applications in geoscience. Recently, deep learning-based methods have been widely used in HSI classification. However, due to the scarcity of labeled samples in HSI, the potential of deep learning-based methods has not been fully exploited. To solve this problem, a self-supervised learning (SSL) method with adaptive distillation is proposed to train the deep neural network with extensive unlabeled samples. The proposed method consists of two modules: adaptive knowledge distillation with spatial–spectral similarity and 3-D transformation on HSI cubes. The SSL with adaptive knowledge distillation uses the self-supervised information to train the network by knowledge distillation, where self-supervised knowledge is the adaptive soft label generated by spatial–spectral similarity measurement. The SSL with adaptive knowledge distillation mainly includes the following three steps. First, the similarity between unlabeled samples and object classes in HSI is generated based on the spatial–spectral joint distance (SSJD) between unlabeled samples and labeled samples. Second, the adaptive soft label of each unlabeled sample is generated to measure the probability that the unlabeled sample belongs to each object class. Third, a progressive convolutional network (PCN) is trained by minimizing the cross-entropy between the adaptive soft labels and the probabilities generated by the forward propagation of the PCN. The SSL with 3-D transformation rotates the HSI cube in both the spectral domain and the spatial domain to fully exploit the labeled samples. Experiments on three public HSI data sets have demonstrated that the proposed method can achieve better performance than existing state-of-the-art methods.
Jun Yue 0004, Leyuan Fang, Hossein Rahmani 0001, Pedram Ghamisi
IEEE Trans. Geosci. Remote. Sens.1
2022 Adaptive Spatial Pyramid Constraint for Hyperspectral Image Classification With Limited Training Samples
abstract
Deep learning-based methods have made significant progress in hyperspectral image (HSI) classification in recent years. However, deep learning-based methods usually rely on a large number of samples, and in many cases, it is difficult to label HSI and only limited training samples are available. To solve this problem, an HSI classification method based on adaptive spatial pyramid constraint (ASPC) is proposed to make full use of the global spatial neighborhood information of the labeled samples, which can improve the generalization ability of the classification model. The main steps of the proposed method are as follows. First, an HSI complexity evaluation method based on edge detection is proposed to assess the homogeneity of the objects in the HSI. Second, an HSI pyramid segmentation method based on spatial pyramid is proposed to generate multiscale subregions, where HSI complexity is used to adaptively determine the scale of the segmentation. Third, a spatial supervised constraint is proposed to generate the loss function of labeled subregions. Fourth, a spatial unsupervised constraint is proposed to generate the loss function of unlabeled subregions. The proposed method fully explores the spatial-spectral correlation between unlabeled samples and labeled samples, and add corresponding constraints to the training objective according to the correlation. By adding the ASPC, the trained model becomes more robust and can make full use of the limited training samples. To verify the effectiveness of the proposed method, three benchmark hyperspectral datasets are used to verify the performance of the proposed method. Experimental results show that the performance of this method is better than the existing state-of-the-art methods.
Jun Yue 0004, Dingshun Zhu, Leyuan Fang, Pedram Ghamisi, Yaowei Wang 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 Deep Bilateral Filtering Network for Point-Supervised Semantic Segmentation in Remote Sensing Images
abstract
Semantic segmentation methods based on deep neural networks have achieved great success in recent years. However, training such deep neural networks relies heavily on a large number of images with accurate pixel-level labels, which requires a huge amount of human effort, especially for large-scale remote sensing images. In this paper, we propose a point-based weakly supervised learning framework called the deep bilateral filtering network (DBFNet) for the semantic segmentation of remote sensing images. Compared with pixel-level labels, point annotations are usually sparse and cannot reveal the complete structure of the objects; they also lack boundary information, thus resulting in incomplete prediction within the object and the loss of object boundaries. To address these problems, we incorporate the bilateral filtering technique into deeply learned representations in two respects. First, since a target object contains smooth regions that always belong to the same category, we perform deep bilateral filtering (DBF) to filter the deep features by a nonlinear combination of nearby feature values, which encourages the nearby and similar features to become closer, thus achieving a consistent prediction in the smooth region. In addition, the DBF can distinguish the boundary by enlarging the distance between the features on different sides of the edge, thus preserving the boundary information well. Experimental results on two widely used datasets, the ISPRS 2-D semantic labeling Potsdam and Vaihingen datasets, demonstrate that our proposed DBFNet can achieve a highly competitive performance compared with state-of-the-art fully-supervised methods. Code is available at https://github.com/Luffy03/DBFNet.
Linshan Wu, Leyuan Fang, Jun Yue 0004, Bob Zhang 0001, Pedram Ghamisi
IEEE Trans. Image Process.3
2022 Spectral-Spatial Latent Reconstruction for Open-Set Hyperspectral Image Classification
abstract
Deep learning-based methods have produced significant gains for hyperspectral image (HSI) classification in recent years, leading to high impact academic achievements and industrial applications. Despite the success of deep learning-based methods in HSI classification, they still lack the robustness of handling unknown object in open-set environment (OSE). Open-set classification is to deal with the problem of unknown classes that are not included in the training set, while in closed-set environment (CSE), unknown classes will not appear in the test set. The existing open-set classifiers almost entirely rely on the supervision information given by the known classes in the training set, which leads to the specialization of the learned representations into known classes, and makes it easy to classify unknown classes as known classes. To improve the robustness of HSI classification methods in OSE and meanwhile maintain the classification accuracy of known classes, a spectral-spatial latent reconstruction framework which simultaneously conducts spectral feature reconstruction, spatial feature reconstruction and pixel-wise classification in OSE is proposed. By reconstructing the spectral and spatial features of HSI, the learned feature representation is enhanced, so as to retain the spectral-spatial information useful for rejecting unknown classes and distinguishing known classes. The proposed method uses latent representations for spectral-spatial reconstruction, and achieves robust unknown detection without compromising the accuracy of known classes. Experimental results show that the performance of the proposed method outperforms the existing state-of-the-art methods in OSE.
Jun Yue 0004, Leyuan Fang
IEEE Trans. Image Process.1
2021 Oriented Spatial Correlative Aligned Feature for Remote Sensing Object Detection
abstract
In the past decade, the object detection in remote sensing has become a research focus with the development of aerial images' acquirement technology. However, it is difficult to accurately describe rotated rigid objects with large aspect ratio through axis-aligned convolutional features. To solve this problem, we propose an oriented spatial correlative feature alignment Net (OSCFA-Net) with a correlative feature alignment module (CFAM). CFAM aligns features by rotating object features and considering the spatial correlation of rigid objects to ensure the feature alignment between the internal feature points. Experiment results demonstrate that OSCFA-Net can achieve state-of-the-art performance in remote sensing object detection.
Guangmiao Guo, Leyuan Fang, Jun Yue 0004
IGARSS3
2020 Large-Scale Few-Shot Learning via Multi-modal Knowledge Discovery
Shuo Wang 0008, Jun Yue 0004, Jianzhuang Liu, Qi Tian 0001, Meng Wang 0001
ECCV (10)2