Jing Yao 0002

dblp:24/5678-2 · DBLP profile ↗
← Back
36ranked-venue papers
8as first author
29since 2021 · last 2025
0000-0003-1301-9758ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 27 · 5 first-author · 24 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2025 AeroGen: Enhancing Remote Sensing Object Detection with Diffusion-Driven Data Generation
abstract
Remote sensing image object detection (RSIOD) aims to identify and locate specific objects within satellite or aerial imagery. However, there is a scarcity of labeled data in current RSIOD datasets, which significantly limits the performance of current detection algorithms. Although existing techniques, e.g., data augmentation and semi-supervised learning, can mitigate this scarcity issue to some extent, they are heavily dependent on high-quality labeled data and perform worse in rare object classes. To address this issue, this paper proposes a layout-controllable diffusion generative model (i.e. AeroGen) tailored for RSIOD. To our knowledge, AeroGen is the first model to simultaneously support horizontal and rotated bounding box condition generation, thus enabling the generation of high-quality synthetic images that meet specific layout and object category requirements. Additionally, we propose an end-to-end data augmentation framework that integrates a diversity-conditioned generator and a filtering mechanism to enhance both the diversity and quality of generated data. Experimental results demonstrate that the synthetic data produced by our method are of high quality and diversity. Furthermore, the synthetic RSIOD data can significantly improve the detection performance of existing RSIOD models, i.e., the mAP metrics on DIOR, DIOR-R, and HRSC datasets are improved by 3.7%, 4.3%, and 2.43%, respectively. The code is available at here.
Datao Tang, Xiangyong Cao, Jing Yao 0002, Xueru Bai, Dongsheng Jiang, Deyu Meng
CVPR5
2025 Towards Satellite Image Road Graph Extraction: A Global-Scale Dataset and A Novel Method
abstract
Recently, road graph extraction has garnered increasing attention due to its crucial role in autonomous driving, navigation, etc. However, accurately and efficiently extracting road graphs remains a persistent challenge, primarily due to the severe scarcity of labeled data. To address this limitation, we collect a global-scale satellite road graph extraction dataset, i.e. Global-Scale dataset. Specifically, the Global-Scale dataset is ∼ 20× larger than the largest existing public road extraction dataset and spans over 13,800 km2globally. Additionally, we develop a novel road graph extraction model, i.e. SAM-Road++, which adopts a node-guided resampling method to alleviate the mismatch issue between training and inference in SAM-Road [17], a pioneering state-of-the-art road graph extraction model. Furthermore, we propose a simple yet effective "extended-line" strategy in SAM-Road++ to mitigate the occlusion issue on the road. Extensive experiments demonstrate the validity of the collected Global-Scale dataset and the proposed SAM-Road++ method, particularly highlighting its superior predictive power in unseen regions. The dataset and code are available at https://github.com/earth-insights/samroadplus.
Pan Yin, Kaiyu Li 0001, Xiangyong Cao, Jing Yao 0002, Lei Liu 0014, Xueru Bai, Feng Zhou 0001, Deyu Meng
CVPR4
2025 ISDAT: An image-semantic dual adversarial training framework for robust image classification
Chenhong Sui, Hao Liu 0019, Qingtao Gong, Jing Yao 0002, Danfeng Hong
Pattern Recognit.6
2025 Mask Approximation Net: A Novel Diffusion Model Approach for Remote Sensing Change Captioning
abstract
Remote sensing (RS) image change description represents an innovative multimodal task within the realm of RS processing. This task not only facilitates the detection of alterations in surface conditions but also provides comprehensive descriptions of these changes, thereby improving human interpretability and interactivity. Current deep learning methods typically adopt a three-stage framework consisting of feature extraction, feature fusion, and change localization, followed by text generation. Most approaches focus heavily on designing complex network modules but lack solid theoretical guidance, relying instead on extensive empirical experimentation and iterative tuning of network components. This experience-driven design paradigm may lead to overfitting and design bottlenecks, thereby limiting the model’s generalizability and adaptability. To address these limitations, this article proposes a paradigm that shifts toward data distribution learning using diffusion models, reinforced by frequency-domain noise filtering, to provide a theoretically motivated and practically effective solution to multimodal RS change description. The proposed method primarily includes a simple multiscale change detection (CD) module, whose output features are subsequently refined by a well-designed diffusion model. Furthermore, we introduce a frequency-guided complex filter module to boost the model’s performance by managing high-frequency noise throughout the diffusion process. We validate the effectiveness of our proposed method across several datasets for RS CD and description, showcasing its superior performance compared to existing techniques. The code will be available athttps://github.com/sundongweiMaskApproxNet
Dongwei Sun, Jing Yao 0002, Wu Xue, Changsheng Zhou, Pedram Ghamisi, Xiangyong Cao
IEEE Trans. Geosci. Remote. Sens.2
2024 SpectralGPT: Spectral Remote Sensing Foundation Model
abstract
The foundation model has recently garnered significant attention due to its potential to revolutionize the field of visual representation learning in a self-supervised manner. While most foundation models are tailored to effectively process RGB images for various visual tasks, there is a noticeable gap in research focused on spectral data, which offers valuable information for scene understanding, especially in remote sensing (RS) applications. To fill this gap, we created for the first time a universal RS foundation model, named SpectralGPT, which is purpose-built to handle spectral RS images using a novel 3D generative pretrained transformer (GPT). Compared to existing foundation models, SpectralGPT 1) accommodates input images with varying sizes, resolutions, time series, and regions in a progressive training fashion, enabling full utilization of extensive RS Big Data; 2) leverages 3D token generation for spatial-spectral coupling; 3) captures spectrally sequential patterns via multi-target reconstruction; and 4) trains on one million spectral RS images, yielding models with over 600 million parameters. Our evaluation highlights significant performance improvements with pretrained SpectralGPT models, signifying substantial potential in advancing spectral RS Big Data applications within the field of geoscience across four downstream tasks: single/multi-label scene classification, semantic segmentation, and change detection.
Danfeng Hong, Bing Zhang 0001, Chenyu Li 0002, Jing Yao 0002, Naoto Yokoya, Hao Li 0019, Pedram Ghamisi, Xiuping Jia, Antonio Plaza, Paolo Gamba, Jón Atli Benediktsson, Jocelyn Chanussot
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 Interpretable Networks for Hyperspectral Anomaly Detection: A Deep Unfolding Solution
abstract
Current hyperspectral anomaly detection (HAD) benchmark datasets suffer from low resolution, simple background, and small size of the anomalies. These factors also limit the performance of the well-known low-rank representation (LRR) models in terms of robustness on the separation of background and target features and the reliance on manual parameter selection. To this end, we build a new HAD benchmark dataset for improving the robustness in complex scenarios, AIR-HAD for short, and propose an interpretable network with deep unfolding a binary subspace learning, named LRR-Net+, which is capable of spectrally decoupling the background structure and object properties in a more generalized fashion and eliminating the bias introduced by vital interference targets simultaneously. In addition, LRR-Net+ integrates the solution process of the alternating direction method of multipliers (ADMM) optimizer with the deep network, guiding its search process and imparting a level of interpretability to parameter optimization. Additionally, the integration of physical models with DL techniques eliminates the need for manual parameter tuning. The manually tuned parameters are seamlessly transformed into trainable parameters for deep neural networks, facilitating a more efficient and automated optimization process. Extensive experiments conducted on the AIR-HAD dataset show the superiority of our LRR-Net+ in terms of detection performance and generalization ability, compared to top-performing competitors. Furthermore, our AIR-HAD benchmark datasets will be made available freely and openly athttps://github.com/danfenghong/IEEE_TGRS_LRR-Net.
Chenyu Li 0002, Bing Zhang 0001, Danfeng Hong, Jing Yao 0002, Xiuping Jia, Antonio Plaza, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.4
2024 Learnable Representative Coefficient Image Denoiser for Hyperspectral Image
abstract
Fully characterizing the spatial-spectral priors of hyperspectral images (HSI) is crucial for HSI denoising tasks. Recently, HSI denoising models based on representative coefficient images (RCIs) under the spectral low-rank decomposition framework have garnered significant attention due to their clever utilization of spatial-spectral information in HSI at a low cost. However, current methods either employ handcrafted classical denoisers or off-the-shelf deep denoisers to denoise RCIs, failing to fully capture the structural information of RCIs. In this paper, we propose a specific optimization framework for learning an RCI denoiser under the low-rank decomposition framework for the first time. Since low-rank decomposition can characterize the global low-rank property of HSI, our RCI denoiser only needs to learn the spatial prior of RCIs. Consequently, our optimization framework is inclined to learn a more powerful RCI denoiser. However, learning an RCI denoiser is not an easy task, primarily due to the lack of paired clean-noisy RCI data. To address this issue, we employ parametric techniques to represent the to-be-restored HSI as a function of RCI denoiser network parameters. In this way, the parameters of the RCI denoiser can thus be updated using noisy-clean HSI pairs. Furthermore, we adopt residual learning and Gaussian whitening techniques to enhance the RCI denoiser’s denoising ability for HSIs with various noise levels and different rank settings. Extensive experiments demonstrate that our method can achieve significant improvements in both denoising effectiveness and speed compared to state-of-the-art methods. The code of our algorithm is released at https://github.com/andrew-pengjj/RCILD.git.
Jiangjun Peng, Hailin Wang 0001, Xiangyong Cao, Qian Zhao 0002, Jing Yao 0002, Hong-Ying Zhang 0001, Deyu Meng
IEEE Trans. Geosci. Remote. Sens.5
2024 IG-GAN: Interactive Guided Generative Adversarial Networks for Multimodal Image Fusion
abstract
Multimodal image fusion has recently garnered increasing interest in the field of remote sensing. By leveraging the complementary information in different modalities, the fused results may be more favorable in characterizing objects of interest, thereby increasing the chance of a more comprehensive and accurate perception of the scene. Unfortunately, most existing fusion methods tend to extract modality-specific features independently without considering intermodal alignment and complementarity, leading to a suboptimal fusion process. To address this issue, we propose a novel interactive generative adversarial network (IG-GAN), for the task of multimodal image fusion. IG-GAN comprises guided dual streams tailored for enhanced learning of details and content, as well as cross-modal consistency. Specifically, a details-guided interactive running-in module (GIR1) and a content-guided interactive running-in module (GIR2) are developed, with the stronger modality serving as guidance for detail richness or content integrity, and the weaker one assisting. To fully integrate multigranularity features from dual-modality, a hierarchical fusion and reconstruction branch is established. Specifically, a shallow interactive fusion (SIF) module followed by a multilevel interactive fusion (MIF) module is designed to aggregate multilevel local and long-range features. Concerning feature decoding and fused image generation, a high-level interactive fusion and reconstruction module (HRM) is further developed. In addition, to empower the fusion network to generate fused images with complete content, sharp edges, and high fidelity without supervision, a loss function facilitating the mutual game between the generator and two discriminators is also formulated. Comparative experiments with 14 state-of-the-art methods are conducted on three datasets. Qualitative and quantitative results indicate that IG-GAN exhibits obvious superiority in terms of both visual effect and quantitative metrics. Moreover, experiments on two RGB-IR object detection datasets are also conducted, which demonstrate that IG-GAN can enhance the accuracy of object detection by integrating complementary information from different modalities. The code will be available athttps://github.com/flower6top.
Chenhong Sui, Guobin Yang, Danfeng Hong, Jing Yao 0002, Peter M. Atkinson, Pedram Ghamisi
IEEE Trans. Geosci. Remote. Sens.5
2023 Decoupled-and-Coupled Networks: Self-Supervised Hyperspectral Image Super-Resolution With Subpixel Fusion
abstract
Enormous efforts have been recently made to super-resolve hyperspectral (HS) images with the aid of high spatial resolution multispectral (MS) images. Most prior works usually perform the fusion task by means of multifarious pixel-level priors. Yet the intrinsic effects of a large distribution gap between HS-MS data due to differences in the spatial and spectral resolution are less investigated. The gap might be caused by unknown sensor-specific properties or highly-mixed spectral information within one pixel (due to low spatial resolution). To this end, we propose a subpixel-level HS super-resolution framework by devising a novel decoupled-and-coupled network, called DC-Net, to progressively fuse HS-MS information from the pixel- to subpixel-level, from the image- to feature-level. As the name suggests, DC-Net first decouples the input into common (or cross-sensor) and sensor-specific components to eliminate the gap between HS-MS images before further fusion, and then thoroughly blends them by a model-guided coupled spectral unmixing (CSU) net. More significantly, we append a self-supervised learning module behind the CSU net by guaranteeing material consistency to enhance the detailed appearance of the restored HS product. Extensive experimental results show the superiority of our method both visually and quantitatively and achieve a significant improvement in comparison with the state-of-the-art.
Danfeng Hong, Jing Yao 0002, Chenyu Li 0002, Deyu Meng, Naoto Yokoya, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.2
2023 LRR-Net: An Interpretable Deep Unfolding Network for Hyperspectral Anomaly Detection
abstract
Considerable endeavors have been expended towards enhancing the representation performance for Hyperspectral Anomaly Detection (HAD) through physical model-based methods and recent deep learning-based approaches. Of these methods, the Low-Rank Representation (LRR) model is widely adopted for its formidable separation capabilities for background and target features, however, its practical applications are limited due to the reliance on manual parameter selection and subpar generalization performance. To this end, this paper presents a new HAD baseline network, referred to as LRR-Net, which synergizes the LRR model with deep learning techniques. LRR-Net leverages the alternating direction method of multipliers (ADMM) optimizer to solve the LRR model efficiently and incorporates the solution as prior knowledge into the deep network to guide the optimization of parameters. Moreover, LRR-Net transforms the regularized parameters into trainable parameters of the deep neural network, thus alleviating the need for manual parameter tuning. Additionally, this paper proposes a sparse neural network embedding to demonstrate the scalability of the LRR-Net framework. Empirical evaluations on eight distinct datasets illustrate the efficacy and superiority of the proposed approach compared to state-of-the-art methods.
Chenyu Li 0002, Bing Zhang 0001, Danfeng Hong, Jing Yao 0002, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.4
2023 Learning Double Subspace Representation for Joint Hyperspectral Anomaly Detection and Noise Removal
abstract
Efforts to enhance the detection accuracy of hyperspectral (HS) anomaly detection (AD) have been significant, but the impact of noise resulting from HS data acquisition and transmission has not been well studied. Furthermore, the separation of denoising and subsequent interpretation makes it challenging to evaluate and control the influence of noise on the detection results. To this end, we proposed a joint anomaly detection and noise removal (ADNR) paradigm called DSR-ADNR, which develops a double subspace representation method to obtain both denoised and detection results simultaneously. DSR-ADNR uses a low-dimensional orthogonal basis to represent HS images and extract distinctive features for AD. The feature matrix is represented by a dictionary-based low-rank subspace that captures the complex nature of the low-dimensional features. In each iteration, DSR-ADNR utilizes the nonlocal self-similarity of the feature matrix to remove noise and improve intermediate detection performance. Meanwhile, the progressive LR representation of the background and anomalies for the feature matrix upgrades the explicit LR expression of nonlocal self-similar patches for better denoising. The well-designed linearized alternating direction method of multipliers with an adaptive penalty (LADMAP) is utilized to solve the proposed DSR-ADNR. Extensive experiments on simulated and real-world data sets demonstrate the effectiveness of DSR-ADNR in the HS AD task under different noise cases.
Danfeng Hong, Bing Zhang 0001, Longfei Ren, Jing Yao 0002, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.5
2023 UCSL: Toward Unsupervised Common Subspace Learning for Cross-Modal Image Classification
abstract
The emerging research line of cross-modal learning focuses on the issue of transferring feature representation manner learned from limited multimodal data with labelings to the testing phase with partial modalities. This is essentially common and practical in the remote sensing community when only modal-incomplete data are in users’ hands due to inevitable imaging or access restrictions under large-scale observation scenarios. However, most of the existing cross-modal learning methods have been designed with exclusive reliance on labeling, which can be either limited or noisy due to their costly production. To address this issue, we explore in this paper the possibility to learn cross-modal feature representation in an unsupervised fashion. By integrating the multimodal data into a fully recombined matrix form, we propose 1) the use of common subspace representation as the regression target instead of conventionally adopted binary labels, and 2) the orthogonality and manifold alignment regularization terms to shrink the solution space whilst preserving the pairwise manifold correlations. Through this manner, the modality-specific and mutual latent representations in this common subspace as well as their corresponding projections can be learned simultaneously and their optimums can be efficiently reached through a nearly one-step computation with the help of Eigen decomposition. Finally, we show the superiority of our method through extensive image classification experiments on three multimodal datasets with four remotely sensed modalities involved (i.e., hyperspectral, multispectral, synthetic aperture radar, and light detection and ranging data). The code and dataset will be made freely available at https://github.com/jingyao16/UCSL after a possible publication to encourage the reproduction of our method and further use.
Jing Yao 0002, Danfeng Hong, Hao Liu 0019, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.1
2023 Extended Vision Transformer (ExViT) for Land Use and Land Cover Classification: A Multimodal Deep Learning Framework
abstract
The recent success of attention mechanism-driven deep models, like Vision Transformer (ViT) as one of the most representative, has intrigued a wave of advanced research to explore their adaptation to broader domains. However, current Transformer-based approaches in the remote sensing (RS) community pay more attention to single-modality data, which might lose expandability in making full use of the ever-growing multimodal Earth observation data. To this end, we propose a novel multimodal deep learning framework by extending conventional ViT with minimal modifications, abbreviated as ExViT, aiming at the task of land use and land cover classification. Unlike common stems that adopt either linear patch projection or deep regional embedder, our approach processes multimodal RS image patches with parallel branches of position-shared ViTs extended with separable convolution modules, which offers an economical solution to leverage both spatial and modality-specific channel information. Furthermore, to promote information exchange across heterogeneous modalities, their tokenized embeddings are then fused through a cross-modality attention module by exploiting pixel-level spatial correlation in RS scenes. Both of these modifications significantly improve the discriminative ability of classification tokens in each modality and thus further performance increase can be finally attained by a full tokens-based decision-level fusion module. We conduct extensive experiments on two multimodal RS benchmark datasets, i.e., the Houston2013 dataset containing hyperspectral and light detection and ranging (LiDAR) data, and Berlin dataset with hyperspectral and synthetic aperture radar (SAR) data, to demonstrate that our ExViT outperforms concurrent competitors based on Transformer or convolutional neural network (CNN) backbones, in addition to several competitive machine learning-based models. The source codes and investigated datasets of this work will be made publicly available at https://github.com/jingyao16/ExViT.
Jing Yao 0002, Bing Zhang 0001, Chenyu Li 0002, Danfeng Hong, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.1
2022 Multimodal Hyperspectral Unmixing via Attention Networks
abstract
Owing to the powerful feature extraction and representation capabilities, deep learning (DL) has been successfully applied in hyperspectral unmixing (HU). However, only relying on hyperspectral data for unmixing fails to distinguish objects with similar spectral information, resulting in the degradation of unmixing performance. To this end, this paper presents a novel multimodal unmixing network, MUNet for short, by considering the height information of light detection and ranging (LiDAR) data in a squeeze-and-excitation (SE) attention fashion to guide the unmixing process toward a more accurate performance. MUNet is capable of efficiently embedding the height information obtained from LiDAR data into the autoencoder unmixing architecture through the attention mechanism, thereby fusing more spatial information to obtain ideal unmixing results. Experimental results conducted on the real multimodal dataset demonstrate the effectiveness and superiority of the proposed MUNet compared to several state-of-the-art deep unmixing approaches.
Zhu Han 0002, Danfeng Hong, Lianru Gao, Jing Yao 0002, Bing Zhang 0001, Jocelyn Chanussot
IGARSS4
2022 Multimodal Remote Sensing Benchmark Datasets for Land Cover Classification
abstract
Over the past few decades, a large collection of feature ex-traction and classification algorithms have been developed for land cover mapping using remote sensing data. Although these methods have shown the gradually-increasing performance, their potential inevitably meets the bottleneck due to the lack of high-quality and diversified remote sensing bench-mark datasets, particularly for the multimodal cases. Accordingly, this, to a larger extent, limits the development of the corresponding methodologies and the practical application of land cover classification. To this end, we aim in this pa-per to introduce and build several multimodal remote sensing benchmark datasets for land cover classification. Further-more, two new multimodal land cover classification bench-mark datasets, i.e., Berlin and Augsburg, are openly available. Experiments are conducted on the two datasets for evaluating the performance of several multimodal feature learning and classification methods.
Jing Yao 0002, Danfeng Hong, Lianru Gao, Jocelyn Chanussot
IGARSS1
2022 PolSAR Image Classification Based on Robust Low-Rank Feature Extraction and Markov Random Field
abstract
Polarimetric synthetic aperture radar (PolSAR) image classification has been investigated vigorously in various remote sensing applications. However, it is still a challenging task nowadays. One significant barrier lies in the speckle effect embedded in the PolSAR imaging process, which greatly degrades the quality of the images and further complicates the classification. To this end, we present a novel PolSAR image classification method that removes speckle noise via low-rank (LR) feature extraction and enforces smoothness priors via the Markov random field (MRF). Especially, we employ the mixture of Gaussian-based robust LR matrix factorization to simultaneously extract discriminative features and remove complex noises. Then, a classification map is obtained by applying a convolutional neural network with data augmentation on the extracted features, where local consistency is implicitly involved, and the insufficient label issue is alleviated. Finally, we refine the classification map by MRF to enforce contextual smoothness. We conduct experiments on two benchmark PolSAR data sets. Experimental results indicate that the proposed method achieves promising classification performance and preferable spatial consistency.
Haixia Bi, Jing Yao 0002, Zhiqiang Wei 0004, Danfeng Hong, Jocelyn Chanussot
IEEE Geosci. Remote. Sens. Lett.2
2022 Deep Unsupervised Blind Hyperspectral and Multispectral Data Fusion
abstract
Hyperspectral images (HSIs) usually have finer spectral resolution but coarser spatial resolution than multispectral images (MSIs). To obtain a desired HSI with higher spatial resolution, great research attention has been paid to achieving hyperspectral super-resolution by fusing the observed HSI with an auxiliary MSI of the same scene. However, most of the existing HSI-MSI fusion methods rely either on prior knowledge of the degradation model or on sufficient training data, hindering their practicality and interpretability. In this letter, we propose a novel unsupervised HSI-MSI fusion network with the ability of degradation adaptive learning, namely, UDALN. Specifically, we propose three modules to straightly encode the spatial and spectral transformations across resolutions, i.e., SpaDnet, SpeUnet, and SpeDnet. Through an elaborately designed three-stage unsupervised training strategy, the estimated network parameters can exhibit clear physical meanings of degradation processes and therefore help guarantee a faithful reconstruction of the desired HSI. The experimental results on two widely used hyperspectral datasets demonstrate the effectiveness of our method in comparison to the state-of-the-art HSI-MSI fusion models. (Code available athttps://github.com/JiaxinLiCAS/UDALN_GRSL.)
Jiaxin Li 0002, Jing Yao 0002, Lianru Gao, Danfeng Hong
IEEE Geosci. Remote. Sens. Lett.3
2022 A Disjoint Samples-Based 3D-CNN With Active Transfer Learning for Hyperspectral Image Classification
abstract
Convolutional Neural Networks (CNNs) have been extensively studied for Hyperspectral Image Classification (HSIC). However, CNNs are critically attributed to a large number of labeled training samples, which outlays high costs in terms of time and resources. Moreover, CNNs are trained on some samples and have been tested on the entire HSI. Perhaps, the entire HSI is taken into account at test time to appropriately generate the ground truth maps. In order to obtain a higher accuracy while considering the limited availability of training samples and disjoint validation and test samples, this work proposes a fast and compact 3D CNN-based Active Learning (AL) for HSIC that integrates both deep transfer learning and AL into a unified framework. In the proposed methodology, a 3D CNN model is trained with very few training samples (i.e., 5%, only) and in the next phase, the most informative and heterogeneous samples are queried from the validation set (candidate set) based on the fuzziness, mutual information and breaking ties of the trained model. The 3D CNN model is later fine-tuned (rather retraining from scratch) with the new training samples (i.e., 200 samples are selected in each iteration) to reduce the computational cost. The proposed method has been compared with the state-of-the-art traditional and deep models proposed for HSIC. Experimental results proved the superiority of our proposed method on several benchmark HSI datasets with significantly fewer labeled samples. Matlab demo can be accessed on GitHub: github.com/mahmad00.
Muhammad Ahmad 0002, Usman Ghous, Danfeng Hong, Adil Khan 0001, Jing Yao 0002, Shaohua Wang 0001, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.5
2022 Multimodal Hyperspectral Unmixing: Insights From Attention Networks
abstract
Deep learning (DL) has aroused wide attention in hyperspectral unmixing (HU) owing to its powerful feature representation ability. As a representative of unsupervised DL approaches, autoencoder (AE) has been proven to be effective to better capture nonlinear components of hyperspectral images than the traditional model-driven linearized methods. However, only using hyperspectral images for unmixing fails to distinguish objects in complex scene, especially for different endmembers with similar materials. To overcome this limitation, we propose a novel multimodal unmixing network for hyperspectral images, called MUNet, by considering the height differences of light detection and ranging (LiDAR) data in a squeeze-and-excitation (SE)-driven attention fashion to guide the unmixing process, yielding performance improvement. MUNet is capable of fusing multimodal information and using the attention map derived by LiDAR to aid network that focuses on more discriminative and meaningful spatial information regarding scenes. Moreover, attribute profile (AP) is adopted to extract the geometrical structures of different objects to better model the spatial information of LiDAR. Experimental results on synthetic and real datasets demonstrate the effectiveness and superiority of the proposed method compared with several state-of-the-art unmixing algorithms. The codes will be available athttps://github.com/hanzhu97702/IEEE_TGRS_MUNet, contributing to the remote sensing community.
Zhu Han 0002, Danfeng Hong, Lianru Gao, Jing Yao 0002, Bing Zhang 0001, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.4
2022 SpectralFormer: Rethinking Hyperspectral Image Classification With Transformers
abstract
Hyperspectral (HS) images are characterized by approximately contiguous spectral information, enabling the fine identification of materials by capturing subtle spectral discrepancies. Owing to their excellent locally contextual modeling ability, convolutional neural networks (CNNs) have been proven to be a powerful feature extractor in HS image classification. However, CNNs fail to mine and represent the sequence attributes of spectral signatures well due to the limitations of their inherent network backbone. To solve this issue, we rethink HS image classification from a sequential perspective with transformers, and propose a novel backbone network called \ul{SpectralFormer}. Beyond band-wise representations in classic transformers, SpectralFormer is capable of learning spectrally local sequence information from neighboring bands of HS images, yielding group-wise spectral embeddings. More significantly, to reduce the possibility of losing valuable information in the layer-wise propagation process, we devise a cross-layer skip connection to convey memory-like components from shallow to deep layers by adaptively learning to fuse "soft" residuals across layers. It is worth noting that the proposed SpectralFormer is a highly flexible backbone network, which can be applicable to both pixel- and patch-wise inputs. We evaluate the classification performance of the proposed SpectralFormer on three HS datasets by conducting extensive experiments, showing the superiority over classic transformers and achieving a significant improvement in comparison with state-of-the-art backbone networks. The codes of this work will be available at https://github.com/danfenghong/IEEE_TGRS_SpectralFormer for the sake of reproducibility.
Danfeng Hong, Zhu Han 0002, Jing Yao 0002, Lianru Gao, Bing Zhang 0001, Antonio Plaza, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.3
2022 Semi-Active Convolutional Neural Networks for Hyperspectral Image Classification
abstract
Owing to the powerful data representation ability of deep learning (DL) techniques, tremendous progress has been recently made in hyperspectral image (HSI) classification. Convolutional neural network (CNN), as a main part of the DL family, has been proven to be considerably effective to extract spatial-spectral features for HSIs. Nevertheless, its classification performance, to a great extent, depends on the quality and quantity of samples in the network training process. To select those samples, either labeled or unlabeled, that can be used to enhance the generalization ability of CNNs and further improve the classification accuracy, we propose an iterative semi-supervised CNNs framework by means of active learning and superpixel segmentation techniques, dubbed as semi-active CNNs (SA-CNNs), for HSI classification. More specifically, we start to pre-train a CNNs-based model on a small-scale unbiased labeled set and infer unlabeled data using the trained model, i.e., generating pseudo-labels. Then, the reliable samples, which consist of two parts: high label-homogeneity and most informativeness, are actively selected from superpixel segments. These selected labeled and unlabeled samples with their labels and pseudo-labels are re-fed into the next-round network training. Moreover, three different schedules, i.e.,log-,exp-, andlinear-schedules, are progressively adopted to fully explore their potentials in sample selection, until a labeling budget is finally reached. Extensive experiments are conducted on three benchmark HSI datasets, demonstrating substantial performance improvements of the proposed SA-CNNs over other similar competitors.
Jing Yao 0002, Xiangyong Cao, Danfeng Hong, Xin Wu 0001, Deyu Meng, Jocelyn Chanussot, Zongben Xu
IEEE Trans. Geosci. Remote. Sens.1
2022 Sparsity-Enhanced Convolutional Decomposition: A Novel Tensor-Based Paradigm for Blind Hyperspectral Unmixing
abstract
Blind hyperspectral unmixing (HU) has long been recognized as a crucial component in analyzing the hyperspectral imagery (HSI) collected by airborne and spaceborne sensors. Due to the highly ill-posed problems of such a blind source separation scheme and the effects of spectral variability in hyperspectral imaging, the ability to accurately and effectively unmixing the complex HSI still remains limited. To this end, this article presents a novel blind HU model, called sparsity-enhanced convolutional decomposition (SeCoDe), by jointly capturing spatial–spectral information of HSI in a tensor-based fashion. SeCoDe benefits from two perspectives. On the one hand, the convolutional operation is employed in SeCoDe to locally model the spatial relation between the targeted pixel and its neighbors, which can be well explained by spectral bundles that are capable of addressing spectral variabilities effectively. It maintains, on the other hand, physically continuous spectral components by decomposing the HSI along with the spectral domain. With sparsity-enhanced regularization, an alternative optimization strategy with alternating direction method of multipliers (ADMM)-based optimization algorithm is devised for efficient model inference. Extensive experiments conducted on three different data sets demonstrate the superiority of the proposed SeCoDe compared to previous state-of-the-art methods. We will also release the code athttps://github.com/danfenghong/IEEE_TGRS_SeCoDeto encourage the reproduction of the given results.
Jing Yao 0002, Danfeng Hong, Lin Xu 0001, Deyu Meng, Jocelyn Chanussot, Zongben Xu
IEEE Trans. Geosci. Remote. Sens.1
2022 Endmember-Guided Unmixing Network (EGU-Net): A General Deep Learning Framework for Self-Supervised Hyperspectral Unmixing
abstract
Over the past decades, enormous efforts have been made to improve the performance of linear or nonlinear mixing models for hyperspectral unmixing (HU), yet their ability to simultaneously generalize various spectral variabilities (SVs) and extract physically meaningful endmembers still remains limited due to the poor ability in data fitting and reconstruction and the sensitivity to various SVs. Inspired by the powerful learning ability of deep learning (DL), we attempt to develop a general DL approach for HU, by fully considering the properties of endmembers extracted from the hyperspectral imagery, called endmember-guided unmixing network (EGU-Net). Beyond the alone autoencoder-like architecture, EGU-Net is a two-stream Siamese deep network, which learns an additional network from the pure or nearly pure endmembers to correct the weights of another unmixing network by sharing network parameters and adding spectrally meaningful constraints (e.g., nonnegativity and sum-to-one) toward a more accurate and interpretable unmixing solution. Furthermore, the resulting general framework is not only limited to pixelwise spectral unmixing but also applicable to spatial information modeling with convolutional operators for spatial-spectral unmixing. Experimental results conducted on three different datasets with the ground truth of abundance maps corresponding to each material demonstrate the effectiveness and superiority of the EGU-Net over state-of-the-art unmixing algorithms. The codes will be available from the website: https://github.com/danfenghong/IEEE_TNNLS_EGU-Net.
Danfeng Hong, Lianru Gao, Jing Yao 0002, Naoto Yokoya, Jocelyn Chanussot, Uta Heiden, Bing Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.3
2021 Multimodal Convolutional Neural Networks with Cross-Channel Reconstruction
abstract
With the ever-growing availability of remote sensing (RS) data from either satellite or airborne sensors, simultaneous processing and analysis of multimodal data have been paid more and more attention by researchers in various RS-related applications. In this paper, we propose a multimodal convolutional neural network with an advanced cross-channel reconstruction module, called CCR-Net. As the name suggests, CCR-Net enables a more compact fusion of different RS data sources by the means of the reconstruction strategy across modalities that can mutually exchange information in a more effective way. Experiment are conducted on a widely-used dataset, including hyperspectral and Light Detection and Ranging (LiDAR) data, i.e., Houston2013, to verify the effectiveness and superiority of the proposed CCR - N et in comparison with several state-of-the-art baseline methods.
Danfeng Hong, Xin Wu 0001, Jing Yao 0002, Lianru Gao, Bing Zhang 0001, Jocelyn Chanussot
IGARSS3
2021 An Enhanced 3-D Discrete Wavelet Transform for Hyperspectral Image Classification
abstract
In the classification of hyperspectral image (HSI), there exists a common issue that the collected HSI data set is always contaminated by various noise (e.g., Gaussian, stripe, and deadline), degrading the classification results. To tackle this issue, we modify the 3-dimensional discrete wavelet transform (3DDWT) method by considering the noise effect on feature quality and propose an enhanced 3DDWT (E-3DDWT) approach to extract the feature and meanwhile alleviate the noise. Specifically, the proposed E-3DDWT method first applies classical 3DDWT method to the HSI data cube and thus can generate eight subcubes in each level. Then, the stripe noise is concentrated into several subcubes due to its spatial vertical property. Finally, we abandon these subcubes and obtain the feature cube by stacking the remaining ones. After acquiring the feature, we then adopt the convolutional neural network (CNN) model with an active learning strategy for classification since CNN has been verified to be a state-of-the-art feature extraction method for HSI classification, and active learning strategy can alleviate the insufficient labeled sample issue to some extent. In addition, we apply the Markov random field to enhance the final categorized results. Experiments on two synthetically striped data sets show that our proposed approach achieves better categorized results than other advanced methods.
Xiangyong Cao, Jing Yao 0002, Xueyang Fu, Haixia Bi, Danfeng Hong
IEEE Geosci. Remote. Sens. Lett.2
2021 Spectral Superresolution of Multispectral Imagery With Joint Sparse and Low-Rank Learning
abstract
Extensive attention has been widely paid to enhance the spatial resolution of hyperspectral (HS) images with the aid of multispectral (MS) images in remote sensing. However, the ability in the fusion of HS and MS images remains to be improved, particularly in large-scale scenes, due to the limited acquisition of HS images. Alternatively, we super-resolve MS images in the spectral domain by the means of partially overlapped HS images, yielding a novel and promising topic: spectral superresolution (SSR) of MS imagery. This is challenging and less investigated task due to its high ill-posedness in inverse imaging. To this end, we develop a simple but effective method, called joint sparse and low-rank learning (J-SLoL), to spectrally enhance MS images by jointly learning low-rank HS-MS dictionary pairs from overlapped regions. J-SLoL infers and recovers the unknown HS signals over a larger coverage by sparse coding on the learned dictionary pair. Furthermore, we validate the SSR performance on three HS-MS data sets (two for classification and one for unmixing) in terms of reconstruction, classification, and unmixing by comparing with several existing state-of-the-art baselines, showing the effectiveness and superiority of the proposed J-SLoL algorithm. Furthermore, the codes and data sets will be available at https://github.com/danfenghong/IEEE_TGRS_J-SLoL, contributing to the remote sensing (RS) community.
Lianru Gao, Danfeng Hong, Jing Yao 0002, Bing Zhang 0001, Paolo Gamba, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.3
2021 More Diverse Means Better: Multimodal Deep Learning Meets Remote-Sensing Imagery Classification
abstract
Classification and identification of the materials lying over or beneath the earth's surface have long been a fundamental but challenging research topic in geoscience and remote sensing (RS), and have garnered a growing concern owing to the recent advancements of deep learning techniques. Although deep networks have been successfully applied in single-modality-dominated classification tasks, yet their performance inevitably meets the bottleneck in complex scenes that need to be finely classified, due to the limitation of information diversity. In this work, we provide a baseline solution to the aforementioned difficulty by developing a general multimodal deep learning (MDL) framework. In particular, we also investigate a special case of multi-modality learning (MML)-cross-modality learning (CML) that exists widely in RS image classification applications. By focusing on “what,” “where,” and “how” to fuse, we show different fusion strategies as well as how to train deep networks and build the network architecture. Specifically, five fusion architectures are introduced and developed, further being unified in our MDL framework. More significantly, our framework is not only limited to pixel-wise classification tasks but also applicable to spatial information modeling with convolutional neural networks (CNNs). To validate the effectiveness and superiority of the MDL framework, extensive experiments related to the settings of MML and CML are conducted on two different multimodal RS data sets. Furthermore, the codes and data sets will be available at https://github.com/danfenghong/IEEE_TGRS_MDL-RS, contributing to the RS community.
Danfeng Hong, Lianru Gao, Naoto Yokoya, Jing Yao 0002, Jocelyn Chanussot, Qian Du 0001, Bing Zhang 0001
IEEE Trans. Geosci. Remote. Sens.4
2021 Graph Convolutional Networks for Hyperspectral Image Classification
abstract
Convolutional neural networks (CNNs) have been attracting increasing attention in hyperspectral (HS) image classification due to their ability to capture spatial-spectral feature representations. Nevertheless, their ability in modeling relations between the samples remains limited. Beyond the limitations of grid sampling, graph convolutional networks (GCNs) have been recently proposed and successfully applied in irregular (or nongrid) data representation and analysis. In this article, we thoroughly investigate CNNs and GCNs (qualitatively and quantitatively) in terms of HS image classification. Due to the construction of the adjacency matrix on all the data, traditional GCNs usually suffer from a huge computational cost, particularly in large-scale remote sensing (RS) problems. To this end, we develop a new minibatch GCN (called miniGCN hereinafter), which allows to train large-scale GCNs in a minibatch fashion. More significantly, our miniGCN is capable of inferring out-of-sample data without retraining networks and improving classification performance. Furthermore, as CNNs and GCNs can extract different types of HS features, an intuitive solution to break the performance bottleneck of a single model is to fuse them. Since miniGCNs can perform batchwise network training (enabling the combination of CNNs and GCNs), we explore three fusion strategies: additive fusion, elementwise multiplicative fusion, and concatenation fusion to measure the obtained performance gain. Extensive experiments, conducted on three HS data sets, demonstrate the advantages of miniGCNs over GCNs and the superiority of the tested fusion strategies with regard to the single CNN or GCN models. The codes of this work will be available at https://github.com/danfenghong/IEEE_TGRS_GCN for the sake of reproducibility.
Danfeng Hong, Lianru Gao, Jing Yao 0002, Bing Zhang 0001, Antonio Plaza, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.3
2021 Multimodal GANs: Toward Crossmodal Hyperspectral-Multispectral Image Segmentation
abstract
This article addresses the problem of semantic segmentation with limited cross-modality data in large-scale urban scenes. Most prior works have attempted to address this issue by using multimodal deep neural networks (DNNs). However, their ability to effectively blending different properties across multimodalities and robustly learning representations from complex scenes remains limited, particularly in the absence of sufficient and well-annotated training images. This leads to a challenge related to cross-modality learning with multimodal DNNs. To this end, we introduce two novel plug-and-play units in the network: self-generative adversarial networks (GANs) module and mutual-GANs module, to learn perturbation-insensitive feature representations and to eliminate the gap between multimodalities, respectively, yielding more effective and robust information transfer. Furthermore, a patchwise progressive training strategy is devised to enable effective network learning with limited samples. We evaluate the proposed network on two multimodal (hyperspectral and multispectral) overhead image data sets and achieve a significant improvement in comparison with several state-of-the-art methods.
Danfeng Hong, Jing Yao 0002, Deyu Meng, Zongben Xu, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.2
2020 Cross-Attention in Coupled Unmixing Nets for Unsupervised Hyperspectral Super-Resolution
Jing Yao 0002, Danfeng Hong, Jocelyn Chanussot, Deyu Meng, Xiao Xiang Zhu 0001, Zongben Xu
ECCV (29)1
2020 Unsupervised Hyperspectral Embedding by Learning a Deep Regression Network
abstract
This work presents a novel hyperspectral embedding technique by learning a deep regression network in an unsupervised fashion, which aims at reducing the computational complexity and storage-costing of traditional manifold embedding methods as well as improving the representation ability of spectral signatures effectively. The proposed method attempts to learn an explicit and unified nonlinear mapping from all patch-wise correspondences of original hyperspectral data and dimension-reduced products generated by some existing manifold learning approaches. This process can be well performed by means of a deep regression model. The learned model is not only capable of locally capturing the manifold structure of the whole hyperspectral image from densely patch-based random sampling but also better applicable to high-efficient out-of-sample inference. Experimental results conducted on the real hyperspectral data demonstrate the effectiveness and superiority of the proposed hyperspectral embedding technique.
Danfeng Hong, Jing Yao 0002, Jocelyn Chanussot, Xiao Xiang Zhu 0001
IGARSS2
2020 Locally Linear Reconstruction for Spectral Enhancement Using Limited Pixel-to-Pixel Multispectral and Hyperspectral Data
abstract
Recently, spectral enhancement of multispectral imagery has attracted a growing interest in the remote sensing (RS) community. Without any prior knowledge, this task is highly ill-conditioned in inverse problems. To this end, we develop a simple but effective method, called locally linear reconstruction (LLR), to spectrally enhance the multispectral imagery (MSI) using partially overlapped hyperspectral data. LLR learns reconstruction coefficients of each pixel from the MSI and shares the same weights to recover the unknown hyperspectral signals over a larger coverage. We validate the performance of the proposed LLR on the real hyperspectral data in comparison with several state-of-the-art baselines, demonstrating its effectiveness and superiority.
Danfeng Hong, Jing Yao 0002, Renlong Hang, Jocelyn Chanussot
IGARSS2
2020 Hyperspectral Image Classification With Convolutional Neural Network and Active Learning
abstract
Deep neural network has been extensively applied to hyperspectral image (HSI) classification recently. However, its success is greatly attributed to numerous labeled samples, whose acquisition costs a large amount of time and money. In order to improve the classification performance while reducing the labeling cost, this article presents an active deep learning approach for HSI classification, which integrates both active learning and deep learning into a unified framework. First, we train a convolutional neural network (CNN) with a limited number of labeled pixels. Next, we actively select the most informative pixels from the candidate pool for labeling. Then, the CNN is fine-tuned with the new training set constructed by incorporating the newly labeled pixels. This step together with the previous step is iteratively conducted. Finally, Markov random field (MRF) is utilized to enforce class label smoothness to further boost the classification performance. Compared with the other state-of-the-art traditional and deep learning-based HSI classification methods, our proposed approach achieves better performance on three benchmark HSI data sets with significantly fewer labeled samples.
Xiangyong Cao, Jing Yao 0002, Zongben Xu, Deyu Meng
IEEE Trans. Geosci. Remote. Sens.2
2019 Nonconvex-Sparsity and Nonlocal-Smoothness-Based Blind Hyperspectral Unmixing
abstract
Blind hyperspectral unmixing (HU), as a crucial technique for hyperspectral data exploitation, aims to decompose mixed pixels into a collection of constituent materials weighted by the corresponding fractional abundances. In recent years, nonnegative matrix factorization (NMF) based methods have become more and more popular for this task and achieved promising performance. Among these methods, two types of properties upon the abundances, namely the sparseness and the structural smoothness, have been explored and shown to be important for blind HU. However, all of previous methods ignores another important insightful property possessed by a natural hyperspectral images (HSI), non-local smoothness, which means that similar patches in a larger region of an HSI are sharing the similar smoothness structure. Based on previous attempts on other tasks, such a prior structure reflects intrinsic configurations underlying a HSI, and is thus expected to largely improve the performance of the investigated HU problem. In this paper, we firstly consider such prior in HSI by encoding it as the nonlocal total variation (NLTV) regularizer. Furthermore, by fully exploring the intrinsic structure of HSI, we generalize NLTV to non-local HSI TV (NLHTV) to make the model more suitable for the bind HU task. By incorporating these two regularizers, together with a non-convex log-sum form regularizer characterizing the sparseness of abundance maps, to the NMF model, we propose novel blind HU models named NLTV/NLHTV and log-sum regularized NMF (NLTV-LSRNMF/NLHTV-LSRNMF), respectively. To solve the proposed models, an efficient algorithm is designed based on alternative optimization strategy (AOS) and alternating direction method of multipliers (ADMM). Extensive experiments conducted on both simulated and real hyperspectral data sets substantiate the superiority of the proposed approach over other competing ones for blind HU task.
Jing Yao 0002, Deyu Meng, Qian Zhao 0002, Wenfei Cao, Zongben Xu
IEEE Trans. Image Process.1
2018 Robust subspace clustering via penalized mixture of Gaussians
Jing Yao 0002, Xiangyong Cao, Qian Zhao 0002, Deyu Meng, Zongben Xu
Neurocomputing1
2018 A tensor-based nonlocal total variation model for multi-channel image recovery
Wenfei Cao, Jing Yao 0002, Jian Sun 0009
Signal Process.2