EDBT 2026 Demo / reviewers in the wild / expert
Pedram Ghamisi
dblp:117/6185
· DBLP profile ↗
144ranked-venue papers
23as first author
74since 2021 · last 2026
0000-0003-1203-741XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 125 · 22 first-author · 57 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Generating Any Changes in the Noise DomainabstractChange detection is essential in Earth observation, yet current models heavily rely on large-scale annotated datasets. Generative models offer a promising alternative by synthesizing training data, but generating temporally coherent image pairs with realistic, semantically meaningful changes remains a significant challenge. Existing approaches typically simulate changes by generating pre- and post-change label maps using either heuristic rules (e.g., copy-pasting) or text prompts. However, the former offers limited change diversity, while the latter often fails to maintain spatial consistency between image pairs. We observe that the noise space of diffusion models encodes strong generative capacity and spatial controllability: localized perturbations in the noise can yield meaningful, interpretable changes in corresponding image regions. Motivated by this, we propose Noise2Change, a framework for simulating change directly in the noise domain. The key idea is to manipulate the semantic composition of the initial noise sampled from the noise domain, such that the diffusion process generates structurally consistent pre- and post-change images reflecting realistic transformations. Since the unperturbed noise is shared between both images, the resulting pairs exhibit strong temporal alignment and semantic coherence, effectively addressing the trade-off between realism and consistency. Concretely, we employ a discrete diffusion model to extract high-level semantics from the initial noise. Guided by these semantics, we introduce a change simulation strategy that optimizes the noise to encode intended changes. The modified noise is then used to drive the diffusion process, yielding pre- and post-change label maps with natural structural transitions. These maps are passed through a unified framework for image generation and label refinement, producing highly aligned image-label pairs. Our framework supports diverse change types across a wide range of scenarios. Extensive experiments on multiple change detection tasks demonstrate that our method achieves superior performance compared to existing generative approaches. Jun Yue 0004, Pedram Ghamisi, Weiying Xie, Leyuan Fang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Real-CD: Change Detection Under Real-World Complex Interference via Dynamic Distribution CorrectionabstractWhile change detection (CD) is crucial for tracking dynamic changes on the Earth's surface, it faces substantial challenges in real-world settings caused by seasonal variations and sensor-related interference. Current CD models often suffer performance degradation under such conditions, mainly due to two key challenges. First, most existing CD datasets lack sufficient temporal and environmental diversity, as they are typically collected over constrained time spans. This limits the models' ability to generalize across varying conditions. Second, many CD methods are heavily data-driven and rely on simplified assumptions, leading to models that are not adequately designed to handle the complex, heterogeneous nature of real-world scenarios. Together, these challenges restrict the robustness and practical applicability of current CD approaches. To overcome these challenges, we make the following contributions in this paper: 1) Regarding data diversity, we construct a comprehensive benchmark by introducing five typical perturbations (fog, snow, motion blur, Gaussian noise, and impulse noise) into three classical CD datasets and supplementing them with a real-world seasonal dataset, resulting in 75 interference-rich scenarios. This enables a systematic evaluation under diverse real-world conditions, revealing that such perturbations induce severe distribution shifts across both temporal phases and hierarchical network layers, leading to substantial performance degradation in existing models. 2) Algorithmically, we propose Real-CD, a novel method specifically designed to address distribution shifts in real-world CD. The core of Real-CD is to leverage bi-temporal correlations to perform adaptive distribution alignment across hierarchical layers and temporal phases. Specifically, we propose the Distribution Shifts Alleviation Module (DSAM) to correct distribution shifts. The DSAM captures bi-temporal differences and similarities to formulate temporal-specific adjustment strategies for each LayerNorm (LN) layer. To stabilize the optimization of DSAM, we propose the Distribution Consistency Optimization Strategy (DCOS), which introduces a flip-based auxiliary task that encourages the model to maintain distributional consistency under complex bi-temporal disturbances. Consequently, our method outperforms other state-of-the-art approaches and achieves the best performance on the proposed dataset. Our datasets and code implementation will be available at https://github.com/fangyee-ISALAB/Real-CD. Leyuan Fang, Pedram Ghamisi |
IEEE Trans. Image Process. | 3 |
| 2026 | Seed-to-Semantics: Few-Shot Prototype-Guided Progressive Learning for Hyperspectral and LiDAR ClassificationabstractDeep learning-based fusion of hyperspectral images (HSI) and LiDAR has achieved strong performance in multimodal remote sensing classification, but its success is heavily constrained by the high cost of pixel-wise annotation. In extremely label-scarce regimes, such as 2-5 labeled samples per class, conventional deep models are prone to severe overfitting, while standard semi-supervised learning (SSL) methods often suffer from confirmation bias because pseudo-labels are generated from unstable early-stage representations. To address these challenges, we propose Prototype-Guided Progressive Learning (PGPL), a unified framework for few-shot HSI-LiDAR classification. Instead of relying solely on model confidence in latent space, PGPL first constructs a reliable initialization pool directly in the original data domain using spectral-angle and elevation-consistency cues, and then progressively expands the training set through class-balanced pseudo-label admission and temporal confidence stabilization. In this way, the framework improves pseudo-label reliability during both initialization and subsequent self-training. Extensive experiments on three benchmark datasets demonstrate that PGPL consistently outperforms state-of-the-art supervised and semi-supervised baselines under the corresponding 2-5-shot settings, achieving overall accuracy gains of 4.64% points on Houston, 1.16% on Trento, and 3.92% on MUUFL over the strongest competing methods, while also yielding higher pseudo-label purity. The source code will be publicly available at https://github.com/zhangyiyan001/PGPL. Hongmin Gao 0001, Weiping Ding 0001, Pedram Ghamisi, Zhonghao Chen, Bing Zhang 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | MSHCCT: A Multiscale Compact Convolutional Network for High-Resolution Aerial Scene ClassificationabstractThe growing popularity of vision transformers (ViTs) in remote sensing image classification is due to their ability to effectively capture long-range dependencies. However, their high computational cost and memory footprint limit their applicability, particularly for small-scale datasets and resource-constrained environments. To address these challenges, we propose the multiscale multihead compact convolutional transformer (MSHCCT), a lightweight yet powerful model that integrates convolutional tokenization with small-scale ViTs to enhance multiscale feature representation while maintaining computational efficiency. Despite a modest increase in parameters and training time, MSHCCT achieves superior classification accuracy and robustness on high-resolution aerial scenes. Importantly, our approach eliminates the need for model pretraining, additional datasets, or multisensor data fusion, ensuring a computationally efficient and practical solution for remote sensing applications. The code will be made publicly available athttps://github.com/aj1365/MSHCCT Ali Jamali, Swalpa Kumar Roy, Bing Lu 0003, Leila Hashemi Beni, Nafiseh Kakhani, Pedram Ghamisi |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2025 | Auto-Prompting SAM for Weakly Supervised Landslide ExtractionabstractWeakly supervised landslide extraction aims to identify landslide regions from remote sensing data using models trained with weak labels, particularly image-level labels. However, it is often challenged by the imprecise boundaries of the extracted objects due to the lack of pixel-wise supervision and the properties of landslide objects. To tackle these issues, we propose a simple yet effective method by auto-prompting the Segment Anything Model (SAM), i.e., APSAM. Instead of depending on high-quality class activation maps (CAMs) for pseudo-labeling or fine-tuning SAM, our method directly yields fine-grained segmentation masks from SAM inference through prompt engineering. Specifically, it adaptively generates hybrid prompts from the CAMs obtained by an object localization network. To provide sufficient information for SAM prompting, an adaptive prompt generation (APG) algorithm is designed to fully leverage the visual patterns of CAMs, enabling the efficient generation of pseudo-masks for landslide extraction. These informative prompts are able to identify the extent of landslide areas (box prompts) and denote the centers of landslide objects (point prompts), guiding SAM in landslide segmentation. Experimental results on high-resolution aerial and satellite datasets demonstrate the effectiveness of our method, achieving improvements of at least 3.0% in F1 score and 3.69% in IoU compared to other state-of-the-art methods. The source codes and datasets will be available at https://github.com/zxk688. Xianping Ma, Weikang Yu, Pedram Ghamisi |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2025 | HZSCM: Hyperspectral Image Zero-Shot Classification via Vision-Language ModelsabstractMost hyperspectral image (HSI) classification methods assume that all classes in the test set are present during training. However, in real-world applications, acquiring labeled training samples is challenging. As a result, it is difficult for the training dataset to cover all possible land cover types, leading to the generalized zero-shot learning (GZSL) problem. Recently, vision-language models (VLMs) have provided rich semantic priors for land cover classes, offering promising potential for GZSL. However, two fundamental gaps hinder their application to HSI classification: the task paradigm gap, arising from the difference between image-level VLMs and the pixel-level HSI classification task; and the knowledge gap, due to the inconsistency between VLM features and HSI spectral–spatial representations. To bridge both gaps, a novel framework leveraging VLM semantic priors for GZSL in HSI classification is proposed, primarily using pseudo-labeling technique to provide knowledge for unseen classes. Specifically, a pseudo-label generation and enhancement module enables a paradigm transition from image-level understanding to pixel-level classification by incorporating HSI’s spatial information. A pseudo-label correction module then refines noisy labels using spectral cues to address the knowledge gap. Finally, a global learning strategy integrates pseudo-label distillation, supervised learning, and feature regularization to classify seen classes while enabling generalization to unseen ones. Experiments on benchmark HSI datasets demonstrate the proposed method’s superiority in generalized zero-shot classification. This work highlights the potential of VLMs in advancing HSI classification in practical applications. Lingbo Huang, Yushi Chen 0002, Zhaokui Li, Pedram Ghamisi, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Mask Approximation Net: A Novel Diffusion Model Approach for Remote Sensing Change CaptioningabstractRemote sensing (RS) image change description represents an innovative multimodal task within the realm of RS processing. This task not only facilitates the detection of alterations in surface conditions but also provides comprehensive descriptions of these changes, thereby improving human interpretability and interactivity. Current deep learning methods typically adopt a three-stage framework consisting of feature extraction, feature fusion, and change localization, followed by text generation. Most approaches focus heavily on designing complex network modules but lack solid theoretical guidance, relying instead on extensive empirical experimentation and iterative tuning of network components. This experience-driven design paradigm may lead to overfitting and design bottlenecks, thereby limiting the model’s generalizability and adaptability. To address these limitations, this article proposes a paradigm that shifts toward data distribution learning using diffusion models, reinforced by frequency-domain noise filtering, to provide a theoretically motivated and practically effective solution to multimodal RS change description. The proposed method primarily includes a simple multiscale change detection (CD) module, whose output features are subsequently refined by a well-designed diffusion model. Furthermore, we introduce a frequency-guided complex filter module to boost the model’s performance by managing high-frequency noise throughout the diffusion process. We validate the effectiveness of our proposed method across several datasets for RS CD and description, showcasing its superior performance compared to existing techniques. The code will be available athttps://github.com/sundongweiMaskApproxNet Dongwei Sun, Jing Yao 0002, Wu Xue, Changsheng Zhou, Pedram Ghamisi, Xiangyong Cao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | HSLabeling: Toward Efficient Labeling for Large-Scale Remote Sensing Image Segmentation With Hybrid Sparse LabelingabstractDense pixel-wise labeling of large-scale remote sensing images (RSI) is very time-consuming, while sparse labels (i.e., points, scribbles, or blocks) can be an efficient way to reduce labeling costs. Most existing sparse label-based methods adopt only one type of label for image segmentation, which cannot reflect the complex land covers in the RSI for training the model, thus leading to inferior segmentation performance. We observe that land covers with different shapes and complexity can be optimally represented by different sparse labels. Inspired by this observation, we propose a novel sparse labeling framework, termed Hybrid Sparse Labeling (HSLabeling), for large-scale RSI segmentation. Our HSLabeling can adaptively select the optimal hybrid sparse labels for different land covers, according to labeling cost and segmentation contribution of different sparse labels. Specifically, we first propose a label segmentation contribution information estimation module that estimates the information of different sparse labels according to the diversity and shape of land covers. After that, we propose an Optimal Hybrid Labeling Strategy (OHLS) to assign optimal types of labels for different land covers. In the OHLS, label assignment is formulated as an optimization problem that trades off label segmentation contribution information and labeling cost. We employ the greedy algorithm to efficiently solve the optimization problem and adaptively assign labels for varied land covers. Extensive experiments on three large-scale RSI datasets have demonstrated that our HSLabeling achieves almost fully supervised performance with extremely low labeling costs. In addition, compared with the single type sparse label, HSLabeling can also utilize much lower labeling costs to obtain the same performance. The source code is available at https://github.com/linjiaxing99/HSLabeling. Jiaxing Lin, Zhen Yang 0026, Yinglong Yan, Pedram Ghamisi, Weiying Xie, Leyuan Fang |
IEEE Trans. Image Process. | 5 |
| 2024 | Towards 3D Hyperspectral ImagingabstractWe argue that traditional 2D hyperspectral imaging is not adapted to many modern challenges. With the rise of high spatial resolution, hyperspectral sensors mounted on different platforms (e.g. drones, terrestrial, satellites) and innovative applications (e.g. urban mapping, mining monitoring), projections, occlusions, perspective effects and data processing limit the use of 2D hyperspectral imaging. We propose that 3D hyperclouds, in which Lidar or photogrammetric point clouds are augmented with hyperspectral attributes, can address numerous of these challenges. We demonstrate the benefits of hyperclouds and dedicated machine learning architectures with several realistic examples. Richard Gloaguen, Aldino Rizaldy, Ahmed J. Afifi, Sandra Lorenz, Samuel T. Thiele, Moritz Kirsch, Pedram Ghamisi |
IGARSS | 7 |
| 2024 | PolSARConvMixer: A Channel and Spatial Mixing Convolutional Algorithm for PolSAR Data ClassificationabstractGiven the exceptional effectiveness of deep Convolutional Neural Networks (CNNs) in computer vision, there has been a recent surge of interest in employing CNNs for various applications in image classification. Additionally, scientists are exploring the potential of vision transformers for Earth observation applications, owing to their recent tremendous success. However, a major challenge with vision transformers is their increased demand for training data compared to CNN classifiers. Furthermore, vision transformers exhibit quadratic complexity and necessitate substantial hardware resources. In the context of PolSAR image classification, we propose the PolSARConvMixer—a fundamental framework that segregates the mixing of spatial and channel dimensions, maintains uniform size and resolution across the network and directly processes PolSAR image patches as input. Our experiments on two PolSAR data benchmarks, namely Flevoland and San Francisco, demonstrate the significant superiority of the developed PolSARConvMixer over several other algorithms, including AlexNet, ResNet, FNet, a 2D CNN, and PolSARFormer. Ali Jamali, Swalpa Kumar Roy, Bing Lu 0003, Avik Bhattacharya, Pedram Ghamisi |
IGARSS | 5 |
| 2024 | Dimensional Dilemma: Navigating the Fusion of Hyperspectral and Lidar Point Cloud Data for Optimal Precision - 2D vs. 3DabstractDespite the extensive body of research conducted on the fusion of lidar and hyperspectral data for land cover classification in urban areas, the predominant approach has been the utilization of rasterized lidar data merged with hyperspectral data. This image-centric methodology tends to overlook the primary advantage inherent in lidar technology—namely, the production of 3D point cloud data. In our work, we present a framework demonstrating how we infer semantic information from 3D point cloud data, comprising both lidar and hyperspectral features—a concept we refer to as a 3D hyperspectral point cloud. We illustrate the generation of hyperspectral point clouds and evaluate the performance of various deep learning models for point learning. Our findings on the original test data of the 2018 IEEE GRSS Data Fusion Challenge, disclosed by IEEE Image Analysis and Data Fusion Technical Committee, indicate that recent deep learning models not only produce better shapes for predicted objects but also yield more precise semantic information. Finally, we plan to release the 3D hyperspectral point cloud data to the community, hoping to inspire future studies on data fusion in the point cloud domain. Aldino Rizaldy, Ahmed J. Afifi, Pedram Ghamisi, Richard Gloaguen |
IGARSS | 3 |
| 2024 | MineNet-CD: Global Mining Change Detection DatasetabstractMining change detection requires dedicated datasets because it includes unique objects such as pits or quarries, tailings dams, overburden, processing plants, haul roads, access roads, buildings/sheds, mining and blasting equipment, and plants, among others. This paper introduces a benchmark, large dataset for mining change detection, termed the MineNet-CD, to facilitate large-scale change detection. The proposed dataset contains a total of 100 high-resolution bi-temporal mining images from all over the world with corresponding ground truth. Unlike existing datasets, the images of MineNet-CD exhibit significant background, topological variation, and well-defined ground truth that ignores insignificant change maps. The work analyzes the efficacy of state-of-the-art deep learning methods for change detection. The results demonstrate that more efficient and advanced networks are required to accurately predict the change maps. The dataset and code are available at https://github.com/EricYu97/MineNet-CD. Weikang Yu, Samiran Das, Aldino Rizaldy, Richard Gloaguen, Pedram Ghamisi |
IGARSS | 6 |
| 2024 | A cross-modal feature aggregation and enhancement network for hyperspectral and LiDAR joint classification
Hongmin Gao 0001, Jun Zhou 0001, Pedram Ghamisi, Shufang Xu, Bing Zhang 0001 |
Expert Syst. Appl. | 5 |
| 2024 | Open Set Recognition in Real World
Zhen Yang 0026, Jun Yue 0004, Pedram Ghamisi, Shiliang Zhang, Jiayi Ma 0001, Leyuan Fang |
Int. J. Comput. Vis. | 3 |
| 2024 | Spatial-Gated Multilayer Perceptron for Land Use and Land Cover MappingabstractDue to its capacity to recognize detailed spectral differences, hyperspectral data have been extensively used for precise Land Use Land Cover (LULC) mapping. However, recent multi-modal methods have shown their superior classification performance over the algorithms that use single data sets. On the other hand, Convolutional Neural Networks (CNNs) are models extensively utilized for the hierarchical extraction of features. Vision transformers (ViTs), through a self-attention mechanism, have recently achieved superior modeling of global contextual information compared to CNNs. However, to harness their image classification strength, ViTs require substantial training datasets. In cases where the available training data is limited, current advanced multi-layer perceptrons (MLPs) can provide viable alternatives to both deep CNNs and ViTs. In this paper, we developed the SGU-MLP, a deep learning algorithm that effectively combines MLPs and spatial gating units (SGUs) for precise Land Use Land Cover (LULC) mapping using multi-modal data from multi-spectral, LiDAR, and hyperspectral data. Results illustrated the superiority of the developed SGU-MLP classification algorithm over several CNN and CNN-ViT-based models, including HybridSN, ResNet, iFormer, EfficientFormer, and CoAtNet. The SGU-MLP classification model consistently outperformed the benchmark CNN and CNN-ViT-based algorithms. The code will be made publicly available at https: //github.com/aj1365/SGUMLP. Ali Jamali, Swalpa Kumar Roy, Danfeng Hong, Peter M. Atkinson, Pedram Ghamisi |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2024 | Attention Graph Convolutional Network for Disjoint Hyperspectral Image ClassificationabstractConvolutional Neural Networks (CNNs) are employed extensively in remote sensing due to their capacity to capture intricate features from a broad range of object patterns, irrespective of object size, shape or color. These networks excel at extracting high-frequency spectral information such as angles, edges and outlines. The classification boundary zone, however, becomes hazy for CNNs because they learn characteristics by means of a fixed shape kernel concentrated on the central pixel, and can perform poorly in image classification at class boundaries. Additionally, CNNs are not designed to capture global relations. Thus, in this letter, we propose an Attention Graph Convolutional Network (Attention-GCN) as a solution to the aforementioned shortcomings. The developed model illustrated a high level of superiority over several CNN and ViT-based models. For example, in the Augsburg data benchmark, the developed algorithm exhibited an average accuracy of 61.11%, substantially outperforming other models such as HybridSN, iFormer, Efficient Former, GCN, CoAtNet, 2D-CNN, 3D-CNN, and ResNet by approximately 9, 13, 14, 15, 18, 24, 25 and 29 percentage points, respectively. The code will be made publicly available at https://github.com/aj1365/AGCN. Ali Jamali, Swalpa Kumar Roy, Danfeng Hong, Peter M. Atkinson, Pedram Ghamisi |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2024 | Neighborhood Attention Makes the Encoder of ResUNet Stronger for Accurate Road ExtractionabstractIn the domain of remote sensing image interpretation, road extraction from high-resolution aerial imagery has already been a hot research topic. Although deep CNNs have presented excellent results for semantic segmentation, the efficiency and capabilities of vision transformers are yet to be fully researched. As such, for accurate road extraction, a deep semantic segmentation neural network that utilizes the abilities of residual learning, HetConvs, UNet, and vision transformers, which is called ResUNetFormer, is proposed in this letter. The developed ResUNetFormer is evaluated on various cutting-edge deep learning-based road extraction techniques on the public Massachusetts road dataset. Statistical and visual results demonstrate the superiority of the ResUNetFormer over the state-of-the-art CNNs and vision transformers for segmentation. The code will be made available publicly at https://github.com/aj1365/ResUNetFormer. Ali Jamali, Swalpa Kumar Roy, Jonathan Li 0001, Pedram Ghamisi |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | LVM-StARS: Large Vision Model Soft Adaption for Remote Sensing Scene ClassificationabstractRecently, both large language models and large vision models (LVMs) have gained significant attention. Trained on large-scale datasets, these large models have showcased remarkable capabilities across various research domains. To enhance the accuracy of remote sensing (RS) scene classification, LVM-based methods are explored in this letter. Due to the differences between RS images and natural images, simply transferring LVMs to RS tasks is impractical. Therefore, we conducted research on relevant techniques and appended learnable prompt tokens to the input tokens while freezing the backbone weights, reducing the parameter scale and making the LVM weights easier to harness and to transfer. In consideration of latent catastrophic forgetting issues induced by ordinary finetuning techniques and the inherent complexity and redundancy of RS images, we introduced soft adaption mechanisms between backbone layers based on prompt tuning technique and implemented the first LVM tuning method, namely, the Large Vision Model Soft Adaption for RS scene classification (LVM-StARS)-Deep and the LVM-StARS-Shallow to make LVMs more suitable for RS scene classification tasks. The proposed methods are evaluated on two popular RS scene classification datasets, and the experimental results indicate that the proposed method outperforms other state-of-the-art methods. The experimental results demonstrate that our proposed method enhances overall accuracy (OA) by 1.71%–3.94%, while updating only 0.1%–0.5% of the parameters compared to full finetuning. Furthermore, our method outperforms the existing methods. Bohan Yang 0013, Yushi Chen 0002, Pedram Ghamisi |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | SpectralGPT: Spectral Remote Sensing Foundation ModelabstractThe foundation model has recently garnered significant attention due to its potential to revolutionize the field of visual representation learning in a self-supervised manner. While most foundation models are tailored to effectively process RGB images for various visual tasks, there is a noticeable gap in research focused on spectral data, which offers valuable information for scene understanding, especially in remote sensing (RS) applications. To fill this gap, we created for the first time a universal RS foundation model, named SpectralGPT, which is purpose-built to handle spectral RS images using a novel 3D generative pretrained transformer (GPT). Compared to existing foundation models, SpectralGPT 1) accommodates input images with varying sizes, resolutions, time series, and regions in a progressive training fashion, enabling full utilization of extensive RS Big Data; 2) leverages 3D token generation for spatial-spectral coupling; 3) captures spectrally sequential patterns via multi-target reconstruction; and 4) trains on one million spectral RS images, yielding models with over 600 million parameters. Our evaluation highlights significant performance improvements with pretrained SpectralGPT models, signifying substantial potential in advancing spectral RS Big Data applications within the field of geoscience across four downstream tasks: single/multi-label scene classification, semantic segmentation, and change detection. Danfeng Hong, Bing Zhang 0001, Chenyu Li 0002, Jing Yao 0002, Naoto Yokoya, Hao Li 0019, Pedram Ghamisi, Xiuping Jia, Antonio Plaza, Paolo Gamba, Jón Atli Benediktsson, Jocelyn Chanussot |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2024 | Tinto: Multisensor Benchmark for 3-D Hyperspectral Point Cloud Segmentation in the GeosciencesabstractThe increasing use of deep learning techniques has reduced interpretation time and, ideally, reduced interpreter bias by automatically deriving geological maps from digital outcrop models. However, accurate validation of these automated mapping approaches is a significant challenge due to the subjective nature of geological mapping and the difficulty in collecting quantitative validation data. Additionally, many state-of-the-art deep learning methods are limited to 2D image data, which is insufficient for 3D digital outcrops, such as hyperclouds. To address these challenges, we present Tinto, a multi-sensor benchmark digital outcrop dataset designed to facilitate the development and validation of deep learning approaches for geological mapping, especially for non-structured 3D data like point clouds. Tinto comprises two complementary sets: 1) a real digital outcrop model from Corta Atalaya (Spain), with spectral attributes and ground-truth data, and 2) a synthetic twin that uses latent features in the original datasets to reconstruct realistic spectral data (including sensor noise and processing artifacts) from the ground-truth. The point cloud is dense and contains 3,242,964 labeled points. We used these datasets to explore the abilities of different deep learning approaches for automated geological mapping. By making Tinto publicly available, we hope to foster the development and adaptation of new deep learning tools for 3D applications in Earth sciences. The dataset can be accessed through this link: https://doi.org/10.14278/rodare.2256. Ahmed J. Afifi, Samuel T. Thiele, Aldino Rizaldy, Sandra Lorenz, Pedram Ghamisi, Raimon Tolosana-Delgado, Moritz Kirsch, Richard Gloaguen, Michael Heizmann |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Dsfer-Net: A Deep Supervision and Feature Retrieval Network for Bitemporal Change Detection Using Modern Hopfield NetworksabstractChange detection, an essential application for high-resolution remote sensing (RS) images, aims to monitor and analyze changes in the land surface over time. Due to the rapid increase in the quantity of high-resolution RS data and the complexity of texture features, several quantitative deep learning-based methods have been proposed. These methods outperform traditional change detection (CD) methods by extracting deep features and combining spatial–temporal information. However, reasonable explanations for how deep features improve detection performance are still lacking. In our investigations, we found that modern Hopfield network (MHN) layers significantly enhance semantic understanding. In this article, we propose a deep supervision and feature retrieval network (Dsfer-Net) for bitemporal CD. Specifically, the highly representative deep features of bitemporal images are jointly extracted through a fully convolutional Siamese network. Based on the sequential geographical information of the bitemporal images, we designed a feature retrieval module to extract difference features and leverage discriminative information in a deeply supervised manner. In addition, we observed that the deeply supervised feature retrieval (DSFR) module provides explainable evidence of the semantic understanding of the proposed network in its deep layers. Finally, our end-to-end network establishes a novel framework by aggregating retrieved features and feature pairs from different layers. Experiments conducted on three public datasets (LEVIR-CD, WHU-CD, and CDD) confirm the superiority of the proposed Dsfer-Net over other state-of-the-art methods. Compared to the best-performing DSAMNet, Dsfer-Net demonstrates significant improvements, with$F1$scores increasing by 4.7%, 5.9%, and 2.3%. Furthermore, compared to our previous FrNet, Dsfer-Net also achieves noteworthy enhancements, with$F1$scores increasing by 2.0%, 1.4%, and 4.5% on three datasets. The code will be available online (https://github.com/ShizhenChang/Dsfer-Net). Shizhen Chang, Michael Kopp 0001, Pedram Ghamisi, Bo Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Diffusion Models Meet Remote Sensing: Principles, Methods, and PerspectivesabstractAs a newly emerging advance in deep generative models, diffusion models have achieved state-of-the-art results in many fields, including computer vision, natural language processing, and molecule design. The remote sensing (RS) community has also noticed the powerful ability of diffusion models and quickly applied them to a variety of tasks for image processing. Given the rapid increase in research on diffusion models in the field of RS, it is necessary to conduct a comprehensive review of existing diffusion model-based RS papers, to help researchers recognize the potential of diffusion models and provide some directions for further exploration. Specifically, this article first introduces the theoretical background of diffusion models, and then systematically reviews the applications of diffusion models in RS, including image generation, enhancement, and interpretation. Finally, the limitations of existing RS diffusion models and worthy research directions for further exploration are discussed and summarized. Yidan Liu, Jun Yue 0004, Shaobo Xia, Pedram Ghamisi, Weiying Xie, Leyuan Fang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Cross Hyperspectral and LiDAR Attention Transformer: An Extended Self-Attention for Land Use and Land Cover ClassificationabstractThe successes of attention-driven deep models like the Vision Transformer (ViT) have sparked interest in cross-domain exploration. However, current transformer-based techniques in remote sensing primarily focus on single-modal data, limiting their potential to exploit the growing array of multimodal Earth observation data fully. Enhancing these models for multimodal integration is crucial for comprehensive remote sensing applications. To achieve this, we extend the traditional self-attention mechanism by introducing Cross Hyperspectral and LiDAR (Cross-HL) attention. We present a novel multimodal deep learning framework that effectively fuses remote sensing (RS) data, intending to improve land use and land cover (LULC) recognition. To enhance the accurate exchange of information across different modalities, we fuse their patch projections using the Cross-HL self-attention module. In this process, LiDAR patch tokens serve as queries (Q), while keys (K) and values (V) are derived from HS patch tokens. To demonstrate the superiority of Cross-HL in the proposed multimodal deep learning framework, we conducted extensive experiments on three multimodal RS benchmark datasets: Houston, Trento, and MUUFL. These datasets contain hyperspectral and light detection and ranging (LiDAR) data. The source code for Cross-HL will be made available publicly at https://github.com/AtriSukul1508/Cross-HL. Swalpa Kumar Roy, Atri Sukul, Ali Jamali, Juan Mario Haut, Pedram Ghamisi |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | IG-GAN: Interactive Guided Generative Adversarial Networks for Multimodal Image FusionabstractMultimodal image fusion has recently garnered increasing interest in the field of remote sensing. By leveraging the complementary information in different modalities, the fused results may be more favorable in characterizing objects of interest, thereby increasing the chance of a more comprehensive and accurate perception of the scene. Unfortunately, most existing fusion methods tend to extract modality-specific features independently without considering intermodal alignment and complementarity, leading to a suboptimal fusion process. To address this issue, we propose a novel interactive generative adversarial network (IG-GAN), for the task of multimodal image fusion. IG-GAN comprises guided dual streams tailored for enhanced learning of details and content, as well as cross-modal consistency. Specifically, a details-guided interactive running-in module (GIR1) and a content-guided interactive running-in module (GIR2) are developed, with the stronger modality serving as guidance for detail richness or content integrity, and the weaker one assisting. To fully integrate multigranularity features from dual-modality, a hierarchical fusion and reconstruction branch is established. Specifically, a shallow interactive fusion (SIF) module followed by a multilevel interactive fusion (MIF) module is designed to aggregate multilevel local and long-range features. Concerning feature decoding and fused image generation, a high-level interactive fusion and reconstruction module (HRM) is further developed. In addition, to empower the fusion network to generate fused images with complete content, sharp edges, and high fidelity without supervision, a loss function facilitating the mutual game between the generator and two discriminators is also formulated. Comparative experiments with 14 state-of-the-art methods are conducted on three datasets. Qualitative and quantitative results indicate that IG-GAN exhibits obvious superiority in terms of both visual effect and quantitative metrics. Moreover, experiments on two RGB-IR object detection datasets are also conducted, which demonstrate that IG-GAN can enhance the accuracy of object detection by integrating complementary information from different modalities. The code will be available athttps://github.com/flower6top. Chenhong Sui, Guobin Yang, Danfeng Hong, Jing Yao 0002, Peter M. Atkinson, Pedram Ghamisi |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | RSAdapter: Adapting Multimodal Models for Remote Sensing Visual Question AnsweringabstractIn recent years, with the rapid advancement of transformer models, transformer-based multimodal architectures have found wide application in various downstream tasks, including, but not limited to, image captioning, visual question answering (VQA), and image–text generation. However, contemporary approaches to remote sensing (RS) VQA often involve resource-intensive techniques, such as full fine-tuning of large models or the extraction of image–text features from pretrained multimodal models, followed by modality fusion using decoders. These approaches demand significant computational resources and time, and a considerable number of trainable parameters are introduced. To address these challenges, we introduce a novel method known as RSAdapter, which prioritizes runtime and parameter efficiency. RSAdapter comprises two key components: the parallel adapter and an additional linear transformation layer inserted after each fully connected (FC) layer within the adapter. This approach not only improves adaptation to pretrained multimodal models but also allows the parameters of the linear transformation layer to be integrated into the preceding FC layers during inference, reducing inference costs. To demonstrate the effectiveness of RSAdapter, we conduct an extensive series of experiments using three distinct RS-VQA datasets and achieve state-of-the-art results on all three datasets. The code for RSAdapter is available online athttps://github.com/Y-D-Wang/RSAdapter. Yuduo Wang, Pedram Ghamisi |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | MaskCD: A Remote Sensing Change Detection Network Based on Mask ClassificationabstractChange detection (CD) from remote sensing (RS) images using deep learning has been widely investigated in the literature. It is typically regarded as a pixelwise labeling task that aims to classify each pixel as changed or unchanged. Although per-pixel classification networks in encoder-decoder structures have shown dominance, they still suffer from imprecise boundaries and incomplete object delineation at various scenes. For high-resolution RS images, partly or totally changed objects are more worthy of attention rather than a single pixel. Therefore, we revisit the CD task from the mask prediction and classification perspective and propose mask classification-based CD (MaskCD) to detect changed areas by adaptively generating categorized masks from input image pairs. Specifically, it utilizes a cross-level change representation perceiver (CLCRP) to learn multiscale change-aware representations and capture spatiotemporal relations from encoded features by exploiting deformable multihead self-attention (DeformMHSA). Subsequently, a masked cross-attention-based detection transformers (MCA-DETRs) decoder is developed to accurately locate and identify changed objects based on masked cross-attention and self-attention (SA) mechanisms. It reconstructs the desired changed objects by decoding the pixelwise representations into learnable mask proposals and making final predictions from these candidates. Experimental results on five benchmark datasets demonstrate the proposed approach outperforms other state-of-the-art models. Codes and pretrained models are available online at:https://github.com/EricYu97/MaskCD. Weikang Yu, Samiran Das, Xiao Xiang Zhu 0001, Pedram Ghamisi |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | MineNetCD: A Benchmark for Global Mining Change Detection on Remote Sensing ImageryabstractMonitoring land changes triggered by mining activities is crucial for industrial control, environmental management, and regulatory compliance, yet it poses significant challenges due to the vast and often remote locations of mining sites. Remote sensing technologies have increasingly become indispensable to detect and analyze these changes over time. We thus introduce MineNetCD, a comprehensive benchmark designed for global mining change detection using remote sensing imagery. The benchmark comprises three key contributions. First, we establish a global mining change detection dataset featuring more than 70k paired patches of bitemporal high-resolution remote sensing images and pixel-level annotations from 100 mining sites worldwide. Second, we develop a novel baseline model based on a change-aware fast Fourier transform (ChangeFFT) module, which enhances various backbones by leveraging essential spectrum components within features in the frequency domain and capturing the channelwise correlation of bitemporal feature differences to learn change-aware representations. Third, we construct a unified change detection (UCD) framework that currently integrates 20 change detection methods. This framework is designed for streamlined and efficient processing, using the cloud platform hosted by HuggingFace. Extensive experiments have been conducted to demonstrate the superiority of the proposed baseline model compared with 19 state-of-the-art change detection approaches. Empirical studies on modularized backbones comprehensively confirm the efficacy of different representation learners on change detection. This benchmark represents significant advancements in the field of remote sensing and change detection, providing a robust resource for future research and applications in global mining monitoring. Dataset and Codes are available via the link. Weikang Yu, Richard Gloaguen, Xiao Xiang Zhu 0001, Pedram Ghamisi |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Prototypical Unknown-Aware Multiview Consistency Learning for Open-Set Cross-Domain Remote Sensing Image ClassificationabstractDeveloping a cross-domain classification model for remote sensing images has drawn significant attention in the literature. By leveraging the open-set unsupervised domain adaptation (UDA) technique, the generalization performance of deep learning models has been improved with the capability to recognize unknown categories. However, it remains challenging to explore distribution patterns in the target domain using uncertain category-wise supervision from unlabeled datasets while reducing negative transfer caused by unknown samples. To develop a robust open-set UDA framework, this article presents prototypical unknown-aware multiview consistency learning (PUMCL) designed for remote sensing scene classification across heterogeneous domains. Specifically, it employs a consistency learning scheme with multiview and multilevel perturbations to improve feature learning from unlabeled target samples. An entropy separation strategy is utilized to facilitate open-set detection and recognition during adaptation, enabling unknown-aware feature alignment. Furthermore, the introduction of prototypical constraints optimizes pseudo-label generation through online denoising and promotes a compact category-wise feature subspace for improved class separation across domains. Experiments conducted on six cross-domain scenarios using AID, NWPU, and UCMD datasets demonstrate the method’s superior performance compared to nine state-of-the-art approaches, achieving a gain of 4.5% to 21.2% in mIoU. More importantly, it shows promising class separability with clear boundaries between different classes and compact clustering of unknown samples in the feature space. The source code will be available athttps://github.com/zxk688. Wanjing Wu, Mi Zhang 0004, Weikang Yu, Pedram Ghamisi |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Context Enhancing Representation for Semantic Segmentation in Remote Sensing ImagesabstractAs the foundation of image interpretation, semantic segmentation is an active topic in the field of remote sensing. Facing the complex combination of multiscale objects existing in remote sensing images (RSIs), the exploration and modeling of contextual information have become the key to accurately identifying the objects at different scales. Although several methods have been proposed in the past decade, insufficient context modeling of global or local information, which easily results in the fragmentation of large-scale objects, the ignorance of small-scale objects, and blurred boundaries. To address the above issues, we propose a contextual representation enhancement network (CRENet) to strengthen the global context (GC) and local context (LC) modeling in high-level features. The core components of the CRENet are the local feature alignment enhancement module (LFAEM) and the superpixel affinity loss (SAL). The LFAEM aligns and enhances the LC in low-level features by constructing contextual contrast through multilayer cascaded deformable convolution and is then supplemented with high-level features to refine the segmentation map. The SAL assists the network to accurately capture the GC by supervising semantic information and relationship learned from superpixels. The proposed method is plug-and-play and can be embedded in any FCN-based network. Experiments on two popular RSI datasets demonstrate the effectiveness of our proposed network with competitive performance in qualitative and quantitative aspects. Leyuan Fang, Peng Zhou 0036, Xinxin Liu 0002, Pedram Ghamisi, Si-Wei Chen 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Hyperspectral Domain Adaptation for the Detection of Material Types in Recycling Streams at the Example of ElectrolyzersabstractHyperspectral datasets obtained from a specific sensor can experience changes in their characteristics due to environmental noise and variations in illumination. Consequently, a segmentation model trained on one dataset may struggle to accurately predict labels and detect objects on a different dataset due to discrepancies between the two domains. To overcome this challenge, domain adaptation techniques can be employed. In the paper, we study hyperspectral domain adaptation for adapting the target domain to align with the source domain in detecting the material type of mm-scale particles from shredded electrolyzers on a conveyor belt for recycling applications. This is necessary due to the non-uniform distribution of particles, variations in material types, and changes in the imaging environment. The results show improvements compared to a pre-trained model using a 2D convolutional neural network. Behnood Rasti, Aayush Jain, Margret C. Fuchs, Pedram Ghamisi, Richard Gloaguen |
IGARSS | 4 |
| 2023 | Unpaired remote sensing image super-resolution with content-preserving weak supervision neural network
Jie Wu 0035, Runmin Cong, Leyuan Fang, Chunle Guo, Bob Zhang 0001, Pedram Ghamisi |
Sci. China Inf. Sci. | 6 |
| 2023 | Guest Editorial: Spectral imaging powered computer visionabstractThe increasing accessibility and affordability of spectral imaging technology have revolutionised computer vision, allowing for data capture across various wavelengths beyond the visual spectrum.This advancement has greatly enhanced the capabilities of computers and AI systems in observing, understanding, and interacting with the world.Consequently, new datasets in various modalities, such as infrared, ultraviolet, fluorescent, multispectral, and hyperspectral, have been constructed, presenting fresh opportunities for computer vision research and applications.Although significant progress has been made in processing, learning, and utilising data obtained through spectral imaging technology, several challenges persist in the field of computer vision.These challenges include the presence of low-quality images, sparse input, high-dimensional data, expensive data labelling processes, and a lack of methods to effectively analyse and utilise data considering their unique properties.Many mid-level and high-level computer vision tasks, such as object segmentation, detection and recognition, image retrieval and classification, and video tracking and understanding, still have not leveraged the advantages offered by spectral information.Additionally, the problem of effectively and efficiently fusing data in different modalities to create robust vision systems remains unresolved.Therefore, there is a pressing need for novel computer vision methods and applications to advance this research area.This special issue aims to provide a venue for researchers to present innovative computer vision methods driven by the spectral imaging technology. Jun Zhou 0001, Fengchao Xiong, Naoto Yokoya, Pedram Ghamisi |
IET Comput. Vis. | 5 |
| 2023 | Transformer-based contrastive prototypical clustering for multimodal remote sensing data
Yaoming Cai, Zijia Zhang 0001, Pedram Ghamisi, Behnood Rasti, Xiaobo Liu 0001, Zhihua Cai |
Inf. Sci. | 3 |
| 2023 | Local Window Attention Transformer for Polarimetric SAR Image ClassificationabstractConvolutional neural networks (CNNs) have recently found great attention in image classification since deep CNNs have exhibited excellent performance in computer vision. Owing to their immense success, of late, scientists are exploring the functionality of transformers in Earth observation applications. Nevertheless, the primary issue with transformers is that they demand significantly more training data than CNN classifiers. Thus, the use of these transformers in remote sensing is considered challenging, notably in utilizing polarimetric synthetic aperture radar (PolSAR) data, due to the insufficient number of existing labeled data. In this letter, we develop and propose a vision transformer (ViT)-based framework that utilizes 3-D and 2-D CNNs as feature extractors and, in addition, local window attention (LWA) for the effective classification of PolSAR data. Extensive experimental results demonstrated that the developed modelPolSARFormerobtained better classification accuracy than the state-of-the-art vision Swin Transformer and FNet algorithms. ThePolSARFormeroutperformed the Swin Transformer and FNet by the margin of 5.86% and 17.63%, in terms of average accuracy (AA) in the San Francisco data benchmark. Moreover, the results over the Flevoland dataset illustrated that thePolSARFormerexceeds several other algorithms, including the ResNet (97.49%), Swin Transformer (96.54%), FNet (95.28%), 2-D CNN (94.57%), and AlexNet (91.83%), with a kappa index (KI) of 99.30%. The code will be made available publicly athttps://github.com/aj1365/PolSARFormer. Ali Jamali, Swalpa Kumar Roy, Avik Bhattacharya, Pedram Ghamisi |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2023 | Backdoor Attacks for Remote Sensing Data With Wavelet TransformabstractRecent years have witnessed the great success of deep learning algorithms in the geoscience and remote sensing realm. Nevertheless, the security and robustness of deep learning models deserve special attention when addressing safety-critical remote sensing tasks. In this paper, we provide a systematic analysis of backdoor attacks for remote sensing data, where both scene classification and semantic segmentation tasks are considered. While most of the existing backdoor attack algorithms rely on visible triggers like squared patches with well-designed patterns, we propose a novel wavelet transform-based attack (WABA) method, which can achieve invisible attacks by injecting the trigger image into the poisoned image in the low-frequency domain. In this way, the high-frequency information in the trigger image can be filtered out in the attack, resulting in stealthy data poisoning. Despite its simplicity, the proposed method can significantly cheat the current state-of-the-art deep learning models with a high attack success rate. We further analyze how different trigger images and the hyper-parameters in the wavelet transform would influence the performance of the proposed method. Extensive experiments on four benchmark remote sensing datasets demonstrate the effectiveness of the proposed method for both scene classification and semantic segmentation tasks and thus highlight the importance of designing advanced backdoor defense algorithms to address this threat in remote sensing scenarios. The code will be available online at https://github.com/ndraeger/waba. Nikolaus Dräger, Yonghao Xu, Pedram Ghamisi |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Hyperspectral Remote Sensing Benchmark Database for Oil Spill Detection With an Isolation Forest-Guided Unsupervised DetectorabstractOil spill detection has attracted increasing attention in recent years since marine oil spill accidents severely affect environments, natural resources, and the lives of coastal inhabitants. Hyperspectral remote sensing images provide rich spectral information which is beneficial for the monitoring of oil spills in complex ocean scenarios. However, most of the existing approaches are based on supervised and semi-supervised frameworks to detect oil spills from hyperspectral images (HSIs), which require a massive amount of effort to annotate a certain number of high-quality training sets. In this study, we make the first attempt to develop an unsupervised oil spill detection method based on isolation forest for HSIs. First, a Gaussian statistical model is designed to remove the bands corrupted by severe noise. Then, kernel principal component analysis (KPCA) is employed to reduce the high dimensionality of the HSIs. Next, the probability of each pixel belonging to one of the classes of seawater and oil spills is estimated with the isolation forest, and a set of pseudo-labeled training samples is automatically produced using the clustering algorithm on the detected probability. Finally, an initial detection map can be obtained by performing the support vector machine (SVM) on the dimension-reduced data, and the initial detection result is further optimized with the extended random walker (ERW) model so as to improve the detection accuracy of oil spills. Experiments on hyperspectral oil spill data (HOSD) created by ourselves demonstrate that the proposed method obtains superior detection performance with respect to other state-of-the-art detection approaches. We will make HOSD and our developed library for oil spill detection publicly available at https://github.com/PuhongDuan/HOSD to further promote this research topic. Puhong Duan, Xudong Kang, Pedram Ghamisi, Shutao Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | PCLDet: Prototypical Contrastive Learning for Fine-Grained Object Detection in Remote Sensing ImagesabstractThe capacity of satellites to supply high-resolution imaging has promoted the fine-grained object detection task in remote sensing images. However, this type of object detection is challenging due to low interclass feature differences in objects. To address this issue, we propose a prototypical contrastive learning-based detector (PCLDet) for fine-grained object detection in remote sensing images. The PCLDet first introduces the prototype to learn the fine-grained objects’ features, and then adopts contrastive learning to compare the target and the learned features, thus improving the differentiability of the fine-grained object. Specifically, we first introduce the prototype, which represents the feature centers of each class, and then construct a prototype bank to store the feature prototypes of each class. Then, we introduce contrastive learning to extract the discriminative features by maximizing the interclass distance and minimizing the intraclass distance. Furthermore, we propose the ProtoCL loss as a part of the model optimization, which enables more representative prototypes to be learned. Finally, to address the long-tail problem in the remote sensing fine-grained object detection dataset, we propose a new proposal sampler, the class-balanced sampler (CBS) that can sample each class equally. Extensive experiments demonstrate that our method can achieve state-of-the-art performance on a commonly used aerial fine-grained object dataset (Fair1M) and aerial fine-grained ship dataset (OFSD) while maintaining high efficiency. The code will be available at https://github.com/G-Naughty/PCLDet. Lihan Ouyang, Guangmiao Guo, Leyuan Fang, Pedram Ghamisi, Jun Yue 0004 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Learning Crop-Type Mapping From Regional Label Proportions in Large-Scale SAR and Optical ImageryabstractThe application of deep learning algorithms to Earth observation (EO) in recent years has enabled substantial progress in fields that rely on remotely sensed data. However, given the data scale in EO, creating large datasets with pixel-level annotations by experts is expensive and highly time-consuming. In this context, priors are seen as an attractive way to alleviate the burden of manual labeling when training deep learning methods for EO. For some applications, those priors are readily available. Motivated by the great success of contrastive-learning methods for self-supervised feature representation learning in many computer-vision tasks, this study proposes an online deep clustering method using crop label proportions as priors to learn a sample-level classifier based on government crop-proportion data for a whole agricultural region. We evaluate the method using two large datasets from two different agricultural regions in Brazil. Extensive experiments demonstrate that the method is robust to different data types (synthetic-aperture radar and optical images), reporting higher accuracy values considering the major crop types in the target regions. Thus, it can alleviate the burden of large-scale image annotation in EO applications. Laura Elena Cue La Rosa, Dário A. B. Oliveira, Pedram Ghamisi |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Changes to Captions: An Attentive Network for Remote Sensing Change CaptioningabstractIn recent years, advanced research has focused on the direct learning and analysis of remote-sensing images using natural language processing (NLP) techniques. The ability to accurately describe changes occurring in multi-temporal remote sensing images is becoming increasingly important for geospatial understanding and land planning. Unlike natural image change captioning tasks, remote sensing change captioning aims to capture the most significant changes, irrespective of various influential factors such as illumination, seasonal effects, and complex land covers. In this study, we highlight the significance of accurately describing changes in remote sensing images and present a comparison of the change captioning task for natural and synthetic images and remote sensing images. To address the challenge of generating accurate captions, we propose an attentive changes-to-captions network, called Chg2Cap for short, for bi-temporal remote sensing images. The network comprises three main components: 1) a Siamese CNN-based feature extractor to collect high-level representations for each image pair; 2) an attentive encoder that includes a hierarchical self-attention block to locate change-related features and a residual block to generate the image embedding; and 3) a transformer-based caption generator to decode the relationship between the image embedding and the word embedding into a description. The proposed Chg2Cap network is evaluated on two representative remote sensing datasets, and a comprehensive experimental analysis is provided. The code and pre-trained models will be available online at https://github.com/ShizhenChang/Chg2Cap. Shizhen Chang, Pedram Ghamisi |
IEEE Trans. Image Process. | 2 |
| 2023 | Txt2Img-MHN: Remote Sensing Image Generation From Text Using Modern Hopfield NetworksabstractThe synthesis of high-resolution remote sensing images based on text descriptions has great potential in many practical application scenarios. Although deep neural networks have achieved great success in many important remote sensing tasks, generating realistic remote sensing images from text descriptions is still very difficult. To address this challenge, we propose a novel text-to-image modern Hopfield network (Txt2Img-MHN). The main idea of Txt2Img-MHN is to conduct hierarchical prototype learning on both text and image embeddings with modern Hopfield layers. Instead of directly learning concrete but highly diverse text-image joint feature representations for different semantics, Txt2Img-MHN aims to learn the most representative prototypes from text-image embeddings, achieving a coarse-to-fine learning strategy. These learned prototypes can then be utilized to represent more complex semantics in the text-to-image generation task. To better evaluate the realism and semantic consistency of the generated images, we further conduct zero-shot classification on real remote sensing data using the classification model trained on synthesized images. Despite its simplicity, we find that the overall accuracy in the zero-shot classification may serve as a good metric to evaluate the ability to generate an image from text. Extensive experiments on the benchmark remote sensing text-image dataset demonstrate that the proposed Txt2Img-MHN can generate more realistic remote sensing images than existing methods. Code and pre-trained models are available online (https://github.com/YonghaoXu/Txt2Img-MHN). Yonghao Xu, Weikang Yu, Pedram Ghamisi, Michael Kopp 0001, Sepp Hochreiter |
IEEE Trans. Image Process. | 3 |
| 2023 | Fully Linear Graph Convolutional Networks for Semi-Supervised and Unsupervised ClassificationabstractThis article presents FLGC, a simple yet effective fully linear graph convolutional network for semi-supervised and unsupervised learning. Instead of using gradient descent, we train FLGC based on computing a global optimal closed-form solution with a decoupled procedure, resulting in a generalized linear framework and making it easier to implement, train, and apply. We show that (1) FLGC is powerful to deal with both graph-structured data and regular data, (2) training graph convolutional models with closed-form solutions improve computational efficiency without degrading performance, and (3) FLGC acts as a natural generalization of classic linear models in the non-Euclidean domain (e.g., ridge regression and subspace clustering). Furthermore, we implement a semi-supervised FLGC and an unsupervised FLGC by introducing an initial residual strategy, enabling FLGC to aggregate long-range neighborhoods and alleviate over-smoothing. We compare our semi-supervised and unsupervised FLGCs against many state-of-the-art methods on a variety of classification and clustering benchmarks, demonstrating that the proposed FLGC models consistently outperform previous methods in terms of accuracy, robustness, and learning efficiency. The core code of our FLGC is released at https://github.com/AngryCai/FLGC . Yaoming Cai, Zijia Zhang 0001, Pedram Ghamisi, Zhihua Cai, Xiaobo Liu 0001, Yao Ding 0010 |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2022 | EOD: The IEEE GRSS Earth Observation DatabaseabstractIn the era of deep learning, annotated datasets have become a crucial asset to the remote sensing community. In the last decade, a plethora of different datasets was published, each designed for a specific data type and with a specific task or application in mind. In the jungle of remote sensing datasets, it can be hard to keep track of what is available already. With this paper, we introduce EOD - the IEEE GRSS Earth Observation Database (EOD) - an interactive online platform for cataloguing different types of datasets leveraging remote sensing imagery. Michael Schmitt 0003, Pedram Ghamisi, Naoto Yokoya, Ronny Hänsch |
IGARSS | 2 |
| 2022 | Unsupervised Deep Hyperspectral Inpainting Using a New Mixing ModelabstractIn this paper, we propose a deep learning-based hyperspectral inpainting (DeepHyIn). The proposed approach is unsupervised since it only utilizes the observed image for training the network. First, we propose a novel model for hyperspectral inpainting in which the degraded hyperspectral image is a linear mixture of endmembers and degraded abundances. The proposed model is subjected to abundance sum to one and nonnegativity constraints. We further assume that the endmembers are known. Then, we propose an optimization problem to estimate the unknown abundance using an image prior. Inspired by deep image prior, we shift the optimization problem to optimize the parameters of a deep network. The proposed method uses a deep convolutional encoder-decoder architecture as a backbone. Finally, we apply the DeepHyIn to the Samson dataset and evaluate the results. DeepHyIn demonstrates considerable quantitative and qualitative improvements compared with the state-of-the-art. DeepHyIn was implemented in Python (3.9) using PyTorch as the platform for the deep network and is available online: https://github.com/BehnoodRasti/DeepHyIn. Behnood Rasti, Pedram Ghamisi, Richard Gloaguen |
IGARSS | 2 |
| 2022 | Hyperspectral Clustering Using Atrous Spatial-Spectral Convolutional NetworkabstractHyperspectral imaging is an important technology in the field of geosciences and remote sensing.However, the highdimensional nature of hyperspectral images (HSIs) together with the limited availability of training/labeled samples challenge an efficient processing of HSIs.To alleviate these challenges, we propose a deep multi-resolution clustering network (DMC-Net) to analyze HSIs.DMC-Net, without requiring training/labeled samples for the training process, captures the non-linear intrinsic relation within data points in an HSI and analyzes the image at various resolutions by applying atrous convolutions.Furthermore, DMC-Net preserves the spectral information by directly incorporating extracted features from the original HSI into the reconstruction phase.In terms of clustering accuracy, experimental results on two real HSIs demonstrate the superior performance of DMC-Net compared to the state-of-the-art deep learning-based clustering approaches. Kasra Rafiezadeh Shahi, Pedram Ghamisi, Behnood Rasti, Paul Scheunders, Richard Gloaguen |
IGARSS | 2 |
| 2022 | Edge-Preserving Filtering-Based Dehazing for Remote Sensing ImagesabstractHaze in remote sensing images severely degrades image visibility, making it hard to identify different land covers. This work proposes an edge-preserving filtering-based image dehazing method for remote sensing images, mainly consisting of following several steps. First, the original image contaminated with haze is decomposed by a multiscale guided filtering into base layers that contain haze components and detail layers that reflect the spatial details of input. Then, an optimized atmospheric scattering model is performed on the base layers to eliminate haze. Next, adaptive nonlinear mapping is exploited to enhance image details. Finally, the resulting image is reconstructed by combining the dehazed base layers and the refined detail layers. Experiments on several images demonstrate that the proposed method has an outstanding haze removal capability and yields better dehazing performance compared to other dehazing approaches in terms of objective indexes and subjective results. Puhong Duan, Pedram Ghamisi |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Complementary Learning-Based Scene Classification of Remote Sensing Images With Noisy LabelsabstractRecently, many deep convolutional neural network (DCNN)-based methods have been proposed for remote sensing (RS) image scene classification (SC). In general, DCNNs obtain good generalization capabilities under the condition of correct labels. Unfortunately, the given samples are sometimes mislabeled. In this letter, the classification of RS images with noisy labels is investigated. First, complementary learning (CL), which learns from complementary labels rather than the original labels, is introduced for RS image classification with noisy labels. CL can decrease the probability of learning from incorrect information, and therefore, it is robust to noisy labels. Then, soft CL, which randomly disturbs the complementary labels of the training samples, is proposed to prevent the overfitting issue in training a DCNN. Moreover, an RS image scene classification framework combining ordinary learning (OL) and CL (RS-COCL) is proposed, which uses CL to obtain a good model and OL to fine-tune the deep model. Additionally, noisy labels filtering is used in RS-COCL (RS-COCL-NLF) to detected and corrected noisy samples. At last, soft CL is used in RS-COCL-NLF to obtain better classification performance. The proposed methods are tested on two widely used datasets (i.e., Northwestern Polytechnical University (NWPU)-RESISC45 and PatternNet) and the obtained results show that the proposed methods provide competitive classification accuracy compared to the state-of-the-art methods. Qingyun Li, Yushi Chen 0002, Pedram Ghamisi |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | OptFus: Optical Sensor Fusion for the Classification of Multisource Data: Application to Mineralogical MappingabstractWe propose a new fusion-based classification technique for optical multisource remote-sensing images called OptFus. OptFus is developed to merge and process optical imagery having different spatial and spectral resolutions. The spatial features are extracted using morphological filters from the RGB data containing high spatial resolution. A feature fusion technique is developed to combine all the sensor data in a subspace using a common set of representative features. Finally, the fused features are classified using a support vector machine to ensure a robust supervised spectral classification. The proposed method is designed to allocate varying weights to the data from various imaging sensors in the fusion process. OptFus is applied to two multisource optical datasets captured from geological drill-core samples. The classification accuracy demonstrates considerable improvements compared to the state-of-the-art. A MATLAB implementation of OptFus is available online:https://github.com/BehnoodRasti/OptFus. Behnood Rasti, Pedram Ghamisi, Richard Gloaguen |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Superpixel Contracted Neighborhood Contrastive Subspace Clustering Network for Hyperspectral ImagesabstractDeep subspace clustering has achieved remarkable performances in the unsupervised classification of hyperspectral images. However, previous models based on pixel-level self-expressiveness of data suffer from the exponential growth of computational complexity and access memory requirements with increasing number of samples, thus leading to poor applicability to large hyperspectral images. This paper presents a Neighborhood Contrastive Subspace Clustering network (NCSC), a scalable and robust deep subspace clustering approach, for unsupervised classification of large hyperspectral images. Instead of using a conventional autoencoder, we devise a novel superpixel pooling autoencoder to learn the superpixel-level latent representation and subspace, allowing a contracted self-expressive layer. To encourage a robust subspace representation, we propose a novel neighborhood contrastive regularization to maximize the agreement between positive samples in subspace. We jointly train the resulting model in an end-to-end fashion by optimizing an adaptively weighted multi-task loss. Extensive experiments on three hyperspectral benchmarks demonstrate the effectiveness of the proposed approach and its substantial advancement of state-of-the-art approaches. Yaoming Cai, Zijia Zhang 0001, Pedram Ghamisi, Yao Ding 0010, Xiaobo Liu 0001, Zhihua Cai, Richard Gloaguen |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Nonnegative-Constrained Joint Collaborative Representation With Union Dictionary for Hyperspectral Anomaly DetectionabstractRecently, many collaborative representation-based (CR) algorithms have been proposed for hyperspectral anomaly detection. CR-based detectors approximate the image by a linear combination of background dictionaries and the coefficient matrix, and derive the detection map by utilizing recovery residuals. However, these CR-based detectors are often established on the premise of precise background features and strong image representation, which are very difficult to obtain. In addition, pursuing the coefficient matrix reinforced by the generall2-min is very time consuming. To address these issues, a nonnegative-constrained joint collaborative representation model is proposed in this paper for the hyperspectral anomaly detection task. To extract reliable samples, a union dictionary consisting of background and anomaly sub-dictionaries is designed, where the background sub-dictionary is obtained at the superpixel level and the anomaly sub-dictionary is extracted by the pre-detection process. And the coefficient matrix is jointly optimized by the Frobenius norm regularization with a nonnegative constraint and a sum-to-one constraint. After the optimization process, the abnormal information is finally derived by calculating the residuals that exclude the assumed background information. To conduct comparable experiments, the proposed nonnegative-constrained joint collaborative representation (NJCR) model and its kernel version (KNJCR) are tested in four HSI datasets and achieve superior results compared with other state-of-the-art detectors. The codes of the proposed method will be available online. Shizhen Chang, Pedram Ghamisi |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Sketched Multiview Subspace Learning for Hyperspectral Anomalous Change DetectionabstractIn recent years, multi-view subspace learning has been garnering increasing attention. It aims to capture the inner relationships of the data that are collected from multiple sources by learning a unified representation. In this way, comprehensive information from multiple views is shared and preserved for the generalization processes. As a special branch of temporal series hyperspectral image (HSI) processing, the anomalous change detection task focuses on detecting very small changes among different temporal images. However, when the volume of datasets is very large or the classes are relatively comprehensive, existing methods may fail to find those changes between the scenes, and end up with terrible detection results. In this paper, inspired by the sketched representation and multi-view subspace learning, a sketched multi-view subspace learning (SMSL) model is proposed for HSI anomalous change detection. The proposed model preserves major information from the image pairs and improves computational complexity by using a sketched representation matrix. Furthermore, the differences between scenes are extracted by utilizing the specific regularizer of the self-representation matrices. To evaluate the detection effectiveness of the proposed SMSL model, experiments are conducted on a benchmark hyperspectral remote sensing dataset and a natural hyperspectral dataset, and compared with other state-of-the-art approaches. The codes of the proposed method will be made available online1. Shizhen Chang, Michael Kopp 0001, Pedram Ghamisi |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Landslide4Sense: Reference Benchmark Data and Deep Learning Models for Landslide DetectionabstractThis study introducesLandslide4Sense, a reference benchmark for landslide detection from remote sensing. The repository features 3,799 image patches fusing optical layers from Sentinel-2 sensors with the digital elevation model and slope layer derived from ALOS PALSAR. The added topographical information facilitates an accurate detection of landslide borders, which recent researches have shown to be challenging using optical data alone. The extensive data set supports deep learning (DL) studies in landslide detection and the development and validation of methods for the systematic update of landslide inventories. The benchmark data set has been collected at four different times and geographical locations: Iburi (September 2018), Kodagu (August 2018), Gorkha (April 2015), and Taiwan (August 2009). Each image pixel is labelled as belonging to a landslide or not, incorporating various sources and thorough manual annotation. We then evaluate the landslide detection performance of 11 state-of-the-art DL segmentation models: U-Net, ResU-Net, PSPNet, ContextNet, DeepLab-v2, DeepLab-v3+, FCN-8s, LinkNet, FRRN-A, FRRN-B, and SQNet. All models were trained from scratch on patches from one quarter of each study area and tested on independent patches from the other three quarters. Our experiments demonstrate that ResU-Net outperformed the other models for the landslide detection task. We make the multi-source landslide benchmark data (Landslide4Sense) and the tested DL models publicly available at https://www.iarai.ac.at/landslide4sense, establishing an important resource for remote sensing, computer vision, and machine learning communities in studies of image classification in general and applications to landslide detection in particular. Omid Ghorbanzadeh, Yonghao Xu, Pedram Ghamisi, Michael Kopp 0001, David P. Kreil |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Dual Graph Convolutional Network for Hyperspectral Image Classification With Limited Training SamplesabstractDue to powerful feature extraction capability, convolutional neural networks (CNNs) have been widely used for hyperspectral image (HSI) classification. However, because of a large number of parameters that need to be trained, sufficient training samples are usually required for deep CNN-based methods. Unfortunately, limited training samples are a common issue in the remote sensing community. In this study, a dual graph convolutional network (DGCN) is proposed for the supervised classification of HSI with limited training samples. The first GCN fully extracts features existing in and among HSI samples, while the second GCN utilizes label distribution learning, and thus, it potentially reduces the number of required training samples. The two GCNs are integrated through several iterations to decrease interclass distances, which leads to a more accurate classification step. Moreover, a new idea entitled multiscale feature cutout is proposed as a regularization technique for HSI classification (DGCN-M). Different from the regularization methods (e.g., dropout and DropBlock), the proposed multiscale feature cutout could randomly mask out multiscale region sizes in a feature map, which further reduces the overfitting problem and yields consistent improvement. Experimental results on the four popular hyperspectral data sets (i.e., Salinas, Indian Pines, Pavia, and Houston) indicate that the proposed method obtains good classification performance compared to state-of-the-art methods, which shows the potential of GCN for HSI classification. Xin He 0004, Yushi Chen 0002, Pedram Ghamisi |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Modality Translation in Remote Sensing Time SeriesabstractModality translation, which aims to translate images from a source modality to a target one, has attracted a growing interest in the field of remote sensing recently. Compared to translation problems in multimedia applications, modality translation in remote sensing often suffers from inherent ambiguities, i.e., a single input image could correspond to multiple possible outputs, and the results may not be valid in the following image interpretation tasks, such as classification and change detection. To address these issues, we make the attempt to utilizing time-series data to resolve the ambiguities. We propose a novel multimodality image translation framework, which exploits temporal information from two aspects: 1) by introducing a guidance image from given temporally neighboring images in the target modality, we employ a feature mask module and transfer semantic information from temporal images to the output without requiring the use of any semantic labels and 2) while incorporating multiple pairs of images in time series, a temporal constraint is formulated during the learning process in order to guarantee the uniqueness of the prediction result. We also build a multimodal and multitemporal dataset that contains synthetic aperture radar (SAR), visible, and short-wave length infrared band (SWIR) image time series of the same scene to encourage and promote research on modality translation in remote sensing. Experiments are conducted on the dataset for two cross-modality translation tasks (SAR to visible and visible to SWIR). Both qualitative and quantitative results demonstrate the effectiveness and superiority of the proposed model. Danfeng Hong, Jocelyn Chanussot, Baojun Zhao, Pedram Ghamisi |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | NFANet: A Novel Method for Weakly Supervised Water Extraction From High-Resolution Remote-Sensing ImageryabstractThe use of deep learning for water extraction requires precise pixel-level labels. However, it is very difficult to label high-resolution remote-sensing images at the pixel level. Therefore, we study how to utilize point labels to extract water bodies and propose a novel method called the neighbor feature aggregation network (NFANet). Compared with pixel-level labels, point labels are much easier to obtain, but they will lose much information. In this article, we take advantage of the similarity between the adjacent pixels of a local water body, and propose a neighbor sampler to resample remote-sensing images. Then, the sampled images are sent to the network for feature aggregation. In addition, we use an improved recursive training algorithm to further improve the extraction accuracy, making the water boundary more natural. Furthermore, our method utilizes neighboring features instead of global or local features to learn more representative features. The experimental results show that the proposed NFANet method not only outperforms other studied weakly supervised approaches, but also obtains similar results as the state-of-the-art ones. Leyuan Fang, Muxing Li, Bob Zhang 0001, Yi Zhang 0018, Pedram Ghamisi |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | UnDIP: Hyperspectral Unmixing Using Deep Image PriorabstractIn this article, we introduce a deep learning-based technique for the linear hyperspectral unmixing problem. The proposed method contains two main steps. First, the endmembers are extracted using a geometric endmember extraction method, i.e., a simplex volume maximization in the subspace of the data set. Then, the abundances are estimated using a deep image prior. The main motivation of this work is to boost the abundance estimation and make the unmixing problem robust to noise. The proposed deep image prior uses a convolutional neural network to estimate the fractional abundances, relying on the extracted endmembers and the observed hyperspectral data set. The proposed method is evaluated on simulated and three real remote sensing data for a range of SNR values (i.e., from 20 to 50 dB). The results show considerable improvements compared to state-of-the-art methods. The proposed method was implemented in Python (3.8) using PyTorch as the platform for the deep network and is available online:https://github.com/BehnoodRasti/UnDIP. Behnood Rasti, Bikram Koirala, Paul Scheunders, Pedram Ghamisi |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Semisupervised Hyperspectral Image Classification Using a Probabilistic Pseudo-Label Generation FrameworkabstractDeep neural networks (DNNs) show impressive performance for hyperspectral image (HSI) classification when abundant labeled samples are available. The problem is that HSI sample annotation is extremely costly and the budget for this task is usually limited. To reduce the reliance on labeled samples, deep semi-supervised learning (SSL), which jointly learns from labeled and unlabeled samples, has been introduced in the literature. However, learning robust and discriminative features from unlabeled data is a challenging task due to various noise effects and ambiguity of unlabeled samples. As a result, recent advances are constrained, mainly in the pre-training or warm-up stage. In this paper, we propose a deep probabilistic framework to generate reliable pseudo labels to explicitly learn discriminative features from unlabeled samples. The generated pseudo labels of our proposed framework can be fed to various DNNs to improve their generalization capacity. Our proposed framework takes only 10 labeled samples per class to represent the label set as an uncertainty-aware distribution in the latent space. The pseudo labels are then generated for those unlabeled samples whose feature values match the distribution with high probability. By performing extensive experiments on four publicly available datasets, we show that our framework can generate reliable pseudo labels to significantly improve the generalization capacity of several state-of-the-art DNNs. In addition, we introduce a new DNN for HSI classification that demonstrates outstanding accuracy results in comparison with its rivals. Majid Seydgar, Shahryar Rahnamayan, Pedram Ghamisi, Azam Asilian Bidgoli |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Asymmetric Hash Code Learning for Remote Sensing Image RetrievalabstractRemote sensing image retrieval (RSIR), aiming at searching for a set of similar items to a given query image, is a very important task in remote sensing applications. Deep hashing learning as the current mainstream method has achieved satisfactory retrieval performance. On one hand, various deep neural networks are used to extract semantic features of remote sensing images. On the other hand, the hashing techniques are subsequently adopted to map the high-dimensional deep features to the low-dimensional binary codes. This kind of method attempts to learn one hash function for both the query and database samples in a symmetric way. However, with the number of database samples increasing, it is typically time-consuming to generate the hash codes of large-scale database images. In this article, we propose a novel deep hashing method, named asymmetric hash code learning (AHCL), for RSIR. The proposed AHCL generates the hash codes of query and database images in an asymmetric way. In more detail, the hash codes of query images are obtained by binarizing the output of the network, while the hash codes of database images are directly learned by solving the designed objective function. In addition, we combine the semantic information of each image and the similarity information of pairs of images as supervised information to train a deep hashing network, which improves the representation ability of deep features and hash codes. The experimental results on three public datasets demonstrate that the proposed method outperforms symmetric methods in terms of retrieval accuracy and efficiency. The source code is available athttps://github.com/weiweisong415/Demo_AHCL_for_TGRS2022. Zhi Gao 0005, Renwei Dian, Pedram Ghamisi, Yongjun Zhang 0002, Jón Atli Benediktsson |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Transferring CNN With Adaptive Learning for Remote Sensing Scene ClassificationabstractAccurate classification of remote sensing (RS) images is perennial topic of interest in the RS community. Recently, transfer learning, especially for fine-tuning pre-trained convolutional neural networks (CNNs), has been proposed as a feasible strategy for RS scene classification. However, because the target domain (i.e., the RS images) and the source domain (e.g., ImageNet) are quite different, simply using the model pre-trained on an ImageNet dataset presents some difficulties. The RS images and the pre-trained models need to be properly adjusted to build a better classification system. In this study, an adaptive learning strategy for transferring a CNN-based model is proposed. First, an adaptive transform is used to adjust the original size of the RS image to a certain size, which is tailored to the input of the subsequent pre-trained model. Then, an adaptive transferring model is proposed to automatically learn what knowledge from the pre-trained model should be transferred to the RS scene classification model. Finally, in combination with a label smoothing approach, adaptive label is presented to generate soft labels based on the statistics of the classification model predictions for each category, which is beneficial for learning the relationships between the target and non-target categories of scenes. In general, the proposed methods adaptively manage the input, model, and label simultaneously, which leads to better classification performance for RS scene classification. The proposed methods are tested on three widely-used data sets and the obtained results show that the proposed methods provide competitive classification accuracy compared to the state-of-the-art methods. Yushi Chen 0002, Pedram Ghamisi |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Universal Adversarial Examples in Remote Sensing: Methodology and BenchmarkabstractDeep neural networks have achieved great success in many important remote sensing tasks. Nevertheless, their vulnerability to adversarial examples should not be neglected. In this study, we systematically analyze the Universal Adversarial Examples in Remote Sensing (UAE-RS) data for the first time, without any knowledge from the victim model. Specifically, we propose a novel black-box adversarial attack method, namely, Mixup-Attack, and its simple variant Mixcut-Attack, for remote sensing data. The key idea of the proposed methods is to find common vulnerabilities among different networks by attacking the features in the shallow layer of a given surrogate model. Despite their simplicity, the proposed methods can generate transferable adversarial examples that deceive most of the state-of-the-art deep neural networks in both scene classification and semantic segmentation tasks with high success rates. We further provide the generated universal adversarial examples in the dataset named UAE-RS, which is the first dataset that provides black-box adversarial samples in the remote sensing field. We hope UAE-RS may serve as a benchmark that helps researchers design deep neural networks with strong resistance toward adversarial attacks in the remote sensing field. Codes and the UAE-RS dataset are available online (https://github.com/YonghaoXu/UAE-RS). Yonghao Xu, Pedram Ghamisi |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Self-Supervised Learning With Adaptive Distillation for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification is an important topic in the community of remote sensing, which has a wide range of applications in geoscience. Recently, deep learning-based methods have been widely used in HSI classification. However, due to the scarcity of labeled samples in HSI, the potential of deep learning-based methods has not been fully exploited. To solve this problem, a self-supervised learning (SSL) method with adaptive distillation is proposed to train the deep neural network with extensive unlabeled samples. The proposed method consists of two modules: adaptive knowledge distillation with spatial–spectral similarity and 3-D transformation on HSI cubes. The SSL with adaptive knowledge distillation uses the self-supervised information to train the network by knowledge distillation, where self-supervised knowledge is the adaptive soft label generated by spatial–spectral similarity measurement. The SSL with adaptive knowledge distillation mainly includes the following three steps. First, the similarity between unlabeled samples and object classes in HSI is generated based on the spatial–spectral joint distance (SSJD) between unlabeled samples and labeled samples. Second, the adaptive soft label of each unlabeled sample is generated to measure the probability that the unlabeled sample belongs to each object class. Third, a progressive convolutional network (PCN) is trained by minimizing the cross-entropy between the adaptive soft labels and the probabilities generated by the forward propagation of the PCN. The SSL with 3-D transformation rotates the HSI cube in both the spectral domain and the spatial domain to fully exploit the labeled samples. Experiments on three public HSI data sets have demonstrated that the proposed method can achieve better performance than existing state-of-the-art methods. Jun Yue 0004, Leyuan Fang, Hossein Rahmani 0001, Pedram Ghamisi |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Adaptive Spatial Pyramid Constraint for Hyperspectral Image Classification With Limited Training SamplesabstractDeep learning-based methods have made significant progress in hyperspectral image (HSI) classification in recent years. However, deep learning-based methods usually rely on a large number of samples, and in many cases, it is difficult to label HSI and only limited training samples are available. To solve this problem, an HSI classification method based on adaptive spatial pyramid constraint (ASPC) is proposed to make full use of the global spatial neighborhood information of the labeled samples, which can improve the generalization ability of the classification model. The main steps of the proposed method are as follows. First, an HSI complexity evaluation method based on edge detection is proposed to assess the homogeneity of the objects in the HSI. Second, an HSI pyramid segmentation method based on spatial pyramid is proposed to generate multiscale subregions, where HSI complexity is used to adaptively determine the scale of the segmentation. Third, a spatial supervised constraint is proposed to generate the loss function of labeled subregions. Fourth, a spatial unsupervised constraint is proposed to generate the loss function of unlabeled subregions. The proposed method fully explores the spatial-spectral correlation between unlabeled samples and labeled samples, and add corresponding constraints to the training objective according to the correlation. By adding the ASPC, the trained model becomes more robust and can make full use of the limited training samples. To verify the effectiveness of the proposed method, three benchmark hyperspectral datasets are used to verify the performance of the proposed method. Experimental results show that the performance of this method is better than the existing state-of-the-art methods. Jun Yue 0004, Dingshun Zhu, Leyuan Fang, Pedram Ghamisi, Yaowei Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Contour Structural Profiles: An Edge-Aware Feature Extractor for Hyperspectral Image ClassificationabstractFeature extraction provides an effective tool to classify hyperspectral images (HSIs). However, most hyperspectral feature extraction methods tend to yield an over-smoothed phenomenon, which leads to inconsistency between the homogeneous regions and the ground objects in the actual scene. To alleviate this problem, an edge-aware feature extractor called contour structural profiles (CSPs) is proposed to extract the discriminative features for hyperspectral images classification (HSIC). The proposed classification method comprises three components. First, the spectral dimension of the HSI is reduced with an averaging-based method. Then, an edge-aware total variation (TV) model is constructed to extract the contour structural profile, in which a learned contour probability map is served as one of the major cues in the feature extraction process. Next, multiscale structural profiles (MSSPs) are constructed using the edge-aware TV model with different parameters so as to fully characterize ground objects with different scales. Finally, the MSSPs are fused with a kernel principal component analysis (KPCA) followed by a spectral classifier to obtain the final classification map. Experimental results on several publicly available hyperspectral datasets illustrate that the proposed method obtains superior classification performance over several state-of-the-art classification approaches, especially when the number of training samples is insufficient. Ying Zhang 0063, Puhong Duan, Jianxu Mao, Xudong Kang, Leyuan Fang, Pedram Ghamisi |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Deep Bilateral Filtering Network for Point-Supervised Semantic Segmentation in Remote Sensing ImagesabstractSemantic segmentation methods based on deep neural networks have achieved great success in recent years. However, training such deep neural networks relies heavily on a large number of images with accurate pixel-level labels, which requires a huge amount of human effort, especially for large-scale remote sensing images. In this paper, we propose a point-based weakly supervised learning framework called the deep bilateral filtering network (DBFNet) for the semantic segmentation of remote sensing images. Compared with pixel-level labels, point annotations are usually sparse and cannot reveal the complete structure of the objects; they also lack boundary information, thus resulting in incomplete prediction within the object and the loss of object boundaries. To address these problems, we incorporate the bilateral filtering technique into deeply learned representations in two respects. First, since a target object contains smooth regions that always belong to the same category, we perform deep bilateral filtering (DBF) to filter the deep features by a nonlinear combination of nearby feature values, which encourages the nearby and similar features to become closer, thus achieving a consistent prediction in the smooth region. In addition, the DBF can distinguish the boundary by enlarging the distance between the features on different sides of the edge, thus preserving the boundary information well. Experimental results on two widely used datasets, the ISPRS 2-D semantic labeling Potsdam and Vaihingen datasets, demonstrate that our proposed DBFNet can achieve a highly competitive performance compared with state-of-the-art fully-supervised methods. Code is available at https://github.com/Luffy03/DBFNet. Linshan Wu, Leyuan Fang, Jun Yue 0004, Bob Zhang 0001, Pedram Ghamisi |
IEEE Trans. Image Process. | 5 |
| 2022 | Consistency-Regularized Region-Growing Network for Semantic Segmentation of Urban Scenes With Point-Level AnnotationsabstractDeep learning algorithms have obtained great success in semantic segmentation of very high-resolution (VHR) remote sensing images. Nevertheless, training these models generally requires a large amount of accurate pixel-wise annotations, which is very laborious and time-consuming to collect. To reduce the annotation burden, this paper proposes a consistency-regularized region-growing network (CRGNet) to achieve semantic segmentation of VHR remote sensing images with point-level annotations. The key idea of CRGNet is to iteratively select unlabeled pixels with high confidence to expand the annotated area from the original sparse points. However, since there may exist some errors and noises in the expanded annotations, directly learning from them may mislead the training of the network. To this end, we further propose the consistency regularization strategy, where a base classifier and an expanded classifier are employed. Specifically, the base classifier is supervised by the original sparse annotations, while the expanded classifier aims to learn from the expanded annotations generated by the base classifier with the region-growing mechanism. The consistency regularization is thereby achieved by minimizing the discrepancy between the predictions from both the base and the expanded classifiers. We find such a simple regularization strategy is yet very useful to control the quality of the region-growing mechanism. Extensive experiments on two benchmark datasets demonstrate that the proposed CRGNet significantly outperforms the existing state-of-the-art methods. Codes and pre-trained models are available online (https://github.com/YonghaoXu/CRGNet). Yonghao Xu, Pedram Ghamisi |
IEEE Trans. Image Process. | 2 |
| 2022 | Normal Assisted Pixel-Visibility Learning With Cost Aggregation for Multiview StereoabstractMultiple-View Stereo (MVS) aims to reconstruct the dense 3D representations of scenes. MVS has potential applications in the fields of autonomous driving (unstructured environment construction) and robotic navigation (visual-inertial navigation). To mitigate the error of depth estimation in low-textured or occluded regions, this work proposes a two-stage multi-view stereo network for fast and accurate depth estimation. The improvements of this work over the state of the art are as follows: 1) Sparse costs are constructed to jointly predict the initial depth map and surface normal by cost regularization, which proves that the surface normals can be estimated in this way with low memory consumption. 2) A new edge refinement block is developed to refine the coarse surface normal to obtain a fine-grained surface normal map. 3) Instead of using the general variance-based metric to equally aggregate cost, a new content-adaptive cost aggregation mechanism based on the similarity of the neighboring surface normal is designed for reliable cost aggregation. To the best of our knowledge, the proposed work is the first trainable network that leverages surface normal as guidance to capture neighboring pixel-visibility, which is an effective supplement to existing depth/normal estimation frameworks. Experimental results indicate that our method can not only achieve accurate depth estimation for scene perception but also make no concession to the real-time performance and limited memory bottleblock. Multiple-view stereo (MVS) aims to reconstruct the dense 3D representations of scenes. It is widely used in the fields of industrial measurement, autonomous driving, and robotic navigation. To mitigate the error of depth estimation in challenging scenarios, this work proposes a two-stage multi-view stereo network for fast and accurate depth estimation. Our method is the first trainable network that leverages surface normal as pixel-visibility guidance to aggregate reliable cost, which could achieve accurate depth estimation and provide the perception ability for the robot. The proposed method has great potential in the fields of 3D reconstruction, industrial measurement, and robotic navigation to estimate real-time and accurate depth with limited memory consumption. Kevin W. Tong, Xiaorong Guan, Jian Kang 0005, Zhao-Hui Sun, Rob Law 0001, Pedram Ghamisi, Qi Wu 0003 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2021 | CDCEO'21 - First Workshop on Complex Data Challenges in Earth ObservationabstractHigh-resolution remote sensing technology for Earth Observation (EO) has radically changed how we monitor the state of our planet around the clock. An effective interpretation of the resulting complex large-scale time series adopts the best machine learning techniques from signal processing, computer vision, pattern recognition, and artificial intelligence. The First Workshop on Complex Data Challenges in Earth Observation was open to both method development and advanced applications in a wide range of related topics, including image and signal processing, gap-filling, data fusion, feature extraction, prediction of spatio-temporal features, and the detection of rules underlying the observed state transitions and causal relationships. The full agenda, featuring keynotes and a selection of high quality contributed talks is available online at www.iarai.ac.at/cdceo21 Aleksandra Gruca, Pedro Herruzo, Pilar Rípodas, Andrzej Kucik, Christian Briese, Michael Kopp 0001, Sepp Hochreiter, Pedram Ghamisi, David P. Kreil |
CIKM | 8 |
| 2021 | Spectral Unmixing Using Deep Convolutional Encoder-DecoderabstractIn this paper, we introduce ‘Unmixing Deep Image Prior’ (UnDIP), a deep learning-based technique for the linear hyperspectral unmixing problem. The proposed method contains two steps. First, the endmembers are extracted using a geometric endmember extraction method, i.e. a simplex volume maximization in a subspace of the dataset. Then, the abundances are estimated using a deep image prior. The proposed deep image prior uses a convolutional neural network to estimate the fractional abundances, relying on the extracted endmembers and the observed hyperspectral dataset. The results show considerable improvements compared to state-of-the-art methods. Behnood Rasti, Bikram Koirala, Paul Scheunders, Pedram Ghamisi |
IGARSS | 4 |
| 2021 | Boosting Hyperspectral Image Unmixing Using Denoising: Four ScenariosabstractWe present an analysis of the influence of noise on the unmixing of hyperspectral data. We propose four scenarios to 1) investigate the effect of noise reduction as a preprocessing step on the performance of hyperspectral unmixing and 2) study the relation between noise and different endmembers selection strategies. Experiments are conducted on a simu-1ated and a real datasets with a wide range of signal to noise ratios (from 10 to 50 dB). Behnood Rasti, Bikram Koirala, Paul Scheunders, Pedram Ghamisi, Richard Gloaguen |
IGARSS | 4 |
| 2021 | When is the Right Time to Apply Denoising?abstractRemote sensing data is contaminated with different types of noise that can severely affect the analysis of this data. Generally, in modern treatment chains of satellite and aerial data, denoising techniques are applied to atmospherically corrected images prior to further analysis (e.g., classification). However, since the noise contaminates the measured radiance at the sensor, it can influence the atmospheric correction in itself and consequently the remaining of the processing chain. In this paper, we compare the performance of a denoising technique, when applied before or after atmospheric correction. Our observations challenge the current de facto paradigm of denoising in a processing chain of spaceborne and airborne remotely sensed images. Kasra Rafiezadeh Shahi, Behnood Rasti, Pedram Ghamisi, Paul Scheunders, Richard Gloaguen |
IGARSS | 3 |
| 2021 | U-IMG2DSM: Unpaired Simulation of Digital Surface Models With Generative Adversarial NetworksabstractHigh-resolution digital surface models (DSMs) provide valuable height information about the Earth's surface, which can be successfully combined with other types of remotely sensed data in a wide range of applications. However, the acquisition of DSMs with high spatial resolution is extremely time-consuming and expensive with their estimation from a single optical image being an ill-possed problem. To overcome these limitations, this letter presents a new unpaired approach to obtain DSMs from optical images using deep learning techniques. Specifically, our new deep neural model is based on variational autoencoders (VAEs) and generative adversarial networks (GANs) to perform image-to-image translation, obtaining DSMs from optical images. Our newly proposed method has been tested in terms of photographic interpretation, reconstruction error, and classification accuracy using three well-known remotely sensed data sets with very high spatial resolution (obtained over Potsdam, Vaihingen, and Stockholm). Our experimental results demonstrate that the proposed approach obtains satisfactory reconstruction rates that allow enhancing the classification results for these images. The source code of our method is available from: https://github.com/mhaut/UIMG2DSM. Mercedes Eugenia Paoletti, Juan Mario Haut, Pedram Ghamisi, Naoto Yokoya, Javier Plaza, Antonio Plaza |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2021 | Multiscale Densely-Connected Fusion Networks for Hyperspectral Images ClassificationabstractConvolutional neural network (CNN) has demonstrated to be a powerful tool for hyperspectral images (HSIs) classification. Previous CNN-based HSI classification methods only adopt the fixed-size patches to train the CNN model, and such single scale patches may not reflect the complex spatial structural information in the HSIs. In addition, although different layers of CNN can extract features of multiple scales, the traditional CNN model can only utilize features from the highest level for the classification task. These features, however, do not fully consider the strong complementary yet correlated information among different layers. To address these issues, in this paper, a multiscale densely-connected convolutional network (MS-DenseNet) framework is proposed to sufficiently exploit multiple scales information for the HSIs classification. Specifically, for each pixel, the MS-DenseNet, first, extracts its surrounding patches of multiple scales. These patches can separately constitute multiple scale training and testing samples. Within each specific scale sample, instead of using the forward convolutional layers, the MS-DenseNet adopts the dense blocks, which can connect each layer to other layers in a feed-forward fashion and thus can exploit the information among different layers for training and testing. Furthermore, since high correlations exist in patches of different scales, the MS-DenseNet introduces several dense blocks to fuse the multiscale information among different layers for the final HSI classification. Experimental results on several real HSIs demonstrate the superiority of the proposed MS-DenseNet over single scale-based CNN classification model and several well-known classification methods. Jie Xie 0002, Nanjun He, Leyuan Fang, Pedram Ghamisi |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Fusion of Dual Spatial Information for Hyperspectral Image ClassificationabstractThe inclusion of spatial information into spectral classifiers for fine-resolution hyperspectral imagery has led to significant improvements in terms of classification performance. The task of spectral-spatial hyperspectral image (HSI) classification has remained challenging because of high intraclass spectrum variability and low interclass spectral variability. This fact has made the extraction of spatial information highly active. In this work, a novel HSI classification framework using the fusion of dual spatial information is proposed, in which the dual spatial information is built by both exploiting pre-processing feature extraction and post-processing spatial optimization. In the feature extraction stage, an adaptive texture smoothing method is proposed to construct the structural profile (SP), which makes it possible to precisely extract discriminative features from HSIs. The SP extraction method is used here for the first time in the remote sensing community. Then, the extracted SP is fed into a spectral classifier. In the spatial optimization stage, a pixel-level classifier is used to obtain the class probability followed by an extended random walker-based spatial optimization technique. Finally, a decision fusion rule is utilized to fuse the class probabilities obtained by the two different stages. Experiments performed on three data sets from different scenes illustrate that the proposed method can outperform other state-of-the-art classification techniques. In addition, the proposed feature extraction method, i.e., SP, can effectively improve the discrimination between different land covers. Puhong Duan, Pedram Ghamisi, Xudong Kang, Behnood Rasti, Shutao Li 0001, Richard Gloaguen |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Hyperspectral Image Classification With Attention-Aided CNNsabstractConvolutional neural networks (CNNs) have been widely used for hyperspectral image classification. As a common process, small cubes are first cropped from the hyperspectral image and then fed into CNNs to extract spectral and spatial features. It is well known that different spectral bands and spatial positions in the cubes have different discriminative abilities. If fully explored, this prior information will help improve the learning capacity of CNNs. Along this direction, we propose an attention-aided CNN model for spectral-spatial classification of hyperspectral images. Specifically, a spectral attention subnetwork and a spatial attention subnetwork are proposed for spectral and spatial classifications, respectively. Both of them are based on the traditional CNN model and incorporate attention modules to aid networks that focus on more discriminative channels or positions. In the final classification phase, the spectral classification result and the spatial classification result are combined together via an adaptively weighted summation method. To evaluate the effectiveness of the proposed model, we conduct experiments on three standard hyperspectral data sets. The experimental results show that the proposed model can achieve superior performance compared with several state-of-the-art CNN-related models. Renlong Hang, Zhu Li 0001, Qingshan Liu 0001, Pedram Ghamisi, Shuvra S. Bhattacharyya |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | Classification of Hyperspectral Images via Multitask Generative Adversarial NetworksabstractDeep learning has shown its huge potential in the field of hyperspectral image (HSI) classification. However, most of the deep learning models heavily depend on the quantity of available training samples. In this article, we propose a multitask generative adversarial network (MTGAN) to alleviate this issue by taking advantage of the rich information from unlabeled samples. Specifically, we design a generator network to simultaneously undertake two tasks: the reconstruction task and the classification task. The former task aims at reconstructing an input hyperspectral cube, including the labeled and unlabeled ones, whereas the latter task attempts to recognize the category of the cube. Meanwhile, we construct a discriminator network to discriminate the input sample coming from the real distribution or the reconstructed one. Through an adversarial learning method, the generator network will produce real-like cubes, thus indirectly improving the discrimination and generalization ability of the classification task. More importantly, in order to fully explore the useful information from shallow layers, we adopt skip-layer connections in both reconstruction and classification tasks. The proposed MTGAN model is implemented on three standard HSIs, and the experimental results show that it is able to achieve higher performance than other state-of-the-art deep learning models. Renlong Hang, Feng Zhou 0006, Qingshan Liu 0001, Pedram Ghamisi |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2020 | Sun Glint Removal of Hyperspectral Images via Texture-Aware Total VariationabstractSun glint, as the spectral reflection of solar radiation on non-flat water surfaces, is a serious confounding factor for coastal shallow-water environments. When the coastal areas are observed with a hyperspectral sensor, the existing sun glint in the produced images can seriously influence the quality of the image interpretation. To solve this issue, in this paper, we propose a novel sun glint removal method based on a variation model for hyperspectral images (HSIs). The proposed method aims to decompose the original HSI into a desired clean image and a sun glint image. To achieve this, we exploit a texture-aware total variation to remove the sun glint in HSIs, where the texture information is imposed on the total variation regularization to highlight sun glint. Experiments on simulated and real datasets demonstrate that our method can obtain outstanding performance with respect to other state-of-the-art approaches. Puhong Duan, Jian Kang 0005, Xudong Kang, Pedram Ghamisi, Shutao Li 0001 |
IGARSS | 4 |
| 2020 | Intrinsic Image Decomposition-Based Resolution Enhancement for Mineral MappingabstractHyperspectral imaging plays an important role for mineral mapping in a nondestructive and noninvasive way. In this paper, a novel resolution enhancement method is proposed based on the principle of intrinsic image decomposition for mineral mapping. This method is based on an assumption that hyperspectral image (HSI) can be decomposed into a reflectance component and an illumination component. Based on this idea, the RGB image is first transformed into Intensity-Hue-Saturation (IHS) space, and the intensity channel is considered as the illumination component of the HSI with an ideal high spatial resolution. Then, the reflectance component of the ideal HSI is estimated with the downsampled HSI image and the downsampled intensity channel. Finally, the HSI with high resolution can be reconstructed by utilizing the estimated illumination and the reflectance components. Experimental results validate the effectiveness of the proposed method qualitatively and quantitatively by outperforming several state-of-the-art approaches. Puhong Duan, Pedram Ghamisi, Robert Jackisch, Xudong Kang, Richard Gloaguen, Shutao Li 0001 |
IGARSS | 2 |
| 2020 | Remote Sensing and Deep Learning for Sustainable MiningabstractThis paper brings together advances in remote sensing and deep learning for mineral mapping in a sustainable way. In more detail, we propose a multisensor feature fusion approach to integrate heterogeneous RGB, multispectral, and hyperspectral images for sustainable mining. The proposed approach is composed of two main steps; Feature extraction and classification. In the feature extraction step, we develop a three-stream convolutional neural network to extract high-level information from the input multisensor data. In the classification step, we develop a multisensor composite kernel approach to perform fusion and mapping simultaneously. The proposed approach produces very high quality classification maps with exceptional results in terms of classification accuracies. Pedram Ghamisi, Hao Li 0019, Robert Jackisch, Behnood Rasti, Richard Gloaguen |
IGARSS | 1 |
| 2020 | Towards 4D Virtual Outcrops with Hyperspectral ImagingabstractAccurately mapping lithology and geological structures remains a challenge in rough terrain or in active mining areas. We propose that the integration of terrestrial and drone-borne multi-sensor remote sensing techniques can significantly boost the reliability, safety, and efficiency of geological activities in exploration and for the monitoring of mining activities. We have now developed a complete procedural chain to jointly and accurately process Structure-from-Motion Multi-View Stereo point clouds and hyperspectral data cubes in the visible to near-infrared (VNIR) and short-wave infrared (SWIR), as well as long-wave infrared (LWIR) ranges acquired by terrestrial sensors. Hyperspectral data are processed using spectroscopic and machine learning algorithms to generate meaningful 2.5D (i.e., surface) maps that are available to geologists on the ground shortly after data acquisition. We classify the geological information content using innovative machine learning techniques. We validate the remote sensing data with in-situ mineralogical and structural measurements. Repeated acquisitions allow then to integrate a time component. Richard Gloaguen, Moritz Kirsch, Sandra Lorenz, René Booysen, Robert Zimmermann, Pedram Ghamisi, Behnood Rasti |
IGARSS | 6 |
| 2020 | Fusion of Multispectral LiDAR and Hyperspectral ImageryabstractThis paper presents a technique for the fusion of multispectral LiDAR and hyperspectral data. The proposed method is based on the fusion of the features of multispectral LiDAR and hyperspectral data projected in two different subspaces. First, the spatial features are extracted from both data using morphological filters. Then, the fused features are estimated by proposing a novel constraint penalized cost function. The estimated fused features are used for the purpose of mapping. The classification accuracies obtained by applying a random forest classifier on the fused data confirm considerable improvements compared with the other methods used in the experiments. Behnood Rasti, Pedram Ghamisi, Richard Gloaguen |
IGARSS | 2 |
| 2020 | Multispectral Change Detection With Bilinear Convolutional Neural NetworksabstractRecently, deep learning has been demonstrated to be an effective tool to detect changes in bitemporal remote sensing images. However, most existing methods based on deep learning obtain the ultimate change map by analyzing the difference image (DI) or the stacked feature vectors of input images, which cannot sufficiently capture the relationship between the two input images to obtain the change information. In this letter, a new method named bilinear convolutional neural networks (BCNNs) is proposed to detect changes in bitemporal multispectral images. The model can be trained end to end with two symmetric convolutional neural networks (CNNs), which are capable of learning the feature representation from bitemporal images and utilizing the relations between the two input images by a linear outer product operation in an effective way. Specifically, two sets of patches obtained from two multispectral images of different times are first input into two CNNs to extract deep features, respectively. Then, the matrix outer product is applied on the output feature maps to obtain the combined bilinear features. Finally, the ultimate change detected result can be produced by applying the softmax classifier on the combined features. Experimental results on real multispectral data sets demonstrate the superiority of the proposed method over several well-known change-detection approaches. Shutao Li 0001, Leyuan Fang, Pedram Ghamisi |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2020 | Hyperspectral Mixed Gaussian and Sparse Noise ReductionabstractHyperspectral images (HSIs) are often degraded by different noise types such as Gaussian and sparse noise. In this letter, a hyperspectral mixed Gaussian and sparse noise reduction technique, the HyMiNoR, is proposed. The proposed technique, hierarchically, removes the mixed noise. First, the Gaussian noise is removed using a recently developed automatic hyperspectral noise removal technique called hyperspectral restoration (HyRes). Then, we develop a novel sparse noise removal technique to remove the sparse noise, including salt and pepper noise, missing pixels, and missing lines. The performance of the proposed approach has been validated using both real and simulated data sets. Results on the simulated data set confirm considerable improvements in terms of signal-to-noise ratio and singular angle distance compared to the state-of-the-art techniques used in the experiments. In addition, visual improvements can be clearly observed in the case of real data set experiments. Behnood Rasti, Pedram Ghamisi, Jón Atli Benediktsson |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2020 | Multichannel Pulse-Coupled Neural Network-Based Hyperspectral Image VisualizationabstractHyperspectral Image (HSI) visualization, which aims at displaying as much material information of original images as possible on a trichromatic monitor with natural color, plays an important role in image interpretation and analysis. However, most of the HSI visualization methods only focus on presenting the detail information of a scene without providing natural colors and distinguishing land covers with similar colors. In order to address this problem, this article proposes a multichannel pulse-coupled neural network (MPCNN)-based HSI visualization method, which consists of the following steps. First, the MPCNN is proposed and explored to fuse the original HSI so as to obtain a fused band with rich spatial details. Then, a color mapping scheme is proposed to determine the weights of red, green, and blue (RGB) channels. Finally, the weighted RGB channels are stacked together for visualization. Experiments performed on four hyperspectral data sets demonstrate that the proposed method not only displays the HSI with nature colors but also improves the details in the image. The effectiveness of the proposed method is demonstrated in terms of both visual effect and objective indexes. Puhong Duan, Xudong Kang, Shutao Li 0001, Pedram Ghamisi |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2020 | Classification of Hyperspectral and LiDAR Data Using Coupled CNNsabstractIn this article, we propose an efficient and effective framework to fuse hyperspectral and light detection and ranging (LiDAR) data using two coupled convolutional neural networks (CNNs). One CNN is designed to learn spectral-spatial features from hyperspectral data, and the other one is used to capture the elevation information from LiDAR data. Both of them consist of three convolutional layers, and the last two convolutional layers are coupled together via a parameter-sharing strategy. In the fusion phase, feature-level and decision-level fusion methods are simultaneously used to integrate these heterogeneous features sufficiently. For the feature-level fusion, three different fusion strategies are evaluated, including the concatenation strategy, the maximization strategy, and the summation strategy. For the decision-level fusion, a weighted summation strategy is adopted, where the weights are determined by the classification accuracy of each output. The proposed model is evaluated on an urban data set acquired over Houston, USA, and a rural one captured over Trento, Italy. On the Houston data, our model can achieve a new record overall accuracy (OA) of 96.03%. On the Trento data, it achieves an OA of 99.12%. These results sufficiently certify the effectiveness of our proposed model. Renlong Hang, Zhu Li 0001, Pedram Ghamisi, Danfeng Hong, Guiyu Xia, Qingshan Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | Heterogeneous Transfer Learning for Hyperspectral Image Classification Based on Convolutional Neural NetworkabstractDeep convolutional neural networks (CNNs) have shown their outstanding performance in the hyperspectral image (HSI) classification. The success of CNN-based HSI classification relies on the availability sufficient training samples. However, the collection of training samples is expensive and time consuming. Besides, there are many pretrained models on large-scale data sets, which extract the general and discriminative features. The proper reusage of low-level and midlevel representations will significantly improve the HSI classification accuracy. The large-scale ImageNet data set has three channels, but HSI contains hundreds of channels. Therefore, there are several difficulties to simply adapt the pretrained models for the classification of HSIs. In this article, heterogeneous transfer learning for HSI classification is proposed. First, a mapping layer is used to handle the issue of having different numbers of channels. Then, the model architectures and weights of the CNN trained on the ImageNet data sets are used to initialize the model and weights of the HSI classification network. Finally, a well-designed neural network is used to perform the HSI classification task. Furthermore, attention mechanism is used to adjust the feature maps due to the difference between the heterogeneous data sets. Moreover, controlled random sampling is used as another training sample selection method to test the effectiveness of the proposed methods. Experimental results on four popular hyperspectral data sets with two training sample selection strategies show that the transferred CNN obtains better classification accuracy than that of state-of-the-art methods. In addition, the idea of heterogeneous transfer learning may open a new window for further research. Xin He 0004, Yushi Chen 0002, Pedram Ghamisi |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | Invariant Attribute Profiles: A Spatial-Frequency Joint Feature Extractor for Hyperspectral Image ClassificationabstractSo far, a large number of advanced techniques have been developed to enhance and extract the spatially semantic information in hyperspectral image processing and analysis. However, locally semantic change, such as scene composition, relative position between objects, spectral variability caused by illumination, atmospheric effects, and material mixture, has been less frequently investigated in modeling spatial information. Consequently, identifying the same materials from spatially different scenes or positions can be difficult. In this article, we propose a solution to address this issue by locally extracting invariant features from hyperspectral imagery (HSI) in both spatial and frequency domains, using a method called invariant attribute profiles (IAPs). IAPs extract the spatial invariant features by exploiting isotropic filter banks or convolutional kernels on HSI and spatial aggregation techniques (e.g., superpixel segmentation) in the Cartesian coordinate system. Furthermore, they model invariant behaviors (e.g., shift, rotation) by the means of a continuous histogram of oriented gradients constructed in a Fourier polar coordinate. This yields a combinatorial representation of spatial-frequency invariant features with application to HSI classification. Extensive experiments conducted on three promising hyperspectral data sets (Houston2013 and Houston2018) to demonstrate the superiority and effectiveness of the proposed IAP method in comparison with several state-of-the-art profile-related techniques. The codes will be available from the website: https://sites.google.com/view/danfeng-hong/data-code. Danfeng Hong, Xin Wu 0001, Pedram Ghamisi, Jocelyn Chanussot, Naoto Yokoya, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | Deep Metric Learning Based on Scalable Neighborhood Components for Remote Sensing Scene CharacterizationabstractWith the development of convolutional neural networks (CNNs), the semantic understanding of remote sensing (RS) scenes has been significantly improved based on their prominent feature encoding capabilities. While many existing deep-learning models focus on designing different architectures, only a few works in the RS field have focused on investigating the performance of the learned feature embeddings and the associated metric space. In particular, two main loss functions have been exploited: the contrastive and the triplet loss. However, the straightforward application of these techniques to RS images may not be optimal in order to capture their neighborhood structures in the metric space due to the insufficient sampling of image pairs or triplets during the training stage and to the inherent semantic complexity of remotely sensed data. To solve these problems, we propose a new deep metric learning approach, which overcomes the limitation on the class discrimination by means of two different components: 1) scalable neighborhood component analysis (SNCA) that aims at discovering the neighborhood structure in the metric space and 2) the cross-entropy loss that aims at preserving the class discrimination capability based on the learned class prototypes. Moreover, in order to preserve feature consistency among all the minibatches during training, a novel optimization mechanism based on momentum update is introduced for minimizing the proposed loss. An extensive experimental comparison (using several state-of-the-art models and two different benchmark data sets) has been conducted to validate the effectiveness of the proposed method from different perspectives, including: 1) classification; 2) clustering; and 3) image retrieval. The related codes of this article will be made publicly available for reproducible research by the community. Jian Kang 0005, Rubén Fernández-Beltran, Zhen Ye 0009, Xiaohua Tong, Pedram Ghamisi, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2019 | Mineral Mapping of Drill Core Hyperspectral Data with Extreme Learning MachinesabstractHyperspectral scanners are increasingly being used in the mining industry as a non-destructive and non-invasive technique to efficiently map minerals in drill core samples. Hyperspectral data allows the characterization of different mineral assemblages, structural features and alteration patterns based on reflectance spectrum profiles. Traditional methods to analysis drill core hyperspectral data include the use of reference spectral libraries by visual analysis or a well established software. However, although these approaches produce good results, they are time-consuming and prone to errors. Therefore, in this paper, we take advantage of the latest and advanced machine learning techniques proposed in different scientific fields and explore the use of extreme learning machines (ELM) to map minerals in drill core hyperspectral data. This is a supervised technique that provides fast and automatic means to characterize hyperspectral data. To be able to implement this technique, a reference map was generated from the drill core hyperspectral data. The obtained results indicate that ELM can successfully map minerals in drill core hyperspectral data producing better quantitative and qualitative results than a typical RF classifier. Isabel Cecilia Contreras Acosta, Mahdi Khodadadzadeh, Pedram Ghamisi, Richard Gloaguen |
IGARSS | 3 |
| 2019 | A Novel Composite Kernel Approach for Multisensor Remote Sensing Data FusionabstractThe increased availability of active and passive data captured over the same scene of interest makes it desirable to jointly utilize multisensor data to perform accurate classification. This paper proposes a novel fusion approach to integrate hyperspectral and LiDAR-derived digital surface model for land-cover classification. In this context, we propose a novel multisensor composite kernel technique based on extreme learning machines (named as multisensor composite kernels (MCKs)), which is capable of combining different methods in the feature fusion level in an effective way. In the proposed approach, we use extinction profiles to extract spatial and elevation features of hyperspectral and LiDAR data. Then, hyperspectral Stein's unbiased risk estimator (HySURE) is applied to identify the subspace (informative features) of spectral, spatial, and elevation features. Finally, MCK is applied to the extracted spectral, spatial, and elevation features to produce the final classification map. Results obtained by the proposed approach reveal the fact that this approach can effectively fuse and classify hyperspectral and LiDAR images and improve the classification accuracy of each data source significantly. In addition, the proposed method is fully automatic. Pedram Ghamisi, Behnood Rasti, Richard Gloaguen |
IGARSS | 1 |
| 2019 | Multi-Source and multi-Scale Imaging-Data Integration to boost Mineral MappingabstractWe propose to develop an efficient and integrated exploration workflow that includes remote sensing data obtained by multiple types of sensors at different altitudes, a combination that has been identified as potentially disruptive technology for the mineral exploration sector. The fusion of multi-source and multi-temporal data is, therefore, a key challenge for a successful data integration. Ultimately, the objective is to boost the competitiveness, growth, sustainability, and attractiveness of the raw material sector. Richard Gloaguen, Margret C. Fuchs, Mahdi Khodadadzadeh, Pedram Ghamisi, Moritz Kirsch, René Booysen, Robert Zimmermann, Sandra Lorenz |
IGARSS | 4 |
| 2019 | Multisensor Feature Fusion Using Low-Rank Modeling and Component AnalysisabstractIn this paper, we propose a framework to fuse features extracted from hyperspectral and Light Detection And Ranging (LiDAR)-derived data. Spatial and elevation features are extracted from multisensor data using extinction profiles (EP). All the features, including the spectral ones, are fused using sparse and smooth low-rank analysis (SSLRA). In terms of classification accuracy, the proposed framework outperforms other studied fusion techniques used in the experiments. Behnood Rasti, Pedram Ghamisi, Richard Gloaguen |
IGARSS | 2 |
| 2019 | LW-ODF: A Light-Weight Object Detection Framework for Optical Remote Sensing ImageryabstractIn this paper, we propose to extract the multi-scaled and rotation-insensitive deep features to address the issues of object multi-solutions and rotations in geospatial object detection. To this end, we develop a novel object detection framework where a rotation-insensitive convolution neural network is applied for extracting multi-scaled and direction-insensitive feature representation and then the learned features can be fed into the ensemble classifier learning with fast feature pyramid. Such a non-end-to-end learning strategy intuitively reduces the computational cost without the additional performance loss, yielding an effective and efficient light-weight object detection framework. Experimental results conducted on the NWPU VHR-10 dataset demonstrate that the proposed framework outperforms several state-of-the-art baselines. Xin Wu 0001, Danfeng Hong, Pedram Ghamisi, Wei Li 0032, Ran Tao 0003 |
IGARSS | 3 |
| 2019 | Multiple convolutional layers fusion framework for hyperspectral image classification
Guangzhe Zhao, Guangyun Liu, Leyuan Fang, Bing Tu, Pedram Ghamisi |
Neurocomputing | 5 |
| 2019 | Multisensor Composite Kernels Based on Extreme Learning MachinesabstractIn this letter, we first propose multisensor composite kernel (MCK) extreme learning machines to fuse hyperspectral and light detection and ranging (LiDAR) features effectively. Then, based on the MCK, we develop a fully automatic fusion framework. In the proposed framework, spatial and elevation features of hyperspectral and LiDAR data are first extracted using extinction profiles. Then, hyperspectral Stein's unbiased risk estimator is utilized to extract the subspace (informative features) of spectral, spatial, and elevation features. The obtained results indicate that the proposed approach can successfully integrate and classify hyperspectral and LiDAR images to provide accurate classification results classification accuracies in an automatic manner. Pedram Ghamisi, Behnood Rasti, Jón Atli Benediktsson |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2019 | LiDAR Data Classification Using Spatial Transformation and CNNabstractLight detection and ranging (LiDAR) is a useful data acquisition technique, which is widely used in a variety of practical applications. The classification of LiDAR-derived rasterized digital surface model (LiDAR-DSM) is a fundamental technique in LiDAR data processing. In recent years, deep learning methods, especially convolutional neural networks (CNNs), have shown their capability in remote sensing areas, including LiDAR data processing. Traditional deep models empirically use a fixed neighborhood system as input to the network. Therefore, the weight and height of the input rectangle may not be optimal. In order to modify such handcrafted setting, a spatial transformation network is used here to identify optimal inputs. The transformed inputs are fed into a well-designed CNN to obtain the final classification results. Furthermore, morphological profiles are combined with spatial transformation CNN to further improve the classification accuracy. The proposed frameworks are tested on two LiDAR-DSMs (i.e., the Recology and Houston data sets). The experimental results show that the proposed models provide competitive results compared to the state-of-the-art methods. Furthermore, the proposed optimal input identification approach can also be found beneficial for other remote sensing applications. Xin He 0004, Aili Wang 0001, Pedram Ghamisi, Yushi Chen 0002 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2019 | Automatic Design of Convolutional Neural Network for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification is a core task in the remote sensing community, and recently, deep learning-based methods have shown their capability of accurate classification of HSIs. Among the deep learning-based methods, deep convolutional neural networks (CNNs) have been widely used for the HSI classification. In order to obtain a good classification performance, substantial efforts are required to design a proper deep learning architecture. Furthermore, the manually designed architecture may not fit a specific data set very well. In this paper, the idea of automatic CNN for the HSI classification is proposed for the first time. First, a number of operations, including convolution, pooling, identity, and batch normalization, are selected. Then, a gradient descent-based search algorithm is used to effectively find the optimal deep architecture that is evaluated on the validation data set. After that, the best CNN architecture is selected as the model for the HSI classification. Specifically, the automatic 1-D Auto-CNN and 3-D Auto-CNN are used as spectral and spectral-spatial HSI classifiers, respectively. Furthermore, the cutout is introduced as a regularization technique for the HSI spectral-spatial classification to further improve the classification accuracy. The experiments on four widely used hyperspectral data sets (i.e., Salinas, Pavia University, Kennedy Space Center, and Indiana Pines) show that the automatically designed data-dependent CNNs obtain competitive classification accuracy compared with the state-of-the-art methods. In addition, the automatic design of the deep learning architecture opens a new window for future research, showing the huge potential of using neural architectures' optimization capabilities for the accurate HSI classification. Yushi Chen 0002, Kaiqiang Zhu, Lin Zhu 0013, Xin He 0004, Pedram Ghamisi, Jón Atli Benediktsson |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2019 | Fusion of Multiple Edge-Preserving Operations for Hyperspectral Image ClassificationabstractIn this article, a novel hyperspectral image (HSI) classification method based on fusing multiple edge-preserving operations (EPOs) is proposed, which consists of the following steps. First, the edge-preserving features are obtained by performing different types of EPOs, i.e., local edge-preserving filtering and global edge-preserving smoothing on the dimension-reduced HSI. Then, with the assistance of a superpixel segmentation method, the edge-preserving features are further improved by considering the inter and intra spectral properties of superpixels. Finally, the spectral and edge-preserving features are fused to form one composite kernel, which is fed into the support vector machine (SVM) followed by a majority voting fusion scheme. Experimental results on three data sets demonstrate the superiority of the proposed method over several state-of-the-art classification approaches, especially when the training sample size is limited. Furthermore, 21 well-known methods, including mathematical morphology-based approaches, sparse representation models, and deep learning-based classifiers, are adopted to be compared with the proposed method on Houston data set with standard sets of training and test samples released during 2013 Data Fusion Contest, which also shows the effectiveness of the proposed method. Puhong Duan, Xudong Kang, Shutao Li 0001, Pedram Ghamisi, Jón Atli Benediktsson |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2019 | Hyperspectral Image Classification With Squeeze Multibias NetworkabstractA convolutional neural network (CNN) has recently demonstrated its outstanding capability for the classification of hyperspectral images (HSIs). Typical CNN-based methods usually adopt image patches as inputs to the network. However, a fixed-size image patch in HSI with complex spatial contexts may contain multiple ground objects of different classes, which will deteriorate the classification performance of the CNN. In addition, traditional convolutional layers adopted in the CNN have a huge amount of parameters needed to be tuned, which will cause high computational cost. To address the above-mentioned issues, a novel squeeze multibias network (SMBN) is proposed for HSI classification. Specifically, the proposed SMBN first introduces the multibias module (MBM), which incorporates multibias into the rectified linear unit layers. The MBM can decouple the feature maps of input patches into multiple response maps (corresponding to different ground objects) and adaptively select the meaningful maps for classification. Furthermore, the proposed SMBN replaces the traditional convolutional layer with a squeeze convolution module, which can greatly reduce the number of parameters in the network, thus saving the running time, while still maintaining high classification accuracy. Experimental results on three real HSIs demonstrate the superiority of the proposed SMBN method over several state-of-the-art classification approaches. Leyuan Fang, Guangyun Liu, Shutao Li 0001, Pedram Ghamisi, Jón Atli Benediktsson |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2019 | Cascaded Recurrent Neural Networks for Hyperspectral Image ClassificationabstractBy considering the spectral signature as a sequence, recurrent neural networks (RNNs) have been successfully used to learn discriminative features from hyperspectral images (HSIs) recently. However, most of these models only input the whole spectral bands into RNNs directly, which may not fully explore the specific properties of HSIs. In this paper, we propose a cascaded RNN model using gated recurrent units to explore the redundant and complementary information of HSIs. It mainly consists of two RNN layers. The first RNN layer is used to eliminate redundant information between adjacent spectral bands, while the second RNN layer aims to learn the complementary information from nonadjacent spectral bands. To improve the discriminative ability of the learned features, we design two strategies for the proposed model. Besides, considering the rich spatial information contained in HSIs, we further extend the proposed model to its spectral-spatial counterpart by incorporating some convolutional layers. To test the effectiveness of our proposed models, we conduct experiments on two widely used HSIs. The experimental results show that our proposed models can achieve better results than the compared models. Renlong Hang, Qingshan Liu 0001, Danfeng Hong, Pedram Ghamisi |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2019 | Deep Learning for Hyperspectral Image Classification: An OverviewabstractHyperspectral image (HSI) classification has become a hot topic in the field of remote sensing. In general, the complex characteristics of hyperspectral data make the accurate classification of such data challenging for traditional machine learning methods. In addition, hyperspectral imaging often deals with an inherently nonlinear relation between the captured spectral information and the corresponding materials. In recent years, deep learning has been recognized as a powerful feature-extraction tool to effectively address nonlinear problems and widely used in a number of image processing tasks. Motivated by those successful applications, deep learning has also been introduced to classify HSIs and demonstrated good performance. This survey paper presents a systematic review of deep learning-based HSI classification literatures and compares several strategies for this topic. Specifically, we first summarize the main challenges of HSI classification which cannot be effectively overcome by traditional machine learning methods, and also introduce the advantages of deep learning to handle these problems. Then, we build a framework that divides the corresponding works into spectral-feature networks, spatial-feature networks, and spectral-spatial-feature networks to systematically review the recent achievements in deep learning-based HSI classification. In addition, considering the fact that available training samples in the remote sensing field are usually very limited and training deep networks require a large number of samples, we include some strategies to improve classification performance, which can provide some guidelines for future studies on this topic. Finally, several representative deep learning-based classification methods are conducted on real HSIs in our experiments. Shutao Li 0001, Leyuan Fang, Yushi Chen 0002, Pedram Ghamisi, Jón Atli Benediktsson |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2019 | Fusion of Heterogeneous Earth Observation Data for the Classification of Local Climate ZonesabstractThis paper proposes a novel framework for fusing multi-temporal, multispectral satellite images and OpenStreetMap (OSM) data for the classification of local climate zones (LCZs). Feature stacking is the most commonly used method of data fusion but does not consider the heterogeneity of multimodal optical images and OSM data, which becomes its main drawback. The proposed framework processes two data sources separately and then combines them at the model level through two fusion models (the landuse fusion model and building fusion model) that aim to fuse optical images with landuse and buildings layers of OSM data, respectively. In addition, a new approach to detecting building incompleteness of OSM data is proposed. The proposed framework was trained and tested using the data from the 2017 IEEE GRSS Data Fusion Contest and further validated on one additional test (AT) set containing test samples that are manually labeled in Munich and New York. The experimental results have indicated that compared with the feature stacking-based baseline framework, the proposed framework is effective in fusing optical images with OSM data for the classification of LCZs with high generalization capability on a large scale. The classification accuracy of the proposed framework outperforms the baseline framework by more than 6% and 2% while testing on the test set of 2017 IEEE GRSS Data Fusion Contest and the AT set, respectively. In addition, the proposed framework is less sensitive to spectral diversities of optical satellite images and thus achieves more stable classification performance than the state-of-the-art frameworks. Guichen Zhang, Pedram Ghamisi, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2018 | The Need for Multi-Source, Multi-Scale Hyperspectral Imaging to Boost Non-Invasive Mineral ExplorationabstractThe high demand for raw materials in our post-industrial societies contrasts the increasing difficulties to find new mineral deposits. In Europe, accessible and high-grade deposits are mostly exhausted or currently mined. Hence, future exploration must focus on the remaining, more remote locations or penetrate much deeper into the Earth's crust. Sustaining mining activities in Europe would allow the development of key technologies but also sustainable and ethical production of technological metals. Thus, we suggest to focus research on advances in multi-scale and multi-sensor remote sensing-based Earth integration techniques. The scale should range from satellite to air- and drone-borne systems and include ground validation. Multi-sensor downscaling methods involving SAR and optical data are particularly promising. We demonstrate that the integration with other sensors and/or measures such as geophysical/geochemical data as well as non-conventional remote sensing features such as textures and geometries are of interest. Thus, ultimately, our objective is to boost the competitiveness, growth, sustainability and attractiveness of the raw material sector in Europe. While we focus on the raw material sector as it is currently of strategic importance, the required methods are transferable to most environmental studies. Richard Gloaguen, Pedram Ghamisi, Sandra Lorenz, Moritz Kirsch, Robert Zimmermann, René Booysen, Louis Andreani, Robert Jackisch, Erik Hermann, Laura Tusa, Gabriel Unger, Isabel Cecilia Contreras Acosta, Mahdi Khodadadzadeh, Margret C. Fuchs |
IGARSS | 2 |
| 2018 | Subspace Multinomial Logistic Regression Ensemble for Classification of Hyperspectral ImagesabstractExploiting multiple complementary classifiers in an ensemble framework has shown to be effective for improving hyperspectral image classification results, specially when the training samples are limited. With a different principle and based on this assumption that hyperspectal feature vectors effectively lie in a low-dimensional subspace, the subspace-based techniques have shown great classification performance. In this work, we propose a new ensemble method for accurate classification of hyperspectral images, which exploits the concept of subspace projection. For this purpose, we extend the subspace multinomial logistic regression classifier (MLRsub) to learn from multiple random subspaces for each class. More specifically, we impose diversity in constructing MLRsub by randomly selecting bootstrap samples from the training set and subsets of the original hyperspectral feature space, which lead to generate different class subspace features. Experimental results, conducted on two real hyperspectral data sets, indicate that the proposed method provides significant classification results in comparison with other state-of-the-art approaches. Mahdi Khodadadzadeh, Pedram Ghamisi, Isabel Cecilia Contreras Acosta, Richard Gloaguen |
IGARSS | 2 |
| 2018 | Feature Importance Analysis of Sentinel-2 Imagery for Large-Scale Urban Local Climate Zone ClassificationabstractThis paper evaluates different spectral-spatial features that can be extracted from Sentinel-2 imagery regarding their relevance for discriminating different Local Climate Zone (LCZ) classes. The features include spectral reflectance, spectral indices, Morphological Profiles (MPs), as well as Global Urban Footprint (GUF), the Open Street Map layers buildings and land use, and their combinations. Using a residual convolutional neural network (ResNet), a systematic analysis of feature importance is performed with a manually generated dataset distributed in Europe. The results of this evaluation are meant to provide guidance about the choice of both spectral and spatial features for the task of LCZ classification on a global scale. The results show that GUF and OSM can contribute to the classification performance, and ResNet relies less on additional features with the highest accuracy provided by the reflectance only. Chunping Qiu, Michael Schmitt 0003, Pedram Ghamisi, Lichao Mou, Xiao Xiang Zhu 0001 |
IGARSS | 3 |
| 2018 | Sparse and Smooth Feature Extraction for Hyperspectral ImageryabstractIn this paper, a hyperspectral feature extraction (FE) method called sparse and smooth low-rank analysis (SSLRA) is proposed. First, we propose a new low-rank model for hyperspectral images (HSIs). In the new model, HSI is decomposed into smooth and sparse unknown features which live in an unknown orthogonal subspace. Then, the sparse and smooth features are simultaneously estimated using a non-convex constrained penalized cost function. In the experiments' SSLRA is applied on a real HSI and the smooth features extracted are used for the HSI classification. The results confirm improvements in classification accuracies compared to state-of-the-art FE methods. Behnood Rasti, Magnus O. Ulfarsson, Pedram Ghamisi |
IGARSS | 3 |
| 2018 | IMG2DSM: Height Simulation From Single Imagery Using Conditional Generative Adversarial NetabstractThis letter proposes a groundbreaking approach in the remote-sensing community to simulating the digital surface model (DSM) from a single optical image. This novel technique uses conditional generative adversarial networks whose architecture is based on an encoder-decoder network with skip connections (generator) and penalizing structures at the scale of image patches (discriminator). The network is trained on scenes where both the DSM and optical data are available to establish an image-to-DSM translation rule. The trained network is then utilized to simulate elevation information on target scenes where no corresponding elevation information exists. The capability of the approach is evaluated both visually (in terms of photographic interpretation) and quantitatively (in terms of reconstruction errors and classification accuracies) on subdecimeter spatial resolution data sets captured over Vaihingen, Potsdam, and Stockholm. The results confirm the promising performance of the proposed framework. Pedram Ghamisi, Naoto Yokoya |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2018 | LiDAR Data Classification Using Morphological Profiles and Convolutional Neural NetworksabstractIn recent years, deep learning-based methods, especially convolutional neural networks (CNNs), have shown their capabilities in remote sensing data processing. The efficacy of light detection and ranging (LiDAR) has been already proven in a wide variety of research areas. Most of the existing methods do not extract the informative features from LiDAR-derived rasterized digital surface models (LiDAR-DSM) data in a deep manner. In order to utilize the advantages of deep models for the classification of LiDAR-derived features, deep CNN is proposed here to hierarchically extract the robust and discriminant features of the input data. Moreover, morphological profiles and multiattribute profiles (MAPs) are investigated to enrich the inputs of the CNN and further to improve the ultimate classification performance. Furthermore, a new activation function, sigmoid-weighted linear units (SiLUs), is introduced. The proposed frameworks are tested on two LiDAR-DSMs (i.e., Bayview Park and Houston data sets). The MAP-CNNs with SiLU outperform original CNNs by 6.62% and 6.88% in terms of overall accuracy on Bayview Park and Houston data sets, respectively, when the number of training samples of each class is 40. Aili Wang 0001, Xin He 0004, Pedram Ghamisi, Yushi Chen 0002 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2018 | Deformable Convolutional Neural Networks for Hyperspectral Image ClassificationabstractConvolutional neural networks (CNNs) have recently been demonstrated to be a powerful tool for hyperspectral image (HSI) classification, since they adopt deep convolutional layers whose kernels can effectively extract high-level spatial-spectral features. However, sampling locations of traditional convolutional kernels are fixed and cannot be changed according to complex spatial structures in HSIs. In addition, the typical pooling layers (e.g., average or maximum operations) in CNNs are also fixed and cannot be learned for feature downsampling in an adaptive manner. In this letter, a novel deformable CNN-based HSI classification method is proposed, which is called deformable HSI classification networks (DHCNet). The proposed network, DHCNet, introduces the deformable convolutional sampling locations, whose size and shape can be adaptively adjusted according to HSIs' complex spatial contexts. Specifically, to create the deformable sampling locations, 2-D offsets are first calculated for each pixel of input images. The sampling locations of each pixel with calculated offsets can cover the locations of other neighboring pixels with similar characteristics. With the deformable sampling locations, deformable feature images are then created by compressing neighboring similar structural information of each pixel into fixed grids. Therefore, applying the regular convolutions on the deformable feature images can reflect complex structures more effectively. Moreover, instead of adopting the pooling layers, the strided convolution is further introduced on the feature images, which can be learned for feature downsampling according to spatial contexts. Experimental results on two real HSI data sets demonstrate that DHCNet can obtain better classification performance than can several well-known classification methods. Jian Zhu 0006, Leyuan Fang, Pedram Ghamisi |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2018 | Extinction Profiles Fusion for Hyperspectral Images ClassificationabstractAn extinction profile (EP) is an effective spatial-spectral feature extraction method for hyperspectral images (HSIs), which has recently drawn much attention. However, the existing methods utilize the EPs in a stacking way, which is hard to fully explore the information in EPs for HSI classification. In this paper, a novel fusion framework termed EPs-fusion (EPs-F) is proposed to exploit the information within and among EPs for HSI classification. In general, EPs-F includes the following two stages. In the first stage, by extracting the EPs from three independent components of an HSI, three complementary groups of EPs can be constructed. For each EP, an adaptive superpixel-based composite kernel strategy is proposed to explore the spatial information within an EP. The weights to create the composite kernel and the number of superpixels are automatically determined based on the spatial information of each EP. In the second stage, since the different EPs contain highly complementary information, a simple yet effective decision fusion method is further applied to obtain the final classification result. Experiments on three real HSI data sets verify the qualitative and quantitative superiority of the proposed EPs-F method over several state-of-the-art HSI classifiers. Leyuan Fang, Nanjun He, Shutao Li 0001, Pedram Ghamisi, Jón Atli Benediktsson |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2018 | Unsupervised Spectral-Spatial Feature Learning via Deep Residual Conv-Deconv Network for Hyperspectral Image ClassificationabstractSupervised approaches classify input data using a set of representative samples for each class, known as training samples. The collection of such samples is expensive and time demanding. Hence, unsupervised feature learning, which has a quick access to arbitrary amounts of unlabeled data, is conceptually of high interest. In this paper, we propose a novel network architecture, fully Conv-Deconv network, for unsupervised spectral-spatial feature learning of hyperspectral images, which is able to be trained in an end-to-end manner. Specifically, our network is based on the so-called encoder-decoder paradigm, i.e., the input 3-D hyperspectral patch is first transformed into a typically lower dimensional space via a convolutional subnetwork (encoder), and then expanded to reproduce the initial data by a deconvolutional subnetwork (decoder). However, during the experiment, we found that such a network is not easy to be optimized. To address this problem, we refine the proposed network architecture by incorporating: 1) residual learning and 2) a new unpooling operation that can use memorized max-pooling indexes. Moreover, to understand the “black box,” we make an in-depth study of the learned feature maps in the experimental analysis. A very interesting discovery is that some specific “neurons” in the first residual block of the proposed network own good description power for semantic visual patterns in the object level, which provide an opportunity to achieve “free” object detection. This paper, for the first time in the remote sensing community, proposes an end-to-end fully Conv-Deconv network for unsupervised spectral-spatial feature learning. Moreover, this paper also introduces an in-depth investigation of learned features. Experimental results on two widely used hyperspectral data, Indian Pines and Pavia University, demonstrate competitive performance obtained by the proposed methodology compared with other studied approaches. Lichao Mou, Pedram Ghamisi, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2018 | Corrections to "Deep Recurrent Neural Networks for Hyperspectral Image Classification"abstractHere, we correct some errors caused by a programming bug (a data type error) in overall accuracies (OAs) reported in[1]. The corrected OAs are underlined and shown in bold inTables I–III. Lichao Mou, Pedram Ghamisi, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2018 | Random Forest Ensembles and Extended Multiextinction Profiles for Hyperspectral Image ClassificationabstractClassification techniques for hyperspectral images based on random forest (RF) ensembles and extended multiextinction profiles (EMEPs) are proposed as a means of improving performance. To this end, five strategies - bagging, boosting, random subspace, rotation-based, and boosted rotation-based - are used to construct the RF ensembles. EPs, which are based on an extrema-oriented connected filtering technique, are applied to the images associated with the first informative components extracted by independent component analysis, leading to a set of EMEPs. The effectiveness of the proposed method is investigated on two benchmark hyperspectral images: the University of Pavia and Indian Pines. Comparative experimental evaluations reveal the superior performance of the proposed methods, especially those employing rotation-based and boosted rotation-based approaches. An additional advantage is that the CPU processing time is acceptable. Junshi Xia, Pedram Ghamisi, Naoto Yokoya, Akira Iwasaki |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2018 | Generative Adversarial Networks for Hyperspectral Image ClassificationabstractA generative adversarial network (GAN) usually contains a generative network and a discriminative network in competition with each other. The GAN has shown its capability in a variety of applications. In this paper, the usefulness and effectiveness of GAN for classification of hyperspectral images (HSIs) are explored for the first time. In the proposed GAN, a convolutional neural network (CNN) is designed to discriminate the inputs and another CNN is used to generate so-called fake inputs. The aforementioned CNNs are trained together: the generative CNN tries to generate fake inputs that are as real as possible, and the discriminative CNN tries to classify the real and fake inputs. This kind of adversarial training improves the generalization capability of the discriminative CNN, which is really important when the training samples are limited. Specifically, we propose two schemes: 1) a well-designed 1D-GAN as a spectral classifier and 2) a robust 3D-GAN as a spectral-spatial classifier. Furthermore, the generated adversarial samples are used with real training samples to fine-tune the discriminative CNN, which improves the final classification performance. The proposed classifiers are carried out on three widely used hyperspectral data sets: Salinas, Indiana Pines, and Kennedy Space Center. The obtained results reveal that the proposed models provide competitive results compared to the state-of-the-art methods. In addition, the proposed GANs open new opportunities in the remote sensing community for the challenging task of HSI classification and also reveal the huge potential of GAN-based methods for the analysis of such complex and inherently nonlinear data. Lin Zhu 0013, Yushi Chen 0002, Pedram Ghamisi, Jón Atli Benediktsson |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2017 | Multiple composite kernel learning for hyperspectral image classificationabstractIn this work, we develop a new framework to combine ensemble learning and composite kernel learning for hyperspectral image classification. We refer it as the multiple composite kernel learning, which is based on an iterative architecture. More specifically, in each iteration, we use the rotation-based ensemble to create rotation matrix, which is used to generate rotated features for both spectral and spatial information (e.g., extinction profiles). Then, the new spectral and spatial features are integrated into the composite kernels based on support vector machines classifier. Different rotation matrices will lead to obtaining various newly spectral and spatial characteristics, thereby they further increase the diversity and the classification performance. Experimental results on Indian Pines benchmark hyperspectral dataset demonstrate the excellent performance of the proposed method. Peijun Du, Junshi Xia, Pedram Ghamisi, Akira Iwasaki, Jón Atli Benediktsson |
IGARSS | 3 |
| 2017 | Feature fusion of hyperspectral and lidar data using extinction profiles and total variationabstractTo improve the classification of hyperspectral images, this paper proposes an approach for multi-sensor data fusion of LiDAR and hyperspectral data using extinction profiles and Orthogonal Total Variation Component Analysis (OTVCA). Results on the benchmark Houston data indicate the superior performance of the proposed approach compared to other approaches used in the experiments based on classification accuracies. Pedram Ghamisi, Behnood Rasti, Xiao Xiang Zhu 0001 |
IGARSS | 1 |
| 2017 | Hyperspectral images classification by fusing extinction profiles featureabstractExtinction profile (EP) is an effective feature extraction method which can well preserve the geometrical characteristics of a hyperspectral image (HSI) and by extracting the EP from first three independent components (ICs) of an HSI, three correlated and complementary groups of EP features can be constructed. In this paper, an EPs fusion (EPs-F) strategy is proposed for HSI classification by exploring spatial-spectral information within and among three EP features. In general, the EPs-F method includes two stages. In the first stage, within each EP feature, a superpixel-based composite kernel strategy is proposed to adaptively fuse the spatial information of EP and the spectral feature of HSI. Then, the obtained adaptive composite kernel is used to create a classification map for each EP. In the second stage, decision fusion is further applied on different classification maps to create the final classification result. Experiments on two real HSIs verify the effectiveness of the proposed EPs-F algorithm. Nanjun He, Leyuan Fang, Shutao Li 0001, Pedram Ghamisi, Jón Atli Benediktsson |
IGARSS | 4 |
| 2017 | Evaluation of polsar similarity measures with spectral clusteringabstractPolarimetric Synthetic Aperture Radar (PolSAR) is a valuable remote sensing data source. It is usually challenging to interpret PolSAR data, especially in urban areas, and hense, spatial clustering comes as a powerful tool for the application of PolSAR data. In data clustering, similarity measurement indexes are of great importance. By far, there are quite some similarity measures of PolSAR data. However, to our knowledge, there has no practical and systematic evaluation of the performances of these measures. In this paper, we evaluate seven different similarity measurements of PolSAR data in the context of clustering using the conventional spectral clustering algorithm. Jingliang Hu, Yuanyuan Wang 0002, Pedram Ghamisi, Xiao Xiang Zhu 0001 |
IGARSS | 3 |
| 2017 | Fully conv-deconv network for unsupervised spectral-spatial feature extraction of hyperspectral imagery via residual learningabstractSupervised approaches classify input data using a set of representative samples for each class, known as training samples. The collection of such samples are expensive and time-demanding. Hence, unsupervised feature learning, which has a quick access to arbitrary amount of unlabeled data, is conceptually of high interest. In this paper, we propose a novel network architecture, fully Conv-Deconv network with residual learning, for unsupervised spectral-spatial feature learning of hyperspectral images, which is able to be trained in an end-to-end manner. Specifically, our network is based on the so-called encoder-decoder paradigm, i.e., the input 3D hyperspectral patch is first transformed into a typically lower-dimensional space via a convolutional sub-network (encoder), and then expanded to reproduce the initial data by a deconvolutional sub-network (decoder). Experimental results on the Pavia University hyperspectral data set demonstrate competitive performance obtained by the proposed methodology compared to other studied approaches. Lichao Mou, Pedram Ghamisi, Xiao Xiang Zhu 0001 |
IGARSS | 2 |
| 2017 | Multimodal, multitemporal, and multisource global data fusion for local climate zones classification based on ensemble learningabstractThis paper presents a new methodology for classification of local climate zones based on ensemble learning techniques. Landsat-8 data and open street map data are used to extract spectral-spatial features, including spectral reflectance, spectral indexes, and morphological profiles fed to subsequent classification methods as inputs. Canonical correlation forests and rotation forests are used for the classification step. The final classification map is generated by majority voting on different classification maps obtained by the two classifiers using multiple training subsets. The proposed method achieved an overall accuracy of 74.94% and a kappa coefficient of 0.71 in the 2017 IEEE GRSS Data Fusion Contest. Naoto Yokoya, Pedram Ghamisi, Junshi Xia |
IGARSS | 2 |
| 2017 | Deep Fusion of Remote Sensing Data for Accurate ClassificationabstractThe multisensory fusion of remote sensing data has obtained a great attention in recent years. In this letter, we propose a new feature fusion framework based on deep neural networks (DNNs). The proposed framework employs deep convolutional neural networks (CNNs) to effectively extract features of multi-/hyperspectral and light detection and ranging data. Then, a fully connected DNN is designed to fuse the heterogeneous features obtained by the previous CNNs. Through the aforementioned deep networks, one can extract the discriminant and invariant features of remote sensing data, which are useful for further processing. At last, logistic regression is used to produce the final classification results. Dropout and batch normalization strategies are adopted in the deep fusion framework to further improve classification accuracy. The obtained results reveal that the proposed deep fusion model provides competitive results in terms of classification accuracy. Furthermore, the proposed deep learning idea opens a new window for future remote sensing data fusion. Yushi Chen 0002, Pedram Ghamisi, Xiuping Jia, Yanfeng Gu |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2017 | Hyperspectral Images Classification With Gabor Filtering and Convolutional Neural NetworkabstractRecently, the capability of deep learning-based approaches, especially deep convolutional neural networks (CNNs), has been investigated for hyperspectral remote sensing feature extraction (FE) and classification. Due to the large number of learnable parameters in convolutional filters, lots of training samples are needed in deep CNNs to avoid the overfitting problem. On the other hand, Gabor filtering can effectively extract spatial information including edges and textures, which may reduce the FE burden of the CNNs. In this letter, in order to make the most of deep CNN and Gabor filtering, a new strategy, which combines Gabor filters with convolutional filters, is proposed for hyperspectral image classification to mitigate the problem of overfitting. The obtained results reveal that the proposed model provides competitive results in terms of classification accuracy, especially when only a limited number of training samples are available. Yushi Chen 0002, Lin Zhu 0013, Pedram Ghamisi, Xiuping Jia |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2017 | LiDAR Data Classification Using Extinction Profiles and a Composite Kernel Support Vector MachineabstractThis letter proposes a novel framework for the classification of light detection and ranging (LiDAR)-derived features. In this context, several features are extracted directly from the LiDAR point cloud data using aggregated local point neighborhoods, including laser echo ratio, variance of point elevation, plane fitting residuals, and echo intensity. Additionally, the LiDAR digital surface model (DSM) is input to our classification. Thus, both the LiDAR raster DSM and also rich geometric and also backscatter 3-D point cloud information aggregated to images are considered in our workflow. These extracted features are characterized as base images to be fed to extinction profiles to model spatial and contextual information. Then, a composite kernel support vector machine is investigated to efficiently integrate the elevation and spatial information suitable for the LiDAR data. Results indicate that the proposed method can obtain high classification accuracy using LiDAR data alone (e.g., more than 86% overall accuracy on the benchmark Houston LiDAR data using the standard set of training and test samples on all 15 classes) in a short CPU processing time. Pedram Ghamisi, Bernhard Höfle |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2017 | Automatic Hyperspectral Image Restoration Using Sparse and Low-Rank ModelingabstractHyperspectral restoration is a preprocessing step for hyperspectral imagery. In this letter, we propose a parameter-free method for the restoration of hyperspectral images (HSIs) called HyRes. The restoration method is based on a sparse low-rank model that uses the ℓ1penalized least squares for estimating the unknown signal. The Stein's unbiased risk estimator is exploited to select all the parameters of the model yielding a fully automatic (parameter free) technique. Experimental results confirm that HyRes outperforms the state-of-the-art techniques in terms of signal-to-noise ratio, structural similarity index, and spectral angle distance for a simulated data set and in terms of noise-level estimation for the real data sets used in this letter. In the experiments, it was noted that HyRes is computationally less expensive compared with competitive techniques. Therefore, HyRes can be used as a reliable automatic preprocessing step for further analysis of HSIs. Behnood Rasti, Magnus O. Ulfarsson, Pedram Ghamisi |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2017 | Deep Recurrent Neural Networks for Hyperspectral Image ClassificationabstractIn recent years, vector-based machine learning algorithms, such as random forests, support vector machines, and 1-D convolutional neural networks, have shown promising results in hyperspectral image classification. Such methodologies, nevertheless, can lead to information loss in representing hyperspectral pixels, which intrinsically have a sequence-based data structure. A recurrent neural network (RNN), an important branch of the deep learning family, is mainly designed to handle sequential data. Can sequence-based RNN be an effective method of hyperspectral image classification? In this paper, we propose a novel RNN model that can effectively analyze hyperspectral pixels as sequential data and then determine information categories via network reasoning. As far as we know, this is the first time that an RNN framework has been proposed for hyperspectral image classification. Specifically, our RNN makes use of a newly proposed activation function, parametric rectified tanh (PRetanh), for hyperspectral sequential data analysis instead of the popular tanh or rectified linear unit. The proposed activation function makes it possible to use fairly high learning rates without the risk of divergence during the training procedure. Moreover, a modified gated recurrent unit, which uses PRetanh for hidden representation, is adopted to construct the recurrent layer in our network to efficiently process hyperspectral data and reduce the total number of parameters. Experimental results on three airborne hyperspectral images suggest competitive performance in the proposed mode. In addition, the proposed network architecture opens a new window for future research, showcasing the huge potential of deep recurrent networks for hyperspectral data analysis. Lichao Mou, Pedram Ghamisi, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2017 | Hyperspectral and LiDAR Fusion Using Extinction Profiles and Total Variation Component AnalysisabstractThe classification accuracy of remote sensing data can be increased by integrating ancillary data provided by multisource acquisition of the same scene. We propose to merge the spectral and spatial content of hyperspectral images (HSIs) with elevation information from light detection and ranging (LiDAR) measurements. In this paper, we propose to fuse the data sets using orthogonal total variation component analysis (OTVCA). Extinction profiles are used to automatically extract spatial and elevation information from HSI and rasterized LiDAR features. The extracted spatial and elevation information is then fused with spectral information using the OTVCA-based feature fusion method to produce the final classification map. The extracted features have high dimension, and therefore OTVCA estimates the fused features in a lower dimensional space. OTVCA also promotes piece-wise smoothness while maintaining the spatial structures. Both attributes are important to provide homogeneous regions in the final classification maps. We benchmark the proposed approach (OTVCA-fusion) with an urban data set captured over an urban area in Houston/USA and a rural region acquired in Trento/Italy. In the experiments, OTVCA-fusion is evaluated using random forest and support vector machine classifiers. Our experiments demonstrate the ability of OTVCA-fusion to produce accurate classification maps while using fewer features compared with other approaches investigated in this paper. Behnood Rasti, Pedram Ghamisi, Richard Gloaguen |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2017 | Fusion of Hyperspectral and LiDAR Data Using Sparse and Low-Rank Component AnalysisabstractThe availability of diverse data captured over the same region makes it possible to develop multisensor data fusion techniques to further improve the discrimination ability of classifiers. In this paper, a new sparse and low-rank technique is proposed for the fusion of hyperspectral and light detection and ranging (LiDAR)-derived features. The proposed fusion technique consists of two main steps. First, extinction profiles are used to extract spatial and elevation information from hyperspectral and LiDAR data, respectively. Then, the sparse and low-rank technique is utilized to estimate the low-rank fused features from the extracted ones that are eventually used to produce a final classification map. The proposed approach is evaluated over an urban data set captured over Houston, USA, and a rural one captured over Trento, Italy. Experimental results confirm that the proposed fusion technique outperforms the other techniques used in the experiments based on the classification accuracies obtained by random forest and support vector machine classifiers. Moreover, the proposed approach can effectively classify joint LiDAR and hyperspectral data in an ill-posed situation when only a limited number of training samples are available. Behnood Rasti, Pedram Ghamisi, Javier Plaza, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2016 | Deep fusion of hyperspectral and LiDAR data for thematic classificationabstractRecently, the fusion of hyperspectral and light detection and ranging (LiDAR) data has obtained a great attention in the remote sensing community. In this paper, we propose a new feature fusion framework using deep neural network (DNN). The proposed framework employs a novel 3D convolutional neural network (CNN) to extract the spectral-spatial features of hyperspectral data, a deep 2D CNN to extract the elevation features of LiDAR data, and then a fully connected deep neural network to fuse the extracted features in the previous CNNs. Through the aforementioned three deep networks, one can extract the discriminant and invariant features of hyperspectral and LiDAR data. At last, logistic regression is used to produce the final classification results. The experimental results reveal that the proposed deep fusion model provides competitive results. Furthermore, the proposed deep fusion idea opens a new window for future research. Pedram Ghamisi, Chunyu Shi, Yanfeng Gu |
IGARSS | 3 |
| 2016 | Extinction profiles: A novel approach for the analysis of remote sensing dataabstractThis paper presents a novel approach named extinction profiles to model the spatial information of remote sensing images. Then, the output of the extinction profile is fed to a grid-search random forest classification method. Results indicate that the proposed approach can effectively extract spatial information from remote sensing gray scale images and provide high classification accuracies in an automatic way. Pedram Ghamisi, Roberto Souza 0001, Letícia Rittner, Jón Atli Benediktsson, Roberto A. Lotufo, Xiao Xiang Zhu 0001 |
IGARSS | 1 |
| 2016 | A Self-Improving Convolution Neural Network for the Classification of Hyperspectral DataabstractIn this letter, a self-improving convolutional neural network (CNN) based method is proposed for the classification of hyperspectral data. This approach solves the so-called curse of dimensionality and the lack of available training samples by iteratively selecting the most informative bands suitable for the designed network via fractional order Darwinian particle swarm optimization. The selected bands are then fed to the classification system to produce the final classification map. Experimental results have been conducted with two well-known hyperspectral data sets: Indian Pines and Pavia University. Results indicate that the proposed approach significantly improves a CNN-based classification method in terms of classification accuracy. In addition, this letter uses the concept of dither for the first time in the remote sensing community to tackle overfitting. Pedram Ghamisi, Yushi Chen 0002, Xiao Xiang Zhu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2016 | Hyperspectral Data Classification Using Extended Extinction ProfilesabstractThis letter proposes a new approach for the spectral-spatial classification of hyperspectral images, which is based on a novel extrema-oriented connected filtering technique, entitled as extended extinction profiles. The proposed approach progressively simplifies the first informative features extracted from hyperspectral data considering different attributes. Then, the classification approach is applied on two well-known hyperspectral data sets, i.e., Pavia University and Indian Pines, and compared with one of the most powerful filtering approaches in the literature, i.e., extended attribute profiles. Results indicate that the proposed approach is able to efficiently extract spatial information for the classification of hyperspectral images automatically and swiftly. In addition, an array-based node-oriented max-tree representation was carried out to efficiently implement the proposed approach. Pedram Ghamisi, Roberto Souza 0001, Jón Atli Benediktsson, Letícia Rittner, Roberto A. Lotufo, Xiao Xiang Zhu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2016 | Deep Feature Extraction and Classification of Hyperspectral Images Based on Convolutional Neural NetworksabstractDue to the advantages of deep learning, in this paper, a regularized deep feature extraction (FE) method is presented for hyperspectral image (HSI) classification using a convolutional neural network (CNN). The proposed approach employs several convolutional and pooling layers to extract deep features from HSIs, which are nonlinear, discriminant, and invariant. These features are useful for image classification and target detection. Furthermore, in order to address the common issue of imbalance between high dimensionality and limited availability of training samples for the classification of HSI, a few strategies such as L2 regularization and dropout are investigated to avoid overfitting in class data modeling. More importantly, we propose a 3-D CNN-based FE model with combined regularization to extract effective spectral-spatial features of hyperspectral imagery. Finally, in order to further improve the performance, a virtual sample enhanced method is proposed. The proposed approaches are carried out on three widely used hyperspectral data sets: Indian Pines, University of Pavia, and Kennedy Space Center. The obtained results reveal that the proposed models with sparse constraints provide competitive results to state-of-the-art methods. In addition, the proposed deep FE opens a new window for further research. Yushi Chen 0002, Hanlu Jiang, Xiuping Jia, Pedram Ghamisi |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2016 | Extinction Profiles for the Classification of Remote Sensing DataabstractWith respect to recent advances in remote sensing technologies, the spatial resolution of airborne and spaceborne sensors is getting finer, which enables us to precisely analyze even small objects on the Earth. This fact has made the research area of developing efficient approaches to extract spatial and contextual information highly active. Among the existing approaches, morphological profile and attribute profile (AP) have gained great attention due to their ability to classify remote sensing data. This paper proposes a novel approach that makes it possible to precisely extract spatial and contextual information from remote sensing images. The proposed approach is based on extinction filters, which are used here for the first time in the remote sensing community. Then, the approach is carried out on two well-known high-resolution panchromatic data sets captured over Rome, Italy, and Reykjavik, Iceland. In order to prove the capabilities of the proposed approach, the obtained results are compared with the results from one of the strongest approaches in the literature, i.e., APs, using different points of view such as classification accuracies, simplification rate, and complexity analysis. Results indicate that the proposed approach can significantly outperform its alternative in terms of classification accuracies. In addition, based on our implementation, profiles can be generated in a very short processing time. It should be noted that the proposed approach is fully automatic. Pedram Ghamisi, Roberto Souza 0001, Jón Atli Benediktsson, Xiao Xiang Zhu 0001, Letícia Rittner, Roberto A. Lotufo |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2015 | An advanced classifier for the joint use of LiDAR and hyperspectral data: Case study in Queensland, AustraliaabstractWith respect to the exponential increase in the number of available remote sensors in recent years, the possibility of having different types of data captured over the same scene, has resulted in many research works related to the joint use of passive and active sensors for the accurate classification of different materials. However, until now, there is a small number of research works related to the integration of highly valuable information obtained from the joint use of LiDAR and hyperspectral data. This paper proposes an efficient classification approach in terms of accuracies and demanded CPU processing time for integrating big data sets (e.g., LiDAR and hyperspectral) to provide land cover mapping capabilities at a range of spatial scales. In addition, the proposed approach is fully automatic and is able to efficiently handle big data containing a huge number of features with very limited number of training samples in few seconds. Pedram Ghamisi, Gabriele Cavallaro, Jón Atli Benediktsson, Stuart R. Phinn, Nicola Falco |
IGARSS | 1 |
| 2015 | Feature Selection Based on Hybridization of Genetic Algorithm and Particle Swarm OptimizationabstractA new feature selection approach that is based on the integration of a genetic algorithm and particle swarm optimization is proposed. The overall accuracy of a support vector machine classifier on validation samples is used as a fitness value. The new approach is carried out on the well-known Indian Pines hyperspectral data set. Results confirm that the new approach is able to automatically select the most informative features in terms of classification accuracy within an acceptable CPU processing time without requiring the number of desired features to be set a priori by users. Furthermore, the usefulness of the proposed method is also tested for road detection. Results confirm that the proposed method is capable of discriminating between road and background pixels and performs better than the other approaches used for comparison in terms of performance metrics. Pedram Ghamisi, Jón Atli Benediktsson |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2015 | A Novel Feature Selection Approach Based on FODPSO and SVMabstractA novel feature selection approach is proposed to address the curse of dimensionality and reduce the redundancy of hyperspectral data. The proposed approach is based on a new binary optimization method inspired by fractional-order Darwinian particle swarm optimization (FODPSO). The overall accuracy (OA) of a support vector machine (SVM) classifier on validation samples is used as fitness values in order to evaluate the informativity of different groups of bands. In order to show the capability of the proposed method, two different applications are considered. In the first application, the proposed feature selection approach is directly carried out on the input hyperspectral data. The most informative bands selected from this step are classified by the SVM. In the second application, the main shortcoming of using attribute profiles (APs) for spectral-spatial classification is addressed. In this case, a stacked vector of the input data and an AP with all widely used attributes are created. Then, the proposed feature selection approach automatically chooses the most informative features from the stacked vector. Experimental results successfully confirm that the proposed feature selection technique works better in terms of classification accuracies and CPU processing time than other studied methods without requiring the number of desired features to be set a priori by users. Pedram Ghamisi, Micael S. Couceiro, Jón Atli Benediktsson |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2015 | A Survey on Spectral-Spatial Classification Techniques Based on Attribute ProfilesabstractJust over a decade has passed since the concept of morphological profile was defined for the analysis of remote sensing images. Since then, the morphological profile has largely proved to be a powerful tool able to model spatial information (e.g., contextual relations) of the image. However, due to the shortcomings of using the morphological profiles, many variants, extensions, and refinements of its definition have appeared stating that the morphological profile is still under continuous development. In this case, recently introduced theoretically sound attribute profiles (APs) can be considered as a generalization of the morphological profile, which is a powerful tool to model spatial information existing in the scene. Although the concept of the AP has been introduced in remote sensing only recently, an extensive literature on its use in different applications and on different types of data has appeared. To that end, the great amount of contributions in the literature that address the application of the AP to many tasks (e.g., classification, object detection, segmentation, change detection, etc.) and to different types of images (e.g., panchromatic, multispectral, and hyperspectral) proves how the AP is an effective and modern tool. The main objective of this survey paper is to recall the concept of the APs along with all its modifications and generalizations with special emphasis on remote sensing image classification and summarize the important aspects of its efficient utilization while also listing potential future works. Pedram Ghamisi, Mauro Dalla Mura, Jón Atli Benediktsson |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2014 | Fusion of hyperspectral and LiDAR data in classification of urban areasabstractIn this paper, the fusion of hyperspectral and Li-DAR data is taken into account in order to develop a new classification framework for the accurate analysis of urban areas. In this method, an attribute profile is considered in order to model the spatial information of LiDAR and hyper-spectral data. In parallel, in order to reduce the redundancy of the hyperspectral data and address the so-called curse of dimensionality, a supervised feature extraction technique is used. Then, the new features obtained by the attribute profile and the supervised feature extraction technique are concatenated into a stacked vector. The final classification map is achieved by using a Random Forest classifier. Results infer that the proposed method can provide very good results in terms of classification accuracy and CPU processing time in an automatic manner. Pedram Ghamisi, Jón Atli Benediktsson, Stuart R. Phinn |
IGARSS | 1 |
| 2014 | Integration of Segmentation Techniques for Classification of Hyperspectral ImagesabstractA new spectral–spatial method for classification of hyperspectral images is introduced. The proposed approach is based on two segmentation methods, fractional-order Darwinian particle swarm optimization and mean shift segmentation. The output of these two methods is classified by support vector machines. Experimental results indicate that the integration of the two segmentation methods can overcome the drawbacks of each other and increase the overall accuracy in classification. Pedram Ghamisi, Micael S. Couceiro, Mathieu Fauvel, Jón Atli Benediktsson |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2014 | Automatic Spectral-Spatial Classification Framework Based on Attribute Profiles and Supervised Feature ExtractionabstractA robust framework for the classification of hyperspectral images which takes into account both spectral and spatial information is proposed. The extended multivariate attribute profile (EMAP) is used for extracting spatial information. Moreover, for solving the so-called curse of dimensionality, supervised feature extraction is carried out on both the original hyperspectral data and the output of the EMAP. After performing the dimensionality reduction, two output vectors of the original data and attributes are concatenated into one stacked vector. The final classification map is achieved by using a random-forest classifier. The main difficulties of using an EMAP is to initialize the attribute parameters. Therefore, a fully automatic scheme of the proposed method is introduced to overcome the shortcomings of using EMAP. The proposed method is tested on two widely known data sets. Experimental results confirm that the proposed method provides an accurate classification map in an acceptable CPU processing time. Pedram Ghamisi, Jón Atli Benediktsson, Johannes R. Sveinsson |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2014 | Spectral-Spatial Classification of Hyperspectral Images Based on Hidden Markov Random FieldsabstractHyperspectral remote sensing technology allows one to acquire a sequence of possibly hundreds of contiguous spectral images from ultraviolet to infrared. Conventional spectral classifiers treat hyperspectral images as a list of spectral measurements and do not consider spatial dependences, which leads to a dramatic decrease in classification accuracies. In this paper, a new automatic framework for the classification of hyperspectral images is proposed. The new method is based on combining hidden Markov random field segmentation with support vector machine (SVM) classifier. In order to preserve edges in the final classification map, a gradient step is taken into account. Experiments confirm that the new spectral and spatial classification approach is able to improve results significantly in terms of classification accuracies compared to the standard SVM method and also outperforms other studied methods. Pedram Ghamisi, Jón Atli Benediktsson, Magnus O. Ulfarsson |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2014 | Multilevel Image Segmentation Based on Fractional-Order Darwinian Particle Swarm OptimizationabstractHyperspectral remote sensing images contain hundreds of data channels. Due to the high dimensionality of the hyperspectral data, it is difficult to design accurate and efficient image segmentation algorithms for such imagery. In this paper, a new multilevel thresholding method is introduced for the segmentation of hyperspectral and multispectral images. The new method is based on fractional-order Darwinian particle swarm optimization (FODPSO) which exploits the many swarms of test solutions that may exist at any time. In addition, the concept of fractional derivative is used to control the convergence rate of particles. In this paper, the so-called Otsu problem is solved for each channel of the multispectral and hyperspectral data. Therefore, the problem of n-level thresholding is reduced to an optimization problem in order to search for the thresholds that maximize the between-class variance. Experimental results are favorable for the FODPSO when compared to other bioinspired methods for multilevel segmentation of multispectral and hyperspectral images. The FODPSO presents a statistically significant improvement in terms of both CPU time and fitness value, i.e., the approach is able to find the optimal set of thresholds with a larger between-class variance in less computational time than the other approaches. In addition, a new classification approach based on support vector machine (SVM) and FODPSO is introduced in this paper. Results confirm that the new segmentation method is able to improve upon results obtained with the standard SVM in terms of classification accuracies. Pedram Ghamisi, Micael S. Couceiro, Fernando M. L. Martins, Jón Atli Benediktsson |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2013 | The spectral-spatial classification of hyperspectral images based on Hidden Markov Random Field and its Expectation-MaximizationabstractIn this work, a new framework for accurate classification of hyperspectral images is proposed. The new method is based on Hidden Markov Random Field and its Expectation Maximization (HMRF-EM) and Support Vector Machine (SVM) classifier. In order to preserve edges in final map, the Sobel edge detector is used. Result confirms that the combination of the spectral and spatial information can significantly improve results compared to the standard SVM method. Pedram Ghamisi, Jón Atli Benediktsson, Magnus O. Ulfarsson |
IGARSS | 1 |
| 2013 | Spectral-spatial classification based on integrated segmentationabstractA new spectral-spatial method for the classification of hyperspectral images is introduced. The proposed approach is based on two segmentation methods, Fractional-Order Darwinian Particle Swarm Optimization and Mean Shift Segmentation and one clustering method, K-means. In parallel, the input data set is classified by Support Vector Machines (SVM). Furthermore, the result of the segmentation and clustering steps are combined with the result of SVM through majority voting within each object. The final classification map is made by using majority voting between three produced classification maps. Experimental results indicate that the proposed method can significantly improve SVM and other studied methods in terms of accuracies. Pedram Ghamisi, Micael S. Couceiro, Mathieu Fauvel, Jón Atli Benediktsson |
IGARSS | 1 |
| 2012 | Use of Darwinian Particle Swarm Optimization technique for the segmentation of Remote Sensing imagesabstractIn this work, a novel method for segmentation of Remote Sensing (RS) images based on the Darwinian Particle Swarm Optimization (DPSO) for determining the n-1 optimal n-level threshold on a given image is proposed. The efficiency of the proposed method is compared with the Particle Swarm Optimization (PSO) based segmentation method. Results show that DPSO-based image segmentation performs better than PSO-based method in a number of different measures. Pedram Ghamisi, Micael S. Couceiro, Nuno M. Fonseca Ferreira, Lalit Kumar 0002 |
IGARSS | 1 |
| 2012 | An efficient method for segmentation of images based on fractional calculus and natural selection
Pedram Ghamisi, Micael S. Couceiro, Jón Atli Benediktsson, Nuno M. Fonseca Ferreira |
Expert Syst. Appl. | 1 |