Xiao Xiang Zhu 0001

dblp:35/8954 · also Xiaoxiang Zhu 0001 · DBLP profile ↗
← Back
346ranked-venue papers
20as first author
193since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 306 · 20 first-author · 164 since 2021Artificial intelligence and machine learning · 34 · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 20 since 2021Databases, data management, data science and information retrieval · 5 · 4 since 2021
YearPublicationVenuePosition
2026 ProPL: Universal Semi-Supervised Ultrasound Image Segmentation via Prompt-Guided Pseudo-Labeling
abstract
Existing approaches for the problem of ultrasound image segmentation, whether supervised or semi-supervised, are typically specialized for specific anatomical structures or tasks, limiting their practical utility in clinical settings. In this paper, we pioneer the task of universal semi-supervised ultrasound image segmentation and propose ProPL, a framework that can handle multiple organs and segmentation tasks while leveraging both labeled and unlabeled data. At its core, ProPL employs a shared vision encoder coupled with prompt-guided dual decoders, enabling flexible task adaptation through a prompting-upon-decoding mechanism and reliable self-training via an uncertainty-driven pseudo-label calibration (UPLC) module. To facilitate research in this direction, we introduce a comprehensive ultrasound dataset spanning 5 organs and 8 segmentation tasks. Extensive experiments demonstrate that ProPL outperforms state-of-the-art methods across various metrics, establishing a new benchmark for universal ultrasound image segmentation.
Yaxiong Chen, Qicong Wang, Jingliang Hu, Yilei Shi, Shengwu Xiong 0001, Xiao Xiang Zhu 0001, Lichao Mou
AAAI7
2026 GEWDiff: Geometric Enhanced Wavelet-based Diffusion Model for Hyperspectral Image Super-resolution
abstract
Improving the quality of hyperspectral images (HSIs), such as through super-resolution, is a crucial research area. However, generative modeling for HSIs presents several challenges. Due to their high spectral dimensionality, HSIs are too memory-intensive for direct input into conventional diffusion models. Furthermore, general generative models lack an understanding of the topological and geometric structures of ground objects in remote sensing imagery. In addition, most diffusion models optimize loss functions at the noise level, leading to a non-intuitive convergence behavior and suboptimal generation quality for complex data. To address these challenges, we propose a Geometric Enhanced Wavelet-based Diffusion Model (GEWDiff), a novel framework for reconstructing hyperspectral images at 4-times super-resolution. A wavelet-based encoder-decoder is introduced that efficiently compresses HSIs into a latent space while preserving spectral-spatial information. To avoid distortion during generation, we incorporate a geometry-enhanced diffusion process that preserves the geometric features. Furthermore, a multi-level loss function was designed to guide the diffusion process, promoting stable convergence and improved reconstruction fidelity. Our model demonstrated state-of-the-art results across multiple dimensions, including fidelity, spectral accuracy, visual realism, and clarity.
Sirui Wang 0010, Natàlia Blasco Andreo, Xiao Xiang Zhu 0001
AAAI4
2026 A deep dive into OpenStreetMap research since its inception (2008-2024): contributors, topics, and future trends
abstract
OpenStreetMap (OSM) has transitioned from a pioneering volunteered geographic information project into a global, multi-disciplinary research nexus. This study presents a bibliometric and systematic analysis of the OSM research landscape, examining its development trajectory and key driving forces. By evaluating 1926 publications from the Web of Science (WoS) Core Collection and 782 State of the Map (SotM) presentations up to June 2024, we quantify publication growth, collaboration patterns, and thematic evolution. Results demonstrate simultaneous consolidation and diversification within the field. While a stable core of contributors continues to anchor OSM research, themes have shifted from initial concerns over data production and quality toward advanced analytical and applied uses. Comparative analysis of OSM-related research in WoS and SotM reveals distinct but complementary agendas between scholars and the OSM community. Building on these findings, we identify six emerging research directions and discuss how evolving partnerships among academia, the OSM community, and industry are poised to shape the future of OSM research. This study establishes a structured reference for understanding the state of OSM studies and offers strategic pathways for navigating its future trajectory.
Yao Sun 0005, Liqiu Meng, Andrés Camero, Stefan Auer, Xiao Xiang Zhu 0001
Int. J. Geogr. Inf. Sci.5
2025 MPTSNet: Integrating Multiscale Periodic Local Patterns and Global Dependencies for Multivariate Time Series Classification
abstract
Multivariate Time Series Classification (MTSC) is crucial in extensive practical applications, such as environmental monitoring, medical EEG analysis, and action recognition. Real-world time series datasets typically exhibit complex dynamics. To capture this complexity, RNN-based, CNN-based, Transformer-based, and hybrid models have been proposed. Unfortunately, current deep learning-based methods often neglect the simultaneous construction of local features and global dependencies at different time scales, lacking sufficient feature extraction capabilities to achieve satisfactory classification accuracy. To address these challenges, we propose a novel Multiscale Periodic Time Series Network (MPTSNet), which integrates multiscale local patterns and global correlations to fully exploit the inherent information in time series. Recognizing the multi-periodicity and complex variable correlations in time series, we use the Fourier transform to extract primary periods, enabling us to decompose data into multiscale periodic segments. Leveraging the inherent strengths of CNN and attention mechanism, we introduce the PeriodicBlock, which adaptively captures local patterns and global dependencies while offering enhanced interpretability through attention integration across different periodic scales. The experiments on UEA benchmark datasets demonstrate that the proposed MPTSNet outperforms 21 existing advanced baselines in the MTSC tasks.
Yang Mu, Muhammad Shahzad 0002, Xiao Xiang Zhu 0001
AAAI3
2025 Q-PART: Quasi-Periodic Adaptive Regression with Test-time Training for Pediatric Left Ventricular Ejection Fraction Regression
abstract
In this work, we address the challenge of adaptive pediatric Left Ventricular Ejection Fraction (LVEF) assessment. While Test-time Training (TTT) approaches show promise for this task, they suffer from two significant limitations. Existing TTT works are primarily designed for classification tasks rather than continuous value regression, and they lack mechanisms to handle the quasi-periodic nature of cardiac signals. To tackle these issues, we propose a novel Quasi-Periodic Adaptive Regression with Test-time Training (Q-PART) framework. In the training stage, the proposed Quasi-Period Network decomposes the echocardiogram into periodic and aperiodic components within latent space by combining parameterized helix trajectories with Neural Controlled Differential Equations. During inference, our framework further employs a variance minimization strategy across image augmentations that simulate common quality issues in echocardiogram acquisition, along with differential adaptation rates for periodic and aperiodic components. Theoretical analysis is provided to demonstrate that our variance minimization objective effectively bounds the regression error under mild conditions. Furthermore, extensive experiments across three pediatric age groups demonstrate that Q-PART not only significantly outperforms existing approaches in pediatric LVEF prediction, but also exhibits strong clinical screening capability with high mAUROC scores (up to 0.9747) and maintains gender-fair performance across all metrics, validating its robustness and practical utility in pediatric echocardiography analysis. The project can be found in Q-PART.
Jie Liu 0044, Tiexin Qin, Hui Liu 0036, Yilei Shi, Lichao Mou, Xiao Xiang Zhu 0001, Shiqi Wang 0001, Haoliang Li
CVPR6
2025 Parametric Point Cloud Completion for Polygonal Surface Reconstruction
abstract
Existing polygonal surface reconstruction methods heavily depend on input completeness and struggle with incomplete point clouds. We argue that while current point cloud completion techniques may recover missing points, they are not optimized for polygonal surface reconstruction, where the parametric representation of underlying surfaces remains overlooked. To address this gap, we introduce parametric completion, a novel paradigm for point cloud completion, which recovers parametric primitives instead of individual points to convey high-level geometric structures. Our presented approach, PaCo, enables high-quality polygonal surface reconstruction by leveraging plane proxies that encapsulate both plane parameters and inlier points, proving particularly effective in challenging scenarios with highly incomplete data. Comprehensive evaluations of our approach on the ABC dataset establish its effectiveness with superior performance and set a new standard for polygonal surface reconstruction from incomplete data. Project page: https://parametric-completion.github.io.
Zhaiyu Chen, Liangliang Nan, Xiao Xiang Zhu 0001
CVPR4
2025 On the Generalization of Representation Uncertainty in Earth Observation
Spyros Kondylatos, Nikolaos-Ioannis Bountos, Dimitrios Michail 0001, Xiao Xiang Zhu 0001, Gustau Camps-Valls, Ioannis Papoutsis
ICCV4
2025 Towards a Unified Copernicus Foundation Model for Earth Vision
abstract
Advances in Earth observation (EO) foundation models have unlocked the potential of big satellite data to learn generic representations from space, benefiting a wide range of downstream applications crucial to our planet. However, most existing efforts remain limited to fixed spectral sensors, focus solely on the Earth's surface, and overlook valuable metadata beyond imagery. In this work, we take a step towards next-generation EO foundation models with three key components: 1) Copernicus-Pretrain, a massive-scale pretraining dataset that integrates 18.7M aligned images from all major Copernicus Sentinel missions, spanning from the Earth's surface to its atmosphere; 2) Copernicus-FM, a unified foundation model capable of processing any spectral or non-spectral sensor modality using extended dynamic hypernetworks and flexible metadata encoding; and 3) Copernicus-Bench, a systematic evaluation benchmark with 15 hierarchical downstream tasks ranging from preprocessing to specialized applications for each Sentinel mission. Our dataset, model, and benchmark greatly improve the scalability, versatility, and multimodal adaptability of EO foundation models, while also creating new opportunities to connect EO, weather, and climate research. Codes, datasets and models are available at https://github.com/zhu-xlab/Copernicus-FM.
Yi Wang 0072, Zhitong Xiong, Chenying Liu 0001, Adam J. Stewart, Thomas Dujardin, Nikolaos-Ioannis Bountos, Angelos Zavras, Franziska Gerken, Ioannis Papoutsis, Laura Leal-Taixé, Xiao Xiang Zhu 0001
ICCV11
2025 Scale-Aware Contrastive Reverse Distillation for Unsupervised Medical Anomaly Detection
abstract
Unsupervised anomaly detection using deep learning has garnered significant research attention due to its broad applicability, particularly in medical imaging where labeled anomalous data are scarce. While earlier approaches leverage generative models like autoencoders and generative adversarial networks (GANs), they often fall short due to overgeneralization. Recent methods explore various strategies, including memory banks, normalizing flows, self-supervised learning, and knowledge distillation, to enhance discrimination. Among these, knowledge distillation, particularly reverse distillation, has shown promise. Following this paradigm, we propose a novel scale-aware contrastive reverse distillation model that addresses two key limitations of existing reverse distillation methods: insufficient feature discriminability and inability to handle anomaly scale variations. Specifically, we introduce a contrastive student-teacher learning approach to derive more discriminative representations by generating and exploring out-of-normal distributions. Further, we design a scale adaptation mechanism to softly weight contrastive distillation losses at different scales to account for the scale variation issue. Extensive experiments on benchmark datasets demonstrate state-of-the-art performance, validating the efficacy of the proposed method. The code will be made publicly available.
Yilei Shi, Jingliang Hu, Xiao Xiang Zhu 0001, Lichao Mou
ICLR4
2025 High-Order Progressive Trajectory Matching for Medical Image Dataset Distillation
Jinghao Bian, Jingyang Hou, Jingliang Hu, Yilei Shi, Weisheng Dong, Xiao Xiang Zhu 0001, Lichao Mou
MICCAI (14)7
2025 REOBench: Benchmarking Robustness of Earth Observation Foundation Models
abstract
Earth observation foundation models have shown strong generalization across multiple Earth observation tasks, but their robustness under real-world perturbations remains underexplored. To bridge this gap, we introduce REOBench, the first comprehensive benchmark for evaluating the robustness of Earth observation foundation models across six tasks and twelve types of image corruptions, including both appearance-based and geometric perturbations. To ensure realistic and fine-grained evaluation, our benchmark focuses on high-resolution optical remote sensing images, which are widely used in critical applications such as urban planning and disaster response. We conduct a systematic evaluation of a broad range of models trained using masked image modeling, contrastive learning, and vision-language pre-training paradigms. Our results reveal that (1) existing Earth observation foundation models experience significant performance degradation when exposed to input corruptions. (2) The severity of degradation varies across tasks, model architectures, backbone sizes, and types of corruption, with performance drop varying from less than 1% to over 25%. (3) Vision-language models show enhanced robustness, particularly in multimodal tasks. REOBench underscores the vulnerability of current Earth observation foundation models to real-world corruptions and provides actionable insights for developing more robust and reliable models.
Xiang Li 0001, Siwei Liu 0001, Zhitong Xiong, Chunbo Luo, Lu Liu 0001, Mykola Pechenizkiy, Xiao Xiang Zhu 0001, Tianjin Huang
NeurIPS9
2025 Learning Generalizable Shape Completion with SIM(3) Equivariance
abstract
3D shape completion methods typically assume scans are pre-aligned to a canonical frame. This leaks pose and scale cues that networks may exploit to memorize absolute positions rather than inferring intrinsic geometry. When such alignment is absent in real data, performance collapses. We argue that robust generalization demands architectural equivariance to the similarity group, SIM(3), so the model remains agnostic to pose and scale. Following this principle, we introduce the first SIM(3)-equivariant shape completion network, whose modular layers successively canonicalize features, reason over similarity-invariant geometry, and restore the original frame. Under a de-biased evaluation protocol that removes the hidden cues, our model outperforms both equivariant and augmentation baselines on the PCN benchmark. It also sets new cross-domain records on real driving and indoor scans, lowering minimal matching distance on KITTI by 17\% and Chamfer distance $\ell1$ on OmniObject3D by 14\%. Perhaps surprisingly, ours under the stricter protocol still outperforms competitors under their biased settings. These results establish full SIM(3) equivariance as an effective route to truly generalizable shape completion.
Zhaiyu Chen, Xiao Xiang Zhu 0001
NeurIPS3
2025 SmallMinesDS: A Multimodal Dataset for Mapping Artisanal and Small-Scale Gold Mines
abstract
The increasing demand for gold, coupled with persistently high market prices over the past decade, has driven a significant rise in small-scale gold production. The expansion of unregularized small-scale gold mines fuels environmental degradation and poses a risk to miners and mining communities. To promote sustainable mining practices, support reclamation initiatives and pave the way for understudying the impacts of mining on human and environmental resources, we presentSmallMinesDS, a dataset derived from multi-sensor satellite imagery covering five districts in southwestern Ghana in two time periods.SmallMinesDSprovides precise reference data for artisanal mining sites, enabling the development of machine learning models for timely, large-scale, and cost-effective monitoring. Notably, foundation models fine-tuned onSmallMinesDSachieve up to 75% intersection-over-union while maintaining a strong balance between minimizing false positives and negatives.
Stella Ofori-Ampofo, Antony Zappacosta, Ridvan Salih Kuzu, Peter Schauer, Martin Willberg, Xiao Xiang Zhu 0001
IEEE Geosci. Remote. Sens. Lett.6
2025 Semi-Supervised Building Footprint Extraction Using Debiased Pseudo-Labels
abstract
Accurate extraction of building footprints from satellite imagery is of high value. Currently, deep learning methods are predominant in this field due to their powerful representation capabilities. However, they generally require extensive pixel-wise annotations, which constrains their practical application. Semi-supervised learning (SSL) significantly mitigates this requirement by leveraging large volumes of unlabeled data for model self-training (ST), thus enhancing the viability of building footprint extraction. Despite its advantages, SSL faces a critical challenge: the imbalanced distribution between the majority background class and the minority building class, which often results in model bias toward the background during training. To address this issue, this article introduces a novel method called DeBiased matching (DBMatch) for semi-supervised building footprint extraction. DBMatch comprises three main components: 1) a basic supervised learning module (SUP) that uses labeled data for initial model training; 2) a classical weak-to-strong ST module that generates pseudo-labels from unlabeled data for further model ST; and 3) a novel logit debiasing (LDB) module that calculates a global logit bias between building and background, allowing for dynamic pseudo-label calibration. To verify the effectiveness of the proposed DBMatch, extensive experiments are performed on three public building footprint extraction datasets covering six global cities in SSL setting. The experimental results demonstrate that our method significantly outperforms some advanced SSL methods in semi-supervised building footprint extraction. Our codes will be publicly provided athttps://github.com/zhu-xlab/SSL_Buildings.
Wei Huang 0068, Ziqi Gu, Yilei Shi, Zhitong Xiong, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 Height-Assisted Semi-Supervised Building Footprint Extraction From Optical Remote Sensing Images
abstract
Automatic building footprint extraction from optical remote sensing (RS) images is popular and crucial for various downstream applications. Current building footprint extraction methods are mainly based on deep learning, which requires large amounts of manually labeled data for model training, limiting their practical deployment. Semi-supervised semantic segmentation (SSS), which leverages limited labeled data for supervised learning and abundant unlabeled data for unsupervised self-training, offers a promising solution to reduce this reliance. Nonetheless, directly applying existing SSS methods to building footprint extraction with limited labels fails to fully exploit the geometric structural features of buildings—key characteristics that distinguish them from background. To tackle this challenge, we propose a semi-supervised learning framework, HeightMatch, which integrates real or synthetic height information with RS images to extract more comprehensive and discriminative feature representations of buildings, particularly in limited-label scenarios. During training, these height maps effectively enhance the model’s ability to capture geometric structures, leading to more accurate pseudo-labels for unlabeled data and thereby enabling more effective self-training. At inference, building predictions rely solely on RS images, ensuring the practicality of the proposed method. Extensive experimental results on five widely-used building footprint extraction datasets demonstrate the effectiveness and superiority of our method in comparison with multiple state-of-the-art SSS methods. Our code is available at https://github.com/zhu-xlab/HeightMatch.
Wei Huang 0068, Yilei Shi, Zhitong Xiong, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.4
2025 DREB-Net: Dual-Stream Restoration Embedding Blur-Feature Fusion Network for High-Mobility UAV Object Detection
abstract
Object detection algorithms are pivotal components of UAV imaging systems, extensively employed in complex fields. However, images captured by high-mobility UAVs often suffer from motion blur cases, which significantly impedes the performance of advanced object detection algorithms. To address these challenges, we propose an innovative object detection algorithm specifically designed for blurry images, named dual-stream restoration embedding blur-feature fusion network (DREB-Net). First, DREB-Net addresses the particularities of blurry image object detection problem by incorporating a blurry image restoration auxiliary branch (BRAB) during the training phase. Second, it fuses the extracted shallow features via multilevel attention-guided feature fusion (MAGFF) module, to extract richer features. Here, the MAGFF module comprises local attention modules and global attention modules, which assign different weights to the branches. Then, during the inference phase, the deep feature extraction of the BRAB can be removed to reduce computational complexity and improve detection speed. In loss function, a combined loss of mean squared error (MSE) and SSIM is added to the BRAB to restore blurry images. Finally, DREB-Net introduces fast Fourier transform in the early stages of feature extraction, via a learnable frequency domain amplitude modulation module (LFAMM), to adjust feature amplitude and enhance feature processing capability. Compared to the baseline, DREB-Net achieved an approximate 7% increase in both mAP50 and mAR50 across two experimental datasets. Experimental results indicate that DREB-Net can still effectively perform object detection tasks under motion blur in captured images, showcasing excellent performance and broad application prospects. Our source code will be available athttps://github.com/EEIC-Lab/DREB-Net.git.
Qingpeng Li, Leyuan Fang, Yuhan Kang, Shutao Li 0001, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.6
2025 CromSS: Cross-Modal Pretraining With Noisy Labels for Remote Sensing Image Segmentation
abstract
We explore the potential of large-scale noisily labeled data to enhance feature learning by pretraining semantic segmentation models within a multimodal framework for geospatial applications. We propose a novel cross-modal sample selection (CromSS) method, a weakly supervised pretraining strategy designed to improve feature representations through cross-modal consistency and noise mitigation techniques. Unlike conventional pretraining approaches, CromSS exploits massive amounts of noisy and easy-to-come-by labels for improved feature learning beneficial to semantic segmentation tasks. We investigate middle and late fusion strategies to optimize the multimodal pretraining architecture design. We also introduce a cross-modal sample selection module to mitigate the adverse effects of label noise, which employs a cross-modal entangling strategy to refine the estimated confidence masks within each modality to guide the sampling process. Additionally, we introduce a spatial–temporal label smoothing technique to counteract overconfidence for enhanced robustness against noisy labels. To validate our approach, we assembled the multimodal dataset, NoLDO-S12, which consists of a large-scale noisy label subset from Google’s Dynamic World (DW) dataset for pretraining and two downstream subsets with high-quality labels from Google DW and OpenStreetMap (OSM) for transfer learning. Experimental results on two downstream tasks and the publicly available DFC2020 dataset demonstrate that when effectively utilized, the low-cost noisy labels can significantly enhance feature learning for segmentation tasks. The data, codes, and pretrained weights are freely available athttps://github.com/zhu-xlab/CromSS.
Chenying Liu 0001, Conrad M. Albrecht, Yi Wang 0072, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.4
2025 Can Uncertainty Quantification Benefit From Label Embeddings? A Case Study on Local Climate Zone Classification
abstract
Modern deep learning models have achieved superior performance in almost all fields of remote sensing. An often neglected aspect of these models is the quantification and evaluation of predictive uncertainties. Regarding a classification task, this means that the focus of the analysis solely lies on performance metrics such as accuracy or the loss. On the other hand, a notion of uncertainty indicates the model’s indecisiveness among the given classes and is essential to understand where the model struggles to classify the data samples. In this work, three levels of uncertainty are distinguished, starting with the typical softmax pseudo-probabilities as level-1 uncertainty. As a next level, the more flexible Dirichlet framework is utilized as model output space, and hereby also, a Bayesian setting with an uninformative prior is considered. For the level-3 uncertainty, an empirical Bayes setting is incorporated where a latent embedding of the label space is iteratively estimated by the marginal likelihood of the fully parameterized label space (see [1]). The estimated embeddings are then learned by the network in three different settings: Two regression losses use the embeddings directly, while the closed-form solution of the Kullback-Leibler (KL-) Divergence uses the embedding parameterized as a Dirichlet distribution. To assess the different levels of uncertainty, the label evaluation subset of the So2Sat LCZ42 dataset, which contains label votes from multiple remote sensing experts, is investigated. The predictive uncertainties are evaluated by means of Out-of-Distribution (OoD) detection and calibration performance. Overall, the embedding-based approaches show strong performance for calibration, while for the OoD experiments, the Bayesian Dirichlet setting with an uninformative prior achieves the best performance. In conclusion, embedded labels offer a flexible framework for incorporating uncertain or ambiguous labels into a supervised training setup. They could be highly beneficial for applications in fields such as urban planning or disaster response.
Christoph Schweden, Katharina Hechinger, Göran Kauermann, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.4
2025 Weak-Strong Graph Contrastive Learning Neural Network for Hyperspectral Image Classification
abstract
Deep learning methods have shown promising results in various hyperspectral image (HSI) analysis tasks. Despite these advancements, existing models still struggle to accurately identify fine-classified land cover types on noisy hyperspectral images. Traditional methods have limited performance when extracting features from noisy hyperspectral data. Graph Neural Networks (GNNs) offer an adaptable and robust structure by effectively extracting both spectral and spatial features. However, supervised models still require large quantities of labeled data for effective training, posing a significant challenge. Contrastive learning, which leverages unlabeled data for pre-training, can mitigate this issue by reducing the dependency on extensive manual annotation. To address the issues, we propose WSGraphCL, a weak-strong graph contrastive learning model for HSI classification, and conduct experiments in a few-shot scenario. First, the image is transformed into K-hop subgraphs through a spectral-spatial adjacency matrix construction method. Second, WSGraphCL leverages contrastive learning to pre-train a graph-based encoder on the unlabeled hyperspectral image. We demonstrate that weak-strong augmentations and false negative pairs filtering stabilize pre-training and get good-quality representations. Finally, we test our model with a lightweight classifier on the features with a handful of labels. Experimental results showcase the superior performance of WSGraphCL compared to several baseline models, thereby emphasizing its efficacy in addressing the identified limitations in HSI classification. The code repository will be published on the GitHub project under the URL: https://github.com/zhu-xlab/WSGraphCL.
Sirui Wang 0010, Nassim Ait Ali Braham, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.3
2025 Global Collinearity-Aware Polygonizer for Polygonal Building Mapping in Remote Sensing
abstract
This paper addresses the challenge of mapping polygonal buildings from remote sensing images and introduces a novel algorithm, the Global Collinearity-aware Polygonizer (GCP). GCP, built upon an instance segmentation framework, processes binary masks produced by any instance segmentation model. The algorithm begins by collecting polylines sampled along the contours of the binary masks. These polylines undergo a refinement process using a transformer-based regression module to ensure they accurately fit the contours of the targeted building instances. Subsequently, a collinearity-aware polygon simplification module simplifies these refined polylines and generate the final polygon representation. This module employs dynamic programming technique to optimize an objective function that balances the simplicity and fidelity of the polygons, achieving globally optimal solutions. Furthermore, the optimized collinearity-aware objective is seamlessly integrated into network training, enhancing the cohesiveness of the entire pipeline. The effectiveness of GCP has been validated on three public benchmarks for polygonal building mapping. Further experiments reveal that applying the collinearity-aware polygon simplification module to arbitrary polylines, without prior knowledge, enhances accuracy over traditional methods such as the Douglas-Peucker algorithm. This finding underscores the broad applicability of GCP. The code for the proposed method will be made available at https://github.com/zhu-xlab/GCP.
Fahong Zhang 0001, Yilei Shi, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.3
2025 RainScaler: A Physics-Inspired Network for Precipitation Correction and Downscaling
abstract
Spatial downscaling of precipitation, in which finegrained regional precipitation patterns are recovered from coarse-resolution images, plays a crucial role in various weather and meteorological analyses. However, the intricate noise information presented in the observation data intertwines with the fine-scale characteristics, which poses challenges for subsequent feature extraction. Regional precipitation suffers from complex spatial patterns. Moreover, the real observatory data contains information inconsistent with the established physical principle, due either to inaccurate or incomplete physical models or limited data quality, thus making the implementation of physicallyinformed deep learning more difficult. For example, strong physical constraints may lead to over-regularization, in which the model becomes too rigid and fails to capture certain complexities in the data. In this work, we propose RainScaler, a physicsinspired deep neural network, to tackle these issues. First, to remove the noise and preserve the vital precipitation patterns effectively, the proposed RainScaler exploits an Inconsistencyaware Denoising Net to explicitly model the spatial variability of noise in the input. In addition, a graph module is designed to learn the geographical-dependent fine-grained patterns in high dimensional feature space at a moderate computation cost. Finally, multi-scale physical constraints are skillfully embedded to incorporate additional insights into the data-driven framework. We test our approach on a public dataset consisting of over 60,000 real low-resolution and high-resolution precipitation map pairs collected by different sensors. Our method produces realisticlooking precipitation maps with better discernment capability and corrects the structural error of precipitation distribution, especially for extreme events. Moreover, we evaluate the potential risks of incorporating physical constraints in real-world data applications. Our method unveils opportunities for multi-source data fusion and provides possible solutions to improve the physical feasibility of data-driven models. The codes are available in https://github.com/zhu-xlab/RainScaler.git.
Shan Zhao 0007, Zhitong Xiong, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.3
2025 Hybrid Quantum Deep Learning With Superpixel Encoding for Earth Observation Data Classification
abstract
Earth observation (EO) has inevitably entered the Big Data era. The computational challenge associated with analyzing large EO data using sophisticated deep learning models has become a significant bottleneck. To address this challenge, there has been a growing interest in exploring quantum computing as a potential solution. However, the process of encoding EO data into quantum states for analysis potentially undermines the efficiency advantages gained from quantum computing. This article introduces a hybrid quantum deep learning model that effectively encodes and analyzes EO data for classification tasks. The proposed model uses an efficient encoding approach called superpixel encoding, which reduces the quantum resources required for large image representation by incorporating the concept of superpixels. To validate the effectiveness of our model, we conducted evaluations on multiple EO benchmarks, including Overhead-MNIST, So2Sat LCZ42, and SAT-6 datasets. In addition, we studied the impacts of different interaction gates and measurements on classification performance to guide model optimization. The experimental results suggest the validity of our model for accurate classification of EO data. Our models and code are available on https://github.com/zhu-xlab/SEQNN.
Yilei Shi, Tobias Guggemos, Xiao Xiang Zhu 0001
IEEE Trans. Neural Networks Learn. Syst.4
2024 Urban Land Cover Classification with Efficient Hybrid Quantum Machine Learning Model
abstract
Urban land cover classification aims to derive crucial information from earth observation data and categorize it into specific land uses. To achieve accurate classification, sophisticated machine learning models trained with large earth observation data are employed, but the required computation power has become a bottleneck. Quantum computing might tackle this challenge in the future. However, representing images into quantum states for analysis with quantum computing is challenging due to the high demand for quantum resources. To tackle this challenge, we propose a hybrid quantum neural network that can effectively represent and classify remote sensing imagery with reduced quantum resources. Our model was evaluated on the Local Climate Zone (LCZ)-based land cover classification task using the TensorFlow Quantum platform, and the experimental results indicate its validity for accurate urban land cover classification.
Yilei Shi, Xiao Xiang Zhu 0001
CEC3
2024 Representation Enhancement-Stabilization: Reducing Bias-Variance of Domain Generalization
Wei Huang 0068, Yilei Shi, Zhitong Xiong, Xiao Xiang Zhu 0001
ECCV (36)4
2024 Decoupling Common and Unique Representations for Multimodal Self-supervised Learning
Yi Wang 0072, Conrad M. Albrecht, Nassim Ait Ali Braham, Chenying Liu 0001, Zhitong Xiong, Xiao Xiang Zhu 0001
ECCV (29)6
2024 Learning Building Energy Efficiency with Semantic Attributes
abstract
Non-intrusive estimation of building energy efficiency has profound applications in advancing sustainability in the built environment. Recent studies often focus on predicting energy performance alone, neglecting the interplay between the performance and related building semantics. This paper investigates whether incorporating semantic attributes benefits energy efficiency estimation. We develop a neural network to estimate energy efficiency, with building age and usage type as additional supervision for multi-task learning. The neural network processes both aerial imagery and airborne LiDAR data to classify buildings as energy-efficient or inefficient. Our results demonstrate the effectiveness of the superimposed semantics, particularly with building age. With the multi-task model achieving a 63.78% F1 score and outperforming that supervised solely with energy efficiency by 2.86%, this paper reveals the potential of integrating semantic attributes in modeling building energy performance.
Zhaiyu Chen, Ziqi Gu, Yilei Shi, Xiao Xiang Zhu 0001
IGARSS4
2024 Building Attributes Recognition with Noisy and Incomplete Labels
abstract
Recognizing building attributes from remote sensing images is crucial for various applications. Recent developments in deep learning have demonstrated promising results in identifying these attributes. Nonetheless, a major challenge is the requirement for extensive and accurate building attribute data. Two primary data sources are commonly considered: Open-StreetMap (OSM), which offers global building information but often lacks completeness and correctness, and cadastral data, known for its high quality but typically restricted to certain areas. These two sources enable comparison between deep learning models trained on noisy and incomplete OSM data and those trained on accurate and complete cadastral data. In this work, comprehensive experiments on buildings in Bavaria, Germany, are conducted, covering diverse attributes such as footprints, use, and height. A large building dataset with corresponding building attribute labels from OSM and cadastral data is created, with OSM data featuring varying levels of incompleteness and noise for different attributes and cadastral data serving as ground truth. Moreover, we evaluate the effectiveness of several prevailing methods designed to handle noisy and incomplete labels, assessing their applicability to real-world scenarios with incomplete and noisy OSM labels.
Ziqi Gu, Zhaiyu Chen, Yilei Shi, Xiao Xiang Zhu 0001
IGARSS4
2024 Disentangling Semi-Supervised Semantic Segmentation of Remote Sensing Images
abstract
In Earth observation, semantic understanding of Remote Sensing (RS) images holds significant importance, yet it is hindered in practice by the need for extensive manual pixel-level labeling. Semi-supervised semantic segmentation (SSS) of RS images would be a promising solution, which fully utilizes unlabeled data for model self-training under the guidance of limited labeled data. The mainstream SSS methods use pseudo-labels of the unlabeled data for model training, however, their performance is bottlenecked because of confirmation bias, i.e., stubborn incorrect pseudo-labels. To counter this, our study introduces a novel disentanglement learning (DL) method tailored for RS-SSS. It separates the predictions of the labeled and unlabeled data by two individual prediction heads during current training, and then integrates them during follow-up training. The experimental results verify its effectiveness on two widely-used RS semantic segmentation datasets in semi-supervised setting.
Wei Huang 0068, Yilei Shi, Zhitong Xiong, Xiao Xiang Zhu 0001
IGARSS4
2024 Task Specific Pretraining with Noisy Labels for Remote Sensing Image Segmentation
abstract
Compared to supervised deep learning, self-supervision provides remote sensing a tool to reduce the amount of exact, human-crafted geospatial annotations. While image-level information for unsupervised pretraining efficiently works for various classification downstream tasks, the performance on pixel-level semantic segmentation lags behind in terms of model accuracy. On the contrary, many easily available label sources (e.g., automatic labeling tools and land cover land use products) exist, which can provide a large amount of noisy labels for segmentation model training. In this work, we propose to exploit noisy semantic segmentation maps for model pretraining. Our experiments provide insights on robustness per network layer. The transfer learning settings test the cases when the pretrained encoders are fine-tuned for different label classes and decoders. The results from two datasets indicate the effectiveness of task-specific supervised pretraining with noisy labels. Our findings pave new avenues to improved model accuracy and novel pretraining strategies for efficient remote sensing image segmentation.
Chenying Liu 0001, Conrad M. Albrecht, Yi Wang 0072, Xiao Xiang Zhu 0001
IGARSS4
2024 Dominant Leaf Type Classification Using Sentinel-1 Time Series
abstract
The classification of dominant leaf types, which distinguishes forests based on their leaf conditions, is beneficial for forest management and policymakers. This paper proposes a model based on U-Net to classify the land into non-tree areas, broadleaf forests, and coniferous forests. The dual-pol Sentinel-1 data from January, May, August, and October of 2018 were stacked as a time series. Due to the class imbalance issue, where the non-tree area category dominates (52.69%) the dataset, the model tends to be biased. Thus, re-weighting is introduced to balance the loss. We tested and compared two types of methods: class-aware and task-aware re-weighting. The results indicate that re-weighting effectively mitigates the class imbalance issue.
Qian Song, Ridvan Salih Kuzu, Xiao Xiang Zhu 0001
IGARSS3
2024 Multi-Label Guided Supervised Contrastive Learning for Earth Observation Pretraining
abstract
Pretraining foundation models on large-scale satellite imagery has raised great interest in Earth observation. While most pretraining is conducted purely self-supervised, many land cover land use products that provide free and global annotations tend to be overlooked. To bridge this gap, we propose to exploit land-cover-generated multi-label annotations to guide supervised contrastive learning for Earth observation. We match the SSL4EO-S12 dataset with Dynamic World land cover maps and integrate image-level multi-label annotations. During pretraining, the label similarities between different images are calculated, and those with high similarity scores are pulled together in the embedding space. Experimental results on classification and segmentation downstream tasks demonstrate the effectiveness of the proposed method.
Yi Wang 0072, Conrad M. Albrecht, Xiao Xiang Zhu 0001
IGARSS3
2024 One for All: Toward Unified Foundation Models for Earth Vision
abstract
Foundation models characterized by extensive parameters and trained on large-scale datasets have demonstrated remarkable efficacy across various downstream tasks for remote sensing data. Current remote sensing foundation models typically specialize in a single modality or a specific spatial resolution range, limiting their versatility for downstream datasets. While there have been attempts to develop multi-modal remote sensing foundation models, they typically employ separate vision encoders for each modality or spatial resolution, necessitating a switch in backbones contingent upon the input data. To address this issue, we introduce a simple yet effective method, termed OFA-Net (One-For-All Network): employing a single, shared Transformer backbone for multiple data modalities with different spatial resolutions. Using the masked image modeling mechanism, we pre-train a single Transformer backbone on a curated multi-modal dataset with this simple design. Then the backbone model can be used in different downstream tasks, thus forging a path towards a unified foundation backbone model in Earth vision. The proposed method is evaluated on 12 distinct downstream tasks and demonstrates promising performance.
Zhitong Xiong, Yi Wang 0072, Fahong Zhang 0001, Xiao Xiang Zhu 0001
IGARSS4
2024 Referring Image Segmentation for Remote Sensing Data
abstract
In this paper, we present a new task: referring image segmentation for remote sensing data, which targets segmenting out specific objects referred to by natural language. Due to the absence of a dataset for this task, we construct a dataset based on the SkyScapes dataset. Our dataset is designed with linguistically structured expressions that focus on object categories, attributes, and spatial relationships, enabling the generation of binary masks from semantic segmentation maps. To benchmark this task, we evaluate and compare the performance of three different convolutional neural network (CNN)-based methods and a Transformer-based method. Experimental results provide valuable insights into the adaptability of these methods to remote sensing data, highlighting the potential of our dataset as a resource for the remote sensing community to further explore vision-language tasks.
Zhenghang Yuan, Lichao Mou, Yuansheng Hua, Xiao Xiang Zhu 0001
IGARSS4
2024 Deep-Learning-Based Large-Scale Forest Height Generation
abstract
The vegetation height has been identified as a key biophysical parameter to justify the role of forests in the carbon cycle and ecosystem productivity. Therefore, consistent and large-scale forest height is essential for managing terrestrial ecosystems, mitigating climate change, and preventing biodiversity loss. Since spaceborne multispectral instruments, Light Detection and Ranging (LiDAR), and Synthetic Aperture Radar (SAR) have been widely used for large-scale earth observation for years, this paper explores the possibility of generating largescale and high-accuracy forest heights with the synergy of the Sentinel-1, Sentinel-2, and ICESat-2 data. A Forest Height Generative Adversarial Network (FH-GAN) is developed to retrieve forest height from Sentinel-1 and Sentinel-2 images sparsely supervised by the ICESat-2 data. This model is made up of a cascade forest height and coherence generator, where the output of the forest height generator is fed into the spatial discriminator to regularize spatial details, and the coherence generator is connected to a coherence discriminator to refine the vertical details. A progressive strategy further underpins the generator to boost the accuracy of multi-source forest height estimation. Results indicated that FH-GAN achieves the best RMSE of 2.10 m at a large scale compared with the LVIS reference and the best RMSE of 6.16 m compared with the ICESat-2 reference.
Yuanyuan Wang 0002, Xiao Xiang Zhu 0001
IGARSS3
2024 Efficient Subseasonal Weather Forecast Using Teleconnection-Informed Transformers
abstract
Subseasonal forecasting, which is pivotal for agriculture, water resource management, and early warning of disasters, faces challenges due to the chaotic nature of the atmosphere. Recent advances in machine learning (ML) have revolutionized weather forecasting by achieving competitive predictive skills to numerical models. However, training such foundation models requires thousands of GPU days, which causes substantial carbon emissions and limits their broader applicability. Moreover, ML models tend to fool the pixel-wise error scores by producing smoothed results which lack physical consistency and meteorological meaning. To deal with the aforementioned problems, we propose a teleconnection-informed transformer. Our architecture leverages the pretrained Pangu model to achieve good initial weights and integrates a teleconnection-informed temporal module to improve predictability in an extended temporal range. Remarkably, by adjusting 1.1% of the Pangu model’s parameters, our method enhances predictability on four surface and five upper-level atmospheric variables at a two-week lead time. Furthermore, the teleconnection-filtered features improve the spatial granularity of outputs significantly, indicating their potential physical consistency. Our research underscores the importance of atmospheric and oceanic teleconnections in driving future weather conditions. Besides, it presents a resource-efficient pathway for researchers to leverage existing foundation models on versatile downstream tasks.
Shan Zhao 0007, Zhitong Xiong, Xiao Xiang Zhu 0001
IGARSS3
2024 Towards Large-Scale Urban Flood Mapping Using Sentinel-1 Data
abstract
Within the realm of deep learning techniques, numerous remote sensing applications can be effectively addressed using deep learning algorithms. However, there is a scarcity of studies in synthetic Aperture radar (SAR)-based urban flood mapping involving deep learning techniques, primarily due to two reasons. First, SAR-based urban flood mapping is inherently rooted in change detection, resulting in a complex multi-modality problem within the imbalance data. This complexity arises from the integration of SAR intensity, InSAR coherence, and even SAR phase information acquired from different polarizations (i.e., VV and VH polarization in Sentinel-1 data) both before and after the event. The second challenge is the absence of a benchmark dataset specifically designed for SAR-based urban flood mapping. In an effort to fill this gap, a benchmark dataset for large-scale flood mapping using Sentinel-1 data, which includes not only SAR intensity but also InSAR coherence, should be created. The SAR pre-processing should be carefully checked at the very beginning. With this aim, we tested the specking filter and kernel size selection for the SAR preprocessing for deep learning models. Through this initiative, we found that despeckling SAR intensity and selecting the kernel size in InSAR coherence calculation do not significantly affect the accuracy in deep learning-based urban flood mapping using Sentinel-1 data. The curated benchmark dataset will be presented in the final paper.
Jie Zhao 0021, Xiao Xiang Zhu 0001
IGARSS2
2024 Ultrasound Image-to-Video Synthesis via Latent Dynamic Diffusion Models
Tingxiu Chen, Yilei Shi, Zixuan Zheng, Bingcong Yan, Jingliang Hu, Xiao Xiang Zhu 0001, Lichao Mou
MICCAI (4)6
2024 CausalCLIPSeg: Unlocking CLIP's Potential in Referring Medical Image Segmentation with Causal Intervention
Yaxiong Chen, Minghong Wei, Zixuan Zheng, Jingliang Hu, Yilei Shi, Shengwu Xiong 0001, Xiao Xiang Zhu 0001, Lichao Mou
MICCAI (3)7
2024 Striving for Simplicity: Simple Yet Effective Prior-Aware Pseudo-labeling for Semi-supervised Ultrasound Image Segmentation
Yaxiong Chen, Zixuan Zheng, Jingliang Hu, Yilei Shi, Shengwu Xiong 0001, Xiao Xiang Zhu 0001, Lichao Mou
MICCAI (9)7
2024 Rethinking Cell Counting Methods: Decoupling Counting and Localization
Zixuan Zheng, Yilei Shi, Jingliang Hu, Xiao Xiang Zhu 0001, Lichao Mou
MICCAI (4)5
2024 Reducing Annotation Burden: Exploiting Image Knowledge for Few-Shot Medical Video Object Segmentation via Spatiotemporal Consistency Relearning
Zixuan Zheng, Yilei Shi, Jingliang Hu, Xiao Xiang Zhu 0001, Lichao Mou
MICCAI (12)5
2024 Medical hyperspectral image classification based weakly supervised single-image global learning network
Lichao Mou, Shihao Shan, Hao Zhang 0113, Yafei Qi, Dexin Yu, Xiao Xiang Zhu 0001, Nianzheng Sun, Xiangrong Zheng, Xiaopeng Ma
Eng. Appl. Artif. Intell.7
2024 Can Land Cover Classification Models Benefit From Distance-Aware Architectures?
abstract
The quantification of predictive uncertainties helps to understand where existing models struggle to find the correct prediction. A useful quality control tool is the task of detecting out-of-distribution (OOD) data by examining the model’s predictive uncertainty. For this task, deterministic single forward pass frameworks have recently been established as deep learning models and have shown competitive performance in certain tasks. The unique combination of spectrally normalized weight matrices and residual connection networks with an approximate Gaussian Process output layer can here offer the best trade-off between performance and complexity. We utilize this framework with a refined version that adds spectral batch normalization and an inducing points approximation of the Gaussian Process for the task of OOD detection in remote sensing image classification. This is an important task in the field of remote sensing because it provides an evaluation of how reliable the model’s predictive uncertainty estimates are. By performing experiments on the benchmark datasetsEurosatandSo2Sat LCZ42, we can show the effectiveness of the proposed adaptions to the residual networks. Depending on the chosen dataset, the proposed methodology achieves OOD detection performance up to 16% higher than previously considered distance-aware networks. Compared to other uncertainty quantification methodologies, the results are on the same level and exceed them in certain experiments by up to 2%. In particular, spectral batch normalization, which normalizes the batched data as opposed to normalizing the network weights by the spectral normalization, plays a crucial role and leads to performance gains of up to 3% in every single experiment. For reproducibility, the code can be found here: https://github.com/ChrisKo94/DUE Land Cover.
Christoph Koller, Peter Jung 0001, Xiao Xiang Zhu 0001
IEEE Geosci. Remote. Sens. Lett.3
2024 Beyond local patches: Preserving global-local interactions by enhancing self-attention via 3D point cloud tokenization
abstract
Transformer-based architectures have recently shown impressive performance on various point cloud understanding tasks such as 3D object shape classification and semantic segmentation. Particularly, this can be attributed to their self-attention mechanism, which has the ability to capture long-range dependencies. However, current methods have constrained it to operate in local patches due to its quadratic memory constraints. This hinders their generalization ability and scaling capacity due to the loss of non-locality in early layers. To tackle this issue, we propose a window-based transformer architecture that captures long-range dependencies while aggregating information in the local patches. We do this by interacting each window with a set of global point cloud tokens — a representative subset of the entire scene — and augmenting the local geometry through a 3D Histogram of Oriented Gradients (HOG) descriptor. Through a series of experiments on segmentation and classification tasks, we show that our model exceeds the state-of-the-art on S3DIS semantic segmentation (+1.67% mIoU), ShapeNetPart part segmentation (+1.03% instance mIoU) and performs competitively on ScanObjectNN 3D object classification.1
Muhammad Shahzad 0002, Saqib Ali Khan, Muhammad Moazam Fraz, Xiao Xiang Zhu 0001
Pattern Recognit.5
2024 Integrating Detailed Features and Global Contexts for Semantic Segmentation in Ultrahigh-Resolution Remote Sensing Images
abstract
Semantic segmentation of ultrahigh-resolution (UHR) remote sensing images is a fundamental task for many downstream applications. Achieving precise pixel-level classification is paramount for obtaining exceptional segmentation results. This challenge becomes even more complex due to the need to address intricate segmentation boundaries and accurately delineate small objects within the remote sensing imagery. To meet these demands effectively, it is critical to integrate two crucial components: global contextual information and spatial detail feature information. In response to this imperative, the multilevel context-aware segmentation network (MCSNet) emerges as a promising solution. MCSNet is engineered to not only model the overarching global context but also extract intricate spatial detail features, thereby optimizing segmentation outcomes. The strength of MCSNet lies in its two pivotal modules, the spatial detail feature extraction (SDFE) module and the refined multiscale feature fusion (RMFF) module. Moreover, to further harness the potential of MCSNet, a multitask learning approach is employed. This approach integrates boundary detection and semantic segmentation, ensuring that the network is well-rounded in its segmentation capabilities. The efficacy of MCSNet is rigorously demonstrated through comprehensive experiments conducted on two established international society for photogrammetry and remote sensing (ISPRS) 2-D semantic labeling datasets: Potsdam and Vaihingen. These experiments unequivocally establish MCSNet stands as a pioneering solution, that delivers state-of-the-art performance, as evidenced by its outstanding mean intersection over union (mIoU) and mean$F1$-score (mF1) metrics. The code is available at:https://github.com/WUTCM-Lab/MCSNet.
Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu, Xiao Xiang Zhu 0001, Lichao Mou
IEEE Trans. Geosci. Remote. Sens.5
2024 PixelDINO: Semi-Supervised Semantic Segmentation for Detecting Permafrost Disturbances in the Arctic
abstract
Arctic permafrost is facing significant changes due to global climate change. As these regions are largely inaccessible, remote sensing plays a crucial rule in better understanding the underlying processes across the Arctic. In this study, we focus on the remote detection of retrogressive thaw slumps (RTSs), a permafrost disturbance comparable to slow landslides. For such remote sensing tasks, deep learning has become an indispensable tool, but limited labeled training data remains a challenge for training accurate models. We present PixelDINO, a semi-supervised learning approach, to improve model generalization across the Arctic with a limited number of labels. PixelDINO leverages unlabeled data by training the model to define its own segmentation categories (pseudoclasses), promoting consistent structural learning across strong data augmentations. This allows the model to extract structural information from unlabeled data, supplementing the learning from labeled data. PixelDINO surpasses both supervised baselines and existing semi-supervised methods, achieving average intersection-over-union (IoU) of 30.2 and 39.5 on the two evaluation sets, representing significant improvements of 13% and 21%, respectively, over the strongest existing models. This highlights the potential for training robust models that generalize well to regions that were not included in the training data.
Konrad Heidler, Ingmar Nitze, Guido Grosse, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 Going Beyond One-Hot Encoding in Classification: Can Human Uncertainty Improve Model Performance in Earth Observation?
abstract
Technological and computational advances continuously drive forward the field of deep learning in remote sensing. In recent years, the derivation of quantities describing the uncertainty in the prediction - which naturally accompanies the modeling process - has sparked interest in the remote sensing community. Often neglected in the machine learning setting is the human uncertainty that influences numerous labeling processes. As the core of this work, the task of Local Climate Zone (LCZ) classification is studied by means of a data set that contains multiple label votes by domain experts for each image. The inherent label uncertainty describes the ambiguity among the domain experts and is explicitly embedded into the training process via distributional labels. We show that incorporating the label uncertainty helps the model to generalize better to the test data and increases model performance. Similar to existing calibration methods, the distributional labels lead to better-calibrated probabilities, which in turn yield more certain and trustworthy predictions. For reproducibility, we provide our code here https://github.com/ChrisKo94/LCZ_LDL and here https://gitlab.lrz.de/ai4eo/WG_Uncertainty/lcz_ldl.
Christoph Koller, Göran Kauermann, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.3
2024 A Review of Building Extraction From Remote Sensing Imagery: Geometrical Structures and Semantic Attributes
abstract
In the remote sensing community, extracting buildings from remote sensing imagery has triggered great interest. While many studies have been conducted, a comprehensive review of these approaches that are applied to optical and synthetic aperture radar (SAR) imagery is still lacking. Therefore, we provide an in-depth review of both early efforts and recent advances, which are aimed at extracting geometrical structures or semantic attributes of buildings, including building footprint generation, building facade segmentation, roof segment and superstructure segmentation, building height retrieval, building type classification, building change detection, and annotation data correction. Furthermore, a list of corresponding benchmark datasets is given. Finally, challenges and outlooks of existing approaches as well as promising applications are discussed to enhance comprehension within this realm of research.
Qingyu Li 0001, Lichao Mou, Yao Sun 0005, Yuansheng Hua, Yilei Shi, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.6
2024 AIO2: Online Correction of Object Labels for Deep Learning With Incomplete Annotation in Remote Sensing Image Segmentation
abstract
While the volume of remote sensing data is increasing daily, deep learning in Earth Observation faces lack of accurate annotations for supervised optimization. Crowdsourcing projects such as OpenStreetMap distribute the annotation load to their community. However, such annotation inevitably generates noise due to insufficient control of the label quality, lack of annotators, frequent changes of the Earth’s surface as a result of natural disasters and urban development, among many other factors. We presentAdaptively trIggered Online Object-wise correction (AIO2)to address annotation noise induced by incomplete label sets. AIO2 features anAdaptive Correction Trigger (ACT)module that avoids label correction when the model training under- or overfits, and anOnline Object-wise Correction (O2C)methodology that employs spatial information for automated label modification. AIO2 utilizes a mean teacher model to enhance training robustness with noisy labels to both stabilize the training accuracy curve for fitting in ACT and provide pseudo labels for correction in O2C. Moreover, O2C is implementedonlinewithout the need to store updated labels every training epoch. We validate our approach on two building footprint segmentation datasets with different spatial resolutions. Experimental results with varying degrees of building label noise demonstrate the robustness of AIO2. Source code will be available at https://github.com/zhu-xlab/AIO2.git.
Chenying Liu 0001, Conrad M. Albrecht, Yi Wang 0072, Qingyu Li 0001, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 HyperLISTA-ABT: An Ultralight Unfolded Network for Accurate Multicomponent Differential Tomographic SAR Inversion
abstract
Deep neural networks based on unrolled iterative algorithms have achieved remarkable success in sparse reconstruction applications, such as synthetic aperture radar (SAR) tomographic inversion (TomoSAR). However, the currently available deep learning-based TomoSAR algorithms are limited to 3-D reconstruction. The extension of deep learning-based algorithms to 4-D imaging, i.e., differential TomoSAR (D-TomoSAR) applications, is impeded mainly due to the high-dimensional weight matrices required by the network designed for D-TomoSAR inversion, which typically contain millions of freely trainable parameters. Learning such huge number of weights requires an enormous number of training samples, resulting in a large memory burden and excessive time consumption. To tackle this issue, we propose an efficient and accurate algorithm called HyperLISTA-ABT. The weights in HyperLISTA-ABT are determined in an analytical way according to a minimum coherence criterion, trimming the model down to an ultra-light one with only three hyperparameters. Additionally, HyperLISTA-ABT improves the global thresholding by utilizing an adaptive blockwise thresholding (ABT) scheme, which applies block-coordinate techniques and conducts thresholding in local blocks, so that weak expressions and local features can be retained in the shrinkage step layer by layer. Simulations were performed and demonstrated the effectiveness of our approach, showing that HyperLISTA-ABT achieves superior computational efficiency with no significant performance degradation compared to the state-of-the-art methods. Real data experiments showed that a high-quality 4-D point cloud could be reconstructed over a large area by the proposed HyperLISTA-ABT with affordable computational resources and in a fast time.
Kun Qian 0020, Yuanyuan Wang 0002, Peter Jung 0001, Yilei Shi, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 Multilabel-Guided Soft Contrastive Learning for Efficient Earth Observation Pretraining
abstract
Self-supervised pretraining on large-scale satellite data has raised great interest in building Earth observation (EO) foundation models. However, many important resources beyond pure satellite imagery, such as land-cover-land-use products that provide free global semantic information, as well as vision foundation models that hold strong knowledge of the natural world, are not widely studied. In this work, we show these free additional resources not only help resolve common contrastive learning bottlenecks but also significantly boost the efficiency and effectiveness of EO pretraining. Specifically, we first propose soft contrastive learning (SoftCon) that optimizes cross-scene soft similarity based on land-cover-generated multilabel supervision, naturally solving the issue of multiple positive samples and too strict positive matching in complex scenes. Second, we revisit and explore cross-domain continual pretraining for both multispectral and synthetic aperture radar (SAR) imagery, building efficient EO foundation models from strongest vision models such as DINOv2. Adapting simple weight-initialization and Siamese masking strategies into our SoftCon framework, we demonstrate impressive continual pretraining performance even when the input modalities are not aligned. Without prohibitive training, we produce multispectral and SAR foundation models that achieve significantly better results in 10 out of 11 downstream tasks than most existing SOTA models. For example, our ResNet50/ViT-S achieve 84.8/85.0 linear probing mAP scores on BigEarthNet-10%, which are better than most existing ViT-L models; under the same setting, our ViT-B sets a new record of 86.8 in multispectral, and 82.5 in SAR, the latter even better than many multispectral models. Dataset and models are available athttps://github.com/zhu-xlab/softcon.
Yi Wang 0072, Conrad M. Albrecht, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.3
2024 Multimodal Co-Learning for Building Change Detection: A Domain Adaptation Framework Using VHR Images and Digital Surface Models
abstract
In this article, we propose a multimodal co-learning framework for building change detection. This framework can be adopted to jointly train a Siamese bitemporal image network and a height difference map (HDiff) network with labeled source data and unlabeled target data pairs. Three co-learning combinations (vanilla co-learning, fusion co-learning, and detached fusion co-learning) are proposed and investigated with two types of co-learning loss functions within our framework. Our experimental results demonstrate that the proposed methods are able to take advantage of unlabeled target data pairs and therefore enhance the performance of single-modal neural networks on the target data. In addition, our synthetic-to-real experiments demonstrate that the recently published synthetic dataset SMARS is feasible to be used in real change detection scenarios, where the optimal result is with the F1 score of 79.29%.
Yuxing Xie, Xiangtian Yuan, Xiao Xiang Zhu 0001, Jiaojiao Tian
IEEE Trans. Geosci. Remote. Sens.3
2024 Self-Supervised Pretraining With Monocular Height Estimation for Semantic Segmentation
abstract
Monocular height estimation (MHE) is key for generating 3-D city models, essential for swift disaster response. Moving beyond the traditional focus on performance enhancement, our study breaks new ground by probing the interpretability of MHE networks. We have pioneeringly discovered that neurons within MHE models demonstrate selectivity for both height and semantic classes. This insight sheds light on the complex inner workings of MHE models and inspires innovative strategies for leveraging elevation data more effectively. Informed by this insight, we propose a pioneering framework that employs MHE as a self-supervised pretraining method for remote sensing (RS) imagery. This approach significantly enhances the performance of semantic segmentation tasks. Furthermore, we develop a disentangled latent transformer (DLT) module that leverages explainable deep representations from pretrained MHE networks for unsupervised semantic segmentation. Our method demonstrates the significant potential of MHE tasks in developing foundation models for sophisticated pixel-level semantic analyses. Additionally, we present a new dataset designed to benchmark the performance of both semantic segmentation and height estimation tasks. The dataset and code will be publicly available athttps://github.com/zhu-xlab/DLT-MHE.pytorch.
Zhitong Xiong, Sining Chen, Yilei Shi, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 Multimodal and Multiresolution Data Fusion for High-Resolution Cloud Removal: A Novel Baseline and Benchmark
abstract
Cloud removal is a significant and challenging problem in remote sensing, and in recent years, there have been notable advancements in this area. However, two major issues remain hindering the development of cloud removal: the unavailability of high-resolution imagery for existing datasets and the absence of evaluation regarding the semantic meaningfulness of the generated structures. In this paper, we introduce M3R-CR, a benchmark dataset for high-resolution Cloud Removal with Multi-Modal and Multi-Resolution data fusion. M3R-CR is the first public dataset for cloud removal to feature globally sampled high-resolution optical observations, paired with radar measurements and pixel-level land cover annotations. With this dataset, we consider the problem of cloud removal in high-resolution optical remote sensing imagery by integrating multi-modal and multi-resolution information. In this context, we have to take into account the alignment errors caused by the multi-resolution nature, along with the more pronounced misalignment issues in high-resolution images due to inherent imaging mechanism differences and other factors. Existing multi-modal data fusion based methods, which assume the image pairs are aligned accurately at pixel-level, are thus not appropriate for this problem. To this end, we design a new baseline named Align-CR to perform the low-resolution SAR image guided high-resolution optical image cloud removal. It gradually warps and fuses the features of the multi-modal and multi-resolution data during the reconstruction process, effectively mitigating concerns associated with misalignment. In the experiments, we evaluate the performance of cloud removal by analyzing the quality of visually pleasing textures using image reconstruction metrics and further analyze the generation of semantically meaningful structures using a well-established semantic segmentation task. The proposed Align-CR method is superior to other baseline methods in both areas. The project is available at https://github.com/zhu-xlab/M3R-CR.
Yilei Shi, Patrick Ebel 0002, Wen Yang 0001, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 Attention-ConvNet Network for Ocean-Front Prediction via Remote Sensing SST Images
abstract
Ocean front is one typical geophysical phenomenon acting as oases in the ocean for fishes and marine mammals. Accurate ocean-front prediction is critical for fishery and navigation safety. However, the formation and evolution of ocean fronts are inherently nonlinear and are influenced by various factors such as ocean currents, wind fields, and temperature changes, making ocean-front prediction a considerable challenge. This study proposes a temporal-sensitive network named Attention-ConvNet to address this challenge. Ocean fronts exhibit significant multiscale characteristics, requiring analysis and prediction across various temporal and spatial scales. The proposed network designs a hierarchical attention mechanism (HAM) that efficiently prioritizes relevant spatial and temporal information to meet the specific requirement. What is more, the proposed network uses a complex hierarchical branching convolutional network (HBCNet) architecture, which allows our network to leverage the complementary strengths of spatial and temporal information, effectively capturing the dynamic and complex variations in ocean fronts. In general, the network prioritizes and focuses on the most relevant information of front dynamics, which ensures its ability to effectively predict the ocean front. External experiments demonstrate that our network significantly outperforms conventional methods, confirming its capability for precise ocean-front prediction. The codes will be publicly available athttps://github.com/yuhudeyue/Ocean-Front-Prediction-Model.
Yuting Yang 0001, Xin Sun 0003, Junyu Dong, Kin-Man Lam 0001, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 MaskCD: A Remote Sensing Change Detection Network Based on Mask Classification
abstract
Change detection (CD) from remote sensing (RS) images using deep learning has been widely investigated in the literature. It is typically regarded as a pixelwise labeling task that aims to classify each pixel as changed or unchanged. Although per-pixel classification networks in encoder-decoder structures have shown dominance, they still suffer from imprecise boundaries and incomplete object delineation at various scenes. For high-resolution RS images, partly or totally changed objects are more worthy of attention rather than a single pixel. Therefore, we revisit the CD task from the mask prediction and classification perspective and propose mask classification-based CD (MaskCD) to detect changed areas by adaptively generating categorized masks from input image pairs. Specifically, it utilizes a cross-level change representation perceiver (CLCRP) to learn multiscale change-aware representations and capture spatiotemporal relations from encoded features by exploiting deformable multihead self-attention (DeformMHSA). Subsequently, a masked cross-attention-based detection transformers (MCA-DETRs) decoder is developed to accurately locate and identify changed objects based on masked cross-attention and self-attention (SA) mechanisms. It reconstructs the desired changed objects by decoding the pixelwise representations into learnable mask proposals and making final predictions from these candidates. Experimental results on five benchmark datasets demonstrate the proposed approach outperforms other state-of-the-art models. Codes and pretrained models are available online at:https://github.com/EricYu97/MaskCD.
Weikang Yu, Samiran Das, Xiao Xiang Zhu 0001, Pedram Ghamisi
IEEE Trans. Geosci. Remote. Sens.4
2024 MineNetCD: A Benchmark for Global Mining Change Detection on Remote Sensing Imagery
abstract
Monitoring land changes triggered by mining activities is crucial for industrial control, environmental management, and regulatory compliance, yet it poses significant challenges due to the vast and often remote locations of mining sites. Remote sensing technologies have increasingly become indispensable to detect and analyze these changes over time. We thus introduce MineNetCD, a comprehensive benchmark designed for global mining change detection using remote sensing imagery. The benchmark comprises three key contributions. First, we establish a global mining change detection dataset featuring more than 70k paired patches of bitemporal high-resolution remote sensing images and pixel-level annotations from 100 mining sites worldwide. Second, we develop a novel baseline model based on a change-aware fast Fourier transform (ChangeFFT) module, which enhances various backbones by leveraging essential spectrum components within features in the frequency domain and capturing the channelwise correlation of bitemporal feature differences to learn change-aware representations. Third, we construct a unified change detection (UCD) framework that currently integrates 20 change detection methods. This framework is designed for streamlined and efficient processing, using the cloud platform hosted by HuggingFace. Extensive experiments have been conducted to demonstrate the superiority of the proposed baseline model compared with 19 state-of-the-art change detection approaches. Empirical studies on modularized backbones comprehensively confirm the efficacy of different representation learners on change detection. This benchmark represents significant advancements in the field of remote sensing and change detection, providing a robust resource for future research and applications in global mining monitoring. Dataset and Codes are available via the link.
Weikang Yu, Richard Gloaguen, Xiao Xiang Zhu 0001, Pedram Ghamisi
IEEE Trans. Geosci. Remote. Sens.4
2024 RRSIS: Referring Remote Sensing Image Segmentation
abstract
Localizing desired objects from remote sensing images is of great use in practical applications. Referring image segmentation, which aims at segmenting out the objects to which a given expression refers, has been extensively studied in natural images. However, almost no research attention is given to this task of remote sensing imagery. Considering its potential for real-world applications, in this paper, we introduce referring remote sensing image segmentation (RRSIS) to fill in this gap and make some insightful explorations. Specifically, we create a new dataset, called RefSegRS, for this task, enabling us to evaluate different methods. Afterward, we benchmark referring image segmentation methods of natural images on the RefSegRS dataset and find that these models show limited efficacy in detecting small and scattered objects. To alleviate this issue, we propose a language-guided cross-scale enhancement (LGCE) module that utilizes linguistic features to adaptively enhance multi-scale visual features by integrating both deep and shallow features. The proposed dataset, benchmarking results, and the designed LGCE module provide insights into the design of a better RRSIS model. The dataset and code will be available at https://gitlab.lrz.de/ai4eo/reasoning/rrsis.
Zhenghang Yuan, Lichao Mou, Yuansheng Hua, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 A Real-Time Unsupervised Hyperspectral Band Selection via Spatial-Spectral Information Fusion-Based Downscaled Region
abstract
Information fusion plays a vital role in hyperspectral band selection as it enables the exploration of the spatial-spectral structure relationship present in bands of hyperspectral images (HSIs). However, the focus of most algorithms primarily lies in the processing of single-band vectors, leaving only a few algorithms to handle spatial-spectral features of HSIs in a complex and inefficient manner. To overcome these limitations, a real-time unsupervised hyperspectral band selection method via spatial-spectral information fusion-based downscaled region (SIFDR) is proposed in this study. In particular, this approach incorporates an energy constraint method for assigning band weights to each detected pixel and estimates band spectral information through average fusion. Furthermore, a band weak redundancy sorting method is introduced, which is based on spectral information peaks, thereby achieving complementary spectral information. By performing regional downscaling of the HSI, spatial-spectral information is effectively fused, resulting in a real-time entire process. To evaluate the effectiveness of the proposed algorithm, experiments were conducted on four hyperspectral datasets, including an ultrahigh-dimensional medical HSIs, which distinguishes itself from previous methods that are typically evaluated exclusively on remote sensing datasets. Comparative results with several state-of-the-art (SOTA) algorithms demonstrate that the proposed algorithm excellently accomplishes hyperspectral band selection tasks in real time. The code of SIFDR has been shared onhttps://github.com/zhangchenglong1116/SIFDR.
Lichao Mou, Xiangrong Zheng, Xiao Xiang Zhu 0001, Xiaopeng Ma
IEEE Trans. Geosci. Remote. Sens.5
2024 Few-Shot Object Detection in Remote Sensing: Lifting the Curse of Incompletely Annotated Novel Objects
abstract
Object detection is an essential and fundamental task in computer vision and satellite image processing. Existing deep learning methods have achieved impressive performance thanks to the availability of large-scale annotated datasets. Yet, in real-world applications the availability of labels is limited. In this context, few-shot object detection (FSOD) has emerged as a promising direction, which aims at enabling the model to detect novel objects with only few of them annotated. However, many existing FSOD algorithms overlook a critical issue: when an input image contains multiple novel objects and only a subset of them are annotated, the unlabeled objects will be considered as background during training. This can cause confusions and severely impact the model’s ability to recall novel objects. To address this issue, we propose a self-training-based FSOD (ST-FSOD) approach, which incorporates the self-training mechanism into the few-shot fine-tuning process. ST-FSOD aims to enable the discovery of novel objects that are not annotated, and take them into account during training. On the one hand, we devise a two-branch region proposal networks (RPN) to separate the proposal extraction of base and novel objects, On another hand, we incorporate the student-teacher mechanism into RPN and the region of interest (RoI) head to include those highly confident yet unlabeled targets as pseudo labels. Experimental results demonstrate that our proposed method outperforms the state-of-the- art in various FSOD settings by a large margin. The codes will be publicly available at https://github.com/zhu-xlab/ST-FSOD.
Fahong Zhang 0001, Yilei Shi, Zhitong Xiong, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 Co-Enhanced Global-Part Integration for Remote-Sensing Scene Classification
abstract
Remote sensing (RS) scene classification aims to classify remote sensing images with similar scene characteristics into one category. Plenty of RS images are complex in background, rich in content, and multi-scale in target, exhibiting the characteristics of both intra-class separation and inter-class convergence. Therefore, discriminative feature representations designed to highlight the differences between classes are the key to RS scene classification. Existing methods represent scene images by extracting either global context or discriminative part features from RS images. However, global-based methods often lack salient details in similar RS scenes, while part-based methods tend to ignore the relationships between local ground objects, thus weakening the discriminative feature representation. In this paper, we propose to combine global context and part-level discriminative features within a unified framework called CGINet for accurate RS scene classification. To be specific, we develop a light context-aware attention block (LCAB) to explicitly model the global context to obtain larger receptive fields and contextual information. A co-enhanced loss module (CELM) is also devised to encourage the model to actively locate discriminative parts for feature enhancement. In particular, CELM is only used during training and not activated during inference, which introduces less computational cost. Benefiting from LCAB and CELM, our proposed CGINet improves the discriminability of features, thereby improving classification performance. Comprehensive experiments over four benchmark datasets show that the proposed method achieves consistent performance gains over state-of-the-art RS scene classification methods.
Yichen Zhao, Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu, Xiao Xiang Zhu 0001, Lichao Mou
IEEE Trans. Geosci. Remote. Sens.5
2024 Hybrid Quantum-Classical Convolutional Neural Network Model for Image Classification
abstract
Image classification plays an important role in remote sensing. Earth observation (EO) has inevitably arrived in the big data era, but the high requirement on computation power has already become a bottleneck for analyzing large amounts of remote sensing data with sophisticated machine learning models. Exploiting quantum computing might contribute to a solution to tackle this challenge by leveraging quantum properties. This article introduces a hybrid quantum-classical convolutional neural network (QC-CNN) that applies quantum computing to effectively extract high-level critical features from EO data for classification purposes. Besides that, the adoption of the amplitude encoding technique reduces the required quantum bit resources. The complexity analysis indicates that the proposed model can accelerate the convolutional operation in comparison with its classical counterpart. The model's performance is evaluated with different EO benchmarks, including Overhead-MNIST, So2Sat LCZ42, PatternNet, RSI-CB256, and NaSC-TG2, through the TensorFlow Quantum platform, and it can achieve better performance than its classical counterpart and have higher generalizability, which verifies the validity of the QC-CNN model on EO data classification tasks.
Yilei Shi, Tobias Guggemos, Xiao Xiang Zhu 0001
IEEE Trans. Neural Networks Learn. Syst.4
2023 Large-Scale Land Cover Mapping with Fine-Grained Classes via Class-Aware Semi-Supervised Semantic Segmentation
abstract
Semi-supervised learning has attracted increasing attention in the large-scale land cover mapping task. However, existing methods overlook the potential to alleviate the class imbalance problem by selecting a suitable set of unlabeled data. Besides, in class-imbalanced scenarios, existing pseudo-labeling methods mostly only pick confident samples, failing to exploit the hard samples during training. To tackle these issues, we propose a unified Class-Aware Semi-Supervised Semantic Segmentation framework. The proposed framework consists of three key components. To construct a better semi-supervised learning dataset, we propose a class-aware unlabeled data selection method that is more balanced towards the minority classes. Based on the built dataset with improved class balance, we propose a Class-Balanced Cross Entropy loss, jointly considering the annotation bias and the class bias to re-weight the loss in both sample and class levels to alleviate the class imbalance problem. Moreover, we propose the Class Center Contrast method to jointly utilize the labeled and unlabeled data. Specifically, we decompose the feature embedding space using the ground truth and pseudo-labels, and employ the embedding centers for hard and easy samples of each class per image in the contrast loss to exploit the hard samples during training. Compared with state-of-the-art class-balanced pseudo-labeling methods, the proposed method improves the mean accuracy and mIoU by 4.28% and 1.70%, respectively, on the large-scale Sentinel-2 dataset with 24 land cover classes.
Runmin Dong, Lichao Mou, Mengxuan Chen, Xin-Yi Tong 0003, Shuai Yuan 0005, Lixian Zhang 0002, Juepeng Zheng, Xiao Xiang Zhu 0001, Haohuan Fu
ICCV9
2023 Semi-Supervised Deep Learning Representations in Earth Observation Based Forest Management
abstract
In this study, we examine the potential of several self-supervised deep learning models in predicting forest attributes and detecting forest changes using ESA Sentinel-1 and Sentinel-2 images. The performance of the proposed deep learning models is compared to established conventional machine learning approaches. Studied use-cases include mapping of forest disturbance (windthrown forests, snowload damages) using deep change vector analysis, forest height mapping using UNet+ based models, Momentum contrast and regression modeling. Study areas were represented by several boreal forest sites in Finland. Our results indicate that developed methods allow to achieve superior classification and prediction accuracies compared to traditional methodologies and mimimize the amount of necessary in-situ forestry data.
Oleg Antropov, Matthieu Molinier, Ridvan Salih Kuzu, Lloyd H. Hughes, Marc Rußwurm, Devis Tuia, Corneliu Octavian Dumitru, Shaojia Ge, Sudipan Saha, Xiao Xiang Zhu 0001
IGARSS10
2023 An Analysis of the Gap Between Hybrid and Real Data for Volcanic Deformation Detection
abstract
Recently deep learning models were applied to detect fast short-term volcanic deformations using interferometric synthetic aperture radar (InSAR) data. However, volcanic deformation detection is limited by the availability of real positive samples. In previous work, we used hybrid synthetic-real InSAR deformation maps set to train an InceptionResNet v2 model capable of detecting deformations down to 5 mm/year in real set. However, our model also reported false positive detections. One possible reason is the data distribution gap between the real and hybrid sets. In this paper, an experiment is conducted to analyze the gap between the hybrid and real sets that resulted in false positives. Three subsets of the fine-tuning set are created based on t-SNE analysis using different sampling strategies. The classification model is fine-tuned using these subsets. The results show that the strategy of removing only the most confusing examples and keeping the larger data set size reduces the false positive rate from 32.29% to 27.01%.
Teo Beker, Qian Song, Xiao Xiang Zhu 0001
IGARSS3
2023 Polyhedron-Based Graph Neural Network for Compact Building Model Reconstruction
abstract
Three-dimensional (3D) building models play a crucial role in shaping digital twin cities and enabling a wide range of urban applications. However, one challenge remains in obtaining a compact representation of buildings from remote sensing. This paper introduces a novel deep learning approach to reconstructing polygonal building models from LiDAR point clouds. Our method leverages a graph neural network to assemble the polyhedra generated through space partitioning, thereby formulating building surface reconstruction as a graph node classification problem. To facilitate network training, we construct a synthetic dataset by simulating aerial LiDAR point clouds on building surface meshes. Experimental results demonstrate the effectiveness of our method, achieving a polyhedral classification accuracy of 96.4%. Moreover, our approach offers high efficiency and interpretability through end-to-end optimization.
Zhaiyu Chen, Yilei Shi, Zhitong Xiong, Xiao Xiang Zhu 0001
IGARSS4
2023 Adaptive Bins for Monocular Height Estimation from Single Remote Sensing Images
abstract
Monocular height estimation is of great importance in generating 3D city models from single remote sensing images, while it is a challenging task due to the ill-posed nature of the problem. To address the issue, we propose to adopt adaptive bins (AdaBins) for the network design, which enhances the representation capability of the network with the classification-regression paradigm and the incorporation of local features and global context via a vision transformer encoder. Besides, to weaken the biases of the trained networks caused by the long-tailed nature of the dataset, a head-tail cut is conducted for different treatments of head and tail pixels. Experiments show that improvements are expected with the proposed network on the proposed GBH dataset.
Sining Chen, Yilei Shi, Zhitong Xiong, Xiao Xiang Zhu 0001
IGARSS4
2023 Exploring Distance-Aware Uncertainty Quantification for Remote Sensing Image Classification
abstract
Deep Learning models for classification often suffer from overconfidence, which naturally results in poor predictive uncertainty estimates. To overcome this, many calibration techniques have been established. These techniques operate on the labels or the output space of the network but ignore the input image space. A recently proposed approach considers the distances between different network inputs explicitly and theoretically propagates the distances through the network. The resulting predictive uncertainties of the model are then able to better reflect these distances. We test this approach in the context of remote sensing image classification for land use. To evaluate the predictive uncertainties, we set up an Out-of Distribution (OoD) detection framework based on class separation.
Christoph Koller, Peter Jung 0001, Xiao Xiang Zhu 0001
IGARSS3
2023 Roof Superstructure Detection from Aerial Imagery
abstract
Identifying suitable building roofs for the installation of photovoltaic (PV) systems is important to sustainable energy planning. However, most existing approaches neglect roof superstructures that can obstruct the installation of PV systems. In this research, we propose a novel method, which can help to deal with this issue by detecting roof superstructures from aerial imagery. Considering that semantic information about roof masks is also informative, we propose to first learns roof segmentation maps that are further used to learn roof superstructure maps. Experiments are conducted on Roof Information Dataset (RID). Our method outperforms the state-of-the-art methods both quantitatively and qualitatively.
Qingyu Li 0001, Sebastian Krapf, Lichao Mou, Yilei Shi, Xiao Xiang Zhu 0001
IGARSS5
2023 High Spatial Resolution for Crop Yield Prediction in Large Farming Systems: A Necessity or Additional Overhead
abstract
The availability of open-access satellite data and advancements in machine learning techniques has exhibited significant potential in crop yield prediction. In the context of large farming systems and county-level predictions, it is customary to rely on coarse-resolution satellite images. However, these images often lack the sufficient textural detail to accurately summarise spatial information. This research aims to evaluate the advantages of enhanced spatial resolution by conducting a comparative analysis between coarse-resolution, high-temporal-frequency MODIS data and relatively high-resolution, low-temporal-frequency Landsat data for predicting corn yield in the USA. We benchmark this comparison against several models in a spatial versus non-spatial input data context. Our results suggest that, the use of high-spatial resolution for county-level yield prediction in large farming systems is not beneficial and the models explored are unable to generalize well to drought-struck years.
Stella Ofori-Ampofo, Ridvan Salih Kuzu, Xiao Xiang Zhu 0001
IGARSS3
2023 Semi-Supervised Learning for Hyperspectral Images by Non Parametrically Predicting View AssignmentCRediT
abstract
Hyperspectral image (HSI) classification is gaining a lot of momentum in present time because of high inherent spectral information within the images. However, these images suffer from the problem of curse of dimensionality and usually require a large number samples for tasks such as classification, especially in supervised setting. Recently, to effectively train the deep learning models with minimal labelled samples, the unlabeled samples are also being leveraged in self-supervised and semi-supervised setting. In this work, we leverage the idea of semi-supervised learning to assist the discriminative self-supervised pretraining of the models. The proposed method takes different augmented views of the unlabeled samples as input and assigns them the same pseudo-label corresponding to the labelled sample from the downstream task. We train our model on two HSI datasets, anemly Houston dataset (from data fusion contest, 2013) and Pavia university dataset, and show that the proposed approach performs better than self-supervised approach and supervised training.
Shivam Pande, Nassim Ait Ali Braham, Yi Wang 0072, Conrad M. Albrecht, Biplab Banerjee, Xiao Xiang Zhu 0001
IGARSS6
2023 3d Point Cloud Simulation for Above-Ground Forest Biomass Estimation
abstract
In this paper, we proposed a new framework for 3D point cloud simulation for forest above-ground biomass estimation. It takes tree variables as input and automatically generates 30m by 30m scenes and simulates their corresponding Li-DAR point clouds. 2000 3D tree models of 10 species are generated, with which 5000 forest scenes representing four types of eco-regions are built. Their corresponding biomass are then calculated with allometric equations. We used the simulated tropical scenes (1000 samples) to test four classical machine learning models’ ability in biomass estimation from point clouds. Experimental results show that data augmentation is able to significantly boost test accuracy; self-supervised learning can improve the estimation results; among the four models, ResNet18 is the best baseline model which has achieved a R-square score of 0.6615, and reduced the root mean square error to 202 kg (mean biomass in our dataset is 2114kg).
Qian Song, Yuanyuan Wang 0002, Xiao Xiang Zhu 0001
IGARSS3
2023 A Dataset for Individual Tree Delineation from 3D Point Cloud data
abstract
LiDAR scanning data, which is able to acquire the vertical structures of forests, is of great potential in forest monitoring and biodiversity quantification. Besides, the derivation of some forest indices, such as biomass, relies on individual tree delineation (ITD). In this paper, we generated a dataset for individual tree delineation using LiDAR-derived point clouds. This dataset can be used to fairly compare different ITD methods and to develop deep learning algorithms for tree segmentation. The acquired LiDAR data consist of 0.94 billion points covering an area of about 31 km2in the Netherlands. We first used a rule-based algorithm to remove non-tree points. And then a mean shift clustering method is utilized to segment the points. Besides, we proposed a method that compares the highest point in the same cluster to evaluate the delineation results. In the future, the derived segmentation result will be compared with existing individual tree delineation algorithms.
Qian Song, Xiao Xiang Zhu 0001
IGARSS2
2023 Towards a Benchmark EO Semantic Segmentation Dataset for Uncertainty Quantification
abstract
In order to achieve the objective of accurate and reliable use of deep neural networks for Earth Observation in large-scale scene understanding and interpretation, a large and diverse dataset with proper quantification of uncertainty is required. In this work, we exemplify the lack of a benchmark dataset and present the progress of a novel benchmark dataset for uncertainty quantification of deep learning models in the classic problem of building segmentation from overhead imagery. We present a synthetic dataset where synthetic UAV images were rendered from 3D mesh models of Berlin, Germany. The building masks were extracted from precise LoD-2 building models of the same area. We compare and contrast the performances of baseline methods for semantic segmentation and various uncertainty quantification techniques on this dataset. The experiments show that U-Net is the most accurate model with mIoU of 0.812. Moreover, the Bayesian model is found to be the most reliable uncertainty quantification method on our dataset, with the least ECE.
Dawood Wasif, Yuanyuan Wang 0002, Muhammad Shahzad 0002, Rudolph Triebel, Xiao Xiang Zhu 0001
IGARSS5
2023 RSSOD-Bench: a Large-Scale Benchmark Dataset for Salient Object Detection in Optical Remote Sensing Imagery
abstract
We present the RSSOD-Bench dataset for salient object detection (SOD) in optical remote sensing imagery. While SOD has achieved success in natural scene images with deep learning, research in SOD for remote sensing imagery (RSSOD) is still in its early stages. Existing RSSOD datasets have limitations in terms of scale, and scene categories, which make them misaligned with real-world applications. To address these shortcomings, we construct the RSSOD-Bench dataset, which contains images from four different cities in the USA1. The dataset provides annotations for various salient object categories, such as buildings, lakes, rivers, highways, bridges, aircraft, ships, athletic fields, and more. The salient objects in RSSOD-Bench exhibit large-scale variations, cluttered backgrounds, and different seasons. Unlike existing datasets, RSSOD-Bench offers uniform distribution across scene categories. We benchmark 23 different state-of-the-art approaches from both the computer vision and remote sensing communities. Experimental results demonstrate that more research efforts are required for the RSSOD task.
Zhitong Xiong, Yanfeng Liu, Qi Wang 0009, Xiao Xiang Zhu 0001
IGARSS4
2023 Multi-Modal Multi-Task Learning for Semantic Segmentation of Land Cover Under Cloudy Conditions
abstract
The majority of the Earth’s surface is covered by clouds, causing optical images to suffer serious degradation of ground information. Synthetic Aperture Radar (SAR) images with the cloud-penetration capability could provide supplementary information to optical images. Thus, the fusion of optical and SAR image can remarkably improve the interpretation accuracy under cloudy conditions. In this paper, we propose to exploit related cloud removal task for accurate multi-modal semantic segmentation of land cover. Towards this goal, we develop an end-to-end learnable architecture which solves the tasks of cloud removal and semantic segmentation of land cover jointly. The cloud removal task encourages to learn knowledgeable features to overcome negative effects of semantic ambiguity. Our experiments show that the proposed algorithm can effectively improve the semantic segmentation accuracy under cloudy conditions.
Yilei Shi, Wen Yang 0001, Xiao Xiang Zhu 0001
IGARSS4
2023 DisasterNets: Embedding Machine Learning in Disaster Mapping
abstract
Disaster mapping is a critical task that often requires on-site experts and is time-consuming. To address this, a comprehensive framework is presented for fast and accurate recognition of disasters using machine learning, termed DisasterNets. It consists of two stages, space granulation and attribute granulation. The space granulation stage leverages supervised/semi-supervised learning, unsupervised change detection, and domain adaptation with/without source data techniques to handle different disaster mapping scenarios. Furthermore, the disaster database with the corresponding geographic information field properties is built by using the attribute granulation stage. The framework is applied to earthquake-triggered landslide mapping and large-scale flood mapping. The results demonstrate a competitive performance for high-precision, high-efficiency, and cross-scene recognition of disasters. To bridge the gap between disaster mapping and machine learning communities, we will provide an openly accessible tool based on DisasterNets. The framework and tool will be available at https://github.com/HydroPML/DisasterNets.
Qingsong Xu 0001, Yilei Shi, Xiao Xiang Zhu 0001
IGARSS3
2023 Overcoming Language Bias in Remote Sensing Visual Question Answering Via Adversarial Training
abstract
The Visual Question Answering (VQA) system offers a user-friendly interface and enables human-computer interaction. However, VQA models commonly face the challenge of language bias, resulting from the learned superficial correlation between questions and answers. To address this issue, in this study, we present a novel framework to reduce the language bias of the VQA for remote sensing data (RSVQA). Specifically, we add an adversarial branch to the original VQA framework. Based on the adversarial branch, we introduce two regularizers to constrain the training process against language bias. Furthermore, to evaluate the performance in terms of language bias, we propose a new metric that combines standard accuracy with the performance drop when incorporating question and random image information. Experimental results demonstrate the effectiveness of our method. We believe that our method can shed light on future work for reducing language bias on the RSVQA task.
Zhenghang Yuan, Lichao Mou, Xiao Xiang Zhu 0001
IGARSS3
2023 The Spatial-Temporal Role of Urban Green Space in Mitigating Land Surface Temperature in Chinese Megacities
abstract
The health impacts of extreme heat intensify, increasing deaths and illnesses. As the most direct and effective way to cool down, urban green space (UGS) has gained great attention. However, due to the limited availability of UGS resources, the critical point is how to increase the cooling effect of UGS in response to the context of climate warming. Since the different spatial distribution of UGS affects its ability to alleviate the urban heat island (UHI) effect, improving the equity of UGS will significantly alter the spatial distribution of UGS. In this study, we analyze UGS and land surface temperature (LST) data from two Chinese megacities between 2000 and 2020 to investigate the potential of enhancing equitable UGS distribution in mitigating the UHI. Our findings indicate that addressing inequities in UGS distribution contributes modestly to UHI mitigation, shedding new light on enhancing the cooling effect of UGS.
Qingsong Xu 0001, Xiao Xiang Zhu 0001
IGARSS3
2023 A Seq2seq-Based Forest Height Estimation for Zero-Baseline Repeat-Pass Polinsar Data
abstract
This paper proposes a forest height estimation method based on Seq2Seq models for zero-baseline repeat-pass PolInSAR acquisitions. In zero-baseline configuration, the polarimetric coherence of RMoG model can be refined as a real function varying with ground to volume ratio (polarization) since the phase information is negligible. The primary decorrelation in this real polarimetric coherence is derived from the temporal changes, which has an intuitive connection with forest height and can be further used for forest height estimation. An initial attempt is to estimate the sequence of polarimetric coherence and ground to volume ratio directly from the PolInSAR data and then extract the forest height from the estimated sequence based on the nonlinear functions originating from the RMoG model. However, large discrepancies between the estimated and the model-based sequences make it hard to retrieve the forest height accurately. Thereby, a trained Seq2Seq model is used so that the characteristics of the output sequence can be easily recognized by the model for forest height estimation. Experiments are conducted on PolInSAR data simulated by PolSARproSim+ and results indicate that forest height can be effectively extracted from zero-baseline repeat-pass data with the proposed method.
Yuanyuan Wang 0002, Xiao Xiang Zhu 0001
IGARSS3
2023 DDM-Former: Global Ocean Wind Speed Retrieval with Transformer Networks
abstract
As a novel remote sensing technique, GNSS reflectometry (GNSS-R) opens a new era of retrieving Earth surface parameters. Several studies employ the combination of deep learning and GNSS-R observable delay-Doppler maps (DDMs) to generate ocean wind speed estimation. Unlike these methods that often use convolutional neural networks (CNNs) with inductive bias, we proposed a Transformer-based model, named DDM-Former, to exploit fine-grained delay-Doppler correlation independently. Our model is evaluated on the Cyclone GNSS (CYGNSS) version 3.0 dataset and shown to outperform the other retrieval methods.
Daixin Zhao, Konrad Heidler, Milad Asgarimehr, Caroline Arnold, Tianqi Xiao, Jens Wickert, Xiao Xiang Zhu 0001, Lichao Mou
IGARSS7
2023 Exploring Geometric Deep Learning for Precipitation Nowcasting
abstract
Precipitation nowcasting (up to a few hours) remains a challenge due to the highly complex local interactions that need to be captured accurately. Convolutional Neural Networks rely on convolutional kernels convolving with grid data and the extracted features are trapped by limited receptive field, typically expressed in excessively smooth output compared to ground truth. Thus they lack the capacity to model complex spatial relationships among the grids. Geometric deep learning aims to generalize neural network models to non-Euclidean domains. Such models are more flexible in defining nodes and edges and can effectively capture dynamic spatial relationship among geographical grids. Motivated by this, we explore a geometric deep learning-based temporal Graph Convolutional Network (GCN) for precipitation nowcasting. The adjacency matrix that simulates the interactions among grid cells is learned automatically by minimizing the L1 loss between prediction and ground truth pixel value during the training procedure. Then, the spatial relationship is refined by GCN layers while the temporal information is extracted by 1D convolution with various kernel lengths. The neighboring information is fed as auxiliary input layers to improve the final result. We test the model on sequences of radar reflectivity maps over the Trento/Italy area. The results show that GCNs improves the effectiveness of modeling the local details of the cloud profile as well as the prediction accuracy by achieving decreased error measures.
Shan Zhao 0007, Sudipan Saha, Zhitong Xiong, Niklas Boers, Xiao Xiang Zhu 0001
IGARSS5
2023 Function Assignment of Plastics based on Hyperspectral Satellite Images and High-Resolution Data Using Deep Learning Algorithms
abstract
Plastic pollution is becoming an increasingly prominent problem and the function of plastics determines whether they need to be recycled or not. In order to explore the possibility of using satellite imagery to classify the functionality of plastics, this study proposes a two-stage workflow: firstly, a classification map is obtained based on hyperspectral satellite imagery to generate plastic types, and then using these identified plastic coverage areas, a deep learning algorithm is used to assign functionality to these classified plastic areas based on sentinel-2 imagery. By comparing five leading-edge image classification models, classification accuracies of up to 74% were achieved, demonstrating the feasibility of using deep learning models trained on satellite images to identify plastic features.
Shanyu Zhou, Lichao Mou, Lixian Zhang 0002, Yuansheng Hua, Hermann Kaufmann 0001, Xiao Xiang Zhu 0001
IGARSS6
2023 GEO-Bench: Toward Foundation Models for Earth Monitoring
abstract
Recent progress in self-supervision has shown that pre-training large neural networks on vast amounts of unsupervised data can lead to substantial increases in generalization to downstream tasks. Such models, recently coined foundation models, have been transformational to the field of natural language processing.Variants have also been proposed for image data, but their applicability to remote sensing tasks is limited.To stimulate the development of foundation models for Earth monitoring, we propose a benchmark comprised of six classification and six segmentation tasks, which were carefully curated and adapted to be both relevant to the field and well-suited for model evaluation. We accompany this benchmark with a robust methodology for evaluating models and reporting aggregated results to enable a reliable assessment of progress. Finally, we report results for 20 baselines to gain information about the performance of existing models.We believe that this benchmark will be a driver of progress across a variety of Earth monitoring tasks.
Alexandre Lacoste, Nils Lehmann, Pau Rodríguez, Evan D. Sherwin, Hannah Kerner, Björn Lütjens, Jeremy Irvin, David Dao, Hamed Alemohammad, Alexandre Drouin, Mehmet Gunturkun, Gabriel Huang, David Vázquez 0001, Dava J. Newman, Yoshua Bengio, Stefano Ermon, Xiao Xiang Zhu 0001
NeurIPS17
2023 MedSat: A Public Health Dataset for England Featuring Medical Prescriptions and Satellite Imagery
abstract
As extreme weather events become more frequent, understanding their impact on human health becomes increasingly crucial. However, the utilization of Earth Observation to effectively analyze the environmental context in relation to health remains limited. This limitation is primarily due to the lack of fine-grained spatial and temporal data in public and population health studies, hindering a comprehensive understanding of health outcomes. Additionally, obtaining appropriate environmental indices across different geographical levels and timeframes poses a challenge. For the years 2019 (pre-COVID) and 2020 (COVID), we collected spatio-temporal indicators for all Lower Layer Super Output Areas in England. These indicators included: i) 111 sociodemographic features linked to health in existing literature, ii) 43 environmental point features (e.g., greenery and air pollution levels), iii) 4 seasonal composite satellite images each with 11 bands, and iv) prescription prevalence associated with five medical conditions (depression, anxiety, diabetes, hypertension, and asthma), opioids and total prescriptions. We combined these indicators into a single MedSat dataset, the availability of which presents an opportunity for the machine learning community to develop new techniques specific to public health. These techniques would address challenges such as handling large and complex data volumes, performing effective feature engineering on environmental and sociodemographic factors, capturing spatial and temporal dependencies in the models, addressing imbalanced data distributions, developing novel computer vision methods for health modeling based on satellite imagery, ensuring model explainability, and achieving generalization beyond the specific geographical region.
Sanja Scepanovic, Ivica Obadic, Sagar Joglekar 0001, Laura Giustarini, Cristiano Nattero, Daniele Quercia, Xiao Xiang Zhu 0001
NeurIPS7
2023 SurroundNet: Towards effective low-light image enhancement
Fei Zhou 0007, Xin Sun 0003, Junyu Dong, Xiao Xiang Zhu 0001
Pattern Recognit.4
2023 Deep Learning for Subtle Volcanic Deformation Detection With InSAR Data in Central Volcanic Zone
abstract
Subtle volcanic deformations point to volcanic activities, and monitoring them helps predict eruptions. Today, it is possible to remotely detect volcanic deformation in mm/year scale thanks to advances in interferometric Synthetic Aperture Radar (InSAR). This paper proposes a framework based on a deep learning model to automatically discriminate subtle volcanic deformations from other deformation types in five-year-long InSAR stacks. Models are trained on a synthetic training set. To better understand and improve the models, explainable AI analyses are performed. In initial models, gradient-weighted Class Activation Mapping (Grad-CAM) linked new-found patterns of slope processes and salt lake deformations to false-positive detections. The models are then improved by fine-tuning with a hybrid synthetic-real data, and additional performance is extracted by low-pass spatial filtering of the real test set. T-SNE latent feature visualization confirmed the similarity and shortcomings of the fine-tuning set, highlighting the problem of elevation components in residual tropospheric noise. After fine-tuning, all the volcanic deformations are detected, including the smallest one, Lazufre, deforming 5 mm/year. The first time confirmed deformation of Cerro El Condor is observed, deforming 9.9-17.5 mm/year. Finally, sensitivity analysis uncovered the model’s minimal detectable deformation of 2 mm/year.
Teo Beker, Homa Ansari, Sina Montazeri, Qian Song, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 Deep Saliency Smoothing Hashing for Drone Image Retrieval
abstract
Deep hashing algorithms are widely exploited in retrieval tasks due to its low storage and retrieval efficiency. Most of which focus on global feature learning, whilst neglecting local fine-grained features and saliency information for drone images. In this paper, we tackle these dilemmas with a novelDeep Saliency Smoothing Hashing(DSSH) algorithm, which can leverage saliency capture mechanism, distribution smoothing term, global features and local fine-grained features to learn effective hash codes for drone image retrieval. The DSSH algorithm first designs information extraction module to capture global features and local fine-grained features for drone images. Meanwhile, a saliency capture module is proposed to perform information interaction attention and visual enhancement attention, which can capture the saliency area of drone images effectively. On top of the two paths, a novel objective function is designed to preserve the similarity of hash codes, smooth the distribution of drone image datasets and reduce the quantization errors between hash codes and hash-like codes concurrently. Extensive experiments on the Drone Action Dataset and ERA Drone Dataset demonstrate that the DSSH algorithm can further improve the retrieval performance compared to other deep hashing algorithms.
Yaxiong Chen, Lichao Mou, Pu Jin, Shengwu Xiong 0001, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.6
2023 HTC-DC Net: Monocular Height Estimation From Single Remote Sensing Images
abstract
3D geo-information is of great significance for understanding the living environment; however, 3D perception from remote sensing data, especially on a large scale, is restricted, mainly due to the high costs of 3D sensors such as LiDAR. To tackle this problem, we propose a method for monocular height estimation from optical imagery, which is currently one of the richest sources of remote sensing data. As an ill-posed problem, monocular height estimation requires well-designed networks for enhanced representations to improve performance. Moreover, the distribution of height values is long-tailed with the low-height pixels, e.g., the background, as the head, and thus trained networks are usually biased and tend to underestimate building heights. To solve the problems, instead of formalizing the problem as a regression task, we propose HTC-DC Net following the classification-regression paradigm, with the head-tail cut (HTC) and the distribution-based constraints (DCs) as the main contributions. HTC-DC Net is composed of the backbone network as the feature extractor, the HTC-AdaBins module, and the hybrid regression process. The HTC-AdaBins module serves as the classification phase to determine bins adaptive to each input image. It is equipped with a vision transformer encoder to incorporate local context with holistic information and involves an HTC to address the long-tailed problem in monocular height estimation for balancing the performances of foreground and background pixels. The hybrid regression process does the regression via the smoothing of bins from the classification phase, which is trained via DCs. The proposed network is tested on datasets of different resolutions, namely, DFC19 (1.3 m) and GBH (3 m). Experimental results show the superiority of the proposed network over existing methods by large margins. Extensive ablation studies demonstrate the effectiveness of each design component. Codes and trained models are published at https://github.com/zhu-xlab/HTC-DC-Net.
Sining Chen, Yilei Shi, Zhitong Xiong, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.4
2023 A Deep Active Contour Model for Delineating Glacier Calving Fronts
abstract
Choosing how to encode a real-world problem as a machine learning task is an important design decision in machine learning. The task of glacier calving front modeling has often been approached as a semantic segmentation task. Recent studies have shown that combining segmentation with edge detection can improve the accuracy of calving front detectors. Building on this observation, we completely rephrase the task as a contour tracing problem and propose a model for explicit contour detection that does not incorporate any dense predictions as intermediate steps. The proposed approach, called “Charting Outlines by Recurrent Adaptation” (COBRA), combines Convolutional Neural Networks (CNNs) for feature extraction and active contour models for the delineation. By training and evaluating on several large-scale datasets of Greenland’s outlet glaciers, we show that this approach indeed outperforms the aforementioned methods based on segmentation and edge-detection. Finally, we demonstrate that explicit contour detection has benefits over pixel-wise methods when quantifying the models’ prediction uncertainties. The project page containing the code and animated model predictions can be found at https://khdlr.github.io/COBRA/.
Konrad Heidler, Lichao Mou, Erik Loebel, Mirko Scheinert, Sébastien Lefèvre, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.6
2023 AdaptMatch: Adaptive Matching for Semisupervised Binary Segmentation of Remote Sensing Images
abstract
There are various binary semantic segmentation tasks in remote sensing (RS) that aim to extract the foreground areas of interest, such as buildings and roads, from the background in satellite images. In particular, semi-supervised learning, which can use limited labeled data to guide a large amount of unlabeled data for model training, can significantly promote the fast applications of these tasks in practice. However, due to the predominance of the background in RS images, the foreground only accounts for a small proportion of the pixels. It poses a challenge: models are biased toward the majority class of the background, leading to poor performance on the minority class of the foreground. To address this issue, this paper proposes a novel and effective semi-supervised learning framework, Adaptive Matching (AdaptMatch), for RS binary segmentation. AdaptMatch calculates individual and adaptive thresholds of the foreground and background based on their convergence difficulty in an online manner at the training stage; the adaptive thresholds are then used to select the high-confidence pseudo-labeled data of the two classes for model self-training in turn. Extensive experiments are conducted on two widely-studied RS binary segmentation tasks, building footprint extraction and road extraction, to demonstrate the effectiveness and generalizability of the proposed method. The results show that the proposed AdaptMatch achieves superior performance compared with some state-of-the-art semi-supervised methods in RS binary segmentation tasks. The codes will be publicly available at https://github.com/zhu-xlab/AdaptMatch.
Wei Huang 0068, Yilei Shi, Zhitong Xiong, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.4
2023 THE Benchmark: Transferable Representation Learning for Monocular Height Estimation
abstract
Generating 3D city models rapidly is crucial for many applications. Monocular height estimation is one of the most efficient and timely ways to obtain large-scale geometric information. However, existing works focus primarily on training and testing models using unbiased datasets, which does not align well with real-world applications. Therefore, we propose a new benchmark dataset to study the transferability of height estimation models in a cross-dataset setting. To this end, we first design and construct a large-scale benchmark dataset for cross-dataset transfer learning on the height estimation task. This benchmark dataset includes a newly proposed large-scale synthetic dataset, a newly collected real-world dataset, and four existing datasets from different cities. Next, a new experimental protocol,few-shot cross-dataset transfer, is designed. Furthermore, in this paper, we propose a scale-deformable convolution module to enhance the window-based Transformer for handling the scale-variation problem in the height estimation task. Experimental results have demonstrated the effectiveness of the proposed methods in traditional and cross-dataset transfer settings. The datasets and codes are publicly available at https://mediatum.ub.tum.de/1662763 and https://thebenchmarkh.github.io/.
Zhitong Xiong, Wei Huang 0068, Jingtao Hu, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.4
2023 UCDFormer: Unsupervised Change Detection Using a Transformer-Driven Image Translation
abstract
Change detection (CD) by comparing two bi-temporal images is a crucial task in remote sensing. With the advantages of requiring no cumbersome labeled change information, unsupervised CD has attracted extensive attention in the community. However, existing unsupervised CD approaches rarely consider the seasonal and style differences incurred by the illumination and atmospheric conditions in multi-temporal images. To this end, we propose a change detection with domain shift setting for remote sensing images. Furthermore, we present a novel unsupervised CD method using a light-weight transformer, called UCDFormer. Specifically, a transformer-driven image translation composed of a light-weight transformer and a domain-specific affinity weight is first proposed to mitigate domain shift between two images with real-time efficiency. After image translation, we can generate the difference map between the translated before-event image and the original after-event image. Then, a novel reliable pixel extraction module is proposed to select significantly changed/unchanged pixel positions by fusing the pseudo change maps of fuzzy c-means clustering and adaptive threshold. Finally, a binary change map is obtained based on these selected pixel pairs and a binary classifier. Experimental results on different unsupervised CD tasks with seasonal and style changes demonstrate the effectiveness of the proposed UCDFormer. For example, compared with several other related methods, UCDFormer improves performance on the Kappa coefficient by more than 12%. In addition, UCDFormer achieves excellent performance for earthquake-induced landslide detection when considering large-scale applications. The code is available at https://github.com/zhu-xlab/UCDFormer.
Qingsong Xu 0001, Yilei Shi, Jianhua Guo 0002, Chaojun Ouyang, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 Universal Domain Adaptation for Remote Sensing Image Scene Classification
abstract
The domain adaptation (DA) approaches available to date are usually not well suited for practical DA scenarios of remote sensing image classification since these methods (such as unsupervised DA) rely on rich prior knowledge about the relationship between label sets of source and target domains, and source data are often not accessible due to privacy or confidentiality issues. To this end, we propose a practical universal DA (UniDA) setting for remote sensing image scene classification that requires no prior knowledge on the label sets. Furthermore, a novel UniDA method without source data is proposed for cases when the source data are unavailable. The architecture of the model is divided into two parts: the source data generation stage and the model adaptation stage. The first stage estimates the conditional distribution of source data from the pretrained model using the knowledge of class separability in the source domain and then synthesizes the source data. With this synthetic source data in hand, it becomes a UniDA task to classify a target sample correctly if it belongs to any category in the source label set or mark it as “unknown” otherwise. In the second stage, a novel transferable weight that distinguishes the shared and private label sets in each domain promotes the adaptation in the automatically discovered shared label set and recognizes the “unknown” samples successfully. Empirical results show that the proposed model is effective and practical for remote sensing image scene classification, regardless of whether the source data are available or not. The code is available athttps://github.com/zhu-xlab/UniDA.
Qingsong Xu 0001, Yilei Shi, Xin Yuan 0002, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.4
2023 Pseudo Features-Guided Self-Training for Domain Adaptive Semantic Segmentation of Satellite Images
abstract
Semantic segmentation is a fundamental and crucial task that is of great importance to real-world satellite image-based applications. Yet a widely acknowledged issue that occurs when applying the semantic segmentation models to unseen scenery is that the model will perform much poorer than when it was applied to scenery similar to the training data. This phenomenon is usually termed as the domain shift problem. To tackle it, this article presents a self-training-based unsupervised domain adaptation (UDA) method. Different from the previous self-training approaches which focus on rectifying and improving the quality of the pseudo labels, we instead seek to exploit feature-level relation among neighboring pixels to structure and regularize the prediction of the adapted model. Based on the assumption that spatial topological relation is maintained despite the impact of the domain shift, we propose a novel self-training mechanism to perform DA by exploiting local relation in the feature space spanned by the teacher model, from which the pseudo labels are generated. Quantitative experiments on four different public benchmarks demonstrate that the proposed method can outperform the other UDA methods. Besides, analytical experiments also intuitively verify the proposed assumption. Codes will be publicly available athttps://github.com/zhu-xlab/PFST.
Fahong Zhang 0001, Yilei Shi, Zhitong Xiong, Wei Huang 0068, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 Adaptive Morphology Filter: A Lightweight Module for Deep Hyperspectral Image Classification
abstract
Deep neural network models significantly outperform classical algorithms in the hyperspectral image (HSI) classification task. These deep models improve generalization but incur significant computational demands. This article endeavors to alleviate the computational distress in a depthwise manner through the use of morphological operations. We propose the adaptive morphology filter (AMF) to effectively extract spatial features like the conventional depthwise convolution layer. Furthermore, we reparameterize AMF into its equivalent form, i.e., a traditional binary morphology filter, which drastically reduces the number of parameters in the inference phase. Finally, we stack multiple AMFs to achieve a large receptive field and construct a lightweight AMNet for classifying HSIs. It is noteworthy that we prove the deep stack of depthwise AMFs to be equivalent to structural element decomposition. We test our model on five benchmark datasets. Experiments show that our approach outperforms state-of-the-art methods with fewer parameters (${\approx }10 k$). The codes will be publicly available athttps://github.com/zhu-xlab/Adaptive-Morphology-Filter.
Fei Zhou 0007, Xin Sun 0003, Chengze Sun, Junyu Dong, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 ReforesTree: A Dataset for Estimating Tropical Forest Carbon Stock with Deep Learning and Aerial Imagery
abstract
Forest biomass is a key influence for future climate, and the world urgently needs highly scalable financing schemes, such as carbon offsetting certifications, to protect and restore forests. Current manual forest carbon stock inventory methods of measuring single trees by hand are time, labour, and cost intensive and have been shown to be subjective. They can lead to substantial overestimation of the carbon stock and ultimately distrust in forest financing. The potential for impact and scale of leveraging advancements in machine learning and remote sensing technologies is promising, but needs to be of high quality in order to replace the current forest stock protocols for certifications. In this paper, we present ReforesTree, a benchmark dataset of forest carbon stock in six agro-forestry carbon offsetting sites in Ecuador. Furthermore, we show that a deep learning-based end-to-end model using individual tree detection from low cost RGB-only drone imagery is accurately estimating forest carbon stock within official carbon offsetting certification standards. Additionally, our baseline CNN model outperforms state-of-the-art satellite-based forest biomass and carbon stock estimates for this type of small-scale, tropical agro-forestry sites. We present this dataset to encourage machine learning research in this area to increase accountability and transparency of monitoring, verification and reporting (MVR) in carbon offsetting projects, as well as scaling global reforestation financing through accurate remote sensing.
Gyri Reiersen, David Dao, Björn Lütjens, Konstantin Klemmer, Kenza Amara, Attila Steinegger, Ce Zhang 0001, Xiao Xiang Zhu 0001
AAAI8
2022 Peaks Fusion assisted Early-stopping Strategy for Overhead Imagery Segmentation with Noisy Labels
abstract
Automatic label generation systems, which are capable to generate huge amounts of labels with limited human efforts, enjoy lots of potential in the deep learning era. These easy-to-come-by labels inevitably bear label noises due to a lack of human supervision and can bias model training to some inferior solutions. However, models can still learn some plausible features, before they start to overfit on noisy patterns. Inspired by this phenomenon, we propose a new Peaks fusion assisted EArly-Stopping (PEAS) approach for imagery segmentation with noisy labels, which is mainly composed of two parts. First, a fitting based early-stopping criterion is used to detect the turning phase from which models are about to mimic noise details. After that, a peaks fusion strategy is applied to select reliable models in the detection zone to generate final fusion results. Here, validation accuracies are utilized as indicators in model selection. The proposed method was evaluated on New York City dataset whose labels were automatically collected by a rule-based label generation system, thus noisy to some extent due to a lack of human supervision. The experimental results showed that the proposed PEAS method can achieve both promising statistical and visual results when trained with noisy labels.
Chenying Liu 0001, Conrad M. Albrecht, Yi Wang 0072, Xiao Xiang Zhu 0001
IEEE Big Data4
2022 Deep Semantic Model Fusion for Ancient Agricultural Terrace Detection
abstract
Discovering ancient agricultural terraces in desert regions is important for the monitoring of long-term climate changes on the Earth’s surface. However, traditional ground surveys are both costly and limited in scale. With the increasing accessibility of aerial and satellite data, machine learning techniques bear large potential for the automatic detection and recognition of archaeological landscapes. In this paper, we propose a deep semantic model fusion method for ancient agricultural terrace detection. The input data includes aerial images and LiDAR generated terrain features in the Negev desert. Two deep semantic segmentation models, namely DeepLabv3+ and UNet, with EfficientNet backbone, are trained and fused to provide segmentation maps of ancient terraces and walls. The proposed method won the first prize in the International AI Archaeology Challenge. Codes are available at https://github.com/wangyi111/international-archaeologyai-challenge.
Yi Wang 0072, Chenying Liu 0001, Arti Tiwari, Micha Silver, Arnon Karnieli, Xiao Xiang Zhu 0001, Conrad M. Albrecht
IEEE Big Data6
2022 DAPHNE: An Open and Extensible System Infrastructure for Integrated Data Analysis Pipelines
Patrick Damme, Marius Birkenbach, Constantinos Bitsakos, Matthias Boehm 0001, Philippe Bonnet, Florina M. Ciorba, Mark Dokter, Pawel Dowgiallo, Ahmed Eleliemy, Christian Färber, Georgios I. Goumas, Dirk Habich, Niclas Hedam, Marlies Hofer, Kevin Innerebner, Vasileios Karakostas, Roman Kern, Tomaz Kosar, Alexander Krause 0001, Daniel Krems, Andreas Laber, Wolfgang Lehner, Eric Mier, Marcus Paradies, Bernhard Peischl, Gabrielle Poerwawinata, Stratos Psomadakis, Tilmann Rabl, Piotr Ratuszniak, Pedro Silva 0011, Nikolai Skuppin, Andreas Starzacher, Benjamin Steinwender, Ilin Tolovski, Pinar Tözün, Wojciech Ulatowski, Yuanyuan Wang 0002, Izajasz P. Wrosz, Ales Zamuda, Ce Zhang 0001, Xiao Xiang Zhu 0001
CIDR42
2022 DynamicEarthNet: Daily Multi-Spectral Satellite Dataset for Semantic Change Segmentation
abstract
Earth observation is a fundamental tool for monitoring the evolution of land use in specific areas of interest. Observing and precisely defining change, in this context, requires both time-series data and pixel-wise segmentations. To that end, we propose the DynamicEarthNet dataset that consists of daily, multi-spectral satellite observations of 75 selected areas of interest distributed over the globe with imagery from Planet Labs. These observations are paired with pixel-wise monthly semantic segmentation labels of 7 land use and land cover (LULC) classes. DynamicEarthNet is the first dataset that provides this unique combination of daily measurements and high-quality labels. In our experiments, we compare several established baselines that either utilize the daily observations as additional training data (semi-supervised learning) or multiple observations at once (spatio-temporal learning) as a point of reference for future research. Finally, we propose a new evaluation metric SCS that addresses the specific challenges associated with time-series semantic change segmentation. The data is available at: https://mediatum.ub.tum.de/1650201.
Aysim Toker, Lukas Kondmann, Mark Weber, Marvin Eisenberger, Andrés Camero, Jingliang Hu, Ariadna Pregel Hoderlein, Çaglar Senaras, Tim Davis 0001, Daniel Cremers, Giovanni Marchisio, Xiao Xiang Zhu 0001, Laura Leal-Taixé
CVPR12
2022 Doubly Deformable Aggregation of Covariance Matrices for Few-Shot Segmentation
Zhitong Xiong, Haopeng Li 0001, Xiao Xiang Zhu 0001
ECCV (20)3
2022 Multi-Sensor Time Series Cloud Removal Fusing Optical and SAR Satellite Information
abstract
On average, about half of all optical satellite data observing Earth is covered by haze or clouds. These atmospheric disturbances hinder the ongoing observation of our planet and prevent the seamless application of established remote sensing methods. Accordingly, to allow for an ongoing monitoring of Earth, approaches to reconstruct optical space-borne observations are required. This work introduces a new data set, SEN12MS-CR-TS, for the purpose of multi-sensor time series cloud removal. SEN12MS-CR-TS consists of co-registered radar and optical satellite data, featuring a se-quence of bi-weekly observations throughout an entire year. Finally, we demonstrate the usability of our novel data set by developing a new multi-sensor time-series cloud removal ar-chitecture. We are positive that our curated data set as well as the proposed model will advance future research in satellite image reconstruction and benefit the expanding adaptation of global and all-weather remote sensing applications.
Patrick Ebel 0002, Yajin Xu, Michael Schmitt 0003, Xiao Xiang Zhu 0001
IGARSS4
2022 Bounding Box Regression Network for Building Height Retrieval Using a Single SAR Image
abstract
In this paper, we propose a bounding box regression network for building height retrieval using a single TerraSAR - X stripmap image. The proposed network employs building footprints from GIS data and exploits the location relationship between a building's footprint and its bounding box, enabling fast computation. Experimental results over Rotterdam show that the proposed network can reduce the computation cost significantly while keeping the height accuracy of individual buildings compared to a Faster R-CNN based method.
Yao Sun 0005, Lichao Mou, Yuanyuan Wang 0002, Xiao Xiang Zhu 0001
IGARSS4
2022 Monitoring Urban Forests from Auto-Generated Segmentation MAPS
abstract
We present and evaluate a weakly-supervised methodology to quantify the spatiotemporal distribution of urban forests based on remotely sensed data with close-to-zero human interaction. Successfully training machine learning models for semantic segmentation typically depends on the availability of high-quality labels. We evaluate the benefit of high-resolution, three-dimensional point cloud data (LiDAR) as source of noisy labels in order to train models for the localization of trees in orthophotos. As proof of concept we sense Hurricane Sandy's impact on urban forests in Coney Island, New York City (NYC) and reference it to less impacted urban space in Brooklyn, NYC.
Conrad M. Albrecht, Chenying Liu 0001, Yi Wang 0072, Levente J. Klein, Xiao Xiang Zhu 0001
IGARSS5
2022 Explainability Analysis of CNN in Detection of Volcanic Deformation Signal
abstract
With improvement in the processing of synthetic aperture radar interferometry (InSAR) data, the detection of long-term volcanic deformations becomes possible. While deep learning (DL) models are considered black-box models, challenging to debug, the advances in explainable AI (XAI) help understand the model and how it makes decisions. In this paper, the model is trained on synthetic InSAR velocity maps to detect slow, sustained deformations. XAI tools, including Grad-CAM and t-SNE, are utilized for understanding and improving the trained model. Grad-CAM helps identify the slope-induced signal and salt lake patterns responsible for the model’s mis-classifications. T-SNE feature representation visualizations are used to estimate data sets and model class separation ability. Additionally, a sensitivity analysis shows the model performance with different intensity deformation data and uncovers the minimal detectable deformations of 1 cm cumulative deformation over five years.
Teo Beker, Homa Ansari, Sina Montazeri, Qian Song, Xiao Xiang Zhu 0001
IGARSS5
2022 Self Supervised Learning for Few Shot Hyperspectral Image Classification
abstract
Deep learning has proven to be a very effective approach for Hyperspectral Image (HSI) classification. However, deep neural networks require large annotated datasets to generalize well. This limits the applicability of deep learning for HSI classification, where manually labelling thousands of pixels for every scene is impractical. In this paper, we propose to leverage Self Supervised Learning (SSL) for HSI classification. We show that by pre-training an encoder on unlabeled pixels using Barlow-Twins, a state-of-the-art SSL algorithm, we can obtain accurate models with a handful of labels. Experimental results demonstrate that this approach significantly outperforms vanilla supervised learning.
Nassim Ait Ali Braham, Lichao Mou, Jocelyn Chanussot, Julien Mairal, Xiao Xiang Zhu 0001
IGARSS5
2022 Siamese Attention U-Net for Multi-Class Change Detection
abstract
Recent developments in deep learning have pushed the capabilities of pixel-wise change detection. This work introduces the winning solution of the DynamicEarthNet Weakly-Supervised Multi-Class Change Detection Challenge held at the EARTHVISION Workshop in CVPR 2021. The proposed approach is a pixel-wise change detection network coined Siamese Attention U-Net that incorporates attention mechanisms in the Siamese U-Net architecture. Moreover, this work finds the location of the attention mechanism within the network is crucial in achieving higher performance. Positioning the attention blocks in the up-sample path of the decoder filters noisy lower resolution features and allows for more fine-grained outputs. The impact of architectural changes, alongside training strategies such as semi-supervised learning are also evaluated on the DynamicEarthNet Challenge dataset.11Code is available at: https://github.com/solcummings/earthvision2021-weakly-supervised.
Sol Cummings, Lukas Kondmann, Xiao Xiang Zhu 0001
IGARSS3
2022 Earth Observation Data Classification with Quantum-Classical Convolutional Neural Network
abstract
Due to the rapid growth of earth observation (EO) data and the complexity of machine learning models, the high requirement on the computation power for EO data analysis becomes a bottleneck. Exploiting quantum computing might tackle this challenge in the future. In this paper, we present a hybrid quantum-classical convolutional neural network (QC-CNN) to classify EO data which can accelerate feature extraction compared with its classical counterpart and handle multi-category classification tasks with reduced quantum resources. The model's validity is verified with the Overhead-MNIST dataset through the TensorFlow Quantum platform.
Yilei Shi, Xiao Xiang Zhu 0001
IGARSS3
2022 Robust Distribution-Shift Aware Sar-Optical data Fusion for Multi-Label Scene Classification
abstract
Out-of-distribution (OOD) detection is an emerging research topic in remote sensing where existing works focus on single sensor analysis. However, many remote sensing works use multi-modal data to benefit from different characteristics of the sensors. Data that is in-domain for one sensor may be OOD for another sensor. In this work, we address such a scenario focusing on Synthetic Aperture Radar (SAR) and optical data fusion for multi-label scene classification. Besides data distribution shifts caused by unknown classes and snow, we also consider cases where only one modality is affected. Optical images acquired with significant cloud coverage are considered as OOD, while their corresponding SAR images can be in-distribution. We propose a weighted feature propagation strategy based on the in-distribution probabilities of the single modalities. We show, that we not only improve the prediction performance on the cloudy samples but also receive a higher predictive uncertainty when both modalities are OOD.
Jakob Gawlikowski, Sudipan Saha, Julia Niebling, Xiao Xiang Zhu 0001
IGARSS4
2022 Analysing the Interactions Between Training Dataset Size, Label Noise and Model Performance in Remote Sensing Data
abstract
In this work we analyse how training datasize affects the ability of a deep neural network to deal with noisy training labels in a semantic segmentation task with labels from OpenStreetMap. To this end, several versions of the training set were created by introducing varying amounts of label noise, and a model was then trained on subsets of varying size of these versions. The results indicate that the relationship between noise level and model performance is largely independent of the datasize except for very small datasizes where adding label noise has an even more deteriorating effect than usual.
Jonas Gütter, Julia Niebling, Xiao Xiang Zhu 0001
IGARSS3
2022 Deep Active Contour Models for Delineating Glacier Calving Fronts
abstract
We present a deep active contour model for detecting and delineating glacier calving fronts from satellite imagery. Contrary to existing deep learning-based calving front detectors, our model does not perform an intermediate segmentation or pixel-wise edge detection, but instead directly predicts the contour parametrized by a fixed number of vertices. The model works by first deriving feature maps from an input image, and then updating an initial contour in an iterative fashion. Evaluating on the CALFIN dataset, which maps calving fronts in Greenland, our model outperforms existing approaches. Code for the experiments and animated predictions can be found at https://github.com/khdlr/deep-acm
Konrad Heidler, Lichao Mou, Erik Loebel, Mirko Scheinert, Sébastien Lefèvre, Xiao Xiang Zhu 0001
IGARSS6
2022 Comparative Analysis of Nitrogen Dioxide (NO2) Levels in Munich Using Sentinel-5P Atmospheric Products and Ground-Based Measurements
abstract
In this paper, we compare trends in air quality measured by official stations on the ground with remote-sensing-based measurements in Munich, Germany. With Earth Observation data from Sentinel-5P and Sentinel-2, this study investigates the discrepancies that may arise in trends and pollutant levels of NO2. We find that the defined fixed measuring stations in Munich cover the worst pollution levels in the city. However, we do also find inconsistencies in the information provided by ground measurements regarding pollutant concentrations in contrast to high spatial coverage remote sensing data.
Ariadna Pregel Hoderlein, Lukas Kondmann, Xiao Xiang Zhu 0001
IGARSS3
2022 Uncertainty-Guided Representation Learning in Local Climate Zone Classification
abstract
A significant leap forward in the performance of remote sensing models can be attributed to recent advances in machine and deep learning. Large data sets particularly benefit from deep learning models, which often comprise millions of parameters. On which part of the data a machine learner focuses on during learning, however, remains an open research question. With the aid of a notion of label uncertainty, we try to address this question in local climate zone (LCZ) classification. Using a deep network as a feature extractor, we identify data samples that are seemingly easy or hard to classify for the model and base our experiments on the relatively more uncertain samples. For training of the network, we make use of distributional (probabilistic) labels to incorporate the voter confusion directly into the training process. The effectiveness of the proposed uncertainty-guided representation learning is shown in context of active learning framework where we show that adding more certain data to the training pool increases model performance even with the limited data.
Christoph Koller, Muhammad Shahrad, Xiao Xiang Zhu 0001
IGARSS3
2022 Feature and Output Consistency Training for Semi-Supervised Building Footprint Generation
abstract
Building footprint maps are important to urban planning and monitoring. However, most existing approaches that fall back on convolutional neural networks (CNNs), require massive annotated samples for network learning. In this research, we propose a novel semi-supervised network, which can help to deal with this issue by leveraging a large amount of unlabeled data. Considering that rich information is also encoded in feature maps, we propose to integrate the consistency of both features and outputs in the end-to-end network training of unlabeled samples on data perturbation, enabling to impose additional constraints. Experiments are conducted on Inria dataset. Our approach is much superior to the state-of-the-art methods in both quantitative and qualitative results.
Qingyu Li 0001, Yilei Shi, Xiao Xiang Zhu 0001
IGARSS3
2022 Controlling Weather Field Synthesis Using Variational Autoencoders
abstract
One of the consequences of climate change is an observed increase in the frequency of extreme climate events. That poses a challenge for weather forecast and generation algorithms, which learn from historical data but should embed an often uncertain bias to create correct scenarios. This paper investigates how mapping climate data to a known distribution using variational autoencoders might help explore such biases and control the synthesis of weather fields towards scenarios with more frequent extreme weather events. We experimented using a monsoon-affected precipitation dataset from southwest India, which should give a roughly stable pattern of rainy days and ease investigating the suitability of our solution. We report compelling results showing that mapping complex weather data to a known distribution implements an efficient control for weather field synthesis towards more (or less) extreme scenarios.
Dário A. B. Oliveira, Jorge Guevara Diaz, Bianca Zadrozny, Campbell D. Watson, Xiao Xiang Zhu 0001
IGARSS5
2022 Complex-Valued Sparse Long Short-Term Memory Unit with Application to Super-Resolving SAR Tomography
abstract
To achieve super-resolution synthetic aperture radar (SAR) tomography (TomoSAR), compressive sensing (CS)-based algorithms are usually employed, which are, however, computationally expensive, and thus is not often applied in large-scale processing. Recently, deep unfolding techniques have provided a good combination of physical model-based algorithms and the ability of neural networks to learn from data. In this vein, iterative CS-based algorithms can usually be un-rolled as neural networks with only 10 to 20 layers. When trained, it shows great computational efficiency for further TomoSAR processing. However, the learning architecture of neural networks built in this approach tends to result in error propagation and information loss, thus degrading the performance. In this paper, we propose to employ complex-valued sparse long short-term memory (CV-SLSTM) units to tackle this problem by incorporating historically updating information into the optimization procedure and preserving full information. Simulations are carried out to validate the performance of the proposed algorithm.
Kun Qian 0020, Yuanyuan Wang 0002, Peter Jung 0001, Yilei Shi, Xiao Xiang Zhu 0001
IGARSS5
2022 Compact Feature Representation for Unsupervised Ood Detection
abstract
Distributional mismatch between training and test data may cause the remote sensing models to behave in unpredictable manner, thus reducing the trustworthiness of such models. Most existing methods for out-of-distribution (OOD) detection rely on availability of OOD samples during training. However, access to OOD data during training is counter intuitive and may be impractical sometimes. Considering this, we propose an unsupervised OOD detection model that does not require training OOD data. The proposed method works by projecting the in-domain samples as a union of 1-dimensional subspaces. Due to the compact feature representation of in-domain samples, OOD samples are less likely to occupy the same feature space, thus they are easily identified. Experimental results demonstrate the capability of the proposed method to detect OOD samples.
Sudipan Saha, Jakob Gawlikowski, Jay Nandy, Xiao Xiang Zhu 0001
IGARSS4
2022 Mitigating Distribution Shift for Multi-Sensor Classification
abstract
Distribution shift may pose significant challenges in Earth observation, especially when dealing with significantly differ-ent sensors like multispectral optical and Synthetic Aperture Radar (SAR). Deep learning models trained for optical image classification generally do not generalize well for SAR images. This is due to very marked differences between them. Though there is a considerable amount of works on domain adaptation, only few deal with such strong differences. Towards this, we propose a co-teaching based domain adaptation method using dual classifier head, a Multi-layer Perceptron (MLP) classi-fier and a Graph Neural Network (GNN) classifier. The two classifier heads teach each other in an iterative manner, thus gradually adapting both of them for target classification. We experimentally demonstrate the efficacy of the proposed approach on Sentinel 2 (optical) as source and Sentinel 1 (SAR) images as target - both product of Copernicus program of European Space Agency.
Sudipan Saha, Shan Zhao 0007, Muhammad Shahzad 0002, Xiao Xiang Zhu 0001
IGARSS4
2022 Building Type Classification with Incomplete Labels
abstract
Buildings can be distinguished by their form or function and maps of building types can be used by authorities for city planning. Training models to perform this classification requires appropriate training data. OpenStreetMap (OSM) data is globaly available and partly provides information on building types. However, this data can be incomplete or wrong. In this work a U-Net is trained to group buildings into one of the three major function classes (commercial/industrial, residential and other) using incomplete OSM data or ground-truth cadastral data. The model achieves overall accuracies of 72 and 75 percent. Given the OSM data has only around 20 percent of the ground truth labels this shows the incomplete data can be used to train for the building classification task.
Nikolai Skuppin, Eike Jens Hoffmann, Yilei Shi, Xiao Xiang Zhu 0001
IGARSS4
2022 Towards Global Forest Biomass Estimators from Tree Height Data
abstract
In order to estimate tree biomass, allometric equations take tree parameters such as tree height, wood density, circumference of trunk, and crown diameter as input parameters. Given that most of these quantities are challenging to be extracted from remote sensing data, we evaluate the option to approximate biomass by tree height only. We study our approach by evaluating linear regression, random forest, and Gaussian process regressor models when applied to the 2016 Jucker dataset. Results indicate that linear models fail to properly capture the relationship between biomass and tree height, but the Gaussian process regressor outperms the other two candidate models.
Qian Song, Conrad M. Albrecht, Zhitong Xiong, Xiao Xiang Zhu 0001
IGARSS4
2022 Self-Supervised Vision Transformers for Joint SAR-Optical Representation Learning
abstract
Self-supervised learning (SSL) has attracted much interest in remote sensing and Earth observation due to its ability to learn task-agnostic representations without human annotation. While most of the existing SSL works in remote sensing utilize ConvNet backbones and focus on a single modality, we explore the potential of vision transformers (ViTs) for joint SAR-optical representation learning. Based on DINO, a state-of-the-art SSL algorithm that distills knowledge from two augmented views of an input image, we combine SAR and optical imagery by concatenating all channels to a unified input. Subsequently, we randomly mask out channels of one modality as a data augmentation strategy. While training, the model gets fed optical-only, SAR-only, and SAR-optical image pairs learning both inner- and intra-modality representations. Experimental results employing the BigEarthNet-MM dataset demonstrate the benefits of both, the ViT backbones and the proposed multimodal SSL algorithm DINO-MM.
Yi Wang 0072, Conrad M. Albrecht, Xiao Xiang Zhu 0001
IGARSS3
2022 Knowledge Transfer for Label-Efficient Monocular Height Estimation
abstract
Estimating height from monocular remote sensing images is one of the most efficient ways for building large-scale 3D city models. However, existing deep learning based methods usu-ally require a large amount of training data, which could be cost-consuming or even not possible to obtain. Towards a label-efficient deep learning model, we propose a new task and dataset for weak-shot monocular height estimation. In this task, only the relative height labels between pairs of a small portion of points are given, which is cheaper and more friendly for humans to annotate. In addition, to enhance the model performance under the sparse and weak-shot super-vision, we propose a Transformer-based network for trans-ferring the learned knowledge from a large-scale synthetic dataset to real-world data. Experimental results have shown the effectiveness of the proposed method on a public dataset under the sparse and weak supervision.
Zhitong Xiong, Xiao Xiang Zhu 0001
IGARSS2
2022 Universal Domain Adaptation without Source Data for Remote Sensing Image Scene Classification
abstract
Existing domain adaptation (DA) approaches are usually not well suited for practical DA scenarios of remote sensing image classification, since these methods (such as unsupervised DA) rely on rich prior knowledge about the relationship between label sets of source and target domains, and source data are usually not accessible in many cases due to the privacy or confidentiality issues. To this end, we propose a novel source data generation-based universal domain adaptation (SDG-UniDA) model, which includes two parts, i.e., the stage of source data generation and the stage of model adaptation. The first stage is to estimate the conditional distribution of source data from the pre-trained model using the knowledge of class-separability in the source domain and then to synthesize the source data. With this synthetic source data in hand, it becomes a universal DA task that requires no prior knowledge on the label sets. A novel transferable weight is proposed to distinguish the shared and private label sets to each domain, thereby promoting the adaptation in the automatically discovered shared label set and recognizing the "unknown" samples successfully. Empirical results show that SDG-UniDA is effective and practical in this challenging setting for remote sensing image scene classification.
Qingsong Xu 0001, Yilei Shi, Xiao Xiang Zhu 0001
IGARSS3
2022 Change-Aware Visual Question Answering
abstract
Change detection has been a hot research topic in the field of remote sensing, and it can provide information on observing changes of Earth's surface. However, segmentation-based change results are not very friendly to end users. Thus, in order to improve user experience and offer them high-level semantic information on change detection, we introduce a new task: change-aware visual question answering (VQA) on multi-temporal aerial images. Specifically, given a pair of multi-temporal aerial images and questions, this task aims to automatically provide natural language answers. By doing so, end users have better access to easy-to-understand change information through natural language. Besides, we also create a dataset made of multi-temporal image-question-answer triplets and a baseline method for this task. Experimental results offer valuable insights for the further research on this task.
Zhenghang Yuan, Lichao Mou, Xiao Xiang Zhu 0001
IGARSS3
2022 Domain-Agnostic Domain Adaption for Building Footprint Extraction
abstract
For global range satellite imaging mission, images captured from different areas may have large distribution biases due to different illuminations, shooting angles and atmospheric conditions. A straightforward idea to mitigate this problem is to categorize the images into different domains according the cities they belong to, and apply domain adaptation approaches. However, categorization by cities becomes unreasonable with the increase of the city number, and the emergence of inter-city similarity and intra-city discrepancy. With such consideration, this paper proposes a novel domain adaptation method named domain-agnostic domain adaptation (DADA) to reduce the distribution biases without explicitly defining the domain each image belongs to. To implement this, we augment the images to the styles of different domains by Generative Adversarial Networks (GAN) and contrastive learning to increase the generalizability of down-stream tasks. Experiments on Planetscope building footprint extraction datasets verify the effectiveness of our method.
Fahong Zhang 0001, Yilei Shi, Xiao Xiang Zhu 0001
IGARSS3
2022 Deep Relearning in the Geospatial Domain for Semantic Remote Sensing Image Segmentation
abstract
We present a classification postprocessing (CPP) technique based on fully convolutional neural networks (CNNs) for semantic remote sensing image segmentation. Conventional CPP techniques aim to enhance the classification accuracy by imposing smoothness priors in the image domain. Contrary to that, here, a relearning strategy is proposed where the initial classification outcome of a CNN model is provided to a subsequent CNN model via an extended input space to guide the learning of discriminative feature representations in an end-to-end fashion. This deep relearning CNN (DRCNN) explicitly accounts for the geospatial domain by taking the spatial alignment of preliminary class labels into account. Hereby, we evaluate to learn the DRCNN in a cumulative and noncumulative way, i.e., extending the input space based on all previous or solely preceding model outputs, respectively, during an iterative procedure. Besides, the DRCNN can also be conveniently coupled with alternative CPP techniques such as object-based voting (OBV). The experimental results obtained from two test sites of WorldView-II imagery underline the beneficial performance properties of the DRCNN models. They can increase the accuracies of the initial CNN models on average from 72.64% to 76.01% and from 92.43% to 94.52% in terms of$\kappa $statistic. An additional increase of 1.65 and 2.84 percentage points can be achieved when combining the DRCNN models with an OBV strategy. From an epistemological point of view, our results underline that CNNs can benefit from the consideration of preliminary model outcomes and that conventional CPP techniques can profit from an upstream relearning strategy.
Christian Geiß, Yue Zhu 0004, Chunping Qiu, Lichao Mou, Xiao Xiang Zhu 0001, Hannes Taubenböck
IEEE Geosci. Remote. Sens. Lett.5
2022 Semantic Segmentation of Remote Sensing Images With Sparse Annotations
abstract
Training convolutional neural networks (CNNs) for very high-resolution images requires a large quantity of high-quality pixel-level annotations, which is extremely labor-intensive and time-consuming to produce. Moreover, professional photograph interpreters might have to be involved in guaranteeing the correctness of annotations. To alleviate such a burden, we propose a framework for semantic segmentation of aerial images based on incomplete annotations, where annotators are asked to label a few pixels with easy-to-draw scribbles. To exploit these sparse scribbled annotations, we propose the FEature and Spatial relaTional regulArization (FESTA) method to complement the supervised task with an unsupervised learning signal that accounts for neighborhood structures both in spatial and feature terms. For the evaluation of our framework, we perform experiments on two remote sensing image segmentation data sets involving aerial and satellite imagery, respectively. Experimental results demonstrate that the exploitation of sparse annotations can significantly reduce labeling costs, while the proposed method can help improve the performance of semantic segmentation when training on such annotations. The sparse labels and codes are publicly available for reproducibility purposes.https://github.com/Hua-YS/Semantic-Segmentation-with-Sparse-Labels
Yuansheng Hua, Diego Marcos, Lichao Mou, Xiao Xiang Zhu 0001, Devis Tuia
IEEE Geosci. Remote. Sens. Lett.4
2022 Exploring Transformer and Multilabel Classification for Remote Sensing Image Captioning
abstract
High-resolution remote sensing images are now available with the progress of remote sensing technology. With respect to popular remote sensing tasks like scene classification, image captioning provides comprehensible information about such images by summarizing the image content in human-readable text. Most existing remote sensing image captioning methods are based on deep learning-based encoder-decoder frameworks, using Convolutional Neural Network or Recurrent Neural Network as the backbone of such frameworks. Such frameworks show a limited capability to analyze sequential data and cope with the lack of captioned remote sensing training images. Recently introduced Transformer architecture exploits self-attention to obtain superior performance for sequence-analysis tasks. Inspired by this, in this work, we employ a Transformer as an encoder-decoder for remote sensing image captioning. Moreover, to deal with the limited training data, an auxiliary decoder is used that further helps the encoder in the training process. The auxiliary decoder is trained for multi-label scene classification due to its conceptual similarity to image captioning and capability of highlighting semantic classes. To the best of our knowledge, this is the first work exploiting multi-label classification to improve remote sensing image captioning. Experimental results on the UC Merced caption data set show the efficacy of the proposed method. The implementation details can be found in https://gitlab.lrz.de/ai4eo/captioningMultilabel.
Hitesh Kandala, Sudipan Saha, Biplab Banerjee, Xiao Xiang Zhu 0001
IEEE Geosci. Remote. Sens. Lett.4
2022 Multitask Learning for Human Settlement Extent Regression and Local Climate Zone Classification
abstract
Human settlement extent (HSE) and local climate zone (LCZ) maps are both essential sources, e.g., for sustainable urban development and Urban Heat Island (UHI) studies. Remote sensing (RS)- and deep learning (DL)-based classification approaches play a significant role by providing the potential for global mapping. However, most of the efforts only focus on one of the two schemes, usually on a specific scale. This leads to unnecessary redundancies since the learned features could be leveraged for both of these related tasks. In this letter, the concept of multitask learning (MTL) is introduced to HSE regression and LCZ classification for the first time. We propose an MTL framework and develop an end-to-end convolutional neural network (CNN), which consists of a backbone network for shared feature learning, attention modules for task-specific feature learning, and a weighting strategy for balancing the two tasks. We additionally propose to exploit HSE predictions as a prior for LCZ classification to enhance the accuracy. The MTL approach was extensively tested with Sentinel-2 data of 13 cities across the world. The results demonstrate that the framework is able to provide a competitive solution for both tasks.
Chunping Qiu, Lukas Liebel, Lloyd H. Hughes, Michael Schmitt 0003, Marco Körner 0001, Xiao Xiang Zhu 0001
IEEE Geosci. Remote. Sens. Lett.6
2022 Patch-Level Unsupervised Planetary Change Detection
abstract
Change detection (CD) is critical for analyzing data collected by planetary exploration missions, e.g., for identification of new impact craters. However, CD is still a relatively new topic in the context of planetary exploration. Sheer variation of planetary data makes CD much more challenging than in the case of Earth observation (EO). Unlike CD for EO, patch-level decision is preferred in planetary exploration as it is difficult to obtain perfect pixelwise alignment/coregistration between the bi-temporal planetary images. Lack of labeled bi-temporal data impedes supervised CD. To overcome these challenges, we propose an unsupervised CD method that exploits a pretrained feature extractor to obtain bi-temporal deep features that are further processed using global max-pooling to obtain patch-level feature description. Bi-temporal patch-level features are further analyzed based on difference to determine whether a patch is changed. Additionally, a self-supervised method is proposed to estimate the decision boundary between the changed and unchanged patches. Experimental results on three planetary CD datasets from two different planetary bodies (Mars and Moon) demonstrate that the proposed method often outperforms supervised planetary CD methods. Code is available athttps://gitlab.lrz.de/ai4eo/cd/-/tree/main/planetaryCDUnsup.
Sudipan Saha, Xiao Xiang Zhu 0001
IEEE Geosci. Remote. Sens. Lett.2
2022 Multitarget Domain Adaptation for Remote Sensing Classification Using Graph Neural Network
abstract
Remote sensing deals with huge variations in geography, acquisition season, and a plethora of sensors. Considering the difficulty of collecting labeled data uniformly representing all scenarios, data-hungry deep learning models are often trained with labeled data in a source domain that is limited in the above-mentioned aspects. Domain adaptation (DA) methods can adapt such model for applying on target domains with different distributions from the source domain. However, most remote sensing DA methods are designed for single-target, thus requiring a separate target classifier to be trained for each target domain. To mitigate this, we propose multitarget DA in which a single classifier is learned for multiple unlabeled target domains. To build a multitarget classifier, it may be beneficial to effectively aggregate features from the labeled source and different unlabeled target domains. Toward this, we exploit coteaching based on the graph neural network that is capable of leveraging unlabeled data. We use a sequential adaptation strategy that first adapts on the easier target domains assuming that the network finds it easier to adapt to the closest target domain. We validate the proposed method on two different datasets, representing geographical and seasonal variation. Code is available athttps://gitlab.lrz.de/ai4eo/da-multitarget-gnn/.
Sudipan Saha, Shan Zhao 0007, Xiao Xiang Zhu 0001
IEEE Geosci. Remote. Sens. Lett.3
2022 SCAF-Net: Scene Context Attention-Based Fusion Network for Vehicle Detection in Aerial Imagery
abstract
In recent years, deep learning methods have achieved great success for vehicle detection tasks in aerial imagery. However, most existing methods focus only on extracting latent vehicle target features, and rarely consider the scene context as vital prior knowledge. In this letter, we propose a scene context attention-based fusion network (SCAF-Net), to fuse the scene context of vehicles into an end-to-end vehicle detection network. First, we propose a novel strategy, patch cover, to keep the original target and scene context information in raw aerial images of a large scale as much as possible. Next, we use an improved YOLO-v3 network as one branch of SCAF-Net, to generate vehicle candidates on each patch. Here, a novel branch for the scene context is utilized to extract the latent scene context of vehicles on each patch without any extra annotations. Then, these two branches above are concatenated together as a fusion network, and we apply an attention-based model to further extract vehicle candidates of each local scene. Finally, all vehicle candidates of different patches, are merged by global nonmax suppress (g-NMS) to output the detection result of the whole original image. Experimental results demonstrate that our proposed method outperforms the comparison methods with both high detection accuracy and speed. Our code is released athttps://github.com/minghuicode/SCAF-Net.
Qingpeng Li, Yunchao Gu, Leyuan Fang, Xiao Xiang Zhu 0001
IEEE Geosci. Remote. Sens. Lett.5
2022 Deep Quadruple-Based Hashing for Remote Sensing Image-Sound Retrieval
abstract
With the rapid progress of earth observation technology, cross-modal remote sensing (RS) image-sound retrieval has attracted much attention from the field of RS data processing. Existing approaches usually learn the pairwise similarity relations between RS images and sounds. However, these approaches ignore relative semantic similarity relationships, which leads to poor performance of cross-modal RS image-sound retrieval. In this article, we address this dilemma with a noveldeep quadruple-based hashing(DQH) approach. We first devise a novel quadruple-based hashing network to learn relative semantic similarity relationships of hash codes. Meanwhile, we propose a quadruple construction hard module, which randomly selects two triplet hard units to directly learn relative semantic similarity relationships. On top of the two paths, we develop a new objective function to perform effective hash codes learning. The new objective function not only captures the relative semantic correlation of hash codes across different modalities and learns the relative semantic correlation of deep features but also enhances category-level semantics of hash codes and reduces the quantization error between hash-like codes and hash codes. The reasonableness and effectiveness of the proposed architecture are well illustrated by comprehensive experiments on diverse RS image-sound datasets.
Yaxiong Chen, Shengwu Xiong 0001, Lichao Mou, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 SEN12MS-CR-TS: A Remote-Sensing Data Set for Multimodal Multitemporal Cloud Removal
abstract
About half of all optical observations collected via spaceborne satellites are affected by haze or clouds. Consequently, cloud coverage affects the remote-sensing practitioner’s capabilities of a continuous and seamless monitoring of our planet. This work addresses the challenge of optical satellite image reconstruction and cloud removal by proposing a novel multimodal and multitemporal data set called SEN12MS-CR-TS. We propose two models highlighting the benefits and use cases of SEN12MS-CR-TS: First, a multimodal multitemporal 3-D convolution neural network that predicts a cloud-free image from a sequence of cloudy optical and radar images. Second, a sequence-to-sequence translation model that predicts a cloud-free time series from a cloud-covered time series. Both approaches are evaluated experimentally, with their respective models trained and tested on SEN12MS-CR-TS. The conducted experiments highlight the contribution of our data set to the remote-sensing community as well as the benefits of multimodal and multitemporal information to reconstruct noisy information. Our data set is available athttps://patrickTUM.github.io/cloud_removal.
Patrick Ebel 0002, Yajin Xu, Michael Schmitt 0003, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 An Advanced Dirichlet Prior Network for Out-of-Distribution Detection in Remote Sensing
abstract
Remote sensing deals with a plethora of sensors, a large number of classes/categories, and a huge variation in geography. Due to the difficulty of collecting labeled data uniformly representing all scenarios, data-hungry deep learning models are often trained with labeled data in a source domain that is limited in the above-mentioned aspects. However, during the test/inference phase, such deep learning models are often subjected to a distributional shift, also called out-of-distribution (OOD) samples, in the form of unseen classes, geographic differences, and multisensor differences. Deep learning models can behave in an unexpected manner when subjected to such distributional uncertainties. Vulnerability to OOD data severely reduces the reliability of deep learning models and trusting on such predictions in the absence of any reliability indicator may lead to wrong policy decisions or mishaps in time-bound remote sensing applications. Motivated by this, in this work, we propose a Dirichlet prior network-based model to quantify the distributional uncertainty of deep learning-based remote sensing models. The approach seeks to maximize the representation gap between the in-domain and OOD examples for better segregation of OOD samples at test time. Extensive experiments on several remote sensing image classification datasets demonstrate that the proposed model can quantify distributional uncertainty. To the best of our knowledge, this is the first work to elaborately study distributional uncertainty in context of remote sensing. The codes are publicly available athttps://gitlab.lrz.de/ai4eo/Uncertainty/-/tree/main/DPN-RS.
Jakob Gawlikowski, Sudipan Saha, Anna M. Kruspe, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 HED-UNet: Combined Segmentation and Edge Detection for Monitoring the Antarctic Coastline
abstract
Deep learning-based coastline detection algorithms have begun to outshine traditional statistical methods in recent years. However, they are usually trained only as single-purpose models to either segment land and water or delineate the coastline. In contrast to this, a human annotator will usually keep a mental map of both segmentation and delineation when performing manual coastline detection. To take into account this task duality, we, therefore, devise a new model to unite these two approaches in a deep learning model. By taking inspiration from the main building blocks of a semantic segmentation framework (UNet) and an edge detection framework (HED), both tasks are combined in a natural way. Training is made efficient by employing deep supervision on side predictions at multiple resolutions. Finally, a hierarchical attention mechanism is introduced to adaptively merge these multiscale predictions into the final model output. The advantages of this approach over other traditional and deep learning-based methods for coastline detection are demonstrated on a data set of Sentinel-1 imagery covering parts of the Antarctic coast, where coastline detection is notoriously difficult. An implementation of our method is available athttps://github.com/khdlr/HED-UNet.
Konrad Heidler, Lichao Mou, Celia A. Baumhoer, Andreas J. Dietz, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 MultiScene: A Large-Scale Dataset and Benchmark for Multiscene Recognition in Single Aerial Images
Yuansheng Hua, Lichao Mou, Pu Jin, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 FuTH-Net: Fusing Temporal Relations and Holistic Features for Aerial Video Classification
abstract
Unmanned aerial vehicles (UAVs) are now widely applied to data acquisition due to its low cost and fast mobility. With the increasing volume of aerial videos, the demand for automatically parsing these videos is surging. To achieve this, current research mainly focuses on extracting a holistic feature with convolutions along both spatial and temporal dimensions. However, these methods are limited by small temporal receptive fields and cannot adequately capture long-term temporal dependencies that are important for describing complicated dynamics. In this article, we propose a novel deep neural network, termed Fusing Temporal relations and Holistic features for aerial video classification (FuTH-Net), to model not only holistic features but also temporal relations for aerial video classification. Furthermore, the holistic features are refined by the multiscale temporal relations in a novel fusion module for yielding more discriminative video representations. More specially, FuTH-Net employs a two-pathway architecture: 1) a holistic representation pathway to learn a general feature of both frame appearances and short-term temporal variations and 2) a temporal relation pathway to capture multiscale temporal relations across arbitrary frames, providing long-term temporal dependencies. Afterward, a novel fusion module is proposed to spatiotemporally integrate the two features learned from the two pathways. Our model is evaluated on two aerial video classification datasets, ERA and Drone-Action, and achieves the state-of-the-art results. This demonstrates its effectiveness and good generalization capacity across different recognition tasks (event classification and human action recognition). To facilitate further research, we release the code athttps://gitlab.lrz.de/ai4eo/reasoning/futh-net.
Pu Jin, Lichao Mou, Yuansheng Hua, Gui-Song Xia, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 Anomaly Detection in Aerial Videos With Transformers
abstract
Unmanned aerial vehicles (UAVs) are widely applied for purposes of inspection, search, and rescue operations by the virtue of low-cost, large-coverage, real-time, and high-resolution data acquisition capacities. Massive volumes of aerial videos are produced in these processes, in which normal events often account for an overwhelming proportion. It is extremely difficult to localize and extract abnormal events containing potentially valuable information from long video streams manually. Therefore, we are dedicated to developing anomaly detection methods to solve this issue. In this paper, we create a new dataset, named Drone-Anomaly, for anomaly detection in aerial videos. This dataset provides 37 training video sequences and 22 testing video sequences from 7 different realistic scenes with various anomalous events. There are 87,488 color video frames (51,635 for training and 35,853 for testing) with the size of 640 × 640 at 30 frames per second. Based on this dataset, we evaluate existing methods and offer a benchmark for this task. Furthermore, we present a new baseline model, ANomaly Detection with Transformers (ANDT), which treats consecutive video frames as a sequence of tubelets, utilizes a Transformer encoder to learn feature representations from the sequence, and leverages a decoder to predict the next frame. Our network models normality in the training phase and identifies an event with unpredictable temporal dynamics as an anomaly in the test phase. Moreover, To comprehensively evaluate the performance of our proposed method, we use not only our Drone-Anomaly dataset but also another dataset. We will make our dataset and code publicly available. A demo video is available at https://youtu.be/ancczYryOBY. We make our dataset and code publicly available1.
Pu Jin, Lichao Mou, Gui-Song Xia, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 Spatial Context Awareness for Unsupervised Change Detection in Optical Satellite Images
Lukas Kondmann, Aysim Toker, Sudipan Saha, Bernhard Schölkopf, Laura Leal-Taixé, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.6
2022 Building Footprint Generation Through Convolutional Neural Networks With Attraction Field Representation
abstract
Building footprint generation is a vital task in a wide range of applications, including, to name a few, land use management, urban planning and monitoring, and geographical database updating. Most existing approaches addressing this problem fall back on convolutional neural networks (CNNs) to learn semantic masks of buildings. However, one limitation of their results is blurred building boundaries. To address this, we propose to learn attraction field representation for building boundaries, which is capable of providing an enhanced representation power. Our method comprises two elemental modules: an Img2AFM module and an AFM2Mask module. More specifically, the former aims at learning an attraction field representation conditioned on an input image, which is capable of enhancing building boundaries and suppressing the background. The latter module predicts segmentation masks of buildings using the learned attraction field map. The proposed method is evaluated on three datasets with different spatial resolutions: the ISPRS dataset, the INRIA dataset, and the Planet dataset. From experimental results, we find that the proposed framework can well preserve geometric shapes and sharp boundaries of buildings, which brings significant improvements over other competitors. The trained model and code are available at https://github.com/lqycrystal/AFM_building.
Qingyu Li 0001, Lichao Mou, Yuansheng Hua, Yilei Shi, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 Semi-Supervised Building Footprint Generation With Feature and Output Consistency Training
abstract
Accurate and reliable building footprint maps are vital to urban planning and monitoring, and most existing approaches fall back on convolutional neural networks (CNNs) for building footprint generation. However, one limitation of these methods is that they require strong supervisory information from massive annotated samples for network learning. State-of-the-art semi-supervised semantic segmentation networks with consistency training can help to deal with this issue by leveraging a large amount of unlabeled data, which encourages the consistency of model output on data perturbation. Considering that rich information is also encoded in feature maps, we propose to integrate the consistency of both features and outputs in the end-to-end network training of unlabeled samples, enabling to impose additional constraints. Prior semi-supervised semantic segmentation networks have established the cluster assumption, in which the decision boundary should lie in the vicinity of low sample density. In this work, we observe that for building footprint generation, the low-density regions are more apparent at the intermediate feature representations within the encoder than the encoder’s input or output. Therefore, we propose an instruction to assign the perturbation to the intermediate feature representations within the encoder, which considers the spatial resolution of input remote sensing imagery and the mean size of individual buildings in the study area. The proposed method is evaluated on three datasets with different resolutions: Planet dataset (3 m/pixel), Massachusetts dataset (1 m/pixel), and Inria dataset (0.3 m/pixel). Experimental results show that the proposed approach can well extract more complete building structures and alleviate omission errors.
Qingyu Li 0001, Yilei Shi, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 A Lightweight Deep Learning-Based Cloud Detection Method for Sentinel-2A Imagery Fusing Multiscale Spectral and Spatial Features
abstract
Clouds are a very important factor in the availability of optical remote sensing images. Recently, deep learning (DL)-based cloud detection methods have surpassed classical methods based on rules and physical models of clouds. However, most of these deep models are very large, which limits their applicability and explainability, while other models do not make use of the full spectral information in multispectral images, such as Sentinel-2. In this article, we propose a lightweight network for cloud detection, fusing multiscale spectral and spatial features (CD-FM3SFs) and tailored for processing all spectral bands in Sentinel-2A images. The proposed method consists of an encoder and a decoder. In the encoder, three input branches are designed to handle spectral bands at their native resolution and extract multiscale spectral features. Three novel components are designed: a mixed depthwise separable convolution (MDSC) and a shared and dilated residual block (SDRB) to extract multiscale spatial features, and a concatenation and sum (CS) operation to fuse multiscale spectral and spatial features with little calculation and no additional parameters. The decoder of CD-FM3SF outputs three cloud masks at the same resolution as input bands to enhance the supervision information of small, middle, and large clouds. To validate the performance of the proposed method, we manually labeled 36 Sentinel-2A scenes evenly distributed over mainland China. The experiment results demonstrate that CD-FM3SF outperforms traditional cloud detection methods and state-of-the-art DL-based methods in both accuracy and speed.
Jun Li 0087, Zhaocong Wu, Zhongwen Hu, Canliang Jian, Shaojie Luo, Lichao Mou, Xiao Xiang Zhu 0001, Matthieu Molinier
IEEE Trans. Geosci. Remote. Sens.7
2022 Extracting Glacier Calving Fronts by Deep Learning: The Benefit of Multispectral, Topographic, and Textural Input Features
abstract
An accurate parameterization of glacier calving is essential for understanding glacier dynamics and constraining ice-sheet models. The increasing availability and quality of remote sensing imagery open the prospect of a continuous and precise mapping of relevant parameters, such as calving front locations. However, it also calls for automated and scalable analysis strategies. Deep neural networks provide powerful tools for processing large quantities of remote sensing data. In this contribution, we assess the benefit of diverse input data for calving front extraction. In particular, we focus on Landsat-8 imagery supplementing single-band inputs with multispectral data, topography, and textural information. We assess the benefit of these three datasets using a dropped-variable approach. The associated reference dataset comprises 728 manually delineated calving front positions of 23 Greenland and two Antarctic outlet glaciers from 2013 to 2021. Resulting feature importance emphasizes both the potential integrating additional input information as well as the significance of their thoughtful selection. We advocate utilizing multispectral features as their integration leads generally to more accurate predictions compared with conventional single-band inputs. This is especially prevalent for challenging ice mélange and illumination conditions. In contrast, the application of both textural and topographic inputs cannot be recommended without reservation, since they may lead to model overfitting. The results of this assessment are not only relevant for advancing automated calving front extraction but also for a wider range of glaciology-related land surface classification tasks using deep neural networks.
Erik Loebel, Mirko Scheinert, Martin Horwath, Konrad Heidler, Julia Christmann, Long Duc Phan, Angelika Humbert, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.8
2022 Detecting Changes by Learning No Changes: Data-Enclosing-Ball Minimizing Autoencoders for One-Class Change Detection in Multispectral Imagery
abstract
Change detection is a long-standing and challenging problem in remote sensing. Very often, features about changes are difficult to model beforehand, thus making the collection of changed samples a challenging task. In comparison, it is much easier to collect numerous no-change samples. It is possible to define a change detection approach by using only easily available annotated no-change samples, which we henceforth call one-class change detection. Autoencoder networks being trained on no-change data are natural candidates for addressing this task due to their superior performance as compared to other one-class classification models. However, the autoencoders usually suffer from the problem of overgeneralization, i.e., they tend to generalize too well, thus risking properly reconstructing changed samples. In this paper, we propose a novel data-enclosing-ball minimizing autoencoder (DebM-AE) that is trained with dual objectives—a reconstruction error criterion and a minimum volume criterion. The network learns a compact latent space, where encodings of no-change samples have low intra-class variance, which as counter part has the identification of changed instances. We conducted extensive experiments on three real-world data sets. Results demonstrate advantages of the proposed method over other competitors. We make our data and code publicly available1.
Lichao Mou, Yuansheng Hua, Sudipan Saha, Francesca Bovolo, Lorenzo Bruzzone, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.6
2022 Deep Reinforcement Learning for Band Selection in Hyperspectral Image Classification
abstract
Band selection refers to the process of choosing the most relevant bands in a hyperspectral image. By selecting a limited number of optimal bands, we aim at speeding up model training, improving accuracy, or both. It reduces redundancy among spectral bands while trying to preserve the original information of the image. By now, many efforts have been made to develop unsupervised band selection approaches, of which the majorities are heuristic algorithms devised by trial and error. In this article, we are interested in training an intelligent agent that, given a hyperspectral image, is capable of automatically learning policy to select an optimal band subset without any hand-engineered reasoning. To this end, we frame the problem of unsupervised band selection as a Markov decision process, propose an effective method to parameterize it, and finally solve the problem by deep reinforcement learning. Once the agent is trained, it learns a band-selection policy that guides the agent to sequentially select bands by fully exploiting the hyperspectral image and previously picked bands. Furthermore, we propose two different reward schemes for the environment simulation of deep reinforcement learning and compare them in experiments. This, to the best of our knowledge, is the first study that explores a deep reinforcement learning model for hyperspectral image analysis, thus opening a new door for future research and showcasing the great potential of deep reinforcement learning in remote sensing applications. Extensive experiments are carried out on four hyperspectral data sets, and experimental results demonstrate the effectiveness of the proposed method. The code is publicly available.
Lichao Mou, Sudipan Saha, Yuansheng Hua, Francesca Bovolo, Lorenzo Bruzzone, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.6
2022 Basis Pursuit Denoising via Recurrent Neural Network Applied to Super-Resolving SAR Tomography
abstract
Finding sparse solutions of underdetermined linear systems commonly requires the solving ofL1regularized least squares minimization problem, which is also known as the basis pursuit denoising (BPDN). They are computationally expensive since they cannot be solved analytically. An emerging technique known asdeep unrollingprovided a good combination of the descriptive ability of neural networks, explainable, and computational efficiency for BPDN. Many unrolled neural networks for BPDN, e.g. learned iterative shrinkage thresholding algorithm and its variants, employ shrinkage functions to prune elements with small magnitude. Through experiments on synthetic aperture radar tomography (TomoSAR), we discover the shrinkage step leads to unavoidable information loss in the dynamics of networks and degrades the performance of the model. We propose a recurrent neural network (RNN) with novel sparse minimal gated units (SMGUs) to solve the information loss issue. The proposed RNN architecture with SMGUs benefits from incorporating historical information into optimization, and thus effectively preserves full information to the final output. Taking TomoSAR inversion as an example, extensive simulations demonstrated that the proposed RNN outperforms the state-of-the-art deep learning-based algorithm in terms of super-resolution power as well as generalization ability. It achieved 10% to 20% higher double scatterers detection rate and is less sensitive to phase and amplitude ratio difference between scatterers. Test on real TerraSAR-X spotlight images also shows high-quality 3-D reconstruction of test site.
Kun Qian 0020, Yuanyuan Wang 0002, Peter Jung 0001, Yilei Shi, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 γ-Net: Superresolving SAR Tomographic Inversion via Deep Learning
abstract
Synthetic aperture radar tomography (TomoSAR) has been extensively employed in 3-D reconstruction in dense urban areas using high-resolution SAR acquisitions. Compressive sensing (CS)-based algorithms are generally considered as the state-of-the art in super-resolving TomoSAR, in particular in the single look case. This superior performance comes at the cost of extra computational burdens, because of the sparse reconstruction, which cannot be solved analytically, and we need to employ computationally expensive iterative solvers. In this article, we propose a novel deep learning-based super-resolving TomoSAR inversion approach,$\boldsymbol {\gamma }$-Net, to tackle this challenge.$\boldsymbol {\gamma }$-Net adopts advanced complex-valued learned iterative shrinkage thresholding algorithm (CV-LISTA) to mimic the iterative optimization step in sparse reconstruction. Simulations show the height estimate from a well-trained$\boldsymbol {\gamma }$-Net approaches the Cramér-Rao lower bound (CRLB) while improving the computational efficiency by one to two orders of magnitude comparing to the first-order CS-based methods. It also shows no degradation in the super-resolution power comparing to the state-of-the-art second-order TomoSAR solvers, which are much more computationally expensive than the first-order methods. Specifically,$\boldsymbol {\gamma }$-Net reaches more than 90% detection rate in moderate super-resolving cases at 25 measurements at 6 dB SNR. Moreover, simulation at limited baselines demonstrates that the proposed algorithm outperforms the second-order CS-based method by a fair margin. Test on real TanDEM-X data with just six interferograms also shows high-quality 3-D reconstruction with high-density detected double scatterers.
Kun Qian 0020, Yuanyuan Wang 0002, Yilei Shi, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 Self-Supervised Multisensor Change Detection
abstract
Most change detection (CD) methods assume that prechange and postchange images are acquired by the same sensor. However, in many real-life scenarios, e.g., natural disasters, it is more practical to use the latest available images before and after the occurrence of incidence, which may be acquired using different sensors. In particular, we are interested in the combination of the images acquired by optical and synthetic aperture radar (SAR) sensors. SAR images appear vastly different from the optical images even when capturing the same scene. Adding to this, CD methods are often constrained to use only target image-pair, no labeled data, and no additional unlabeled data. Such constraints limit the scope of traditional supervised machine learning and unsupervised generative approaches for multisensor CD. The recent rapid development of self-supervised learning methods has shown that some of them can even work with only few images. Motivated by this, in this work, we propose a method for multisensor CD using only the unlabeled target bitemporal images that are used for training a network in a self-supervised fashion by using deep clustering and contrastive learning. The proposed method is evaluated on four multimodal bitemporal scenes showing change, and the benefits of our self-supervised approach are demonstrated. Code is available athttps://gitlab.lrz.de/ai4eo/cd/-/tree/main/sarOpticalMultisensorTgrs2021.
Sudipan Saha, Patrick Ebel 0002, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 Unsupervised Single-Scene Semantic Segmentation for Earth Observation
abstract
Earth observation data has huge potential to enrich our knowledge about our planet. An important step in many Earth observation tasks is semantic segmentation. Generally, a large number of pixelwise labeled images are required to train deep models for supervised semantic segmentation. On the contrary, strong inter-sensor and geographic variations impede the availability of annotated training data in Earth observation. In practice, most Earth observation tasks use only the target scene without assuming availability of any additional scene, labeled or unlabeled. Keeping in mind such constraints, we propose a semantic segmentation method that learns to segment from a single scene, without using any annotation. Earth observation scenes are generally larger than those encountered in typical computer vision datasets. Exploiting this, the proposed method samples smaller unlabeled patches from the scene. For each patch an alternate view is generated by simple transformations, e.g., addition of noise. Both views are then processed through a two-stream network and weights are iteratively refined using deep clustering, spatial consistency, and contrastive learning in the pixel space. The proposed model automatically segregates the major classes present in the scene and produces the segmentation map. Extensive experiments on four Earth observation datasets collected by different sensors show the effectiveness of the proposed method. Implementation is available at https://gitlab.lrz.de/ai4eo/cd/-/tree/main/unsupContrastiveSemanticSeg.
Sudipan Saha, Muhammad Shahzad 0002, Lichao Mou, Qian Song, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 Learning to Generate SAR Images With Adversarial Autoencoder
abstract
Deep learning-based synthetic aperture radar (SAR) target recognition often suffers from sparsely distributed training samples and rapid angular variations due to scattering scintillation. Thus, data-driven SAR target recognition is considered a typical few-shot learning (FSL) task. This article first reviews the key issues of FSL and provides a definition of the FSL task. A novel adversarial autoencoder (AAE) is then proposed as an SAR representation and generation network. It consists of a generator network that decodes target knowledge to SAR images and an adversarial discriminator network that not only learns to discriminate “fake” generated images from real ones but also encodes the input SAR image back to target knowledge. The discriminator employs progressively expanding convolution layers and a corresponding layer-by-layer training strategy. It uses two cyclic loss functions to enforce consistency between the inputs and outputs. Moreover, rotated cropping is introduced as a mechanism to address the challenge of representing the target orientation. The moving and stationary Target recognition (MSTAR) 7-target dataset is used to evaluate the AAE’s performance, and the results demonstrate its ability to generate SAR images with aspect angular diversity. Using only 90 training samples with at least 25° of orientation interval, the trained AAE is able to generate the remaining 1748 samples of other orientation angles with an unprecedented level of fidelity. Thus, it can be used for data augmentation in SAR target recognition FSL tasks. Our experimental results show that the AAE could boost the test accuracy by 5.77%.
Qian Song, Feng Xu 0001, Xiao Xiang Zhu 0001, Ya-Qiu Jin
IEEE Trans. Geosci. Remote. Sens.3
2022 CG-Net: Conditional GIS-Aware Network for Individual Building Segmentation in VHR SAR Images
abstract
Object retrieval and reconstruction from very-high-resolution (VHR) synthetic aperture radar (SAR) images are of great importance for urban SAR applications, yet highly challenging due to the complexity of SAR data. This article addresses the issue of individual building segmentation from a single VHR SAR image in large-scale urban areas. To achieve this, we introduce building footprints from geographic information system (GIS) data as a complementary information and propose a novel conditional GIS-aware network (CG-Net). The proposed model learns multilevel visual features and employs building footprints to normalize the features for predicting building masks in the SAR image. We validate our method using a high-resolution spotlight TerraSAR-X image collected over Berlin. Experimental results show that the proposed CG-Net effectively brings improvements with variant backbones. We further compare two representations of building footprints, namely, complete building footprints and sensor-visible footprint segments, for our task, and conclude that the use of the former leads to better segmentation results. Moreover, we investigate the impact of inaccurate GIS data on our CG-Net, and this study shows that CG-Net is robust against positioning errors in the GIS data. In addition, we propose an approach of ground truth generation of buildings from an accurate digital elevation model (DEM), which can be used to generate large-scale SAR image data sets. The segmentation results can be applied to reconstruct 3-D building models at level-of-detail (LoD) 1, which is demonstrated in our experiments.
Yao Sun 0005, Yuansheng Hua, Lichao Mou, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 An Unsupervised Remote Sensing Change Detection Method Based on Multiscale Graph Convolutional Network and Metric Learning
abstract
As a fundamental application, change detection (CD) is widespread in the remote sensing (RS) community. With the increase in the spatial resolution of RS images, high-resolution remote sensing (HRRS) image CD tasks receive growing attention. The change information hidden in multitemporal HRRS images could help discover our planet comprehensively. In the current deep learning era, convolutional neural networks (CNNs) have become one of the most powerful tools for a wide range of RS tasks including HRRS image CD, due to their superb feature learning capacity. However, most of them need a large amount of labeled data to accomplish the CD process, which is challenging or even impractical in many RS applications. Also, given the limited valid receptive field, CNNs can only capture short-range context within HRRS images, which is probably not enough to fully explore change information from the images. To overcome these limitations, in this article, we propose an unsupervised CD method, termed GMCD, based on graph convolutional network (GCN) and metric learning. GMCD consists of a Siamese fully convolution network (FCN), a multiscale dynamic GCN (Mlt-GCN), and a pseudolabel generation mechanism based on metric learning. The Siamese FCN contains a Siamese encoder and a pyramid-shaped decoder, aiming to extract multiscale features and integrate them to generate reliable difference images (DIs). Mlt-GCN focuses on capturing the short- and long-range contextual patterns at feature map level to extract changed and unchanged areas completely. The pseudolabel generation mechanism aims to produce reliable pseudolabels (changed, unchanged, and uncertain) to help accomplish the model training in an unsupervised way. Experiments on four HRRS image CD datasets demonstrate that GMCD outperforms the existing state-of-the-art methods.
Xu Tang 0004, Lichao Mou, Fang Liu 0034, Xiangrong Zhang, Xiao Xiang Zhu 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.6
2022 SCIDA: Self-Correction Integrated Domain Adaptation From Single- to Multi-Label Aerial Images
abstract
Most publicly available datasets for image classification are with single labels, while images are inherently multilabeled in our daily life. Such an annotation gap makes many pretrained single-label classification models fail in practical scenarios. For aerial images, this annotation issue is more concerned: Aerial data naturally cover a relatively large land area with multiple labels, while annotated aerial datasets currently publicly available (e.g., UCM and AID) are single-labeled. As manually annotating multilabel aerial images (MAIs) would be time-/ labor-consuming, we propose a novel self-correction integrated domain adaptation (SCIDA) method for automatic multilabel learning. SCIDA is weakly supervised, i.e., automatically learning the multilabel image classification model from using massive, publicly available single-label images. To achieve this goal, we propose a novel labelwise self-correction (LWC) module to better explore underlying label correlations. This module also makes the unsupervised domain adaptation (UDA) from single-label to multilabel data possible. For model training, the proposed method uses single-label information yet requires no prior knowledge of multilabeled data and predicts labels for MAIs. Through extensive evaluations, the proposed model, which is trained with single-labeled MAI-AID-s and MAI-UCM-s datasets, achieves much better performances than comparative methods on our collected multiscene aerial image dataset. The code and data are available on GitHub (https://github.com/Ryan315/Single2multi-DA).
Tianze Yu, Jianzhe Lin, Lichao Mou, Yuansheng Hua, Xiao Xiang Zhu 0001, Z. Jane Wang 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 From Easy to Hard: Learning Language-Guided Curriculum for Visual Question Answering on Remote Sensing Data
abstract
Visual question answering (VQA) for remote sensing scene has great potential in intelligent human-computer interaction system. Although VQA in computer vision has been widely researched, VQA for remote sensing data (RSVQA) is still in its infancy. There are two characteristics that need to be specially considered for the RSVQA task. 1) No object annotations are available in RSVQA datasets, which makes it difficult for models to exploit informative region representation; 2) There are questions with clearly different difficulty levels for each image in the RSVQA task. Directly training a model with questions in a random order may confuse the model and limit the performance. To address these two problems, in this paper, a multi-level visual feature learning method is proposed to jointly extract language-guided holistic and regional image features. Besides, a self-paced curriculum learning (SPCL)-based VQA model is developed to train networks with samples in an easy-to-hard way. To be more specific, a language-guided SPCL method with a soft weighting strategy is explored in this work. The proposed model is evaluated on three public datasets, and extensive experimental results show that the proposed RSVQA framework can achieve promising performance. Code will be available at https://gitlab.lrz.de/ai4eo/reasoning/VQA-easy2hard.
Zhenghang Yuan, Lichao Mou, Qi Wang 0009, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 Change Detection Meets Visual Question Answering
abstract
The Earth’s surface is continually changing, and identifying changes plays an important role in urban planning and sustainability. Although change detection techniques have been successfully developed for many years, these techniques are still limited to experts and facilitators in related fields. In order to provide every user with flexible access to change information and help them better understand land-cover changes, we introduce a novel task: change detection-based visual question answering (CDVQA) on multi-temporal aerial images. In particular, multi-temporal images can be queried to obtain high level change-based information according to content changes between two input images. We first build a CDVQA dataset including multi-temporal image-question-answer triplets using an automatic question-answer generation method. Then, a baseline CDVQA framework is devised in this work, and it contains four parts: multi-temporal feature encoding, multi-temporal fusion, multi-modal fusion, and answer prediction. In addition, we also introduce a change enhancing module to multi-temporal feature encoding, aiming at incorporating more change-related information. Finally, effects of different backbones and multi-temporal fusion strategies are studied on the performance of CDVQA task. The experimental results provide useful insights for developing better CDVQA models, which are important for future research on this task. The dataset will be available at https://github.com/YZHJessica/CDVQA.
Zhenghang Yuan, Lichao Mou, Zhitong Xiong, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 SAR4LCZ-Net: A Complex-Valued Convolutional Neural Network for Local Climate Zones Classification Using Gaofen-3 Quad-Pol SAR Data
abstract
The recent local climate zones (LCZ) classification scheme provides spatially fine granular descriptions of inner urban morphology. It is universally applicable to cities worldwide and capable of supporting various urban studies. Although optical and dual-pol synthetic aperture radar (SAR) data continue to push the frontiers of this task, the potential of quad-pol SAR data for LCZ classification is not yet explored. In this article, we propose a novel complex-valued convolutional neural network (CNN),SAR4LCZ-Net, to tackle this challenge. SAR4LCZ-Net improves the state-of-the-art by exploiting two facts of this specific task: the semantic hierarchical structure of the LCZ classification scheme and the complex-valued nature of quad-pol SAR data. To validate the performance of our algorithm, we generate a Chinese Gaofen-3 quad-pol SAR dataset for LCZ which covers 31 cities around the world. Results show that the proposed SAR4LCZ-Net improves 2.4% on overall accuracy (OA) and 4.5% on average accuracy (AA) compared with the real-valued CNN with the same structure. Gaofen-3 quad-pol SAR data also showed its advantage over the dual-pol Sentinel-1 data. It enhanced 5.0% on OA and 7.2% on AA in LCZ classification, under a fair comparison with a model trained by Sentinel-1 of the same area.
Rui Zhang 0100, Yuanyuan Wang 0002, Jingliang Hu, Wei Yang 0004, Jie Chen 0009, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.6
2022 Adversarial Shape Learning for Building Extraction in VHR Remote Sensing Images
abstract
Building extraction in VHR RSIs remains a challenging task due to occlusion and boundary ambiguity problems. Although conventional convolutional neural networks (CNNs) based methods are capable of exploiting local texture and context information, they fail to capture the shape patterns of buildings, which is a necessary constraint in the human recognition. To address this issue, we propose an adversarial shape learning network (ASLNet) to model the building shape patterns that improve the accuracy of building segmentation. In the proposed ASLNet, we introduce the adversarial learning strategy to explicitly model the shape constraints, as well as a CNN shape regularizer to strengthen the embedding of shape features. To assess the geometric accuracy of building segmentation results, we introduced several object-based quality assessment metrics. Experiments on two open benchmark datasets show that the proposed ASLNet improves both the pixel-based accuracy and the object-based quality measurements by a large margin. The code is available at: https://github.com/ggsDing/ASLNet.
Lei Ding 0008, Hao Tang 0005, Yilei Shi, Xiao Xiang Zhu 0001, Lorenzo Bruzzone
IEEE Trans. Image Process.5
2021 Compact Neural Architecture Search for Local Climate Zones Classification
abstract
State-of-the-art Computer Vision models achieve impressive performance but with an increasing complexity.Great advances have been made towards automatic model design, but accounting for model performance and low complexity is still an open challenge.In this study, we propose a neural architecture search strategy for high performance low complexity classification models, that combines an efficient search algorithm with mechanisms for reducing complexity.We tested our proposal on a real World remote sensing problem, the Local Climate Zone classification.The results show that our proposal achieves state-of-the-art performance, while being at least 91.8% more compact in terms of size and FLOPs.
René Traoré, Andrés Camero, Xiao Xiang Zhu 0001
ESANN3
2021 InSAR Displacement Time Series Mining: A Machine Learning Approach
abstract
Interferometric Synthetic Aperture Radar (InSAR)-derived surface displacement time series enable a wide range of applications from urban structural monitoring to geohazard assessment. With systematic data acquisitions becoming the new norm for SAR missions, millions of time series are continuously generated. Machine Learning provides a framework for the efficient mining of such big data. Here, we focus on unsupervised mining of the data via clustering the similar temporal patterns and data-driven displacement signal reconstruction from the InSAR time series. We propose a deep Long Short Term Memory (LSTM) autoencoder model which can exploit temporal relations in contrast to the commonly used shallow learning methods, such as Uniform Manifold Approximation and Projection (UMAP). We also modify the loss function to allow the quantification of uncertainties in the time series data. The two approaches are applied to the Lazufre Volcanic Complex located at the central volcanic zone of the Andes and thereby compared.
Homa Ansari, Marc Rußwurm, Sina Montazeri, Alessandro Parizzi, Xiao Xiang Zhu 0001
IGARSS6
2021 Mask-Height R-CNN: An End-to-End Network for 3D Building Reconstruction from Monocular Remote Sensing Imagery
abstract
3D building reconstruction from monocular remote sensing imagery is a promising and economical way to generate 3D city models at a large scale, yet the task is rarely touched. The paper tackles the problem via an end-to-end network. The goal is achieved by a modified network, named Mask-Height R-CNN, based on Mask R-CNN, with an additional height prediction head in the Region Proposal Network (RPN). Unlike most deep learning based methods, the height estimation is done on the instance level instead of pixel level, which does not require the assembly of the height maps and building masks. The proposed network gains good performances on ISPRS datasets, with 3D F1 scores of over 0.8.
Sining Chen, Lichao Mou, Qingyu Li 0001, Yao Sun 0005, Xiao Xiang Zhu 0001
IGARSS5
2021 Internal Learning for Sequence-to-Sequence Cloud Removal via Synthetic Aperture Radar Prior Information
abstract
Many observations acquired via optical satellites are polluted by cloud coverage, impeding a continuous and on-demand monitoring of the Earth. Recent advances in the field of cloud removal consider multi-temporal data to reconstruct pixels covered by clouds at a time point of interest. Yet, the limitation of preceding work is that information gets integrated over time, removing any temporal resolution from the de-clouded end products. In this work we consider a sequence-to-sequence approach, translating cloudy time series to a series of cloud-free multi-spectral images without the need of any external cloud-free data set. Our network is guided by synthetic aperture radar (SAR) information providing a strong prior for the reconstruction of cloud-covered information. We analyze the proposed method by visual inspection of predictions and in terms of error metrics to highlight its benefits. Finally, an ablation study is conducted in which the our network is compared against a baseline model and the effectiveness of the proposed SAR prior is demonstrated.
Patrick Ebel 0002, Michael Schmitt 0003, Xiao Xiang Zhu 0001
IGARSS3
2021 Towards Out-of-Distribution Detection for Remote Sensing
abstract
In remote sensing, distributional mismatch between the training and test data may arise due to several reasons, including unseen classes in the test data, differences in the geographic area, and multi-sensor differences. Deep learning based models may behave in unexpected manners when subjected to test data that has such distributional shifts from the training data, also called out-of-distribution (OOD) examples. Vulnerability to OOD data severely reduces the reliability of deep learning based models. In this work, we address this issue by proposing a model to quantify distributional uncertainty of deep learning based remote sensing models. In particular, we adopt a Dirichlet Prior Network for remote sensing data. The approach seeks to maximize the representation gap between the in-domain and OOD examples for a better identification of unknown examples at test time. Experimental results on three exemplary test scenarios show that the proposed model can detect OOD images in remote sensing.
Jakob Gawlikowski, Sudipan Saha, Anna M. Kruspe, Xiao Xiang Zhu 0001
IGARSS4
2021 An OpenStreetMap-Based Dataset of Building Footprints for Analysing Different Types of Label Noise
abstract
We present a dataset consisting of OpenStreetMap imagery and corresponding building footprint labels. Multiple label sets are provided, each containing a different type of label noise. The purpose of the dataset is to enable a systematic analysis of different label noise types in the earth observation domain and to provide a benchmark dataset for noise removal techniques. We also present some preliminary results from experiments on the effect of different label noise types on model performance.
Jonas Gütter, Anna M. Kruspe, Xiao Xiang Zhu 0001
IGARSS3
2021 Hed-Unet: A Multi-Scale Framework for Simultaneous Segmentation and Edge Detection
abstract
Segmentation models for remote sensing imagery are usually trained on the segmentation task alone. However, for many applications, the class boundaries carry semantic value. To account for this, we propose a new approach that unites both tasks within a single deep learning model. The proposed network architecture follows the successful encoder-decoder approach, and is improved by employing deep supervision at multiple resolution levels, as well as merging these resolution levels into a final prediction using a hierarchical attention mechanism. This framework is trained to detect the coastline in Sentinel-1 images of the Antarctic coastline. Its performance is then compared to conventional single-task approaches, and shown to outperform these methods. The code is available at https://github.com/khdlr/HED-UNet.
Konrad Heidler, Lichao Mou, Celia A. Baumhoer, Andreas J. Dietz, Xiao Xiang Zhu 0001
IGARSS5
2021 Seeing the Bigger Picture: Enabling Large Context Windows in Neural Networks by Combining Multiple Zoom Levels
abstract
When adopting deep learning methods for remote sensing applications, the data usually needs to be cut into patches due to hardware limitations. Clearly, this practice discards a lot of contextual information as the model's information is limited to imagery from the given patch. We propose a memory-efficient way around this limitation by using multiple patches of varying spatial extents on different resolution levels. Finally, this new approach is evaluated for the task of automated sea ice charting, where the added contextual information is shown to be beneficial to model performance.
Konrad Heidler, Lichao Mou, Xiao Xiang Zhu 0001
IGARSS3
2021 Zooming into Uncertainties: Towards Fusing Multi Zoom Level Imagery for Urban Land Use Segmentation
abstract
Urban land use prediction is an ill-posed problem from a remote sensing perspective. Some areas are easy to predict with aerial images, e.g. residential areas or industrial areas, whereas it is nearly impossible to predict land use in dense urban centers with highly mixed land use. In this study, we use a fully convolutional, Bayesian neural network for urban land use segmentation that yields predictions and pixel-wise uncertainty values side-by-side. By adding aleatoric uncertainty to the output of our model, we can assess how much the model benefits from the provided data. We train our network using a dataset from four metropolitan areas in the U.S. on two different zoom levels. Our results show that adding aleatoric uncertainty can improve the IoU scores if a sufficient amount of informative data is provided.
Eike Jens Hoffmann, Xiao Xiang Zhu 0001
IGARSS3
2021 An Overview of Multimodal Remote Sensing Data Fusion: From Image to Feature, From Shallow to Deep
abstract
With the ever-growing availability of different remote sensing (RS) products from both satellite and airborne platforms, simultaneous processing and interpretation of multimodal RS data have shown increasing significance in the RS field. Different resolutions, contexts, and sensors of multimodal RS data enable the identification and recognition of the materials lying on the earth's surface at a more accurate level by describing the same object from different points of the view. As a result, the topic on multimodal RS data fusion has gradually emerged as a hotspot research direction in recent years. This paper aims at presenting an overview of multimodal RS data fusion in several mainstream applications, which can be roughly categorized by 1) image pansharpening, 2) hyperspectral and multispectral image fusion, 3) multimodal feature learning, and (4) crossmodal feature learning. For each topic, we will briefly describe what is the to-be-addressed research problem related to multimodal RS data fusion and give the representative and state-of-the-art models from shallow to deep perspectives.
Danfeng Hong, Jocelyn Chanussot, Xiao Xiang Zhu 0001
IGARSS3
2021 Unconstrained Aerial Scene Recognition with Deep Neural Networks and a New Dataset
abstract
Aerial scene recognition is a fundamental research problem in interpreting high-resolution aerial imagery. Over the past few years, most studies focus on classifying an image into one scene category, while in real-world scenarios, it is more often that a single image contains multiple scenes. Therefore, in this paper, we investigate a more practical yet underexplored task-multi-scene recognition in single images. To this end, we create a large-scale dataset, called Mul-tiScene dataset, composed of 100,000 unconstrained images each with multiple labels from 36 different scenes. Among these images, 14,000 of them are manually interpreted and assigned ground-truth labels, while the remaining images are provided with crowdsourced labels, which are generated from low-cost but noisy OpenStreetMap (OSM) data. By doing so, our dataset allows two branches of studies: 1) developing novel CNNs for multi-scene recognition and 2) learning with noisy labels. We experiment with extensive baseline models on our dataset to offer a benchmark for multi-scene recognition in single images. Aiming to expedite further researches, we will make our dataset and pre-trained models available11https://github.com/Hua-YS/Multi-Scene-Recognition.
Yuansheng Hua, Lichao Mou, Pu Jin, Xiao Xiang Zhu 0001
IGARSS4
2021 Temporal Relations Matter: A Two-Pathway Network for Aerial Video Recognition
abstract
With the increasing volume of aerial videos, the demand for automatically parsing these videos is surging. To achieve this, current researches mainly focus on extracting a holistic feature with convolutions along both spatial and temporal dimensions. However, these methods are limited by small temporal receptive fields and cannot adequately capture long-term temporal dependencies which are important for describing complicated dynamics. In this paper, we propose a novel two-pathway network to model not only holistic features, but also temporal relations for aerial video classification. More specially, our model employs a two-pathway architecture: (1) a holistic representation pathway to learn a general feature of frame appearances and short-term temporal variations and (2) a temporal relation pathway to capture multi-scale temporal relations across arbitrary frames, providing long-term temporal dependencies. Our model is evaluated on event recognition dataset, ERA, and achieves the state-of-the-art results. This demonstrates its effectiveness and good generalization capacity.
Pu Jin, Lichao Mou, Yuansheng Hua, Gui-Song Xia, Xiao Xiang Zhu 0001
IGARSS5
2021 Anomaly Detection in Aerial Videos Via Future Frame Prediction Networks
abstract
By the virtue of high flexibility, low-cost, real-time, and high-resolution data acquisition capacity, unmanned aerial vehicles (UAVs) can be exploited for a wide range of applications, especially in surveillance, inspection, and search fields. Such applications aim to detect potential suspicious events, violent human actions from an untrimmed and lengthy UAV video. Anomaly detection methods are highly in demand because it is unrealistic for human experts to manually detect all abnormal events in image scene. However, anomaly detection methods in aerial videos are rarely studied in the remote sensing community. In this paper, We propose a future frame prediction network based on convolutional variational autoencoder networks to detect anomalous events. Compared to several models, our network has a superior performance.
Pu Jin, Lichao Mou, Gui-Song Xia, Xiao Xiang Zhu 0001
IGARSS4
2021 Blinded by the Light: Monitoring Local Economic Development Over Time With Nightlight Emissions
abstract
Nighttime light (NTL) emissions are widely used across disciplines to map the spatial distribution of a variety of socioeconomic variables. For economic studies, NTLs allow to proxy for levels of economic indicators at the country level as well as the local level. Further, multi-temporal differences in NTL intensity are also related to GDP differences on the country level. In this study, we investigate if this relation in temporal differences also holds for the local level. We test this with DMSP as well as VIIRS NTLs data from 2010–2015 in Nigeria, Tanzania, and Uganda. Even though we successfully map local levels of socio-economic status with NTLs, we find multi-temporal changes in NTLs at this local level to be uncorrelated with socio-economic development over time. We conclude that luminosity values based on current DMSP and VIIRS sensors are no silver bullet in measuring local economic changes over time and should, if at all, only be used with caution for this.
Lukas Kondmann, Hannes Taubenböck, Xiao Xiang Zhu 0001
IGARSS3
2021 End-to-End Semantic Segmentation and Boundary Regularization of Buildings from Satellite Imagery
abstract
Building footprint generation is a vital task of satellite imagery interpretation. However, the segmentation masks of buildings obtained by existing semantic segmentation networks often have blurred boundaries and irregular shapes. In this research, we propose a new boundary regularization network for building footprint generation in satellite images. More specifically, we consider semantic segmentation and boundary regularization in an end-to-end generative adversarial network (GAN). The learned building footprints are regularized by the interplay between the generator and discriminator. By doing so, the straight boundaries and geometric details of the building could be preserved. Experiments are conducted on a collected dataset of Planetscope satellite imagery (spatial resolution: 4.77 m/pixel). Our approach is much superior to the state-of-the-art methods in both quantitative and qualitative results.
Qingyu Li 0001, Stefano Zorzi, Yilei Shi, Friedrich Fraundorfer, Xiao Xiang Zhu 0001
IGARSS5
2021 Improving Land Cover Classification with a Shift-Invariant Center-Focusing Convolutional Neural Network
abstract
Convolutional neural networks (CNNs) are widely employed in remote sensing community. The CNN-based, also known as patch-based land cover classification method has gained increasing attention. However, this method very often requires the aid of post-processing, otherwise it is difficult to obtain accurate boundaries separating different land cover classes. In this paper, we discuss the reason of this phenomenon and propose a shift-invariant center-focusing (SICF) network to deliver more accurate boundaries to improve the patch-based land cover classification. The principle of SICF is calculating the class score from a center-focusing area based on a shift-invariant feature extraction module to calibrate prediction. We employ three modern CNNs to build corresponding SICF networks, the evaluation results indicate that compared with the conventional CNNs, the improvements made by SICF for delivering accurate boundaries in land cover classification are significant.
Yuansheng Hua, Lichao Mou, Xiao Xiang Zhu 0001
IGARSS4
2021 Mitigating Spatial and Spectral Differences for Change Detection Using Super-Resolution and Unsupervised Learning
abstract
Change detection (CD) is one of the most researched areas in remote sensing. However, most CD methods assume that the pre-change and post-change images are acquired by the same sensor, having the same set of spectral bands and same spatial resolution. This severely limits the applicability of CD methods. It is not trivial to apply the existing CD methods in multisensor scenario. Towards this direction, we propose an unsupervised CD method that can handle large differences in spatial resolution and can work with completely different set of spectral bands. The proposed method uses a self-supervised super-resolution strategy to upsample the lower resolution image, thus mitigating differences in spatial resolution. To mitigate spectral differences, a self-supervised learning strategy is used that ingests both images as input and trains a network using self-supervised loss accounting for the spectral differences in both images. Once trained this network is used in deep change vector analysis framework for change detection. We validated the proposed method in an experimental setup where the pre-change and post-change images have different spatial resolution (10m and 20 m/pixel) and completely disjoint set of spectral bands.
Jonathan Prexl, Sudipan Saha, Xiao Xiang Zhu 0001
IGARSS3
2021 Super-Resolving Sar Tomography Using Deep Learning
abstract
Synthetic aperture radar tomography (TomoSAR) has been widely employed in 3-D urban mapping. However, state-of-the-art super-resolving TomoSAR algorithms are computationally expensive, because conventional numerical solvers need to solve the$l_{2^{-}}l_{1}$mix norm minimization. This paper proposes a computationally efficient super-resolving To-moSAR inversion algorithm based on deep learning. We studied the potential of deep learning to mimic a conventional$l_{2}-l_{1}$mix norm solver, i.e. iterative shrinkage thresholding algorithm (ISTA), and proposed several improvements of the complex-valued learned ISTA for TomoSAR inversion. Investigation on the super-resolution ability and estimator efficiency of the proposed algorithm shows that the proposed algorithm approaches the Cramer Rao lower bound (CRLB) with a computational efficiency more than 100 times better than the conventional solver.
Kun Qian 0020, Yuanyuan Wang 0002, Yilei Shi, Xiao Xiang Zhu 0001
IGARSS4
2021 Trusting Small Training Dataset for Supervised Change Detection
abstract
Deep learning (DL) based supervised change detection (CD) models require large labeled training data. Due to the difficulty of collecting labeled multi-temporal data, unsupervised methods are preferred in the CD literature. However, unsupervised methods cannot fully exploit the potentials of data-driven deep learning and thus they are not absolute alternative to the supervised methods. This motivates us to look deeper into the supervised DL methods and investigate how they can be adopted intelligently for CD by minimizing the requirement of labeled training data. Towards this, in this work we show that geographically diverse training dataset can yield significant improvement over less diverse training datasets of the same size. We propose a simple confidence indicator for verifying the trustworthiness/confidence of supervised models trained with small labeled dataset. Moreover, we show that for the test cases where supervised CD model is found to be less confident/trustworthy, unsupervised methods often produce better result than the supervised ones.
Sudipan Saha, Biplab Banerjee, Xiao Xiang Zhu 0001
IGARSS3
2021 Generation of Large Scale 3-D City Models Using Insar and Optical Data
abstract
Interferometric synthetic aperture radar (InSAR) techniques are powerful tool for reconstructing the 3-D position of scatterers, especially for the urban areas. Since the estimation accuracy depends on the inverse of number of interferograms and signal-to-noise ratio (SNR), it is necessary to use as many as possible interferograms in order to achieve more accurate result. However, the number of interferograms of TanDEM-X data is generally limited for most areas. Therefore, in order to maintain the estimation accuracy, one feasible way is to increase the SNR. In this work, we proposed a novel framework, which integrates the non-local procedure into SAR tomography inversion and combines the robust estimation. A large-scale demonstration has been carried out with five TanDEM-X bistatic data, which covers the entire city of Munich, Germany. Quantitative evaluation of the reconstructed result with the LiDAR reference exhibits the standard deviation of the height difference is within two meters, which implies the proposed framework has great potential for high quality large-scale 3-D urban modeling.
Yilei Shi, Richard Bamler, Yuanyuan Wang 0002, Xiao Xiang Zhu 0001
IGARSS4
2021 Conditional GIS-Aware Network for Individual Building Segmentation in a VHR SAR Image
abstract
In this paper, we propose a network for individual building segmentation from a single VHR SAR image. The proposed network employs building footprints from GIS data in learning multi-level visual features to predict building masks in the SAR image. Experimental results over Berlin show that the proposed network effectively brings improvements with variant backbones. In addition, we propose an approach for generating building labels from an accurate digital elevation model (DEM), which can be used to generate large-scale SAR image datasets.
Yao Sun 0005, Yuansheng Hua, Lichao Mou, Xiao Xiang Zhu 0001
IGARSS4
2021 Efficient SAR Tomographic Inversion via Sparse Bayesian Learning
abstract
SAR tomographic inversion (TomoSAR) has been widely employed for 3-D urban mapping. Existing algorithms are mostly based on an explicit inversion of the SAR imaging model, which are often computationally expensive for large scale processing. This is especially true for compressive sensing-based TomoSAR algorithms. Previous literature showed perspective of using data-driven methods like PCA and kernel PCA to decompose the signal and reduce the computational complexity of parameter inversion. This paper gives a preliminary demonstration of a data-driven TomoSAR method based on sparse Bayesian learning. Experiments on simulated data show the proposed algorithm can provide moderate detection rate and super-resolution power, comparing to the state-of-the-art compressive sensing based algorithms. As the proposed algorithm is purely based on conventional (non-superresolving) estimators, it is much more computationally efficient than compressive sensing based ones. This gives us a perspective of employing it for large scale TomoSAR processing. Experiments on real data will be given in the final paper.
Yuanyuan Wang 0002, Kun Qian 0020, Xiao Xiang Zhu 0001
IGARSS3
2021 Self-Paced Curriculum Learning for Visual Question Answering on Remote Sensing Data
abstract
Answering questions with natural language by extracting information from image has great potential in various applications. Although visual question answering (VQA) for natural image has been broadly studied, VQA for remote sensing data is still in the early research stage. For the same remote sensing image, there exist questions with dramatically different difficulty-levels. Treating these questions equally may mislead the model and limit the VQA model performance. Considering this problem, in this work, we propose a self-paced curriculum learning (SPCL) based VQA model with hard and soft weighting strategies for remote sensing data. Like human learning process, the model is trained from easy to hard question samples gradually. Extensive experimental results on two datasets demonstrate that the proposed training method can achieve promising performance.
Zhenghang Yuan, Lichao Mou, Xiao Xiang Zhu 0001
IGARSS3
2021 Whats Next in AI4EO?
abstract
Surfing the wave of artificial intelligence (AI) and thanks to open Earth observation (EO) data, there has been a buzz in the community for some time about exploiting AI methods like deep learning (DL) to solve problems in EO that have been solved before (“we can also do it with deep learning”) yet to achieve superior performances. There is no doubt that AI/ML/DL can do a great job in classification and detection tasks, which are only small fractions of EO problems. Is this all? Are you getting bored by applying AI methods in EO and playing with toy data? As the introduction talk of the invited session “AI4EO: Reasoning, Uncertainty and Ethics”, two important RD lines are discussed, namely developing cutting-edge next generation AI4EO methods, and generating large scale and representative benchmarks enabling these developments. With all the advancements in EO and AI technology, so far non-existing geo-information has and will become openly available that may raise ethical issues, like mis-use, bias or stigmatization. Hence, ethical accompanying AI4EO research is becoming increasingly important.
Xiao Xiang Zhu 0001
IGARSS1
2021 Semisupervised Change Detection Using Graph Convolutional Network
abstract
Most change detection (CD) methods are unsupervised as collecting substantial multitemporal training data is challenging. Unsupervised CD methods are driven by heuristics and lack the capability to learn from data. However, in many real-world applications, it is possible to collect a small amount of labeled data scattered across the analyzed scene. Such a few scattered labeled samples in the pool of unlabeled samples can be effectively handled by graph convolutional network (GCN) that has recently shown good performance in semisupervised single-date analysis, to improve change detection performance. Based on this, we propose a semisupervised CD method that encodes multitemporal images as a graph via multiscale parcel segmentation that effectively captures the spatial and spectral aspects of the multitemporal images. The graph is further processed through GCN to learn a multitemporal model. Information from the labeled parcels is propagated to the unlabeled ones over training iterations. By exploiting the homogeneity of the parcels, the model is used to infer the label at a pixel level. To show the effectiveness of the proposed method, we tested it on a multitemporal Very High spatial Resolution (VHR) data set acquired by Pleiades sensor over Trento, Italy.
Sudipan Saha, Lichao Mou, Xiao Xiang Zhu 0001, Francesca Bovolo, Lorenzo Bruzzone
IEEE Geosci. Remote. Sens. Lett.3
2021 Joint and Progressive Subspace Analysis (JPSA) With Spatial-Spectral Manifold Alignment for Semisupervised Hyperspectral Dimensionality Reduction
abstract
Conventional nonlinear subspace learning techniques (e.g., manifold learning) usually introduce some drawbacks in explainability (explicit mapping) and cost effectiveness (linearization), generalization capability (out-of-sample), and representability (spatial-spectral discrimination). To overcome these shortcomings, a novel linearized subspace analysis technique with spatial-spectral manifold alignment is developed for a semisupervised hyperspectral dimensionality reduction (HDR), called joint and progressive subspace analysis (JPSA). The JPSA learns a high-level, semantically meaningful, joint spatial-spectral feature representation from hyperspectral (HS) data by: 1) jointly learning latent subspaces and a linear classifier to find an effective projection direction favorable for classification; 2) progressively searching several intermediate states of subspaces to approach an optimal mapping from the original space to a potential more discriminative subspace; and 3) spatially and spectrally aligning a manifold structure in each learned latent subspace in order to preserve the same or similar topological property between the compressed data and the original data. A simple but effective classifier, that is, nearest neighbor (NN), is explored as a potential application for validating the algorithm performance of different HDR approaches. Extensive experiments are conducted to demonstrate the superiority and effectiveness of the proposed JPSA on two widely used HS datasets: 1) Indian Pines (92.98%) and 2) the University of Houston (86.09%) in comparison with previous state-of-the-art HDR methods. The demo of this basic work (i.e., ECCV2018) is openly available at https://github.com/danfenghong/ECCV2018_J-Play.
Danfeng Hong, Naoto Yokoya, Jocelyn Chanussot, Jian Xu 0008, Xiao Xiang Zhu 0001
IEEE Trans. Cybern.5
2021 SDFL-FC: Semisupervised Deep Feature Learning With Feature Consistency for Hyperspectral Image Classification
abstract
Semisupervised deep learning methods (DLMs) can mitigate the dependence on large amounts of labeled samples using a small number of labeled samples. However, for semisupervised deep feature learning (SDFL), the quality of extracted features cannot be well ensured without a certain amount of labeled samples. To address this issue, we develop the SDFL method with feature consistency (SDFL-FC) for the hyperspectral image (HSI) classification. The SDFL-FC first adopts the convolutional neural network (CNN) to extract spectral–spatial features of HSI and then uses the fully connected layers (FCLs) to model the feature consistency. Moreover, two constraints that enforce both the feature consistency of single pixel (FCS) and feature consistency of group pixels (FCG) are introduced to obtain the representative and discriminative features. The FCS is achieved by the generative adversarial network (GAN) regularization, which can reconstruct the original data from extracted features. The FCG is based on the assumption that the features of group pixels should have similar characteristics within a superpixel, which is embedded in each FCL. The final FCL outputs the class labels, and the cross-entropy (CE) loss is calculated with the labeled samples, while the two losses of FCS and FCG are calculated with all the training samples (both labeled and unlabeled). SDFL-FC integrates the FCS, FCG, and CE loss into a unified objective function and uses a customized iterative optimization algorithm to optimize it. Experiments demonstrate that the SDFL-FC can outperform the related state-of-the-art HSI classification methods.
Yuebin Wang, Junhuan Peng, Chunping Qiu, Lei Ding 0008, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.6
2021 Multisensor Data Fusion for Cloud Removal in Global and All-Season Sentinel-2 Imagery
abstract
The majority of optical observations acquired via spaceborne Earth imagery are affected by clouds. While there is numerous prior work on reconstructing cloud-covered information, previous studies are, oftentimes, confined to narrowly defined regions of interest, raising the question of whether an approach can generalize to a diverse set of observations acquired at variable cloud coverage or in different regions and seasons. We target the challenge of generalization by curating a large novel data set for training new cloud removal approaches and evaluate two recently proposed performance metrics of image quality and diversity. Our data set is the first publically available to contain a global sample of coregistered radar and optical observations, cloudy and cloud-free. Based on the observation that cloud coverage varies widely between clear skies and absolute coverage, we propose a novel model that can deal with either extreme and evaluate its performance on our proposed data set. Finally, we demonstrate the superiority of training models on real over synthetic data, underlining the need for a carefully curated data set of real observations. To facilitate future research, our data set is made available online.
Patrick Ebel 0002, Andrea Meraner, Michael Schmitt 0003, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.4
2021 Single-Look Multi-Master SAR Tomography: An Introduction
abstract
This article addresses the general problem of single-look multi-master SAR tomography. For this purpose, we establish the single-look multi-master data model, analyze its implications for the single and double scatterers, and propose a generic inversion framework. The core of this framework is the nonconvex sparse recovery, for which we develop two algorithms: one extends the conventional nonlinear least squares (NLS) to the single-look multi-master data model and the other is based on bi-convex relaxation and alternating minimization (BiCRAM). We provide two theorems for the objective function of the NLS subproblem, which lead to its analytic solution up to a constant phase angle in the 1-D case. We also report our findings from the experiments on different acceleration techniques for BiCRAM. The proposed algorithms are applied to a real TerraSAR-X data set and validated with the height ground truth made available by an SAR imaging geodesy and simulation framework. This shows empirically that the single-master approach, if applied to a single-look multi-master stack, can be insufficient for layover separation, and the multi-master approach can indeed perform slightly better (despite being computationally more expensive) even in the case of single scatterers. In addition, this article also sheds light on the special case of single-look bistatic SAR tomography, which is relevant for the current and future SAR missions such as TanDEM-X and Tandem-L.
Nan Ge, Richard Bamler, Danfeng Hong, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.4
2021 A Multispectral and Multiangle 3-D Convolutional Neural Network for the Classification of ZY-3 Satellite Images Over Urban Areas
abstract
The recent availability of high-resolution multiview ZY-3 satellite images, with angular information, can provide an opportunity to capture 3-D structural features for classification. In high-resolution image classification over urban areas, objects with diverse vertical structures make urban landscape more heterogeneous in 3-D space and consequently can make the classification challenging. In this article, a novel multiangle gray-level cooccurrence tensor feature is proposed based on the multiview bands of the ZY-3 imagery, namely, GLCMMA–T. The GLCMMA–Tfeature captures the distributions of the gray-level spatial variation under different viewing angles, which can depict the 3-D textures and structures of urban objects. The spectral and GLCMMA–Ttensor features are interpreted by two 3-D convolutional neural network (CNN) streams and then concatenated as the input to the fully connected layer. This novel multispectral and multiangle 3-D convolutional neural network (M2-3-DCNN) combines the spectral and angular information, and the fused feature has the potential to provide a comprehensive description of urban objects with complex vertical structures. The experimental results on ZY-3 multiview images from four test areas indicate that the proposed method can significantly improve the classification accuracy when compared with several state-of-the-art multiangle features and deep-learning-based image classification methods.
Xin Huang 0002, Jiayi Li 0001, Xiuping Jia, Jun Li 0009, Xiao Xiang Zhu 0001, Jón Atli Benediktsson
IEEE Trans. Geosci. Remote. Sens.6
2021 Attention-Aware Pseudo-3-D Convolutional Neural Network for Hyperspectral Image Classification
abstract
Convolutional neural networks (CNNs) have been applied for hyperspectral image classification recently. Among this class of deep models, 3-D CNN has been shown to be more effective by learning discriminative features from abundant spectral signatures and spatial contexts in hyperspectral imagery (HSI). However, by simply imposing 3-D CNN to HSI, a large amount of initial information might be lost in this CNN pipeline. The proposed attention-aware pseudo-3-D (AP3D) convolutional network for HSI classification is motivated by two observations. First, each dimension of the 3-D HSI is not equally important, different attention should be paid to different dimensions of the initial HSI image, especially in the first convolution operation. Second, intermediate representations of the 3-D input image at different stages in the 3-D CNN pipeline represent different levels of features and should not be neglected and abandoned. Instead, a 2-D matrix of scores for each feature map should be fed to the final softmax layer. Quantitative and qualitative results demonstrate that the proposed AP3D model outperforms the state-of-the-art HSI classification methods in agricultural and rural/urban data sets: Indian Pines, Pavia University, and Salinas Scene.
Jianzhe Lin, Lichao Mou, Xiao Xiang Zhu 0001, Xiangyang Ji, Z. Jane Wang 0001
IEEE Trans. Geosci. Remote. Sens.3
2021 Unifying Top-Down Views by Task-Specific Domain Adaptation
abstract
In this article, we aim to learn a unified representation of images from satellite/aerial/ground views by exploring their underlying correlations. Inspired by recent advances in domain adaptation (DA), we propose a novel task-specific DA method for this purpose. Different from traditional DA methods, this proposed method not only applies task-specific classifiers1but also introduces domain-specific tasks for different domains during the adaptation process. The experiments are conducted on two newly proposed ground-/satellite-to-aerial scene adaptation (GSSA) data sets. Since the semantic gap between the ground/satellite scenes and the aerial scenes is much larger than that between ground scenes, the DA task between these scenes is more challenging than traditional DA tasks. On GSSA data sets, we not only demonstrate the proposed unsupervised DA method but also explore the few-shot DA in the discussion section. The proposed method is easy to implement, and our method substantially outperforms the state-of-the-art methods on the studied data sets.We hope that the proposed method for the novel GSSA data sets can be a good baseline for future researchers. The related data sets/codes will be available online.
Jianzhe Lin, Tianze Yu, Lichao Mou, Xiao Xiang Zhu 0001, Rabab K. Ward, Z. Jane Wang 0001
IEEE Trans. Geosci. Remote. Sens.4
2021 Endmember Bundle Extraction Based on Multiobjective Optimization
abstract
A number of endmember extraction methods have been developed to identify pure pixels in hyperspectral images (HSIs). The majority of them use only one spectrum to represent one kind of material, which ignores the spectral variability problem that particularly characterizes a HSI with high spatial resolution. Only a few algorithms have been developed to identify multiple endmembers representing the spectral variability within each class, called endmember bundle extraction (EBE). This article introduces multiobjective particle swarm optimization for the identification of multiple endmember spectra with variability. Unlike existing convex geometry-based EBE methods, which operate on a single geometry of the dataspace, the proposed method divides the observed data into subsets along the spectral dimension and simultaneously operates on multiple dataspaces to obtain candidate endmembers based on multiobjective particle swarm optimization. The candidate endmembers are then refined by spatial post-processing and sequential forward floating selection to produce the final result. Experiments are conducted on both synthetic and real hyperspectral data to demonstrate the effectiveness of the proposed method in comparison with several state-of-the-art methods.
Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.2
2021 SAR Tomography via Nonlinear Blind Scatterer Separation
abstract
Layover separation has been fundamental to many synthetic aperture radar applications, such as building reconstruction and biomass estimation. Retrieving the scattering profile along the mixed dimension (elevation) is typically solved by inversion of the synthetic aperture radar (SAR) imaging model, a process known as SAR tomography. This article proposes a nonlinear blind scatterer separation method to retrieve the phase centers of the layovered scatterers, avoiding the computationally expensive tomographic inversion. We demonstrate that conventional linear separation methods, for example, principle component analysis (PCA), can only partially separate the scatterers under good conditions. These methods produce systematic phase bias in the retrieved scatterers due to the nonorthogonality of the scatterers' steering vectors, especially when the intensities of the sources are similar or the number of images is low. The proposed method artificially increases the dimensionality of the data using kernel PCA, hence mitigating the aforementioned limitations. In the processing, the proposed method sequentially deflates the covariance matrix using the estimate of the brightest scatterer from kernel PCA. Simulations demonstrate the superior performance of the proposed method over conventional PCA-based methods in various respects. Experiments using TerraSAR-X data show an improvement in height reconstruction accuracy by a factor of one to three, depending on the used number of looks.
Yuanyuan Wang 0002, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.2
2020 Cross-Task Transfer for Geotagged Audiovisual Aerial Scene Recognition
Di Hu 0001, Xuhong Li 0002, Lichao Mou, Pu Jin, Liping Jing, Xiao Xiang Zhu 0001, Dejing Dou
ECCV (24)7
2020 Cross-Attention in Coupled Unmixing Nets for Unsupervised Hyperspectral Super-Resolution
Jing Yao 0002, Danfeng Hong, Jocelyn Chanussot, Deyu Meng, Xiao Xiang Zhu 0001, Zongben Xu
ECCV (29)5
2020 A Novel Actor Dual-Critic Model for Remote Sensing Image Captioning
abstract
We deal with the problem of generating textual captions from optical remote sensing (RS) images using the notion of deep reinforcement learning. Due to the high inter-class similarity in reference sentences describing remote sensing data, jointly encoding the sentences and images encourages prediction of captions that are semantically more precise than the ground truth in many cases. To this end, we introduce an Actor Dual-Critic training strategy where a second critic model is deployed in the form of an encoder-decoder RNN to encode the latent information corresponding to the original and generated captions. While all actor-critic methods use an actor to predict sentences for an image and a critic to provide rewards, our proposed encoder-decoder RNN guarantees high-level comprehension of images by sentence-to-image translation. We observe that the proposed model generates sentences on the test data highly similar to the ground truth and is successful in generating even better captions in many critical cases. Extensive experiments on the benchmark Remote Sensing Image Captioning Dataset (RSICD) and the UCM-captions dataset confirm the superiority of the proposed approach in comparison to the previous state-of-the-art where we obtain a gain of sharp increments in both the ROUGE-L and CIDEr measures.
Ruchika Chavhan, Biplab Banerjee, Xiao Xiang Zhu 0001, Subhasis Chaudhuri
ICPR3
2020 ADVANCING DEEP LEARNING FOR EARTH SCIENCES: FROM HYBRID MODELING TO INTERPRETABILITY
abstract
Machine learning and deep learning in particular have made a huge impact in many fields of science and engineering. In the last decade, advanced deep learning methods have been developed and applied to remote sensing and geoscientific data problems extensively. Applications on classification and parameter retrieval are making a difference: methods are very accurate, can handle large amounts of data, and can deal with spatial and temporal data structures efficiently. Nevertheless, several important challenges need still to be addressed. First, current standard deep architectures cannot deal with long-range dependencies so distant driving processes (in space or time) are not captured, and they cannot cope with non-Euclidean spaces efficiently. Second, as other data-driven techniques, deep learning models do not necessarily respect physical or causal relations. Finally, deep learning models are still obscure and resistant to interpretability. Advances are needed to cope with arbitrary signal structures and data relations, physical plausibility and interpretability. This paper discusses about ways forward to develop new DL methods for the Earth sciences in all three directions.
Gustau Camps-Valls, Markus Reichstein, Xiao Xiang Zhu 0001, Devis Tuia
IGARSS3
2020 Cloud Removal in Unpaired Sentinel-2 Imagery Using Cycle-Consistent GAN and SAR-Optical Data Fusion
abstract
The majority of optical images acquired via spaceborne remote sensing are affected by clouds. Recent advances in cloud removal combine multimodal data with deep neural networks recovering the affected areas. To relax the requirements on the data the network is trained on previous approaches utilized generative models no longer necessitating strict pixel-wise correspondences between cloudy input and cloud-free target images. However, such models are often-times prone to fiction, i.e. the generation of content systematically differing from the structure of the target images. In this work we combine the fusion of optical and radar imagery with the advantages of generative models trainable on unpaired optical data, while reducing fiction by reconstructing optical information only where it need be-over cloud-covered areas. We evaluate our approach qualitatively and quantitatively and demonstrate its effectiveness.
Patrick Ebel 0002, Michael Schmitt 0003, Xiao Xiang Zhu 0001
IGARSS3
2020 On the Fusion Strategies of Sentinel-1 and Sentinel-2 Data for Local Climate Zone Classification
abstract
Local Climate Zone (LCZ) classification is the most commonly used scheme to analyze how local urban morphology affects the climate of local areas. Classification methods are often based on remote sensing data or on a fusion of several data sources. In this study, the effects of different fusion strategies of optical and synthetic aperture radar (SAR) data on the accuracy of LCZ classifications are investigated. The data processing is implemented with a convolutional neural network (CNN), where until a fusion layer, separate data sources are processed separately on branches. Strategies of splitting the data into branches and the effects of different fusion stages are compared, together with approaches based on sums of independent classifiers. For our setting, the stage of fusion does not seem to have a big influence on the accuracy. The results of this study contribute to a better understanding of cooperative usage of multispectral and SAR data.
Jakob Gawlikowski, Michael Schmitt 0003, Anna M. Kruspe, Xiao Xiang Zhu 0001
IGARSS4
2020 Unsupervised Hyperspectral Embedding by Learning a Deep Regression Network
abstract
This work presents a novel hyperspectral embedding technique by learning a deep regression network in an unsupervised fashion, which aims at reducing the computational complexity and storage-costing of traditional manifold embedding methods as well as improving the representation ability of spectral signatures effectively. The proposed method attempts to learn an explicit and unified nonlinear mapping from all patch-wise correspondences of original hyperspectral data and dimension-reduced products generated by some existing manifold learning approaches. This process can be well performed by means of a deep regression model. The learned model is not only capable of locally capturing the manifold structure of the whole hyperspectral image from densely patch-based random sampling but also better applicable to high-efficient out-of-sample inference. Experimental results conducted on the real hyperspectral data demonstrate the effectiveness and superiority of the proposed hyperspectral embedding technique.
Danfeng Hong, Jing Yao 0002, Jocelyn Chanussot, Xiao Xiang Zhu 0001
IGARSS4
2020 Learning Multi-Label Aerial Image Classification Under Label Noise: A Regularization Approach Using Word Embeddings
abstract
Training deep neural networks requires well-annotated datasets. However, real world datasets are often noisy, especially in a multi-label scenario, i.e. where each data point can be attributed to more than one class. To this end, we propose a regularization method to learn multi-label classification networks from noisy data. This regularization is based on the assumption that semantically close classes are more likely to appear together in a given image. Hereby, we encode label correlations with prior knowledge and regularize noisy network predictions using label correlations. To evaluate its effectiveness, we perform experiments on a mutli-label aerial image dataset contaminated with controlled levels of label noise. Results indicate that networks trained using the proposed method outperform those directly learned from noisy labels and that the benefits increase proportionally to the amount of noise present.
Yuansheng Hua, Sylvain Lobry, Lichao Mou, Devis Tuia, Xiao Xiang Zhu 0001
IGARSS5
2020 Instance Segmentation of Buildings Using Keypoints
abstract
Building segmentation is of great importance in the task of remote sensing imagery interpretation. However, the existing semantic segmentation and instance segmentation methods often lead to segmentation masks with blurred boundaries. In this paper, we propose a novel instance segmentation network for building segmentation in high-resolution remote sensing images. More specifically, we consider segmenting an individual building as detecting several keypoints. The detected keypoints are subsequently reformulated as a closed polygon, which is the semantic boundary of the building. By doing so, the sharp boundary of the building could be preserved. Experiments are conducted on selected Aerial Imagery for Roof Segmentation (AIRS) dataset, and our method achieves better performance in both quantitative and qualitative results with comparison to the state-of-the-art methods. Our network is a bottom-up instance segmentation method that could well preserve geometric details.
Qingyu Li 0001, Lichao Mou, Yuansheng Hua, Yao Sun 0005, Pu Jin, Yilei Shi, Xiao Xiang Zhu 0001
IGARSS7
2020 Event and Activity Recognition in Aerial Videos Using Deep Neural Networks and a New Dataset
abstract
Unmanned aerial vehicles (UAVs) are now widespread available. Yet the more UAVs there are in the skies, the more video data they create. It is unrealistic for humans to screen such big data and understand their contents. Hence methodological research on UAV video content understanding is of great importance. In this paper, we introduce a novel task of event recognition in unconstrained aerial videos in the remote sensing community and present a dataset for this task. Organized in a rich semantic taxonomy, the proposed dataset covers a wide range of events involving diverse environments and scales. We report results of plenty of deep networks in two ways: single-frame classification and video classification. The dataset and trained models can be downloaded from https://1cmou.github.io/ERA_Dataset/.
Lichao Mou, Yuansheng Hua, Pu Jin, Xiao Xiang Zhu 0001
IGARSS4
2020 Model and Data Uncertainty for Satellite Time Series Forecasting with Deep Recurrent Models
abstract
Deep Learning is often criticized as being a black-box method that provides accurate predictions, but a limited explanation of the underlying processes and no indication when to not trust those predictions. Equipping existing deep learning models with an (general) notion of uncertainty can help mitigate both these issues. The Bayesian deep learning community has developed model-agnostic methodology to estimate both data and model uncertainty that can be implemented on top of existing deep learning models. In this work, we test this methodology for deep recurrent satellite time series forecasting and test its assumptions on data and model uncertainty. We tested its effectiveness on an application on climate change where the activity of seasonal vegetation decreased over multiple years.
Marc Rußwurm, Xiao Xiang Zhu 0001, Yarin Gal, Marco Körner 0001
IGARSS3
2020 A Novel Approach to Unsupervised Segmentation of Multitemporal VHR Images based on Deep Learning
abstract
Very-high-resolution (VHR) multi-temporal images are important in remote sensing to monitor the dynamics of the Earth surface. Image semantic segmentation classifies pixels and assigns them label from meaningful object groups. It has been extensively studied in context of single image analysis, however not explored for multi-temporal one. In this paper we propose to extend supervised semantic segmentation to the unsupervised joint segmentation of multi-temporal images. The proposed method processes multi-temporal images by separately feeding them to a deep network comprising of trainable convolutional layers. The training process does not involve any external label. Segmentation labels are obtained from argmax classification of the final layer. Multi-temporal segmentation labels and weights of the trainable layers are jointly optimized in iterations. We tested the method on a VHR dataset from Trento, Italy. Both quantitative and qualitative results demonstrated the effectiveness of the proposed approach.
Sudipan Saha, Lichao Mou, Chunping Qiu, Xiao Xiang Zhu 0001, Francesca Bovolo, Lorenzo Bruzzone
IGARSS4
2020 Building Extraction by Gated Graph Convolutional Neural Network with Deep Structured Feature Embedding
abstract
Building footprint information is an essential ingredient for 3-D reconstruction of urban models. The automatic generation of building footprints from satellite images presents a considerable challenge due to the complexity of building shapes. Recent developments in deep convolutional neural networks (DCNNs) have enabled accurate pixel-level labeling tasks. One central issue remains, which is the precise delineation of boundaries. Deep architectures generally fail to produce fine-grained segmentation with accurate boundaries due to progressive downsampling. In this work, we we introduce a generic framework to overcome the issue, integrating the gated graph convolutional network (GGCN) and deep structured feature embedding (DSFE) into an end-to-end workflow.
Yilei Shi, Qinyu Li, Xiao Xiang Zhu 0001
IGARSS3
2020 Vision-Based Scattering Key-Frame Extraction for VideoSAR Summarization
abstract
Video synthetic aperture radar (VideoSAR) presents significant potential for improving the performance of information interpretation. Key frames represent the aspect-dependent electromagnetic energy, which frequently obscures other scattering physics dominated by specular returns. In this paper, we propose a vision-based background subtraction approach for capturing VideoSAR scattering key-frame information in simultaneously single-channel and single-pass configurations. The spatiotemporal key-frame extractor combines subaperture energy gradient with modified statistical and knowledge-based object tracker. It can robustly discriminate the scattering features of the alternation between transient persistence and disappearance. We evaluate the proposed method using several measured airborne data. Experimental results and performance comparison have demonstrated that the scattering key-frame extractor can achieve a high accuracy for VideoSAR summarization.
Ying Zhang 0049, Lichao Mau, Daiyin Zhu, Xiao Xiang Zhu 0001
IGARSS4
2020 Dual Adversarial Network for Unsupervised Ground/Satellite-to-Aerial Scene Adaptation
abstract
Recent domain adaptation work tends to obtain a uniformed representation in an adversarial manner through joint learning of the domain discriminator and feature generator. However, this domain adversarial approach could render sub-optimal performances due to two potential reasons: First, it might fail to consider the task at hand when matching the distributions between the domains. Second, it generally treats the source and target domain data in the same way. In our opinion, the source domain data which serves the feature adaption purpose should be supplementary, whereas the target domain data mainly needs to consider the task-specific classifier. Motivated by this, we propose a dual adversarial network for domain adaptation, where two adversarial learning processes are conducted iteratively, in correspondence with the feature adaptation and the classification task respectively. The efficacy of the proposed method is first demonstrated on Visual Domain Adaptation Challenge (VisDA) 2017 challenge, and then on two newly proposed Ground/Satellite-to-Aerial Scene adaptation tasks. For the proposed tasks, the data for the same scene is collected not only by the traditional camera on the ground, but also by satellite from the out space and unmanned aerial vehicle (UAV) at the high-altitude. Since the semantic gap between the ground/satellite scene and the aerial scene is much larger than that between ground scenes, the newly proposed tasks are more challenging than traditional domain adaptation tasks. The datasets/codes can be found at https://github.com/jianzhelin/DuAN.
Jianzhe Lin, Lichao Mou, Tianze Yu, Xiao Xiang Zhu 0001, Z. Jane Wang 0001
ACM Multimedia4
2020 Learning-Shared Cross-Modality Representation Using Multispectral-LiDAR and Hyperspectral Data
abstract
Due to the ever-growing diversity of the data source, multimodality feature learning has attracted more and more attention. However, most of these methods are designed by jointly learning feature representation from multimodalities that exist in both training and test sets, yet they are less investigated in the absence of certain modality in the test phase. To this end, in this letter, we propose to learn a shared feature space across multimodalities in the training process. By this way, the out-of-sample from any of multimodalities can be directly projected onto the learned space for a more effective cross-modality representation. More significantly, the shared space is regarded as a latent subspace in our proposed method, which connects the original multimodal samples with label information to further improve the feature discrimination. Experiments are conducted on the multispectral-Light Detection and Ranging (LIDAR) and hyperspectral data set provided by the 2018 IEEE GRSS Data Fusion Contest to demonstrate the effectiveness and superiority of the proposed method in comparison with several popular baselines.
Danfeng Hong, Jocelyn Chanussot, Naoto Yokoya, Jian Kang 0005, Xiao Xiang Zhu 0001
IEEE Geosci. Remote. Sens. Lett.5
2020 Fusing Multiseasonal Sentinel-2 Imagery for Urban Land Cover Classification With Multibranch Residual Convolutional Neural Networks
abstract
Exploiting multitemporal Sentinel-2 images for urban land cover classification has become an important research topic, since these images have become globally available at relatively fine temporal resolution, thus offering great potential for large-scale land cover mapping. However, appropriate exploitation of the images needs to address problems such as cloud cover inherent to optical satellite imagery. To this end, we propose a simple yet effective decision-level fusion approach for urban land cover prediction from multiseasonal Sentinel-2 images, using the state-of-the-art residual convolutional neural networks (ResNet). We extensively tested the approach in a cross-validation manner over a seven-city study area in central Europe. Both quantitative and qualitative results demonstrated the superior performance of the proposed fusion approach over several baseline approaches, including observation- and feature-level fusion.
Chunping Qiu, Lichao Mou, Michael Schmitt 0003, Xiao Xiang Zhu 0001
IEEE Geosci. Remote. Sens. Lett.4
2020 Invariant Attribute Profiles: A Spatial-Frequency Joint Feature Extractor for Hyperspectral Image Classification
abstract
So far, a large number of advanced techniques have been developed to enhance and extract the spatially semantic information in hyperspectral image processing and analysis. However, locally semantic change, such as scene composition, relative position between objects, spectral variability caused by illumination, atmospheric effects, and material mixture, has been less frequently investigated in modeling spatial information. Consequently, identifying the same materials from spatially different scenes or positions can be difficult. In this article, we propose a solution to address this issue by locally extracting invariant features from hyperspectral imagery (HSI) in both spatial and frequency domains, using a method called invariant attribute profiles (IAPs). IAPs extract the spatial invariant features by exploiting isotropic filter banks or convolutional kernels on HSI and spatial aggregation techniques (e.g., superpixel segmentation) in the Cartesian coordinate system. Furthermore, they model invariant behaviors (e.g., shift, rotation) by the means of a continuous histogram of oriented gradients constructed in a Fourier polar coordinate. This yields a combinatorial representation of spatial-frequency invariant features with application to HSI classification. Extensive experiments conducted on three promising hyperspectral data sets (Houston2013 and Houston2018) to demonstrate the superiority and effectiveness of the proposed IAP method in comparison with several state-of-the-art profile-related techniques. The codes will be available from the website: https://sites.google.com/view/danfeng-hong/data-code.
Danfeng Hong, Xin Wu 0001, Pedram Ghamisi, Jocelyn Chanussot, Naoto Yokoya, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.6
2020 Relation Network for Multilabel Aerial Image Classification
abstract
Multilabel classification plays a momentous role in perceiving intricate contents of an aerial image and triggers several related studies over the last years. However, most of them deploy few efforts in exploiting label relations, while such dependencies are crucial for making accurate predictions. Although an long short term memory (LSTM) layer can be introduced to modeling such label dependencies in a chain propagation manner, the efficiency might be questioned when certain labels are improperly inferred. To address this, we propose a novel aerial image multilabel classification network, attention-aware label relational reasoning network. Particularly, our network consists of three elemental modules: 1) a label-wise feature parcel learning module; 2) an attentional region extraction module; and 3) a label relational inference module. To be more specific, the label-wise feature parcel learning module is designed for extracting high-level label-specific features. The attentional region extraction module aims at localizing discriminative regions in these features without region proposal generation, yielding attentional label-specific features. The label relational inference module finally predicts label existences using label relations reasoned from outputs of the previous module. The proposed network is characterized by its capacities of extracting discriminative label-wise features and reasoning about label relations naturally and interpretably. In our experiments, we evaluate the proposed model on two multilabel aerial image data sets, of which one is newly produced. Quantitative and qualitative results on these two data sets demonstrate the effectiveness of our model. To facilitate progress in the multilabel aerial image classification, our produced data set will be made publicly available.
Yuansheng Hua, Lichao Mou, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.3
2020 Multipass SAR Interferometry Based on Total Variation Regularized Robust Low Rank Tensor Decomposition
abstract
Multipass SAR interferometry (InSAR) techniques based on meter-resolution spaceborne SAR satellites, such as TerraSAR-X or COSMO-SkyMed, provide 3D reconstruction and the measurement of ground displacement over large urban areas. Conventional methods such as persistent scatterer interferometry (PSI) usually requires a fairly large SAR image stack (usually in the order of tens) to achieve reliable estimates of these parameters. Recently, low rank property in multipass InSAR data stack was explored and investigated in our previous work (J. Kang et al., “Object-based multipass InSAR via robust low-rank tensor decomposition,” IEEE Trans. Geosci. Remote Sens., vol. 56, no. 6, 2018). By exploiting this low rank prior, a more accurate estimation of the geophysical parameters can be achieved, which in turn can effectively reduce the number of interferograms required for a reliable estimation. Based on that, this article proposes a novel tensor decomposition method in a complex domain, which jointly exploits low rank and variational prior of the interferometric phase in InSAR data stacks. Specifically, a total variation (TV) regularized robust low rank tensor decomposition method is exploited for recovering outlier-free InSAR stacks. We demonstrate that the filtered InSAR data stacks can greatly improve the accuracy of geophysical parameters estimated from real data. Moreover, this article demonstrates for the first time in the community that tensor-decomposition-based methods can be beneficial for large-scale urban mapping problems using multipass InSAR. Two TerraSAR-X data stacks with large spatial areas demonstrate the promising performance of the proposed method.
Jian Kang 0005, Yuanyuan Wang 0002, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.3
2020 Building Footprint Generation by Integrating Convolution Neural Network With Feature Pairwise Conditional Random Field (FPCRF)
abstract
Building footprint maps are vital to many remote sensing (RS) applications, such as 3-D building modeling, urban planning, and disaster management. Due to the complexity of buildings, the accurate and reliable generation of the building footprint from RS imagery is still a challenging task. In this article, an end-to-end building footprint generation approach that integrates convolution neural network (CNN) and graph model is proposed. CNN serves as the feature extractor, while the graph model can take spatial correlation into consideration. Moreover, we propose to implement the feature pairwise conditional random field (FPCRF) as a graph model to preserve sharp boundaries and fine-grained segmentation. Experiments are conducted on four different data sets: 1) Planetscope satellite imagery of the cities of Munich, Paris, Rome, and Zurich; 2) ISPRS Benchmark data from the city of Potsdam; 3) Dstl Kaggle data set; and 4) Inria Aerial Image Labeling data of Austin, Chicago, Kitsap County, Western Tyrol, and Vienna. It is found that the proposed end-to-end building footprint generation framework with the FPCRF as the graph model can further improve the accuracy of building footprint generation by using only CNN, which is the current state of the art.
Qingyu Li 0001, Yilei Shi, Xin Huang 0002, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.4
2020 Relation Matters: Relational Context-Aware Fully Convolutional Network for Semantic Segmentation of High-Resolution Aerial Images
abstract
Most current semantic segmentation approaches fall back on deep convolutional neural networks (CNNs). However, their use of convolution operations with local receptive fields causes failures in modeling contextual spatial relations. Prior works have sought to address this issue by using graphical models or spatial propagation modules in networks. But such models often fail to capture long-range spatial relationships between entities, which leads to spatially fragmented predictions. Moreover, recent works have demonstrated that channel-wise information also acts a pivotal part in CNNs. In this article, we introduce two simple yet effective network units, the spatial relation module, and the channel relation module to learn and reason about global relationships between any two spatial positions or feature maps, and then produce Relation-Augmented (RA) feature representations. The spatial and channel relation modules are general and extensible, and can be used in a plug-and-play fashion with the existing fully convolutional network (FCN) framework. We evaluate relation module-equipped networks on semantic segmentation tasks using two aerial image data sets, namely International Society for Photogrammetry and Remote Sensing (ISPRS) Vaihingen and Potsdam data sets, which fundamentally depend on long-range spatial relational reasoning. The networks achieve very competitive results, a mean F1score of 88.54% on the Vaihingen data set and a mean F1score of 88.01% on the Potsdam data set, bringing significant improvements over baselines.
Lichao Mou, Yuansheng Hua, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.3
2020 Nonlocal Graph Convolutional Networks for Hyperspectral Image Classification
abstract
Over the past few years making use of deep networks, including convolutional neural networks (CNNs) and recurrent neural networks (RNNs), classifying hyperspectral images has progressed significantly and gained increasing attention. In spite of being successful, these networks need an adequate supply of labeled training instances for supervised learning, which, however, is quite costly to collect. On the other hand, unlabeled data can be accessed in almost arbitrary amounts. Hence it would be conceptually of great interest to explore networks that are able to exploit labeled and unlabeled data simultaneously for hyperspectral image classification. In this article, we propose a novel graph-based semisupervised network called nonlocal graph convolutional network (nonlocal GCN). Unlike existing CNNs and RNNs that receive pixels or patches of a hyperspectral image as inputs, this network takes the whole image (including both labeled and unlabeled data) in. More specifically, a nonlocal graph is first calculated. Given this graph representation, a couple of graph convolutional layers are used to extract features. Finally, the semisupervised learning of the network is done by using a cross-entropy error over all labeled instances. Note that the nonlocal GCN is end-to-end trainable. We demonstrate in extensive experiments that compared with state-of-the-art spectral classifiers and spectral-spatial classification networks, the nonlocal GCN is able to offer competitive results and high-quality classification maps (with fine boundaries and without noisy scattered points of misclassification).
Lichao Mou, Xiaoqiang Lu, Xuelong Li 0001, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.4
2020 Learning to Pay Attention on Spectral Domain: A Spectral Attention Module-Based Convolutional Network for Hyperspectral Image Classification
abstract
Over the past few years, hyperspectral image classification using convolutional neural networks (CNNs) has progressed significantly. In spite of their effectiveness, given that hyperspectral images are of high dimensionality, CNNs can be hindered by their modeling of all spectral bands with the same weight, as probably not all bands are equally informative and predictive. Moreover, the usage of useless spectral bands in CNNs may even introduce noises and weaken the performance of networks. For the sake of boosting the representational capacity of CNNs for spectral-spatial hyperspectral data classification, in this work, we improve networks by discriminating the significance of different spectral bands. We design a network unit, which is termed as the spectral attention module, that makes use of a gating mechanism to adaptively recalibrate spectral bands by selectively emphasizing informative bands and suppressing less useful ones. We theoretically analyze and discuss why such a spectral attention module helps in a CNN for hyperspectral image classification. We demonstrate using extensive experiments that in comparison with state-of-the-art approaches, the spectral attention module-based convolutional networks are able to offer competitive results. Furthermore, this work sheds light on how a CNN interacts with spectral bands for the purpose of classification.
Lichao Mou, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.2
2020 Unsupervised Deep Joint Segmentation of Multitemporal High-Resolution Images
abstract
High/very-high-resolution (HR/VHR) multitemporal images are important in remote sensing to monitor the dynamics of the Earth's surface. Unsupervised object-based image analysis provides an effective solution to analyze such images. Image semantic segmentation assigns pixel labels from meaningful object groups and has been extensively studied in the context of single-image analysis, however not explored for multitemporal one. In this article, we propose to extend supervised semantic segmentation to the unsupervised joint semantic segmentation of multitemporal images. We propose a novel method that processes multitemporal images by separately feeding to a deep network comprising of trainable convolutional layers. The training process does not involve any external label, and segmentation labels are obtained from the argmax classification of the final layer. A novel loss function is used to detect object segments from individual images as well as establish a correspondence between distinct multitemporal segments. Multitemporal semantic labels and weights of the trainable layers are jointly optimized in iterations. We tested the method on three different HR/VHR data sets from Munich, Paris, and Trento, which shows the method to be effective. We further extended the proposed joint segmentation method for change detection (CD) and tested on a VHR multisensor data set from Trento.
Sudipan Saha, Lichao Mou, Chunping Qiu, Xiao Xiang Zhu 0001, Francesca Bovolo, Lorenzo Bruzzone
IEEE Trans. Geosci. Remote. Sens.4
2020 SAR Tomography at the Limit: Building Height Reconstruction Using Only 3-5 TanDEM-X Bistatic Interferograms
abstract
Multibaseline interferometric synthetic aperture radar (InSAR) techniques are effective approaches for retrieving the 3-D information of urban areas. In order to obtain a plausible reconstruction, it is necessary to use more than 20 interferograms. Hence, these methods are commonly not appropriate for large-scale 3-D urban mapping using TanDEM-X data, where only a few acquisitions are available in average for each city. This article proposes a new SAR tomographic processing framework to work with those extremely small stacks, which integrates the nonlocal filtering into SAR tomography inversion. The applicability of the algorithm is demonstrated using a TanDEM-X multibaseline stack with five bistatic interferograms over the whole city of Munich, Germany. A systematic comparison of our result with TanDEM-X raw digital elevation models (DEMs) and airborne LiDAR data shows that the relative height accuracy of two-third buildings is within 2 m, which outperforms the TanDEM-X raw DEM. The promising performance of the proposed algorithm paved the first step toward high-quality large-scale 3-D urban mapping.
Yilei Shi, Richard Bamler, Yuanyuan Wang 0002, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.4
2019 A Relation-Augmented Fully Convolutional Network for Semantic Segmentation in Aerial Scenes
abstract
Most current semantic segmentation approaches fall back on deep convolutional neural networks (CNNs). However, their use of convolution operations with local receptive fields causes failures in modeling contextual spatial relations. Prior works have sought to address this issue by using graphical models or spatial propagation modules in networks. But such models often fail to capture long-range spatial relationships between entities, which leads to spatially fragmented predictions. Moreover, recent works have demonstrated that channel-wise information also acts a pivotal part in CNNs. In this work, we introduce two simple yet effective network units, the spatial relation module and the channel relation module, to learn and reason about global relationships between any two spatial positions or feature maps, and then produce relation-augmented feature representations. The spatial and channel relation modules are general and extensible, and can be used in a plug-and-play fashion with the existing fully convolutional network (FCN) framework. We evaluate relation module-equipped networks on semantic segmentation tasks using two aerial image datasets, which fundamentally depend on long-range spatial relational reasoning. The networks achieve very competitive results, bringing significant improvements over baselines.
Lichao Mou, Yuansheng Hua, Xiao Xiang Zhu 0001
CVPR3
2019 The Challenge of Creating The Sarptical Dataset
abstract
The SARptical dataset1consists of about 10,000 matching pairs of very high resolution optical and SAR image patches extracted from TerraSAR-X very high-resolution spotlight images and aerial UltraCAM optical images in dense urban areas of the city of Berlin, Germany. This dataset is distinct from any other existing SAR optical dataset, because the 3D location of the center pixels of the SAR and optical were explicitly matched. Still, creating such dataset poses a fundamental challenge. The reason is that a pixel level matching between SAR and optical images is generally impossible to achieve a without the assistance of a precise 3-D model. This is especially true in dense urban areas, because of the inevitable layover caused by the side-looking SAR imaging geometry. Such misalignment will affect applications like joint classification using SAR and optical images, and raise the difficulty in designing machine learning algorithms. This paper will discuss the challenge of jointly using SAR and optical images for remote sensing applications and propose possible methods to mitigate those misalignment errors.
Yuanyuan Wang 0002, Xiao Xiang Zhu 0001
IGARSS2
2019 Building Type Classification from Social Media Texts via Geo-Spatial Textmining
abstract
In this work, we present a model for building type classification from Twitter text messages (tweets) by employing geo-spatial textmining methods. First, we apply standard text pre-processing methods and convert the tweets into sentence vectors using fastText. For classification, we apply a feedforward network with two fully connected hidden layers and feed the generated sentence vectors as linguistic features. Classification results suggest that the classes are distinguishable to a certain extent with pure text even with unbalanced class distributions and a very small sample size. However, these findings also undermine, that building type classification with pure text data is a challenging task.
Matthias Häberle, Martin Werner 0001, Xiao Xiang Zhu 0001
IGARSS3
2019 Uavsar Tomography of Munich
abstract
In May-June 2015 UAVSAR was flown to Europe to collect data in support of experiments in Iceland, Norway and Ger-many. The deployment in Germany was focused on PolIn-SAR and tomographic data collections at the Traunstein Forest and in the Munich urban area. In this paper we describe tomographic processing of the Munich data and comparison with in situ ground truth data.
Scott Hensley, Brian P. Hawkins, Thierry Michel, Ronald Muellerschoen, Xiao Xiang Zhu 0001, Andreas Reigber, Gustavo D. Martín del Campo-Becerra
IGARSS5
2019 Mutual Information Analysis of Social Media Images and Building Functions
abstract
Understanding urban dynamics requires detailed insights into urban land use. On the most fine-grained level this classification is done on single building instance levels. This level of detail can hardly be solved using remote sensing only, but requires complementary data. Social media images are a promising additional image data source since they are captured on a global scale in vast volumes.In this study we investigate the relation between objects showing up in geotagged social media images and functions of buildings proximate to the image location. We propose a rasterization approach to embed features from images and labels from a target domain to calculate mutual information both domains share. In our study area of Los Angeles, USA, we show that using object detection is a valuable way of extracting features from social media images to predict building functions. Furthermore, we present the most significant object types for five types of buildings.
Eike Jens Hoffmann, Martin Werner 0001, Xiao Xiang Zhu 0001
IGARSS3
2019 WU-Net: A Weakly-Supervised Unmixing Network for Remotely Sensed Hyperspectral Imagery
abstract
Recently, enormous efforts have been made to improve the performance of the linear or nonlinear mixing model for hyperspectral unmixing, yet their ability to handle spectral variability and extract physically meaningful endmembers remains limited. Based on the powerful learning ability of deep learning, we propose a weakly-supervised unmixing network, called WU-Net, to break the bottleneck. Beyond the autoencoder-like architecture, WU-Net learns an additional network from the pure or nearly-pure endmembers to correct the weights of another unmixing network towards a more accurate and interpretable unmixing solution, thus yielding a two-stream deep network. Experimental results conducted on two different datasets, one fully artificial simulation dataset and one simulated EnMap dataset generated from a real HyMap dataset, demonstrate the effectiveness and superiority of WU-Net over several state-of-the-art algorithms.
Danfeng Hong, Jocelyn Chanussot, Naoto Yokoya, Uta Heiden, Wieke Heldens, Xiao Xiang Zhu 0001
IGARSS6
2019 A Topological Data Analysis Guided Fusion Algorithm: Mapper-Regularized Manifold Alignment
abstract
Hyperspectral images and polarimetric synthetic aperture radar (PolSAR) data are two important data sources, yet they barely appear under the same scope, even though multi-modal data fusion is attracting more and more attention. To our best knowledge, this paper investigates for the first time semi-supervised manifold alignment (SSMA) for the fusion of the hyperspectral image and PolSAR data. The SSMA searches a latent space where different data sources are aligned, which is accomplished by using the label information and the topological structure of the data. This paper is the first attempt to apply topological data analysis (TDA), a recent mathematic sub-field of data analysis, in remote sensing. It aims to reveal relevant information from the shape of a data in its feature space, and has been proven powerful in medicine. The paper also proposes a novel algorithm, MAPPER-regularized manifold alignment, which embeds the TDA into a semi-supervised manifold alignment for the fusion of the hyper-spectral image and PolSAR data. The proposed algorithm exhibits superior performance in fusing a simulated EnMAP data set and a Sentinel-1 data set for an image of Berlin.
Jingliang Hu, Danfeng Hong, Yuanyuan Wang 0002, Xiao Xiang Zhu 0001
IGARSS4
2019 Label Relation Inference for Multi-Label Aerial Image Classification
abstract
Multi-label aerial image classification is a challenging visual task and obtaining increasing attention recently. Most of the existing methods resort to training independent classifier for each label, while underlying label correlations are not fully exploited while making predictions. To this end, we propose an innovative inference network, which takes advantage of pairwise label relations to infer multiple object labels of a high-resolution aerial image. Specifically, we first employ a feature extraction module to extract high-level feature representations of an aerial image, and then, feed them into a relational inference module to predict the presence of each object label. We evaluate our network on the UCM multilabel dataset and experiment with various popular convolutional neural networks (CNNs) as the backbone of the feature extraction module. Experimental results demonstrate that the proposed network behaves superiorly in comparison with other existing methods.
Yuansheng Hua, Lichao Mou, Xiao Xiang Zhu 0001
IGARSS3
2019 Mitigation of Positioning Bias in PSI Point Clouds
abstract
In this paper, we propose a framework for geocoding error correction of point clouds obtained from Persistent Scatterer Interferometry (PSI) in urban areas. The horizontal positioning bias of Persistent Scatterers (PS) is mitigated by applying geodetic and atmospheric corrections to the PS azimuth and range timings. Furthermore, the vertical positioning bias, caused due to the unknown height of the PSI reference point, is estimated and compensated for by the use of Synthetic Aperture Radar (SAR)-based Ground Control Point (GCP)s. Experimental results based on cross-heading pairs of TerraSAR-X (TS-X) high resolution spotlight images over the city of Berlin, Germany and in Oulu, Finland are presented. For selected test sites, the localization accuracy of the point clouds is analyzed with respect to a reference LiDAR data, which demonstrates the applicability of the proposed correction approach.
Sina Montazeri, Fernando Rodríguez González, Xiao Xiang Zhu 0001
IGARSS3
2019 Spatial Relational Reasoning in Networks for Improving Semantic Segmentation of Aerial Images
abstract
Most current semantic segmentation approaches rely on deep convolutional neural networks (CNNs). However, their use of convolution operations with local receptive fields causes failures in modeling contextual spatial relations. Prior works have tried to address this issue by using graphical models or spatial propagation modules in networks. But such models often fail to capture long-range spatial relationships between entities, which leads to spatially fragmented predictions. In this work, we introduce a simple yet effective network unit, the spatial relation module, to learn and reason about global relationships between any two spatial positions, and then produce relation-enhanced feature representations. The spatial relation module is general and extensible, and can be used in a plug-and-play fashion with the existing fully convolutional network (FCN) framework. We evaluate spatial relation module-equipped networks on semantic segmentation tasks using two aerial image datasets. The networks achieve very competitive results, bringing significant improvements over baselines.
Lichao Mou, Yuansheng Hua, Xiao Xiang Zhu 0001
IGARSS3
2019 Fusing Multi-Seasonal Sentinel-2 Images with Residual Convolutional Neural Networks for Local Climate Zone-Derived Urban Land Cover Classification
abstract
This paper proposes a framework to fuse multi-seasonal Sentinel-2 images, with application on LCZ-derived urban land cover classification. Cross-validation over a seven-city study area in central Europe demonstrates its consistently better performance over several previous approaches, with the same experimental setup. Based on our previous work, we can conclude that decision-level fusion is better than feature-level fusion for similar tasks at similar scale with multi-seasonal Sentinel-2 images. With the framework, urban land cover maps of several cities are produced. The visualization of two exemplary areas shows urban structures that are consistent with existing datasets. This framework can be also generally beneficial for other types of urban mapping.
Chunping Qiu, Michael Schmitt 0003, Xiao Xiang Zhu 0001
IGARSS3
2019 Non-Local SAR Tomography for Large-Scale Urban Mapping
abstract
Multi-baseline synthetic aperture radar (SAR) interferometric techniques, such as SAR tomography, is well established for 3-D reconstruction in the urban area. These methods usually require fairly large interferometric stacks (> 20 images) for a reliable reconstruction. Hence, they are usually not directly applicable for large-scale 3-D urban mapping using TanDEM-X data where only a few acquisitions are available in average for each city. This work proposes a new SAR tomographic processing framework to those extremely small stacks. The applicability of the algorithm is demonstrated using a TanDEM-X multi-baseline stack with five bistatic interferograms over the whole city of Munich, Germany. Systematic comparison of our result with TanDEM-X raw digital elevation models (DEM) and airborne LiDAR data shows that the relative height accuracy is two meters, which outperforms the TanDEM-X raw DEM. The promising performance of the proposed algorithm paved the first step towards high quality large-scale 3-D urban mapping.
Yilei Shi, Yuanyuan Wang 0002, Xiao Xiang Zhu 0001, Richard Bamler
IGARSS3
2019 Building Footprint Extraction with Graph Convolutional Network
abstract
Building footprint information is an essential ingredient for 3-D reconstruction of urban models. The automatic generation of building footprints from satellite images presents a considerable challenge due to the complexity of building shapes. Recent developments in deep convolutional neural networks (DCNNs) have enabled accurate pixel-level labeling tasks. One central issue remains, which is the precise delineation of boundaries. Deep architectures generally fail to produce fine-grained segmentation with accurate boundaries due to progressive downsampling. In this work, we have proposed a end-to-end framework to overcome this issue, which uses the graph convolutional network (GCN) for building footprint extraction task. Our proposed framework outperforms state-of-the-art methods.
Yilei Shi, Qingyu Li 0001, Xiao Xiang Zhu 0001
IGARSS3
2019 Automatic Registration of SAR Image and GIS Building Footprints Data in Dense Urban Area
abstract
In this paper, we propose a framework for the automatic registration of GIS building footprint polygons to a corresponding SAR image through the corresponding features of building walls in the two data. To extract feature lines, the Potts model is adopted for SAR image segmentation, and visibility test is performed on both data. The feature lines are then sampled to two point sets, and are registered using Iterative Closest Point (ICP) algorithm. The test result shows a registration accuracy of 0.67 m in azimuth direction, and 1.64m in range direction.
Yao Sun 0005, Yuanyuan Wang 0002, Xiao Xiang Zhu 0001
IGARSS3
2019 GeoSay: A geometric saliency for extracting buildings in remote sensing images
Gui-Song Xia, Nan Xue 0001, Qikai Lu, Xiao Xiang Zhu 0001
Comput. Vis. Image Underst.5
2019 Building Footprint Generation Using Improved Generative Adversarial Networks
abstract
Building footprint information is an essential ingredient for 3-D reconstruction of urban models. The automatic generation of building footprints from satellite images presents a considerable challenge due to the complexity of building shapes. In this letter, we have proposed improved generative adversarial networks (GANs) for the automatic generation of building footprints from satellite images. We used a conditional GAN (CGAN) with a cost function derived from the Wasserstein distance and added a gradient penalty term. The achieved results indicated that the proposed method can significantly improve the quality of building footprint generation compared to CGANs, the U-Net, and other networks. In addition, our method nearly removes all hyperparameters tuning.
Yilei Shi, Qingyu Li 0001, Xiao Xiang Zhu 0001
IEEE Geosci. Remote. Sens. Lett.3
2019 Bistatic-Like Differential SAR Tomography
abstract
Motivated by prospective synthetic aperture radar (SAR) satellite missions, this paper addresses the problem of differential SAR tomography (D-TomoSAR) in urban areas using spaceborne bistatic or pursuit monostatic acquisitions. A bistatic or pursuit monostatic interferogram is not subject to significant temporal decorrelation or atmospheric phase screen and, therefore, ideal for elevation reconstruction. We propose a framework that incorporates this reconstructed elevation as deterministic prior to deformation estimation, which uses conventional repeat-pass interferograms generated from bistatic or pursuit monostatic pairs. By means of theoretical and empirical analyses, we show that this framework is, in the pursuit monostatic case, both statistically and computationally more efficient than the standard D-TomoSAR. In the bistatic case, its theoretical bound is no worse by a factor of 2. We also show that reasonable results can be obtained by using merely six TerraSAR-X add-on for digital elevation measurements (TanDEM-X) pursuit monostatic pairs, if additional spatial prior is introduced. The proposed framework can be easily extended for multistatic configurations or external sources of scatterer's elevation.
Nan Ge, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.2
2019 CoSpace: Common Subspace Learning From Hyperspectral-Multispectral Correspondences
abstract
With a large amount of open satellite multispectral (MS) imagery (e.g., Sentinel-2 and Landsat-8), considerable attention has been paid to global MS land cover classification. However, its limited spectral information hinders further improving the classification performance. Hyperspectral imaging enables discrimination between spectrally similar classes but its swath width from space is narrow compared to MS ones. To achieve accurate land cover classification over a large coverage, we propose a cross-modality feature learning framework, called common subspace learning (CoSpace), by jointly considering subspace learning and supervised classification. By locally aligning the manifold structure of the two modalities, CoSpace linearly learns a shared latent subspace from hyperspectral-MS (HS-MS) correspondences. The MS out-of-samples can be then projected into the subspace, which are expected to take advantages of rich spectral information of the corresponding hyperspectral data used for learning, and thus leads to a better classification. Extensive experiments on two simulated HS-MS data sets (University of Houston and Chikusei), where HS-MS data sets have tradeoffs between coverage and spectral resolution, are performed to demonstrate the superiority and effectiveness of the proposed method in comparison with previous state-of-the-art methods.
Danfeng Hong, Naoto Yokoya, Jocelyn Chanussot, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.4
2019 MIMA: MAPPER-Induced Manifold Alignment for Semi-Supervised Fusion of Optical Image and Polarimetric SAR Data
abstract
Multi-modal data fusion has recently been shown promise in classification tasks in remote sensing. Optical data and radar data, two important yet intrinsically different data sources, are attracting more and more attention for potential data fusion. It is already widely known that a machine learning-based methodology often yields excellent performance. However, the methodology relies on a large training set, which is very expensive to achieve in remote sensing. The semi-supervised manifold alignment (SSMA), a multi-modal data fusion algorithm, has been designed to amplify the impact of an existing training set by linking labeled data to unlabeled data via unsupervised techniques. In this paper, we explore the potential of SSMA in fusing optical data and polarimetric synthetic aperture radar (SAR) data, which are multi-sensory data sources. Furthermore, we propose a MAPPER-induced manifold alignment (MIMA) for the semi-supervised fusion of multi-sensory data sources. Our proposed method unites SSMA with MAPPER, which is developed from the emerging topological data analysis (TDA) field. To the best of our knowledge, this is the first time that SSMA has been applied on fusing optical data and SAR data, and also the first time that TDA has been applied in remote sensing. The conventional SSMA derives a topological structure using k-nearest neighbor (kNN), while MIMA employs MAPPER, which considers the field knowledge and derives a novel topological structure through the spectral clustering in a data-driven fashion. The experimental results on data fusion with respect to land cover land use classification and local climate zone classification suggest superior performance of MIMA.
Jingliang Hu, Danfeng Hong, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.3
2019 R3-Net: A Deep Network for Multioriented Vehicle Detection in Aerial Images and Videos
abstract
Vehicle detection is a significant and challenging task in aerial remote sensing applications. Most existing methods detect vehicles with regular rectangle boxes and fail to offer the orientation of vehicles. However, the orientation information is crucial for several practical applications, such as the trajectory and motion estimation of vehicles. In this paper, we propose a novel deep network, called a rotatable region-based residual network (R3-Net), to detect multioriented vehicles in aerial images and videos. More specially, R3-Net is utilized to generate rotatable rectangular target boxes in a half coordinate system. First, we use a rotatable region proposal network (R-RPN) to generate rotatable region of interests (R-RoIs) from feature maps produced by a deep convolutional neural network. Here, a proposed batch averaging rotatable anchor strategy is applied to initialize the shape of vehicle candidates. Next, we propose a rotatable detection network (R-DN) for the final classification and regression of the R-RoIs. In R-DN, a novel rotatable position-sensitive pooling is designed to keep the position and orientation information simultaneously while downsampling the feature maps of R-RoIs. In our model, R-RPN and R-DN can be trained jointly. We test our network on two open vehicle detection image data sets, namely, DLR 3K Munich Data set and VEDAI Data set, demonstrating the high precision and robustness of our method. In addition, further experiments on aerial videos show the good generalization capability of the proposed method and its potential for vehicle tracking in aerial videos. The demo video is available athttps://youtu.be/xCYD-tYudN0.
Qingpeng Li, Lichao Mou, Qizhi Xu, Yun Zhang 0014, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.5
2019 Learning Spectral-Spatial-Temporal Features via a Recurrent Convolutional Neural Network for Change Detection in Multispectral Imagery
abstract
Change detection is one of the central problems in earth observation and was extensively investigated over recent decades. In this paper, we propose a novel recurrent convolutional neural network (ReCNN) architecture, which is trained to learn a joint spectral-spatial-temporal feature representation in a unified framework for change detection in multispectral images. To this end, we bring together a convolutional neural network and a recurrent neural network into one end-to-end network. The former is able to generate rich spectral-spatial feature representations, while the latter effectively analyzes temporal dependence in bitemporal images. In comparison with previous approaches to change detection, the proposed network architecture possesses three distinctive properties: 1) it is end-to-end trainable, in contrast to most existing methods whose components are separately trained or computed; 2) it naturally harnesses spatial information that has been proven to be beneficial to change detection task; and 3) it is capable of adaptively learning the temporal dependence between multitemporal images, unlike most of the algorithms that use fairly simple operation like image differencing or stacking. As far as we know, this is the first time that a recurrent convolutional network architecture has been proposed for multitemporal remote sensing image analysis. The proposed network is validated on real multispectral data sets. Both visual and quantitative analyses of the experimental results demonstrate competitive performance in the proposed mode.
Lichao Mou, Lorenzo Bruzzone, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.3
2019 Buildings Detection in VHR SAR Images Using Fully Convolution Neural Networks
abstract
This paper addresses the highly challenging problem of automatically detecting man-made structures especially buildings in very high-resolution (VHR) synthetic aperture radar (SAR) images. In this context, this paper has two major contributions. First, it presents a novel and generic workflow that initially classifies the spaceborne SAR tomography (TomoSAR) point clouds-generated by processing VHR SAR image stacks using advanced interferometric techniques known as TomoSAR-into buildings and nonbuildings with the aid of auxiliary information (i.e., either using openly available 2-D building footprints or adopting an optical image classification scheme) and later back project the extracted building points onto the SAR imaging coordinates to produce automatic large-scale benchmark labeled (buildings/nonbuildings) SAR data sets. Second, these labeled data sets (i.e., building masks) have been utilized to construct and train the state-of-the-art deep fully convolution neural networks with an additional conditional random field represented as a recurrent neural network to detect building regions in a single VHR SAR image. Such a cascaded formation has been successfully employed in computer vision and remote sensing fields for optical image classification but, to our knowledge, has not been applied to SAR images. The results of the building detection are illustrated and validated over a TerraSAR-X VHR spotlight SAR image covering approximately 39 km2-almost the whole city of Berlin- with the mean pixel accuracies of around 93.84%.
Muhammad Shahzad 0002, Michael Maurer, Friedrich Fraundorfer, Yuanyuan Wang 0002, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.5
2019 Nonlocal Compressive Sensing-Based SAR Tomography
abstract
Tomographic synthetic aperture radar (TomoSAR) inversion of urban areas is an inherently sparse reconstruction problem and, hence, can be solved using compressive sensing (CS) algorithms. This paper proposes solutions for two notorious problems in this field. First, TomoSAR requires a high number of data sets, which makes the technique expensive. However, it can be shown that the number of acquisitions and the signal-to-noise ratio (SNR) can be traded off against each other, because it is asymptotically only the product of the number of acquisitions and SNR that determines the reconstruction quality. We propose to increase SNR by integrating nonlocal (NL) estimation into the inversion and show that a reasonable reconstruction of buildings from only seven interferograms is feasible. Second, CS-based inversion is computationally expensive and therefore, barely suitable for large-scale applications. We introduce a new fast and accurate algorithm for solving the NL L1-L2-minimization problem, central to CS-based reconstruction algorithms. The applicability of the algorithm is demonstrated using simulated data and TerraSAR-X high-resolution spotlight images over an area in Munich, Germany.
Yilei Shi, Xiao Xiang Zhu 0001, Richard Bamler
IEEE Trans. Geosci. Remote. Sens.2
2019 Fusion of Heterogeneous Earth Observation Data for the Classification of Local Climate Zones
abstract
This paper proposes a novel framework for fusing multi-temporal, multispectral satellite images and OpenStreetMap (OSM) data for the classification of local climate zones (LCZs). Feature stacking is the most commonly used method of data fusion but does not consider the heterogeneity of multimodal optical images and OSM data, which becomes its main drawback. The proposed framework processes two data sources separately and then combines them at the model level through two fusion models (the landuse fusion model and building fusion model) that aim to fuse optical images with landuse and buildings layers of OSM data, respectively. In addition, a new approach to detecting building incompleteness of OSM data is proposed. The proposed framework was trained and tested using the data from the 2017 IEEE GRSS Data Fusion Contest and further validated on one additional test (AT) set containing test samples that are manually labeled in Munich and New York. The experimental results have indicated that compared with the feature stacking-based baseline framework, the proposed framework is effective in fusing optical images with OSM data for the classification of LCZs with high generalization capability on a large scale. The classification accuracy of the proposed framework outperforms the baseline framework by more than 6% and 2% while testing on the test set of 2017 IEEE GRSS Data Fusion Contest and the AT set, respectively. In addition, the proposed framework is less sensitive to spectral diversities of optical satellite images and thus achieves more stable classification performance than the state-of-the-art frameworks.
Guichen Zhang, Pedram Ghamisi, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.3
2019 An Augmented Linear Mixing Model to Address Spectral Variability for Hyperspectral Unmixing
abstract
Hyperspectral imagery collected from airborne or satellite sources inevitably suffers from spectral variability, making it difficult for spectral unmixing to accurately estimate abundance maps. The classical unmixing model, the linear mixing model (LMM), generally fails to handle this sticky issue effectively. To this end, we propose a novel spectral mixture model, called the augmented linear mixing model (ALMM), to address spectral variability by applying a data-driven learning strategy in inverse problems of hyperspectral unmixing. The proposed approach models the main spectral variability (i.e., scaling factors) generated by variations in illumination or typography separately by means of the endmember dictionary. It then models other spectral variabilities caused by environmental conditions (e.g., local temperature and humidity, atmospheric effects) and instrumental configurations (e.g., sensor noise), as well as material nonlinear mixing effects, by introducing a spectral variability dictionary. To effectively run the data-driven learning strategy, we also propose a reasonable prior knowledge for the spectral variability dictionary, whose atoms are assumed to be low-coherent with spectral signatures of endmembers, which leads to a well-known low-coherence dictionary learning problem. Thus, a dictionary learning technique is embedded in the framework of spectral unmixing so that the algorithm can learn the spectral variability dictionary and estimate the abundance maps simultaneously. Extensive experiments on synthetic and real datasets are performed to demonstrate the superiority and effectiveness of the proposed method in comparison with previous state-of-the-art methods.
Danfeng Hong, Naoto Yokoya, Jocelyn Chanussot, Xiao Xiang Zhu 0001
IEEE Trans. Image Process.4
2018 Joint and Progressive Learning from High-Dimensional Data for Multi-label Classification
Danfeng Hong, Naoto Yokoya, Jian Xu 0008, Xiao Xiang Zhu 0001
ECCV (8)4
2018 Fusing Spaceborne SAR Interferometry and Street View Images for 4D Urban Modeling
abstract
Obtaining city models in a large scale is usually achieved by means of remote sensing techniques, such as synthetic aperture radar (SAR) interferometry and optical image stereogrammetry. Despite the controlled quality of these products, such observation is restricted by the characteristics of their sensor platform, such as revisit time and spatial resolution. Over the last decade, the rapid development of online geographic information systems, such as Google map, has accumulated vast amount of online images. Despite their uncontrolled quality, these images constitute a set of redundant spatial-temporal observations of our dynamic 3D urban environment. These images contain useful information that can complement the remote sensing data, especially the SAR images. This paper presents a one of the first studies of fusing online street view images and spaceborne SAR images, for the reconstruction of spatial-temporal (hence 4D) city models. We describe a general approach to geometrically combine the information of these two types of images that are nearly impossible to even coregister without a precise 3D city model due to their distinct imaging geometry. It is demonstrated that, one can obtain a new kind of city model that includes high resolution optical texture for better scene understanding and the dynamics of individual buildings up to the precision of millimeter retrieved from SAR interferometry.
Yuanyuan Wang 0002, Jian Kang 0005, Xiao Xiang Zhu 0001
FUSION3
2018 Multi-Pass SAR Interferometry for 3D Reconstruction of Complex Mountainous Areas Based on Robust Low Rank Tensor Decomposition
abstract
During the past decades, multi-pass SAR interferometry (In-SAR) techniques have been developed for retrieving geophysical parameters such as elevation, over large areas. Conventional method such as periodogram usually requires a fairly large SAR image stack (usually in the order of tens), in order to achieve reliable estimates of these parameters. However, when it comes to large-area processing, it is time-consuming and luxury to obtain a sufficient number of SAR images for the reconstruction. In this paper, we demonstrate a novel multi-pass InSAR method for 3D reconstruction using low rank tensor decomposition. By exploiting the low rank prior knowledge in the multi-pass InSAR stack, simulations show that the proposed method can improve the accuracy of elevation estimates by a factor of two, compared to the state-of-the-art InSAR filtering methods, such as SqueeSAR. The capability of the proposed algorithm is also demonstrated on real data using one TanDEM-X InSAR stack of a complex mountainous area.
Jian Kang 0005, Yuanyuan Wang 0002, Xiao Xiang Zhu 0001
IGARSS3
2018 Urban TanDEM-X Raw DEM Fusion Based ON TV-L1 and Huber Models
abstract
Recently, the TanDEM-X DEM has been produced as a global DEM with unprecedented relative accuracy. One important step of the chain of global DEM generation is to mosaic multiple raw DEM tiles by DEM fusion methods to reach the best possible target accuracy. Currently, Weighted Averaging (WA) is used as a fast and simple method for TanDEM-X raw DEM fusion in which the weights are computed from height error maps delivered from the Interferometric TanDEM-X Processor (ITP). In this paper, we investigate the efficiency of variational models such as TV-L1 and Huber model for the TanDEM-X raw DEM fusion task in comparison to WA. The results illustrate that using variational models can improve the quality of DEM fusion outputs especially for areas with high-frequency contents and more complex morphological features like urban areas. Using variational models could improve the DEM quality by up to about 1m.
Hossein Bagheri, Michael Schmitt 0003, Xiao Xiang Zhu 0001
IGARSS3
2018 A Conditional Generative Adversarial Network to Fuse Sar And Multispectral Optical Data For Cloud Removal From Sentinel-2 Images
abstract
In this paper, we present the first conditional generative adversarial network (cGAN) architecture that is specifically designed to fuse synthetic aperture radar (SAR) and optical multi-spectral (MS) image data to generate cloud- and haze-free MS optical data from a cloud-corrupted MS input and an auxiliary SAR image. Experiments on Sentinel-2 MS and Sentinel-l SAR data confirm that our extended SAR-Opt-cGAN model utilizes the auxiliary SAR information to better reconstruct MS images than an equivalent model which uses the same architecture but only single-sensor MS data as input.
Claas Grohnfeldt, Michael Schmitt 0003, Xiao Xiang Zhu 0001
IGARSS3
2018 Exploring Sentinel-L Data for Local Climate Zone Classification
abstract
Local climate zone (LCZ) is a categorical scheme describing the morphology of urban area, which is a valuable not only for the original purpose of temperature study, but also for other urban oriented studies like population density estimation and economical development monitoring. Standard LCZ production works only on individual cities using merely optical data, mostly LandSat-8 data. Our goal is to develop a framework that 1) can potentially work on a large number of cities, i.e., training on number of cities and testing on number of other cities; 2) exploits Synthetic Aperture Radar (SAR) data. In this paper, we investigated the potential of Sentinel-l Dual-Pol data on producing LCZ maps in general. It shows the Sentinel-1 data could improve the classification accuracy of several LCZ classes. Joint use of LandSat-8 data, Open Street Map (OSM) data and Sentinel-1 data provide 62.05% overall accuracy, which is higher than 51.20% achieved by using only LandSat-8 and OSM data.
Jingliang Hu, Xiao Xiang Zhu 0001
IGARSS2
2018 LAHNet: A Convolutional Neural Network Fusing Low- and High-Level Features for Aerial Scene Classification
abstract
In this paper, we proposed an innovative end-to-end convolutional neural network (CNN), which is trained to learn how to fuse multi-level features for aerial scene classification. Instead of using only coarse semantic features as conventional CNNs, we resort to first hierarchically extracting dense high-level features and then element-wise fusing them with low-level features to build a comprehensive feature representation, which contains not only high-level semantic information but also fine-grained low-level details, for scene classification. The network is evaluated on two broadly used aerial scene datasets, UCM and AID. The experimental results indicate that the proposed LAHNet performs superiorly compared to the existing benchmark methods. Furthermore, visualization of the fused features presents an intuitive illustration of the remarkable improvement.
Yuansheng Hua, Lichao Mou, Xiao Xiang Zhu 0001
IGARSS3
2018 Classification of Settlement Types from Tweets Using LDA and LSTM
abstract
Land use reflects the interrelation between the physically built environment and the activity patterns of people. It is indispensable information for decision-makes, but up-to-date and accurate land use information is often absent. Unlike approaches that make use of remote sensing data, in this work, we are interested in a novel data source, tweets, and explore its potential for land use classification in urban areas. Specifically, we propose a general framework for classifying settlement land-use types by extracting location, time, quantity and text features of twitter data. To do so, we apply latent Dirichlet allocation (LDA) and long short-term memory (LSTM) and then combines those features with spatial-temporal feature using Fused SVM and a two-stream convolutional neural network (CNN) for classification. For the case of classifying individual tweets by the land-use classes relevant in this study - residential, non-residential and mixed usage -, we reach overall accuracy (OA), average accuracy (AA), and Kappa coefficient with 72.35%, 73.76%, and 58.43%, respectively. As for the case of classifying block settlement types, we reach 61.90%, 63.33%, and 42.84%, respectively.
Hannes Taubenböck, Lichao Mou, Xiao Xiang Zhu 0001
IGARSS4
2018 Generative Adversarial Networks for Hard Negative Mining in CNN-Based SAR-Optical Image Matching
abstract
In this paper we propose a deep generative framework, based on a generative adversarial network (GAN) and an auto encoder (AE), for generating non-corresponding SAR patches to be used in hard negative mining in situations of limited data quantities. We evaluate the effectiveness of this formulation of hard negative mining for reducing the false positive rate (FPR) and improving network determinability in a SAR-optical patch matching application. Our generative network is trained to generate realistic SAR images using an existing SAR-optical matching dataset. These generated images are then used as non-corresponding, hard negative samples for training a SAR-optical matching network. Our results show that we are able to generate realistic SAR images which exhibit many SAR-like features, such as layover and speckle. We further show that by fine tuning the original matching network using these hard negative samples we are able to improve the overall performance of the original SAR-optical matching network.
Lloyd H. Hughes, Michael Schmitt 0003, Xiao Xiang Zhu 0001
IGARSS3
2018 Hierarchical Region Based Convolution Neural Network for Multiscale Object Detection in Remote Sensing Images
abstract
In this paper, we propose a novel Faster R-CNN based method to detect multiscale objects in very high resolution optical remote sensing images. Firstly, a pre-trained CNN is used to extract features from an input image; and then a set of object candidates are generated. To efficiently detect objects with various scales, we design a hierarchical selective filtering (HSF) layer to map features in different scales to the same scale space. The HSF layer can be applied on both region proposal and the subsequent detection network. More importantly, it can be plugged into Faster R-CNN network without modifying its architecture, meanwhile boosting the performance on detecting objects with varying scales. The proposed model can be trained in an end-to-end manner. We test our network on three datasets containing different multiscale objects, including airplanes, ships and buildings, which are collected from Google Earth images and GaoFen-2 images. Experiments demonstrate high precision and robustness of our method.
Qingpeng Li, Lichao Mou, Kaiyu Jiang, Qingjie Liu 0001, Yunhong Wang 0001, Xiao Xiang Zhu 0001
IGARSS6
2018 A Recurrent Convolutional Neural Network for Land Cover Change Detection in Multispectral Images
abstract
In this paper, we propose a novel network architecture, a recurrent convolutional neural network, which is trained to learn a joint spectral-spatial-temporal feature representation in a unified framework for change detection of multispectral images. To this end, we bring together a convolutional neural network (CNN) and a recurrent neural network (RNN) into one end-to-end network. The former is able to generate rich spectral-spatial feature representations while the latter effectively analyzes temporal dependency in bi-temporal images. Although both CNN and RNN are well-established techniques for remote sensing applications, to the best of our knowledge, we are the first to combine them for multitemporal data analysis in the remote sensing community. Both visual and quantitative analysis of experimental results demonstrates competitive performance in the proposed mode.
Lichao Mou, Xiao Xiang Zhu 0001
IGARSS2
2018 Feature Importance Analysis of Sentinel-2 Imagery for Large-Scale Urban Local Climate Zone Classification
abstract
This paper evaluates different spectral-spatial features that can be extracted from Sentinel-2 imagery regarding their relevance for discriminating different Local Climate Zone (LCZ) classes. The features include spectral reflectance, spectral indices, Morphological Profiles (MPs), as well as Global Urban Footprint (GUF), the Open Street Map layers buildings and land use, and their combinations. Using a residual convolutional neural network (ResNet), a systematic analysis of feature importance is performed with a manually generated dataset distributed in Europe. The results of this evaluation are meant to provide guidance about the choice of both spectral and spatial features for the task of LCZ classification on a global scale. The results show that GUF and OSM can contribute to the classification performance, and ResNet relies less on additional features with the highest accuracy provided by the reflectance only.
Chunping Qiu, Michael Schmitt 0003, Pedram Ghamisi, Lichao Mou, Xiao Xiang Zhu 0001
IGARSS5
2018 Extraction of Buildings in VHR SAR Images Using Fully Convolution Neural Networks
abstract
Modern spaceborne synthetic aperture radar (SAR) sensors, such as TerraSAR-X/TanDEM-X and COSMO-SkyMed, can deliver very high resolution (VHR) data beyond the inherent spatial scales (on the order of 1m) of buildings, constituting invaluable data source for large-scale urban mapping. Processing this VHR data with advanced interferometric techniques, such as SAR tomography (TomoSAR), enables the generation of 3-D (or even 4-D) TomoSAR point clouds from space. In this paper, we present a novel and generic workflow that exploits these TomoSAR point clouds in a way that is capable to automatically produce benchmark annotated (buildings/non-buildings) SAR datasets. These annotated datasets (building masks) have been utilized to construct and train the state-of-the-art deep Fully Convolution Neural Networks with an additional Conditional Random Field represented as a Recurrent Neural Network to detect building regions in a single VHR SAR image. The results of building detection are illustrated and validated over TerraSAR-X VHR spotlight SAR image covering approximately 39 km2- almost the whole city of Berlin - with mean pixel accuracies of around 93.84%.
Muhammad Shahzad 0002, Michael Maurer, Friedrich Fraundorfer, Yuanyuan Wang 0002, Xiao Xiang Zhu 0001
IGARSS5
2018 SAR Tomography Using Non-Local Sparse Reconstruction
abstract
Synthetic Aperture Radar Tomography is an advanced remote sensing method that is able to reconstruct the 3D distribution of scatterers. One promising solution is to use sophisticated reconstruction algorithms, like the compressive sensing based algorithm. However, it suffers from the high demand of number of images and the computational expenses. Therefore, it is hard to apply for large scale practice. In this work, a complete work-flow for Non-Local Compressive Sensing based SAR Tomography method is presented. Moreover, the applicability of the algorithm is demonstrated by exploiting TerraSAR-X high resolution spotlight images over a test site in Munich, Germany.
Yilei Shi, Xiao Xiang Zhu 0001, Richard Bamler
IGARSS2
2018 The SARptical Dataset for Joint Analysis of SAR and Optical Image in Dense Urban Area
abstract
The joint interpretation of very high resolution SAR and optical images in dense urban area are not trivial due to the distinct imaging geometry of the two types of images. Especially, the inevitable layover caused by the side-looking SAR imaging geometry renders this task even more challenging. Only until recently, the “SARptical” framework [1], [2] proposed a promising solution to tackle this. SARptical can trace individual SAR scatterers in corresponding high-resolution optical images, via rigorous 3-D reconstruction and matching. This paper introduces the SARptical dataset1, which is a dataset of over 10,000 pairs of corresponding SAR, and optical image patches extracted from TerraSAR-X high-resolution spotlight images and aerial UltraCAM optical images. This dataset opens new opportunities of multisensory data analysis. One can analyze the geometry, material, and other properties of the imaged object in both SAR and optical image domain. More advanced applications such as SAR and optical image matching via deep learning [3], [4] is now also possible.
Yuanyuan Wang 0002, Xiao Xiang Zhu 0001
IGARSS2
2018 Identifying Corresponding Patches in SAR and Optical Images With a Pseudo-Siamese CNN
abstract
In this letter, we propose a pseudo-siamese convolutional neural network architecture that enables to solve the task of identifying corresponding patches in very high-resolution optical and synthetic aperture radar (SAR) remote sensing imagery. Using eight convolutional layers each in two parallel network streams, a fully connected layer for the fusion of the features learned in each stream, and a loss function based on binary cross entropy, we achieve a one-hot indication if two patches correspond or not. The network is trained and tested on an automatically generated data set that is based on a deterministic alignment of SAR and optical imagery via previously reconstructed and subsequently coregistered 3-D point clouds. The satellite images, from which the patches comprising our data set are extracted, show a complex urban scene containing many elevated objects (i.e., buildings), thus providing one of the most difficult experimental environments. The achieved results show that the network is able to predict corresponding patches with high accuracy, thus indicating great potential for further development toward a generalized multisensor key-point matching procedure.
Lloyd H. Hughes, Michael Schmitt 0003, Lichao Mou, Yuanyuan Wang 0002, Xiao Xiang Zhu 0001
IEEE Geosci. Remote. Sens. Lett.5
2018 A Nonlocal InSAR Filter for High-Resolution DEM Generation From TanDEM-X Interferograms
abstract
This paper presents a nonlocal interferometric synthetic aperture radar (InSAR) filter with the goal of generating digital elevation models (DEMs) of higher resolution and accuracy from bistatic TanDEM-X strip map interferograms than with the processing chain used in production. The currently employed boxcar multilooking filter naturally decreases the resolution and has inherent limitations on what level of noise reduction can be achieved. The proposed filter is specifically designed to account for the inherent diversity of natural terrain by setting several filtering parameters adaptively. In particular, it considers the local fringe frequency and scene heterogeneity, ensuring proper denoising of interferograms with considerable underlying topography as well as urban areas. A comparison using synthetic and TanDEM-X bistatic strip map data sets with existing InSAR filters shows the effectiveness of the proposed techniques, most of which could readily be integrated into existing nonlocal filters. The resulting DEMs outclass the ones produced with the existing global TanDEM-X DEM processing chain by effectively increasing the resolution from 12 to 6 m and lowering the noise level by roughly a factor of two.
Gerald Baier, Cristian Rossi, Marie Lachaise, Xiao Xiang Zhu 0001, Richard Bamler
IEEE Trans. Geosci. Remote. Sens.4
2018 Nonlocal Tensor Completion for Multitemporal Remotely Sensed Images' Inpainting
abstract
Remotely sensed images may contain some missing areas because of poor weather conditions and sensor failure. Information of those areas may play an important role in the interpretation of multitemporal remotely sensed data. This paper aims at reconstructing the missing information by a nonlocal low-rank tensor completion method. First, nonlocal correlations in the spatial domain are taken into account by searching and grouping similar image patches in a large search window. Then, low rankness of the identified fourth-order tensor groups is promoted to consider their correlations in spatial, spectral, and temporal domains, while reconstructing the underlying patterns. Experimental results on simulated and real data demonstrate that the proposed method is effective both qualitatively and quantitatively. In addition, the proposed method is computationally efficient compared with other patch-based methods such as the recently proposed patch matching-based multitemporal group sparse representation method.
Teng-Yu Ji, Naoto Yokoya, Xiao Xiang Zhu 0001, Ting-Zhu Huang
IEEE Trans. Geosci. Remote. Sens.3
2018 Object-Based Multipass InSAR via Robust Low-Rank Tensor Decomposition
abstract
The most unique advantage of multipass synthetic aperture radar interferometry (InSAR) is the retrieval of long-term geophysical parameters, e.g., linear deformation rates, over large areas. Recently, an object-based multipass InSAR framework has been proposed by Kang, as an alternative to the typical single-pixel methods, e.g., persistent scatterer interferometry (PSI), or pixel-cluster-based methods, e.g., SqueeSAR. This enables the exploitation of inherent properties of InSAR phase stacks on an object level. As a follow-on, this paper investigates the inherent low rank property of such phase tensors and proposes a Robust Multipass InSAR technique via Object-based low rank tensor decomposition. We demonstrate that the filtered InSAR phase stacks can improve the accuracy of geophysical parameters estimated via conventional multipass InSAR techniques, e.g., PSI, by a factor of 10-30 in typical settings. The proposed method is particularly effective against outliers, such as pixels with unmodeled phases. These merits, in turn, can effectively reduce the number of images required for a reliable estimation. The promising performance of the proposed method is demonstrated using high-resolution TerraSAR-X image stacks.
Jian Kang 0005, Yuanyuan Wang 0002, Michael Schmitt 0003, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.4
2018 HSF-Net: Multiscale Deep Feature Embedding for Ship Detection in Optical Remote Sensing Imagery
abstract
Ship detection is an important and challenging task in remote sensing applications. Most methods utilize specially designed hand-crafted features to detect ships, and they usually work well only on one scale, which lack generalization and impractical to identify ships with various scales from multiresolution images. In this paper, we propose a novel deep feature-based method to detect ships in very high-resolution optical remote sensing images. In our method, a regional proposal network is used to generate ship candidates from feature maps produced by a deep convolutional neural network. To efficiently detect ships with various scales, a hierarchical selective filtering layer is proposed to map features in different scales to the same scale space. The proposed method is an end-to-end network that can detect both inshore and offshore ships ranging from dozens of pixels to thousands. We test our network on a large ship data set which will be released in the future, consisting of Google Earth images, GaoFen-2 images, and unmanned aerial vehicle data. Experiments demonstrate high precision and robustness of our method. Further experiments on aerial images show its good generalization to unseen scenes.
Qingpeng Li, Lichao Mou, Qingjie Liu 0001, Yunhong Wang 0001, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.5
2018 Automatic Detection and Positioning of Ground Control Points Using TerraSAR-X Multiaspect Acquisitions
abstract
Geodetic stereo synthetic aperture radar (SAR) is capable of absolute 3-D localization of natural persistent scatterers, which allows for ground control point (GCP) generation using only SAR data. The prerequisite for the method to achieve high-precision results is the correct detection of common scatterers in SAR images acquired from different viewing geometries. In this contribution, we describe three strategies for automatic detection of identical targets in SAR images of urban areas taken from different orbit tracks. Moreover, a complete workflow for automatic generation of large number of GCPs using SAR data is presented and its applicability is shown by exploiting TerraSAR-X high-resolution spotlight images over the city of Oulu, Finland, and a test site in Berlin, Germany.
Sina Montazeri, Christoph Gisinger, Michael Eineder, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.4
2018 Unsupervised Spectral-Spatial Feature Learning via Deep Residual Conv-Deconv Network for Hyperspectral Image Classification
abstract
Supervised approaches classify input data using a set of representative samples for each class, known as training samples. The collection of such samples is expensive and time demanding. Hence, unsupervised feature learning, which has a quick access to arbitrary amounts of unlabeled data, is conceptually of high interest. In this paper, we propose a novel network architecture, fully Conv-Deconv network, for unsupervised spectral-spatial feature learning of hyperspectral images, which is able to be trained in an end-to-end manner. Specifically, our network is based on the so-called encoder-decoder paradigm, i.e., the input 3-D hyperspectral patch is first transformed into a typically lower dimensional space via a convolutional subnetwork (encoder), and then expanded to reproduce the initial data by a deconvolutional subnetwork (decoder). However, during the experiment, we found that such a network is not easy to be optimized. To address this problem, we refine the proposed network architecture by incorporating: 1) residual learning and 2) a new unpooling operation that can use memorized max-pooling indexes. Moreover, to understand the “black box,” we make an in-depth study of the learned feature maps in the experimental analysis. A very interesting discovery is that some specific “neurons” in the first residual block of the proposed network own good description power for semantic visual patterns in the object level, which provide an opportunity to achieve “free” object detection. This paper, for the first time in the remote sensing community, proposes an end-to-end fully Conv-Deconv network for unsupervised spectral-spatial feature learning. Moreover, this paper also introduces an in-depth investigation of learned features. Experimental results on two widely used hyperspectral data, Indian Pines and Pavia University, demonstrate competitive performance obtained by the proposed methodology compared with other studied approaches.
Lichao Mou, Pedram Ghamisi, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.3
2018 Corrections to "Deep Recurrent Neural Networks for Hyperspectral Image Classification"
abstract
Here, we correct some errors caused by a programming bug (a data type error) in overall accuracies (OAs) reported in[1]. The corrected OAs are underlined and shown in bold inTables I–III.
Lichao Mou, Pedram Ghamisi, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.3
2018 Vehicle Instance Segmentation From Aerial Image and Video Using a Multitask Learning Residual Fully Convolutional Network
abstract
Object detection and semantic segmentation are two main themes in object retrieval from high-resolution remote sensing images, which have recently achieved remarkable performance by surfing the wave of deep learning and, more notably, convolutional neural networks. In this paper, we are interested in a novel, more challenging problem of vehicle instance segmentation, which entails identifying, at a pixel level, where the vehicles appear as well as associating each pixel with a physical instance of a vehicle. In contrast, vehicle detection and semantic segmentation each only concern one of the two. We propose to tackle this problem with a semantic boundary-aware multitask learning network. More specifically, we utilize the philosophy of residual learning to construct a fully convolutional network that is capable of harnessing multilevel contextual feature representations learned from different residual blocks. We theoretically analyze and discuss why residual networks can produce better probability maps for pixelwise segmentation tasks. Then, based on this network architecture, we propose a unified multitask learning network that can simultaneously learn two complementary tasks, namely, segmenting vehicle regions and detecting semantic boundaries. The latter subproblem is helpful for differentiating “touching” vehicles that are usually not correctly separated into instances. Currently, data sets with a pixelwise annotation for vehicle extraction are the ISPRS data set and the IEEE GRSS DFC2015 data set over Zeebrugge, which specializes in a semantic segmentation. Therefore, we built a new, more challenging data set for vehicle instance segmentation, called the Busy Parking Lot Unmanned Aerial Vehicle Video data set, and we make our data set available at http://www.sipeo.bgu.tum.de/downloads so that it can be used to benchmark future vehicle instance segmentation algorithms.
Lichao Mou, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.2
2018 A Fast and Accurate Basis Pursuit Denoising Algorithm With Application to Super-Resolving Tomographic SAR
abstract
$L_{1}$regularization is used for finding sparse solutions to an underdetermined linear system. As sparse signals are widely expected in remote sensing, this type of regularization scheme and its extensions have been widely employed in many remote sensing problems, such as image fusion, target detection, image super-resolution, and others, and have led to promising results. However, solving such sparse reconstruction problems is computationally expensive and has limitations in its practical use. In this paper, we proposed a novel efficient algorithm for solving the complex-valued$L_{1}$regularized least squares problem. Taking the high-dimensional tomographic synthetic aperture radar (TomoSAR) as a practical example, we carried out extensive experiments, both with the simulation data and the real data, to demonstrate that the proposed approach can retain the accuracy of the second-order methods while dramatically speeding up the processing by one or two orders. Although we have chosen TomoSAR as the example, the proposed method can be generally applied to any spectral estimation problems.
Yilei Shi, Xiao Xiang Zhu 0001, Wotao Yin, Richard Bamler
IEEE Trans. Geosci. Remote. Sens.2
2018 InSAR-BM3D: A Nonlocal Filter for SAR Interferometric Phase Restoration
abstract
The block-matching 3-D (BM3D) algorithm, based on the nonlocal approach, is one of the most effective methods to date for additive white Gaussian noise image denoising. Likewise, its extension to synthetic aperture radar (SAR) amplitude images, SAR-BM3D, is a state-of-the-art SAR despeckling algorithm. In this paper, we further extend BM3D to address the restoration of SAR interferometric phase images. While keeping the general structure of BM3D, its processing steps are modified to take into account the peculiarities of the SAR interferometry signal. Experiments on simulated and real-world Tandem-X SAR interferometric pairs prove the effectiveness of the proposed method.
Francescopaolo Sica, Davide Cozzolino, Xiao Xiang Zhu 0001, Luisa Verdoliva, Giovanni Poggi
IEEE Trans. Geosci. Remote. Sens.3
2017 Learning a low-coherence dictionary to address spectral variability for hyperspectral unmixing
abstract
This paper presents a novel spectral mixture model to address spectral variability in inverse problems of hyperspectral unmixing. Based on the linear mixture model (LMM), our model introduces a spectral variability dictionary to account for any residuals that cannot be explained by the LMM. Atoms in the dictionary are assumed to be low-coherent with spectral signatures of endmembers. A dictionary learning technique is proposed to learn the spectral variability dictionary while solving unmixing problems simultaneously. Experimental results on synthetic and real datasets demonstrate that the performance of the proposed method is superior to state-of-the-art methods.
Danfeng Hong, Naoto Yokoya, Jocelyn Chanussot, Xiao Xiang Zhu 0001
ICIP4
2017 Robust linear unmixing with enhanced sparsity
abstract
Spectral unmixing is a central problem in hyperspectral imagery. It is usually assuming a linear mixture model. Solving this inverse problem, however, can be seriously impacted by a wrong estimation of the number of endmembers, a bad estimation of the endmembers themselves, the spectral variability of the endmembers or the presence of nonlinearities. These problems can result in a too large number of retained endmembers. We propose to tackle this problem by introducing a new formulation for robust linear unmixing enhancing sparsity. With a single tuning parameter the optimization leads to a range of behaviors: from the standard linear model (low sparsity) to a hard classification (maximal sparsity : only one endmember is retained per pixel). We solve the proposed new functional using a computationally efficient proximal primal dual method. The experimental study, including both realistic simulated data and real data demonstrates the versatility of the proposed approach.
Alexandre Tiard, Laurent Condat, Lucas Drumetz, Jocelyn Chanussot, Wotao Yin, Xiao Xiang Zhu 0001
ICIP6
2017 Fusion of SAR and optical remote sensing data - Challenges and recent trends
abstract
In this paper, we summarize challenges, proposed solutions and recent trends in the field of SAR-optical remote sensing data fusion. Although being a pre-processing step before the actual fusion-by-estimation, it is shown that matching and coregistration is one of the core challenges in that regard, which is mainly due to the strongly different geometric and radiometric properties of the two observation types. We then review some of the published fusion methods and discuss the future trends of this topic.
Michael Schmitt 0003, Florence Tupin, Xiao Xiang Zhu 0001
IGARSS3
2017 Fusion of TanDEM-X and Cartosat-1 DEMS using TV-norm regularization and ANN-predicted weights
abstract
This paper deals with TanDEM-X and Cartosat-1 DEM fusion over urban areas with support of weight maps predicted by an artificial neural network (ANN). Although the TanDEM-X DEM is a global elevation dataset of unprecedented accuracy (following HRTI-3 standard), its quality decreases over urban areas because of artifacts intrinsic to the SAR imaging geometry. DEM fusion techniques can be used to improve the TanDEM-X DEM in problematic areas. In this investigation, Cartosat-1 elevation data were fused with the TanDEM-X DEM by weighted averaging and total variation (TV)-based regularization, resorting to weight maps derived by a specifically trained ANN. The results show that the proposed fusion strategy can significantly improve the final DEM quality.
Hossein Bagheri, Michael Schmitt 0003, Xiao Xiang Zhu 0001
IGARSS3
2017 Nonlocal InSAR filtering for high resolution DEM generation from TanDEM-X interferograms
abstract
We investigate the feasibility of generating highly accurate digital elevation models (DEM) from TanDEM-X interferograms by using nonlocal filters for phase denoising. Some of the shortcomings of existing nonlocal filters that render them not applicable to our goal are briefly described and a new filter is proposed that alleviates these problems. The most significant new properties are addressing the slope dependent denoising performance of existing nonlocal InSAR filters and several measures to bolster denoising near edgelike features. We evaluate the proposed filter using synthetic interferograms and by visual inspection of a DEM generated from a TanDEM-X interferogram.
Gerald Baier, Cristian Rossi, Marie Lachaise, Xiao Xiang Zhu 0001, Richard Bamler
IGARSS4
2017 Feature fusion of hyperspectral and lidar data using extinction profiles and total variation
abstract
To improve the classification of hyperspectral images, this paper proposes an approach for multi-sensor data fusion of LiDAR and hyperspectral data using extinction profiles and Orthogonal Total Variation Component Analysis (OTVCA). Results on the benchmark Houston data indicate the superior performance of the proposed approach compared to other approaches used in the experiments based on classification accuracies.
Pedram Ghamisi, Behnood Rasti, Xiao Xiang Zhu 0001
IGARSS3
2017 Evaluation of polsar similarity measures with spectral clustering
abstract
Polarimetric Synthetic Aperture Radar (PolSAR) is a valuable remote sensing data source. It is usually challenging to interpret PolSAR data, especially in urban areas, and hense, spatial clustering comes as a powerful tool for the application of PolSAR data. In data clustering, similarity measurement indexes are of great importance. By far, there are quite some similarity measures of PolSAR data. However, to our knowledge, there has no practical and systematic evaluation of the performances of these measures. In this paper, we evaluate seven different similarity measurements of PolSAR data in the context of clustering using the conventional spectral clustering algorithm.
Jingliang Hu, Yuanyuan Wang 0002, Pedram Ghamisi, Xiao Xiang Zhu 0001
IGARSS4
2017 Improve multi-baseline InSAR parameter retrieval by semantic information from optical images
abstract
One of the most unique benefits of multi-baseline synthetic aperture radar interferometry (InSAR) is the long-term monitoring of subtle ground deformation over large areas. Most state-of-the-art algorithms for retrieving such parameter are based on single pixels, e.g. Permanent Scatterer InSAR [1] or clusters of ergodic pixels with stationary phases e.g. SqueeSAR [2]. None of the studies has addressed the joint inversion in an object level, where the true interferometric phase may be varying subject to topography and deformation. Recently, one study has investigated SAR and optical data fusion in order to make use of the rich semantic information from optical images [3]. Based on that work, we seek to investigate the possibility of an object-level multi-baseline InSAR deformation reconstruction given the semantic information from the corresponding optical images. In this paper, we introduced the tensor model for the multi-baseline InSAR inversion and proposed a maximum a posteriori estimator of the deformation parameters by including a spatial prior function in the objective function. Substantial improvement in the deformation estimation is observed in the experiments using both simulated and the real SAR data.
Jian Kang 0005, Yuanyuan Wang 0002, Marco Körner 0001, Xiao Xiang Zhu 0001
IGARSS4
2017 Automatic positioning of SAR ground control points from multi-aspect TerraSAR-X acquisitions
abstract
Geodetic stereo SAR is capable of absolute 3-D localization of natural persistent scatterers (PS)s which allows for Ground Control Point (GCP) generation using only SAR data. The prerequisite for the method to achieve high precision results is the correct detection of common scatterers in SAR images acquired from different viewing geometries. In this contribution, we describe three strategies for automatic detection of identical point targets in SAR images of urban areas taken from different orbit tracks. Moreover, a complete work-flow for automatic generation of large number of GCPs using SAR data is presented and its applicability is shown by exploiting TerraSAR-X high resolution spotlight images over the city of Oulu, Finland and a test site in Berlin, Germany.
Sina Montazeri, Christoph Gisinger, Xiao Xiang Zhu 0001, Michael Eineder, Richard Bamler
IGARSS3
2017 Identifying corresponding patches in SAR and optical imagery with a convolutional neural network
abstract
In this paper, we investigate making use of a convolutional neural network (CNN) to solve the task of identifying corresponding patches in very high resolution (VHR) optical and SAR imagery of complicated urban scenery. By doing so, the binary decision function is learnt directly from automatically generated training data and does not resort to any hand-crafted features. First evaluations show great potential for further studies towards a generalized multi-sensor matching procedure.
Lichao Mou, Michael Schmitt 0003, Yuanyuan Wang 0002, Xiao Xiang Zhu 0001
IGARSS4
2017 Fully conv-deconv network for unsupervised spectral-spatial feature extraction of hyperspectral imagery via residual learning
abstract
Supervised approaches classify input data using a set of representative samples for each class, known as training samples. The collection of such samples are expensive and time-demanding. Hence, unsupervised feature learning, which has a quick access to arbitrary amount of unlabeled data, is conceptually of high interest. In this paper, we propose a novel network architecture, fully Conv-Deconv network with residual learning, for unsupervised spectral-spatial feature learning of hyperspectral images, which is able to be trained in an end-to-end manner. Specifically, our network is based on the so-called encoder-decoder paradigm, i.e., the input 3D hyperspectral patch is first transformed into a typically lower-dimensional space via a convolutional sub-network (encoder), and then expanded to reproduce the initial data by a deconvolutional sub-network (decoder). Experimental results on the Pavia University hyperspectral data set demonstrate competitive performance obtained by the proposed methodology compared to other studied approaches.
Lichao Mou, Pedram Ghamisi, Xiao Xiang Zhu 0001
IGARSS3
2017 Comparative evaluation of signal-based and descriptor-based similarity measures for SAR-optical image matching
abstract
This paper compares different similarity measures for the matching of very-high-resolution SAR and optical images over urban areas. It is meant to provide guidance about the performance of both signal-based and descriptor-based similarity measures in the context of this non-trivial case of multi-sensor correspondence matching. Using an automatically generated training dataset, thresholds for the distinction between correct matches and wrong matches are determined. It is shown that descriptor-based similarity measures outperform signal-based similarity measures significantly.
Chunping Qiu, Michael Schmitt 0003, Xiao Xiang Zhu 0001
IGARSS3
2017 Robust blind scatterer separation in multibaseline InSAR
abstract
The side-looking imaging geometry of synthetic aperture radar (SAR) causes inevitable layover in SAR images. Separating the contributions from different scatterers has been the fundamental for many applications. It is typically solved by explicit inversion of the SAR imaging model to retrieve the scattering profile along the mixed dimension (elevation), which is otherwise known as SAR tomography. This paper proposed a robust blind scatterer separation method to demix the layovered scatterers, avoiding the computationally expensive tomographic inversion. We demonstrate that the state-of-the-art principle component decomposition-based methods are heavily influenced by the nonergodicity of the selected samples, especially in urban area, such as point scatterers appearing often on facades. The proposed method is shown to be more robust than the state-of-the-art. Real data example shows that the proposed method outperforms the state-of-the-art by a factor of three in terms of the accuracy of the retrieved phase.
Yuanyuan Wang 0002, Xiao Xiang Zhu 0001
IGARSS2
2017 L1-Regularization-Based SAR Imaging and CFAR Detection via Complex Approximated Message Passing
abstract
Synthetic aperture radar (SAR) is a widely used active high-resolution microwave imaging technique that has alltime and all-weather reconnaissance ability. Compared with traditionally matched filtering (MF)-based methods, Lq(0 ≤ q ≤ 1) regularization technique can efficiently improve SAR imaging performance e.g., suppressing sidelobes and clutter. However, conventional Lq-regularization-based SAR imaging approach requires transferring the 2-D echo data into a vector and reconstructing the scene via 2-D matrix operations. This leads to significantly more computational complexity compared with MF, and makes it very difficult to apply in high-resolution and wide-swath imaging. Typical Lqregularization recovery algorithms, e.g., iterative thresholding algorithm, can improve imaging performance of bright targets, but not preserve the image background distribution well. Thus, image background statistical-property-based applications, such as constant false alarm rate (CFAR) detection, cannot be applied to regularization recovered SAR images. On the other hand, complex approximated message passing (CAMP), an iterative recovery algorithm for L1regularization reconstruction, can achieve not only the sparse estimation of the original signal as typical regularization recovery algorithms but also a nonsparse solution simultaneously. In this paper, two novel CAMP-based SAR imaging algorithms are proposed for raw data and complex radar image data, respectively, along with CFAR detection via the CAMP recovered nonsparse result. The proposed method for raw data can not only improve SAR image performance as conventional L1regularization technique but also reduce the computational cost efficiently. While only when we have MF recovered SAR complex image rather than raw data, the proposed method for complex image data can achieve a similar reconstructed image quality as the regularization-based SAR imaging approach using the full raw data. The most important contribution of this paper is that the proposed CAMP-based methods make CFAR detection based on the regularization reconstruction SAR image possible using their nonsparse scene estimations, which has a similar background statistical distribution as the MF recovered images. The experimental results validated the effectiveness of the proposed methods and the feasibility of the recovered nonsparse images being used for CFAR detection.
Hui Bi 0001, Bingchen Zhang, Xiao Xiang Zhu 0001, Wen Hong, Jinping Sun, Yirong Wu
IEEE Trans. Geosci. Remote. Sens.3
2017 Extended Chirp Scaling-Baseband Azimuth Scaling-Based Azimuth-Range Decouple L1 Regularization for TOPS SAR Imaging via CAMP
abstract
This paper proposes a novel azimuth-range decouple-based L1regularization imaging approach for the focusing in terrain observation by progressive scans (TOPS) synthetic aperture radar (SAR). Since conventional L1regularization technique requires transferring the (2-D) echo data into a vector and reconstructing the scene via 2-D matrix operations leading to significantly more computational complexity, it is very difficult to apply in high-resolution and wide-swath SAR imaging, e.g., TOPS. The proposed method can achieve azimuth-range decouple by constructing an approximated observation operator to simulate the raw data, the inverse of matching filtering (MF) procedure, which makes large-scale sparse reconstruction, or called compressive sensing reconstruction of surveillance region with full- or downsampled raw data in TOPS SAR possible. Compared with MF algorithm, e.g., extended chirp scaling-baseband azimuth scaling, it shows huge potential in image performance improvement; while compared with conventional L1regularization technique, it significantly reduces the computational cost, and provides similar image features. Furthermore, this novel approach can also obtain a nonsparse estimation of considered scene retaining a similar background statistical distribution as the MF-based image, which can be used to the further application of SAR images with precondition being preserving image statistical properties, e.g., constant false alarm rate detection. Experimental results along with a performance analysis validate the proposed method.
Hui Bi 0001, Bingchen Zhang, Xiao Xiang Zhu 0001, Chenglong Jiang, Wen Hong
IEEE Trans. Geosci. Remote. Sens.3
2017 Robust Object-Based Multipass InSAR Deformation Reconstruction
abstract
Deformation monitoring by multipass synthetic aperture radar (SAR) interferometry (InSAR) is, so far, the only imaging-based method to assess millimeter-level deformation over large areas from space. Past research mostly focused on the optimal retrieval of deformation parameters on the basis of a single pixel or a pixel cluster. Only until recently, the first demonstration of object-based urban infrastructure monitoring by fusing InSAR and the semantic classification labels derived from optical images was presented by Wanget al.Given such classification labels in the SAR image, we propose a general framework for object-based InSAR parameter retrieval, where the parameters of the whole object are jointly estimated by the inversion of a regularized tensor model instead of pixelwise. Our approach does not assume the stationarity of each sample in the object, which is usually assumed in other pixel cluster-based methods, such as SqueeSAR. In addition, to handle outliers in real data, a robust phase recovery step prior to parameter retrieval is also introduced. In typical settings, the proposed method outperforms the current pixelwise estimators, e.g., periodogram, by a factor of several tens in the accuracy of the linear deformation estimates. Last but not least, for a practical demonstration on bridge monitoring, we present a full workflow of long-term bridge monitoring using the proposed approach.
Jian Kang 0005, Yuanyuan Wang 0002, Marco Körner 0001, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.4
2017 Deep Recurrent Neural Networks for Hyperspectral Image Classification
abstract
In recent years, vector-based machine learning algorithms, such as random forests, support vector machines, and 1-D convolutional neural networks, have shown promising results in hyperspectral image classification. Such methodologies, nevertheless, can lead to information loss in representing hyperspectral pixels, which intrinsically have a sequence-based data structure. A recurrent neural network (RNN), an important branch of the deep learning family, is mainly designed to handle sequential data. Can sequence-based RNN be an effective method of hyperspectral image classification? In this paper, we propose a novel RNN model that can effectively analyze hyperspectral pixels as sequential data and then determine information categories via network reasoning. As far as we know, this is the first time that an RNN framework has been proposed for hyperspectral image classification. Specifically, our RNN makes use of a newly proposed activation function, parametric rectified tanh (PRetanh), for hyperspectral sequential data analysis instead of the popular tanh or rectified linear unit. The proposed activation function makes it possible to use fairly high learning rates without the risk of divergence during the training procedure. Moreover, a modified gated recurrent unit, which uses PRetanh for hidden representation, is adopted to construct the recurrent layer in our network to efficiently process hyperspectral data and reduce the total number of parameters. Experimental results on three airborne hyperspectral images suggest competitive performance in the proposed mode. In addition, the proposed network architecture opens a new window for future research, showcasing the huge potential of deep recurrent networks for hyperspectral data analysis.
Lichao Mou, Pedram Ghamisi, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.3
2017 Fusing Meter-Resolution 4-D InSAR Point Clouds and Optical Images for Semantic Urban Infrastructure Monitoring
abstract
Using synthetic aperture radar (SAR) interferometry to monitor long-term millimeter-level deformation of urban infrastructures, such as individual buildings and bridges, is an emerging and important field in remote sensing. In the state-of-the-art methods, deformation parameters are retrieved and monitored on a pixel basis solely in the SAR image domain. However, the inevitable side-looking imaging geometry of SAR results in undesired occlusion and layover in urban area, rendering the current method less competent for a semantic-level monitoring of different urban infrastructures. This paper presents a framework of a semantic-level deformation monitoring by linking the precise deformation estimates of SAR interferometry and the semantic classification labels of optical images via a 3-D geometric fusion and semantic texturing. The proposed approach provides the first “SARptical” point cloud of an urban area, which is the SAR tomography point cloud textured with attributes from optical images. This opens a new perspective of InSAR deformation monitoring. Interesting examples on bridge and railway monitoring are demonstrated.
Yuanyuan Wang 0002, Xiao Xiang Zhu 0001, Bernhard Zeisl, Marc Pollefeys
IEEE Trans. Geosci. Remote. Sens.2
2017 Multisensor Coupled Spectral Unmixing for Time-Series Analysis
abstract
We present a new framework, called multisensor coupled spectral unmixing (MuCSUn), that solves unmixing problems involving a set of multisensor time-series spectral images in order to understand dynamic changes of the surface at a subpixel scale. The proposed methodology couples multiple unmixing problems based on regularization on graphs between the time-series data to obtain robust and stable unmixing solutions beyond data modalities due to different sensor characteristics and the effects of nonoptimal atmospheric correction. Atmospheric normalization and cross calibration of spectral response functions are integrated into the framework as a preprocessing step. The proposed methodology is quantitatively validated using a synthetic data set that includes seasonal and trend changes on the surface and the residuals of nonoptimal atmospheric correction. The experiments on the synthetic data set clearly demonstrate the efficacy of MuCSUn and the importance of the preprocessing step. We further apply our methodology to a real time-series data set composed of 11 Hyperion and 22 Landsat-8 images taken over Fukushima, Japan, from 2011 to 2015. The proposed methodology successfully obtains robust and stable unmixing results and clearly visualizes class-specific changes at a subpixel scale in the considered study area.
Naoto Yokoya, Xiao Xiang Zhu 0001, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.2
2016 GPU-based nonlocal filtering for large scale SAR processing
abstract
In the past few years nonlocal filters have emerged as a serious contender for denoising synthetic aperture radar (SAR) images, offering superior noise reduction and detail preservation compared to many other filters. In this manuscript we analyze how nonlocal filters, whose computational costs were so far prohibitive for large scale processing, can be implemented efficiently on graphics processing units (GPU). As a case study NL-SAR, a state of the art SAR filter, is implemented to run on a NVIDIA Tesla K40. We describe the appeal of GPUs, or any other coprocessor, for nonlocal filters. Nonlocal filtering of TanDEM-X interferograms for generating digital elevation models with a higher resolution and accuracy is given as an application that benefits from efficient and fast nonlocal filtering.
Gerald Baier, Xiao Xiang Zhu 0001
IGARSS2
2016 Extinction profiles: A novel approach for the analysis of remote sensing data
abstract
This paper presents a novel approach named extinction profiles to model the spatial information of remote sensing images. Then, the output of the extinction profile is fed to a grid-search random forest classification method. Results indicate that the proposed approach can effectively extract spatial information from remote sensing gray scale images and provide high classification accuracies in an automatic way.
Pedram Ghamisi, Roberto Souza 0001, Letícia Rittner, Jón Atli Benediktsson, Roberto A. Lotufo, Xiao Xiang Zhu 0001
IGARSS6
2016 Local manifold learning with robust neighbors selection for hyperspectral dimensionality reduction
abstract
Manifold learning has been successfully applied to hyperspectral dimensionality reduction to embed nonlinear and nonconvex manifolds in the data. However, dimensionality reduction by manifold learning is sensitive to non-uniform data distribution and the selection of neighbors. To address the two issues to some extents, in this work a new manifold framework based on locality linear embedding (LLE), namely local normalization and local feature selection (LNLFS), is proposed. Classification is explored as a potential application to validate the proposed algorithm. Classification accuracy using data obtained using different dimensionality reduction methods is evaluated and compared, while applying two kinds of strategies for selecting the training and test samples: random sampling and region-based sampling. Experimental results show the classification accuracy obtained with LNLFS is superior to state-of-the-art dimensionality reduction methods.
Danfeng Hong, Naoto Yokoya, Xiao Xiang Zhu 0001
IGARSS3
2016 Object-based InSAR deformation reconstruction with application to bridge monitoring
abstract
Deformation monitoring by multi-baseline synthetic aperture radar (SAR) interferometry is so far the only imaging-based method to assess millimeter-level deformation over large areas from space. Past research mostly focused on optimal deformation parameters retrieval on a pixel-basis. Only until recently, the first demonstration of object-based urban infrastructures monitoring by fusing SAR interferometry (InSAR) and the semantic classification labels derived from optical images was presented in [1]–[3]. This paper proposes an algorithm for object-based joint InSAR deformation reconstruction using these classification labels. We derive an object-based multi-baseline InSAR reconstruction model, and propose an efficient algorithm for bridge detection in optical images.
Jian Kang 0005, Yuanyuan Wang 0002, Marco Körner 0001, Xiao Xiang Zhu 0001
IGARSS4
2016 SAR ground control point identification with the aid of high resolution optical data
abstract
Only until recently, it has been demonstrated that absolute localization with centimeter accuracy can be achieved for manually matched Persistent Scatterer (PS)s from TerraSAR-X images acquired from cross-heading geometries [1]. This paper describes an automatic algorithm for absolute localization of natural PSs in SAR images, where the detection of potential PSs is aided by high resolution optical data. As the focus of the study is on urban area, the target detection part relies on identification of lamp posts using template matching. These targets are, most probably, the only ones visible in SAR images acquired from both ascending and descending orbits. Thus, the methodology includes identification of lamp posts from high resolution optical data and retrieves the precise absolute three-dimensional coordinates of the points from corrected TerraSAR-X timing measurements using the stereo SAR method [1]. Preliminary results for a test site in the city of Berlin acquired from TerraSAR-X high resolution spotlight mode are shown.
Sina Montazeri, Xiao Xiang Zhu 0001, Ulrich Balss, Christoph Gisinger, Yuanyuan Wang 0002, Michael Eineder, Richard Bamler
IGARSS2
2016 Spatiotemporal scene interpretation of space videos via deep neural network and tracklet analysis
abstract
Spaceborne remote sensing videos are becoming indispensable resources, opening up opportunities for new remote sensing applications. To exploit this new type of data, we need sophisticated algorithms for semantic scene interpretation. The main difficulties are: 1) Due to the relatively poor spatial resolution of the video acquired from space, moving objects, like cars, are very difficult to detect, not to mention track; 2) camera movement handicaps scene interpretation. To address these challenges, in this paper we propose a novel framework that fuses multispectral images and space videos for spatiotemporal analysis. Taking a multispectral image and a spaceborne video as input, an innovative deep neural network is proposed to fuse them in order to achieve a fine-resolution spatial scene labeling map. Moreover, a sophisticated approach is proposed to analyze activities and estimate traffic density from 150,000+ tracklets produced by a Kanade-Lucas-Tomasi keypoint tracker. The proposed framework is validated using data provided for the 2016 IEEE GRSS data fusion contest, including a video acquired from the International Space Station and a DEIMOS-2 multispectral image. Both visual and quantitative analysis of the experimental results demonstrates the effectiveness of our approach.
Lichao Mou, Xiao Xiang Zhu 0001
IGARSS2
2016 Forest remote sensing on the individual tree level by airborne millimeterwave SAR
abstract
This paper presented experimental results discussing the potential of millimeterwave SAR for forest remote sensing on the individual tree level. As can be seen from the experimental results, although there is a certain amount of canopy penetration, a significant part of the signal response is received from the tree crowns. This provides both interesting perspectives for an analysis of forest volumes by continuous TomoSAR models as well as the reconstruction of individual tree models by utilization of discrete TomoSAR models.
Michael Schmitt 0003, Muhammad Shahzad 0002, Xiao Xiang Zhu 0001
IGARSS3
2016 Robust multibaseline InSAR optimization
abstract
Multibaseline SAR interferometry may face unmodeled interferometric phase such as unmodeled motion phase and uncompensated atmospheric phase, as well as non-Gaussian statistics in the context of distributed scatterer. We developed the robust InSAR optimization (RIO) [1] framework to systematically tackle these issues. Experiments show that RIO outperforms the current multibaseline InSAR methods in terms of the variance of the phase history parameters estimates for contaminated observations, while still keeping a relative efficiency of 80% for outlier-free observations.
Yuanyuan Wang 0002, Xiao Xiang Zhu 0001
IGARSS2
2016 A Self-Improving Convolution Neural Network for the Classification of Hyperspectral Data
abstract
In this letter, a self-improving convolutional neural network (CNN) based method is proposed for the classification of hyperspectral data. This approach solves the so-called curse of dimensionality and the lack of available training samples by iteratively selecting the most informative bands suitable for the designed network via fractional order Darwinian particle swarm optimization. The selected bands are then fed to the classification system to produce the final classification map. Experimental results have been conducted with two well-known hyperspectral data sets: Indian Pines and Pavia University. Results indicate that the proposed approach significantly improves a CNN-based classification method in terms of classification accuracy. In addition, this letter uses the concept of dither for the first time in the remote sensing community to tackle overfitting.
Pedram Ghamisi, Yushi Chen 0002, Xiao Xiang Zhu 0001
IEEE Geosci. Remote. Sens. Lett.3
2016 Hyperspectral Data Classification Using Extended Extinction Profiles
abstract
This letter proposes a new approach for the spectral-spatial classification of hyperspectral images, which is based on a novel extrema-oriented connected filtering technique, entitled as extended extinction profiles. The proposed approach progressively simplifies the first informative features extracted from hyperspectral data considering different attributes. Then, the classification approach is applied on two well-known hyperspectral data sets, i.e., Pavia University and Indian Pines, and compared with one of the most powerful filtering approaches in the literature, i.e., extended attribute profiles. Results indicate that the proposed approach is able to efficiently extract spatial information for the classification of hyperspectral images automatically and swiftly. In addition, an array-based node-oriented max-tree representation was carried out to efficiently implement the proposed approach.
Pedram Ghamisi, Roberto Souza 0001, Jón Atli Benediktsson, Letícia Rittner, Roberto A. Lotufo, Xiao Xiang Zhu 0001
IEEE Geosci. Remote. Sens. Lett.6
2016 Unambiguous SAR Imaging for Nonuniform DPC Sampling: ℓq Regularization Method Using Filter Bank
abstract
The displaced phase center antenna (DPCA) technique is a classical method for achieving high-resolution wide-swath synthetic aperture radar (SAR) imaging. For optimum performance, the pulse repetition frequency (PRF) of DPCA SAR systems should satisfy the azimuth uniform sampling condition as far as possible. However, this rigid PRF selection may conflict with the timing diagram for some incidence angles, which usually results in a nonuniform sampling of the synthetic aperture. According to the sparse signal processing theory, this letter proposes a novel DPCA imaging algorithm for the nonuniform displaced phase center sampling. By combining the DPCA data processing operator based on a filter bank with the ℓqregularization scheme, the algorithm can efficiently recover the backscattering coefficients of the observed scene. The experimental results have shown that it is capable of resolving ambiguity and suppressing clutter effectively and is meanwhile insensitive to additive noise.
Xiangyin Quan, Bingchen Zhang, Xiao Xiang Zhu 0001, Yirong Wu
IEEE Geosci. Remote. Sens. Lett.3
2016 Demonstration of Single-Pass Millimeterwave SAR Tomography for Forest Volumes
abstract
In this letter, for the first time, the potential of millimeterwave synthetic aperture radar (SAR) is investigated with respect to a tomographic analysis of forest volumes. Exploiting both parametric and nonparametric SAR tomography (TomoSAR) methods designed for both discrete and continuous reflectivity profiles, it is shown that even Ka-band signals with a wavelength of only 8.55 mm can penetrate the tree canopy to a certain extent and allow a separation of ground and tree crowns. First experimental results exploiting airborne multiantenna data are evaluated with respect to LiDAR ground truth and indicate a promising perspective.
Michael Schmitt 0003, Xiao Xiang Zhu 0001
IEEE Geosci. Remote. Sens. Lett.2
2016 Extinction Profiles for the Classification of Remote Sensing Data
abstract
With respect to recent advances in remote sensing technologies, the spatial resolution of airborne and spaceborne sensors is getting finer, which enables us to precisely analyze even small objects on the Earth. This fact has made the research area of developing efficient approaches to extract spatial and contextual information highly active. Among the existing approaches, morphological profile and attribute profile (AP) have gained great attention due to their ability to classify remote sensing data. This paper proposes a novel approach that makes it possible to precisely extract spatial and contextual information from remote sensing images. The proposed approach is based on extinction filters, which are used here for the first time in the remote sensing community. Then, the approach is carried out on two well-known high-resolution panchromatic data sets captured over Rome, Italy, and Reykjavik, Iceland. In order to prove the capabilities of the proposed approach, the obtained results are compared with the results from one of the strongest approaches in the literature, i.e., APs, using different points of view such as classification accuracies, simplification rate, and complexity analysis. Results indicate that the proposed approach can significantly outperform its alternative in terms of classification accuracies. In addition, based on our implementation, profiles can be generated in a very short processing time. It should be noted that the proposed approach is fully automatic.
Pedram Ghamisi, Roberto Souza 0001, Jón Atli Benediktsson, Xiao Xiang Zhu 0001, Letícia Rittner, Roberto A. Lotufo
IEEE Trans. Geosci. Remote. Sens.4
2016 Three-Dimensional Deformation Monitoring of Urban Infrastructure by Tomographic SAR Using Multitrack TerraSAR-X Data Stacks
abstract
Differential synthetic aperture radar tomography (D-TomoSAR), similar to its conventional counterparts such as differential interferometric SAR and persistent scatterer interferometry, is only capable of capturing 1-D deformation along the satellite's line of sight. In this paper, we propose a method based on L1-norm minimization within local spatial cubes to reconstruct 3-D displacement vectors from TomoSAR point clouds available from at least three different viewing geometries. The methodology is applied on two pairs of cross-heading-combination of ascending and descending-TerraSAR-X (TS-X) spotlight image stacks over the city of Berlin. The linear deformation rate and the amplitude of seasonal deformation are decomposed, and the results from two test sites with remarkable deformation pattern are discussed in detail. The results, to our knowledge, demonstrate the first attempt for motion decomposition using TomoSAR data from multiple viewing geometries.
Sina Montazeri, Xiao Xiang Zhu 0001, Michael Eineder, Richard Bamler
IEEE Trans. Geosci. Remote. Sens.2
2016 Automatic Detection and Reconstruction of 2-D/3-D Building Shapes From Spaceborne TomoSAR Point Clouds
abstract
Modern spaceborne synthetic aperture radar (SAR) sensors, such as TerraSAR-X/TanDEM-X and COSMO-SkyMed, can deliver very high resolution (VHR) data beyond the inherent spatial scales of buildings. Processing these VHR data with advanced interferometric techniques, such as SAR tomography (TomoSAR), allows for the generation of four-dimensional point clouds, containing not only the 3-D positions of the scatterer location but also the estimates of seasonal/temporal deformation on the scale of centimeters or even millimeters, making them very attractive for generating dynamic city models from space. Motivated by these chances, the authors have earlier proposed approaches that demonstrated first attempts toward reconstruction of building facades from this class of data. The approaches work well when high density of facade points exists, and the full shape of the building could be reconstructed if data are available from multiple views, e.g., from both ascending and descending orbits. However, there are cases when no or only few facade points are available. This usually happens for lower height buildings and renders the detection of facade points/regions very challenging. Moreover, problems related to the visibility of facades mainly facing toward the azimuth direction (i.e., facades orthogonally oriented to the flight direction) can also cause difficulties in deriving the complete structure of individual buildings. These problems motivated us to reconstruct full 2-D/3-D shapes of buildings via exploitation of roof points. In this paper, we present a novel and complete data-driven framework for the automatic (parametric) reconstruction of 2-D/3-D building shapes (or footprints) using unstructured TomoSAR point clouds particularly generated from one viewing angle only. The proposed approach is illustrated and validated by examples using TomoSAR point clouds generated using TerraSAR-X high-resolution spotlight data stacks acquired from ascending orbit covering two different test areas, with one containing simple moderate-sized buildings in Las Vegas, USA and the other containing relatively complex building structures in Berlin, Germany.
Muhammad Shahzad 0002, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.2
2016 Robust Estimators for Multipass SAR Interferometry
abstract
This paper introduces a framework for robust parameter estimation in multipass interferometric synthetic aperture radar (InSAR), such as persistent scatterer interferometry, SAR tomography, small baseline subset, and SqueeSAR. These techniques involve estimation of phase history parameters with or without covariance matrix estimation. Typically, their optimal estimators are derived on the assumption of stationary complex Gaussian-distributed observations. However, their statistical robustness has not been addressed with respect to observations with nonergodic and non-Gaussian multivariate distributions. The proposed robust InSAR optimization (RIO) framework answers two fundamental questions in multipass InSAR: 1) how to optimally treat images with a large phase error, e.g., due to unmolded motion phase, uncompensated atmospheric phase, etc.; and 2) how to estimate the covariance matrix of a non-Gaussian complex InSAR multivariate, particularly those with nonstationary phase signals. For the former question, RIO employs a robust M-estimator to effectively downweight these images; and for the latter, we propose a new method, i.e., the rank M-estimator, which is robust against non-Gaussian distribution. Furthermore, it can work without the assumption of sample stationarity, which is a topic that has not previously been addressed. We demonstrate the advantages of the proposed framework for data with large phase error and heavily tailed distribution, by comparing it with state-of-the-art estimators for persistent and distributed scatterers. Substantial improvement can be achieved in terms of the variance of estimates. The proposed framework can be easily extended to other multipass InSAR techniques, particularly to those where covariance matrix estimation is vital.
Yuanyuan Wang 0002, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.2
2016 Exploiting Joint Sparsity for Pansharpening: The J-SparseFI Algorithm
abstract
Recently, sparse signal representation of image patches has been explored to solve the pansharpening problem. Although these proposed sparse-reconstruction-based methods lead to promising results, three issues remained unsolved: 1) high computational cost; 2) no consideration given to the possibility of mutually correlated information in different multispectral channels; and 3) requirement that the spectral responses of the panchromatic (Pan) image and the multispectral image cover the same wavelength range, which is not necessarily valid for most sensors. In this paper, we propose a sophisticated sparse image fusion algorithm, which is named “jointly sparse fusion of images” (J-SparseFI). It is based on the earlier proposed sparse fusion of images (SparseFI) algorithm and overcomes the aforementioned three drawbacks of the existing sparse image fusion algorithms. The computational problem is handled by reducing the problem size and by proposing a fully parallelizable scheme. Moreover, J-SparseFI exploits the possible signal structure correlations between multispectral channels by introducing the joint sparsity model (JSM) and sharpening the highly correlated adjacent multispectral channels together. This is done by exploiting the distributed compressive sensing theory that restricts the solution of an underdetermined system by considering an ensemble of signals being jointly sparse. J-SparseFI also offers a practical solution to overcome spectral range mismatch between the Pan and multispectral images. By means of sensor spectral response and channel mutual correlation analysis, the multispectral channels are assigned to primary groups of joint channels, secondary groups of joint channels, and individual channels. Primary groups of joint channels, individual channels, and secondary groups of joint channels are then reconstructed sequentially, by the JSM or by modified SparseFI, using a dictionary trained from the Pan image or previously reconstructed high-resolution multispectral channels. A recipe of how to choose appropriate algorithm parameters, including the most crucial regularization parameter, is provided. The algorithm is evaluated and validated using WorldView-2-like images that are simulated using very high resolution airborne HySpex hyperspectral imagery and further practically demonstrated using real WorldView-2 images. The algorithm's performance is compared with other state-of-the-art methods. Visual and quantitative analyses demonstrate the high quality of the proposed method. In particular, the analysis of the difference images suggests that J-SparseFI is superior in image resolution recovery.
Xiao Xiang Zhu 0001, Claas Grohnfeldt, Richard Bamler
IEEE Trans. Geosci. Remote. Sens.1
2016 Geodetic SAR Tomography
abstract
In this paper, we propose a framework referred to as “geodetic synthetic aperture radar (SAR) tomography” that fuses the SAR imaging geodesy and tomographic SAR inversion (TomoSAR) approaches to obtain absolute 3-D positions of a large amount of natural scatterers. The methodology is applied on four very high resolution TerraSAR-X spotlight image stacks acquired over the city of Berlin. Since all the TomoSAR estimates are relative to the same reference point object whose absolute 3-D positions are retrieved by means of stereo SAR, the point clouds reconstructed using data acquired from different viewing angles can be geodetically fused. To assess the accuracy of the position estimates, the resulting absolute shadow-free 3-D TomoSAR point clouds are compared with a digital surface model obtained by airborne LiDAR. It is demonstrated that an absolute positioning accuracy of around 20 cm and a meter-order relative positioning accuracy can be achieved by the proposed framework using TerraSAR-X data.
Xiao Xiang Zhu 0001, Sina Montazeri, Christoph Gisinger, Ramon F. Hanssen, Richard Bamler
IEEE Trans. Geosci. Remote. Sens.1
2015 Region growing based nonlocal filtering for InSAR
abstract
This paper proposes a nonlocal filter variant that replaces the conventional static search window of nonlocal InSAR filters with an adaptive region growing based search window. The region growing approach has the allure that it preselects only similar pixels for the averaging process and that it may find a larger number of statistically homogeneous pixels than a traditional, fixed search window. A Monte-Carlo simulation shows the possible benefits that could be realized with the region growing approach for InSAR filtering. The proposed method is also experimentally evaluated for rural and urban test sites.
Gerald Baier, Xiao Xiang Zhu 0001
IGARSS2
2015 Sparse pixel-wise spectral unmixing - Which algorithm to use and how to improve the results
abstract
Recently, many sparse approximation methods have been applied to solve spectral unmixing problems. These methods in contrast to traditional methods for spectral unmixing are designed to work with large a-prori given spectral dictionaries containing hundreds of labelled material spectra enabling to skip the expensive endmember extraction and labelling step. However, it has been shown that sparse approximation methods sometimes have problems with selection of correct spectra from the dictionary when these are similar. In this paper we study the detection and approximation accuracy of different sparse approximation methods as well as the influence of the proposed modifications.
Jakub Bieniarz, Rupert Müller, Xiao Xiang Zhu 0001, Peter Reinartz
IGARSS3
2015 Towards a combined sparse representation and unmixing based hybrid hyperspectral resolution enhancement method
abstract
The fusion of hyperspectral data with a corresponding higher resolution multispectral image has become an increasingly active research field. The goal is to create a hyperspectral image that has the spatial resolution of the multispectral image. This work aims at combining two established fusion algorithms, namely J-SparseFI-HM and CNMF, to a new method which features their individual advantages. The sparse representation based J-SparseFI-HM algorithm is used to pre-process those hyperspectral channels that have a strong spectral overlap with the multispectral instrument. Then, three modified versions of the matrix factorization and unmixing based CNMF method are used for post-processing. The results are assessed and compared to the individual products of J-SparseFI-HM and CNMF, revealing a great potential for performance improvement.
Claas Grohnfeldt, Xiao Xiang Zhu 0001
IGARSS2
2015 Compressive sensing for neutrospheric water vapor tomography using GNSS and InSAR observations
abstract
This paper presents the innovative Compressive Sensing (CS) concept for tomographic reconstruction of 3D neutrospheric water vapor fields using data from Global Navigation Satellite Systems (GNSS) and Interferometric Synthetic Aperture Radar (InSAR). The Precipitable Water Vapor (PWV) input data are derived from simulations of the Weather Research and Forecasting modeling system. We apply a Compressive Sensing based approach for tomographic inversion. Using the Cosine transform, a sparse representation of the water vapor field is obtained. The new aspects of this work include both the combination of GNSS and InSAR data for water vapor tomography and the sophisticated CS estimation: The combination of GNSS and InSAR data shows a significant improvement in 3D water vapor reconstruction; and the CS estimation produces better results than a traditional Tikhonov regulari-zation with l2norm penalty term.
Marion Heublein, Xiao Xiang Zhu 0001, Fadwa Alshawaf, Michael Mayer, Richard Bamler, Stefan Hinz
IGARSS2
2015 Automatic coastline detection in non-locally filtered tandem-X data
abstract
The detection of coastlines in SAR imagery has been studied for more than two decades now. Whereas the first works were based on the exploitation of amplitude imagery and the corresponding need to deal with speckle noise [1, 2, 3], with the ERS-1/2 tandem configuration also coherence maps started to be used as input. Based on the insights gained on these experiments, later the authors began to exploit both amplitude and coherence imagery simultaneously, finally giving way to the first approach using the original complex SAR data for statistically motivated coastline extraction.
Michael Schmitt 0003, Lingyun Wei, Xiao Xiang Zhu 0001
IGARSS3
2015 Semantic interpretation of InSAR point clouds
abstract
This paper presents a step towards a better interpretation of the scattering mechanism of different objects and their deformation histories in SAR interferometry (InSAR). The proposed technique traces individual SAR scatterer in high resolution optical images where their geometries, materials, and other properties can be better analyzed and classified. And hence scatterers of a same object can be analyzed in group, which brings us to a new level of InSAR deformation monitoring.
Yuanyuan Wang 0002, Xiao Xiang Zhu 0001
IGARSS2
2015 Exploiting sparsity in remote sensing and earth observation: Theory, applications and future trends
abstract
Sparse signals are commonly expected in remote sensing and Earth observation. Along with the significant development of the compressive sensing theory, exploitation of sparsity in remote sensing became a very relevant and active field. Breakthroughs are brought in different remote sensing problems covering synthetic aperture radar, multispectral and hyperspectral image analysis, and LiDAR. Tailored to this special session, this tutorial gives a review, to the best knowledge of the session chair, on recent advances in sparsity exploitation in remote sensing and Earth observation, regarding the theory, applications and future trends.
Xiao Xiang Zhu 0001, Richard Bamler
IGARSS1
2015 Joint Sparsity Model for Multilook Hyperspectral Image Unmixing
abstract
Recent work on hyperspectral image (HSI) unmixing has addressed the use of overcomplete dictionaries by employing sparse models. In essence, this approach exploits the fact that HSI pixels can be associated with a small number of constituent pure materials. However, unlike traditional least-squares-based methods, sparsity-based techniques do not require a preselection of endmembers and are thus able to simultaneously estimate the underlying active materials along with their respective abundances. In addition, this perspective has been extended so as to exploit the spatial homogeneity of abundance vectors. As a result, these techniques have been reported to provide improved estimation accuracy. In this letter, we present an alternative approach that is able to relax, yet exploit, the assumption of spatial homogeneity by introducing a model that captures both similarities and differences between neighboring abundances. In order to validate this approach, we analyze our model using simulated as well as real hyperspectral data acquired by the HyMap sensor.
Jakub Bieniarz, Esteban Aguilera, Xiao Xiang Zhu 0001, Rupert Müller, Peter Reinartz
IEEE Geosci. Remote. Sens. Lett.3
2015 Precise Three-Dimensional Stereo Localization of Corner Reflectors and Persistent Scatterers With TerraSAR-X
abstract
This paper reports on the absolute 3-D localization of radar corner reflectors and persistent scatterers (PSs) by stereo synthetic aperture radar (SAR) carried out with TerraSAR-X and TanDEM-X. The novel combination of rigorously linearized range-Doppler equations with thorough sensor modeling, consideration of atmospheric delays, geodynamic signals, satellite dynamics, and geometrical calibration allows a direct target localization at the centimeter level. Therefore, there is no need for the application of geocoding since our approach delivers 3-D positions directly in the International Terrestrial Reference Frame (ITRF) 2008. Four radar corner reflectors located at the Wettzell Geodetic Observatory, Metsähovi Geodetic Observatory, and German Antarctic Receiving Station O'Higgins were localized in 3-D with a precision better than 4 cm. By introducing an updated geometrical calibration of TerraSAR-X and TanDEM-X, the absolute accuracy of the same level becomes possible, which was demonstrated by externally validating the reflector positions at Metsähovi and Wettzell against results of terrestrial geodetic surveying. Furthermore, PSs located in Berlin were retrieved with a precision of 10 cm, making the method a suitable tool for radar “positioning” in urban environments. Potential fields of application for our approach are the joint analysis with phase-based methods, geometrical calibration of radar satellites, and direct integration of SAR observations into Global Navigation Satellite System networks.
Christoph Gisinger, Ulrich Balss, Roland Pail, Xiao Xiang Zhu 0001, Sina Montazeri, Stefan Gernhardt, Michael Eineder
IEEE Trans. Geosci. Remote. Sens.4
2015 Robust Reconstruction of Building Facades for Large Areas Using Spaceborne TomoSAR Point Clouds
abstract
With data provided by modern meter-resolution synthetic aperture radar (SAR) sensors and advanced multipass interferometric techniques such as tomographic SAR inversion (TomoSAR), it is now possible to reconstruct the shape and monitor the undergoing motion of urban infrastructures on the scale of centimeters or even millimeters from space in very high level of details. The retrieval of rich information allows us to take a step further toward generation of 4-D (or even higher dimensional) dynamic city models, i.e., city models that can incorporate temporal (motion) behavior along with the 3-D information. Motivated by these opportunities, the authors proposed an approach that first attempts to reconstruct facades from this class of data. The approach works well for small areas containing only a couple of buildings. However, towards automatic reconstruction for the whole city area, a more robust and fully automatic approach is needed. In this paper, we present a complete extended approach for automatic (parametric) reconstruction of building facades from 4-D TomoSAR point cloud data and put particular focus on robust reconstruction of large areas. The proposed approach is illustrated and validated by examples using TomoSAR point clouds generated from a stack of TerraSAR-X high-resolution spotlight images from ascending orbit covering an approximately 2- km2high-rise area in the city of Las Vegas.
Muhammad Shahzad 0002, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.2
2014 Hyperspectral image resolution enhancement based on joint sparsity spectral unmixing
abstract
Relatively low spatial resolution of the space-borne hyper-spectral images (HSI) is the main drawback to derive value added products. Recently, several techniques have been proposed in order to enhance the spatial resolution HSI by means of fusion with higher spatial resolution multispectral images. This paper presents an alternative approach based on the joint sparsity model for spectral unmixing with the use of a-priori spectral dictionary. To assess the results, we compare our algorithm with the state of the art methods.
Jakub Bieniarz, Rupert Müller, Xiao Xiang Zhu 0001, Peter Reinartz
IGARSS3
2014 Automatic large area reconstruction of building façades from spaceborne TomoSAR point clouds
abstract
Improved resolution of SAR sensors and advanced multipass interferometric techniques, such as tomographic SAR inversion (TomoSAR), opens up new possibilities of 4D (or even higher dimensional) imaging that can be potentially used to reconstruct dynamic models of entire cities, i.e., city models that can incorporate temporal (motion) behaviour along with the 3D information. Motivated by these chances, this paper presents an approach that systematically allows automatic reconstruction of building façades from 4D point cloud generated from tomographic SAR processing and put particular focus on robust reconstruction of large areas. The approach is modular and is illustrated/validated by examples using TomoSAR point clouds generated from a stack of TerraSAR-X high-resolution spotlight images from ascending orbit covering approx. 2 km2high rise area in the city of Las Vegas.
Muhammad Shahzad 0002, Xiao Xiang Zhu 0001
IGARSS2
2014 Beyond the 12m TanDEM-X DEM
abstract
The standard TanDEM-X product meats HRTI-3 DEM specification and comes with a sample spacing of 12 m. We apply non-local means (NL) interferogram filtering to the TanDEM-X data. In this paper, we present modifications of the original NL filter which render it more appropriate and efficient for massive processing of TanDEM-X data. Further, we investigate the noise reduction properties as well as the resolution and the coherence estimation accuracy of the new NL filter. Simulations and tests with TanDEM-X data hint that the improved DEMs possess a quality close to the HRTI-4 standard. Also future global InSAR missions like Tandem-L will greatly benefit from this type of filters.
Xiao Xiang Zhu 0001, Marie Lachaise, Fathalrahman Adam, Yilei Shi, Michael Eineder, Richard Bamler
IGARSS1
2014 Geodetic TomoSAR - Fusion of SAR imaging geodesy and TomoSAR for 3D absolute scatterer positioning
abstract
In this paper, we propose a framework referred to as “geodetic TomoSAR“ that fuses the SAR image geodesy and TomoSAR approaches to obtain absolute 3D positions of a large amount of natural scatterers. The methodology is applied on four Very High Resolution (VHR) TerraSAR-X spotlight image stacks acquired over the city of Berlin. Since the TomoSAR estimates are referred to the identical reference point whose absolute 3D positions are retrieved by means of Stereo-SAR, the point clouds from ascending and descending orbits are automatically fused. To assess the accuracy of the position estimates, the resulting absolute shadow-free 3D TomoSAR point clouds are compared to a DSM obtained by airborne LiDAR.
Xiao Xiang Zhu 0001, Sina Montazeri, Christoph Gisinger, Ramon F. Hanssen, Richard Bamler
IGARSS1
2014 An Efficient Tomographic Inversion Approach for Urban Mapping Using Meter Resolution SAR Image Stacks
abstract
This letter describes an efficient approach of multidimensional synthetic aperture radar (SAR) imaging for urban mapping. The proposed approach is an integration of tomographic SAR inversion and the well-known persistent scatterer interferometry (PSI). It consists of three steps: first, a global estimation of the topography and motion parameters using efficient algorithms such as PSI; second, a single and double scatterer discrimination step based on the results of the first step; finally, a tomographic SAR inversion, which is performed on the preclassified double scatterers, using the prior knowledge obtained in the first step, retrieving the topography and motion parameters of both scatterers. The proposed approach has been tested on a dozen of TerraSAR-X high-resolution spotlight image stacks. In this letter, examples from Las Vegas and Berlin are presented. The results are comparable with the one obtained by the most computationally expensive tomographic SAR algorithms (e.g., SLIMMER) only and saves computational time by a factor of 50.
Yuanyuan Wang 0002, Xiao Xiang Zhu 0001, Richard Bamler
IEEE Geosci. Remote. Sens. Lett.2
2014 Facade Reconstruction Using Multiview Spaceborne TomoSAR Point Clouds
abstract
Recent advances in very high resolution tomographic synthetic aperture radar inversion (TomoSAR) using multiple data stacks from different viewing angles enables us to generate 4-D (space-time) point clouds of the illuminated area from space with a point density comparable to LiDAR. They can be potentially used for facade reconstruction and deformation monitoring in urban environment. In this paper, we present the first attempt to reconstruct facades from this class of data: First, the facade region is extracted using the density estimates of the points projected to the ground plane, the extracted facade points are then clustered into individual facades by means of orientation analysis, surface (flat or curved) model parameters of the segmented building facades are further estimated, and the geometric primitives such as intersection points of the adjacent facades are determined to complete the reconstruction process. The proposed approach is illustrated and validated by examples using TomoSAR point clouds generated from stacks of TerraSAR-X high-resolution spotlight images from two viewing angles, i.e., both ascending and descending orbits. The performance of the proposed approach is systematically analyzed. To explore the possible applications, we refine the elevation estimate of each raw TomoSAR point by using its more accurate azimuth and range coordinates and the corresponding reconstructed building facade model. Compared to the raw TomoSAR point clouds, significantly improved elevation positioning accuracy is achieved. Finally, a first example of the reconstructed 4-D city model is presented.
Xiao Xiang Zhu 0001, Muhammad Shahzad 0002
IEEE Trans. Geosci. Remote. Sens.1
2013 Jointly sparse fusion of hyperspectral and multispectral imagery
abstract
In this paper we apply the recently proposed J-SparseFI data fusion method to the fusion of a low-resolution hyperspectral image and a high-resolution multispectral image. The high correlation of signals in adjacent hyperspectral channels is exploited by assuming signals in different channels are jointly sparse in suitable dictionaries that are created from the multispectral image. First experimental results using airborne HySpex hyperspectral data and synthesized WorldView-2 imagery are presented.
Claas Grohnfeldt, Xiao Xiang Zhu 0001, Richard Bamler
IGARSS2
2013 Reconstruction of building façades using spaceborne multiview TomoSAR point clouds
abstract
In this paper we present an approach that allows automatic reconstruction of building façades from 4D point cloud generated from tomographic SAR processing. The approach is modular and works by extracting façade points from the point density projected onto the ground plane. Individual façades are segmented using an unsupervised clustering procedure. Surface (flat or curved) model parameters of the segmented building façades are further estimated and finally the geometric primitives such as intersection points of the adjacent façades are determined to complete the reconstruction process. The proposed approach is illustrated and validated by examples using TomoSAR point clouds generated from TerraSAR-X high resolution spotlight images.
Muhammad Shahzad 0002, Xiao Xiang Zhu 0001
IGARSS2
2013 Feature-based fusion of tomosar point clouds from multiview TerraSAR-X data stacks
abstract
This article presents a technique of fusing point clouds from multiple view angles generated using synthetic aperture radar (SAR) tomography. Using TerraSAR-X high resolution spotlight data stacks, one such point has a population of about 2×107points, with a density of around 106points / km2. Such large point population leads to a high computational cost while doing the fusion in 3D space. Therefore, we introduce a feature-based unsupervised technique for point clouds fusion by detecting and matching building contour end points and aligning flat roofs in the two point clouds. The same idea can also be exploited as a general way to evaluate the fusion accuracy of other fusion techniques.
Yuanyuan Wang 0002, Xiao Xiang Zhu 0001
IGARSS2
2013 Collaborative sparse reconstruction for pan-sharpening
abstract
In this paper, we extend the Sparse Fusion of Images (SparseFI, pronounced “sparsify”) algorithm, proposed by the authors before, to a Jointly Sparse Fusion of Images (J-SparseFI) algorithm by exploiting the possible signal structural correlations between different multispectral channels. The algorithm is evaluated using airborne UltraCam data. The superior performance of the proposed methods has been demonstrated by a statistic assessment. Moreover, first experimental result using airborne HySpex hyperspectral data is presented.
Xiao Xiang Zhu 0001, Claas Grohnfeldt, Richard Bamler
IGARSS1
2013 A Sparse Image Fusion Algorithm With Application to Pan-Sharpening
abstract
Data provided by most optical Earth observation satellites such as IKONOS, QuickBird, and GeoEye are composed of a panchromatic channel of high spatial resolution (HR) and several multispectral channels at a lower spatial resolution (LR). The fusion of an HR panchromatic and the corresponding LR spectral channels is called “pan-sharpening.” It aims at obtaining an HR multispectral image. In this paper, we propose a new pan-sharpening method named Sparse F usion of Images (SparseFI, pronounced as “sparsify”). SparseFI is based on the compressive sensing theory and explores the sparse representation of HR/LR multispectral image patches in the dictionary pairs cotrained from the panchromatic image and its downsampled LR version. Compared with conventional methods, it “learns” from, i.e., adapts itself to, the data and has generally better performance than existing methods. Due to the fact that the SparseFI method does not assume any spectral composition model of the panchromatic image and due to the super-resolution capability and robustness of sparse signal reconstruction algorithms, it gives higher spatial resolution and, in most cases, less spectral distortion compared with the conventional methods.
Xiao Xiang Zhu 0001, Richard Bamler
IEEE Trans. Geosci. Remote. Sens.1
2012 Façade structure reconstruction using spaceborne TomoSAR point clouds
abstract
Very high resolution SAR tomography using multiple data stacks from different viewing angles enables us for the first time to generate 4D point clouds of the illuminated area from space with a point density comparable to LiDAR. They can be potentially used for façade reconstruction and monitoring in urban environment. In this paper, we propose an approach for façade detection and reconstruction from such point clouds. Firstly, the façade region is extracted by thresholding the point density on the ground plane. The extracted façades points are then clustered into segments corresponding to individual façades by means of slope analysis. Surface (flat or curved) model parameters of the segmented building façades are further estimated. Finally, the elevation estimates of each raw TomoSAR point is refined by using its more accurate azimuth and range coordinates, and the corresponding reconstructed surface model of the façade. The proposed approach is illustrated and validated by examples using TomoSAR point clouds generated from a stack of 25 TerraSAR-X high spotlight images.
Muhammad Shahzad 0002, Xiao Xiang Zhu 0001, Richard Bamler
IGARSS2
2012 Operational TomoSAR processing using TerraSAR-X high resolution spotlight stacks from multiple view angles
abstract
With the availability of meter resolution space-borne SAR systems, urban monitoring using SAR Tomography (TomoSAR) becomes increasingly popular, because of its layover separation capability. However, compared to Persistent Scatterer Interferometry (PSI), TomoSAR applications are much more computationally expensive. This article introduces a TomoSAR processing system for long-term large urban area mapping and monitoring. Two new features were introduced: 1. PSI was integrated into the currently available TomoSAR algorithms (e.g. TSVD, SVD-Wiener, and SL1MMER) to increase the overall computational efficiency; 2. results from multiple view angles were fused to provide full coverage of each building façade. This processing system handles an entire TerraSAR-X high resolution spotlight scene in an affordable time, achieving scatterer density up to 1 million scatterer/km2from a single stack, comparing to 60 thousand to 100 thousand scatterer/km2for PSI.
Yuanyuan Wang 0002, Xiao Xiang Zhu 0001, Yilei Shi, Richard Bamler
IGARSS2
2012 Sparse tomographic SAR reconstruction from mixed TerraSAR-X/TanDEM-X data stacks
abstract
This paper and the conference presentation present the first demonstration of high precision very high resolution tomographic SAR inversion with the assistance of TanDEM-X data. The data quality of TerraSAR-X and TanDEM-X is investigated. TomoSAR algorithms such as SVD-Wiener, Nonlinear Least Squares and SLIMMER are extended for mixed repeat- and single-pass data stacks. A systematic approach is proposed for the fusion of TerraSAR-X and TanDEM-X data in which the different data quality provided by the TerraSAR-X and TanDEM-X data are taken into account by introducing a weighting according to the noise covariance matrix. The proposed approach is evaluated with simulated data. The simulation result shows that the reconstruction accuracy of tomographic SAR inversion can be improved significantly by using jointly fused TerraSAR-X and TanDEM-X data.
Xiao Xiang Zhu 0001, Richard Bamler
IGARSS1
2012 Super-Resolution Power and Robustness of Compressive Sensing for Spectral Estimation With Application to Spaceborne Tomographic SAR
abstract
We address the problem of resolving two closely spaced complex-valued points from N irregular Fourier do- main samples. Although this is a generic super-resolution (SR) problem, our target application is SAR tomography (TomoSAR), where typically the number of acquisitions is N = 10 - 100 and SNR = 0-10 dB. As the TomoSAR algorithm, we introduce "Scale-down by LI norm Minimization, Model selection, and Estimation Reconstruction" (SL1MMER), which is a spectral estimation algorithm based on compressive sensing, model order selection, and final maximum likelihood parameter estimation. We investigate the limits of SLIMMER concerning the following questions. How accurately can the positions of two closely spaced scatterers be estimated? What is the closest distance of two scat- terers such that they can be separated with a detection rate of 50% by assuming a uniformly distributed phase difference? How many acquisitions N are required for a robust estimation (i.e., for separating two scatterers spaced by one Rayleigh resolution unit with a probability of 90%)? For all of these questions, we provide numerical results, simulations, and analytical approxima- tions. Although we take TomoSAR as the preferred application, the SLIMMER algorithm and our results on SR are generally applicable to sparse spectral estimation, including SR SAR focus- ing of point-like objects. Our results are approximately applicable to nonlinear least-squares estimation, and hence, although it is derived experimentally, they can be considered as a fundamental bound for SR of spectral estimators. We show that SR factors are in the range of 1.5-25 for the aforementioned parameter ranges of N and SNR.
Xiao Xiang Zhu 0001, Richard Bamler
IEEE Trans. Geosci. Remote. Sens.1
2012 Demonstration of Super-Resolution for Tomographic SAR Imaging in Urban Environment
abstract
Tomographic synthetic aperture radar (SAR) inversion, including SAR tomography and differential SAR tomography, is essentially a spectral analysis problem. The resolution in the elevation direction depends on the elevation aperture size, i.e., on the spread of orbit tracks. Since the orbits of modern meter-resolution spaceborne SAR systems, such as TerraSAR-X, are tightly controlled, the tomographic elevation resolution is at least an order of magnitude lower than in range and azimuth. Hence, super-resolution (SR) reconstruction algorithms are desired. Considering the sparsity of the signal in elevation, a compressive sensing based super-resolving algorithm, named “Scale-down by L1norm Minimization, Model selection, and Estimation Reconstruction” (SL1MMER, pronounced “slimmer”), was proposed by the authors in a previous paper. The ultimate bounds of the technique on localization accuracy and SR power were investigated. In this paper, the essential role of SR for layover separation in urban infrastructure monitoring is indicated by geometric and statistical analysis. It is shown that double scatterers with small elevation distances are more frequent than those with large elevation distances. Furthermore, the SR capability of SL1MMER is demonstrated using TerraSAR-X real data examples. For a high rise building complex, the percentage of detected double scatterers is almost doubled compared to classical linear estimators. Among them, half of the detected double scatterer pairs have elevation distances smaller than the Rayleigh elevation resolution. This confirms the importance of SR for this type of applications.
Xiao Xiang Zhu 0001, Richard Bamler
IEEE Trans. Geosci. Remote. Sens.1
2011 Optimal estimation of distributed scatterer phase history parameters from meter-resolution SAR data
abstract
Measuring the long-term line-of-sight deformation using a multi pass SAR data stack by standard persistent scatterer technique has been explored since the late 1990s. Researches have been continuously conducted on increasing the data coverage at non-PS rich areas. The recently developed SqueeSAR™ technique has validated the potential of extracting useful information from distributed scatterers. With the availability of high resolution TerraSAR-X spotlight data, this technique can benefit greatly from its higher data density and quality. This article presents an algorithm of parameter estimation at distributed scatterers by maximum likelihood estimator in high resolution TS-X spotlight data. Different to SqueeSAR™, this article pays particular attention to the accurate covariance matrix estimation for phase history retrieval on each individual distributed scatterer. Solutions are presented for adaptive sample selection by a different statistical test (Anderson-Darling). An adaptive multi resolution defringe algorithm is introduced to cope with the problem of accurate fringe removal and in turn accurate covariance matrix estimation. And finally maximum likelihood estimator was employed to estimate the model parameters by weighting each measurement according to its coherence. By combining both the persistent scatterers and distributed scatterers, the increase in the capable monitoring area is phenomenal.
Yuanyuan Wang 0002, Xiao Xiang Zhu 0001, Richard Bamler
IGARSS2
2011 Within the resolution cell: Super-resolution in tomographic SAR imaging
abstract
We address the problem of resolving two closely spaced complex-valued points from N irregular Fourier domain samples. Although this is a generic super-resolution problem, our target application is SAR tomography where typically the number of acquisitions is N = 10... 100 and SNR = 0 ...10dB. In this paper, a compressive sensing based algorithm is introduced for tomographic SAR inversion. It is named "Scale-down by LI. norm Minimization. Model selection, and Estimation Reconstruction" (SLIMMER, pronounced "slimmer"). SLIMMER combines the advantage of compressive sensing, e.g. high localization accuracy and super-resolution, and the radiometric accuracy of the linear estimator. Moreover, a systematic performance assessment of the SLIMMER algorithm is carried out regarding the elevation estimation accuracy and super-resolution. It is proven that SLIMMER is an efficient estimator; its super resolution factors are in the range of 1.5 to 25 for the aforementioned parameter ranges of TV and SNR. Our results are approximately applicable to nonlinear least-squares estimation, and hence can be considered as fundamental bounds for super-resolution of spectral estimators.
Xiao Xiang Zhu 0001, Richard Bamler
IGARSS1
2011 Multi-component nonlinear motion estimation in differential SAR tomography - the time-warp method
abstract
In the differential SAR Tomography (D-TomoSAR) system model the motion history appears as a phase term. In the case of nonlinear motion this phase term is no longer linear, and hence can not be retrieved by spectral estimation. We propose the "time warp" method that rearranges the acquisition dates such that a linear motion is pretended. The multi-component generalization of time warp rewrites the D-TomoSAR system model to an M+l-dimensional standard spectral estimation problem, where M indicates the user defined motion model order, and hence enables the motion estimation for all possible complex motion models. Both simulations and real data (from TerraSAR-X spotlight) examples demonstrate the applicability of the method and show that linear and periodic (seasonal) motion components can be separated and retrieved.
Xiao Xiang Zhu 0001, Richard Bamler
IGARSS1
2011 Compressive sensing for image fusion - with application to pan-sharpening
abstract
Data provided by most optic earth observation satellites such as IKONOS, Quick Bird and GeoEye are composed of a panchromatic channel of high spatial resolution (HR) and several multispectral channels at a lower spatial resolution (LR). The fusion of a HR panchromatic and the corresponding LR spectral channels is called "pan-sharpening". It aims at obtaining a high resolution multispectral image. In this paper, we propose a new sophisticated pan-sharpening method named Sparse Fusion of Images (SparseFI, pronounced as sparsify). SparseFI is based on the compressive sensing theory and explore the sparse representation of HR/LR multispectral image patches in the dictionaries pairs co-trained from the panchromatic image and its corresponding down-sampled version. Compared to other methods it "learns" from, i.e. adapts itself to, the data and has better performance than existing methods. Due to the fact that the SparseFI algorithm does not assume any model of the panchromatic image and thanks to the super-resolution capability and robustness of compressive sensing, it gives higher spatial and spectral resolution with less spectral distortion compared to the conventional methods.
Xiao Xiang Zhu 0001, Richard Bamler
IGARSS1
2011 Tomographic Imaging and Monitoring of Buildings With Very High Resolution SAR Data
abstract
Layover is frequent in imaging and monitoring with synthetic aperture radar (SAR) areas characterized by a high density of scatterers with steep topography, e.g., in urban environment. Using medium-resolution SAR data tomographic techniques has been proven to be capable of separating multiple scatterers interfering (in layover) in the same pixel. With the advent of the new generation of high-resolution sensors, the layover effect on buildings becomes more evident. In this letter, we exploit the potential of the 4-D imaging applied to a set of TerraSAR-X spotlight acquisitions. Results show that the combination of high-resolution data and advanced coherent processing techniques can lead to impressive reconstruction and monitoring capabilities of the whole 3-D structure of buildings.
Diego Reale, Gianfranco Fornaro, Antonio Pauciullo, Xiao Xiang Zhu 0001, Richard Bamler
IEEE Geosci. Remote. Sens. Lett.4
2011 Let's Do the Time Warp: Multicomponent Nonlinear Motion Estimation in Differential SAR Tomography
abstract
In the differential synthetic aperture radar tomography (D-TomoSAR) system model, the motion history appears as a phase term. In the case of nonlinear motion, this phase term is no longer linear and, hence, cannot be retrieved by spectral estimation. We propose the “time warp” method that rearranges the acquisition dates such that a linear motion is pretended. The multicomponent generalization of time warp rewrites the D-TomoSAR system model to an (M+ 1)-dimensional standard spectral estimation problem, whereMindicates the user-defined motion model order and, hence, enables the motion estimation for all possible complex motion models. Both simulations and real data (from TerraSAR-X spotlight) examples demonstrate the applicability of the method and show that linear and periodic (seasonal) motion components can be separated and retrieved.
Xiao Xiang Zhu 0001, Richard Bamler
IEEE Geosci. Remote. Sens. Lett.1
2010 Advanced techniques and new high resolution SAR sensors for monitoring urban areas
abstract
In the last years MultiDimensional (3D and 4D) Synthetic Aperture Radar (SAR) techniques, also known as SAR tomography and differential SAR tomography, are emerging in the field of coherent combination of multibaseline/multitemporal SAR data. With respect to the classical differential interferometric processing, these techniques improve the capability of detection and monitoring of the ground targets. Moreover they were proven to be effective in resolving the signal interference due to the layover effect, that may occur in areas with high density of scatterers located on vertical structures, such as urban areas. Beside the development of these advanced techniques the new generation of sensor, such as TerraSAR-X and COSMO-SKYMED with very high spatial resolution offer new perspectives in the imaging and monitoring of urban areas. In this paper we address the application of the SAR tomography to real spaceborne data. Particularly, we show and discuss the first results of the application of this technique to high resolution TerraSAR-X data.
Diego Reale, Gianfranco Fornaro, Gianfranco Pauciullo, Xiao Xiang Zhu 0001, Nico Adam, Richard Bamler
IGARSS4
2010 Compressive sensing for high resolution differential SAR tomography - the SL1MMER algorithm
abstract
Differential SAR tomography extends the synthetic aperture principle into the elevation and time directions for 4-D imaging. With modern meter-resolution space-borne SAR systems like TerraSAR-X (TS-X), systematic tomographic imaging of urban infrastructure and its deformations becomes feasible. We demonstrate the potential of TS-X data for this purpose and introduce several novel concepts. Since building deformation in general is nonlinear, e.g. due to thermal dilation, we start from a tomographic system formulation that is general enough to allow for the inclusion of motion models (linear, periodic, etc.). By appropriate warping of the time axis we map the motion model function to become linear and lead to a peak in the spectral domain. For the differential tomographic inversion itself we propose a 2-D compressive sensing (CS) based approach - “SL1MMER”. We demonstrate the super-resolution power and the robustness of SL1MMER both with simulated and with real data. We also show that it provides an attractive compromise between parametric and non-parametric methods. A full reconstruction of a building complex and its seasonal deformation from a stack of TS-X spotlight data is finally presented.
Xiao Xiang Zhu 0001, Richard Bamler
IGARSS1
2010 Tomographic SAR Inversion by L1 -Norm Regularization - The Compressive Sensing Approach
abstract
Synthetic aperture radar (SAR) tomography (TomoSAR) extends the synthetic aperture principle into the elevation direction for 3-D imaging. The resolution in the elevation direction depends on the size of the elevation aperture, i.e., on the spread of orbit tracks. Since the orbits of modern meter-resolution spaceborne SAR systems, like TerraSAR-X, are tightly controlled, the tomographic elevation resolution is at least an order of magnitude lower than in range and azimuth. Hence, super-resolution reconstruction algorithms are desired. The high anisotropy of the 3-D tomographic resolution element renders the signals sparse in the elevation direction; only a few pointlike reflections are expected per azimuth-range cell. This property suggests using compressive sensing (CS) methods for tomographic reconstruction. This paper presents the theory of 4-D (differential, i.e., space-time) CS TomoSAR and compares it with parametric (nonlinear least squares) and nonparametric (singular value decomposition) reconstruction methods. Super-resolution properties and point localization accuracies are demonstrated using simulations and real data. A CS reconstruction of a building complex from TerraSAR-X spotlight data is presented.
Xiao Xiang Zhu 0001, Richard Bamler
IEEE Trans. Geosci. Remote. Sens.1
2010 Very High Resolution Spaceborne SAR Tomography in Urban Environment
abstract
Synthetic aperture radar tomography (TomoSAR) extends the synthetic aperture principle into the elevation direction for 3-D imaging. It uses stacks of several acquisitions from slightly different viewing angles (the elevation aperture) to reconstruct the reflectivity function along the elevation direction by means of spectral analysis for every azimuth-range pixel. The new class of meter-resolution spaceborne SAR systems (TerraSAR-X and COSMO-Skymed) offers a tremendous improvement in tomographic reconstruction of urban areas and man-made infrastructure. The high resolution fits well to the inherent scale of buildings (floor height, distance of windows, etc.). This paper demonstrates the tomographic potential of these SARs and the achievable quality on the basis of TerraSAR-X spotlight data of urban environment. A new Wiener-type regularization to the singular-value decomposition method-equivalent to a maximum a posteriori estimator-for TomoSAR is introduced and is extended to the differential case (4-D, i.e., space-time). Different model selection schemes for the estimation of the number of scatterers in a resolution cell are compared and proven to be applicable in practice. Two parametric estimation algorithms of the scatterers' elevation and their velocities are evaluated. First 3-D and 4-D reconstructions of an entire building complex (including its radar reflectivity) with very high level of detail from spaceborne SAR data by pixelwise TomoSAR are presented.
Xiao Xiang Zhu 0001, Richard Bamler
IEEE Trans. Geosci. Remote. Sens.1
2009 Techniques and Examples for the 3D Reconstruction of Complex Scattering Situations using TerraSAR-X
abstract
The German radar satellite TerraSAR-X was launched in June 2007. Since then, it is continuously providing high resolution space-borne radar data which are perfectly suitable for sophisticated interferometric applications. I.e. the mission concept and the SAR sensor support the coherent stacking of radar scenes which is the basis for advanced processing techniques e.g. Persistent Scatterer Interferometry (PSI) and SAR tomography. In particular, the short repeat cycle of eleven days and the highly reproducible scene repetition of the spotlight acquisitions support the stacking and consequently the time series analysis of the radar data. Furthermore, the sensor's orbital tube is precisely controlled to be in the order of 200 m which basically allows to utilize the baseline spread of the stacked acquisitions. However, this small spread is actually limiting the resolution in the SAR tomography. Interferometric applications could be demonstrated already in a very early stage of the TerraSAR-X mission. Because the resolution is 0.6 m in slant range and 1.1 m in azimuth in the high resolution spotlight mode the PSI and the SAR tomography processing results were impressive. Urban areas and single buildings could be mapped from space in three dimensions. Even the structural stress of single buildings caused by thermal dilation could be demonstrated. However, extended layover areas are caused by typical buildings and as a consequence complicated scattering situations need to be resolved. DLR's operational In-SAR processing system GENESIS had already been adapted to cope with the new sensor modes of TerraSAR-X and their new specific spectral characteristics. Now, the new image characteristics e.g. the extended layover areas and the long time coherent distributed scatterer need better to be supported. Subject is to optimally exploit the available information e.g. the radar reflectivity. Several algorithms of the processing system can take advantage of this, e.g. the scatterer configuration detection. As a matter of fact, the scatterer configuration has now become a very important characteristic for each resolution cell. It influences e.g. the estimation data extraction, the estimation of the 3D location and basically the estimation precision. A typical resolution cell can be composed of a single dominant point scatterer surrounded by clutter, two or more dominant point scatterers in clutter and of distributed scatterers with a specific phase stability over time. The paper provides technical details and a processing example of a newly developed algorithm to retrieve the 3D location of point scatterers from the scene's intensity which finally also provides the information on the scatterer configuration in a resolution cell.
Nico Adam, Xiao Xiang Zhu 0001, Christian Minet, Werner Liebhart, Michael Eineder, Richard Bamler
IGARSS (3)2
2009 3D Analysis of Scattering Effects based on Ray Tracing Techniques
abstract
The side-looking geometry of SAR sensors hampers the interpretation of SAR images of urban areas. Simulation tools for illuminating 3D models of man-made objects by means of a virtual sensor support the interpretation of scattering effects by providing artificial images in the azimuth-range plane. In this paper, a simulation approach is presented which extends SAR simulation to three dimensions in order to focus detected intensity contributions in azimuth, range and elevation. Based on the simulation output, a concept for creating scatterer histograms displaying the number of scatterers within one resolution cell is introduced. Methods for analyzing simulated elevation data by means of selected slices are presented for an urban test site. Eventually, the number of scatterers extracted for a selected pixel by tomographic analysis, using a stack of spotlight TerraSAR-X images, is confirmed by results provided by the simulator.
Stefan Auer, Xiao Xiang Zhu 0001, Stefan Hinz, Richard Bamler
IGARSS (3)2
2009 Space-borne High resolution Tomographic Interferometry
abstract
SAR tomography (TomoSAR) is a way of overcoming the limitations of standard 2-D imaging by achieving focused 3D images. Its capability of 3-D reflectivity reconstruction and multiple scatterers separation has been demonstrated in the urban environment in absence of significant deformation. In this paper, an example which reveals the distortion of 3-D reconstruction result by coupled deformation information is presented. Differential SAR Tomography (or 4-D SAR focusing), a new interferometric mode crossing differential SAR interferometry and the 3-D multi-baseline tomography concepts, are implemented. In this paper, space-borne high resolution differential SAR tomography is demonstrated using TerraSAR-X high resolution spotlight data. High resolution SAR is proven to be very attractive for city mapping.
Xiao Xiang Zhu 0001, Nico Adam, Richard Bamler
IGARSS (4)1