EDBT 2026 Demo / reviewers in the wild / expert
Lichao Mou
dblp:167/0614
· DBLP profile ↗
97ranked-venue papers
17as first author
61since 2021 · last 2026
0000-0001-8407-6413ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 84 · 16 first-author · 55 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ProPL: Universal Semi-Supervised Ultrasound Image Segmentation via Prompt-Guided Pseudo-LabelingabstractExisting approaches for the problem of ultrasound image segmentation, whether supervised or semi-supervised, are typically specialized for specific anatomical structures or tasks, limiting their practical utility in clinical settings. In this paper, we pioneer the task of universal semi-supervised ultrasound image segmentation and propose ProPL, a framework that can handle multiple organs and segmentation tasks while leveraging both labeled and unlabeled data. At its core, ProPL employs a shared vision encoder coupled with prompt-guided dual decoders, enabling flexible task adaptation through a prompting-upon-decoding mechanism and reliable self-training via an uncertainty-driven pseudo-label calibration (UPLC) module. To facilitate research in this direction, we introduce a comprehensive ultrasound dataset spanning 5 organs and 8 segmentation tasks. Extensive experiments demonstrate that ProPL outperforms state-of-the-art methods across various metrics, establishing a new benchmark for universal ultrasound image segmentation. Yaxiong Chen, Qicong Wang, Jingliang Hu, Yilei Shi, Shengwu Xiong 0001, Xiao Xiang Zhu 0001, Lichao Mou |
AAAI | 8 |
| 2026 | Semantic consistency-aware pseudo-temporal framework for multimodal remote sensing image segmentation
Yuejiang Li, Weisheng Dong, Peng Wu 0015, Lichao Mou, Xin Li 0005 |
Neural Networks | 6 |
| 2025 | Q-PART: Quasi-Periodic Adaptive Regression with Test-time Training for Pediatric Left Ventricular Ejection Fraction RegressionabstractIn this work, we address the challenge of adaptive pediatric Left Ventricular Ejection Fraction (LVEF) assessment. While Test-time Training (TTT) approaches show promise for this task, they suffer from two significant limitations. Existing TTT works are primarily designed for classification tasks rather than continuous value regression, and they lack mechanisms to handle the quasi-periodic nature of cardiac signals. To tackle these issues, we propose a novel Quasi-Periodic Adaptive Regression with Test-time Training (Q-PART) framework. In the training stage, the proposed Quasi-Period Network decomposes the echocardiogram into periodic and aperiodic components within latent space by combining parameterized helix trajectories with Neural Controlled Differential Equations. During inference, our framework further employs a variance minimization strategy across image augmentations that simulate common quality issues in echocardiogram acquisition, along with differential adaptation rates for periodic and aperiodic components. Theoretical analysis is provided to demonstrate that our variance minimization objective effectively bounds the regression error under mild conditions. Furthermore, extensive experiments across three pediatric age groups demonstrate that Q-PART not only significantly outperforms existing approaches in pediatric LVEF prediction, but also exhibits strong clinical screening capability with high mAUROC scores (up to 0.9747) and maintains gender-fair performance across all metrics, validating its robustness and practical utility in pediatric echocardiography analysis. The project can be found in Q-PART. Jie Liu 0044, Tiexin Qin, Hui Liu 0036, Yilei Shi, Lichao Mou, Xiao Xiang Zhu 0001, Shiqi Wang 0001, Haoliang Li |
CVPR | 5 |
| 2025 | Scale-Aware Contrastive Reverse Distillation for Unsupervised Medical Anomaly DetectionabstractUnsupervised anomaly detection using deep learning has garnered significant research attention due to its broad applicability, particularly in medical imaging where labeled anomalous data are scarce. While earlier approaches leverage generative models like autoencoders and generative adversarial networks (GANs), they often fall short due to overgeneralization. Recent methods explore various strategies, including memory banks, normalizing flows, self-supervised learning, and knowledge distillation, to enhance discrimination. Among these, knowledge distillation, particularly reverse distillation, has shown promise. Following this paradigm, we propose a novel scale-aware contrastive reverse distillation model that addresses two key limitations of existing reverse distillation methods: insufficient feature discriminability and inability to handle anomaly scale variations. Specifically, we introduce a contrastive student-teacher learning approach to derive more discriminative representations by generating and exploring out-of-normal distributions. Further, we design a scale adaptation mechanism to softly weight contrastive distillation losses at different scales to account for the scale variation issue. Extensive experiments on benchmark datasets demonstrate state-of-the-art performance, validating the efficacy of the proposed method. The code will be made publicly available. Yilei Shi, Jingliang Hu, Xiao Xiang Zhu 0001, Lichao Mou |
ICLR | 5 |
| 2025 | High-Order Progressive Trajectory Matching for Medical Image Dataset Distillation
Jinghao Bian, Jingyang Hou, Jingliang Hu, Yilei Shi, Weisheng Dong, Xiao Xiang Zhu 0001, Lichao Mou |
MICCAI (14) | 8 |
| 2024 | Evaluating Feature Impact on Ocean Wind Speed Predictions: An Application of Explainable AI to GNSS Reflectometry DataabstractArtificial intelligence (AI) models developed for the Global Navigation Satellite System Reflectometry (GNSS-R) observations are capable of estimating geophysical parameters, especially ocean surface wind speeds. Understanding the decision-making process of deep learning models can be as significant as improving output accuracy in practical applications. This study explores the exploitation of Explainable Artificial Intelligence (XAI) for interpreting complex deep learning models. With the assistance of SHAP (SHapley Additive exPlanations) Gradient Explainer, this study evaluates the impact of both Delay-Doppler Map (DDM) pixels and ancillary parameters on model predictions, which can further help in understanding the role of specific input parameters. Additionally, this study investigates the potential of applying XAI to enhance the accuracy of deep learning models, which can be extended for climate-related applications. Tianqi Xiao, Milad Asgarimehr, Jens Wickert, Daixin Zhao, Lichao Mou, Caroline Arnold |
IGARSS | 5 |
| 2024 | Referring Image Segmentation for Remote Sensing DataabstractIn this paper, we present a new task: referring image segmentation for remote sensing data, which targets segmenting out specific objects referred to by natural language. Due to the absence of a dataset for this task, we construct a dataset based on the SkyScapes dataset. Our dataset is designed with linguistically structured expressions that focus on object categories, attributes, and spatial relationships, enabling the generation of binary masks from semantic segmentation maps. To benchmark this task, we evaluate and compare the performance of three different convolutional neural network (CNN)-based methods and a Transformer-based method. Experimental results provide valuable insights into the adaptability of these methods to remote sensing data, highlighting the potential of our dataset as a resource for the remote sensing community to further explore vision-language tasks. Zhenghang Yuan, Lichao Mou, Yuansheng Hua, Xiao Xiang Zhu 0001 |
IGARSS | 2 |
| 2024 | Ultrasound Image-to-Video Synthesis via Latent Dynamic Diffusion Models
Tingxiu Chen, Yilei Shi, Zixuan Zheng, Bingcong Yan, Jingliang Hu, Xiao Xiang Zhu 0001, Lichao Mou |
MICCAI (4) | 7 |
| 2024 | CausalCLIPSeg: Unlocking CLIP's Potential in Referring Medical Image Segmentation with Causal Intervention
Yaxiong Chen, Minghong Wei, Zixuan Zheng, Jingliang Hu, Yilei Shi, Shengwu Xiong 0001, Xiao Xiang Zhu 0001, Lichao Mou |
MICCAI (3) | 8 |
| 2024 | Striving for Simplicity: Simple Yet Effective Prior-Aware Pseudo-labeling for Semi-supervised Ultrasound Image Segmentation
Yaxiong Chen, Zixuan Zheng, Jingliang Hu, Yilei Shi, Shengwu Xiong 0001, Xiao Xiang Zhu 0001, Lichao Mou |
MICCAI (9) | 8 |
| 2024 | Rethinking Cell Counting Methods: Decoupling Counting and Localization
Zixuan Zheng, Yilei Shi, Jingliang Hu, Xiao Xiang Zhu 0001, Lichao Mou |
MICCAI (4) | 6 |
| 2024 | Reducing Annotation Burden: Exploiting Image Knowledge for Few-Shot Medical Video Object Segmentation via Spatiotemporal Consistency Relearning
Zixuan Zheng, Yilei Shi, Jingliang Hu, Xiao Xiang Zhu 0001, Lichao Mou |
MICCAI (12) | 6 |
| 2024 | Medical hyperspectral image classification based weakly supervised single-image global learning network
Lichao Mou, Shihao Shan, Hao Zhang 0113, Yafei Qi, Dexin Yu, Xiao Xiang Zhu 0001, Nianzheng Sun, Xiangrong Zheng, Xiaopeng Ma |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | Integrating Detailed Features and Global Contexts for Semantic Segmentation in Ultrahigh-Resolution Remote Sensing ImagesabstractSemantic segmentation of ultrahigh-resolution (UHR) remote sensing images is a fundamental task for many downstream applications. Achieving precise pixel-level classification is paramount for obtaining exceptional segmentation results. This challenge becomes even more complex due to the need to address intricate segmentation boundaries and accurately delineate small objects within the remote sensing imagery. To meet these demands effectively, it is critical to integrate two crucial components: global contextual information and spatial detail feature information. In response to this imperative, the multilevel context-aware segmentation network (MCSNet) emerges as a promising solution. MCSNet is engineered to not only model the overarching global context but also extract intricate spatial detail features, thereby optimizing segmentation outcomes. The strength of MCSNet lies in its two pivotal modules, the spatial detail feature extraction (SDFE) module and the refined multiscale feature fusion (RMFF) module. Moreover, to further harness the potential of MCSNet, a multitask learning approach is employed. This approach integrates boundary detection and semantic segmentation, ensuring that the network is well-rounded in its segmentation capabilities. The efficacy of MCSNet is rigorously demonstrated through comprehensive experiments conducted on two established international society for photogrammetry and remote sensing (ISPRS) 2-D semantic labeling datasets: Potsdam and Vaihingen. These experiments unequivocally establish MCSNet stands as a pioneering solution, that delivers state-of-the-art performance, as evidenced by its outstanding mean intersection over union (mIoU) and mean$F1$-score (mF1) metrics. The code is available at:https://github.com/WUTCM-Lab/MCSNet. Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu, Xiao Xiang Zhu 0001, Lichao Mou |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | A Review of Building Extraction From Remote Sensing Imagery: Geometrical Structures and Semantic AttributesabstractIn the remote sensing community, extracting buildings from remote sensing imagery has triggered great interest. While many studies have been conducted, a comprehensive review of these approaches that are applied to optical and synthetic aperture radar (SAR) imagery is still lacking. Therefore, we provide an in-depth review of both early efforts and recent advances, which are aimed at extracting geometrical structures or semantic attributes of buildings, including building footprint generation, building facade segmentation, roof segment and superstructure segmentation, building height retrieval, building type classification, building change detection, and annotation data correction. Furthermore, a list of corresponding benchmark datasets is given. Finally, challenges and outlooks of existing approaches as well as promising applications are discussed to enhance comprehension within this realm of research. Qingyu Li 0001, Lichao Mou, Yao Sun 0005, Yuansheng Hua, Yilei Shi, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | RRSIS: Referring Remote Sensing Image SegmentationabstractLocalizing desired objects from remote sensing images is of great use in practical applications. Referring image segmentation, which aims at segmenting out the objects to which a given expression refers, has been extensively studied in natural images. However, almost no research attention is given to this task of remote sensing imagery. Considering its potential for real-world applications, in this paper, we introduce referring remote sensing image segmentation (RRSIS) to fill in this gap and make some insightful explorations. Specifically, we create a new dataset, called RefSegRS, for this task, enabling us to evaluate different methods. Afterward, we benchmark referring image segmentation methods of natural images on the RefSegRS dataset and find that these models show limited efficacy in detecting small and scattered objects. To alleviate this issue, we propose a language-guided cross-scale enhancement (LGCE) module that utilizes linguistic features to adaptively enhance multi-scale visual features by integrating both deep and shallow features. The proposed dataset, benchmarking results, and the designed LGCE module provide insights into the design of a better RRSIS model. The dataset and code will be available at https://gitlab.lrz.de/ai4eo/reasoning/rrsis. Zhenghang Yuan, Lichao Mou, Yuansheng Hua, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | A Real-Time Unsupervised Hyperspectral Band Selection via Spatial-Spectral Information Fusion-Based Downscaled RegionabstractInformation fusion plays a vital role in hyperspectral band selection as it enables the exploration of the spatial-spectral structure relationship present in bands of hyperspectral images (HSIs). However, the focus of most algorithms primarily lies in the processing of single-band vectors, leaving only a few algorithms to handle spatial-spectral features of HSIs in a complex and inefficient manner. To overcome these limitations, a real-time unsupervised hyperspectral band selection method via spatial-spectral information fusion-based downscaled region (SIFDR) is proposed in this study. In particular, this approach incorporates an energy constraint method for assigning band weights to each detected pixel and estimates band spectral information through average fusion. Furthermore, a band weak redundancy sorting method is introduced, which is based on spectral information peaks, thereby achieving complementary spectral information. By performing regional downscaling of the HSI, spatial-spectral information is effectively fused, resulting in a real-time entire process. To evaluate the effectiveness of the proposed algorithm, experiments were conducted on four hyperspectral datasets, including an ultrahigh-dimensional medical HSIs, which distinguishes itself from previous methods that are typically evaluated exclusively on remote sensing datasets. Comparative results with several state-of-the-art (SOTA) algorithms demonstrate that the proposed algorithm excellently accomplishes hyperspectral band selection tasks in real time. The code of SIFDR has been shared onhttps://github.com/zhangchenglong1116/SIFDR. Lichao Mou, Xiangrong Zheng, Xiao Xiang Zhu 0001, Xiaopeng Ma |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Co-Enhanced Global-Part Integration for Remote-Sensing Scene ClassificationabstractRemote sensing (RS) scene classification aims to classify remote sensing images with similar scene characteristics into one category. Plenty of RS images are complex in background, rich in content, and multi-scale in target, exhibiting the characteristics of both intra-class separation and inter-class convergence. Therefore, discriminative feature representations designed to highlight the differences between classes are the key to RS scene classification. Existing methods represent scene images by extracting either global context or discriminative part features from RS images. However, global-based methods often lack salient details in similar RS scenes, while part-based methods tend to ignore the relationships between local ground objects, thus weakening the discriminative feature representation. In this paper, we propose to combine global context and part-level discriminative features within a unified framework called CGINet for accurate RS scene classification. To be specific, we develop a light context-aware attention block (LCAB) to explicitly model the global context to obtain larger receptive fields and contextual information. A co-enhanced loss module (CELM) is also devised to encourage the model to actively locate discriminative parts for feature enhancement. In particular, CELM is only used during training and not activated during inference, which introduces less computational cost. Benefiting from LCAB and CELM, our proposed CGINet improves the discriminability of features, thereby improving classification performance. Comprehensive experiments over four benchmark datasets show that the proposed method achieves consistent performance gains over state-of-the-art RS scene classification methods. Yichen Zhao, Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu, Xiao Xiang Zhu 0001, Lichao Mou |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Large-Scale Land Cover Mapping with Fine-Grained Classes via Class-Aware Semi-Supervised Semantic SegmentationabstractSemi-supervised learning has attracted increasing attention in the large-scale land cover mapping task. However, existing methods overlook the potential to alleviate the class imbalance problem by selecting a suitable set of unlabeled data. Besides, in class-imbalanced scenarios, existing pseudo-labeling methods mostly only pick confident samples, failing to exploit the hard samples during training. To tackle these issues, we propose a unified Class-Aware Semi-Supervised Semantic Segmentation framework. The proposed framework consists of three key components. To construct a better semi-supervised learning dataset, we propose a class-aware unlabeled data selection method that is more balanced towards the minority classes. Based on the built dataset with improved class balance, we propose a Class-Balanced Cross Entropy loss, jointly considering the annotation bias and the class bias to re-weight the loss in both sample and class levels to alleviate the class imbalance problem. Moreover, we propose the Class Center Contrast method to jointly utilize the labeled and unlabeled data. Specifically, we decompose the feature embedding space using the ground truth and pseudo-labels, and employ the embedding centers for hard and easy samples of each class per image in the contrast loss to exploit the hard samples during training. Compared with state-of-the-art class-balanced pseudo-labeling methods, the proposed method improves the mean accuracy and mIoU by 4.28% and 1.70%, respectively, on the large-scale Sentinel-2 dataset with 24 land cover classes. Runmin Dong, Lichao Mou, Mengxuan Chen, Xin-Yi Tong 0003, Shuai Yuan 0005, Lixian Zhang 0002, Juepeng Zheng, Xiao Xiang Zhu 0001, Haohuan Fu |
ICCV | 2 |
| 2023 | Amodal Segmentation Considering Visible and Non-Visible Elements of Urban SurfacesabstractThis study addresses the challenge of amodal segmentation in computer vision, a change in basic assumptions towards perceiving objects holistically, even when partially occluded, deviating from the traditional modal perspective that predominantly focuses on visible elements. Thus, we propose a new approach for the amodal segmentation of top-view aerial images, with particular attention to the first layer of elements, constituted by asphalt and natural soils, normally occluded by different objects (trees, buildings, and vehicles). This proposed methodology is data-centric, assigning weights to specific image sections and distinguishing non-visible elements. The best model used the U-Net architecture with Efficient-net-B7 as the backbone and can accurately classify occluded segments, achieving an Intersection over Union (IoU) greater than 80% for most classes. The developed method provides a basis for exploring amodal segmentation based on data-centric models, impacting our understanding of complex and occlusion-prone environments, such as urban environments. Osmar Luiz Ferreira de Carvalho, Anesmar Olino de Albuquerque, Osmar Abílio de Carvalho Jr., Lichao Mou, Daniel G. Silva |
IGARSS | 4 |
| 2023 | A Data-Centric Approach for Rapid Dataset Generation Using Iterative Learning and Sparse AnnotationsabstractThis study investigates the application of iterative sparse annotations for semantic segmentation in remote-sensing imagery, focusing on minimizing the laborious and expensive data labeling process. By leveraging Geographic Information Systems (GIS), we implemented circular polygon shapefiles to label portions of each class, attributing a value of -1 outside these polygons. The model training used the simplified BSB Aerial Dataset with eight classes. The semantic segmentation model was U-Net architecture with the Efficient-net-B7 backbone and a modified cross-entropy loss function. Our results showed promising improvement, particularly in error-prone classes, with the iterative addition of more samples. This approach suggests a quicker method for dataset creation using sparse, iteratively enhanced annotations. Future work will aim to implement further iterative rounds to approximate the results of continuous labeling, thereby enhancing the efficiency of semantic segmentation in large-scale remote-sensing images. Osmar Luiz Ferreira de Carvalho, Anesmar Olino de Albuquerque, Argélica Saiaka Luiz, Pedro Henrique Guimarães Ferreira, Lichao Mou, Daniel G. Silva, Osmar Abílio de Carvalho Jr. |
IGARSS | 5 |
| 2023 | Roof Superstructure Detection from Aerial ImageryabstractIdentifying suitable building roofs for the installation of photovoltaic (PV) systems is important to sustainable energy planning. However, most existing approaches neglect roof superstructures that can obstruct the installation of PV systems. In this research, we propose a novel method, which can help to deal with this issue by detecting roof superstructures from aerial imagery. Considering that semantic information about roof masks is also informative, we propose to first learns roof segmentation maps that are further used to learn roof superstructure maps. Experiments are conducted on Roof Information Dataset (RID). Our method outperforms the state-of-the-art methods both quantitatively and qualitatively. Qingyu Li 0001, Sebastian Krapf, Lichao Mou, Yilei Shi, Xiao Xiang Zhu 0001 |
IGARSS | 3 |
| 2023 | Overcoming Language Bias in Remote Sensing Visual Question Answering Via Adversarial TrainingabstractThe Visual Question Answering (VQA) system offers a user-friendly interface and enables human-computer interaction. However, VQA models commonly face the challenge of language bias, resulting from the learned superficial correlation between questions and answers. To address this issue, in this study, we present a novel framework to reduce the language bias of the VQA for remote sensing data (RSVQA). Specifically, we add an adversarial branch to the original VQA framework. Based on the adversarial branch, we introduce two regularizers to constrain the training process against language bias. Furthermore, to evaluate the performance in terms of language bias, we propose a new metric that combines standard accuracy with the performance drop when incorporating question and random image information. Experimental results demonstrate the effectiveness of our method. We believe that our method can shed light on future work for reducing language bias on the RSVQA task. Zhenghang Yuan, Lichao Mou, Xiao Xiang Zhu 0001 |
IGARSS | 2 |
| 2023 | DDM-Former: Global Ocean Wind Speed Retrieval with Transformer NetworksabstractAs a novel remote sensing technique, GNSS reflectometry (GNSS-R) opens a new era of retrieving Earth surface parameters. Several studies employ the combination of deep learning and GNSS-R observable delay-Doppler maps (DDMs) to generate ocean wind speed estimation. Unlike these methods that often use convolutional neural networks (CNNs) with inductive bias, we proposed a Transformer-based model, named DDM-Former, to exploit fine-grained delay-Doppler correlation independently. Our model is evaluated on the Cyclone GNSS (CYGNSS) version 3.0 dataset and shown to outperform the other retrieval methods. Daixin Zhao, Konrad Heidler, Milad Asgarimehr, Caroline Arnold, Tianqi Xiao, Jens Wickert, Xiao Xiang Zhu 0001, Lichao Mou |
IGARSS | 8 |
| 2023 | Function Assignment of Plastics based on Hyperspectral Satellite Images and High-Resolution Data Using Deep Learning AlgorithmsabstractPlastic pollution is becoming an increasingly prominent problem and the function of plastics determines whether they need to be recycled or not. In order to explore the possibility of using satellite imagery to classify the functionality of plastics, this study proposes a two-stage workflow: firstly, a classification map is obtained based on hyperspectral satellite imagery to generate plastic types, and then using these identified plastic coverage areas, a deep learning algorithm is used to assign functionality to these classified plastic areas based on sentinel-2 imagery. By comparing five leading-edge image classification models, classification accuracies of up to 74% were achieved, demonstrating the feasibility of using deep learning models trained on satellite images to identify plastic features. Shanyu Zhou, Lichao Mou, Lixian Zhang 0002, Yuansheng Hua, Hermann Kaufmann 0001, Xiao Xiang Zhu 0001 |
IGARSS | 2 |
| 2023 | Deep Saliency Smoothing Hashing for Drone Image RetrievalabstractDeep hashing algorithms are widely exploited in retrieval tasks due to its low storage and retrieval efficiency. Most of which focus on global feature learning, whilst neglecting local fine-grained features and saliency information for drone images. In this paper, we tackle these dilemmas with a novelDeep Saliency Smoothing Hashing(DSSH) algorithm, which can leverage saliency capture mechanism, distribution smoothing term, global features and local fine-grained features to learn effective hash codes for drone image retrieval. The DSSH algorithm first designs information extraction module to capture global features and local fine-grained features for drone images. Meanwhile, a saliency capture module is proposed to perform information interaction attention and visual enhancement attention, which can capture the saliency area of drone images effectively. On top of the two paths, a novel objective function is designed to preserve the similarity of hash codes, smooth the distribution of drone image datasets and reduce the quantization errors between hash codes and hash-like codes concurrently. Extensive experiments on the Drone Action Dataset and ERA Drone Dataset demonstrate that the DSSH algorithm can further improve the retrieval performance compared to other deep hashing algorithms. Yaxiong Chen, Lichao Mou, Pu Jin, Shengwu Xiong 0001, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | A Deep Active Contour Model for Delineating Glacier Calving FrontsabstractChoosing how to encode a real-world problem as a machine learning task is an important design decision in machine learning. The task of glacier calving front modeling has often been approached as a semantic segmentation task. Recent studies have shown that combining segmentation with edge detection can improve the accuracy of calving front detectors. Building on this observation, we completely rephrase the task as a contour tracing problem and propose a model for explicit contour detection that does not incorporate any dense predictions as intermediate steps. The proposed approach, called “Charting Outlines by Recurrent Adaptation” (COBRA), combines Convolutional Neural Networks (CNNs) for feature extraction and active contour models for the delineation. By training and evaluating on several large-scale datasets of Greenland’s outlet glaciers, we show that this approach indeed outperforms the aforementioned methods based on segmentation and edge-detection. Finally, we demonstrate that explicit contour detection has benefits over pixel-wise methods when quantifying the models’ prediction uncertainties. The project page containing the code and animated model predictions can be found at https://khdlr.github.io/COBRA/. Konrad Heidler, Lichao Mou, Erik Loebel, Mirko Scheinert, Sébastien Lefèvre, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Bounding Box Regression Network for Building Height Retrieval Using a Single SAR ImageabstractIn this paper, we propose a bounding box regression network for building height retrieval using a single TerraSAR - X stripmap image. The proposed network employs building footprints from GIS data and exploits the location relationship between a building's footprint and its bounding box, enabling fast computation. Experimental results over Rotterdam show that the proposed network can reduce the computation cost significantly while keeping the height accuracy of individual buildings compared to a Faster R-CNN based method. Yao Sun 0005, Lichao Mou, Yuanyuan Wang 0002, Xiao Xiang Zhu 0001 |
IGARSS | 2 |
| 2022 | Self Supervised Learning for Few Shot Hyperspectral Image ClassificationabstractDeep learning has proven to be a very effective approach for Hyperspectral Image (HSI) classification. However, deep neural networks require large annotated datasets to generalize well. This limits the applicability of deep learning for HSI classification, where manually labelling thousands of pixels for every scene is impractical. In this paper, we propose to leverage Self Supervised Learning (SSL) for HSI classification. We show that by pre-training an encoder on unlabeled pixels using Barlow-Twins, a state-of-the-art SSL algorithm, we can obtain accurate models with a handful of labels. Experimental results demonstrate that this approach significantly outperforms vanilla supervised learning. Nassim Ait Ali Braham, Lichao Mou, Jocelyn Chanussot, Julien Mairal, Xiao Xiang Zhu 0001 |
IGARSS | 2 |
| 2022 | Deep Active Contour Models for Delineating Glacier Calving FrontsabstractWe present a deep active contour model for detecting and delineating glacier calving fronts from satellite imagery. Contrary to existing deep learning-based calving front detectors, our model does not perform an intermediate segmentation or pixel-wise edge detection, but instead directly predicts the contour parametrized by a fixed number of vertices. The model works by first deriving feature maps from an input image, and then updating an initial contour in an iterative fashion. Evaluating on the CALFIN dataset, which maps calving fronts in Greenland, our model outperforms existing approaches. Code for the experiments and animated predictions can be found at https://github.com/khdlr/deep-acm Konrad Heidler, Lichao Mou, Erik Loebel, Mirko Scheinert, Sébastien Lefèvre, Xiao Xiang Zhu 0001 |
IGARSS | 2 |
| 2022 | Change-Aware Visual Question AnsweringabstractChange detection has been a hot research topic in the field of remote sensing, and it can provide information on observing changes of Earth's surface. However, segmentation-based change results are not very friendly to end users. Thus, in order to improve user experience and offer them high-level semantic information on change detection, we introduce a new task: change-aware visual question answering (VQA) on multi-temporal aerial images. Specifically, given a pair of multi-temporal aerial images and questions, this task aims to automatically provide natural language answers. By doing so, end users have better access to easy-to-understand change information through natural language. Besides, we also create a dataset made of multi-temporal image-question-answer triplets and a baseline method for this task. Experimental results offer valuable insights for the further research on this task. Zhenghang Yuan, Lichao Mou, Xiao Xiang Zhu 0001 |
IGARSS | 2 |
| 2022 | Deep Relearning in the Geospatial Domain for Semantic Remote Sensing Image SegmentationabstractWe present a classification postprocessing (CPP) technique based on fully convolutional neural networks (CNNs) for semantic remote sensing image segmentation. Conventional CPP techniques aim to enhance the classification accuracy by imposing smoothness priors in the image domain. Contrary to that, here, a relearning strategy is proposed where the initial classification outcome of a CNN model is provided to a subsequent CNN model via an extended input space to guide the learning of discriminative feature representations in an end-to-end fashion. This deep relearning CNN (DRCNN) explicitly accounts for the geospatial domain by taking the spatial alignment of preliminary class labels into account. Hereby, we evaluate to learn the DRCNN in a cumulative and noncumulative way, i.e., extending the input space based on all previous or solely preceding model outputs, respectively, during an iterative procedure. Besides, the DRCNN can also be conveniently coupled with alternative CPP techniques such as object-based voting (OBV). The experimental results obtained from two test sites of WorldView-II imagery underline the beneficial performance properties of the DRCNN models. They can increase the accuracies of the initial CNN models on average from 72.64% to 76.01% and from 92.43% to 94.52% in terms of$\kappa $statistic. An additional increase of 1.65 and 2.84 percentage points can be achieved when combining the DRCNN models with an OBV strategy. From an epistemological point of view, our results underline that CNNs can benefit from the consideration of preliminary model outcomes and that conventional CPP techniques can profit from an upstream relearning strategy. Christian Geiß, Yue Zhu 0004, Chunping Qiu, Lichao Mou, Xiao Xiang Zhu 0001, Hannes Taubenböck |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Semantic Segmentation of Remote Sensing Images With Sparse AnnotationsabstractTraining convolutional neural networks (CNNs) for very high-resolution images requires a large quantity of high-quality pixel-level annotations, which is extremely labor-intensive and time-consuming to produce. Moreover, professional photograph interpreters might have to be involved in guaranteeing the correctness of annotations. To alleviate such a burden, we propose a framework for semantic segmentation of aerial images based on incomplete annotations, where annotators are asked to label a few pixels with easy-to-draw scribbles. To exploit these sparse scribbled annotations, we propose the FEature and Spatial relaTional regulArization (FESTA) method to complement the supervised task with an unsupervised learning signal that accounts for neighborhood structures both in spatial and feature terms. For the evaluation of our framework, we perform experiments on two remote sensing image segmentation data sets involving aerial and satellite imagery, respectively. Experimental results demonstrate that the exploitation of sparse annotations can significantly reduce labeling costs, while the proposed method can help improve the performance of semantic segmentation when training on such annotations. The sparse labels and codes are publicly available for reproducibility purposes.https://github.com/Hua-YS/Semantic-Segmentation-with-Sparse-Labels Yuansheng Hua, Diego Marcos, Lichao Mou, Xiao Xiang Zhu 0001, Devis Tuia |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Deep Quadruple-Based Hashing for Remote Sensing Image-Sound RetrievalabstractWith the rapid progress of earth observation technology, cross-modal remote sensing (RS) image-sound retrieval has attracted much attention from the field of RS data processing. Existing approaches usually learn the pairwise similarity relations between RS images and sounds. However, these approaches ignore relative semantic similarity relationships, which leads to poor performance of cross-modal RS image-sound retrieval. In this article, we address this dilemma with a noveldeep quadruple-based hashing(DQH) approach. We first devise a novel quadruple-based hashing network to learn relative semantic similarity relationships of hash codes. Meanwhile, we propose a quadruple construction hard module, which randomly selects two triplet hard units to directly learn relative semantic similarity relationships. On top of the two paths, we develop a new objective function to perform effective hash codes learning. The new objective function not only captures the relative semantic correlation of hash codes across different modalities and learns the relative semantic correlation of deep features but also enhances category-level semantics of hash codes and reduces the quantization error between hash-like codes and hash codes. The reasonableness and effectiveness of the proposed architecture are well illustrated by comprehensive experiments on diverse RS image-sound datasets. Yaxiong Chen, Shengwu Xiong 0001, Lichao Mou, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Bi-Temporal Semantic Reasoning for the Semantic Change Detection in HR Remote Sensing ImagesabstractSemantic change detection (SCD) extends the multiclass change detection (MCD) task to provide not only the change locations but also the detailed land-cover/land-use (LCLU) categories before and after the observation intervals. This fine-grained semantic change information is very useful in many applications. Recent studies indicate that the SCD can be modeled through a triple-branch convolutional neural network (CNN), which contains two temporal branches and a change branch. However, in this architecture, the communications between the temporal branches and the change branch are insufficient. To overcome the limitations in existing methods, we propose a novel CNN architecture for the SCD, where the semantic temporal features are merged in a deep CD unit. Furthermore, we elaborate on this architecture to reason the bi-temporal semantic correlations. The resulting bi-temporal semantic reasoning network (Bi-SRNet) contains two types of semantic reasoning blocks to reason both single-temporal and cross-temporal semantic correlations, as well as a novel loss function to improve the semantic consistency of change detection results. Experimental results on a benchmark dataset show that the proposed architecture obtains significant accuracy improvements over the existing approaches, while the added designs in the Bi-SRNet further improve the segmentation of both semantic categories and the changed areas. The codes in this article are accessible athttps://github.com/ggsDing/Bi-SRNet. Lei Ding 0008, Haitao Guo, Sicong Liu 0001, Lichao Mou, Jing Zhang 0023, Lorenzo Bruzzone |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | HED-UNet: Combined Segmentation and Edge Detection for Monitoring the Antarctic CoastlineabstractDeep learning-based coastline detection algorithms have begun to outshine traditional statistical methods in recent years. However, they are usually trained only as single-purpose models to either segment land and water or delineate the coastline. In contrast to this, a human annotator will usually keep a mental map of both segmentation and delineation when performing manual coastline detection. To take into account this task duality, we, therefore, devise a new model to unite these two approaches in a deep learning model. By taking inspiration from the main building blocks of a semantic segmentation framework (UNet) and an edge detection framework (HED), both tasks are combined in a natural way. Training is made efficient by employing deep supervision on side predictions at multiple resolutions. Finally, a hierarchical attention mechanism is introduced to adaptively merge these multiscale predictions into the final model output. The advantages of this approach over other traditional and deep learning-based methods for coastline detection are demonstrated on a data set of Sentinel-1 imagery covering parts of the Antarctic coast, where coastline detection is notoriously difficult. An implementation of our method is available athttps://github.com/khdlr/HED-UNet. Konrad Heidler, Lichao Mou, Celia A. Baumhoer, Andreas J. Dietz, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | MultiScene: A Large-Scale Dataset and Benchmark for Multiscene Recognition in Single Aerial Images
Yuansheng Hua, Lichao Mou, Pu Jin, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | FuTH-Net: Fusing Temporal Relations and Holistic Features for Aerial Video ClassificationabstractUnmanned aerial vehicles (UAVs) are now widely applied to data acquisition due to its low cost and fast mobility. With the increasing volume of aerial videos, the demand for automatically parsing these videos is surging. To achieve this, current research mainly focuses on extracting a holistic feature with convolutions along both spatial and temporal dimensions. However, these methods are limited by small temporal receptive fields and cannot adequately capture long-term temporal dependencies that are important for describing complicated dynamics. In this article, we propose a novel deep neural network, termed Fusing Temporal relations and Holistic features for aerial video classification (FuTH-Net), to model not only holistic features but also temporal relations for aerial video classification. Furthermore, the holistic features are refined by the multiscale temporal relations in a novel fusion module for yielding more discriminative video representations. More specially, FuTH-Net employs a two-pathway architecture: 1) a holistic representation pathway to learn a general feature of both frame appearances and short-term temporal variations and 2) a temporal relation pathway to capture multiscale temporal relations across arbitrary frames, providing long-term temporal dependencies. Afterward, a novel fusion module is proposed to spatiotemporally integrate the two features learned from the two pathways. Our model is evaluated on two aerial video classification datasets, ERA and Drone-Action, and achieves the state-of-the-art results. This demonstrates its effectiveness and good generalization capacity across different recognition tasks (event classification and human action recognition). To facilitate further research, we release the code athttps://gitlab.lrz.de/ai4eo/reasoning/futh-net. Pu Jin, Lichao Mou, Yuansheng Hua, Gui-Song Xia, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Anomaly Detection in Aerial Videos With TransformersabstractUnmanned aerial vehicles (UAVs) are widely applied for purposes of inspection, search, and rescue operations by the virtue of low-cost, large-coverage, real-time, and high-resolution data acquisition capacities. Massive volumes of aerial videos are produced in these processes, in which normal events often account for an overwhelming proportion. It is extremely difficult to localize and extract abnormal events containing potentially valuable information from long video streams manually. Therefore, we are dedicated to developing anomaly detection methods to solve this issue. In this paper, we create a new dataset, named Drone-Anomaly, for anomaly detection in aerial videos. This dataset provides 37 training video sequences and 22 testing video sequences from 7 different realistic scenes with various anomalous events. There are 87,488 color video frames (51,635 for training and 35,853 for testing) with the size of 640 × 640 at 30 frames per second. Based on this dataset, we evaluate existing methods and offer a benchmark for this task. Furthermore, we present a new baseline model, ANomaly Detection with Transformers (ANDT), which treats consecutive video frames as a sequence of tubelets, utilizes a Transformer encoder to learn feature representations from the sequence, and leverages a decoder to predict the next frame. Our network models normality in the training phase and identifies an event with unpredictable temporal dynamics as an anomaly in the test phase. Moreover, To comprehensively evaluate the performance of our proposed method, we use not only our Drone-Anomaly dataset but also another dataset. We will make our dataset and code publicly available. A demo video is available at https://youtu.be/ancczYryOBY. We make our dataset and code publicly available1. Pu Jin, Lichao Mou, Gui-Song Xia, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Building Footprint Generation Through Convolutional Neural Networks With Attraction Field RepresentationabstractBuilding footprint generation is a vital task in a wide range of applications, including, to name a few, land use management, urban planning and monitoring, and geographical database updating. Most existing approaches addressing this problem fall back on convolutional neural networks (CNNs) to learn semantic masks of buildings. However, one limitation of their results is blurred building boundaries. To address this, we propose to learn attraction field representation for building boundaries, which is capable of providing an enhanced representation power. Our method comprises two elemental modules: an Img2AFM module and an AFM2Mask module. More specifically, the former aims at learning an attraction field representation conditioned on an input image, which is capable of enhancing building boundaries and suppressing the background. The latter module predicts segmentation masks of buildings using the learned attraction field map. The proposed method is evaluated on three datasets with different spatial resolutions: the ISPRS dataset, the INRIA dataset, and the Planet dataset. From experimental results, we find that the proposed framework can well preserve geometric shapes and sharp boundaries of buildings, which brings significant improvements over other competitors. The trained model and code are available at https://github.com/lqycrystal/AFM_building. Qingyu Li 0001, Lichao Mou, Yuansheng Hua, Yilei Shi, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | A Lightweight Deep Learning-Based Cloud Detection Method for Sentinel-2A Imagery Fusing Multiscale Spectral and Spatial FeaturesabstractClouds are a very important factor in the availability of optical remote sensing images. Recently, deep learning (DL)-based cloud detection methods have surpassed classical methods based on rules and physical models of clouds. However, most of these deep models are very large, which limits their applicability and explainability, while other models do not make use of the full spectral information in multispectral images, such as Sentinel-2. In this article, we propose a lightweight network for cloud detection, fusing multiscale spectral and spatial features (CD-FM3SFs) and tailored for processing all spectral bands in Sentinel-2A images. The proposed method consists of an encoder and a decoder. In the encoder, three input branches are designed to handle spectral bands at their native resolution and extract multiscale spectral features. Three novel components are designed: a mixed depthwise separable convolution (MDSC) and a shared and dilated residual block (SDRB) to extract multiscale spatial features, and a concatenation and sum (CS) operation to fuse multiscale spectral and spatial features with little calculation and no additional parameters. The decoder of CD-FM3SF outputs three cloud masks at the same resolution as input bands to enhance the supervision information of small, middle, and large clouds. To validate the performance of the proposed method, we manually labeled 36 Sentinel-2A scenes evenly distributed over mainland China. The experiment results demonstrate that CD-FM3SF outperforms traditional cloud detection methods and state-of-the-art DL-based methods in both accuracy and speed. Jun Li 0087, Zhaocong Wu, Zhongwen Hu, Canliang Jian, Shaojie Luo, Lichao Mou, Xiao Xiang Zhu 0001, Matthieu Molinier |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Detecting Changes by Learning No Changes: Data-Enclosing-Ball Minimizing Autoencoders for One-Class Change Detection in Multispectral ImageryabstractChange detection is a long-standing and challenging problem in remote sensing. Very often, features about changes are difficult to model beforehand, thus making the collection of changed samples a challenging task. In comparison, it is much easier to collect numerous no-change samples. It is possible to define a change detection approach by using only easily available annotated no-change samples, which we henceforth call one-class change detection. Autoencoder networks being trained on no-change data are natural candidates for addressing this task due to their superior performance as compared to other one-class classification models. However, the autoencoders usually suffer from the problem of overgeneralization, i.e., they tend to generalize too well, thus risking properly reconstructing changed samples. In this paper, we propose a novel data-enclosing-ball minimizing autoencoder (DebM-AE) that is trained with dual objectives—a reconstruction error criterion and a minimum volume criterion. The network learns a compact latent space, where encodings of no-change samples have low intra-class variance, which as counter part has the identification of changed instances. We conducted extensive experiments on three real-world data sets. Results demonstrate advantages of the proposed method over other competitors. We make our data and code publicly available1. Lichao Mou, Yuansheng Hua, Sudipan Saha, Francesca Bovolo, Lorenzo Bruzzone, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Deep Reinforcement Learning for Band Selection in Hyperspectral Image ClassificationabstractBand selection refers to the process of choosing the most relevant bands in a hyperspectral image. By selecting a limited number of optimal bands, we aim at speeding up model training, improving accuracy, or both. It reduces redundancy among spectral bands while trying to preserve the original information of the image. By now, many efforts have been made to develop unsupervised band selection approaches, of which the majorities are heuristic algorithms devised by trial and error. In this article, we are interested in training an intelligent agent that, given a hyperspectral image, is capable of automatically learning policy to select an optimal band subset without any hand-engineered reasoning. To this end, we frame the problem of unsupervised band selection as a Markov decision process, propose an effective method to parameterize it, and finally solve the problem by deep reinforcement learning. Once the agent is trained, it learns a band-selection policy that guides the agent to sequentially select bands by fully exploiting the hyperspectral image and previously picked bands. Furthermore, we propose two different reward schemes for the environment simulation of deep reinforcement learning and compare them in experiments. This, to the best of our knowledge, is the first study that explores a deep reinforcement learning model for hyperspectral image analysis, thus opening a new door for future research and showcasing the great potential of deep reinforcement learning in remote sensing applications. Extensive experiments are carried out on four hyperspectral data sets, and experimental results demonstrate the effectiveness of the proposed method. The code is publicly available. Lichao Mou, Sudipan Saha, Yuansheng Hua, Francesca Bovolo, Lorenzo Bruzzone, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Unsupervised Single-Scene Semantic Segmentation for Earth ObservationabstractEarth observation data has huge potential to enrich our knowledge about our planet. An important step in many Earth observation tasks is semantic segmentation. Generally, a large number of pixelwise labeled images are required to train deep models for supervised semantic segmentation. On the contrary, strong inter-sensor and geographic variations impede the availability of annotated training data in Earth observation. In practice, most Earth observation tasks use only the target scene without assuming availability of any additional scene, labeled or unlabeled. Keeping in mind such constraints, we propose a semantic segmentation method that learns to segment from a single scene, without using any annotation. Earth observation scenes are generally larger than those encountered in typical computer vision datasets. Exploiting this, the proposed method samples smaller unlabeled patches from the scene. For each patch an alternate view is generated by simple transformations, e.g., addition of noise. Both views are then processed through a two-stream network and weights are iteratively refined using deep clustering, spatial consistency, and contrastive learning in the pixel space. The proposed model automatically segregates the major classes present in the scene and produces the segmentation map. Extensive experiments on four Earth observation datasets collected by different sensors show the effectiveness of the proposed method. Implementation is available at https://gitlab.lrz.de/ai4eo/cd/-/tree/main/unsupContrastiveSemanticSeg. Sudipan Saha, Muhammad Shahzad 0002, Lichao Mou, Qian Song, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | CG-Net: Conditional GIS-Aware Network for Individual Building Segmentation in VHR SAR ImagesabstractObject retrieval and reconstruction from very-high-resolution (VHR) synthetic aperture radar (SAR) images are of great importance for urban SAR applications, yet highly challenging due to the complexity of SAR data. This article addresses the issue of individual building segmentation from a single VHR SAR image in large-scale urban areas. To achieve this, we introduce building footprints from geographic information system (GIS) data as a complementary information and propose a novel conditional GIS-aware network (CG-Net). The proposed model learns multilevel visual features and employs building footprints to normalize the features for predicting building masks in the SAR image. We validate our method using a high-resolution spotlight TerraSAR-X image collected over Berlin. Experimental results show that the proposed CG-Net effectively brings improvements with variant backbones. We further compare two representations of building footprints, namely, complete building footprints and sensor-visible footprint segments, for our task, and conclude that the use of the former leads to better segmentation results. Moreover, we investigate the impact of inaccurate GIS data on our CG-Net, and this study shows that CG-Net is robust against positioning errors in the GIS data. In addition, we propose an approach of ground truth generation of buildings from an accurate digital elevation model (DEM), which can be used to generate large-scale SAR image data sets. The segmentation results can be applied to reconstruct 3-D building models at level-of-detail (LoD) 1, which is demonstrated in our experiments. Yao Sun 0005, Yuansheng Hua, Lichao Mou, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | An Unsupervised Remote Sensing Change Detection Method Based on Multiscale Graph Convolutional Network and Metric LearningabstractAs a fundamental application, change detection (CD) is widespread in the remote sensing (RS) community. With the increase in the spatial resolution of RS images, high-resolution remote sensing (HRRS) image CD tasks receive growing attention. The change information hidden in multitemporal HRRS images could help discover our planet comprehensively. In the current deep learning era, convolutional neural networks (CNNs) have become one of the most powerful tools for a wide range of RS tasks including HRRS image CD, due to their superb feature learning capacity. However, most of them need a large amount of labeled data to accomplish the CD process, which is challenging or even impractical in many RS applications. Also, given the limited valid receptive field, CNNs can only capture short-range context within HRRS images, which is probably not enough to fully explore change information from the images. To overcome these limitations, in this article, we propose an unsupervised CD method, termed GMCD, based on graph convolutional network (GCN) and metric learning. GMCD consists of a Siamese fully convolution network (FCN), a multiscale dynamic GCN (Mlt-GCN), and a pseudolabel generation mechanism based on metric learning. The Siamese FCN contains a Siamese encoder and a pyramid-shaped decoder, aiming to extract multiscale features and integrate them to generate reliable difference images (DIs). Mlt-GCN focuses on capturing the short- and long-range contextual patterns at feature map level to extract changed and unchanged areas completely. The pseudolabel generation mechanism aims to produce reliable pseudolabels (changed, unchanged, and uncertain) to help accomplish the model training in an unsupervised way. Experiments on four HRRS image CD datasets demonstrate that GMCD outperforms the existing state-of-the-art methods. Xu Tang 0004, Lichao Mou, Fang Liu 0034, Xiangrong Zhang, Xiao Xiang Zhu 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | SCIDA: Self-Correction Integrated Domain Adaptation From Single- to Multi-Label Aerial ImagesabstractMost publicly available datasets for image classification are with single labels, while images are inherently multilabeled in our daily life. Such an annotation gap makes many pretrained single-label classification models fail in practical scenarios. For aerial images, this annotation issue is more concerned: Aerial data naturally cover a relatively large land area with multiple labels, while annotated aerial datasets currently publicly available (e.g., UCM and AID) are single-labeled. As manually annotating multilabel aerial images (MAIs) would be time-/ labor-consuming, we propose a novel self-correction integrated domain adaptation (SCIDA) method for automatic multilabel learning. SCIDA is weakly supervised, i.e., automatically learning the multilabel image classification model from using massive, publicly available single-label images. To achieve this goal, we propose a novel labelwise self-correction (LWC) module to better explore underlying label correlations. This module also makes the unsupervised domain adaptation (UDA) from single-label to multilabel data possible. For model training, the proposed method uses single-label information yet requires no prior knowledge of multilabeled data and predicts labels for MAIs. Through extensive evaluations, the proposed model, which is trained with single-labeled MAI-AID-s and MAI-UCM-s datasets, achieves much better performances than comparative methods on our collected multiscene aerial image dataset. The code and data are available on GitHub (https://github.com/Ryan315/Single2multi-DA). Tianze Yu, Jianzhe Lin, Lichao Mou, Yuansheng Hua, Xiao Xiang Zhu 0001, Z. Jane Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | From Easy to Hard: Learning Language-Guided Curriculum for Visual Question Answering on Remote Sensing DataabstractVisual question answering (VQA) for remote sensing scene has great potential in intelligent human-computer interaction system. Although VQA in computer vision has been widely researched, VQA for remote sensing data (RSVQA) is still in its infancy. There are two characteristics that need to be specially considered for the RSVQA task. 1) No object annotations are available in RSVQA datasets, which makes it difficult for models to exploit informative region representation; 2) There are questions with clearly different difficulty levels for each image in the RSVQA task. Directly training a model with questions in a random order may confuse the model and limit the performance. To address these two problems, in this paper, a multi-level visual feature learning method is proposed to jointly extract language-guided holistic and regional image features. Besides, a self-paced curriculum learning (SPCL)-based VQA model is developed to train networks with samples in an easy-to-hard way. To be more specific, a language-guided SPCL method with a soft weighting strategy is explored in this work. The proposed model is evaluated on three public datasets, and extensive experimental results show that the proposed RSVQA framework can achieve promising performance. Code will be available at https://gitlab.lrz.de/ai4eo/reasoning/VQA-easy2hard. Zhenghang Yuan, Lichao Mou, Qi Wang 0009, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Change Detection Meets Visual Question AnsweringabstractThe Earth’s surface is continually changing, and identifying changes plays an important role in urban planning and sustainability. Although change detection techniques have been successfully developed for many years, these techniques are still limited to experts and facilitators in related fields. In order to provide every user with flexible access to change information and help them better understand land-cover changes, we introduce a novel task: change detection-based visual question answering (CDVQA) on multi-temporal aerial images. In particular, multi-temporal images can be queried to obtain high level change-based information according to content changes between two input images. We first build a CDVQA dataset including multi-temporal image-question-answer triplets using an automatic question-answer generation method. Then, a baseline CDVQA framework is devised in this work, and it contains four parts: multi-temporal feature encoding, multi-temporal fusion, multi-modal fusion, and answer prediction. In addition, we also introduce a change enhancing module to multi-temporal feature encoding, aiming at incorporating more change-related information. Finally, effects of different backbones and multi-temporal fusion strategies are studied on the performance of CDVQA task. The experimental results provide useful insights for developing better CDVQA models, which are important for future research on this task. The dataset will be available at https://github.com/YZHJessica/CDVQA. Zhenghang Yuan, Lichao Mou, Zhitong Xiong, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Mask-Height R-CNN: An End-to-End Network for 3D Building Reconstruction from Monocular Remote Sensing Imageryabstract3D building reconstruction from monocular remote sensing imagery is a promising and economical way to generate 3D city models at a large scale, yet the task is rarely touched. The paper tackles the problem via an end-to-end network. The goal is achieved by a modified network, named Mask-Height R-CNN, based on Mask R-CNN, with an additional height prediction head in the Region Proposal Network (RPN). Unlike most deep learning based methods, the height estimation is done on the instance level instead of pixel level, which does not require the assembly of the height maps and building masks. The proposed network gains good performances on ISPRS datasets, with 3D F1 scores of over 0.8. Sining Chen, Lichao Mou, Qingyu Li 0001, Yao Sun 0005, Xiao Xiang Zhu 0001 |
IGARSS | 2 |
| 2021 | Hed-Unet: A Multi-Scale Framework for Simultaneous Segmentation and Edge DetectionabstractSegmentation models for remote sensing imagery are usually trained on the segmentation task alone. However, for many applications, the class boundaries carry semantic value. To account for this, we propose a new approach that unites both tasks within a single deep learning model. The proposed network architecture follows the successful encoder-decoder approach, and is improved by employing deep supervision at multiple resolution levels, as well as merging these resolution levels into a final prediction using a hierarchical attention mechanism. This framework is trained to detect the coastline in Sentinel-1 images of the Antarctic coastline. Its performance is then compared to conventional single-task approaches, and shown to outperform these methods. The code is available at https://github.com/khdlr/HED-UNet. Konrad Heidler, Lichao Mou, Celia A. Baumhoer, Andreas J. Dietz, Xiao Xiang Zhu 0001 |
IGARSS | 2 |
| 2021 | Seeing the Bigger Picture: Enabling Large Context Windows in Neural Networks by Combining Multiple Zoom LevelsabstractWhen adopting deep learning methods for remote sensing applications, the data usually needs to be cut into patches due to hardware limitations. Clearly, this practice discards a lot of contextual information as the model's information is limited to imagery from the given patch. We propose a memory-efficient way around this limitation by using multiple patches of varying spatial extents on different resolution levels. Finally, this new approach is evaluated for the task of automated sea ice charting, where the added contextual information is shown to be beneficial to model performance. Konrad Heidler, Lichao Mou, Xiao Xiang Zhu 0001 |
IGARSS | 2 |
| 2021 | Unconstrained Aerial Scene Recognition with Deep Neural Networks and a New DatasetabstractAerial scene recognition is a fundamental research problem in interpreting high-resolution aerial imagery. Over the past few years, most studies focus on classifying an image into one scene category, while in real-world scenarios, it is more often that a single image contains multiple scenes. Therefore, in this paper, we investigate a more practical yet underexplored task-multi-scene recognition in single images. To this end, we create a large-scale dataset, called Mul-tiScene dataset, composed of 100,000 unconstrained images each with multiple labels from 36 different scenes. Among these images, 14,000 of them are manually interpreted and assigned ground-truth labels, while the remaining images are provided with crowdsourced labels, which are generated from low-cost but noisy OpenStreetMap (OSM) data. By doing so, our dataset allows two branches of studies: 1) developing novel CNNs for multi-scene recognition and 2) learning with noisy labels. We experiment with extensive baseline models on our dataset to offer a benchmark for multi-scene recognition in single images. Aiming to expedite further researches, we will make our dataset and pre-trained models available11https://github.com/Hua-YS/Multi-Scene-Recognition. Yuansheng Hua, Lichao Mou, Pu Jin, Xiao Xiang Zhu 0001 |
IGARSS | 2 |
| 2021 | Temporal Relations Matter: A Two-Pathway Network for Aerial Video RecognitionabstractWith the increasing volume of aerial videos, the demand for automatically parsing these videos is surging. To achieve this, current researches mainly focus on extracting a holistic feature with convolutions along both spatial and temporal dimensions. However, these methods are limited by small temporal receptive fields and cannot adequately capture long-term temporal dependencies which are important for describing complicated dynamics. In this paper, we propose a novel two-pathway network to model not only holistic features, but also temporal relations for aerial video classification. More specially, our model employs a two-pathway architecture: (1) a holistic representation pathway to learn a general feature of frame appearances and short-term temporal variations and (2) a temporal relation pathway to capture multi-scale temporal relations across arbitrary frames, providing long-term temporal dependencies. Our model is evaluated on event recognition dataset, ERA, and achieves the state-of-the-art results. This demonstrates its effectiveness and good generalization capacity. Pu Jin, Lichao Mou, Yuansheng Hua, Gui-Song Xia, Xiao Xiang Zhu 0001 |
IGARSS | 2 |
| 2021 | Anomaly Detection in Aerial Videos Via Future Frame Prediction NetworksabstractBy the virtue of high flexibility, low-cost, real-time, and high-resolution data acquisition capacity, unmanned aerial vehicles (UAVs) can be exploited for a wide range of applications, especially in surveillance, inspection, and search fields. Such applications aim to detect potential suspicious events, violent human actions from an untrimmed and lengthy UAV video. Anomaly detection methods are highly in demand because it is unrealistic for human experts to manually detect all abnormal events in image scene. However, anomaly detection methods in aerial videos are rarely studied in the remote sensing community. In this paper, We propose a future frame prediction network based on convolutional variational autoencoder networks to detect anomalous events. Compared to several models, our network has a superior performance. Pu Jin, Lichao Mou, Gui-Song Xia, Xiao Xiang Zhu 0001 |
IGARSS | 2 |
| 2021 | Improving Land Cover Classification with a Shift-Invariant Center-Focusing Convolutional Neural NetworkabstractConvolutional neural networks (CNNs) are widely employed in remote sensing community. The CNN-based, also known as patch-based land cover classification method has gained increasing attention. However, this method very often requires the aid of post-processing, otherwise it is difficult to obtain accurate boundaries separating different land cover classes. In this paper, we discuss the reason of this phenomenon and propose a shift-invariant center-focusing (SICF) network to deliver more accurate boundaries to improve the patch-based land cover classification. The principle of SICF is calculating the class score from a center-focusing area based on a shift-invariant feature extraction module to calibrate prediction. We employ three modern CNNs to build corresponding SICF networks, the evaluation results indicate that compared with the conventional CNNs, the improvements made by SICF for delivering accurate boundaries in land cover classification are significant. Yuansheng Hua, Lichao Mou, Xiao Xiang Zhu 0001 |
IGARSS | 3 |
| 2021 | Conditional GIS-Aware Network for Individual Building Segmentation in a VHR SAR ImageabstractIn this paper, we propose a network for individual building segmentation from a single VHR SAR image. The proposed network employs building footprints from GIS data in learning multi-level visual features to predict building masks in the SAR image. Experimental results over Berlin show that the proposed network effectively brings improvements with variant backbones. In addition, we propose an approach for generating building labels from an accurate digital elevation model (DEM), which can be used to generate large-scale SAR image datasets. Yao Sun 0005, Yuansheng Hua, Lichao Mou, Xiao Xiang Zhu 0001 |
IGARSS | 3 |
| 2021 | Self-Paced Curriculum Learning for Visual Question Answering on Remote Sensing DataabstractAnswering questions with natural language by extracting information from image has great potential in various applications. Although visual question answering (VQA) for natural image has been broadly studied, VQA for remote sensing data is still in the early research stage. For the same remote sensing image, there exist questions with dramatically different difficulty-levels. Treating these questions equally may mislead the model and limit the VQA model performance. Considering this problem, in this work, we propose a self-paced curriculum learning (SPCL) based VQA model with hard and soft weighting strategies for remote sensing data. Like human learning process, the model is trained from easy to hard question samples gradually. Extensive experimental results on two datasets demonstrate that the proposed training method can achieve promising performance. Zhenghang Yuan, Lichao Mou, Xiao Xiang Zhu 0001 |
IGARSS | 2 |
| 2021 | Semisupervised Change Detection Using Graph Convolutional NetworkabstractMost change detection (CD) methods are unsupervised as collecting substantial multitemporal training data is challenging. Unsupervised CD methods are driven by heuristics and lack the capability to learn from data. However, in many real-world applications, it is possible to collect a small amount of labeled data scattered across the analyzed scene. Such a few scattered labeled samples in the pool of unlabeled samples can be effectively handled by graph convolutional network (GCN) that has recently shown good performance in semisupervised single-date analysis, to improve change detection performance. Based on this, we propose a semisupervised CD method that encodes multitemporal images as a graph via multiscale parcel segmentation that effectively captures the spatial and spectral aspects of the multitemporal images. The graph is further processed through GCN to learn a multitemporal model. Information from the labeled parcels is propagated to the unlabeled ones over training iterations. By exploiting the homogeneity of the parcels, the model is used to infer the label at a pixel level. To show the effectiveness of the proposed method, we tested it on a multitemporal Very High spatial Resolution (VHR) data set acquired by Pleiades sensor over Trento, Italy. Sudipan Saha, Lichao Mou, Xiao Xiang Zhu 0001, Francesca Bovolo, Lorenzo Bruzzone |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2021 | Attention-Aware Pseudo-3-D Convolutional Neural Network for Hyperspectral Image ClassificationabstractConvolutional neural networks (CNNs) have been applied for hyperspectral image classification recently. Among this class of deep models, 3-D CNN has been shown to be more effective by learning discriminative features from abundant spectral signatures and spatial contexts in hyperspectral imagery (HSI). However, by simply imposing 3-D CNN to HSI, a large amount of initial information might be lost in this CNN pipeline. The proposed attention-aware pseudo-3-D (AP3D) convolutional network for HSI classification is motivated by two observations. First, each dimension of the 3-D HSI is not equally important, different attention should be paid to different dimensions of the initial HSI image, especially in the first convolution operation. Second, intermediate representations of the 3-D input image at different stages in the 3-D CNN pipeline represent different levels of features and should not be neglected and abandoned. Instead, a 2-D matrix of scores for each feature map should be fed to the final softmax layer. Quantitative and qualitative results demonstrate that the proposed AP3D model outperforms the state-of-the-art HSI classification methods in agricultural and rural/urban data sets: Indian Pines, Pavia University, and Salinas Scene. Jianzhe Lin, Lichao Mou, Xiao Xiang Zhu 0001, Xiangyang Ji, Z. Jane Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Unifying Top-Down Views by Task-Specific Domain AdaptationabstractIn this article, we aim to learn a unified representation of images from satellite/aerial/ground views by exploring their underlying correlations. Inspired by recent advances in domain adaptation (DA), we propose a novel task-specific DA method for this purpose. Different from traditional DA methods, this proposed method not only applies task-specific classifiers1but also introduces domain-specific tasks for different domains during the adaptation process. The experiments are conducted on two newly proposed ground-/satellite-to-aerial scene adaptation (GSSA) data sets. Since the semantic gap between the ground/satellite scenes and the aerial scenes is much larger than that between ground scenes, the DA task between these scenes is more challenging than traditional DA tasks. On GSSA data sets, we not only demonstrate the proposed unsupervised DA method but also explore the few-shot DA in the discussion section. The proposed method is easy to implement, and our method substantially outperforms the state-of-the-art methods on the studied data sets.We hope that the proposed method for the novel GSSA data sets can be a good baseline for future researchers. The related data sets/codes will be available online. Jianzhe Lin, Tianze Yu, Lichao Mou, Xiao Xiang Zhu 0001, Rabab K. Ward, Z. Jane Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | Cross-Task Transfer for Geotagged Audiovisual Aerial Scene Recognition
Di Hu 0001, Xuhong Li 0002, Lichao Mou, Pu Jin, Liping Jing, Xiao Xiang Zhu 0001, Dejing Dou |
ECCV (24) | 3 |
| 2020 | Learning Multi-Label Aerial Image Classification Under Label Noise: A Regularization Approach Using Word EmbeddingsabstractTraining deep neural networks requires well-annotated datasets. However, real world datasets are often noisy, especially in a multi-label scenario, i.e. where each data point can be attributed to more than one class. To this end, we propose a regularization method to learn multi-label classification networks from noisy data. This regularization is based on the assumption that semantically close classes are more likely to appear together in a given image. Hereby, we encode label correlations with prior knowledge and regularize noisy network predictions using label correlations. To evaluate its effectiveness, we perform experiments on a mutli-label aerial image dataset contaminated with controlled levels of label noise. Results indicate that networks trained using the proposed method outperform those directly learned from noisy labels and that the benefits increase proportionally to the amount of noise present. Yuansheng Hua, Sylvain Lobry, Lichao Mou, Devis Tuia, Xiao Xiang Zhu 0001 |
IGARSS | 3 |
| 2020 | Instance Segmentation of Buildings Using KeypointsabstractBuilding segmentation is of great importance in the task of remote sensing imagery interpretation. However, the existing semantic segmentation and instance segmentation methods often lead to segmentation masks with blurred boundaries. In this paper, we propose a novel instance segmentation network for building segmentation in high-resolution remote sensing images. More specifically, we consider segmenting an individual building as detecting several keypoints. The detected keypoints are subsequently reformulated as a closed polygon, which is the semantic boundary of the building. By doing so, the sharp boundary of the building could be preserved. Experiments are conducted on selected Aerial Imagery for Roof Segmentation (AIRS) dataset, and our method achieves better performance in both quantitative and qualitative results with comparison to the state-of-the-art methods. Our network is a bottom-up instance segmentation method that could well preserve geometric details. Qingyu Li 0001, Lichao Mou, Yuansheng Hua, Yao Sun 0005, Pu Jin, Yilei Shi, Xiao Xiang Zhu 0001 |
IGARSS | 2 |
| 2020 | Event and Activity Recognition in Aerial Videos Using Deep Neural Networks and a New DatasetabstractUnmanned aerial vehicles (UAVs) are now widespread available. Yet the more UAVs there are in the skies, the more video data they create. It is unrealistic for humans to screen such big data and understand their contents. Hence methodological research on UAV video content understanding is of great importance. In this paper, we introduce a novel task of event recognition in unconstrained aerial videos in the remote sensing community and present a dataset for this task. Organized in a rich semantic taxonomy, the proposed dataset covers a wide range of events involving diverse environments and scales. We report results of plenty of deep networks in two ways: single-frame classification and video classification. The dataset and trained models can be downloaded from https://1cmou.github.io/ERA_Dataset/. Lichao Mou, Yuansheng Hua, Pu Jin, Xiao Xiang Zhu 0001 |
IGARSS | 1 |
| 2020 | A Novel Approach to Unsupervised Segmentation of Multitemporal VHR Images based on Deep LearningabstractVery-high-resolution (VHR) multi-temporal images are important in remote sensing to monitor the dynamics of the Earth surface. Image semantic segmentation classifies pixels and assigns them label from meaningful object groups. It has been extensively studied in context of single image analysis, however not explored for multi-temporal one. In this paper we propose to extend supervised semantic segmentation to the unsupervised joint segmentation of multi-temporal images. The proposed method processes multi-temporal images by separately feeding them to a deep network comprising of trainable convolutional layers. The training process does not involve any external label. Segmentation labels are obtained from argmax classification of the final layer. Multi-temporal segmentation labels and weights of the trainable layers are jointly optimized in iterations. We tested the method on a VHR dataset from Trento, Italy. Both quantitative and qualitative results demonstrated the effectiveness of the proposed approach. Sudipan Saha, Lichao Mou, Chunping Qiu, Xiao Xiang Zhu 0001, Francesca Bovolo, Lorenzo Bruzzone |
IGARSS | 2 |
| 2020 | Dual Adversarial Network for Unsupervised Ground/Satellite-to-Aerial Scene AdaptationabstractRecent domain adaptation work tends to obtain a uniformed representation in an adversarial manner through joint learning of the domain discriminator and feature generator. However, this domain adversarial approach could render sub-optimal performances due to two potential reasons: First, it might fail to consider the task at hand when matching the distributions between the domains. Second, it generally treats the source and target domain data in the same way. In our opinion, the source domain data which serves the feature adaption purpose should be supplementary, whereas the target domain data mainly needs to consider the task-specific classifier. Motivated by this, we propose a dual adversarial network for domain adaptation, where two adversarial learning processes are conducted iteratively, in correspondence with the feature adaptation and the classification task respectively. The efficacy of the proposed method is first demonstrated on Visual Domain Adaptation Challenge (VisDA) 2017 challenge, and then on two newly proposed Ground/Satellite-to-Aerial Scene adaptation tasks. For the proposed tasks, the data for the same scene is collected not only by the traditional camera on the ground, but also by satellite from the out space and unmanned aerial vehicle (UAV) at the high-altitude. Since the semantic gap between the ground/satellite scene and the aerial scene is much larger than that between ground scenes, the newly proposed tasks are more challenging than traditional domain adaptation tasks. The datasets/codes can be found at https://github.com/jianzhelin/DuAN. Jianzhe Lin, Lichao Mou, Tianze Yu, Xiao Xiang Zhu 0001, Z. Jane Wang 0001 |
ACM Multimedia | 2 |
| 2020 | Fusing Multiseasonal Sentinel-2 Imagery for Urban Land Cover Classification With Multibranch Residual Convolutional Neural NetworksabstractExploiting multitemporal Sentinel-2 images for urban land cover classification has become an important research topic, since these images have become globally available at relatively fine temporal resolution, thus offering great potential for large-scale land cover mapping. However, appropriate exploitation of the images needs to address problems such as cloud cover inherent to optical satellite imagery. To this end, we propose a simple yet effective decision-level fusion approach for urban land cover prediction from multiseasonal Sentinel-2 images, using the state-of-the-art residual convolutional neural networks (ResNet). We extensively tested the approach in a cross-validation manner over a seven-city study area in central Europe. Both quantitative and qualitative results demonstrated the superior performance of the proposed fusion approach over several baseline approaches, including observation- and feature-level fusion. Chunping Qiu, Lichao Mou, Michael Schmitt 0003, Xiao Xiang Zhu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2020 | Relation Network for Multilabel Aerial Image ClassificationabstractMultilabel classification plays a momentous role in perceiving intricate contents of an aerial image and triggers several related studies over the last years. However, most of them deploy few efforts in exploiting label relations, while such dependencies are crucial for making accurate predictions. Although an long short term memory (LSTM) layer can be introduced to modeling such label dependencies in a chain propagation manner, the efficiency might be questioned when certain labels are improperly inferred. To address this, we propose a novel aerial image multilabel classification network, attention-aware label relational reasoning network. Particularly, our network consists of three elemental modules: 1) a label-wise feature parcel learning module; 2) an attentional region extraction module; and 3) a label relational inference module. To be more specific, the label-wise feature parcel learning module is designed for extracting high-level label-specific features. The attentional region extraction module aims at localizing discriminative regions in these features without region proposal generation, yielding attentional label-specific features. The label relational inference module finally predicts label existences using label relations reasoned from outputs of the previous module. The proposed network is characterized by its capacities of extracting discriminative label-wise features and reasoning about label relations naturally and interpretably. In our experiments, we evaluate the proposed model on two multilabel aerial image data sets, of which one is newly produced. Quantitative and qualitative results on these two data sets demonstrate the effectiveness of our model. To facilitate progress in the multilabel aerial image classification, our produced data set will be made publicly available. Yuansheng Hua, Lichao Mou, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | Relation Matters: Relational Context-Aware Fully Convolutional Network for Semantic Segmentation of High-Resolution Aerial ImagesabstractMost current semantic segmentation approaches fall back on deep convolutional neural networks (CNNs). However, their use of convolution operations with local receptive fields causes failures in modeling contextual spatial relations. Prior works have sought to address this issue by using graphical models or spatial propagation modules in networks. But such models often fail to capture long-range spatial relationships between entities, which leads to spatially fragmented predictions. Moreover, recent works have demonstrated that channel-wise information also acts a pivotal part in CNNs. In this article, we introduce two simple yet effective network units, the spatial relation module, and the channel relation module to learn and reason about global relationships between any two spatial positions or feature maps, and then produce Relation-Augmented (RA) feature representations. The spatial and channel relation modules are general and extensible, and can be used in a plug-and-play fashion with the existing fully convolutional network (FCN) framework. We evaluate relation module-equipped networks on semantic segmentation tasks using two aerial image data sets, namely International Society for Photogrammetry and Remote Sensing (ISPRS) Vaihingen and Potsdam data sets, which fundamentally depend on long-range spatial relational reasoning. The networks achieve very competitive results, a mean F1score of 88.54% on the Vaihingen data set and a mean F1score of 88.01% on the Potsdam data set, bringing significant improvements over baselines. Lichao Mou, Yuansheng Hua, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | Nonlocal Graph Convolutional Networks for Hyperspectral Image ClassificationabstractOver the past few years making use of deep networks, including convolutional neural networks (CNNs) and recurrent neural networks (RNNs), classifying hyperspectral images has progressed significantly and gained increasing attention. In spite of being successful, these networks need an adequate supply of labeled training instances for supervised learning, which, however, is quite costly to collect. On the other hand, unlabeled data can be accessed in almost arbitrary amounts. Hence it would be conceptually of great interest to explore networks that are able to exploit labeled and unlabeled data simultaneously for hyperspectral image classification. In this article, we propose a novel graph-based semisupervised network called nonlocal graph convolutional network (nonlocal GCN). Unlike existing CNNs and RNNs that receive pixels or patches of a hyperspectral image as inputs, this network takes the whole image (including both labeled and unlabeled data) in. More specifically, a nonlocal graph is first calculated. Given this graph representation, a couple of graph convolutional layers are used to extract features. Finally, the semisupervised learning of the network is done by using a cross-entropy error over all labeled instances. Note that the nonlocal GCN is end-to-end trainable. We demonstrate in extensive experiments that compared with state-of-the-art spectral classifiers and spectral-spatial classification networks, the nonlocal GCN is able to offer competitive results and high-quality classification maps (with fine boundaries and without noisy scattered points of misclassification). Lichao Mou, Xiaoqiang Lu, Xuelong Li 0001, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | Learning to Pay Attention on Spectral Domain: A Spectral Attention Module-Based Convolutional Network for Hyperspectral Image ClassificationabstractOver the past few years, hyperspectral image classification using convolutional neural networks (CNNs) has progressed significantly. In spite of their effectiveness, given that hyperspectral images are of high dimensionality, CNNs can be hindered by their modeling of all spectral bands with the same weight, as probably not all bands are equally informative and predictive. Moreover, the usage of useless spectral bands in CNNs may even introduce noises and weaken the performance of networks. For the sake of boosting the representational capacity of CNNs for spectral-spatial hyperspectral data classification, in this work, we improve networks by discriminating the significance of different spectral bands. We design a network unit, which is termed as the spectral attention module, that makes use of a gating mechanism to adaptively recalibrate spectral bands by selectively emphasizing informative bands and suppressing less useful ones. We theoretically analyze and discuss why such a spectral attention module helps in a CNN for hyperspectral image classification. We demonstrate using extensive experiments that in comparison with state-of-the-art approaches, the spectral attention module-based convolutional networks are able to offer competitive results. Furthermore, this work sheds light on how a CNN interacts with spectral bands for the purpose of classification. Lichao Mou, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | Unsupervised Deep Joint Segmentation of Multitemporal High-Resolution ImagesabstractHigh/very-high-resolution (HR/VHR) multitemporal images are important in remote sensing to monitor the dynamics of the Earth's surface. Unsupervised object-based image analysis provides an effective solution to analyze such images. Image semantic segmentation assigns pixel labels from meaningful object groups and has been extensively studied in the context of single-image analysis, however not explored for multitemporal one. In this article, we propose to extend supervised semantic segmentation to the unsupervised joint semantic segmentation of multitemporal images. We propose a novel method that processes multitemporal images by separately feeding to a deep network comprising of trainable convolutional layers. The training process does not involve any external label, and segmentation labels are obtained from the argmax classification of the final layer. A novel loss function is used to detect object segments from individual images as well as establish a correspondence between distinct multitemporal segments. Multitemporal semantic labels and weights of the trainable layers are jointly optimized in iterations. We tested the method on three different HR/VHR data sets from Munich, Paris, and Trento, which shows the method to be effective. We further extended the proposed joint segmentation method for change detection (CD) and tested on a VHR multisensor data set from Trento. Sudipan Saha, Lichao Mou, Chunping Qiu, Xiao Xiang Zhu 0001, Francesca Bovolo, Lorenzo Bruzzone |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | Incorporating Metric Learning and Adversarial Network for Seasonal Invariant Change DetectionabstractChange detection by comparing two bitemporal images is one of the most fundamental challenges for dynamic monitoring of the Earth surface. In this article, we propose a metric learning-based generative adversarial network (GAN) (MeGAN) to automatically explore seasonal invariant features for pseudochange suppressing and real change detection. To achieve this purpose, a seasonal invariant term is introduced to maximally suppress pseudochanges, whereas the MeGAN explores the transition patterns between adjacent images in a self-learning fashion. Different from the previous works on bitemporal imagery change detection, the proposed MeGAN have the following contributions: 1) it automatically explores change patterns from the complex bitemporal background without human intervention and 2) it aims to maximally exclude pseudochanges from the seasonal transition term and map out real changes efficiently. To our best knowledge, this is the first time we incorporate the seasonal transition term and GAN for change detection between bitemporal images. At last, to demonstrate the robustness of the proposed method, we included two data sets which are the Google Earth data and the Landsat data, for bitemporal change detection and evaluation. The experimental results indicated that the proposed method is able to perform change detection with precision can be as high as 81% and 88% for the Google Earth and Landsat data set, respectively. Wenzhi Zhao, Lichao Mou, Jiage Chen, Yanchen Bo, William J. Emery |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | A Relation-Augmented Fully Convolutional Network for Semantic Segmentation in Aerial ScenesabstractMost current semantic segmentation approaches fall back on deep convolutional neural networks (CNNs). However, their use of convolution operations with local receptive fields causes failures in modeling contextual spatial relations. Prior works have sought to address this issue by using graphical models or spatial propagation modules in networks. But such models often fail to capture long-range spatial relationships between entities, which leads to spatially fragmented predictions. Moreover, recent works have demonstrated that channel-wise information also acts a pivotal part in CNNs. In this work, we introduce two simple yet effective network units, the spatial relation module and the channel relation module, to learn and reason about global relationships between any two spatial positions or feature maps, and then produce relation-augmented feature representations. The spatial and channel relation modules are general and extensible, and can be used in a plug-and-play fashion with the existing fully convolutional network (FCN) framework. We evaluate relation module-equipped networks on semantic segmentation tasks using two aerial image datasets, which fundamentally depend on long-range spatial relational reasoning. The networks achieve very competitive results, bringing significant improvements over baselines. Lichao Mou, Yuansheng Hua, Xiao Xiang Zhu 0001 |
CVPR | 1 |
| 2019 | Label Relation Inference for Multi-Label Aerial Image ClassificationabstractMulti-label aerial image classification is a challenging visual task and obtaining increasing attention recently. Most of the existing methods resort to training independent classifier for each label, while underlying label correlations are not fully exploited while making predictions. To this end, we propose an innovative inference network, which takes advantage of pairwise label relations to infer multiple object labels of a high-resolution aerial image. Specifically, we first employ a feature extraction module to extract high-level feature representations of an aerial image, and then, feed them into a relational inference module to predict the presence of each object label. We evaluate our network on the UCM multilabel dataset and experiment with various popular convolutional neural networks (CNNs) as the backbone of the feature extraction module. Experimental results demonstrate that the proposed network behaves superiorly in comparison with other existing methods. Yuansheng Hua, Lichao Mou, Xiao Xiang Zhu 0001 |
IGARSS | 2 |
| 2019 | Spatial Relational Reasoning in Networks for Improving Semantic Segmentation of Aerial ImagesabstractMost current semantic segmentation approaches rely on deep convolutional neural networks (CNNs). However, their use of convolution operations with local receptive fields causes failures in modeling contextual spatial relations. Prior works have tried to address this issue by using graphical models or spatial propagation modules in networks. But such models often fail to capture long-range spatial relationships between entities, which leads to spatially fragmented predictions. In this work, we introduce a simple yet effective network unit, the spatial relation module, to learn and reason about global relationships between any two spatial positions, and then produce relation-enhanced feature representations. The spatial relation module is general and extensible, and can be used in a plug-and-play fashion with the existing fully convolutional network (FCN) framework. We evaluate spatial relation module-equipped networks on semantic segmentation tasks using two aerial image datasets. The networks achieve very competitive results, bringing significant improvements over baselines. Lichao Mou, Yuansheng Hua, Xiao Xiang Zhu 0001 |
IGARSS | 1 |
| 2019 | R3-Net: A Deep Network for Multioriented Vehicle Detection in Aerial Images and VideosabstractVehicle detection is a significant and challenging task in aerial remote sensing applications. Most existing methods detect vehicles with regular rectangle boxes and fail to offer the orientation of vehicles. However, the orientation information is crucial for several practical applications, such as the trajectory and motion estimation of vehicles. In this paper, we propose a novel deep network, called a rotatable region-based residual network (R3-Net), to detect multioriented vehicles in aerial images and videos. More specially, R3-Net is utilized to generate rotatable rectangular target boxes in a half coordinate system. First, we use a rotatable region proposal network (R-RPN) to generate rotatable region of interests (R-RoIs) from feature maps produced by a deep convolutional neural network. Here, a proposed batch averaging rotatable anchor strategy is applied to initialize the shape of vehicle candidates. Next, we propose a rotatable detection network (R-DN) for the final classification and regression of the R-RoIs. In R-DN, a novel rotatable position-sensitive pooling is designed to keep the position and orientation information simultaneously while downsampling the feature maps of R-RoIs. In our model, R-RPN and R-DN can be trained jointly. We test our network on two open vehicle detection image data sets, namely, DLR 3K Munich Data set and VEDAI Data set, demonstrating the high precision and robustness of our method. In addition, further experiments on aerial videos show the good generalization capability of the proposed method and its potential for vehicle tracking in aerial videos. The demo video is available athttps://youtu.be/xCYD-tYudN0. Qingpeng Li, Lichao Mou, Qizhi Xu, Yun Zhang 0014, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | Learning Spectral-Spatial-Temporal Features via a Recurrent Convolutional Neural Network for Change Detection in Multispectral ImageryabstractChange detection is one of the central problems in earth observation and was extensively investigated over recent decades. In this paper, we propose a novel recurrent convolutional neural network (ReCNN) architecture, which is trained to learn a joint spectral-spatial-temporal feature representation in a unified framework for change detection in multispectral images. To this end, we bring together a convolutional neural network and a recurrent neural network into one end-to-end network. The former is able to generate rich spectral-spatial feature representations, while the latter effectively analyzes temporal dependence in bitemporal images. In comparison with previous approaches to change detection, the proposed network architecture possesses three distinctive properties: 1) it is end-to-end trainable, in contrast to most existing methods whose components are separately trained or computed; 2) it naturally harnesses spatial information that has been proven to be beneficial to change detection task; and 3) it is capable of adaptively learning the temporal dependence between multitemporal images, unlike most of the algorithms that use fairly simple operation like image differencing or stacking. As far as we know, this is the first time that a recurrent convolutional network architecture has been proposed for multitemporal remote sensing image analysis. The proposed network is validated on real multispectral data sets. Both visual and quantitative analyses of the experimental results demonstrate competitive performance in the proposed mode. Lichao Mou, Lorenzo Bruzzone, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2018 | LAHNet: A Convolutional Neural Network Fusing Low- and High-Level Features for Aerial Scene ClassificationabstractIn this paper, we proposed an innovative end-to-end convolutional neural network (CNN), which is trained to learn how to fuse multi-level features for aerial scene classification. Instead of using only coarse semantic features as conventional CNNs, we resort to first hierarchically extracting dense high-level features and then element-wise fusing them with low-level features to build a comprehensive feature representation, which contains not only high-level semantic information but also fine-grained low-level details, for scene classification. The network is evaluated on two broadly used aerial scene datasets, UCM and AID. The experimental results indicate that the proposed LAHNet performs superiorly compared to the existing benchmark methods. Furthermore, visualization of the fused features presents an intuitive illustration of the remarkable improvement. Yuansheng Hua, Lichao Mou, Xiao Xiang Zhu 0001 |
IGARSS | 2 |
| 2018 | Classification of Settlement Types from Tweets Using LDA and LSTMabstractLand use reflects the interrelation between the physically built environment and the activity patterns of people. It is indispensable information for decision-makes, but up-to-date and accurate land use information is often absent. Unlike approaches that make use of remote sensing data, in this work, we are interested in a novel data source, tweets, and explore its potential for land use classification in urban areas. Specifically, we propose a general framework for classifying settlement land-use types by extracting location, time, quantity and text features of twitter data. To do so, we apply latent Dirichlet allocation (LDA) and long short-term memory (LSTM) and then combines those features with spatial-temporal feature using Fused SVM and a two-stream convolutional neural network (CNN) for classification. For the case of classifying individual tweets by the land-use classes relevant in this study - residential, non-residential and mixed usage -, we reach overall accuracy (OA), average accuracy (AA), and Kappa coefficient with 72.35%, 73.76%, and 58.43%, respectively. As for the case of classifying block settlement types, we reach 61.90%, 63.33%, and 42.84%, respectively. Hannes Taubenböck, Lichao Mou, Xiao Xiang Zhu 0001 |
IGARSS | 3 |
| 2018 | Hierarchical Region Based Convolution Neural Network for Multiscale Object Detection in Remote Sensing ImagesabstractIn this paper, we propose a novel Faster R-CNN based method to detect multiscale objects in very high resolution optical remote sensing images. Firstly, a pre-trained CNN is used to extract features from an input image; and then a set of object candidates are generated. To efficiently detect objects with various scales, we design a hierarchical selective filtering (HSF) layer to map features in different scales to the same scale space. The HSF layer can be applied on both region proposal and the subsequent detection network. More importantly, it can be plugged into Faster R-CNN network without modifying its architecture, meanwhile boosting the performance on detecting objects with varying scales. The proposed model can be trained in an end-to-end manner. We test our network on three datasets containing different multiscale objects, including airplanes, ships and buildings, which are collected from Google Earth images and GaoFen-2 images. Experiments demonstrate high precision and robustness of our method. Qingpeng Li, Lichao Mou, Kaiyu Jiang, Qingjie Liu 0001, Yunhong Wang 0001, Xiao Xiang Zhu 0001 |
IGARSS | 2 |
| 2018 | A Recurrent Convolutional Neural Network for Land Cover Change Detection in Multispectral ImagesabstractIn this paper, we propose a novel network architecture, a recurrent convolutional neural network, which is trained to learn a joint spectral-spatial-temporal feature representation in a unified framework for change detection of multispectral images. To this end, we bring together a convolutional neural network (CNN) and a recurrent neural network (RNN) into one end-to-end network. The former is able to generate rich spectral-spatial feature representations while the latter effectively analyzes temporal dependency in bi-temporal images. Although both CNN and RNN are well-established techniques for remote sensing applications, to the best of our knowledge, we are the first to combine them for multitemporal data analysis in the remote sensing community. Both visual and quantitative analysis of experimental results demonstrates competitive performance in the proposed mode. Lichao Mou, Xiao Xiang Zhu 0001 |
IGARSS | 1 |
| 2018 | Feature Importance Analysis of Sentinel-2 Imagery for Large-Scale Urban Local Climate Zone ClassificationabstractThis paper evaluates different spectral-spatial features that can be extracted from Sentinel-2 imagery regarding their relevance for discriminating different Local Climate Zone (LCZ) classes. The features include spectral reflectance, spectral indices, Morphological Profiles (MPs), as well as Global Urban Footprint (GUF), the Open Street Map layers buildings and land use, and their combinations. Using a residual convolutional neural network (ResNet), a systematic analysis of feature importance is performed with a manually generated dataset distributed in Europe. The results of this evaluation are meant to provide guidance about the choice of both spectral and spatial features for the task of LCZ classification on a global scale. The results show that GUF and OSM can contribute to the classification performance, and ResNet relies less on additional features with the highest accuracy provided by the reflectance only. Chunping Qiu, Michael Schmitt 0003, Pedram Ghamisi, Lichao Mou, Xiao Xiang Zhu 0001 |
IGARSS | 4 |
| 2018 | Identifying Corresponding Patches in SAR and Optical Images With a Pseudo-Siamese CNNabstractIn this letter, we propose a pseudo-siamese convolutional neural network architecture that enables to solve the task of identifying corresponding patches in very high-resolution optical and synthetic aperture radar (SAR) remote sensing imagery. Using eight convolutional layers each in two parallel network streams, a fully connected layer for the fusion of the features learned in each stream, and a loss function based on binary cross entropy, we achieve a one-hot indication if two patches correspond or not. The network is trained and tested on an automatically generated data set that is based on a deterministic alignment of SAR and optical imagery via previously reconstructed and subsequently coregistered 3-D point clouds. The satellite images, from which the patches comprising our data set are extracted, show a complex urban scene containing many elevated objects (i.e., buildings), thus providing one of the most difficult experimental environments. The achieved results show that the network is able to predict corresponding patches with high accuracy, thus indicating great potential for further development toward a generalized multisensor key-point matching procedure. Lloyd H. Hughes, Michael Schmitt 0003, Lichao Mou, Yuanyuan Wang 0002, Xiao Xiang Zhu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2018 | HSF-Net: Multiscale Deep Feature Embedding for Ship Detection in Optical Remote Sensing ImageryabstractShip detection is an important and challenging task in remote sensing applications. Most methods utilize specially designed hand-crafted features to detect ships, and they usually work well only on one scale, which lack generalization and impractical to identify ships with various scales from multiresolution images. In this paper, we propose a novel deep feature-based method to detect ships in very high-resolution optical remote sensing images. In our method, a regional proposal network is used to generate ship candidates from feature maps produced by a deep convolutional neural network. To efficiently detect ships with various scales, a hierarchical selective filtering layer is proposed to map features in different scales to the same scale space. The proposed method is an end-to-end network that can detect both inshore and offshore ships ranging from dozens of pixels to thousands. We test our network on a large ship data set which will be released in the future, consisting of Google Earth images, GaoFen-2 images, and unmanned aerial vehicle data. Experiments demonstrate high precision and robustness of our method. Further experiments on aerial images show its good generalization to unseen scenes. Qingpeng Li, Lichao Mou, Qingjie Liu 0001, Yunhong Wang 0001, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2018 | Unsupervised Spectral-Spatial Feature Learning via Deep Residual Conv-Deconv Network for Hyperspectral Image ClassificationabstractSupervised approaches classify input data using a set of representative samples for each class, known as training samples. The collection of such samples is expensive and time demanding. Hence, unsupervised feature learning, which has a quick access to arbitrary amounts of unlabeled data, is conceptually of high interest. In this paper, we propose a novel network architecture, fully Conv-Deconv network, for unsupervised spectral-spatial feature learning of hyperspectral images, which is able to be trained in an end-to-end manner. Specifically, our network is based on the so-called encoder-decoder paradigm, i.e., the input 3-D hyperspectral patch is first transformed into a typically lower dimensional space via a convolutional subnetwork (encoder), and then expanded to reproduce the initial data by a deconvolutional subnetwork (decoder). However, during the experiment, we found that such a network is not easy to be optimized. To address this problem, we refine the proposed network architecture by incorporating: 1) residual learning and 2) a new unpooling operation that can use memorized max-pooling indexes. Moreover, to understand the “black box,” we make an in-depth study of the learned feature maps in the experimental analysis. A very interesting discovery is that some specific “neurons” in the first residual block of the proposed network own good description power for semantic visual patterns in the object level, which provide an opportunity to achieve “free” object detection. This paper, for the first time in the remote sensing community, proposes an end-to-end fully Conv-Deconv network for unsupervised spectral-spatial feature learning. Moreover, this paper also introduces an in-depth investigation of learned features. Experimental results on two widely used hyperspectral data, Indian Pines and Pavia University, demonstrate competitive performance obtained by the proposed methodology compared with other studied approaches. Lichao Mou, Pedram Ghamisi, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2018 | Corrections to "Deep Recurrent Neural Networks for Hyperspectral Image Classification"abstractHere, we correct some errors caused by a programming bug (a data type error) in overall accuracies (OAs) reported in[1]. The corrected OAs are underlined and shown in bold inTables I–III. Lichao Mou, Pedram Ghamisi, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2018 | Vehicle Instance Segmentation From Aerial Image and Video Using a Multitask Learning Residual Fully Convolutional NetworkabstractObject detection and semantic segmentation are two main themes in object retrieval from high-resolution remote sensing images, which have recently achieved remarkable performance by surfing the wave of deep learning and, more notably, convolutional neural networks. In this paper, we are interested in a novel, more challenging problem of vehicle instance segmentation, which entails identifying, at a pixel level, where the vehicles appear as well as associating each pixel with a physical instance of a vehicle. In contrast, vehicle detection and semantic segmentation each only concern one of the two. We propose to tackle this problem with a semantic boundary-aware multitask learning network. More specifically, we utilize the philosophy of residual learning to construct a fully convolutional network that is capable of harnessing multilevel contextual feature representations learned from different residual blocks. We theoretically analyze and discuss why residual networks can produce better probability maps for pixelwise segmentation tasks. Then, based on this network architecture, we propose a unified multitask learning network that can simultaneously learn two complementary tasks, namely, segmenting vehicle regions and detecting semantic boundaries. The latter subproblem is helpful for differentiating “touching” vehicles that are usually not correctly separated into instances. Currently, data sets with a pixelwise annotation for vehicle extraction are the ISPRS data set and the IEEE GRSS DFC2015 data set over Zeebrugge, which specializes in a semantic segmentation. Therefore, we built a new, more challenging data set for vehicle instance segmentation, called the Busy Parking Lot Unmanned Aerial Vehicle Video data set, and we make our data set available at http://www.sipeo.bgu.tum.de/downloads so that it can be used to benchmark future vehicle instance segmentation algorithms. Lichao Mou, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2017 | Identifying corresponding patches in SAR and optical imagery with a convolutional neural networkabstractIn this paper, we investigate making use of a convolutional neural network (CNN) to solve the task of identifying corresponding patches in very high resolution (VHR) optical and SAR imagery of complicated urban scenery. By doing so, the binary decision function is learnt directly from automatically generated training data and does not resort to any hand-crafted features. First evaluations show great potential for further studies towards a generalized multi-sensor matching procedure. Lichao Mou, Michael Schmitt 0003, Yuanyuan Wang 0002, Xiao Xiang Zhu 0001 |
IGARSS | 1 |
| 2017 | Fully conv-deconv network for unsupervised spectral-spatial feature extraction of hyperspectral imagery via residual learningabstractSupervised approaches classify input data using a set of representative samples for each class, known as training samples. The collection of such samples are expensive and time-demanding. Hence, unsupervised feature learning, which has a quick access to arbitrary amount of unlabeled data, is conceptually of high interest. In this paper, we propose a novel network architecture, fully Conv-Deconv network with residual learning, for unsupervised spectral-spatial feature learning of hyperspectral images, which is able to be trained in an end-to-end manner. Specifically, our network is based on the so-called encoder-decoder paradigm, i.e., the input 3D hyperspectral patch is first transformed into a typically lower-dimensional space via a convolutional sub-network (encoder), and then expanded to reproduce the initial data by a deconvolutional sub-network (decoder). Experimental results on the Pavia University hyperspectral data set demonstrate competitive performance obtained by the proposed methodology compared to other studied approaches. Lichao Mou, Pedram Ghamisi, Xiao Xiang Zhu 0001 |
IGARSS | 1 |
| 2017 | Deep Recurrent Neural Networks for Hyperspectral Image ClassificationabstractIn recent years, vector-based machine learning algorithms, such as random forests, support vector machines, and 1-D convolutional neural networks, have shown promising results in hyperspectral image classification. Such methodologies, nevertheless, can lead to information loss in representing hyperspectral pixels, which intrinsically have a sequence-based data structure. A recurrent neural network (RNN), an important branch of the deep learning family, is mainly designed to handle sequential data. Can sequence-based RNN be an effective method of hyperspectral image classification? In this paper, we propose a novel RNN model that can effectively analyze hyperspectral pixels as sequential data and then determine information categories via network reasoning. As far as we know, this is the first time that an RNN framework has been proposed for hyperspectral image classification. Specifically, our RNN makes use of a newly proposed activation function, parametric rectified tanh (PRetanh), for hyperspectral sequential data analysis instead of the popular tanh or rectified linear unit. The proposed activation function makes it possible to use fairly high learning rates without the risk of divergence during the training procedure. Moreover, a modified gated recurrent unit, which uses PRetanh for hidden representation, is adopted to construct the recurrent layer in our network to efficiently process hyperspectral data and reduce the total number of parameters. Experimental results on three airborne hyperspectral images suggest competitive performance in the proposed mode. In addition, the proposed network architecture opens a new window for future research, showcasing the huge potential of deep recurrent networks for hyperspectral data analysis. Lichao Mou, Pedram Ghamisi, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2016 | Spatiotemporal scene interpretation of space videos via deep neural network and tracklet analysisabstractSpaceborne remote sensing videos are becoming indispensable resources, opening up opportunities for new remote sensing applications. To exploit this new type of data, we need sophisticated algorithms for semantic scene interpretation. The main difficulties are: 1) Due to the relatively poor spatial resolution of the video acquired from space, moving objects, like cars, are very difficult to detect, not to mention track; 2) camera movement handicaps scene interpretation. To address these challenges, in this paper we propose a novel framework that fuses multispectral images and space videos for spatiotemporal analysis. Taking a multispectral image and a spaceborne video as input, an innovative deep neural network is proposed to fuse them in order to achieve a fine-resolution spatial scene labeling map. Moreover, a sophisticated approach is proposed to analyze activities and estimate traffic density from 150,000+ tracklets produced by a Kanade-Lucas-Tomasi keypoint tracker. The proposed framework is validated using data provided for the 2016 IEEE GRSS data fusion contest, including a video acquired from the International Space Station and a DEIMOS-2 multispectral image. Both visual and quantitative analysis of the experimental results demonstrates the effectiveness of our approach. Lichao Mou, Xiao Xiang Zhu 0001 |
IGARSS | 1 |
| 2016 | Video parsing via spatiotemporally analysis with images
Xuelong Li 0001, Lichao Mou, Xiaoqiang Lu |
Multim. Tools Appl. | 2 |
| 2015 | Scene Parsing From an MAP PerspectiveabstractScene parsing is an important problem in the field of computer vision. Though many existing scene parsing approaches have obtained encouraging results, they fail to overcome within-category inconsistency and intercategory similarity of superpixels. To reduce the aforementioned problem, a novel method is proposed in this paper. The proposed approach consists of three main steps: 1) posterior category probability density function (PDF) is learned by an efficient low-rank representation classifier (LRRC); 2) prior contextual constraint PDF on the map of pixel categories is learned by Markov random fields; and 3) final parsing results are yielded up to the maximum a posterior process based on the two learned PDFs. In this case, the nature of being both dense for within-category affinities and almost zeros for intercategory affinities is integrated into our approach by using LRRC to model the posterior category PDF. Meanwhile, the contextual priori generated by modeling the prior contextual constraint PDF helps to promote the performance of scene parsing. Experiments on benchmark datasets show that the proposed approach outperforms the state-of-the-art approaches for scene parsing. Xuelong Li 0001, Lichao Mou, Xiaoqiang Lu |
IEEE Trans. Cybern. | 2 |
| 2015 | Semi-Supervised Multitask Learning for Scene RecognitionabstractScene recognition has been widely studied to understand visual information from the level of objects and their relationships. Toward scene recognition, many methods have been proposed. They, however, encounter difficulty to improve the accuracy, mainly due to two limitations: 1) lack of analysis of intrinsic relationships across different scales, say, the initial input and its down-sampled versions and 2) existence of redundant features. This paper develops a semi-supervised learning mechanism to reduce the above two limitations. To address the first limitation, we propose a multitask model to integrate scene images of different resolutions. For the second limitation, we build a model of sparse feature selection-based manifold regularization (SFSMR) to select the optimal information and preserve the underlying manifold structure of data. SFSMR coordinates the advantages of sparse feature selection and manifold regulation. Finally, we link the multitask model and SFSMR, and propose the semi-supervised learning method to reduce the two limitations. Experimental results report the improvements of the accuracy in scene recognition. Xiaoqiang Lu, Xuelong Li 0001, Lichao Mou |
IEEE Trans. Cybern. | 3 |
| 2015 | Scene Recognition by Manifold Regularized Deep Learning ArchitectureabstractScene recognition is an important problem in the field of computer vision, because it helps to narrow the gap between the computer and the human beings on scene understanding. Semantic modeling is a popular technique used to fill the semantic gap in scene recognition. However, most of the semantic modeling approaches learn shallow, one-layer representations for scene recognition, while ignoring the structural information related between images, often resulting in poor performance. Modeled after our own human visual system, as it is intended to inherit humanlike judgment, a manifold regularized deep architecture is proposed for scene recognition. The proposed deep architecture exploits the structural information of the data, making for a mapping between visible layer and hidden layer. By the proposed approach, a deep architecture could be designed to learn the high-level features for scene recognition in an unsupervised fashion. Experiments on standard data sets show that our method outperforms the state-of-the-art used for scene recognition. Yuan Yuan 0001, Lichao Mou, Xiaoqiang Lu |
IEEE Trans. Neural Networks Learn. Syst. | 2 |