EDBT 2026 Demo / reviewers in the wild / expert
Yuansheng Hua
dblp:224/0267
· DBLP profile ↗
25ranked-venue papers
7as first author
16since 2021 · last 2024
0000-0001-9238-2920ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 24 · 7 first-author · 16 since 2021Artificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Referring Image Segmentation for Remote Sensing DataabstractIn this paper, we present a new task: referring image segmentation for remote sensing data, which targets segmenting out specific objects referred to by natural language. Due to the absence of a dataset for this task, we construct a dataset based on the SkyScapes dataset. Our dataset is designed with linguistically structured expressions that focus on object categories, attributes, and spatial relationships, enabling the generation of binary masks from semantic segmentation maps. To benchmark this task, we evaluate and compare the performance of three different convolutional neural network (CNN)-based methods and a Transformer-based method. Experimental results provide valuable insights into the adaptability of these methods to remote sensing data, highlighting the potential of our dataset as a resource for the remote sensing community to further explore vision-language tasks. Zhenghang Yuan, Lichao Mou, Yuansheng Hua, Xiao Xiang Zhu 0001 |
IGARSS | 3 |
| 2024 | A Review of Building Extraction From Remote Sensing Imagery: Geometrical Structures and Semantic AttributesabstractIn the remote sensing community, extracting buildings from remote sensing imagery has triggered great interest. While many studies have been conducted, a comprehensive review of these approaches that are applied to optical and synthetic aperture radar (SAR) imagery is still lacking. Therefore, we provide an in-depth review of both early efforts and recent advances, which are aimed at extracting geometrical structures or semantic attributes of buildings, including building footprint generation, building facade segmentation, roof segment and superstructure segmentation, building height retrieval, building type classification, building change detection, and annotation data correction. Furthermore, a list of corresponding benchmark datasets is given. Finally, challenges and outlooks of existing approaches as well as promising applications are discussed to enhance comprehension within this realm of research. Qingyu Li 0001, Lichao Mou, Yao Sun 0005, Yuansheng Hua, Yilei Shi, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | RRSIS: Referring Remote Sensing Image SegmentationabstractLocalizing desired objects from remote sensing images is of great use in practical applications. Referring image segmentation, which aims at segmenting out the objects to which a given expression refers, has been extensively studied in natural images. However, almost no research attention is given to this task of remote sensing imagery. Considering its potential for real-world applications, in this paper, we introduce referring remote sensing image segmentation (RRSIS) to fill in this gap and make some insightful explorations. Specifically, we create a new dataset, called RefSegRS, for this task, enabling us to evaluate different methods. Afterward, we benchmark referring image segmentation methods of natural images on the RefSegRS dataset and find that these models show limited efficacy in detecting small and scattered objects. To alleviate this issue, we propose a language-guided cross-scale enhancement (LGCE) module that utilizes linguistic features to adaptively enhance multi-scale visual features by integrating both deep and shallow features. The proposed dataset, benchmarking results, and the designed LGCE module provide insights into the design of a better RRSIS model. The dataset and code will be available at https://gitlab.lrz.de/ai4eo/reasoning/rrsis. Zhenghang Yuan, Lichao Mou, Yuansheng Hua, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Function Assignment of Plastics based on Hyperspectral Satellite Images and High-Resolution Data Using Deep Learning AlgorithmsabstractPlastic pollution is becoming an increasingly prominent problem and the function of plastics determines whether they need to be recycled or not. In order to explore the possibility of using satellite imagery to classify the functionality of plastics, this study proposes a two-stage workflow: firstly, a classification map is obtained based on hyperspectral satellite imagery to generate plastic types, and then using these identified plastic coverage areas, a deep learning algorithm is used to assign functionality to these classified plastic areas based on sentinel-2 imagery. By comparing five leading-edge image classification models, classification accuracies of up to 74% were achieved, demonstrating the feasibility of using deep learning models trained on satellite images to identify plastic features. Shanyu Zhou, Lichao Mou, Lixian Zhang 0002, Yuansheng Hua, Hermann Kaufmann 0001, Xiao Xiang Zhu 0001 |
IGARSS | 4 |
| 2022 | Semantic Segmentation of Remote Sensing Images With Sparse AnnotationsabstractTraining convolutional neural networks (CNNs) for very high-resolution images requires a large quantity of high-quality pixel-level annotations, which is extremely labor-intensive and time-consuming to produce. Moreover, professional photograph interpreters might have to be involved in guaranteeing the correctness of annotations. To alleviate such a burden, we propose a framework for semantic segmentation of aerial images based on incomplete annotations, where annotators are asked to label a few pixels with easy-to-draw scribbles. To exploit these sparse scribbled annotations, we propose the FEature and Spatial relaTional regulArization (FESTA) method to complement the supervised task with an unsupervised learning signal that accounts for neighborhood structures both in spatial and feature terms. For the evaluation of our framework, we perform experiments on two remote sensing image segmentation data sets involving aerial and satellite imagery, respectively. Experimental results demonstrate that the exploitation of sparse annotations can significantly reduce labeling costs, while the proposed method can help improve the performance of semantic segmentation when training on such annotations. The sparse labels and codes are publicly available for reproducibility purposes.https://github.com/Hua-YS/Semantic-Segmentation-with-Sparse-Labels Yuansheng Hua, Diego Marcos, Lichao Mou, Xiao Xiang Zhu 0001, Devis Tuia |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | MultiScene: A Large-Scale Dataset and Benchmark for Multiscene Recognition in Single Aerial Images
Yuansheng Hua, Lichao Mou, Pu Jin, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | FuTH-Net: Fusing Temporal Relations and Holistic Features for Aerial Video ClassificationabstractUnmanned aerial vehicles (UAVs) are now widely applied to data acquisition due to its low cost and fast mobility. With the increasing volume of aerial videos, the demand for automatically parsing these videos is surging. To achieve this, current research mainly focuses on extracting a holistic feature with convolutions along both spatial and temporal dimensions. However, these methods are limited by small temporal receptive fields and cannot adequately capture long-term temporal dependencies that are important for describing complicated dynamics. In this article, we propose a novel deep neural network, termed Fusing Temporal relations and Holistic features for aerial video classification (FuTH-Net), to model not only holistic features but also temporal relations for aerial video classification. Furthermore, the holistic features are refined by the multiscale temporal relations in a novel fusion module for yielding more discriminative video representations. More specially, FuTH-Net employs a two-pathway architecture: 1) a holistic representation pathway to learn a general feature of both frame appearances and short-term temporal variations and 2) a temporal relation pathway to capture multiscale temporal relations across arbitrary frames, providing long-term temporal dependencies. Afterward, a novel fusion module is proposed to spatiotemporally integrate the two features learned from the two pathways. Our model is evaluated on two aerial video classification datasets, ERA and Drone-Action, and achieves the state-of-the-art results. This demonstrates its effectiveness and good generalization capacity across different recognition tasks (event classification and human action recognition). To facilitate further research, we release the code athttps://gitlab.lrz.de/ai4eo/reasoning/futh-net. Pu Jin, Lichao Mou, Yuansheng Hua, Gui-Song Xia, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Building Footprint Generation Through Convolutional Neural Networks With Attraction Field RepresentationabstractBuilding footprint generation is a vital task in a wide range of applications, including, to name a few, land use management, urban planning and monitoring, and geographical database updating. Most existing approaches addressing this problem fall back on convolutional neural networks (CNNs) to learn semantic masks of buildings. However, one limitation of their results is blurred building boundaries. To address this, we propose to learn attraction field representation for building boundaries, which is capable of providing an enhanced representation power. Our method comprises two elemental modules: an Img2AFM module and an AFM2Mask module. More specifically, the former aims at learning an attraction field representation conditioned on an input image, which is capable of enhancing building boundaries and suppressing the background. The latter module predicts segmentation masks of buildings using the learned attraction field map. The proposed method is evaluated on three datasets with different spatial resolutions: the ISPRS dataset, the INRIA dataset, and the Planet dataset. From experimental results, we find that the proposed framework can well preserve geometric shapes and sharp boundaries of buildings, which brings significant improvements over other competitors. The trained model and code are available at https://github.com/lqycrystal/AFM_building. Qingyu Li 0001, Lichao Mou, Yuansheng Hua, Yilei Shi, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Detecting Changes by Learning No Changes: Data-Enclosing-Ball Minimizing Autoencoders for One-Class Change Detection in Multispectral ImageryabstractChange detection is a long-standing and challenging problem in remote sensing. Very often, features about changes are difficult to model beforehand, thus making the collection of changed samples a challenging task. In comparison, it is much easier to collect numerous no-change samples. It is possible to define a change detection approach by using only easily available annotated no-change samples, which we henceforth call one-class change detection. Autoencoder networks being trained on no-change data are natural candidates for addressing this task due to their superior performance as compared to other one-class classification models. However, the autoencoders usually suffer from the problem of overgeneralization, i.e., they tend to generalize too well, thus risking properly reconstructing changed samples. In this paper, we propose a novel data-enclosing-ball minimizing autoencoder (DebM-AE) that is trained with dual objectives—a reconstruction error criterion and a minimum volume criterion. The network learns a compact latent space, where encodings of no-change samples have low intra-class variance, which as counter part has the identification of changed instances. We conducted extensive experiments on three real-world data sets. Results demonstrate advantages of the proposed method over other competitors. We make our data and code publicly available1. Lichao Mou, Yuansheng Hua, Sudipan Saha, Francesca Bovolo, Lorenzo Bruzzone, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Deep Reinforcement Learning for Band Selection in Hyperspectral Image ClassificationabstractBand selection refers to the process of choosing the most relevant bands in a hyperspectral image. By selecting a limited number of optimal bands, we aim at speeding up model training, improving accuracy, or both. It reduces redundancy among spectral bands while trying to preserve the original information of the image. By now, many efforts have been made to develop unsupervised band selection approaches, of which the majorities are heuristic algorithms devised by trial and error. In this article, we are interested in training an intelligent agent that, given a hyperspectral image, is capable of automatically learning policy to select an optimal band subset without any hand-engineered reasoning. To this end, we frame the problem of unsupervised band selection as a Markov decision process, propose an effective method to parameterize it, and finally solve the problem by deep reinforcement learning. Once the agent is trained, it learns a band-selection policy that guides the agent to sequentially select bands by fully exploiting the hyperspectral image and previously picked bands. Furthermore, we propose two different reward schemes for the environment simulation of deep reinforcement learning and compare them in experiments. This, to the best of our knowledge, is the first study that explores a deep reinforcement learning model for hyperspectral image analysis, thus opening a new door for future research and showcasing the great potential of deep reinforcement learning in remote sensing applications. Extensive experiments are carried out on four hyperspectral data sets, and experimental results demonstrate the effectiveness of the proposed method. The code is publicly available. Lichao Mou, Sudipan Saha, Yuansheng Hua, Francesca Bovolo, Lorenzo Bruzzone, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | CG-Net: Conditional GIS-Aware Network for Individual Building Segmentation in VHR SAR ImagesabstractObject retrieval and reconstruction from very-high-resolution (VHR) synthetic aperture radar (SAR) images are of great importance for urban SAR applications, yet highly challenging due to the complexity of SAR data. This article addresses the issue of individual building segmentation from a single VHR SAR image in large-scale urban areas. To achieve this, we introduce building footprints from geographic information system (GIS) data as a complementary information and propose a novel conditional GIS-aware network (CG-Net). The proposed model learns multilevel visual features and employs building footprints to normalize the features for predicting building masks in the SAR image. We validate our method using a high-resolution spotlight TerraSAR-X image collected over Berlin. Experimental results show that the proposed CG-Net effectively brings improvements with variant backbones. We further compare two representations of building footprints, namely, complete building footprints and sensor-visible footprint segments, for our task, and conclude that the use of the former leads to better segmentation results. Moreover, we investigate the impact of inaccurate GIS data on our CG-Net, and this study shows that CG-Net is robust against positioning errors in the GIS data. In addition, we propose an approach of ground truth generation of buildings from an accurate digital elevation model (DEM), which can be used to generate large-scale SAR image data sets. The segmentation results can be applied to reconstruct 3-D building models at level-of-detail (LoD) 1, which is demonstrated in our experiments. Yao Sun 0005, Yuansheng Hua, Lichao Mou, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | SCIDA: Self-Correction Integrated Domain Adaptation From Single- to Multi-Label Aerial ImagesabstractMost publicly available datasets for image classification are with single labels, while images are inherently multilabeled in our daily life. Such an annotation gap makes many pretrained single-label classification models fail in practical scenarios. For aerial images, this annotation issue is more concerned: Aerial data naturally cover a relatively large land area with multiple labels, while annotated aerial datasets currently publicly available (e.g., UCM and AID) are single-labeled. As manually annotating multilabel aerial images (MAIs) would be time-/ labor-consuming, we propose a novel self-correction integrated domain adaptation (SCIDA) method for automatic multilabel learning. SCIDA is weakly supervised, i.e., automatically learning the multilabel image classification model from using massive, publicly available single-label images. To achieve this goal, we propose a novel labelwise self-correction (LWC) module to better explore underlying label correlations. This module also makes the unsupervised domain adaptation (UDA) from single-label to multilabel data possible. For model training, the proposed method uses single-label information yet requires no prior knowledge of multilabeled data and predicts labels for MAIs. Through extensive evaluations, the proposed model, which is trained with single-labeled MAI-AID-s and MAI-UCM-s datasets, achieves much better performances than comparative methods on our collected multiscene aerial image dataset. The code and data are available on GitHub (https://github.com/Ryan315/Single2multi-DA). Tianze Yu, Jianzhe Lin, Lichao Mou, Yuansheng Hua, Xiao Xiang Zhu 0001, Z. Jane Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | Unconstrained Aerial Scene Recognition with Deep Neural Networks and a New DatasetabstractAerial scene recognition is a fundamental research problem in interpreting high-resolution aerial imagery. Over the past few years, most studies focus on classifying an image into one scene category, while in real-world scenarios, it is more often that a single image contains multiple scenes. Therefore, in this paper, we investigate a more practical yet underexplored task-multi-scene recognition in single images. To this end, we create a large-scale dataset, called Mul-tiScene dataset, composed of 100,000 unconstrained images each with multiple labels from 36 different scenes. Among these images, 14,000 of them are manually interpreted and assigned ground-truth labels, while the remaining images are provided with crowdsourced labels, which are generated from low-cost but noisy OpenStreetMap (OSM) data. By doing so, our dataset allows two branches of studies: 1) developing novel CNNs for multi-scene recognition and 2) learning with noisy labels. We experiment with extensive baseline models on our dataset to offer a benchmark for multi-scene recognition in single images. Aiming to expedite further researches, we will make our dataset and pre-trained models available11https://github.com/Hua-YS/Multi-Scene-Recognition. Yuansheng Hua, Lichao Mou, Pu Jin, Xiao Xiang Zhu 0001 |
IGARSS | 1 |
| 2021 | Temporal Relations Matter: A Two-Pathway Network for Aerial Video RecognitionabstractWith the increasing volume of aerial videos, the demand for automatically parsing these videos is surging. To achieve this, current researches mainly focus on extracting a holistic feature with convolutions along both spatial and temporal dimensions. However, these methods are limited by small temporal receptive fields and cannot adequately capture long-term temporal dependencies which are important for describing complicated dynamics. In this paper, we propose a novel two-pathway network to model not only holistic features, but also temporal relations for aerial video classification. More specially, our model employs a two-pathway architecture: (1) a holistic representation pathway to learn a general feature of frame appearances and short-term temporal variations and (2) a temporal relation pathway to capture multi-scale temporal relations across arbitrary frames, providing long-term temporal dependencies. Our model is evaluated on event recognition dataset, ERA, and achieves the state-of-the-art results. This demonstrates its effectiveness and good generalization capacity. Pu Jin, Lichao Mou, Yuansheng Hua, Gui-Song Xia, Xiao Xiang Zhu 0001 |
IGARSS | 3 |
| 2021 | Improving Land Cover Classification with a Shift-Invariant Center-Focusing Convolutional Neural NetworkabstractConvolutional neural networks (CNNs) are widely employed in remote sensing community. The CNN-based, also known as patch-based land cover classification method has gained increasing attention. However, this method very often requires the aid of post-processing, otherwise it is difficult to obtain accurate boundaries separating different land cover classes. In this paper, we discuss the reason of this phenomenon and propose a shift-invariant center-focusing (SICF) network to deliver more accurate boundaries to improve the patch-based land cover classification. The principle of SICF is calculating the class score from a center-focusing area based on a shift-invariant feature extraction module to calibrate prediction. We employ three modern CNNs to build corresponding SICF networks, the evaluation results indicate that compared with the conventional CNNs, the improvements made by SICF for delivering accurate boundaries in land cover classification are significant. Yuansheng Hua, Lichao Mou, Xiao Xiang Zhu 0001 |
IGARSS | 2 |
| 2021 | Conditional GIS-Aware Network for Individual Building Segmentation in a VHR SAR ImageabstractIn this paper, we propose a network for individual building segmentation from a single VHR SAR image. The proposed network employs building footprints from GIS data in learning multi-level visual features to predict building masks in the SAR image. Experimental results over Berlin show that the proposed network effectively brings improvements with variant backbones. In addition, we propose an approach for generating building labels from an accurate digital elevation model (DEM), which can be used to generate large-scale SAR image datasets. Yao Sun 0005, Yuansheng Hua, Lichao Mou, Xiao Xiang Zhu 0001 |
IGARSS | 2 |
| 2020 | Learning Multi-Label Aerial Image Classification Under Label Noise: A Regularization Approach Using Word EmbeddingsabstractTraining deep neural networks requires well-annotated datasets. However, real world datasets are often noisy, especially in a multi-label scenario, i.e. where each data point can be attributed to more than one class. To this end, we propose a regularization method to learn multi-label classification networks from noisy data. This regularization is based on the assumption that semantically close classes are more likely to appear together in a given image. Hereby, we encode label correlations with prior knowledge and regularize noisy network predictions using label correlations. To evaluate its effectiveness, we perform experiments on a mutli-label aerial image dataset contaminated with controlled levels of label noise. Results indicate that networks trained using the proposed method outperform those directly learned from noisy labels and that the benefits increase proportionally to the amount of noise present. Yuansheng Hua, Sylvain Lobry, Lichao Mou, Devis Tuia, Xiao Xiang Zhu 0001 |
IGARSS | 1 |
| 2020 | Instance Segmentation of Buildings Using KeypointsabstractBuilding segmentation is of great importance in the task of remote sensing imagery interpretation. However, the existing semantic segmentation and instance segmentation methods often lead to segmentation masks with blurred boundaries. In this paper, we propose a novel instance segmentation network for building segmentation in high-resolution remote sensing images. More specifically, we consider segmenting an individual building as detecting several keypoints. The detected keypoints are subsequently reformulated as a closed polygon, which is the semantic boundary of the building. By doing so, the sharp boundary of the building could be preserved. Experiments are conducted on selected Aerial Imagery for Roof Segmentation (AIRS) dataset, and our method achieves better performance in both quantitative and qualitative results with comparison to the state-of-the-art methods. Our network is a bottom-up instance segmentation method that could well preserve geometric details. Qingyu Li 0001, Lichao Mou, Yuansheng Hua, Yao Sun 0005, Pu Jin, Yilei Shi, Xiao Xiang Zhu 0001 |
IGARSS | 3 |
| 2020 | Event and Activity Recognition in Aerial Videos Using Deep Neural Networks and a New DatasetabstractUnmanned aerial vehicles (UAVs) are now widespread available. Yet the more UAVs there are in the skies, the more video data they create. It is unrealistic for humans to screen such big data and understand their contents. Hence methodological research on UAV video content understanding is of great importance. In this paper, we introduce a novel task of event recognition in unconstrained aerial videos in the remote sensing community and present a dataset for this task. Organized in a rich semantic taxonomy, the proposed dataset covers a wide range of events involving diverse environments and scales. We report results of plenty of deep networks in two ways: single-frame classification and video classification. The dataset and trained models can be downloaded from https://1cmou.github.io/ERA_Dataset/. Lichao Mou, Yuansheng Hua, Pu Jin, Xiao Xiang Zhu 0001 |
IGARSS | 2 |
| 2020 | Relation Network for Multilabel Aerial Image ClassificationabstractMultilabel classification plays a momentous role in perceiving intricate contents of an aerial image and triggers several related studies over the last years. However, most of them deploy few efforts in exploiting label relations, while such dependencies are crucial for making accurate predictions. Although an long short term memory (LSTM) layer can be introduced to modeling such label dependencies in a chain propagation manner, the efficiency might be questioned when certain labels are improperly inferred. To address this, we propose a novel aerial image multilabel classification network, attention-aware label relational reasoning network. Particularly, our network consists of three elemental modules: 1) a label-wise feature parcel learning module; 2) an attentional region extraction module; and 3) a label relational inference module. To be more specific, the label-wise feature parcel learning module is designed for extracting high-level label-specific features. The attentional region extraction module aims at localizing discriminative regions in these features without region proposal generation, yielding attentional label-specific features. The label relational inference module finally predicts label existences using label relations reasoned from outputs of the previous module. The proposed network is characterized by its capacities of extracting discriminative label-wise features and reasoning about label relations naturally and interpretably. In our experiments, we evaluate the proposed model on two multilabel aerial image data sets, of which one is newly produced. Quantitative and qualitative results on these two data sets demonstrate the effectiveness of our model. To facilitate progress in the multilabel aerial image classification, our produced data set will be made publicly available. Yuansheng Hua, Lichao Mou, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | Relation Matters: Relational Context-Aware Fully Convolutional Network for Semantic Segmentation of High-Resolution Aerial ImagesabstractMost current semantic segmentation approaches fall back on deep convolutional neural networks (CNNs). However, their use of convolution operations with local receptive fields causes failures in modeling contextual spatial relations. Prior works have sought to address this issue by using graphical models or spatial propagation modules in networks. But such models often fail to capture long-range spatial relationships between entities, which leads to spatially fragmented predictions. Moreover, recent works have demonstrated that channel-wise information also acts a pivotal part in CNNs. In this article, we introduce two simple yet effective network units, the spatial relation module, and the channel relation module to learn and reason about global relationships between any two spatial positions or feature maps, and then produce Relation-Augmented (RA) feature representations. The spatial and channel relation modules are general and extensible, and can be used in a plug-and-play fashion with the existing fully convolutional network (FCN) framework. We evaluate relation module-equipped networks on semantic segmentation tasks using two aerial image data sets, namely International Society for Photogrammetry and Remote Sensing (ISPRS) Vaihingen and Potsdam data sets, which fundamentally depend on long-range spatial relational reasoning. The networks achieve very competitive results, a mean F1score of 88.54% on the Vaihingen data set and a mean F1score of 88.01% on the Potsdam data set, bringing significant improvements over baselines. Lichao Mou, Yuansheng Hua, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | A Relation-Augmented Fully Convolutional Network for Semantic Segmentation in Aerial ScenesabstractMost current semantic segmentation approaches fall back on deep convolutional neural networks (CNNs). However, their use of convolution operations with local receptive fields causes failures in modeling contextual spatial relations. Prior works have sought to address this issue by using graphical models or spatial propagation modules in networks. But such models often fail to capture long-range spatial relationships between entities, which leads to spatially fragmented predictions. Moreover, recent works have demonstrated that channel-wise information also acts a pivotal part in CNNs. In this work, we introduce two simple yet effective network units, the spatial relation module and the channel relation module, to learn and reason about global relationships between any two spatial positions or feature maps, and then produce relation-augmented feature representations. The spatial and channel relation modules are general and extensible, and can be used in a plug-and-play fashion with the existing fully convolutional network (FCN) framework. We evaluate relation module-equipped networks on semantic segmentation tasks using two aerial image datasets, which fundamentally depend on long-range spatial relational reasoning. The networks achieve very competitive results, bringing significant improvements over baselines. Lichao Mou, Yuansheng Hua, Xiao Xiang Zhu 0001 |
CVPR | 2 |
| 2019 | Label Relation Inference for Multi-Label Aerial Image ClassificationabstractMulti-label aerial image classification is a challenging visual task and obtaining increasing attention recently. Most of the existing methods resort to training independent classifier for each label, while underlying label correlations are not fully exploited while making predictions. To this end, we propose an innovative inference network, which takes advantage of pairwise label relations to infer multiple object labels of a high-resolution aerial image. Specifically, we first employ a feature extraction module to extract high-level feature representations of an aerial image, and then, feed them into a relational inference module to predict the presence of each object label. We evaluate our network on the UCM multilabel dataset and experiment with various popular convolutional neural networks (CNNs) as the backbone of the feature extraction module. Experimental results demonstrate that the proposed network behaves superiorly in comparison with other existing methods. Yuansheng Hua, Lichao Mou, Xiao Xiang Zhu 0001 |
IGARSS | 1 |
| 2019 | Spatial Relational Reasoning in Networks for Improving Semantic Segmentation of Aerial ImagesabstractMost current semantic segmentation approaches rely on deep convolutional neural networks (CNNs). However, their use of convolution operations with local receptive fields causes failures in modeling contextual spatial relations. Prior works have tried to address this issue by using graphical models or spatial propagation modules in networks. But such models often fail to capture long-range spatial relationships between entities, which leads to spatially fragmented predictions. In this work, we introduce a simple yet effective network unit, the spatial relation module, to learn and reason about global relationships between any two spatial positions, and then produce relation-enhanced feature representations. The spatial relation module is general and extensible, and can be used in a plug-and-play fashion with the existing fully convolutional network (FCN) framework. We evaluate spatial relation module-equipped networks on semantic segmentation tasks using two aerial image datasets. The networks achieve very competitive results, bringing significant improvements over baselines. Lichao Mou, Yuansheng Hua, Xiao Xiang Zhu 0001 |
IGARSS | 2 |
| 2018 | LAHNet: A Convolutional Neural Network Fusing Low- and High-Level Features for Aerial Scene ClassificationabstractIn this paper, we proposed an innovative end-to-end convolutional neural network (CNN), which is trained to learn how to fuse multi-level features for aerial scene classification. Instead of using only coarse semantic features as conventional CNNs, we resort to first hierarchically extracting dense high-level features and then element-wise fusing them with low-level features to build a comprehensive feature representation, which contains not only high-level semantic information but also fine-grained low-level details, for scene classification. The network is evaluated on two broadly used aerial scene datasets, UCM and AID. The experimental results indicate that the proposed LAHNet performs superiorly compared to the existing benchmark methods. Furthermore, visualization of the fused features presents an intuitive illustration of the remarkable improvement. Yuansheng Hua, Lichao Mou, Xiao Xiang Zhu 0001 |
IGARSS | 1 |