VLDB 2026 Research / reviewers in the wild / expert
Yonghao Xu
dblp:210/5089
· DBLP profile ↗
24ranked-venue papers
14as first author
16since 2021 · last 2025
0000-0002-6857-0152ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 8 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | On the Adversarial Vulnerabilities of Transfer Learning in Remote SensingabstractTransfer learning with pretrained models from general computer vision tasks is widely adopted in remote sensing, offering reduced training costs and improved performance across applications like scene classification and semantic segmentation. However, this practice exposes downstream tasks to significant vulnerabilities, as adversaries can exploit publicly available pre-trained models to craft attacks that compromise model integrity. This paper introduces Adversarial Neuron Manipulation (ANM), a novel attack strategy with two variants—ANM-S (Single Neuron) and ANM-M (Multiple Neurons)—that generates highly transferable perturbations by selectively targeting fragile neurons in pretrained models. Unlike existing adversarial attacks, ANM requires no domain-specific knowledge, enhancing its applicability and efficiency across diverse settings. Evaluated on various benchmark datasets, ANM significantly degrades model performance. ANM-M, in particular, reduces accuracy by up to 84.24% in scene classification, 42.47% in semantic segmentation, and 75.33% in object detection (mAP). These results reveal critical weaknesses in deep learning models, especially their susceptibility to transferable attacks across CNN and Transformer architectures. By exposing the security risks inherent in transfer learning for remote sensing, ANM underscores the urgent need for robust defense mechanisms to safeguard safety-critical applications. Our findings advocate for enhanced model resilience and motivate future research into countering sophisticated adversarial threats. Xingjian Tian, Yonghao Xu, Bihan Wen |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Flexible Distribution Alignment: Towards Long-Tailed Semi-supervised Learning with Proper Calibration
Emanuel Sanchez Aimar, Nathaniel Helgesen, Yonghao Xu, Marco Kuhlmann, Michael Felsberg |
ECCV (54) | 3 |
| 2024 | SEN2FIRE: A Challenging Benchmark Dataset for Wildfire Detection Using Sentinel DataabstractUtilizing satellite imagery for wildfire detection presents substantial potential for practical applications. To advance the development of machine learning algorithms in this domain, our study introduces the Sen2Fire dataset–a challenging satellite remote sensing dataset tailored for wildfire detection. This dataset is curated from Sentinel-2 multi-spectral data and Sentinel-5P aerosol product, comprising a total of 2466 image patches. Each patch has a size of 512 512 pixels with 13 bands. Given the distinctive sensitivities× of various wavebands to wildfire responses, our research focuses on optimizing wildfire detection by evaluating different wavebands and employing a combination of spectral indices, such as normalized burn ratio (NBR) and normalized difference vegetation index (NDVI). The results suggest that, in contrast to using all bands for wildfire detection, selecting specific band combinations yields superior performance. Additionally, our study underscores the positive impact of integrating Sentinel-5 aerosol data for wildfire detection. The code and dataset are available online (https://zenodo.org/records/10881058). Yonghao Xu, Amanda Berg, Leif Haglund |
IGARSS | 1 |
| 2024 | Stealthy Adversarial Examples for Semantic Segmentation in Remote SensingabstractDeep learning methods have been proven effective in remote sensing image analysis and interpretation, where semantic segmentation plays a vital role. These deep segmentation methods are susceptible to adversarial attacks, while most of the existing attack methods tend to manipulate the image globally, leading to noticeable perturbations and chaotic segmentation. In this work, we propose a novel Stealthy Attack for Semantic Segmentation (SASS), which can largely increase the effectiveness and stealthiness from the existing attack methods on remote sensing images. SASS manipulates specific victim classes or objects of interest while preserving the original segmentation results for other classes or objects. In practice, as different inference mechanisms, overlapped inference, can be applied in segmentation, the efficacy of SASS may be degraded. To this end, we further introduce the Masked Stealthy Attack for Semantic Segmentation (MSASS), which generates augmented adversarial perturbations that only affect victim areas. We evaluate the effectiveness of SASS and MSASS using four state-of-the-art semantic segmentation models on the Vaihingen and Zurich Summer datasets. Extensive experiments demonstrate that our SASS and MSASS methods achieve superior attack performances on victim areas while maintaining high accuracies of other areas (drop less than 2%). The detection success rates of adversarial examples for segmentation, as characterized by Xiaoet al. [1], significantly drop from 97.78% for the untargeted PGD attack to 28.71% for our MSASS method on Zurich Summer dataset. Our work contributes to the field of adversarial attacks in semantic segmentation for remote sensing images by improving stealthiness, flexibility, and robustness. We anticipate that our findings will inspire the development of defense methods to enhance the security and reliability of semantic segmentation models against our stealthy attack. Yonghao Xu, Bihan Wen |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | S³ANet: Spatial-Spectral Self-Attention Learning Network for Defending Against Adversarial Attacks in Hyperspectral Image ClassificationabstractDeep neural networks have demonstrated impressive capabilities in hyperspectral image (HSI) classification tasks. However, they are highly vulnerable to adversarial attacks, raising significant security concerns, especially in the remote sensing community. Even subtle adversarial perturbations that are imperceptible to human observers can mislead deep learning (DL) models and result in incorrect predictions. Therefore, ensuring the robustness of DL models has become a critical focus in addressing security-related remote sensing tasks. Considerable progress has been made in defending against adversarial attacks in HSI classification. Nevertheless, existing methods primarily concentrate on spatial relationships between pixels while overlooking the valuable spectral information present in HSI. Besides, these methods are usually limited to a specific scale and cannot accommodate the precise classification demands for ground objects with various scales. To address these limitations, we propose a spatial–spectral self-attention learning network (S3ANet) for defending against adversarial attacks in HSI classification. Our S3ANet incorporates a pyramid spatial attention learning module to effectively capture spatial dependency at multiple scales. In addition, it utilizes a global spectral transformer to establish correlations between pixels in the spectral dimension. By employing the defense method of spatial–spectral fusion, our model can effectively address adversarial attacks from a comprehensive perspective, seamlessly integrating both spatial and spectral information. Extensive experiments conducted on four benchmark HSI datasets illustrate that the proposed S3ANet achieves competitive performance compared to state-of-the-art methods when faced with adversarial attacks. The code is available online athttps://github.com/YichuXu/S3ANet. Yichu Xu, Yonghao Xu, Hongzan Jiao, Zhi Gao 0005, Lefei Zhang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Robust Self-Ensembling Network for Hyperspectral Image ClassificationabstractRecent research has shown the great potential of deep learning algorithms in the hyperspectral image (HSI) classification task. Nevertheless, training these models usually requires a large amount of labeled data. Since the collection of pixel-level annotations for HSI is laborious and time-consuming, developing algorithms that can yield good performance in the small sample size situation is of great significance. In this study, we propose a robust self-ensembling network (RSEN) to address this problem. The proposed RSEN consists of two subnetworks including a base network and an ensemble network. With the constraint of both the supervised loss from the labeled data and the unsupervised loss from the unlabeled data, the base network and the ensemble network can learn from each other, achieving the self-ensembling mechanism. To the best of our knowledge, the proposed method is the first attempt to introduce the self-ensembling technique into the HSI classification task, which provides a different view on how to utilize the unlabeled data in HSI to assist the network training. We further propose a novel consistency filter to increase the robustness of self-ensembling learning. Extensive experiments on three benchmark HSI datasets demonstrate that the proposed algorithm can yield competitive performance compared with the state-of-the-art methods. Yonghao Xu, Bo Du 0001, Liangpei Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Backdoor Attacks for Remote Sensing Data With Wavelet TransformabstractRecent years have witnessed the great success of deep learning algorithms in the geoscience and remote sensing realm. Nevertheless, the security and robustness of deep learning models deserve special attention when addressing safety-critical remote sensing tasks. In this paper, we provide a systematic analysis of backdoor attacks for remote sensing data, where both scene classification and semantic segmentation tasks are considered. While most of the existing backdoor attack algorithms rely on visible triggers like squared patches with well-designed patterns, we propose a novel wavelet transform-based attack (WABA) method, which can achieve invisible attacks by injecting the trigger image into the poisoned image in the low-frequency domain. In this way, the high-frequency information in the trigger image can be filtered out in the attack, resulting in stealthy data poisoning. Despite its simplicity, the proposed method can significantly cheat the current state-of-the-art deep learning models with a high attack success rate. We further analyze how different trigger images and the hyper-parameters in the wavelet transform would influence the performance of the proposed method. Extensive experiments on four benchmark remote sensing datasets demonstrate the effectiveness of the proposed method for both scene classification and semantic segmentation tasks and thus highlight the importance of designing advanced backdoor defense algorithms to address this threat in remote sensing scenarios. The code will be available online at https://github.com/ndraeger/waba. Nikolaus Dräger, Yonghao Xu, Pedram Ghamisi |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Txt2Img-MHN: Remote Sensing Image Generation From Text Using Modern Hopfield NetworksabstractThe synthesis of high-resolution remote sensing images based on text descriptions has great potential in many practical application scenarios. Although deep neural networks have achieved great success in many important remote sensing tasks, generating realistic remote sensing images from text descriptions is still very difficult. To address this challenge, we propose a novel text-to-image modern Hopfield network (Txt2Img-MHN). The main idea of Txt2Img-MHN is to conduct hierarchical prototype learning on both text and image embeddings with modern Hopfield layers. Instead of directly learning concrete but highly diverse text-image joint feature representations for different semantics, Txt2Img-MHN aims to learn the most representative prototypes from text-image embeddings, achieving a coarse-to-fine learning strategy. These learned prototypes can then be utilized to represent more complex semantics in the text-to-image generation task. To better evaluate the realism and semantic consistency of the generated images, we further conduct zero-shot classification on real remote sensing data using the classification model trained on synthesized images. Despite its simplicity, we find that the overall accuracy in the zero-shot classification may serve as a good metric to evaluate the ability to generate an image from text. Extensive experiments on the benchmark remote sensing text-image dataset demonstrate that the proposed Txt2Img-MHN can generate more realistic remote sensing images than existing methods. Code and pre-trained models are available online (https://github.com/YonghaoXu/Txt2Img-MHN). Yonghao Xu, Weikang Yu, Pedram Ghamisi, Michael Kopp 0001, Sepp Hochreiter |
IEEE Trans. Image Process. | 1 |
| 2023 | Can Spectral Information Work While Extracting Spatial Distribution? - An Online Spectral Information Compensation Network for HSI ClassificationabstractIn the past few years, deep learning-based methods have shown commendable performance for hyperspectral image (HSI) classification. Many works focus on designing independent spectral and spatial branches and then fusing the output features from two branches for category prediction. In this way, the correlation that exists between spectral and spatial information is not completely explored, and spectral information extracted from one branch is always not sufficient. Some studies also try to directly extract spectral-spatial features using 3D convolutions but are accompanied by the severe over-smoothing phenomenon and poor representation ability of spectral signatures. Unlike the above-mentioned approaches, in this paper, we propose a novel online spectral information compensation network (OSICN) for HSI classification, which consists of a candidate spectral vector mechanism, progressive filling process, and multi-branch network. To the best of our knowledge, this paper is the first to online supplement spectral information into the network when spatial features are extracted. The proposed OSICN makes the spectral information participate in network learning in advance to guide spatial information extraction, which truly processes spectral and spatial features in HSI as a whole. Accordingly, OSICN is more reasonable and more effective for complex HSI data. Experimental results on three benchmark datasets demonstrate that the proposed approach has more outstanding classification performance compared with the state-of-the-art methods, even with a limited number of training samples. Jiaqi Yang 0005, Bo Du 0001, Yonghao Xu, Liangpei Zhang 0001 |
IEEE Trans. Image Process. | 3 |
| 2023 | Self-Ensembling GAN for Cross-Domain Semantic SegmentationabstractDeep neural networks (DNNs) have greatly contributed to the performance gains in semantic segmentation. Nevertheless, training DNNs generally requires large amounts of pixel-level labeled data, which is expensive and time-consuming to collect in practice. To mitigate the annotation burden, this paper proposes a self-ensembling generative adversarial network (SE-GAN) exploiting cross-domain data for semantic segmentation. In SE-GAN, a teacher network and a student network constitute a self-ensembling model for generating semantic segmentation maps, which together with a discriminator, forms a GAN. Despite its simplicity, we find SE-GAN can significantly boost the performance of adversarial training and enhance the stability of the model, the latter of which is a common barrier shared by most adversarial training-based methods. We theoretically analyze SE-GAN and provide an$\mathcal {O}(1/\sqrt{N})$generalization bound ($N$is the training sample size), which suggests controlling the discriminator's hypothesis complexity to enhance the generalizability. Accordingly, we choose a simple network as the discriminator. Extensive and systematic experiments in two standard settings demonstrate that the proposed method significantly outperforms current state-of-the-art approaches. Yonghao Xu, Fengxiang He, Bo Du 0001, Dacheng Tao, Liangpei Zhang 0001 |
IEEE Trans. Multim. | 1 |
| 2022 | Landslide4Sense: Reference Benchmark Data and Deep Learning Models for Landslide DetectionabstractThis study introducesLandslide4Sense, a reference benchmark for landslide detection from remote sensing. The repository features 3,799 image patches fusing optical layers from Sentinel-2 sensors with the digital elevation model and slope layer derived from ALOS PALSAR. The added topographical information facilitates an accurate detection of landslide borders, which recent researches have shown to be challenging using optical data alone. The extensive data set supports deep learning (DL) studies in landslide detection and the development and validation of methods for the systematic update of landslide inventories. The benchmark data set has been collected at four different times and geographical locations: Iburi (September 2018), Kodagu (August 2018), Gorkha (April 2015), and Taiwan (August 2009). Each image pixel is labelled as belonging to a landslide or not, incorporating various sources and thorough manual annotation. We then evaluate the landslide detection performance of 11 state-of-the-art DL segmentation models: U-Net, ResU-Net, PSPNet, ContextNet, DeepLab-v2, DeepLab-v3+, FCN-8s, LinkNet, FRRN-A, FRRN-B, and SQNet. All models were trained from scratch on patches from one quarter of each study area and tested on independent patches from the other three quarters. Our experiments demonstrate that ResU-Net outperformed the other models for the landslide detection task. We make the multi-source landslide benchmark data (Landslide4Sense) and the tested DL models publicly available at https://www.iarai.ac.at/landslide4sense, establishing an important resource for remote sensing, computer vision, and machine learning communities in studies of image classification in general and applications to landslide detection in particular. Omid Ghorbanzadeh, Yonghao Xu, Pedram Ghamisi, Michael Kopp 0001, David P. Kreil |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Universal Adversarial Examples in Remote Sensing: Methodology and BenchmarkabstractDeep neural networks have achieved great success in many important remote sensing tasks. Nevertheless, their vulnerability to adversarial examples should not be neglected. In this study, we systematically analyze the Universal Adversarial Examples in Remote Sensing (UAE-RS) data for the first time, without any knowledge from the victim model. Specifically, we propose a novel black-box adversarial attack method, namely, Mixup-Attack, and its simple variant Mixcut-Attack, for remote sensing data. The key idea of the proposed methods is to find common vulnerabilities among different networks by attacking the features in the shallow layer of a given surrogate model. Despite their simplicity, the proposed methods can generate transferable adversarial examples that deceive most of the state-of-the-art deep neural networks in both scene classification and semantic segmentation tasks with high success rates. We further provide the generated universal adversarial examples in the dataset named UAE-RS, which is the first dataset that provides black-box adversarial samples in the remote sensing field. We hope UAE-RS may serve as a benchmark that helps researchers design deep neural networks with strong resistance toward adversarial attacks in the remote sensing field. Codes and the UAE-RS dataset are available online (https://github.com/YonghaoXu/UAE-RS). Yonghao Xu, Pedram Ghamisi |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Consistency-Regularized Region-Growing Network for Semantic Segmentation of Urban Scenes With Point-Level AnnotationsabstractDeep learning algorithms have obtained great success in semantic segmentation of very high-resolution (VHR) remote sensing images. Nevertheless, training these models generally requires a large amount of accurate pixel-wise annotations, which is very laborious and time-consuming to collect. To reduce the annotation burden, this paper proposes a consistency-regularized region-growing network (CRGNet) to achieve semantic segmentation of VHR remote sensing images with point-level annotations. The key idea of CRGNet is to iteratively select unlabeled pixels with high confidence to expand the annotated area from the original sparse points. However, since there may exist some errors and noises in the expanded annotations, directly learning from them may mislead the training of the network. To this end, we further propose the consistency regularization strategy, where a base classifier and an expanded classifier are employed. Specifically, the base classifier is supervised by the original sparse annotations, while the expanded classifier aims to learn from the expanded annotations generated by the base classifier with the region-growing mechanism. The consistency regularization is thereby achieved by minimizing the discrepancy between the predictions from both the base and the expanded classifiers. We find such a simple regularization strategy is yet very useful to control the quality of the region-growing mechanism. Extensive experiments on two benchmark datasets demonstrate that the proposed CRGNet significantly outperforms the existing state-of-the-art methods. Codes and pre-trained models are available online (https://github.com/YonghaoXu/CRGNet). Yonghao Xu, Pedram Ghamisi |
IEEE Trans. Image Process. | 1 |
| 2021 | Adaptive Spectral-Spatial Multiscale Contextual Feature Extraction for Hyperspectral Image ClassificationabstractIn this article, we propose an end-to-end adaptive spectral-spatial multiscale network to extract multiscale contextual information for hyperspectral image (HSI) classification, which contains spectral feature extraction (FE) and spatial FE subnetworks. In spectral FE aspect, different from previous methods where features are obtained in a single scale, which limits the accuracy improvement, we propose two schemes based on band grouping strategy, and the long short-time memory (LSTM) model is used for perceiving spectral multiscale information. In spatial subnetwork, on the foundation of existing multiscale architecture, the spatial contextual features which are usually ignored by previous literature are successfully obtained under the aid of convolutional LSTM (ConvLSTM) model. Besides, a new spatial grouping strategy is proposed for convenience of ConvLSTM to extract the more discriminative features. Then, a novel adaptive feature combining way is proposed considering the different importance of spectral and spatial parts. Experiments on three public data sets in HSI community demonstrate that our methods achieve competitive results compared with other state-of-the-art methods. Di Wang 0023, Bo Du 0001, Liangpei Zhang 0001, Yonghao Xu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | Assessing the Threat of Adversarial Examples on Deep Neural Networks for Remote Sensing Scene Classification: Attacks and DefensesabstractDeep neural networks, which can learn the representative and discriminative features from data in a hierarchical manner, have achieved state-of-the-art performance in the remote sensing scene classification task. Despite the great success that deep learning algorithms have obtained, their vulnerability toward adversarial examples deserves our special attention. In this article, we systematically analyze the threat of adversarial examples on deep neural networks for remote sensing scene classification. Both targeted and untargeted attacks are performed to generate subtle adversarial perturbations, which are imperceptible to a human observer but may easily fool the deep learning models. Simply adding these perturbations to the original high-resolution remote sensing (HRRS) images, adversarial examples can be generated, and there are only slight differences between the adversarial examples and the original ones. An intriguing discovery in our study shows that most of these adversarial examples may be misclassified into the wrong category by the state-of-the-art deep neural networks with very high confidence. This phenomenon, undoubtedly, may limit the practical deployment of these deep learning models in the safety-critical remote sensing field. To address this problem, the adversarial training strategy is further investigated in this article, which significantly increases the resistibility of deep models toward adversarial examples. Extensive experiments on three benchmark HRRS image data sets demonstrate that while most of the well-known deep neural networks are sensitive to adversarial perturbations, the adversarial training strategy helps to alleviate their vulnerability toward adversarial examples. Yonghao Xu, Bo Du 0001, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | Self-Attention Context Network: Addressing the Threat of Adversarial Attacks for Hyperspectral Image ClassificationabstractDeep learning models have shown their great capability for the hyperspectral image (HSI) classification task in recent years. Nevertheless, their vulnerability towards adversarial attacks could not be neglected. In this study, we systematically analyze the influence of adversarial attacks on the HSI classification task for the first time. While existing research of adversarial attacks focuses on the generation of adversarial examples in the RGB domain, the experiments in this study show such adversarial examples could also exist in the hyperspectral domain. Although the difference between the generated adversarial image and the original hyperspectral data is imperceptible to the human visual system, most of the existing state-of-the-art deep learning models could be fooled by the adversarial image to make wrong predictions. To address this challenge, a novel self-attention context network (SACNet) is further proposed. We discover that the global context information contained in HSI can significantly improve the robustness of deep neural networks when confronted with adversarial attacks. Extensive experiments on three benchmark HSI datasets demonstrate that the proposed SACNet possesses stronger resistibility towards adversarial examples compared with the existing state-of-the-art deep learning models. Yonghao Xu, Bo Du 0001, Liangpei Zhang 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | Beyond the Patchwise Classification: Spectral-Spatial Fully Convolutional Networks for Hyperspectral Image ClassificationabstractIn recent years, patchwise classification methods are commonly adopted when dealing with the hyperspectral image (HSI) classification. Despite their promising results from the perspective of accuracy, the efficiency of these methods can hardly be ensured since there are redundant computations between adjacent patches. In this paper, we propose a spectral-spatial fully convolutional network for HSI classification with an end-to-end, pixel-to-pixel architecture. Compared with patchwise methods, the proposed framework can avoid the patch extraction and is more efficient. Since the training samples in HSIs are highly sparse, the training strategy in original fully convolutional networks is no longer feasible for HSIs. To solve this problem, we propose a novel mask matrix to assist the back-propagation in the training stage. Considering the importance of spectral and spatial features may vary for different objects and scenes, we combine both features with two weighting factors which can be adaptively learned during the network training. Besides, the dense conditional random field (CRF) is introduced into the framework to further balance the local and global information. Experiments on three benchmark HSI data sets demonstrate that the proposed method can yield competitive results with less time costs compared with patchwise methods. Yonghao Xu, Bo Du 0001, Liangpei Zhang 0001 |
IEEE Trans. Big Data | 1 |
| 2020 | Learning From Synthetic Images via Active Pseudo-LabelingabstractSynthetic visual data refers to the data automatically rendered by the mature computer graphic algorithms. With the rapid development of these techniques, we can now collect photo-realistic synthetic images with accurate pixel-level annotations without much effort. However, due to the domain gaps between synthetic data and real data, in terms of not only visual appearance but also label distribution, directly applying models trained on synthetic images to real ones can hardly yield satisfactory performance. Since the collection of accurate labels for real images is very laborious and time-consuming, developing algorithms which can learn from synthetic images is of great significance. In this paper, we propose a novel framework, namely Active Pseudo-Labeling (APL), to reduce the domain gaps between synthetic images and real images. In APL framework, we first predict pseudo-labels for the unlabeled real images in the target domain by actively adapting the style of the real images to source domain. Specifically, the style of real images is adjusted via a novel task guided generative model, and then pseudo-labels are predicted for these actively adapted images. Lastly, we fine-tune the source-trained model in the pseudo-labeled target domain, which helps to fit the distribution of the real data. Experiments on both semantic segmentation and object detection tasks with several challenging benchmark data sets demonstrate the priority of our proposed method compared to the existing state-of-the-art approaches. Liangchen Song, Yonghao Xu, Lefei Zhang, Bo Du 0001, Qian Zhang 0009, Xinggang Wang |
IEEE Trans. Image Process. | 2 |
| 2019 | Self-Ensembling Attention Networks: Addressing Domain Shift for Semantic SegmentationabstractRecent years have witnessed the great success of deep learning models in semantic segmentation. Nevertheless, these models may not generalize well to unseen image domains due to the phenomenon of domain shift. Since pixel-level annotations are laborious to collect, developing algorithms which can adapt labeled data from source domain to target domain is of great significance. To this end, we propose self-ensembling attention networks to reduce the domain gap between different datasets. To the best of our knowledge, the proposed method is the first attempt to introduce selfensembling model to domain adaptation for semantic segmentation, which provides a different view on how to learn domain-invariant features. Besides, since different regions in the image usually correspond to different levels of domain gap, we introduce the attention mechanism into the proposed framework to generate attention-aware features, which are further utilized to guide the calculation of consistency loss in the target domain. Experiments on two benchmark datasets demonstrate that the proposed framework can yield competitive performance compared with the state of the art methods. Yonghao Xu, Bo Du 0001, Lefei Zhang, Qian Zhang 0009, Guoli Wang 0004, Liangpei Zhang 0001 |
AAAI | 1 |
| 2019 | Fast Spatio-Temporal Residual Network for Video Super-ResolutionabstractRecently, deep learning based video super-resolution (SR) methods have achieved promising performance. To simultaneously exploit the spatial and temporal information of videos, employing 3-dimensional (3D) convolutions is a natural approach. However, straight utilizing 3D convolutions may lead to an excessively high computational complexity which restricts the depth of video SR models and thus undermine the performance. In this paper, we present a novel fast spatio-temporal residual network (FSTRN) to adopt 3D convolutions for the video SR task in order to enhance the performance while maintaining a low computational load. Specifically, we propose a fast spatio-temporal residual block (FRB) that divide each 3D filter to the product of two 3D filters, which have considerably lower dimensions. Furthermore, we design a cross-space residual learning that directly links the low-resolution space and the high-resolution space, which can greatly relieve the computational burden on the feature fusion and up-scaling parts. Extensive evaluations and comparisons on benchmark datasets validate the strengths of the proposed approach and demonstrate that the proposed network significantly outperforms the current state-of-the-art methods. Sheng Li 0001, Fengxiang He, Bo Du 0001, Lefei Zhang, Yonghao Xu, Dacheng Tao |
CVPR | 5 |
| 2019 | Simultaneous Segmentation and Edge Detection for Hyperspectral Image Via a Deep Supervised and boundary-constrained NetworkabstractRecent research has shown the great potential of convolutional neural networks (CNNs) in hyperspectral image (HSI) classification. Nevertheless, CNN based approaches may lead to over-smoothing effect due to the spatial information loss during the convolution and pooling operations. To address this problem, in this paper, we propose a deep supervised and boundary-constrained network (DSBC-Net) which takes the boundary information into consideration. With the well-designed architecture, DSBC-Net can achieve segmentation and edge detection for HSI simultaneously. Besides, a novel deep supervision strategy is proposed to improve the training of the deep neural network. Experimental results on two benchmark HSI datasets demonstrate that the DSBC-Net can better maintain the boundaries of different objects with higher classification accuracy compared with previous methods. Yonghao Xu, Bo Du 0001, Liangpei Zhang 0001 |
IGARSS | 1 |
| 2018 | Multi-Source Remote Sensing Data Classification via Fully Convolutional Networks and Post-Classification ProcessingabstractThis paper presents a new data fusion methodology named Fusion-FCN for the classification of multi-source remote sensing data using fully convolutional networks (FCNs). Three different types of data including LiDAR data, hyperspectral images and very high resolution images are utilized in the proposed framework. Considering the confusions between similar categories (e.g., road and highway), we further implement post-classification processing with the topological relationship among different objects based on the result yielded by the proposed Fusion-FCN. The proposed method achieved an overall accuracy of 80.78% and a kappa coefficient of 0.80, which ranked first in the 2018 IEEE GRSS Data Fusion Contest. Yonghao Xu, Bo Du 0001, Liangpei Zhang 0001 |
IGARSS | 1 |
| 2018 | Can We Generate Good Samples for Hyperspectral Classification? - A Generative Adversarial Network Based MethodabstractThe insufficiency of training samples is really a great challenge for hyperspectral image (HSI) classification. Samples generation is a commonly used technique in deep learning based remote sensing field which can extend the training set. However, previous methods ignore the real distribution of the training samples in the feature space and thus can hardly ensure that the generated samples possess the same patterns with the real ones. In this paper, we propose a generative adversarial network based method (SpecGAN) to handle this problem. Different from traditional GAN framework where the generated samples have no categories, for the first time we take the label information into consideration for hyperspectral images. Feeding a random noise z and a class label vector y into the generator, we can get a spectral sample of the corresponding category. The experiments on the Pavia University data set demonstrate the potential of the proposed SpecGAN in spectral samples generation. Yonghao Xu, Bo Du 0001, Liangpei Zhang 0001 |
IGARSS | 1 |
| 2018 | Spectral-Spatial Unified Networks for Hyperspectral Image ClassificationabstractIn this paper, we propose a spectral–spatial unified network (SSUN) with an end-to-end architecture for the hyperspectral image (HSI) classification. Different from traditional spectral–spatial classification frameworks where the spectral feature extraction (FE), spatial FE, and classifier training are separated, these processes are integrated into a unified network in our model. In this way, both FE and classifier training will share a uniform objective function and all the parameters in the network can be optimized at the same time. In the implementation of the SSUN, we propose a band grouping-based long short-term memory model and a multiscale convolutional neural network as the spectral and spatial feature extractors, respectively. In the experiments, three benchmark HSIs are utilized to evaluate the performance of the proposed method. The experimental results demonstrate that the SSUN can yield a competitive performance compared with existing methods. Yonghao Xu, Liangpei Zhang 0001, Bo Du 0001, Fan Zhang 0006 |
IEEE Trans. Geosci. Remote. Sens. | 1 |