Renlong Hang

dblp:149/7523 · DBLP profile ↗
← Back
44ranked-venue papers
13as first author
28since 2021 · last 2026
0000-0001-6046-3689ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 31 · 10 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021
YearPublicationVenuePosition
2026 More realistic and accurate precipitation nowcasting with Conditional Rectified Flow Transformers
Yunlong Zhou, Fanfan Ji, Renlong Hang, Qingshan Liu 0001, Xiao-Tong Yuan
Eng. Appl. Artif. Intell.4
2026 Adaptive frequency collaboration for remote sensing change detection
Feng Zhou 0006, Hui Shuai, Qingshan Liu 0001, Renlong Hang
Neural Networks5
2026 Graph Transformer With Structural Embedding and Training for Hyperspectral Image
abstract
Graph transformer networks have received more attention in hyperspectral image (HSI) classification. However, they overlooked the influence of graph connectivity strength in positional encoding and distribution. In order to address the above deficiencies, we proposed the novel graph transformer with structural embedding and training (GTSET) for HSI classification. Specifically, the structural embedding module firstly aimed at extracting effectively local and non-local feature information via patch-based distance encoding and centrality correlation coefficients based on graph connectivity strength, alleviating spectral variability. Secondly, the structural training module aimed at addressing imbalanced structural position distribution of labeled samples by leveraging the topological graph connectivity to determine their structural position distribution and reweighting the influence of labeled samples on the graph transformer training stage, exploring the guiding role of labeled samples in low spatial resolution of HSI. Next, we further refine training weights based on the spectral feature smoothness of labeled samples. Finally, comprehensive experiments on three real-world HSI datasets demonstrate that the GTSET achieves superior performance in HSI classification with limited labeled samples, compared to other popular classification methods. Implementation of GTSET, along with examples, can be found on the GitHub repository: https://github.com/xuchengchao0/GTSET.
Yun Ding, Chengchao Xu, Pi-Jing Wei, Renlong Hang, Chun-Hou Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.5
2025 A Remote Sensing Change Detection Network Using Visual-Prompt Enhanced CLIP
abstract
Remote Sensing Change Detection (RSCD) plays a crucial role in various earth observation tasks. Recently, deep learning-based methods have been widely employed for CD due to their exceptional performance. Although, existing approaches can detect obviously changed regions easily, they suffer from difficulties in dealing with pseudo-changes caused by lighting condition changes, season changes and complex land cover conditions. To tackle this challenge, we propose a CD network using visual-prompt enhanced CLIP (CVNet), which incorporates the foundation model CLIP into the RSCD task to leverage its semantic information for identifying pseudochange regions. Specifically, we use CLIP visual encoder with transformer-structure to enhance single-time features extracted by ResNet. To ensure effective transfer ability for downstream tasks while considering computational cost, we fine-tune CLIP using a visual prompt. In addition, to efficiently enhance features extracted by ResNet, we design a CLIP-guided feature refinement (CGFR) module that adaptively integrates both types of features. Furthermore, a transformer encoder structure is introduced to get change information for dual-time images and a transformer decoder is introduced to propagate change information back. To test the performance of our proposed model, we conduct experiments on two datasets, including LEVIR-CD and WHU-CD. The experimental results show that our model can outperform several state-of-the-art models on both datasets. The code is available at https://github.com/Hyper-Baller/CVNet.
Yuhao Liu 0013, Zhiyong Zheng, Renlong Hang
IEEE Geosci. Remote. Sens. Lett.3
2025 Text-Augmented Semantic Feature Extraction and Difference Information Learning for Remote Sensing Image Change Captioning
abstract
Remote sensing image change captioning (RSICC) aims to generate sentence descriptions about land cover changes in bitemporal images. The effective acquisition of semantic-level change information is critical for this task. However, due to the effects of illumination interference, appearance similarities and scale differences between different objects, it is difficult to accurately extract change information from bitemporal images. In this article, we attempt to take advantage of the high-level semantic information inherent in text and propose a text-augmented semantic feature extraction and difference information learning model for RSICC. Specifically, we first pre-define some text prompts for each remote sensing image and use the contrastive language-image pretraining (CLIP) model to select the most suitable text descriptions for them. Then, we adopt a refined segment anything model (SAM) to learn fine-grained visual features from each image, which is further enhanced via a designed selective text-image fusion (STIF) module. After that, to extract the semantic differences between bitemporal images, we propose a text-guided difference capture (TGDC) module capable of extracting multiscale difference information under the guidance of text differences between different-time images. Finally, a transformer-based caption generator is applied to generate sentence descriptions from the extracted difference information. In order to test the performance of our proposed model, we conduct comprehensive experiments on two widely used RSICC datasets, including LEVIR-CC and Dubai-CC. The experimental results show that our proposed model is able to outperform several state-of-the-art models, which validates the effectiveness of it. The codes of our proposed model will be released at https://github.com/Richardkimyo/TACC.
Renlong Hang, Jinyu Luo, Qingshan Liu 0001
IEEE Trans. Geosci. Remote. Sens.1
2025 Remote Sensing Object Counting With Online Knowledge Learning
abstract
Efficient models for remote sensing object counting are urgently required for applications in scenarios with limited computing resources, such as drones or embedded systems. A straightforward yet powerful technique to achieve this is knowledge distillation (KD), which steers the learning of student networks by leveraging the experience of already-trained teacher networks. However, it faces a pair of challenges. First, due to its two-stage training nature, a longer training period is essential, especially as the training samples increase. Second, despite the proficiency of teacher networks in transmitting assimilated knowledge, they tend to overlook the latent insights gained during their learning process. To address these challenges, we introduce an online distillation learning method for remote sensing object counting. It builds an end-to-end training framework that seamlessly integrates two distinct networks into a unified one. It comprises a shared shallow module, a teacher branch, and a student branch. The shared module serving as the foundation for both branches is dedicated to learning some primitive information. The teacher branch utilizes prior knowledge to reduce the difficulty of learning and guides the student branch in online learning. In parallel, the student branch achieves parameter reduction and rapid inference capabilities by means of channel reduction. This design empowers the student branch not only to receive privileged insights from the teacher branch but also to tap into the latent reservoir of knowledge held by the teacher branch during the learning process. Moreover, we propose a relation-in-relation distillation (RiRD) method that allows the student branch to effectively comprehend the evolution of the relationship of intralayer teacher features among different interlayer features. Extensive experiments on two challenging datasets demonstrate the effectiveness of our method, which achieves comparable performance to state-of-the-art (SOTA) methods despite using far fewer parameters.
Shengqin Jiang, Yuan Gao 0053, Fengna Cheng, Renlong Hang, Qingshan Liu 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 DiFormer: A Difference Transformer Network for Remote Sensing Change Detection
abstract
Change detection (CD) is one of the most important methods for monitoring land surface changes. Recently, transformer-based models have been employed to CD. However, most of them focus on modeling the correlation within each image, while cannot well model the difference between bi-temporal images. In this paper, we propose a difference transformer network (DiFormer) to address this issue. Specifically, we propose a Token Exchange-based Difference Evaluation (TEDE) module to generate the inconsistency between the changed region and the surrounding context to highlight the difference between the bi-temporal images. In addition, to obtain semantically rich exchangeable tokens, we design a Multi-scale Semantic Perception (MSP) module, which provides assistance for difference modeling. In order to test the performance of DiFormer, we conduct qualitative and quantitative experiments on two public datasets, including LEVIR-CD and S2Looking. Experimental results show that our proposed DiFormer is able to achieve better results than several state-of-the-art models, with F1 of 92.15% on the LEVIR-CD dataset and 66.31% on the S2Looking dataset.
Renlong Hang, Shanmin Wang, Qingshan Liu 0001
IEEE Geosci. Remote. Sens. Lett.2
2024 Deep Precipitation Nowcasting With Dual Regions Displacement Information and Global Spatiotemporal Representations Learning
abstract
The deep precipitation nowcasting using radar echo map prediction can mitigate the socio-economic impact of extreme precipitation events. Existing methods employ long short-term memory (LSTM) to extract rich precipitation features. However, existing methods often combine the learning and modeling of rain and nonrain regions in a single module, without clearly distinguishing their different features and motion patterns, which impairs the spatial distribution and precipitation intensity prediction of rainfall. Moreover, these LSTMs only capture local spatiotemporal features, while ignoring the global spatiotemporal features, resulting in prediction results lacking structural and strength consistency. Therefore, we propose a dual regions center displacement (DRCD) module, which separately learns and models the spatial information of rainfall and nonrainfall regions and employs this module to estimate the locations and intensity residuals of the future regions. Moreover, we also introduce a novel Global LSTM module (GLSTM) that learns the global spatiotemporal features from the sequences, which can estimate the structure and intensity of dual regions. Extensive experiments demonstrate that our method has superior or competitive performance over the state-of-the-art precipitation nowcasting methods and has the potential to be implemented as an alternative product globally.
Fanfan Ji, Yunlong Zhou, Renlong Hang, Qingshan Liu 0001, Xiao-Tong Yuan
IEEE Geosci. Remote. Sens. Lett.4
2024 A Regionally Indicated Visual Grounding Network for Remote Sensing Images
abstract
Visual grounding (VG) is essential to promote the human-computer interaction in object detection tasks. Most of the current VG methods mainly focus on grounding the target objects in natural images with simple language expressions. They cannot generalize well to remote sensing images, where the target objects only cover a small fraction (e.g., 0.34%) of the whole scene and the language expression is complex. To address these challenges, we propose a regionally indicated network (RINet) for remote sensing VG in this article. Specifically, RINet first exploits DarkNet-53 and BERT to extract visual and language features, respectively. Then, these features are fed into a regional indication generator (RIG) to generate an initial indication map, which indicates the possibility of each region containing the target object. This indication map is fine-tuned by taking advantage of a high-resolution detailed feature via a comprehensive alignment module (CAM) and a correction gate (CG). In CAM, a word contribution learner is designed to evaluate the importance of each word and make it pay more attention to the words easily ignored before. The whole fine-tuning process is repeated several rounds so that the complex language information can be fully explored and the region containing the target object is located more accurately. Finally, a detection head is adopted to ground the target object. To test the performance of our proposed model, we conduct experiments on two public remote sensing datasets, including RSVG and DIOR-RSVG. The experimental results show that our proposed RINet can outperform several state-of-the-art models significantly, which validates its effectiveness. The source code of our proposed model will be released athttps://github.com/KevinDaldry/RINet.
Renlong Hang, Siqi Xu, Qingshan Liu 0001
IEEE Trans. Geosci. Remote. Sens.1
2024 AANet: An Ambiguity-Aware Network for Remote-Sensing Image Change Detection
abstract
Remote sensing image change detection (CD) task plays an important role in land-use survey, city construction investigation and other vital industries. Recently, deep learning has become a mainstream method for this task due to its satisfactory performance in most cases. However, it often suffers from difficulties in dealing with ambiguity regions, where pseudo-changes happen or real changes are corrupted. In this article, we propose an ambiguity-aware network (AANet) to address the aforementioned issue. Specifically, our network firstly adopts convolutional layers to learn features from dual-temporal images. After that, an ambiguity refinement module (ARM) is designed to extract the ambiguity regions and then difference features are generated based on it. Considering that the scales of different changed objects vary, a weight rearrangement module (WRM) is proposed to fuse the difference features from different layers. In order to test the performance of our proposed model, we conduct experiments on three benchmark datasets, including SYSU-CD, SVCD, and LEVIR-CD. The experimental results show that our model can outperform several state-of-the-art models on all three datasets, which validates the effectiveness of it. The source code of our proposed model will be released at https://github.com/KevinDaldry/AANet.
Renlong Hang, Siqi Xu, Panli Yuan, Qingshan Liu 0001
IEEE Trans. Geosci. Remote. Sens.1
2024 A Closed-Loop Circular Regression Network for 2-m Air Temperature Downscaling Over Southwestern China
abstract
High-quality meteorological grid data are essential for meteorological research and applications, especially in regional scales. Statistical downscaling (SD) is an efficient method to provide more detailed information at spatial scale, and has already been implemented in many regions. In recent years, deep convolutional neural networks have exhibited promising performance in SD, effectively learning non-linear mappings from the low-resolution (LR) meteorological data to its corresponding high-resolution (HR) one. Nevertheless, existing deep-learning-based SD approaches may encounter two potential limitations. First, most of the previous deep-learning-based downscaling algorithms utilize a supervised learning framework, which necessitates the formation of data pairs consisting of HR labels and LR data for model training. However, the acquisition of meteorological data at regional scale is more challenging, making it difficult to meet the training requirements of traditional supervised-learning-based downscaling models in some cases. Second, learning the non-linear mapping between LR and HR meteorology data is typically an ill-posed issue, which means that there are infinite HR solutions for the same LR sample, making it harder to find the optimal solution within the large solution space, especially in the case of insufficient HR training labels. In this study, we propose a closed-loop circular regression network for simultaneous restoration of medium-resolution (MR) and HR 2m air temperature over Sichuan and surrounding areas, China. The model leverages the circular structure consistency to train both the downscaling and upscaling networks simultaneously. Specifically, in terms of insufficient HR labels, we introduce an additional constraint of MR supervision information to reduce the space of possible functions, forming a gradual downscaling process from LR to MR to HR data. Besides, we also establish an extra upscaling mapping from HR to MR to LR, which forms a circular consistency constraint on LR and MR data to provide additional supervision. Extensive experiments demonstrate that the proposed algorithm attains a Root Mean Square Error (RMSE) of 0.84 when utilizing 50% of the training data and 0.77 when using 75%. This performance surpasses that of many classic supervised-learning-based SD methods, even the complete supervised information is not utilized.
Guangyu Liu 0002, Renlong Hang, Rui Zhang 0049, Chunxiang Shi, Qingshan Liu 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 A Short-Long Term Sequence Learning Network for Precipitation Nowcasting
abstract
Precipitation nowcasting is a critical task for various applications, such as disaster mitigation, water resource management, traffic safety, and agricultural planning. In recent years, deep learning methods equipped with long short-term memory (LSTM) have become a mainstream method for this task. Typically, these methods take as input a radar echo sequence and focus on learning the temporal features of precipitation. However, due to the limited spatial modeling ability of the current LSTM modules, they cannot sufficiently learn the spatial–temporal features of precipitation. To address this issue, we propose in this article a short-long term sequence learning (SLTSL) network. SLTSL mainly consists of the short-term sequence learning module (SSLM) and the long-term sequence learning module (LSLM). SSLM weights and integrates the spatial distribution features of precipitation at each location of the short-term sequence by various weighted operations, where the short-term sequence is obtained by SSLM through matrix concatenation of the feature maps of four adjacent moments. LSLM integrates all feature maps into a long-term sequence through matrix fusion and then captures the temporal features of precipitation at all moments from the long-term sequence by means of multistate transitions and aggregation. In order to test the performance of the proposed network, we carry out experiments on three widely used datasets, including RadarCIKM, TAASRAD19, and RadarKNMI. The experimental results demonstrate that our proposed network can achieve superior or comparable performance to several state-of-the-art baseline methods.
Renlong Hang, Qingshan Liu 0001, Xiao-Tong Yuan
IEEE Trans. Geosci. Remote. Sens.2
2024 Estimating Tropical Cyclone Intensity Using an STIA Model From Himawari-8 Satellite Images in the Western North Pacific Basin
abstract
Analyzing the temporal evolution of historical tropical cyclone (TC) structures is essential for accurate TC intensity estimation. In this article, a novel spatiotemporal interaction attention (STIA) model is proposed to estimate TC intensity using Himawari-8 data in the western North Pacific (WNP) basin. The model incorporates a spatial feature extraction module and a spatiotemporal interaction module, which leverage historical satellite images. Based on a sequence of observed satellite images, the spatial feature extraction module is expected to extract spatial features of each TC frame. After that, the spatiotemporal interaction module comprising the temporal–spatial (TS) module and the spatial–temporal (ST) module is responsible for fusing the temporal and spatial features of each frame. The experimental data are composed of Himawari-8 infrared (IR) and water vapor (WV) images from 2015 to 2020 with a time interval of 1 h. The model is trained on images from 2015 to 2018 and evaluated on images from 2019 to 2020. Ablation experiments are conducted to analyze the impact of the number of frames, and the ST and TS modules. The results demonstrate that using 18-frame inputs yields the best performance, achieving an overall root-mean-square error (RMSE) of 3.61 m/s and a mean absolute error (MAE) of 2.83 m/s. In addition, the ST and TS modules significantly contribute to enhancing the accuracy of TC intensity estimation. The performance of the STIA model already surpasses the state-of-the-art benchmarks, demonstrating its excellence in TC intensity estimation.
Rui Zhang 0049, Luhui Yue, Qingshan Liu 0001, Renlong Hang
IEEE Trans. Geosci. Remote. Sens.5
2024 Masked Spectral-Spatial Feature Prediction for Hyperspectral Image Classification
abstract
Transformer has emerged as a preferred method for hyperspectral (HS) image classification due to its ability to model long-range dependency. Whereas the transformer contains numerous parameters and further available labeled HS data is limited, which makes it difficult to get a well-trained transformer. Accordingly, we propose a novel HS image classification method called masked spectral–spatial feature prediction (MSSFP). It aims at helping the transformer understand the complicated spectral–spatial structures without labeled HS data, further improving the classification performance. Specifically, the input HS cube is first divided into two sequences along spectral and spatial dimensions, respectively. Then, a portion of these two sequences are masked out and we train a transformer-based encoder–decoder network to predict the hand-crafted features of masked regions. After pretraining, the encoder is fine-tuned to derive two classification results from input spectral and spatial sequences. Finally, spectral and spatial results are aggregated adaptively based on uncertainty comparison. In comparison experiments, MSSFP outperforms several state-of-the-art HS image classification methods on three benchmark datasets including Indian Pines (IP), Houston (HU), and Pavia University (PUS).
Feng Zhou 0006, Guowei Yang 0002, Renlong Hang, Qingshan Liu 0001
IEEE Trans. Geosci. Remote. Sens.4
2023 MSNet: Multi-Resolution Synergistic Networks for Adaptive Inference
abstract
Adaptive inference with multiple networks has attracted much attention for resource-limited image classification. It assumes that a large portion of test samples can be correctly classified by small networks with fewer layers or channels, which poses a great challenge for them. In this paper, we argue that large networks have abilities to help the small ones address this challenge if fully explored. To this end, we propose a multi-resolution synergistic network (MSNet) using two different kinds of fusion modules. The first one is a cross-branch aggregation module, which aims to transfer the high-resolution features to the low-resolution ones between neighboring branches. The other one is an adaptive distillation module, whose purpose is feeding the discriminative ability of the large network to the other ones. Via these two modules, the small networks will be powerful enough to correctly classify large numbers of test samples, thus improving the classification accuracy and inference efficiency. We evaluate MSNet on three benchmark datasets: CIFAR-10, CIFAR-100, and ImageNet. Experimental results show that our network can obtain better results than several state-of-the-art networks in both anytime classification and budgeted batch classification settings. The code is available athttps://github.com/bigdata-qian/MSNet-Pytorch.
Renlong Hang, Xuwei Qian, Qingshan Liu 0001
IEEE Trans. Circuits Syst. Video Technol.1
2023 Mining Joint Intraimage and Interimage Context for Remote Sensing Change Detection
abstract
Recent deep learning methods for change detection focus on excavating more discriminative context within individual images. However, due to seasonal change, noise, and so on, the appearance of objects tends to be more heterogeneous among various scenes. Consequently, the above intra-image context is inadequate to represent specific-category objects and pseudo changes would be inevitable in detection results. To deal with this issue, we propose a context aggregation network (CANet) to mine inter-image context over all training images for further enhancing intra-image context. Specifically, a Siamese network attached with temporal attention modules is served as a feature encoder to extract multi-scale temporal features from bitemporal images. Then, a context extraction module is devised to capture long-range spatial-channel context within individual images. Meanwhile, context representations of underlying categories in the scene are inferred using all training images in an unsupervised manner. Finally, these two kinds of contextual information are aggregated to one which is subsequently fed into a multi-scale fusion module to produce the detection map. CANet is compared with several state-of-the-art methods on three benchmark datasets, including the season-varying change detection (SVCD) dataset, the Sun Yat-sen University change detection (SYSU-CD) dataset, and the Learning Vision and Remote Sensing Laboratory building change detection (LEVIR-CD) dataset. It is demonstrated that our method outperforms all comparison methods in terms of F1, overall accuracy (OA), and Intersection-of-Union (IoU). The results of CANet on three datasets are available at https://github.com/NuistZF/CANet-for-change-detection and codes will be public soon.
Feng Zhou 0006, Renlong Hang, Rui Zhang 0049, Qingshan Liu 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 ReX: An Efficient Approach to Reducing Memory Cost in Image Classification
abstract
Exiting simple samples in adaptive multi-exit networks through early modules is an effective way to achieve high computational efficiency. One can observe that deployments of multi-exit architectures on resource-constrained devices are easily limited by high memory footprint of early modules. In this paper, we propose a novel approach named recurrent aggregation operator (ReX), which uses recurrent neural networks (RNNs) to effectively aggregate intra-patch features within a large receptive field to get delicate local representations, while bypassing large early activations. The resulting model, named ReXNet, can be easily extended to dynamic inference by introducing a novel consistency-based early exit criteria, which is based on the consistency of classification decisions over several modules, rather than the entropy of the prediction distribution. Extensive experiments on two benchmark datasets, i.e., Visual Wake Words, ImageNet-1k, demonstrate that our method consistently reduces the peak RAM and average latency of a wide variety of adaptive models on low-power devices.
Xuwei Qian, Renlong Hang, Qingshan Liu 0001
AAAI2
2022 Deep Encoder-Decoder Networks for Classification of Hyperspectral and LiDAR Data
abstract
Deep learning (DL) has been garnering increasing attention in remote sensing (RS) due to its powerful data representation ability. In particular, deep models have been proven to be effective for RS data classification based on a single given modality. However, with one single modality, the ability in identifying the materials remains limited due to the lack of feature diversity. To overcome this limitation, we present a simple but effective multimodal DL baseline by following a deep encoder–decoder network architecture, EndNet for short, for the classification of hyperspectral and light detection and ranging (LiDAR) data. EndNet fuses the multimodal information by enforcing the fused features to reconstruct the multimodal input in turn. Such a reconstruction strategy is capable of better activating the neurons across modalities compared with some conventional and widely used fusion strategies, e.g., early fusion, middle fusion, and late fusion. Extensive experiments conducted on two popular hyperspectral and LiDAR data sets demonstrate the superiority and effectiveness of the proposed EndNet in comparison with several state-of-the-art baselines in the hyperspectral-LiDAR classification task. The codes will be available athttps://github.com/danfenghong/IEEE_GRSL_EndNet, contributing to the RS community.
Danfeng Hong, Lianru Gao, Renlong Hang, Bing Zhang 0001, Jocelyn Chanussot
IEEE Geosci. Remote. Sens. Lett.3
2022 Spectral-Spatial Correlation Exploration for Hyperspectral Image Classification via Self-Mutual Attention Network
abstract
Recently, deep learning methods have been widely used to extract spectral-spatial features for hyperspectral image (HSI) classification, and dramatically boost the performance. However, most of them usually take the original HSI cube as the input, where spectral-spatial information are mixed together. Consequently, they cannot explicitly model the inherent correlation (e.g., complementary relation) between spectral and spatial domains, limiting the classification performance. To alleviate this issue, a spectral-spatial self-mutual attention network (S3MANet) is proposed in this letter. It respectively extracts spectral and spatial features via the corresponding feature module. Subsequently, a self-mutual attention module is designed to enhance these features. More concretely, it performs feature interaction to emphasize the correlation of spectral and spatial domains via mutual attention while self attention is applied to each domain for learning long-range dependencies. Finally, we infer two classification results from the enhanced spectral and spatial features, and a weighted summation is further applied to obtain a joint spectral-spatial results. Experimental results on two public HSI datasets validate that the proposed S3MANet could achieve more satisfactory performance in comparison with several state-of-the-art methods.
Feng Zhou 0006, Renlong Hang
IEEE Geosci. Remote. Sens. Lett.2
2022 Hierarchical Context Network for Airborne Image Segmentation
abstract
Most of the recent methods focus on capturing contextual information by measuring relations (e.g., feature similarity) between each pixel and all the others for airborne image segmentation. Nevertheless, these methods have difficulty in handling confusing objects with a partially similar appearance. In this article, we attempt to simultaneously explore pixel-to-pixel (P2P) and pixel-to-object (P2O) relations to learn contextual information. For this purpose, a hierarchical context network (HCNet) is proposed. It consists of a P2P subnetwork and a P2O subnetwork. The P2P subnetwork learns the P2P relation (detail-grained context) for better preservation of the details (e.g., boundary) of the objects. Meanwhile, the P2O subnetwork models the P2O relation (semantic-grained context), aiming at improving the intraobject semantic consistency. When inferring the segmentation results, outputs of these two subnetworks are aggregated to obtain the hierarchical contextual information. Experimental results demonstrate that the proposed model achieves competitive performance on three challenging benchmarks.
Feng Zhou 0006, Renlong Hang, Hui Shuai, Qingshan Liu 0001
IEEE Trans. Geosci. Remote. Sens.2
2022 Cross-Modality Contrastive Learning for Hyperspectral Image Classification
abstract
Deep learning has attracted much attention in the field of hyperspectral image classification recently, due to its powerful representation and generalization abilities. Most of current deep learning models are trained in a supervised manner, which require large amounts of labeled samples to achieve state-of-the-art performance. Unfortunately, pixel-level labeling in hyperspectral imageries is difficult, time-consuming, and human-dependent. To address this issue, we propose an unsupervised feature learning model using multi-modal data, hyperspectral and LiDAR in particular. It takes advantage of the relationship between hyperspectral and LiDAR data to extract features, without using any label information. After that, we design a dual fine-tuning strategy to transfer the extracted features for hyperspectral image classification with small numbers of training samples. Such strategy is able to explore not only the semantic information but also the intrinsic structure information of training samples. In order to test the performance of our proposed model, we conduct comprehensive experiments on three hyperspectral and LiDAR datasets. Experimental results show that our proposed model can achieve better performance than several state-of-the-art deep learning models.
Renlong Hang, Xuwei Qian, Qingshan Liu 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 Multiscale Progressive Segmentation Network for High-Resolution Remote Sensing Imagery
abstract
Semantic segmentation of high-resolution remote sensing imageries (HRSIs) is a critical task for a wide range of applications, such as precision agriculture and urban planning. Although convolutional neural networks (CNNs) have made great progress in accomplishing this task recently, there still exist some challenges to address, one of which is simultaneously segmenting objects with large scale variations in a HRSI. Targeting at this challenge, previous CNNs often adopt multiple convolution kernels in one layer or skip-layer connections between different layers to extract multiscale representations. However, due to the limited learning capacity of each CNN, it tends to make trade-offs in segmenting different-scale objects. This would lead to unsatisfactory segmentation results for some objects, especially the small or the large ones. In this paper, we propose a multiscale progressive segmentation network to address this issue. Instead of forcing one network to deal with all scales of objects, our network attempts to cascade three subnetworks for gradually segmenting objects with small scales, large scales, and other scales. In order to make the subnetwork focus on the specific scale objects, a scale guidance module is designed. It takes advantage of segmentation results from the preceding subnetwork to guide the feature learning of the succeeding one. Additionally, to acquire the final segmentation results, we propose a position sensitive module for adaptively combining the outputs of the three subnetworks. This module is capable of assigning combination weights of different subnetworks according to their importance. Experiments on two benchmark datasets named Vaihingen and Potsdam indicate that our proposed network can achieve considerable improvements in comparison with several state-of-the-art segmentation models.
Renlong Hang, Feng Zhou 0006, Qingshan Liu 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 Global Tropical Cyclone Precipitation Estimation via a Multitask Convolutional Neural Network Based on HURSAT-B1 Data
abstract
Fast and accurate global tropical cyclone (TC) precipitation estimation from satellite observations is still a challenging issue. In this article, we propose an effective model based on a multitask convolutional neural network (CNN) to estimate near-real-time global TC precipitation from HURSAT-B1 data. Our network mainly consists of three modules: the feature extraction module, the wind grade classification module, and the precipitation estimation module. The first module aims at extracting the spatial features of satellite imageries, the second module focuses on classifying the wind grades of the satellite imageries into six categories that are used to assist in estimating TC precipitation, and the third module is to estimate TC precipitation. To evaluate the effectiveness of our proposed model, we compare it with multiple linear regression (MLR) and random forest (RF) models based on integrated multisatellite retrievals for the global precipitation measurement (GPM) mission (IMERG). Besides, four typical TC events are selected to specifically analyze the temporal and spatial distribution of TC precipitation estimation. Experimental results show that the probability of detection and accuracy achieved by our proposed model are 0.68 and 0.81, while the correlation coefficient (CC) and MSE are 0.61 and 7.80, respectively. In terms of the four TC events, our proposed model obtains a more consistent and continuous spatial distribution of precipitation than MLR and RF. More importantly, our proposed model can achieve high spatiotemporal results, which has the potential to serve as an operational algorithm for global TC precipitation estimation.
Mei Xue, Renlong Hang, Xiao-Tong Yuan, Qingshan Liu 0001
IEEE Trans. Geosci. Remote. Sens.2
2022 Predicting Tropical Cyclogenesis Using a Deep Learning Method From Gridded Satellite and ERA5 Reanalysis Data in the Western North Pacific Basin
abstract
This article proposes a deep learning model to predict tropical cyclogenesis (TCG) from gridded satellite and ERA5 reanalysis data in the western North Pacific basin. The proposed model contains two modules. First, convolutional neural network (CNN)-based deep features are extracted for each predictor, and then, the extracted features are fused with two fully connected layers to differentiate and investigate the relationship between predictors and TCG. The experimental data of this study are composed of 3232 developing tropical cluster clouds and 6657 nondeveloping ones; 90% of the collected data are utilized to train the model, and the rest are used to evaluate the trained model. Totally, nine predictors have been considered for the study, and the results show that the brightness temperature (IR), relative vorticity (Vo), and geopotential height (Z) perform better than the other predictors. A combined model with six predictors [IR, Z, RH (relative humidity), Vo, WS10 m(wind speed at the height of ten meters above the surface of the Earth), and mslp (mean sea-level pressure)] achieves the best TCG predicting performance, i.e., 97.1% of developing tropical cyclones are detected at a probability threshold of 0.13 with a false alarm rate of 20.3%. The experimental results demonstrate that the proposed method is superior to the existing methods and also indicate that the fusion of satellite and reanalysis data is a promising method to predict TCG.
Rui Zhang 0049, Qingshan Liu 0001, Renlong Hang, Guangcan Liu
IEEE Trans. Geosci. Remote. Sens.3
2021 Hyperspectral Image Classification With Attention-Aided CNNs
abstract
Convolutional neural networks (CNNs) have been widely used for hyperspectral image classification. As a common process, small cubes are first cropped from the hyperspectral image and then fed into CNNs to extract spectral and spatial features. It is well known that different spectral bands and spatial positions in the cubes have different discriminative abilities. If fully explored, this prior information will help improve the learning capacity of CNNs. Along this direction, we propose an attention-aided CNN model for spectral-spatial classification of hyperspectral images. Specifically, a spectral attention subnetwork and a spatial attention subnetwork are proposed for spectral and spatial classifications, respectively. Both of them are based on the traditional CNN model and incorporate attention modules to aid networks that focus on more discriminative channels or positions. In the final classification phase, the spectral classification result and the spatial classification result are combined together via an adaptively weighted summation method. To evaluate the effectiveness of the proposed model, we conduct experiments on three standard hyperspectral data sets. The experimental results show that the proposed model can achieve superior performance compared with several state-of-the-art CNN-related models.
Renlong Hang, Zhu Li 0001, Qingshan Liu 0001, Pedram Ghamisi, Shuvra S. Bhattacharyya
IEEE Trans. Geosci. Remote. Sens.1
2021 Classification of Hyperspectral Images via Multitask Generative Adversarial Networks
abstract
Deep learning has shown its huge potential in the field of hyperspectral image (HSI) classification. However, most of the deep learning models heavily depend on the quantity of available training samples. In this article, we propose a multitask generative adversarial network (MTGAN) to alleviate this issue by taking advantage of the rich information from unlabeled samples. Specifically, we design a generator network to simultaneously undertake two tasks: the reconstruction task and the classification task. The former task aims at reconstructing an input hyperspectral cube, including the labeled and unlabeled ones, whereas the latter task attempts to recognize the category of the cube. Meanwhile, we construct a discriminator network to discriminate the input sample coming from the real distribution or the reconstructed one. Through an adversarial learning method, the generator network will produce real-like cubes, thus indirectly improving the discrimination and generalization ability of the classification task. More importantly, in order to fully explore the useful information from shallow layers, we adopt skip-layer connections in both reconstruction and classification tasks. The proposed MTGAN model is implemented on three standard HSIs, and the experimental results show that it is able to achieve higher performance than other state-of-the-art deep learning models.
Renlong Hang, Feng Zhou 0006, Qingshan Liu 0001, Pedram Ghamisi
IEEE Trans. Geosci. Remote. Sens.1
2021 Class-Guided Feature Decoupling Network for Airborne Image Segmentation
abstract
Contextual information has been demonstrated to be helpful for airborne image segmentation. However, most of the previous works focus on the exploitation of spatially contextual information, which is difficult to segment isolated objects, mainly surrounded by uncorrelated objects. To alleviate this issue, we attempt to take advantage of the co-occurrence relations between different classes of objects in the scene. Especially, similar to other works, convolutional features are first extracted to capture the spatially contextual information. Then, a feature decoupling module is designed to encode the class co-occurrence relations into the convolutional features; thus, the most discriminative features can be decoupled. Finally, the segmentation result is inferred from the decoupled features. The whole process is integrated to form an end-to-end network, named class-guided feature decoupling network (CGFDN). Experimental results on two widely used benchmark data sets show that CGFDN obtains competitive results (>90% overall accuracy (OA) on 5-cm-resolution Potsdam and >91% OA on 9-cm-resolution Vaihingen) in comparison with several state-of-the-art models.
Feng Zhou 0006, Renlong Hang, Qingshan Liu 0001
IEEE Trans. Geosci. Remote. Sens.2
2021 Spectral Super-Resolution Network Guided by Intrinsic Properties of Hyperspectral Imagery
abstract
Hyperspectral imagery (HSI) contains rich spectral information, which is beneficial to many tasks. However, acquiring HSI is difficult because of the limitations of current imaging technology. As an alternative method, spectral super-resolution aims at reconstructing HSI from its corresponding RGB image. Recently, deep learning has shown its power to this task, but most of the used networks are transferred from other domains, such as spatial super-resolution. In this paper, we attempt to design a spectral super-resolution network by taking advantage of two intrinsic properties of HSI. The first one is the spectral correlation. Based on this property, a decomposition subnetwork is designed to reconstruct HSI. The other one is the projection property, i.e., RGB image can be regarded as a three-dimensional projection of HSI. Inspired from it, a self-supervised subnetwork is constructed as a constraint to the decomposition subnetwork. These two subnetworks constitute our end-to-end super-resolution network. In order to test the effectiveness of it, we conduct experiments on three widely used HSI datasets (i.e., CAVE, NUS, and NTIRE2018). Experimental results show that our proposed network can achieve competitive reconstruction performance in comparison with several state-of-the-art networks.
Renlong Hang, Qingshan Liu 0001, Zhu Li 0001
IEEE Trans. Image Process.1
2020 Prinet: A Prior Driven Spectral Super-Resolution Network
abstract
Spectral super-resolution aims to reconstruct hyperspectral images from RGB images directly. In recent years, convolutional networks have been successfully employed to this task. However, few of them take into account the specific properties of hyperspectral images. In this paper, we attempt to design a super-resolution network, named PriNET, based on two prior knowledge about hyperspectral images. The first one is spectral correlation. According to this property, we design a decomposition network to reconstruct hyperspectral images. In this network, the whole spectral bands of hyperspectral images are divided into several groups, and multiple residual networks are proposed to reconstruct them separately. The second knowledge is that the hyperspectral image should be able to generate its corresponding RGB image. Inspired from it, we design a self-supervised network to fine-tune the reconstruction results of the decomposition network. Finally, these two networks are combined together to constitute PriNET. Experimental results on two hyperspectral datasets demonstrate that the proposed PriNET can achieve better performance than several state-of-the-art networks.
Renlong Hang, Zhu Li 0001, Qingshan Liu 0001, Shuvra S. Bhattacharyya
ICME1
2020 Locally Linear Reconstruction for Spectral Enhancement Using Limited Pixel-to-Pixel Multispectral and Hyperspectral Data
abstract
Recently, spectral enhancement of multispectral imagery has attracted a growing interest in the remote sensing (RS) community. Without any prior knowledge, this task is highly ill-conditioned in inverse problems. To this end, we develop a simple but effective method, called locally linear reconstruction (LLR), to spectrally enhance the multispectral imagery (MSI) using partially overlapped hyperspectral data. LLR learns reconstruction coefficients of each pixel from the MSI and shares the same weights to recover the unknown hyperspectral signals over a larger coverage. We validate the performance of the proposed LLR on the real hyperspectral data in comparison with several state-of-the-art baselines, demonstrating its effectiveness and superiority.
Danfeng Hong, Jing Yao 0002, Renlong Hang, Jocelyn Chanussot
IGARSS3
2020 Classification of Hyperspectral and LiDAR Data Using Coupled CNNs
abstract
In this article, we propose an efficient and effective framework to fuse hyperspectral and light detection and ranging (LiDAR) data using two coupled convolutional neural networks (CNNs). One CNN is designed to learn spectral-spatial features from hyperspectral data, and the other one is used to capture the elevation information from LiDAR data. Both of them consist of three convolutional layers, and the last two convolutional layers are coupled together via a parameter-sharing strategy. In the fusion phase, feature-level and decision-level fusion methods are simultaneously used to integrate these heterogeneous features sufficiently. For the feature-level fusion, three different fusion strategies are evaluated, including the concatenation strategy, the maximization strategy, and the summation strategy. For the decision-level fusion, a weighted summation strategy is adopted, where the weights are determined by the classification accuracy of each output. The proposed model is evaluated on an urban data set acquired over Houston, USA, and a rural one captured over Trento, Italy. On the Houston data, our model can achieve a new record overall accuracy (OA) of 96.03%. On the Trento data, it achieves an OA of 99.12%. These results sufficiently certify the effectiveness of our proposed model.
Renlong Hang, Zhu Li 0001, Pedram Ghamisi, Danfeng Hong, Guiyu Xia, Qingshan Liu 0001
IEEE Trans. Geosci. Remote. Sens.1
2020 Tropical Cyclone Intensity Estimation Using Two-Branch Convolutional Neural Network From Infrared and Water Vapor Images
abstract
This article proposes a two-branch convolutional neural network model (TCIENet) to estimate the intensity of tropical cyclone (TC) from infrared and water vapor images in the northwest Pacific basin. Three different sizes of input images are explored to train the TCIENet model, and the size of 60 × 60 pixels (radius 450 km) achieves the best performance with an overall root mean square error (RMSE) of 5.13 m/s and mean absolute error (MAE) of 4.03 m/s. TCs are divided into six categories whose RMSEs range from 4.07 to 6.05 m/s. In addition, the TCs in the year 2017 are used to analyze the correlation between the rainfall intensity from the global precipitation measurement (GPM) mission and the estimation errors of the TCIENet model. Preliminary results suggest that the model performs the best at the categories of tropical storm and super typhoon, but it degrades in performance for moderate intense categories and the weakest category of the tropical depression. The correlation coefficient between the estimation error and the rainfall intensity is 0.19. It is far from certain that the rainfall intensity accounts for the error achieved by the TCIENet model.
Rui Zhang 0049, Qingshan Liu 0001, Renlong Hang
IEEE Trans. Geosci. Remote. Sens.3
2019 Low Resolution Recognition of Aerial Images
abstract
Remote sensing for classification has been widely studied and is useful for a lot of applications like precision agriculture, surveillance, and military applications. Recently, due to tremendous results achieved by deep learning using Convolutional Neural Networks (CNN) for Imagenet dataset, there have been a large number of works which use deep learning for aerial image classification. Most of the works concentrate on original resolution and there are no works on low-resolution recognition of aerial images. This work is critical because aerial images are taken from a very high distance from the ground and the cost of installing high definition cameras is high, so it is hard to get a high resolution of the image. In this paper, we explore how we can do the better classification of aerial images for original spatial resolution and low spatial resolution in deep learning by using texture information. In our framework, we use YUV color space which is generally used for video coding and we also use Laplacian of Gaussian (LOG) information to exploit the texture information. We decouple RGB information into luminance information (Y channel), color information (UV) and texture information (LOG) and we train a separate CNN for each feature and combine them using autoencoder and with our results, we show that we do better than RGB images in original resolution and low resolution.
Raghunath Sai Puttagunta, Renlong Hang, Zhu Li 0001, Shuvra S. Bhattacharyya
VCIP2
2019 Privacy-Preserving Fall Detection with Deep Learning on mmWave Radar Signal
abstract
Fall is one of the main reasons for body injuries among seniors. Traditional fall detection methods are mainly achieved by wearable and non-wearable techniques, which may cause skin discomfort or invasion of privacy to users. In this paper, we propose an automatic fall detection method with the assist of the mmWave radar signal to solve the aforementioned issues. The radar devices are capable to record the reflection from objects in both the spatial and temporal domain, which can be used to depict the activities of users with the support of a recurrent neural network (RNN) with long-short-term memory (LSTM) units. First, we employ the radar low-dimension embedding (RLDE) algorithm to preprocess the Range-angle reflection heatmap sequence converted from the raw radar signal for reducing the redundancy in the spatial domain. Then, the processed sequence is split into frames for inputting LSTM units one by one. Eventually, the output from the last LSTM unit is fed in a Softmax layer for classifying different activities. To validate the effectiveness of our proposed method, we construct a radar dataset with the assist of market radar module devices, to implement several experiments. The experimental results demonstrate that, compared to LSTM only and the widely used 3-D convolutional neural network (3-D CNN), combining RLDE and LSTM can achieve the best detection results with much less computational time consumption. In addition, we extend the proposed method to classify multiple human activities simultaneously and the satisfied performances are observed.
Yangfan Sun, Renlong Hang, Zhu Li 0001, Mouqing Jin, Kelvin Xu
VCIP2
2019 Hyperspectral image classification using spectral-spatial LSTMs
Feng Zhou 0006, Renlong Hang, Qingshan Liu 0001, Xiao-Tong Yuan
Neurocomputing2
2019 Retrieving Soil Moisture Over Continental U.S. via Multi-View Multi-Task Learning
abstract
Soil moisture (SM) is an essential variable in the hydrological cycle. Quantifying the magnitude of SM is crucial for the climate system. In this letter, we present a new model with multi-view multi-task learning (MVMTL) to estimate SM over continental U.S. Specifically, the multi-view component is used to make full use of spatial and temporal features of each grid cell (0.25° × 0.25°). Meanwhile, the multi-task component aims to capture the spatial correlations and to perform coestimations between different grid cells in the study area. To evaluate the effectiveness of MVMTL, we compare it with several retrieval methods in terms of the SM product from the European Center for Medium-Range Weather Forecasts Reanalysis Interim (ERA-Interim) and in situ SM measurements. The experimental results show that the MVMTL model can achieve higher performance than the other methods.
Lingling Ge, Renlong Hang, Qingshan Liu 0001
IEEE Geosci. Remote. Sens. Lett.2
2019 Cascaded Recurrent Neural Networks for Hyperspectral Image Classification
abstract
By considering the spectral signature as a sequence, recurrent neural networks (RNNs) have been successfully used to learn discriminative features from hyperspectral images (HSIs) recently. However, most of these models only input the whole spectral bands into RNNs directly, which may not fully explore the specific properties of HSIs. In this paper, we propose a cascaded RNN model using gated recurrent units to explore the redundant and complementary information of HSIs. It mainly consists of two RNN layers. The first RNN layer is used to eliminate redundant information between adjacent spectral bands, while the second RNN layer aims to learn the complementary information from nonadjacent spectral bands. To improve the discriminative ability of the learned features, we design two strategies for the proposed model. Besides, considering the rich spatial information contained in HSIs, we further extend the proposed model to its spectral-spatial counterpart by incorporating some convolutional layers. To test the effectiveness of our proposed models, we conduct experiments on two widely used HSIs. The experimental results show that our proposed models can achieve better results than the compared models.
Renlong Hang, Qingshan Liu 0001, Danfeng Hong, Pedram Ghamisi
IEEE Trans. Geosci. Remote. Sens.1
2019 Learning-Based Sphere Nonlinear Interpolation for Motion Synthesis
abstract
Motion synthesis technology can produce natural and coordinated motion data without a motion capture process, which is complex and costly. Current motion synthesis methods usually provide a few interfaces to avoid the arbitrariness of the synthesis process, but this actually reduces the understandability of the synthesis process. In this paper, we propose a learning-based Sphere nonlinear interpolation (Snerp) model that can generate natural in-between motions in terms of a given start-end frame pair. Variety of the input frame pairs will enrich the diversity of the generated motions. The angle speed of natural human motion is not uniform and presents different change rules (we call them motion patterns) for different motions, so we first extract the motion patterns and then build the relation between motion pattern space and frame pair space via a paired dictionary learning process. After learning, we estimate the motion pattern according to the representation of a given start-end frame pair on the frame pair dictionary. We select several different types of start-end frame pairs from the real motion sequences as the testing data and good results of both objective and subjective evaluations on the generated motions demonstrate the superior performance of Snerp.
Guiyu Xia, Huaijiang Sun, Qingshan Liu 0001, Renlong Hang
IEEE Trans. Ind. Informatics4
2018 Integrating Convolutional Neural Network and Gated Recurrent Unit for Hyperspectral Image Spectral-Spatial Classification
Feng Zhou 0006, Renlong Hang, Qingshan Liu 0001, Xiao-Tong Yuan
PRCV (4)2
2018 Multi-component group sparse RPCA model for motion object detection under complex dynamic background
Yubao Sun, Renlong Hang, Qingshan Liu 0001, Guangcan Liu
Neurocomputing3
2018 Learning Multiscale Deep Features for High-Resolution Satellite Image Scene Classification
abstract
In this paper, we propose a multiscale deep feature learning method for high-resolution satellite image scene classification. Specifically, we first warp the original satellite image into multiple different scales. The images in each scale are employed to train a deep convolutional neural network (DCNN). However, simultaneously training multiple DCNNs is time-consuming. To address this issue, we explore DCNN with spatial pyramid pooling (SPP-net). Since different SPP-nets have the same number of parameters, which share the identical initial values, and only fine-tuning the parameters in fully connected layers ensures the effectiveness of each network, thereby greatly accelerating the training process. Then, the multiscale satellite images are fed into their corresponding SPP-nets, respectively, to extract multiscale deep features. Finally, a multiple kernel learning method is developed to automatically learn the optimal combination of such features. Experiments on two difficult data sets show that the proposed method achieves favorable performance compared with other state-of-the-art methods.
Qingshan Liu 0001, Renlong Hang, Huihui Song 0002
IEEE Trans. Geosci. Remote. Sens.2
2018 Nonlinear Low-Rank Matrix Completion for Human Motion Recovery
abstract
Human motion capture data has been widely used in many areas, but it involves a complex capture process and the captured data inevitably contains missing data due to the occlusions caused by the actor's body or clothing. Motion recovery, which aims to recover the underlying complete motion sequence from its degraded observation, still remains as a challenging task due to the nonlinear structure and kinematics property embedded in motion data. Low-rank matrix completion based methods have shown promising performance in short-time-missing motion recovery problems. However, low-rank matrix completion, which is designed for linear data, lacks the theoretic guarantee when applied to the recovery of nonlinear motion data. To overcome this drawback, we propose a tailored nonlinear matrix completion model for human motion recovery. Within the model, we first learn a combined low-rank kernel via multiple kernel learning. By exploiting the learned kernel, we embed the motion data into a high dimensional Hilbert space where motion data is of desirable low-rank and we then use the low-rank matrix completion to recover motions. In addition, we add two kinematic constraints to the proposed model to preserve the kinematics property of human motion. Extensive experiment results and comparisons with five other state-of-the-art methods demonstrate the advantage of the proposed method.
Guiyu Xia, Huaijiang Sun, Beijia Chen, Qingshan Liu 0001, Lei Feng 0003, Guoqing Zhang 0002, Renlong Hang
IEEE Trans. Image Process.7
2017 Multicore implementation of the multi-scale adaptive deep pyramid matching model for remotely sensed image classification
abstract
Artificial neural networks (ANNs) have been widely used in the analysis of remotely sensed imagery. In particular, convolutional neural networks (CNNs) are gaining more and more attention. Unlike traditional CNNs methods, where the relevant information to classify the elements of a remotely sensed image is extracted only from the last fully-connected layer, the new adaptive deep pyramid matching (ADPM) model [1] takes advantage of the features from all of the convolutional layers. This model allows the optimal fusing weights for different convolutional layers be learned from the data itself. In addition, the combination of CNNs with spatial pyramid pooling (SPP-net) to create the basic deep network allows the use of images with multiple scales, which results in better learning process thanks to the complementary information. The original ADPM method is divided in two parts: the multi-scale deep feature extraction and the ADPM core. In this paper we present a computational improvement of the ADPM core, coding a parallel-multicore version. This strategy is shown to significantly enhance performance in the analysis of remotely sensed data.
Mercedes Eugenia Paoletti, Juan Mario Haut, Javier Plaza, Antonio Plaza, Qingshan Liu 0001, Renlong Hang
IGARSS6
2016 Matrix-Based Discriminant Subspace Ensemble for Hyperspectral Image Spatial-Spectral Feature Fusion
abstract
Spatial-spectral feature fusion is well acknowledged as an effective method for hyperspectral (HS) image classification. Many previous studies have been devoted to this subject. However, these methods often regard the spatial-spectral high-dimensional data as 1-D vector and then extract informative features for classification. In this paper, we propose a new HS image classification method. Specifically, matrix-based spatial-spectral feature representation is designed for each pixel to capture the local spatial contextual and the spectral information of all the bands, which can well preserve the spatial-spectral correlation. Then, matrix-based discriminant analysis is adopted to learn the discriminative feature subspace for classification. To further improve the performance of discriminative subspace, a random sampling technique is used to produce a subspace ensemble for final HS image classification. Experiments are conducted on three HS remote sensing data sets acquired by different sensors, and experimental results demonstrate the efficiency of the proposed method.
Renlong Hang, Qingshan Liu 0001, Huihui Song 0002, Yubao Sun
IEEE Trans. Geosci. Remote. Sens.1