Feng Zhou 0006

dblp:21/6430-6 · DBLP profile ↗
← Back
13ranked-venue papers
9as first author
10since 2021 · last 2026
0000-0002-1519-1001ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Causal-SAM: Enhancing segment anything for remote sensing instance segmentation via causal representation learning
Yunhui Zhu, Buliao Huang, Feng Zhou 0006
Knowl. Based Syst.3
2026 Adaptive frequency collaboration for remote sensing change detection
Feng Zhou 0006, Hui Shuai, Qingshan Liu 0001, Renlong Hang
Neural Networks1
2024 Masked Spectral-Spatial Feature Prediction for Hyperspectral Image Classification
abstract
Transformer has emerged as a preferred method for hyperspectral (HS) image classification due to its ability to model long-range dependency. Whereas the transformer contains numerous parameters and further available labeled HS data is limited, which makes it difficult to get a well-trained transformer. Accordingly, we propose a novel HS image classification method called masked spectral–spatial feature prediction (MSSFP). It aims at helping the transformer understand the complicated spectral–spatial structures without labeled HS data, further improving the classification performance. Specifically, the input HS cube is first divided into two sequences along spectral and spatial dimensions, respectively. Then, a portion of these two sequences are masked out and we train a transformer-based encoder–decoder network to predict the hand-crafted features of masked regions. After pretraining, the encoder is fine-tuned to derive two classification results from input spectral and spatial sequences. Finally, spectral and spatial results are aggregated adaptively based on uncertainty comparison. In comparison experiments, MSSFP outperforms several state-of-the-art HS image classification methods on three benchmark datasets including Indian Pines (IP), Houston (HU), and Pavia University (PUS).
Feng Zhou 0006, Guowei Yang 0002, Renlong Hang, Qingshan Liu 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 Mining Joint Intraimage and Interimage Context for Remote Sensing Change Detection
abstract
Recent deep learning methods for change detection focus on excavating more discriminative context within individual images. However, due to seasonal change, noise, and so on, the appearance of objects tends to be more heterogeneous among various scenes. Consequently, the above intra-image context is inadequate to represent specific-category objects and pseudo changes would be inevitable in detection results. To deal with this issue, we propose a context aggregation network (CANet) to mine inter-image context over all training images for further enhancing intra-image context. Specifically, a Siamese network attached with temporal attention modules is served as a feature encoder to extract multi-scale temporal features from bitemporal images. Then, a context extraction module is devised to capture long-range spatial-channel context within individual images. Meanwhile, context representations of underlying categories in the scene are inferred using all training images in an unsupervised manner. Finally, these two kinds of contextual information are aggregated to one which is subsequently fed into a multi-scale fusion module to produce the detection map. CANet is compared with several state-of-the-art methods on three benchmark datasets, including the season-varying change detection (SVCD) dataset, the Sun Yat-sen University change detection (SYSU-CD) dataset, and the Learning Vision and Remote Sensing Laboratory building change detection (LEVIR-CD) dataset. It is demonstrated that our method outperforms all comparison methods in terms of F1, overall accuracy (OA), and Intersection-of-Union (IoU). The results of CANet on three datasets are available at https://github.com/NuistZF/CANet-for-change-detection and codes will be public soon.
Feng Zhou 0006, Renlong Hang, Rui Zhang 0049, Qingshan Liu 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 Spectral-Spatial Correlation Exploration for Hyperspectral Image Classification via Self-Mutual Attention Network
abstract
Recently, deep learning methods have been widely used to extract spectral-spatial features for hyperspectral image (HSI) classification, and dramatically boost the performance. However, most of them usually take the original HSI cube as the input, where spectral-spatial information are mixed together. Consequently, they cannot explicitly model the inherent correlation (e.g., complementary relation) between spectral and spatial domains, limiting the classification performance. To alleviate this issue, a spectral-spatial self-mutual attention network (S3MANet) is proposed in this letter. It respectively extracts spectral and spatial features via the corresponding feature module. Subsequently, a self-mutual attention module is designed to enhance these features. More concretely, it performs feature interaction to emphasize the correlation of spectral and spatial domains via mutual attention while self attention is applied to each domain for learning long-range dependencies. Finally, we infer two classification results from the enhanced spectral and spatial features, and a weighted summation is further applied to obtain a joint spectral-spatial results. Experimental results on two public HSI datasets validate that the proposed S3MANet could achieve more satisfactory performance in comparison with several state-of-the-art methods.
Feng Zhou 0006, Renlong Hang
IEEE Geosci. Remote. Sens. Lett.1
2022 Hierarchical Context Network for Airborne Image Segmentation
abstract
Most of the recent methods focus on capturing contextual information by measuring relations (e.g., feature similarity) between each pixel and all the others for airborne image segmentation. Nevertheless, these methods have difficulty in handling confusing objects with a partially similar appearance. In this article, we attempt to simultaneously explore pixel-to-pixel (P2P) and pixel-to-object (P2O) relations to learn contextual information. For this purpose, a hierarchical context network (HCNet) is proposed. It consists of a P2P subnetwork and a P2O subnetwork. The P2P subnetwork learns the P2P relation (detail-grained context) for better preservation of the details (e.g., boundary) of the objects. Meanwhile, the P2O subnetwork models the P2O relation (semantic-grained context), aiming at improving the intraobject semantic consistency. When inferring the segmentation results, outputs of these two subnetworks are aggregated to obtain the hierarchical contextual information. Experimental results demonstrate that the proposed model achieves competitive performance on three challenging benchmarks.
Feng Zhou 0006, Renlong Hang, Hui Shuai, Qingshan Liu 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 Multiscale Progressive Segmentation Network for High-Resolution Remote Sensing Imagery
abstract
Semantic segmentation of high-resolution remote sensing imageries (HRSIs) is a critical task for a wide range of applications, such as precision agriculture and urban planning. Although convolutional neural networks (CNNs) have made great progress in accomplishing this task recently, there still exist some challenges to address, one of which is simultaneously segmenting objects with large scale variations in a HRSI. Targeting at this challenge, previous CNNs often adopt multiple convolution kernels in one layer or skip-layer connections between different layers to extract multiscale representations. However, due to the limited learning capacity of each CNN, it tends to make trade-offs in segmenting different-scale objects. This would lead to unsatisfactory segmentation results for some objects, especially the small or the large ones. In this paper, we propose a multiscale progressive segmentation network to address this issue. Instead of forcing one network to deal with all scales of objects, our network attempts to cascade three subnetworks for gradually segmenting objects with small scales, large scales, and other scales. In order to make the subnetwork focus on the specific scale objects, a scale guidance module is designed. It takes advantage of segmentation results from the preceding subnetwork to guide the feature learning of the succeeding one. Additionally, to acquire the final segmentation results, we propose a position sensitive module for adaptively combining the outputs of the three subnetworks. This module is capable of assigning combination weights of different subnetworks according to their importance. Experiments on two benchmark datasets named Vaihingen and Potsdam indicate that our proposed network can achieve considerable improvements in comparison with several state-of-the-art segmentation models.
Renlong Hang, Feng Zhou 0006, Qingshan Liu 0001
IEEE Trans. Geosci. Remote. Sens.3
2021 Robust Lightweight Facial Expression Recognition Network with Label Distribution Training
abstract
This paper presents an efficiently robust facial expression recognition (FER) network, named EfficientFace, which holds much fewer parameters but more robust to the FER in the wild. Firstly, to improve the robustness of the lightweight network, a local-feature extractor and a channel-spatial modulator are designed, in which the depthwise convolution is employed. As a result, the network is aware of local and global-salient facial features. Then, considering the fact that most emotions occur as combinations, mixtures, or compounds of the basic emotions, we introduce a simple but efficient label distribution learning (LDL) method as a novel training strategy. Experiments conducted on realistic occlusion and pose variation datasets demonstrate that the proposed EfficientFace is robust under occlusion and pose variation conditions. Moreover, the proposed method achieves state-of-the-art results on RAF-DB, CAER-S, and AffectNet-7 datasets with accuracies of 88.36%, 85.87%, and 63.70%, respectively, and a comparable result on the AffectNet-8 dataset with an accuracy of 59.89%. The code is public available at https://github.com/zengqunzhao/EfficientFace.
Zengqun Zhao, Qingshan Liu 0001, Feng Zhou 0006
AAAI3
2021 Classification of Hyperspectral Images via Multitask Generative Adversarial Networks
abstract
Deep learning has shown its huge potential in the field of hyperspectral image (HSI) classification. However, most of the deep learning models heavily depend on the quantity of available training samples. In this article, we propose a multitask generative adversarial network (MTGAN) to alleviate this issue by taking advantage of the rich information from unlabeled samples. Specifically, we design a generator network to simultaneously undertake two tasks: the reconstruction task and the classification task. The former task aims at reconstructing an input hyperspectral cube, including the labeled and unlabeled ones, whereas the latter task attempts to recognize the category of the cube. Meanwhile, we construct a discriminator network to discriminate the input sample coming from the real distribution or the reconstructed one. Through an adversarial learning method, the generator network will produce real-like cubes, thus indirectly improving the discrimination and generalization ability of the classification task. More importantly, in order to fully explore the useful information from shallow layers, we adopt skip-layer connections in both reconstruction and classification tasks. The proposed MTGAN model is implemented on three standard HSIs, and the experimental results show that it is able to achieve higher performance than other state-of-the-art deep learning models.
Renlong Hang, Feng Zhou 0006, Qingshan Liu 0001, Pedram Ghamisi
IEEE Trans. Geosci. Remote. Sens.2
2021 Class-Guided Feature Decoupling Network for Airborne Image Segmentation
abstract
Contextual information has been demonstrated to be helpful for airborne image segmentation. However, most of the previous works focus on the exploitation of spatially contextual information, which is difficult to segment isolated objects, mainly surrounded by uncorrelated objects. To alleviate this issue, we attempt to take advantage of the co-occurrence relations between different classes of objects in the scene. Especially, similar to other works, convolutional features are first extracted to capture the spatially contextual information. Then, a feature decoupling module is designed to encode the class co-occurrence relations into the convolutional features; thus, the most discriminative features can be decoupled. Finally, the segmentation result is inferred from the decoupled features. The whole process is integrated to form an end-to-end network, named class-guided feature decoupling network (CGFDN). Experimental results on two widely used benchmark data sets show that CGFDN obtains competitive results (>90% overall accuracy (OA) on 5-cm-resolution Potsdam and >91% OA on 9-cm-resolution Vaihingen) in comparison with several state-of-the-art models.
Feng Zhou 0006, Renlong Hang, Qingshan Liu 0001
IEEE Trans. Geosci. Remote. Sens.1
2020 Flow driven attention network for video salient object detection
abstract
Salient object detection has been revolutionised by convolutional neural network (CNN) recently. However, it is hard to transfer the state‐of‐the‐art still‐image based saliency detectors to videos directly, owing to the neglect of temporal contexts between frames. In this study, the authors propose a flow‐driven attention network (FDAN) to exploit motion information for video salient object detection. FDAN consists of an appearance feature extractor, a motion‐guided attention module and a saliency map regression module. It extracts the appearance feature per frame, refines appearance feature with optical flow and infers the ultimate saliency map, respectively. Motion‐guided attention module is the core of FDAN, which extracts motion information in the form of attention. This attention mechanism is a two‐branch CNN, fusing optical flow and appearance features. In addition, a shortcut connection is applied to the attention multiplied feature map for noise suppression intensively. Experimental results show that the proposed method can achieve performance on par with the state‐of‐the‐art method flow‐guided recurrent neural encoder on challenging benchmarks of Densely Annotated Video Segmentation and Freiburg–Berkeley Motion Segmentation while being two times faster in detection.
Feng Zhou 0006, Hui Shuai, Qingshan Liu 0001, Guodong Guo
IET Image Process.1
2019 Hyperspectral image classification using spectral-spatial LSTMs
Feng Zhou 0006, Renlong Hang, Qingshan Liu 0001, Xiao-Tong Yuan
Neurocomputing1
2018 Integrating Convolutional Neural Network and Gated Recurrent Unit for Hyperspectral Image Spectral-Spatial Classification
Feng Zhou 0006, Renlong Hang, Qingshan Liu 0001, Xiao-Tong Yuan
PRCV (4)1