EDBT 2026 Demo / reviewers in the wild / expert
Peng Zhu 0004
dblp:63/5448-4
· DBLP profile ↗
13ranked-venue papers
2as first author
11since 2021 · last 2025
0000-0001-9616-5149ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Source-Free Cross-Domain Scene Classification of Remote Sensing Images via Statistics Matching and Noise AdaptationabstractIn recent years, in order to alleviate the performance degradation problem caused by domain drift in scene classification tasks, some unsupervised domain adaptation methods have been introduced into the field of remote sensing images. Such methods require simultaneous access to both source and target domain data during training. However, the large storage and transmission costs of remote sensing images limit the further application of these methods. To address this challenge, we investigate the task of source-free cross-domain scene classification for remote sensing images. In the model adaptation process, only the target domain dataset and the trained source domain model are used. Our approach consists of two parts: a distribution alignment strategy based on source domain model statistics matching and a noise adaptation strategy. In order to fully utilize the knowledge of the pre-trained source domain model, we fix the classifier to get the feature distribution of the source domain, so that the target domain feature distribution is close to the feature distribution of the source domain. The noise adaptation layer is inserted after the classifier in order to improve the robustness of the model to noise-containing pseudo-labels, and the sample-wise noise transfer matrix is learned. Experimental results on 12 transfer tasks on the cross-scene dataset, and 2 transfer tasks on the cross-sensor dataset, to validate the effectiveness of our approach. Compared to traditional unsupervised domain adaptive methods, our method is able to achieve better performance under the condition of not accessing the source domain data. Peng Zhu 0004, Xiangrong Zhang, Xiao Han 0012, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | ESDINet: Efficient Shallow-Deep Interaction Network for Semantic Segmentation of High-Resolution Aerial ImagesabstractSemantic segmentation of high-resolution remote sensing images is essential in many fields. Nevertheless, in practical applications, constrained by limited computational resources and complex network structures, many advanced models on semantic segmentation often fail to show efficient performance, prompting research on lightweight models. For lightweight semantic segmentation models, the two-branch architecture has been shown to work well in speed and performance. However, such two-branch architectures usually do not utilize enough information for shallow structures to efficiently provide richer multiscale information for the two branches. The lightweight modules it uses are difficult to extract the global context information of the features effectively. Compared with the current advanced semantic segmentation models, lightweight models still have some differences in performance. In order to solve these problems, we propose a new lightweight dual-branch architecture efficient shallow-deep interaction network (ESDINet), which can quickly extract low-level spatial and high-level semantic information of images through the detail branch and semantic branch. Specifically, we have constructed an efficient double-branch structure with shallow and deep different interactions to achieve multiscale information interaction. At the same time, we optimize the semantic branch and propose a new linear attention block to effectively improve the global perception of the semantic branch. We performed extensive experiments and the results show that our model achieves a good balance between segmentation accuracy and inference speed. In particular, ESDINet achieves 82.03% mean intersection over union (mIoU) on the Vaihingen test set, while the proposed model achieves an inference speed of 116 frames/s (FPS) for$512\times512$inputs on a single NVIDIA GTX 2080Ti GPU. Xiangrong Zhang, Zhenhang Weng, Peng Zhu 0004, Xiao Han 0012, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | High-Resolution Remote Sensing Image Segmentation With Global-Guided Normalization and Local Affinity DistillationabstractIn recent years, high-resolution (HR) remote sensing images (RSIs) segmentation has received growing attention. The huge number of pixels poses a challenge to the semantic segmentation algorithm, which is limited by the storage of GPUs, so the current methods for processing HR RSIs are categorized into two main categories, i.e., global methods and local methods. The former downsamples the original image and loses a lot of feature details. The latter crops the original image and fails to obtain global contextual information. Both types of methods lead to limited segmentation accuracy. In this article, we propose an end-to-end framework, called global injection network (GINet), which explores two levels of feature distribution and feature relationship to achieve tradeoff between global context and local details. In concrete terms, we propose the global-guided normalization (GGN) module, which injects global context information into local branch and modulates local features using global features to enhance the global perception of local branch. In addition, to constrain the spatial consistency of two branches, inspired by the knowledge distillation technique, we propose local affinity distillation (LAD) loss, which distills the relations in local features into global features to keep the similarity of the relationships corresponding to patches in the two branches. The comprehensive experimental results on three large-scale land-cover classification datasets, DeepGlobe ($2448 \times 2448$), Inria Aerial ($5000 \times 5000$), and GID-15 ($7200 \times 6800$), confirm the effectiveness and superiority of our method in HR semantic segmentation tasks. Peng Zhu 0004, Xiangrong Zhang, Xiao Han 0012, Puhua Chen, Xu Tang 0004, Xina Cheng, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Global-Local Representation Coupling Network for Remote Sensing Image Change DetectionabstractChange detection is one of the important tasks in remote sensing image processing, and the powerful feature extraction ability of convolutional networks has achieved some success in change detection. However, the problem of the limited field size of pure convolutional networks makes the change detection accuracy of high-resolution remote sensing images limited. The introduction of transformers can link the concept of long-distance in space and time. Therefore, in order to maximize the respective advantages of transformers and CNNs, we propose a new parallel architecture. To accomplish the above goal, we propose a new network, which consists of a local detail branch and transformer global spatial-temporal feature branch and a feature fusion module. The experimental results reach the current sota level. Fanghan Yang, Xiangrong Zhang, Peng Zhu 0004, Zhenhang Weng, Puhua Chen |
IGARSS | 3 |
| 2023 | High-Quality Angle Prediction for Oriented Object Detection in Remote Sensing ImagesabstractOriented object detection is a challenging task in remote sensing, where the detected objects can be represented by oriented bounding boxes (OBBs). Angle prediction in oriented object detection has been widely studied, due to its crucial role in object detection. However, the precision of angle prediction is severely limited by misalignments in most of the existing methods, including representation-, evaluation-, and optimization-based misalignments. To alleviate these misalignments, this paper presents a novel angle prediction method, called Angle Quality Estimation (AQE). Specifically, our proposed AQE transforms the angle prediction task into a distribution estimation task to address the representation misalignment problem and implicitly measure the quality of the predicted angles. Based on the estimated angle quality, we then propose a new metric to comprehensively evaluate the quality of OBBs. Then we propose an object aspect ratio based loss function to optimize angle prediction for addressing the optimization misalignment. Our proposed AQE is a plug-and-play method, which can be embedded on any existing oriented object detector. Experimental results on three public benchmarks, including DOTA, HRSC2016, and ICDAR2015 datasets, show that our method achieves better performance than the other state-of-the-art. Guanchun Wang, Xiangrong Zhang, Peng Zhu 0004, Xu Tang 0004, Puhua Chen, Licheng Jiao, Huiyu Zhou 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Semantic-Aware Context Modeling for Road Extraction in Remote Sensing ImagesabstractRoad extraction faces the great challenges of occlusion, large span, and complex backgrounds in remote sensing images. Many existing methods receive context from regions near the road non-differently, and the context from irrelevant regions instead harms the semantics of features and leads to the mis-classification of the network. To address the above problem, we propose a Semantic-Aware Context Module (SACM) that encourages the network to model the context of different se-mantics supervised by a soft foreground map. And Strip Pooling Module (SPM) is introduced to match the fact that roads tend to be strip-shaped, contributing to the suppression of contamination information in irrelevant regions. Both SACM and SPM enable the network to obtain more specific semanti-cally relevant context. The experimental results on the Deep-Globe dataset show that the proposed method tremendously improves the performance of the network. Xiangrong Zhang, Xiaoqian Zhu, Peng Zhu 0004, Xu Tang 0004, Licheng Jiao |
IGARSS | 4 |
| 2022 | Semantic Attention and Scale Complementary Network for Instance Segmentation in Remote Sensing ImagesabstractIn this article, we focus on the challenging multicategory instance segmentation problem in remote sensing images (RSIs), which aims at predicting the categories of all instances and localizing them with pixel-level masks. Although many landmark frameworks have demonstrated promising performance in instance segmentation, the complexity in the background and scale variability instances still remain challenging, for instance, segmentation of RSIs. To address the above problems, we propose an end-to-end multicategory instance segmentation model, namely, the semantic attention (SEA) and scale complementary network, which mainly consists of a SEA module and a scale complementary mask branch (SCMB). The SEA module contains a simple fully convolutional semantic segmentation branch with extra supervision to strengthen the activation of interest instances on the feature map and reduce the background noise's interference. To handle the undersegmentation of geospatial instances with large varying scales, we design the SCMB that extends the original single mask branch to trident mask branches and introduces complementary mask supervision at different scales to sufficiently leverage the multiscale information. We conduct comprehensive experiments to evaluate the effectiveness of our proposed method on the iSAID dataset and the NWPU Instance Segmentation dataset and achieve promising performance. Tianyang Zhang 0002, Xiangrong Zhang, Peng Zhu 0004, Xu Tang 0004, Chen Li 0011, Licheng Jiao, Huiyu Zhou 0001 |
IEEE Trans. Cybern. | 3 |
| 2022 | Spatial Pooling Graph Convolutional Network for Hyperspectral Image ClassificationabstractGraph convolution networks (GCNs) have been applied in a variety of fields due to their powerful ability in processing graph-like data. However, the massive number of hyperspectral pixels makes it challenging to define general graph structures on hyperspectral images (HSIs). On the other hand, convolutional neural networks (CNNs) take in regular image regions with fixed square size, and have demonstrated impressive accuracy while being efficient in computation. Inspired by the classification framework of CNNs, we develop a GCN-based model that generates effective local spectral–spatial features for HSI classification. Specifically, graph convolutions are performed separately on every local region, which significantly limits the graph’s size. While graph convolution extracts features of every pixel, it does not reduce the number of them. To fuse suitable representations for the classification task, we develop a graph pooling operation to preserve classification-specific features and reduce redundant pixels. Based on local regions of HSIs, pooling in the graph domain is equivalent to spatial pooling in the spatial domain. The proposed method is thus named the spatial pooling graph convolutional network (SPGCN). Experimental results on several typical datasets demonstrated that the proposed SPGCN provides competitive results compared with other state-of-the-art CNN-based methods. Xiangrong Zhang, Peng Zhu 0004, Xu Tang 0004, Jie Feng 0003, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Foreground Refinement Network for Rotated Object Detection in Remote Sensing ImagesabstractObject detection has been a fundamental task in the field of remote sensing and has made considerable progress in recent years. However, the high background complexity in remote sensing images (RSIs) remains challenging. In this article, we propose a refined rotation detector, namely, the Foreground Refinement Network (FoRDet), to alleviate the above problem by leveraging the information of foreground regions from the perspectives of feature and optimization. Specifically, we propose a foreground relation module (FRL) that aggregates the foreground-contextual representations from the coarse stage and improves the discrimination of foreground regions on feature maps in the refined stage. Besides, considering the risk of the potential foreground anchors being overwhelmed in the training phase, we design a foreground anchor reweighting (FRW) loss that integrates the classification confidence and localization accuracy of each foreground anchor from the coarse stage to dynamically regulate their contributions in the refined stage, which highlights the potential foreground anchors. The comprehensive experimental results on three public datasets for rotated object detection DOTA, HRSC2016, and UCAS-AOD demonstrate the effectiveness of our proposed method. Tianyang Zhang 0002, Xiangrong Zhang, Peng Zhu 0004, Puhua Chen, Xu Tang 0004, Chen Li 0011, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Adaptive Affinity Loss and Erroneous Pseudo-Label Refinement for Weakly Supervised Semantic SegmentationabstractSemantic segmentation has been continuously investigated in the last ten years, and majority of the established technologies are based on supervised models. In recent years, image-level weakly supervised semantic segmentation (WSSS), including single- and multi-stage process, has attracted large attention due to data labeling efficiency. In this paper, we propose to embed affinity learning of multi-stage approaches in a single-stage model. To be specific, we introduce an adaptive affinity loss to thoroughly learn the local pairwise affinity. As such, a deep neural network is used to deliver comprehensive semantic information in the training phase, whilst improving the performance of the final prediction module. On the other hand, considering the existence of errors in the pseudo labels, we propose a novel label reassign loss to mitigate over-fitting. Extensive experiments are conducted on the PASCAL VOC 2012 dataset to evaluate the effectiveness of our proposed approach that outperforms other standard single-stage methods and achieves comparable performance against several multi-stage methods. Xiangrong Zhang, Zelin Peng, Peng Zhu 0004, Tianyang Zhang 0002, Chen Li 0011, Huiyu Zhou 0001, Licheng Jiao |
ACM Multimedia | 3 |
| 2021 | GRS-Det: An Anchor-Free Rotation Ship Detector Based on Gaussian-Mask in Remote Sensing ImagesabstractShip detection is a significant and challenging task in remote sensing. Due to the arbitrary-oriented property and large aspect ratio of ships, most of the existing detectors adopt rotation boxes to represent ships. However, manual-designed rotation anchors are needed in these detectors, which causes multiplied computational cost and inaccurate box regression. To address the abovementioned problems, an anchor-free rotation ship detector, named GRS-Det, is proposed, which mainly consists of a feature extraction network with selective concatenation module (SCM), a rotation Gaussian-Mask model, and a fully convolutional network-based detection module. First, a U-shape network with SCM is used to extract multiscale feature maps. With the help of SCM, the channel unbalance problem between different-level features in feature fusion is solved. Then, a rotation Gaussian-Mask is designed to model the ship based on its geometry characteristics, which aims at solving the mislabeling problem of rotation bounding boxes. Meanwhile, the Gaussian-Mask leverages context information to strengthen the perception of ships. Finally, multiscale feature maps are fed to the detection module for classification and regression of each pixel. Our proposed method, evaluated on ship detection benchmarks, including HRSC2016 and DOTA Ship data sets, achieves state-of-the-art results. Xiangrong Zhang, Guanchun Wang, Peng Zhu 0004, Tianyang Zhang 0002, Chen Li 0011, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | Adaptive Feature Aggregation Network for Object Detection in Remote Sensing ImagesabstractObject detection in remote sensing images is a challenging task because of the large scale variations across the geospatial objects. The feature pyramid network (FPN) is widely used to alleviate the scale variations problem, however, it only fuses the features from adjacent levels and lacks the information of the entire feature hierarchy. In this paper, we propose a novel and effective feature pyramid aggregation network, called Adaptive Feature Aggregation Network (AFANet). Specifically, we propose the Adaptive Feature Aggregation (AFA) module to adaptively aggregate multi-level features of FPN and introduce the Bottom-up Path to enhance the location information of the entire feature levels. In addition, we use the Receptive Field Block (RFB) module to capture different receptive field features for each level feature map. We evaluate the effectiveness of our AFANet on the DOTA dataset and achieves noticeable performance compared with the baseline. Wenliang Sun, Xiangrong Zhang, Tianyang Zhang 0002, Peng Zhu 0004, Xu Tang 0004 |
IGARSS | 4 |
| 2020 | Discriminative Feature Pyramid Network For Object Detection In Remote Sensing ImagesabstractMulti-class geospatial object detection in remote sensing images suffer great challenges, such as large scales variability and complex background. Although feature pyramid network (FPN) can alleviate the problem of scale variation to some extent, it causes the loss of spatial and semantic information which is not conducive to object location. To address the above problem, this paper proposes a discriminative feature pyramid network (DFPN) by introducing a global guidance module (GGM) and a feature aggregation module (FAM). Specifically, the global guidance module delivers the high-level semantic information to lower layers, so as to obtain feature maps with stronger semantic information to eliminate the interference caused by complex background. The feature aggregation module enhances the interflow of information between different layers and better captures the discrimination information at each layer. We validate the effectiveness of our method on the NWPU VHR-10 and RSOD datasets, the results outperform baseline by 2.06 and 3.88 points respectively. Xiaoqian Zhu, Xiangrong Zhang, Tianyang Zhang 0002, Peng Zhu 0004, Xu Tang 0004, Chen Li 0011 |
IJCNN | 4 |