Shunyi Zheng

dblp:01/1405 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
7since 2021 · last 2025
0000-0001-5594-3493ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2025 AMMUNet: Multiscale Attention Map Merging for Remote Sensing Image Segmentation
abstract
The advancement of deep learning has driven notable progress in remote sensing semantic segmentation. Multihead self-attention (MSA) mechanisms have been widely adopted in semantic segmentation tasks. Network architectures exemplified by Vision Transformers have implemented window-based operations in the spatial domain to reduce computational costs. However, this approach comes at the expense of a weakened capacity to capture long-range dependencies, potentially limiting their efficacy in remote sensing image processing. In this letter, we propose AMMUNet, a UNet-based framework that employs multiscale attention map (AM) merging, comprising two key innovations: the attention map merging mechanism (AMMM) module and the granular multihead self-attention (GMSA). AMMM effectively combines multiscale AMs into a unified representation using a fixed mask template, enabling the modeling of a global attention mechanism. By integrating precomputed AMs in preceding layers, AMMM reduces computational costs while preserving global correlations. The proposed GMSA efficiently acquires global information while substantially mitigating computational costs in contrast to the global MSA mechanism. This is accomplished through the strategic alignment of granularity and the reduction of relative position bias parameters, thereby optimizing computational efficiency. Experimental evaluations highlight the superior performance of our approach, achieving remarkable mean intersection over union (mIoU) scores of 75.48% on the challenging Vaihingen dataset and an exceptional 77.90% on the Potsdam dataset, demonstrating the superiority of our method in precise remote sensing semantic segmentation. Codes are available athttps://github.com/interpretty/AMMUNet.
Shunyi Zheng, Xiqi Wang, Zhao Liu 0004
IEEE Geosci. Remote. Sens. Lett.2
2024 Few-Shot Semantic Segmentation via Mask Aggregation
abstract
Abstract Few-shot semantic segmentation aims to recognize novel classes with only very few labelled data. This challenging task requires mining of the correlation between the query image and the support images. Previous works have typically regarded it as a pixel-wise classification problem. Therefore, various models have been designed to explore the correlation of pixels between the query image and the support images. However, they focus only on pixel-wise correspondence and ignore the overall correlation of objects. In this paper, we introduce a mask-based classification method for addressing this problem. The mask aggregation network, which is a simple mask classification model, is proposed to simultaneously generate a fixed number of masks and their probabilities of being targets. Then, the final segmentation result is obtained by aggregating all the masks according to their locations. Experiments on both the PASCAL- $$5^i$$ 5 i and COCO- $$20^i$$ 20 i datasets show that our method performs comparably to the state-of-the-art pixel-based methods. This competitive performance demonstrates the potential of mask classification as an alternative baseline method for few-shot semantic segmentation.
Shunyi Zheng
Neural Process. Lett.2
2023 Efficient large-scale oblique image matching based on cascade hashing and match data scheduling
Shunyi Zheng, Ce Zhang 0005, Xiqi Wang, Rui Li 0036
Pattern Recognit.2
2023 Few-Shot Aerial Image Semantic Segmentation Leveraging Pyramid Correlation Fusion
abstract
Few-shot semantic segmentation has gained significant attention due to its ability to segment novel objects using only a limited number of labeled samples, thereby addressing the problem of overfitting caused by a lack of training data. Although this technique is widely studied in the field of computer vision, there are few methods for remote sensing images. Prevalent few-shot semantic segmentation methods can achieve remarkable results for natural images, but they are difficult to apply to remote sensing image processing because existing methods rarely take into consideration the large scale and resolution differences in remote sensing images. Consequently, it is hard for them to obtain correct semantic guidance from few annotated remote sensing images. To tackle these problems, this article proposes the pyramid correlation fusion network (PCFNet) to promote the ability to mine helpful information by calculating multi-scale pixel-wise semantic correspondence. Particularly, the dual distance correlation (DDC) module is designed to simultaneously compute the cosine similarity and Euclidean distance between query features and support features, producing adequate guidance information to determine the category of each pixel. Moreover, to improve segmentation accuracy for small objects, the scale-aware cross-entropy loss (SACELoss) is introduced to dynamically assign loss weights according to the actual sizes of objects. This enables smaller objects to be assigned larger weight values and thus receive more attention during training. Comprehensive experiments on both the iSAID-5iand DLRSD-5idatasets demonstrate that our method outperforms state-of-the-art few-shot semantic segmentation methods. Our code is available at https://github.com/TinyAway/PCFNet.
Shunyi Zheng, Zhi Gao 0005
IEEE Trans. Geosci. Remote. Sens.2
2022 MACU-Net for Semantic Segmentation of Fine-Resolution Remotely Sensed Images
abstract
Semantic segmentation of remotely sensed images plays an important role in land resource management, yield estimation, and economic assessment. U-Net, a deep encoder–decoder architecture, has been used frequently for image segmentation with high accuracy. In this letter, we incorporate multiscale features generated by different layers of U-Net and design a multiscale skip connected and asymmetric-convolution-based U-Net (MACU-Net), for segmentation using fine-resolution remotely sensed images. Our design has the following advantages: (1) the multiscale skip connections combine and realign semantic features contained in both low-level and high-level feature maps; (2) the asymmetric convolution block strengthens the feature representation and feature extraction capability of a standard convolution layer. Experiments conducted on two remotely sensed data sets captured by different satellite sensors demonstrate that the proposed MACU-Net transcends the U-Net, U-Netpyramid pooling layers (PPL), U-Net 3+, among other benchmark approaches. Code is available athttps://github.com/lironui/MACU-Net.
Rui Li 0036, Chenxi Duan, Shunyi Zheng, Ce Zhang 0005, Peter M. Atkinson
IEEE Geosci. Remote. Sens. Lett.3
2022 Multistage Attention ResU-Net for Semantic Segmentation of Fine-Resolution Remote Sensing Images
abstract
The attention mechanism can refine the extracted feature maps and boost the classification performance of the deep network, which has become an essential technique in computer vision and natural language processing. However, the memory and computational costs of the dot-product attention mechanism increase quadratically with the spatiotemporal size of the input. Such growth hinders the usage of attention mechanisms considerably in application scenarios with large-scale inputs. In this letter, we propose a linear attention mechanism (LAM) to address this issue, which is approximately equivalent to dot-product attention with computational efficiency. Such a design makes the incorporation between attention mechanisms and deep networks much more flexible and versatile. Based on the proposed LAM, we refactor the skip connections in the raw U-Net and design a multistage attention ResU-Net (MAResU-Net) for semantic segmentation from fine-resolution remote sensing images. Experiments conducted on the Vaihingen data set demonstrated the effectiveness and efficiency of our MAResU-Net. Our code is available athttps://github.com/lironui/MAResU-Net.
Rui Li 0036, Shunyi Zheng, Chenxi Duan, Jianlin Su, Ce Zhang 0005
IEEE Geosci. Remote. Sens. Lett.2
2022 Multiattention Network for Semantic Segmentation of Fine-Resolution Remote Sensing Images
abstract
Semantic segmentation of remote sensing images plays an important role in a wide range of applications, including land resource management, biosphere monitoring, and urban planning. Although the accuracy of semantic segmentation in remote sensing images has been increased significantly by deep convolutional neural networks, several limitations exist in standard models. First, for encoder–decoder architectures such as U-Net, the utilization of multiscale features causes the underuse of information, where low-level features and high-level features are concatenated directly without any refinement. Second, long-range dependencies of feature maps are insufficiently explored, resulting in suboptimal feature representations associated with each semantic class. Third, even though the dot-product attention mechanism has been introduced and utilized in semantic segmentation to model long-range dependencies, the large time and space demands of attention impede the actual usage of attention in application scenarios with large-scale input. This article proposed a multiattention network (MANet) to address these issues by extracting contextual dependencies through multiple efficient attention modules. A novel attention mechanism of kernel attention with linear complexity is proposed to alleviate the large computational demand in attention. Based on kernel attention and channel attention, we integrate local feature maps extracted by ResNet-50 with their corresponding global dependencies and reweight interdependent channel maps adaptively. Numerical experiments on two large-scale fine-resolution remote sensing datasets demonstrate the superior performance of the proposed MANet. Code is available athttps://github.com/lironui/Multi-Attention-Network.
Rui Li 0036, Shunyi Zheng, Ce Zhang 0005, Chenxi Duan, Jianlin Su, Peter M. Atkinson
IEEE Trans. Geosci. Remote. Sens.2
2019 An image thresholding approach based on Gaussian mixture model
Like Zhao, Shunyi Zheng, Haitao Wei
Pattern Anal. Appl.2
2007 Generation of 3D Surface Model of Complex Objects Based on Non-Metric Camera
abstract
A reconstruction method for complex objects was developed using non-metric images. Firstly, a rotatable platform was designed to capture images and a planar grid board on its top was used for camera calibration. With the help of the platform, images were captured and then processed with a multi-baseline image matching algorithm to produce corresponding point features. The initial external elements of the images can be obtained with traditional photogrammetry approach. These parameters and coordinates of corresponding points were input to block bundle adjustment as initial values, and then point clouds and accurate image parameters were obtained. In the end, a visibility constrained Delaunay triangulation algorithm was used to produce surface models. The methods include artful hardware design, efficient image matching, and rational 3D surface reconstruction. It is low-cost and automatic, and experiment results demonstrate its feasibility and effectiveness.
Shunyi Zheng, Ruifang Zhai, Zuxun Zhang
ICIP (3)1