Zhuoyi Zhao

dblp:185/7864 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2026 A dual-stream foreground-aware enhancement network with spiralscan-Mamba for vision-based occupancy prediction in autonomous driving
Nannan Liu, Yanyin Guo, Chuiyi Deng, Zhuoyi Zhao, Junwei Li 0009
Eng. Appl. Artif. Intell.5
2025 Optimizing Age of Information without Knowing the Age of Information
Zhuoyi Zhao, Igor Kadota
INFOCOM1
2025 Optimizing Age of Information in Networks with Large and Small Updates
abstract
Modern sensing and monitoring applications typically consist of sources transmitting updates of different sizes, ranging from a few bytes (position, temperature, etc.) to multiple megabytes (images, video frames, LIDAR point scans, etc.). Existing approaches to wireless scheduling for information freshness typically ignore this mix of large and small updates, leading to suboptimal performance. In this paper, we consider a single-hop wireless broadcast network with sources transmitting updates of different sizes to a base station over unreliable links. Some sources send large updates spanning many time slots while others send small updates spanning only a few time slots. Due to medium access constraints, only one source can transmit to the base station at any given time, thus requiring careful design of scheduling policies that takes the sizes of updates into account. First, we derive a lower bound on the achievable Age of Information (AoI) by any transmission scheduling policy. Second, we develop optimal randomized policies that consider both switching and no-switching during the transmission of large updates. Third, we introduce a novel Lyapunov function and associated analysis to propose an AoI-based Max-Weight policy that has provable constant factor optimality guarantees. Finally, we evaluate and compare the performance of our proposed scheduling policies through simulations, which show that our Max-Weight policy achieves near-optimal AoI performance.
Zhuoyi Zhao, Vishrant Tripathi, Igor Kadota
WiOpt1
2025 GSFANet: Global Spatial-Frequency Attention Network for Infrared Small Target Detection
abstract
Infrared small target detection (IRSTD) has progressed significantly in spatial-domain learning. However, single-frame images’ limited spatial semantics impair discrimination between targets and similar noise while complicating integrity detection of large-scale target. To address these, we propose Global Spatial-Frequency Attention Network (GSFANet), which enhances the distribution difference between targets and noise from a frequency-domain perspective while preserving spatial information integrity. The core innovations consist of three modules: 1) Parametric Wavelet Downsampling (PWD), preserving small target details during frequency refinement to prevent feature fragmentation; 2) Hierarchical Gated Kernel Attention (HGKA), capturing cross-level frequency relationships through Cross-channel Kernel Attention (C2K) and maintaining spatial coherence via Cross-spatial Gate Attention (CSG), effectively bridging semantic gaps across layers; 3) Adaptive Frequency-Decoupled Fusion (AdaFD), dynamically fusing target-associated frequency components while suppressing noise. We further develop AdaFL Loss to balance multi-scale target gradients and stabilize training. Experiments on three benchmark datasets demonstrate GSFANet’s superior detection performance and enhanced segmentation robustness in complex scenarios compared to state-of-the-art methods. Our code will be made public at https://github.com/dengfa02/GSFANet_IRSTD.
Chuiyi Deng, Zhuoyi Zhao, Xiang Xu 0002, Yixin Xia, Junwei Li 0009, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.2
2024 Hyperspectral Image Classification Using Groupwise Separable Convolutional Vision Transformer Network
abstract
Recently, Vision Transformer (ViT)-based deep learning models have achieved remarkable performance gains in hyperspectral image classification (HSIC) due to their abilities to model long-range dependencies and extract global spatial features. However, ViT is built with a stack of Transformer blocks and faces the challenge of learning a large number of parameters when processing hyperspectral data. Besides, the inherent modeling of global correlation in Transformer ignores the effective representation of local spatial and spectral features. To address these issues, we propose a lightweight ViT network known as Groupwise Separable Convolutional Vision Transformer (GSC-ViT). Firstly, a Groupwise Separable Convolution (GSC) module, which is a combination of grouped pointwise convolution and group convolution, is designed to significantly decrease the number of convolutional kernel parameters, and effectively capture local spectral-spatial information in hyperspectral image. Secondly, a Groupwise Separable Multi-Head Self-Attention (GSSA) module is employed to substitute the conventional Multi-Head Self-Attention (MSA) in ViT, in which the Groupwise Self-Attention(GSA) provides local spatial feature extraction, and the Pointwise Self-Attention(PWSA) provides global spatial feature extraction. Thirdly, a simple pointwise layer with enhanced skip connection mechanism is employed to substitute the Multi-Layer Perceptron (MLP) layer in all Transformer blocks of ViT, so as to eliminate unnecessary nonlinear transformations and facilitate the fusion of features derived from GSC and GSSA modules. Extensive experiments on four benchmark hyperspectral datasets reveal that our GSC-ViT can achieve surprising classification performance with relatively few training samples as compared with some existing HSIC approaches. The source code is available at https://github.com/flyzzie/TGRS-GSC-VIT.
Zhuoyi Zhao, Xiang Xu 0002, Shutao Li 0001, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.1
2023 Decadal Evolution of Retrogressive Thaw Slumps Retrieved from Landsat Imagery Via Heatmap Regression: A Case Study of the Beiluhe Region in Central Tibet
abstract
Retrogressive thaw slumps (RTSs) have become more common in the Circum-Arctic and the Qinghai-Tibet Plateau. Due to the longevity of the RTSs, it is necessary to conduct observation for more than ten years. With decades of continuous service, Landsat can potentially be used to serve the purpose, but the low resolution of its images can be challenging for delineating such a relatively small landform. Thus, designing an algorithm that can capture the development of RTSs from Landsat images is key to understanding the long-term spatial-temporal evolution of RTSs over a large extent. In this paper, we propose a heatmap regression-based deep learning method to locate and estimate the area of RTSs. Our method is applied to the Beiluhe study area using Landsat 8 images from 2013 to 2022 and the results reveal the abrupt increase of RTSs after 2016 and the estimated areas match well with the ground truth data.
Zhuoyi Zhao, Zhuoxuan Xia, Lin Liu 0010
IGARSS1
2023 Text-based person search via local-relational-global fine grained alignment
Junfeng Zhou, Baigang Huang, Wenjiao Fan, Ziqian Cheng, Zhuoyi Zhao
Knowl. Based Syst.5
2023 Gabor-Modulated Grouped Separable Convolutional Network for Hyperspectral Image Classification
abstract
Nowadays, convolutional neural network (CNN)-based deep learning models have been popularized in hyperspectral image classification (HSIC) and achieved significant accuracy gains, which is due to their hierarchical and nonlinear feature learning patterns. However, too deeper network structures may induce a huge amount of parameters and excessive computing overhead, leading to the need for plenty of labeled samples for training. Besides, highly abstract semantic features may not be the most suitable for hyperspectral land-cover classification tasks. To address these issues, we propose a fairly lightweight network model for HSIC, which is built on a type of exquisitely designed convolution module, namelygrouped separable convolution. Compared with the standard convolution, the designed grouped separable convolution module combines grouped convolution with point-wise convolution, which not only greatly reduces the number of parameters of convolution kernels, but also caters to the inherent 3D cube style of hyperspectral image data. Moreover, Gabor filters are introduced to modulate the grouped separable convolution kernels, so as to further use relatively few convolution kernels with additional prior orientation and scale information for feature extraction. The experiments are carried out on four real hyperspectral datasets, and the experimental results reveal that the proposed model has low training cost and memory overhead. Compared with some existing deep network models that have been applied to HSIC, our proposed model can achieve competitive classification accuracy with fewer training samples.
Zhuoyi Zhao, Xiang Xu 0002, Jun Li 0009, Shutao Li 0001, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.1
2021 Automatic Detection of Widely Distributed Local-Scale Subsidence Bowls in Rapidly Urbanizing Metropolitan Region Using Time-Series InSAR and Deep Learning Methods
abstract
Multi-temporal interferometric synthetic aperture radar (MT-InSAR) has been used to produce deformation velocity map for investigating the surface subsidence in the rapidly urbanizing metropolitan regions. However, simple analysis techniques like thresholding cannot detect and locate the widely distributed local-scale subsidence reliably. In this study, we propose a deep-learning based method to automatically detect the local-scale subsidence bowls in the deformation velocity map. To test our method, we choose the Guangdong-Hong Kong-Macao Greater Bay Area (GBA) as the study region, where widespread local-scale subsidence bowls exist associated with the urbanization. Using deformation velocity maps spanning 2015–2017 derived from MT-InSAR, our method detects several subsidence bowls due to dewatering, excavation of foundation pits and subways, and other engineering works. The results demonstrate the potential applicability of the proposed method to automatically detect and analyze the local-scale subsidence bowls in the built-up regions.
Zherong Wu, Zhuoyi Zhao, Yi Zheng 0012, Peifeng Ma
IGARSS2
2016 Crossing-Line Crowd Counting with Two-Phase Deep Neural Networks
Zhuoyi Zhao, Hongsheng Li 0001, Rui Zhao 0001, Xiaogang Wang 0001
ECCV (8)1