Guangzhen Yao

dblp:383/0560 · DBLP profile ↗
← Back
22ranked-venue papers
3as first author
22since 2021 · last 2027
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 2 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2027 Multi-scale asymmetric graph contrastive anomaly detection
Wenxin Zhang 0005, Xi Xuan, Guangzhen Yao, Renda Han, Xiangxiang Lang, Feng Zhou 0011, Cuicui Luo
Inf. Process. Manag.3
2026 V-Pruner: A Fast and Globally-informed Token Pruning Framework for Vision Transformer
abstract
Vision Transformer (ViT) has become one of the cornerstones of the computer vision field, demonstrating exceptional performance. However, its inherent high computational complexity and inference latency still pose significant obstacles for deployment in resource-constrained environments. Token pruning, by removing less informative tokens, offers an effective strategy to reduce computational overhead. However, existing pruning methods largely rely on static or local token importance scores. This myopic approach fundamentally overlooks the sequential dependency of pruning decisions and fails to capture the interaction effects between pruning decisions across layers, often neglecting the global interactions between mask variables. To address this limitation, we propose V-Pruner, a fast and globally-informed token pruning framework for Vision Transformer. V-Pruner first leverages Fisher information to perform an initial assessment of token importance, providing a principled initial prior for pruning decisions. Building on this, V-Pruner introduces a Reinforcement Learning (RL) Proximal Policy Optimization (PPO) algorithm, refining token pruning into a global sequential decision process. The algorithm combines a composite reward signal that incorporates both model performance and computational cost to guide policy exploration, effectively evaluating the long-term impact of different pruning decision combinations on global model performance. Extensive experiments on ViT-L, DeiT-B, DeiT-S, and DeiT-T demonstrate that V-Pruner achieves a better balance between accuracy, GFLOPs, inference speed, and training time, surpassing existing mainstream ViT pruning algorithms in overall performance.
Guangzhen Yao, Jiayun Zheng, Zezhou Wang, Wenxin Zhang 0005, Renda Han, Chuangxin Zhao, Zeyu Zhang 0006
AAAI1
2026 A Graph Foundation Model for Unified Anomaly Detection
Renda Han, Xiaobao Wang, Luzhi Wang, Wenxin Zhang 0005, Guangzhen Yao, Hongxiang Liang
WWW5
2026 Federated graph-level clustering network with adaptive knowledge compensation
Renda Han, Guangzhen Yao, Wenxin Zhang 0005, Ronghao Fu, Zeyu Zhang 0006
Neural Networks3
2026 Attribute-incomplete graph anomaly detection network
Renda Han, Xiaobao Wang, Guangzhen Yao, Wenxin Zhang 0005, Ronghao Fu, Dayu Hu, Zeyu Zhang 0006, Kaiming Wang
Pattern Recognit.5
2025 Unlocking the Full Potential of Separable Convolutions on Tensor Cores
Aodie Cui, Chuangxin Zhao, Gaozhe Jiang, Guangzhen Yao, Renda Han, Wenxin Zhang 0005, Xi Xuan
ICIC (16)6
2025 SecureNT: Smart Topology Obfuscation for Privacy-Aware Network Monitoring
Chengze Du 0001, Jibin Shi, Guangzhen Yao
ICIC (15)4
2025 FreCT: Frequency-Augmented Convolutional Transformer for Robust Time Series Anomaly Detection
Wenxin Zhang 0005, Guangzhen Yao, Xiaojian Lin, Renxiang Guan, Chengze Du 0001, Renda Han, Xi Xuan, Cuicui Luo
ICIC (16)3
2025 GuidedLatent: Defending VAEs against Membership Inference Attacks via Distribution-Guided Privacy
abstract
Variational autoencoders (VAEs) have been deployed in many privacy-sensitive domains, and their vulnerability to membership inference attacks (MIAs) poses giant privacy risks. While some existing privacy protection methods like differential privacy often compromise generative models’ utility, we present GuidedLatent, a novel mechanism that enhances membership privacy and preserves their generative performance. GuidedLatent allows the model to adjust latent representations dynamically based on distribution similarities, coupled with a two-phase training strategy that gradually incorporates privacy constraints. We also establish bounds on the privacy-utility trade-off theoretically and prove our mechanism reduces the performance of membership inference attacks compared to other baseline approaches. Extensive experiments demonstrate that our method maintains high-quality generation capabilities while minimizing degradation in quality metrics. Our method performs effectively across various VAE variants and architectures, providing a practical solution for privacy-preserving generative models.1
Chengze Du 0001, Guangzhen Yao, Jibin Shi, Renda Han
IJCNN2
2025 Dual Boost-Driven Graph-Level Clustering Network
abstract
Graph-level clustering remains a pivotal yet formidable challenge in graph learning. Recently, the integration of deep learning with representation learning has demonstrated notable advancements, yielding performance enhancements to a certain degree. However, existing methods suffer from at least one of the following issues: 1) the original graph structure has noise, and 2) during feature propagation and pooling processes, noise is gradually aggregated into the graph-level embeddings through information propagation. Consequently, these two limitations mask clustering-friendly information, leading to suboptimal graph-level clustering performance. To this end, we propose a novel Dual Boost-Driven Graph-Level Clustering Network (DBGCN) to alternately promote graph-level clustering and filtering out interference information in a unified framework. Specifically, in the pooling step, we evaluate the contribution of features at the global and optimize them using a learnable transformation matrix to obtain high-quality graph-level representation, such that the model’s reasoning capability can be improved. Moreover, to enable reliable graph-level clustering, we first identify and suppress information detrimental to clustering by evaluating similarities between graph-level representations, providing more accurate guidance for multi-view fusion. Extensive experiments demonstrated that DBGCN outperforms the state-of-the-art graph-level clustering methods on six benchmark datasets.
Renda Han, Wenxuan Tu, Wenxin Zhang 0005, Jingxin Liu 0006, Jieren Cheng, Huajie Lei, Guangzhen Yao, Lingren Wang, Yu Li 0047
IJCNN9
2025 Multi-Relation Graph-Kernel Strengthen Network for Graph-Level Clustering
abstract
Graph-level clustering is a fundamental task of data mining, aiming at dividing unlabeled graphs into distinct groups. However, existing deep methods that are limited by pooling have difficulty extracting diverse and complex graph structure features, while traditional graph kernel methods rely on exhaustive substructure search, unable to adaptively handle multi-relational data. This limitation hampers producing robust and representative graph-level embeddings. To address this issue, we propose a novel Multi-Relation Graph-Kernel Strengthen Network for Graph-Level Clustering (MGSN), which integrates Multi-Relation Modeling (MRM) with graph kernel to fully employ their respective advantages. Specifically, MGSN constructs multi-relation graphs to capture diverse semantic relationships between nodes and graphs, which employ graph kernel methods to extract graph affinity, enriching the representation space. Moreover, a Relation-aware Embedding Strengthening Strategy (RESS) is designed, which adaptively aligns multi-relation information across views while strengthening graph-level features through a progressive fusion process. Extensive experiments on multiple benchmark datasets demonstrate the superiority of MGSN over state-of-the-art methods. The results highlight its ability to leverage multi-relation structures and graph kernel features, establishing a new paradigm for robust graph-level clustering.
Renda Han, Guangzhen Yao, Wenxin Zhang 0005, Yu Li 0047, Wen Xin, Huajie Lei, Zeyu Zhang 0006, Chengze Du 0001, Yahe Tian
IJCNN2
2025 JTFM: Joint Time-Frequency Method For Long-term Time Series Forecasting
abstract
Long-term Time Series Forecasting (LTSF) is an important task with extensive applications across diverse domains. While contemporary methodologies have achieved notable results through the integration of time and frequency domain features, significant challenges persist. Current approaches frequently disregard the information degradation inherent in Fast fourier transform (FFT) and inverse Fast fourier transform (IFFT) operations, substantially compromising predictive accuracy. Furthermore, conventional weighting mechanisms demonstrate limitations in their capacity to capture the intricate relationships between temporal and frequency representations, leading to suboptimal feature fusion and consequent information loss. To address these limitations, we present the Joint Time-Frequency Method (JTFM), a novel framework that simultaneously extracts sequence features from both temporal and frequency domains, thereby transcending single-domain constraints and enhancing feature comprehensiveness. Additionally, we introduce the Dynamic Harmonic Accumulation Weighting Mechanism (DHAWM), which surpasses traditional weighting approaches by dynamically modulating the relative contributions of temporal and frequency domain features based on sequence-specific characteristics. This adaptive mechanism strengthens the model’s feature representation capabilities and enhances forecasting precision. Empirical validation on eight real-world datasets demonstrates the JTFM’s superior performance compared to state-of-the-art baseline methods, establishing its efficacy in long-term time series forecasting applications.
Yu Li 0047, Wenxin Zhang 0005, Renda Han, Guangzhen Yao, Zeyu Zhang 0006, Cuicui Luo
IJCNN4
2025 MedConv: Convolutions Beat Transformers on Long-Tailed Bone Density Prediction
abstract
Bone density prediction via CT scans to estimate T-scores is crucial, providing a more precise assessment of bone health compared to traditional methods like X-ray bone density tests, which lack spatial resolution and the ability to detect localized changes. However, CT-based prediction faces two major challenges: the high computational complexity of transformer-based architectures, which limits their deployment in portable and clinical settings, and the imbalanced, long-tailed distribution of real-world hospital data that skews predictions. To address these issues, we introduce MedConv, a convolutional model for bone density prediction that outperforms transformer models with lower computational demands. We also adapt Bal-CE loss and post-hoc logit adjustment to improve class balance. Extensive experiments on our AustinSpine dataset shows that our approach achieves up to 21% improvement in accuracy and 20% in ROC AUC over previous state-of-the-art methods. Code will be available at https://github.com/Richardqiyi/MedConv.
Xuyin Qi, C. Zeyu Zhang, Huazhan Zheng, Mingxi Chen, Numan Kutaiba, Ruth Lim, Cherie Chiang, Zi En Tham, Xuan Ren, Wenxin Zhang 0005, Wenbing Lv, Guangzhen Yao, Renda Han, Kangsheng Wang, Hongtao Mao, Yu Li 0047, Zhibin Liao, Yang Zhao 0019, Minh-Son To
IJCNN14
2025 Enhancing Intra-Modality Compactness in Text-to-Image Person ReID
abstract
In ReID tasks, multimodal data helps address cross-view variations and occlusions. However, methods like CLIP and ViLBERT overlook intra-modal relationships, limiting their effectiveness. Additionally, current ReID models are trained on limited datasets, leaving their performance to improve further. To address this, we introduce a novel framework for person ReID that incorporates both text-to-text and image-to-image similarity modeling alongside traditional image-text alignment. We further introduce Similarity Distribution Matching (SDM), based on KL divergence, to align similarity distributions across and within modalities. We use MLM to generate additional text descriptions for each image, creating more image-text pairs to train a more generalized ReID model. Extensive experiments on benchmark datasets demonstrate our method achieves state-of-the-art performance in these tasks, validating its effectiveness and establishing a new benchmark for person ReID.
Zhanghao Qin, Hongtao Mao, Guangzhen Yao
IJCNN6
2025 GLFormer: A Lightweight Vision Transformer for Balancing Global and Local Information
abstract
In recent years, Vision Transformers (ViT) have achieved significant success in various complex visual tasks, but they also come with substantial computational costs and memory overheads. To address this issue, lightweight Vision Transformers have become an important research direction. Current research on lightweight ViTs mainly focuses on combining CNNs and Transformers, leveraging the advantages of CNNs in local feature extraction while utilizing Transformers’ ability to model global context. However, existing lightweight models often suffer from an imbalance in processing low-frequency global information and high-frequency local information. While sparse attention mechanisms effectively capture global context and reduce computational load, they typically adopt relatively simple strategies for handling high-frequency local information, failing to fully exploit the details of local features. To address this issue, we introduce a new lightweight Vision Transformer model, A Lightweight Vision Transformer for Balancing Global and Local Information (GLFormer). GLFormer combines dynamic weight adjustment with context-aware mechanisms to effectively aggregate high-frequency local information, optimizing the balance between global and local information. Additionally, we introduce a Depth Perception Feed-Forward Network (DPFFN), which further enhances feature fusion and detail refinement, thus enhancing the model’s performance and its capacity to generalize. Based on GLFormer and DPFFN, we design a novel visual backbone network—GLNet. Extensive experimental results show that GLNet consistently demonstrates excellent performance across various tasks, while maintaining a relatively low computational cost.
Zezhou Wang, Yuping Yuan, Suyang Chen, Guangzhen Yao, Chengze Du 0001, Renda Han, Bobin Xie, Sandong Zhu
IJCNN6
2025 RL-Pruner: Retraining-Free Global Exploration Pruning Method Based on Reinforcement Learning
abstract
Large language models (LLMs) have achieved significant success in complex tasks across various domains, but these achievements come with high computational costs and long inference delays. Pruning, as an effective optimization technique, simplifies model structures by removing redundant components, thereby improving model generalization and operational efficiency. Although existing pruning retraining-free algorithms perform excellently in pruning time, these algorithms often focus on local optimal solutions in encoder-based language models, lacking comprehensive exploration of global optimal solutions, which may affect the overall model performance. To address this issue, we propose a novel retraining-free structured pruning algorithm, named RL-Pruner. The algorithm consists of two main stages: the Mask Rearrangement Based on Asynchronous Advantage Actor-Critic (MA3C) stage and the BiConjugate Gradient Solver for Mask Tuning (BGMT) stage. It aims to explore the intra-layer interactions of mask variables and efficiently find the global optimal solution without requiring retraining. We evaluate this method using BERTBASEand DistilBERT models on the GLUE and SQuAD benchmark tests. Experimental results show that RL-Pruner significantly improves accuracy on the SQuAD1.1benchmark. Under a 60% FLOPs constraint, compared with existing pruning retraining-free algorithms, the F1 score increases by 4.25%.
Guangzhen Yao, Wenxin Zhang 0005, Xaioyu Deng, Chengze Du 0001, Renda Han, Zhanghao Qin, Yu Li 0047, Bobin Xie, Haiming Peng, Sandong Zhu, Zezhou Wang, Zeyu Zhang 0006
IJCNN1
2025 DConAD: A Differencing-based Contrastive Representation Learning Framework for Time Series Anomaly Detection
abstract
Time series anomaly detection holds notable importance for risk identification and fault detection across diverse application domains. Unsupervised learning methods have become popular because they have no requirement for labels. However, due to the challenges posed by the multiplicity of abnormal patterns, the sparsity of anomalies, and the growth of data scale and complexity, these methods often fail to capture robust and representative dependencies within the time series for identifying anomalies. To enhance the ability of models to capture normal patterns of time series and avoid the retrogression of modeling ability triggered by the dependencies on high-quality prior knowledge, we propose a differencing-based contrastive representation learning framework for time series anomaly detection (DConAD). Specifically, DConAD generates differential data to provide additional information about time series and utilizes transformer-based architecture to capture spatiotemporal dependencies, which enhances the robustness of unbiased representation learning ability. Furthermore, DConAD implements a novel KL divergence-based contrastive learning paradigm that only uses positive samples to avoid deviation from reconstruction and deploys the stop-gradient strategy to compel convergence. Extensive experiments on five public datasets show the superiority and effectiveness of DConAD compared with nine baselines. The code is available at https://github.com/shaieesss/DConAD.
Wenxin Zhang 0005, Xiaojian Lin, Guangzhen Yao, Jingxing Zhong, Yu Li 0047, Renda Han, Songcheng Xu, Cuicui Luo
IJCNN4
2025 Dual-channel Heterophilic Message Passing for Graph Fraud Detection
abstract
Fraudulent activities have significantly increased across various domains, such as e-commerce, online review platforms, and social networks, making fraud detection a critical task. Spatial Graph Neural Networks (GNNs) have been successfully applied to fraud detection tasks due to their strong inductive learning capabilities. However, existing spatial GNN-based methods often enhance the graph structure by excluding heterophilic neighbors during message passing to align with the homophilic bias of GNNs. Unfortunately, this approach can disrupt the original graph topology and increase uncertainty in predictions. To address these limitations, this paper proposes a novel framework, Dual-channel Heterophilic Message Passing (DHMP), for fraud detection. DHMP leverages a heterophily separation module to divide the graph into homophilic and heterophilic subgraphs, mitigating the low-pass inductive bias of traditional GNNs. It then applies shared weights to capture signals at different frequencies independently and incorporates a customized sampling strategy for training. This allows nodes to adaptively balance the contributions of various signals based on their labels. Extensive experiments on three real-world datasets demonstrate that DHMP outperforms existing methods, highlighting the importance of separating signals with different frequencies for improved fraud detection. The code is available at https://github.com/shaieesss/DHMP.
Wenxin Zhang 0005, Jingxing Zhong, Guangzhen Yao, Renda Han, Xiaojian Lin, Zeyu Zhang 0006, Cuicui Luo
IJCNN3
2025 A Reinforcement Learning-Based Retraining-Free Pruning for Encoder-Based Language Models
abstract
Natural Language Processing (NLP) has achieved significant success in complex tasks across various domains, yet it also brings high computational costs and inference delays. Pruning, as a model optimization technique, can effectively reduce model complexity and enhance its generalization capability and efficiency. However, current encoder-based language model pruning algorithms often lack robust dynamic adaptability and tend to focus only on short-term optimal solutions, without fully considering the interactions between different solutions. This limits their ability to find global optima, thereby potentially impacting overall model performance. To address these challenges, we propose a structured pruning algorithm based on reinforcement learning, named RLM (Reinforcement Learning Masking), which includes QLOM (Q-Learning Optimization Mask) and QRMT (Quasi-Minimal Residual Mask Tuning) components. This algorithm aims to rapidly and effectively find global optima without the need for retraining. We evaluated this method using BERTBASEand DistilBERT models on the GLUE and SQuAD benchmarks. Experimental results show that RLM significantly enhances model accuracy in the SQuAD benchmark. Under a 60% FLOPs constraint, RLM achieves a 8.45% increase in F1 score compared to existing retraining-free pruning algorithms, demonstrating its effectiveness in improving performance while managing computational resources efficiently.
Bobin Xie, Renda Han, Guangzhen Yao, Haiming Li, Sandong Zhu
ISCAS3
2025 A Lightweight Hybrid Network for Object Detection in Remote Sensing Images Balancing Global and Local Information
abstract
In recent years, hybrid convolutional neural networks (CNNs) and Transformer-based object detection technologies have achieved remarkable success. In the field of remote sensing image detection, since remote sensing systems rely on large-scale deployment of edge devices, detection models need to be lightweight with low parameter complexity to adapt to resource-constrained environments. However, existing lightweight models often struggle with an imbalance in extracting low-frequency global and high-frequency local information. In particular, when processing high-frequency local information (such as edges, textures, and fine structures), these models often lack in-depth analysis, leading to insufficient extraction of local features and reduced detection accuracy. To address the imbalance between low-frequency global information and high-frequency local information in lightweight remote sensing models, we propose an efficient and lightweight hybrid network detection framework, which mainly consists of the Global-Local Balance (GLB) module and the Detail-Aware Feature Fusion (DAFF) module. The GLB module adopts dynamic weight adjustment and context-aware mechanisms to effectively aggregate high-frequency local information in the image. The DAFF module further enhances feature fusion and detail refinement, improving the model’s performance and generalization ability. Experimental results on remote sensing datasets, including RSOD, NWPU VHR-10, and LEVIR, demonstrate that our proposed method achieves a well-balanced trade-off between model size and detection accuracy, reaching state-of-the-art performance.
Shuting Huang, Huanzun Zhang, Guangzhen Yao, Sandong Zhu, Jun Kong 0004
IEEE Geosci. Remote. Sens. Lett.5
2025 SOD-YOLOv10: Small Object Detection in Remote Sensing Images Based on YOLOv10
abstract
YOLOv10, known for its efficiency in object detection methods, quickly and accurately detects objects in images. However, when detecting small objects in remote sensing imagery, traditional algorithms often encounter challenges like background noise, missing information, and complex multiobject interactions, which can affect detection performance. To address these issues, we propose an enhanced algorithm for detecting small objects, named SOD-YOLOv10. We design the Multidimensional Information Interaction for the Transformer Backbone (TransBone) Network, which enhances global perception capabilities and effectively integrates both local and global information, thereby improving the detection of small object features. We also propose a feature fusion technology using an attention mechanism, called aggregated attention in a gated feature pyramid network (AA-GFPN). This technology uses an efficient feature aggregation network and re-parameterization techniques to optimize information interaction between feature maps of different scales. Additionally, by incorporating the aggregated attention (AA) mechanism, it accurately identifies essential features of small objects. Moreover, we propose the adaptive focal powerful IoU (AFP-IoU) loss function, which not only prevents excessive expansion of the anchor box area but also significantly accelerates model convergence. To evaluate our method, we conduct thorough tests on the RSOD, NWPU VHR-10, VisDrone2019, and AI-TOD datasets. The findings indicate that our SOD-YOLOv10 model attains 95.90%, 92.46%, 55.61%, and 59.47% for [email protected] and 73.42%, 66.84%, 39.03%, and 42.67% for [email protected]:0.95.
Guangzhen Yao, Sandong Zhu, Jun Kong 0004
IEEE Geosci. Remote. Sens. Lett.2
2024 DP-Prune: Global Optimal Strategy for Retraining-Free Pruning of Transformer Models
abstract
Transformer models have achieved significant success in various complex tasks, but their high computational costs and longer inference latency serve as limiting factors. To effectively reduce these costs, pruning has been widely adopted as an efficient method for Transformer models. Despite the excellent pruning speed demonstrated by existing retraining-free pruning algorithms, these methods often only find local optima when assessing the importance of attention heads and feed-forward networks. This limitation may lead to unstable solutions, thus affecting the overall performance of the model. To address these challenges, we propose DP-Prune (Dynamic Programming-Prune), a retraining-free structured pruning algorithm that employs a global optimization strategy. The algorithm consists of two parts: DPMO (Dynamic Programming Mask Optimization) and GSMT (GCROTMK Solver Mask Tuning), designed to quickly and effectively find global optima. We evaluate this method using BERTBASEand DistilBERT models on the GLUE and SQuAD benchmark tests. Experimental results demonstrate significant accuracy improvements on the SQuAD2.0task test without any further training. Under a 60% FLOPs constraint, DP-Prune achieves an 8.42% increase in F1 score compared with some existing retraining-free pruning algorithms.
Guangzhen Yao, Sandong Zhu, Miao Qi
IPCCC1