Weiyao Zhang

dblp:281/4524 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 5 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SpecDetect: Simple, Fast, and Training-Free Detection of LLM-Generated Text via Spectral Analysis
abstract
The proliferation of high-quality text from Large Language Models (LLMs) demands reliable and efficient detection methods. While existing training-free approaches show promise, they often rely on surface-level statistics and overlook fundamental signal properties of the text generation process. In this work, we reframe detection as a signal processing problem, introducing a novel paradigm that analyzes the sequence of token log-probabilities in the frequency domain. By systematically analyzing the signal's spectral properties using the global Discrete Fourier Transform (DFT) and the local Short-Time Fourier Transform (STFT), we find that human-written text consistently exhibits significantly higher spectral energy. This higher energy reflects the larger-amplitude fluctuations inherent in human writing compared to the suppressed dynamics of LLM-generated text. Based on this key insight, we construct SpecDetect, a detector built on a single, robust feature from the global DFT: DFT total energy. We also propose an enhanced version, SpecDetect++, which incorporates a sampling discrepancy mechanism to further boost robustness. Extensive experiments show that our approach outperforms the state-of-the-art model while running in nearly half the time. Our work introduces a new, efficient, and interpretable pathway for LLM-generated text detection, showing that classical signal processing techniques offer a surprisingly powerful solution to this modern challenge.
Haitong Luo, Weiyao Zhang, Suhang Wang, Wenji Zou, Chungang Lin, Xuying Meng, Yujun Zhang 0001
AAAI2
2026 Heterogeneity-Oblivious Robust Federated Learning
abstract
Federated Learning (FL) remains highly vulnerable to poisoning attacks, especially under real-world hyper-heterogeneity, where clients differ significantly in data distributions, communication capabilities, and model architectures. Such heterogeneity not only undermines the effectiveness of aggregation strategies but also makes attacks more difficult to detect. Furthermore, high-dimensional models expand the attack surface. To address these challenges, we propose Horus, a heterogeneity-oblivious robust FL framework centered on low-rank adaptations (LoRAs). Rather than aggregating full model parameters, Horus inserts LoRAs into empirically stable layers and aggregates only LoRAs to reduce the attack uncover a key empirical observation that the input projection (LoRA-A) is markedly more stable than the output projection (LoRA-B) under heterogeneity and poisoning. Leveraging this, we design a Heterogeneity-Oblivious Poisoning Score using the features from LoRA-A to filter poisoned clients. For the remaining benign clients, we propose projection-aware aggregation mechanism to preserve collaborative signals while suppressing drifts, which reweights client updates by consistency with the global directions. Extensive experiments across diverse datasets, model architectures, and attacks demonstrate that Horus consistently outperforms state-of-the-art baselines in both robustness and accuracy.
Weiyao Zhang, Chungang Lin, Haitong Luo, Xuying Meng
INFOCOM1
2026 Dynamic Federated Edge Anomaly Detection Based on Dual Fuzzing Strategies
abstract
With the rapid proliferation of edge computing, mobile edges interconnected via networks are increasingly ex posed to a wide range of attacks. Considering privacy concerns, federated edge anomaly detection, collaboratively training a global detection model across multiple edge clients without centralizing sensitive data, is a promising paradigm to detecting such attacks. However, mobile edge environment is inherently dynamic, with frequent evolving in both network traffic and participating clients. Most existing methods are pre-defined for specific attacks and fixed client settings, which significantly limits their adaptability in dynamic mobile edges. To address these issues, we define the federated adaptability from both data-level and client-level perspectives for the first time, and derive the decisive factors for improving adaptability, i.e., high smoothness and low gradient similarity. Driven by this theoretical foundation, we propose DIFF, a federated edge anomaly detection framework based on dual fuzzing strategies. DIFF incorporates two novel designs for improving adaptability: (i) a classless fuzzing strategy to improve data-level adaptability by fuzzing the boundary between normal and abnormal samples, thereby encouraging model smoothness and enhancing adaptability to unseen traffic and emerging attacks; and (ii) a direction fuzzing strategy to enhance client-level adaptability by perturbing the optimization directions of local models, enabling the aggregated global model to adapt effectively to new clients. Experiments on public datasets demonstrate the superior adapt ability of DIFF compared to the state-of-art methods.
Weiyao Zhang, Jinyang Li 0009, Botao Peng, Xuying Meng, Yujun Zhang 0001
IEEE Trans. Mob. Comput.1
2025 TraGe: A Generic Packet Representation for Traffic Classification Based on Header-Payload Differences
abstract
Traffic classification has a significant impact on maintaining the Quality of Service (QoS) of the network. Since traditional methods heavily rely on feature extraction and largescale labeled data, some recent pre-trained models manage to reduce the dependency by utilizing different pre-training tasks to train generic representations for network packets. However, existing pre-trained models typically adopt pre-training tasks developed for image or text data, which are not tailored to traffic data. As a result, the obtained traffic representations fail to fully reflect the information contained in the traffic, and may even disrupt the protocol information. To address this, we propose TraGe, a novel generic packet representation model for traffic classification. Based on the differences between the header and payload-the two fundamental components of a network packetwe perform differentiated pre-training according to the byte sequence variations (continuous in the header vs. discontinuous in the payload). A dynamic masking strategy is further introduced to prevent overfitting to fixed byte positions. Once the generic packet representation is obtained, TraGe can be finetuned for diverse traffic classification tasks using limited labeled data. Experimental results demonstrate that TraGe significantly outperforms state-of-the-art methods on two traffic classification tasks, with up to a 6.97% performance improvement. Moreover, TraGe exhibits superior robustness under parameter fluctuations and variations in sampling configurations.
Chungang Lin, Yilong Jiang, Weiyao Zhang, Xuying Meng, Tianyu Zuo, Yujun Zhang 0001
IWQoS3
2025 LogOW: A semi-supervised log anomaly detection model in open-world setting
Jingwei Ye, Zhaojun Gu, Xuying Meng, Weiyao Zhang, Yujun Zhang 0001
J. Syst. Softw.6
2024 Spectral-Based Graph Neural Networks for Complementary Item Recommendation
abstract
Modeling complementary relationships greatly helps recommender systems to accurately and promptly recommend the subsequent items when one item is purchased. Unlike traditional similar relationships, items with complementary relationships may be purchased successively (such as iPhone and Airpods Pro), and they not only share relevance but also exhibit dissimilarity. Since the two attributes are opposites, modeling complementary relationships is challenging. Previous attempts to exploit these relationships have either ignored or oversimplified the dissimilarity attribute, resulting in ineffective modeling and an inability to balance the two attributes. Since Graph Neural Networks (GNNs) can capture the relevance and dissimilarity between nodes in the spectral domain, we can leverage spectral-based GNNs to effectively understand and model complementary relationships. In this study, we present a novel approach called Spectral-based Complementary Graph Neural Networks (SComGNN) that utilizes the spectral properties of complementary item graphs. We make the first observation that complementary relationships consist of low-frequency and mid-frequency components, corresponding to the relevance and dissimilarity attributes, respectively. Based on this spectral observation, we design spectral graph convolutional networks with low-pass and mid-pass filters to capture the low-frequency and mid-frequency components. Additionally, we propose a two-stage attention mechanism to adaptively integrate and balance the two attributes. Experimental results on four e-commerce datasets demonstrate the effectiveness of our model, with SComGNN significantly outperforming existing baseline models.
Haitong Luo, Xuying Meng, Suhang Wang, Hanyun Cao, Weiyao Zhang, Yequan Wang, Yujun Zhang 0001
AAAI5
2023 Multi-Layer Collaborative Bandit for Multivariate Time Series Anomaly Detection
abstract
Multivariate Time Series Anomaly Detection (MTSAD) detects abnormal indicators from Multivariate Time Series (MTS), and provides the rank of the multiple abnormal indicators to meet the expert's detection interest in current environment, which underpin the security and stability of intelligent cyber-physical systems. However, popular integration-based methods, which are pre-defined, fall short in locating the exact abnormal indicator, nor can they perceive the environmental dynamic and evolve accordingly. Let alone meeting the expert's interest. As a result, the expert's workload is exaggerated. These issues motivate us to propose a novel multi-layer collaborative bandit framework MULA for MTSAD. MULA decomposes MTS and pairs individual time series with a bandit arm, which locates the abnormal indicator directly. Then, MULA sorts the indicators by abnormal scores computed based on the expert's feedback, which facilitates the experts. Besides, to address the adaptability issue, we devise a dual signal to comprehensively monitor environmental changes, and design a multi-layer collaborative mechanism for MULA to adapt to the dynamic environment. Theoretical analysis and experiments on public datasets demonstrate the superiority of MULA compared to the state-of-art.
Weiyao Zhang, Xuying Meng, Jinyang Li 0009, Yequan Wang, Yujun Zhang 0001
IWQoS1
2022 Dual-track Protocol Reverse Analysis Based on Share Learning
abstract
Private protocols, whose specifications are agnostic, are widely used in the Industrial Internet. While providing customized service, they also raise essential security concerns as well, due to their agnostic nature. The Protocol Reverse Analysis (PRA) techniques are developed to infer the specifications of private protocols. However, the conventional PRA techniques are far from perfection for the following reasons: (i) Error propagation: Canonical solutions strictly follow the "from keyword extraction to message clustering" serial structure, which deteriorates the performance for ignoring the interplay between the sub-tasks, and the error will flow and accumulate through the sequential workflow. (ii) Increasing diversity: As the protocols’ diversities of characteristics increase, tailoring for specific types of protocols becomes infeasible. To address these issues, we design a novel dual-track framework SPRA, and propose Share Learning, a new concept of protocol reverse analysis. Particularly, based on the share layer for protocol learning, SPRA builds a parallel workflow to co-optimize both the generative model for keyword extraction and the probability-based model for message clustering, which delivers automatic and robust syntax inference across diverse protocols and greatly improves the performance. Experiments on five real-world datasets demonstrate that the proposed SPRA achieves better performance compared with the state-of-art PRA methods.
Weiyao Zhang, Xuying Meng, Yujun Zhang 0001
INFOCOM1