Yaqiang Zhang

dblp:196/6861 · DBLP profile ↗
← Back
11ranked-venue papers
8as first author
8since 2021 · last 2026
0000-0001-9935-0606ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 8 · 7 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Enabling Memory-Disaggregated Cloud Infrastructure for LLMs: An Adaptive CXL-based KV Cache Scheduling Approach
Yaqian Zhao, Yaqiang Zhang, Guangyuan Xu
INFOCOM3
2025 Proactive Fault-tolerance Driven Task Scheduling System for IoV Edge Networks
abstract
The emergence of Internet of Vehicles (IoV) technology provides a wider range of application scenarios for edge computing based on Vehicle-to-everything (V2X). It is essential to ensure the high availability and reliability of services in IoV systems. Currently, cloud service providers have established a data center level of fault tolerance, such as redundancy and checkpoints, guaranteeing the reliability of cloud infrastructure and reducing phenomena such as service termination or downtime. However, current computing systems reactively handle failures. Especially in edge computing, this approach not only lacks flexibility but also consumes excessive system resources, which is not conducive to ensuring the reliability in resource-constrained systems and poses security risks to end users. To mitigate this problem, we propose a Proactive Fault-tolerance Driven Task Scheduling System. Different from the traditional reactive strategies, the proposed framework predicts the possible system crashes by monitoring the critical state indicators of the computing system. According to the prediction results, a class of tasks or services that are most likely to be terminated are rescheduled in advance. Extensive experiments are conducted, and evaluation results demonstrate that our proposed proactive fault tolerance framework can effectively improve the long-term performance of the IoV edge system.
Yaqiang Zhang, RenGang Li, Yaqian Zhao, Hongzhi Shi, Guangyuan Xu
ICNP1
2025 Are GNNs Actually Effective for Multimodal Fault Diagnosis in Microservice Systems?
abstract
Graph Neural Networks (GNNs) are widely used for fault diagnosis in microservice systems, but their true benefit is often conflated with that of complex preprocessing pipelines. To isolate the GNN's contribution, we propose DiagMLP, a minimal, topology-agnostic MLP baseline. We conduct an ablation study by replacing GNN modules with DiagMLP in existing state-of-the-art frameworks. Across five datasets, this simple baseline achieves performance parity with GNN-based methods in fault detection, localization, and classification. These findings challenge the assumption that GNNs are indispensable, suggesting their contribution is marginal and that performance is primarily driven by preprocessing that already encodes critical dependency information. Our work advocates for a systematic re-evaluation of model complexity and the adoption of rigorous baselines to validate future innovations.
Ruyue Xin, Yaqiang Zhang
ICWS4
2025 Asymptotically Optimal Repair of Reed-Solomon Codes with Small Sub-Packetization under Rack-Aware Model
abstract
This paper presents a comprehensive study on the asymptotically optimal repair of Reed-Solomon (RS) codes with small sub-packetization, specifically tailored for rack-aware distributed storage systems. Through the utilization of multibase expansion, we introduce a novel approach that leverages monomials to construct linear repair schemes for RS codes. Our repair schemes which adapt to all admissible parameters achieve asymptotically optimal repair bandwidth while significantly reducing the sub-packetization compared with existing schemes. Furthermore, our approach is capable of repairing RS codes with asymptotically optimal repair bandwidth under the homogeneous storage model, achieving smaller sub-packetization than existing methods.
Zhongyan Liu, RenGang Li, Yaqian Zhao, Yaqiang Zhang
ITW5
2023 Multi-agent deep reinforcement learning for online request scheduling in edge cooperation networks
Yaqiang Zhang, Ruyang Li, Yaqian Zhao, RenGang Li, Zhangbing Zhou
Future Gener. Comput. Syst.1
2022 Deep Reinforcement Learning based Mobility-Aware Service Migration for Multi-access Edge Computing Environment
abstract
Multi-access Edge Computing (MEC) plays an im-portant role for providing end users with high reliability and low latency services at the edge of mobile network. In the scenario of Internet of Vehicles (IoV), vehicle users continually access nearby base stations to offload real-time tasks for reducing their computing overhead, while the ongoing services on current deployed edge nodes may be far away from users with the vehicles moving, potentially resulting in a high delay of data transmission. To address this challenge, in this paper, we propose a Deep Reinforcement Learning (DRL)-based mobility-aware service migration mechanism for effectively reducing the service delay and migration delay of the network. The proposed technique is adopted by re-calibrating required services at edge locations near the mobile user. Edge network state and user movement information are considered to ensure the generation of real-time service migration decision. Extensive experiments are conducted, and evaluation results demonstrate that our proposed DRL-based technique can effectively reduce the long-term average delay of the MEC system, compared with the state-of-the-art techniques.
Yaqiang Zhang, RenGang Li, Yaqian Zhao, Ruyang Li
ISCC1
2022 Online Decentralized Task Allocation Optimization for Edge Collaborative Networks
abstract
In centralized task allocation strategies, real-time status information needs to be collected from distributed edge nodes. Therefore, the overloaded transmission on backbone network appears and leads to devastating decrease in the per-formance of centralized strategies. To address this issue, this paper proposes a multi-agent deep reinforcement learning based online decentralized task allocation mechanism, where each edge node makes task allocation decisions based on local network-state information. A centralized-training distributed-execution method is adopted to decrease data transmission load, and a value decomposition-based technique is applied at training stage for improving long-term performance of task allocation in edge col-laborative networks. Extensive experiments are conducted, and evaluation results demonstrate that our mechanism outperforms other three baseline algorithms in reducing the long-term average system delay and improving request completion rate.
Yaqiang Zhang, Ruyang Li, Yaqian Zhao, RenGang Li, Xuelei Li
ISCC1
2021 Deep Reinforcement Learning for DAG-based Concurrent Requests Scheduling in Edge Networks
Yaqiang Zhang, Ruyang Li, Zhangbing Zhou, Yaqian Zhao, RenGang Li
WASA (3)1
2019 QoE-Constrained Concurrent Request Optimization Through Collaboration of Edge Servers
abstract
Cloud computing, which is claimed to provide plentiful storage, computational, and other resources, has become a promising platform to support resource-intensive applications. Due to the wide adoption of smart things to support domain applications and considering the delay-sensitivity of certain requests and limited network capacity compared with huge data packets to be transmitted, the quality of experience (QoE) may be hard to be satisfied when requests are solely supported by cloud computing. In this setting, edge computing has become an infrastructure to facilitate request satisfaction at the network edge. This article proposes a mechanism to optimize the collaboration of heterogeneous edge servers with certain QoE constraints. Specifically, concurrent requests, which are usually represented in terms of SQL queries, are rewritten as atomic queries, and these atomic queries are optimally assigned to edge servers through adopting an algorithm inspired by the minimum spanning tree, where QoE factors, including the delay, size of data packets, and number of operators, are considered. Evaluation results indicate that the proposed mechanism can effectively improve the QoE of requests compared with the state-of-the-art's mechanisms.
Yaqiang Zhang, Lin Meng 0001, Xiao Xue 0001, Zhangbing Zhou, Hiroyuki Tomiyama
IEEE Internet Things J.1
2018 Boundary Region Detection for Continuous Objects in Wireless Sensor Networks
abstract
Industrial Internet of Things has been widely used to facilitate disaster monitoring applications, such as liquid leakage and toxic gas detection. Since disasters are usually harmful to the environment, detecting accurate boundary regions for continuous objects in an energy‐efficient and timely fashion is a long‐standing research challenge. This article proposes a novel mechanism for continuous object boundary region detection in a fog computing environment, where sensing holes may exist in the deployed network region. Leveraging sensory data that have been gathered, interpolation algorithms have been applied to estimate sensory data at certain geographical locations, in order to estimate a more accurate boundary line. To examine whether estimated sensory data reflect that fact, mobile sensors are adopted to traverse these locations for gathering their sensory data, and the boundary region is calibrated accordingly. Experimental evaluation shows that this technique can generate a precise object boundary region with certain time constraints, and the network lifetime can be prolonged significantly.
Yaqiang Zhang, Zhenhua Wang 0005, Lin Meng 0001, Zhangbing Zhou
Wirel. Commun. Mob. Comput.1
2017 A Genetic Algorithm Based Mechanism for Scheduling Mobile Sensors in Hybrid WSNs Applications
Yaqiang Zhang, Zhangbing Zhou, Deng Zhao, Yunchuan Sun, Xiao Xue 0001
WASA1