Zengyin Yang

dblp:199/1746 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
5since 2021 · last 2025
0000-0001-6307-7310ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 4 · 4 since 2021Computer networks · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Identifying Performance Issues in Cloud Service Systems Based on Relational-Temporal Features
abstract
Cloud systems, typically comprised of various components (e.g., microservices), are susceptible to performance issues, which may cause service-level agreement violations and financial losses. Identifying performance issues is thus of paramount importance for cloud vendors. In current practice, crucial metrics, i.e., Key Performance Indicators (KPIs), are monitored periodically to provide insight into the operational status of components. Identifying performance issues is often formulated as an anomaly detection problem, which is tackled by analyzing each metric independently. However, this approach overlooks the complex dependencies existing among cloud components. Some graph neural network-based methods take both temporal and relational information into account; however, the correlation violations in the metrics that serve as indicators of underlying performance issues are difficult for them to identify. Furthermore, a large volume of components in a cloud system results in a vast array of noisy metrics. This complexity renders it impractical for engineers to fully comprehend the correlations, making it challenging to identify performance issues accurately. To address these limitations, we propose Identifying Performance Issues based on Relational-Temporal Features (ISOLATE), a learning-based approach that leverages both the relational and temporal features of metrics to identify performance issues. In particular, it adopts a graph neural network with attention to characterizing the relations among metrics and extracts long-term and multi-scale temporal patterns using a GRU and a convolution network, respectively. The learned graph attention weights can be further used to localize the correlation-violated metrics. Moreover, to relieve the impact of noisy data, ISOLATE utilizes a Positive Unlabeled (PU) Learning strategy that tags pseudo-labels based on a small portion of confirmed negative examples. Extensive evaluation on both public and industrial datasets shows that ISOLATE outperforms all baseline models with 0.945 F1 score and 0.920 Hit rate@3. The ablation study also proves the effectiveness of the relational-temporal features and the PU-Learning strategy. Furthermore, we share the success stories of leveraging ISOLATE to identify performance issues in Huawei Cloud, which demonstrates its superiority in practice.
Wenwei Gu, Jinyang Liu 0002, Zhuangbin Chen, Jianping Zhang 0002, Yuxin Su 0001, Jiazhen Gu, Zengyin Yang, Yongqiang Yang, Michael R. Lyu
ACM Trans. Softw. Eng. Methodol.8
2024 LubeRDMA: A Fail-safe Mechanism of RDMA
abstract
Recent years have witnessed a wide adoption of Remote Direct Memory Access (RDMA) to accelerate distributed systems. As the scale of distributed applications keeps increasing, network failures become more prominent. Although some link/switch failures can be circumvented by in-network rerouting, failures like NIC failure are still fatal in RDMA networks and may cause the entire system to fail.
Shengkai Lin, Qinwei Yang, Zengyin Yang, Shizhen Zhao
APNet3
2024 Demystifying and Extracting Fault-indicating Information from Logs for Failure Diagnosis
abstract
Logs are imperative in the maintenance of online service systems, which often encompass important information for effective failure mitigation. While existing anomaly detection methodologies facilitate the identification of anomalous logs within extensive runtime data, manual investigation of log messages by engineers remains essential to comprehend faults, which is labor-intensive and error-prone. Upon examining the log-based troubleshooting practices at CloudA1, we find that engineers typically prioritize two categories of log information for diagnosis. These include fault-indicating descriptions, which record abnormal system events, and fault-indicating parameters, which specify the associated entities. Motivated by this finding, we propose an approach to automatically extract such fault-indicating information from logs for fault diagnosis, named LoFI. LoFI comprises two key stages. In the first stage, LoFI performs coarse-grained filtering to collect logs related to the faults based on semantic similarity. In the second stage, LoFI leverages a pre-trained language model with a novel prompt-based tuning method to extract fine-grained information of interest from the collected logs. We evaluate LoFI on logs collected from Apache Spark and an industrial dataset from CloudA. The experimental results demonstrate that LoFI outperforms all baseline methods by a significant margin, achieving an absolute improvement of 25.8˜37.9 in F1 over the best baseline method, ChatGPT. This highlights the effectiveness of LoFI in recognizing fault-indicating information. Furthermore, the successful deployment of LoFI at CloudA and user studies validate the utility of our method2.
Junjie Huang 0008, Jinyang Liu 0002, Yintong Huo, Jiazhen Gu, Zhuangbin Chen, Zengyin Yang, Michael R. Lyu
ISSRE9
2023 Practical Anomaly Detection over Multivariate Monitoring Metrics for Online Services
abstract
As modern software systems continue to grow in terms of complexity and volume, anomaly detection on multivariate monitoring metrics, which profile systems’ health status, becomes more and more critical and challenging. In particular, the dependency between different metrics and their historical patterns plays a critical role in pursuing prompt and accurate anomaly detection. Existing approaches fall short of industrial needs for being unable to capture such information efficiently. To fill this significant gap, in this paper, we propose CMAnomaly, an anomaly detection framework on multivariate monitoring metrics based on collaborative machine. The proposed collaborative machine is a mechanism to capture the pairwise interactions along with feature and temporal dimensions with linear time complexity. Cost-effective models can then be employed to leverage both the dependency between monitoring metrics and their historical patterns for anomaly detection. The proposed framework is extensively evaluated with both public data and industrial data collected from a large-scale online service system of Huawei Cloud. The experimental results demonstrate that compared with state-of-the-art baseline models, CMAnomaly achieves an average F1 score of 0.9494, outperforming baselines by 6.77% ~ 10.68%, and runs 10× ~ 20× faster. Furthermore, we also share our experience of deploying CMAnomaly in Huawei Cloud.
Jinyang Liu 0002, Zhuangbin Chen, Yuxin Su 0001, Zengyin Yang, Michael R. Lyu
ISSRE6
2023 Prism: Revealing Hidden Functional Clusters from Massive Instances in Cloud Systems
abstract
Ensuring the reliability of cloud systems is critical for both cloud vendors and customers. Cloud systems often rely on virtualization techniques to create instances of hardware resources, such as virtual machines. However, virtualization hinders the observability of cloud systems, making it challenging to diagnose platform-level issues. To improve system observability, we propose to infer functional clusters of instances, i.e., groups of instances having similar functionalities. We first conduct a pilot study on a large-scale cloud system, i.e., Huawei Cloud, demonstrating that instances having similar functionalities share similar communication and resource usage patterns. Motivated by these findings, we formulate the identification of functional clusters as a clustering problem and propose a non-intrusive solution called Prism. Prism adopts a coarse-to-fine clustering strategy. It first partitions instances into coarse-grained chunks based on communication patterns. Within each chunk, Prism further groups instances with similar resource usage patterns to produce fine-grained functional clusters. Such a design reduces noises in the data and allows Prism to process massive instances efficiently. We evaluate Prism on two datasets collected from the real-world production environment of Huawei Cloud. Our experiments show that Prism achieves a v-measure of ∼0.95, surpassing existing state-of-the-art solutions. Additionally, we illustrate the integration of Prism within monitoring systems for enhanced cloud reliability through two real-world use cases.
Jinyang Liu 0002, Jiazhen Gu, Junjie Huang 0008, Zhuangbin Chen, Zengyin Yang, Yongqiang Yang, Michael R. Lyu
ASE7
2017 Analyzing and optimizing BGP stability in future space-based internet
abstract
Future Space-Based Internet (FSBI) aims to provide global Internet access by interconnecting geosynchronous orbit (GEO), Medium Earth Orbit (MEO), Low Earth Orbit (LEO) satellites, and gateways on the ground. Border Gateway Protocol (BGP) is considered as a feasible routing protocol for interconnecting these independent facilities in FSBI. However, the vast Mobility-Related Topology Changes (MRTCs) in FSBI will put great stress on the BGP stability. We carry out a deep analysis on it and find that the routing updates of BGP will be triggered frequently. Many BGP routing updates will occur even before the accomplishment of previous updates and then the network will be instable for most of the time. To solve this problem, we build a Discrete-Time Topology Changes Aggregation (DT-TCA) scheme to improve BGP stability by making as many MRTCs as possible be triggered at the same time. Using a high-fidelity testbed and a virtualization-based emulator of FSBI, we show that DT-TCA dramatically improves the BGP stability in small-scale and large-scale scenarios.
Zengyin Yang, Hewu Li, Qian Wu 0001
IPCCC1