Zeyu Che

dblp:357/7012 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2025
0009-0005-9369-7592ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 5 · 5 since 2021
YearPublicationVenuePosition
2025 LLM-Powered Multi-Agent Collaboration for Intelligent Industrial On-Call Automation
abstract
In large-scale enterprises, on-call engineers (OCEs) are critical for ensuring service availability and reliability. However, as incidents grow in volume and complexity, traditional manual on-call processes are becoming increasingly inadequate. Recent advances in large language models (LLMs) have demonstrated remarkable capabilities in reasoning and multi-agent collaboration, presenting new opportunities for automation. We propose OncallX, an end-to-end automated on-call system designed for real-world industrial scenarios that integrates LLMs with multi-agent cooperation to enable intelligent and efficient incident management. OncallX first enhances user queries by leveraging external knowledge bases and multi-turn dialogue interactions. Subsequently, multiple expert agents collaborate through tree-search-based mechanisms to generate effective responses and solutions. When incidents cannot be resolved automatically, OncallX accurately assigns them to the most appropriate teams. Comprehensive experiments conducted in the real-world production environment of a top-tier global online video service provider demonstrate that OncallX efficiently responds to incidents and accurately triages tickets, significantly outperforming existing methods in both automated metrics and human evaluations. Furthermore, OncallX has been successfully deployed in production for two months, during which it has substantially enhanced on-call efficiency, reducing average incident response time to just 21 seconds and average triage time to 4 seconds—representing a transformative improvement in operational excellence.
Ruowei Fu, Yang Zhang 0103, Zeyu Che, Zhenyu Zhong, Zhiqiang Ren, Shenglin Zhang, Feng Wang 0054, Yongqian Sun, Yu Zhang 0209
ASE3
2025 Efficient Multivariate Time Series Anomaly Detection through Transfer Learning for Large-Scale Software Systems
abstract
Timely anomaly detection of multivariate time series (MTS) is of vital importance for managing large-scale software systems. However, many deep learning-based MTS anomaly detection models require long-term MTS training data to achieve optimal performance, which often conflicts with the frequent pattern changes observed in software systems. Moreover, the training overhead of vast MTS in large-scale software systems is unacceptably high. To address these issues, we design OmniTransfer , a model-agnostic framework that combines weighted hierarchical agglomerative clustering with an adaptive transfer learning strategy, making many state-of-the-art (SOTA) MTS anomaly detection models efficient and effective. Extensive experiments using real-world data from a large web content service provider and a network operator show that OmniTransfer significantly reduces the model initialization time by 46.49% and the training cost by 74.51%, while maintaining high accuracy in detecting anomalies.
Yongqian Sun, Minghan Liang, Shenglin Zhang, Zeyu Che, Zhiyao Luo, Dongwen Li, Dan Pei, Lemeng Pan, Liping Hou
ACM Trans. Softw. Eng. Methodol.4
2024 LabelEase: A Semi-Automatic Tool for Efficient and Accurate Trace Labeling in Microservices
abstract
Trace data is crucial for system observability and maintainability within microservices architectures, and many operation algorithms depend heavily on trace data, including anomaly detection, root cause analysis, etc. However, the actual performance of these algorithms might be unsatisfactory due to the absence of high-quality labeled datasets for effective training and evaluation. Since billions of traces could be generated daily for large-scale microservices, labeling overhead is the main hurdle to obtaining high-quality trace datasets.In this paper, we propose LabelEase, a novel semi-automatic trace labeling tool, which uses active learning techniques to achieve efficient and accurate trace labeling. For anomaly trace labeling, LabelEase clusters similar traces with a graph-based trace representation technique and selects a few representative traces for human labeling, avoiding labeling most of the traces. For root cause labeling, LabelEase aggregates the labeled anomalous traces and identifies the service’s failures for operators to label. Our systematic experiments on two large-scale datasets show that LabelEase achieves over 0.98 F1-score in anomaly trace labeling and 0.89 precision of failure detection in root cause labeling, LabelEase can reduce operators’ labeling overhead by more than 99.9%. To the best of our knowledge, we are the first to propose a semi-automatic trace labeling tool capable of achieving efficient and accurate trace labeling.
Shenglin Zhang, Zeyu Che, Zhongjie Pan, Xiaohui Nie, Yongqian Sun, Lemeng Pan, Dan Pei
ISSRE2
2023 Efficient Multivariate Time Series Anomaly Detection Through Transfer Learning for Large-Scale Web Services
abstract
Timely anomaly detection of multivariate time series (MTS) is of vital importance for managing large-scale Web services. However, many deep learning-based MTS anomaly detection models require long-term MTS training data to achieve good performance, which conflicts with frequent pattern changes in Web services entities. Moreover, the training overhead of vast MTS in large-scale Web services is unacceptable. To address these issues, we design OmniTransfer, a model-agnostic framework that combines improved hierarchical agglomerative clustering with an adaptive transfer learning strategy, making many state-of-the-art (SOTA) MTS anomaly detection models efficient and effective. Extensive experiments using real-world data from a large Web content service provider show that OmniTransfer significantly reduces the model initialization time by 59.72% and the training cost by 85.01%, while maintaining high accuracy in detecting anomalies.
Yongqian Sun, Minghan Liang, Zeyu Che, Dongwen Li, Tinghua Zheng, Shenglin Zhang, Pengtian Zhu, Dan Pei
ICWS3
2023 An Empirical Analysis of Anomaly Detection Methods for Multivariate Time Series
abstract
Using multivariate time series (MTS) data for anomaly detection is widely adopted in service systems, such as web services and financial businesses. Researchers have recently proposed some well-performed algorithms for MTS anomaly detection from different perspectives. When applied to the real world, we observe that none of the algorithms is adaptable to all scenarios due to the complex data and anomaly characteristics. Moreover, there is currently a lack of comprehensive analysis work of these algorithms to guide operators in selecting the appropriate one in practice. To bridge this gap, we conduct an empirical study using various real-world data to gain an in-depth understanding of state-of-the-art anomaly detection algorithms. First, we provide general recommendations to guide operators in selecting suitable models based on the volume of training data, computational resources, and effectiveness requirements. Then, we summarize the typical data characteristics and types of anomalies and offer tailored model selection suggestions for different data characteristics and anomaly types. At last, we apply the summarized model selection suggestions to all the datasets we collected. The results show that most of our suggestions can achieve better than any single algorithm alone, demonstrating the effectiveness and generalization of our recommendations.
Dongwen Li, Shenglin Zhang, Yongqian Sun, Zeyu Che, Zhenyu Zhong, Minghan Liang, Minyi Shao, Mingjie Li 0005, Dan Pei
ISSRE5