VLDB 2026 Research / reviewers in the wild / expert
Zeyan Li 0001
dblp:187/8216
· DBLP profile ↗
19ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0002-3529-5879ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 6 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 6 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Systems, architecture and hardware · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ChatTS: Aligning Time Series with LLMs via Synthetic Data for Enhanced Understanding and ReasoningabstractUnderstanding time series is crucial for its application in real-world scenarios. Recently, large language models (LLMs) have been increasingly applied to time series tasks, leveraging their strong language capabilities to enhance various applications. However, research on multimodal LLMs (MLLMs) for time series understanding and reasoning remains limited, primarily due to the scarcity of high-quality datasets that align time series with textual information. This paper introduces ChatTS, a novel MLLM designed for time series analysis. ChatTS treats time series as a modality, similar to how vision MLLMs process images, enabling it to perform both understanding and reasoning with time series. To address the scarcity of training data, we propose an attribute-based method for generating synthetic time series and Time Series Evol-Instruct to generates diverse Q&As for enhanced reasoning capabilities. To the best of our knowledge, ChatTS is the first MLLM that takes multivariate time series as input for understanding and reasoning, which is fine-tuned exclusively on synthetic datasets. We evaluate its performance using benchmark datasets with real-world data, including six alignment tasks and four reasoning tasks. Our results show that ChatTS significantly outperforms existing vision-based MLLMs (e.g., GPT-4o) and text/agent-based LLMs, achieving a 46.0% improvement in alignment tasks and a 25.8% improvement in reasoning tasks. We have open-sourced the source code, model checkpoint and datasets at https://github.com/NetManAIOps/ChatTS. Zhe Xie, Zeyan Li 0001, Xiao He 0008, Longlong Xu, Xidao Wen, Tieying Zhang, Jianjun Chen 0001, Dan Pei |
Proc. VLDB Endow. | 2 |
| 2024 | SparseRCA: Unsupervised Root Cause Analysis in Sparse Microservice Testing TracesabstractMicroservice architecture has become a predominant paradigm in the software industry. This architecture necessitates robust end-to-end testing to ensure seamless integration of all components before deployment. Rapidly pinpointing issues when test cases fail is crucial for enhancing software development efficiency. However, in testing environments, the available trace is often sparse, and the system is continuously upgrading, which renders existing microservice-based root cause analysis (RCA) ineffective. To address these challenges, we propose SparseRCA. By assessing the abnormality of the exclusive latency, SparseRCA directly determines the probability of the root cause, solving the challenge of not being able to fully obtain the fault propagation information, such as call relationships in sparse trace scenarios. At the same time, by reconstructing the exclusive latency using the decoupled atomic span units, it solves the problem of latency prediction for new traces caused by frequent upgrades. We evaluate SparseRCA on real-world datasets from a large e-commerce system’s testing environment, where it demonstrates significant improvements over existing models. Our findings underscore the effectiveness of SparseRCA in addressing the challenges of RCA in microservice testing environments. Zhenhe Yao, Haowei Ye, Changhua Pei, Guangpei Wang, Hang Cui 0004, Zeyan Li 0001, Gaogang Xie, Dan Pei |
ISSRE | 9 |
| 2024 | A survey on intelligent management of alerts and incidents in IT services
Qingyang Yu, Nengwen Zhao, Mingjie Li 0005, Zeyan Li 0001, Honglin Wang, Wenchi Zhang, Kaixin Sui, Dan Pei |
J. Netw. Comput. Appl. | 4 |
| 2024 | Diagnosing Performance Issues for Large-Scale Microservice Systems With Heterogeneous GraphabstractThe availability of microservice systems is critical to business operations and corporate reputation. However, the dynamics and complexity of microservice systems introduce significant challenges to the performance issue diagnosis of large-scale microservice systems. After investigating hundreds of real-world performance issue cases in Tencent, we find that previous troubleshooting approaches fail to accurately localize root causes because they overlook the inconsistency between causality and calling relationships. Therefore, we propose a novel approach, MicroDig, to diagnose performance issues for large-scale microservice systems. Specifically, MicroDig constructs a heterogeneous propagation graph to capture the causal relationships between calls and microservices. It then conducts a heterogeneity-oriented random walk (HORW) to pinpoint the culprit microservice. Extensive evaluation experiments have been conducted to evaluate MicroDig's performance on 60 real-world performance issues collected from Tencent, 80 manually injected ones collected from a widely used open-source microservice system and 128 performance issues collected from an e-commerce system used by a top-tier global commercial bank. MicroDig achieves 94.1%, 85.5% and 93.8% top-3 accuracy on the three datasets, respectively, significantly outperforming six popular baseline methods. Additionally, we have shared our success stories and learned lessons from the deployment of MicroDig in Tencent. Xianglin Lu, Shenglin Zhang, Jiaqi Luan, Yingke Li, Mingjie Li 0005, Zeyan Li 0001, Qingyang Yu, Hucheng Xie, Chenyuan Hu, Canqun Yang, Dan Pei |
IEEE Trans. Serv. Comput. | 7 |
| 2023 | CMDiagnostor: An Ambiguity-Aware Root Cause Localization Approach Based on Call Metric DataabstractThe availability of online services is vital as its strong relevance to revenue and user experience. To ensure online services’ availability, quickly localizing the root causes of system failures is crucial. Given the high resource consumption of traces, call metric data are widely used by existing approaches to construct call graphs in practice. However, ambiguous correspondences between upstream and downstream calls may exist and result in exploring unexpected edges in the constructed call graph. Conducting root cause localization on this graph may lead to misjudgments of real root causes. To the best of our knowledge, we are the first to investigate such ambiguity, which is overlooked in the existing literature. Inspired by the law of large numbers and the Markov properties of network traffic, we propose a regression-based method (named AmSitor) to address this problem effectively. Based on AmSitor, we propose an ambiguity-aware root cause localization approach based on Call Metric Data named CMDiagnostor, containing metric anomaly detection, ambiguity-free call graph construction, root cause exploration, and candidate root cause ranking modules. The comprehensive experimental evaluations conducted on real-world datasets show that our CMDiagnostor can outperform the state-of-the-art approaches by 14% on the top-5 hit rate. Moreover, AmSitor can also be applied to existing baseline approaches separately to improve their performances one step further. The source code is released at https://github.com/NetManAIOps/CMDiagnostor. Qingyang Yu, Changhua Pei, Mingjie Li 0005, Zeyan Li 0001, Shenglin Zhang, Xianglin Lu, Jiaqi Li 0021, Dan Pei |
WWW | 5 |
| 2023 | Generic and robust root cause localization for multi-dimensional data in online service systems
Zeyan Li 0001, Junjie Chen 0003, Yiwei Zhao 0001, Yongqian Sun, Kaixin Sui, Xiping Wang, Dan Pei |
J. Syst. Softw. | 1 |
| 2022 | Generic and Robust Performance Diagnosis via Causal Inference for OLTP Database SystemsabstractOnline transaction processing (OLTP) database systems provide an effective solution to data support for online applications with high concurrency and low latency. An interruption or performance degradation of OLTP database systems may impact the availability of services and bring substantial economic loss. Thus, diagnosing the issue timely and mitigating it rapidly are essential for database administrators (DBAs). However, performance diagnosis for database systems is challenging due to numerous abnormal metrics, complex failure propagation, and high-performance requirements. Existing works relying on anomaly detection or causal graph construction cannot handle all these challenges simultaneously. In this paper, we propose an unsupervised learning-based method, CauseRank, to perform root cause localization with superior efficiency, high accuracy, and good interpretability. Two key techniques in CauseRank are a novel causal discovery algorithm named Group-based Greedy Equivalent Search (G-GES) incorporated with domain knowledge which treats metric groups as nodes to capture failure propagation and a simple yet effective ranking method named Causal Oriented Personalized PageRank (COPP). Extensive experiments on 97 real-world failure cases collected from a large-scale Oracle database demonstrate the effectiveness of CauseRank, achieving 82.5% top-3 accuracy and 93.8% top-5 accuracy and outperforming baseline approaches. The core idea and framework of CauseRank are generic and can be applied to other large-scale system components. Xianglin Lu, Zhe Xie, Zeyan Li 0001, Mingjie Li 0005, Xiaohui Nie, Nengwen Zhao, Qingyang Yu, Shenglin Zhang, Kaixin Sui, Dan Pei |
CCGRID | 3 |
| 2022 | Causal Inference-Based Root Cause Analysis for Online Service Systems with Intervention RecognitionabstractFault diagnosis is critical in many domains, as faults may lead to safety threats or economic losses. In the field of online service systems, operators rely on enormous monitoring data to detect and mitigate failures. Quickly recognizing a small set of root cause indicators for the underlying fault can save much time for failure mitigation. In this paper, we formulate the root cause analysis problem as a new causal inference task namedintervention recognition. We proposed a novel unsupervised causal inference-based method namedCausal Inference-based Root Cause Analysis (CIRCA). The core idea is a sufficient condition for a monitoring variable to be a root cause indicator,i.e., the change of probability distribution conditioned on the parents in the Causal Bayesian Network (CBN). Towards the application in online service systems, CIRCA constructs a graph among monitoring metrics based on the knowledge of system architecture and a set of causal assumptions. The simulation study illustrates the theoretical reliability of CIRCA. The performance on a real-world dataset further shows that CIRCA can improve the recall of the top-1 recommendation by 25% over the best baseline method. Mingjie Li 0005, Zeyan Li 0001, Kanglin Yin, Xiaohui Nie, Wenchi Zhang, Kaixin Sui, Dan Pei |
KDD | 2 |
| 2022 | Actionable and interpretable fault localization for recurring failures in online service systemsabstractFault localization is challenging in an online service system due to its monitoring data's large volume and variety and complex dependencies across/within its components (e.g., services or databases). Furthermore, engineers require fault localization solutions to be actionable and interpretable, which existing research approaches cannot satisfy. Therefore, the common industry practice is that, for a specific online service system, its experienced engineers focus on localization for recurring failures based on the knowledge accumulated about the system and historical failures. More specifically, 1) they can identify the underlying root causes and take mitigation actions when pinpointing a group of indicative metrics on the faulty component; 2) their diagnosis knowledge is roughly based on how one failure might affect the components in the whole system. Zeyan Li 0001, Nengwen Zhao, Mingjie Li 0005, Xianglin Lu, Dongdong Chang, Xiaohui Nie, Wenchi Zhang, Kaixin Sui, Guoqiang Duan, Dan Pei |
ESEC/SIGSOFT FSE | 1 |
| 2021 | Practical Root Cause Localization for Microservice Systems via Trace AnalysisabstractMicroservice architecture is applied by an increasing number of systems because of its benefits on delivery, scalability, and autonomy. It is essential but challenging to localize root-cause microservices promptly when a fault occurs. Traces are helpful for root-cause microservice localization, and thus many recent approaches utilize them. However, these approaches are less practical due to relying on supervision or other unrealistic assumptions. To overcome their limitations, we propose a more practical root-cause microservice localization approach named TraceRCA. The key insight of TraceRCA is that a microservice with more abnormal and less normal traces passing through it is more likely to be the root cause. Based on it, TraceRCA is composed of trace anomaly detection, suspicious microservice set mining and microservice ranking. We conducted experiments on hundreds of injected faults in a widely-used open-source microservice benchmark and a production system. The results show that TraceRCA is effective in various situations. The top-1 accuracy of TraceRCA outperforms the state-of-the-art unsupervised approaches by 44.8%. Besides, TraceRCA is applied in a large commercial bank, and it helps operators localize root causes for real-world faults accurately and efficiently. We also share some lessons learned from our real-world deployment. Zeyan Li 0001, Junjie Chen 0003, Nengwen Zhao, Shuwei Zhang, Long Jiang, Leiqin Yan, Zikai Wang 0010, Zhekang Chen, Wenchi Zhang, Xiaohui Nie, Kaixin Sui, Dan Pei |
IWQoS | 1 |
| 2021 | An empirical investigation of practical log anomaly detection for online service systemsabstractLog data is an essential and valuable resource of online service systems, which records detailed information of system running status and user behavior. Log anomaly detection is vital for service reliability engineering, which has been extensively studied. However, we find that existing approaches suffer from several limitations when deploying them into practice, including 1) inability to deal with various logs and complex log abnormal patterns; 2) poor interpretability; 3) lack of domain knowledge. To help understand these practical challenges and investigate the practical performance of existing work quantitatively, we conduct the first empirical study and an experimental study based on large-scale real-world data. We find that logs with rich information indeed exhibit diverse abnormal patterns (e.g., keywords, template count, template sequence, variable value, and variable distribution). However, existing approaches fail to tackle such complex abnormal patterns, producing unsatisfactory performance. Motivated by obtained findings, we propose a generic log anomaly detection system named LogAD based on ensemble learning, which integrates multiple anomaly detection approaches and domain knowledge, so as to handle complex situations in practice. About the effectiveness of LogAD, the average F1-score achieves 0.83, outperforming all baselines. Besides, we also share some success cases and lessons learned during our study. To our best knowledge, we are the first to investigate practical log anomaly detection in the real world deeply. Our work is helpful for practitioners and researchers to apply log anomaly detection to practice to enhance service reliability. Nengwen Zhao, Honglin Wang, Zeyan Li 0001, Zhu Pan, Xidao Wen, Wenchi Zhang, Kaixin Sui, Dan Pei |
ESEC/SIGSOFT FSE | 3 |
| 2020 | ZeroWall: Detecting Zero-Day Web Attacks through Encoder-Decoder Recurrent Neural NetworksabstractThe following topics are dealt with: learning (artificial intelligence); optimisation; telecommunication traffic; Internet; cloud computing; computational complexity; mobile computing; resource allocation; security of data; and telecommunication network routing. Ruming Tang, Zeyan Li 0001, Weibin Meng, Haixin Wang 0003, Qi Li 0002, Yongqian Sun, Dan Pei, Tao Wei 0002, Yanfei Xu, Yan Liu 0069 |
INFOCOM | 3 |
| 2020 | Integrated Control-Fluidic Codesign Methodology for Paper-Based Digital Microfluidic BiochipsabstractPaper-based digital microfluidic biochips (P-DMFBs) have recently emerged as a promising low-cost and fast-responsive platform for biochemical assays. In P-DMFBs, electrodes and control lines are printed on a piece of photograph paper using an inkjet printer and carbon nanotubes (CNTs) conductive ink. Compared with traditional digital microfluidic biochips (DMFBs), P-DMFBs enjoy significant advantages, such as faster in-place fabrication with printer and ink, lower costs, and better disposability. Since electrodes and CNT control lines are printed on the same side of this paper, a critical design challenge for P-DMFB is to prevent control interference between moving droplets and the voltages on CNT control lines. Control interference may result in unexpected droplet movements and thus incorrect assay outputs. To address this design challenge, a control-fluidic codesign methodology is proposed in this paper, along with two demonstrative design flows integrating both fluidic design and control design, i.e., the droplet-oriented codesign flow and the electrode-oriented codesign flow. The droplet-oriented flow is suitable for designing biochips with sparse electrodes and relatively larger number of droplets, whereas the electrode-oriented flow is suitable for biochips with dense electrodes and smaller number of droplets. The computational simulation results of real-life bioassays demonstrate the effectiveness of the proposed codesign flows. Qin Wang 0005, Ulf Schlichtmann, Yici Cai, Weiqing Ji, Zeyan Li 0001, Haena Cheong, Oh-Sun Kwon, Hailong Yao 0002, Tsung-Yi Ho, Kwanwoo Shin, Bing Li 0005 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2019 | Unsupervised Anomaly Detection for Intricate KPIs via Adversarial Training of VAEabstractTo ensure the reliability of the Internet-based application services, KPIs (Key Performance Monitors) are closely monitored in real time and the anomalies presented in the KPIs must be discovered in time. While anomaly detection for the seasonal smooth service-level KPIs (e.g., number of transactions per minute) have been solved reasonably well in the literature, the intricate KPIs at the machine level (e.g., the number of I/O requests on a server monitored per second) has been little studied. These intricate KPIs are prevalent and important, but exhibit non-Gaussian noises and complex data distribution that are hard to model. In this paper, we propose an adversarial training method in the Bayesian network based on partition analysis with solid theoretical proof. Based on it, we propose the first unsupervised anomaly detection algorithmBuzz for intricate KPIs with high performance. Its best F-scores on the data from a global Internet company range from 0.92 to 0.99, significantly outperforming a state-of-art VAE-based unsupervised approach without adversarial training and a state-of-art supervised approach. Wenxiao Chen, Zeyan Li 0001, Dan Pei, Honglin Qiao, Zhaogang Wang |
INFOCOM | 3 |
| 2019 | Generic and Robust Localization of Multi-dimensional Root CausesabstractOperators of online software services periodically collect various measures with many attributes. When a measure becomes abnormal, indicating service problems such as reliability degrade, operators would like to rapidly and accurately localize the root cause attribute combinations within a huge multi-dimensional search space. Unfortunately, previous approaches are not generic or robust in that they all suffer from impractical root cause assumptions, handling only directly collected measures but not derived ones, handling only anomalies with signicant magnitudes but not those insignicant but important ones, requiring manual parameter ne-tuning, or being too slow. This paper proposes a generic and robust multi-dimensional root cause localization approach, Squeeze, that overcomes all above limitations, the first in the literature. Through our novel bottom-up then top-down searching strategy and the techniques based on our proposed generalized ripple effect and generalized potential score, Squeeze is able to reach a good trade off between search speed and accuracy in a generic and robust manner. Case studies in several banks and an Internet company show that Squeeze can localize root causes much more rapidly and accurately than the traditional manual analysis. Furthermore, our extensive experiments on semi-synthetic datasets show that the F1-score of Squeeze outperforms previous approaches by 0.4 on average, while its localization time is only about 10 seconds. Zeyan Li 0001, Dan Pei, Yiwei Zhao 0001, Yongqian Sun, Kaixin Sui, Xiping Wang |
ISSRE | 1 |
| 2018 | Robust and Unsupervised KPI Anomaly Detection Based on Conditional Variational AutoencoderabstractTo ensure undisrupted web-based services, operators need to closely monitor various KPIs (Key Performance Indicator, such as CPU usages, network throughput, page views, number of online users, and etc), detect anomalies in them, and trigger timely troubleshooting or mitigation. There can be hundreds of thousands to even millions of KPIs to be monitored, thus operators need automatic anomaly detection approaches. However, neither traditional statistical approaches nor supervised ensemble approaches satisfy this requirement in practice when facing large number of KPIs. A state-of-art unsupervised approach Donut offering promising results, but it is not a sequential model thus cannot deal with the time information related anomalies. Thus, in this paper we propose Bagel, a robust and unsupervised anomaly detection algorithm for KPI that can handle time information related anomalies, using CVAE to incorporate time information and dropout layer to avoid overfitting. Our experiments using real data from Internet companies show that, compared to Donut, Bagel improves the anomaly detection best F1-score by 0.08 to 0.43. Zeyan Li 0001, Wenxiao Chen, Dan Pei |
IPCCC | 1 |
| 2018 | The Frame Latency of Personalized Livestreaming Can Be Significantly Slowed Down by WiFiabstractThe popular personalized livestreaming (PL) in China, arguably the largest PL market in the world, is more monetized than PL in US and hence demands much lower interactive latencies to ensure a good quality of user experience. However, our pilot experiment shows that the video frame latency, dominant component of PL's interactive latency, can be significantly slowed down by WiFi, the primary Internet access method for PL. Understanding and further improving the frame latency over WiFi, however, have difficulties in 1) measuring end-to-end latency; 2) parsing encrypted PL's traffic and 3) modeling complex relationships between WiFi radio factors and the latency. To tackle these challenges, we design and prototype Latency Doctor (LTDr), a practical system which aims to model and optimize PL's video frame latency over WiFi. We deploy LTDr in our campus and obtain several key observations based on 13.9M video frames extracted from 12K individual views on three leading PLs in China. We observe that 40% frame latencies over WiFi hop are more than 30ms, and channel utilization should be less than 64% for low latency. Then we build a predictive model based on the dataset using the machine learning methodologies. Two real cases show that the median frame latencies are decreased by LTDr from 130ms to 22ms, and 50ms to 12ms respectively over WiFi networks. Guoshun Nan, Xiuquan Qiao, Jiting Wang, Zeyan Li 0001, Jiahao Bu, Changhua Pei, Mengyu Zhou, Dan Pei |
IPCCC | 4 |
| 2018 | Unsupervised Anomaly Detection via Variational Auto-Encoder for Seasonal KPIs in Web ApplicationsabstractTo ensure undisrupted business, large Internet companies need to closely monitor various KPIs (e.g., Page Views, number of online users, and number of orders) of its Web applications, to accurately detect anomalies and trigger timely troubleshooting/mitigation. However, anomaly detection for these seasonal KPIs with various patterns and data quality has been a great challenge, especially without labels. In this paper, we proposed Donut, an unsupervised anomaly detection algorithm based on VAE. Thanks to a few of our key techniques, Donut greatly outperforms a state-of-arts supervised ensemble approach and a baseline VAE approach, and its best F-scores range from 0.75 to 0.9 for the studied KPIs from a top global Internet company. We come up with a novel KDE interpretation of reconstruction for Donut, making it the first VAE-based anomaly detection algorithm with solid theoretical explanation. Wenxiao Chen, Nengwen Zhao, Zeyan Li 0001, Jiahao Bu, Zhihan Li 0002, Ying Liu 0024, Youjian Zhao, Dan Pei, Zhaogang Wang, Honglin Qiao |
WWW | 4 |
| 2016 | Control-fluidic CoDesign for paper-based digital microfluidic biochipsabstractPaper-based digital microfluidic biochips (P-DMFBs) have recently emerged as a promising low-cost and fast-responsive platform for biochemical assays. In P-DMFBs, electrodes and control lines are printed on a piece of photo paper using inkjet printer and conductive ink of carbon nanotubes (CNTs). Compared with traditional digital microfluidic biochips (DMFBs), P-DMFBs enjoy notable advantages, such as faster in-place fabrication with printer and ink, lower costs, better disposability, etc. Because electrodes and CNT control lines are printed on the same side of a paper, a new design challenge for P-DMFB is to prevent the interference between moving droplets and the voltages on CNT control lines. These interactions may result in unexpected droplet movements and thus incorrect assay outputs. To address the new challenges in automated design of P-DMFBs, this paper proposes the first control-fluidic codesign flow, which simultaneously adjusts the control line routing and fluidic droplet scheduling to achieve an optimized solution. As the control line routing may not be able to address all the interferences between moving droplets and the voltages on control lines, droplet rescheduling is performed to effectively deal with the remaining interferences in the routing solution. Computational simulation results on real-life bioassays show that the proposed codesign method successfully eliminates all the interferences, while a state-of-the-art maze routing method cannot solve any of the benchmarks without conflicts. Qin Wang 0005, Zeyan Li 0001, Haena Cheong, Oh-Sun Kwon, Hailong Yao 0002, Tsung-Yi Ho, Kwanwoo Shin, Bing Li 0005, Ulf Schlichtmann, Yici Cai |
ICCAD | 2 |