VLDB 2026 Research / reviewers in the wild / expert
Mingjie Li 0005
dblp:48/10103-5
· DBLP profile ↗
9ranked-venue papers
2as first author
8since 2021 · last 2024
0000-0002-4778-4098ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A survey on intelligent management of alerts and incidents in IT services
Qingyang Yu, Nengwen Zhao, Mingjie Li 0005, Zeyan Li 0001, Honglin Wang, Wenchi Zhang, Kaixin Sui, Dan Pei |
J. Netw. Comput. Appl. | 3 |
| 2024 | Diagnosing Performance Issues for Large-Scale Microservice Systems With Heterogeneous GraphabstractThe availability of microservice systems is critical to business operations and corporate reputation. However, the dynamics and complexity of microservice systems introduce significant challenges to the performance issue diagnosis of large-scale microservice systems. After investigating hundreds of real-world performance issue cases in Tencent, we find that previous troubleshooting approaches fail to accurately localize root causes because they overlook the inconsistency between causality and calling relationships. Therefore, we propose a novel approach, MicroDig, to diagnose performance issues for large-scale microservice systems. Specifically, MicroDig constructs a heterogeneous propagation graph to capture the causal relationships between calls and microservices. It then conducts a heterogeneity-oriented random walk (HORW) to pinpoint the culprit microservice. Extensive evaluation experiments have been conducted to evaluate MicroDig's performance on 60 real-world performance issues collected from Tencent, 80 manually injected ones collected from a widely used open-source microservice system and 128 performance issues collected from an e-commerce system used by a top-tier global commercial bank. MicroDig achieves 94.1%, 85.5% and 93.8% top-3 accuracy on the three datasets, respectively, significantly outperforming six popular baseline methods. Additionally, we have shared our success stories and learned lessons from the deployment of MicroDig in Tencent. Xianglin Lu, Shenglin Zhang, Jiaqi Luan, Yingke Li, Mingjie Li 0005, Zeyan Li 0001, Qingyang Yu, Hucheng Xie, Chenyuan Hu, Canqun Yang, Dan Pei |
IEEE Trans. Serv. Comput. | 6 |
| 2023 | An Empirical Analysis of Anomaly Detection Methods for Multivariate Time SeriesabstractUsing multivariate time series (MTS) data for anomaly detection is widely adopted in service systems, such as web services and financial businesses. Researchers have recently proposed some well-performed algorithms for MTS anomaly detection from different perspectives. When applied to the real world, we observe that none of the algorithms is adaptable to all scenarios due to the complex data and anomaly characteristics. Moreover, there is currently a lack of comprehensive analysis work of these algorithms to guide operators in selecting the appropriate one in practice. To bridge this gap, we conduct an empirical study using various real-world data to gain an in-depth understanding of state-of-the-art anomaly detection algorithms. First, we provide general recommendations to guide operators in selecting suitable models based on the volume of training data, computational resources, and effectiveness requirements. Then, we summarize the typical data characteristics and types of anomalies and offer tailored model selection suggestions for different data characteristics and anomaly types. At last, we apply the summarized model selection suggestions to all the datasets we collected. The results show that most of our suggestions can achieve better than any single algorithm alone, demonstrating the effectiveness and generalization of our recommendations. Dongwen Li, Shenglin Zhang, Yongqian Sun, Zeyu Che, Zhenyu Zhong, Minghan Liang, Minyi Shao, Mingjie Li 0005, Dan Pei |
ISSRE | 10 |
| 2023 | CMDiagnostor: An Ambiguity-Aware Root Cause Localization Approach Based on Call Metric DataabstractThe availability of online services is vital as its strong relevance to revenue and user experience. To ensure online services’ availability, quickly localizing the root causes of system failures is crucial. Given the high resource consumption of traces, call metric data are widely used by existing approaches to construct call graphs in practice. However, ambiguous correspondences between upstream and downstream calls may exist and result in exploring unexpected edges in the constructed call graph. Conducting root cause localization on this graph may lead to misjudgments of real root causes. To the best of our knowledge, we are the first to investigate such ambiguity, which is overlooked in the existing literature. Inspired by the law of large numbers and the Markov properties of network traffic, we propose a regression-based method (named AmSitor) to address this problem effectively. Based on AmSitor, we propose an ambiguity-aware root cause localization approach based on Call Metric Data named CMDiagnostor, containing metric anomaly detection, ambiguity-free call graph construction, root cause exploration, and candidate root cause ranking modules. The comprehensive experimental evaluations conducted on real-world datasets show that our CMDiagnostor can outperform the state-of-the-art approaches by 14% on the top-5 hit rate. Moreover, AmSitor can also be applied to existing baseline approaches separately to improve their performances one step further. The source code is released at https://github.com/NetManAIOps/CMDiagnostor. Qingyang Yu, Changhua Pei, Mingjie Li 0005, Zeyan Li 0001, Shenglin Zhang, Xianglin Lu, Jiaqi Li 0021, Dan Pei |
WWW | 4 |
| 2022 | Generic and Robust Performance Diagnosis via Causal Inference for OLTP Database SystemsabstractOnline transaction processing (OLTP) database systems provide an effective solution to data support for online applications with high concurrency and low latency. An interruption or performance degradation of OLTP database systems may impact the availability of services and bring substantial economic loss. Thus, diagnosing the issue timely and mitigating it rapidly are essential for database administrators (DBAs). However, performance diagnosis for database systems is challenging due to numerous abnormal metrics, complex failure propagation, and high-performance requirements. Existing works relying on anomaly detection or causal graph construction cannot handle all these challenges simultaneously. In this paper, we propose an unsupervised learning-based method, CauseRank, to perform root cause localization with superior efficiency, high accuracy, and good interpretability. Two key techniques in CauseRank are a novel causal discovery algorithm named Group-based Greedy Equivalent Search (G-GES) incorporated with domain knowledge which treats metric groups as nodes to capture failure propagation and a simple yet effective ranking method named Causal Oriented Personalized PageRank (COPP). Extensive experiments on 97 real-world failure cases collected from a large-scale Oracle database demonstrate the effectiveness of CauseRank, achieving 82.5% top-3 accuracy and 93.8% top-5 accuracy and outperforming baseline approaches. The core idea and framework of CauseRank are generic and can be applied to other large-scale system components. Xianglin Lu, Zhe Xie, Zeyan Li 0001, Mingjie Li 0005, Xiaohui Nie, Nengwen Zhao, Qingyang Yu, Shenglin Zhang, Kaixin Sui, Dan Pei |
CCGRID | 4 |
| 2022 | Mining Fluctuation Propagation Graph Among Time Series with Active Learning
Mingjie Li 0005, Minghua Ma, Xiaohui Nie, Kanglin Yin, Xidao Wen, Zhiyun Yuan, Duogang Wu, Guoying Li, Dan Pei |
DEXA (1) | 1 |
| 2022 | Causal Inference-Based Root Cause Analysis for Online Service Systems with Intervention RecognitionabstractFault diagnosis is critical in many domains, as faults may lead to safety threats or economic losses. In the field of online service systems, operators rely on enormous monitoring data to detect and mitigate failures. Quickly recognizing a small set of root cause indicators for the underlying fault can save much time for failure mitigation. In this paper, we formulate the root cause analysis problem as a new causal inference task namedintervention recognition. We proposed a novel unsupervised causal inference-based method namedCausal Inference-based Root Cause Analysis (CIRCA). The core idea is a sufficient condition for a monitoring variable to be a root cause indicator,i.e., the change of probability distribution conditioned on the parents in the Causal Bayesian Network (CBN). Towards the application in online service systems, CIRCA constructs a graph among monitoring metrics based on the knowledge of system architecture and a set of causal assumptions. The simulation study illustrates the theoretical reliability of CIRCA. The performance on a real-world dataset further shows that CIRCA can improve the recall of the top-1 recommendation by 25% over the best baseline method. Mingjie Li 0005, Zeyan Li 0001, Kanglin Yin, Xiaohui Nie, Wenchi Zhang, Kaixin Sui, Dan Pei |
KDD | 1 |
| 2022 | Actionable and interpretable fault localization for recurring failures in online service systemsabstractFault localization is challenging in an online service system due to its monitoring data's large volume and variety and complex dependencies across/within its components (e.g., services or databases). Furthermore, engineers require fault localization solutions to be actionable and interpretable, which existing research approaches cannot satisfy. Therefore, the common industry practice is that, for a specific online service system, its experienced engineers focus on localization for recurring failures based on the knowledge accumulated about the system and historical failures. More specifically, 1) they can identify the underlying root causes and take mitigation actions when pinpointing a group of indicative metrics on the faulty component; 2) their diagnosis knowledge is roughly based on how one failure might affect the components in the whole system. Zeyan Li 0001, Nengwen Zhao, Mingjie Li 0005, Xianglin Lu, Dongdong Chang, Xiaohui Nie, Wenchi Zhang, Kaixin Sui, Guoqiang Duan, Dan Pei |
ESEC/SIGSOFT FSE | 3 |
| 2018 | The DevOps Lab Platform for Managing Diversified Projects in Educating Agile Software EngineeringabstractThis Research Work-in-Progress paper presents the design of a Software Engineering (SE) course to support project-based practical training. Group projects, especially projects from industry partners, are deemed to be necessary for students to gain hands-on experiences. With projects from the real world, students learn not only practical engineering solutions, but also the context, constraints, and social aspects of SE. For a course having over 100 students with different interests and experiences, it is desired to provide diversified choices of projects to stimulate enthusiasm for learning. However, management and evaluation of diversified projects are challenging. Following the Agile principles, we need to continuously track progress and activities of each group, to provide quick feedback of deliveries, and to periodically evaluate students’ performance. Therefore, we built a DevOps platform based on GitLab version control and continuous integration framework. Commits to GitLab code repositories automatically trigger build, testing, and analysis functions (which provide both qualitative and quantitative feedback to the students). This system has been in operations since 2014 for an undergraduate SE course, with over 500 students participating in over 130 project teams in total. The preliminary research showed promising results in improving SE education. Xiaoying Bai, Dan Pei, Mingjie Li 0005 |
FIE | 3 |