VLDB 2026 Research / reviewers in the wild / expert
Kaixin Sui
dblp:167/3797
· DBLP profile ↗
35ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0003-4545-7621ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 14 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 11 · 6 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Systems, architecture and hardware · 3 · 1 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RPM-MCTS: Knowledge-Retrieval as Process Reward Model with Monte Carlo Tree Search for Code GenerationabstractTree search-based methods have made significant progress in enhancing the code generation capabilities of large language models. However, due to the difficulty in effectively evaluating intermediate algorithmic steps and the inability to locate and timely correct erroneous steps, these methods often generate incorrect code and incur increased computational costs. To tackle these problems, we propose RPM-MCTS, an effective method that utilizes Knowledge-Retrieval as Process Reward Model based on Monte Carlo Tree Search to evaluate intermediate algorithmic steps. By utilizing knowledge base retrieval, RPM-MCTS avoids the complex training of process reward models. During the expansion phase, similarity filtering is employed to remove redundant nodes, ensuring diversity in reasoning paths. Furthermore, our method utilizes sandbox execution feedback to locate erroneous algorithmic steps during generation, enabling timely and targeted corrections. Extensive experiments on four public code generation benchmarks demonstrate that RPM-MCTS outperforms current state-of-the-art methods while achieving an approximately 15% reduction in token consumption. Furthermore, full fine-tuning of the base model using the data constructed by RPM-MCTS significantly enhances its code capabilities. Xiangyu Ouyang, Kaixin Sui |
AAAI | 4 |
| 2025 | LogSage: An LLM-Based Framework for CI/CD Failure Detection and Remediation with Industrial ValidationabstractContinuous Integration and Deployment (CI/CD) pipelines are critical to modern software engineering, yet diagnosing and resolving their failures remains complex and labor-intensive. We present LogSage, the first end-to-end LLM-powered framework for root cause analysis (RCA) and automated remediation of CI/CD failures. LogSage employs a token-efficient log preprocessing pipeline to filter noise and extract critical errors, then performs structured diagnostic prompting for accurate RCA. For solution generation, it leverages retrieval-augmented generation (RAG) to reuse historical fixes and invokes automation fixes via LLM tool-calling.On a newly curated benchmark of 367 GitHub CI/CD failures, LogSage achieves over 98% precision, near-perfect recall, and an F1 improvement of more than 38% points in the RCA stage, compared with recent LLM-based baselines. In a yearlong industrial deployment at ByteDance, it processed over 1.07M executions, with end-to-end precision exceeding 80%. These results demonstrate that LogSage provides a scalable and practical solution for automating CI/CD failure management in real-world DevOps workflows. Weiyuan Xu, Juntao Luo, Kaixin Sui, Qijun Ma, Isami Akasaka, Xiaoxue Shi |
ASE | 4 |
| 2024 | A survey on intelligent management of alerts and incidents in IT services
Qingyang Yu, Nengwen Zhao, Mingjie Li 0005, Zeyan Li 0001, Honglin Wang, Wenchi Zhang, Kaixin Sui, Dan Pei |
J. Netw. Comput. Appl. | 7 |
| 2023 | Multi-stage Location for Root-Cause Metrics in Online Service SystemsabstractThe failure of the online service system will seriously affect the user experience and bring huge economic losses. Therefore, the operators usually monitor service-level metrics and machine-level metrics to help quickly find failures, locate root-cause metrics, and reduce MTTR(mean time to repair). Many methods have emerged in recent years to automatically locate root-cause metrics. However, the existing methods cannot meet the requirements of efficiency, accuracy, and ease of deployment at the same time, and are difficult to use in practice. To overcome their limitations, we propose MetricMiner- a multi-stage location method for root-cause metrics in online service systems. Our approach is based on a key observation from numerous real-world cases: root-cause metrics tend to be unique in both the time dimension and the machine dimension. Therefore, we divide the root-cause metrics localization into three stages: first, quickly filter out normal metrics with limited historical data; second, obtain sufficient historical data to eliminate abnormal metrics; finally, according to the clustering of abnormal metrics between machines to sort and locate root-cause metrics. Experimental results on two real-world datasets with 194 cases show that our method can significantly outperform the state-of-the-art methods. Moreover, MetricMiner has been deployed to multiple banking services for more than six months, and we also shared some lessons learned from real deployment. Wenchi Zhang, Shize Zhang, Kaixin Sui, Enhuan Dong, Jiahai Yang 0001 |
NOMS | 5 |
| 2023 | Generic and robust root cause localization for multi-dimensional data in online service systems
Zeyan Li 0001, Junjie Chen 0003, Yiwei Zhao 0001, Yongqian Sun, Kaixin Sui, Xiping Wang, Dan Pei |
J. Syst. Softw. | 7 |
| 2022 | Generic and Robust Performance Diagnosis via Causal Inference for OLTP Database SystemsabstractOnline transaction processing (OLTP) database systems provide an effective solution to data support for online applications with high concurrency and low latency. An interruption or performance degradation of OLTP database systems may impact the availability of services and bring substantial economic loss. Thus, diagnosing the issue timely and mitigating it rapidly are essential for database administrators (DBAs). However, performance diagnosis for database systems is challenging due to numerous abnormal metrics, complex failure propagation, and high-performance requirements. Existing works relying on anomaly detection or causal graph construction cannot handle all these challenges simultaneously. In this paper, we propose an unsupervised learning-based method, CauseRank, to perform root cause localization with superior efficiency, high accuracy, and good interpretability. Two key techniques in CauseRank are a novel causal discovery algorithm named Group-based Greedy Equivalent Search (G-GES) incorporated with domain knowledge which treats metric groups as nodes to capture failure propagation and a simple yet effective ranking method named Causal Oriented Personalized PageRank (COPP). Extensive experiments on 97 real-world failure cases collected from a large-scale Oracle database demonstrate the effectiveness of CauseRank, achieving 82.5% top-3 accuracy and 93.8% top-5 accuracy and outperforming baseline approaches. The core idea and framework of CauseRank are generic and can be applied to other large-scale system components. Xianglin Lu, Zhe Xie, Zeyan Li 0001, Mingjie Li 0005, Xiaohui Nie, Nengwen Zhao, Qingyang Yu, Shenglin Zhang, Kaixin Sui, Dan Pei |
CCGRID | 9 |
| 2022 | Causal Inference-Based Root Cause Analysis for Online Service Systems with Intervention RecognitionabstractFault diagnosis is critical in many domains, as faults may lead to safety threats or economic losses. In the field of online service systems, operators rely on enormous monitoring data to detect and mitigate failures. Quickly recognizing a small set of root cause indicators for the underlying fault can save much time for failure mitigation. In this paper, we formulate the root cause analysis problem as a new causal inference task namedintervention recognition. We proposed a novel unsupervised causal inference-based method namedCausal Inference-based Root Cause Analysis (CIRCA). The core idea is a sufficient condition for a monitoring variable to be a root cause indicator,i.e., the change of probability distribution conditioned on the parents in the Causal Bayesian Network (CBN). Towards the application in online service systems, CIRCA constructs a graph among monitoring metrics based on the knowledge of system architecture and a set of causal assumptions. The simulation study illustrates the theoretical reliability of CIRCA. The performance on a real-world dataset further shows that CIRCA can improve the recall of the top-1 recommendation by 25% over the best baseline method. Mingjie Li 0005, Zeyan Li 0001, Kanglin Yin, Xiaohui Nie, Wenchi Zhang, Kaixin Sui, Dan Pei |
KDD | 6 |
| 2022 | Actionable and interpretable fault localization for recurring failures in online service systemsabstractFault localization is challenging in an online service system due to its monitoring data's large volume and variety and complex dependencies across/within its components (e.g., services or databases). Furthermore, engineers require fault localization solutions to be actionable and interpretable, which existing research approaches cannot satisfy. Therefore, the common industry practice is that, for a specific online service system, its experienced engineers focus on localization for recurring failures based on the knowledge accumulated about the system and historical failures. More specifically, 1) they can identify the underlying root causes and take mitigation actions when pinpointing a group of indicative metrics on the faulty component; 2) their diagnosis knowledge is roughly based on how one failure might affect the components in the whole system. Zeyan Li 0001, Nengwen Zhao, Mingjie Li 0005, Xianglin Lu, Dongdong Chang, Xiaohui Nie, Wenchi Zhang, Kaixin Sui, Guoqiang Duan, Dan Pei |
ESEC/SIGSOFT FSE | 10 |
| 2021 | Identifying Root-Cause Metrics for Incident Diagnosis in Online Service SystemsabstractIncidents in online service systems could incur poor user experience and tremendous economic loss. To reduce the influence of incidents and guarantee service reliability, it is critical to identify root-cause metrics for engineers with clues to assist incident diagnosis. However, it is a challenging task due to the complicated dependencies and huge volume of various metrics in large-scale systems. Existing approaches are based on either anomaly detection or correlation analysis, performing not well in terms of accuracy or efficiency. To better understand the problem of root-cause metric identification, we conduct a preliminary study based on real-world data analysis and interactions with engineers. The key observation is that root-cause metrics should satisfy two requirements. One is that the metric is expected to behave abnormally during the incident; the other is that the anomaly pattern should meet physical meaning and engineers' demand. Motivated by the findings obtained from the study, we propose an effective approach named PatternMatcher to identifying root-cause metrics accurately. Specifically, PatternMatcher contains three steps, where coarse-grained anomaly detection aiming to filter out normal metrics, anomaly pattern classification aiming to filter out unimportant anomaly patterns, and root-cause metric ranking. An extensive study on four real-world datasets including 113 incident cases from a large commercial bank demonstrates that PatternMatcher outperforms all baseline approaches, achieving top-3 average accuracy of 0.91. Moreover, we have deployed PatternMatcher in practice and shared some successful cases from real deployment. Canhua Wu, Nengwen Zhao, Xiaoqin Yang, ShiNing Li, Xidao Wen, Xiaohui Nie, Wenchi Zhang, Kaixin Sui, Dan Pei |
ISSRE | 11 |
| 2021 | Practical Root Cause Localization for Microservice Systems via Trace AnalysisabstractMicroservice architecture is applied by an increasing number of systems because of its benefits on delivery, scalability, and autonomy. It is essential but challenging to localize root-cause microservices promptly when a fault occurs. Traces are helpful for root-cause microservice localization, and thus many recent approaches utilize them. However, these approaches are less practical due to relying on supervision or other unrealistic assumptions. To overcome their limitations, we propose a more practical root-cause microservice localization approach named TraceRCA. The key insight of TraceRCA is that a microservice with more abnormal and less normal traces passing through it is more likely to be the root cause. Based on it, TraceRCA is composed of trace anomaly detection, suspicious microservice set mining and microservice ranking. We conducted experiments on hundreds of injected faults in a widely-used open-source microservice benchmark and a production system. The results show that TraceRCA is effective in various situations. The top-1 accuracy of TraceRCA outperforms the state-of-the-art unsupervised approaches by 44.8%. Besides, TraceRCA is applied in a large commercial bank, and it helps operators localize root causes for real-world faults accurately and efficiently. We also share some lessons learned from our real-world deployment. Zeyan Li 0001, Junjie Chen 0003, Nengwen Zhao, Shuwei Zhang, Long Jiang, Leiqin Yan, Zikai Wang 0010, Zhekang Chen, Wenchi Zhang, Xiaohui Nie, Kaixin Sui, Dan Pei |
IWQoS | 14 |
| 2021 | Identifying bad software changes via multimodal anomaly detection for online service systemsabstractIn large-scale online service systems, software changes are inevitable and frequent. Due to importing new code or configurations, changes are likely to incur incidents and destroy user experience. Thus it is essential for engineers to identify bad software changes, so as to reduce the influence of incidents and improve system re- liability. To better understand bad software changes, we perform the first empirical study based on large-scale real-world data from a large commercial bank. Our quantitative analyses indicate that about 50.4% of incidents are caused by bad changes, mainly be- cause of code defect, configuration error, resource contention, and software version. Besides, our qualitative analyses show that the current practice of detecting bad software changes performs not well to handle heterogeneous multi-source data involved in soft- ware changes. Based on the findings and motivation obtained from the empirical study, we propose a novel approach named SCWarn aiming to identify bad changes and produce interpretable alerts accurately and timely. The key idea of SCWarn is drawing support from multimodal learning to identify anomalies from heterogeneous multi-source data. An extensive study on two datasets with various bad software changes demonstrates our approach significantly outperforms all the compared approaches, achieving 0.95 F1-score on average and reducing MTTD (mean time to detect) by 20.4%∼60.7%. In particular, we shared some success stories and lessons learned from the practical usage. Nengwen Zhao, Junjie Chen 0003, Zhaoyang Yu 0002, Honglin Wang, Jiesong Li, Bin Qiu, Hongyu Xu, Wenchi Zhang, Kaixin Sui, Dan Pei |
ESEC/SIGSOFT FSE | 9 |
| 2021 | An empirical investigation of practical log anomaly detection for online service systemsabstractLog data is an essential and valuable resource of online service systems, which records detailed information of system running status and user behavior. Log anomaly detection is vital for service reliability engineering, which has been extensively studied. However, we find that existing approaches suffer from several limitations when deploying them into practice, including 1) inability to deal with various logs and complex log abnormal patterns; 2) poor interpretability; 3) lack of domain knowledge. To help understand these practical challenges and investigate the practical performance of existing work quantitatively, we conduct the first empirical study and an experimental study based on large-scale real-world data. We find that logs with rich information indeed exhibit diverse abnormal patterns (e.g., keywords, template count, template sequence, variable value, and variable distribution). However, existing approaches fail to tackle such complex abnormal patterns, producing unsatisfactory performance. Motivated by obtained findings, we propose a generic log anomaly detection system named LogAD based on ensemble learning, which integrates multiple anomaly detection approaches and domain knowledge, so as to handle complex situations in practice. About the effectiveness of LogAD, the average F1-score achieves 0.83, outperforming all baselines. Besides, we also share some success cases and lessons learned during our study. To our best knowledge, we are the first to investigate practical log anomaly detection in the real world deeply. Our work is helpful for practitioners and researchers to apply log anomaly detection to practice to enhance service reliability. Nengwen Zhao, Honglin Wang, Zeyan Li 0001, Zhu Pan, Xidao Wen, Wenchi Zhang, Kaixin Sui, Dan Pei |
ESEC/SIGSOFT FSE | 11 |
| 2020 | Practical and White-Box Anomaly Detection through Unsupervised and Active LearningabstractTo ensure quality of service and user experience, large Internet companies often monitor various Key Performance Indicators (KPIs) of their systems so that they can detect anomalies and identify failure in real time. However, due to a large number of various KPIs and the lack of high-quality labels, existing KPI anomaly detection approaches either perform well only on certain types of KPIs or consume excessive resources. Therefore, to realize generic and practical KPI anomaly detection in the real world, we propose a KPI anomaly detection framework named iRRCF-Active, which contains an unsupervised and white-box anomaly detector based on Robust Random Cut Forest (RRCF), and an active learning component. Specifically, we novelly propose an improved RRCF (iRRCF) algorithm to overcome the drawbacks of applying original RRCF in KPI anomaly detection. Besides, we also incorporate the idea of active learning to make our model benefit from high-quality labels given by experienced operators. We conduct extensive experiments on a large-scale public dataset and a private dataset collected from a large commercial bank. The experimental resulta demonstrate that iRRCF-Active performs better than existing traditional statistical methods, unsupervised learning methods and supervised learning methods. Besides, each component in iRRCF-Active has also been demonstrated to be effective and indispensable. Zhaowei Wang 0003, Zejun Xie, Nengwen Zhao, Junjie Chen 0003, Wenchi Zhang, Kaixin Sui, Dan Pei |
ICCCN | 7 |
| 2020 | Root-Cause Metric Location for Microservice Systems via Log Anomaly DetectionabstractMicroservice systems are typically fragile and failures are inevitable in them due to their complexity and large scale. However, it is challenging to localize the root-cause metric due to its complicated dependencies and the huge number of various metrics. Existing methods are based on either correlation between metrics or correlation between metrics and failures. All of them ignore the key data source in microservice, i.e., logs. In this paper, we propose a novel root-cause metric localization approach by incorporating log anomaly detection. Our approach is based on a key observation, the value of root-cause metric should be changed along with the change of the log anomaly score of the system caused by the failure. Specifically, our approach includes two components, collecting anomaly scores by log anomaly detection algorithm and identifying root-cause metric by robust correlation analysis with data augmentation. Experiments on an open-source benchmark microservice system have demonstrated our approach can identify root-cause metrics more accurately than existing methods and only require a short localization time. Therefore, our approach can assist engineers to save much effort in diagnosing and mitigating failures as soon as possible. Lingzhi Wang 0002, Nengwen Zhao, Junjie Chen 0003, Pinnong Li, Wenchi Zhang, Kaixin Sui |
ICWS | 6 |
| 2020 | Automatically and Adaptively Identifying Severe Alerts for Online Service SystemsabstractIn large-scale online service system, to enhance the quality of services, engineers need to collect various monitoring data and write many rules to trigger alerts. However, the number of alerts is way more than what on-call engineers can properly investigate. Thus, in practice, alerts are classified into several priority levels using manual rules, and on-call engineers primarily focus on handling the alerts with the highest priority level (i.e., severe alerts). Unfortunately, due to the complex and dynamic nature of the online services, this rule-based approach results in missed severe alerts or wasted troubleshooting time on non-severe alerts. In this paper, we propose AlertRank, an automatic and adaptive framework for identifying severe alerts. Specifically, AlertRank extracts a set of powerful and interpretable features (textual and temporal alert features, univariate and multivariate anomaly features for monitoring metrics), adopts XGBoost ranking algorithm to identify the severe alerts out of all incoming alerts, and uses novel methods to obtain labels for both training and testing. Experiments on the datasets from a top global commercial bank demonstrate that AlertRank is effective and achieves the F1-score of 0.89 on average, outperforming all baselines. The feedback from practice shows AlertRank can significantly save the manual efforts for on-call engineers. Nengwen Zhao, Panshi Jin, Xiaoqin Yang, Wenchi Zhang, Kaixin Sui, Dan Pei |
INFOCOM | 7 |
| 2020 | Real-time incident prediction for online service systemsabstractIncidents in online service systems could dramatically degrade system availability and destroy user experience. To guarantee service quality and reduce economic loss, it is essential to predict the occurrence of incidents in advance so that engineers can take some proactive actions to prevent them. In this work, we propose an effective and interpretable incident prediction approach, called eWarn, which utilizes historical data to forecast whether an incident will happen in the near future based on alert data in real time. More specifically, eWarn first extracts a set of effective features (including textual features and statistical features) to represent omen alert patterns via careful feature engineering. To reduce the influence of noisy alerts (that are not relevant to the occurrence of incidents), eWarn then incorporates the multi-instance learning formulation. Finally, eWarn builds a classification model via machine learning and generates an interpretable report about the prediction result via a state-of-the-art explanation technique (i.e., LIME). In this way, an early warning signal along with its interpretable report can be sent to engineers to facilitate their understanding and handling for the incoming incident. An extensive study on 11 real-world online service systems from a large commercial bank demonstrates the effectiveness of eWarn, outperforming state-of-the-art alert-based incident prediction approaches and the practice of incident prediction with alerts. In particular, we have applied eWarn to two large commercial banks in practice and shared some success stories and lessons learned from real deployment. Nengwen Zhao, Junjie Chen 0003, Xiaohui Nie, Wenchi Zhang, Kaixin Sui, Dan Pei |
ESEC/SIGSOFT FSE | 11 |
| 2019 | Neural Feature Search: A Neural Architecture for Automated Feature EngineeringabstractFeature engineering is a crucial step for developing effective machine learning models. Traditionally, feature engineering is performed manually, which requires much domain knowledge and is time-consuming. In recent years, many automated feature engineering methods have been proposed. These methods improve the accuracy of a machine learning model by automatically transforming the original features into a set of new features. However, existing methods either lack ability to perform high-order transformations or suffer from the feature space explosion problem. In this paper, we present Neural Feature Search (NFS), a novel neural architecture for automated feature engineering. We utilize a recurrent neural network based controller to transform each raw feature through a series of transformation functions. The controller is trained through reinforcement learning to maximize the expected performance of the machine learning algorithm. Extensive experiments on public datasets illustrate that our neural architecture is effective and outperforms the existing state-of-the-art automated feature engineering methods. Our architecture can efficiently capture potentially valuable high-order transformations and mitigate the feature explosion problem. Xiangning Chen, Bo Qiao 0001, Wei Wu 0011, Murali Chintalapati, Dongmei Zhang 0001, Qingwei Lin, Chuan Luo 0002, Hongyu Zhang 0002, Yong Xu 0010, Yingnong Dang, Kaixin Sui, Xu Zhang 0024 |
ICDM | 13 |
| 2019 | Generic and Robust Localization of Multi-dimensional Root CausesabstractOperators of online software services periodically collect various measures with many attributes. When a measure becomes abnormal, indicating service problems such as reliability degrade, operators would like to rapidly and accurately localize the root cause attribute combinations within a huge multi-dimensional search space. Unfortunately, previous approaches are not generic or robust in that they all suffer from impractical root cause assumptions, handling only directly collected measures but not derived ones, handling only anomalies with signicant magnitudes but not those insignicant but important ones, requiring manual parameter ne-tuning, or being too slow. This paper proposes a generic and robust multi-dimensional root cause localization approach, Squeeze, that overcomes all above limitations, the first in the literature. Through our novel bottom-up then top-down searching strategy and the techniques based on our proposed generalized ripple effect and generalized potential score, Squeeze is able to reach a good trade off between search speed and accuracy in a generic and robust manner. Case studies in several banks and an Internet company show that Squeeze can localize root causes much more rapidly and accurately than the traditional manual analysis. Furthermore, our extensive experiments on semi-synthetic datasets show that the F1-score of Squeeze outperforms previous approaches by 0.4 on average, while its localization time is only about 10 seconds. Zeyan Li 0001, Dan Pei, Yiwei Zhao 0001, Yongqian Sun, Kaixin Sui, Xiping Wang |
ISSRE | 6 |
| 2019 | FluxRank: A Widely-Deployable Framework to Automatically Localizing Root Cause Machines for Software Service Failure MitigationabstractThe failures of software service directly affect user experiences and service revenue. Thus operators monitor both service-level KPIs (e.g., response time) and machine-level KPIs (e.g., CPU usage) on each machine underlying the service. When a service fails, the operators must localize the root cause machines, and mitigate the failure as quickly as possible. Existing approaches have limited application due to the difficulty to obtain the required additional measurement data. As a result, failure localization is largely manual and very time-consuming. This paper presents FluxRank, a widely-deployable framework that can automatically and accurately localize the root cause machines, so that some actions can be triggered to mitigate the service failure. Our evaluation using historical cases from five real services (with tens of thousands of machines) of a top search company shows that the root cause machines are ranked top 1 (top 3) for 55 (66) cases out of 70 cases. Comparing to existing approaches, FluxRank cuts the localization time by more than 80% on average. FluxRank has been deployed online at one Internet service and six banking services for three months, and correctly localized the root cause machines as the top 1 for 55 cases out of 59 cases. Xiaohui Nie, Jing Zhu 0007, Shenglin Zhang, Kaixin Sui, Dan Pei |
ISSRE | 6 |
| 2019 | Cross-dataset Time Series Anomaly Detection for Cloud Systems
Xu Zhang 0024, Qingwei Lin, Yong Xu 0010, Si Qin, Hongyu Zhang 0002, Bo Qiao 0001, Yingnong Dang, Xinsheng Yang, Murali Chintalapati, Youjiang Wu, Ken Hsieh, Kaixin Sui, Yaohai Xu, Wenchi Zhang, Furao Shen, Dongmei Zhang 0001 |
USENIX ATC | 13 |
| 2019 | Dynamic TCP Initial Windows and Congestion Control Schemes Through Reinforcement LearningabstractDespite many years of improvements to it, TCP still suffers from an unsatisfactory performance. For services dominated by short flows (e.g., web search and e-commerce), TCP suffers from the flow startup problem and cannot fully utilize the available bandwidth in the modern Internet: TCP starts from a conservative and static initial window (IW, 2-4 or 10), while most of the web flows are too short to converge to the best sending rate before the session ends. For services dominated by long flows (e.g., video streaming and file downloading), the congestion control (CC) scheme manually and statically configured might not offer the best performance for the latest network conditions. To address these two challenges, we propose TCP-RL, which uses reinforcement learning (RL) techniques to dynamically configure IW and CC in order to improve the performance of TCP flow transmission. Basing on the latest network conditions observed at the server side of a web service, TCP-RL dynamically configures a suitable IW for short flows through group-based RL, and dynamically configures a suitable CC scheme for long flows through deep RL. Our extensive experiments show that for short flows, TCP-RL can reduce the average transmission time by about 23%; and for long flows, compared with the performance of 14 CC schemes, TCP-RL's performance ranks top 5 for about 85% of the 288 given static network conditions, whereas for about 90% of conditions, its performance drops by less than 12% compared with that of the best-performing CC schemes for the same network conditions. Xiaohui Nie, Youjian Zhao, Zhihan Li 0002, Guo Chen 0001, Kaixin Sui, Zijie Ye, Dan Pei |
IEEE J. Sel. Areas Commun. | 5 |
| 2018 | Reducing Web Latency Through Dynamically Setting TCP Initial Window with Reinforcement LearningabstractLatency, which directly affects the user experience and revenue of web services, is far from ideal in reality, due to the well-known TCP flow startup problem. Specifically, since TCP starts from a conservative and static initial window (IW, 2~4 or 10), most of the web flows are too short to have enough time to find its best congestion window before the session ends. As a result, TCP cannot fully utilize the available bandwidth in the modern Internet. In this paper, we propose to use group-based reinforcement learning (RL) to enable a web server, through trial-and-error, to dynamically set a suitable IW for a web flow before its transmission starts. Our proposed system, SmartIW, collects TCP flow performance metrics (e.g., transmission time, loss rate, RTT) in real-time without any client assistance. Then these metrics are aggregated into groups with similar features (subnet, ISP, province, etc.) to satisfy RL's requirement. SmartIW has been deployed in one of the top global search engines for more than a year. Our online and testbed experiments show that, compared to the common practice of IW=10, SmartIW can reduce the average transmission time by 23% to 29%. Xiaohui Nie, Youjian Zhao, Dan Pei, Guo Chen 0001, Kaixin Sui |
IWQoS | 5 |
| 2018 | BigIN4: Instant, Interactive Insight Identification for Multi-Dimensional Big DataabstractThe ability to identify insights from multi-dimensional big data is important for business intelligence. To enable interactive identification of insights, a large number of dimension combinations need to be searched and a series of aggregation queries need to be quickly answered. The existing approaches answer interactive queries on big data through data cubes or approximate query processing. However, these approaches can hardly satisfy the performance or accuracy requirements for ad-hoc queries demanded by interactive exploration. In this paper, we present BigIN4, a system for instant, interactive identification of insights from multi-dimensional big data. BigIN4 gives insight suggestions by enumerating subspaces and answers queries by combining data cube and approximate query processing techniques. If a query cannot be answered by the cubes, BigIN4 decomposes it into several low dimensional queries that can be directly answered by the cubes through an online constructed Bayesian Network and gives an approximate answer within a statistical interval. Unlike the related works, BigIN4 does not require any prior knowledge of queries and does not assume a certain data distribution. Our experiments on ten real-world large-scale datasets show that BigIN4 can successfully identify insights from big data. Furthermore, BigIN4 can provide approximate answers to aggregation queries effectively (with less than 10% error on average) and efficiently (50x faster than sampling-based methods). Qingwei Lin, Weichen Ke, Jian-Guang Lou, Hongyu Zhang 0002, Kaixin Sui, Yong Xu 0010, Bo Qiao 0001, Dongmei Zhang 0001 |
KDD | 5 |
| 2018 | Predicting Node failure in cloud service systemsabstractIn recent years, many traditional software systems have migrated to cloud computing platforms and are provided as online services. The service quality matters because system failures could seriously affect business and user experience. A cloud service system typically contains a large number of computing nodes. In reality, nodes may fail and affect service availability. In this paper, we propose a failure prediction technique, which can predict the failure-proneness of a node in a cloud service system based on historical data, before node failure actually happens. The ability to predict faulty nodes enables the allocation and migration of virtual machines to the healthy nodes, therefore improving service availability. Predicting node failure in cloud service systems is challenging, because a node failure could be caused by a variety of reasons and reflected by many temporal and spatial signals. Furthermore, the failure data is highly imbalanced. To tackle these challenges, we propose MING, a novel technique that combines: 1) a LSTM model to incorporate the temporal data, 2) a Random Forest model to incorporate spatial data; 3) a ranking model that embeds the intermediate results of the two models as feature inputs and ranks the nodes by their failure-proneness, 4) a cost-sensitive function to identify the optimal threshold for selecting the faulty nodes. We evaluate our approach using real-world data collected from a cloud service system. The results confirm the effectiveness of the proposed approach. We have also successfully applied the proposed approach in real industrial practice. Qingwei Lin, Ken Hsieh, Yingnong Dang, Hongyu Zhang 0002, Kaixin Sui, Yong Xu 0010, Jian-Guang Lou, Chenggang Li, Youjiang Wu, Randolph Yao, Murali Chintalapati, Dongmei Zhang 0001 |
ESEC/SIGSOFT FSE | 5 |
| 2018 | Improving Service Availability of Cloud Systems by Predicting Disk Error
Yong Xu 0010, Kaixin Sui, Randolph Yao, Hongyu Zhang 0002, Qingwei Lin, Yingnong Dang, Peng Li 0062, Keceng Jiang, Wenchi Zhang, Jian-Guang Lou, Murali Chintalapati, Dongmei Zhang 0001 |
USENIX ATC | 2 |
| 2017 | TCP WISE: One initial congestion window is not enoughabstractCurrent TCP is very inefficient for web services. Web transactions are often very short-lived. TCP flow starts with a conservative initial congestion window (IW), which causes multiple round-trip times to finish the transmission even if the end-to-end bandwidth is sufficient for the transaction to be finished in one round-trip time. Previous research efforts have been focusing on finding the overall best IW for the entire Internet or a service company. However, we observe that one-IW-fits-all is suboptimal after one year of online measurement in Baidu, one of the top global search engine companies. To reduce the TCP latency, we propose TCP WISE, which dynamically assigns suitable IWs for different user cluster at different times on the server side. The values of users' IWs are proactively learned based on the historical experience on the server-side. Our testbed experiments show that our learning algorithm can handle the network changes and converge to the best IW. We have deployed TCP WISE in one of Baidu's production data center, and results shows that the 80thpercentile latency of the HTTP responses has been reduced by 10.4% compared with current TCP with a fixed IW of 10. Xiaohui Nie, Youjian Zhao, Guo Chen 0001, Kaixin Sui, Yazheng Chen, Dan Pei |
IPCCC | 4 |
| 2017 | You can hide, but your periodic schedule can'tabstractThe enterprise Wi-Fi networks enable the collection of large-scale users' trajectory datasets, which are highly desired for both research and commercial purposes. Meanwhile, releasing these mobility data also raises serious privacy concerns. A large body of work tries to achieve k-anonymity as the first step to solve the privacy problem and it has been qualitatively recognized that k-anonymity is still risky when the diversity of sensitive information in the k-anonymity set is low. However, there lacks a study that provides a quantitative understanding for trajectory data. In this work, we investigate the schedule-leakage risk for the first time, by presenting a large-scale measurement based analysis of the high schedule-leakage risk over sixteen weeks of trajectory data collected from Tsinghua University, a campus with 2,670 access points deployed in 111 buildings. Using this dataset, we recognize the high risk of the schedule-leakage, i.e., even when 4-anonymity is satisfied, 28% of individuals' schedules are totally disclosed, and 56% are partly disclosed. Minghua Ma, Kaixin Sui, Yong Li 0008, Dan Pei |
IWQoS | 3 |
| 2016 | EDUM: classroom education measurements via large-scale WiFi networksabstractBehavior in classroom-based courses is hard to measure at large-scale. In this paper, we propose the EDUM (EDUcation Measurement) system to help characterize educational behavior through data collected from WLANs (WiFi networks) on campuses. EDUM characterizes students' punctuality (attendances, late arrivals, and early departures) for lectures using longitudinal WLAN data, and further characterizes the attractiveness of lectures using mobile phone's interactive states at minute-scale granularity. EDUM is easy to deploy and extensible for new types of data. We deploy EDUM at Tsinghua University where ~700 volunteer students' data are measured during a 9-week period by ~2,800 APs and two popular mobile apps. Our results show that EDUM makes it possible to obtain large-scale observations on punctuality, distraction and study performance, and quantitatively confirm or disprove numerous assumptions about educational behavior. Mengyu Zhou, Minghua Ma, Yangkun Zhang, Kaixin Sui, Dan Pei, Thomas Moscibroda |
UbiComp | 4 |
| 2016 | FOCUS: Shedding light on the high search response time in the wildabstractResponse time plays a key role in Web services, as it significantly impacts user engagement, and consequently the Web providers' revenue. Using a large search engine as a case study, we propose a machine learning based analysis framework, called FOCUS, as the first step to automatically debug high search response time (HSRT) in search logs. The output of FOCUS offers a promising starting point for operators' further investigation. FOCUS has been deployed in one of the largest search engines for 2.5 months and analyzed about one billion search logs. Compared with a previous approach, FOCUS generates 90% less items for investigation and achieves both higher recall and higher precision. The results of FOCUS enable us to make several interesting observations. For example, we find that popular queries are more image-intensive (e.g., TV series and shopping), but they have relatively low SRT because they are cached well by servers. Additionally, as suggested by the first-month analysis results of FOCUS, we conduct an optimization on image transmission time. A one-month real-world deployment shows that we successfully reduce the 80th percentile of search response time by 253ms, and reduce the fraction of HSRT by one third. Youjian Zhao, Kaixin Sui, Dan Pei, Qingqian Tao, Xiyang Chen, Dai Tan |
INFOCOM | 3 |
| 2016 | Mining causality graph for automatic web-based service diagnosisabstractIt is crucial for Internet company to provide highly reliable web-based services. The web-based services always have many components running in the large-scale infrastructure with complex interactions. As an indispensable part of high reliability, the diagnosis remains to be a thorny problem. With the growth of system scale and complexity, it becomes even more difficult. In this paper, we propose an automatic diagnosis system based on causality graph to help system operators find the root causes. The causality graph is mainly extracted from the historical data of the monitoring system, and the method consists of four steps. 1) It utilizes a data mining method to extract the initial causality graph. 2) Once a failure happens, it lists top-k suspects with a ranking algorithm based on the causality graph. 3) Then system operators check the suspects and label them either right or wrong. 4) A supervised learning algorithm takes the labels as the input to tune the causality graph, in order to improve the diagnosis accuracy on step 2 iteratively. This method requires neither knowledge about the design and implementation details of the web-based service, nor instrumenting the services' source code. Our controlled experiments show that the root causes can be ranked in top 3 with 100% accuracy after countable learning iterations. Xiaohui Nie, Youjian Zhao, Kaixin Sui, Dan Pei, Xianping Qu |
IPCCC | 3 |
| 2016 | Your trajectory privacy can be breached even if you walk in groupsabstractThe enterprise Wi-Fi networks enable the collection of large-scale users' mobility information at an indoor level. The collected trajectory data is very valuable for both research and commercial purposes, but the use of the trajectory data also raises serious privacy concerns. A large body of work tries to achieve k-anonymity (hiding each user in an anonymity set no smaller than k) as the first step to solve the privacy problem. Yet it has been qualitatively recognized that k-anonymity is still risky when the diversity of the sensitive information in the k-anonymity set is low. There, however, still lacks a study that provides a quantitative understanding of that risk in the trajectory dataset. In this work, we present a large-scale measurement based analysis of the low-diversity risk over four weeks of trajectory data collected from Tsinghua, a campus that covers an area of 4 km2, on which 2,670 access points are deployed in 111 buildings. Using this dataset, we highlight the high risk of the low diversity. For example, we find that even when 5-anonymity is satisfied, the sensitive attributes of 25% of individuals can be easily guessed. We also find that although a larger k increases the size of anonymity sets, the corresponding improvement on the diversity of anonymity sets is very limited (decayed exponentially). These results suggest that diversity-oriented solutions are necessary. Kaixin Sui, Youjian Zhao, Minghua Ma, Zimu Li, Dan Pei |
IWQoS | 1 |
| 2016 | Understanding the Impact of AP Density on WiFi Performance Through Real-World Deploymentabstract802.11 (WiFi) networks have become increasingly important for our daily lives. However, previous work has shown that enterprise WiFi performance is often unsatisfactory and that over-utilization and interference from rogue APs are the two primary reasons. To address the above problem, this paper proposes to improve the capacity of WiFi infrastructures by increasing the enterprise AP deployment density, as well as disabling the wired Internet access in buildings to eliminate rogue APs and their interference. We deployed several WiFi networks with different AP density and vendors on Tsinghua campus. Based on the measurement results from our real-world deployments, we made three main observations: 1) in general, higher AP density improves WiFi performance; 2) over-dense deployment with unnecessarily high transmission power can worsen WiFi performance. 3) choice of AP vendors also has an impact on WiFi performance. Kaixin Sui, Yousef Azzabi, Xiaoping Zhang 0004, Youjian Zhao, Jilong Wang 0001, Zimu Li, Dan Pei |
LANMAN | 1 |
| 2016 | Characterizing and Improving WiFi Latency in Large-Scale Operational NetworksabstractWiFi latency is a key factor impacting the user experience of modern mobile applications, but it has not been well studied at large scale. In this paper, we design and deploy WiFiSeer, a framework to measure and characterize WiFi latency at large scale. WiFiSeer comprises a systematic methodology for modeling the complex relationships between WiFi latency and a diverse set of WiFi performance metrics, device characteristics, and environmental factors. WiFiSeer was deployed on Tsinghua campus to conduct a WiFi latency measurement study of unprecedented scale with more than 47,000 unique user devices. We observe that WiFi latency follows a long tail distribution and the 90th (99th) percentile is around 20 ms (250 ms). Furthermore, our measurement results quantitatively confirm some anecdotal perceptions about impacting factors and disapprove others. We deploy three practical solutions for improving WiFi latency in Tsinghua, and the results show significantly improved WiFi latencies. In particular, over 1,000 devices use our AP selection service based on a predictive WiFi latency model for 2.5 months, and 72% of their latencies are reduced by over half after they re-associate to the suggested APs. Kaixin Sui, Mengyu Zhou, Minghua Ma, Dan Pei, Youjian Zhao, Zimu Li, Thomas Moscibroda |
MobiSys | 1 |
| 2015 | How bad are the rogues' impact on enterprise 802.11 network performance?abstractEnterprise 802.11 Network (EWLAN) is an important infrastructure to the Mobile Internet, but its performance is being significantly impacted by the ever-increasing Rogue access points (RAPs). For example, in the university EWLAN we studied, the number of RAPs is more than seven times that of the enterprise APs. In this paper, we propose a generic methodology to measure RAP's carrier sense interference and hidden terminal interference, and it only uses readily available SNMP metrics, without any additional measurement hardware. Our results show that, on average, the carrier sense interference due to RAPs causes only 5% access delay increase at the MAC layer, because of careful engineering and software optimization. However, hidden terminal interference due to RAPs causes (a much more severe) up to 30% MAC layer loss rate increase on average, because no existing approach has explicitly dealt with the hidden terminal impact from rogue APs. Overall, the RAP interference would increase the IP layer delay at the WiFi hop by up to 50%. Kaixin Sui, Youjian Zhao, Dan Pei, Zimu Li |
INFOCOM | 1 |
| 2015 | Learning thresholds for PV change detection from operators' labelsabstractPage Views (PVs) are very crucial for search engines due to their close relationship to the revenue. When PVs change significantly, operators must be informed so that they can diagnose and fix the problem quickly, and prevent further loss. In reality, PVs can be counted in many ways (e.g., PVs originated from different ISPs), and different PVs are of different interest to operators (e.g., the PVs of a larger ISP is more important). As a result, different PVs often require different detection standards, or thresholds. However, attempts to tune a number of thresholds have been hampered by the cost of the manual effort involved. To address the above problem, we propose a practical framework, called PTL (practical threshold learning). Operators only need to provide a few simple labels about the detection results, then PTL will automatically tune the thresholds for different PVs. Using 4-month PVs from a global top search engine, our evaluation demonstrates that PTL can improve the accuracy of detection dramatically. More importantly, it introduces very little labeling overhead for operators. For example, when detecting the PVs of 103 ISPs, PTL can reduce the overall false negative rate from 96% to 9% using only 29 labels per week on average. Youjian Zhao, Kaixin Sui, Shiwen Cheng, Dan Pei, Chengbin Quan, Jiao Luo, Xiaowei Jing, Mei Feng |
IPCCC | 3 |