Xiaohui Nie

dblp:119/6181 · DBLP profile ↗
← Back
27ranked-venue papers
5as first author
19since 2021 · last 2026
0000-0002-0371-854XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 11 · 1 first-author · 9 since 2021Computer networks · 10 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Systems, architecture and hardware · 3 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021
YearPublicationVenuePosition
2026 RCAgentBench: An Agent-Oriented Benchmark for Multimodal Root Cause Analysis in Microservices
Hengyue Jiang, Xiaohui Nie, Changhua Pei
IWQoS3
2026 SwiftShift: Accelerating QUIC Migration for Ultra-Low-Latency Interactive Media
Fangshuo Han, Dongbiao He, Xiaohui Nie, Yanbiao Li 0001
NOSSDAV5
2025 DeST: An Unsupervised Decoupled Spatio-Temporal Framework for Microservice Incident Management
abstract
Effective incident management in large-scale microservice systems demands both accurate anomaly detection (AD) and precise root cause localization (RCL) across heterogeneous data modalities. However, existing approaches often treat these tasks in isolation, resulting in redundant maintenance, delayed response, and the absence of shared diagnostic context. While recent efforts have explored unified frameworks to support both tasks, these approaches often suffer from high falsealarm rates due to cross-modal interference. To address these issues, we propose DeST, an unsupervised decoupled spatiotemporal framework that jointly performs anomaly detection and root cause localization. DeST proposes a multi-stage fusion strategy that decouples temporal and spatial feature learning to mitigate cross-modal interference and prevent cross-modal interference. Furthermore, it incorporates task-specific modal routing to direct learned representations to different tasks, enhancing both detection and localization accuracy. To ensure robustness against transient noise, DeST designs a Differential Multi-Scale Convolutional Network (DMCN) for noise-resistant temporal feature representation. We evaluate DeST on two real-world microservice benchmarks, where it achieves a perfect F1-score of $\mathbf{1. 0 0}$ for anomaly detection and outperforms existing methods in root cause localization accuracy. Ablation studies highlight the effectiveness of key components. Our unified framework reduces false alarms in anomaly detection and streamlines root cause localization, providing a robust and practical solution for microservice incident management.
Xiaohui Nie, Hang Cui 0004, Changhua Pei, Haotian Si, Ke Xiang, Yanbiao Li 0001, Gaogang Xie, Dan Pei
ISSRE1
2025 TrioXpert: An Automated Incident Management Framework for Microservice System
abstract
Automated incident management plays a pivotal role in large-scale microservice systems. However, many existing methods rely solely on single-modal data (e.g., metrics, logs, and traces) and struggle to simultaneously address multiple downstream tasks, including anomaly detection (AD), failure triage (FT), and root cause localization (RCL). Moreover, the lack of clear reasoning evidence in current techniques often leads to insufficient interpretability. To address these limitations, we propose TrioXpert, an end-to-end incident management framework capable of fully leveraging multimodal data. TrioXpert designs three independent data processing pipelines based on the inherent characteristics of different modalities, comprehensively characterizing the operational status of microservice systems from both numerical and textual dimensions. It employs a collaborative reasoning mechanism using large language models (LLMs) to simultaneously handle multiple tasks while providing clear reasoning evidence to ensure strong interpretability. We conducted extensive evaluations on two microservice system datasets, and the experimental results demonstrate that TrioXpert achieves outstanding performance in AD (improving by 4.7% to 57.7%), FT (improving by 2.1% to 40.6%), and RCL (improving by 1.6% to 163.1%) tasks. TrioXpert has also been deployed in Lenovo’s production environment, demonstrating substantial gains in diagnostic efficiency and accuracy.
Yongqian Sun, Yu Luo 0011, Xidao Wen, Yuan Yuan 0034, Xiaohui Nie, Shenglin Zhang
ASE5
2025 AIOpsArena: Scenario-Oriented Evaluation and Leaderboard for AIOps Algorithms in Microservices
abstract
AIOps algorithms playa crucial role in the mainte-nance of microservice systems. Many previous benchmarks' per-formance leaderboard provides valuable guidance for selecting appropriate algorithms. However, existing AIOps benchmarks mainly utilize offline static datasets to evaluate algorithms. They cannot consistently evaluate the performance of algorithms using real-time datasets, and the operation scenarios for evaluation are static, which is insufficient for effective algorithm selection. To address these issues, we propose an evaluation-consistent and scenario-oriented evaluation framework named AIOpsArena. The core idea is to build a live microservice benchmark to generate real-time datasets and consistently simulate the specific operation scenarios on it. AIOpsArena supports different leaderboards by selecting specific algorithms and datasets according to the operation scenarios. It also supports the deployment of various types of algorithms, enabling algorithms hot-plugging. At last, we test AIOpsArena with typical microservice operation scenarios to demonstrate its efficiency and usability. Platform and a video demonstrating the functioning of AIOpsArena is available from https://github.com/AIOpsArena/aiopsarena.
Yongqian Sun, Jiaju Wang, Zhengdan Li, Xiaohui Nie, Minghua Ma, Shenglin Zhang, Yuhe Ji, Wen Long, Hengmao Chen, Yongnan Luo, Dan Pei
SANER4
2025 A Comprehensive Benchmark and Empirical Study of Trace Anomaly Detection
abstract
The growing complexity of modern Internet applications and the widespread use of microservice architectures have amplified the need for efficient trace anomaly detection to maintain system stability. Despite the fact that many trace anomaly detection algorithms have been proposed to identify abnormal behaviors, a comprehensive evaluation of these methods is lacking, which makes it difficult for developers to choose the most suitable algorithm for real-world applications. To address this gap, we presentTADBench, a comprehensive and extensible benchmark for trace anomaly detection.TADBenchconsolidates diverse publicly available trace datasets and algorithms into a unified repository, standardizes data formats, and incorporates manual anomaly labels. To ensure reproducibility and fair comparisons, we propose a modular evaluation framework supporting end-to-end model assessment. Additionally, we provide practical guidance for algorithm selection based on specific data attributes by evaluating their performance across datasets with different characteristics, thereby effectively bridging the gap between academic research and industrial deployment. To the best of our knowledge, this is the first comprehensive empirical study of trace anomaly detection algorithms. Our findings aim to facilitate the adoption of these methods in production environments, offering actionable insights for developers and researchers.
Yongqian Sun, Minyi Shao, Xiaohui Nie, Xingda Li, Shenglin Zhang, Changhua Pei, Dongbiao He, Yanbiao Li 0001, Dan Pei
IEEE Trans. Serv. Comput.3
2024 LabelEase: A Semi-Automatic Tool for Efficient and Accurate Trace Labeling in Microservices
abstract
Trace data is crucial for system observability and maintainability within microservices architectures, and many operation algorithms depend heavily on trace data, including anomaly detection, root cause analysis, etc. However, the actual performance of these algorithms might be unsatisfactory due to the absence of high-quality labeled datasets for effective training and evaluation. Since billions of traces could be generated daily for large-scale microservices, labeling overhead is the main hurdle to obtaining high-quality trace datasets.In this paper, we propose LabelEase, a novel semi-automatic trace labeling tool, which uses active learning techniques to achieve efficient and accurate trace labeling. For anomaly trace labeling, LabelEase clusters similar traces with a graph-based trace representation technique and selects a few representative traces for human labeling, avoiding labeling most of the traces. For root cause labeling, LabelEase aggregates the labeled anomalous traces and identifies the service’s failures for operators to label. Our systematic experiments on two large-scale datasets show that LabelEase achieves over 0.98 F1-score in anomaly trace labeling and 0.89 precision of failure detection in root cause labeling, LabelEase can reduce operators’ labeling overhead by more than 99.9%. To the best of our knowledge, we are the first to propose a semi-automatic trace labeling tool capable of achieving efficient and accurate trace labeling.
Shenglin Zhang, Zeyu Che, Zhongjie Pan, Xiaohui Nie, Yongqian Sun, Lemeng Pan, Dan Pei
ISSRE4
2024 Guardian of the Resiliency: Detecting Erroneous Software Changes Before They Make Your Microservice System Less Fault-Resilient
abstract
The microservice system’s resilience is crucial for ensuring the quality of service. Nowadays, software changes are frequent and error-prone, and erroneous software changes could reduce microservice systems’ resilience to handle faults, leading to service failures and negatively impacting user experience. To better understand erroneous software changes, we conducted an empirical study on 256 real-world incidents from four famous microservice systems. Our quantitative results indicate that 37.87% of erroneous software changes make the microservice systems less fault-resilient; that is, when a fault (e.g., network fluctuation, high CPU usage, etc.) happens in the system after the software change, the services are more likely to experience failures. We refer to these software changes as Erroneous Software Changes that Reduce fault Resilience(ESCR). Traditional methods struggle to detect ESCRs effectively because the occurrence of faults is unpredictable and can hardly be in their post-change monitoring windows. In this paper, we propose a novel framework named ResilienceGuardian, aiming to detect ESCRs before they make microservice systems less fault-resilient. The key idea is utilizing fault injection techniques to evaluate systems’ fault resilience in the staging environment and then training lightweight classifiers of KPI segment pairs to detect ESCRs. The performance of ResilienceGuardian is systematically evaluated on three datasets with various faults and erroneous software changes. The results show that ResilienceGuardian significantly outperforms all the baselines with a 0.9 F1-score in identifying ESCRs and reduces the training time by 56.23% to 97.53%. Besides, ResilienceGuardian can achieve minute-level ESCR detection in large-scale microservice systems.
Guanglei He, Xiaohui Nie, Ruming Tang, Zhaoyang Yu 0002, Xidao Wen, Kanglin Yin, Dan Pei
IWQoS2
2024 End-to-End AutoML for Unsupervised Log Anomaly Detection
abstract
As modern software systems evolve towards greater complexity, ensuring their reliable operation has become a critical challenge. Log data analysis is vital in maintaining system stability, with anomaly detection being a key aspect. However, existing log anomaly detection methods heavily rely on manual effort from experts, lacking transferability across systems. This has led to the situation where to perform anomaly detection on a new dataset, the operators must have a high level of understanding of the dataset, make multiple attempts, and spend a lot of time to deploy an algorithm that performs well successfully. This paper proposes LogCraft, an end-to-end unsupervised log anomaly detection framework based on automated machine learning (AutoML). LogCraft automates feature engineering, model selection, and anomaly detection, reducing the need for specialized knowledge and lowering the threshold for algorithm deployment. Extensive evaluations on five public datasets demonstrate LogCraft's effectiveness, achieving an average F1 score of 0.899, which outperforms the second-best average F1 score of 0.847 obtained by existing unsupervised algorithms. According to our knowledge, LogCraft is the first attempt to extract fixed-dimensional vectors as latent representations from a complete log dataset. The proposed meta-feature extractor also exhibits promising potential for measuring log dataset similarity and guiding future log analytics research.
Shenglin Zhang, Yuhe Ji, Jiaqi Luan, Xiaohui Nie, Minghua Ma, Yongqian Sun, Dan Pei
ASE4
2024 Microservice Root Cause Analysis With Limited Observability Through Intervention Recognition in the Latent Space
abstract
Many failure root cause analysis (RCA) algorithms for microservices have been proposed with the widespread adoption of microservices systems. Existing algorithms generally focus on RCA with ranking single-level (e.g. metric-level or service-level) root cause candidates (RCCs) with comprehensive monitoring metrics. However, many heterogeneous RCCs exist with limited observability in real-world microservices systems. Further, we find that the limited observability may result in inaccurate RCA through real-world failures in eBay. In this paper, for the first time, we propose to "model RCCs as latent variables". The core idea is to infer the status of RCCs as latent variables with related monitoring metrics instead of directly extracting features from only the observable metrics. Based on this, we propose LatentScope, an unsupervised RCA framework with heterogeneous RCCs under limited observability. A dual-space graph is proposed to model both observable and unobservable variables, with many-to-many relationships between spaces. To achieve fast inference of latent variables and RCA, we propose the LatentRegressor algorithm, which includes Regression-based Latent-space Intervention Recognition (RLIR) to achieve intervention recognition-based RCA in latent space. LatentScope has been deployed in eBay's production environment and evaluated on both eBay's real-world failures and a testbed dataset. The evaluation results show that, compared with baseline algorithms, our model significantly improves the Top-1 recall by 9.7%-57.9%. The source code of LatentScope and the dataset are available at https://github.com/NetManAIOps/LatentScope.
Zhe Xie, Shenglin Zhang, Yitong Geng, Yao Zhang 0009, Minghua Ma, Xiaohui Nie, Zhenhe Yao, Longlong Xu, Yongqian Sun, Dan Pei
KDD6
2022 Generic and Robust Performance Diagnosis via Causal Inference for OLTP Database Systems
abstract
Online transaction processing (OLTP) database systems provide an effective solution to data support for online applications with high concurrency and low latency. An interruption or performance degradation of OLTP database systems may impact the availability of services and bring substantial economic loss. Thus, diagnosing the issue timely and mitigating it rapidly are essential for database administrators (DBAs). However, performance diagnosis for database systems is challenging due to numerous abnormal metrics, complex failure propagation, and high-performance requirements. Existing works relying on anomaly detection or causal graph construction cannot handle all these challenges simultaneously. In this paper, we propose an unsupervised learning-based method, CauseRank, to perform root cause localization with superior efficiency, high accuracy, and good interpretability. Two key techniques in CauseRank are a novel causal discovery algorithm named Group-based Greedy Equivalent Search (G-GES) incorporated with domain knowledge which treats metric groups as nodes to capture failure propagation and a simple yet effective ranking method named Causal Oriented Personalized PageRank (COPP). Extensive experiments on 97 real-world failure cases collected from a large-scale Oracle database demonstrate the effectiveness of CauseRank, achieving 82.5% top-3 accuracy and 93.8% top-5 accuracy and outperforming baseline approaches. The core idea and framework of CauseRank are generic and can be applied to other large-scale system components.
Xianglin Lu, Zhe Xie, Zeyan Li 0001, Mingjie Li 0005, Xiaohui Nie, Nengwen Zhao, Qingyang Yu, Shenglin Zhang, Kaixin Sui, Dan Pei
CCGRID5
2022 Mining Fluctuation Propagation Graph Among Time Series with Active Learning
Mingjie Li 0005, Minghua Ma, Xiaohui Nie, Kanglin Yin, Xidao Wen, Zhiyun Yuan, Duogang Wu, Guoying Li, Dan Pei
DEXA (1)3
2022 Effective Attribute Selection for Multi-dimensional Root Cause Analysis
abstract
Using large-scale multi-dimensional data for root cause analysis (MDRCA) is vitally important for online software services. It helps operators narrow down the scope of anomalies and failures quickly and localize the root cause to a finer granularity. However, most existing MDRCA algorithms can only solve low-dimensional problems. When dealing with high-dimensional data, the complexity of these algorithms would significantly increase, and even some algorithms would no longer work. Intuitively, passing only a subset of attributes rather than full attributes can improve the performance of these MDRCA algorithms. However, it is challenging due to data imbalance and novel root cause attributes. To better understand the problem of root-cause-oriented attribute selection (RCOAS), we conduct a preliminary study based on real-world data. We find that there exist several straightforward rules to filter out some attributes. In addition, we reveal that existing approaches do not fit the requirements of RCOAS. Motivated by the study, we propose an RCOAS approach, RC-LIR, to select a subset of attributes for downstream algorithms. RC-LIR first performs rule-based selection. Then it improves a feature selection algorithm by two strategies, i.e., scaling up imbalanced data and considering the redundant cost. Experiments on 1000 real-world fault cases demonstrate that RC-LIR can achieve an F1-score of 0.88, outper-forming the baseline approaches by at least 0.15. Furthermore, our experiments with four widely adopted MDRCA algorithms show that integrating RC-LIR can lead to more effective and efficient MDRCA.
Yiran Cheng, Pengxiang Jin, Yongqian Sun, Xiaohui Nie, Nengwen Zhao, Shenglin Zhang, Dan Pei
ISSRE5
2022 Causal Inference-Based Root Cause Analysis for Online Service Systems with Intervention Recognition
abstract
Fault diagnosis is critical in many domains, as faults may lead to safety threats or economic losses. In the field of online service systems, operators rely on enormous monitoring data to detect and mitigate failures. Quickly recognizing a small set of root cause indicators for the underlying fault can save much time for failure mitigation. In this paper, we formulate the root cause analysis problem as a new causal inference task namedintervention recognition. We proposed a novel unsupervised causal inference-based method namedCausal Inference-based Root Cause Analysis (CIRCA). The core idea is a sufficient condition for a monitoring variable to be a root cause indicator,i.e., the change of probability distribution conditioned on the parents in the Causal Bayesian Network (CBN). Towards the application in online service systems, CIRCA constructs a graph among monitoring metrics based on the knowledge of system architecture and a set of causal assumptions. The simulation study illustrates the theoretical reliability of CIRCA. The performance on a real-world dataset further shows that CIRCA can improve the recall of the top-1 recommendation by 25% over the best baseline method.
Mingjie Li 0005, Zeyan Li 0001, Kanglin Yin, Xiaohui Nie, Wenchi Zhang, Kaixin Sui, Dan Pei
KDD4
2022 Actionable and interpretable fault localization for recurring failures in online service systems
abstract
Fault localization is challenging in an online service system due to its monitoring data's large volume and variety and complex dependencies across/within its components (e.g., services or databases). Furthermore, engineers require fault localization solutions to be actionable and interpretable, which existing research approaches cannot satisfy. Therefore, the common industry practice is that, for a specific online service system, its experienced engineers focus on localization for recurring failures based on the knowledge accumulated about the system and historical failures. More specifically, 1) they can identify the underlying root causes and take mitigation actions when pinpointing a group of indicative metrics on the faulty component; 2) their diagnosis knowledge is roughly based on how one failure might affect the components in the whole system.
Zeyan Li 0001, Nengwen Zhao, Mingjie Li 0005, Xianglin Lu, Dongdong Chang, Xiaohui Nie, Wenchi Zhang, Kaixin Sui, Guoqiang Duan, Dan Pei
ESEC/SIGSOFT FSE7
2021 Identifying Root-Cause Metrics for Incident Diagnosis in Online Service Systems
abstract
Incidents in online service systems could incur poor user experience and tremendous economic loss. To reduce the influence of incidents and guarantee service reliability, it is critical to identify root-cause metrics for engineers with clues to assist incident diagnosis. However, it is a challenging task due to the complicated dependencies and huge volume of various metrics in large-scale systems. Existing approaches are based on either anomaly detection or correlation analysis, performing not well in terms of accuracy or efficiency. To better understand the problem of root-cause metric identification, we conduct a preliminary study based on real-world data analysis and interactions with engineers. The key observation is that root-cause metrics should satisfy two requirements. One is that the metric is expected to behave abnormally during the incident; the other is that the anomaly pattern should meet physical meaning and engineers' demand. Motivated by the findings obtained from the study, we propose an effective approach named PatternMatcher to identifying root-cause metrics accurately. Specifically, PatternMatcher contains three steps, where coarse-grained anomaly detection aiming to filter out normal metrics, anomaly pattern classification aiming to filter out unimportant anomaly patterns, and root-cause metric ranking. An extensive study on four real-world datasets including 113 incident cases from a large commercial bank demonstrates that PatternMatcher outperforms all baseline approaches, achieving top-3 average accuracy of 0.91. Moreover, we have deployed PatternMatcher in practice and shared some successful cases from real deployment.
Canhua Wu, Nengwen Zhao, Xiaoqin Yang, ShiNing Li, Xidao Wen, Xiaohui Nie, Wenchi Zhang, Kaixin Sui, Dan Pei
ISSRE9
2021 Practical Root Cause Localization for Microservice Systems via Trace Analysis
abstract
Microservice architecture is applied by an increasing number of systems because of its benefits on delivery, scalability, and autonomy. It is essential but challenging to localize root-cause microservices promptly when a fault occurs. Traces are helpful for root-cause microservice localization, and thus many recent approaches utilize them. However, these approaches are less practical due to relying on supervision or other unrealistic assumptions. To overcome their limitations, we propose a more practical root-cause microservice localization approach named TraceRCA. The key insight of TraceRCA is that a microservice with more abnormal and less normal traces passing through it is more likely to be the root cause. Based on it, TraceRCA is composed of trace anomaly detection, suspicious microservice set mining and microservice ranking. We conducted experiments on hundreds of injected faults in a widely-used open-source microservice benchmark and a production system. The results show that TraceRCA is effective in various situations. The top-1 accuracy of TraceRCA outperforms the state-of-the-art unsupervised approaches by 44.8%. Besides, TraceRCA is applied in a large commercial bank, and it helps operators localize root causes for real-world faults accurately and efficiently. We also share some lessons learned from our real-world deployment.
Zeyan Li 0001, Junjie Chen 0003, Nengwen Zhao, Shuwei Zhang, Long Jiang, Leiqin Yan, Zikai Wang 0010, Zhekang Chen, Wenchi Zhang, Xiaohui Nie, Kaixin Sui, Dan Pei
IWQoS13
2021 Jump-Starting Multivariate Time Series Anomaly Detection for Online Service Systems
Minghua Ma, Shenglin Zhang, Junjie Chen 0003, Jim Xu, Yongliang Lin, Xiaohui Nie, Dan Pei
USENIX ATC7
2021 BDS+: An Inter-Datacenter Data Replication System With Dynamic Bandwidth Separation
abstract
Many important cloud services require replicating massive data from one datacenter (DC) to multiple DCs. While the performance of pair-wise inter-DC data transfers has been much improved, prior solutions are insufficient to optimize bulk-data multicast, as they fail to explore the rich inter-DC overlay paths that exist in geo-distributed DCs, as well as the remaining bandwidth reserved for online traffic under fixed bandwidth separation scheme. To take advantage of these opportunities, we present BDS+, a near-optimal network system for large-scale inter-DC data replication. BDS+ is an application-level multicast overlay network with a fully centralized architecture, allowing a central controller to maintain an up-to-date global view of data delivery status of intermediate servers, in order to fully utilize the available overlay paths. Furthermore, in each overlay path, it leverages dynamic bandwidth separation to make use of the remaining available bandwidth reserved for online traffic. By constantly estimating online traffic demand and rescheduling bulk-data transfers accordingly, BDS+ can further speed up the massive data multicast. Through a pilot deployment in one of the largest online service providers and large-scale real-trace simulations, we show that BDS+ can achieve 3- 5× speedup over the provider's existing system and several well-known overlay routing baselines of static bandwidth separation. Moreover, dynamic bandwidth separation can further reduce the completion time of bulk data transfers by 1.2 to 1.3 times.
Yuchao Zhang 0004, Xiaohui Nie, Junchen Jiang, Wendong Wang 0003, Ke Xu 0002, Youjian Zhao, Martin J. Reed, Kai Chen 0005, Guang Yao
IEEE/ACM Trans. Netw.2
2020 Real-time incident prediction for online service systems
abstract
Incidents in online service systems could dramatically degrade system availability and destroy user experience. To guarantee service quality and reduce economic loss, it is essential to predict the occurrence of incidents in advance so that engineers can take some proactive actions to prevent them. In this work, we propose an effective and interpretable incident prediction approach, called eWarn, which utilizes historical data to forecast whether an incident will happen in the near future based on alert data in real time. More specifically, eWarn first extracts a set of effective features (including textual features and statistical features) to represent omen alert patterns via careful feature engineering. To reduce the influence of noisy alerts (that are not relevant to the occurrence of incidents), eWarn then incorporates the multi-instance learning formulation. Finally, eWarn builds a classification model via machine learning and generates an interpretable report about the prediction result via a state-of-the-art explanation technique (i.e., LIME). In this way, an early warning signal along with its interpretable report can be sent to engineers to facilitate their understanding and handling for the incoming incident. An extensive study on 11 real-world online service systems from a large commercial bank demonstrates the effectiveness of eWarn, outperforming state-of-the-art alert-based incident prediction approaches and the practice of incident prediction with alerts. In particular, we have applied eWarn to two large commercial banks in practice and shared some success stories and lessons learned from real deployment.
Nengwen Zhao, Junjie Chen 0003, Xiaohui Nie, Wenchi Zhang, Kaixin Sui, Dan Pei
ESEC/SIGSOFT FSE9
2019 FluxRank: A Widely-Deployable Framework to Automatically Localizing Root Cause Machines for Software Service Failure Mitigation
abstract
The failures of software service directly affect user experiences and service revenue. Thus operators monitor both service-level KPIs (e.g., response time) and machine-level KPIs (e.g., CPU usage) on each machine underlying the service. When a service fails, the operators must localize the root cause machines, and mitigate the failure as quickly as possible. Existing approaches have limited application due to the difficulty to obtain the required additional measurement data. As a result, failure localization is largely manual and very time-consuming. This paper presents FluxRank, a widely-deployable framework that can automatically and accurately localize the root cause machines, so that some actions can be triggered to mitigate the service failure. Our evaluation using historical cases from five real services (with tens of thousands of machines) of a top search company shows that the root cause machines are ranked top 1 (top 3) for 55 (66) cases out of 70 cases. Comparing to existing approaches, FluxRank cuts the localization time by more than 80% on average. FluxRank has been deployed online at one Internet service and six banking services for three months, and correctly localized the root cause machines as the top 1 for 55 cases out of 59 cases.
Xiaohui Nie, Jing Zhu 0007, Shenglin Zhang, Kaixin Sui, Dan Pei
ISSRE3
2019 Dynamic TCP Initial Windows and Congestion Control Schemes Through Reinforcement Learning
abstract
Despite many years of improvements to it, TCP still suffers from an unsatisfactory performance. For services dominated by short flows (e.g., web search and e-commerce), TCP suffers from the flow startup problem and cannot fully utilize the available bandwidth in the modern Internet: TCP starts from a conservative and static initial window (IW, 2-4 or 10), while most of the web flows are too short to converge to the best sending rate before the session ends. For services dominated by long flows (e.g., video streaming and file downloading), the congestion control (CC) scheme manually and statically configured might not offer the best performance for the latest network conditions. To address these two challenges, we propose TCP-RL, which uses reinforcement learning (RL) techniques to dynamically configure IW and CC in order to improve the performance of TCP flow transmission. Basing on the latest network conditions observed at the server side of a web service, TCP-RL dynamically configures a suitable IW for short flows through group-based RL, and dynamically configures a suitable CC scheme for long flows through deep RL. Our extensive experiments show that for short flows, TCP-RL can reduce the average transmission time by about 23%; and for long flows, compared with the performance of 14 CC schemes, TCP-RL's performance ranks top 5 for about 85% of the 288 given static network conditions, whereas for about 90% of conditions, its performance drops by less than 12% compared with that of the best-performing CC schemes for the same network conditions.
Xiaohui Nie, Youjian Zhao, Zhihan Li 0002, Guo Chen 0001, Kaixin Sui, Zijie Ye, Dan Pei
IEEE J. Sel. Areas Commun.1
2018 BDS: a centralized near-optimal overlay network for inter-datacenter data replication
abstract
Many important cloud services require replicating massive data from one datacenter (DC) to multiple DCs. While the performance of pair-wise inter-DC data transfers has been much improved, prior solutions are insufficient to optimize bulk-data multicast, as they fail to explore the capability of servers to store-and-forward data, as well as the rich inter-DC overlay paths that exist in geo-distributed DCs. To take advantage of these opportunities, we present BDS, an application-level multicast overlay network for large-scale inter-DC data replication. At the core of BDS is a fully centralized architecture, allowing a central controller to maintain an up-to-date global view of data delivery status of intermediate servers, in order to fully utilize the available overlay paths. To quickly react to network dynamics and workload churns, BDS speeds up the control algorithm by decoupling it into selection of overlay paths and scheduling of data transfers, each can be optimized efficiently. This enables BDS to update overlay routing decisions in near realtime (e.g., every other second) at the scale of multicasting hundreds of TB data over tens of thousands of overlay paths. A pilot deployment in one of the largest online service providers shows that BDS can achieve 3-5 x speedup over the provider's existing system and several well-known overlay routing baselines.
Yuchao Zhang 0004, Junchen Jiang, Ke Xu 0002, Xiaohui Nie, Martin J. Reed, Guang Yao, Kai Chen 0005
EuroSys4
2018 Reducing Web Latency Through Dynamically Setting TCP Initial Window with Reinforcement Learning
abstract
Latency, which directly affects the user experience and revenue of web services, is far from ideal in reality, due to the well-known TCP flow startup problem. Specifically, since TCP starts from a conservative and static initial window (IW, 2~4 or 10), most of the web flows are too short to have enough time to find its best congestion window before the session ends. As a result, TCP cannot fully utilize the available bandwidth in the modern Internet. In this paper, we propose to use group-based reinforcement learning (RL) to enable a web server, through trial-and-error, to dynamically set a suitable IW for a web flow before its transmission starts. Our proposed system, SmartIW, collects TCP flow performance metrics (e.g., transmission time, loss rate, RTT) in real-time without any client assistance. Then these metrics are aggregated into groups with similar features (subnet, ISP, province, etc.) to satisfy RL's requirement. SmartIW has been deployed in one of the top global search engines for more than a year. Our online and testbed experiments show that, compared to the common practice of IW=10, SmartIW can reduce the average transmission time by 23% to 29%.
Xiaohui Nie, Youjian Zhao, Dan Pei, Guo Chen 0001, Kaixin Sui
IWQoS1
2017 TCP WISE: One initial congestion window is not enough
abstract
Current TCP is very inefficient for web services. Web transactions are often very short-lived. TCP flow starts with a conservative initial congestion window (IW), which causes multiple round-trip times to finish the transmission even if the end-to-end bandwidth is sufficient for the transaction to be finished in one round-trip time. Previous research efforts have been focusing on finding the overall best IW for the entire Internet or a service company. However, we observe that one-IW-fits-all is suboptimal after one year of online measurement in Baidu, one of the top global search engine companies. To reduce the TCP latency, we propose TCP WISE, which dynamically assigns suitable IWs for different user cluster at different times on the server side. The values of users' IWs are proactively learned based on the historical experience on the server-side. Our testbed experiments show that our learning algorithm can handle the network changes and converge to the best IW. We have deployed TCP WISE in one of Baidu's production data center, and results shows that the 80thpercentile latency of the HTTP responses has been reduced by 10.4% compared with current TCP with a fixed IW of 10.
Xiaohui Nie, Youjian Zhao, Guo Chen 0001, Kaixin Sui, Yazheng Chen, Dan Pei
IPCCC1
2016 Mining causality graph for automatic web-based service diagnosis
abstract
It is crucial for Internet company to provide highly reliable web-based services. The web-based services always have many components running in the large-scale infrastructure with complex interactions. As an indispensable part of high reliability, the diagnosis remains to be a thorny problem. With the growth of system scale and complexity, it becomes even more difficult. In this paper, we propose an automatic diagnosis system based on causality graph to help system operators find the root causes. The causality graph is mainly extracted from the historical data of the monitoring system, and the method consists of four steps. 1) It utilizes a data mining method to extract the initial causality graph. 2) Once a failure happens, it lists top-k suspects with a ranking algorithm based on the causality graph. 3) Then system operators check the suspects and label them either right or wrong. 4) A supervised learning algorithm takes the labels as the input to tune the causality graph, in order to improve the diagnosis accuracy on step 2 iteratively. This method requires neither knowledge about the design and implementation details of the web-based service, nor instrumenting the services' source code. Our controlled experiments show that the root causes can be ranked in top 3 with 100% accuracy after countable learning iterations.
Xiaohui Nie, Youjian Zhao, Kaixin Sui, Dan Pei, Xianping Qu
IPCCC1
2016 PieBridge: A Cross-DR scale Large Data Transmission Scheduling System
abstract
Cross-DR WAN (Datacenter Region Wide Area Network) with various services are deployed to provide timely data information and analytics for users in a wide range of geographical locations. For its reliability and performance, data duplication synchronization is essential among different IDCs (Internet datacenters). However, this problem poses a challenge. First, data duplication requires huge amount of bandwidth whereas the bandwidth of cross-DR links and the upload/download rates of server interfaces are limited. Second, data transmissions are time sensitive, but the current network cannot complete such tasks in a timely manner. In this work, we present PieBridge, a cross-RD data duplicate transmission platform that accommodates hundreds of TBs of data generated from user applications online data analytics. We deployed PieBridge on the IDCs of Baidu and obtained promising performance results in comparison with the prevalent approaches.
Yuchao Zhang 0004, Ke Xu 0002, Guang Yao, Xiaohui Nie
SIGCOMM5