Ji Xue

dblp:147/6124 · DBLP profile ↗
← Back
11ranked-venue papers
7as first author
1since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-authorSecurity and privacy · 3 · 2 first-authorComputer networks · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Storage systems · 100%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems › storage reliability
failure characterization
0.412019
SSD failures in the field: symptoms, causes, and prediction models · SC 2019
Storage systems
flash and SSD
0.412019
SSD failures in the field: symptoms, causes, and prediction models · SC 2019
Storage systems › flash and SSD › SSD reliability
SSD failure prediction
0.412019
SSD failures in the field: symptoms, causes, and prediction models · SC 2019
Storage systems
storage reliability
0.412019
SSD failures in the field: symptoms, causes, and prediction models · SC 2019

Methods — techniques the papers use, named apart from their topics

machine learning · 0.8failure prediction models · 0.8
YearPublicationVenuePosition
2023 Bridging Resource Prediction and System Management: A Case Study in Cloud Systems
abstract
In recent years, there has been a significant amount of research focused on predicting resources in order to enhance the performance of cloud systems. Many researchers believe that the more accurate the prediction, the more effective resource management will be in ensuring reliable performance. However, our study in this paper demonstrates that there is a gap between resource demand prediction and system performance. Furthermore, our experiment results have demonstrated that the accurate and fine-grained prediction helps to achieve a more reliable and efficient system performance, especially in CPU utilization rate.
Justin Kur, Ji Xue, Jingshu Chen, Jun Huang 0001
CNSM2
2019 Session details: Fairness and Performance
abstract
No abstract available.
Ji Xue
HPDC1
2019 SSD failures in the field: symptoms, causes, and prediction models
abstract
In recent years, solid state drives (SSDs) have become a staple of high-performance data centers for their speed and energy efficiency. In this work, we study the failure characteristics of 30,000 drives from a Google data center spanning six years. We characterize the workload conditions that lead to failures and illustrate that their root causes differ from common expectation but remain difficult to discern. Particularly, we study failure incidents that result in manual intervention from the repair process. We observe high levels of infant mortality and characterize the differences between infant and non-infant failures. We develop several machine learning failure prediction models that are shown to be surprisingly accurate, achieving high recall and low false positive rates. These models are used beyond simple prediction as they aid us to untangle the complex interaction of workload characteristics that lead to failures and identify failure root causes from monitored symptoms.
Jacob Alter, Ji Xue, Alma Dimnaku, Evgenia Smirni
SC2
2018 Machine Learning Models for GPU Error Prediction in a Large Scale HPC System
abstract
GPUs are widely deployed on large-scale HPC systems to provide powerful computational capability for scientific applications from various domains. As those applications are normally long-running, investigating the characteristics of GPU errors becomes imperative for reliability. In this paper, we first study the system conditions that trigger GPU errors using six-month trace data collected from a large-scale, operational HPC system. Then, we use machine learning to predict the occurrence of GPU errors, by taking advantage of temporal and spatial dependencies of the trace data. The resulting machine learning prediction framework is robust and accurate under different workloads.
Bin Nie, Ji Xue, Saurabh Gupta 0002, Tirthak Patel, Christian Engelmann, Evgenia Smirni, Devesh Tiwari
DSN2
2018 Spatial-Temporal Prediction Models for Active Ticket Managing in Data Centers
abstract
Performance ticket handling is an expensive operation in data centers, where physical boxes host multiple virtual machines (VMs). A large body of tickets arise from resource usage warnings, e.g., CPU and RAM usages that exceed predefined thresholds. The transient nature of CPU and RAM usage as well as their strong correlation across time among co-located VMs within boxes drastically increase the complexity of ticket management. Based on large resource usage data collected from production data centers, with 6K physical boxes and more than 80K VMs, we first discover patterns of spatial and temporal dependencies among/within the usage series of co-located resources. Leveraging our key findings, we develop an active ticket managing (ATM) system that aims to drastically reduce usage tickets. ATM consists of: 1) a spatial-temporal dependency-based time series prediction methodology and 2) a proactive capacity planning policy for CPU and RAM resources for VMs co-located within a box and boxes within a single data center client, that aims to drastically reduce usage tickets. ATM exploits the spatial-temporal dependency across/within multiple resources of co-located VMs and single-client boxes for usage prediction, and then actuates proactive capacity planning. Evaluation results on traces of 6K physical boxes from operating data centers show that ATM is able to provide accurate prediction of usage series in cloud data centers with low computational overhead. At the same time ATM achieves significant ticket reduction up to 60% for both VM and box usage series.
Ji Xue, Robert Birke, Lydia Y. Chen, Evgenia Smirni
IEEE Trans. Netw. Serv. Manag.1
2017 Fill-in the gaps: Spatial-temporal models for missing data
abstract
Effective workload characterization and prediction are instrumental for efficiently and proactively managing large systems. System management primarily relies on the workload information provided by underlying system tracing mechanisms that record system-related events in log files. However, such tracing mechanisms may temporarily fail due to various reasons, yielding “holes” in data traces. This missing data phenomenon significantly impedes the effectiveness of data analysis. In this paper, we study real-world data traces collected from over 80K virtual machines (VMs) hosted on 6K physical boxes in the data centers of a service provider. We discover that the usage series of VMs co-located on the same physical box exhibit strong correlation with one another, and that most VM usage series show temporal patterns. By taking advantage of the observed spatial and temporal dependencies, we propose a data-filling method to predict the missing data in the VM usage series. Detailed evaluation using trace data in the wild shows that the proposed method is sufficiently accurate as it achieves an average of 20% absolute percentage errors. We also illustrate its usefulness via a use case.
Ji Xue, Bin Nie, Evgenia Smirni
CNSM1
2017 Characterizing Temperature, Power, and Soft-Error Behaviors in Data Center Systems: Insights, Challenges, and Opportunities
abstract
GPUs have become part of the mainstream high performance computing facilities that increasingly require more computational power to simulate physical phenomena quickly and accurately. However, GPU nodes also consume significantly more power than traditional CPU nodes, and high power consumption introduces new system operation challenges, including increased temperature, power/cooling cost, and lower system reliability. This paper explores how power consumption and temperature characteristics affect reliability, provides insights into what are the implications of such understanding, and how to exploit these insights toward predicting GPU errors using neural networks.
Bin Nie, Ji Xue, Saurabh Gupta 0002, Christian Engelmann, Evgenia Smirni, Devesh Tiwari
MASCOTS2
2016 Managing Data Center Tickets: Prediction and Active Sizing
abstract
Performance ticket handling is an expensive operation in highly virtualized cloud data centers where physical boxes host multiple virtual machines (VMs). A large body of tickets arise from the resource usage warnings, e.g., CPU and RAM usages that exceed predefined thresholds. The transient nature of CPU and RAM usage as well as their strong correlation across time among co-located VMs drastically increase the complexity in ticket management. Based on a large resource usage data collected from production data centers, amount to 6K physical machines and more than 80K VMs, we first discover patterns of spatial dependency among co-located virtual resources. Leveraging our key findings, we develop an Active Ticket Managing(ATM) system that consists of (i) a novel time series prediction methodology and (ii) a proactive VM resizing policy for CPU and RAM resources for co-located VMs on a physical box that aims to drastically reduce usage tickets. ATM exploits the spatial dependency across multiple resources of co-located VMs for usage prediction and proactive VM resizing. Evaluation results on traces of 6K physical boxes and a prototype of a MediaWiki system show that ATM is able to achieve excellent prediction accuracy of a large number of VM time series and significant usage ticket reduction, i.e., up to 60%, at low computational overhead.
Ji Xue, Robert Birke, Lydia Y. Chen, Evgenia Smirni
DSN1
2016 Tale of Tails: Anomaly Avoidance in Data Centers
abstract
It is a common practice that today's cloud data centers guard the performance by monitoring the resource usage, e.g., CPU and RAM, and issuing anomaly tickets whenever detecting usages exceeding predefined target values. Ensuring free of such usage anomaly can be extremely challenging, while catering to a large amount of virtual machines (VMs) showing bursty workloads on a limited amount of physical resource. Using resource usage data from production data centers that consist of more than 6K physical machines hosting more than 80K VMs, we identify statistic properties of anomaly instances (AIs) on physical servers, highlighting their burst duration and potential root causes. To strike a tradeoff between a strong performance guarantee and resource provisions, we propose a tail-driven anomaly avoidance policy for boxes, TailGuard, which allows a small fraction of AIs, e.g., 5% of usages can be above the target value, and still avoid severe performance degradation, typically caused by a burst of continuous AI. Specifically, TailGuard first introduces a novel usage tail prediction that explores the similarity patterns across a great number of boxes within a very recent history, and then redistributes the server load in an online fashion by proactive VM cloning and reactive load balancing. Evaluation results show that TailGuard can not only achieve an accuracy comparable with prediction methodology that relies on long history of usage data but also dramatically reduce the number of CPU AIs by 60%, with a tenfold reduction of their duration, from more than 25 time windows to only 2.
Ji Xue, Robert Birke, Lydia Y. Chen, Evgenia Smirni
SRDS1
2016 PROST: Predicting Resource Usages with Spatial and Temporal Dependencies
abstract
We present a tool, PROST, which can achieve scalable and accurate prediction of server workload time series in data centers. As several virtual machines are typically co-located on physical servers, the CPU and RAM show strong temporal and spatial dependencies. PROST is able to leverage the spatial dependency among co-located VMs to improve the scalability of prediction models solely based on temporal features, such as neural network. We show the benefits of PROST in obtaining accurate prediction of resource usage series and designing effective VM sizing strategies for the private data centers.
Ji Xue, Evgenia Smirni, Thomas Scherer, Robert Birke, Lydia Y. Chen
ICPE1
2015 PRACTISE: Robust prediction of data center time series
abstract
We analyze workload traces from production data centers and focus on their VM usage patterns of CPU, memory, disk, and network bandwidth. Burstiness is a clear characteristic of many of these time series: there exist peak loads within clear periodic patterns but also within patterns that do not have clear periodicity. We present PRACTISE, a neural network based framework that can efficiently and accurately predict future loads, peak loads, and their timing. Extensive experimentation using traces from IBM data centers illustrates PRACTISE's superiority when compared to ARIMA and baseline neural network models, with average prediction errors that are significantly smaller. Its robustness is also illustrated with respect to the prediction window that can be short-term (i.e., hours) or long-term (i.e., a week).
Ji Xue, Feng Yan 0001, Robert Birke, Lydia Y. Chen, Thomas Scherer, Evgenia Smirni
CNSM1