Jacob Alter

dblp:252/4596 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
1since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 first-authorSecurity and privacy · 2 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Storage systems · 94% Cloud and datacenter computing · 6%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems
flash and SSD
1.022023
Lifespan and Failures of SSDs and HDDs: Similarities, Differences, and Prediction Models · IEEE Trans. Dependable Secur. Comput. 2023
SSD failures in the field: symptoms, causes, and prediction models · SC 2019
Storage systems
storage reliability
1.022023
Lifespan and Failures of SSDs and HDDs: Similarities, Differences, and Prediction Models · IEEE Trans. Dependable Secur. Comput. 2023
SSD failures in the field: symptoms, causes, and prediction models · SC 2019
Storage systems › storage reliability
failure characterization
0.412019
SSD failures in the field: symptoms, causes, and prediction models · SC 2019
Storage systems › flash and SSD › SSD reliability
SSD failure prediction
0.412019
SSD failures in the field: symptoms, causes, and prediction models · SC 2019
Cloud and datacenter computing
datacenter storage
0.212023
Lifespan and Failures of SSDs and HDDs: Similarities, Differences, and Prediction Models · IEEE Trans. Dependable Secur. Comput. 2023

Methods — techniques the papers use, named apart from their topics

machine learning · 0.8failure prediction models · 0.8machine learning failure prediction · 0.7
YearPublicationVenuePosition
2023 Lifespan and Failures of SSDs and HDDs: Similarities, Differences, and Prediction Models
abstract
Data center downtime typically centers around IT equipment failure. Storage devices are the most frequently failing components in data centers. We present a comparative study of hard disk drives (HDDs) and solid state drives (SSDs) that constitute the typical storage in data centers. Using six-year field data of 100,000 HDDs of different models from the same manufacturer from the Backblaze dataset and six-year field data of 30,000 SSDs of three models from a Google data center, we characterize the workload conditions that lead to failures. We illustrate that their root failure causes differ from common expectations and that they remain difficult to discern. For the case of HDDs we observe that young and old drives do not present many differences in their failures. Instead, failures may be distinguished by discriminating drives based on the time spent for head positioning. For SSDs, we observe high levels of infant mortality and characterize the differences between infant and non-infant failures. We develop several machine learning failure prediction models that are shown to be surprisingly accurate, achieving high recall and low false positive rates. These models are used beyond simple prediction as they aid us to untangle the complex interaction of workload characteristics that lead to failures and identify failure root causes from monitored symptoms.
Riccardo Pinciroli, Lishan Yang 0001, Jacob Alter, Evgenia Smirni
IEEE Trans. Dependable Secur. Comput.3
2020 Mining Multivariate Discrete Event Sequences for Knowledge Discovery and Anomaly Detection
abstract
Modern physical systems deploy large numbers of sensors to record at different time-stamps the status of different systems components via measurements such as temperature, pressure, speed, but also the component's categorical state. Depending on the measurement values, there are two kinds of sequences: continuous and discrete. For continuous sequences, there is a host of state-of-the-art algorithms for anomaly detection based on time-series analysis, but there is a lack of effective methodologies that are tailored specifically to discrete event sequences. This paper proposes an analytics framework for discrete event sequences for knowledge discovery and anomaly detection. During the training phase, the framework extracts pairwise relationships among discrete event sequences using a neural machine translation model by viewing each discrete event sequence as a "natural language". The relationship between sequences is quantified by how well one discrete event sequence is "translated" into another sequence. These pairwise relationships among sequences are aggregated into a multivariate relationship graph that clusters the structural knowledge of the underlying system and essentially discovers the hidden relationships among discrete sequences. This graph quantifies system behavior during normal operation. During testing, if one or more pairwise relationships are violated, an anomaly is detected. The proposed framework is evaluated on two real-world datasets: a proprietary dataset collected from a physical plant where it is shown to be effective in extracting sensor pairwise relationships for knowledge discovery and anomaly detection, and a public hard disk drive dataset where its ability to effectively predict upcoming disk failures is illustrated.
Bin Nie, Jianwu Xu, Jacob Alter, Evgenia Smirni
DSN3
2019 SSD failures in the field: symptoms, causes, and prediction models
abstract
In recent years, solid state drives (SSDs) have become a staple of high-performance data centers for their speed and energy efficiency. In this work, we study the failure characteristics of 30,000 drives from a Google data center spanning six years. We characterize the workload conditions that lead to failures and illustrate that their root causes differ from common expectation but remain difficult to discern. Particularly, we study failure incidents that result in manual intervention from the repair process. We observe high levels of infant mortality and characterize the differences between infant and non-infant failures. We develop several machine learning failure prediction models that are shown to be surprisingly accurate, achieving high recall and low false positive rates. These models are used beyond simple prediction as they aid us to untangle the complex interaction of workload characteristics that lead to failures and identify failure root causes from monitored symptoms.
Jacob Alter, Ji Xue, Alma Dimnaku, Evgenia Smirni
SC1