Alma Dimnaku

dblp:252/4632 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Storage systems · 100%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems › storage reliability
failure characterization
0.412019
SSD failures in the field: symptoms, causes, and prediction models · SC 2019
Storage systems
flash and SSD
0.412019
SSD failures in the field: symptoms, causes, and prediction models · SC 2019
Storage systems › flash and SSD › SSD reliability
SSD failure prediction
0.412019
SSD failures in the field: symptoms, causes, and prediction models · SC 2019
Storage systems
storage reliability
0.412019
SSD failures in the field: symptoms, causes, and prediction models · SC 2019

Methods — techniques the papers use, named apart from their topics

machine learning · 0.8failure prediction models · 0.8
YearPublicationVenuePosition
2019 SSD failures in the field: symptoms, causes, and prediction models
abstract
In recent years, solid state drives (SSDs) have become a staple of high-performance data centers for their speed and energy efficiency. In this work, we study the failure characteristics of 30,000 drives from a Google data center spanning six years. We characterize the workload conditions that lead to failures and illustrate that their root causes differ from common expectation but remain difficult to discern. Particularly, we study failure incidents that result in manual intervention from the repair process. We observe high levels of infant mortality and characterize the differences between infant and non-infant failures. We develop several machine learning failure prediction models that are shown to be surprisingly accurate, achieving high recall and low false positive rates. These models are used beyond simple prediction as they aid us to untangle the complex interaction of workload characteristics that lead to failures and identify failure root causes from monitored symptoms.
Jacob Alter, Ji Xue, Alma Dimnaku, Evgenia Smirni
SC3