Eric Pershey

dblp:201/4897 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
2since 2021 · last 2025
0000-0003-2613-2768ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Hardware reliability and fault tolerance · 56% Cloud and datacenter computing · 44%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing
log analysis
0.412019
Exploring Properties and Correlations of Fatal Events in a Large-Scale HPC System · IEEE Trans. Parallel Distributed Syst. 2019
Hardware reliability and fault tolerance
system reliability
0.412019
Exploring Properties and Correlations of Fatal Events in a Large-Scale HPC System · IEEE Trans. Parallel Distributed Syst. 2019
Hardware reliability and fault tolerance › reliability analysis
correlated failures
0.112019
Exploring Properties and Correlations of Fatal Events in a Large-Scale HPC System · IEEE Trans. Parallel Distributed Syst. 2019

Methods — techniques the papers use, named apart from their topics

temporal correlation analysis · 0.4log filtering · 0.4distribution fitting · 0.4
YearPublicationVenuePosition
2025 A Big Data Approach for Efficient Processing of Machine Operational Data
Eric Pershey, Ben Lenard, Brian R. Toonen, Peter Upton, Alexander Rasin
SSDBM1
2023 An Approach for Efficient Processing of Machine Operational Data
Ben Lenard, Eric Pershey, Zachary Nault, Alexander Rasin
DEXA (1)2
2019 Characterizing and Understanding HPC Job Failures Over The 2K-Day Life of IBM BlueGene/Q System
abstract
An in-depth understanding of the failure features of HPC jobs in a supercomputer is critical to the large-scale system maintenance and improvement of the service quality for users. In this paper, we investigate the features of hundreds of thousands of jobs in one of the most powerful supercomputers, the IBM Blue Gene/Q Mira, based on 2001 days of observations with a total of over 32.44 billion core-hours. We study the impact of the system's events on the jobs' execution in order to understand the system's reliability from the perspective of jobs and users. The characterization involves a joint analysis based on multiple data sources, including the reliability, availability, and serviceability (RAS) log; job scheduling log; the log regarding each job's physical execution tasks; and the I/O behavior log. We present 22 valuable takeaways based on our in-depth analysis. For instance, 99,245 job failures are reported in the job-scheduling log, a large majority (99.4%) of which are due to user behavior (such as bugs in code, wrong configuration, or misoperations). The job failures are correlated with multiple metrics and attributes, such as users/projects and job execution structure (number of tasks, scale, and core-hours). The best-fitting distributions of a failed job's execution length (or interruption interval) include Weibull, Pareto, inverse Gaussian, and Erlang/exponential, depending on the types of errors (i.e., exit codes). The RAS events affecting job executions exhibit a high correlation with users and core-hours and have a strong locality feature. In terms of the failed jobs, our similarity-based event-filtering analysis indicates that the mean time to interruption is about 3.5 days.
Sheng Di, Hanqi Guo 0001, Eric Pershey, Marc Snir, Franck Cappello
DSN3
2019 Exploring Properties and Correlations of Fatal Events in a Large-Scale HPC System
abstract
In this paper, we explore potential correlations of fatal system events for one of the most powerful supercomputers-IBM Blue Gene/Q Mira, which is deployed at Argonne National Laboratory, based on its 5-year reliability, availability, and serviceability (RAS) log. Our contribution is two-fold. (1) We design an efficient log analysis tool, namely LogAider, with a novel filtering method to effectively extract fatal events from masses of system messages that are heavily duplicated in the log. LogAider exhibits a very precise detection of temporal-correlation with a high similarity (up to 95 percent) to the ground-truth (i.e., compared to the failure records reported by the administrators). The total number of fatal events can be reduced to about 1,255 compared with originally 2.6 million duplicated fatal messages. (2) We analyze the 5-year RAS log of the MIRA system using LogAider, and summarize six important “takeaways” which can help system vendors and administrators better understand an extreme-scale system's fatal events. Specifically, we find that the distribution or proportion of the fatal system events follow a Pareto-like principle in general. The temporal correlation among fatal events is much stronger than that of warn messages and info messages, and the correlated events tend to constitute a few clusters. The mean time between fatal events (MTBFE) of the Mira system is about 1.3 days from the perspective of the system, and the MTTI is 2-4 days from the perspective of users. The most error-prone item value with respect to any key attribute appears likely in the log every 2-10 days. Weibull, Gamma, and Pearson6 are the three best-fit distributions for the fatal event intervals. The overall correlation of fatal events on the 5D torus network is not prominent, whereas the small-region locality correlation (e.g., the fatal events inside racks) is relatively strong. We believe our work will be interesting to large-scale HPC system administrators and vendors and to fault tolerance researchers, enabling them to better understand fatal events and mitigate such events accordingly.
Sheng Di, Hanqi Guo 0001, Rinku Gupta, Eric Pershey, Marc Snir, Franck Cappello
IEEE Trans. Parallel Distributed Syst.4
2017 LogAider: A tool for mining potential correlations of HPC log events
abstract
Today's large-scale supercomputers are producing a huge amount of log data. Exploring various potential correlations of fatal events is crucial for understanding their causality and improving the working efficiency for system administrators. To this end, we developed a toolkit, named LogAider, that can reveal three types of potential correlations: across-field, spatial, and temporal. Across-field correlation refers to the statistical correlation across fields within a log or across multiple logs based on probabilistic analysis. For analyzing the spatial correlation of events, we developed a generic, easy-to-use visualizer that can view any events queried by userson a system machine graph. LogAider can also mine spatial correlations by an optimized K-meaning clustering algorithm over a Torus network topology. It is also able to disclose the temporal correlations (or error propagations) over a certain period inside a log or across multiple logs, based on an effective similarity analysis strategy. We assessed LogAider using theone-year reliability-availability-serviceability (RAS) log of Mira system (one of the world's most powerful supercomputers), as well as its job log. We find that LogAider very helpful for revealing the potential correlations of fatal system events and job events, with an accurate mining of across-field correlation with both precision and recall of 99.9-100%, as well as precisedetection of temporal-correlation with a high similarity (up to 95%) to the ground-truth.
Sheng Di, Rinku Gupta, Marc Snir, Eric Pershey, Franck Cappello
CCGrid4