Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Sam Zeltner

dblp:417/4316 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
0009-0003-4347-2462ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Hardware reliability and fault tolerance · 30% High-performance computing · 30% Distributed systems · 30%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing › supercomputing
exascale computing
0.912025
Fine-grained Automated Failure Management for Extreme-Scale GPU Accelerated Systems · SC 2025
Distributed systems
fault tolerance
0.912025
Fine-grained Automated Failure Management for Extreme-Scale GPU Accelerated Systems · SC 2025
GPUs and heterogeneous computing › GPU computing
GPU-accelerated systems
0.312025
Fine-grained Automated Failure Management for Extreme-Scale GPU Accelerated Systems · SC 2025

Methods — techniques the papers use, named apart from their topics

failure analysis · 0.9
YearPublicationVenuePosition
2025 Fine-grained Automated Failure Management for Extreme-Scale GPU Accelerated Systems
abstract
As high-performance computing (HPC) systems scale in size, system wide hardware failure rates increase. Historical data from previous large-scale HPC installations illustrate this trend, with the mean time between failures (MTBF) decreasing steadily over the past decade. Recent studies from artificial intelligence and machine-learning (AI/ML) training extrapolate MTBF declining even further for future GPU accelerated systems. As MTBF decreases, mean time to repair (MTTR) becomes more pronounced, highlighting the need for efficient recovery strategies.
Yonatan Levitt, Richard Barella, Sam Zeltner, Thomas Musta, Lance C. Cheney, Gustavo Espinosa, Olivier Franza, Balazs Gerofi
SC3